Command Palette
Search for a command to run...
FOUNDATIONS OF PROACTIVE AGENTS: PRINCIPLES, TECHNICAL LAYERS, AND PROACTIVITY-GYM
FOUNDATIONS OF PROACTIVE AGENTS: PRINCIPLES, TECHNICAL LAYERS, AND PROACTIVITY-GYM
Jio Oh Seunghyun Do Young-Jun Lee Steven Euijong Whang Dongyeop Kang
Abstract
Proactive LLM agents can turn idle compute into useful support before users ask. Yet even correct work can misread user context, impose review costs, or undermine trust. This work proposes foundations for designing, realizing, and evaluating proactive LLM agents around three joint principles (3T): Task Capability, anticipating relevant needs and correctly performing useful work; Temporal Allocation, allocating compute according to resource availability and when results are needed; and Trust, sustaining users’ confidence and appropriate reliance on the agent. We connect these objectives to a design space organized around five dimensions: task scope, anticipation horizon, activation trigger, processing timing, and intervention depth, and specify the situation and system modeling needed to support its choices, including user and environment representations, backbone LLMs, and agent harnesses. Lastly, we propose PROACTIVITY-GYM, a simulationbased evaluation testbed including multi-day scenarios, stateful environments, and persona-conditioned simulated users that can evaluate the consequences of proactive assistance across interactions. Evaluations across 23 model-harness configurations uncover substantial performance gaps across 3T and reveal that LLM judges often conflate task capability and trust. A human study with 30 participants demonstrates the importance of the joint 3T optimization: participants show sharp trust declines after intervention misalignment despite correct outcomes, and prefer sleep-time assistance, even when imperfect, to preserve ongoing focus. Together, these findings support designing and evaluating proactive agents through the joint consideration of useful work, compute allocation, and evolving user trust.
One-sentence Summary
Researchers from KAIST and the University of Minnesota propose foundations for designing, realizing, and evaluating proactive LLM agents around the joint 3T principles of Task Capability, Temporal Allocation, and Trust, and they introduce PROACTIVITY-GYM, a simulation-based testbed whose evaluations across 23 model-harness configurations and a 30-participant human study reveal sharp trust declines after intervention misalignment and a preference for sleep-time assistance even when imperfect.
Key Contributions
- Introduces a blueprint for proactive LLM agents built on three joint objectives: Task Capability, Temporal Allocation, and Trust (3T), which are often conflated or overlooked in existing work.
- Formalizes the proactivity design space along five dimensions: task scope, anticipation horizon, activation trigger, processing timing, and intervention depth, and specifies the situation and system modeling needed to realize these choices, including user and environment representations, backbone LLMs, and agent harnesses.
- Presents PROACTIVITY-GYM, a simulation-based evaluation testbed with multi-day scenarios, stateful environments, and persona-conditioned simulated users; evaluations across 23 model-harness configurations and a 30-participant human study reveal substantial performance gaps across the 3T objectives and show that task capability alone is insufficient for sustaining appropriate trust.
Introduction
Open-source agent frameworks and open-weight models now make it practical to run personal AI agents on consumer hardware, creating an opportunity for proactive assistance that uses idle compute before the user asks. However, current agent research and deployment remain mostly reactive, and existing evaluations often collapse failures in timing, intervention depth, and user trust into simple task success, which hides why proactive help may be ignored or rejected. The authors address this gap by proposing three joint design principles, Task Capability, Temporal Allocation, and Trust (3T), and by introducing PROACTIVITY-GYM, a multi-day simulation testbed with persona-conditioned feedback; they evaluate 23 model-harness configurations and conduct a 30-participant study showing that task performance alone is insufficient.
Dataset
The authors introduce PROACTIVITY-GYM as a synthetic, dynamic evaluation environment for proactive assistance. The data is purpose-built rather than collected from an existing corpus.
- Composition and scale: It contains 10 multi-day scenarios. Each scenario spans 7 to 10 simulated days and covers professional or everyday domains such as research, shopping, and business.
- Scenario schema: Each scenario includes a user goal, timestamped events and tasks, relevant tools organized into multiple sessions, latent user needs, and competing tasks with resource, availability, and deadline constraints.
- Persona variants: For each scenario, the authors construct three user personas that differ in preferred intervention depth. This yields 30 persona-scenario configurations.
- User simulator: A persona-conditioned user simulator provides implicit and explicit feedback, which can change later interactions and should be used to infer trust state.
- Dynamic and NOOP cases: Scenarios are stateful and action-dependent. Follow-up events depend on earlier user and agent actions and the resulting state. The authors also include NOOP cases where a proactive opportunity is irrelevant or already resolved and no intervention is needed.
- Processing and filtering: No filtering is described because the scenarios are manually constructed for evaluation. The main processing design is the addition of dynamic constraints and persona-conditioned feedback.
- Usage: The authors use this environment to evaluate proactive assistance under the 3T objectives, combining a simulated clock, stateful tool environments, and scheduled events. It tests how proactive actions affect later work, task progress, resource availability, and user state over time.
Method
The authors define proactivity as an agent's anticipatory behavior intended to address and fill the gaps of user needs without an explicit request. To ensure useful assistance, they introduce three design objectives (3T): Task Capability (TC), Temporal Allocation (TA), and Trust (TR). Task Capability is the ability to anticipate relevant user needs and correctly perform useful work. Temporal Allocation is the ability to allocate compute over time according to resource availability and when results are needed, distinguishing between interaction time and sleep time. Trust is the extent to which a user is confident in and willing to rely on an agent's recommendations, often proxied by intervention depth.
As shown in the figure below, these objectives jointly shape assistance in a workflow, such as deferring compute-heavy tasks to sleep time to avoid disrupting the user.
To guide agent design and evaluation, the authors combine these objectives into a weighted formulation:
A∈AmaxλTCSTC(A)+λTASTA(A)+λTRSTR(A)where A is the set of agent designs, S scores the design on each objective, and λ are weights summing to 1.
The authors organize proactive assistance into a design space spanning three connected decisions: what work to pursue, when to initiate and process it, and how far to proceed. Refer to the framework diagram for the five dimensions.
Identifying useful work involves choosing the task scope (within-task or out-of-task) and anticipation horizon (immediate, within-interaction, or out-of-interaction). Initiating and scheduling work depends on the activation trigger (user, event, or agent) and processing timing (interaction time or sleep time). Choosing intervention depth determines autonomy levels: prepare, suggest, or execute, reflecting estimated trust.
To realize such agents, the authors propose a system architecture connecting context, decisions, and feedback through situation and system modeling.
Situation modeling integrates interaction history, memory, and external observations to represent the current user and environment state, tracking goals, attention, preferences, trust, and resource availability. System modeling specifies how the backbone LLM and harness support proactive work. The harness orchestrates model calls, tools, memory, and permissions, tracking pending tasks and allocating compute.
At decision step t, the components select an action from the available context ct:
zt=R(ct),at∼πA(⋅∣zt,ct),ct+1=U(ct,at,ot+1)Here, R builds the situation representation zt, πA selects an action, and U updates the context with the action trace and new observations ot+1, including user feedback.
Experiment
The experiments evaluate nine models across OpenClaw, Claude Code, and Codex harnesses using task completion, temporal allocation, intervention-depth alignment, and trust ratings, and include a 30-person human study with scenario reviews and simulated week-long logs. Benchmark results show that even the strongest agents leave substantial gaps in temporal allocation and intervention alignment, and that task completion gains do not consistently improve those dimensions; LLM-based trust ratings also appear tolerant of misaligned interventions. The human study reinforces that people value correctness but strongly prefer intervention alignment, that deferring assistance to periods like sleep time can make proactive help more acceptable, and that trust losses from misalignment are larger and harder to recover than gains.
The example scenario contrasts an agent that asks for confirmation before adding a language review session and reminder with an agent that performs the same correct actions without approval. Participants valued task correctness highly, but they also valued alignment with the user's standing request, and trust was sensitive to intervention misalignment. Proactive assistance was more acceptable when scheduled for low-disruption periods such as sleep time, even if it required later correction. When content quality was held constant, participants preferred agents that respected the user's approval preference. Misaligned intervention behavior reduced trust even when the agent's task outcomes remained correct. Trust declined with repeated misalignment and recovered only partially after a return to aligned behavior. Participants often preferred sleep-time assistance over immediate assistance, even when the immediate action was correct and required little revision.
This experiment compared agents that requested confirmation before adding a language review session and reminder against agents that acted identically but without approval. It examined how proactive assistance, timing, and alignment with user preferences affect trust and acceptability. Participants valued both task correctness and respect for a standing request; trust dropped with repeated misaligned interventions even when outcomes were correct, and recovered only partially after realignment. They often preferred low-disruption sleep-time assistance over immediate correct actions that required little revision.