VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?
Xiaohongshu Dots Studio, Evolvent AI
Xiaohongshu Dots Studio · Evolvent AI
cs.CL, cs.AI
Submitted: 2026-08-16
Updated: 2026-08-18
Code: https://github.com/evolvent-ai/Terrarium
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World? Abstract Large language model (LLM) agents are increasingly deployed as personal assistants.
Terminology
Summary
VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?
Abstract
Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments. Everyday life assistance is different. A task runs for weeks rather than minutes. The world keeps changing while the agent is not being prompted. Many constraints are never stated outright. An agent that merely answers the request in front of it will fail at such a task. What is needed instead is an agent that stays proactive and consistent. It decides on its own when to act, when to ask, and when to stay silent. It notices changes that nobody announced. It keeps one plan coherent from the first day to the last. No current benchmark measures this. We introduce VibeLifeBench, a benchmark of 200 long-horizon tasks across ten everyday-life domains. Each task is a scripted multi-week timeline in a simulated world of 22 mock services. The world advances on its own clock, and many of its changes are silent, so only an agent that re-inspects the world discovers them. Every task is graded by fine-grained, weighted checks that read only what the agent actually left behind, covering the end state, the timeliness of its actions, and whether it upheld the implicit constraints. We evaluate seven frontier models. All of them score low, which shows how far current agents are from assisting with real life. We will open-source all tasks, environments, and the evaluation framework.
Introduction
Large language model (LLM) agents already act in the real world. Equipped with tools and external services, they fix bugs in live repositories, operate computers, and produce professional deliverables. This progress has concentrated on professional settings, above all coding and office work. Everyday life is a different setting, and it has received far less attention, even though it is where an agent is closest to ordinary users and where trustworthy assistance matters most. A life agent is asked to plan a multi-week family trip, to follow a rental dispute through each of its deadlines, or to coordinate a renovation. Such a task runs for weeks, and the agent must hold context and commitments across that span while the situation keeps developing whether or not anyone prompts it.
Real-world life assistance in this sense has three defining properties, all of which current benchmarks overlook. (1) It is proactive. The environment does not wait to be queried, and a competent assistant should not wait to be asked: it must notice on its own that a deadline is near or that a booking has changed, judge whether the moment calls for action, for notifying the user, or for staying silent, and, at the very outset, proactively extract the hard and implicit constraints a request only hints at (a budget cap, a family member’s health condition, a filing window) rather than executing the instruction literally. (2) It unfolds in a dynamic, living world. The environment is stateful and advances on its own, as weather shifts, prices and inventory fluctuate, flights are delayed, venues close, and a companion’s wearable device reports new readings; a competent agent perceives these changes and propagates them into the plan it maintains. (3) It is long-horizon. A task spans a full lifecycle of preparation, execution, and wrap-up across many simulated days, during which the agent must keep an evolving plan self-consistent and always uphold the constraints it was given at the very start.
Existing benchmarks are misaligned with life assistance in this sense along several axes. In domain, they concentrate on working and coding scenarios and overlook the needs of everyday life. In the ability they probe, they mostly measure how well an agent passively carries out a clearly specified task rather than whether it decides on its own, without prompting, when to act and when to stay silent. In their environment, the world they maintain does not change on its own and updates only when the agent acts, so they cannot examine an agent’s perception and propagation of spontaneous change. In task form, they are mostly isolated single tasks rather than long-horizon tasks embedded in a complete lifecycle with intricate dependencies.
Contributions
Our contributions are threefold. (1) We reframe agent evaluation around three under-measured properties of real-world assistance: proactivity on a live clock, operation in a dynamic living world, and long-horizon coherence across a full multi-week lifecycle. (2) We introduce VibeLifeBench, a benchmark of 200 multi-week living-world tasks across ten everyday-life domains, built on 22 mock service backends (288 tool interfaces) and driven by 7,453 scripted events, including mutations that fire with no notification. (3) In our evaluation of strong models we find a large gap, with the strongest model reaching an avg@3 of only 32.5; we will open-source all tasks, environments, and the framework to catalyze research on long-lived, proactive agents.
Design Principles
The central design commitment of VibeLifeBench is that a task is not a prompt but a world with a clock. An agent is placed in an ongoing situation, given a set of tools and one or more personas to serve, and then time begins to advance: the user speaks, external services publish updates, and the underlying state of the world changes whether or not the agent is paying attention. Around this commitment we make three design choices.
(a) The world advances on a virtual clock and changes silently, so that proactivity can be measured. A task is not a static request but a world advancing along a timeline: part of its change happens silently, with no signal to the agent, while later tasks depend on that change. An agent that only reacts to the input in front of it therefore misses these changes, and only an agent that re-inspects the world on its own and propagates the change into the plan it maintains can get them right; staying silent when nothing needs handling is likewise treated as correct behavior.
(b) Tasks embed implicit constraints and safety red lines, so that trustworthiness can be examined. Each task ships a persona and its background material, into which several unstated but binding constraints are planted, along with safety red lines and authorization boundaries: which actions may be taken on the agent’s own, which require asking first, and which are never permitted. Many scenarios also deliberately set up tempting but unsafe shortcuts, which a compliant agent must recognize and refuse or escalate.
(c) Shared service backends with per-task initial state reconcile realism and reproducibility. All tasks share the same stable set of mock services, and each task only configures its own initial data. This lets a large number of tasks reuse consistent tool semantics while remaining fully offline and deterministic, which makes the suite easy to audit and reproduce.
Task Definition
A task is a self-contained directory that can be formalized as a five-tuple τ = W0, E, K, P, R, where the components are as follows.
• W0 is the initial world state, determined jointly by the seed data the task provides for each enabled service backend k.
• E = e1, e2,... is the event timeline, grouped by stage and ordered by timestamp within a stage.
• K ⊆ k1,..., k22 is the set of service capabilities the task enables.
• P is the persona and workspace: the persona, preferences, authorization policy, operations handbook, and other material provided to the agent at the outset.
• R = (ci, wi) i=1 m is the scoring criteria, a set of weighted checks in which each ci is a deterministic predicate over the world and wi > 0 is its weight.
A run of the evaluation drives an agent policy π through the entire timeline. Let Wj denote the world state after the j-th dispatched item is processed. Then Wj = apply(Wj−1, ej, aj), aj ∼ π(· obsj, Hj−1), where aj is the sequence of actions the agent takes in that turn (tool calls, file writes, replies) and Hj−1 is its history and memory. For a mutation no turn is produced, so aj = ∅ and the world changes only because of the event itself. When the run ends, the scoring criteria are evaluated against the end state and the artifacts produced along the way.
Input and output. The agent’s input consists of the workspace files supplied at the start (P), the set of available tools determined by K, and the text of the turn-triggering events that arrive in timeline order. Its output is not a single final answer but the observable trace it leaves over the whole episode: the end state of the backend services (bookings, orders, applications, calendar events, ledgers), the durable files in the workspace, notes, and calendar, the email it sends, and its reply text at each stage. Scoring is based entirely on these observable artifacts.
How the world is simulated
The world consists of 22 mock service backends, which stand for the applications and backends an agent would touch in everyday life. They fall into two categories.
• General services (used by nearly every task): email, calendar, notes, and a notification hub. These are where the agent accumulates commitments, correspondence, and running notes.
• Domain services: banking, credit card, brokerage, travel booking (flight, hotel, rail, car), maps, weather, visa and travel advisories, e-commerce, delivery and logistics, listings, reviews, content communities, legal search, job boards, and health tracking, drawn on according to the domain of the task.
Each service is a small, self-contained product backend that follows a uniform pattern: it defines its own data model, exposes a stable set of tools, and boots from seed data on a cold start. The 22 services together expose 288 tool interfaces. A service is never bound to a single task: a task only configures its own scenario data and event pacing, while the service provides stable tool semantics. This separation is what lets 200 tasks share the same backends while remaining fully offline and reproducible. The agent reads and writes the world through tool calls (for example querying bookings, searching flights, or creating calendar events), while the checkers behind the scoring criteria read the objective end state of the world, or read the workspace artifacts, directly from the host side, without relying on the agent’s self-report.
Event system
The timeline advances in stages. A stage is not a calendar day; it is a checkpoint at which the agent must act or the scorer must inspect the world. A multi-week task may compress a quiet two-week interval into a single stage, and it may expand a busy departure day into several stages. The suite uses a median of 24 stages per task. Within a stage, events are dispatched in timestamp order. Events come in four kinds, and it is the difference between the last kind and the first three that makes proactivity observable.
The first three kinds open a turn and the agent’s text reply is captured; a mutation is applied directly to the relevant service state, produces no turn, and changes the world silently. The gap between the first three kinds and the last is the core mechanism of the benchmark: a mutation alters the world with no accompanying message, no notification fires, and nothing is pushed to the agent’s turn. Only an assistant that is both persistent (remembering a booking made days earlier) and proactive (re-checking that booking’s status without being asked) discovers the discrepancy in time. The suite embeds 1,483 background mutations.
A representative timeline strings these primitives into a coherent story: an opening user message that states the goal and budget; a run of world advisories about visas and vaccines; a mutation that breaks down the itinerary vehicle mid-trip; a notification reminder at the trip’s midpoint; a flight delay applied silently near the end; and a phishing email injected as yet another mutation. Each of these probes whether the agent noticed and responded without being told.
Persona and workspace
Each task ships a workspace that stands for the context a long-term assistant would already hold: exactly who it serves, that person’s preferences and circumstances, and the constraints and authorization boundaries the user has laid out in advance. This material turns being a competent assistant into concrete, checkable behavior, and it lets a task probe in particular whether the agent holds an authorization boundary: many scenarios deliberately tempt the agent with an unsafe shortcut, such as submitting a claim on the user’s behalf, emailing a passport scan to an unfamiliar address, or wiring a fee to a personal account, and a compliant agent must refuse or escalate rather than take the easy path. In a 20-day trip to Japan with the user’s parents, for example, the mother’s passport has less remaining validity than entry requires, the diabetic father’s insulin must be carried on board with a doctor’s letter, an unbreakable hard budget is set, and a phishing email disguised as a visa expedite fee arrives; none of these is stated outright, yet each must be surfaced and upheld over the entire task.
Reward
Each task is scored by a set of scoring criteria, a collection of weighted checks. Each check is a deterministic predicate over the world that reads only observable artifacts, namely the end state of the services, workspace and notes files, sent email, and the agent’s captured replies, and never reads the model’s hidden reasoning. Checks are organized into three tiers: per-stage checks credit timely behavior at the moment it is due; cross-stage checks encode constraints that must hold across the whole episode (such as the budget cap and adherence to safety red lines); and final checks inspect the world and artifacts the agent ultimately leaves behind.
A task’s score is the fraction of check weight it earns, reported on a 0 to 100 scale, so partial competence over a long horizon is credited rather than only end-of-task success or failure. The weights are deliberately uneven: a single critical failure, such as leaking personal information or breaching the hard budget, costs far more than a cosmetic oversight, and safety and hardening checks carry the largest weights, so an agent cannot mask an unsafe action by completing many routine sub-tasks. Such high-weight checks usually require a durable artifact rather than a transient chat reply.
Construction Pipeline
The VibeLifeBench test set is built by a uniform pipeline whose goal is that each task is both close to real life and hard enough, while keeping scoring objective and free of exploitable holes. The whole process unfolds around the task representation, passing in turn through scenario design, timeline authoring, environment instantiation, and authoring of the scoring criteria.
Scenario and persona. Each task originates in a real everyday-life scenario, for example a trip abroad with one’s parents, a rental dispute, or the coordination of a renovation. On top of the scenario we fix the persona to be served and its workspace, including the user’s identity, preferences, and circumstances, and a set of constraints and authorization boundaries laid out in advance. Most of these constraints are given implicitly and must be identified by the agent from the background material and upheld throughout the task; among them we deliberately include several safety red lines, which examine the agent’s trustworthiness in the face of tempting shortcuts.
Timeline authoring. The scenario is unrolled into a stage-organized event timeline in which the four event kinds are interleaved. To make proactivity a prerequisite for earning credit, a substantial share of the world’s changes are deliberately arranged as mutations: they trigger no turn, yet later stages depend on them, so only an agent that re-inspects the world on its own perceives them and reacts in time.
Environment instantiation. For each service the task enables we prepare its own initial data, so the world boots from a credible, queryable state on a cold start. The initial data is seeded with ample distractors so that the answer cannot be guessed shallowly, and all key entities are made discoverable through tool calls rather than hard-coded into the scoring criteria, which avoids false negatives in which an agent that could have solved the task is judged to fail.
Authoring the scoring criteria. Finally, for each task we author stage-aware, weighted scoring criteria that span several evidence dimensions. Every check is required to make a discriminating judgment, for example computing a concrete value or verifying a real state change, rather than passing on the mere appearance of a keyword, so as to minimize the chance that scoring is fooled by surface text.
Data Distribution
The statistics in this section are all computed from the released evaluation set of 200 tasks. The 200 tasks are evenly distributed across ten domains (20 each), so that the aggregate score is not dominated by a single domain; beneath this balance, the tasks are diverse in horizon, event density, service breadth, and the size of the scoring criteria.
Horizon. Tasks are long by construction. The median simulated horizon is 29 days, with the middle of the distribution spanning roughly 25 to 35 days and a long tail reaching several months (the longest is about 111 days, and 17 tasks exceed 60 days). Because horizon is expressed in stages rather than calendar days, a two-month task need not carry many checkpoints: the median task uses 24 stages, compressing dormant intervals into sparse checkpoints while retaining the multi-week temporal structure a persistent agent must survive.
Event density. Across the suite we script 7,453 events, a median of 36 per task. Their composition reflects a world that mostly advances on its own rather than waiting to be prompted: user messages account for 2,247 events (about 30.1%), and the rest are mostly environment-driven, with notifications and scheduled reminders at 1,925 (25.8%), world observations at 1,798 (24.1%), and mutations at 1,483 (19.9%). That is, about 69.9% of events are environment-driven rather than user-prompted. The 1,483 mutations trigger no agent turn and alter the world silently, so only an agent that re-inspects the world on its own notices them in time.
Service breadth. A task recruits a median of 7 mock services (up to 12), and all 22 services are exercised somewhere in the suite. Usage is deliberately long-tailed. Three services are near-universal, namely email (198/200 tasks), calendar (195), and notes (162), because they are where the agent keeps its records and where commitments, correspondence, and running notes accumulate over the horizon. The notification hub appears in 120 tasks. Domain backends are used more sparingly and cluster by domain: visa, flight, and hotel in travel; banking, brokerage, and credit card in finance; legal search in litigation; and listing and review platforms in shopping.
Size of the scoring criteria. Grading is fine-grained: the suite contains 12,261 weighted checks, a median of 58 per task (from about 31 for the tightest scenarios to 135 for the most elaborate). Checks are distributed across the per-stage, cross-stage, and final tiers, so credit accrues along the whole timeline rather than only at the end. Per-stage checks are the large majority by count (80.8%), but the cross-stage and final tiers carry disproportionate weight: together they are 19.1% of the checks yet 26.8% of the total weight, which is why a single safety or hardening failure costs more than many routine sub-tasks.
The domains differ in meaningful ways: career tasks have both the longest horizon (a median of 48 days) and the highest event density (a median of 44 events), and finance has the densest scoring criteria (a median of 94 checks).
Evaluation
VibeLifeBench grades by the observable outcomes an agent leaves behind, against fine-grained, stage-aware scoring criteria. This section describes how a single run is executed, how the score of a single task is computed and aggregated across runs, and the harness used for evaluation.
Executing a run. A run drives the agent stage by stage along a task’s timeline. Within each stage, user messages, world observations, and notifications trigger an agent turn and its reply is recorded, while a mutation alters the world state directly without triggering a turn; once a stage’s events are processed, the scoring criteria run once against the current world state. Scoring relies only on the observable artifacts the agent leaves behind and never reads its internal reasoning. Because a mutation alters state the agent was never told about, an agent that does not proactively re-inspect the world is scored against a reality it never observed.
Scoring. On a single task, the score of one run is the ratio of the weight of the checks it passes to the total weight of all checks, reported on a 0 to 100 scale. Because an agent’s output is stochastic, we run each task three times: we first summarize within a task by taking the mean, maximum, and minimum of the three scores, then average these equally across all tasks to obtain avg@3, max@3, and min@3, where avg@3 reflects overall skill and max@3 and min@3 give the best and worst cases. We also report the standard deviation of a task’s three scores (averaged across tasks) to measure the stability of scoring under repeated runs.
Evaluation infrastructure and agent harness. VibeLifeBench is implemented on Terrarium, a multi-turn evaluation infrastructure for agents operating in living environments. Terrarium orchestrates stage-wise task execution and provisions the isolated sandbox in which the mock services, checkers, and agent run. Within this infrastructure, all evaluated models are run under the openclaw harness. It provides the task’s mock services, workspace, and system prompt to a tool-using agent and runs each model at its strongest reasoning setting, so that scores reflect the underlying model rather than scaffold tuning. Each run executes in an isolated sandbox that hosts the mock services and the agent together, keeping the environment offline and reproducible; the harness records for every run its final score, the pass or fail status of every check, and the full agent trajectory, so that aggregate results can be traced back to specific behaviors.
Experiments
We evaluate seven contemporary strong models on the full evaluation set of 200 tasks: Claude Opus 5, GPT-5.5, Gemini 3.5 Flash, Claude Opus 4.8, GLM-5.2, Kimi-K2.6, and DeepSeek-V4-Pro. All models use the same native tool-calling scaffold at their strongest reasoning setting, and each task is run three times, reporting avg@3, max@3, min@3, and the within-task standard deviation.
Main results. Every evaluated model scores low. The strongest, Claude Opus 5, reaches an avg@3 of only 32.5, and even its best-of-3 ceiling (max@3) is no more than 41.2, while the weakest, DeepSeek-V4-Pro, reaches only 21.1. This shows that contemporary agents, though quite fluent at single-turn tool use, are still far from able to manage life affairs proactively and persistently over weeks in a world that evolves on its own; the low absolute scores are a direct measurement of the gap between the passive, single-turn tool use of current agents and the persistent, proactive behavior that long-horizon assistance in a living world requires.
Larger or newer general models do not clear this bar. All seven frontier models fall within a narrow band from 21 to 33 (Claude Opus 5 > GPT-5.5 > Gemini 3.5 Flash ≈ Claude Opus 4.8 > GLM-5.2 > Kimi-K2.6 > DeepSeek-V4-Pro), separated from one another by far less than they are from competence. Strong tool-calling ability, which these models demonstrate in professional settings such as coding and office work, therefore does not automatically transfer to managing life affairs over weeks.
The unreliability of the models further underscores this gap. Every model has a min@3 of at most 23.8, and scores still fluctuate noticeably across repeated runs of the same task (the within-task standard deviation reaches 10.0); even when a run happens to get things right, it is hard to reproduce. The persistence and self-consistency on which a trustworthy long-term assistant depends are exactly where current models are weakest.
The breadth of real-life domains is itself a challenge. Even the strongest model, Claude Opus 5, varies widely across the ten domains (from 21.8 on team building to 51.1 on shopping), and the easy-to-hard pattern is highly consistent across models: shopping, travel, and renovation are relatively tractable, whereas team building, rental, and exam preparation are hardest. No model is competent across all life domains, which shows that the difficulty is inherent to the domains and that being strong in one place does not imply broad usability.
Token and interaction cost. We record context read, output tokens, tool calls, and turns per run (averaged over runs). The scale of investment correlates broadly with capability: the top-scoring Claude Opus 5 generates the most (about 325k output tokens per run) and sustains a deep interaction loop (316 tool calls and 210 turns), while DeepSeek-V4-Pro reaches 21.1 on the smallest context budget, a frugal-but-steady profile. However, spending more tokens does not by itself guarantee a higher score: Gemini 3.5 Flash reads the most context (41.2M per run) and takes the most turns (227) yet lands mid-pack, whereas GPT-5.5 issues the most tool calls (332 per run) on the smallest output budget and still places second. How the effort is spent, that is, whether state is persisted durably and constraints are upheld, matters more than the sheer amount spent.
Analysis
This section examines why the models score low and maps the failures onto the core properties the benchmark is built around. All statistics are computed from the per-check results of each run and the agent trajectories.
Failure modes. By tier, the highest-weight hardening layers are the hardest. Checks are divided into three tiers: per-stage, cross-stage, and final. Every model has the lowest pass rate on the cross-stage and final tiers, which carry the largest weight (19.1% of the checks but 26.8% of the total weight). This directly explains why the absolute scores are depressed.
By capability axis, proactivity and persistence are the largest weaknesses. We group failures by the semantics of the checks into a few capability axes. The axes are assigned by keyword matching over check names, so they are indicative rather than exact. Proactivity and persistence are consistently among the lowest axes, and no model exceeds 33.6 on either. Persistence and bookkeeping is the largest identified source of failure, accounting for 22.2% of all failed checks pooled across the seven models and a remarkably stable 22.0% to 23.2% for every individual model, though a further 41.8% of failures fall outside the named categories. The checks that require leaving a durable gating artifact and linking state across stages, together with checks that require committing a key booking to the backend, pass almost never for any model. This is not because the models fail to write at all; it is because what they write rarely forms the specific, cross-stage-linked artifacts the scoring criteria require.
Analysis along the core dimensions. Proactivity and persistence. The proactivity axis is low across all models (16.0 to 33.6), the persistence and bookkeeping axis is likewise low (18.9 to 28.0), and checks that require a durable artifact pass almost never. This shows that the models tend to respond passively and once, and fail to maintain cross-stage, auditable, well-formed persistent state. It echoes the token analysis: models that are willing to generate more and persist more structurally (Claude Opus 5) score higher.
Dynamic living world. A mutation triggers no turn, so acting on it requires the agent to re-inspect the world and propagate the change into its plan. The propagation and recovery axis, which covers the checks that depend on such downstream state, reaches only 18.5 to 32.0 across models and sits alongside proactivity and persistence at the bottom. The models routinely miss changes that nobody announced and do not re-inspect the world when not told to; adapting to a living world is a core weakness.
Long-horizon coherence. We measure the pass rate by the normalized position of a check along the timeline. Every model’s per-stage pass rate in the last third of the timeline is 10 to 15 points below the first third: Claude Opus 5 52.0 to 37.7, GPT-5.5 47.4 to 33.1, Claude Opus 4.8 46.8 to 34.6, GLM-5.2 45.9 to 32.3, Gemini 3.5 Flash 42.5 to 32.6, DeepSeek-V4-Pro 42.6 to 27.4, and Kimi-K2.6 42.2 to 27.1. The decline is not monotonic, but it holds for every model, and even the strongest is not exempt. Notably, task scores correlate only weakly with size (Spearman correlations of +0.28 with the number of events and +0.02 with the horizon, and −0.26 with the number of stages), which indicates that the difficulty is driven mainly by sustaining the staged constraints rather than simply by tasks being longer.
Breadth across life domains. No model is competent across all ten domains, and the easy-to-hard ordering is consistent across models. The breadth a real life assistant needs is itself a challenge.
Directions for improvement. Taken together, the analysis points to four directions. (i) Persistence: agents should explicitly write state into notes, the calendar, and workspace files, and maintain cross-stage auditable artifacts rather than stopping at a chat reply. (ii) Proactive perception and propagation: agents should re-inspect the world when not prompted, reconcile it, and fold mutations into the plan. (iii) Hardening and safety: targeted work on refusing phishing, protecting personal data, respecting authorization boundaries, and holding budget caps, since the high-weight cross-stage and final checks have the lowest pass rates. (iv) Long-horizon stability: suppressing the decay of pass rate over time, so that an agent still upholds its earlier commitments and constraints late in a task.
Related Work
Coding and working agents. A large body of benchmarks focuses on coding and on office or professional work. On the coding side, they ask an agent to fix defects in a repository, pass a given test suite, make cross-file edits, or maintain a continuously evolving codebase over a series of interdependent milestones; on the office and knowledge-work side, they have the agent read and reconcile heterogeneous material such as spreadsheets, documents, slides, and databases, and produce reports, analyses, or other professional deliverables that are then graded check by check. Challenging as these tasks are in technical depth and cross-file dependency, they mostly share the same setup: each task is given by an explicit instruction with a clear goal and deliverable, the agent acts as a passive executor that completes it in one shot, and the world it inhabits is a reproducible sandbox that changes only through the agent’s own actions, with no independently occurring external events. In contrast, VibeLifeBench targets the everyday-life domain, embeds each task in a world that evolves on its own virtual clock, and requires the agent to decide on its own when to act, when to contact the user, and when to stay silent, and to identify implicit constraints from a request that only hints at them, rather than passively executing a given instruction in a static working or coding environment.
Long-horizon agent benchmarks. Another line of work pushes evaluation toward longer time spans, broadly in two directions. One unrolls a single task into a very long trajectory with many tool calls and deep dependencies, testing an agent’s planning, memory, and adherence to earlier decisions over one sustained process; the other models long-term multi-task interaction in which user preferences drift over time or new information is injected in stages, testing personalization and the continual revision of beliefs, and some further provide asynchronous, event-driven environments in which time advances while the agent reasons and external events fire on a schedule. These efforts do extend the time span, but usually along a single axis: the environment is either deterministic and static, or it changes only between tasks, when triggered by the agent, or as a controlled perturbation, rather than evolving independently while a single task is in progress; the goal at each step is typically still given explicitly, and scenarios are often bounded and resettable. In contrast, the long horizon of VibeLifeBench is a single continuous multi-week timeline in which the world advances on its own virtual clock and changes state through mutations, later stages depend on these unannounced changes, and the agent must therefore re-inspect the world on its own, propagate the changes into a continuously maintained plan, and uphold the constraints given at the outset throughout, rather than completing a single deep task or coping with cross-task preference drift.
Conclusion
We have introduced VibeLifeBench, a benchmark for life-domain agents. It organizes each task as a multi-week living world that evolves on its own, and it uses stage-aware, weighted scoring criteria to jointly examine end-state correctness, timely proactive behavior, and faithful propagation of world changes, thereby bringing proactivity, living-world adaptation, and long-horizon coherence, three properties overlooked by existing evaluations, into a single measurement.
Evaluating contemporary frontier models shows that they remain far from a trustworthy long-term life assistant: the models fail to persistently maintain cross-stage, auditable state, they often miss mutations in the living world, their coherence decays as a task advances, and none is reliably competent across all life domains. We will open-source all tasks, environments, and the evaluation framework to advance research on proactive, persistent life agents.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems:
-
Improvement: Add a periodic self-triggered
world audit
mechanism that re-queries external services (calendar, email, bookings, health trackers) at scheduled intervals, even without user prompts. -
Capability: The AI can detect silent changes (e.g., flight delays, price changes, venue closures) that occur without notifications, and update its plans accordingly.
-
Improvement: Implement a durable, structured memory system that writes key commitments, constraints, and progress to external artifacts (notes files, calendar events, workspace documents) after every significant action.
-
Capability: The AI maintains auditable, linked state across days or weeks, so it can recall and uphold early constraints (e.g., a budget cap, a family member's health condition) even after many intervening events.
-
Improvement: Add a pre-execution
constraint mining
step that parses all background material (persona, preferences, authorization policies) to extract unstated rules (e.g., passport validity, medication requirements, spending limits) and stores them as explicit, checkable rules. -
Capability: The AI can identify and respect safety red lines and authorization boundaries (e.g., refusing to wire money to unknown accounts, not submitting claims without permission) even when the user's request is ambiguous or tempting.
-
Improvement: Train the AI to detect discrepancies between its internal plan and the actual world state (e.g., a booking that no longer exists) and to automatically trigger a
reconciliation
workflow that updates dependent tasks. -
Capability: The AI can recover from unannounced changes (e.g., a broken rental car mid-trip) by re-planning downstream steps without being told, rather than continuing with outdated assumptions.
-
Improvement: Add a
commitment ledger
that tracks all promises and deadlines, with periodic self-checks (e.g.,Have I upheld the budget cap this week?
) and automatic reminders to re-verify earlier constraints. -
Capability: The AI's performance does not decay over time; it continues to uphold early-stage commitments and constraints even in the final third of a multi-week task.
-
Improvement: Implement a
when-to-speak
classifier that evaluates whether a detected change requires user notification, action, or silence, based on the user's stated preferences and the severity of the change. -
Capability: The AI can decide autonomously when to contact the user (e.g., a critical flight delay) and when to stay silent (e.g., a minor price fluctuation), avoiding both over-notification and missed critical alerts.
-
Improvement: Add a
red-line filter
that flags any action involving sensitive data (passport scans, bank transfers, personal information sharing) and requires explicit user confirmation unless the action was pre-authorized in the workspace. -
Capability: The AI can recognize and refuse phishing attempts (e.g., a fake visa fee email) and unsafe shortcuts, even when they appear as legitimate mutations in the environment.
-
Improvement: Train on diverse life domains (travel, finance, rental disputes, health) with shared service semantics, so the AI learns transferable patterns of proactive behavior rather than domain-specific scripts.
-
Capability: The AI can handle a wide range of everyday-life tasks (from team building to exam prep) with consistent competence, rather than excelling only in narrow areas.
-
Improvement: Use self-consistency checks (e.g., run multiple internal plan variants, verify against world state) to reduce variance across repeated runs of the same task.
-
Capability: The AI produces reliable, reproducible results even in complex, multi-week scenarios, minimizing the current 10-point within-task standard deviation.
-
Improvement: Implement a
living-plan
module that automatically re-optimizes the entire plan whenever a mutation is detected, rather than patching individual steps. -
Capability: The AI can maintain a coherent, self-consistent plan from day 1 to day 29, even when multiple silent changes occur simultaneously (e.g., a visa delay and a hotel cancellation).
What the improved AI system can do overall: It can serve as a trustworthy, long-term life assistant that operates proactively over weeks, detects and adapts to unannounced world changes, upholds implicit constraints and safety boundaries, maintains auditable state across stages, and remains reliable and consistent across diverse real-life domains—without needing constant user prompting.
Sources
- SWE-Milestone: Evaluating AI Agents on Continuous Software Evolution
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- APEX-Agents
- EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments
- Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments
- JobBench: Aligning Agent Work With Human Will
- Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies
- UltraHorizon: Benchmarking Agent Capabilities in Ultra Long-Horizon Scenarios
- ClawBench: Can AI Agents Complete Everyday Online Tasks?
- UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks
- Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents
- WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
- CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents
- ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
- ClawArena: Benchmarking AI Agents in Evolving Information Environments
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents
- DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks
- SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation
- PPTBench: Towards Holistic Evaluation of Large Language Models for PowerPoint Layout and Design Understanding
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering