Ask Now, Use Later: Benchmarking the Proactivity Gap in Long-Lived LLM Agents
cs.CL
Submitted: 2026-05-27
Updated: 2026-09-15
Comments: Accepted to EMNLP 2026 Main Conference
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Springdrift: An Auditable Persistent Runtime for LLM Agents with Case-Based Memory, Normative Safety, and Ambient Self-Perception
- PIRA-Bench: A Transition from Reactive GUI Agents to GUI-based Proactive Intent Recommendation Agents
- KnowU-Bench: Towards Interactive, Proactive, and Personalized Mobile Agent Evaluation
- BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental Design
- LifeSim: Long-Horizon User Life Simulator for Personalized Assistant Evaluation
- PersonalAlign: Hierarchical Implicit Intent Alignment for Personalized GUI Agent with Long-Term User-Centric Records
- Latent Preference Modeling for Multi-Session Personalized Tool Calling
- Proactive Agent Research Environment: Simulating Active Users to Evaluate Proactive Assistants
- $\pi$-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows
- MemGPT: Towards LLMs as Operating Systems
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering