What Happens During Autonomous Deep Research After the User Steps Away?
cs.AI
Submitted: 2026-09-27
Updated: 2026-09-27
Terminology
Sources
- DeepResearchGym: A Free, Transparent, and Reproducible Evaluation Sandbox for Deep Research
- Characterizing Deep Research: A Benchmark and Formal Definition
- Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale
- IntentRL: Training Proactive User-intent Agents for Open-ended Deep Research via Reinforcement Learning
- ClarifyMT-Bench: Benchmarking and Improving Multi-Turn Clarification for Conversational Large Language Models
- AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
- AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios
- PersonaFeedback: A Large-scale Human-annotated Benchmark For Personalization
- HiL-Bench (Human-in-Loop Benchmark): Do Agents Know When to Ask for Help?
- AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment
- ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry
- Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs
- PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants
- Counterfactual Trace Auditing of LLM Agent Skills
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection