AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?
cs.AI, cs.LG
Submitted: 2026-07-31
Updated: 2026-09-27
Code: https://github.com/microsoft/Sico
Terminology
Sources
- Position: Agentic Evolution is the Path to Evolving LLMs
- A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
- MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory
- AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution
- Memento-Skills: Let Agents Design Agents
- Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses
- Humanity's Last Exam
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
- OpenAI GPT-5 System Card
- Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning
- Collaborative Multi-Agent Test-Time Reinforcement Learning for Reasoning
- TTSR: Test-Time Self-Evolving via Reflection
- TTCS: Test-Time Curriculum Synthesis for Self-Evolving
- Test-Time Learning with an Evolving Library
- Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents
- TARSE: Test-Time Adaptation via Retrieval of Skills and Experience for Reasoning Agents
- TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking
- Self-Improving LLM Agents at Test-Time
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection