The Right Lesson at the Right Step: Deriving Control Updates for Self-Evolving Agents
cs.CL
Submitted: 2026-09-28
Updated: 2026-09-28
Terminology
Sources
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
- Recursive Partitioning for Heterogeneous Causal Effects
- Program Synthesis with Large Language Models
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- CausalFlow: Causal Attribution and Counterfactual Repair for LLM Agent Failures
- Reinforcement Learning for Long-Horizon Interactive LLM Agents
- Evaluating Large Language Models Trained on Code
- AutoManual: Constructing Instruction Manuals by LLM Agents via Interactive Environmental Learning
- Trace is the Next AutoDiff: Generative Optimization with Rich Feedback, Execution Traces, and LLMs
- Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution
- AutoGuide: Automated Generation and Selection of Context-Aware Guidelines for Large Language Model Agents
- Automated Design of Agentic Systems
- DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
- Self-Refine: Iterative Refinement with Self-Feedback
- CLIN: A Continually Learning Language Agent for Rapid Task Adaptation and Generalization
- Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs
- Symbolic Learning Enables Self-Evolving Agents
- ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory
- EvoClawBench: Can Agents Learn Reusable Skills from Their Own Runs?
- AgentSquare: Automatic LLM Agent Search in Modular Design Space
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering