Shockingly Simple Self-retrospection Improves Agentic Models Without RL
cs.AI, cs.CL
Submitted: 2026-09-28
Updated: 2026-09-28
Terminology
Sources
- L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
- A General Language Assistant as a Laboratory for Alignment
- SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents
- Improving Code Generation by Training with Natural Language Feedback
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- FrogNano: Training a 4B Coding Agent via Online Task Synthesis
- Dialogue Learning With Human-In-The-Loop
- AvalonBench: Evaluating LLMs Playing the Game of Avalon
- Strategist: Self-improvement of LLM Decision Making via Bi-Level Tree Search
- DISC: Dynamic Decomposition Improves LLM Inference Scaling
- HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation
- Procedural Memory Distillation: Online Reflection for Self-Improving Language Models
- Self-Refine: Iterative Refinement with Self-Feedback
- Privileged Information Distillation for Language Models
- Training Language Models with Language Feedback at Scale
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Experiential Reinforcement Learning
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Learning by Distilling Context
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection