Smarter by the Moment: Environment-Driven Dynamic Policies for Continual LLM Improvement
cs.CL
Submitted: 2026-09-15
Updated: 2026-09-20
Comments: 25 pages, 13 figures. Accepted to the Conference on Language Modeling (COLM) 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large Language Models (LLMs) have achieved remarkable progress across diverse domains, but continual adaptation to evolving tasks and environments remains a key challenge.
Terminology
Abstract
Large Language Models (LLMs) have achieved remarkable progress across diverse domains, but continual adaptation to evolving tasks and environments remains a key challenge. Existing memory-augmented approaches retrieve individual past examples as direct references, but do not explicitly synthesize actionable strategies from them, causing the same types of errors to recur. We propose Dynamic Retrieval-based Policy Generation (DRPG), a framework that integrates memory-based retrieval with a dynamic policy generator, leveraging historical data and environment feedback to produce task-specific policies for continual LLM improvement. We evaluate DRPG across six benchmarks spanning text-to-SQL, question answering, medical diagnosis, and Python programming, using seven LLMs from both proprietary and open-weight families. DRPG outperforms strong baselines across most datasets and models. Further analysis demonstrates that DRPG's policy generation is robust to retrieval strategy, operates effectively without prior policy continuity, and can leverage smaller or cross-family models as cost-efficient policy generators. We also find that the benefit of policy-level guidance depends on task characteristics, offering practical insights into when and under what conditions this mechanism is most effective.
Sources
- Learning From Mistakes Makes LLM Better Reasoner
- A Survey on the Optimization of Large Language Model-based Agents
- The Llama 3 Herd of Models
- Gemini: A Family of Highly Capable Multimodal Models
- Gemma 4 Technical Report
- Scaling Laws for Forgetting When Fine-Tuning Large Language Models
- Self-evolving Agents with reflective and memory-augmented abilities
- Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?
- The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities
- Revisiting Dynamic Evaluation: Online Adaptation for Large Language Models
- Proximal Policy Optimization Algorithms
- How to Alleviate Catastrophic Forgetting in LLMs Finetuning? Hierarchical Layer-Wise and Element-Wise Regularization
- Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory
- Two-stage LLM Fine-tuning with Less Specialization and More Generalization
- Larger language models do in-context learning differently
- HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
- Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task
- CoSQL: A Conversational Text-to-SQL Challenge Towards Cross-Domain Natural Language Interfaces to Databases
- In-Context Principle Learning from Mistakes
- Agent-Pro: Learning to Evolve via Policy-Level Reflection and Optimization
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering