CLEAR: Context Augmentation from Contrastive Learning of Experience via Agentic Reflection
cs.AI
Submitted: 2026-04-08
Updated: 2026-09-11
Code: https://github.com/awslabs/CLEAR
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Large language model agents rely on effective model context to obtain task-relevant information for decision-making.
Terminology
Abstract
Large language model agents rely on effective model context to obtain task-relevant information for decision-making. Many existing context engineering approaches primarily rely on the context generated from the past experience and retrieval mechanisms that reuse these context. However, retrieved context from past tasks must be adapted by the execution agent to fit new situations, placing additional reasoning burden on the underlying LLM. To address this limitation, we propose a generative context augmentation framework using Contrastive Learning of Experience via Agentic Reflection (CLEAR). CLEAR first employs a reflection agent to perform contrastive analysis over past execution trajectories and summarize useful context for each observed task. These summaries are then used as supervised fine-tuning data to train a context augmentation model (CAM). Then we further optimize CAM using reinforcement learning, where the reward signal is obtained by running the task execution agent. By learning to generate task-specific knowledge rather than retrieve knowledge from the past, CAM produces context that is better tailored to the current task. We conduct comprehensive evaluations on the AppWorld and WebShop benchmarks. Experimental results show that CLEAR consistently outperforms strong baselines. It improves task completion rate from 72.62% to 81.15% on AppWorld test set and averaged reward from 0.68 to 0.74 on a subset of WebShop, compared with baseline agent. Our code is publicly available at https://github.com/awslabs/CLEAR.
Sources
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
- Experiential Reflective Learning for Self-Improving LLM Agents
- How to Train Your Advisor: Steering Black-Box LLMs with Advisor Models
- A General Language Assistant as a Laboratory for Alignment
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- Learning Hierarchical Procedural Memory for LLM Agents through Bayesian Selection and Contrastive Refinement
- Retrieval-Augmented Generation for Large Language Models: A Survey
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- QualityFlow: An Agentic Workflow for Program Synthesis Controlled by LLM Quality Checks
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Scaling Laws for Neural Language Models
- Scalable agent alignment via reward modeling: a research direction
- DeepSeek-V3 Technical Report
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- MigrationBench: Repository-Level Code Migration Benchmark from Java 8
- A Survey of Context Engineering for Large Language Models
- WebGPT: Browser-assisted question-answering with human feedback
- Olmo 3
- ToolRL: Reward is All Tool Learning Needs
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection