ExpHarness: Model-Agnostic Experience Learning through a Trainable Harness
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "ExpHarness: Model-Agnostic Experience Learning through a Trainable Harness".
Tom: I will meticulously combine the provided text fragments to construct a detailed, accurate summary of ExpGraph,
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So we’re starting with the title and who wrote this thing. "ExpHarness: Model-Agnostic Experience Learning through a Trainable Harness." It sounds like they are introducing a way to harness experience in a flexible way that doesn't tie the learning process down to one specific model.
Jane: It basically suggests they’ve created some sort of system where the experience can be used by different executors, and that executor itself stays frozen.
Lu: The paper is really pushing the idea that you can decouple how an agent learns from how powerful the underlying LLM executor actually is.
Meng: I’m wondering if this means we could swap out our reasoning engine mid-development and still get performance gains from the experience we've already collected, even if the new engine is fundamentally different.
Lalam: It sounds like they are building a sort of universal memory system that any agent can plug into for improvement.
The paper's summary: Tom: Okay, looking at the actual summary of "ExpHarness," they’re talking about how you summarize past task trajectories into reusable skills and failure lessons. These are then organized in a structure called an experience graph.
Jane: That graph is key because it connects related experiences together, so when the agent needs help on a new task, it doesn't just look at one random old example; it can find related strategies.
Lu: They emphasize that this isn't just storing data; they are encoding transferable strategies and shared sub-goals directly into the structure of those connections within the graph.
Meng: So instead of having a massive database of text, they’ve made it a structured map where the agent can navigate relationships between past failures and successes.
Lalam: It sounds like they are creating a persistent, organized way for an agent to remember *how* things worked, not just *what* the answer was.
The paper's improvements: Tom: Now for the actual improvements they point out in "ExpHarness." They highlight four main components working together: graph-structured experience, graph diffusion, utility-aware ranking, and adaptive retrieval.
Jane: That adaptive retrieval part is what I find really interesting. It’s not just a simple search; it’s a mechanism that decides which experience is most useful for the current situation based on real performance feedback.
Lu: They use reinforcement learning to optimize this selection process, guided by utility-grounded feedback which compares the agent's performance when it uses an experience versus when it doesn't.
Meng: That optimization loop is what makes it smarter; it learns how to pick the right experience based on what actually helps the executor perform better in that specific context.
Lalam: It moves beyond just finding things that are semantically similar; this system is designed to select experiences that actually lead to a measurable improvement in the agent’s actual task success.
Conclusion: Tom: So, wrapping up on "ExpHarness," the main implication is that this approach lets frozen and replaceable LLM executors improve through external experience reuse without needing any retraining of those specific models.
Jane: It suggests a path where we can evolve our agents by feeding them structured memory instead of constantly having to fine-tune the core model itself for every new capability.
Lu: The results show that when you combine all these elements—the graph structure, the ranking, and the adaptive retrieval—you get really solid gains in agentic tasks.
Meng: I see a practical implication in deployment: we can deploy a stable experience system once and let it serve many different executor models efficiently over time.
Lalam: It makes the agent improvement process continuous; it’s not just about one big training run, but about constantly curating and retrieving useful past interactions.
Tom: So, "ExpHarness" gives us a flexible way to keep our agents sharp by building that external experience memory and letting an adaptive copilot manage how they pull from it.
Tao Feng, Chongrui Ye, Tianyang Luo, Jingjun Xu, Xueqiang Xu, Haozhen Zhang, Zhigang Hua, Yan Xie, Shuang Yang
University of Illinois Urbana-Champaign
cs.CL
Submitted: 2026-05-29
Updated: 2026-10-04
Code: https://github.com/ulab-uiuc/ExpGraph
Importance score: 92/100
The gist: I will meticulously combine the provided text fragments to construct a detailed, accurate summary of ExpGraph, ensuring all key technical contributions are captured while addressing the context
Key concepts
- Trainable Harness
- A trainable harness is an external structure built around an LLM executor. Instead of changing the core LLM, you train this surrounding structure. This harness learns how to effectively use past experiences to guide the LLM's decision-making process, making it more capable without needing to retrain the original model.
- Experience Learning
- This is a method where an AI system improves by reusing past interactions. Instead of learning from scratch every time, it stores successful and unsuccessful task outcomes as 'experiences.' The system then learns how to select the most relevant past experiences to inform its current actions, leading to better performance over time.
- Model-Agnostic
- Model-agnostic means the learning framework works with any LLM executor. It doesn't require you to change or retrain the specific LLM being used. The harness is designed to be flexible enough to work with different types of LLMs, allowing for easy swapping of executors without rewriting the learning logic.
Terminology
Summary
I will meticulously combine the provided text fragments to construct a detailed, accurate summary of ExpGraph, ensuring all key technical contributions are captured while addressing the context provided by the contrasting snippet (B).
Here is the comprehensive summary:
Detailed Summary of ExpGraph: Model-Agnostic Experience Learning with Graph-Structured Memory for LLM Agents
ExpGraph is presented as a sophisticated, model-agnostic experience learning framework specifically designed to enable frozen and replaceable Large Language Model (LLM) executors to improve their performance through external experience reuse without requiring any modification or retraining of the underlying executor parameters. The core philosophy of ExpGraph is to decouple the process of experience learning from the training lifecycle of the executor LLM itself, treating the executor as a replaceable task solver that learns how to leverage useful experiences provided via context.
Core Architectural Components and Mechanism:
- Experience Graph (Self-Evolving Memory):
-
ExpGraph functions by summarizing historical task trajectories into structured, reusable components: skills and failure lessons.
-
These summarized experiences are organized as nodes within a self-evolving experience graph. Crucially, the framework connects related experiences to facilitate retrieval mechanisms that extend far beyond simple flat nearest-neighbor matching. This graph serves as the persistent, structured memory where past successes and failures are stored.
- Retrieval Copilot (Adaptive Selection):
-
For any given task instance, a lightweight retrieval copilot is employed to adaptively control the process of graph diffusion.
-
The copilot's primary function is to perform utility-aware ranking, retrieving experiences that are simultaneously task-relevant and historically useful for the current state of the frozen executor.
- Reinforcement Learning Optimization:
-
The effectiveness of the retrieval copilot is optimized using reinforcement learning. This optimization is guided by utility-grounded feedback, which quantitatively compares the performance of the executor with retrieved experiences against its performance without them.
-
This feedback loop allows the system to learn how to adapt experience selection strategies across different tasks and potentially different executor models.
Key Advantages and Performance Metrics:
The framework's design provides significant practical benefits:
-
Executor Agnosticism: The most significant advantage is the decoupling of learning from training. When a stronger or entirely different LLM becomes available, the same external experience system (the graph and copilot) can be reused or adapted without necessitating retraining of the original executor LLM.
-
Efficiency Gains: Empirical results demonstrate substantial improvements over strong baselines:
-
On static tasks, ExpGraph improves performance by 12.2% with a smaller executor and 4.7% with a larger one.
-
In agentic environments, these gains increase to 21.4% and 12.7%, respectively.
-
Furthermore, the framework reduces the average number of interaction steps by 12.7% and 21.6%.
Ablation Study Findings:
Ablation studies confirm that the synergistic combination of three key elements—graph-structured experience, utility-aware ranking, and adaptive retrieval—is jointly responsible for enabling effective experience reuse across a diverse spectrum of tasks and executor architectures. This confirms that the framework provides a practical, executor-agnostic pathway for LLM agents to learn from their interactions efficiently.
**In essence, ExpGraph transforms the LLM agent into an entity capable of continuous, data-driven improvement by maintaining an external, structured memory and a learning mechanism (the copilot) that intelligently curates past experiences for the frozen model.
Improvements for AI systems
-
Model-Agnostic Experience Learning for Frozen Executors: ExpGraph allows
frozen and replaceable LLM executors to improve through external experience reuse without modifying their parameters.
This enables improvement across rapidly evolving LLM capabilities without requiring executor-specific retraining, decoupling learning from training. -
Graph-Structured Relational Memory: The framework "summarizes historical trajectories into reusable skills and failure lessons, organizes them as nodes in a self-evolving experience graph, and connects related experiences to support retrieval beyond flat nearest-neighbor matching.
This moves beyond isolated records by encoding
transferable strategies, shared sub-goal, or failure patterns" through edges in the graph. -
Utility-Aware Adaptive Retrieval: ExpGraph employs a
lightweight retrieval copilot adaptively controls graph diffusion and utility-aware ranking,
optimizing with RL usingutility-grounded feedback that compares executor performance with and without retrieved experiences.
This ensures retrieval selects experiences thattruly improve the executor rather than merely matching the task semantically.
-
Flexible Executor Transfer: The system enables superior generalization across different LLM scales by showing that
Graph+Copilot Transfer performs closest to target-specific ExpGraph
in zero-shot transfer settings, suggestingtransferring both components together is consistently more effective.
Sources
- Evaluating Large Language Models Trained on Code
- Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- LightMem: Lightweight and Efficient Memory-Augmented Generation
- The Llama 3 Herd of Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Steer2Adapt: Dynamically Composing Steering Vectors Elicits Efficient Adaptation of LLMs
- Adaptation of Agentic AI: A Survey of Post-Training, Memory, and Skills
- Process Reward Models That Think
- Training Language Models to Self-Correct via Reinforcement Learning
- Search-o1: Agentic Search-Enhanced Large Reasoning Models
- DeepSeek-V3 Technical Report
- Understanding R1-Zero-Like Training: A Critical Perspective
- Decoupled Weight Decay Regularization
- Let's reward step by step: Step-Level reward model as the Navigators for Reasoning
- ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory
- WebCanvas: Benchmarking Web Agents in Online Environments
- Proximal Policy Optimization Algorithms
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering