ExpHarness: Model-Agnostic Experience Learning through a Trainable Harness
summary
The gist
I will meticulously combine the provided text fragments to construct a detailed, accurate summary of ExpGraph, ensuring all key technical contributions are captured while addressing the context
In short
ExpHarness introduces a trainable harness that allows Large Language Models (LLMs) to learn from experience without retraining the LLM itself. It attempts to solve this by creating a framework where an external, structured memory stores past interactions. The result is a system that uses this memory and an adaptive retrieval mechanism to improve the performance of frozen LLM executors through experience reuse.
Key concepts
- Trainable Harness
- A trainable harness is an external structure built around an LLM executor. Instead of changing the core LLM, you train this surrounding structure. This harness learns how to effectively use past experiences to guide the LLM's decision-making process, making it more capable without needing to retrain the original model.
- Experience Learning
- This is a method where an AI system improves by reusing past interactions. Instead of learning from scratch every time, it stores successful and unsuccessful task outcomes as 'experiences.' The system then learns how to select the most relevant past experiences to inform its current actions, leading to better performance over time.
- Model-Agnostic
- Model-agnostic means the learning framework works with any LLM executor. It doesn't require you to change or retrain the specific LLM being used. The harness is designed to be flexible enough to work with different types of LLMs, allowing for easy swapping of executors without rewriting the learning logic.
Terminology used across episodes
This episode discusses
- ExpHarness: Model-Agnostic Experience Learning through a Trainable Harness · Paper Radio
- Evaluating Large Language Models Trained on Code
- Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- LightMem: Lightweight and Efficient Memory-Augmented Generation
- The Llama 3 Herd of Models · Paper Radio
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Steer2Adapt: Dynamically Composing Steering Vectors Elicits Efficient Adaptation of LLMs
- Adaptation of Agentic AI: A Survey of Post-Training, Memory, and Skills
- Process Reward Models That Think
- Training Language Models to Self-Correct via Reinforcement Learning
- Search-o1: Agentic Search-Enhanced Large Reasoning Models
- DeepSeek-V3 Technical Report
- Understanding R1-Zero-Like Training: A Critical Perspective
- Decoupled Weight Decay Regularization
- Let's reward step by step: Step-Level reward model as the Navigators for Reasoning
- ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory
- WebCanvas: Benchmarking Web Agents in Online Environments
- Proximal Policy Optimization Algorithms
The paper
ExpHarness: Model-Agnostic Experience Learning through a Trainable Harness · Read on arXiv
Tao Feng, Chongrui Ye, Tianyang Luo, Jingjun Xu, Xueqiang Xu, Haozhen Zhang, Zhigang Hua, Yan Xie, Shuang Yang
University of Illinois Urbana-Champaign
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "ExpHarness: Model-Agnostic Experience Learning through a Trainable Harness".
Tom: I will meticulously combine the provided text fragments to construct a detailed, accurate summary of ExpGraph,
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So we’re starting with the title and who wrote this thing. "ExpHarness: Model-Agnostic Experience Learning through a Trainable Harness." It sounds like they are introducing a way to harness experience in a flexible way that doesn't tie the learning process down to one specific model.
Jane: It basically suggests they’ve created some sort of system where the experience can be used by different executors, and that executor itself stays frozen.
Lu: The paper is really pushing the idea that you can decouple how an agent learns from how powerful the underlying LLM executor actually is.
Meng: I’m wondering if this means we could swap out our reasoning engine mid-development and still get performance gains from the experience we've already collected, even if the new engine is fundamentally different.
Lalam: It sounds like they are building a sort of universal memory system that any agent can plug into for improvement.
The paper's summary: Tom: Okay, looking at the actual summary of "ExpHarness," they’re talking about how you summarize past task trajectories into reusable skills and failure lessons. These are then organized in a structure called an experience graph.
Jane: That graph is key because it connects related experiences together, so when the agent needs help on a new task, it doesn't just look at one random old example; it can find related strategies.
Lu: They emphasize that this isn't just storing data; they are encoding transferable strategies and shared sub-goals directly into the structure of those connections within the graph.
Meng: So instead of having a massive database of text, they’ve made it a structured map where the agent can navigate relationships between past failures and successes.
Lalam: It sounds like they are creating a persistent, organized way for an agent to remember *how* things worked, not just *what* the answer was.
The paper's improvements: Tom: Now for the actual improvements they point out in "ExpHarness." They highlight four main components working together: graph-structured experience, graph diffusion, utility-aware ranking, and adaptive retrieval.
Jane: That adaptive retrieval part is what I find really interesting. It’s not just a simple search; it’s a mechanism that decides which experience is most useful for the current situation based on real performance feedback.
Lu: They use reinforcement learning to optimize this selection process, guided by utility-grounded feedback which compares the agent's performance when it uses an experience versus when it doesn't.
Meng: That optimization loop is what makes it smarter; it learns how to pick the right experience based on what actually helps the executor perform better in that specific context.
Lalam: It moves beyond just finding things that are semantically similar; this system is designed to select experiences that actually lead to a measurable improvement in the agent’s actual task success.
Conclusion: Tom: So, wrapping up on "ExpHarness," the main implication is that this approach lets frozen and replaceable LLM executors improve through external experience reuse without needing any retraining of those specific models.
Jane: It suggests a path where we can evolve our agents by feeding them structured memory instead of constantly having to fine-tune the core model itself for every new capability.
Lu: The results show that when you combine all these elements—the graph structure, the ranking, and the adaptive retrieval—you get really solid gains in agentic tasks.
Meng: I see a practical implication in deployment: we can deploy a stable experience system once and let it serve many different executor models efficiently over time.
Lalam: It makes the agent improvement process continuous; it’s not just about one big training run, but about constantly curating and retrieving useful past interactions.
Tom: So, "ExpHarness" gives us a flexible way to keep our agents sharp by building that external experience memory and letting an adaptive copilot manage how they pull from it.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language