GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents

arXiv:2605.29668 · cs.AI, cs.CL · Submitted 2026-08-21 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents".

Jane: The paper was written by Johannes Moll, Jean-Philippe Corbeil, Jiazhen Pan, Martin Hadamitzky, Daniel Rueckert et al. from Technical University of Munich and TUM University Hospital, 2 Microsoft Healthcare & Life Sciences.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary and Implications: Tom: The core of this approach is really summarized by their results on the MedAgentBench benchmark, which shows gpt-oss-120b jumping from forty point six percent to a massive eighty-eight point eight percent.

Jane: That jump is so impressive because it indicates that they are encoding actual procedural knowledge into those skills, not just getting lucky with some sudden burst of general reasoning ability in the model.

Lu: I’m particularly interested in their findings regarding cross-model transfer, where a stronger model can write skills that then improve even weaker executors.

Meng: That makes sense if the library is truly capturing environment-specific facts; it suggests that the writer’s inherent strength is less important than the structure of the how-to guide itself.

Lalam: It feels like we are moving toward a future where AI agents aren't just trained on general data but are taught highly specific, verifiable operational logic that transcends their own model architecture.

Tom: And it’s not just medical environments, they also showed this effect in non-medical setups like ALFWorld and WebShop, which have very clear task structures.

Jane: The way GRASP generalizes is impressive; it is finding patterns that allow these skills to move across different domains of applicability while keeping the same logic intact.

Lu: I think the asymmetry—where a stronger writer helps a weaker executor but not vice versa—is the most telling evidence that this is about distilling procedural knowledge, not just leveraging model-specific strengths.

Meng: If we can reliably transfer this structured knowledge, it suggests a viable path toward deploying AI agents across different company systems without needing to retrain them from scratch every single one.

Improvements and Deep Dive: Tom: So, we have seen how GRASP tackles the problem of self-improvement in complex environments by moving away from monolithic memory and towards a gated library of skills.

Jane: It’s clear that the success they found is driven by this validation gate, not just in generating more candidate edits; it's the gate that makes the the difference.

Lu: The entire process, from identifying a failure to creating those structured, verifiable edits, demonstrates a level of sophisticated control that really pushes what we expect from these systems.

Meng: It seems like achieving operational reliability is no longer just about good prompt engineering but having an an explicit validated architecture for managing knowledge updates.

Lalam: The whole team is excited because this provides a robust framework, which gives us confidence that AI agents can evolve and perform reliably within the complexity of real-world systems.

Tom: We've had a great discussion on GRASP, Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents.

Jane: We're looking forward to seeing how this translates into other applications, but we're certainly taking a lot of ideas away from this one, too.

Lu: The creative potential here is massive; the ability to distill complex procedural knowledge into transferable skills opens up so many new avenues for highly specialized AI systems.

Meng: It provides a concrete path for building agents that can actually scale and maintain consistency without constant manual intervention or catastrophic failures in deployment.

Lalam: I believe the ultimate impact will be in creating digital workers who possess verifiable operational competence, allowing us to build more reliable and self-correcting software environments for everyone.

Paper discussion segment 3: Tom: The results are truly impressive because of how much better gpt-oss-120b performed on the MedAgentBench benchmark, leaping from about forty point six percent to a massive eighty-eight point eight percent.

Jane: That huge jump shows that the agent isn't just getting lucky; it’s really encoding genuine procedural knowledge into those skills rather than relying on some fleeting moment of general reasoning ability.

Lu: I’m particularly captivated by the cross-model transfer results, where a stronger model can write a skill and then improve an even weaker executor using that same instruction set.

Meng: That's quite interesting because if the library is actually capturing environment-specific facts, then the strength of the writer shouldn't matter as much as the inherent structure of the how-to guide itself.

Lalam: It seems like we are heading toward a future where AI agents aren't just trained on broad, general data but are taught specific, verifiable operational logic that transcends their own model architecture.

Tom: And this capability applies outside medical settings too, since they demonstrated that effect in non-medical environments like ALFWorld and WebShop where tasks have very clear structures.

Jane: The way G RASP generalizes is impressive; it's finding patterns that allow these behavioral skills to move between different domains of applicability.

Lu: I think the asymmetry—where a superior writer helps a weaker executor but not the reverse—is the most telling evidence that this is about distilling procedural knowledge, not just leveraging model-specific strengths.

Meng: If we can reliably transfer this structured knowledge across systems, it suggests a viable path toward deploying AI agents throughout different company workflows without needing to retrain them from scratch.

Lalam: This allows us to build digital workers who possess verifiable operational competence, which is a massive step toward creating more reliable and self-correcting software environments.

Conclusion: Tom: So, we've seen how GRASP tackles the problem of self-improvement in complex environments by moving away from monolithic memory toward a gated library of skills and validation checks that keep things safe.

Jane: It’s clear that the success isn't just about generating more candidates; it's fundamentally tied to this rigorous validation gate that prevents regressions, which is exactly what we see in operational environments like databases.

Lu: The entire process, from failure identification to creating those structured edits, shows a sophisticated level of control and architecture that really pushes the boundaries of what we expect from these systems.

Meng: It seems like the key to operational reliability is no longer just in prompt engineering but in having an explicit, validated framework for managing knowledge updates without unbounded growth.

Lalam: This provides a robust path forward, giving us confidence that AI agents can evolve and perform reliably within the complex culture of real-world systems.

Tom: We've spent a lot of time with GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents, and it's clear it is making significant progress in this area.

Jane: I think we are all very excited about how this work will be used in the future, but before we wrap up, I want to hear one final thought from everyone else.

Lu: I am especially interested in seeing how this capability scales across different types of complex tasks beyond the clinical environment they studied.

Meng: From a practical standpoint, I hope to see more evidence that this approach allows us to manage and maintain existing software infrastructure without constant manual intervention.

Lalam: I believe the ultimate impact will be in creating digital workers who possess verifiable operational competence, enabling self-corrective systems for everyone.

Tom: That’s a powerful vision, Lu, and it really highlights how far we've come from thinking of LLMs as just conversational tools.

Johannes Moll, Jean-Philippe Corbeil, Jiazhen Pan, Martin Hadamitzky, Daniel Rueckert, Lisa Adams, Keno Bressem

Technical University of Munich and TUM University Hospital, 2 Microsoft Healthcare & Life Sciences

cs.AI, cs.CL

Submitted: 2026-08-21

Updated: 2026-08-24

Comments: Accepted at EMNLP 2026 (Main Conference). Code and data: https://github.com/jomoll/GRASP

Code: https://github.com/jomoll/GRASP

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 86/100

The gist: " LLM agents operating within structured environments, such as electronic health records (EHR) or databases, exhibit failures that are "operational rather than conversational." Reliability in these

Key concepts

Gated Library of Skills
GRASP moves away from monolithic memory by using a gated library of skills. This allows the agent to manage knowledge updates through an explicit, validated architecture that prevents regressions and ensures operational reliability in complex environments.
Cross-Model Transfer
This is the ability a stronger AI model can write a skill or instruction set that then improves an even weaker executor. The effectiveness relies on the structure of the 'how-to guide' itself, not just the inherent strength of the writer.
Procedural Knowledge
The system captures verifiable operational logic rather than relying on general reasoning. This knowledge transcends a model's architecture, allowing AI agents to perform specific tasks reliably across different domains.

Terminology

Summary

"

LLM agents operating within structured environments, such as electronic health records (EHR) or databases, exhibit failures that are operational rather than conversational. Reliability in these settings depends critically on procedural knowledge of the environment rather than general language ability. Historically, self-improvement methods have failed because they accumulate natural language guidance at inference... without a check against existing behavior. This lack of validation means that a correction applied to fix one trajectory can silently regress another, leading to the dilution of task-relevant content and degradation of the agent’s context over time.

To address these issues, the authors introduce GRASP (Gated Regression-Aware Skill Proposer), which fundamentally changes how agent self-improvement is viewed. GRASP treat[s] agent self-improvement as a sequence of validated edits to a small library of behavioral instructions, or skills, rather than a monotonic accumulation of guidance.

A skill is defined as "a structured behavioral instruction injected into the agent’s context at inference... specifying a trigger condition, a behavioral rule, an optional verification step for actions with side effects, and a contrastive example." These skills are stored in an auditable library (Markdown documents with YAML frontmatter).

The GRASP mechanism operates through several distinct stages after each batch of training episodes:

  1. Failure Classification and Proposal Generation:
  • Failed traces are grouped by failure mode.

  • A skill-writing model proposes candidate edits (ADD, MODIFY, or REMOVE) for the most frequent modes. The framework generates K proposals per update cycle by cycling over failure-mode groups from largest to smallest.

  1. ** Regression-Aware Selection (The Gate):**
  • This is the core of the GRASP mechanism: Each candidate is re-run on a balanced probe of previously failing and previously passing examples.

  • The probe construction involves selecting up to N/2 previously-failing and N/2 previously-passing samples from episodes completed earlier in the current epoch.

  • A candidate is accepted only if it provides a net improvement while strictly adhering to a hard regression budget. The acceptance criterion is defined as:

F(c) - F 0 - R(c) - R 0 > 0 and R(c) R 0

where F and R are the number of fixes and regressions for a candidate, and F 0 and R 0 are the baseline counts.

  1. Contrastive Revision: If a winning candidate causes at least one regression, the skill-writer is invoked again to produce a narrower version that preserves the original fixes while exempting the regression patterns. This revision is only accepted if its adjusted score is strictly higher and still satisfies the hard budget.

The evaluation was conducted across five base models (gpt-oss-120b, DeepSeek V4 Flash, Gemini 3.1 Flash Lite, GPT-4.1, and GPT5.4) on two FHIR-based clinical benchmarks: MedAgentBench and MedAgentBenchv2.

  • Performance: On MedAgentBench, G RASP lifts gpt-oss-120b from 40.6% to 88.8%, exceeding the strongest of five self-improvement baselines by 21.0 points.

  • Attribution: Ablations showed that the gain is attributed to comparative proposal generation, the acceptance gate, and the hard regression budget, rather than simply skill writing itself (w/o validation, skill writing was no better than using no skills).

  • Generalization: The mechanism generalizes beyond clinical tasks: The mechanism generalizes beyond the clinical domain, improving agents on three of four non-clinical environments.

The study also examined cross-model transfer. A significant asymmetry was observed: Frozen libraries transfer across models, where skills from a stronger model improve weaker executors beyond what they learn for themselves while the reverse does not. This indicates that the libraries encode environment-specific procedural knowledge rather than model-specific patterns.

In conclusion, G RASP successfully converts agent failures into validated edits to a bounded library of behavioral skills. The authors conclude that "Treat[ing] agent self-improvement as gated editing of a bounded library, rather than monotonic accumulation of guidance, offers a path toward iterative refinement of LLM agents in environments where procedural reliability matters."

Improvements for AI systems

Based on a rigorous analysis of the GRASP methodology—a framework designed to rectify the fundamental flaw in current LLM self-improvement paradigms—the following improvements and resulting capabilities can be integrated into advanced AI systems.

The core improvement is shifting from unbounded, monotonic accumulation of guidance (which leads to catastrophic forgetting and dilution) to gated, validated editing of a bounded library.

Instead of injecting raw natural language critiques or flat memory blocks into the prompt context, implement a structured, versioned knowledge base (the Skill Library).

  • Specific Change: The agent’s context at inference will not contain accumulated text, but rather than 5–34k tokens of notes/memories. It will contain a small set of highly specific behavioral instructions (Skills) defined by YAML frontmatter.

  • Structure: Each skill must explicitly define:

  • A Trigger Condition (e.g., Instruction contains 'clean').

  • A Behavioral Rule (the corrective action, e.g., Enforce cleaning before placement).

  • An optional Verification Step (to confirm side effects).

  • A Contrastive Example (a successful trajectory).

Integrate a formalized, automated validation loop into the self-improvement training phase. This replaces trial and error learning with targeted refinement.

  • Specific Change: After every batch of episodes, failures are grouped by mechanism (Failure-Mode Classification). A specialized Skill Writer model generates candidate edits (Add, Modify, or Remove) for the most frequent failure modes.

  • The Validation Gate (GRASP): Before any candidate is applied to the live agent library, it must pass a rigorous check against a held-out Probe.

  • Step 1: Baseline Measurement (F 0, R 0): Measure the performance of the current stable library on the probe (how many failures F 0, how many regressions R 0).

  • Step 2: Candidate Scoring: Test a candidate library (S c) against the same probe to calculate its new failure count (F(c)) and regression count (R(c)).

  • Step 3: Acceptance Criterion (The Gate): The candidate is accepted only if it meets two criteria:

  1. Net Improvement: F(c) - F 0 > R 0 (It fixes more failures than it causes regressions).

  2. Hard Regression Budget: R(c) R 0 (It does not introduce any new regressions beyond the existing baseline).

By implementing this architecture, the improved AI system gains several distinct advantages over current state-of-the-art LLM agents:

The system will achieve a quantifiable level of reliability in structured environments that is currently unattainable by other methods. It eliminates silent regression—the scenario where a fix for one failure mode inadvertently causes another—which has historically plagued self-improving LLMs.

  • Capability: The system can maintain high accuracy (e.g., 85%+ on complex tasks) over long training periods without the performance degradation (dilution) seen in unbounded memory or flat note accumulation methods.

The system learns not just what is wrong, but precisely how to fix it via a structured behavioral instruction rather than a vague natural language critique.

  • Capability: The agent can solve highly specific, recurring operational problems (e.g., Always verify the opioid order has a matching naloxone prescription or Resolve the MRN to Patient ID before any dependent FHIR request) and generalize that procedural knowledge across different input samples, making it reliable in environments where tasks recur.

The system learns procedural knowledge tied to specific environment interfaces, not just model-specific patterns.

  • Capability: The knowledge encoded in the skill library is transferable. A frozen library trained on one benchmark (e.g., MedAgentBench) can be successfully applied to a different executor model, improving its performance significantly beyond what the executor could achieve by self-training alone—a phenomenon not observed in current state-of-the-art baselines.

The system maintains efficiency and transparency.

  • Capability: The active skill set is kept small (typically under 10 skills) compared to the massive token bloat of memory approaches, significantly reducing inference costs while remaining highly auditable because every single behavioral change is a versioned, justified edit tied to a specific failure mode.

Abstract

LLM agents acting in structured environments fail in operational rather than conversational ways, and reliability depends on procedural knowledge of the environment. Prior self-improvement methods accumulate natural-language guidance without checking that each new item preserves previously correct behavior, so a note that fixes one trajectory can silently regress another. We introduce GRASP (Gated Regression-Aware Skill Proposer), which treats agent improvement as a sequence of edits to a bounded skill library, admitting each candidate only if it produces a net improvement on a balanced held-out probe under a hard regression budget. We evaluate GRASP across five base models on two FHIR-based clinical benchmarks, which score procedural reliability against FHIR state rather than clinical correctness or patient outcomes. On MedAgentBench, GRASP lifts gpt-oss-120b from 40.6% to 88.8%, exceeds the strongest of five self-improvement baselines by 21.0 points, and improves every other base model by 17.2 to 40.3 points. Ablations attribute the gain to comparative proposal generation, the acceptance gate, and the hard regression budget rather than to skill writing itself, which without validation is no better than using no skills. Granting the same acceptance gate to all five baselines lifts each of them in-domain and none of them out of distribution, isolating the gain to the gate applied to a bounded, editable library rather than to held-out validation itself. The mechanism helps in non-clinical environments where tasks recur with verifiable structure and is flat where the action space is open-ended. Frozen libraries transfer across models and across benchmarks that share a tool-calling convention and degrade under interface mismatch.

Sources

Related papers