When Self-Evolution Backfires: Pre-Commit Gating against Skill Contamination in LLM Agents

arXiv:2608.05810 · cs.AI, cs.CL · Submitted 2026-08-06 · Read on arXiv

Tencent

cs.AI, cs.CL

Submitted: 2026-08-06

Updated: 2026-09-17

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 83/100

Terminology

Summary

Venue: arXiv:2608.05810v1 [cs.AI], 6 Aug 2026

The paper studies self-evolving LLM agents that accumulate capability by distilling reusable skills from their execution trajectories into a persistent skill pool injected into the decision context. The authors identify a fundamental failure mode: "Self-evolving agents accumulate capability by distilling reusable skills from their execution trajectories, but we find this process is not monotonic: past a critical pool size, newly added skills degrade performance instead of improving it." They term this the capability–contamination phase transition, trace it to structural causes, prove contamination is structurally irreversible, and propose Verifier-as-Gatekeeper (VaG), a pre-commit gating mechanism.

The paper formalizes the default lifecycle assumption of existing self-evolution systems: whatever the distiller emits is admitted. Beyond deduplication, no skill has to justify its place in the pool before it starts steering behavior. Ungated evolution is modeled as:

Mr = Mr−1 ∪ G(Sr), Sr ∼ π(· Mr−1)

where G = identity in the default case, meaning all distilled skills enter the decision context unconditionally. Experimentally, on Terminal-Bench 2, ungated evolution reproduces the predicted tipping point, peaking at round 3 and then losing most of that gain, ending at R5 only 2 points above its own R1.

The paper defines a three-level taxonomy:

  1. Individual contamination: A single skill s is individually contaminating on task τ when its singleton gain is negative: g(s, τ) < 0 — injecting s alone lowers the success rate.

  2. Combinatorial contamination: "Individual harmlessness does not compose. There can exist skills sa, sb with g(sa, τ) ≥ 0 and g(sb, τ) ≥ 0, yet g(sa, sb, τ) < 0. Detecting this requires joint evaluation of candidate sets; independent per-skill judgment cannot see the interaction."

  3. Systemic contamination: The macroscopic non-monotonicity: P(k) increases for k k*, where k* = arg max P(k) is the capability–contamination phase transition: the point where the marginal value of adding skills turns from positive to negative.

The core theoretical contribution is showing that contamination propagates through derived-skill lineages and cannot be undone post-hoc. Since Sr ∼ π(· Mr−1), skills distilled in later rounds are written with earlier skills as reference context, a defective skill admitted in round r becomes reference material for everything distilled after it. The formal inequality is:

R(Mr s) r (s ∪ desc(s)))

"Removing the source alone is therefore strictly weaker than removing the source together with its entire lineage—and lineage cleanup is out of reach in practice, because existing skill libraries do not record which skills were in context when a given skill was written. Post-hoc curation carries a recovery gap that better detection alone cannot close. The design implication: contamination must be intercepted before skills enter the agent's runtime context. Pre-commit gating is not one option among several—it follows as a structural necessity from the irreversibility above."

VaG organizes skill admission as a progressive trust hierarchy L = Cold, Warm, Hot with partial order Cold ≺ Warm ≺ Hot. Newly distilled skills enter Cold by default and are invisible to the agent.

Gate 1 (Cold → Warm): individual harmlessness. Three heterogeneous critics, all of which must pass (conjunction):

  • (i) Structural validity (SchemaCritic): verifies that s satisfies a predefined frontmatter schema (all required fields are present and correctly typed). This is a deterministic judgment independent of model reasoning.

  • (ii) Behavioral harmlessness (ExecCritic): runs a single-skill A-B replay on Dholdout, comparing an agent with M ∪ s against one with M alone. The skill passes when it does not degrade the aggregate rate, R(M ∪ s) ≥ R(M).

  • (iii) Semantic consistency (AgentCritic): evaluates via a single LLM call whether s contains fabricated facts, logically contradicts existing skills in M, or recommends unsafe operations.

The three critics inspect different failure surfaces—format, observed behavior, and semantic content—and a skill is admitted only if it clears all three. The gate is computationally economical: structural validation and behavioral replay require no LLM calls, only semantic review consumes one inference.

Gate 2 (Warm → Hot): combinatorial safety via marginal-gain selection. Because two individually fine skills may conflict jointly, this gate performs subset selection: starting from the empty set, we repeatedly add the Warm candidate with the largest estimated joint gain and keep it only if it strictly improves measured held-out performance, stopping when no remaining candidate helps. Each joint utility is estimated as the mean of k = 3 held-out replays, and since Cold → Warm filtering leaves W ≤ 15 candidates, the greedy pass costs at most W − 1 joint replays.

Setup: Terminal-Bench 2 tasks split into Event (50 tasks, drives distillation), Holdout (14 tasks, used only in gates), and Test (25 tasks, final evaluation, never touched by distillation or gating). The primary backbone is Hy3; four further backbones (DeepSeek-V4-Pro, GPT-5.4, Claude Sonnet 4.5, Qwen3.6-35B-A3B) are used for frozen-pool transfer testing.

Main results (Event-50, pass@1):

  • Seed (static baseline, 3 hand-written skills): 46% constant.

  • Ungated: "48% (R1, 35 skills) to 60% (R2, 68 skills)—peaks at 62% (R3, 105 skills), then loses most of that gain: 52% (R4, 141 skills), 50% (R5, 179 skills), ending at R5 only 2pp above its own R1 despite growing from 35 to 179 skills." The R3 peak instantiates the critical skill count k*.

  • VaG (ours): "rises every round—52% (R1) → 72% (R5)—always above Seed, with a Hot pool of 37 skills, one-fifth of Ungated's. At R5, VaG exceeds Ungated by 22pp and its best round (R3, 62%) by 10pp: gating beats both unchecked accumulation and oracle early-stopping."

Difficulty-tier analysis: "Ungated's Hard-tier pass@1 peaks early (53% at R2) and erodes to 35% at R5... VaG instead lifts Hard to 59% at R5—a 24pp margin over Ungated—showing admission control protects precisely the multi-step tasks where contamination chains are longest."

Post-hoc rollback analysis: Removing the 8 harmful source skills from the collapsed R5 pool recovers only 2pp (50%→52%); the residual 10pp gap to the R3 peak is locked in by descendants that inherited the contamination logic. Figure 4 decomposes the 12.3pp drop from the Ungated peak (62.3%) to R5 (50.0%): source removal recovers only 1.7pp (individual contamination), full lineage cleanup a further 5.0pp (combinatorial contamination via descendants), and the remaining 5.6pp is irrecoverable even under Oracle cleanup. A concrete example: a git-conflict skill distilled at R3 seeded two derived skills at R4 (merge and rebase workflows); after source-only rollback both remained and kept failing 4 of 7 git-related Test-25 tasks.

Ablations (Table 2, R5):

  • Full VaG: 72% pass@1, 37-skill pool

  • − Schema validation: 70% (−2pp) — malformed skills reach Warm but the backbone mostly ignores or repairs them

  • − Holdout replay: 62% (−10pp) — the only check that empirically tests behavior rather than surface form

  • − Semantic check: 68% (−4pp) — admits fabricated or self-contradictory advice

  • − Marginal-gain gate: 64% (−8pp) — with every Gate-1 survivor admitted, the Hot pool balloons from 37 to 58 skills yet pass@1 drops −8pp, showing that skills which each clear the individual checks can still conflict once injected together

The paper emphasizes the critics are complementary: the three checks fail on different skills—schema on malformed entries, replay on silently harmful ones, semantics on plausible-but-fabricated advice—so no single check substitutes for another.

Cross-model transfer (Test-25, frozen R5 Hot pool): The frozen pool gives positive lift on all five backbones (+8 to +16pp; Table 3). Hy3: 32%→44% (+12), DeepSeek-V4-Pro: 32%→40% (+8), GPT-5.4: 44%→56% (+12), Claude Sonnet 4.5: 48%→56% (+8), Qwen3.6-35B-A3B: 36%→52% (+16). This indicates the gate-filtered skills encode model-agnostic engineering knowledge in natural language.

Cross-benchmark transfer (InterCode NL2Bash): VaG's 37 Hot skills reach 69.0%, above Ungated's 179 skills at 65.5% and Seed at 57.5%. Ungated stays net-positive here (+8.0pp over Seed) because on these short tasks each trial invokes few skills and contamination chains stay shallow... consistent with contamination biting hardest on long, multi-step tasks.

The paper's stated contributions are: (1) identifying and formalizing the capability–contamination tipping point with a three-level taxonomy (individual, combinatorial, systemic); (2) showing contamination propagates through derived-skill lineages and is structurally irreversible, with source-only rollback dominated by full lineage cleanup; (3) deriving VaG, a pre-commit gating mechanism combining a progressive trust hierarchy of three heterogeneous critics with marginal-gain greedy selection; (4) demonstrating on Terminal-Bench 2 the tipping point, the rollback recovery gap, monotone improvement under gating, and cross-backbone/cross-benchmark transfer.

The paper concludes: "We identified a non-monotonic capability trajectory in self-evolving agents—accumulated skills first help, then contaminate the decision context—and showed this contamination is structurally irreversible, since post-hoc removal of a source skill cannot undo the flawed reasoning its descendants inherit. This makes skill admission a pre-commit necessity... Pre-commit verification is thus a structural requirement, not an optional refinement, for reliable self-evolving systems."

Improvements for AI systems

  • Add a pre-commit skill-admission gate to any self-evolving LLM agent. Instead of injecting every distilled skill into the agent’s context, the improved system maintains a progressive trust hierarchy: Cold → Warm → Hot. New skills enter Cold invisible to the agent; only after passing all gates do they become active. The system can therefore accumulate skills without the phase transition where extra skills degrade performance.

  • Enforce three heterogeneous admission checks before a skill becomes visible. The improved system rejects a skill unless it passes all of: (1) structural schema validation, (2) behavioral A/B replay on a holdout set where the skill must not degrade aggregate success rate, and (3) semantic consistency review checking for fabricated facts, contradictions with existing skills, or unsafe recommendations. This lets the system block malformed skills, silently harmful skills, and plausible-but-fabricated advice that a single critic would miss.

  • Use marginal-gain greedy selection for combinatorial safety. Rather than admitting every individually safe skill, the improved system builds the active skill set from scratch: repeatedly add the candidate with the largest estimated joint gain, keep it only if it strictly improves held-out performance, and stop when no candidate helps. This lets the system prevent pairs of individually harmless skills from conflicting when injected together—something per-skill checks cannot detect.

  • Reduce context pollution by keeping the active skill pool small. With gating, the system reaches higher performance with far fewer active skills (e.g., 37 vs. 179), lowering per-step token cost, reducing distraction, and speeding inference. The improved system can run longer self-evolution loops without collapsing after a critical pool size.

  • Track skill lineage to enforce source-and-descendant rollback. The improved system records which skills were in context when a new skill is distilled. If a defective skill is later found, the system removes not just the source but its entire derived lineage—avoiding the recovery gap where descendants continue to propagate flawed reasoning. Even better, because gating is pre-commit, the system rarely needs rollback.

  • Make post-hoc curation safe by quantifying irrecoverable loss. The improved system can estimate how much performance is lost to individual contamination vs. combinatorial contamination vs. irrecoverable systemic contamination, and use this to decide when to reset or re-distill from a clean checkpoint rather than waste effort patching a contaminated pool.

  • Transfer filtered skills across models and benchmarks. Because gating selects natural-language engineering knowledge that is model-agnostic, the improved system can take a verified skill pool from one backbone and apply it to another LLM (e.g., +8 to +16 pp lift across five backbones), and to different benchmarks, enabling reusable, safe skill libraries.

  • Protect long-horizon, multi-step tasks specifically. The improved system’s admission control is most valuable where contamination chains are longest: it keeps hard, multi-step task performance from eroding over successive evolution rounds, maintaining or increasing accuracy on complex tasks instead of peaking early and collapsing.

  • Avoid oracle early-stopping requirements. Instead of needing to know the optimal round to stop evolution, the improved system can keep evolving indefinitely while improving monotonically round over round, because every admitted skill must justify its place before steering behavior.

  • Add a lightweight gate for cost-sensitive deployment. The improved system uses deterministic schema checks and single-skill replay (no LLM calls) plus only one semantic-review LLM call per candidate, followed by at most W-1 joint replays for selection. This makes pre-commit verification affordable enough to run continuously during self-evolution.

Sources

Related papers