SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback

arXiv:2608.13120 · cs.AI · Submitted 2026-08-13 · Read on arXiv

Qianxi Yan, Chunrong Chen, Jiuzhou Zhao, Min Zhang, Yongzhou Xu, Xiaochuan Xu

Tencent Cloud Andon · Zhejiang University

cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: The paper addresses the challenge of continual improvement for Agent Skills—portable modules that encapsulate domain knowledge and handling procedures in customer support systems.

Terminology

Summary

The paper addresses the challenge of continual improvement for Agent Skills—portable modules that encapsulate domain knowledge and handling procedures in customer support systems. The authors argue that existing approaches to skill evolution are fundamentally limited: Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no closed loop through which they might improve from the interaction failures they actually cause.

The central claim is that the quality of Skill self-evolution is governed by the quality of the feedback signal and by controllable evolution governance, rather than by editing capability or the number of iterations. The paper identifies a critical asymmetry in existing methods: once the first round has patched the gaps that a single exchange can reveal, the evolution gradient decays, the defects that surface only across multiple turns remain invisible, and evolution stalls.

The authors further critique existing governance mechanisms: Governance in these systems is likewise driven by an end-to-end verification score—a scalar gate that can reject a degraded candidate but can neither localize nor repair its structural cause.

SkillEvo rests on two pillars:

  1. Trustworthy feedback generation: "The first component recasts multi-turn user simulation from an evaluation endpoint into a feedback generator: follow-up questions expose defects layer by layer, so that every round of revision both consumes feedback and produces new feedback."

  2. Controllable governance: "The second replaces the passive rejection of a scalar gate with an independent governance layer that actively repairs factual degradation and structural bloat, preventing the gradient from drifting as degradation accumulates."

The paper formalizes three necessary conditions for trustworthy feedback:

An intent state machine tracks for each intent, whether it has been raised and whether it has been substantively addressed. Normal termination is permitted only when all intents satisfy both conditions. Intent coverage is measured as cU = Kasked/K, where K is the set of key intents and Kasked is the subset actually raised.

SkillEvo evaluates the simulator side and agent side separately: a failure occurring where an intent was never raised is attributed to simulation distortion rather than to the agent's Skill. The agent-side metric is exposed-intent response accuracy, computed as a weighted score where key intents receive weight alpha = 0.7 and minor intents receive 1 − alpha.

Failures are classified into three categories: "Knowledge Gap, where a stable fact present in the human reference was omitted or answered incorrectly; Capability Limit, covering permission and tooling restrictions, poor delivery, and infrastructure faults; and Evaluation Noise, covering false negatives and scenario distortion. Only Knowledge Gap is projected into the feedback; the other classes are isolated from the revision loop. Multiple failures pointing to the same gap are merged by semantic similarity into a single feedback signal... focusing revision on cross-sample commonalities instead of per-instance noise."

Fact consistency requires a revised Skill to retain at least the stable facts of the production baseline: Facts(St) ⊇ Facts(S0) ∩ Sstable. The constraint is checked against dual anchors: S0 detects fact loss accumulated across rounds, and St−1 detects factual errors newly introduced in the current round, so that degradation is attributed to the correct revision source.

Three violation classes are detected: "knowledge loss against S0, where a stable fact has been deleted; process errors against St−1, comprising factual errors newly introduced this round; and self-contradiction globally, where the revised Skill asserts conflicting statements."

The Skill is treated as a structured knowledge system forming a directed graph in which routing nodes point to knowledge nodes and knowledge nodes cross-reference one another. Three degradation dimensions are identified:

  • Knowledge bloat: redundancy grows within nodes, diluting routing precision

  • Reference breakage: inter-node connectivity is severed, as dangling references or orphan files

  • Factual over-generalization: concrete values, versions, and rules decay into vague statements

Unlike fact consistency which rejects candidates, structural consistency operates as a soft constraint: rather than discarding the revision, it merges governance recommendations with the attributed gaps and injects them into the next round.

The evaluation uses production technical-support scenarios of Tencent Cloud, spanning six categories of cloud services, 9 production Skills, and 98 skill-reference files. The dataset consists of tickets escalated to human agents—the failure set of the existing Skills—each ticket corresponds to a knowledge gap that a real user has already exposed and the current Skill fails to cover.

Tickets are split chronologically: The first three constitute the development set and drive the evolution loop... The fourth is held out as the evaluation set, is fed back into no stage of the loop, and serves solely for measurement and reporting.

Baselines include:

  • Original Skill: hand-authored, never updated

  • Self-Reflection: Model self-reflects and edits the Skill directly, no evaluation feedback

  • Single-turn QA: Single-turn QA evaluation-driven evolution (representing SkillForge and similar methods)

The per-round TSR results show:

Method Init R1 R2 R3 R4


Original Skill 30.0 — — — —

Self-Reflection 30.0 59.2 58.7 57.4 58.8

Single-turn QA 30.0 58.9 64.5 65.7 66.4

SkillEvo 30.0 59.4 71.3 77.9 81.8

The paper explains the differences: "Self-Reflection possesses no gradient in the absence of evaluation, so multi-round blind editing merely oscillates around its first-round level. Single-turn QA supplies a first-round gradient, but a single question–answer pair reaches only the gaps present in the user's opening statement; once those are patched the gradient decays."

For SkillEvo: "Follow-up questions and clarification expose defects layer by layer: knowledge patched in the current round lets the dialogue proceed further and reach the next layer of defects, previously masked by shallower failures. Each round of revision therefore not only consumes gradients but generates new ones."

Two ablations are reported:

  • (a) Single-turn QA: removes multi-turn interaction with single-turn QA evaluation while leaving attribution, revision, and governance untouched → TSR drops to 66.4, rendering it substantively equivalent to the Single-turn QA baseline

  • (b) w/o Governance: removes the governance layer while leaving feedback, attribution, and revision untouched → TSR falls to 78.6 (−3.2)

The paper concludes: removing multi-turn interaction closes the gap entirely, and SkillEvo's 15.4-point lead is therefore attributable to the feedback source itself. The value of governance lies not in raising the score but in preventing degradation from accumulating across rounds.

Dual-sided orthogonal evaluation results:

  • Intent coverage (cU): 98.9%

  • Fidelity (rho): 95.3% (human-rated similarity, validated by two domain experts blindly comparing 200 simulated dialogues against real tickets)

  • Exposed-intent accuracy (sC): 71.1%

The paper notes: The value reflects the Skill's knowledge-coverage gaps and is the direct source of the evolution gradient; it is not a whole-dialogue resolution rate.

Cross-round regression rate (RegR) declines across transitions: 28.2% (R1→2), 24.4% (R2→3), 21.1% (R3→4), a first-to-last change of −7.1%.

Knowledge bloat comparison:

  • With governance: +2.8% cumulative growth, Growth concentrates in the first round and then tapers off

  • Without governance: +16.2%, Bloat accumulates round after round with no dissolution mechanism

The paper emphasizes: That TSR improves by 51.8 points while volume barely changes indicates that the capability gain arises from revising existing knowledge correctly rather than from expanding the text.

  1. Formalization of three necessary conditions for trustworthy feedback—coverage, accuracy, and attributability—and recast multi-turn user simulation from an evaluation endpoint into a feedback generator

  2. Identification of a Skill as a structured knowledge system whose multi-round revision incurs degradations that a scalar score cannot diagnose: knowledge bloat, reference breakage, and factual over-generalization

  3. Empirical results: SkillEvo improves over the original Skills by 51.8 points, over self-reflection-based evolution by 23.0 points, and over single-turn-QA-driven evolution by 15.4 points

The paper concludes: "Reliable skill evolution thus rests on two requirements: a high-quality feedback loop that exposes interaction defects and isolates simulation distortion, and a governance mechanism that actively maintains the Skill as a structured knowledge system. The framework is deployed in the Tencent Cloud production environment, evidencing its effectiveness under real operating conditions."

Improvements for AI systems

Based on this paper, I can make the following specific improvements to AI systems:

1. Multi-turn feedback generation for iterative self-improvement

  • Implement a feedback loop where the AI simulates follow-up user questions to expose defects layer-by-layer, rather than relying on single-turn evaluation

  • After each revision, generate new feedback from deeper interaction layers that were previously masked by shallower failures

  • Track intent coverage (which user intents were raised and substantively addressed) to ensure complete problem-space exploration before declaring success

2. Dual-sided orthogonal evaluation for accurate failure attribution

  • Separate evaluation of the AI's response quality from the simulation/input quality

  • Only attribute failures to the AI when the relevant intent was actually raised in the interaction

  • Weight key intents (α=0.7) more heavily than minor intents in accuracy scoring

3. Collective attribution with failure classification

  • Classify failures into Knowledge Gap, Capability Limit, and Evaluation Noise

  • Only feed Knowledge Gap failures back into the revision loop; isolate the other classes

  • Merge semantically similar failures into single feedback signals to focus on cross-sample commonalities rather than per-instance noise

4. Governance layer with hard and soft constraints

  • Enforce fact consistency as a hard constraint: revised versions must retain at least the stable facts of the production baseline

  • Check against dual anchors (original version for cumulative loss, previous version for newly introduced errors)

  • Detect three violation classes: knowledge loss, process errors, and self-contradictions

5. Structural consistency maintenance

  • Treat the AI's knowledge as a directed graph with routing nodes and knowledge nodes

  • Monitor and repair three degradation dimensions: knowledge bloat (redundancy), reference breakage (dangling links), and factual over-generalization (concrete values decaying into vague statements)

  • Apply structural fixes as soft constraints that merge with attributed gaps for the next revision round, rather than rejecting candidates outright

The improved AI system can:

  • Continuously improve from real interaction failures rather than plateauing after initial patches

  • Distinguish between its own knowledge gaps and external limitations (tooling, permissions, infrastructure)

  • Maintain factual accuracy across multiple revision rounds without accumulating errors or contradictions

  • Keep its knowledge base compact and well-structured even as it grows

  • Diagnose and repair structural degradation (bloat, broken references, over-generalization) that scalar scores cannot detect

  • Generate increasingly deeper feedback signals with each revision round, enabling sustained improvement (81.8% success vs 66.4% for single-turn methods)

Abstract

Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no closed loop through which they might improve from the interaction failures they actually cause. Recent work does close this loop, but derives its feedback from single-turn question-answering evaluation. The consequence is a sharp asymmetry: once the first round has patched the gaps that a single exchange can reveal, the evolution gradient decays, the defects that surface only across multiple turns remain invisible, and evolution stalls. Governance in these systems is likewise driven by an end-to-end verification score, a scalar gate that can reject a degraded candidate but can neither localize nor repair its structural cause. We argue that the binding constraint on sustained skill evolution is neither editing capability nor the number of iterations, but whether the evaluation feedback keeps supplying trustworthy evolution gradients. We introduce SkillEvo, in which trustworthy feedback generates the gradient and controllable governance constrains its direction. The first component recasts multi-turn user simulation from an evaluation endpoint into a feedback generator: follow-up questions expose defects layer by layer, so that every round of revision both consumes feedback and produces new feedback. The second replaces the passive rejection of a scalar gate with an independent governance layer that actively repairs factual degradation and structural bloat, preventing the gradient from drifting as degradation accumulates. Across six categories of cloud services, 9 production Skills, and 98 skill-reference files, SkillEvo surpasses self-reflection-based evolution by 23.0 points and single- turn-QA-driven evolution by 15.4 points.

Sources

Related papers