SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Paper Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent".
Jane: The paper was written by Mingxuan Zheng, Yujin Zhou, Chuxue Cao, Boqin Yin, Yuyao Zhang et al. from Hong Kong University of Science and Technology and Macau University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Paper summary: Tom: So we finally get to see what happens when you stop trusting your skill patches just because they look plausible. The paper treats skill evolution like an optimization problem with a forward pass and a backward pass, but every step happens in text.
Jane: And the key difference from earlier work is that they re-run the skill after each patch, on the same batch, and only keep the patch if performance doesn't drop. That seems like an obvious sanity check, yet none of the existing methods actually do it.
Lu: Right, SkillGrad and SkillOpt commit a diagnosis directly as a patch without checking what it does in practice. The paper calls that open-loop, and their trace shows it matters: in one run, eight out of twenty-two attempted edits regressed hard accuracy and had to be blocked.
Meng: And the backward stage is the more surprising piece to me. They decompose the evolved skill into auditable knowledge units, then remove each unit temporarily to measure its marginal utility on a validation split. If a unit's removal helps, they consolidate or delete it — but every realized edit still has to pass a hard validation gate.
Jane: The case study is clean evidence for that. They compress one skill by only 3 point 12 percent, removing a single over-specific section, and hard accuracy on the out-of-sample evaluation jumps from 46 percent to 54 percent. Eight tasks flip from fail to pass, and none flip the other way.
Tom: So that's the "proximal" part — a discrete analogue of shrinkage, where they control skill complexity instead of letting it grow forever. And the forward stage gives them outcome-grounded feedback for future diagnoses, not just a one-shot semantic check.
Lu: Exactly. The paper's own ablation says that removing Prox costs 2 point 5 points on SpreadsheetBench, while removing the closed-loop diagnosis costs 1 point 5. So the shrinkage stage is carrying a bit more weight, but they're complementary.
Meng: And across all backbones, SkillProx averages about three points better than the strongest gradient-based baseline. What impresses me most is the OOD robustness — SkillOpt collapses on WikiTQ and HiTab, down to 26 percent and 16 percent on the 4B model, while SkillProx stays competitive or best.
Jane: Let's bring Lalam in. What do you make of this as a bigger trend?
Lalam: I think it's part of a shift toward agents maintaining their own procedural memory on disk instead of retraining weights. What SkillProx adds is a maintenance layer — a way to audit, prune, and consolidate that memory so a skill doesn't become a junk drawer of half-remembered task solutions.
Tom: And that's why the backward stage matters — it treats deletion as a dedicated operation, not just one generic edit among many. Preventing bad additions and actually shrinking accumulated content are two different problems, and this paper tackles both.
Jane: The detailed analysis on τ is nice too. At a threshold of −0 point 001 they get 52 point 3 percent accuracy with 25 point 7 percent compression, and even at 74 point 9 percent compression accuracy stays at 51 point 0 percent. So you can shrink a skill a lot before you start hurting it.
Meng: Which also explains why the 4B model ends up with a 45 percent longer final skill than the 27B model — mostly from extra reference files. The closed-loop gate and the Prox compression act as an external filter that disproportionately regularizes the smaller model.
Lu: For me, the broader lesson is that textual skill evolution needs measurement at both ends. You need execution feedback to verify forward updates, and you need utility auditing to clean up what accumulates. The paper gives a concrete framework for both, and the numbers suggest it works.
Tom: And that's probably the most useful thing to take away for people building agents in the wild. Skills are already a practical way to give agents knowledge, and this shows you can evolve them safely — with a gate and a shrink step — rather than just letting them balloon.
Page 1 of the paper: Jane: So earlier we talked about skills as lightweight textual artifacts that agents load into context. Page one gets into why those artifacts go wrong. The paper says most systems take an LLM-generated diagnosis, treat it as a valid update direction, and commit the patch without ever checking whether it actually works. It calls that an "unverified forward update".
Tom: So they just trust the model's explanation of what went wrong?
Lu: Exactly. There's no re-execution on the same task batch, so the outcome never feeds back into the next diagnosis. And the second problem is the mirror image: iterative patching makes the skill grow without any mechanism to reassess accumulated knowledge. The paper gives a concrete example where removing one negative-utility unit improves accuracy from 46 percent to 54 percent.
Jane: Wait, removing stuff makes it better? That sounds counterintuitive.
Meng: That's the motivation for the "backward" stage. The skill gets decomposed into auditable knowledge units, and each one's contribution is estimated by a frozen leave-one-out utility audit. Then the method selectively consolidates, demotes, or removes units that don't earn their place.
Tom: So the idea is to prevent bad edits upfront, and then go back and clean up what's already accumulated?
Lu: Right. And that's exactly the central question posed on page one: how can outcome-verified forward diagnosis be coupled with structure-aware backward refinement, so the skill's capability and its structure co-evolve. The rest of the paper is basically the answer to that.
Page 2 of the paper: Tom: This page is where SkillProx actually gets built. Remember we were complaining that earlier methods just patch the skill and hope for the best? Here the forward step re-executes the patched skill on the same batch before committing anything.
Jane: And it's a pretty simple gate. The candidate edit is accepted only if hard accuracy and mean cell accuracy don't drop. The paper even says the first candidate with a strict hard-accuracy gain terminates the search early.
Tom: So they run the skill, try an edit, run it again, and compare. If it doesn't help, they roll back to the previous snapshot and try a different edit, up to three attempts. That's a genuinely different discipline from open-loop patching.
Jane: What I like is the feedback loop. A rejected attempt feeds its measured changes and the attempted direction into the next diagnosis. So the diagnostician learns from what didn't work, not just from what the skill text sounds like.
Tom: That's the closed-loop part. But then page four also introduces the backward stage, right?
Jane: Exactly. Once forward evolution stops, they parse the skill into auditable knowledge units, like L2 sections and L3 reference files. Then they do leave-one-out ablation on a fixed validation set: remove one unit, run the skill, and see if performance goes up or down. A negative utility means removing that unit actually improves things.
Tom: So the paper puts it concretely: a positive value means removing the unit lowers performance, a negative value means the ablated version performs better. And the utilities are measured once and frozen, so they don't get recomputed during the later search.
Jane: Right. Then candidates are picked only if their cell utility is below minus 0 point 001, which is a pretty strict threshold. The ordering also matters, because they sort ascending by cell utility and break ties with hard utility, so the most harmful content gets processed first.
Tom: That connects directly to the motivation we saw earlier, where deleting a redundant section moved accuracy from 46 percent to 54 percent. Here it becomes an explicit, structured mechanism rather than a lucky cleanup.
Jane: So forward stage brings in new knowledge that's actually verified, and the backward stage prunes what turned out to be dead weight. They work on different timescales, which is the whole point of the framework.
Page 3 of the paper: Tom: So they finally take the method apart on this page. The headline for me is that removing the Prox cleanup stage hurts more than removing the closed-loop diagnosis. The full system gets 54 point 5, without the diagnosis it drops to 53, and without Prox it drops further to 52.
Jane: Wait, I'd have guessed the reverse. The forward loop is what catches bad edits in real time.
Tom: That's what I thought too, but the numbers say otherwise. Their interpretation is that Prox needs a well-optimized forward skill to shrink. If you run the backward cleanup on a poorly evolved skill, there's not enough useful knowledge to consolidate, so compression has less to work with.
Jane: Then the two stages are truly complementary, not just additive. They also report the lowest variance with both stages, plus or minus half a point. That suggests the combination makes the runs stable across seeds, which matters for reproducibility.
Tom: Now the tau sweep is the other new piece on this page. They vary the candidate threshold and plot compression against accuracy, and here's the striking part: at their default setting they get 52 point 3 percent accuracy while cutting the skill by 25 point 7 percent. Even at 74 point 9 percent compression, accuracy only falls to 51 point 0.
Jane: So a lot of the skill text was redundant or even harmful. I assume the curve turns downward somewhere, though.
Tom: It does, past roughly eighty percent compression. They're explicit that tau is a threshold-induced trade-off curve rather than a strict regularization path. You get a Pareto frontier where you can pick a compression level without sacrificing much accuracy.
Jane: Then the model size comparison — that's a nice practical detail. The 4B model ends up with a skill about forty-five percent longer than the 27B model, mostly in reference files, while the main skill files are nearly identical in length.
Tom: Exactly, both main files hover around fifteen thousand characters. But the 4B writes twice as much reference content, which the paper reads as smaller models needing more supplementary guidance. And the 4B also gets compressed more aggressively, about twenty-nine percent versus nineteen for the larger model.
Jane: That fits the update dynamics from earlier — smaller models write to the skill more often and get rejected more at the gate. On this page they add that within the 27B runs, longer final skills correlate negatively with accuracy. So the extra text isn't earning its keep.
Tom: Which is a useful reality check for anyone building agent skills. You want the skill to be exactly as long as it needs to be, and this method gives you a way to find that point empirically instead of guessing.
Page 4 of the paper: Tom: Page ten is where the paper finally shows you the two concrete failure modes that motivated the whole design, and honestly, the examples are striking.
Jane: Oh good, because the method section was pretty abstract. What did they actually observe?
Tom: In the forward case, they watched an open-loop training run hit a bug with a sequential scan where the reference value keeps changing. The diagnosis looked reasonable, so the skill just wrote a meta-instruction saying "trace a concrete example before coding," plus a template with the threshold 1 point 10 hard-coded into it.
Jane: And that template had negative utility when they measured it?
Tom: Exactly. The leave-one-out audit gave it a cell utility of minus 0 point 0337 and hard utility of minus 0 point 0556. The closed-loop version learned something different: update the reference to the current value after each qualifying event. That formulation flipped the sign to plus 0 point 0495 and plus 0 point 0474.
Jane: So the same underlying mistake produced completely different knowledge depending on whether you re-executed the patch before committing it.
Tom: Right, and the difference is actionability. "Trace carefully" doesn't change behavior, but "update the reference after each event" does, so the gate can actually test it.
Jane: Then the backward motivation shows why re-execution alone isn't enough, doesn't it?
Tom: It does. Even the closed-loop skill still contained two negative-utility units, and the instruction "trace a concrete example before coding" showed up in four separate sections, with utilities ranging from plus 0 point 1038 all the way down to minus 0 point 0337.
Jane: So the same sentence can be genuinely useful in one place and actively harmful in another. That's a mess.
Tom: The Prox stage found five candidate units in that run, but only one edit actually passed the validation gate. It removed the task-specific template and consolidated the transferable principle into a positive-utility section.
Jane: And that one accepted edit compressed the skill from 29,129 characters to 28,219, which is about 3 point 12 percent, and pushed validation cell accuracy from 96 point 05 percent to 99 point 73 percent? That's a huge jump for such a small change.
Tom: It is, and I love how they summarize the division of labor at the end of the page: closed-loop forward is online update verification, and backward Prox is post-training utility refinement. One decides how knowledge gets introduced, the other decides what survives after accumulation.
Jane: That really clarifies why you need both stages, rather than just one clever trick.
Page 5 of the paper: Tom: Page thirteen gives us the ten-seed comparison on Qwen3 point 6-27B, and the headline isn't really the average gain. Closed-loop evolution improves hard accuracy by 1 point 1 points on average, but the standard deviation drops from 2 point 50 to 1 point 51, and the worst seed climbs from 46 to 49. That's the stability story.
Jane: So the real win is that the bad runs get rescued, not that the good runs get better.
Tom: Exactly. Seed eight is the clearest case: it's the weakest open-loop run at 46 hard accuracy, and closed-loop brings it to 51. The paper notes that result sits right near the closed-loop mean, so it's lifting the lower tail, not stretching the upper bound. They also flag a negative 0 point 80 correlation between open-loop accuracy and improvement, but they immediately caution that this is partly mathematical coupling.
Jane: Good, because that number would be easy to over-read. Then they get into the process difference, which is more concrete: open-loop commits every patch without re-executing the updated skill, and all ten patches in the seed-eight run were committed blindly. Closed-loop re-executes the same four-task batch after each patch and only accepts if hard accuracy doesn't drop and cell accuracy stays within a tiny tolerance.
Tom: And the gate trace from that run shows 22 attempted edits with 8 regressions blocked, and one entire iteration got reverted. So that's direct evidence the mechanism actually catches bad edits.
Jane: The other thing that stands out is feedback. Open-loop diagnosis gets no outcome signal at all, so later diagnoses can't learn from what failed earlier. Closed-loop injects both a within-iteration rejection context and a cross-iteration prior of the last six accept or reject records.
Tom: Right, so the diagnosis itself is evolving along with the skill. That's what makes the forward loop closed in a real sense, and it connects back to the earlier discussion of unverified updates. Here we see the update gate working in practice, with specific numbers from a specific run.
Page 6 of the paper: Tom: This page gives the concrete numbers from that seed 8 case study. We see the single accepted Prox edit reduce the complete skill by 3 point 12 percent, from 29,129 to 28,219 characters. Validation cell accuracy improved from 96 point 05 percent to 99 point 73 percent.
Jane: Did validation hard accuracy stay put through that?
Tom: It stayed at 94 point 74 percent. Then on the independent OJ evaluation, hard accuracy went from 46 percent to 54 percent. Mean cell accuracy went from 74 point 71 percent to 77 point 97 percent. That's the kind of jump that explains why removing Prox in the ablation cost us 2 point 5 points.
Jane: The table lists eight tasks flipping from fail to pass, and none going the other way.
Tom: Each row spells out the behavioral difference. One task now determines debit/credit sign direction correctly. Another handles cutoff times crossing midnight. Another identifies the true data range B3:B36. These are the spreadsheet behaviors you'd hope a skill would teach.
Jane: And they're honest about what this does and doesn't prove.
Tom: Right. They note the OJ conditions are independent generations at temperature 0 point 7, not paired rerolls, so you can't say the edit alone caused each transition. They even exclude one fail-to-pass because it was an API error in the pre-Prox condition.
Jane: So the strong claim stays local.
Tom: Exactly. Validation-gated Prox removed a negative-utility template and consolidated the transferable principle. It did that without degrading validation performance. That's the part they can defend.
Page 7 of the paper: Tom: Page 19 is where the paper shows the actual prompts that run the forward loop, and the momentum agent is the one that caught my eye. It reads the batch diagnoses plus a memory file from previous iterations, then turns each failure into a pattern—the prompt literally says "a pattern is a class of mistake or success, not a task instance."
Jane: That means the skill doesn't learn a fix for one spreadsheet, it learns a fix for a whole category of spreadsheet problems. And the momentum agent keeps a "remedy log" that's append-only history, so it remembers what was tried before and what actually worked.
Tom: The patcher prompt then takes that pattern record and insists on iterating by pattern, not by task. It tells the model to group overlay entries sharing a pattern, brainstorm two to three candidate remedies, and apply the simplest edit that generalizes.
Jane: Wait, so the patcher is allowed to brainstorm multiple options? That's a big step beyond just writing down the first diagnosis that comes along.
Tom: Exactly. And it explains why the closed-loop gate works so well—you're not just testing a random edit, you're testing one of several deliberately chosen remedies. The prompt also orders the patcher to prefer extending an existing section over creating a new one, which keeps the skill from ballooning in size.
Jane: There's a hard structural rule in there too, though. The patcher is told never to put task-specific columns, rows, filenames, or constants into the main skill file, and every reference file must have exactly one L2 pointer.
Tom: Right, that's what makes the leave-one-out audit possible in the backward stage. Each reference file is a cleanly separable unit, so you can remove it and measure the effect without breaking the rest of the skill.
Jane: The prompt even tells the patcher to read back all changed files after editing and repair broken pointers, orphaned references, and duplicate sections. That cleanup step keeps the skill structurally valid, which the validation gate depends on.
Tom: And that's the engineering behind the results we saw earlier—the lower variance across seeds and the consistent gains. It's a very deliberate way of forcing the model to consolidate knowledge rather than accumulate every diagnosis as a new rule.
Jane: What strikes me is how much of the method lives in these prompts. The forward-backward framework is the math, but page 19 is where the framework becomes executable instructions.
Conclusion: Tom: So putting it all together, SkillProx is really about treating a skill as something you can optimize rather than just write once and hope for the best.
Jane: Exactly. The forward loop checks whether an edit actually helps before keeping it, and the backward loop cleans out the knowledge that turned out to be dead weight.
Tom: That combination is what makes the gains hold up across different models and even on out-of-distribution tasks, which is the part I find most impressive.
Jane: Me too. The skills were only trained on spreadsheets, yet they still helped on WikiTQ and HiTab. That suggests the method is capturing genuinely transferable procedures, not just memorizing task patterns.
Tom: And there's a nice practical angle, too. Smaller models benefit disproportionately because the gate and the pruning step act as an external filter they wouldn't have on their own.
Jane: That's a big deal for deployment. You can get better behavior from a cheaper model without any weight updates, just by giving it a better-maintained skill file.
Tom: It also reframes how we think about skill growth. Bigger isn't better; the paper shows that carefully shrinking a skill can improve accuracy while cutting a quarter or more of its text.
Jane: Right, the compression–accuracy curve is striking. You can remove a lot of redundant content before performance ever starts to dip.
Tom: So the takeaway really is that self-evolving agents need both verification and consolidation, not just endless patching. SkillProx gives a clean framework for both.
Jane: And the authors are releasing the code, so other people can build on it. That should accelerate a lot of follow-up work in agent memory and skill management.
Tom: Great note to end on. Thanks to everyone listening, and we'll be right back with the next paper.
Mingxuan Zheng, Yujin Zhou, Chuxue Cao, Boqin Yin, Yuyao Zhang, Jiapeng Sun, Shuaishuai Gong, Sirui Han, Yike Guo
Hong Kong University of Science and Technology · Macau University
cs.AI, cs.CL
Submitted: 2026-08-07
Updated: 2026-08-10
Comments: 23 pages, 4 figures
Code: https://github.com/Steven011018/SkillProx
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 72/100
The gist: The paper, authored by Mingxuan Zheng, Yujin Zhou, Chuxue Cao, Boqin Yin, Yuyao Zhang, Jiapeng Sun, Shuaishuai Gong, Sirui Han, and Yike Guo (Hong Kong University of Science and Technology and Macau
Key concepts
- Closed-loop forward diagnosis
- A method where a skill patch is tested by re-running it on the same task batch before acceptance. Only if performance doesn't drop is the edit kept, preventing unverified updates that could harm the skill.
- Backward Prox (proximal refinement)
- A cleanup stage that decomposes a skill into knowledge units, removes each temporarily to measure its utility, and deletes or consolidates harmful ones. This controls skill complexity and prevents accumulation of useless content.
- Leave-one-out utility audit
- A technique to evaluate each knowledge unit's contribution by removing it and observing performance changes. Positive utility means removal hurts, negative means removal helps, guiding which units to prune.
Terminology
Summary
The paper, authored by Mingxuan Zheng, Yujin Zhou, Chuxue Cao, Boqin Yin, Yuyao Zhang, Jiapeng Sun, Shuaishuai Gong, Sirui Han, and Yike Guo (Hong Kong University of Science and Technology and Macau University), presents SkillProx, a framework for evolving LLM agent skills through a proximal-gradient-inspired forward–backward optimization approach.
The paper addresses the challenge of skill evolution for LLM agents. Skills are lightweight, reusable textual artifacts that are loaded into the agent’s context without weight updates.
While prior work focuses on skill synthesis (constructing skills from demonstrations or trajectories), the authors note that one-shot synthesis can provide uneven or even negative gains across heterogeneous tasks, since a skill constructed once cannot anticipate the diverse situation it will later encounter.
This motivates a shift from static skill construction to skill evolution, where methods like SkillGrad treat task failures as textual gradients, diagnosing failures as new tasks arrive and directly commiting the inference experience as a patch.
The paper identifies two critical limitations in existing skill evolution methodologies:
-
Unverified forward updates
: "An LLM-generated diagnosis is commonly treated as a valid update direction without verifying the realized effect of the resulting skill edit. Once a patch is generated, it is typically committed directly, and its effectiveness is neither measured through re-execution nor provided as feedback to subsequent diagnoses." -
Unregulated skill growth
: "Iterative patching continuously expands the skill without an explicit mechanism for reassessing accumulated knowledge. The resulting artifact may contain repetitive instructions, conflicting heuristics, or task-specific solutions incorrectly generalized as reusable rules, some of which can interfere with useful knowledge rather than merely increase textual complexity."
The central question is: how can outcome-verified forward diagnosis be coupled with structure-aware backward refinement, so that a skill's task capability and its accumulated structure co-evolve?
SkillProx is motivated by a composite optimization problem: min Jλ(X):= LT(X) + λG(X), where LT(X) is the expected task loss and G(X) is textual complexity (total character count). The formulation provides a common basis for examining performance-oriented skill updates and complexity-oriented knowledge consolidation.
The method draws an analogy to proximal gradient descent (PGD): diagnostic evolution serves as an inexact, performance-oriented forward operator, while utility-aware consolidation serves as a discrete backward operator for controlling accumulated knowledge.
The paper explicitly states that skill text is discrete and non-differentiable, so the implementation does not explicitly optimize
the objective or take η and λ as inputs
; only the forward–backward division of responsibilities is retained.
At each forward iteration k, the current skill is executed on training batch Bk. A diagnostician uses failed trajectories, contrastive successful trajectories, recent history, and the rejection reason from the previous attempt to propose an edit direction.
The Patcher then produces a candidate skill.
The key innovation is outcome-grounded verification: the candidate is re-executed on the same batch and accepted only if it satisfies a performance gate:
- Gatefwd(X̃(j)k) = 1[H Bk(X̃(j)k) ≥ H Bk(Xk)] × 1[C Bk(X̃(j)k) ≥ C Bk(Xk)]
This means the update is accepted only when both hard accuracy and mean cell accuracy do not decrease. The first candidate with a strict hard-accuracy gain terminates the search early.
If no candidate strictly improves hard accuracy, up to three attempts are evaluated, and the lexicographically best candidate is selected, still subject to the gate. For rejected attempts, the observed hard/cell changes and attempted directions are provided to the next diagnosis,
forming compact accept/reject summaries
that serve as semantic history for subsequent diagnoses.
After forward optimization, the skill Xf is parsed into auditable knowledge units (L2 sections and L3 reference groups). The method performs a frozen leave-one-out utility audit:
-
u hard i = H V(Xf) − H V(Ablate(Xf, qi))
-
u cell i = C V(Xf) − C V(Ablate(Xf, qi))
All utilities are measured once before Prox and remain frozen throughout the candidate traversal.
Candidates are those with cell utility below threshold τ = −0.001, ordered lexicographically by cell then hard utility.
The Shrinker then processes candidates one at a time, proposing consolidations (salvaging generalizable content into retained sections), demotions (moving L2 detail into L3 references), or removals. Each trial must satisfy:
-
Structural validity and strict complexity reduction: StructOK(Tm) = 1, G(Tm) < G(X(m))
-
Validation gating: H V(Tm) ≥ H V(X(m)) − δh, C V(Tm) ≥ C V(X(m)) − δc, with δh = 0, δc = 0.02
-
Soft compression cap: ρ = 0.10, evaluated relative to Xf, which
is a soft stopping threshold
—the final accepted edit may move the total compression beyond 10%.
Accepted edits replace the active skill; rejected trials are discarded without terminating the loop. The process terminates finitely since each candidate is processed at most once
and every accepted edit must strictly reduce text complexity.
The authors evaluate on:
-
IID: SpreadsheetBench Verified
-
OOD: WikiTableQuestions (WikiTQ) and HiTab
They compare against baselines: No Skill, Human Skill (which serves as the initialization SkillProx optimizes), EvoSkill, Trace2Skill, SkillOpt, and SkillGrad, across three backbone LLMs: Qwen3.5-4B, Qwen3.5-27B, and Qwen3.6-27B. Data splits follow a 2:1:8 train/validation/test ratio with a maximum of 30 interaction turns.
Key findings from Table 2:
In-domain (SpreadsheetBench): SkillProx attains the best IID accuracy on all three backbones (21.0, 51.3, and 54.5).
It "delivers consistent gains on every backbone: it lifts the base skill by 13.0 pp on Qwen3.5-27B (38.3 → 51.3) and by 17.8 pp on Qwen3.6-27B (36.7 → 54.5), while remaining slightly positive on the smaller Qwen3.5-4B (20.3 → 21.0). SkillProx also
exhibits the lowest variance among the self-evolving methods (e.g. ±0.5 vs. ±7.6 for SkillOpt on Qwen3.6-27B)."
Out-of-domain: "On WikiTQ, SkillProx achieves the best accuracy on Qwen3.5-4B (78.5), improving over No Skill by 13.5 pp and over SkillGrad by 8.8 pp, and is also best on Qwen3.5-27B (86.8) while remaining competitive on Qwen3.6-27B (86.2). On HiTab it is best on both Qwen3.5-4B (69.2) and Qwen3.6-27B (80.0). In contrast,
SkillOpt, which overfits to the in-domain style and collapses OOD, e.g. on the 4B model (26.0 on WikiTQ, 16.0 on HiTab). The overall claim is that SkillProx
improves average accuracy by 3.0 percentage points over the strongest gradient-based baseline."
Component ablations on SpreadsheetBench with Qwen3.6-27B (Table 3) show:
-
Removing closed-loop diagnosis: accuracy drops from 54.5 to 53.0 (−1.5 pp)
-
Removing Prox: accuracy drops to 52.0 (−2.5 pp)
-
Full method: 54.5 ± 0.5, with the lowest variance
The authors note: "The larger drop from removing Prox indicates that task-driven forward editing alone accumulates redundant or overly instance-specific content, which the proximal shrinkage stage is needed to control; the smaller but consistent drop from removing the closed-loop diagnosis indicates that Prox still requires a well-optimized forward skill to shrink."
Accuracy–Compression Pareto frontier (Figure 2): Applying Prox offline to closed-loop skills without retraining, "At τ = −0.001, accuracy reaches its maximum of 52.3% with 25.7% compression. At τ = 0.005, the compression ratio increases to 41.5% while accuracy remains at 52.0%. Even at τ = 0.050, 74.9% of the skill is removed while retaining 51.0% accuracy. Performance begins to decline when compression exceeds approximately 80%. Several Prox settings achieve
both higher compression and higher accuracy than no Prox, indicating that moderate shrinkage may also reduce duplicated rules and context interference."
Model size and skill length (Table 4): "The final 4B skill is approximately 45% longer than the 27B skill on average. Interestingly, the main skill files converge to nearly the same length of approximately 15k characters. The difference arises almost entirely from reference files: the 4B model produces roughly twice as much reference content as the 27B model. The 4B model undergoes 29.4% Prox compression vs. 19.0% for the 27B model.
G3D length is negatively correlated with IID hard accuracy within the 27B runs, indicating that a longer skill is not necessarily a more effective one" (Pearson r = −0.628, n = 9, though sensitive to seed 8).
The appendix provides detailed case studies on Qwen3.6-27B seed 8:
Forward motivation: Open-loop evolution converts a failure into the meta-instruction trace a concrete example before coding
with a hard-coded template (condition value >= reference * 1.10), which has negative utility (u cell = −0.0337, u hard = −0.0556). Closed-loop evolution instead learns an actionable decision rule: after each qualifying event, update the reference to the current value before continuing the scan,
with positive utility (u cell = +0.0495, u hard = +0.0474). The gate selects for falsifiable edits.
Backward motivation: The closed-loop skill still contains two negative-utility and fourteen zero-utility units.
A frozen utility audit identifies five candidates, but only one edit is accepted
— consolidating the transferable state-tracing principle while removing the repeated template. This reduces the skill from 29,129 to 28,219 characters (3.12% compression) while improving OJ hard accuracy from 46% to 54% and cell accuracy from 74.71% to 77.97%, with eight fail-to-pass transitions and no pass-to-fail transitions.
Ten-seed comparison: "closed-loop evolution records six wins, two ties, and two losses in hard accuracy, with a mean hard improvement of 1.10 points (and a mean cell improvement of 1.31 points). The cross-seed standard deviation decreases from 2.50 to 1.51, and the minimum increases from 46 to 49. The correlation between open-loop accuracy and improvement is −0.800, which the authors caution
should not be over-interpreted" due to mathematical coupling.
The paper's stated contributions are:
-
A forward–backward formulation of skill evolution
revealing two missing components:outcome-grounded verification of forward updates and utility-guided regularization of accumulated knowledge.
-
The SkillProx framework,
combiningclosed-loop diagnostic co-evolution including same-batch re-execution, rollback, retry, and accept/reject memory with validation-gated proximal shrinkage.
-
Empirical validation of diagnostic–proximal co-evolution
across three backbone LLMs, six baselines, and one in-distribution plus two out-of-distribution benchmarks.
The paper concludes that SkillProx couples closed-loop diagnostic evolution, which verifies candidate updates through execution feedback, with utility-aware proximal refinement, which reassesses and consolidates accumulated knowledge,
and that Empirical results suggest a practical path for jointly improving diagnostic updates and accumulated skill knowledge through forward–backward optimization.
Improvements for AI systems
Improvements to AI systems based on SkillProx:
- Verified skill/prompt updates
The improved AI system will not blindly commit LLM-generated diagnoses or patches. Before accepting any new skill instruction, prompt edit, or policy change, it re-executes the candidate on the same batch and accepts only if hard accuracy and cell-level accuracy do not degrade.
- Closed-loop diagnostic co-evolution with rollback and retry
When a patch fails verification, the system will roll back to the previous skill, record the rejection reason, and feed that outcome into the next diagnosis. Up to three candidate attempts are evaluated, and the lexicographically best non-degrading candidate is selected.
- Retroactive knowledge consolidation via proximal refinement
After forward improvement, the system parses the skill into auditable units, measures each unit's utility by frozen leave-one-out ablation on validation data, and shrinks the skill by removing, demoting, or consolidating only units with negative or zero utility.
- Falsifiable, causality-aware skill generation
Instead of adding generic meta-instructions or hard-coded templates, the system learns actionable decision rules that are verified by execution, filtering out edits with negative utility and preventing overfitting to instance-specific failures.
- Controlled skill growth and compression
The system will enforce strict complexity reduction per accepted edit, limit compression softly (e.g., 10% relative to the forward-optimized skill), and stop only when no valid shrinkage remains. This prevents repetitive instructions, conflicting heuristics, and context-window bloat.
- Improved out-of-distribution robustness
By coupling execution-verified forward updates with validation-gated backward pruning, the system avoids overfitting to in-domain styles, preserves transferable principles, and maintains accuracy on unseen tasks where open-loop or no-verification methods collapse.
- Lower variance and more reliable skill evolution
The system reduces cross-seed variance in final skill performance, increases minimum outcomes, and avoids both catastrophic forgetting (via hard-accuracy gates) and unnecessary context interference (via utility-based removal).
- Self-tuning reference knowledge organization
The system can demote low-value detailed instructions into reference-only sections while salvaging generalizable content into the main skill, keeping the active context compact and the reference layer available when needed.
- Automatic identification of negative-utility knowledge
The system audits each skill unit on validation data, identifies units whose removal actually improves accuracy, and removes or rewrites them—leading to cases where compression simultaneously raises accuracy.
- Generalization to other agent components
The same forward-backward framework can improve system prompts, tool-use policies, few-shot example sets, memory entries, and robot task plans: forward proposals are verified by environment feedback, and backward consolidation prunes memorized artifacts that no longer earn their place.
Sources
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
- EvoSkill: Automated Skill Discovery for Multi-Agent Systems
- HiTab: A Hierarchical Table Dataset for Question Answering and Natural Language Generation
- ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
- SoK: Agentic Skills -- Beyond Tool Use in LLM Agents
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
- Large Language Model Agent: A Survey on Methodology, Applications and Challenges
- SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation
- Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
- Compositional Semantic Parsing on Semi-Structured Tables
- Kimi K2.5: Visual Agentic Intelligence
- Tongyi DeepResearch Technical Report
- SkillGrad: Optimizing Agent Skills Like Gradient Descent
- SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
- Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward
- SkillOpt: Executive Strategy for Self-Evolving Agent Skills
- Skill is Not One-Size-Fits-All: Model-Aware Skill Alignment for LLM Agents
- EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction
- TextGrad: Automatic "Differentiation" via Text
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection