SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure

arXiv:2608.11079 · cs.AI · Submitted 2026-08-16 · Read on arXiv

Xiaofan Bai, Hongqiang Lin, Chao Liu, Yantao Zhang, Xuan Jin, Xipeng Cao, Yuhong Li

Alibaba Group · Zhejiang University · Duke University

cs.AI

Submitted: 2026-08-16

Updated: 2026-08-18

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 50/100

The gist: SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure Abstract Summary The paper addresses the problem of skill bloat in self-evolving agents.

Terminology

Summary

SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure

Abstract Summary

The paper addresses the problem of skill bloat in self-evolving agents. These agents accumulate reusable skills by appending successful procedures and failure fixes, but over time, the same requirements are restated in multiple branches, examples, and warnings, while common action sequences are copied rather than reused. This makes the skill expensive to inject and difficult to maintain. Generic prompt compression is ill-suited because a skill is not a flat passage: its name and description define when it applies, its workflow controls execution, its tool and output contracts constrain validity, and rare exceptions may remain essential even when no sampled task activates them. Evaluation-guided compression can test these behaviors but introduces rollouts, cost, and dependence on the compression-time evaluation set.

The paper presents SkillZip, an evaluation-free method that compresses a skill by finding its shortest faithful structural explanation. The intuition is explain once, reference many: state a repeated rule once at the scope where it applies, factor a repeated action sequence into a shared procedure, and keep only the differences as explicit exceptions. This is formalized as a typed minimum-description-length objective over a skill contract and a residual, subject to a hard coverage constraint for every extracted trigger, workflow edge, tool requirement, obligation, and output field. The formulation provides simple sharing thresholds, preserves unique rare rules by construction, and supports efficient local updates. SkillZip has a one-shot mode with one structured extraction call and deterministic optimization, and a continual Zip-on-Write mode that integrates each self-evolution patch without replaying tasks or reparsing the full history. Comprehensive experimental evaluations demonstrate the effectiveness and superiority of SkillZip in compression performance, generalizability, and cost overhead.

Introduction Summary

The paper notes that a self-evolving agent improves by turning experience into reusable instructions. When a tool fails, it appends a warning; when an answer violates a format, it adds an example; when a rare branch succeeds, it records the successful procedure. Each update is locally reasonable, but the accumulated skill is usually edited as an append-only notebook rather than maintained as a coherent program. After enough rounds, never overwrite the source file may appear in the introduction, three workflow branches, and an example, while the same validate–repair–verify sequence is copied repeatedly with only minor differences.

This creates a systematic mismatch between textual growth and procedural growth. New text keeps accumulating even after the number of genuinely new requirements begins to saturate. Because a loaded skill occupies the context on every invocation, redundant text increases prefill cost and can obscure the instructions that actually govern execution. Figure 1 illustrates this phenomenon.

The paper distinguishes between different types of long skills. A conventional one-shot or community-authored skill can mix operational rules with background exposition, templates, examples, or even content unrelated to execution; deleting or deferring such material is often appropriate. A multi-round evolved skill is different. Recent systems accept bounded edits only after rollout feedback or validation, or aggregate recurrent successful and failed trajectories into updates. Its additional text is therefore usually more knowledge-dense: even a singleton warning may encode an expensive failure that the agent learned to avoid. The main redundancy is less often irrelevant topic and more often repeated representation—the same invariant copied into several branches, a workflow restated after each failure, or a general rule followed by progressively narrower exceptions. This distinction changes the compression goal from content filtering to knowledge consolidating.

The paper argues that prompt compression is the wrong abstraction. Prompt compressors typically decide which tokens are useful for a current query or likely answer, while skills should be reusable for all queries of certain tasks. More importantly, its meaning is not distributed uniformly across words. The name and description determine when the skill is selected; temporal phrases determine action order; tool arguments define valid calls; branch guards determine where a rule applies; and an output schema defines when the procedure is complete. A rare exception may be activated only once, yet deleting it can be more damaging than removing several paragraphs of rationale. Task-based validation is also not a complete solution. SkillReducer demonstrates that structure-aware rewriting and progressive disclosure can reduce skill cost, and its feedback loop uses generated tasks to recover missed content. Such validation is valuable, but it makes compression expensive and couples the compressed results to compression-time tasks. A method can repeatedly repair the branches observed by that set while still losing an untested guard or output constraint.

The paper asks a stricter question: Can a skill be shortened using only the structure already present in the skill, without observing tasks, rewards, trajectories, or verifiers?

The answer begins with a simple observation: a skill resembles a compact operating manual more than a generic prompt. It contains an interface (name, purpose, triggers, exclusions), a procedure (steps, branches, loops, fallbacks), contracts for tools and outputs, and rules whose scope can be global or branch-specific. This structure reveals redundancy that token importance cannot see. A rule repeated in every branch can be stated once before the branch. A repeated action sequence can be named once and reused. Several guarded variants can be written as one common rule plus explicit exceptions. Conversely, a unique requirement cannot be removed merely because it is short or rare.

The key compression principle of SkillZip is: prefer the shortest faithful explanation of the skill. Informally, the compressed representation pays once for every shared structure and pays separately only for genuine differences. This is the same intuition as replacing three copies of a code fragment by one function and three calls. Formally, it is an instance of minimum description length (MDL): minimize the cost of a compact skill contract plus the cost of the source details not explained by that contract, while requiring every extracted normative requirement to remain covered.

The paper designs two modes. One-shot SkillZip compresses an existing evolved skill checkpoint with one structured extraction call followed by deterministic optimization. Zip-on-Write participates in self-evolution: when a new self-evolution patch arrives, SkillZip compares it only with compatible parts of the current compact contract, then either absorbs it, refines an existing rule, adds a new requirement, or triggers a local refactoring. Periodic repacking captures patterns that become reusable only after several updates.

The main contributions are:

  • Formulating evaluation-free skill compression as contract representation tailored for evolved skills.

  • Deriving a shortest faithful explanation objective that unifies semantic sharing, scope lifting, workflow reuse, and exception encoding. A hard coverage constraint gives a rare-rule preservation guarantee independent of task frequency.

  • Designing one-shot SkillZip and Zip-on-Write. The first uses one structured extraction call followed by deterministic optimization; the second maintains a compact state during self-evolution and avoids replaying tasks or reparsing the full history.

  • Performing comprehensive empirical evaluations showing that SkillZip delivers substantial gains in compression performance, robust generalizability, and low cost overhead.

Related Work Summary

The paper reviews related work in four areas:

  1. Self-Evolving Agents and Persistent Skills: Agents increasingly convert interaction experience into reflections, memories, programs, or reusable skills. Systems such as ACE, SkillRL, and SkillClaw explicitly maintain and evolve persistent skill artifacts across tasks or users; SkillRevise and SkillGrad use execution traces to diagnose and improve an existing skill. These works focus on how procedural knowledge is acquired. As the artifact grows, acquisition and maintenance become different problems: a useful new patch can still duplicate an old invariant or repeat an existing workflow. SkillZip addresses this consolidation problem without replaying the experiences that produced the skill.

  2. Prompt and Context Compression: Prompt compression prunes or rewrites context using perplexity, query relevance, learned token labels, summarization, or latent representations. These techniques usually treat input as a sequence whose importance is estimated relative to a query, answer, or representative task distribution. Skills violate both assumptions: future tasks are unknown, and procedural meaning depends on typed relations such as before, branch guards, tool arguments, and required output fields. Query-conditioned compression may be effective for a single request, but a reusable skill must retain requirements that no compression-time query activates. SkillZip therefore compresses repeated procedural structure rather than low-salience tokens.

  3. Skill Compression and Efficient Execution: SkillReducer is the closest textual baseline. It minimizes routing descriptions with delta debugging, classifies body content, moves supplementary material to on-demand references, and uses faithfulness checks and task-based feedback. It establishes that skill content should not be treated as homogeneous text and is especially well matched to ordinary skills that contain background, examples, and reference material. Multi-round evolved skills present a complementary regime: accepted updates are often operationally relevant, while redundancy appears as repeated constraints, overlapping workflows, and accumulated exceptions. SkillReducer's generated tasks and feedback loop participate directly in candidate selection. Repeatedly tuning an artifact on a finite evaluation sample can create sample-specific selection effects. SkillZip removes this source of dependence by never exposing compression to tasks, and it keeps a single portable text skill rather than changing the loading architecture. SKIM and TokMem encode procedural knowledge into learned soft or memory tokens, while Skill-to-LoRA transfers a textual skill into model parameters. Their representations can be compact, but they are model-dependent and less convenient to inspect, diff, or update during continual evolution. SkillZip uses a structured sidecar only while maintaining the artifact and renders an ordinary human-readable skill for deployment. Skill and SkillRT treat skills as executable or compilable artifacts, supporting the premise that a skill has an interface, control flow, and contracts, not merely topical text.

  4. Minimum Description Length: The MDL principle selects the model that gives the shortest joint description of a model and the data it explains. Grammar-based methods such as SEQUITUR and Re-Pair similarly replace repeated subsequences with reusable rules. SkillZip adopts this intuitive principle—a shared explanation is useful when defining it once and referencing it is cheaper than repeating it—and adapts it to typed procedural knowledge. Unlike sequence compression, the objective may not merge identical-looking clauses if they occur under incompatible guards, and it may not delete a unique rule even when that rule has zero repetition. The hard coverage is therefore as important as the skill length.

From Skill Text to a Compact Contract Summary

The paper defines evaluation-free compression: Let S be one textual skill and Se its compressed form. During compression, a method may read S, files referenced by S, and—in the continual setting—the sequence of skill patches. It may not access downstream tasks, execution trajectories, rewards, or behavioral verifiers. These resources are used only after compression to measure generalization. Three properties are sought: (1) Se should be substantially shorter under the deployment tokenizer; (2) it should preserve what the skill requires, including rare conditions; (3) it should remain an ordinary text artifact that can be inspected, versioned, and used by different agent backbones.

The second requirement is deliberately structural. Without running tasks, the compressor cannot prove that arbitrary natural language induces identical model behavior. It can, however, require the compressor to preserve every operational element it extracts from the source, and to retain ambiguous source spans verbatim. This makes the boundary of the guarantee explicit rather than hiding it behind a finite evaluation suite.

The paper describes the contract hidden inside a skill. Figure 2 shows the representation used by SkillZip. The extracted contract is written as:

C(S) = ⟨I, G, T, C, O, E⟩

where:

  • I is the interface: skill name, purpose, positive triggers, and exclusions;

  • G is the workflow: actions, order, decisions, loops, fallbacks, and stop conditions;

  • T is the tool protocol: tool names, required arguments, preconditions, expected observations, and error handling;

  • C is the set of scoped rules: what the agent must, must not, or preferably should do, together with the scope and guard under which each rule applies;

  • O is the output contract: response type, required fields, ordering, validation, and completion conditions;

  • E is supporting evidence: examples, templates, and rationale linked to the contract elements they express.

This decomposition changes which compressions are safe. Two sentences about the same tool cannot merge if they require different arguments. A prohibition repeated in all branches may move to the common parent scope, but a prohibition attached to only one guarded branch may not. Two examples can be removed if they merely illustrate an explicit output schema; an example that is the sole source of a required field must first be converted into an explicit output rule.

The paper defines typed units, scope, and coverage. The parser maps source spans to typed units a ∈ A. A unit records:

a = (τ, σ, g, m, p, P)

where τ is its type, σ its scope, g an optional guard, m its modality, p its normalized content, and P the supporting source spans. A workflow unit additionally records incoming and outgoing edges; a tool unit records its argument signature; an output unit records field and validation information.

If a source span cannot be interpreted with sufficient confidence, it becomes a locked residual: it is copied verbatim and excluded from deletion. This conservative fallback is important because the structural parser, rather than the later optimizer, is the main source of semantic uncertainty.

Coverage is defined as: a ⪯ K when compact representation K covers unit a. Coverage can be direct, or structural: a branch-local copy can be covered by the same rule placed at the closest common ancestor; a repeated action sequence can be covered by a shared procedure whose expansion contains the original edges. Coverage is type-sensitive. A tool name does not cover its required arguments, and a general rule does not cover a conflicting guarded exception.

Contract coverage is formally defined: A compact representation K covers a parsed skill if ∀a ∈ Areq(S), a ⪯ K, where Areq contains interface conditions, workflow nodes and edges, tool requirements, scoped rules, and output requirements.

Theoretical Analysis Summary

The paper defines the shortest faithful explanation. A compressed skill is represented by a library K of reusable contract elements and a residual R for unique, exceptional, or uncertain content. The library may contain shared rules, workflow fragments, tool contracts, output fields, and interface entries. The residual preserves explicit exceptions that cannot be safely normalized.

SkillZip selects the shortest representation that still covers every required contract unit:

(K*, R*) = arg min (K,R) ∈ H(S) [L(K) + L(RK)]

s.t. a ⪯ (K,R), ∀a ∈ Areq(S)

Here, H(S) is the finite set of candidate representations proposed from the parsed skill, and L(·) measures the rendered token cost, including the overhead of definitions, references, and scope annotations. This is an instance of minimum description length, but its operational meaning is simple: compression may change how a requirement is written, but not whether it remains represented. In particular, a unique requirement cannot be removed merely as it is short or unsupported by frequently sampled tasks.

The same objective governs four forms of reusable structure:

  • Equivalent requirements: paraphrases are represented once when a shared rule plus any residual differences is shorter than keeping separate copies.

  • Repeated rules across scopes: a rule is moved to the nearest common scope only when it applies to every relevant path; conflicting branch-local behavior remains as an explicit exception.

  • Repeated workflows: a recurring action sequence becomes a shared procedure only when the saved repetitions exceed the cost of defining and calling it.

  • Guarded variants: related clauses may be represented as one common rule plus guarded deltas, but only when the exception structure is both faithful and shorter than listing the variants separately.

These decisions are not independent rewrite heuristics. They are alternative ways of covering the same typed contract, compared under one length objective.

The paper proves Proposition IV.1 (Parsed-contract preservation): If every normative source span is represented by a typed unit or residual, any feasible solution of Eq. (4) preserves all extracted requirements. This follows directly from the hard coverage constraint.

Corollary IV.2 (Rare-rule preservation): The preservation of a unique requirement does not depend on how often its branch appears in any compression-time task distribution. This is the key theoretical benefit of evaluation-free compression. A guard, tool argument, exception, or output field is protected because it belongs to the parsed contract, not because sampled tasks happen to activate it. The guarantee is intentionally limited to the extracted contract; uncertain spans are therefore kept verbatim rather than silently discarded.

The objective also yields an efficient continual form. A new patch usually changes only a small type–scope neighborhood of the current compact contract. If the patch creates no profitable abstraction that crosses this neighborhood, all other cost terms remain constant, so local optimization gives the same update as rerunning the batch objective. When several patches collectively create new cross-scope reuse, occasional global repacking restores the missed saving. Thus local updates provide efficiency, while repacking recovers long-range structure.

SkillZip Method Summary

The paper describes two modes:

One-Shot Compression has five steps:

  1. Scan the SKILL.MD before using a model: A deterministic scanner parses front matter, headings, nested lists, code blocks, tables, and file references. It copies the skill name and description as interface candidates, converts Markdown nesting into a preliminary scope tree, and records stable source-block identifiers. Numbered lists and temporal markers provide high-confidence workflow hints.

  2. Recover the typed contract once: A schema-constrained model receives the numbered blocks and returns the contract: interface entries, workflow nodes and edges, tool calls and required arguments, scoped rules with modality and guard, output fields, and links from examples to the requirements they specify. Every extracted unit must cite source blocks. The host rejects unsupported citations, polarity mismatches, unknown tool names, and invalid workflow references. Ambiguous spans are placed in the locked residual. The extractor is deliberately not asked to compress. Separating interpretation from optimization has two benefits: contract recovery can be evaluated against human annotations, and the optimization is deterministic once the extracted units are fixed.

  3. Propose only type-compatible reuse: Candidate generation uses hard structural blocking: (a) interface entries compare with the same trigger or exclusion role; (b) rules compare when modality, predicate family, and scope ancestry are compatible; (c) tool units require the same tool and compatible argument signatures; (d) output units compare within the same response type and field namespace; (e) workflow reuse is proposed from repeated guarded action sequences with identical entry and exit behavior. Exact matches are found by hashing. Near duplicates are retrieved with an embedding index and verified by a frozen relation checker that predicts equivalence, implication, conflict, or unrelatedness. Conflicts create exception candidates rather than merge candidates.

  4. Select the shortest covering explanation: Each candidate h receives a saving save(h) = L(separate form) − L(form using h). Non-positive candidates are discarded. Equivalent units are clustered after conflict filtering. Rule placement is solved by dynamic programming over the scope tree. Repeated workflow candidates are selected by non-overlapping weighted packing, using saving per covered token as the greedy order and pairwise exchange as a refinement. After each selection, the optimizer checks that all source units remain covered.

  5. Render with fixed templates and audit structurally: The renderer produces a normal skill with a concise purpose and triggers, global rules, a numbered workflow, nested guarded branches, explicit tool requirements, and an output checklist or schema. A shared workflow is named only when references save tokens; otherwise it remains inline. Exceptions are placed after their base rule. An optional structural audit reparses only the compressed skill and compares it with the selected contract. If a trigger, guard, workflow edge, tool argument, polarity, or output field is missing, SkillZip restores the shortest source span that covers it and locks that span.

Continual Compression: Zip-on-Write works as follows. A self-evolving agent usually produces a small patch rather than a complete rewrite. SkillZip stores a sidecar, skillzip.json, containing the current contract, source provenance, scope tree, workflow graph, and candidate indices. The rendered SKILL.MD remains the only artifact loaded by the agent.

For every patch unit, the updater compares four interpretations under the same objective:

  • ABSORB: the patch restates an existing requirement and adds no new contract content;

  • REFINE: the patch adds a guard, tool argument, validation, or explicit exception to an existing unit;

  • EXTEND: the patch introduces a genuinely new requirement;

  • REFACTOR: the patch makes a shared rule or workflow newly worthwhile.

The host selects the feasible operation with the smallest increase in Eq. (4). Candidate search is restricted to the matching type, current scope, ancestor scopes, and adjacent workflow nodes. Thus a patch with d extracted units compares against O(dk) retrieved candidates rather than the complete history.

Local updates may miss reuse that becomes profitable only after several patches. The sidecar therefore tracks approximate counts for rule families and action n-grams. A global repack is triggered when the estimated recoverable saving exceeds θ repack, the contract grows by more than ρ since the last repack, or B patches have arrived. Repacking operates on the compact contract rather than all historical prose.

Skill compression can also be packaged as a skill. The model receives the new patch and a retrieved slice of the sidecar, then proposes one of the four operations with source citations. A deterministic host validates the schema, recomputes the saving, enforces coverage, writes a transaction log, and atomically replaces the skill only after rendering succeeds. The model proposes structure but never directly mutates persistent state.

Crucially, Δt is frozen before compression: the evolver decides what knowledge is learned, while SkillZip decides only how that knowledge is represented. Absorb is feasible only when the patch adds no uncovered contract unit; Refine preserves the old unit and records its new guard, argument, validation, or exception; Extend adds a new required unit; and Refactor changes representation without changing coverage. No operation is accepted because it improves a task score.

Experiments Summary

The paper evaluates SkillZip along two claims: self-evolution introduces substantially more surface text than new procedural knowledge, and this redundancy can be removed without consulting downstream evaluations.

Experimental Setup: Three agent model backbones are evaluated: Qwen3.7-Max, Qwen3.6-Plus, and Kimi K2.6. For every model–benchmark pair, all skill conditions use the same model snapshot, agent scaffold, system prompt, tool definitions, maximum interaction budget, and decoding configuration. Three benchmarks are considered: BFCL-v4 Web Search (multi-step web retrieval and reasoning), LiveMathematicianBench (theorem-grounded multiple-choice problems), and SpreadsheetBench (spreadsheet manipulation).

For each model–benchmark pair, a manually authored human skill is created, then SkillOpt is applied to improve this seed skill using the designated evolution split. The resulting skill is frozen after evolution and used as the common input to all compression methods. The evolution split is disjoint from the final benchmark test set. For one-shot SkillZip, a single fixed compressor model (Qwen3.7-max) is used across all skills, whereas for continual Zip-on-Write the agent's backbone model compresses its own skill. In both modes, the model is invoked only for schema-constrained contract extraction and relation/merge adjudication (greedy decoding, temperature 0); the minimum-cost covering and rendering are deterministic.

Five baselines are compared: (1) No Skill, (2) Human Skill, (3) Evolved Skill (the complete, uncompressed skill produced by SkillOpt), (4) SkillReducer (applied once to the frozen evolved skill), and (5) SkillZip. SkillReducer is the primary compression baseline. SkillZip receives exactly the same evolved skill but does not access benchmark tasks, execution trajectories, rewards, or behavioral verifiers during compression.

RQ1: Skill Growth in Self-Evolution: Figure 4 shows that skill length increases monotonically with self-evolution rounds across all benchmarks. By Round 5, the skills reach approximately 5.6×, 3.1×, and 6.7× their initial sizes on BFCL-V4, LiveMath, and SpreadsheetBench, respectively, with an average growth of about 5.2×. The growth persists across domains and is especially pronounced for tasks that continually accumulate tool-use procedures, failure corrections, and output constraints. Although each update may be locally useful, repeated rules, overlapping workflows, and increasingly specific exceptions accumulate without global consolidation, producing substantial skill bloat and increasing the context cost of every subsequent invocation.

RQ2: Can SkillZip Preserve Skill Fidelity?: Table I compares the uncompressed SkillOpt-evolved skills with their compressed ones and other skill settings. The evolved skill outperforms the no-skill and human-skill conditions in eight of nine settings, confirming that multi-round evolution generally accumulates useful procedural knowledge. SkillZip retains this knowledge while achieving compression rates of 27.1%–36.9% (31.2% on average). Its macro-average score is 0.577, slightly exceeding the uncompressed evolved skill at 0.570, and it matches or improves the evolved skill in five of nine settings. These results suggest that consolidating repeated rules, scopes, and workflows can reduce skill length without systematic behavioral degradation, even though SkillZip uses no tasks, rollouts, or verifiers during compression.

Compared with SkillReducer, SkillZip achieves both higher compression (31.2% vs. 9.2% on average) and higher task performance (0.577 vs. 0.544). This difference reflects the distinct compression regimes targeted by the two methods. SkillReducer is well suited to an initial debloating and quality-control pass over heterogeneous public skills, which may contain verbose background, redundant examples, or content unrelated to execution. In contrast, self-evolved skills are typically knowledge-dense because their updates arise from execution feedback; their main redundancy lies in repeatedly stated constraints, overlapping scopes, and copied workflows. Consequently, structural consolidation is better aligned with evolved-skill compression than content filtering or deferral.

RQ3: Compression Efficiency: SkillZip completes compression faster on all three datasets, reducing average time cost from 1331 to 207 seconds on LiveMath, from 1082 to 332 seconds on Spreadsheet, and from 587 to 318 seconds on BFCL-V4. Averaged across datasets, SkillZip requires 286 seconds, corresponding to a 3.5× speedup. SkillReducer uses fewer direct compression-model calls, but additionally requires 40–80 task rollouts for candidate validation and repair. In contrast, SkillZip uses several structured LLM calls but requires no task rollout on any dataset. The results indicate that environment interaction, rather than the number of compression calls alone, dominates the end-to-end cost of evaluation-guided compression. Consequently, the evaluation-free design of SkillZip substantially reduces latency while avoiding potential overfitting and dependence on executable tasks and behavioral verifiers.

RQ4: Cross-Model Generalization: Figure 6 evaluates whether a skill compressed from one source model can be executed by a different target model. On LiveMath, SkillZip achieves an overall retention of 0.97, compared with 0.91 for SkillReducer. Since their same-model results are comparable, the improvement mainly comes from off-diagonal source–target pairs, suggesting that preserving explicit rules, guards, and output constraints produces a more model-independent skill representation.

RQ5: Continual Zip-on-Write Compression: Figure 5 evaluates Zip-on-Write inside a 16-round self-evolution loop on LiveMath with three backbones. Without compression, the SkillOpt evolver exhibits the bloat pathology consistently across all three models: the length of skill grows monotonically to 2.5×, 3.1×, and 3.7× its seed length, respectively. Activating Zip-on-Write from round 1 bounds this growth for the entire trajectory, capping the skill at roughly 1.6×–1.9× across all three models—a 38%–50% reduction relative to the uncompressed endpoint—because each write is immediately absorbed into the typed contract library and periodically repacked.

Two further observations yield practical guidance. First, activation time matters: switching compression on only at round 8 recovers part of the accumulated redundancy but never catches up with the early-activation trajectory (e.g., 2.6× vs. 1.9× on Kimi-k2.6), showing that redundancy is cheaper to prevent than to remove. Second, compression does not trade accuracy for compactness: the final held-out test accuracy of the round-1 configuration matches or slightly exceeds the uncompressed skill on all three backbones.

Conclusion Summary

Self-evolving agents need a mechanism for forgetting repetition without forgetting procedure. SkillZip treats a skill as a typed contract and compresses it using a simple principle: explain shared structure once, reference it where needed, and keep genuine differences explicit. The resulting shortest-faithful-explanation objective is evaluation-free, protects rare requirements through hard coverage, and unifies rule sharing, scope placement, workflow reuse, and exception handling. One-shot compression requires one structured extraction call followed by deterministic optimization; Zip-on-Write integrates the same principle into continual skill evolution. The proposed experiments separate structural correctness, held-out behavior, evaluation-set overfitting, and cost, providing a practically reliable skill compression method tailored for self-evolving agents.

Improvements for AI systems

Based on this paper, here are the specific improvements I can make to AI systems and what the improved systems can do:

Improvement: Implement a structural compression layer that runs after every self-evolution update, using the typed contract representation (interface, workflow, tools, rules, output) and minimum-description-length objective.

What the improved system can do:

  • Detect when a new patch duplicates an existing rule, workflow, or constraint rather than adding new knowledge

  • Automatically refactor repeated action sequences into shared procedures

  • Move branch-local rules to common parent scopes when they apply universally

  • Maintain a sidecar JSON contract while rendering a compact human-readable skill for deployment

  • Prevent skill bloat from growing 5.2× on average, capping growth at 1.6–1.9× instead

Improvement: Replace evaluation-guided compression (which requires rollouts and task sampling) with a hard-coverage constraint that guarantees preservation of all extracted contract units, regardless of task frequency.

Improvement: Integrate compression directly into the self-evolution loop, processing each patch through four operations (ABSORB, REFINE, EXTEND, REFACTOR) with local optimization and periodic global repacking.

Improvement: Use explicit rules, guards, and output constraints rather than model-specific soft tokens or latent representations, making compressed skills model-agnostic.

Improvement: Use a two-phase approach: schema-constrained extraction that separates interpretation from optimization, with a locked residual for ambiguous spans.

Improvement: Activate compression from the first evolution round rather than after bloat accumulates.

Abstract

Self-evolving agents accumulate reusable skills by appending successful procedures and failure fixes. Over time, the same requirement is often restated in several branches, examples, and warnings, while common action sequences are copied rather than reused. The resulting skill becomes expensive to inject and difficult to maintain. Generic prompt compression is ill-suited to this setting because a skill is not a flat passage: its name and description define when it applies, its workflow controls execution, its tool and output contracts constrain validity, and rare exceptions may remain essential even when no sampled task activates them. Evaluation-guided compression can test these behaviors, but it introduces rollouts, cost, and dependence on the compression-time evaluation set. We present SkillZip, an evaluation-free method that compresses a skill by finding its shortest faithful structural explanation. The intuition is explain once, reference many: state a repeated rule once at the scope where it applies, factor a repeated action sequence into a shared procedure, and keep only the differences as explicit exceptions. We formalize this intuition as a typed minimum description-length objective over a skill contract and a residual, subject to a hard coverage constraint for every extracted trigger, workflow edge, tool requirement, obligation, and output field. The formulation provides simple sharing thresholds, preserves unique rare rules by construction, and supports efficient local updates. SkillZip has a one-shot mode with one structured extraction call and deterministic optimization, and a continual Zip-on-Write mode that integrates each self-evolution patch without replaying tasks or reparsing the full history. Through comprehensive experimental evaluations, we demonstrate the effectiveness and superiority of SkillZip in compression performance, generalizability, and cost overhead.

Sources

Related papers