DIVE: Unlocking Self-Improvement in Frozen Language Models Through Diversity-Driven Skill Evolution
Siheng Xiong, Ali Payani, Oguzhan Gungordu, Faramarz Fekri
Georgia Institute of Technology · Cisco Research
cs.CL
Submitted: 2026-08-12
Updated: 2026-08-14
Code: https://github.com/project-numina/aimo-progressprize
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: Large language models (LLMs) "cannot retain postdeployment experience without parameter updates." The paper introduces DIVE, "a diversity-driven framework that enables frozen LLMs to improve by
Terminology
Summary
Large language models (LLMs) cannot retain postdeployment experience without parameter updates.
The paper introduces DIVE, a diversity-driven framework that enables frozen LLMs to improve by evolving persistent natural-language skills from task experience and verifier feedback.
These skills encode reusable reasoning procedures, verification strategies, common failure modes, and output constraints and are both executed and revised by the same underlying model without access to a teacher model.
The authors note that "Fine-tuning can internalize such experience, but it requires access to model weights, substantial computation, and a carefully designed training pipeline. These requirements are increasingly restrictive as many capable language models are accessible only through APIs or deployed under limited adaptation budgets."
The paper identifies three key challenges in self-improvement of frozen models:
-
Noisy self-generated revisions:
a locally beneficial edit may remove useful guidance, overfit to a small set of failures, or amplify an incorrect reflection, making greedy single-candidate refinement brittle.
-
Context budget limitations:
accumulated experience quickly exceeds the context budget, requiring demonstrations, traces, and feedback to be distilled into compact, reusable abstractions rather than appended indefinitely.
-
Path-dependency of skill evolution: "different initial skills, sampled experiences, and revision trajectories can lead to substantially different solutions. Maintaining only a single evolving trajectory can therefore prematurely discard promising alternatives or converge to a suboptimal solution."
The central principle of DIVE is to preserve diversity throughout skill evolution and exploit the resulting complementarity at inference time.
Rather than refining a single incumbent skill, DIVE independently evolves multiple skill populations from bootstrapped experience.
For each population k ∈ [K], DIVE independently samples a bootstrapped experience subset D exp(k) and a bootstrapped reflection subset D ref(k) from D evo.
The experience subset provides verifier-labeled trajectories from which reusable task knowledge is extracted,
while the reflection subset is used to evaluate candidate skills, diagnose failures, and guide subsequent revisions.
Each population is initialized with an initial seed set S 0 by distilling reusable task knowledge from the verifier-labeled trajectories collected on D exp.
The seeds capture complementary aspects of the experience, including successful reasoning patterns, verification strategies, recurring failure modes, and output constraints.
DIVE maintains a portfolio of heterogeneous operators that support complementary forms of skill evolution.
The initial operator portfolio includes:
-
Reflective Repair (1 parent):
Locally revises a skill to address systematic failures revealed by verifier-labeled trajectories while preserving useful existing guidance.
-
Exploratory Revision (1 parent):
Encourages exploration of a substantially different solution strategy, decomposition, or reasoning approach rather than incrementally modifying the parent skill.
-
Compression (1 parent):
Removes redundant, overly specific, or conflicting guidance and consolidates the remaining knowledge into a more concise and reusable skill.
-
Recombination (2 parents):
Synthesizes complementary strengths from two parent skills into a coherent child skill while resolving redundancy or conflicting guidance.
DIVE uses an upper-confidence-bound (UCB) score
to balance exploitation and exploration:
UCB a(t) = μ̂ a(t) + β √(log t / N a(t))
where N a(t) denotes the number of times operator a has been selected before step t
and μ̂ a(t) is its empirical mean parent-relative reward.
The reward R τ is defined as the improvement of the proposed skill over its best-performing parent at step τ.
At a predefined step t new, DIVE uses the evolution history to generate N new operators that target recurring failure patterns or transformations insufficiently addressed by the existing operator portfolio.
The newly generated operators are added to the operator pool as untried operators.
After evolution, candidate skills from all populations are evaluated on a shared validation set and jointly selected to form a compact, complementary skill set.
The selection process uses greedy marginal improvement:
-
First skill:
selected according to its validation utility
-
Subsequent skills:
compute its marginal contribution to the validation utility of the complete skill set
and select the one with maximum marginal improvement -
Constraint:
at most one skill is selected from each population
At test time, every skill in S final independently generates a candidate response. The candidates are then ranked by f θ, and the highest-ranked candidate is returned as the final response.
For HMMT, we extract the final answer from each candidate response and select the answer using majority voting.
For other datasets, we present all candidate responses to the same frozen model in a single listwise ranking call.
The paper provides theoretical analysis for two central components:
"A heterogeneous portfolio is valuable because the most productive proposal mechanism may change as the population changes. Reflective Repair may be effective once a strong but imperfect skill has been identified, Exploratory Revision may be more useful when current strategies share a common failure mode, and Recombination may become useful only after complementary skills have emerged."
Independent evolution increases the probability that the overall candidate pool contains at least one low-risk skill, with diminishing returns as the number of populations grows.
Furthermore, a skill with lower standalone accuracy may be more useful than a higher-accuracy skill when its correct predictions occur on different instances.
DIVE is evaluated on six mathematical and logical reasoning tasks: HMMT, Equational Theories, Sudoku, Cryptarithm, Calcudoku, and Futoshiki.
For each dataset, we allocate 1,000 training examples to the evolution set D evo and use the remaining as the validation set D val.
Experiments use GPT-5-nano, DeepSeek-v4-flash, and Qwen3.5 under standardized inference settings.
Additional comparisons use Qwen3-8B
for parameter-based methods.
Comparisons include:
-
Standard prompting: ICL, Self-Consistency (SC), Tree-of-Thought (ToT)
-
Experience/memory-based: Experience RAG, ExpeL
-
Skill-learning: Direct Skill Generation, SkillOpt
-
Prompt optimization: MIPROv2, GEPA
-
Parameter-based: SFT, GRPO
Table 1 shows DIVE consistently achieves strong performance without parameter updates
across all model families and tasks. For example, with GPT-5-nano, DIVE (M=10) achieves 82.9% on HMMT vs. 64.3% for zero-shot, and 96.0% on Sudoku vs. 71.7% for zero-shot. The average across all six tasks improves from 52.3% (zero-shot) to 81.5% (DIVE M=10).
Figure 1 shows DIVE improves rapidly with additional experience, while SFT and GEPA plateau at substantially lower performance and GRPO improves more gradually despite requiring considerably more rollouts.
This demonstrates a favorable performance–rollout trade-off.
Table 2 shows GPT-5-nano with DIVE achieves higher average performance than GPT-5 with ICL, while reducing inference cost by 42.5%.
This demonstrates that skill evolution can substantially enhance the reasoning performance of a smaller model, enabling it to rival a stronger model under conventional prompting without parameter updates and at lower inference cost.
Table 3 shows skills developed for Qwen3.5-9B transfer effectively to both Qwen3.5-27B, a larger model in the same family, and DeepSeek-v4-flash, a model from a different family.
For example, Qwen3.5-27B improves from 59.6% average (zero-shot) to 81.2% using skills from Qwen3.5-9B, and to 86.1% using skills evolved for itself.
-
Skill set size:
Increasing the number of selected skills improves performance with diminishing returns, highlighting the benefit of complementary skill hypotheses.
-
Evolution operators:
Heterogeneous operators outperform a single operator, UCB improves over uniform allocation, and adaptive operator generation yields further gains.
-
Skill selection: Table 4 shows
jointly selecting complementary skills achieves the strongest overall performance, outperforming random or individually ranked skill selection.
The paper reports tuning experiments showing:
-
Evolution budget B:
Performance generally improves as the evolution budget increases
with diminishing returns. -
UCB coefficient β:
Performance initially improves as β increases, reaching its highest level, and then declines
- the broad peak indicates DIVEbenefits from balancing operator exploitation and exploration.
-
Parent-sampling temperature τ p:
Performance improves as τ p increases from zero, reaches its highest level at a moderate temperature, and then gradually declines.
The paper's main contributions are:
-
"We formulate self-improvement of frozen LLMs as the evolution of persistent natural-language skills, enabling models to accumulate reusable task knowledge from experience and verifier feedback without weight updates or a stronger teacher model."
-
"We introduce DIVE, a diversity-driven framework that independently evolves multiple skill populations from bootstrapped experience and adaptively allocates the evolution budget across heterogeneous transformation operators, preserving diverse skill hypotheses and evolution trajectories."
-
"We jointly select a complementary set of skills, translating diversity in the evolution process into improved heldout performance. Experiments across diverse reasoning tasks and model families further demonstrate effective self-improvement and cross-model skill transfer."
The paper concludes that "DIVE consistently improves performance, achieves a favorable performance–rollout trade-off, and produces skills that transfer across model scales and families. These results highlight diversity-driven skill evolution as a promising approach for enabling frozen language models to improve from experience. Future work may explore
cross-task skill transfer and integration with long-term memory mechanisms to support more scalable and continual self-improvement."
Improvements for AI systems
Improvements to AI Systems Based on DIVE:
-
Persistent Skill Memory Without Weight Updates: AI systems can now accumulate and refine reusable, natural-language
skills
(reasoning procedures, verification strategies, failure modes, output constraints) from task experience and verifier feedback, without requiring fine-tuning or access to model weights. This enables continual improvement for API-only or frozen models. -
Diversity-Preserving Self-Improvement: Instead of greedy single-trajectory refinement, the system maintains multiple independent skill populations, each evolving from different bootstrapped experience subsets. This prevents premature convergence, avoids overfitting to local failures, and preserves complementary solution strategies.
-
Adaptive Operator Portfolio: The system dynamically selects among heterogeneous evolution operators (reflective repair, exploratory revision, compression, recombination) using an upper-confidence-bound allocation. It also generates new operators targeting recurring failure patterns, ensuring the evolution mechanism adapts to the current skill landscape.
-
Complementary Skill Selection: At inference, the system jointly selects a compact set of skills that maximize marginal validation utility, ensuring diverse skills cover different error regions. This yields higher ensemble performance than selecting individually top-ranked skills.
-
Efficient Inference-Time Ranking: The system generates candidate responses from multiple skills and ranks them via a single listwise call to the frozen model (or majority voting for answer-extraction tasks), improving accuracy without additional training.
-
Cross-Model Skill Transfer: Skills evolved on one model can be directly applied to larger or different-family models, enabling knowledge transfer and boosting performance of stronger models without re-evolution.
What the Improved AI System Can Do:
-
Self-improve from experience on mathematical and logical reasoning tasks (e.g., HMMT, Sudoku, Cryptarithm, Calcudoku, Futoshiki) with no parameter updates, achieving up to +29% accuracy over zero-shot baselines.
-
Outperform larger models: A small model (GPT-5-nano) with DIVE surpasses a much larger model (GPT-5) with standard prompting, while reducing inference cost by 42.5%.
-
Adapt efficiently: Achieve better performance–rollout trade-offs than SFT, GRPO, and prompt optimization methods, improving rapidly with limited experience.
-
Transfer skills across models: Evolve skills on a smaller model and deploy them on larger or cross-family models, improving their accuracy by up to +21.6% without additional evolution.
-
Maintain robustness: Avoid brittle, overfit skills by preserving diverse hypotheses and adaptively allocating evolution effort, leading to consistent gains across tasks and model families.
Abstract
Large language models (LLMs) cannot retain post-deployment experience without parameter updates. We introduce DIVE, a diversity-driven framework that enables frozen LLMs to improve by evolving persistent natural-language skills from task experience and verifier feedback. These skills encode reusable reasoning procedures, verification strategies, common failure modes, and output constraints and are both executed and revised by the same underlying model without access to a teacher model. Since natural-language skill evolution is a stochastic, non-convex search process, optimizing a single skill trajectory can overfit to sampled experience or converge to a suboptimal solution. DIVE mitigates this optimization variance by independently evolving multiple skill populations from bootstrapped experience, adaptively refining them through diverse transformations, and jointly selecting a complementary set of skills. Across six mathematical and logical reasoning tasks and multiple model families, DIVE consistently outperforms existing reasoning methods, prompt-optimization approaches, skill-development frameworks, and memory-based baselines. It achieves rapid self-improvement from accumulated experience, obtaining substantially larger performance gains with fewer rollouts than parameter-based methods such as SFT and GRPO, and prompt optimization with GEPA. Further, the resulting skills transfer across model scales and families, enabling smaller models such as GPT-5-nano to match or outperform larger counterparts, i.e., GPT-5, under conventional prompting. These results establish diversity-driven skill evolution as an effective, interpretable, and parameter-free approach to LLM self-improvement.
Sources
- PromptWizard: Task-Aware Prompt Optimization Framework
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
- Conflict-Aware Fusion: Mitigating Logic Inertia in Large Language Models via Structured Cognitive Priors
- RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization
- The Equational Theories Project: Advancing Collaborative Mathematical Research at Scale
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- Natural Language Reinforcement Learning
- Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution
- Reinforced Self-Training (ReST) for Language Modeling
- PathWise: Planning through World Model for Automated Heuristic Design via Self-Evolving LLMs
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- LLMs Are In-Context Bandit Reinforcement Learners
- ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- HybridFlow: A Flexible and Efficient RLHF Framework
- Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models
- OpenAI GPT-5 System Card
- Qwen3.5-Omni Technical Report
- Inducing Programmatic Skills for Agentic Tasks
- Agent Workflow Memory
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering