Lightning Weave: Improving the Accuracy-Efficiency Frontier of Reasoning Models through Capability Composition
cs.AI
Submitted: 2026-09-13
Updated: 2026-09-16
Comments: Work in progress
Code: https://github.com/jet-ai-projects/Lightning-Weave
License: http://creativecommons.org/licenses/by/4.0/
The gist: A core goal of efficient reasoning is to improve the accuracy-efficiency frontier.
Terminology
Abstract
A core goal of efficient reasoning is to improve the accuracy-efficiency frontier. However, jointly improving reasoning accuracy and inference efficiency can be challenging, as the two objectives can favor different reasoning behaviors. Independently post-trained models already offer distinct strengths in accuracy and efficiency. We introduce Lightning Weave, a post-training framework that extracts and composes these independently learned capabilities in a single student through on-policy distillation. Each acquired capability is represented by the policy shift from the model before post-training to the resulting specialist. Lightning Weave combines aligned log-ratio shifts at shared student token states and uses Tilted-Target DOPD to convert the cached signals into a stable learning target. Each anchor pair scores the cached trajectories once, enabling subsequent student training without serving multiple live anchor models concurrently. Across diverse student models and benchmarks in mathematics and code, Lightning Weave substantially improves upon the base students and achieves a state-of-the-art accuracy-efficiency frontier. On Qwen3.5-4B, it raises HMMT 2025 accuracy from 59.2% to 64.0% with 10.7% fewer response tokens, and LiveCodeBench v5 accuracy from 41.7% to 54.2% with 9.6% fewer response tokens. Adjusting the relative strengths of the anchor signals yields a strong empirical accuracy-efficiency Pareto frontier. These results establish Lightning Weave as a new practical route to efficient reasoning through capability composition. Code will be released soon.
Sources
- Counteraction-Aware Multi-Teacher On-Policy Distillation for General Capability Recovery with Domain Preservation
- Weak-to-Strong Generalization via Direct On-Policy Distillation
- Deep Think with Confidence
- Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
- Skywork Open Reasoner 1 Technical Report
- Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models
- ORBIT: On-policy Exploration-Exploitation for Controllable Multi-Budget Reasoning
- DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning
- MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining
- MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training
- Contrastive On-Policy Distillation
- A Survey of On-Policy Distillation for Large Language Models
- Olmo 3
- Cross-Tokenizer On-Policy Distillation via Byte-Prefix Marginalization
- Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation
- Unlocking Efficient Long-to-Short LLM Reasoning with Model Merging
- The Art of Efficient Reasoning: Data, Reward, and Optimization
- Revisiting Model Interpolation for Efficient Reasoning
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection