SPEAR: Distilling Domain-Adaptive Reasoning Skeletons via Sequential Symbolic Alignment in Reinforcement Learning
cs.CL
Submitted: 2026-08-27
Updated: 2026-09-01
Code: https://github.com/zhuochunli/SPEAR
License: http://creativecommons.org/licenses/by/4.0/
The gist: Reinforcement learning-based knowledge distillation has the potential to transfer complex reasoning from teacher to student models, yet it currently faces a critical dilemma: researchers must choose
Terminology
Abstract
Reinforcement learning-based knowledge distillation has the potential to transfer complex reasoning from teacher to student models, yet it currently faces a critical dilemma: researchers must choose between sparse outcome-based rewards, which provide insufficient logical guidance, or expensive neural Process Reward Models (PRMs) for dense signals. We resolve this by introducing SPEAR (Symbolic Process Evaluation and Alignment Reward), a training-free and plug-and-play process reward method for sequence-level on-policy distillation. SPEAR projects natural-language reasoning traces into domain-adaptive symbolic milestones, providing an efficient proxy for process-level reasoning alignment. By utilizing the longest common subsequence (LCS) to align student explorations with teacher milestones, SPEAR provides a dense, order-aware reward signal that enforces logical consistency without the need for an external neural verifier. Our experiments across math, science, and commonsense reasoning tasks demonstrate that SPEAR effectively bridges the reasoning gap between student and teacher models via sequence-level distillation with efficient dense process rewards. Our code and data are available at: https://github.com/zhuochunli/SPEAR.
Sources
- The Llama 3 Herd of Models
- The False Promise of Imitating Proprietary LLMs
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Measuring Mathematical Problem Solving With the MATH Dataset
- Deep Learning for Symbolic Mathematics
- Process Reward Model with Q-Value Rankings
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- Training Verifiers to Solve Math Word Problems
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- Enhancing Large Language Models through Structured Reasoning
- Understanding R1-Zero-Like Training: A Critical Perspective
- ReFT: Reasoning with Reinforced Fine-Tuning
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning
- A Survey of On-Policy Distillation for Large Language Models
- Reasoning Scaffolding: Distilling the Flow of Thought from LLMs
- MiMo-V2-Flash Technical Report
- Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning
- KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
- RLKD: Distilling LLMs' Reasoning via Reinforcement Learning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering