TeacherGRPO: Closing the Capacity Gap in Reasoning Distillation via Teacher Alignment
cs.CL
Submitted: 2026-09-27
Updated: 2026-09-27
Code: https://github.com/LzyFischer/TeacherGRPO
Terminology
Sources
- Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- MiCoTA: Bridging the Learnability Gap with Intermediate CoT and Teacher Assistants
- On Designing Effective RL Reward at Training Time for LLM Reasoning
- Measuring Mathematical Problem Solving With the MATH Dataset
- Distilling the Knowledge in a Neural Network
- SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
- DRP: Distilled Reasoning Pruning with Skill-aware Step Decomposition for Efficient Large Reasoning Models
- The Impact of Reasoning Step Length on Large Language Models
- The Signal is in the Steps: Local Scoring for Reasoning Data Selection
- DistiLLM: Towards Streamlined Distillation for Large Language Models
- Training Language Models to Self-Correct via Reinforcement Learning
- Less is More: Selective Reflection for Compatible and Efficient Knowledge Distillation in Large Language Models
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Gemma 3 Technical Report
- LightReasoner: Can Small Language Models Teach Large Language Models Reasoning?
- Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort
- Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
- A Survey on Knowledge Distillation of Large Language Models
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering