Critique-Guided Distillation for Robust Reasoning via Refinement
cs.CL, cs.LG
Submitted: 2025-05-16
Updated: 2026-05-19
Comments: Accepted to ICML 2026
Journal ref: Proceedings of the 43rd International Conference on Machine Learning, PMLR 270, 2026
Code: https://github.com/hkust-nlp/simpleRL-reason
License: http://creativecommons.org/licenses/by/4.0/
The gist: Supervised fine-tuning with expert demonstrations often produces models that imitate outputs without internalizing the reasoning processes needed for robust generalization.
Terminology
Abstract
Supervised fine-tuning with expert demonstrations often produces models that imitate outputs without internalizing the reasoning processes needed for robust generalization. While critique-based approaches show promise, training models to generate critiques directly, such as Critique Fine-Tuning (CFT), can lead to output-format drift and degradation of general capabilities. We propose Critique-Guided Distillation (CGD), a training framework that decouples critique consumption from critique generation. During fine-tuning, the student is trained to refine flawed responses conditioned on teacher critiques. CGD treats critiques as a training-time-only supervision signal, encouraging internalization of error-aware reasoning: critiques guide learning but are absent at inference. Controlled ablations confirm that these reasoning gains are directly driven by the specificity and relevance of the teacher's feedback. Across five model families, CGD consistently outperforms CFT and standard distillation on mathematical reasoning benchmarks, yielding 7% average improvements and gains of up to +15.0% on AMC23 and +12.2% on MATH-500. On challenging competition problems such as AIME24 and AIME25, CGD achieves substantially higher Pass@1 and stronger performance at low Pass@k, indicating improved reasoning quality per sample. Importantly, CGD preserves general instruction-following capabilities where CFT degrades significantly (- 21.3% on IFEval). These results position CGD as a practical and compute-efficient intermediate training paradigm for reasoning-centric tasks without introducing architectural inference-time overhead.
Sources
- On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
- RL4F: Generating Natural Language Feedback with Reinforcement Learning for Repairing Model Outputs
- SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Training Verifiers to Solve Math Word Problems
- Let Me Teach You: Pedagogical Foundations of Feedback for Language Models
- LLM Self-Correction with DeCRIM: Decompose, Critique, and Refine for Enhanced Following of Instructions with Multiple Constraints
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
- Evaluating Large Language Models Trained on Code
- Measuring Mathematical Problem Solving With the MATH Dataset
- Distilling the Knowledge in a Neural Network
- Training Language Models to Self-Correct via Reinforcement Learning
- Solving Quantitative Reasoning Problems with Language Models
- s1: Simple test-time scaling
- On the Robustness Tradeoff in Fine-Tuning
- TruthfulQA: Measuring How Models Mimic Human Falsehoods
- What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning
- Orca: Progressive Learning from Complex Explanation Traces of GPT-4
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- An Overview of Multi-Task Learning in Deep Neural Networks
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering