Geometric Self-Distillation for Reasoning Generalization
summary
The gist
On-policy distillation provides dense teacher supervision for large language models by having a model supervise its own generated trajectories, but this privileged context often introduces mismatches
In short
GEOSD addresses mismatches in AI radio's self-distillation by using a geometric approach. It modifies teacher guidance by weighting preferences based on token overlap and penalizing updates that drift too far from recent checkpoints. This method improves out-of-distribution reasoning while keeping the model's in-distribution performance high.
Key concepts
- Geometric Self-Distillation
- This is a new objective that uses geometry to measure how much a student model should follow its own generated examples (teachers). It measures distances in the 'space of next-token distributions' rather than just raw probabilities, allowing for more meaningful guidance.
- Hellinger Loss
- This loss function is used instead of standard ones because it scales each teacher's preference by how much the student already agrees with it. This attenuates the teacher's pull on tokens the student cannot yet support, effectively weighting preferences based on overlap.
- Drift Control Mechanism
- This part penalizes updates that move too far from a recent stable checkpoint. It uses a proximal term based on Fisher–Rao distance to ensure that small, individually good updates do not accumulate into large, incorrect movements over training.
Terminology used across episodes
This episode discusses
- Geometric Self-Distillation for Reasoning Generalization · Paper Radio
- Distilling the Knowledge in a Neural Network
- Reinforcement Learning via Self-Distillation
- Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- Self-Distilled Policy Gradient
- Olmo 3
- Privileged Information Distillation for Language Models
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- On the Geometry of On-Policy Distillation
- A Survey of On-Policy Distillation for Large Language Models
- GATES: Self-Distillation under Privileged Context with Consensus Gating
- Trust Region On-Policy Distillation
- TIP: Token Importance in On-Policy Distillation
- Qwen3 Technical Report
- Self-Distilled RLVR
- On-Policy Context Distillation for Language Models
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
- Fine-Tuning Language Models from Human Preferences
The paper
Geometric Self-Distillation for Reasoning Generalization · Read on arXiv
Josip Jukic´, Ivan Titov
University of Amsterdam · University of Edinburgh
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Geometric Self-Distillation for Reasoning Generalization".
Jane: On-policy distillation provides dense teacher supervision for large language models by having a model supervise its own generated trajectories, but this privileged context often introduces mismatches that degrade out-of-distribution reasoning.
Tom: First, who's behind it and why it matters.
Paper summary: Jane: So to wrap up on Geometric Self-Distillation for Reasoning Generalization, the authors are proposing a specific geometric objective to manage the drift that happens when an AI model tries to learn from its own generated mistakes.
Tom: They introduce GEOSD, which uses a Hellinger loss scaled by overlap and a Fisher–Rao proximal term to control both how much teachers influence us and how far we drift from our starting point.
Lu: The main implication is that this approach maintains the gains from on-policy distillation while substantially improving out-of-distribution reasoning, with empirical accuracy improvements ranging from five point seven to eight point six points depending on the model size.
Meng: For someone looking at practical application, it means we can build models that are highly proficient in their intended tasks but also more robust when faced with novel situations outside the training data scope.
Jane: It’s about pairing strong performance within the known domain with a much better ability to reason when things get new, and they did this through geometric principles of token distributions on a hypersphere.
Tom: The authors are essentially showing that how we measure and correct movement in prediction space matters more than just the raw numbers we use for distillation loss.
Jane: That's the focus—it’s about making sure the self-supervision process leads to better generalization, not just better imitation of the teacher.
Conclusion: Tom: So we’ve seen how this geometric distillation objective works—it basically tells the AI to learn from its own mistakes but in a way that keeps it grounded in what it already knows, and how far you drift from your training data is tracked geometrically on a sphere.
Jane: Exactly. The title, "Geometric Self-Distillation for Reasoning Generalization," points right to that idea of using geometry—think distances and shapes—to guide the learning process instead of just looking at raw probabilities.
Lu: I think what’s really wild here is that they aren't just doing standard distillation; they're measuring movement in this single, unified space of next-token distributions, which makes the updates much more principled than just tweaking parameters randomly.
Meng: From an engineering standpoint, it’s interesting because it handles those low-overlap situations differently than other methods. It’s not just about getting a higher score on the test; it's about controlling *how* the model gets there.
Lalam: If we think about culture, this means our models could become much more reliable when they encounter things that aren't exactly what they were trained on, which could make AI assistants feel much more trustworthy in complex situations.
Tom: Right, so it’s moving past just making the model look good on the training set to actually making it work better out there in the real world.
Jane: And these authors have shown that this specific geometric setup successfully keeps those initial performance gains while significantly boosting how well the AI reasons when things get new.
Lu: It really shows that we can use deep mathematical concepts, like information geometry, to solve very practical problems in training large models.
Tom: So, the big picture here is taking a technique that’s already working and adding this geometric layer to make it much more robust for real-world reasoning tasks.
Jane: And as we look ahead, we need to figure out if this approach can be easily applied across different types of AI problems beyond just mathematical reasoning.
More episodes
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck
- 2407.14562-Thought-Like-Pro: Enhancing Reasoning of Large Language Models through Self-Bootstrapped Prolog-based Chain-of-Thought