Geometric Self-Distillation for Reasoning Generalization

arXiv:2607.06855 · cs.LG, cs.CL · Submitted 2026-07-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Geometric Self-Distillation for Reasoning Generalization".

Jane: On-policy distillation provides dense teacher supervision for large language models by having a model supervise its own generated trajectories, but this privileged context often introduces mismatches that degrade out-of-distribution reasoning.

Tom: First, who's behind it and why it matters.

Paper summary: Jane: So to wrap up on Geometric Self-Distillation for Reasoning Generalization, the authors are proposing a specific geometric objective to manage the drift that happens when an AI model tries to learn from its own generated mistakes.

Tom: They introduce GEOSD, which uses a Hellinger loss scaled by overlap and a Fisher–Rao proximal term to control both how much teachers influence us and how far we drift from our starting point.

Lu: The main implication is that this approach maintains the gains from on-policy distillation while substantially improving out-of-distribution reasoning, with empirical accuracy improvements ranging from five point seven to eight point six points depending on the model size.

Meng: For someone looking at practical application, it means we can build models that are highly proficient in their intended tasks but also more robust when faced with novel situations outside the training data scope.

Jane: It’s about pairing strong performance within the known domain with a much better ability to reason when things get new, and they did this through geometric principles of token distributions on a hypersphere.

Tom: The authors are essentially showing that how we measure and correct movement in prediction space matters more than just the raw numbers we use for distillation loss.

Jane: That's the focus—it’s about making sure the self-supervision process leads to better generalization, not just better imitation of the teacher.

Conclusion: Tom: So we’ve seen how this geometric distillation objective works—it basically tells the AI to learn from its own mistakes but in a way that keeps it grounded in what it already knows, and how far you drift from your training data is tracked geometrically on a sphere.

Jane: Exactly. The title, "Geometric Self-Distillation for Reasoning Generalization," points right to that idea of using geometry—think distances and shapes—to guide the learning process instead of just looking at raw probabilities.

Lu: I think what’s really wild here is that they aren't just doing standard distillation; they're measuring movement in this single, unified space of next-token distributions, which makes the updates much more principled than just tweaking parameters randomly.

Meng: From an engineering standpoint, it’s interesting because it handles those low-overlap situations differently than other methods. It’s not just about getting a higher score on the test; it's about controlling *how* the model gets there.

Lalam: If we think about culture, this means our models could become much more reliable when they encounter things that aren't exactly what they were trained on, which could make AI assistants feel much more trustworthy in complex situations.

Tom: Right, so it’s moving past just making the model look good on the training set to actually making it work better out there in the real world.

Jane: And these authors have shown that this specific geometric setup successfully keeps those initial performance gains while significantly boosting how well the AI reasons when things get new.

Lu: It really shows that we can use deep mathematical concepts, like information geometry, to solve very practical problems in training large models.

Tom: So, the big picture here is taking a technique that’s already working and adding this geometric layer to make it much more robust for real-world reasoning tasks.

Jane: And as we look ahead, we need to figure out if this approach can be easily applied across different types of AI problems beyond just mathematical reasoning.

Josip Jukic´, Ivan Titov

University of Amsterdam · University of Edinburgh

cs.LG, cs.CL

Submitted: 2026-07-07

Updated: 2026-10-05

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 91/100

The gist: On-policy distillation provides dense teacher supervision for large language models by having a model supervise its own generated trajectories, but this privileged context often introduces mismatches

Key concepts

Geometric Self-Distillation
This is a new objective that uses geometry to measure how much a student model should follow its own generated examples (teachers). It measures distances in the 'space of next-token distributions' rather than just raw probabilities, allowing for more meaningful guidance.
Hellinger Loss
This loss function is used instead of standard ones because it scales each teacher's preference by how much the student already agrees with it. This attenuates the teacher's pull on tokens the student cannot yet support, effectively weighting preferences based on overlap.
Drift Control Mechanism
This part penalizes updates that move too far from a recent stable checkpoint. It uses a proximal term based on Fisher–Rao distance to ensure that small, individually good updates do not accumulate into large, incorrect movements over training.

Terminology

Summary

On-policy distillation provides dense teacher supervision for large language models by having a model supervise its own generated trajectories, but this privileged context often introduces mismatches that degrade out-of-distribution reasoning. This paper introduces GEOSD, a geometric self-distillation objective that counters the resulting predictive drift in the space of next-token distributions by attenuating teacher pulls at low-overlap states and regulating how far accumulated updates carry the student over training.

The gist: GEOSD is a geometric self-distillation objective that weights each teacher preference by teacher–student overlap and penalizes predictive drift from a recent checkpoint, optimizing both as distances in a single geometry via natural gradient<ref:2607.06855#pg14>.

How it works

GEOSD counters the drift caused by privileged context mismatch in two complementary ways: attenuating the teacher’s pull at low-overlap states and regulating how far those pulls carry the student over training<ref:2607.06855#pg3>. First, it replaces standard distillation with a Hellinger loss, which scales each teacher preference by the overlap the student already shares with it, attenuating the pull on tokens the student cannot yet support (Page 3). This attenuation is built into the gradient rather than applied externally by clipping or filtering<ref:2607.06855#pg5>.

Second, it penalizes accumulated movement from a recent checkpoint with a proximal term to a recent checkpoint, measured as a Fisher–Rao distance<ref:2607.06855#pg14>. This addresses the problem that many small, individually reasonable pulls can still carry the student far over training (Page 3). Both terms share a single geometry of next-token distributions, allowing for a natural-gradient update that adjust[s] each step by its effect on the student’s predictions rather than by its size in parameter space (Page 3).

The Geometry of Distillation

The paper establishes that the movement is poorly captured in raw probabilities or parameters, so it is measured in a single geometry of next-token distributions<ref:2607.06855#pg4>. The student and teacher next-token distributions live on a hypersphere where statistical distance becomes ordinary distance when parameterized by the square root embedding p 7→√p<ref:2607.06855#pg15>. This sphere carries two canonical distances: the chord, which yields the squared Hellinger divergence, and the arc, which is the Fisher–Rao distance<ref:2607.06855#pg15>.

Key Distinctions in Loss Functions

The choice of divergence is a statement about behavior away from agreement<ref:2607.06855#pg14>. The Hellinger loss is selected because it sits exactly halfway between the two KL directions and is the self-dual member of the family, satisfying Dα(p∥q) = D−α(q∥p), with α = 0 being Hellinger, D0(p∥q) = 4 DH(p, q)<ref:2607.06855#pg18>. The gradient of the Hellinger term is weighted by the geometric mean overlap p pθ(i) qi, which vanishes as pθ(i)→0 regardless of teacher confidence (Page 4).

Drift Control Mechanism

The proximal term penalizes displacement from a recent checkpoint θckpt, and a second-order expansion shows that the loss is locally approximated by the quadratic form: Lprox(θk; θckpt) ≈ ∆⊤ k Fckpt∆k<ref:2607.06855#pg19>. The natural-gradient preconditioner then yields the approximate checkpoint pullback: −ηλ F −1 k ∇θLprox(θk; θckpt) ≈ −2ηλ F −1 k Fckpt∆k<ref:2607.06855#pg20>. This ensures that drift is measured in the checkpoint’s local geometry and returned as a restoring force in the current one (Page 4).

Empirical Results

GEOSD preserves the in-distribution (ID) gains of standard OPSD while substantially improving out-of-distribution (OOD) reasoning, with average OOD accuracy by 5.7–8.6 points over the base model across three model families and scales from 1.7B to 32B<ref:2607.06855#pg4>. GEOSD is the only method that both retains the ID gains of self-distillation and improves OOD reasoning substantially, pairing strong ID performance with the highest OOD accuracy (avg@16 and pass@16) (Page 3).

Analysis of Failure Modes

Standard matching fails out of distribution because it wins agreement with the teacher by draining mass from alternatives at high-entropy states, resulting in confident agreement on wrong answers (Page 4). GEOSD avoids this collapse because its overlap weighting keeps those alternatives in reach (Page 3), which results in a lower false-consensus rate and higher majority-answer accuracy under strong consensus<ref:2607.06855#pg10>.

Limitations

The domain scope is limited to mathematical reasoning, and the study does not qualitatively evaluate other privileged signals like hints or feedback<ref:2607.06855#pg12>. Furthermore, the objective regulates the magnitude of predictive movement rather than its direction or merit<ref:2607.06855#pg3>.

REFERENCES

Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, and Olivier Bachem. On-policy distillation of language models: Learning from selfgenerated mistakes. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=3zKtaqxLhW<ref:2607.06855#pg12>

Shun-ichi Amari. Natural gradient works efficiently in learning. Neural computation, 10(2):251–276, 1998. URL https://neuralcomputation.org/articles/conf neural computation vol30 no2 p251<ref:2607.06855#pg15>

Shun-ichi Amari and Hiroshi Nagaoka. Methods of information geometry, volume 191. American Mathematical Soc., 2000. URL https://amsbook.org/books/methods-of-information-geometry<ref:2607.06855#pg15>

Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang. MiniLLM: Knowledge distillation of large language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=5h0qf7IBZZ<ref:2607.06855#pg12>

Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, et al. Deepseek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning. Nature, 645(8081):633–638, 2025. URL http://dx.doi.org/10.1038/s41586-025-09422-z<ref:2607.06855#pg12>

Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015. URL https://arxiv.org/abs/1503.02531<ref:2607.06855#pg5>

Jonas Hubotter, Frederike L ¨ ubeck, Lejs Behric, Anton Baumann, Marco Bagatella, Daniel Marta, ¨Ido Hakimi, Idan Shenfeld, Thomas Kleine Buening, Carlos Guestrin, and Andreas Krause. Reinforcement learning via self-distillation. arXiv preprint arXiv:2601.20802, 2026. URL https://arxiv.org/abs/2601.20802<ref:2607.06855#pg12>

Woogyeol Jin, Taywon Min, Yongjin Yang, Dennis Wei, Yi Zhou, Swanand Ravindra Kadhe, Nathalie Baracaldo, and Kimin Lee. Entropy-aware on-policy distillation of language models. In Forty-third International Conference on Machine Learning, 2026. URL https://openreview.net/forum?id=J5i09faOOf<ref:2607.06855#pg12>

Jeonghye Kim, Xufang Luo, Minbeom Kim, Sangmook Lee, Dohyung Kim, Jiwon Jeon, Dongsheng Li, and Yuqing Yang. Why does self-distillation (sometimes) degrade the reasoning capability of LLMs?, 2026. URL https://arxiv.

Improvements for AI systems

  1. Bold header: Hellinger Loss for Overlap-Aware Pulls

This objective attenuates each teacher preference by the overlap because tokens the teacher assigns high probability but the student finds unlikely exert little pull, while shared preferences transfer strongly, preventing strong gradients from tokens where support is weak.

  1. Bold header: Fisher-Rao Proximal Term for Drift Control

The proximal term penalizes how far the student’s predictions drift from a recent checkpoint, measured as a Fisher–Rao distance, ensuring that many small, individually reasonable pulls can still carry the student far over training is controlled.

  1. Bold header: Natural Gradient Update for Geometry-Aware Learning

The objective is optimized via a natural-gradient update that adjusts each step by its effect on the student’s predictions rather than by its raw size in parameters, meaning each step makes the smallest change to the student’s predictions that improves the objective within a single geometry.

  1. Bold header: Improved Out-of-Distribution (OOD) Reasoning

The system can achieve improving average OOD accuracy by 5.7–8.6 points over the base model across various mathematical reasoning benchmarks, specifically by countering the failure mode where standard matching wins agreement with the teacher by draining mass from alternatives at high-entropy states.

  1. Bold header: Robustness to Privileged Context Variation

The system demonstrates that GEOSD improves monotonically across the same range of revealed privileged information, as it prevents divergence baselines from copying teacher confidence and degrading performance when more solution detail is provided.

Abstract

On-policy distillation provides dense teacher supervision on a language model's own trajectories. In self-distillation with privileged context, this supervision comes from the model itself, conditioned on a hint or solution trace hidden from the student. When the teacher's preferences hinge on privileged information, it can assign higher probability to continuations the student cannot infer from its own context. Matching these preferences throughout training can induce predictive drift and degrade out-of-distribution (OOD) reasoning. We propose GeoSD, a self-distillation method that controls this drift through two complementary geometric terms. A Hellinger loss weights each teacher preference by the student--teacher overlap, reducing the influence of tokens to which the student assigns low probability. Because these influences can still accumulate, a Fisher--Rao penalty regulates predictive distance from a copy of the student refreshed periodically during training. Both terms compare next-token distributions in Fisher--Rao geometry and are jointly optimized with a preconditioner motivated by the natural gradient. Across three model families, GeoSD retains strong in-distribution gains while improving average mathematical OOD accuracy by 5.7--8.6 points over the base model. OOD gains hold across five model scales from 1.7B to 32B and transfer to code generation, where GeoSD improves code accuracy by 1.9 points on average despite distilling on mathematics alone. Our analysis of mathematical reasoning shows that standard matching rapidly concentrates probability mass at high-entropy states and that its samples confidently agree on incorrect answers. In contrast, GeoSD preserves alternative token mass and reduces false consensus.

Sources

Related papers