BRIDGE: Bridging Reasoning In Distillation Gap Elimination via Structure-Aware Masking

arXiv:2602.17686 · cs.LG, cs.AI · Submitted 2026-08-20 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "BRIDGE: Bridging Reasoning In Distillation Gap Elimination via Structure-Aware Masking".

Jane: The paper was written by Bowen Yu, Sheng Zhang, Binhao Wang, Yi Wen, Jingtong Gao et al. from City University of Hong Kong and Mohamed bin Zayed University of Artificial Intelligence.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: We're looking at "BRIDGE: Bridging Reasoning In Distillation Gap Elimination via Structure-Aware Masking" today.

Jane: That title is quite a mouthful, Tom, but the authors from City University of Hong Kong and MBZUAI are tackling a massive headache in AI.

Tom: They're focusing on that frustrating gap that happens when you try to teach a small model using a massive, smart teacher.

Jane: It's like trying to teach a toddler a college lecture; the kid just can't process that much information at once.

Lu: I think it's brilliant because they aren't just giving the student more data, they're actually changing how the student perceives the logic.

Meng: I've seen this happen in our own training runs where the small models just start repeating themselves because they're overwhelmed.

Lu: Exactly, Meng, and that's why this curriculum approach is so much more creative than just standard fine-tuning.

Meng: I wonder if this actually solves the memory issues we see when we try to cram long reasoning chains into a tiny parameter count.

Lalam: If they succeed, it means we can put high-level reasoning into much smaller devices, which changes how people interact with technology in their daily lives.

Jane: That would make intelligence much more accessible to everyone, wouldn't it?

Lalam: It definitely would, as it moves us toward a future where smart assistance isn't locked behind massive, expensive servers.

Tom: It sounds like they've found a way to make the learning process more digestible.

Jane: Let's look at how they actually structured that learning process.

Summary: Tom: So, Jane, how do they actually build this "bridge" between the big teacher and the small student?

Jane: They use a three-stage curriculum that starts with a "warmup" phase where the student learns to reconstruct scrambled reasoning steps.

Tom: So they aren't just handing over the answers, they're making the student piece the puzzle together?

Jane: Yes, they shuffle the steps and mask some out so the model has to understand the logical skeleton before it even tries to generate its own answers.

Lu: That's such a clever way to force the model to learn the underlying structure rather than just memorizing words.

Meng: I'm curious about the second stage, though, because how do they stop the model from just being too brief and losing the logic?

Jane: That's where they use GRPO, which is a reinforcement learning method that rewards the model for being both correct and concise.

Meng: Using a reward that balances accuracy against brevity sounds like a nightmare to stabilize in training.

Jane: It can be, which is why they use a hierarchical reward that makes sure the model is correct before it ever gets a bonus for being short.

Lu: And then the third stage is the real magic, where they use the teacher to help the student with the hardest problems.

Lalam: It reminds me of a student reading a summary of a difficult book to help them grasp the main points before writing their own essay.

Jane: That's a perfect analogy, Lalam, because the student uses the teacher's solution as a scaffold to internalize the logic.

Tom: It's a very progressive way to build up competence.

Jane: Now let's see if those theoretical stages actually lead to better performance in the real world.

Improvements: Tom: The results for the Qwen2 point 5-3B model seem to back up the whole BRIDGE framework.

Jane: They saw an eleven point two nine percent accuracy improvement on the GSM8K math benchmark.

Tom: And they didn't just get smarter, they got much faster too, right?

Jane: They actually reduced the output length by twenty-seven point four percent compared to the original model.

Lu: I was particularly impressed that it worked on SVAMP and MATH-five hundred even though they didn't train on those specific datasets.

Meng: That kind of zero-shot generalization is what we actually need for production-ready models.

Lu: It shows the model is learning actual reasoning patterns instead of just memorizing math templates.

Meng: I did notice in the error analysis that they still have some issues, like the model occasionally skipping important conditions in a problem.

Jane: They found that condition omission was actually the biggest error type, happening about forty-five percent of the time in their failure cases.

Meng: That makes sense, because if you're pushing a model to be as brief as possible, it might try to cut out details it thinks are unnecessary.

Lalam: Even with those errors, the fact that it's producing much tighter, more efficient reasoning is a huge step forward for efficient AI.

Tom: It's a massive leap from the models that just fall into endless repetition loops when they get confused.

Jane: Let's wrap this all up and talk about what this means for the field.

Conclusion: Tom: We've covered a lot of ground with "BRIDGE: Bridging Reasoning In Distillation Gap Elimination via Structure-Aware Masking."

Jane: It really shows that how we teach is just as important as what we teach.

Tom: Lu, you've been looking at the big picture, what's your final thought?

Lu: I see this as a blueprint for creating specialized, tiny models that can perform tasks we once thought required massive supercomputers.

Meng: From my side, the practical takeaway is that we can finally start getting reliable reasoning out of these smaller, more efficient architectures.

Lalam: I believe this will lead to a more seamless integration of intelligence into our culture, making it a quiet, efficient background tool for everyone.

Tom: Thanks to the whole team for joining us.

Jane: We'll see you next time for another look at the latest research.

Tom: Goodbye for now!

Bowen Yu, Sheng Zhang, Binhao Wang, Yi Wen, Jingtong Gao, Bowen Liu, Zimo Zhao, Shanshan Ye, Wanyu Wang, Maolin Wang, Xiangyu Zhao

City University of Hong Kong · Mohamed bin Zayed University of Artificial Intelligence

cs.LG, cs.AI

Submitted: 2026-08-20

Updated: 2026-08-21

Code: https://github.com/Applied-Machine-Learning-Lab/SDM2026_BRIDGE

Importance score: 76/100

The gist: The paper addresses the "capacity mismatch between teacher and student" in Chain-of-Thought (CoT) distillation, noting that "when compact students (e.g., 3B models) attempt to reproduce these lengthy

Key concepts

Distillation Gap
This refers to the performance gap that occurs when attempting to teach a small AI model (the student) using knowledge from a much larger, more capable model (the teacher). The framework aims to bridge this frustrating gap.
Structure-Aware Masking
A method used in the initial learning phase where steps in reasoning are shuffled and masked. This forces the small model to understand the underlying logical structure of a problem rather than just memorizing answers.
GRPO
A reinforcement learning method used in the second stage of training. It rewards the model for producing answers that are both accurate and concise, helping to balance correctness against brevity.
Curriculum Approach
The overall three-stage learning methodology employed by BRIDGE. It progressively builds competence by starting with simple steps (warmup), moving to balanced performance (GRPO), and finishing with complex scaffolding.

Terminology

Summary

The paper addresses the capacity mismatch between teacher and student in Chain-of-Thought (CoT) distillation, noting that when compact students (e.g., 3B models) attempt to reproduce these lengthy sequences via standard supervised fine-tuning, they lack the representational bandwidth to process or memorize such content effectively. This mismatch manifests as truncated outputs, repetition loops, or superficial mimicry without genuine understanding. Existing remedies are criticized because implicit reasoning methods... trade away interpretability and verifiability, while heuristic compression strategies... destroy logical integrity.

To address this, the authors propose "BRIDGE, a curriculum framework that first establishes structural understanding via masked reconstruction, then uses GRPO-based reinforcement learning to guide students in self-discovering the optimal balance between accuracy and brevity, and finally internalizes complex reasoning through teacher-guided rewriting on failure cases."

The BRIDGE framework is composed of three stages:

Stage 1: Structure-Aware Warmup

This stage establishes a structural foundation through masked reconstruction, training the student to recognize logical dependencies. To prevent the student from exploiting positional shortcuts, the authors implement Step Shuffling, which permutes the order of steps to force the student to recognize causal dependencies between steps. Additionally, Step Masking is applied to approximately 15% of the reasoning steps to force the student to infer the missing intermediate logic from the surrounding context, requiring genuine comprehension of the reasoning flow. The objective is a generative reconstruction task that forces the model to internalize the semantic topology of the reasoning chain.

Stage 2: GRPO-Based Compression

This stage introduces explicit optimization for compression while maintaining correctness. It employs Group Relative Policy Optimization (GRPO) to guide the student in self-discovering the optimal balance between accuracy and brevity. To prevent reward hacking, where a model might produce minimal outputs that fail to solve problems, the authors design a hierarchical reward mechanism: an incorrect output receives zero efficiency reward regardless of its brevity, while a correct output is further rewarded for being brief. The reward is formulated as R(r i) = R base(r i) + I[Correct(r i)] times R eff(r i).

Stage 3: Teacher-Guided Internalization

For difficult queries where the student struggles (identified as failure cases from Stage 2), the framework utilizes teacher-guided rewriting to progressively and effectively internalize complex reasoning deeply into concise form. The student is provided with the teacher’s complete solution as a scaffold and prompted to rewrite the reasoning concisely. This process allows the student to absorb the teacher’s logical structure while adapting it to its own capacity constraints. The reward for this stage, R(r i) = R base(r i) + I[Correct(r i)] times R comp(r i), explicitly encourages shorter-than-teacher outputs.

Experimental Results

The authors demonstrate that on GSM8K, BRIDGE enables Qwen2.5-3B to achieve 11.29% accuracy improvement and 27.4% token reduction over the original model, outperforming instruction-tuned variants and distillation baselines. Furthermore, zero-shot transfer experiments on SVAMP and MATH-500 further confirm the generalization of internalized reasoning, with BRIDGE achieving 83.33% (vs. 79.33% Base, +4.0%) on SVAMP and 38.20% (vs. 36.40% Base, +1.8%) on MATH-500.

Improvements for AI systems

Improvements to AI Systems

  1. Implementation of a Three-Stage Structural Curriculum for Reasoning Distillation:
  • Stage 1: Structure-Aware Warmup (SFT): Replace standard Supervised Fine-Tuning (SFT) on verbose teacher Chain-of-Thought (CoT) with a reconstruction task. This involves applying Step Shuffling (permuting the order of reasoning steps to eliminate positional shortcuts) and Stochastic Step Masking (masking about 15% of steps) to force the model to learn the underlying logical skeleton and causal dependencies rather than verbatim token memorization.

  • Stage 2: Hierarchical GRPO Compression (RL): Integrate Group Relative Policy Optimization (GRPO) using a gated hierarchical reward function. The reward must be structured as R = R base + I[Correct] times R eff, where the efficiency bonus (R eff) is strictly gated by a correctness indicator (I[Correct]). This prevents reward hacking where the model optimizes for brevity at the expense of accuracy.

  • Stage 3: Teacher-Guided Internalization (Hard-Case RL): For samples where the student fails in Stage 2, implement a scaffolded rewriting phase. Provide the teacher's full, verbose CoT as an input prompt and use GRPO to reward the student for producing a version that is both correct and significantly shorter than the teacher's original length (r i < r T).

  1. Optimization of Reward Functions for the Accuracy-Brevity Trade-off:
  • Transition from linear reward combinations to multiplicative gating mechanisms in reinforcement learning. This ensures the model only receives incentives for compression when the logical reasoning is verified as correct, aligning the optimization landscape with actual problem-solving capability.

Capabilities of the Improved AI System

  • High-Fidelity Reasoning in Compact Architectures: Enables small-scale models (e.g., 3B parameters) to achieve mathematical reasoning accuracy levels previously reserved for much larger models (e.g., 14B+), effectively overcoming the capacity mismatch bottleneck.

  • Pareto-Optimal CoT Generation: The system can generate reasoning chains that are simultaneously more accurate and significantly more concise (e.g., about 27% reduction in token count), reducing inference latency and computational overhead.

  • Elimination of Degenerate Output Patterns: The model will be resistant to common small-model failures during distillation, such as repetitive loops, truncated reasoning chains, and superficial mimicry of teacher verbosity.

  • Enhanced Zero-Shot Generalization: By internalizing logical structures rather than linguistic templates, the system can transfer reasoning strategies to unseen mathematical benchmarks (e.g., SVAMP, MATH-500) without additional fine-tuning.

Sources

Related papers