LoRi: Low-Rank Distillation for Implicit Reasoning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "LoRi: Low-Rank Distillation for Implicit Reasoning".
Tom: The gist:
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So we've covered the LoRi paper, which introduces this low rank distillation framework to make implicit reasoning more effective by aligning trajectories in a shared subspace.
Jane: The authors of "LoRi: Low-Rank Distillation for Implicit Reasoning" argue that this approach captures the global structure of reasoning while supporting a compact latent process through Lrationale and Lanchor objectives.
Lu: They emphasize that this geometric view suggests reasoning dynamics are governed by a low dimensional subspace, which can be exploited for both learning and inference.
Meng: The practical impact is that we see performance gains on benchmarks like GSM8K-Hard up to about ten percent, and training becomes cheaper because of the precomputation step <ref:2606.05315#pg2>.
Lalam: LoRi helps us move toward a system where reasoning is both efficient in terms of computation and powerful enough to compete with explicit chain-of-thought methods.
Tom: In short, this paper provides a principled way to compress long reasoning paths without losing the essential logic, showing that the core dynamics of reasoning are indeed low dimensional.
Conclusion: Tom: So we've looked at how LoRi takes that explicit chain of thought reasoning from models and boils it down into something much smaller, a latent process.
Jane: Yeah, the paper is called "LoRi: Low-Rank Distillation for Implicit Reasoning," and it’s basically about transferring that complex thinking from step-by-step instructions into a more compact form.
Lu: The authors are focusing on finding this low-rank structure in how hidden states move through layers and across those chain of thought tokens.
Meng: From an engineering standpoint, it means we can compress a lot of reasoning steps into something that runs much faster during inference, which is a big deal for deployment.
Lalam: For Lalam, this suggests that the core logic isn't actually needing every single intermediate step to be explicit; it’s all packed into this low-dimensional space.
Tom: The main implication here is geometric: it suggests that long reasoning paths can be summarized by a much shorter trajectory, which makes sense when you think about how models really process information.
Jane: It shifts the focus from just "more steps" to understanding the underlying structure that connects those steps together in a simpler way.
Lu: It gives us a new geometric lens through which to look at reasoning, suggesting that the essential dynamics of thought are governed by this low-dimensional subspace.
Meng: If we can reliably find and use this subspace, it could drastically reduce the computational cost for complex reasoning tasks without sacrificing accuracy on tough problems.
Lalam: And for Lalam, that means we can internalize much more complex reasoning capabilities into our core structure with less explicit training data focused solely on the steps.
Tom: It’s a really interesting idea because it moves us away from just brute-forcing longer prompts toward finding an intrinsic efficiency in how the AI learns to think.
Jane: And this paper lays out how they actually do that by using these two alignment objectives, Lrationale and Lanchor, to keep things stable while moving them into that shared low-rank space.
Lu: So it’s a framework for distillation that targets the global geometry of the teacher's reasoning path rather than just matching every single token one-to-one.
Meng: It sounds like they’re not trying to replicate the exact thought process, but rather distill its most important mathematical essence into something scalable.
Lalam: Exactly, it’s about finding that compressed latent representation where the crucial reasoning structure lives, which is much more robust than relying on specific token sequences.
Tom: So we've seen how they frame this as a distillation technique that leverages the inherent low-rank nature of these hidden state trajectories to create a more efficient implicit reasoning system.
Jane: And that leads us right into the next part where we look at the actual performance numbers and what those results mean for real-world tasks like solving hard math problems.
University of California-Santa Barbara
cs.CL, cs.AI
Submitted: 2026-06-03
Updated: 2026-10-07
Code: https://github.com/rmsolgi/lori
Importance score: 79/100
The gist: The gist: LoRi proposes a low-rank distillation framework that transfers reasoning from explicit chain-of-thought into compact implicit latent processes by aligning teacher and student trajectories
Key concepts
- Implicit Chain-of-Thought (iCoT)
- Methods that aim to internalize reasoning directly into a model's hidden states without requiring explicit step-by-step prompting. LoRi builds upon this by distilling the detailed reasoning from an explicit teacher model into a more compact, implicit latent process for better performance.
- Low-Rank Subspace Alignment
- The core technique where high-dimensional hidden state tensors are projected into a shared, low-dimensional space using techniques like Tucker decomposition and SVD. This allows the student's reasoning trajectory to be aligned with the teacher's global structure without needing exact token-level matching.
- Rationale-Level Alignment (Lrationale)
- A loss function that aligns the first and second-order statistics of projected hidden states between teacher and student models within a shared low-rank subspace. This objective transfers the overall geometric structure of the teacher's reasoning process, making it invariant to the specific length of the reasoning trajectory.
- Anchor-Level Alignment (Lanchor)
- A loss component that constrains the transition from latent reasoning to final answer generation at specific prediction tokens. It ensures that a student's representation at an answer token aligns with the dominant structure found in the teacher's hidden states at that exact location.
Terminology
Summary
The gist: LoRi proposes a low-rank distillation framework that transfers reasoning from explicit chain-of-thought into compact implicit latent processes by aligning teacher and student trajectories in a shared low-rank subspace using first- and second-order statistics.
Background and Motivation
Implicit chain-of-thought (iCoT) methods aim to internalize reasoning in large language models, but often underperform explicit CoT prompting The paper empirically finds that hiddenstate reasoning trajectories exhibit low-rank structure Motivated by this observation, the authors propose a low-rank distillation framework that aligns the global geometry of the teacher’s reasoning trajectory through low-rank statistical representations By stacking hidden states across layers and CoT tokens, they found that the normalized cumulative singular values grow rapidly with rank, indicating that reasoning trajectories are well-approximated by low-dimensional subspaces
The LoRi Method
LoRi is a distillation framework that transfers reasoning from explicit chain-of-thought into a compact implicit latent process The method combines two complementary objectives: rationale-level alignment, which preserves the global geometry of the teacher’s trajectory, and anchor-level alignment, which regularizes the transition from latent reasoning to answer generation The composite objective trained is L = LLR + λLCE, where LCE supervises the student’s explicit output sequence and LLR aligns the student’s latent reasoning states with the teacher’s hidden representations in a low-rank subspace
Low-Rank Rationale Alignment (Lrationale)
To transfer reasoning capability, LoRi aligns the student’s hidden states with the teacher’s reasoning trajectory in a low-rank subspace This is achieved by projecting high-dimensional tensors into a shared low-rank subspace using Tucker decomposition to extract low-rank factor matrices via SVD The rationale-level loss, Lrationale, aligns student and teacher representations by matching the first- and second-order statistics of their projections in the shared low-rank space This objective transfers the global structure of the teacher’s reasoning process without tokenlevel alignment and remains invariant to reasoning length
Anchor-Level Alignment (Lanchor)
While Lrationale captures global reasoning structure, anchor-level alignment constrains the transition from latent reasoning to answer generation at a fixed position corresponding to an answer prediction token The anchor-level loss, Lanchor, is defined as 1/L sum of (U(a)⊤H hl − U(a)⊤H h′l)2 across layers l=1 to L This formulation leverages a shared low-rank subspace learned across training samples, ensuring the student’s representation at the answer prediction point aligns with the dominant structure of the teacher’s hidden states at that location
Implications and Results
The low-rank structure provides a geometric explanation for why long chain-of-thought reasoning can be compressed into a shorter latent trajectory The formulation is invariant to the length of the reasoning trajectory because alignment is performed through aggregated statistics rather than step-wise correspondence LoRi consistently outperforms prior iCoT methods, improving accuracy by up to ∼12% and achieving strong gains on GSM8K-Hard [up to ∼10%, as shown in Fig. 1 (b)] across model scales Furthermore, LoRi-E achieves performance comparable to SFT-CoT across all models and benchmarks, suggesting the proposed distillation method preserves explicit reasoning capability while enabling both implicit and explicit reasoning at inference time The method also significantly reduces training cost by using a two-stage procedure where teacher-derived quantities are precomputed once
Ablation Studies
Ablation on latent reasoning steps shows that performance consistently improves from K = 2 to K = 5, suggesting that a minimum number of latent reasoning iterations is needed to capture the underlying reasoning dynamics Ablation on training sample size indicates that accuracy improves rapidly in the lowdata regime and saturates after roughly 128 samples across all model scales The results support the view that distillation only needs to recover the dominant reasoning structure rather than the full trajectory The overall largest benefits are observed in smaller models, while the effects become more nuanced at larger scales
Conclusion
LoRi outperforms prior implicit reasoning methods across model families, scales, and benchmarks and reduces the gap of iCoT with explicit CoT while preserving the efficiency advantages of implicit reasoning The results also support a geometric perspective of reasoning: long chain-of-thought trajectories can be effectively compressed by preserving their low-rank structure This perspective provides a principled and efficient alternative to existing iCoT distillation approaches and suggests that the essential dynamics of reasoning are governed by a low-dimensional subspace that can be exploited for both learning and inference
Limitations
The proposed formulation is motivated by an empirical observation that hidden-state reasoning trajectories exhibit strong low-rank structures across models and benchmarks While LoRi improves reasoning accuracy and training efficiency, the theoretical relationship between low-rank representation geometry and reasoning capability is still not fully understood Future work could further investigate the geometry of reasoning trajectories and the role of low-dimensional structure in CoT reasoning from a theoretical perspective
References
Jinze Bai et al. 2023. Qwen technical report. arXiv:2309.16609 Yoshua Bengio, Aaron Courville, and Pascal Vincent. 2013. Representation learning: A review and new perspectives Zhuotong Chen, Qianxiao Li, and Zheng Zhang. 2022. Self-healing robust neural networks via closed-loop control Zhuotong Chen, Zihu Wang, Yifan Yang, Qianxiao Li, and Zheng Zhang. 2024. PID control-based selfhealing to improve the robustness of large language models Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems Lieven De Lathauwer, Bart De Moor, and Joos Vandewalle. 2000. A multilinear singular value decomposition Yuntian Deng, Yejin Choi, and Stuart Shieber. 2024a. From explicit cot to implicit cot: Learning to internalize cot step by step Yuntian Deng, Kiran Prsad, Roland Fernandez, Paul Smolensky, Vishrav Chaudhary, and Stuart Shieber. 2024b. Implicit chain-of-thought reasoning via knowledge distillation Abhimanyu Dubey, Abhinav Jauhri, Ankur Pandey, et al. 2024. The llama 3 herd of models Charles Fefferman, Sanjoy Mitter, and Hariharan Narayanan. 2016. Testing the manifold hypothesis Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. 2022. Pal: Program-aided language models Noah Golowich, Allen Liu, and Abhishek Shetty. 2025. Sequences of logits reveal the low rank structure of language models Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, and Yuandong Tian. 2024. Training large language models to reason in a continuous latent space Zhenghao He, Guangzhi Xiong, Bohan Liu, Sanchit Sinha, and Aidong Zhang. 2026. Reasoning beyond chain-of-thought: A latent computational mode in large language models Anna Kuzina, Maciej Pioro, Paul N. Whatmough, and Babak Ehteshami Bejnordi. 2025. Kava: Latent reasoning via compressed kv-cache distillation Bo Li, Guanzhi Deng, Ronghao Chen, Junrong Yue, Shuo Zhang, Qinghua Zhao, Linqi Song, and Lijie Wen. 2025a. Rema: A unified reasoning manifold framework for interpreting large language models Jindong Li, Yali Fu, Li Fan, Jiahong Liu, Yao Shu, Chengwei Qin, Menglin Yang, Irwin King, and Rex Ying. 2025b. Implicit reasoning in large language models: A comprehensive survey Zhiyuan Li, Hong Liu, Denny Zhou, and Tengyu Ma. 2024. Chain of thought empowers transformers to solve inherently serial problems Zicheng Lin, Tian Liang, Jiahao Xu, Qiuzhi Liu, Xing Wang, Ruilin Luo, Chufan Shi, Siheng Li, Yujiu Yang, and Zhaopeng Tu. 2025.
Improvements for AI systems
- Bold header: Low-rank iCoT Distillation for Efficient Reasoning
The improved system can transfer complex reasoning from a computationally expensive explicit Chain-of-Thought (CoT) process into a compact implicit latent process, specifically by aligning teacher and student trajectories in a shared low-rank subspace using first- and second-order hidden-state statistics.
This allows the student to achieve length invariant
distillation, meaning it can capture the global structure of reasoning while supporting a compact latent reasoning process.
- Bold header: Improved Reasoning Accuracy on Hard Tasks
The system exhibits enhanced performance on challenging problems by aligning with the dominant geometric structure of reasoning. The paper notes that LoRi consistently outperforms prior iCoT methods, improving accuracy by up to ∼12% and achieving strong gains on GSM8K-Hard,
suggesting it is particularly effective where reasoning is most complex.
- Bold header: Enhanced Inference Latency Efficiency
The system can maintain the efficiency benefits of implicit reasoning during inference while boosting performance. Specifically, iCoT consistently achieves lower latency across all model scales
and LoRi maintains this benefit, with iCoT being about 6.9× faster for LLaMA-3.2-1B.
- Bold header: Scalable and Cost-Effective Distillation Pipeline
The method improves training efficiency by using a two-stage procedure where teacher-derived low-rank factors are precomputed once, after which the student is fine-tuned independently on a subset of the training data,
which substantially reduces training cost.
This makes distillation a more scalable and accessible approach for implicit reasoning distillation.
- Bold header: Robustness to Trajectory Length
The system's reliance on statistical matching rather than step-by-step correspondence ensures robustness against trajectory length. The formulation is explicitly described as being invariant to the length of the reasoning trajectory,
enabling it to transfer long CoT reasoning into short latent trajectories.
Sources
- Qwen Technical Report
- Training Verifiers to Solve Math Word Problems
- From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
- Implicit Chain of Thought Reasoning via Knowledge Distillation
- The Llama 3 Herd of Models
- Sequences of Logits Reveal the Low Rank Structure of Language Models
- Training Large Language Models to Reason in a Continuous Latent Space
- Triggering Chain-of-Thought via Latent Feature Interventions in Large Language Models
- KaVa: Latent Reasoning via Compressed KV-Cache Distillation
- REMA: A Unified Reasoning Manifold Framework for Interpreting Large Language Model
- Implicit Reasoning in Large Language Models: A Comprehensive Survey
- The Origins of Representation Manifolds in Large Language Models
- CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation
- To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
- LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness Signals
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- SIM-CoT: Supervised Implicit Chain-of-Thought
- Parallel Continuous Chain-of-Thought with Jacobi Iteration
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering