T-LoopFormer: Token-Level Elastic-Depth Looped Transformers for Latent Reasoning with Dynamic Routing
cs.AI
Submitted: 2026-09-14
Updated: 2026-09-25
Code: https://github.com/YuMingQian1234/T-LoopFormer
Project page: http://skylion007.github.io/OpenWebTextCorpus
License: http://creativecommons.org/licenses/by/4.0/
The gist: Looped Transformers have recently demonstrated strong performance in both reasoning and language tasks by reusing a shared set of parameters across multiple iterations, achieving parameter efficiency
Terminology
Abstract
Looped Transformers have recently demonstrated strong performance in both reasoning and language tasks by reusing a shared set of parameters across multiple iterations, achieving parameter efficiency without sacrificing representational power. Besides, looped Transformers perform inference directly in the latent space (latent reasoning) to reduce the number of tokens consumed during inference, thereby achieving improved sample efficiency. However, these models typically apply a fixed recursion depth uniformly to every token, leading to suboptimal compute allocation and leaving significant efficiency gains on the table. In this work, we propose dynamic token-choice routing for looped transformers, enabling each token to adaptively determine its own number of loop iterations based on its hidden state. We use a dynamic router to decide whether a token should continue recursing or exit early, allowing simple tokens to bypass unnecessary computation while hard tokens receive deeper processing. To ensure that this adaptive mechanism does not compromise decoding efficiency, we further introduce recursion-wise KV caching, which maintains an independent key-value cache for each recursion loop. This design ensures that tokens at different depths only attend to their corresponding cached states, effectively eliminating redundant computations for exited tokens and enabling fast autoregressive decoding. Extensive experiments show that T-LoopFormer reaches the sota performance under the same parameters on PPL and 10 zero-shot reasoning tasks, even surpassing the base model at 24x FLOPs and our model could reach the lowest inference latency, which validate the effectiveness of token-choice router and recursion-wise KV cache. Code: https://github.com/YuMingQian1234/T-LoopFormer.
Sources
- PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
- xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction
- Loop as a Bridge: Can Looped Transformers Truly Link Representation Space and Natural Language Outputs?
- Compressed Chain of Thought: Efficient Reasoning Through Dense Representations
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Universal Transformers
- ForesightKV: Optimizing KV Cache Eviction for Reasoning Models by Learning Long-Term Contribution
- CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective
- Training Large Language Models to Reason in a Continuous Latent Space
- Training Compute-Optimal Large Language Models
- LoopFormer: Elastic-Depth Looped Transformers for Latent Reasoning via Shortcut Modulation
- Less is More: Recursive Reasoning with Tiny Networks
- Step-resolved data attribution for looped transformers
- ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
- A Survey on Large Language Model Acceleration based on KV Cache Management
- CommVQ: Commutative Vector Quantization for KV Cache Compression
- Let's Think Dot by Dot: Hidden Computation in Transformer Language Models
- Parcae: Scaling Laws For Stable Looped Language Models
- Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
- On Expressive Power of Looped Transformers: Theoretical Analysis and Enhancement via Timestep Encoding
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection