Retrofitting Recurrent Depth into a Pretrained Language Model: Installation, Extrapolation, Transfer, and Retention at Two Parameter Budgets
Mark Shapiro
cs.CL, cs.AI, cs.LG
Submitted: 2026-07-31
Updated: 2026-08-13
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: A dense, pretrained language model can be retrofitted with recurrent depth and learn an iterative latent transition that persists after outcome-only annealing.
Terminology
Abstract
A dense, pretrained language model can be retrofitted with recurrent depth and learn an iterative latent transition that persists after outcome-only annealing. Qwen2.5-0.5B-Instruct is split into a Prelude, a weight-tied Recurrent Block, and a Coda, with an identity-preserving one-loop path and a re-entry bridge on later loops. At loop 1 the retrofit remains non-inferior to its base on a preregistered ARC battery. Three findings. First, the mechanism is a reusable procedure rather than terminal-answer lookup, and installs at two budgets: 6M trained parameters over frozen base weights and 180M full-block. With intermediate-step supervision, the model computes one task step per loop and persists when only final answers are graded. The adapter matched the full block overall (83.8% versus 84.0%), led through depth 11, and trailed beyond. Verbal fine-tuning reached 79-86% on controlled verbal renderings (zero-shot transfer was minimal), and adapter verbal training begun from the installed mechanism outpaced matched fresh training by 18.6 points, including on a held-out test set. Second, the operation extrapolates to roughly 1.5 times its supervised depth, holding 70% accuracy through depth 18. Third, a same-size scratchpad-trained model matched the recurrent model within its learned horizon but collapsed beyond it. The recurrent model won overall, 84% versus 72%, retained 53% versus 2.5% beyond depth 10, and answered 7.6 times faster. An iterative transformer can therefore perform deeper reasoning in latent space faster than comparable or larger models fine-tuned on the same task, in a system-level comparison. A second task, running the rule in reverse, exposed the limits: the inverse was learnable in isolation, but no continuation acquired it while preserving the installed mechanism and general capability, a catastrophic-interference boundary. Learned depth selection remains open.
Sources
- Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA
- Generative Recursive Reasoning
- PonderNet: Learning to Ponder
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
- Compressed Chain of Thought: Efficient Reasoning Through Dense Representations
- Universal Transformers
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
- Adaptive Computation Time for Recurrent Neural Networks
- Training Large Language Models to Reason in a Continuous Latent Space
- LoRA: Low-Rank Adaptation of Large Language Models
- Less is More: Recursive Reasoning with Tiny Networks
- Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers
- Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
- Show Your Work: Scratchpads for Intermediate Computation with Language Models
- Parcae: Scaling Laws For Stable Looped Language Models
- Qwen2.5 Technical Report
- Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models
- CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation
- Hierarchical Reasoning Model
- Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering