LoopCD: Loop-wise Contrastive Decoding for Improving Reasoning in Looped Language Models
cs.CL
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: Accepted to EMNLP 2026 Main Conference
Code: https://github.com/hoeng4/LoopCD
License: http://creativecommons.org/licenses/by/4.0/
The gist: Looped Language Models (LoopLMs) perform "latent reasoning" by recursively refining internal latent representations with shared weights, offering a more effective alternative to explicit verbal
Terminology
Abstract
Looped Language Models (LoopLMs) perform "latent reasoning" by recursively refining internal latent representations with shared weights, offering a more effective alternative to explicit verbal reasoning. Despite their effectiveness, we find that LoopLMs remain prone to loop instability: unstable refinement across iterations can produce localized uncertain "hard" tokens associated with reasoning errors. To address this, we propose LoopCD, loop-wise contrastive decoding that enhances the reasoning performance of LoopLMs by intervening on these tokens at inference time. Specifically, we exploit the internal dynamics of LoopLMs and contrast the logits from earlier iterations with logits from the last refined iteration to form the final sampling distribution. We find that this strategy is highly efficient, introducing only negligible inference overhead and requiring no additional training, while effectively improving reasoning performance by naturally refining reasoning-critical hard tokens. Extensive experiments show that our method improves the performance of recent representative LoopLMs across various reasoning tasks.
Sources
- Program Synthesis with Large Language Models
- Evaluating Large Language Models Trained on Code
- Training Verifiers to Solve Math Word Problems
- DeepSeek-V3 Technical Report
- Deep Think with Confidence
- Universal Reasoning Model
- Training Large Language Models to Reason in a Continuous Latent Space
- Measuring Mathematical Problem Solving With the MATH Dataset
- Classifier-Free Diffusion Guidance
- Less is More: Recursive Reasoning with Tiny Networks
- Solving Quantitative Reasoning Problems with Language Models
- Large Language Model Guided Tree-of-Thought
- Contrastive Decoding Improves Reasoning in Large Language Models
- OpenAI GPT-5 System Card
- Denoising Diffusion Implicit Models
- Let's Verify Step by Step
- Hierarchical Reasoning Model
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
- Scaling Latent Reasoning via Looped Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering