A Model with No Head and Many Thoughts
cs.LG, cs.CL
Submitted: 2026-08-31
Updated: 2026-08-31
Comments: Accepted to Findings of EMNLP 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models decode by projecting hidden states through a large vocabulary head at every step.
Terminology
Abstract
Large language models decode by projecting hidden states through a large vocabulary head at every step. This operation is computationally costly and forces all reasoning to be expressed in discrete tokens. We introduce Soft Latent Thinking, a method that replaces the LM head during reasoning with a lightweight projector, enabling autoregressive rollout in embedding space where reasoning steps remain continuous rather than tokenized. Experiments on DeepSeek-Qwen-1.5B and LLaMA-3.2-3B show that Soft Latent Thinking consistently improves pass@k across all k while reducing per-step compute during chain-of-thought. Our method achieves the highest pass@32 among all soft-thinking approaches, demonstrating that effective reasoning can be carried out in continuous space without discrete token generation.
Sources
- Soft Tokens, Hard Truths
- Evaluating Large Language Models Trained on Code
- Training Verifiers to Solve Math Word Problems
- The Llama 3 Herd of Models
- OpenThoughts: Data Recipes for Reasoning Models
- Training Large Language Models to Reason in a Continuous Latent Space
- Measuring Mathematical Problem Solving With the MATH Dataset
- LoRA: Low-Rank Adaptation of Large Language Models
- Large Language Models are Zero-Shot Reasoners
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- LLMs are Single-threaded Reasoners: Demystifying the Working Mechanism of Soft Thinking
- Diffusion of Thoughts: Chain-of-Thought Reasoning in Diffusion Language Models
- Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space
- SGLang: Efficient Execution of Structured Language Model Programs
- SofT-GRPO: Surpassing Discrete-Token LLM Reinforcement Learning via Gumbel-Reparameterized Soft-Thinking Policy Optimization
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks