A Model with No Head and Many Thoughts

arXiv:2608.31069 · cs.LG, cs.CL · Submitted 2026-08-31 · Read on arXiv

cs.LG, cs.CL

Submitted: 2026-08-31

Updated: 2026-08-31

Comments: Accepted to Findings of EMNLP 2026

License: http://creativecommons.org/licenses/by/4.0/

The gist: Large language models decode by projecting hidden states through a large vocabulary head at every step.

Terminology

Abstract

Large language models decode by projecting hidden states through a large vocabulary head at every step. This operation is computationally costly and forces all reasoning to be expressed in discrete tokens. We introduce Soft Latent Thinking, a method that replaces the LM head during reasoning with a lightweight projector, enabling autoregressive rollout in embedding space where reasoning steps remain continuous rather than tokenized. Experiments on DeepSeek-Qwen-1.5B and LLaMA-3.2-3B show that Soft Latent Thinking consistently improves pass@k across all k while reducing per-step compute during chain-of-thought. Our method achieves the highest pass@32 among all soft-thinking approaches, demonstrating that effective reasoning can be carried out in continuous space without discrete token generation.

Sources

Related papers