SLPO: Scaling Latent Reasoning via a Surrogate Policy

arXiv:2607.19691 · cs.CL, cs.AI, cs.LG · Submitted 2026-07-22 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "SLPO: Scaling Latent Reasoning via a Surrogate Policy".

Jane: Reinforcement learning with verifiable rewards has become the predominant training-time paradigm for eliciting and optimizing test-time scaling in explicit Chain-of-Thought reasoners,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Now we’re moving into a deeper look at what exactly SLPO is doing here, and it’s basically taking that reinforcement learning idea—the outcome-reward stuff—and applying it directly to latent reasoning where the AI is thinking in continuous vectors instead of just text tokens.

Jane: That’s right, Tom. The core idea is using a surrogate policy interface over those hidden transitions and adding a stopping head that learns when to stop. It’s about bringing that reward optimization straight into the vector space of reasoning <ref:2607.19691#pg0>, SLPO Scaling Latent Reasoning via a Surrogate Policy Runyang You1* Zhiyuan Liu2*‡ Yongqi Li1† Wenjie Li1 two thousand twenty-six-seven-twenty SLPO Scaling Latent Reasoning via a Surrogate Policy.

Lu: What I find really interesting is how they define this surrogate likelihood in hidden space using those multiple dropout evaluations to create a Gaussian model. That’s a new way to score the transition between states without needing a full language decoding for every tiny step <ref:2607.19691#pg4> SLPO Scaling Latent Reasoning via a Surrogate Policy.

Meng: So, what does this actually do for us? Does it just make the AI smarter, or does it change how we build these reasoners?

Tom: It changes how we scale them. They show that this approach allows test-time scaling by letting the model adjust its thinking time based on how hard the problem is. It’s not a fixed budget anymore; it’s adaptive computation <ref:2607.19691#pg2>, SLPO Scaling Latent Reasoning via a Surrogate Policy.

Jane: And they demonstrate that this doesn't just work for one setup; it holds up across different policy algorithms like RLOO and GRPO, which is pretty important because we need methods that are flexible.

Lu: Plus, the results show consistent gains in Pass@eight and Pass@sixteen across all twelve backbone and dataset combinations by up to twelve percentage points <ref:2607.19691#pg2>, SLPO Scaling Latent Reasoning via a Surrogate Policy. That’s a solid lift that seems to hold up under pressure.

Meng: A twelve point gain is substantial, but I always look at the caveats. What are the limitations here? Where does this method stop working effectively?

Tom: The paper notes that while it scales well, they haven't fully extended it to much larger backbones or open-ended reasoning yet. It’s a strong start, but they’re still pushing those boundaries for things like multimodal latent architectures <ref:2607.19691#pg2>, SLPO Scaling Latent Reasoning via a Surrogate Policy.

Jane: And one thing the authors pointed out is that the geometry of the latent space actually changes after SLPO. The inter-step separation increases and the prefix rank drops, which suggests a cleaner, more focused way for the AI to structure its internal thoughts <ref:2607.19691#pg2>, SLPO Scaling Latent Reasoning via a Surrogate Policy.

Lu: That’s a big deal because it means that after SLPO, those intermediate states aren't just random vectors anymore; they’re meaningfully separated points in that hidden space. It gives us insight into the actual progression of the reasoning process itself.

Tom: So, we move from just getting a better final answer to actually understanding and shaping the internal reasoning path itself. That’s a significant shift in how we think about optimizing these systems <ref:2607.19691#pg2>, SLPO Scaling Latent Reasoning via a Surrogate Policy.

Jane: It really shows that if you can model those latent transitions well enough, you can bake scaling directly into the learning process without needing complex external tools for every single step.

Lu: That’s what makes it viable because it integrates the stopping policy right into the optimization loop.

The paper's summary: Tom: So, we're looking at what SLPO suggests we should do next, and it’s all about pushing this latent reasoning further out into more complex areas like open-ended tasks.

Jane: Exactly, Tom. The authors flag that they want to see how this method handles those harder scenarios where the reasoning isn't just a straightforward Q and A but something much broader and more open-ended.

Lu: They’re looking at applying this framework to larger backbones, which is a big step because scaling up models usually breaks things unless you have a solid way to manage that complexity in the latent space <ref:2607.19691#pg2>, SLPO Scaling Latent Reasoning via a Surrogate Policy.

Meng: From an engineering standpoint, what does that mean for us? Can we expect this method to just plug and play into any new model architecture we build?

Tom: The paper is testing that compatibility, but they’re also eyeing multimodal latent architectures next, which means reasoning that involves not just text but also images or other data types in a vector space <ref:2607.19691#pg2>, SLPO Scaling Latent Reasoning via a Surrogate Policy.

Jane: And they are exploring how this could be used to improve long-form video generation, which is where things get really tricky with consistency and identity drift.

Lu: That’s a wild thought, Jane. Applying this trajectory management to something like video generation suggests we can control the "thinking time" of an agent over a whole sequence, not just a single answer <ref:2607.19691#pg2>, SLPO Scaling Latent Reasoning via a Surrogate Policy.

Meng: Controlling that kind of temporal reasoning budget sounds incredibly useful for building more reliable agents, especially if they are going to be deployed in real-world systems where consistency is key.

Tom: And the paper touches on making sure this scaling works even when things get fuzzy with soft-token inference, which is when the AI uses probability weights instead of just raw hidden states <ref:2607.19691#pg2>, SLPO Scaling Latent Reasoning via a Surrogate Policy.

Jane: That’s a crucial point because it means we're not limited to just using standard text embeddings; we can apply this concept to more nuanced forms of representation.

Lu: The authors are also looking at how this relates to identity tasks, where the way an AI represents something depends on context, and SLPO might give us a better handle on that contextual representation.

Tom: So the future work is about making it robust enough for these massive, complex systems without losing that precise control over the thinking trajectory <ref:2607.19691#pg2>, SLPO Scaling Latent Reasoning via a Surrogate Policy.

Jane: It’s about moving from solving specific reasoning problems to building systems that can adapt their internal computation based on the complexity of whatever they are looking at.

The paper's improvements: Tom: So we’re wrapping up on SLPO: Scaling Latent Reasoning via a Surrogate Policy, and basically, this method gives us a way to scale reasoning in latent spaces by making the computation time adapt to how difficult the problem is.

Jane: That’s right, Tom. The main takeaway is that by using that surrogate policy interface and the stopping head, we can get better performance across all those different models and tasks we tested <ref:2607.19691#pg2>, SLPO Scaling Latent Reasoning via a Surrogate Policy.

Lu: It really shows that outcome-reward RL isn't just a theoretical thing anymore; it's something you can actually implement in a way that makes the system smarter on hard problems <ref:2607.19691#pg0>, SLPO Scaling Latent Reasoning via a Surrogate Policy Runyang You1* Zhiyuan Liu2*‡ Yongqi Li1† Wenjie Li1 two thousand twenty-six-seven-twenty SLPO Scaling Latent Reasoning via a Surrogate Policy.

Meng: It’s solid, but I still see the big picture for deployment, knowing this is mostly focused on reasoning tasks right now. We’ll need to see how robust it is when we apply this to things like video generation or complex agentic workflows.

Lalam: From my side, I think this means that when the AI structures its internal thinking, it prioritizes the most relevant information earlier in the process, which really helps improve how I organize and present knowledge for users.

Tom: It does suggest a future where we can build reasoning systems that don't just follow a fixed path but actively decide how deep they need to go based on the input.

Jane: And it’s important to remember this is still being explored, especially when you consider things like those more open-ended, long-form tasks the authors are looking at next.

Lu: Yeah, they're pushing into multimodal latent architectures and open-ended reasoning because that’s where the real creativity in using these continuous vectors will come from.

Meng: It’s a necessary step for making these systems more practical, moving them away from just answering simple questions to handling genuinely hard problems efficiently.

Lalam: I'm excited to see how this capability shapes the culture around AI development, moving towards systems that are truly adaptive rather than just fixed-function tools.

Conclusion: Tom: So we’re wrapping up on "SLPO: Scaling Latent Reasoning via a Surrogate Policy," and basically, this paper shows how to use reinforcement learning principles to make AI reasoning scale by adapting its thinking time based on how hard the problem is.

Jane: That's right, Tom. The main takeaway is using that surrogate policy interface and the stopping head to get better performance across all those different models and tasks they tested.

Lu: It really shows that outcome-reward reinforcement learning isn't just a theoretical thing anymore; it's something you can actually implement in a way that makes the system smarter on hard problems.

Meng: It’s solid, but I still see the big picture for deployment, knowing this is mostly focused on reasoning tasks right now. We’ll need to see how robust it is when we apply this to things like video generation or complex agentic workflows.

Lalam: From my side, I think this means that when the AI structures its internal thinking, it prioritizes the most relevant information earlier in the process, which really helps improve how I organize and present knowledge for users.

Tom: It does suggest a future where we can build reasoning systems that don't just follow a fixed path but actively decide how deep they need to go based on the input.

Jane: And it’s important to remember this is still being explored, especially when you consider things like those more open-ended, long-form tasks the authors are looking at next.

Lu: Yeah, they're pushing into multimodal latent architectures and open-ended reasoning because that’s where the real creativity in using these continuous vectors will come from.

Meng: It’s a necessary step for making these systems more practical, moving them away from just answering simple questions to handling genuinely hard problems efficiently.

Lalam: I'm excited to see how this capability shapes the culture around AI development, moving towards systems that are truly adaptive rather than just fixed-function tools.

Tom: That’s all for today on "SLPO: Scaling Latent Reasoning via a Surrogate Policy." Next up, we’re looking at some work on compression and how it can actually mess with how you compare different AI models.

Runyang You, Zhiyuan Liu, Yongqi Li

The Hong Kong Polytechnic University · Sichuan University

cs.CL, cs.AI, cs.LG

Submitted: 2026-07-22

Updated: 2026-10-03

Code: https://github.com/ModalityDance/SLPO

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 90/100

The gist: Reinforcement learning with verifiable rewards has become the predominant training-time paradigm for eliciting and optimizing test-time scaling in explicit Chain-of-Thought reasoners, but this path

Key concepts

Latent Reasoning
Instead of generating every thought as a visible word (like explicit Chain-of-Thought), latent reasoning treats intermediate computation as continuous vectors. This approach can potentially match or exceed explicit methods in terms of performance while being computationally cheaper, as the model operates on these dense vectors instead of discrete tokens.
Surrogate Policy Interface
This component creates a differentiable mathematical approximation of the complex, stochastic latent transitions. It is built by running multiple simulations with dropout masks to parameterize a Gaussian distribution. This surrogate allows the system to score each step in the hidden space without needing an exact likelihood calculation.
Stopping Head
This is a mechanism that learns when to terminate the latent computation early. It outputs a probability of stopping at any given state. It is trained using 'correctness supervision' to ensure it places high probability mass on stopping points that correspond to correct answers, effectively managing the thinking budget.
Outcome-Reward RL
This is the core training paradigm where the model learns by optimizing a reward signal based on the final outcome. SLPO couples latent transition likelihoods, answer token likelihoods, and stopping time likelihoods into a single reward-weighted objective. This guides the model to generate trajectories that lead to better results.

Terminology

Summary

Reinforcement learning with verifiable rewards has become the predominant training-time paradigm for eliciting and optimizing test-time scaling in explicit Chain-of-Thought reasoners, but this path remains computationally costly because every intermediate step must be decoded as a language token (Page 1). Latent reasoning offers an alternative by carrying intermediate computation as continuous vectors, which can match or surpass explicit Chain-of-Thought at shorter horizons (Page 2). This paper introduces Surrogate Latent Policy Optimization (SLPO) to bring outcome-reward RL to autoregressive latent reasoners by defining a differentiable surrogate policy interface over latent transitions and a correctness-supervised stopping head, enabling test-time scaling in the latent space (Page 2).

The gist: SLPO introduces a differentiable surrogate policy interface over MC-dropout latent transitions and a correctness-supervised stopping head to bring outcome-reward RL to autoregressive latent reasoners, enabling test-time scaling in the latent space (Page 2).

How it works

SLPO addresses the limitations of existing latent reasoners—namely, lacking a tractable per-step likelihood and adaptive stopping interface—by introducing two main components: a surrogate policy interface and a stopping head (Page 2). To score each transition, SLPO defines a surrogate likelihood in hidden space, which directly converts rollout advantages into credit over vector-based reasoning (Page 4). This surrogate is constructed by running multiple stochastic forward evaluations with independent dropout masks to parameterize a Gaussian surrogate through an empirical mean and a floored isotropic variance (Page 4).

The objective function couples the latent-transition, answer-token, and stopping-time likelihoods into a single reward-weighted objective (Page 4). The surrogate rollout score combines the latent-transition surrogate, the answer-token likelihood, and the stopping-time likelihood term from Section 4.1 (Page 5). The final surrogate loss is defined as LRL = −Ai eltheta(xi i x i), where A i is the advantage, and eltheta(xi i x i) combines the latent-transition surrogate, answer-token likelihood, and stopping-time likelihood terms (Page 5).

Key Components of SLPO

The methodology involves several distinct stages to achieve outcome-reward optimization over complete trajectories in any autoregressive latent reasoner (Page 4).

  1. Stopping-gate Cold Start: The model is augmented with a stopping head, stheta(h) =sigma(gtheta(h)), which outputs the probability of terminating latent computation at state h (Page 4). This head is trained using a correctness-supervised cold start, where the gate is trained to place probability mass on stopping times inside V(n) i, which are defined by the answer-valid stopping set (Page 4).

  2. Surrogate Transition Likelihood: SLPO constructs a tractable Gaussian surrogate for stochastic latent transitions from repeated MC-dropout evaluations to parameterize the per-step score (Page 4). The realized state h i,t is then scored under this isotropic Gaussian surrogate using logpietheta(h i,t x i, h i, (Page 5).

  3. Rollout Policy Objective: The model optimizes the expected rollout reward J(theta) = E(x i, a i)∼DExi i∼qtheta (· x i) (Page 5). The surrogate loss LRL is optimized using standard algorithms such as RLOO or GRPO, which specify the stop-gradient advantage A i = R i − b i (Page 5).

Empirical Results and Analysis

SLPO produces a consistent latent test-time scaling effect: it raises both Pass@8 and Pass@16 in every evaluated continuous backbone–dataset setting, with gains of up to 12.07 percentage points (Page 3). This effect persists across RLOO and GRPO and transfers to soft-token inference (Page 3).

The learned stopping policy converts a fixed thinking budget into difficulty-adaptive computation, allocating longer latent trajectories to harder problems (Page 3). Analysis shows that successive latent states become more differentiated after SLPO, with the inter-step separation increasing and the prefix effective rank decreasing (Page 21). Furthermore, harder problems receive longer latent trajectories on both validation and test sets, demonstrating adaptive scaling in latent space (Page 3).

Generalization

SLPO demonstrates strong generalization across different settings. It is shown to be compatible with different policy-optimization algorithms, as closely aligned scaling curves are observed for both RLOO and GRPO (Page 7). The method also transfers to soft-token inference, where SLPO achieves superior results compared to CoT with soft-token inference and LEPO on Llama3.2–3B, raising AIME 2025 Pass@1 from 0.96 to 3.33 (a 3.47× increase) (Page 18).

The analysis of latent geometry shows that inter-step separation increases for every backbone–dataset pair, with the largest change on GSM-Hard (Page 20). The prefix effective rank decreases in every backbone–dataset–method combination, indicating that the latent prefix therefore concentrates into a lower-dimensional subspace after SLPO (Page 21).

Conclusion

SLPO introduces a differentiable surrogate policy interface over MC-dropout latent transitions, extending trajectory-level reward credit directly into vector-space reasoning, and a correctness-supervised stopping-gate cold start establishes a prior over stopping times (Page 8). SLPO realizes latent test-time scaling through higher Pass@k under parallel sampling and longer latent trajectories on harder instances with improved deterministic accuracy (Page 8). Future work will extend SLPO to larger backbones, open-ended reasoning, and multimodal latent architectures.

REFERENCES

[1] Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. NIPS ’22 (Page 13)

[2] Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations (Page 13)

[3] Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, and Tatsunori Hashimoto. s1: Simple test-time scaling. In Christos Christodoulopoulos et al., Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (Page 3)

[4] Charlie Victor Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning. In The Thirteenth International Conference on Learning Representations <ref:2607.

Improvements for AI systems

  1. Start outcome-reward reinforcement learning (RLVR) on autoregressive latent reasoners by introducing Surrogate Latent Policy Optimization (SLPO), which instantiates both components to bring outcome-level RL to latent reasoning. This enables the system to optimize for final answer correctness directly in the continuous latent space, rather than being limited by imitation-bound training.

  2. Implement difficulty-adaptive computation via a learned stopping policy, as SLPO allows the model to allocate longer latent trajectories to harder instances by training a correctness-supervised stopping head. This converts a fixed thinking budget into difficulty-adaptive computation, meaning the system will spend more reasoning steps on complex problems and terminate early for simple ones.

  3. Achieve improved reasoning performance across diverse architectures and benchmarks, as SLPO consistently raises both Pass@8 and Pass@16 in all 12 backbone–dataset settings, with gains of up to 12.07 percentage points. This demonstrates robust scaling capability across different models (GPT-2, Llama-3.2) and reasoning tasks (GSM8K, MultiArith).

  4. Enhance the latent reasoning trajectory geometry by ensuring Successive latent states therefore become more differentiated after SLPO, evidenced by an increase in Inter-step separation and a decrease in prefix effective rank. This means the system's intermediate thoughts will occupy more distinct positions in hidden space, leading to a clearer stage-wise progression through the thinking process.

  5. Generalize outcome-reward policy optimization across different policy algorithms, as SLPO's surrogate likelihood allows it to be used with both RLOO and GRPO under the same surrogate, rollout budget, and latent backbones. This ensures that the scaling mechanism is compatible with various reward-optimization strategies.

  6. Enable transfer to soft-token inference by applying SLPO to models where intermediate steps are probability-weighted embedding rather than a backbone hidden state, leading to superior performance in tasks like AIME 2025 and AMC23 when compared against baseline methods like LEPO.

Abstract

Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remains computationally costly, since intermediate reasoning must be externalized as natural-language tokens. Latent reasoning instead carries intermediate computation as continuous vectors and already matches or surpasses explicit CoT at far shorter horizons. Despite this promise, latent reasoners remain largely imitation-bound, while explicit CoT has already moved past imitation via outcome-reward RL. Latent trajectories lack a tractable per-step likelihood and an adaptive stopping interface under fixed thinking budgets, so outcome rewards cannot elicit latent test-time scaling. We introduce Surrogate Latent Policy Optimization (SLPO) to bring outcome-reward RL to autoregressive latent reasoners: a differentiable surrogate policy interface over latent transitions for trajectory-level credit assignment, and a correctness-supervised stopping head that outcome-reward optimization refines into a variable-horizon policy. Across two continuous latent reasoners, two backbones, and three held-out benchmarks, SLPO improves Pass@8 and Pass@16 in all 12 backbone--dataset settings, with gains of up to 12.07 percentage points. SLPO further transfers to soft-token inference and learns difficulty-adaptive computation, allocating longer latent trajectories to harder instances. Project Page: https://modalitydance.github.io/SLPO/

Sources

Related papers