Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning

arXiv:2608.19669 · cs.CV, cs.LG · Submitted 2026-08-20 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning".

Tom: Scaffolding Minds proposes a two-stage training paradigm to optimize latent visual representations for multimodal reasoning,

Jane: First, who's behind it and why it matters.

Title and authors: Tom: Let's talk a bit more about who put this paper together and what they named their approach. The authors are Haoqiang Kang, Yinpeng Chen, Luyang Liu, Jesper Sparre Andersen, Abhijit Ogale, Baochen Sun, Lichan Hong. They come from Google DeepMind and UC San Diego.

Jane: It's interesting that the team is coming from such a strong research background; they’re tackling foundational problems in latent reasoning right out of the gate. Their approach is essentially proposing a two-stage training paradigm to optimize those latent visual representations for multimodal reasoning, which is what the paper calls Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning.

Lu: What's key here is that they aren't just trying to make one model better; they are proposing a system where the training process itself helps build an optimized target representation through two distinct stages, which is what makes this work unique compared to existing frameworks.

Meng: So, if I understand correctly, instead of just feeding the AI a helper image and hoping it learns something useful from that frozen encoder, they are building a dedicated component to learn the best possible latent target representation for the reasoning task. That sounds like a significant architectural shift.

Lalam: Exactly. They identify that relying on an off-the-shelf vision encoder in the supervised fine-tuning stage often results in latent representations that aren't perfectly aligned with what the downstream reasoning task actually needs, which limits how much useful information the model can absorb during that initial training phase.

The paper's summary: Tom: So, to summarize this paper, what they propose is a new way to handle latent targets by using a dedicated scaffolding encoder in the supervised fine-tuning stage to learn an optimized target representation, and then using a learned Gaussian policy in the reinforcement learning stage for direct exploration.

Jane: That makes sense when you break it down. They identify that existing methods have two main issues: first, they use an off-the-shelf vision encoder which gives suboptimal latent representations during supervised fine-tuning, and second, their reinforcement learning methods only use deterministic regularization, which means they can't explore different ways the latent space could be used for reasoning.

Lu: The core mechanism is that in Stage one they train this scaffolding encoder to map the helper image into a target latent token block that is specifically optimized for reasoning using a loss function like LCE y x, z∗. That sets the bar high for what the latent representation should look like before any actual policy learning happens.

Meng: And then in Stage two instead of just sticking to that fixed target or using simple constraints, they introduce Scaffolding RL where they sample residual latent actions from a learned Gaussian distribution, with the mean and variance depending on the current action. That’s how they enable direct exploration in the continuous latent block.

Lalam: It means they are not just optimizing for a single correct answer; because of that learned Gaussian policy, the model can actively search alternative latent trajectories during reinforcement learning to discover more useful reasoning paths that weren't obvious from the initial supervised training.

The paper's improvements: Tom: When we look at the results presented in "Scaffolding Minds," the authors show some pretty tangible gains across different types of reasoning tasks. They tested this method on spatial planning, like FrozenLake, and nine visual-centric reasoning benchmarks.

Jane: And they report some impressive numbers there. On FrozenLake spatial planning, they show an improvement over the strongest prior latent baseline by nine point five percent, and that gain gets even bigger to nineteen percent when testing on the harder thirty-two times thirty-two grids. That’s a pretty substantial difference in performance metrics for navigation tasks.

Lu: Beyond spatial reasoning, they also tested nine visual-centric reasoning benchmarks, where they achieved an average accuracy increase of five point two percent. This shows that the improvements aren't limited to just one type of task; the scaffolding approach is broadly effective across different visual reasoning challenges.

Meng: I noticed in their ablation studies that Target optimization, not just swapping out the objective function, actually drives most of the improvement they see. That suggests that getting a good starting point for the latent target representation is more important than some other specific tweaks they tried during development.

Lalam: And specifically regarding reinforcement learning, Scaffolding RL showed a gain of three point zero percent compared to other RL methods when using their scaffolding encoder, which confirms that having that structure for exploring the latent space directly makes a noticeable difference in how much the model can learn from trial and error in this context.

Conclusion: Tom: So, to wrap up our discussion on "Scaffolding Minds," we’re seeing a framework where they build an optimized latent target representation first and then use learned Gaussian policies during reinforcement learning to directly explore the latent space for better reasoning paths.

Jane: That really boils down to making sure the AI isn't just following a pre-set path, but actively discovering better paths in its internal representation when it needs to make a decision. It shows that both the quality of what the AI is learning and its ability to explore that learned space are essential for high-quality multimodal reasoning.

Lu: The implication here is that future multimodal systems might benefit from having this two-stage approach where one stage focuses on creating a high-quality, task-specific latent target, and the second stage actively samples from a distribution over residual actions based on the current state. It’s a sophisticated way to integrate representation learning with policy optimization.

Meng: For practical AI development, I see this meaning we need more robust ways to create those specialized latent targets before we even start running complex RL loops, otherwise, the exploration phase might be wasted on a poor foundation.

Lalam: Ultimately, Scaffolding Minds gives us a solid direction for improving how we design these models so that they can find more useful reasoning trajectories by exploring the latent space in a guided way. I'm really excited to see how this concept helps improve the overall culture of AI development by pushing us toward more intentional design choices.

Tom: Well, that’s a lot to take in, and it sounds like this paper gives us some very concrete directions for improving how we approach latent reasoning in multimodal systems. We'll be diving into something else next time.

Haoqiang Kang, Yinpeng Chen, Luyang Liu, *Jesper Sparre Andersen*, Abhijit Ogale, Baochen Sun, *Lichan Hong*, *Ed H. Chi*

Google DeepMind

cs.CV, cs.LG

Submitted: 2026-08-20

Updated: 2026-09-29

Importance score: 83/100

The gist: Scaffolding Minds proposes a two-stage training paradigm to optimize latent visual representations for multimodal reasoning, addressing limitations in existing methods where latent targets are

Key concepts

Scaffolding Minds
A two-stage training paradigm proposed to optimize latent visual representations for multimodal reasoning. Stage one trains a dedicated scaffolding encoder to learn an optimized target representation, and stage two uses reinforcement learning with a learned Gaussian policy for direct exploration.
Scaffolding Encoder
A component trained in the supervised fine-tuning stage that maps a helper image into a target latent token block specifically optimized for reasoning using a loss function like LCE y x, z*.
Scaffolding RL
A reinforcement learning method introduced in Stage two that samples residual latent actions from a learned Gaussian distribution. This allows the model to actively search alternative latent trajectories during trial and error exploration.

Terminology

Summary

Scaffolding Minds proposes a two-stage training paradigm to optimize latent visual representations for multimodal reasoning, addressing limitations in existing methods where latent targets are suboptimal and RL lacks explicit exploration capabilities. This framework introduces a dedicated scaffolding encoder to learn an optimized target representation during the supervised fine-tuning stage and a learned Gaussian policy for sampling residual actions during reinforcement learning. The proposed method demonstrates substantial gains over strong baselines on spatial planning and visual-centric reasoning benchmarks, proving that both the quality of the latent target and direct exploration in latent space are crucial for effective multimodal reasoning.

Limitations Addressed

The paper identifies two key limitations in current latent visual reasoning frameworks. First, in the Supervised Fine-Tuning (SFT) stage, methods typically rely on an off-the-shelf vision encoder to encode the helper image, which yields suboptimal latent representations that may not be well aligned with the downstream reasoning task. This results in a latent generator being supervised toward a convenient but suboptimal target, limiting how much useful reasoning information can be learned. Second, existing Reinforcement Learning (RL) methods either treat the latent component only through deterministic regularization, which constrains policy drift but does not create alternative latent trajectories for exploration, or they focus solely on text tokens. This means RL objectives do not directly explore alternative actions in the continuous latent block.

Stage 1: SFT via Scaffolding Encoder

The first stage focuses on learning an optimized latent target representation, denoted as a scaffolding encoder (denoted as a learned function, e.g., psi). This is achieved through two phases:

  1. The Scaffolding Phase trains the dedicated scaffolding encoder to map the helper image into a target latent tokens that are optimized for reasoning. The objective is defined by matching the generated latent block to this optimized target, using a loss function like LCE: Lscaffolding (psi) = LCE y x, z∗.

  2. The Generation Phase fine-tunes the base VLM to produce its own latent block from the input alone, supervised by a term that balances matching the optimized target and fitting the answer: Lgeneration(theta) = lambdalatent (1/K ∑ zk − z∗k2) + lambdatask LCE y x, z.

Stage 2: Scaffolding RL

The second stage introduces Scaffolding RL to enable direct exploration of latent reasoning trajectories. Instead of deterministic regularization or text-only optimization, this method samples residual latent actions from a learned distribution. The mean and variance of the Gaussian sampler are not fixed but are learned as functions of the base action ztheta, predicted by two MLP heads on shared VLM hidden states: deltaz ∼ N(μ(ztheta), σ(ztheta)2) (5). This dynamic sampling allows exploration to scale adaptively—searching a broader continuous space when the prior is uncertain, and safely exploiting the prior when it is already optimal. The final objective combines latent importance ratios with text importance ratios in a GRPO-style clipped objective: LRL = −1/G ∑ min ρ lat(theta)Ai, clip ρ lat(theta), 1−ϵ, 1+ϵ + 1/T∑ min ρ text(theta)Ai, clip ρ text(theta), 1−ϵ, 1+ϵ + βKL G ∑ (DKL πθ(· si,t) πref(· si,t)) (8).

Empirical Results and Contributions

The method's effectiveness is demonstrated across two primary settings: spatial planning (FrozenLake) and nine visual-centric reasoning benchmarks. Empirically, the approach improves over the strongest prior latent baseline by +9.5% on FrozenLake spatial planning, with this gain widening to +19% at 32×32 grid map. Across nine visual-centric reasoning benchmarks, it achieves an average accuracy increase of +5.2%. Ablation studies confirm that Target optimization, not the objective swap, drives the bulk of the improvement, and Scaffolding RL attains a largest and most stable gain of +3.0% compared to other RL methods under the scaffolding encoder. Furthermore, ablation on latent token count shows that K=4 is the best setting at every level for FrozenLake, balancing expressiveness and learnability. The framework achieves strong reasoning gains without requiring explicit intermediate image generation at inference time, maintaining a low overhead of "+4.6% over the base VLM.

Improvements for AI systems

Based on the scientific paper Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning, here are the specific improvements you can implement in AI systems and what those improved systems will be able to do:


The core improvements involve moving away from static or off-the-shelf visual representations toward a dynamically optimized, learned latent space guided by task objectives, coupled with an active exploration strategy during reasoning.

Here are the specific technical implementations derived from the paper:

  1. Learn a dedicated Scaffolding Encoder (Stage 1) that produces an optimized latent target representation for downstream reasoning tasks directly through end-to-end training loss, rather than relying on a frozen off-the-shelf vision encoder.

  2. Refine the latent representation during the Reinforcement Learning (RL) stage using a Scaffolding RL mechanism that replaces deterministic regularization with an explicit Gaussian sampler over residual latent actions, where the mean and variance of this sampler are learned functions conditioned on the current latent state.

These improvements allow for significant enhancements in multimodal reasoning capabilities:

  1. The improved system can perform multi-step spatial planning (e.g., navigating complex mazes like FrozenLake) with significantly higher accuracy, achieving gains of up to +19% on the hardest 32x32 grids compared to prior latent baselines.

  2. The system will demonstrate superior fine-grained visual reasoning across diverse benchmarks (V, BLINK, CVBench), showing improvements of up to +5.2% on average, specifically excelling in tasks requiring spatial discrimination and precise regional evidence extraction (e.g., MME-RealWorld-Lite).

  3. The overall system will exhibit a more robust and generalizable latent reasoning capability by optimizing the latent target directly for the task loss, leading to performance gains that are consistent across different reasoning modalities, rather than being tied to specific helper images or pre-trained feature spaces.

Sources

Related papers