SCOUT: Unlocking Enhanced Spatial Reasoning via Structured Chain-of-Thought and Multi-Objective Process Reward

arXiv:2608.12220 · cs.CV, cs.AI · Submitted 2026-08-12 · Read on arXiv

Zile Zhou, Huining Yuan, Weichen Zhang, Xinlei Chen, Xiao-ping Zhang

Shenzhen International Graduate School, Tsinghua University

cs.CV, cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: 26 pages, 5 figures

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: SCOUT (Structured Chain-Of-Thought Utilizing Process-Supervised RL Training) is a framework proposed to enhance the spatial reasoning capabilities of Vision-Language Models (VLMs).

Terminology

Summary

SCOUT (Structured Chain-Of-Thought Utilizing Process-Supervised RL Training) is a framework proposed to enhance the spatial reasoning capabilities of Vision-Language Models (VLMs). The paper identifies two critical bottlenecks in existing approaches: (1) recent reinforcement learning (RL) methods with verifiable outcomes suffer from poor credit assignment across intermediate reasoning steps, and (2) structured reasoning approaches overlook critical depth perception necessary for comprehensive 3D understanding.

To address these challenges, SCOUT introduces a structured Chain-of-Thought (CoT) framework that explicitly models 3D environmental perception. The reasoning trajectory is encapsulated within a block and follows a sequential format: `............ `. The module generates a global semantic description, the module quantitatively describes key objects with their bounding boxes and depths, and the phase performs explicit logical reasoning and deduction based on the extracted depth values.

The paper also introduces a novel RL algorithm featuring multi-objective process rewards and a tailored advantage estimation method. The multi-objective process rewards include: (1) a Regularized Grounding Reward that aligns predicted objects with ground-truth objects using the Hungarian algorithm, incorporating semantic similarity, Efficient IoU, and depth consistency, with a cardinality penalty to prevent over-generation; (2) a Depth Reward that penalizes deviations in the z-axis; (3) a Reasoning Consistency Reward that uses blind verification to ensure the reasoning path logically entails the final prediction; (4) an Accuracy Reward based on whether the prediction matches the ground-truth answer; and (5) a Format Reward for correctly enclosing reasoning within the specified tags.

The advantage estimation method assigns distinct credit to three functional segments of the CoT: Perception (from to), Analysis (from to), and Answer (from to). The method aggregates normalized rewards into stage-specific base advantages and blends local process supervision with the global outcome advantage using parameters α1 = 0.3 and α2 = 0.3. Token-level advantages are then allocated based on functional segments, and the policy is optimized using a clipped surrogate objective with KL regularization.

To support the framework, the authors developed SCOUT-24k, a comprehensive structured spatial reasoning CoT dataset synthesized through a customized pipeline. The dataset comprises four categories: Spatial Relation Understanding, Relative Distance Prediction, Perspective Transformation Reasoning, and Object-Centric Spatial Reasoning. The construction pipeline involves extracting spatial and semantic information from source images (from EmbSpatial and STVQA datasets) using Qwen-VL-Max and Depth-Anything-3, synthesizing reasoning steps based on templates, and refining the generated reasoning processes with human expert review.

Experimental results demonstrate that SCOUT-3B improves upon baseline models by 16.85% on general spatial benchmarks and 6.3% on complex spatial reasoning tasks. The larger SCOUT-7B outperforms GPT-4o by a margin of 4.28% on general spatial benchmarks and 0.87% on complex spatial reasoning tasks. Specifically, SCOUT-7B achieves an overall score of 79.66 on general spatial benchmarks compared to GPT-4o's 75.38, and 61.79 on complex spatial reasoning tasks compared to GPT-4o's 60.92. Despite being trained exclusively on single images, SCOUT-7B exhibits robust out-of-domain generalization to multi-image and video scenarios, achieving gains of 2.46% on ViewSpatial and 3.13% on the multiple-choice section of VSI-Bench.

Ablation studies validate the necessity of each reward component. The full method achieves the highest average accuracy of 67.94%, significantly outperforming standard GRPO baselines trained without process rewards (65.15%) or without fine-grained credit assignment (65.24%). Excluding the perception advantage (α1 = 0) causes grounding and depth rewards to fail to optimize entirely, while excluding the reasoning advantage (α2 = 0) leads to a catastrophic collapse in the consistency reward. Sensitivity analysis shows that α = 0.3 achieves the best overall performance, with accuracy degrading as α increases.

The paper concludes that SCOUT establishes a new approach for cultivating spatial reasoning in VLMs, paving a pathway toward the next generation of spatially aware AI systems. The authors note limitations including computational constraints restricting experiments to 3B and 7B parameter scales, reliance on bounding box and label annotations, and the requirement for a strictly structured CoT format. Future work directions include scaling to larger models, integrating multi-image and video modalities, and relaxing the strict dependency on structured CoT formats.

Improvements for AI systems

Improvements to AI systems:

  1. Implement a structured, multi-stage reasoning pipeline with explicit 3D perception modules. The AI system will generate reasoning in three distinct phases: (a) global scene captioning, (b) quantitative object grounding with bounding boxes and depth values, and (c) logical deduction based on those depth estimates. This enables the system to perform spatial tasks like relative distance prediction, perspective transformation, and object-centric spatial reasoning with verifiable intermediate steps.

  2. Integrate a multi-objective process reward mechanism during reinforcement learning. The AI system will be trained using five simultaneous reward signals: semantic grounding accuracy (via Hungarian matching with IoU and depth consistency), depth deviation penalties, logical entailment between reasoning and final answer, final answer correctness, and output format compliance. This improves credit assignment by rewarding each reasoning step individually, rather than relying solely on outcome-based rewards.

  3. Deploy a segment-aware advantage estimation for fine-grained policy updates. The AI system will assign distinct credit to three functional segments of its reasoning trajectory—perception, analysis, and answer—by blending normalized process rewards with global outcome advantages (using α1=0.3 and α2=0.3). This allows the system to identify and reinforce which specific part of its reasoning (e.g., depth extraction vs. logical deduction) contributed to success or failure, leading to more targeted learning and faster convergence.

  4. Use a blind verification mechanism for reasoning consistency. The AI system will internally check whether its reasoning path logically entails the final prediction, independent of the ground-truth answer. This prevents the system from generating plausible but logically disconnected reasoning, improving robustness in out-of-distribution spatial scenarios.

What the improved AI system can do:

  • Perform 3D spatial reasoning on single images with explicit depth-aware object localization, achieving a 16.85% improvement over baseline VLMs on general spatial benchmarks and 6.3% on complex spatial reasoning tasks (e.g., relative distance, perspective change, object-centric queries).

  • Outperform GPT-4o on general spatial benchmarks (79.66 vs. 75.38) and complex spatial reasoning (61.79 vs. 60.92) at the 7B parameter scale.

  • Generalize to multi-image and video inputs despite training only on single images, with gains of 2.46% on ViewSpatial and 3.13% on VSI-Bench multiple-choice tasks—enabling applications in autonomous navigation, augmented reality, and video scene understanding.

  • Self-correct its reasoning by isolating failures to specific stages (e.g., depth estimation vs. logical deduction) and adjusting its policy accordingly, reducing cascading errors in long-horizon spatial tasks.

  • Produce interpretable, structured outputs (caption → scene graph with depths → analysis → answer) that can be audited by humans or downstream systems, making it suitable for safety-critical applications like robotics and medical imaging analysis.

Abstract

Existing Vision-Language Models (VLMs) exhibits a critical bottleneck in robust spatial reasoning. Recent reinforcement learning (RL) methods aim to close this gap with verifiable outcomes, yet they suffer from poor credit assignment across intermediate reasoning steps. Concurrently, structured reasoning approaches overlook the critical depth perception necessary for comprehensive 3D understanding. To address these challenges, we propose SCOUT (Structured Chain-Of-Thought Utilizing Process-Supervised RL Training). Specifically, we design a structured Chain-of-Thought (CoT) framework that explicitly models 3D environmental perception to ensure robust spatial understanding and reasoning. Furthermore, we introduce a novel RL algorithm featuring multi-objective process rewards and a tailored advantage estimation method, facilitating fine-grained credit assignment across distinct segments of the reasoning trajectory. To support our framework, we develop SCOUT-24k, a structured spatial reasoning CoT dataset synthesized through a customized pipeline. Extensive evaluations demonstrate that SCOUT-3B improves upon baseline models by 16.85% and 6.3% on general spatial benchmarks and complex spatial reasoning tasks respectively. Notably, our larger SCOUT-7B even outperforms GPT-4o by a margin of 4.28%. Moreover, despite being trained exclusively on single image, SCOUT-7B exhibits robust out-of-domain generalization to multi-image and video scenarios. These empirical results render SCOUT as a critical step towards next generation of spatially-aware VLMs.

Sources

Related papers