Self-Evolving Spatial Reasoning in Vision Language Models via Geometric Logic Consistency

arXiv:2605.18162 · cs.CV, cs.AI · Submitted 2026-05-18 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Self-Evolving Spatial Reasoning in Vision Language Models via Geometric Logic Consistency".

Jane: The gist The proposed Spatial Alignment via Geometric Evolution (SAGE) framework enforces logical consistency in Vision-Language Models through geometric and linguistic duality operations to improve their spatial reasoning capabilities.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: , so to wrap up this talk on Self-Evolving Spatial Reasoning in Vision Language Models via Geometric Logic Consistency, the paper by Junming Liu et al. introduces SAGE as a way to enforce logical consistency using duality operations during training ><ref:2605.18162#pg1>>

Jane: , it really boils down to making sure the AI doesn't just get the answer right for one picture, but that it understands *why* that answer is correct even when you change the picture slightly, which is what this framework aims to achieve ><ref:2605.18162#pg2>>

Lu: , this approach shows how you can use geometric constraints to incrementally discover and enforce equivariance with respect to a growing subgroup of operations, which is mathematically sound ><ref:2605.18162#pg3>>

Meng: , for practical application, this means we can take existing models and apply SAGE as a lightweight post-training stage, which is pretty flexible since it doesn't require you to build the whole system from scratch ><ref:2605.18162#pg1>>

Tom: , so while they show substantial gains on benchmarks like SPAR and MindCube, the limitation they point out is that their current setup focuses specifically on spatial duality; they haven't explicitly mentioned extending this self-evolving mechanism to other types of reasoning invariance yet ><ref:2605.18162#pg1>>

Jane: , so what it changes for us is seeing how task-specific rewards, like standard accuracy, can be nicely complemented by a consistency reward that exposes those deeper logical failures that simple accuracy might miss ><ref:2605.18162#pg3>>

Lu: , the future work they suggest involves extending this operation pool through automated discovery and looking into three dee multi-view consistency, which opens up a lot of possibilities for richer spatial understanding ><ref:2605.18162#pg1>>

Tom: , that’s the gist—it’s not about making the model bigger, it’s about giving it a more rigorous way to check its internal logic when dealing with space and language simultaneously ><ref:2605.18162#pg3>>

Conclusion: Tom: So we're wrapping up this look at "Self-Evolving Spatial Reasoning in Vision Language Models via Geometric Logic Consistency." Basically, these authors are trying to build a way for vision language models to check their own spatial logic using geometric rules.

Jane: They’re using something called duality operations, which means they take an input and its opposite—like a picture reflected horizontally or an option flipped—and see if the model gets the answers right for both.

Lu: The core idea is that if a model truly understands space, it has to respect both the original information and its dual, like a reflection. If it fails on one, it’s showing us where its understanding breaks down.

Meng: From an engineering side, this sounds like they’re not just training on standard answers; they're adding a specific kind of check that pushes the model to become more consistent across different spatial setups.

Lalam: For me, this means when I process complex scenes, the AI won't just guess based on patterns; it will be forced to verify its understanding against these geometric rules, which should make its reasoning much more solid.

Tom: It changes how we think about training these models—instead of just asking if the final answer is right, we’re asking if that answer makes sense when you flip the input around.

Jane: That’s a big shift because it suggests that accuracy alone isn't enough to prove a model has real spatial intuition.

Lu: They showed this works on several different benchmarks, meaning it’s not just an idea; it’s actually improving performance on tasks involving things like video understanding and structured reasoning.

Meng: And they found that the visual part of the method actually gave stronger improvements when testing those spatially grounded benchmarks.

Tom: It’s interesting because they also pointed out that this framework is pretty flexible, so you can apply it to any existing vision language model without needing a completely new setup.

Jane: That flexibility is huge for the community; it means researchers don't have to reinvent the wheel just to add more rigorous checks.

Lu: The future work they mentioned involves making the set of operations even bigger through automated discovery, which could mean uncovering spatial reasoning rules we haven't even thought of yet.

Tom: That’s where things get really exciting—if you can let a model discover its own geometric constraints, the possibilities for spatial AI are really wide open.

Shanghai Artificial Intelligence Laboratory

cs.CV, cs.AI

Submitted: 2026-05-18

Updated: 2026-10-08

Importance score: 83/100

The gist: The gist The proposed Spatial Alignment via Geometric Evolution (SAGE) framework enforces logical consistency in Vision-Language Models through geometric and linguistic duality operations to improve

Key concepts

Duality Operations
These are formal operations that pair an input transformation with its induced answer mapping. The core idea is that if the correct answer to one transformation is 'a', the correct answer to its dual transformation must deterministically be 'phi(a)'. This forces the model to respect complementary spatial structures.
Consistency Score (CT)
This metric measures a model's logical consistency. It compares the model's prediction on an original input with its prediction on the dual input. A low CT score indicates that while the model answers correctly for one transformation, it fails on its complement, signaling a lack of true spatial understanding.
Self-Evolving Operation Pool (P)
This is a dynamic set of transformations (operations) that the model learns over time. Each operation has a state (candidate, active, mastered). The framework prioritizes operations with low consistency and novelty to incrementally discover and enforce geometric equivariance within the model.
Consistency Reward ($r_{cons}$)
This auxiliary reward is added to the training objective based on duality consistency. It rewards the model when its answer for a query and its answer for its dual query are logically coherent, ensuring that both related spatial concepts are correctly understood together.

Terminology

Summary

The gist The proposed Spatial Alignment via Geometric Evolution (SAGE) framework enforces logical consistency in Vision-Language Models through geometric and linguistic duality operations to improve their spatial reasoning capabilities.

How it works

  1. SAGE formalizes paired transformations as duality operations, which are triples T = (T, ϕ, S), where T is an input transformation, ϕ is the induced answer mapping, and S is the applicability domain such that a∗(T(v, q)) = ϕa∗(v, q) for all (v, q) ∈ S The key insight is that Eq. (1) partitions the answer space into complementary pairs: if the correct answer to (v, q) is a, the correct answer to the dual T(v, q) is deterministically ϕ(a) A model that truly understands the underlying spatial structure must respect both the original and its complement

How it works

The framework operates in three stages to enforce this principle. Stage 1 is Duality Probing, which evaluates model predictions on original–dual input pairs to identify reasoning inconsistencies. It probes the model with two families of operations: Visual duality, involving geometric transformations like horizontal reflection and vertical reflection, and Linguistic duality, including option permutation and logical negation. The consistency of a model Mθ with respect to T is defined as CT (θ) = E(v,q)∼DSϕaˆθ(v, q) = ˆaθT(v, q) A model with CT (θ) ≪ 1 answers the original correctly but fails on the complement, which is the hallmark of pseudo-understanding

How it works

Stage 2 involves a self-evolving operation pool P = O = T1,..., TM with a lifecycle for each operation. Each Ti ∈ P maintains a state si ∈ [CANDIDATE, ACTIVE, MASTERED]. The framework evaluates CTi(θt) on a probe set DprobeSi filtered by each operation’s applicability domain and updates states based on transitions defined by Eqs. (3)–(5) The priority score pi = (1 − CTi) + γ · 1[ni < 3] favors operations with low consistency and a novelty bonus γ for under-explored ones

How it works

Stage 3 integrates duality consistency as an auxiliary reward within GRPO training. The total reward for completion oi is defined as r(oi) = racc(oi)accuracy + rfmt(oi)format + rcons(oi)consistency The consistency reward measures whether the model’s answer and its complement are logically coherent: rcons(oi) = λ · 1/G/2 X G/2j=1 1ϕans(oi) = ans(o'j) The policy update follows GRPO, maximizing J (θ) = E h 1 G X G i=1 Aˆi logMθ(oi v, q) i − β DKLMθ∥Mθref

How it works

The self-evolving mechanism is controlled by hyperparameters such as the consistency weight λ = 0.3, maximum active operations K = 3, evaluation interval E = 100, and mastery threshold τ = 0.75. The anti-forgetting spot-check mechanism with probability pf = 0.2 periodically reactivates previously mastered operations when their consistency degrades below 0.8τ

How it works

Experiments on six video understanding benchmarks and seven spatial reasoning benchmarks demonstrate consistent improvements on the majority of benchmarks. SAGE improves duality consistency and achieves competitive or superior performance using substantially less training data than prior GRPO-based methods The results suggest that duality consistency provides a complementary post-training signal to task-specific rewards and helps expose reasoning failures that standard accuracy alone may overlook

How it works

The ablation study shows that the visual branch yields stronger gains on spatially grounded benchmarks such as VSI-Bench and VideoMMME, while the linguistic branch provides more balanced but relatively smaller improvements. The full SFT+SAGE model achieves the best overall performance on most benchmarks, demonstrating the complementarity of supervised initialization and duality-driven optimization This indicates that SFT may introduce representation biases that partially overlap with or interfere with the duality-based objectives

How it works

The analysis of mathematical constraints shows that duality inconsistent models are provably suboptimal and that duality constraints reduce the effective hypothesis space. Proposition 1 states that under M duality operations with non-overlapping orbits of total size NT, the duality-feasible class satisfies HT ≤ C N−NT /2 and VCdim(HT) ≤ N − NT /2 The self-evolving mechanism incrementally discovers and enforces equivariance with respect to a growing subgroup

How it works

The framework is model-agnostic and applicable as a lightweight post-training stage to any existing VLM. Future directions include extending the operation pool through automated discovery, incorporating 3D multi-view consistency, and applying the self-evolving mechanism to other forms of reasoning invariance beyond spatial duality The framework is model-agnostic and applicable as a lightweight post-training stage to any existing VLM. The framework is model-agnostic and applicable as a lightweight post-training stage to any existing VLM

How it works

SAGE consistently improves the spatial reasoning ability of Qwen2.5-VL-7B across all benchmarks, achieving substantial gains on challenging tasks such as SPAR, MindCube, and OmniSp. The improvements on SPAR (+12.6) and MindCube (+15.0) indicate that SAGE is especially effective in enhancing relational and structured spatial reasoning Compared with strong baselines such as InternVL-2.5-8B and SpatialLadder-3B, our method achieves overall superior or competitive performance across most benchmarks, despite being built on a relatively lightweight backbone The results demonstrate that SAGE provides a strong and generalizable inductive bias for spatial reasoning, rather than overfitting to specific benchmark distributions

How it works

The framework is model-agnostic and applicable as a lightweight post-training stage to any existing VLM. Future directions include extending the operation pool through automated discovery, incorporating 3D multi-view consistency, and applying the self-evolving mechanism to other forms of reasoning invariance beyond spatial duality The framework is model-agnostic and applicable as a lightweight post-training stage to any existing VLM

How it works

The framework is model-agnostic and applicable as a lightweight post-training stage to any existing VLM.

Improvements for AI systems

  1. System can achieve robust spatial reasoning by enforcing logical coherence across paired inputs, as SAGE enforces this principle through three stages: (1) duality probing, which evaluates model predictions on original–dual input pairs to identify reasoning inconsistencies. This helps models move beyond surface-level statistical patterns over genuine logical invariances.

  2. System will possess enhanced generalization to unseen data by utilizing a self-evolving operation pool that periodically probes the model for inconsistencies, identifies problem types that appear solved on the original input but fail under dual formulations, and promotes those operations for active training. This allows training to focus on current reasoning vulnerabilities.

  3. System can be deployed as a lightweight post-training stage by integrating duality consistency into GRPO training, providing supervision beyond standard answer accuracy via the reward function that measures whether the model’s answer and its complement are logically coherent.

Sources

Related papers