SpaRRTa: A Synthetic Benchmark for Evaluating Spatial Intelligence in Visual Foundation Models
cs.CV, cs.LG
Submitted: 2026-01-16
Updated: 2026-09-17
Comments: Project page is available at https://sparrta.gmum.net/
Code: https://github.com/gmum/SpaRRTa
Project page: https://sparrta.gmum.net
License: http://creativecommons.org/licenses/by/4.0/
The gist: Visual Foundation Models (VFMs), such as DINO and CLIP, excel in semantic understanding of images but exhibit limited spatial reasoning capabilities, which limits their applicability to embodied
Terminology
Abstract
Visual Foundation Models (VFMs), such as DINO and CLIP, excel in semantic understanding of images but exhibit limited spatial reasoning capabilities, which limits their applicability to embodied systems. As a result, recent work incorporates some 3D tasks (such as depth estimation) into VFM training. However, VFM performance remains inconsistent across other spatial tasks, raising the question of whether these models truly have spatial awareness or overfit to specific 3D objectives. To address this question, we introduce the Spatial Relation Recognition Task (SpaRRTa) benchmark, which evaluates the ability of VFMs to identify relative positions of objects in the image. Unlike traditional 3D objectives that focus on precise metric prediction (e.g., surface normal estimation), SpaRRTa probes a fundamental capability underpinning more advanced forms of human-like spatial understanding. SpaRRTa generates an arbitrary number of photorealistic images with diverse scenes and fully controllable object arrangements, along with freely accessible spatial annotations. Evaluating a range of state-of-the-art VFMs, we reveal significant disparities between their spatial reasoning abilities. Through our analysis, we provide insights into the mechanisms that support or hinder spatial awareness in modern VFMs. We hope that SpaRRTa will serve as a useful tool for guiding the development of future spatially aware visual models.
Sources
- DINOv2: Learning Robust Visual Features without Supervision
- DINOv3
- Learning Transferable Visual Models From Natural Language Supervision
- Microsoft COCO: Common Objects in Context
- Qwen2.5-VL Technical Report
- CroCo v2: Improved Cross-view Completion Pre-training for Stereo Matching and Optical Flow
- AI-Driven Rapid Identification of Bacterial and Fungal Pathogens in Blood Smears of Septic Patients
- SPA: 3D Spatial-Awareness Enables Effective Embodied Representation
- Revisiting Feature Prediction for Learning Visual Representations from Video
- IntPhys: A Framework and Benchmark for Visual Intuitive Physics Reasoning
- Physion: Evaluating Physical Prediction from Vision in Humans and Machines
- IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments
- MultiMAE: Multi-modal Multi-task Masked Autoencoders
- EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation
- Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?
- FlySearch: Exploring how vision-language models explore
- Beyond [cls]: Exploring the true potential of Masked Image Modeling representations
- Vision Transformers Need Registers
- Masked Feature Prediction for Self-Supervised Visual Pre-Training
- NAVI: Category-Agnostic Image Collections with High-Quality 3D Shape and Pose Annotations
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models