BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs".
Jane: Zero-shot geometric comprehension in Vision-Language Models (VLMs) is being rigorously tested by introducing BareBones,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Let's talk about the title of this paper, "BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs," and who put it together. It’s interesting how they named it so directly—it really sets the stage for what they are testing.
Jane: The authors include people from different backgrounds, which is always interesting when you see a new benchmark proposed; it suggests a broad effort to stress-test this area across various model types and datasets.
Lu: The title itself is precise because it focuses on "Zero-Shot Geometric Comprehension," which implies that the models aren't just recognizing things they've seen before, but understanding shapes they haven't been explicitly trained for in a zero-shot setting. This pushes the boundary of what we expect from these multimodal systems.
Meng: I noticed the authors are listed from places like UCF and Calgary, which tells us this isn't just an internal effort; it’s a collaborative piece coming from different research groups, which is good for validating the findings across different model architectures.
Lalam: I see how the authors chose to name it BareBones because it sounds stripped down, like the bare minimum structure needed to test pure geometry, which makes me think about how we can distill complex understanding down to its most essential components.
The paper's summary: Tom: So, let's get into what the paper actually summarizes; they’re showing us that when you take these models and strip away all the color, texture, and background information, their performance collapses severely across every single benchmark. That consistent drop is the main thing they are highlighting.
Jane: Exactly; it shows that for many of these AI systems, recognizing an object relies heavily on those visual surface cues—the texture or color—rather than just its underlying spatial layout. This moves the discussion away from simple pattern recognition toward actual geometric understanding.
Lu: The summary points out that this isn't just a minor hiccup; it’s a universal degradation across all architectures and parameter scales, which suggests the problem is rooted in how the vision encoders are fundamentally processing input, not just being under-parameterized.
Meng: I see what you mean; if it fails regardless of model size, then scaling up models won't fix this specific type of failure; we're looking at a deeper representation issue within the vision backbone itself.
Lalam: The summary really emphasizes that when the vision encoders fail to parse silhouettes, the language decoder defaults to pretraining statistics, which suggests a reliance on learned text associations rather than true visual reasoning capabilities.
The paper's improvements: Tom: Moving into what they suggest as improvements, they are proposing a way to evaluate this more rigorously using these pure pixel-level silhouettes and the two distinct modes: High-Quality and Silhouette. This is their core methodological suggestion for testing this hypothesis.
Jane: They introduce the BareBones benchmark itself as the improvement because it provides that rigorous yardstick by removing all surface patterns, giving us a clean way to measure geometric understanding directly, rather than relying on more complex, texture-rich datasets alone.
Lu: The specific improvement they suggest is training models on these silhouettes directly, which they warn against because they think it reduces a deep structural intelligence test down to something much simpler and less meaningful—just 2D pattern matching <ref:2604.10528#pg0>.
Meng: That warning is significant for practical implementation; if we train only on silhouettes, we might just be creating models that are very good at drawing outlines but don't possess the reasoning needed for complex real-world scenarios.
Lalam: I think the implication here is that to get better results, we need to design training methods that specifically encourage the AI to focus on continuous boundaries rather than surface textures, which sounds like a necessary shift in how we teach these systems visual concepts.
Conclusion: Tom: So, wrapping up this discussion on BareBones: the core message is that achieving genuine geometric grounding requires models to parse continuous structural boundaries without any surface pattern cue. This test confirms that current architectures have significant blindspots when it comes to pure shape comprehension.
Jane: It’s clear that the consistent performance collapse under RGB deprivation, especially across all tested datasets, points toward a fundamental deficit in how these models interpret visual structure when texture is removed from the equation.
Lu: The implication for the future is a need to develop shape-centric encoders that can robustly represent 2D topological information regardless of lighting or surface detail; this could unlock entirely new ways of reasoning about visual data <ref:2604.10528#pg0>.
Meng: For practical application, it means any future system we build needs to be explicitly tested on these silhouette challenges before we assume it has solid geometric understanding in ambiguous conditions. We need to bake this kind of structural robustness into the training pipeline from the start.
Lalam: I think this work provides a vital tool for us in developing better benchmarks, allowing us to move beyond just semantic correlation testing and actually quantify the emergence of structural intelligence in future multimodal AI systems.
Tom: That’s a lot to take in; we've really seen how this paper establishes such a rigorous yardstick for geometric grounding with the BareBones benchmark. What a vital piece of research we’ve been discussing today.
University of Central Florida · Independent Researcher
cs.CV
Submitted: 2026-04-12
Updated: 2026-10-05
Importance score: 81/100
The gist: Zero-shot geometric comprehension in Vision-Language Models (VLMs) is being rigorously tested by introducing BareBones, a new benchmark designed to expose a severe performance collapse when models
Key concepts
- BareBones Benchmark
- A new test designed to measure pure shape comprehension by removing all color, texture, and background information from images. It forces models to recognize objects based only on their structural silhouette rather than visual appearance.
- Texture Bias Cliff
- The phenomenon where performance drops severely when RGB textures are removed. This indicates that models exploit superficial visual shortcuts (like colors or patterns) instead of learning the underlying, continuous geometric structure of an object.
- Zero-Shot Setting Constraints
- Strict rules applied during testing: models must perform without prior fine-tuning (zero-shot). They are constrained by a low temperature and a token limit to ensure they provide exact semantic retrieval based on the provided shape alone.
Terminology
Summary
Zero-shot geometric comprehension in Vision-Language Models (VLMs) is being rigorously tested by introducing BareBones, a new benchmark designed to expose a severe performance collapse when models are deprived of RGB textures. This research matters because it isolates whether VLMs genuinely understand geometric structure or merely exploit superficial visual shortcuts, thereby establishing a Texture Bias Cliff
that reveals fundamental architectural blindspots in current architectures.
The gist
A consistent, severe performance collapse under complete RGB deprivation is observed across all architectures and datasets when models are tested on pure pixel-level silhouettes, revealing a deficit rooted in the inability to parse continuous geometric boundaries without surface pattern cues.
BareBones Benchmark Construction
The BareBones benchmark is designed to stress-test pure geometric shape comprehension by removing all color, texture, and background information. This is achieved through a Shape-Only Canonicalization
process where the dominant segmentation mask for each source image is binarized to a pure structural silhouette (white foreground on black background). The evaluation spans six high-fidelity segmentation taxonomies: ImageNet-S, DIS5K, ThinObject5K, PASCAL VOC, CUB-200, and the novel flagship collection WTP-Bench.
Evaluation Modalities and Constraints
The evaluation is conducted in two explicit visual modes: (1) High-Quality (HQ) Reference—where textures, colors, and shading are provided to establish the texture upper-bound baseline; and (2) Silhouette (Shape) Mode—where only the binarized structural mask is provided. To ensure a fair comparison, all silhouettes are cropped to their bounding box and resized to a 200px maximum dimension. Furthermore, evaluations adhere to strict zero-shot settings: temperature is set to T=0.0 for APIs, and a hard ceiling of 100 generated output tokens is enforced to penalize verbose safety-refusals or descriptive hallucinations, compelling exact semantic retrieval. Accuracy is measured as top-1 string match.
Key Findings on Performance Collapse
The analysis reveals a Texture Bias Cliff,
where removing RGB information induces a severe, consistent performance collapse across all architectures and datasets. This phenomenon is characterized by:
-
Universal degradation: The cliff is
universal across architectures
andinvariant to parameter scale.
-
Architectural failure: Parameter scaling does not resolve the issue; models from 1B to 26B cluster between 1–5% silhouette accuracy, confirming the failure is
architectural, not a function of scale.
-
Geometric Blindness: The collapse is rooted in
hallucination of pretraining priors rather than downstream reasoning failures,
as vision encoders default to textual priors when they fail to parse silhouettes.
Morphological and Typological Degradation
The difficulty of the task is stratified by visual domain and object complexity. Figure 9 shows that fine-grained biological targets (Birds, Animals) collapse to single digits
on silhouettes because feathers and fur are rich texture sources. Coarse object categories (Vehicles, Indoor objects) degrade more gracefully. Furthermore, the elemental and morphological stratification in Figure 10 confirms a clear geometric hierarchy of difficulty,
where Ghost-, Bug-, and Dragon-type silhouettes (amorphous geometry) are consistently the hardest to recognize.
Pre-training Bias and Hallucinations
When vision encoders fail to parse silhouettes, the language decoder falls back on pretraining corpus statistics. Table 8 quantifies this behavior, showing that 79.2% of open-weight misclassifications on WTP-Bench silhouettes default to a Generation 1 target,
despite Gen 1 comprising only 19.6% of the evaluation set. This demonstrates that the failure is rooted in pre-training corpus frequency, not objective geometric difficulty.
The consistent regression from Generation 1 to Generation 8 confirms models rely on memorized priors rather than geometric reasoning.
Critical Geometric Archetypes
Table 4 identifies targets yielding exactly 0.0% accuracy across all evaluated models in silhouette format. These universally impenetrable classes share three recurring geometric archetypes: (1) Gigantamax forms with extensive negativespace arms, tentacles, and floating disconnected geometry
; (2) regional variants that differ from their base forms only in minor appendage repositioning invisible at silhouette scale
; and (3) near-radially-symmetric forms that collapse into indistinguishable blobs.
These findings suggest that the failure is not due to label ambiguity but rather a deficit in representing geometric transformations.
Conclusion on Geometric Grounding
BareBones establishes a rigorous yardstick for genuine geometric grounding
by eradicating 2D surface patterns, localized textures, shading, and contextual backgrounds. The results confirm that precise geometric grounding requires models to parse continuous structural boundaries without any surface pattern cue. Training directly on BareBones silhouettes is warned against because it reduces a profound structural intelligence test to trivial 2D pattern-matching.
Improvements for AI systems
Based on the findings of BareBones, here are specific improvements for AI systems and what those improved systems will be able to do:
-
Improved Geometric Grounding in Vision-Language Models (VLMs):
-
Improved System Capability: VLMs will gain genuine geometric reasoning capabilities rather than relying on superficial texture shortcuts.
-
Architectural Shift to Shape-Centric Encoders:
-
System Capability: VLMs will develop robust representations of 2D structural topology, allowing them to accurately identify and retrieve objects based purely on contour and boundary geometry, regardless of lighting, color, or surface texture (the
Texture Bias Cliff
will be mitigated). -
Enhanced Fine-Grained Classification Precision:
-
System Capability: VLMs will achieve near-perfect discrimination between geometrically similar classes (e.g., different bird species with similar plumage patterns) by focusing on subtle structural differences in appendages, beak curvature, and body plans that are preserved in the silhouette.
-
Robustness Against Distribution Shift and Hallucination:
-
System Capability: VLMs will significantly reduce the tendency to default to high-frequency textual priors (like
American Crow
orAfghan Hound
) when visual input is ambiguous or noisy, leading to more grounded and less hallucinatory predictions in zero-shot settings. -
Improved Morphological Reasoning:
-
System Capability: VLMs will accurately model complex geometric transformations, specifically distinguishing between base forms and their Mega/Gigantamax evolutionary counterparts by recognizing the structural modifications (e.g., external appendages, enlarged regions) rather than treating them as entirely novel entities.
-
Scalability-Invariant Geometric Intelligence:
-
System Capability: The performance gap between small and large models on silhouette tasks will be closed, indicating that architectural improvements in the vision encoder itself—rather than merely parameter count—are the primary bottleneck for geometric understanding.
-
Data-Driven Benchmarking for Geometric Fidelity:
-
System Capability: Researchers can develop new benchmarks (like BareBones) that rigorously test
pure geometric comprehension,
providing a quantifiable yardstick to measure the emergence of structural intelligence in future multimodal models, moving beyond semantic correlation testing.
Sources
- Pixtral 12B
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- LLaVA-OneVision: Easy Visual Task Transfer
- SmolVLM: Redefining small and efficient multimodal models
- Verification Learning: Make Unsupervised Neuro-Symbolic System Feasible
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models