ShapeY: A Principled Framework for Measuring Shape Recognition Capacity via Nearest-Neighbor Matching

arXiv:2604.25065 · cs.CV · Submitted 2026-04-27 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "ShapeY: A Principled Framework for Measuring Shape Recognition Capacity via Nearest-Neighbor Matching".

Jane: The gist The ShapeY framework introduces a novel and principled benchmarking system designed to evaluate shape-based recognition capability in object recognition systems using nearest-neighbor matching.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: We’re diving into the specifics of ShapeY now, which is titled "ShapeY: A Principled Framework for Measuring Shape Recognition Capacity via Nearest-Neighbor Matching." This paper introduces this benchmarking system to evaluate shape recognition capability in object recognition systems using nearest-neighbor matching.

Jane: Essentially, the core thesis is that object recognition in humans relies heavily on shape cues and the ability to recognize objects across varying three dee viewpoints. The authors show how deep networks often rely too much on non-shape cues like texture and background, which creates problems when we test for generalization and robustness.

Lu: To test this gap, ShapeY uses sixty-eight thousand two hundred grayscale images of two hundred three dimensional objects rendered from multiple viewpoints and optionally subjected to non-shape “appearance” changes. They use a nearest-neighbor matching task to probe the fine details of an OR system’s embedding space by evaluating whether object views are clustered by three dee shape similarity across varying three dee viewpoints and other nonshape changes.

Meng: The paper claims ShapeY provides a suite of quantitative and qualitative performance readouts, including error rate graphs, viewpoint tuning curves, histograms of positive and negative matching scores, and grids showing ordered best matches. They are trying to be very comprehensive in how they measure performance here.

Lalam: This task specifically probes the fine-grained structure of an OR system’s embedding space by checking if object views are clustered by three dee shape similarity across varying three dee viewpoints and other nonshape changes, which is a much more detailed test than just a simple classification score.

Tom: The authors are asking themselves a big question here: can we use this method to see if fine-tuning an AI actually teaches it a general sense of three dimensional shape? They want to know if the system learns something useful beyond just memorizing training data.

Jane: They suggest that while fine-tuning helps the system align with the images, conventional networks still have trouble handling shape when you mix in viewpoint changes or lighting shifts.

Lu: The paper points out that some systems just collapse when they focus too much on surface features like color and texture similarity instead of actual geometry.

Meng: That makes sense from an engineering standpoint; if the reconstruction loss focuses on those easy visual cues, it doesn't necessarily build a good shape representation for distinguishing objects.

Lalam: And even with the best architectures out there, they found that OCD errors persist, meaning nearby points in the AI's embedding space are still confusingly close to completely different objects.

Tom: So what does this mean for us regarding the paper ShapeY? It means it’s a way to get a clear picture of where these systems are succeeding and exactly where their weaknesses lie across many different challenges.

Jane: It gives us those quantitative and qualitative reports, showing us how the performance dips when we introduce things like viewpoint changes or when objects start looking too similar.

Lu: The implication is that we need to move past just accuracy scores and start measuring robustness against these kinds of geometric ambiguities in object recognition systems.

Meng: For me, it means we need to design training sets and evaluation benchmarks that specifically test this kind of shape-based discrimination, not just general object recognition on standard datasets.

Lalam: It points toward a future where we can have a more principled way to evaluate if an AI truly understands the three dimensional structure of what it sees.

Conclusion: Tom: We’ve seen how ShapeY sets up a rigorous test for shape recognition using nearest-neighbor matching across different views and appearance changes, so now let's talk about the paper itself, "ShapeY: A Principled Framework for Measuring Shape Recognition Capacity via Nearest-Neighbor Matching."

Jane: It really boils down to giving us a structured way to judge if an AI system is actually grasping three dimensional shape similarity when it looks at things from different angles.

Lu: The authors developed this framework using sixty-eight thousand two hundred grayscale images of two hundred three dimensional objects, specifically designed to probe the fine structure of an AI's internal space.

Meng: They use a matching task where the AI has to rank views by shape similarity, and they look at how those scores change when you move the viewpoint or change something like lighting.

Lalam: The main point is that this framework forces us to check if the system consistently links a view of an object to another view of that same object, regardless of surface changes or pose shifts.

Tom: What this means for us is that we’re moving past just looking at total accuracy scores and starting to measure how well a system handles the subtle geometry of three dimensional shapes.

Jane: They give us a whole suite of reports, both numbers and pictures, showing exactly where the performance drops when those geometric ambiguities start to show up.

Lu: The researchers also look closely at what kinds of errors happen visually in those matches, like when objects that look very similar get confused or when the system just can't resolve the shape detail.

Meng: This suggests that for practical applications, we need to design training sets and benchmarks that specifically test this kind of shape-based discrimination, not just general object recognition on standard datasets.

Lalam: It points toward a future where we have a principled way to evaluate if an AI truly understands the physical three dimensional structure of what it sees.

Tom: Exactly. ShapeY is laying out the rules for how we should judge if an AI system has actually grasped the concept of three dimensional shape recognition.

cs.CV

Submitted: 2026-04-27

Updated: 2026-10-07

Importance score: 92/100

The gist: The gist The ShapeY framework introduces a novel and principled benchmarking system designed to evaluate shape-based recognition capability in object recognition systems using nearest-neighbor

Key concepts

Nearest-Neighbor Matching Task
This task requires an object recognition system to rank all images in a database by how similar they are to a reference image in the system's internal representation (embedding space). Success is measured by whether the closest match is another view of the same object, indicating good shape understanding.
Desirable Image Database Properties
The ideal test set must contain objects that are rigid, isolated from clutter, and shot under varied conditions. This ensures that any recognition performance reflects the system's ability to judge 3D shape accurately rather than relying on color or texture cues.
Embedding Space Offset
When an object's viewpoint changes, it causes a corresponding shift in its position within the system's mathematical embedding space. Analyzing these offsets helps reveal how fine-grained geometric changes affect the system's ability to distinguish between similar 3D shapes.
OCD Errors (Out-of-Distribution)
These errors occur when images of completely different objects are incorrectly clustered closely together in the system's embedding space. The persistence of OCD errors even in top architectures suggests that shape representation remains fragile despite good performance.

Terminology

Summary

The gist The ShapeY framework introduces a novel and principled benchmarking system designed to evaluate shape-based recognition capability in object recognition systems using nearest-neighbor matching.

How it works

ShapeY comprises 68,200 grayscale images of 200 3D objects rendered from multiple viewpoints and optionally subjected to non-shape “appearance” changes. Using a nearest-neighbor matching task, ShapeY specifically probes the fine-grained structure of an OR system’s embedding space by evaluating whether object views are clustered by 3D shape similarity across varying 3D viewpoints and other nonshape changes. ShapeY provides a suite of quantitative and qualitative performance readouts, including error rate graphs, viewpoint tuning curves, histograms of positive and negative matching scores, and grids showing ordered best matches.

Desirable properties of the image database

Regarding the ideal image database to test 3D shape-based recognition capability, seven properties were considered important. These properties include:

  1. The images should depict views of the same or similarly shaped 3D objects shot under different viewing conditions (pose, lighting, etc.).

  2. The objects should be rigid to avoid the same-different shape ambiguity that arises when shape deformations are allowed.

  3. The objects should be isolated (i.e. not occluded, camouflaged, or embedded in cluttered scenes) in order that recognition performance be a pure reflection of shape representation capability.

  4. The images should be shot under favorable imaging and lighting conditions to ensure that recognition performance reflects a system’s representational capabilities, rather than its capabilities at the pixel-processing-level.

  5. All non-shape information, including color, texture and lighting of both object and background surfaces should be manipulable, both to thwart the use of non-shape cues in the recognition process, or to allow such cues to be used as distractors.

  6. The images should be organized into basic level categories, making it possible to assess an OR system’s ability to distinguish objects with varying degrees of shape inter-similarity.

  7. The set of objects and views contained in the database should be large enough so that there is a reasonable chance of detecting failure cases where nearby points in the embedding space correspond to obviously completely different (OCD) objects, if such failures exist.

The nearest-neighbor-based ShapeY matching task

Given a reference image, the ShapeY matching task asks an OR system to rank order all images in the database based on their similarity score to the reference image in the OR system’s embedding space using cosine similarity or another similar measure. A response is scored as “correct” if the closest match is to another (eligible) view of the same object; “categorically correct” if the closest match is to an eligible view of a different object within the same category; and “incorrect” if the closest match is to a distractor from a different object category.

Geometric perspective of ShapeY performance measures

When a view of an object or scene changes in any way, this leads to an offset in an OR system’s embedding space. A more fine-grained geometric view provides further insight. Specifically, if the ball contains one or more exemplars from other categories, a category-level error occurs.

Qualitative assessment of ShapeY matching errors

Visual inspection of an OR system’s matching errors provides additional insights into the fine-scale structure of the system’s embedding space. These errors are categorized into four main types:

  1. Unreasonably similar objects.

  2. Degenerate viewpoint errors.

  3. Shape resolution limit errors.

  4. OCD errors.

Can fine-tuning a network with ShapeY image set teach it a generalizable sense of 3D shape?

Fine-tuning is beneficial in that we can align the DN to our images, eliminating the OOD concern completely However, we conclude that conventional networks face significant challenges in handling shape independently of 3D viewpoint and other non-shape appearance changes.

Can other network architectures support better shape understanding?

The best-performing network, DINOv2 [37], utilized a larger curated dataset, a ViT-based architecture, and distillation-based self-supervision.

Poor performing systems suffer from dimensional collapse

In some cases, some global shape similarity is evident in addition to color/texture similarity (top row), while in most cases, shape-based matching is clearly de-emphasized. This makes sense intuitively since the MAE focuses on reconstruction, which may have caused the network to learn features for interpolation which aren’t useful for distinguishing based on semantics or 3D shape.

Beyond traditional benchmarks: towards principled evaluation of robustness in object recognition systems

ShapeY’s design is based on the idea that an object recognition system should consistently judge that a view of an object is most similar to another view of the same object, regardless of changes in the object’s surface properties, 3D pose, or viewing conditions, compared to any other view of any other object under any viewing conditions. ShapeY’s comprehensive set of numerical and graphical reports provides a detailed overview of the capabilities and weaknesses of an OR system across a spectrum of task difficulties.

B. poor performance of ResNet50 and many DN-based OR systems may stem from their architectures

Interestingly, focusing on image contours alone does not improve performance on our matching task: ResNet50 trained to be more shape-biased using Stylized ImageNet [7] performs worse than its ImageNet-trained counterpart (27% accuracy, re=2 in pr).

C. OCD errors persist in the best performing DN architecture

This indicates that even in the best-performing network, images in the embedding space remain closely surrounded by distractors, undermining its robustness.

REFERENCES

[1] K. Grill-Spector, Z. Kourtzi, and N. Kanwisher, “The lateral occipital complex and its role in object recognition,” Vision Research, vol. 41, no. 10-11, pp.

Improvements for AI systems

  1. ShapeY provides a principled benchmarking framework designed to evaluate shape-based recognition capability in OR systems, allowing researchers to assess whether an object recognition system can consistently judge that a view of an object is most similar to another view of the same object, regardless of changes in the object’s surface properties, 3D pose, or viewing conditions.

  2. ShapeY enables the analysis of embedding space structure by using nearest-neighbor matching while excluding views based on viewpoint transformations and appearance changes (viewpoint tuning curves, histograms of positive and negative matching scores). This allows systems to be evaluated not just on classification accuracy, but on the quality of the embedding space by quantifying how quickly positive match scores decay relative to the top negative match.

  3. The framework reveals that state-of-the-art models often suffer from severely entangled embedding space, meaning they exhibit OCD errors, where a best match is to an object whose shape is “obviously completely different,” signaling a breakdown in robust shape understanding.

  4. Fine-tuning strategies, particularly using the SimSiam objective (negative cosine similarity, maximizing similarity between embedding vectors of the two views of the same object), can be tested to see if they make the network understand 3D shape, although results suggest this only leads to disentanglement at the object-category level, rather than developing object-specific encoding.

  5. The framework allows for architecture selection by demonstrating that ViTs trained with masked-autoencoder reconstruction objective perform the worst, suggesting that architectures and training objectives must be carefully chosen to mitigate reliance on non-shape cues like texture.

Sources

Related papers