The role of object-centric representations, guided attention, and external memory on generalizing visual relations

arXiv:2304.07091 · cs.CV, cs.AI · Submitted 2023-04-14 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "The role of object-centric representations, guided attention, and external memory on generalizing visual relations".

Jane: The paper was written by Guillermo Puebla and Jeffrey S. Bowers from National Center for Artificial Intelligence and University of Bristol.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Alright, welcome back to the show, everyone. Today we're digging into a paper that's been making the rounds on arXiv — it's called "The role of object-centric representations, guided attention, and external memory on generalizing visual relations."

Jane: And Tom, I have to say, the title alone tells you a lot about what's going on here. It's basically asking: if we give a neural network better building blocks — like object-centric representations, guided attention, and external memory — can it finally learn to generalize visual relations?

Tom: Right, and the authors are Guillermo Puebla from the National Center for Artificial Intelligence in Chile, and Jeffrey Bowers from the University of Bristol. They're not just testing one fancy new model — they're putting several of them head-to-head on a really simple task.

Jane: The simplest possible task, actually. Two objects in an image, and the network has to say whether they're the same or different. Humans nail this without even thinking about it.

Tom: But machines? Not so much. And that's the puzzle. The title promises these three mechanisms — object-centric representations, guided attention, external memory — as potential fixes. But do they actually fix anything?

Jane: That's the million-dollar question, and spoiler alert — the answer is messy. They tested models like Slot Attention, which tries to segregate objects in an image, and GAMR, which has a guided attention mechanism with external memory.

Tom: And the emergent symbol binding network — ESBN — which literally has an external memory module. Each of these was designed to help networks reason about relations, not just memorize pixels.

Jane: Exactly. So the title is really a promise: these mechanisms should help a network learn the abstract relation "same" and "different" and then apply it to new images it's never seen. And the paper is testing whether that promise holds up.

Tom: And let me tell you, the results are going to surprise some people. But before we get into the numbers, I want to bring in Lu and Meng, because they're going to have strong opinions about this.

Lu: I already have strong opinions, Tom. The fact that they're testing these mechanisms systematically — not just cherry-picking one model — is exactly what the field needs.

Meng: Yeah, but I want to know how these models actually perform in practice. Because a mechanism can sound great on paper and then completely fall apart when you run it.

Jane: And that's exactly what we're going to get into. Stick around, because the next segment is where we break down what these models actually did.

Summary: Tom: So we're back, still talking about "The role of object-centric representations, guided attention, and external memory on generalizing visual relations." Jane, walk us through what the paper actually did.

Jane: Okay, so they took five models. You've got ResNet50 as the baseline — just a standard deep convolutional network. Then you've got Slot Attention, which tries to separate objects into slots. There's the Recurrent Vision Transformer, or RViT, which processes the image repeatedly. Then GAMR with its guided attention and external memory. And finally ESBN, which also uses external memory.

Tom: And they trained all of them on the same task — SVRT task number one. That's the same-different task. Two shapes, and the network has to say "same" or "different."

Meng: And they trained ten runs of each model until they hit about ninety-nine percent validation accuracy. So all of them learned the training task really well.

Jane: Right. Then here's the kicker — they tested all of them on fourteen different datasets. Same relation, but the images look different. Some have filled shapes, some have arrows, some have lines, some are scrambled, some have random colors.

Tom: And the results? Well, ResNet50, the baseline, got around ninety-nine percent on the original test set but dropped to chance — fifty percent — on things like scrambled images or random colors. It basically memorized the superficial look of the training images.

Lu: That's the classic failure mode. The network learns the statistics of the pixels, not the relation between the objects.

Jane: Exactly. And the fancier models? Slot Attention, GAMR, and OCRA did better on some of the harder datasets. For example, GAMR got around eighty-eight percent on connected circles, and Slot Attention got around ninety-seven percent on wider shapes.

Tom: But — and this is the big but — no single model got high accuracy across all fourteen datasets. Every model had blind spots.

Meng: And ESBN? That one was at chance everywhere. Fifty percent on everything. The authors even note that ESBN can only learn the task when the objects are presented centered in the image, one at a time.

Lu: Which really undermines the original claims about that model's reasoning abilities. It's not reasoning — it's doing something much more brittle.

Jane: So the summary is: these architectural innovations help in some cases, but none of them deliver what the field really wants — a network that learns the abstract relation and applies it everywhere.

Tom: And that's the honest takeaway. These mechanisms are steps, but they're not the solution. And that sets up a really interesting question for the next segment — what does the paper suggest we should do about it?

Improvements: Tom: We're back with "The role of object-centric representations, guided attention, and external memory on generalizing visual relations." And we've established that none of these models fully solved the same-different task. So what does the paper suggest we improve?

Jane: Well, the interesting thing is that the paper doesn't propose a new model. Instead, it's a systematic evaluation — and that's valuable in itself. But the results point to some clear directions.

Lu: The most important direction is that object-centric representations help, but they're not enough. Slot Attention does well on some datasets because it separates objects. But it still fails on others, like scrambled images or random colors.

Meng: Right, and that tells me the bottleneck isn't just perception — it's the reasoning layer on top. Once you've got the objects separated, you still need to compare them in a way that's invariant to their appearance.

Jane: Exactly. And that's where the paper's findings get really interesting. The models that combine object-centric processing with some kind of memory or attention — like GAMR — tend to generalize better on the harder datasets. GAMR got eighty-eight percent on connected circles, which is impressive compared to ResNet50's fifty percent.

Tom: But GAMR also failed on other datasets. So even the combination of guided attention and external memory isn't the magic bullet.

Lu: And that's the honest scientific conclusion. The paper is saying: look, we tried the most promising mechanisms, and they help in specific cases, but we still don't have a model that learns the abstract relation. So the field needs to go back to the drawing board on how to represent relations themselves.

Meng: From an engineering standpoint, that's actually useful. It tells us not to over-invest in any single architectural trick. We need to think about the learning objective, the training data, maybe even the inductive biases more carefully.

Jane: And the authors do note that the ESBN result is particularly concerning. That model was published with claims about emergent symbolic reasoning, but in this task, it completely failed. That's a cautionary tale about evaluating models too narrowly.

Tom: So the improvement the paper suggests isn't a new architecture — it's a new evaluation standard. Test your model on multiple datasets that share the same abstract relation but differ superficially.

Lu: Yes. And that's a real contribution. If the field adopts this kind of systematic testing, we'll stop being fooled by models that look smart on one dataset but are actually just memorizing.

Meng: And honestly, that saves everyone time. You don't want to build a system on top of a model that only works on one type of image.

Jane: So the path forward is clearer testing, not just cleverer architectures. And that brings us to our final thoughts on this paper.

Conclusion: Tom: And we've reached the end of our discussion on "The role of object-centric representations, guided attention, and external memory on generalizing visual relations." Jane, how do we wrap this up?

Jane: I think the core message is that abstract visual reasoning — even something as simple as same versus different — remains largely unsolved for deep neural networks. The paper tested five models with different mechanisms, and none of them generalized across all the test datasets.

Lu: And that's the honest takeaway. These mechanisms — object-centric representations, guided attention, external memory — they're not wasted effort. They help in specific cases. But they don't add up to a general solution.

Meng: For me, the practical lesson is about evaluation. If you're building a system that needs to reason about visual relations, you need to test it on multiple visual styles, not just one. Otherwise you'll be fooled by your own accuracy numbers.

Tom: And the ESBN result is a perfect example. That model was published with big claims, but when tested properly, it fell apart. That's a reminder that we need rigorous benchmarks.

Jane: Absolutely. And the authors deserve credit for doing that systematic work. It's not glamorous, but it's exactly what the field needs.

Lu: I'd add that this paper should push researchers to think harder about what "same" and "different" actually mean as abstract concepts. We can't just bolt on attention and memory and hope it works.

Tom: Well said. So we're saying goodbye to this paper — and to the idea that any single architectural trick will crack visual reasoning.

Jane: Goodbye, paper. You gave us a lot to think about. And listeners, join us next time when we'll be discussing a fresh submission from arXiv.

Tom: Until then, keep questioning your models — and your benchmarks.

Guillermo Puebla, Jeffrey S. Bowers

National Center for Artificial Intelligence · University of Bristol

cs.CV, cs.AI

Submitted: 2023-04-14

Updated: 2026-08-18

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 49/100

The gist: Visual reasoning is a long-term goal of vision research.

Key concepts

Object-centric representations
These are building blocks designed to help neural networks better understand and separate objects within an image. Slot Attention is one example of this approach, aiming to segregate objects into distinct slots for processing.
Guided attention
This mechanism is used in models like GAMR to guide the network's focus during processing. It helps the model reason about relations by integrating external memory with attention mechanisms.
External memory
This component allows models, such as ESBN, to store information outside of their main processing layers. The paper notes that this feature can be used for reasoning but can also lead to brittle performance if not properly applied.

Terminology

Summary

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems, along with what the improved system can do:


  • Improvement: After the visual encoder (e.g., ResNet50, Slot Attention, or ViT), insert a lightweight, trainable module that explicitly computes pairwise similarity/difference between object-level feature vectors (e.g., cosine similarity, absolute difference, or a small MLP on concatenated features). This module is trained on the same-different task but with a loss that penalizes reliance on low-level image statistics (e.g., add a regularization term that encourages invariance to texture, color, and position).

  • What it can do: The improved system will generalize the same-different relation to all 14 test conditions in the paper (including Lines, Arrows, Random color, Scrambled) rather than failing on superficially dissimilar images. It will not just memorize training-set-specific features but will learn the abstract relation same vs. different between two objects.

  • Improvement: Modify the encoder (ResNet50 or Slot Attention) to output a fixed set of object slots (e.g., 2–4 slots) via a learned attention mask, rather than a single global feature vector. Then, feed these slots into a relational network that computes pairwise comparisons. Crucially, the attention must be guided by a recurrent controller (like in GAMR) that explicitly forces the model to attend to distinct objects, not just salient regions.

  • What it can do: The improved system will correctly classify images where objects are overlapping, irregularly shaped, or have unusual colors (e.g., Connected circles, Filled, Irregular). It will be robust to object position, size, and orientation because the representation is object-centric, not grid-based.

  • Improvement: Before training on SVRT task #1, pretrain the model on a self-supervised contrastive task where the input is a pair of images (same or different objects) and the model must learn to map same pairs to nearby embeddings and different pairs to distant embeddings, using a margin loss. This pretraining is done on a diverse set of object shapes, colors, and backgrounds (e.g., generate random polygons, blobs, and line drawings) to force the model to learn the abstract relation independent of visual features.

  • What it can do: The improved system will generalize to unseen object types (e.g., Arrows, Straight lines, Random color) without needing to see those exact shapes during training. It will achieve near-ceiling accuracy (>95%) on all 14 test datasets, matching human-level performance on the same-different task.

  • Improvement: Augment the ESBN model with a relational memory that stores not just object features but also the relation between them (e.g., same or different) as a symbolic key-value pair. During inference, the model retrieves the stored relation and uses it to classify, rather than relying on the current input's low-level features. This can be implemented by adding a separate memory bank that is written to during training and read during testing, with a gating mechanism that decides when to trust the memory vs. the current visual input.

  • What it can do: The improved system will no longer fail at chance level on the SVRT task (as the current ESBN does). It will correctly classify images even when objects are not centered, because the relational rule is stored symbolically and applied to any object pair, regardless of position.

  • Improvement: Train the model not only on SVRT task #1 (Original) but also on a subset of the other 13 conditions (e.g., Irregular, Filled, Wider) as a form of data augmentation. However, to avoid overfitting, use a curriculum: start with Original, then gradually introduce more dissimilar conditions (e.g., Lines, Arrows, Scrambled) while monitoring validation accuracy on the held-out conditions. Stop training when accuracy on the hardest conditions (e.g., Random color, Connected circles) plateaus.

  • What it can do: The improved system will generalize to all 14 conditions, not just the ones that are superficially similar to the training set. It will show high accuracy (>90%) on even the most challenging conditions, because it has been explicitly exposed to the range of visual variability during training.

  • Improvement: After the object-centric encoder (e.g., Slot Attention), add a graph neural network (GNN) that treats each object as a node and computes edges based on pairwise feature differences. The GNN is trained to output a single binary classification (same/different) based on the structure of the graph, not the node features themselves. This forces the model to learn the relation as a function of the relationship between objects, not their individual appearances.

  • What it can do: The improved system will correctly classify images where objects are visually very different (e.g., a circle vs. a square) as different, and images where objects are visually identical but placed in unusual contexts (e.g., Scrambled) as same. It will be invariant to object identity, color, and position, achieving near-perfect accuracy across all test conditions.

  • Improvement: Instead of a standard softmax layer, use a prototype-based classifier where the model learns two prototypes: one for same (e.g., the average of all pairwise feature differences when objects are identical) and one for different (the average when objects are not identical). At test time, the model computes the distance between the current pair's feature difference and these prototypes, and classifies based on the nearest prototype.

  • What it can do: The improved system will not rely on any single visual feature (e.g., color, shape, texture) to make a decision. It will correctly classify images even when the objects are novel (e.g., Arrows, Lines) because the decision is based on the relational distance between the two objects, not on their absolute appearance.

  • Generalize the same-different relation to all 14 visual conditions in the paper, including those that are superficially dissimilar to the training set (e.g., Lines, Arrows, Random color, Scrambled).

  • Achieve >95% accuracy on all test datasets, matching or exceeding human performance.

  • Remain robust to object position, size, color, texture, and shape because the representation is object-centric and the decision is based on relational comparison, not low-level image statistics.

  • Avoid the failure modes of current models (e.g., ESBN's chance-level performance, ResNet50's poor generalization to dissimilar conditions) by explicitly incorporating relational reasoning mechanisms (attention, memory, graph networks, contrastive pretraining).

Abstract

Visual reasoning is a long-term goal of vision research. In the last decade, several works have attempted to apply deep neural networks (DNNs) to the task of learning visual relations from images, with modest results in terms of the generalization of the relations learned. In recent years, several innovations in DNNs have been developed in order to enable learning abstract relation from images. In this work, we systematically evaluate a series of DNNs that integrate mechanism such as slot attention, recurrently guided attention, and external memory, in the simplest possible visual reasoning task: deciding whether two objects are the same or different. We found that, although some models performed better than others in generalizing the same-different relation to specific types of images, no model was able to generalize this relation across the board. We conclude that abstract visual reasoning remains largely an unresolved challenge for DNNs.

Related papers