Emergent Neural Network Mechanisms for Generalization to Objects in Novel Orientations
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Emergent Neural Network Mechanisms for Generalization to Objects in Novel Orientations".
Jane: The paper was written by Avi Cooper, Daniel Harari, Tomotake Sasaki, Spandan Madan, Hanspeter Pfister et al. from Fujitsu Research of America and Weizmann Institute of Science and Massachusetts Institute of Technology and Fujitsu Limited and Harvard University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back, everyone. Today we're looking at a paper that's been making the rounds, and it's called "Emergent Neural Network Mechanisms for Generalization to Objects in Novel Orientations." Jane, I gotta say, the title alone got me hooked.
Jane: Oh, absolutely, Tom. And for our listeners, let's break that down. It's about whether AI systems, specifically deep neural networks, can recognize objects when they're shown at angles they've never seen before. Think of it like this — you've only ever seen a coffee mug from the front, but then someone hands it to you tilted sideways. Can you still figure out it's a mug?
Tom: Right, and the paper is asking that exact question, but for machines. And the authors — Avi Cooper, Daniel Harari, Tomotake Sasaki, Spandan Madan, Hanspeter Pfister, Pawan Sinha, and Xavier Boix — they've put together a really clever setup to test this. They trained networks on some objects from every angle, and other objects from just a few angles.
Jane: And what they found is genuinely surprising. The networks actually *can* generalize, but only to certain novel orientations. It's not random. There's a pattern. And that pattern depends on what the network saw during training.
Tom: Exactly. And the really wild part is that this behavior shows up across different network architectures — the convolutional ones, the transformers, even a brain-inspired one called CORnet. So it's not a fluke of one particular design.
Jane: It's like the networks are developing their own internal rules for dealing with the unknown. And that's what we're going to dig into today — what those rules are, why they emerge, and what it means for building smarter AI.
Tom: And we've got Lu, Meng, and Lalam joining us later to really get into the weeds. But first, let's talk about why this matters. Because recognizing objects in new orientations is something humans do effortlessly, but machines have always struggled with it.
Jane: Right, and that's the heart of it. If we want AI that works in the real world — in self-driving cars, in robots, in medical imaging — it has to handle the unexpected. And this paper gives us a window into how these systems might be doing that, or failing to.
Tom: So stick around. We're going to unpack the methods, the results, and what it all means for the future of machine perception.
Summary: Tom: So, Jane, we've set the stage. Now let's get into the actual summary of "Emergent Neural Network Mechanisms for Generalization to Objects in Novel Orientations." What did the researchers actually do?
Jane: Okay, so imagine you have a bunch of toy airplanes. Some of them, the network gets to see from every possible angle — those are called fully-seen instances. Others, it only sees from a narrow range of angles — those are partially-seen. Then they test the network on those partially-seen airplanes at angles it never saw during training.
Tom: And the key finding is that the network doesn't just fail across the board. It actually succeeds at some novel angles and fails at others. And the pattern of success is really specific.
Jane: Right. They found that the network generalizes well to what they call "in-plane" rotations — that's like spinning the image around in 2D, like turning a photo of a plane upside down. But it struggles with "out-of-plane" rotations, where you're actually moving around the object in three dee, like walking around the plane to see it from the side.
Tom: And here's the kicker — the more fully-seen instances you give the network, the better it gets at those in-plane generalizations. But the out-of-plane ones stay hard. So it's not just about having more data; it's about the kind of data diversity you have.
Jane: Exactly. And they built a mathematical model that can predict, given the training angles, which novel angles the network will generalize to. It's got three components — small angle rotations, in-plane rotations, and silhouette similarity. And that model matches the network's behavior really well, with a correlation coefficient above zero point eight.
Tom: That's a strong result. And it means the network's behavior isn't random — it's following rules, even if those rules are implicit and emergent.
Jane: And that's the exciting part for me. The network isn't told these rules. They emerge from the training process itself. And the researchers even found that individual neurons in the network develop responses that are invariant to certain orientations — and that invariance seems to spread from the fully-seen objects to the partially-seen ones.
Tom: So it's like the network is learning a general principle about orientation, not just memorizing specific images.
Jane: Precisely. And that's a big deal for understanding how these systems work under the hood.
Improvements: Tom: Alright, Jane, so we've covered what the paper found. But what does it suggest we should do differently? What are the improvements it points toward?
Jane: Well, Tom, one of the most interesting implications is about how we train networks. The paper suggests that if we want better out-of-distribution generalization, we need to think about the diversity of instances, not just the sheer volume of data.
Tom: Right, because they showed that increasing the number of fully-seen instances — objects seen from every angle — directly improves generalization to novel orientations for the partially-seen ones. But just adding more images of the same objects doesn't help.
Jane: And that's a practical insight for anyone building vision systems. Instead of just scraping more images, you might want to curate your dataset to include more distinct object instances, each seen from many angles.
Tom: The paper also suggests that we might be able to predict *which* novel orientations a network will handle well, just by knowing the training distribution. That's huge for safety-critical applications.
Jane: Absolutely. Imagine you're deploying a robot in a warehouse. You could use this model to figure out which orientations of objects it will struggle with, and either add more training data for those, or design the robot's path to avoid those viewpoints.
Tom: And there's a deeper point here about architecture. The fact that this behavior is consistent across ResNet, DenseNet, ViT, and CORnet suggests it's not something you can just engineer away with a better architecture. It's a fundamental property of how these networks learn.
Jane: Right, which means the improvement has to come from training strategies, data curation, or maybe new ways of incorporating temporal information — because the paper notes that biological systems use motion over time to build invariance, and that might be a path forward for machines too.
Tom: So the takeaway is that we shouldn't just throw more data at the problem. We need to be smarter about what data we use and how we structure the learning process.
Jane: And that's a message that could really shift how the field approaches out-of-distribution robustness.
First Page: Tom: Jane, let's zoom in on the very first page of "Emergent Neural Network Mechanisms for Generalization to Objects in Novel Orientations." There's a lot packed into that opening.
Jane: There really is. The paper opens by framing object recognition in novel orientations as a fundamental challenge for both biological and artificial intelligence. It's something that's been studied for decades, but we still don't fully understand the mechanisms.
Tom: And they make a really important point early on — that previous work has mostly measured average performance across all orientations, which hides the interesting structure. This paper instead looks at per-orientation accuracy, and that's where the patterns show up.
Jane: Right, and they introduce this beautiful visualization — a cube where each cell represents a specific orientation, and the color shows the network's accuracy. You can literally see the "figure eight" pattern of high accuracy around the training orientations.
Tom: That figure eight is the in-plane rotations we talked about. And the fact that it shows up so clearly in the heatmaps is what motivated their predictive model.
Jane: The first page also sets up the connection to neuroscience. They're explicitly borrowing analytical tools from studies of the brain — looking at how individual neurons respond to different orientations and whether they develop invariance.
Tom: And that's a theme throughout the paper — that the mechanisms driving generalization in DNNs might be similar to those in biological vision. They cite work on the inferior temporal cortex and the idea that neurons become tuned to features while becoming invariant to orientation.
Jane: So the first page is really laying out a research program. It's saying: we're going to treat the network like a brain, analyze its neurons, and see if we can find the same kinds of mechanisms that support generalization in humans and monkeys.
Tom: And that's a bold and exciting approach. It bridges two fields that often don't talk to each other.
Jane: It does. And it sets the stage for the deep dive into neural mechanisms that we've been discussing. The first page promises a journey from behavior down to individual neurons, and the rest of the paper delivers on that.
Conclusion: Tom: Well, Jane, we've reached the end of our discussion on "Emergent Neural Network Mechanisms for Generalization to Objects in Novel Orientations." What a ride.
Jane: It really has been. Let's pull it all together. The paper shows that deep neural networks can generalize to novel orientations, but only in specific, predictable ways — mainly in-plane rotations. And that generalization gets stronger as you increase the diversity of fully-seen instances.
Tom: And they built a model that predicts exactly which orientations will be generalizable, with a correlation above zero point eight. That's a powerful tool for anyone working with vision systems.
Jane: But the deeper finding is the neural mechanism. They showed that individual neurons develop orientation invariance for fully-seen objects, and that invariance spreads to partially-seen objects through shared features. It's a dissemination process that mirrors what we see in biological brains.
Tom: And that's the part that could have lasting impact. It suggests that the way networks generalize isn't arbitrary — it's governed by principles we can discover and potentially harness.
Jane: The limitations are real, though. The paper focuses on supervised classification with relatively small networks. And it leaves open the question of why out-of-plane rotations remain so hard — likely because of self-occlusion.
Tom: But even with those limits, this paper gives us a new way to think about out-of-distribution generalization. It's not just about accuracy numbers; it's about understanding the structure of what the network has learned.
Jane: And that understanding could lead to better training strategies, better data curation, and maybe even new architectures that build on these emergent properties.
Tom: So as we say goodbye to this paper, I think the message is clear — the path to robust object recognition lies in understanding the mechanisms, not just chasing benchmarks.
Jane: Well said, Tom. Thanks to everyone who joined us today. We'll be back soon with another paper to break down. Until then, keep looking at the world from new angles.
Tom: See you next time.
Avi Cooper, Daniel Harari, Tomotake Sasaki, Spandan Madan, Hanspeter Pfister, Pawan Sinha, Xavier Boix
Fujitsu Research of America · Weizmann Institute of Science · Massachusetts Institute of Technology · Fujitsu Limited · Harvard University
cs.CV, cs.AI, cs.LG, q-bio.NC, stat.ML
Submitted: 2026-08-11
Updated: 2026-08-12
Journal ref: Transactions on Machine Learning Research (TMLR), 2025, ISSN 2835-8856
Code: https://github.com/avicooper1/OOD_Orientation_
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 69/100
Terminology
Summary
Summary
This paper investigates the generalization capabilities of Deep Neural Networks (DNNs) to recognize objects in orientations that are outside the training data distribution (out-of-distribution, or OoD, orientations). The authors state that The capability of Deep Neural Networks (DNNs) to recognize objects in orientations outside the training data distribution is not well understood,
and they aim to investigate the limitations of DNNs’ generalization capacities by systematically inspecting DNNs’ patterns of success and failure across out-of-distribution (OoD) orientations.
The study employs a learning paradigm inspired by biological brains, where some instances of an object category... are seen from all orientations during training (fully-seen instances), while other instances are only seen in a subset of all orientations (partially-seen instances).
During test time, the networks are evaluated on their ability to classify partially-seen instances in OoD orientations. The authors note this paradigm facilitates analyzing the impact of several key factors that may influence OoD generalization, such as the number of fully-seen instances and the in-distribution orientations of the partially-seen instances.
The experiments use three object categories: Airplanes, Cars, and Shepard&Metzler objects. The first two were curated from ShapeNet, and the third was procedurally generated. The datasets consist of 200k images, with 4k images for each of 50 object instances. The authors vary the number of fully-seen instances (between 10 and 40) while keeping the total number of instances and training examples constant. They also employ four different seed
orientations (in-distribution orientations for partially-seen instances), denoted α̂, β̂, γ̂, and α̂′, which are small ranges of rotation along two axes and the full range along a third axis.
The authors trained several DNN architectures from scratch on a supervised instance classification task, including ResNet18, DenseNet121, ViT-Base, and CORnet-S. They report that We observed the same behavior for all analyses for all the architectures when analyzing the results of airplane datasets,
and that The same conclusions drawn in the main paper may be drawn from the results with other architectures as well.
The paper's first key finding is that DNNs exhibit structured generalization patterns, which the authors visualize using per-orientation accuracy heatmaps. These heatmaps reveal a structured pattern of generalization in the form of increased classification accuracy for OoD (i.e., novel) orientations.
For example, with seed orientations at the center of the heatmap, the network yields the highest accuracy for adjacent orientations (small 3D perturbations) and for orientations depicting in-plane (2D) rotations of the seed. The authors note that an increase in number of fully-seen instances leads to stronger OoD generalization in the aforementioned orientations,
and they quantify this by reporting that OoD accuracy increases as data diversity (i.e., the number of fully-seen instances) increases, under various conditions, including different seed orientations, different image datasets and across datasets.
The second key finding is a predictive model of DNN generalization. The authors formulate a model, denoted fw(θ), with three components: A(θ) captures small angle rotations around θ; E(θ) captures in-plane (2D) rotations; and S(θ) captures object silhouette
projections. They evaluate the model's performance using the Pearson correlation coefficient between the model's predictions and the network's actual accuracy. The results show that in all experiments our model highly predicts the network’s behavior, indicating that indeed the networks generalization patterns for OoD orientations follow the model’s partitioning rules,
with correlation coefficients greater than 0.8
for all experimental controls. The authors also analyze the contribution of each component, finding that The model’s component A(θ) (‘small-angle’ rotations), is the best predictor for the network’s OoD behaviour, for highly articulated objects such as the SM objects,
while The model’s component E(θ) (‘in-plane’ rotations), is a better predictor for non-articulated objects with inherent symmetries.
Furthermore, they find that the increase in OoD generalization is primarily to ‘in-plane’ rotations,
as the predictive power of the in-plane component correlates with increasing OoD accuracy as the number of fully-seen instances increases.
The third key finding concerns the internal neural mechanisms. The authors analyze the activations of neurons in the penultimate layer and define an invariance score
(Eq. 6) that measures the similarity of a neuron's response to different sets of orientations (e.g., in-distribution vs. generalizable OoD). They find a clear correlation between increasing levels of classification accuracy and increasing invariance score for the partially-seen instances,
and that significantly higher invariance scores are measured for the generalizable orientations
compared to non-generalizable ones. The authors also find a tight correlation between the invariance scores of the fully-seen and partially-seen instances,
which suggests that partially-seen invariance emerges due to development of fully-seen invariance.
This leads to their hypothesis that OoD generalization is driven by dissemination of internal network invariances from fully-seen to partially-seen instances through shared features.
The authors conclude that "the network disseminates orientation-invariance of fully-seen instances to partially-seen instances using brain-like mechanism similar to those reported by (Logothetis & Sheinberg, 1996; Poggio & Anselmi, 2016). They explain that
Neurons are feature detectors, and during training neurons are tuned to detect the features of fully-seen objects at multiple orientations — i.e., the neurons become selective to the feature, but invariant to the orientation. Some features that neurons are tuned to are shared between fully-seen and partially-seen instances... Therefore the invariance that develops for features of fully-seen instances are gained 'for free' for partially-seen instances in the same orientations."
The paper acknowledges limitations, noting that This study was limited to a supervised classification setting, and to fairly small DNNs,
and that a key open question is why DNNs disseminate orientation-invariance only to in-plane orientations,
speculating that this may be because orientations that are not in-plane are affected by self-occlusion, which poses a particular challenge for DNNs.
The authors also suggest that biological agents may overcome these difficulties by leveraging the temporal dimension to associate orientations and learn invariant representations.
Improvements for AI systems
Based on the paper, here are specific improvements to AI systems:
Improvement: Instead of reporting only average accuracy, modify the evaluation pipeline to track accuracy across the full 3D orientation space (α, β, γ Euler angles). Implement a visualization tool that generates per-orientation heatmaps (as in Fig. 3) to identify which orientations the model fails on.
What the improved system can do: During deployment, the system can flag specific orientations where accuracy drops below a threshold, enabling targeted data collection or human review for those specific viewpoints rather than treating all failures equally.
Improvement: Add a pre-deployment module that uses the paper's predictive model f w(θ) (combining small-angle, in-plane, and silhouette components) to estimate which out-of-distribution orientations the trained model will handle well before actual deployment.
Improvement: Modify the training data selection strategy to maximize the number of fully-seen instances (instances seen from all orientations) rather than just total image count. The paper shows OoD accuracy increases with more fully-seen instances (Fig. 4a).
Improvement: During training, compute the invariance score (Eq. 6) between in-distribution and predicted-generalizable orientations for penultimate layer neurons. Monitor the correlation between fully-seen and partially-seen instance invariances (Fig. 5c).
Improvement: For objects with known symmetry (e.g., airplanes, cars), explicitly augment training with silhouette-preserving rotations (π rotations around the γ axis) as described in the paper's third model component.
Improvement: Add a post-training classification layer that uses the model's f w(θ) predictions to label each orientation as generalizable
or non-generalizable
(using the 10% threshold method from Section 2.5).
Improvement: Implement the neural analysis from Section 3.3 to identify which penultimate-layer neurons are active for both fully-seen and partially-seen instances at generalizable orientations. Use this to identify shared features that enable dissemination.
Abstract
The capability of Deep Neural Networks (DNNs) to recognize objects in orientations outside the distribution of the training data is not well understood. We present evidence that DNNs are capable of generalizing to objects in novel orientations by disseminating orientation-invariance obtained from familiar objects seen from many viewpoints. This capability strengthens when training the DNN with an increasing number of familiar objects, but only in orientations that involve 2D rotations of familiar orientations. We show that this dissemination is achieved via neurons tuned to common features between familiar and unfamiliar objects. These results implicate brain-like neural mechanisms for generalization.
Sources
- ShapeNet: An Information-Rich 3D Model Repository
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Adam: A Method for Stochastic Optimization
- Not all Views are Created Equal: Analyzing Viewpoint Instabilities in Vision Foundation Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models