Disassociating performance from compositional feature learning
cs.LG, cs.AI
Submitted: 2025-05-14
Updated: 2026-09-21
Comments: Accepted by IEEE Transactions of Cognitive and Development Systems
License: http://creativecommons.org/licenses/by/4.0/
The gist: Out-of-distribution (OOD) generalisation is considered a hallmark of human and animal intelligence.
Terminology
Abstract
Out-of-distribution (OOD) generalisation is considered a hallmark of human and animal intelligence. To achieve OOD through composition, a system must discover the environment-invariant properties of experienced input-output mappings and transfer them to novel inputs. This can be realised if an intelligent system can identify appropriate, task-invariant, and composable input features, as well as the composition methods, thus allowing it to act based not on the interpolation between learnt data points but on the task-invariant composition of those features. We propose that in order to confirm that an algorithm does indeed learn compositional structures from data, it is not enough to just test on an OOD setup, but one also needs to confirm that the features identified are indeed compositional. We showcase this by exploring two tasks with clearly defined OOD metrics that are not OOD solvable by three commonly used neural networks: a Multi-Layer Perceptron (MLP), a Convolutional Neural Network (CNN), and a Transformer. In addition, we develop two novel network architectures imbued with biases that allow them to be successful in OOD scenarios. We show that even with correct biases and almost perfect OOD performance, an algorithm can still fail to learn the correct features for compositional generalisation.
Sources
- On the Measure of Intelligence
- DreamCoder: Growing generalizable, interpretable knowledge with wake-sleep Bayesian program learning
- A Complexity-Based Theory of Compositionality
- Neurosymbolic AI: The 3rd Wave
- Inductive Biases for Deep Learning of Higher-Level Cognition
- Axial Attention in Multidimensional Transformers
- Amortizing intractable inference in large language models
- Compositionality decomposed: how do neural networks generalise?
- Contrastive Learning of Structured World Models
- A Survey on Compositional Generalization in Applications
- From Frege to chatGPT: Compositionality in language, cognition, and deep neural networks
- A Survey on Compositional Learning of AI Models: Theoretical and Experimental Practices
- Attention Is All You Need
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks