Reachability Is Not Generalization: Understanding Verb--Noun Decomposition in Assembly Action Recognition
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Reachability Is Not Generalization".
Jane: Assembly actions are compositional, combining a manipulation (verb) with a part or tool (noun), yet traditional atomic classifiers assign zero probability to unseen combinations by construction.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, Jane, we're diving into this paper called "Reachability Is Not Generalization: Understanding Verb--Noun Decomposition in Assembly Action Recognition." It sounds like it’s tackling a really fundamental tension between being able to *reach* a new action and actually *generalizing* to it.
Jane: Exactly, Tom. The title itself suggests that just because a model can theoretically construct an unseen action by combining a verb and a noun doesn't mean it will perform well on that unseen combination in the real world.
Lu: It’s about exposing how decomposition, which lets you reach anything outside the training set, doesn't automatically translate into good performance across all novel combinations because of the way we train things.
Meng: From an engineering standpoint, this sounds like it's digging into why our current decomposition methods aren't robust enough when we see something completely new that was never in the training data.
Lalam: It suggests that just being able to combine known parts isn't the same as truly understanding how those parts interact in a novel way, which is a really important distinction for how we build our AI systems.
The paper's summary: Tom: So, what does the paper actually say about verb-noun decomposition? It seems to be focusing on the fact that atomic classifiers give zero probability to anything unseen, but decomposition lets you get past that ceiling.
Jane: Right, and then it immediately sets up a challenge: while this decomposition makes unseen compositions reachable—meaning they can be predicted—the paper shows it doesn't generalize perfectly beyond what was seen during training.
Lu: The core finding is that the success of this decomposition isn't just about the recombination itself; it’s strongly tied to the co-occurrence structure present in the training data.
Meng: So, if I understand correctly, even if we can mathematically form a new action, its success depends on how frequently those specific verb and noun pairs appeared together in our training set.
Lalam: That means the model is essentially relying on patterns it’s already seen rather than learning a deep compositional rule for entirely novel pairings.
The paper's improvements: Tom: The paper points out that the way we approach this needs to be more thoughtful, suggesting improvements based on analyzing where things go wrong during unseen testing.
Jane: They suggest looking at specific conditions required for decomposition to actually generalize, like visual separability and statistical factorization of the joint action distribution.
Lu: They test several formulations—atomic prediction, shared encoder multi-head architectures, and independent recombination—to see which one behaves the best under compositional shift.
Meng: It seems they pinpoint component entanglement as a big issue when we use shared training methods, where the model gets overly reliant on co-occurrence patterns that don't transfer well to new compositions.
Lalam: So the paper suggests we need better ways to separate those components or rethink how we factorize the joint probability so it actually learns compositional reasoning instead of just structured retrieval.
Conclusion: Tom: To wrap up, this paper really hammers home that decomposition gets us into the action space, but generalization is still constrained by how much structure is in our training data.
Jane: It points out that performance relies heavily on primitive support and vocabulary asymmetry, meaning larger vocabulary components often dictate where errors concentrate when we test unseen actions.
Lu: The implication for future work seems to be that we need to focus less on just adding more layers or changing the objective, and more on finding better ways to factorize the training prior itself.
Meng: From a practical side, this means if we want better zero-shot performance, we might need to be smarter about allocating resources toward the components that are inherently harder for the AI to learn on its own.
Lalam: Ultimately, this work suggests that moving forward, we have to focus on genuine compositional reasoning rather than just structured reuse of observed co-occurrences when trying to make our systems truly flexible.
Changyi Li, Yu Xiao
Aalto University
cs.CV
Submitted: 2026-09-04
Updated: 2026-09-04
Comments: Accepted by BMVC 2026
Code: https://github.com/hisalaheiyo/assembly
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 89/100
The gist: Assembly actions are compositional, combining a manipulation (verb) with a part or tool (noun), yet traditional atomic classifiers assign zero probability to unseen combinations by construction.
Key concepts
- Visual Separability
- This concept assumes that the visual information needed to recognize one part of an assembly action can be represented without being too closely tied to the other part. If this holds, a component's score should remain stable regardless of what other parts are present, allowing for reliable decomposition.
- Statistical Factorization
- This requires that the joint probability of a verb and noun action is accurately approximated by multiplying their individual probabilities. If this holds, the system can successfully predict an unseen combination by combining the known predictions for each separate component.
- Component Entanglement
- This occurs when training on joint actions forces components to rely too much on co-occurrence patterns rather than independent features. This entanglement means that while seen actions work well, these learned patterns transfer poorly to novel, unseen compositions.
Terminology
Summary
Assembly actions are compositional, combining a manipulation (verb) with a part or tool (noun), yet traditional atomic classifiers assign zero probability to unseen combinations by construction. The central finding is that while verb–noun decomposition makes unseen compositions reachable, its generalization extends only partially beyond it, revealing that performance is governed by training-data structure rather than unconstrained recombination.
The Gist
Decomposition makes unseen compositions reachable but not necessarily generalizable because success remains strongly conditioned on primitive support and vocabulary asymmetry.
Testing the Assumptions of Decomposition
The paper systematically tests several implicit assumptions required for decomposition to generalize, using three representative formulations: atomic prediction (the unreachable baseline), shared-encoder multi-head architecture (the field standard), and independent recombination of separately trained component models. The analysis proceeds from phenomenon to localization, mechanism, and intervention across four research questions.
-
Whether unseen-composition success reflects genuine recombination or merely interpolation within densely observed regions of the compositional space (RQ1).
-
Where failures occur and whether the dominant source of error can be predicted before training (RQ2).
-
Why the field-standard shared encoder generalizes worse than independent recombination (RQ3).
-
Whether these limitations can be alleviated through alternative factorization granularities, architectures, or objectives (RQ4).
What Must Hold for Decomposition to Generalize?
Decomposition succeeds only if several conditions hold. The first is visual separability: the assumption that the visual evidence required to recognize one component can be represented without being inseparably tied to the other,
meaning a component head's score should remain stable across different noun contexts. Failure of this leads to shared representations introducing a shortcut, quantified by the metric: ∆ent = dropjoint seen→unseen −dropsolo seen→unseen.
The second condition is statistical factorization, requiring that component predictions provide a useful approximation to the joint action distribution: P(v,n x) ≈ Pˆ(v x)Pˆ(n x).
This is tested by comparing unseen accuracy with the amount of training support available for its constituent primitives, measured by ρ = corrAcc(vu,nu), f(vu)· f(nu),
where a high ρ suggests success is driven by interpolation near densely observed regions.
Which Component Becomes the Bottleneck?
The analysis localized failures to specific components based on error decomposition. The hypothesis is that unseen errors systematically concentrate on the larger-vocabulary component.
This asymmetry is predictable based on dataset structure, not component type; for instance, in IMPACT, where the verb vocabulary exceeds its noun vocabulary, the bottleneck reverses with it across all five views.
This is quantified by comparing verb-only and object-only errors. Furthermore, shared training introduces component entanglement,
which encourages reliance on co-occurrence patterns that transfer poorly to novel compositions. This entanglement is measured by the excess unseen drop: a large ∆ent value indicates that shared training provides little benefit on seen compositions while substantially increasing the unseen generalization gap.
Why Shared Training Underperforms (RQ3)
The gap between shared and independent decomposition is attributed to component entanglement rather than capacity alone. While independent recombination generally generalizes better, the shared formulation amplifies coverage limitations. This is evidenced by comparing the seen margin
(joint minus independent accuracy on seen compositions) against the excess unseen drop ∆ent.
The fact that seen margins remain close to zero while ∆ent is consistently large suggests that joint training encourages reliance on co-occurrence structure that benefits familiar combinations but transfers poorly to novel ones. This implies that decomposition, in its current shared form, behaves more like structured retrieval than genuine compositional reasoning.
Can the Limitation Be Engineered Away? (RQ4)
The final question evaluates interventions: factorization granularity, architecture, and training objectives. The results show that finer factorizations do not improve compositional generalization,
as the standard two-component decomposition (k=2) provides the best unseen performance. Increasing granularity beyond k=2 primarily improves seen accuracy while degrading transfer to unseen compositions. Architecture has the largest effect; Independent recombination consistently outperforms the shared formulation, supporting the entanglement mechanism identified in RQ3.
Crucially, none of these interventions change the asymmetry identified in RQ2: the largervocabulary component remains the dominant source of unseen error across architectures.
The remaining headroom is largely a ranking problem rather than a representation problem. The paper suggests that subtracting the factorized training prior at inference, without retraining, improves unseen accuracy by up to +4.2pp on IMPACT.
Discussion: Implications for Compositional Recognition
The overall conclusion is that decomposition improves reachability but not generalization because performance is shaped by training co-occurrence structure. The diagnosis links primitive support, vocabulary asymmetry, and component entanglement to these failures.
This suggests current methods rely more on "structured reuse of observed co-occurrences than on genuine compositional reasoning.
Improvements for AI systems
Here are the specific improvements to AI systems derived from this research, categorized by the mechanism of improvement:
) Primitive Support & Vocabulary Asymmetry Allocation:
Allocate model capacity (parameters, training data augmentation budget) asymmetrically toward the larger-vocabulary component (verb or noun, depending on which is larger in a given dataset). This targets the known bottleneck identified in RQ2 and Table 3.
The improved system will exhibit higher unseen accuracy by focusing more learning resources on the component that is inherently harder to learn or cover, rather than treating both components with equal weight.
) Architectural Choice Over Objective Design:
Prioritize architectural designs that enforce component separation (e.g., independent recombination of trained sub-models) over simply adding compositional objectives (like C2C).
The improved system will be more robust to compositional shift because it preserves the necessary component separation, mitigating the performance degradation caused by
component entanglementidentified in RQ3.
) Inference-Time Prior Subtraction:
Implement a technique where, at inference time, the factorized training prior is subtracted from the final recombined score without requiring retraining.
The improved system will achieve an immediate boost in unseen accuracy (up to +4.2pp on IMPACT) by correcting the bias introduced during factorization, effectively moving it closer to genuine compositional reasoning rather than structured retrieval.
) Enhanced Recombination Strategy (Addressing Interpolation vs. Composition):
When dealing with unseen compositions, the system should employ a strategy that explicitly tests for genuine recombination rather than relying solely on interpolation within densely observed regions of the training grid.
The improved system will be more sensitive to true compositional structure, leading to better performance on low-support/novel combinations where current models default to interpolation.
) Granularity Optimization (Finding the Pareto Frontier):
Instead of sticking to a standard two-way decomposition, the system should dynamically select a factorization granularity (number of heads, k) that maximizes unseen accuracy while minimizing degradation in seen accuracy.
The improved system will adapt its internal representation strategy based on the specific dataset structure, selecting the optimal level of verb/noun/tool decomposition to maximize generalization capability for unseen actions.
This overall improved AI system can perform:
-
Maximize performance on assembly actions that have never been explicitly seen during training (i.e., improve compositional zero-shot learning).
-
Be more robust to changes in the underlying data distribution (compositional shift) by correctly identifying and mitigating the sources of failure (primitive support, vocabulary imbalance, and entanglement).
-
Operate as a more sophisticated compositional reasoning engine rather than a simple structured retrieval system.
Sources
- IMPACT: A Dataset for Multi-Granularity Human Procedural Action Understanding in Industrial Assembly
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models