Compositional Benchmark Synthesis for Hierarchical Human Action Recognition
Farnaz Soleimani, Abdelghani Chibani, Yacine Amirat, Ghazaleh Khodabandelou
cs.AI, cs.CV
Submitted: 2026-08-11
Updated: 2026-08-12
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
Terminology
Summary
Core contribution: The paper proposes "a benchmark-generation and evaluation framework that synthesizes a four-level hierarchical-intention benchmark, spanning actions, activities, low-level intentions (LLIs), and high-level intentions (HLIs), from a flat single-label action corpus while retaining real pre-extracted features at the action level."
Motivation: Large corpora provide isolated, atomically labeled clips without temporal composition, whereas recorded composite-activity corpora offer shallow, domain-narrow, fixed hierarchies.
The authors note that No existing resource pairs a deep intention hierarchy with the scale and subject structure of a large single-label corpus.
Rather than recording a new dataset, they synthesize hierarchical benchmarks from existing flat action corpora.
Ontology: Behavior is organized over four typed levels: "An action is an atomic, single-label clip from the source corpus; an activity is a short composition of actions; a low-level intention (LLI) is a goal-directed unit of activities; and a high-level intention (HLI) is a long-horizon episode of LLIs. The instantiated ontology comprises
8 HLIs, 35 LLIs, 50 activities, and 85 source action classes." Four principles govern the ontology: semantic realism, full coverage (every source action class appears in at least one activity), modality-agnostic activities, and flexible templates. Four situational LLIs (brief drink break, phone check, walking transit, seated rest) are multi-parent, attached to several HLIs, which renders the compositional held-out evaluation non-degenerate.
Synthesis framework: Each episode is a tree of typed nodes rooted at one HLI, expanded recursively: an HLI samples 2–5 LLIs, each LLI samples 2–4 activities, and each activity samples 1–2 source action clips.
Action-level features are pre-extracted once using PoseC3D SlowOnly-R50 backbone (2048-dimensional descriptors for skeleton) and VideoMAE-family extractors (1024-dimensional descriptors for RGB, depth, infrared). Because descriptors are referenced by clip key, episodes remain lightweight and no source data are duplicated.
A transition model encodes preferred temporal orderings as soft constraints, with controlled label-preserving perturbations (activity substitution, order perturbation, filler insertion, each with probability 0.1). All clips in an episode share one performer for subject consistency.
Coverage-aware sampling: The source corpus is markedly skewed: across 60 performers the mean coverage is 44.7 of 85 classes (range 32–85), a clip-availability Gini of 0.566.
The coverage-aware sampler reduces the realized subject-usage Gini to 0.248 (most-used performer in 364 episodes, least in 64).
Eligible-subject pools per HLI range from 40 to 60 performers.
Anti-circularity validity design: A synthesized benchmark whose generator, labels, and evaluation rules coincide can be solved by recovering the generator.
The framework addresses this by keeping sequence-generation rules disjoint from the first-order-logic rules used at evaluation.
Evaluation uses 25 typed first-order-logic rules: "Group X (3 rules) are hierarchy-compatibility rules the generator also enforces, limited to cross-level consistency. Group Y (22 rules) are semantic constraints the generator does not use (temporal precedence, mutual exclusion, cardinality, and co-occurrence implications). The generated ground-truth labels violate Group Y at an intrinsic rate of 0.025 while satisfying Group X at essentially zero, confirming
the Group Y rules are therefore genuinely independent of the generator rather than satisfied by construction."
Benchmark statistics: The instantiated benchmark contains 15,002 episodes with 330,145 forward typed edges.
Episodes average 6.95 actions (median 6, range 4–17), 2.80 LLIs, and 6.30 activities. The action level is long-tailed: the most frequent class supplies 10,996 nodes whereas the rarest dozen supply fewer than 30 each.
A plausibility check found the majority-vote plausibility rate was 0.86 (43∕50) with Fleiss' kappa = 0.725 (substantial agreement).
Reference evaluation: Four baselines from different model families characterize difficulty (not as recognition methods): a hierarchical transformer (HT), sequential transformer (seq-T), order-free bag-of-actions MLP (bag-MLP), and relational graph network (R-GCN). All use ordinary cross-entropy without logical constraints. Results on the cross-subject test split (macro-F1): R-GCN achieves the best HLI performance at 75.4%, followed by HT at 73.2%, seq-T at 70.7%, and bag-MLP at 56.3%. The action level is the lowest of the four for every baseline, consistent with its long tail.
Compositional generalization gap: On the compositional held-out split, HLI macro-F1 falls by 0.13 to 0.17 across all four model families and every seed.
The gaps are: HT 16.7, seq-T 14.3, bag-MLP 12.7, R-GCN 17.2. "Crucially, the graph-aware R-GCN is the strongest recognizer yet shows the largest gap, so a model built for the typed graph structure does not close it; the gap is a structural property of the benchmark rather than a capacity artifact. The drop is
far larger on the 1,854 episodes containing a held-out LLI→HLI pairing in novel context (0.47 on average) than on the 321 episodes held out by ordering pattern alone (0.30)."
Logic-violation rates: Logic-free baselines violate the held-out semantic (Group Y) rules above their 0.025 intrinsic data rate
(rates of 0.028 to 0.042 across baselines), confirming the rules are not satisfied by construction and the benchmark resists circular shortcuts.
Order-destroying control: Destroying temporal order (permuting siblings while holding every label, performer, clip, length, and membership fixed) changes macro-F1 by less than one standard deviation at every level and baseline.
The control functions as a generator-consistency check rather than evidence about temporal order; order flexibility at the intention level is a designed property.
Contributions: (i) a regenerable synthesis protocol preserving real action-level features while adding higher-level structure; (ii) a coverage-aware subject sampler reducing performer imbalance; (iii) an anti-circularity validation design; (iv) a compositional held-out split at the LLI–HLI association level; and (v) an empirical protocol combining recognition, logical-violation rates, and order-destroying controls.
Limitations acknowledged: The intentions are imposed by construction rather than observed
; the order-destroying control is conservative due to unshuffleable single-child parents; the action level is long-tailed; the framework transfers to other corpora only if they have per-performer metadata, sufficient per-subject class coverage, and fine action granularity.
Data availability: "NTU RGB+D 120 licensing prohibits redistribution of source data or extracted features. The release therefore comprises the ontology, synthesis code, FOL rule set, and analysis scripts; the benchmark can be regenerated by any user with NTU RGB+D 120 access."
Key conclusion: The novelty is not any isolated component but their integration into a reproducible protocol that makes hierarchy, coverage, circularity, and compositional generalization measurable in one setting.
The intended use is controlled evaluation of hierarchical and compositional action-recognition models, not deployment for inferring real human intentions in sensitive settings.
Improvements for AI systems
Improvements to AI systems:
-
Hierarchical intention-aware action recognition: Build a model that explicitly predicts all four levels (action → activity → LLI → HLI) jointly, using the typed graph structure (episode trees) as an inductive bias. The improved system can recognize long-horizon human behavior (e.g., a cooking session, a workout routine) from raw video, not just isolated actions, and can output a structured explanation of why a high-level intention is inferred (via intermediate LLI/activity nodes).
-
Compositional generalization via disentangled hierarchy learning: Train a model with separate modules for (a) action-level feature extraction, (b) activity composition rules, and (c) LLI–HLI association patterns, using the benchmark’s compositional held-out split (novel LLI→HLI pairings) as a training curriculum. The improved system can recognize unseen combinations of sub-goals and intentions (e.g., a new sequence of activities that forms a previously unseen high-level goal) without retraining, by recombining learned components rather than memorizing full episodes.
-
Coverage-aware subject balancing for fair evaluation: Use the coverage-aware sampler’s Gini-reduction technique (from 0.566 to 0.248) as a data-augmentation or re-weighting strategy in any multi-subject action recognition pipeline. The improved system becomes robust to performer bias—it no longer overfits to a few high-coverage subjects and generalizes better to new, sparsely-observed performers.
-
Anti-circularity validation for synthetic data: Adopt the paper’s two-rule-set design (generator-enforced vs. evaluation-only FOL rules) as a standard practice for any AI system trained on synthetic or procedurally-generated data. The improved system can be validated without the risk of “shortcut solving” (e.g., recovering the generator’s random seed), ensuring that reported performance reflects genuine semantic understanding rather than data-generation artifacts.
-
Logic-constrained decoding with violation monitoring: Integrate the 25 typed first-order-logic rules (temporal precedence, mutual exclusion, cardinality, co-occurrence) as soft constraints during inference, and track logic-violation rates as a secondary metric alongside accuracy. The improved system can produce predictions that are semantically consistent (e.g., never predicting two mutually exclusive activities in the same episode) and can flag its own uncertainty when a prediction violates a known rule, enabling safer deployment in human-robot interaction or surveillance.
-
Order-flexible intention recognition: Leverage the order-destroying control finding (temporal order is not critical at the intention level) to design a model that is permutation-invariant at the LLI/HLI level but order-sensitive at the action/activity level. The improved system can recognize high-level goals even when the exact sequence of sub-actions is shuffled (e.g., a person making coffee but pouring water before grinding beans), while still distinguishing fine-grained action order where it matters.
-
Long-tail action handling via hierarchical priors: Use the benchmark’s long-tailed action distribution (most frequent class 10,996 nodes vs. rarest <30) to train a hierarchical prior that maps rare actions to higher-level activities/LLIs as fallback. The improved system can recognize rare or unseen actions by inferring them from the surrounding activity context (e.g., identifying an obscure yoga pose from the sequence of preceding and following poses within a known LLI like “stretching routine”).
-
Cross-modal hierarchical fusion: Since the framework provides pre-extracted features from PoseC3D (skeleton) and VideoMAE (RGB/depth/infrared), build a multi-modal hierarchical fusion model that learns to weight modalities differently per level (e.g., skeleton for actions, RGB for activities, depth for LLIs). The improved system can handle modality-agnostic activities (as the ontology requires) and degrade gracefully when one sensor stream is missing or noisy.
-
Regenerable benchmark-driven continual learning: Use the synthesis protocol to generate unlimited, controlled variations of episodes (by perturbing activities, order, fillers) for continual learning. The improved system can be trained incrementally on new HLI types or new performer distributions without catastrophic forgetting, because the benchmark can be regenerated on-demand with specified difficulty and coverage properties.
-
Explainable intention inference for human-robot collaboration: Combine the hierarchical predictions with the FOL rule set to generate natural-language explanations (e.g., “I inferred the HLI ‘preparing a meal’ because I saw the LLI ‘chopping vegetables’ followed by ‘cooking,’ and these co-occur in 90% of training episodes”). The improved system can justify its intention predictions to human users, making it suitable for assistive robotics or ambient intelligence where trust and transparency are critical.
Sources
- Vision and Intention Boost Large Language Model in Long-Term Action Anticipation
- What Has Been Lost with Synthetic Evaluation?
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection