RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Multi-Trajectory Semantic Retrieval
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Multi-Trajectory Semantic Retrieval".
Tom: Referring Atomic Video Action Recognition (RAVAR) aims to recognize fine-grained, atomic-level actions of a specific person conditioned on natural language descriptions,
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So, looking at the title and the authors of "RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Multi-Trajectory Semantic Retrieval," it’s clear they are tackling a specific problem in video understanding where we need to link language directly to fine-grained human actions for specific people.
Jane: I agree, Tom; the approach seems to be about making the connection between what's said and what's happening on screen much more precise than previous methods, especially when dealing with multiple individuals in a single video.
Lu: The authors are taking existing concepts and weaving them together using this new multi-trajectory semantic retrieval mechanism, which suggests they’ve found a way to better align visual information with linguistic cues across different levels of detail.
Meng: I see the complexity immediately; it’s not just one single model doing everything, but a framework that uses multiple pathways for retrieving and aggregating meaning from the text and the video simultaneously.
Lalam: This level of fine-grained control over action recognition really matters because if we can pinpoint an exact interaction or movement, it opens up possibilities for much more nuanced AI applications in areas like assistive technologies or detailed behavioral analysis.
The paper's summary: Tom: The core summary of this paper explains that RefAtomNet++ introduces a novel framework that advances cross-modal token aggregation by using a multi-hierarchical semantic-aligned cross-attention mechanism combined with multi-trajectory Mamba modeling at the partial keyword, scene-attribute, and holistic sentence levels.
Jane: That sounds incredibly intricate; essentially, they are breaking down the language—from the whole sentence to just keywords and even objects in the scene—and using those different pieces of text to guide how the model pulls relevant visual information from the video at different times.
Lu: The way they structure these three hierarchies of semantics, ranging from global context to very specific entity cues, seems like a smart way to ensure that both broad scene understanding and tiny details are being considered in the final prediction.
Meng: The idea of using Mamba modeling on these retrieved visual trajectories to capture long-range temporal dependencies is interesting because it addresses one of the usual weaknesses in sequence modeling when dealing with complex video dynamics.
Lalam: Capturing those continuous dynamics over time, guided by multiple semantic inputs at once, suggests an AI that can track a person’s subtle actions across a long sequence much more reliably than systems that only look at one level of context.
The paper's improvements: Tom: Regarding the improvements they propose, the main thing is this multi-hierarchical semantic retrieval, which involves deriving tokens from three levels: holistic sentence, partial keyword, and scene attribute.
Jane: That means they aren't just relying on one type of textual input; they are layering different types of meaning—the big picture sentence, the specific keywords like "red T-shirt," and the physical objects present—to guide the visual token selection.
Lu: The integration of Mamba modeling for aggregating these retrieved tokens into semantically aligned visual trajectories is a key methodological advancement because it creates these targeted paths for each semantic level to follow through time.
Meng: I’m looking at the scene-attribute level using a pretrained object detector like DETR, which means they are grounding the language in real visual objects present in the frame, which should make the localization part much stronger.
Lalam: When you combine that multi-trajectory Mamba aggregation with their multi-hierarchical cross-attention module, it implies a system that can dynamically weigh how much importance to give to each semantic level as it processes the video sequence, which is really powerful for handling messy real-world data.
Conclusion: Tom: So, to wrap up on "RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Multi-Trajectory Semantic Retrieval," the authors have successfully extended their dataset to RefAVA++ with over two point nine five million frames and over seventy-five thousand annotated persons, achieving results like forty-three point seven one percent mIOU and seventy-one point two seven percent AUROC on the validation sets.
Jane: It sounds like they’ve demonstrated that this multi-trajectory approach leads to tangible improvements, specifically showing gains of over six percent in mIOU compared to prior models on the RefAVA++ set.
Lu: What this means for future work is that we can expect more complex language queries and more nuanced action recognition capabilities because they have established a robust way to handle the temporal dependencies across those different semantic layers.
Meng: From a practical perspective, while the complexity is high, achieving these scores on such a large dataset suggests that the framework has the potential to be very effective in real-world applications where precise individual tracking and action identification are necessary.
Lalam: This work really pushes the envelope for how we can make AI systems understand human behavior not just in broad strokes, but down to the specific atomic interactions, which is a significant step toward truly interactive and intelligent video analysis tools.
Institute for Anthropomatics and Robotics · Karlsruhe Institute of Technology · RISE Research Institutes of Sweden · KTH Royal Institute of Technology · School of Artificial Intelligence and Robotics · National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University
cs.CV, cs.MM, cs.RO, eess.IV
Submitted: 2025-10-18
Updated: 2026-10-03
Comments: Extended version of ECCV 2024 paper arXiv:2407.01872. The dataset and code are available at https://github.com/KPeng9510/refAVA2
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 82/100
The gist: Referring Atomic Video Action Recognition (RAVAR) aims to recognize fine-grained, atomic-level actions of a specific person conditioned on natural language descriptions, which is critical for
Key concepts
- RefAtomNet++
- A novel framework designed for referring atomic video action recognition. It uses a multi-hierarchical semantic alignment approach to fuse visual and language tokens effectively, allowing the model to understand actions based on detailed textual prompts.
- Multi-Hierarchical Semantic Retrieval
- This process derives semantic tokens from three distinct levels: the holistic sentence level (full text), the partial keyword level (masked keywords), and the scene attribute level (object detection). These tokens guide which visual features to retrieve for alignment.
- Multi-Trajectory Mamba Modeling
- The model uses Mamba to aggregate semantically aligned visual trajectories. For each semantic token, it selects the nearest visual token at every time step, creating trajectories that capture long-range temporal dependencies efficiently.
Terminology
Summary
Referring Atomic Video Action Recognition (RAVAR) aims to recognize fine-grained, atomic-level actions of a specific person conditioned on natural language descriptions, which is critical for interactive human action analysis in complex multi-person scenarios. The core contribution of this work is the introduction of RefAtomNet++, a novel framework that advances cross-modal token aggregation through a multi-hierarchical semantic-aligned cross-attention mechanism combined with multi-trajectory Mamba modeling at the partial-keyword, scene-attribute, and holistic-sentence levels to achieve state-of-the-art results.
The Dataset and Benchmark
The work introduces RefAVA++, an enhanced dataset for RAVAR which comprises over 2.95 million frames and more than 75.1k annotated persons in total, significantly extending the previous RefAVA dataset. This large scale enables robust training and evaluation for fine-grained, language-guided action recognition in complex scenarios. The researchers established a comprehensive benchmark by evaluating 15 strong baselines spanning action recognition and localization, video question answering, and vision-language retrieval on this extended dataset.
The Model Architecture: RefAtomNet++
RefAtomNet++ is built upon the vision-language model BLIPv2 [39] and incorporates a novel multi-trajectory semantic-retrieval Mamba to aggregate visual and linguistic tokens, alongside a multihierarchical semantic-aligned cross-attention module. The framework considers three hierarchies of semantics: the holistic-sentence level, the partial-keyword level, and the scene-attribute level.
Multi-Hierarchical Semantic Retrieval
The model derives semantic tokens from these three levels to guide token retrieval:
-
Holistic Sentence Level: Semantic tokens are derived by encoding the complete sentence with the BLIPv2 textual encoder [39].
-
Partial Keyword Level: Stop words are masked, and remaining keywords are passed through an LLM-based textual encoder to obtain fine-grained keyword-level embeddings (partial semantic tokens).
-
Scene Attribute Level: A pretrained object detector, DETR [40], identifies objects in the key frame to extract bounding box coordinates and categorical embeddings. These are combined with spatial information for scene-attribute semantic tokens.
Multi-Trajectory Mamba Modeling
For each semantic hierarchy, the model retrieves relevant visual tokens by selecting the nearest visual token at each time step based on the corresponding semantic token (e.g., partial-keyword or scene-attribute). This forms semantically aligned visual trajectories. These trajectories are then aggregated using Mamba [41], which provides a continuous and memory-efficient mechanism to aggregate multiple semantic trajectories by capturing long-range temporal dependencies with stable dynamic transitions.
This results in keyword-semantic-retrieval tokens, denoted as tKW.
Multi-Hierarchical Semantic-Aligned CrossAttention (MHS-CA)
The aggregated multi-trajectory tokens are fused with the original spatio-temporal visual features via a two-branch architecture employing MHS-CA. For each semantic hierarchy (holistic sentence, keyword, scene attribute), a learnable query embedding is concatenated with the text-derived query before performing cross-attention. This mechanism enables multilevel, semantic-aligned visual token aggregation
by allowing queries from different semantic hierarchies to interact with the retrieved tokens and original features across both spatial and temporal perspectives.
Final Prediction
The aggregated temporal representation (zT) is forwarded to a bounding box regression head (fregT) and an action recognition head (fclsT), while the spatial representation (zS) is processed similarly. The final predictions are obtained by fusing the outputs from both branches: bˆ = AVGbˆT + bˆS, yˆ = AVGyˆT + yˆS.
This joint modeling of spatial and temporal branches ensures that accurate localization depends on spatial precision as well as temporal consistency, while robust action recognition relies on temporal dynamics contextualized by spatial layouts.
Key Results
RefAtomNet++ establishes new state-of-the-art results, achieving 43.71%/42.52% (mIOU), 56.83%/59.81% (mAP), and 71.27%/75.72% (AUROC) on the validation/test sets of RefAVA, and 39.12%/38.58%, 58.24%/58.84%, and 72.58%/73.28% on RefAVA++. The model outperforms RefAtomNet by +6.10% mIOU, +2.29% mAP, and +1.77% AUROC on the validation/test sets of RefAVA++, demonstrating superior performance in complex multi-person scenarios compared to prior models. The optimal configuration for learnable queries was found to be 6 queries, achieving the highest overall scores.
Improvements for AI systems
Here are the specific improvements that can be made to existing AI systems based on the proposed RefAtomNet++ framework, and what those improved systems will be capable of doing:
-
The proposed system will move beyond simple object detection or general action recognition by performing
Referring Atomic Video Action Recognition (RAVAR).
-
It will enable the system to recognize fine-grained, atomic-level actions (e.g.,
Sit,
Carry/Hold,
Talk to
) of a specific person of interest conditioned on natural language descriptions, rather than just detecting general actions or localizing bounding boxes for all people in a scene. -
The improved AI system will perform robust localization and fine-grained action prediction simultaneously, directly from the textual reference without requiring manual cropping or post-hoc selection of Regions of Interest (ROIs).
-
The system will achieve high accuracy on complex, cluttered scenes by utilizing a novel multi-hierarchical semantic-aligned cross-attention mechanism (MHS-CA) that fuses information from three distinct semantic levels:
-
The system will dynamically aggregate visual tokens based on:
-
a) Holistic sentence semantics (global context),
-
b) Partial keyword semantics (localized entity cues like
red T-shirt
), and -
c) Scene-attribute semantics (contextual objects like
phone
orchair
). -
The system will employ a multi-trajectory semantic-retrieval Mamba modeling component to capture long-range temporal dependencies smoothly, which is essential for understanding the continuous dynamics of fine-grained human actions over time.
-
By jointly modeling spatial and temporal branches with separate regression and classification heads, the system will accurately predict both:
-
The precise 2D bounding box location of the referred individual (spatial localization), and
-
The specific atomic action category (e.g., Person Movement, Object Manipulation) they are performing at each timestep (fine-grained action recognition).
-
The system will demonstrate enhanced robustness to linguistic variations through test-time rephrasing capabilities, allowing it to maintain high performance even when the natural language query is slightly modified by a large language model (like ChatGPT).
-
The improved AI system will be highly efficient, achieving state-of-the-art performance while maintaining computational efficiency compared to prior models, specifically through parameter reduction and optimized state-space modeling.
Sources
- InternVideo: General Video Foundation Models via Generative and Discriminative Learning
- End-to-End Spatio-Temporal Action Localisation with Video Transformers
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Weakly-Supervised Spatio-Temporally Grounding Natural Sentence in Video
- Video ChatCaptioner: Towards Enriched Spatiotemporal Descriptions
- EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering
- Revealing Single Frame Bias for Video-and-Language Learning
- VideoChat: Chat-Centric Video Understanding
- Microsoft COCO Captions: Data Collection and Evaluation Server
- OPT: Open Pre-trained Transformer Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models