Language-Augmented Video Action Anticipation: Design Fundamentals, Benchmarks, and Open Challenges
cs.CV
Submitted: 2026-09-15
Updated: 2026-09-30
Terminology
Sources
- A Short Note on the Kinetics-700-2020 Human Action Dataset
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- A Survey on Deep Learning Techniques for Action Anticipation
- Human Action Anticipation: A Survey
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
- On the Efficacy of Text-Based Input Modalities for Action Anticipation
- Vision and Intention Boost Large Language Model in Long-Term Action Anticipation
- VideoLLM: Modeling Video Sequence with Large Language Models
- LEAP: LLM-Generation of Egocentric Action Programs
- TR-LLM: Integrating Trajectory Data for Scene-Aware LLM-Based Human Action Prediction
- Anticipate & Act : Integrating LLMs and Classical Planning for Efficient Task Execution in Household Environments
- Hierarchical Planning with Latent World Models
- StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition
- Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
- Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
- SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
- Video Token Sparsification for Efficient Multimodal LLMs in Autonomous Driving
- PruneVid: Visual Token Pruning for Efficient Video Large Language Models
- AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models