Temporally-Grounded Language Generation: Towards Real-Time Vision-Language Models
cs.CV, cs.AI
Submitted: 2025-05-16
Updated: 2026-08-27
Terminology
Sources
- Can Foundation Models Watch, Talk and Guide You Step by Step to Make a Cake?
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- The Kinetics Human Action Video Dataset
- StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- LVBench: An Extreme Long Video Understanding Benchmark
- Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge
- SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding
- Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
- MLVU: Benchmarking Multi-task Long Video Understanding
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models