TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding
cs.CL, cs.CV
Submitted: 2026-09-25
Updated: 2026-09-25
Code: https://github.com/om-ai-lab/trace-bench
Terminology
Sources
- VideoLLM-online: Online Video Large Language Model for Streaming Video
- LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale
- MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- Online Video Understanding: OVBench and VideoChat-Online
- IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams
- OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
- StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
- Thinking in Streaming Video
- AURA: Always-On Understanding and Real-Time Assistance via Video Streams
- OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants
- EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding
- RIVER: A Real-Time Interaction Benchmark for Video LLMs
- MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention
- S-EMBER: A Large-Scale Benchmark for Streaming Egocentric Memory Retrieval
- ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models
- OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts
- LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
- JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering