ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding
cs.CV, cs.CL
Submitted: 2026-09-02
Updated: 2026-09-30
Code: https://github.com/CURRENTF/ShallowStream
Terminology
Sources
- StreamReady: Learning What to Answer and When in Long Streaming Videos
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanism
- StreamingTOM: Streaming Token Compression for Efficient Video Understanding
- StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
- Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
- What Should a Streaming Video Model Remember?
- A Token is Worth over 1,000 Tokens: Efficient Knowledge Distillation through Low-Rank Clone
- Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
- ResidualKV: Residual-Based KV Cache Compression for Efficient Long-Context Inference
- InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding
- LLaVA-OneVision: Easy Visual Task Transfer
- SCBench: A KV Cache-Centric Analysis of Long-Context Methods
- OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning
- StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
- Vista: Scene-Aware Optimization for Streaming Video Question Answering under Post-Hoc Queries
- LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
- Streaming Long Video Understanding with Large Language Models
- LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models