EGSD: Event-Grounded Self-Distillation for Streaming Video Understanding
cs.CV
Submitted: 2026-09-29
Updated: 2026-10-05
Terminology
Sources
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?
- PaLM-E: An Embodied Multimodal Language Model
- LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning
- MiniLLM: On-Policy Distillation of Large Language Models
- VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
- Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
- video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLM
- MSJoE: Jointly Evolving MLLM and Sampler for Efficient Long-Form Video Understanding
- InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
- FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding
- Self-Distilled RLVR
- MiMo-VL Technical Report
- StepOPSD: Step-Aware Online Preference Self-Distillation for Agent Reinforcement Learning
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
- OPLD: On-Policy Latent Distillation for Multimodal Reasoning
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models