Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies
cs.CV, cs.AI, cs.CL, cs.IR, cs.LG
Submitted: 2026-09-29
Updated: 2026-09-29
Terminology
Sources
- BoT-SORT: Robust Associations Multi-Pedestrian Tracking
- Towards Multimodal Lifelong Understanding: A Dataset and Agentic Baseline
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- Bridging Modalities, Spanning Time: Structured Memory for Ultra-Long Agentic Video Reasoning
- Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
- EgoGraph: Temporal Knowledge Graph for Egocentric Video Understanding
- InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
- VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models