APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants
cs.CV
Submitted: 2026-09-29
Updated: 2026-09-29
Terminology
Sources
- Realtime-Venus: A full-duplex interaction system with asynchronous delegation
- Qwen3-VL Technical Report
- Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity
- StreamTTT: Reconciling Real-Time Perception and Long-Term Memory in Streaming VLMs
- ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding
- What Should a Streaming Video Model Remember?
- Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
- OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
- StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
- An Efficient Streaming Video Understanding Framework with Agentic Control
- PhoStream: Benchmarking Real-World Streaming for Omnimodal Assistants in Mobile Scenarios
- HD-EPIC: A Highly-Detailed Egocentric Video Dataset
- EgoPro-Bench: Benchmarking Personalized Proactive Interaction in Egocentric Video Streams
- A Simple Baseline for Streaming Video Understanding
- Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos
- video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLM
- Think While Watching: Online Streaming Segment-Level Memory for Multi-Turn Video Reasoning in Multimodal Large Language Models
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models