S-EMBER: A Large-Scale Benchmark for Streaming Egocentric Memory Retrieval
cs.CV, cs.AI
Submitted: 2026-07-02
Updated: 2026-08-26
Code: https://github.com/facebookresearch/S-EMBER
License: http://creativecommons.org/licenses/by/4.0/
The gist: As wearable devices enable continuous first-person recording, AI assistants must reason across long time horizons to recall past experiences-a capability known as episodic memory.
Terminology
Abstract
As wearable devices enable continuous first-person recording, AI assistants must reason across long time horizons to recall past experiences-a capability known as episodic memory. Current benchmarks often rely on offline evaluation with access to entire video files, failing to simulate the streaming reality of wearable intelligence. We introduce S-EMBER (Streaming Egocentric Memory Benchmark for Episodic Retrieval), a large-scale benchmark comprising 3,141 videos totaling 388 hours of organic activity captured via Ray-Ban Meta smart glasses. S-EMBER formalizes grounded streaming episodic retrieval, a paradigm shift from global offline search to causal, active recall triggered by visual events in a continuous stream. We provide 9,448 QA pairs requiring manual visual proof through precise temporal localization and supporting flexible response lengths to simulate natural human-AI interaction. Our extensive benchmarking of frontier models reveals a grounded recall gap: models answer and localize with moderate competence in isolation, yet fall furthest short of human performance when both must hold for the same query, the strongest reaching less than half the human rate. S-EMBER establishes a hardware-authentic foundation for developing grounded, reliable episodic memory in the next generation of wearable AI agents.
Sources
- Qwen3-VL Technical Report
- AMEGO: Active Memory from long EGOcentric videos
- Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos
- OpenAI GPT-5 System Card
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Can I Trust Your Answer? Visually Grounded Video Question Answering
- SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models