EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory

arXiv:2608.13113 · cs.CV, cs.AI · Submitted 2026-08-13 · Read on arXiv

Weitao Chen, Hu Jiaxin, Xie Tianyidan, Yang Li, Yuyi Qian, Banghao Xu, Ziheng Tang, Shenyi Wang, Mingyue Yu, Duo Li, Jiacheng Shi, Gao Wang, Zhan Xu, Zhicheng Qiu, Xuanfu Li, Jian Yang, Lanjun Wang, Zili Yi

Nanjing University · Huawei Technologies Co., Ltd. · Tianjin University

cs.CV, cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: 21 pages, 4 figures, 6 tables, including appendices

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory EgoMonth is introduced as the first month-level egocentric video understanding benchmark for evaluating

Terminology

Summary

EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory

EgoMonth is introduced as the first month-level egocentric video understanding benchmark for evaluating long-term spatiotemporal memory in real-world daily-life settings. The paper states: "We introduce EgoMonth, the first month-level egocentric video understanding benchmark. EgoMonth comprises over 300 hours of first-person daily-life recordings from 20 participants spanning 20 to 120 days, paired with 1,443 human-crafted multiple-choice QA pairs."

The dataset construction is described as follows: "We initially collect more than 400 hours of first-person daily-life video from 30 volunteers, with recording spans ranging from 20 to 120 days... After screening, we retain over 300 hours of high-quality video from 20 participants, with at least 1K resolution and a frame rate of at least 25 fps. Privacy is handled through an anonymization pipeline: we deploy an anonymization pipeline that integrates Grounding DINO 1.5 for privacy-sensitive region detection and SAM 2 for precise instance segmentation and masking. Strong Gaussian blurring or masking is applied to detected regions, including faces of bystanders and non-participant individuals, personal identifiers such as license plates and home addresses, and private content on digital screens or physical documents."

The task taxonomy is a cognitively grounded framework: We design a cognitively grounded 14-task evaluation framework organized into three hierarchical cognitive levels: Schema Consolidation, Episodic Indexing, and Cascading Reasoning. The levels are described: "Level 1: Schema Consolidation... evaluates whether models can infer stable behavioral schemas, such as habits or personality traits, from repeated cues across weeks... Level 2: Episodic Indexing... requires models to locate specific, less redundant evidence within the long video stream, such as a particular object state, location, event time, spatial relation, or short temporal segment... Level 3: Cascading Reasoning... further requires models to retrieve, maintain, and compose multiple pieces of evidence across different times, locations, viewpoints, or action sequences."

The 14 tasks are listed in Table 1: "Level 1: Habit Inference, Personality Inference; Level 2: Detail Retrieval, Spatial Relation, Self-localization, Temporal Ordering, Event Time, Object Location; Level 3: Procedure Planning, Event Counting, Object Counting, Route Reasoning, Cross-view Spatial Reasoning, Direction Judgement."

Dataset statistics are provided: "EgoMonth contains 738 video clips from 20 participants, totaling 18,072 minutes (301 hours) of egocentric video. Recordings span 20 to 120 days per participant, with an average clip duration of approximately 24.5 minutes. Based on this corpus, we construct 1,443 multiple-choice QA pairs. The QA pairs include cross-video questions: including cross-video questions whose answers may require evidence from multiple videos."

The evaluation protocol is described: "We evaluate twelve representative open-source and closed-source MLLMs: Chat-UniVi-V1.5 (7B, 256 frames), LLaVA-NeXT-Video (7B, 64 frames), MiniCPM-V 4.5 (8B, 256 frames), Qwen2-VL (7B, 256 frames), Qwen2.5-VL (32B, 256 frames), Qwen3-VL (8B, 256 frames), Qwen3-VL-30B-A3B (30B, 256 frames), ShareGPT4Video (8B, 64 frames), ST-LLM (256 frames), VideoLLaMA3 (7B, 512 frames), VITA-1.5 (7B, 16 frames), and Gemini 2.5 Pro (1 fps). The protocol uses multiple-choice format with four options and reports Avg (macro-average of per-task accuracy) and Acc (micro-accuracy: total correct answers divided by total questions). The random-chance baseline is 25%."

Human baseline is established: "Three trained annotators (not involved in QA creation) independently answer all 1,443 questions after watching the corresponding full videos. We report the macro-average human performance (Avg = 94.2%, Acc = 95.1%; inter-annotator Fleiss' κ = 0.78)."

Main results show: The best-performing model, Gemini 2.5 Pro, achieves 71.8% macro-average accuracy, still 22.4 percentage points below the human baseline of 94.2%. Among open-source models: Qwen2.5-VL (32B) achieves the best overall performance, with strong results on several Level 3 tasks, including Procedure Planning (78.6%), Object Counting (50.9%), and Route Reasoning (52.4%). Task-specific failures are noted: Tasks such as Cross-view Spatial Reasoning, Self-localization, and Direction Judgement remain challenging for most models, with several results close to the 25% random-chance level.

Analysis findings include: "Simply Increasing Frame Density Does Not Guarantee More Accurate Memory... VITA-1.5 uses only 16 input frames but achieves 41.6% on Event Counting, outperforming several open-source models with substantially larger frame budgets, such as Qwen2-VL (7B) (36.8%, 256 frames). Another finding: Larger Scale Helps Evidence Integration, but Accurate Indexing Remains Critical... Many Level 2 tasks depend less on broad reasoning ability and more on precisely indexing a specific moment, object state, location, or temporal segment from a long video stream. A third finding: More Structured Spatiotemporal Representations Are Needed for Long-Term Memory... current MLLMs may capture local events, objects, and scenes, but they often fail to organize these observations into persistent temporal and spatial structures."

The paper concludes: "current models remain far from human-level performance: the best-performing model, Gemini 2.5 Pro, achieves 71.8% macro-average accuracy, compared with 94.2% for humans. Our analysis further shows that accurate long-term memory depends not only on frame density or parameter scale, but also on accurate evidence indexing and more structured spatiotemporal representations. EgoMonth thus provides a challenging testbed and diagnostic framework for developing future video MLLMs with faithful long-term spatiotemporal memory."

Limitations are acknowledged: "the current version is built from 20 participants... a larger and more diverse participant pool would further improve demographic coverage... EgoMonth focuses on month-level to multi-month daily-life recordings. Extending the temporal span to seasonal or year-long recordings would allow future benchmarks to evaluate even longer-term memory."

Improvements for AI systems

Improvements to AI Systems Based on EgoMonth:

  1. Hierarchical Memory Architecture: Implement a three-tier memory system mirroring the benchmark's cognitive levels—(a) a schema layer that aggregates repeated cues across days to infer habits/personality, (b) an episodic index that stores timestamped, location-tagged event snapshots for precise retrieval, and (c) a reasoning layer that composes multiple indexed evidence pieces across time/space. This replaces flat frame-buffering with structured long-term storage.

  2. Spatiotemporal Evidence Indexing Module: Add a dedicated component that, during video ingestion, generates a searchable index of object states, spatial relations, and temporal segments (e.g., object X at location Y between timestamps T1–T2). This enables targeted retrieval for tasks like Object Location, Temporal Ordering, and Self-localization, reducing reliance on dense frame sampling.

  3. Cross-Video Memory Fusion: Extend the model's context window to explicitly link and compare evidence across multiple video clips from the same participant. This supports Route Reasoning and Cross-view Spatial Reasoning by maintaining a persistent world map and timeline that updates as new videos are processed.

  4. Adaptive Frame Selection Based on Task Type: Replace uniform frame sampling with a task-aware selector. For Level 1 tasks (habits/personality), sample sparse, widely spaced frames to capture repeated patterns; for Level 2 (detail retrieval), use dense sampling only around indexed moments; for Level 3 (counting/planning), use a hybrid that prioritizes event boundaries and object appearances.

  5. Privacy-Preserving Representation Learning: Train the model on anonymized data (using the paper's masking pipeline) to learn features robust to blurred/masked regions. This improves generalization to real-world egocentric data where privacy filtering is mandatory, while preventing over-reliance on faces or text identifiers.

  6. Temporal Abstraction Layer: Introduce a module that compresses long recordings into hierarchical event summaries (e.g., morning routine, commute, work session) with duration and frequency statistics. This aids Event Counting and Procedure Planning by providing high-level temporal structure without storing every frame.

  7. Diagnostic Task-Specific Heads: Add lightweight output heads for the 14 task types (e.g., a counting head for Object/Event Counting, a spatial relation head for Direction Judgement). This allows the model to specialize its reasoning per task, improving performance on currently failing tasks like Cross-view Spatial Reasoning and Self-localization.

  8. Confidence-Aware Answering with Temporal Grounding: Train the model to output not just an answer but also the timestamp(s) and video segment(s) it used as evidence. This enables self-correction (e.g., if evidence is weak, fall back to schema-level inference) and provides interpretability for debugging long-term memory failures.

What the Improved AI System Can Do:

  • Answer questions about habits and personality from weeks of egocentric video (e.g., Does this person typically eat lunch at home?) with near-human accuracy.

  • Retrieve specific object states or locations from months of recordings (e.g., Where did I last leave my keys?) by querying the episodic index.

  • Reason across multiple days and viewpoints to solve route planning or cross-view spatial puzzles (e.g., Which path did I take to the gym on Tuesday?).

  • Count events or objects across long spans (e.g., How many times did I meet with my advisor this month?) using temporal abstraction and counting heads.

  • Operate effectively on privacy-masked video, maintaining performance despite missing facial or text details.

  • Provide evidence-grounded answers, allowing users to verify the model's reasoning and trust its long-term memory claims.

Sources

Related papers