EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory
Le Zhang, Hao Chen, Vlad Roznyatovskiy, Jianzhong Zhang, Ke Sun
University of Michigan, Ann Arbor
cs.CV, cs.AI, cs.CL, cs.HC
Submitted: 2026-08-18
Updated: 2026-08-20
Code: https://github.com/RayNeo-AI-2025/LifeDialBench
Project page: https://egocite.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: This paper introduces EgoCITE (Egocentric Context-augmented Indexing and Time-aware Evidence retrieval), a long-horizon agentic memory framework for egocentric question answering (QA).
Terminology
Summary
This paper introduces EgoCITE (Egocentric Context-augmented Indexing and Time-aware Evidence retrieval), a long-horizon agentic memory framework for egocentric question answering (QA). The system is designed to transform continuous first-person video and audio captured by wearable devices into a searchable record of daily life, enabling an agentic assistant to answer questions about past experiences.
The paper identifies two fundamental bottlenecks in existing egocentric memory systems:
-
Context-poor indexing: "Memory entries are typically derived from short, independently generated video captions and speech transcripts, which often omit the local context needed to interpret people, objects, places, and elliptical utterances. Once such context-poor content is indexed, neither a richer memory structure nor a stronger retrieval agent can reconstruct information that was never represented."
-
Temporal intent ignored in retrieval: "Existing systems typically let an agent search captions or entities by semantic similarity and rerank the returned candidates. Yet semantic similarity captures what an event is about, not when or how often it occurred. Questions over long-horizon egocentric memory frequently express temporal intent through qualifiers such as 'last,' 'first,' 'usually,' or 'this morning.' The most semantically similar event may therefore be neither the requested occurrence nor representative of a recurring habit."
Insight 1: Long-horizon egocentric agentic search is limited by what is indexed.
The paper's analysis found that 10% of WorldMM's extracted entities contain unresolved pronouns, and 17% of LoCoMo memory entries contain unresolved verbatim quotes.
Insight 2: Retrieval must model temporal intent in addition to semantic relevance.
The paper notes that More than 76% of EgoLifeQA and 48% of EgoMem questions contain temporal intent.
Furthermore, accuracy on time-related questions is 9.8–10.3% lower for WorldMM and 13.3% lower for VideoRAG than on time-unrelated questions.
EgoCITE comprises a four-stage pipeline:
Converts egocentric video and audio into dense visual captions and speech transcripts, resolves speaker and participant identities, and fuses outputs into multimodal captions.
-
EgoScheme:
Uses local multimodal context to turn fragmentary video captions and speech transcripts into self-contained atomic memory indices.
It resolves coreferences and ellipsis within a context window (T = 5 minutes by default for actions/utterances, T′ = 30 minutes for activities/conversations). -
EgoIndex:
Organizes complementary action, activity, utterance, and conversation representations into searchable multi-view memory indices at multiple granularities.
The four views are: actions (fine-grained physical behavior), activities (coarse-grained physical behavior), utterances (fine-grained spoken interaction), and conversations (coarse-grained spoken interaction).
Combines semantic search with question-conditioned temporal relevance scoring. It uses a dual-agent design:
-
Drafting agent:
Iteratively retrieves specific evidence using semantic and temporal queries
across multiple rounds, accumulating candidates in an index pool. -
Sampling agent:
Reasons over the accumulated evidence in context to curate a coherent set aligned with the question's temporal intent.
The temporal relevance score is computed as:
-
Ri = 1 if tstart ≤ τi ≤ tend
-
Ri = λ(tstart−τi) if τi tend
where λ = 0.99 (default) controls the time-decay factor. The final retrieval score combines semantic and temporal relevance multiplicatively: score(mi) = Si × Ri.
The response agent receives curated atomic memory indices, maps timestamps back to multimodal captions, and produces the final multiple-choice answer.
EgoScheme follows two principles: (1) Human-centered: each atomic memory index is anchored to a specific person
; (2) Decoupled: each index captures one coherent behavior, activity, utterance, or conversational topic, preventing compounded semantics from confusing similarity-based retrieval.
Evaluated on three benchmarks: EgoLifeQA, EgoMem, and EgoR1-Bench.
-
EgoCITE-GPT outperforms long-context LLM-agent baselines by 3.6–8.9% on EgoLifeQA, 0.9–4.1% on EgoMem, and 2.6–4.7% on EgoR1-Bench.
-
Compared with agentic memory baselines, it achieves average gains of at least 14.2%, 4.4%, and 9.0% on the three benchmarks respectively.
-
EgoCITE achieves 36× lower cost than long-context LLM agents while requiring 23× fewer input tokens.
-
EgoCITE-GPT achieves hit rates of 49.6%, 89.6%, and 62.7% on EgoLifeQA, EgoMem, and EgoR1 respectively, improving over the best baseline by up to 15.1%, 17.0%, and 22.7%.
-
EgoCITE-GPT achieves 62.6% accuracy on time-aware questions, outperforming the strongest long-context LLM baseline (Gemini-3.1-Pro agent) by 3.9% and the strongest agentic memory baseline (WorldMM-GPT) by 15.5%.
-
EgoCITE shows only a 4.8% decline from DAY1 to DAY7, compared with an 11.9% drop for the GPT-5.4 agent, demonstrating better scalability with longer memory horizons.
Starting from caption RAG (61.3% accuracy), adding EgoIndex, EgoScheme, and EgoRetrv progressively improves accuracy to 69.7%, 71.3%, and 75.3%, with hit rate improving from 40.7% to 54.3%, 59.3%, and 62.7%.
Using VLM-generated narrations (Gemini-3-Flash and Gemma-4-31B) instead of human-annotated dense captions results in only 1.2% and 4.4% accuracy drops respectively, demonstrating practical deployment capability.
-
Identification of two fundamental failures in current egocentric memory systems: fragmented/elliptical memory entries and retrieval interfaces that don't integrate temporal intent.
-
Introduction of EgoCITE with context-augmented, multi-view atomic memory indices and dual-agent, time-aware retrieval.
-
Comprehensive evaluation demonstrating significant improvements in accuracy, retrieval alignment, and cost efficiency.
The paper suggests testing under noisier perception, open-world identity resolution, longer personal recordings, cross-modal index construction directly over synchronized multimodal signals, and integration with richer memory structures beyond vector databases.
Improvements for AI systems
Improvements to AI Systems:
-
Context-Augmented Memory Indexing (EgoScheme): Implement a preprocessing layer that resolves coreferences, ellipsis, and pronoun references within a sliding temporal context window (e.g., 5–30 minutes) before storing any memory entry. This ensures that all indexed events are self-contained and interpretable without external context, reducing ambiguity in downstream retrieval.
-
Multi-View Memory Organization (EgoIndex): Structure memory into four parallel, granularity-separated indices—actions, activities, utterances, and conversations—each anchored to a specific person. This decoupling prevents semantic contamination between fine-grained and coarse-grained events, enabling more precise similarity matching.
-
Time-Aware Retrieval Scoring: Augment semantic similarity scores with a temporal relevance function that decays exponentially for events outside the query’s time window (λ = 0.99). Combine both scores multiplicatively (score = semantic × temporal) so that retrieval prioritizes events that match both what and when the user asks about.
-
Dual-Agent Retrieval Pipeline: Deploy a drafting agent that iteratively queries the memory indices using both semantic and temporal queries across multiple rounds, accumulating candidate evidence. Then, a sampling agent reasons over the accumulated candidates in context to select a coherent, temporally aligned subset—avoiding the common failure of retrieving the most semantically similar but temporally wrong event.
-
Temporal Intent Parsing: Add a lightweight classifier or prompt-based module that explicitly detects temporal qualifiers (e.g.,
last,
first,
usually,
this morning
) in user queries and converts them into structured time constraints (start/end timestamps) that feed directly into the retrieval scoring function. -
Cost-Efficient Long-Horizon Memory: Replace full-context LLM prompting with indexed retrieval that maps timestamps back to original multimodal captions only at the final reasoning stage. This reduces input token usage by 23× and cost by 36× while maintaining or improving accuracy, enabling deployment on resource-constrained edge devices.
What the Improved AI System Can Do:
-
Answer time-sensitive questions about past experiences (e.g.,
What did I eat for dinner last Tuesday?
orWhen did I first meet John?
) with high accuracy, even when the semantically similar event occurred at a different time. -
Resolve ambiguous references in memory—such as
he,
that place,
orthe one we talked about
—by using local context from the same time window, making the system robust to fragmented or elliptical user recordings. -
Distinguish between recurring habits and one-off events (e.g.,
usually
vs.last
) by jointly modeling semantic similarity and temporal frequency, avoiding the common failure of retrieving the most common but not the requested occurrence. -
Scale to weeks or months of continuous egocentric data without significant performance degradation (only 4.8% accuracy drop over 7 days vs. 11.9% for baselines), making it viable for long-term personal assistant applications.
-
Operate on noisy, real-world perception inputs (e.g., VLM-generated captions instead of human annotations) with minimal accuracy loss (1.2–4.4%), enabling deployment with off-the-shelf vision-language models.
-
Provide explainable retrieval by returning both the semantic and temporal relevance scores for each retrieved memory, allowing users or downstream systems to verify why a specific event was selected.
Sources
- SuperMemory-VQA: An Egocentric Visual Question-Answering Benchmark for Long-Horizon Memory
- Ellipsis Resolution as Question Answering: An Evaluation
- Zep: A Temporal Knowledge Graph Architecture for Agent Memory
- EgoGraph: Temporal Knowledge Graph for Egocentric Video Understanding
- Gemini: A Family of Highly Capable Multimodal Models
- Gemma 4 Technical Report
- Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
- EgoSelf: From Memory to Personalized Egocentric Assistant
- EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
- Qwen3 Technical Report
- ReAct: Synergizing Reasoning and Acting in Language Models
- Evaluating Memory Capability in Continuous Lifelog Scenario
- MLLMs are Deeply Affected by Modality Bias
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models