R4DSG: Relative 4D Scene Graph Memory for Object-Centric Question Answering in Long Egocentric Video
Ke Ma, Yamin Mao, Weiming Li, Shuai Tan, Yijie Zhong, Hao Chen, Haofen Wang, Meng Wang
Tongji University · Samsung R&D Institute China Beijing · Shanghai Jiao Tong University · Samsung Research America
cs.CV, cs.AI, cs.HC, cs.MM
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: 10 pages, 3 figures, ACM Multimedia 2026, egocentric video; 3D scene graph; temporal memory; graph retrieval; object-state reasoning; multimodal question answering
Project page: https://dualtransparency.github.io/R4DSG
License: http://creativecommons.org/licenses/by-sa/4.0/
Importance score: 75/100
The gist: R4DSG introduces a relative 4D scene graph memory for long egocentric video, designed to support object-centric question answering under weak sensing conditions.
Terminology
Summary
R4DSG introduces a relative 4D scene graph memory for long egocentric video, designed to support object-centric question answering under weak sensing conditions. The paper states: "Long-horizon egocentric video is a rich substrate for wearable AI assistants, but object-centric questions such as where an item was moved, when it last changed state, or why it was relocated remain difficult because caption- and transcript-based memories rarely preserve persistent object identity or structured spatial change. The method addresses limitations of existing approaches:
Existing long-video QA methods mainly emphasize temporal grounding and clip retrieval, while prior 3D scene-graph methods typically assume stronger geometry than free-motion wearable RGB video provides, including point clouds, RGB-D input, posed views, sparse reconstruction, or reconstructed scenes."
The core idea is described as: "The main idea is to separate stable anchors from dynamic objects, maintain persistent object identity across frames, and represent object state through anchor-relative transitions rather than a globally aligned world model. The paper explains:
Static anchors such as tables, shelves, counters, drawers, sofas, or fridges provide stable relational references. Dynamic objects accumulate state changes by forming new spatial relations to different anchors over time. Instead of forcing all observations into one global frame, the method associates object instances across frames, retains only anchor-relative transitions that remain stable over time, and promotes those transitions into memory entries."
The method builds on recent RGB-only advances: Built on recent RGB-only advances in promptable video segmentation, temporal propagation, and relative 3D lifting, the method produces a retrieval-ready memory directly usable for long-horizon question answering.
The pipeline consists of five stages: semantic episode construction, a 2D and 3D visual front-end with SAM 3-style consistency, frame graph construction and persistent association, memory writing, and question answering over retrieval-plus memory. The paper notes: "The semantic episode constructor narrows the local object inventory, the visual front-end extracts relative evidence within those windows, and the graph association stage decides persistence and anchor roles before memory writing."
The representation is formalized as: At each sampled time step t, the method constructs a frame graph Gt = (Vt, Et), where Vt = vtspace, vtact ∪ Vtobj contains one space node, one activity node, and a set of object nodes.
Each object node stores "semantic and geometric attributes, at (v) = [ct, mt, bt3d, xt3d, st], where ct is the semantic class, mt is the mask or crop reference inherited from the segmentation stage, bt3d is a coarse 3D extent, xt3d is a relative 3D location, and st is a textual state descriptor. The representation becomes 4D after linking frame nodes into persistent tracks:
Here, '4D' denotes a time-indexed relative 3D scene-graph memory, not a globally consistent 4D reconstruction. For a dynamic track p, the anchor-relative state at time t is
zt (p) = (at, dt, ut, st), where at is the dominant nearby anchor, dt is a coarse relative distance, ut is a relative direction or layout cue, and st is the current textual state. A temporal event is written when
a dynamic object's dominant anchor changes and the new anchor-relative state remains stable across time, ei = (p, [ts, te], a−, a+, ri), where [ts, te] is the event span, a− and a+ are the source and destination anchors, and ri is a short local rationale drawn from nearby action or interaction context."
The memory is written at the segment-window level: "For an activity-centered segment window Sk, the corresponding memory entry is represented as mk = (tauk, lk, alphak, Ok, Zk, Rk, Ik, Tk, Yk), where tauk is the time span, lk is the place, alphak is the activity, Ok is the salient object set, Zk stores anchor-relative object states, Rk stores edge and relation summaries, Ik stores interaction or lexical-variant cues, Tk stores retrieval tokens, and Yk stores optional explanation-oriented fields for why questions. The document is written by a deterministic memory-writing operator h:
mk = h(Gt t∈Sk, ei ei∩Sk≠∅)."
For evaluation, the paper reports: "Evaluation on a 255-question object-related subset from EgoLifeQA shows, under question-only retrieval, a 6.7-point overall gain over EgoRAG-Text and a 12.5-point gain on when questions, which highlights the value of temporally organized object memory." Specifically, Table 2 shows R4DSG achieves 39.6% overall and 43.1% on when questions under question-only retrieval, compared to EgoRAG-Text's 32.9% and 30.6%, respectively. Under option-blind retrieval / no-why, R4DSG achieves 37.3% overall and 43.1% on when, exceeding the AMEGO-inspired baseline by 1.2 and 5.6 points. The No-Transition control is lower than R4DSG despite sharing cached visual evidence and local relations, which is consistent with a benefit from persistent identity and anchor-transition memory rather than object/relation serialization alone.
An exploratory why-oriented extension adds four explanation fields—reason, purpose, utility, and why summary—and shows an increase from 33.3 to 40.0 on the 15-question why subset and a small overall lift.
A memory-granularity ablation shows Retrieval+ has the highest observed overall and when accuracy compared to Event-only, Low-compression, Episodic-only, and Hybrid variants. The memory scale is compact: The 828 clips produce 2,476 frame-info entries and 134 Retrieval+ documents in a 0.58-MB JSON memory.
The paper concludes: "R4DSG reframes long egocentric QA as memory construction over persistent objects, static anchors, and retrieval-ready documents. On object-centric EgoLifeQA, gains on when questions and the exploratory why-oriented extension suggest the value of temporal and explanatory memory. Relative scene graphs are a practical substrate for wearable assistants and embodied multimedia agents. Limitations include:
Errors may come from missed episodes; fragmented or duplicated tracks; cluttered anchors; retrieval misses; or insufficient causal or social context. Outputs are not yet a human-verified perception benchmark. Memory is retrospective and offline with limited cross-day persistence; online updates, stronger identity maintenance, and multi-subject validation remain future work."
Improvements for AI systems
Based on the paper, here are specific improvements I can implement in AI systems:
-
Improvement: Replace global world-model representations with anchor-relative spatial memory that tracks objects relative to stable reference points (tables, shelves, counters) rather than absolute coordinates.
-
Capability: The AI can answer
where did X go?
orwhat moved from the counter to the fridge?
even with noisy, free-motion RGB video where global reconstruction fails. -
Improvement: Implement a track-association module that maintains consistent object IDs across frames using SAM 3-style segmentation and temporal propagation, separating static anchors from dynamic objects.
-
Capability: The AI can answer
when did the cup last change position?
orwhich object was moved multiple times?
without losing identity during occlusions or viewpoint changes. -
Improvement: Add a memory-writing operator that logs events only when a dynamic object's dominant anchor changes and the new state stabilizes, storing source/destination anchors and a rationale.
-
Capability: The AI can answer
why was the laptop relocated?
by retrieving the event span, the anchors involved, and the nearby action context that explains the move. -
Improvement: Generate compact, queryable memory entries per activity segment, containing time span, place, activity, salient objects, anchor-relative states, relation summaries, interaction cues, and retrieval tokens.
-
Capability: The AI can perform question-only retrieval (without seeing video) and achieve higher accuracy on
when
questions (43.1% vs 30.6% baseline) by directly indexing temporal object states. -
Improvement: Extend memory entries with four explanation fields—reason, purpose, utility, and why summary—to capture causal and functional context during memory construction.
-
Capability: The AI can answer
why
questions (e.g.,why is the book on the floor?
) with a 6.7-point accuracy gain on why-subsets, providing human-interpretable rationales. -
Improvement: Implement a five-stage pipeline (episode construction → 2D/3D front-end → graph association → memory writing → QA) that works with RGB-only input, no depth, no point clouds, no posed views.
-
Capability: The AI can operate on standard wearable camera footage (e.g., glasses, chest-mounted) and still produce reliable object-centric answers, unlike methods requiring strong geometry.
-
Improvement: Use the anchor-relative transition representation to compress long videos into small memory footprints (0.58 MB for 828 clips, 134 documents).
-
Capability: The AI can run on-device with limited storage, enabling real-time wearable assistants that retain weeks of episodic memory without cloud dependency.
-
Improvement: Add a control variant that shares cached visual evidence but omits anchor-transition memory, to isolate the benefit of persistent identity.
-
Capability: The AI can self-diagnose whether performance gains come from temporal transitions versus static relations, improving interpretability of its own memory decisions.
-
Improvement: Use activity-centered segment windows rather than frame-level or full-episode memory, with a Retrieval+ variant that balances compression and detail.
-
Capability: The AI can answer questions about specific activities (e.g.,
during cooking, what did you move?
) with higher accuracy than event-only or episodic-only variants. -
Improvement: Implement error-tolerant association that can recover from missed episodes, duplicated tracks, and ambiguous anchor assignments by promoting only stable transitions into memory.
-
Capability: The AI remains robust in real-world cluttered environments (e.g., kitchen with many similar objects) and can still answer correctly when some frames are lost or noisy.
Abstract
Long-horizon egocentric video is a rich substrate for wearable AI assistants, but object-centric questions such as where an item was moved, when it last changed state, or why it was relocated remain difficult because caption- and transcript-based memories rarely preserve persistent object identity or structured spatial change. Existing long-video QA methods mainly emphasize temporal grounding and clip retrieval, while prior 3D scene-graph methods typically assume stronger geometry than free-motion wearable RGB video provides, including point clouds, RGB-D input, posed views, sparse reconstruction, or reconstructed scenes. R4DSG introduces a relative 4D scene graph memory for long egocentric video. Instead of storing raw graph sequences, R4DSG converts video into compact queryable memory entries indexed by time, place, persistent objects, anchor-relative change, and local interaction context. The main idea is to separate stable anchors from dynamic objects, maintain persistent object identity across frames, and represent object state through anchor-relative transitions rather than a globally aligned world model. Built on recent RGB-only advances in promptable video segmentation, temporal propagation, and relative 3D lifting, the method produces a retrieval-ready memory directly usable for long-horizon question answering. Evaluation on a 255-question object-related subset from EgoLifeQA shows, under question-only retrieval, a 6.7-point overall gain over EgoRAG-Text and a 12.5-point gain on when questions, which highlights the value of temporally organized object memory. These results position relative 4D scene graphs as a practical memory substrate for wearable assistants, AR systems, and embodied multimedia agents. GitHub Page: https://dualtransparency.github.io/R4DSG/.
Sources
- SAM 3: Segment Anything with Concepts
- SAM 3D: 3Dfy Anything in Images
- Project Aria: A New Tool for Egocentric Multi-Modal AI Research
- Imaging for All-Day Wearable Smart Glasses
- TransBiolab: A Real-World Multi-View Dataset of Cluttered Transparent Biomedical Objects
- SAM 2: Segment Anything in Images and Videos
- ResilPhase: Plug-and-Play Phase Mapping and Noise-Resilient Macro-Trajectory Extrapolation for Diffusion Acceleration
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models