Linguistic Trajectory Encoding for Efficient Long-Horizon Spatial Memory in Embodied Agents
cs.CV, cs.AI
Submitted: 2026-09-04
Updated: 2026-09-29
Project page: https://egolife-ai.github.io/13
License: http://creativecommons.org/licenses/by/4.0/
The gist: Embodied agents performing long-horizon tasks require a memory representation in which the state transitions of dynamic objects remain queryable in natural language across hours-to-days observation
Terminology
Abstract
Embodied agents performing long-horizon tasks require a memory representation in which the state transitions of dynamic objects remain queryable in natural language across hours-to-days observation horizons. Existing systems either drop fine-grained motion (clip-level video-language embeddings), keep it only as raw coordinates (geometric SLAM), or organise it around immediate task context (agent working memories). None of them gives the agent a per-object timeline whose state transitions are themselves queryable in language. Our key contribution is Linguistic Trajectory Encoding (LTE), which compresses dynamic object motion histories via a hybrid representation combining natural language descriptions, sparse spatial anchors, and visual anchors. LTE adapts compression to motion complexity by anchoring periods without reliable observations to the last seen location, while representing motion with geometric waypoints and linguistic descriptions to preserve accuracy. To evaluate these capabilities across extended time horizons, we construct the Spatial Memory Benchmark (SMB) from EgoLife multi-day recordings, targeting capabilities absent in existing benchmarks: semantic trajectory retrieval and long-horizon object retrieval. On SMB, the LTE-based system achieves 45.3% success in semantic trajectory retrieval and 48.7% in long-horizon object retrieval, outperforming structured-memory and VLM baselines (best prior: 31.9% and 34.4%). LTE achieves trajectory compression by factors of 8.7 times to 26.1 times with sub-second query latency on 24,h video. On Ego4D natural-language queries, the system reaches 28.75% / 55.10% R@1/R@5, +15.80 / +31.30 pts over EgoVLPv2.
Sources
- Qwen3-VL Technical Report
- SAM 3: Segment Anything with Concepts
- OSGNet @ Ego4D Episodic Memory Challenge 2025
- GroundNLQ @ Ego4D Natural Language Queries Challenge 2023
- ViPE: Video Pose Engine for 3D Geometric Perception
- DINOv2: Learning Robust Visual Features without Supervision
- EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation
- Khronos: A Unified Approach for Spatio-Temporal Metric-Semantic SLAM in Dynamic Environments
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Qwen3 Technical Report
- Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models