Retrieved but Not Delivered: Multimodal Memory Delivery for Long-Term Agents
cs.CV, cs.AI, cs.CL, cs.IR
Submitted: 2026-09-26
Updated: 2026-09-26
Project page: https://avalon-s.github.io/DeliverMem
Terminology
Sources
- Argus-Retriever: Vision-LLM Late-Interaction Retrieval with Region-Aware Query-Conditioned MoE for Visual Document Retrieval
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory
- Naive Visual Memory is Not Enough: A Failure-Mode Study of GUI Agents
- MemArbiter: Decision-Time Memory Arbitration for Long-Horizon LLM Agents
- Memory-T1: Reinforcement Learning for Temporal Reasoning in Multi-session Agents
- M2A: Multimodal Memory Agent with Dual-Layer Hybrid Memory for Long-Term Personalized Interactions
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory
- M$^3$Exam: Benchmarking Multimodal Memory for Realistic User-Agent Interactions
- Enabling Chatbots with Eyes and Ears: An Immersive Multimodal Conversation System for Dynamic Interactions
- Presentation, Not Mechanism: A Render Confound in Deprecation-Aware Memory Evaluation
- V-Mem: Modality-Routed Retrieval for Long-Term Multimodal Agentic Memory
- MemOS: An Operating System for Memory-Augmented Generation (MAG) in Large Language Models
- MemVerse: Multimodal Memory for Lifelong Learning Agents
- Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory
- MMA: Multimodal Memory Agent
- Agent Memory: Characterization and System Implications of Stateful Long-Horizon Workloads
- MemGPT: Towards LLMs as Operating Systems
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models