DMV-Bench: Diagnosing Long-Horizon Multimodal Agents' Visual Memory with Incidental Cue Injection
cs.CV, cs.AI, cs.CL
Submitted: 2026-06-25
Updated: 2026-08-31
Comments: Accepted to EMNLP 2026 Main Conference. 17 pages, 8 figures, 18 tables
Code: https://github.com/yyyujintang/DMV-Bench
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Qwen2.5-VL Technical Report
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- M2A: Multimodal Memory Agent with Dual-Layer Hybrid Memory for Long-Term Personalized Interactions
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- Gemini: A Family of Highly Capable Multimodal Models
- MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory
- MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks
- EMemBench: Interactive Benchmarking of Episodic Memory for VLM Agents
- MemGUI-Bench: Benchmarking Memory of Mobile GUI Agents in Dynamic Environments
- MemVerse: Multimodal Memory for Lifelong Learning Agents
- Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory
- MMA: Multimodal Memory Agent
- MemGPT: Towards LLMs as Operating Systems
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- MIRIX: Multi-Agent Memory System for LLM-Based Agents
- Agent Workflow Memory
- Auto-scaling Continuous Memory for GUI Agent
- A-MEM: Agentic Memory for LLM Agents
- FindingDory: A Benchmark to Evaluate Memory in Embodied Agents
- Hybrid Self-evolving Structured Memory for GUI Agents
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models