MM-ContextFold: Context Folding for Multimodal Agentic Retrieval
cs.CV, cs.AI, cs.IR, cs.MM
Submitted: 2026-09-19
Updated: 2026-09-25
Code: https://github.com/iLearn-Lab/MM26-MMContextFold
License: http://creativecommons.org/licenses/by/4.0/
The gist: Multimodal Agentic Retrieval (MAR) requires agents to solve complex information-seeking tasks by iteratively invoking external tools.
Terminology
Abstract
Multimodal Agentic Retrieval (MAR) requires agents to solve complex information-seeking tasks by iteratively invoking external tools. Typical frameworks such as ReAct maintain raw multimodal inputs and the accumulating interaction history in a single, ever-growing context, leading to the context explosion problem. While existing methods alleviate this issue by compressing redundant text, effective strategies for managing token-intensive visual content remain largely underexplored. To address this gap, we first conduct a systematic empirical study of approximately 10,000 trajectories. The results show that as visual cues are progressively extracted through external tools and textualized into the context, raw images become increasingly redundant. Continued image retention is associated with higher output entropy and can even degrade task accuracy. Motivated by these findings, we propose MM-ContextFold, a training-free framework that loads raw images only when needed. It maintains a persistent, text-only main context for high-level planning and spawns ephemeral branch contexts for image-dependent subtasks. Within each branch, the agent loads the relevant images, completes the subtask, and folds the result back into the main context as a concise textual summary; the images and branch trace are then discarded. Experiments on seven MAR benchmarks across five backbone models show that MM-ContextFold improves average accuracy by 6.3 percentage points over ReAct while reducing the working context length by 27.5%.
Sources
- MindWatcher: Toward Smarter Multimodal Tool-Integrated Reasoning
- SenseNova-MARS: Empowering Multimodal Agentic Reasoning and Search via Reinforcement Learning
- REDSearcher: A Scalable and Cost-Efficient Framework for Long-Horizon Search Agents
- DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories
- Language Models (Mostly) Know What They Know
- WebSailor: Navigating Super-human Reasoning for Web Agent
- MemOS: An Operating System for Memory-Augmented Generation (MAG) in Large Language Models
- A Survey of Context Engineering for Large Language Models
- WebResearcher: Unleashing unbounded reasoning capability in Long-Horizon Agents
- MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling
- Tongyi DeepResearch Technical Report
- VimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory Graph
- ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization
- Vision-DeepResearch Benchmark: Rethinking Visual and Textual Search for Multimodal Large Language Models
- VSearcher: Long-Horizon Multimodal Search Agent via Reinforcement Learning
- MLDocRAG: Multimodal Long-Context Document Retrieval Augmented Generation
- WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models