AMIGO: Agentic Multi-Image Grounding Oracle Benchmark
cs.LG, cs.AI
Submitted: 2026-03-30
Updated: 2026-09-16
License: http://creativecommons.org/licenses/by/4.0/
The gist: Agentic vision-language models increasingly act through extended interactions, but most evaluations still focus on single-image, single-turn correctness.
Terminology
Abstract
Agentic vision-language models increasingly act through extended interactions, but most evaluations still focus on single-image, single-turn correctness. We introduce AMIGO (Agentic Multi-Image Grounding Oracle Benchmark), a long-horizon benchmark for hidden-target identification over galleries of visually similar images. In AMIGO, the oracle privately selects a target image, and the model must recover it by asking a sequence of attribute-focused Yes/No questions under a strict protocol that returns Yes/No/Unsure feedback and penalizes invalid actions with Skip. This setting stresses (i) question selection under uncertainty, (ii) consistent constraint tracking across turns, and (iii) fine-grained discrimination as evidence accumulates. We instantiate AMIGO with the Guess My Preferred Dress task and evaluate open-source VLMs with metrics covering identification success, evidence verification, efficiency, protocol compliance, robustness to controlled feedback perturbations, and trajectory-level diagnostics. The benchmarking results show that final-answer accuracy alone overstates evidence-grounded performance: models can guess correctly without verification-passing evidence, waste turns through invalid questions, or fail to preserve the upload protocol. Strong AMIGO performance depends on the combination of visual discrimination, informative question selection, constraint tracking, efficient stopping, sustained protocol following, and recovery from controlled feedback noise; model scale alone does not guarantee reliable long-horizon interactive grounding.
Sources
- Kimi K2.5: Visual Agentic Intelligence
- MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
- MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
- MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
- mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
- MANTIS: Interleaved Multi-Image Instruction Tuning
- MMCR: Advancing Visual Language Model in Multimodal Multi-Turn Contextual Reasoning
- MMMT-IF: A Challenging Multimodal Multi-Turn Instruction Following Benchmark
- Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
- InfoQuest: Evaluating Multi-Turn Dialogue Agents for Open-Ended Conversations with Hidden Context
- Multi-Turn Multi-Modal Question Clarification for Enhanced Conversational Understanding
- M$^3$Searcher: Modular Multimodal Information Seeking Agency with Retrieval-Oriented Reasoning
- MIRAGE: A Benchmark for Multimodal Information-Seeking and Reasoning in Agricultural Expert-Guided Conversations
- Intern-S1: A Scientific Multimodal Foundation Model
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks