ATP-Bench: Towards Agentic Tool Planning for MLLM Interleaved Generation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "ATP-Bench: Towards Agentic Tool Planning for MLLM Interleaved Generation".
Jane: The paper was written by the authors from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Summary: Tom: So, now that we understand the goal, let’s look at what the paper actually found when they tested this concept. They built this massive dataset of seven thousand seven hundred two QA pairs across eight categories like Travel and Renovation to test how well these models can handle real-world scenarios.
Jane: And it’s not just simple knowledge questions; the they included over one thousand five hundred ninety-two VQA pairs, which is really important for making sure the models are actually looking at the images provided in context.
Lu: The initial findings were quite striking because it seems like most of our current state-of-the-art MLLMs struggle to plan this sequence coherently; they often fail to decide when a visual element is truly necessary.
Meng: The results show that models exhibit very different behavioral tendencies in their tool usage, which is something I’m interested in because it suggests that the underlying architecture of these commercial models varies wildly.
Lalam: It's sad to hear they struggle so much with coherent planning, but seeing these variations allows us to identify where we need to improve our cultural tools and guides moving forward.
Tom: That really highlights the problem, Tom; it’s not just about having a model that can generate images anymore.
Jane: And that’s what leads into the core of how they propose solving this problem in Segment three moving beyond simply what was found to *how* we fix it.
Improvements: Tom: Since the initial results showed these models are struggling with coherent planning, the next logical step is looking at how ATP-Bench proposes a new way to evaluate and improve this capability. They’re suggesting a whole new paradigm called Agentic Tool Planning as the fix.
Jane: The authors want to move beyond just having one image generation tool or using external retrieval; they want the model to act as a central controller that autonomously invokes various tools, which is much more sophisticated than what we've seen before.
Lu: This is where Multi-Agent MLLM-as-a-Judge, or MAM, comes in and it’s a huge improvement because it allows us to evaluate the planning ability itself without needing to run the entire execution pipeline.
Meng: That MAM system is very useful for engineering because we can assess tool-call precision—did the model pick the right tool? Did it use the parameters correctly?—without getting bogged down in end-to-end execution errors.
Lalam: I think this multi-agent judging system will allow us to ensure that our future cultural applications are not just visually appealing, but functionally correct and strategically sound.
Tom: It’s a way to measure the quality of the *intent* itself, Tom; which is a massive improvement over simply looking at the final output.
Jane: Which naturally leads us into the final wrap-up and what this all means for our listeners in Segment four.
Conclusion: Tom: So, we've seen that current MLLMs struggle with planning, but we also saw the solutions through ATP-Bench and MAM. We need to summarize what this all means for the world now that we’ve walked through these sections.
Jane: It’s a monumental step forward in how multimodal models can organize their thoughts, allowing us to create much more sophisticated and intuitive content.
Lu: I'm tremendously excited about the possibility of combining these tool-planning capabilities with generative AI; think of the creative freedom this gives you to design complex visual narratives.
Meng: Practically, this means we can build much more reliable systems that can handle dynamic queries, adapting to specific tool requirements rather than being limited by a single fixed output.
Lalam: I hope this work helps us build tools that genuinely improve how people interact with information and find joy in learning and culture for the future.
Tom: It's definitely a powerful foundation for the next time we want to create an interleaved response that truly makes sense, Tom; it’s about unifying factuality with creativity through structured planning.
Jane: That really sums up the essence of ATP-Bench: Towards Agentic Tool Planning for MLLM Interleaved Generation.
Lu: I think the path forward is wide open now that we can measure this planning ability so clearly defined.
Meng: The practical application possibilities are endless, and it’s time to start building those systems based on this structure.
Lalam: And I am hopeful that the success rate of these models will grow with the same level of tooling and feedback we've seen in this research.
Final Wrap-Up: Tom: Well, that’s all the time we have for today on our show. We’ve discussed how ATP-Bench is pushing us toward Agentic Tool Planning, a method where MLLMs are truly deciding to use tools to make sense of real-world queries.
Jane: It’s a huge leap from just relying on retrieval; the complexity of the planning phase is what's really making this paper so impactful.
Lu: I hope future iterations can handle even more abstract and highly creative tool usage, building on these foundations.
Meng: My main concern is making sure that our real-world deployments of these models prioritize this structured reasoning over simple output fidelity.
Lalam: I wish us all the best as we continue to innovate and make these advancements serve humanity through the power of AI.
Tom: We've got a lot to look forward to with the capabilities demonstrated by ATP-Bench: Towards Agentic Tool Planning for MLLM Interleaved Generation, and that's all for today.
cs.AI
Submitted: 2026-08-22
Updated: 2026-08-25
Code: https://github.com/Qwen-Applications/ATP-Bench
Importance score: 90/100
The gist: Interleaved text-and-image generation is described as a "burgeoning frontier for multimodal large language models (MLLMs)," offering a way to convey complex information more intuitively than text
Key concepts
- ATP-Bench / Agentic Tool Planning
- This framework evaluates how well MLLMs can perform coherent, step-by-step planning. It requires the model to act as an autonomous central controller that strategically invokes various tools to handle complex, real-world queries and generate interleaved content.
- Multi-Agent MLLM-as-a-Judge (MAM)
- MAM is an evaluation system that assesses a model's planning ability directly. It allows researchers to measure tool-call precision—whether the model picked the right tool and used its parameters correctly—without needing to run the entire execution pipeline.
- MLLMs (Multimodal Large Language Models)
- These are advanced models capable of processing and understanding multiple types of data, such as text and images. The episode focuses on improving these models' ability to organize thoughts and plan structured responses that unify factuality with creativity.
Terminology
Summary
Interleaved text-and-image generation is described as a burgeoning frontier for multimodal large language models (MLLMs),
offering a way to convey complex information more intuitively than text alone. However, the paper notes that existing research has converged on two separate paradigms—either image generation or retrieval augmentation—which typically treat the two as mutually exclusive paths, failing to unify factuality with creativity.
The authors argue that the next critical milestone in this field is Agentic Tool Planning, where the model serves as a central controller that autonomously determines when, where, and which tools to invoke to produce interleaved responses for visual-critical queries.
To systematically evaluate this new paradigm, the researchers introduce ATP-Bench (Agentic Tool Planning Bench). This novel benchmark is designed to jointly model both reference and generation capabilities. Key characteristics of ATP-Bench include:
-
A total of 7,702 QA pairs (including 1,592 VQA pairs).
-
Coverage across eight categories and 25 visual-critical intents.
*The dataset is the "first to jointly support hybrid image sourcing (Reference & Generation) and dual query types (QA & VQA), with expert-level annotations for both queries and ground truths."
Furthermore, to evaluate agentic planning independently of end-to-end execution, the authors propose a Multi-Agent MLLM-as-a-Judge (MAM) system. MAM is designed to assess tool planning by:
evaluating tool-call precision, identifying missed opportunities for tool use, and assessing overall response quality without requiring ground truth.
The study’s extensive experiments on 10 state-of-the-art MLLMs yielded three key insights:
-
Existing MLLMs struggle to generate coherent interleaved tool plan, particularly for Travel and Renovation.
-
Gemini 3 Pro achieves the leading performance under our task setting.
-
Models exhibit distinct behavioral tendencies in terms of tool-call frequency and preference.
The paper concludes by detailing its contributions: (1) proposing the new paradigm, Agentic Tool Planning; (2) introducing ATP-Bench to enable systematic study of tool planning capabilities; and (3) proposing the MAM system to assess tool-call precision.
Improvements for AI systems
Based on the architecture and findings presented in ATP-Bench, the primary failure mode of current MLLMs is not a lack of capability, but a lack of systematic, agentic planning. The current systems are reactive; they either generate or retrieve. We must transition to a proactive, multi-stage tool orchestration framework.
Here are the specific improvements we must implement and what the resulting AI system can achieve:
Improvement: We must move beyond simple prompt-response chains and establish the LLM as a dedicated Agentic Controller. This controller does not just answer; it executes a dynamic, multi-step plan. The core logic is: Analyze Query to Identify Visual Need (Gap Detection) to Select Tool(s) Tool(s) Refine Output.
Specific Mechanism: The model must be fine-tuned on the Tool Selection Strategy (as defined in the prompt template, e.g., distinguishing when to use Reference vs. Search). This requires explicit training on the intent of tool usage, not just its syntax.
What the Improved System Can Do:
-
Guaranteed Intent Alignment: The system will reliably choose tools based on semantic need (e.g., recognizing that a
hairstyle
query requiresEditorDiffusion, whereas anarchitecture
query requiresSearch). -
Dynamic Orchestration: It can seamlessly integrate real-world data, generated concepts, and provided context within a single coherent response without forcing the user to switch between distinct modules.
Improvement: We must implement a dedicated Visual Gap Detector (VGD) module upstream of the tool selection phase. This module analyzes the textual content against the query and source documents to identify where visual support is necessary but absent.
Specific Mechanism: The VGD uses contextual embeddings to flag segments where a text-only explanation falls below an established threshold of Visual Criticality.
This preemptively forces the tool planning mechanism to consider calling Reference, Search, or Diffusion before the final text generation occurs.
Improvement: We must integrate a continuous, automated verification loop modeled after the Multi-Agent MLLM-as-a-Judge (MAM) system into the deployment pipeline. This moves evaluation from a post-mortem check to an in-process constraint.
Specific Mechanism: Before any final output is presented, the generated tool calls must pass three internal checks:
-
Precision Check (Reference/Tool Choice): Does this tool actually solve the identified visual gap? (Checking
2a. Tool Choiceand2b. Parameter Accuracy). -
Pacing Check (Structural Integrity): Is the image placed immediately after the relevant semantic paragraph, or does it disrupt the flow? (Checking
1c. Structural Integrity). -
Recall Check: Did the model miss any other obvious visual gaps in this paragraph?
Improvement: We must enforce strict adherence to the Tool Schema (JSON format and specific parameter requirements) across all tool invocations.
Specific Mechanism: Implement a pre-execution validator that checks:
-
img index: Must be a string identifier (e.g.,IMG#1-1
), never an integer (to prevent hallucinated references). -
prompt: ForDiffusion, the prompt must be semantically detailed, not vague. -
parameters: All parameters must align with the defined capability boundaries for that specific tool.
Sources
- OpenLEAF: Open-Domain Interleaved Image-Text Generation and Evaluation
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment
- ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation
- WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
- AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn
- LLM-I: LLMs are Naturally Interleaved Multimodal Creators
- GPT-4o System Card
- Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
- LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- Seedream 4.0: Toward Next-generation Multimodal Image Generation
- Emu: Generative Pretraining in Multimodality
- Chameleon: Mixed-Modal Early-Fusion Foundation Models
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
- MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection