FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs

arXiv:2606.04282 · cs.CV · Submitted 2026-06-02 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs".

Jane: Multimodal large language models (MLLMs) are rapidly expanding into structured computer vision tasks like object detection, but there is currently no standardized benchmark to systematically evaluate these capabilities at scale.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Alright, let's talk about what this paper is calling "FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs." It really highlights that the main hurdle right now is that most models are trained to be generalists, meaning they can handle many things at once, but that flexibility makes it hard to evaluate their precision on specific vision tasks.

Jane: That's a big point, Tom. They introduce this benchmark because current evaluations don't account for the ambiguity that comes from promptable bounding box generation; they need a way to standardize inputs and outputs so we can compare models fairly across different localization needs.

Lu: The authors are addressing the fact that MLLMs often generate bounding boxes as part of their free-form text output based on a prompt, which means the spatial prediction is guided by internal representations rather than strict algorithmic rules, which creates evaluation challenges (<ref:2606.04282#pg1>).

Meng: So what's the actual structure they propose to fix that ambiguity? Are we talking about a new kind of training or just a better way to run the tests on existing models?

Lalam: They are proposing a unified framework that standardizes inputs, makes sure bounding box outputs are parsable, and defines clear evaluation protocols across four core task categories: object detection, referring expression detection, instance-level detection, and video-based detection (<ref:2606.04282#pg0>).

The paper's summary: Tom: So the summary of "FindIt" is essentially that they built a comprehensive benchmark covering four main vision tasks—object detection, referring expression detection, instance localization, and video object detection—and they did this by spanning a grid over different data and output formats.

Jane: That means we're not just looking at one kind of task; we’re testing if the model understands language for descriptions, if it can find specific instances in images, and if it can handle sequences from video inputs as well (<ref:2606.04282#pg2>).

Lu: What's really interesting is how they look at the output formats too; they vary everything from corner-based bounding boxes to JSON variants with different keys, and plain text responses (<ref:2606.04282#pg1>).

Meng: That variability in output format sounds like a huge practical hurdle for deployment because, as we see, models are currently very sensitive to those formatting constraints; if they don't get the format exactly right, the score tanks.

Lalam: The paper highlights that this sensitivity means that even if a model is technically accurate in its localization results, it might fail just because the output structure didn't match what was prompted for (<ref:2606.04282#pg1>).

The paper's improvements: Tom: The paper suggests a multi-stage format search process to figure out the best configuration for each model before doing the final evaluation, which involves selecting a representation, then picking the best JSON key name, and finally choosing the prompt that yields the highest average F1 score per task family.

Jane: That systematic approach sounds like they are trying to automate that messy trial-and-error process we see happening when we try to fine-tune these models for specific outputs. It’s about finding the sweet spot for each model's strengths (<ref:2606.04282#pg1>).

Lu: They show that most models tend to specialize in one preferred format and will output that format regardless of what the prompt specifically asks for, which is a significant limitation they are pointing out.

Meng: So, if a model has a preference, we might actually get near-zero F1 scores unless we use its preferred format instead of what the user asked for. That points to some serious issues with reliability in real applications where input requirements change constantly.

Lalam: They also found that open-source models generally show better localization abilities than closed-source models, but closed models are actually more robust when it comes to handling variations in the output format (<ref:2606.04282#pg1>).

Conclusion: Tom: So, wrapping up the discussion on "FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs," the main implication is that we need models that can reliably generate standardized, interpretable outputs across different tasks and prompting conditions.

Jane: That means for agentic systems to actually work in the real world, we need models that can generate consistent spatial data without needing constant retraining just to fix their output structure.

Lu: This benchmark sets a foundation for future research on promptable, generalist localization because it provides the necessary framework to systematically test these capabilities at scale across all those task categories we discussed.

Meng: For practical engineering, this means we can finally start comparing models on a consistent metric that accounts for both how well they localize and how well they stick to the required output specifications.

Lalam: I think this benchmark is crucial because it moves us toward building AI systems that aren't just capable of vision but are reliable enough to function as dependable tools in complex decision-making processes, which is a big step forward for the field.

Tuebingen AI Center, University of Tuebingen

cs.CV

Submitted: 2026-06-02

Updated: 2026-10-06

Code: https://github.com/esh04/FindIt

Importance score: 86/100

The gist: Multimodal large language models (MLLMs) are rapidly expanding into structured computer vision tasks like object detection, but there is currently no standardized benchmark to systematically evaluate

Key concepts

Promptable Localization
This refers to the ability of a multimodal LLM to accurately find and describe an object based on a natural language instruction. The benchmark tests this skill across different visual tasks, ranging from simple detection to complex instance localization.
Format Adherence (FA)
This metric checks if the model's output strictly follows the specific structure requested in the prompt, such as JSON or plain text. It is a prerequisite for most other performance metrics because models are evaluated based on their ability to adhere to these formatting constraints.
Bounding Box Representation
This involves testing different ways of describing a detected object's location, like using coordinates (xyxy), center-width-size (xywh), or four corners. The study found that different models prefer specific formats, and prompting for a format they don't prefer often results in poor performance.
Multi-stage Format Search
This is the process used to find the best configuration for each model before testing. It involves systematically searching through various bounding box formats, JSON keys, and output types to determine which setting yields the highest overall accuracy for a specific task.

Terminology

Summary

Multimodal large language models (MLLMs) are rapidly expanding into structured computer vision tasks like object detection, but there is currently no standardized benchmark to systematically evaluate these capabilities at scale. This work introduces the first comprehensive benchmark specifically designed to assess the promptable localization abilities of generalist MLLMs by spanning four core task categories: object detection, referring expression detection, instance-level detection, and video-based detection.

Benchmark Design and Scope

The benchmark is designed to address the ambiguity introduced by promptable bounding box generation by proposing a unified framework that standardizes inputs, enforces parsable bounding box outputs, and defines transparent evaluation protocols across tasks. It leverages classical datasets for four common localization tasks: (1) object detection, (2) referring expression detection, (3) instance localization, and (4) video object detection. The evaluation spans a grid over three axes: (i) task and data spanning four task families and thirteen datasets; (ii) bounding-box representation comprising corner-based, center-width-size, and four-corner formats; and (iii) output format covering plain text and JSON variants with different keys.

Task Categories

The benchmark covers four distinct task families:

  1. Object detection: Uses one or multiple class labels to indicate which objects to localize.

  2. Referring expression detection: Uses a free-form natural-language description for the object, testing language understanding capabilities.

  3. Instance detection: Requires localization of one specific object instance conditioned on a reference image of that instance, reflecting scenarios relevant to reasoning in high-resolution visual space or robotic scenarios.

  4. Video object detection: Extends object and instance detection to multi-frame inputs from video sequences.

Evaluation Metrics and Format Sensitivity

The evaluation measures both how well a model localizes and how strongly its score depends on the format instruction, as models are evaluated at the configuration where they score highest. Key metrics include:

** mIoU averages the IoU over every matched pair, where the matching cost is defined as 1 − IoU. For multi-label queries, label agreement always dominates IoU when assigning a prediction.**

Format Adherence (FA) is the fraction of responses that parse as the prompted format; it is a precondition for all other metrics. The study highlights that current systems are highly sensitive to formatting constraints and often fail to generalize even to minor variations, with results showing that the format itself can be correct, while localization results can still be wrong.

Performance Analysis Across Models

The evaluation compares open-source and closed-source MLLMs. Results indicate that:

Open-source models can have better localization abilities compared to closed-source models, but also struggle to deviate from a specific output format and that closed-source models are more robust to format variations.

GLM-4.6V provides the best overall scores, mainly driven by its high performance on instance detection tasks. However, most models struggle with the task of instance detection, which has the lowest performance of all four tasks.

Bounding Box Representation and Output Format Sensitivity

The evaluation systematically varies bounding-box representations (e.g., xyxy vs. xywh) and output formats (plain text vs. JSON). The findings show that the preferred bbox for most models is xyxy, yxyx for Gemma-4 and Gemini-2.5-Flash, and xywh for GPT-5.4 and Sonnet-4.5 in JSON. Crucially, the study demonstrates that most models specialize in one preferred format and will output this format independent of the given prompt instructions, meaning prompting for a different format often results in near-zero F1 scores unless the model's preferred format is used.

Conclusion and Implications

The work concludes that although most models exhibit strong localization, their performance remains highly sensitive to specific output formats. This points toward the need for models that can reliably generate standardized, interpretable outputs and generalize across tasks and prompting conditions, which is critical for improving multimodal reasoning in agentic systems. The benchmark serves as a foundation for future research on promptable, generalist localization.

How it works

The benchmark employs a multi-stage format search to select the optimal configuration for each model before full evaluation. This process involves:

  1. Stage 1: Selecting the bounding-box representation based on the models text vs JSON output, sweeping across formats observed in the tested models (e.g., xyxy, xywh).

  2. Stage 2: Fixing that representation and then sweeping the JSON key name to find the best performing key for that format.

  3. Stage 3: Taking the winner from Stage 1 text, the winner from Stage 2 JSON, and the unconstrained prompt to select the configuration with the highest average F1@0.5 score per task family.

Improvements for AI systems

Here are specific improvements for AI systems based on the FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs paper, along with what those improved systems can achieve:


  1. Improve MLLM Localization Reliability via Output Format Specialization:

Large language models (MLLMs) are currently highly sensitive to prompt formatting and output structure. The research demonstrates that current models often fail to generalize even to minor variations in bounding box definitions (e.g., switching from XYXY to XYWH) or JSON key names, leading to near-zero F1 scores despite high format adherence.

Improving AI systems: Implement a pre-processing or fine-tuning layer specifically designed to guide the MLLM's output toward a standardized, robust bounding box format (e.g., consistent XYXY). This involves training the model on explicit instruction following for specific JSON schemas, rather than relying solely on free-form prompting.

Improved AI System Capability: The system will exhibit significantly higher and more stable localization accuracy across different user prompts and application environments (agentic systems), reducing the format sensitivity bottleneck that currently limits real-world deployment.

  1. Develop Format-Aware Prompting for Generalist Models:

The paper shows that models often adhere to a format they were trained on, even when explicitly prompted otherwise (e.g., Qwen3.5-Thinking struggles with center-format definitions).

Improving AI systems: Integrate a Format Selection Module into the LLM pipeline. Before executing the core localization task, this module analyzes the prompt and selects the output format (bounding box representation and JSON key) that maximizes predicted performance for that specific model configuration, effectively performing a multi-stage format search internally.

Improved AI System Capability: The system will dynamically choose the optimal spatial output representation (e.g., corner vs. center-width-height) on a per-query basis, ensuring that the localization accuracy is maximized for every unique user request, rather than relying on a single, potentially sub-optimal default format.

  1. Enhance Instance and Video Localization in Complex Scenarios:

Instance detection and video object detection are identified as areas where most generalist MLLMs struggle compared to classical detectors like YOLO or DETR. Furthermore, the paper notes performance drops when extending tasks from two to eight frames in video inputs (Table 6).

Improving AI systems: Develop a specialized Temporal-Spatial Reasoning Module for video inputs. This module would explicitly handle multi-frame prompts by maintaining object identities across frames and using predicted bounding boxes from previous frames to inform predictions in subsequent frames, rather than treating each frame as an independent image.

Improved AI System Capability: The system will achieve robust tracking and detection in dynamic scenes (e.g., autonomous driving or robotic manipulation) where objects move between frames, significantly outperforming models that treat video as a sequence of static images.

  1. Create a Scalable, Task-Specific Localization Benchmark:

The lack of standardized benchmarks for localization is a major gap in evaluation. The FindIt suite (covering object detection, referring expressions, instance localization, and video detection) provides the necessary framework.

Improving AI systems: Adopt the FindIt benchmark as the mandatory validation standard for all new MLLM development cycles focused on spatial output generation. This involves systematically testing models across all four task families and varying their output formats (JSON vs. Text).

Improved AI System Capability: This standardization will allow researchers to compare models directly on a consistent, multi-faceted metric that accounts for both localization accuracy AND adherence to structured output specifications, leading to faster identification of state-of-the-art performance.

  1. Mitigate Grounding Mismatch in Caption/Description Tasks:

The analysis suggests that adapting grounded datasets (like iGround) introduces a potential ground-truth mismatch when treating nouns as labels, though the delta is small (2–5 mIoU).

Improving AI systems: Implement a Caption-Aware Grounding Correction mechanism. If the model is prompted using a caption from an adapted dataset, this module can use auxiliary knowledge to weigh predictions against the original context, mitigating biases introduced by class-labeling in detection tasks.

Improved AI System Capability: The system will perform more accurately on complex referring expression and grounded captioning tasks (like those found in Synthetic Visual Genome), where the model must understand semantic relationships rather than just raw pixel alignment.

Sources

Related papers