FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs
summary
The gist
Multimodal large language models (MLLMs) are rapidly expanding into structured computer vision tasks like object detection, but there is currently no standardized benchmark to systematically evaluate
In short
This work created a comprehensive benchmark to evaluate how well generalist multimodal models localize objects using text prompts across four tasks: object detection, referring expression detection, instance localization, and video object detection. It found that model performance is highly dependent on the output format requested in the prompt.
Key concepts
- Promptable Localization
- This refers to the ability of a multimodal LLM to accurately find and describe an object based on a natural language instruction. The benchmark tests this skill across different visual tasks, ranging from simple detection to complex instance localization.
- Format Adherence (FA)
- This metric checks if the model's output strictly follows the specific structure requested in the prompt, such as JSON or plain text. It is a prerequisite for most other performance metrics because models are evaluated based on their ability to adhere to these formatting constraints.
- Bounding Box Representation
- This involves testing different ways of describing a detected object's location, like using coordinates (xyxy), center-width-size (xywh), or four corners. The study found that different models prefer specific formats, and prompting for a format they don't prefer often results in poor performance.
- Multi-stage Format Search
- This is the process used to find the best configuration for each model before testing. It involves systematically searching through various bounding box formats, JSON keys, and output types to determine which setting yields the highest overall accuracy for a specific task.
Terminology used across episodes
This episode discusses
- FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs · Paper Radio
- Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- KnowDR-REC: A Benchmark for Referring Expression Comprehension with Real-World Knowledge
- Microsoft COCO: Common Objects in Context
- Improved Baselines with Visual Instruction Tuning
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
- Qwen2.5-VL Technical Report
- Qwen3-VL Technical Report
- Ultralytics YOLO Evolution: An Overview of YOLO27, YOLO26, YOLO11, YOLOv8, and YOLOv5 Object Detectors for Computer Vision and Pattern Recognition · Paper Radio
- CogVLM: Visual Expert for Pretrained Language Models
- Personalize Segment Anything Model with One Shot
The paper
FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs · Read on arXiv
Tuebingen AI Center, University of Tuebingen
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs".
Jane: Multimodal large language models (MLLMs) are rapidly expanding into structured computer vision tasks like object detection, but there is currently no standardized benchmark to systematically evaluate these capabilities at scale.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Alright, let's talk about what this paper is calling "FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs." It really highlights that the main hurdle right now is that most models are trained to be generalists, meaning they can handle many things at once, but that flexibility makes it hard to evaluate their precision on specific vision tasks.
Jane: That's a big point, Tom. They introduce this benchmark because current evaluations don't account for the ambiguity that comes from promptable bounding box generation; they need a way to standardize inputs and outputs so we can compare models fairly across different localization needs.
Lu: The authors are addressing the fact that MLLMs often generate bounding boxes as part of their free-form text output based on a prompt, which means the spatial prediction is guided by internal representations rather than strict algorithmic rules, which creates evaluation challenges (<ref:2606.04282#pg1>).
Meng: So what's the actual structure they propose to fix that ambiguity? Are we talking about a new kind of training or just a better way to run the tests on existing models?
Lalam: They are proposing a unified framework that standardizes inputs, makes sure bounding box outputs are parsable, and defines clear evaluation protocols across four core task categories: object detection, referring expression detection, instance-level detection, and video-based detection (<ref:2606.04282#pg0>).
The paper's summary: Tom: So the summary of "FindIt" is essentially that they built a comprehensive benchmark covering four main vision tasks—object detection, referring expression detection, instance localization, and video object detection—and they did this by spanning a grid over different data and output formats.
Jane: That means we're not just looking at one kind of task; we’re testing if the model understands language for descriptions, if it can find specific instances in images, and if it can handle sequences from video inputs as well (<ref:2606.04282#pg2>).
Lu: What's really interesting is how they look at the output formats too; they vary everything from corner-based bounding boxes to JSON variants with different keys, and plain text responses (<ref:2606.04282#pg1>).
Meng: That variability in output format sounds like a huge practical hurdle for deployment because, as we see, models are currently very sensitive to those formatting constraints; if they don't get the format exactly right, the score tanks.
Lalam: The paper highlights that this sensitivity means that even if a model is technically accurate in its localization results, it might fail just because the output structure didn't match what was prompted for (<ref:2606.04282#pg1>).
The paper's improvements: Tom: The paper suggests a multi-stage format search process to figure out the best configuration for each model before doing the final evaluation, which involves selecting a representation, then picking the best JSON key name, and finally choosing the prompt that yields the highest average F1 score per task family.
Jane: That systematic approach sounds like they are trying to automate that messy trial-and-error process we see happening when we try to fine-tune these models for specific outputs. It’s about finding the sweet spot for each model's strengths (<ref:2606.04282#pg1>).
Lu: They show that most models tend to specialize in one preferred format and will output that format regardless of what the prompt specifically asks for, which is a significant limitation they are pointing out.
Meng: So, if a model has a preference, we might actually get near-zero F1 scores unless we use its preferred format instead of what the user asked for. That points to some serious issues with reliability in real applications where input requirements change constantly.
Lalam: They also found that open-source models generally show better localization abilities than closed-source models, but closed models are actually more robust when it comes to handling variations in the output format (<ref:2606.04282#pg1>).
Conclusion: Tom: So, wrapping up the discussion on "FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs," the main implication is that we need models that can reliably generate standardized, interpretable outputs across different tasks and prompting conditions.
Jane: That means for agentic systems to actually work in the real world, we need models that can generate consistent spatial data without needing constant retraining just to fix their output structure.
Lu: This benchmark sets a foundation for future research on promptable, generalist localization because it provides the necessary framework to systematically test these capabilities at scale across all those task categories we discussed.
Meng: For practical engineering, this means we can finally start comparing models on a consistent metric that accounts for both how well they localize and how well they stick to the required output specifications.
Lalam: I think this benchmark is crucial because it moves us toward building AI systems that aren't just capable of vision but are reliable enough to function as dependable tools in complex decision-making processes, which is a big step forward for the field.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization