Multimodal Large Language Models as Image Classifiers
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Multimodal Large Language Models as Image Classifiers".
Jane: The gist The MLLM classification performance depends critically on evaluation protocol and ground truth quality,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So we're looking at this paper today titled Multimodal Large Language Models as Image Classifiers by Kisel et al., and what they're really pointing out is that how we test these models matters way more than just the raw model performance.
Jane: Exactly, Tom. The main thesis here is that the performance of these multimodal large language models isn't just about the model itself, but critically dependent on the evaluation protocol and how good our ground truth labels are.
Lu: It’s interesting because they look at common evaluation setups—Open-World, Multiple-Choice, and Closed-World—and find that issues like model outputs falling outside a list or poor mapping methods actually skew the results.
Meng: That makes sense from an engineering standpoint. If the system is designed to map a free description to a class list and that mapping is weak, the whole output is just noise regardless of how smart the language model is.
Lalam: I see what they're saying about how these models need structure when we ask them to classify things, especially when they are trying to select a label from a predefined set.
Tom: Right, and then they introduce this multilabel reannotation called ReGT for six hundred twenty-five ImageNet-1k classes which shows that MLLMs actually benefit quite a bit from having corrected labels, up to plus ten point eight percent <ref:2603.06578#pg3>.
Jane: That's a big finding because it suggests that a lot of the performance gap we see between these multimodal models and traditional supervised models might just be because the original ground truth was noisy or flawed.
Lu: It really shifts the focus from just tweaking the model weights to cleaning up our data and refining how we set up those classification tasks, which is a major point for researchers in this area.
Meng: So they're saying that models that don't lean too heavily on those traditional supervised training signals are actually way more sensitive to the quality of the labels we give them.
Lalam: It implies that if we want these multimodal models to perform well, we need to focus a lot of our effort on getting better, cleaner annotations instead of just chasing bigger model sizes.
Tom: Speaking of testing, they compare three main ways to test these models: Open-World where the model makes a free description and you have to map it somehow; Multiple-Choice where the model picks one from a set with distractors; and Closed-World where you give it a full list of one thousand classes <ref:2603.06578#pg1>.
Jane: That’s the framework they use to structure their comparison, and they even introduce CW+, which is this lightweight post-processing step to handle those out-of-prompt predictions in the closed world without needing complicated constrained decoding.
Paper summary: Lu: They also found that Open-World evaluation can actually outperform Closed-World for about half of the models tested when using embedding space mapping, which challenges some previous findings on how these tasks compare.
Meng: From a practical side, I noticed they quantified the impact of things like batch size and image ordering—apparently random in-batch ordering helps because class-grouped batches cause both models to assign the same label to every image.
Lalam: That's a very concrete detail for engineers because it shows that even small design choices in how we feed data into the model can cause huge variations in the final classification accuracy.
Tom: And they also found that confusion matrix distractors actually cause a ten to fifteen percent drop over random ones when evaluating in a multiple-choice setup, which is something you have to be careful about when setting up those benchmarks.
Jane: So what this means for us listening right now is that we shouldn't just look at the raw accuracy score; we need to consider the entire pipeline—from how the data was labeled to how the model was asked to classify it.
Lu: The implication is that MLLMs are more robust than some of their counterparts when dealing with annotation noise, specifically showing those gains on reannotated data for VLMs and MLLMs.
Meng: So, if you're building a system, you need to account for this sensitivity to annotation quality because the model itself isn't the only variable causing the performance dip.
Lalam: It’s about building a pipeline where we actively try to correct those label errors because that correction really helps narrow that gap with supervised models.
Tom: Moving into the conclusion, they discuss how these findings impact our understanding of multimodal AI classification and what it means for the future of visual recognition research.
Jane: The authors emphasize that the performance gains from corrected labels are substantial, up to plus ten point eight percent, which really narrows that perceived gap with supervised models when we use those reannotated sets <ref:2603.06578#pg3>.
Lu: They also pointed out a clear structure in the correlation matrices: supervised models and this nearest-neighbor variant of DINOv3 form a tight cluster, but MLLMs show lower similarity to traditional vision models, suggesting their ways of failing are fundamentally different.
Meng: That suggests that we can't just expect these multimodal models to behave exactly like standard vision transformers when it comes to error modes.
Lalam: It means that for future work, we should probably focus on developing better methods for correcting these labels and finding more principled ways to map the free-form outputs into those predefined classes.
Tom: The overall message from Multimodal Large Language Models as Image Classifiers is that we need cleaner benchmarks and more careful evaluation protocols to get a fair picture of how well these models are actually performing.
Conclusion: Tom: So we're wrapping up this discussion on "Multimodal Large Language Models as Image Classifiers" and talking about what this whole paper actually means for how we build these vision systems.
Jane: It basically shows that just having a big language model isn't enough for image classification; the way you test it, and the quality of the labels you give it, really matters.
Lu: The authors are pointing out that these models perform best when they get corrected labels, which can boost accuracy by up to ten point eight percent compared to what they used originally.
Meng: From a practical standpoint, this means we need better ways to curate our training data before we even let the AI see it. If the source material is noisy, the model will learn that noise instead of real patterns.
Lalam: And my perspective on this is that if we can use these models as assistants to help fix those mistakes in our datasets, we can actually improve how culture and content are recognized by these systems.
Tom: The main implication here is that the gap between these multimodal models and traditional supervised ones shrinks significantly when you use better data.
Jane: They also found that the way they set up their tests—Open-World versus Closed-World—changes how much difference you see in performance, depending on how the model is prompted.
Lu: The authors show that while some settings give a slight edge to closed-world testing, those gains disappear when you move to an open-world setup with proper mapping.
Meng: So for people actually building these products, this suggests we shouldn't rely on just one evaluation method; you have to design your pipeline around the specific task.
Lalam: If we can get better at correcting these labels, it means the AI becomes a more reliable partner in understanding visual information across different contexts.
Tom: It really makes you wonder how much more time we need to spend on data cleaning and protocol design versus just training bigger models.
Czech Technical University in Prague
cs.CV
Submitted: 2026-03-06
Updated: 2026-10-08
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 93/100
The gist: The gist The MLLM classification performance depends critically on evaluation protocol and ground truth quality, showing that corrected labels can narrow the performance gap with supervised models
Key concepts
- Open-World (OW) Evaluation
- In this setting, the MLLM generates a free-form text description of an image. This description is then mapped to dataset classes using simple rules or text embeddings. It tests how well the model can generalize and map novel visual concepts to existing labels without being constrained by a fixed list.
- Multiple-Choice (MC) Evaluation
- Here, the MLLM must select one correct class from a predefined set of candidate names, which includes exactly one ground truth label and several incorrect distractors. This protocol assesses the model's ability to distinguish between closely related classes based on visual input.
- Ground Truth Reannotation (ReGT)
- The authors created a new dataset with corrected labels (ReGT) by having humans review and fix errors in the original labels. MLLMs showed up to a 10.8% accuracy improvement when tested on this corrected data, proving that noisy ground truth is a major source of performance issues.
- Closed-World (CW) Prompting
- The model is given an exhaustive list of all possible classes and must select the single correct label from that list. A variant called CW+ was introduced to handle predictions outside the prompt, often by using nearest-neighbor mapping from Open-World results.
Terminology
Summary
The gist The MLLM classification performance depends critically on evaluation protocol and ground truth quality, showing that corrected labels can narrow the performance gap with supervised models and that model outputs must be mapped to predefined classes through protocols like Open-World (OW), Multiple-Choice (MC), or Closed-World (CW) >
Evaluation Protocols and Ground Truth Quality
The study identifies key issues across common evaluation protocols, including model outputs that fall outside the provided class list and are discarded, inflated results from weak multiple-choice distractors, and an open-world setting that underperforms only due to poor output mapping > The authors introduce a multilabel reannotation of 625 ImageNet-1k classes called ReGT, which reveals that MLLMs benefit most from corrected labels (up to +10.8%), substantially narrowing the perceived gap with supervised models > Much of the reported MLLM underperformance on classification is thus an artifact of noisy ground truth and flawed evaluation protocol rather than genuine model deficiency > Models less reliant on supervised training signals prove most sensitive to annotation quality >
Task Formulations
The paper compares MLLM performance across three main classification tasks:
-
Open-World (OW): The model generates a free-form image description, which is then mapped to dataset classes, either through simple heuristics or nearest-neighbor search in a text-embedding space >
-
Multiple-Choice (MC): The model selects from a set of candidate class names (up to 26 in prior work), including exactly one ground-truth label and several distractors >
-
Closed-World (CW): The model is prompted with an exhaustive list of all 1000 dataset classes, and the model is asked to output a single one from the list > The authors introduce CW+, which resolves out-of-prompt predictions by adopting the embedding-space nearest neighbour mapping from OW, making MC unnecessary unless constrained by input token length >
Impact of Design Choices and Model Sensitivity
The researchers quantify the impact of commonly overlooked design choices, showing they substantially affect accuracy, including batch size, image ordering, and text encoder selection > Specifically, confusion matrix distractors cause a 10–15% drop over random ones in MC evaluation > Preliminary experiments show that LLaVA-OV experiences a significant decrease in accuracy as the batch size increases, consistent with its sensitivity to larger batch sizes > Furthermore, random in-batch ordering is adopted for all experiments with models that process batches of images per request because class-grouped batches frequently cause both MLLMs to assign the same label to every image, inflating ImGT and ReGT accuracy >
MLLMs as Annotation Assistants
In a controlled case study, human annotators confirmed or integrated the MLLM prediction in approximately 50% of difficult cases, demonstrating that MLLMs can serve as powerful annotation assistants > Annotators reviewed each image alongside anonymized predictions from GPT-4o, SigLIP 2, ImGT, and ReGT to identify and correct erroneous labels > However, the study concludes that error corrections made by trained annotators are not entirely reliable, being completely incorrect in 50.6% of S− and 8.7% of M− images >
Model Comparison and Robustness
The results show that MLLMs benefit most from corrected labels (up to +10.8%), substantially narrowing the perceived performance gap with supervised models > The analysis of correlation matrices reveals a clear structure where supervised models and the k-NN variant of DINOv3 form a tight, highly correlated cluster, while MLLMs exhibit lower similarity to traditional vision models, indicating that their failure modes differ substantially from both supervised and self-supervised approaches > The performance on reannotated data indicates that VLMs and MLLMs are more robust to annotation noise than supervised models, as the gains are most pronounced for these model families >
Conclusion
Overall, the findings show both the promise and the current limitations of MLLMs for visual recognition, underscoring the need for cleaner benchmarks, principled evaluation protocols, and careful integration of model assistance in dataset construction > The gap between MLLMs and supervised models got reduced significantly for better performing GPT4o by roughly 6% on reannotated labels > Although closed-source systems retain an advantage under strict Closed-World prompting, this gap largely disappears in Open-World settings > The work demonstrates that MLLMs can meaningfully assist human annotators in a controlled curation pipeline >
Supplementary Material Details
The supplementary material provides detailed breakdowns of out-of-prompt predictions across label categories, showing that harder label categories like M– and S– exhibit a higher ratio of OOP predictions > Table 14 compares two response formats for MLLMs across a randomly sampled subset of 625 images, concluding that the Class Name setup yields higher accuracy for all models > The analysis of embedding space mapping shows that SigLIP 2 encoding performs best for PaliGemma 2 and GPT-4o, while Qwen3-Embedding-8B encoding achieves the highest performance for LLaVA-OV, InternVL3.5, and Qwen3-VL > The study highlights that the correct mapping rate scales with the overall OOP rate >
Weasel Family Case Study
The reannotated dataset, as introduced in Sec. 2.1, only contains 625 out of the original 1000 classes, where the majority of wildlife is excluded because it is notoriously hard to annotate for non-experts > Across model families, accuracy generally increases when evaluated on the reannotated labels, with substantially larger gains on the 159-image WeaselGT subset, indicating that many of their apparent errors under the original ImGT labels stem from annotation noise > The shift from ImGT to WeaselGT indicates that MLLMs and VLMs are more robust to annotation noise >
Class Equivalence
Figure 10 and Figure 11 list all class pairs treated as equal for evaluation, such as “notebook computer” versus “laptop computer,” noting that the image content of the classes is the same in modern context > The study also notes that distinguishing between a cassette player and tape player from images is often unreliable because the terms are commonly used interchangeably, as the visual differences are minimal > The paper concludes that careful reannotation is essential for reliably evaluating model performance on closely related wildlife categories >
Prompt Design
The prompt design varies significantly by task, with GPT-4o using a strict classification rule to return only the single best class name from the provided list > LLaVA-OV, InternVL3.5, and Qwen3-VL use a JSON output structure for the Closed World setup > The Open World prompt for LLaVA-OV and InternVL3.
Improvements for AI systems
-
Increased performance via ground truth correction: MLLMs
benefit most from corrected labels (up to +10.8%), substantially narrowing the perceived gap with supervised models,
meaning systems trained or evaluated on ReGT will show up to 10.8% better performance than those using original ImageNet-1k labels. -
Mitigation of evaluation protocol bias: Systems can be made more robust by addressing
model outputs that fall outside the provided class list and are discarded,
as well as fixinginflated results from weak multiple-choice distractors.
-
Improved handling of ambiguity in classification: MLLMs can be leveraged to assist human annotators, as they can
confirm or integrate the MLLM prediction in approximately 50% of difficult cases,
allowing forlarge-scale dataset curation
by flagging residual annotation errors. -
Enhanced robustness against fine-grained category noise: Models show
substantial improvement on the reannotated labels compared to the original labels
when evaluated on subsets like the weasel family, indicating that MLLMs aremore sensitive to fine-grained semantic corrections introduced by the reannotations.
-
Resolution of closed-world limitations: The introduction of
CW+
resolves issues by adopting anembedding-space nearest neighbour mapping from OW,
which addressesout-of-prompt (OOP) predictions
without requiring costly constrained decoding, enabling afull 1,000-class Closed-World evaluation.
-
Optimized prediction mapping: The system can utilize model-specific language models for prediction mapping, as the results show that
SigLIP 2 performs best for PaliGemma 2 and GPT-4o,
leading to better semantic alignment between model outputs and class names. -
Refined multiple-choice distractor strategy: Performance is improved by using
distractor selection
based on theconfusion matrix of the EVA-02 model
or selecting classesclosest to c in the BERT embedding space,
which yields distractors that aremore challenging than completely random selection.
-
Improved evaluation metrics for open-world tasks: By adopting an embedding-based mapping strategy, systems can outperform previous methods, as
embedding-based mapping to be the most effective
and showing that this approach outperforms prior string matching methods.
Abstract
Multimodal Large Language Model (MLLM) classification performance depends critically on evaluation protocol and ground truth quality. Studies comparing MLLMs with supervised and Vision-Language Models (VLMs) report conflicting conclusions, and we show these conflicts stem from protocols that either inflate or underestimate performance. Across the most common evaluation protocols, we identify and fix key issues: model outputs that fall outside the provided class list and are discarded, inflated results from weak multiple-choice distractors, and open-world setting that underperforms only due to poor output mapping. We additionally quantify the impact of commonly overlooked design choices - batch size, image ordering, and text encoder selection - showing they substantially affect accuracy. Evaluating on ReGT, our multilabel reannotation of 625 ImageNet-1k classes, reveals that MLLMs benefit most from corrected labels (up to +10.8%), substantially narrowing the perceived gap with supervised models. Much of the reported MLLM underperformance on classification is thus an artifact of noisy ground truth and flawed evaluation protocol rather than genuine model deficiency. Models less reliant on supervised training signals prove most sensitive to annotation quality. Finally, we show that MLLMs can assist human annotators: in a controlled case study, annotators confirmed or integrated MLLM predictions in approximately 50% of difficult cases, demonstrating their potential for large-scale dataset curation. This work is part of the Aiming for Perfect ImageNet-1k project, see https://klarajanouskova.github.io/ImageNet/.
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models