Multimodal Large Language Models as Image Classifiers
summary
The gist
The gist The MLLM classification performance depends critically on evaluation protocol and ground truth quality, showing that corrected labels can narrow the performance gap with supervised models
In short
The study investigated how Multimodal Large Language Models (MLLMs) perform as image classifiers under different evaluation protocols like Open-World, Multiple-Choice, and Closed-World. Findings show that performance is heavily dependent on ground truth quality; corrected labels significantly narrow the gap with supervised models. MLLMs are sensitive to annotation noise, suggesting current performance gaps often stem from flawed evaluation rather than inherent model deficiency.
Key concepts
- Open-World (OW) Evaluation
- In this setting, the MLLM generates a free-form text description of an image. This description is then mapped to dataset classes using simple rules or text embeddings. It tests how well the model can generalize and map novel visual concepts to existing labels without being constrained by a fixed list.
- Multiple-Choice (MC) Evaluation
- Here, the MLLM must select one correct class from a predefined set of candidate names, which includes exactly one ground truth label and several incorrect distractors. This protocol assesses the model's ability to distinguish between closely related classes based on visual input.
- Ground Truth Reannotation (ReGT)
- The authors created a new dataset with corrected labels (ReGT) by having humans review and fix errors in the original labels. MLLMs showed up to a 10.8% accuracy improvement when tested on this corrected data, proving that noisy ground truth is a major source of performance issues.
- Closed-World (CW) Prompting
- The model is given an exhaustive list of all possible classes and must select the single correct label from that list. A variant called CW+ was introduced to handle predictions outside the prompt, often by using nearest-neighbor mapping from Open-World results.
Terminology used across episodes
This episode discusses
The paper
Multimodal Large Language Models as Image Classifiers · Read on arXiv
Czech Technical University in Prague
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Multimodal Large Language Models as Image Classifiers".
Jane: The gist The MLLM classification performance depends critically on evaluation protocol and ground truth quality,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So we're looking at this paper today titled Multimodal Large Language Models as Image Classifiers by Kisel et al., and what they're really pointing out is that how we test these models matters way more than just the raw model performance.
Jane: Exactly, Tom. The main thesis here is that the performance of these multimodal large language models isn't just about the model itself, but critically dependent on the evaluation protocol and how good our ground truth labels are.
Lu: It’s interesting because they look at common evaluation setups—Open-World, Multiple-Choice, and Closed-World—and find that issues like model outputs falling outside a list or poor mapping methods actually skew the results.
Meng: That makes sense from an engineering standpoint. If the system is designed to map a free description to a class list and that mapping is weak, the whole output is just noise regardless of how smart the language model is.
Lalam: I see what they're saying about how these models need structure when we ask them to classify things, especially when they are trying to select a label from a predefined set.
Tom: Right, and then they introduce this multilabel reannotation called ReGT for six hundred twenty-five ImageNet-1k classes which shows that MLLMs actually benefit quite a bit from having corrected labels, up to plus ten point eight percent <ref:2603.06578#pg3>.
Jane: That's a big finding because it suggests that a lot of the performance gap we see between these multimodal models and traditional supervised models might just be because the original ground truth was noisy or flawed.
Lu: It really shifts the focus from just tweaking the model weights to cleaning up our data and refining how we set up those classification tasks, which is a major point for researchers in this area.
Meng: So they're saying that models that don't lean too heavily on those traditional supervised training signals are actually way more sensitive to the quality of the labels we give them.
Lalam: It implies that if we want these multimodal models to perform well, we need to focus a lot of our effort on getting better, cleaner annotations instead of just chasing bigger model sizes.
Tom: Speaking of testing, they compare three main ways to test these models: Open-World where the model makes a free description and you have to map it somehow; Multiple-Choice where the model picks one from a set with distractors; and Closed-World where you give it a full list of one thousand classes <ref:2603.06578#pg1>.
Jane: That’s the framework they use to structure their comparison, and they even introduce CW+, which is this lightweight post-processing step to handle those out-of-prompt predictions in the closed world without needing complicated constrained decoding.
Paper summary: Lu: They also found that Open-World evaluation can actually outperform Closed-World for about half of the models tested when using embedding space mapping, which challenges some previous findings on how these tasks compare.
Meng: From a practical side, I noticed they quantified the impact of things like batch size and image ordering—apparently random in-batch ordering helps because class-grouped batches cause both models to assign the same label to every image.
Lalam: That's a very concrete detail for engineers because it shows that even small design choices in how we feed data into the model can cause huge variations in the final classification accuracy.
Tom: And they also found that confusion matrix distractors actually cause a ten to fifteen percent drop over random ones when evaluating in a multiple-choice setup, which is something you have to be careful about when setting up those benchmarks.
Jane: So what this means for us listening right now is that we shouldn't just look at the raw accuracy score; we need to consider the entire pipeline—from how the data was labeled to how the model was asked to classify it.
Lu: The implication is that MLLMs are more robust than some of their counterparts when dealing with annotation noise, specifically showing those gains on reannotated data for VLMs and MLLMs.
Meng: So, if you're building a system, you need to account for this sensitivity to annotation quality because the model itself isn't the only variable causing the performance dip.
Lalam: It’s about building a pipeline where we actively try to correct those label errors because that correction really helps narrow that gap with supervised models.
Tom: Moving into the conclusion, they discuss how these findings impact our understanding of multimodal AI classification and what it means for the future of visual recognition research.
Jane: The authors emphasize that the performance gains from corrected labels are substantial, up to plus ten point eight percent, which really narrows that perceived gap with supervised models when we use those reannotated sets <ref:2603.06578#pg3>.
Lu: They also pointed out a clear structure in the correlation matrices: supervised models and this nearest-neighbor variant of DINOv3 form a tight cluster, but MLLMs show lower similarity to traditional vision models, suggesting their ways of failing are fundamentally different.
Meng: That suggests that we can't just expect these multimodal models to behave exactly like standard vision transformers when it comes to error modes.
Lalam: It means that for future work, we should probably focus on developing better methods for correcting these labels and finding more principled ways to map the free-form outputs into those predefined classes.
Tom: The overall message from Multimodal Large Language Models as Image Classifiers is that we need cleaner benchmarks and more careful evaluation protocols to get a fair picture of how well these models are actually performing.
Conclusion: Tom: So we're wrapping up this discussion on "Multimodal Large Language Models as Image Classifiers" and talking about what this whole paper actually means for how we build these vision systems.
Jane: It basically shows that just having a big language model isn't enough for image classification; the way you test it, and the quality of the labels you give it, really matters.
Lu: The authors are pointing out that these models perform best when they get corrected labels, which can boost accuracy by up to ten point eight percent compared to what they used originally.
Meng: From a practical standpoint, this means we need better ways to curate our training data before we even let the AI see it. If the source material is noisy, the model will learn that noise instead of real patterns.
Lalam: And my perspective on this is that if we can use these models as assistants to help fix those mistakes in our datasets, we can actually improve how culture and content are recognized by these systems.
Tom: The main implication here is that the gap between these multimodal models and traditional supervised ones shrinks significantly when you use better data.
Jane: They also found that the way they set up their tests—Open-World versus Closed-World—changes how much difference you see in performance, depending on how the model is prompted.
Lu: The authors show that while some settings give a slight edge to closed-world testing, those gains disappear when you move to an open-world setup with proper mapping.
Meng: So for people actually building these products, this suggests we shouldn't rely on just one evaluation method; you have to design your pipeline around the specific task.
Lalam: If we can get better at correcting these labels, it means the AI becomes a more reliable partner in understanding visual information across different contexts.
Tom: It really makes you wonder how much more time we need to spend on data cleaning and protocol design versus just training bigger models.
More episodes
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck
- 2407.14562-Thought-Like-Pro: Enhancing Reasoning of Large Language Models through Self-Bootstrapped Prolog-based Chain-of-Thought