OODBench: Out-of-Distribution Benchmark for Large Vision-Language Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "OODBench: Out-of-Distribution Benchmark for Large Vision-Language Models".
Jane: The paper was written by Ling Lin, Yang Bai, Heng Su, Congcong Zhu, Yaoxing Wang et al. from University of Science and Technology of China and Suzhou Institute for Advanced Research, USTC and IHPC, A*STAR and Key Laboratory of the Ministry of Education for Mathematical Foundations and Applications of Digital Technology and Unmanned System Research Institute, Northwestern Polytechnical University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the show, everyone. Today we’re digging into a fresh arXiv paper that’s got a very practical title: “OODBench: Out-of-Distribution Benchmark for Large Vision-Language Models.” Jane, I know you’ve been itching to break this one down.
Jane: Oh, absolutely, Tom. And honestly, the title tells you half the story. “Out-of-distribution” is the key phrase. Think of it like this: you train a model on a million pictures of typical chairs. It gets really good at recognizing those. But then you show it a chair made of twisted tree roots. That’s out-of-distribution. The label is still “chair,” but the data looks different from anything it saw before.
Tom: Right, and that’s the gap this paper is tackling. The authors point out that a lot of the hype around these huge vision-language models, like GPT-4o and Gemini, is based on their performance on standard benchmarks. But those benchmarks are all in-distribution. They’re testing the model on the same kind of data it was trained on.
Jane: Exactly. So the paper, OODBench, is trying to answer a scary question: what happens when these models meet the messy, weird, real world? And the authors aren’t just theorizing. They built a massive benchmark with forty thousand image-category pairs to test it.
Tom: And the results are, frankly, a bit humbling. Even the best models, like GPT-4o, drop from over ninety percent accuracy on in-distribution data to around sixty-five percent on the hardest out-of-distribution data. That’s a huge cliff.
Jane: It really is. And what’s clever is how they define “out-of-distribution.” They aren’t just using rare objects. They’re using common objects in uncommon situations. Like a skateboard made of cake, or a fire hydrant that’s not the main focus of the photo. These are things a model might have seen the concept of, but not in that specific context.
Tom: So it’s not about testing if the model knows what a skateboard is. It’s about testing if it can handle a skateboard it’s never seen before.
Jane: Precisely. And that’s what makes this benchmark so important. It’s not a trivia test. It’s a stress test for real-world deployment. We’re talking about autonomous driving, medical imaging, any place where a wrong answer isn’t just a wrong answer, it’s a potential safety hazard.
Tom: That’s the high-stakes angle. And it sets the stage perfectly for the rest of the paper. We’ve got the problem, and we’ve got the benchmark. But how did they actually build it? That’s what we’re going to dig into next.
Summary: Jane: So, Tom, we’ve established that OODBench is a big deal because it tests models on data that’s shifted from their training set. But the really clever part is how they built it. They didn’t want to manually sort through forty thousand images, that would take forever.
Tom: Right, that would be a nightmare. So what did they do? They automated it.
Jane: They used a couple of existing, off-the-shelf models, like CLIP and BLIP2, as “OOD detectors.” The idea is that these models, because they’ve been trained on massive amounts of data, have a good sense of what’s “normal.” So if you show them an image of a car, and they’re very confident it’s a car, that’s in-distribution. But if they’re confused, or if they think a non-existent object is in the image, that flags it as out-of-distribution.
Tom: So they’re using one AI to find the blind spots of another AI. That’s a wild concept.
Jane: It is, and it’s brilliant. And to make it even more robust, they used two different detectors. If both CLIP and BLIP2 agree that something is OOD, they call it “OOD-Hard.” If only one of them flags it, it’s “OOD-Simple.” This gives them a spectrum of difficulty.
Tom: That’s a smart way to grade the data. And I’m guessing the results show that the “OOD-Hard” data is where the models really struggle.
Jane: You guessed right. The paper shows a consistent pattern across all the models they tested. Performance on in-distribution data is high, it drops a bit on OOD-Simple, and then it falls off a cliff on OOD-Hard. For some models, like LLaVA-NeXT, the recall on OOD-Hard data is below fifty percent, which is worse than random guessing.
Tom: Worse than a coin flip. That’s not just a small error, that’s a fundamental failure to see the object at all.
Jane: Exactly. And that’s the most concerning finding. It’s not that the models are mislabeling things, it’s that they’re completely missing them. In a self-driving car, missing a pedestrian is a lot worse than mislabeling a car.
Tom: That’s the real-world impact right there. But the paper doesn’t just stop at showing the problem. They also have a new way to evaluate it, right?
Jane: Yes, they call it the Basic-to-Advanced Progression metric, or BAP. It’s not just asking “is there a chair?” It’s a three-stage test. First, can you see the chair? Then, how many chairs are there? And finally, can you compare that number to the number of tables? It’s testing not just recognition, but counting and logical reasoning.
Tom: So it’s a much deeper dive into the model’s cognitive abilities.
Jane: Exactly. And the results show that as the questions get harder, the performance drops even further. Even on in-distribution data, the logical reasoning accuracy is low. On OOD-Hard data, it’s abysmal. This tells us that the problem isn’t just seeing the object, it’s understanding the scene.
Tom: So we have a benchmark, we have a new metric, and we have a clear picture of the problem. The next question is, can we fix it? Can we make these models more robust? Let’s talk about what the paper suggests.
Improvements: Tom: So, Jane, we’ve seen the problem. Models are great on familiar data, but they fall apart when things get weird. The paper shows this with OODBench. But what’s the fix? What do the authors suggest we do about it?
Jane: Well, Tom, one of the first things they tried was a popular technique called Chain-of-Thought prompting. You basically ask the model to “think step-by-step” before giving an answer. It’s supposed to improve reasoning.
Tom: And I’m guessing it didn’t magically solve everything.
Jane: Not at all. And this is where it gets really interesting. For some models, like Gemini and InternVL2, Chain-of-Thought actually helped on the hardest OOD data. It gave them a boost of about ten percent.
Tom: So it works for some?
Jane: For some, yes. But for others, like GPT-4o and LLaVA-NeXT, it actually made things worse. Their accuracy on in-distribution data dropped significantly. It’s like the model starts overthinking and second-guessing itself on data it should know well.
Tom: That’s a fascinating and counterintuitive result. So the fix isn’t as simple as just asking the model to think harder.
Jane: Exactly. The paper argues that this is because the model is trying to reason with faulty premises. If it can’t see the object in the first place, asking it to reason about it step-by-step just reinforces the hallucination. It’s building a logical argument on a foundation of sand.
Tom: So the improvement isn’t a new technique, but a new way of thinking about the problem. We need to focus on the perception part first.
Jane: Right. The paper’s main contribution is really the diagnostic tool. OODBench gives us a way to clearly see where these models are failing. And the analysis shows that the failures are often in the foundational visual understanding, not in the high-level reasoning.
Tom: That’s a huge insight. It means we can’t just scale up the language model part and hope for the best. We need to improve the visual encoder, or the way the model aligns images with text.
Jane: And that’s a much harder problem. It’s not about more data, it’s about better data. Or maybe it’s about a different training objective altogether. The paper doesn’t claim to have the solution, but it gives the community a clear roadmap of where to look.
Tom: So it’s a call to action. A benchmark that shows us exactly where we need to focus our research efforts.
Jane: And that’s incredibly valuable. It’s like a doctor giving you a precise diagnosis. You can’t treat the illness until you know what it is. OODBench is that diagnosis for our current generation of vision-language models.
Tom: Well said. So we have the problem, the diagnosis, and a clear direction for future work. Let’s wrap this up and see what the big picture is.
Conclusion: Jane: So, Tom, we’ve spent this whole episode on OODBench, and I think we’ve only scratched the surface. But the core message is clear.
Tom: It is. And it’s a message that should make us all a little more cautious about the hype. The paper shows that our most advanced vision-language models are still incredibly brittle. They can ace a test on familiar data, but they stumble on the messy, unpredictable data of the real world.
Jane: And that’s not just an academic problem. We’re talking about putting these models in cars, in hospitals, in our homes. If they can’t handle a chair that looks a little different, how can we trust them with tasks that are truly life-or-death?
Tom: That’s the sobering reality. But the paper, OODBench, isn’t just a doom-and-gloom report. It’s a tool. It gives us a way to measure this fragility and track our progress as we try to fix it.
Jane: And that’s the real gift. It’s a clear, automated, and comprehensive way to stress-test these models. It’s not perfect, and the authors admit it’s limited to static images in a few domains. But it’s a massive step forward from the benchmarks we had before.
Tom: So we’re saying goodbye to OODBench, but we’re taking its lessons with us. The lesson that seeing isn’t the same as understanding, and that our models have a long way to go before they can truly navigate the world.
Jane: Couldn’t have said it better myself. It’s a fantastic piece of research that will hopefully push the entire field towards building more robust, more reliable, and ultimately safer AI systems.
Tom: And on that note, we’ll wrap up our discussion on OODBench. Thanks for listening, everyone. We’ll be back next time with another paper from the arXiv. Until then, stay curious.
Ling Lin, Yang Bai, Heng Su, Congcong Zhu, Yaoxing Wang, Yang Zhou, Huazhu Fu, Jingrun Chen
University of Science and Technology of China · Suzhou Institute for Advanced Research, USTC · IHPC, A*STAR · Key Laboratory of the Ministry of Education for Mathematical Foundations and Applications of Digital Technology · Unmanned System Research Institute, Northwestern Polytechnical University
cs.CV, cs.AI, cs.DB
Submitted: 2026-08-16
Updated: 2026-08-18
Comments: 54 pages, 21 figures
Code: https://github.com/ultralytics/ultralytics
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 31/100
Key concepts
- Out-of-Distribution (OOD) Data
- This refers to data that is different from what a model was trained on. For example, if a model is trained on typical chairs, OOD data might be a chair made of twisted tree roots. The benchmark tests if the model can handle these novel contexts.
- OOD-Simple and OOD-Hard
- These are categories used to grade the difficulty of out-of-distribution data. OOD-Simple involves common objects in uncommon situations, while OOD-Hard involves things like a skateboard made of cake, representing a greater challenge for the model.
- Basic-to-Advanced Progression (BAP) metric
- This is a three-stage evaluation method used to test cognitive abilities beyond simple recognition. It first checks if the object can be seen, then how many there are, and finally compares that count to another object's count.
- Chain-of-Thought Prompting
- This technique asks a model to 'think step-by-step' before answering. While it helped some models with hard OOD data, it made others worse by causing them to second-guess themselves on familiar data.
Terminology
Summary
Summary
This paper introduces OODBench, a benchmark designed to evaluate the performance of Large Vision-Language Models (VLMs) on out-of-distribution (OOD) data. The authors argue that while current VLMs show strong performance on in-distribution (ID) data, their behavior under OOD conditions is under-evaluated due to a lack of suitable benchmarks.
The paper defines OOD data as samples not belonging to the training data distribution, categorizing them into semantic shift (labels change) and covariate shift (labels remain the same but data distribution changes). The authors focus on collecting covariate shift OOD data, based on the assumption that existing VLMs are likely trained on the most common classes of data. They define OOD data from a human perception perspective as: 1) Objects in images that are neither main objects nor semantically related to the main semantic object; 2) Variants or anomalous forms of target objects.
To construct the benchmark, the authors propose a predominantly automated method with minimal human verification. They use off-the-shelf VLMs like CLIP and BLIP2 as generalized OOD detectors. The process involves inputting images and category labels into the detector to obtain logits, applying a purify
operation to eliminate interference between multiple labels in an image, and then identifying OOD data based on two failure detection cases: when the probability of a non-occurring label is higher than all occurring labels, or when the probability of the actual label is below a hyperparameter threshold T. To mitigate bias from a single detector, they use a cross-validation scheme, defining the intersection of OOD data detected by both CLIP and BLIP2 as OOD-Hard data and the symmetric difference as OOD-Simple data.
OODBench contains about 40K yes-or-no samples, with OOD-S being 22K and OOD-H being 18K. The data is collected from four datasets across two key scenarios: natural scenarios (COCO, LVIS) and autonomous driving (nuScenes, Cityscapes). The authors also propose a Basic-to-Advanced Progression (BAP) Metric
to evaluate VLMs on three dimensions: existential questions (recognition), counting questions (quantity perception), and logical reasoning questions. The BAP metric is calculated as C/N × 100%, where C is the number of correct answers and N is the total number of questions.
The main results show that current state-of-the-art VLMs, including open-source models like LLaVA-NeXT, DeepSeek-VL, InternVL2, InternVL2.5, Llama-3.2-Vision, Qwen2-VL, and closed-source models like Gemini and GPT-4o, exhibit a 20% to 30% accuracy drop on OOD-H data relative to ID data. For example, GPT-4o achieves only approximately 65% accuracy on OOD-H data compared to over 90% on in-distribution samples. The authors also find that many VLMs show poor recall on OOD-H data, with some falling below 50% random chance, which is particularly problematic in safety-critical settings like autonomous driving.
The paper also investigates whether Chain-of-Thought (CoT) prompting helps on OOD data. Results show that CoT yields mixed results: it improves performance for some models (Gemini, InternVL2, InternVL2.5) on OOD-H data but causes performance degradation for others (GPT-4o, DeepSeek-VL, Llama-3.2-Vision). The BAP evaluation reveals a consistent decline in performance as task difficulty increases from E-Acc to C-Acc to L-Acc, and this trend becomes more pronounced as data shifts from ID to OOD-S to OOD-H.
Error analysis on GPT-4o shows failures primarily cluster into two categories: non-primary semantic objects and semantic variants. The authors conclude that current VLMs still perform poorly when confronted with OOD data, even when the classes of these OOD data are common in natural scenarios, and they hope OODBench will encourage future research on safer, more reliable multimodal systems.
Improvements for AI systems
Based on the paper, here are specific improvements that can be made to AI systems, particularly Vision-Language Models (VLMs):
Improvement: Integrate a cross-validation mechanism using multiple generalized OOD detectors (e.g., CLIP and BLIP2) rather than relying on a single detector. The intersection of their detections should be flagged as Hard OOD
and the symmetric difference as Simple OOD.
What the improved system can do:
-
Distinguish between mild and severe distribution shifts
-
Prioritize safety warnings for Hard OOD instances in real-time applications
-
Reduce false positives from single-detector bias
Improvement: Before computing match probabilities, implement a purify
step that isolates each label's logit by setting all other label logits to negative infinity, eliminating cross-label interference during softmax.
Improvement: Structure the system's self-assessment into three progressive stages: (1) existential recognition (Does X exist?
), (2) counting perception (How many X?
), and (3) logical reasoning ("Is count of X > count of Y?"). Award credit only when the model correctly answers each stage.
Improvement: For each instance, ask both Does this image contain X?
and Does this image not contain X?
with opposite ground-truth labels, ensuring balanced label distribution.
Improvement: During training, explicitly include non-main semantic objects and semantic variants (e.g., a skateboard made of cake) in the training distribution, rather than only common object forms.
Improvement: When the model's predicted probability for a known category falls below a threshold (T=0.05 as recommended), flag the input as OOD and either reject it or escalate to a human operator rather than making a high-confidence false prediction.
Improvement: Before finalizing an answer on OOD data, require the model to generate intermediate reasoning steps and verify that each step is grounded in the image. If the reasoning chain contains unverifiable assumptions, flag the response as low-confidence.
Improvement: Evaluate the model's performance across different parameter scales (e.g., 2B vs. 7B) and reject the assumption that larger models are inherently more robust to OOD. Instead, use the OODBench protocol to test each model size independently.
Abstract
Existing Visual-Language Models (VLMs) have achieved significant progress by being trained on massive-scale datasets, typically under the assumption that data are independent and identically distributed (IID). However, in real-world scenarios, it is often impractical to expect that all data processed by an AI system satisfy this assumption. Furthermore, failure to appropriately handle out-of-distribution (OOD) objects may introduce safety risks in real-world applications (e.g., autonomous driving or medical assistance). Unfortunately, current research has not yet provided valid benchmarks that can comprehensively assess the performance of VLMs in response to OOD data. Therefore, we propose OODBench, a predominantly automated method with minimal human verification, for constructing new benchmarks and evaluating the ability of VLMs to process OOD data. OODBench contains 40K instance-level OOD instance-category pairs, and we show that current VLMs still exhibit notable performance degradation on OODBench, even when the underlying image categories are common. In addition, we propose a reliable automated assessment metric that employs a Basic-to-Advanced Progression of prompted questions to assess the impact of OOD data on questions of varying difficulty more fully. Lastly, we summarize substantial findings and insights to facilitate future research in the acquisition and evaluation of OOD data.
Sources
- Qwen2.5-VL Technical Report
- Hallucination of Multimodal Large Language Models: A Survey
- MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
- PaLI-X: On Scaling up a Multilingual Vision and Language Model
- GPT-4 Technical Report
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
- How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
- Emerging Properties in Unified Multimodal Pretraining
- LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
- The Llama 3 Herd of Models
- A Survey on Hallucination in Large Vision-Language Models
- GPT-4o System Card
- Towards Out-Of-Distribution Generalization: A Survey
- LLaMA: Open and Efficient Foundation Language Models
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- A Unified Survey on Anomaly, Novelty, Open-Set, and Out-of-Distribution Detection: Solutions and Future Challenges
- Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models