Unified Hallucination Fuzzing for Multimodal Large Language Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Unified Hallucination Fuzzing for Multimodal Large Language Models".
Jane: The paper was written by Pengfei Zhou, Jiajun Song, Zhiwei Tang, Yixing Ma, Xiaopeng Peng et al. from National University of Singapore and DAMO Academy, Alibaba Group and Renmin University of China and University of California, Berkeley and Zhejiang University and Hupan Lab and Rochester Institute of Technology and The Hong Kong University of Science and Technology.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everybody. We are diving headfirst into a new paper today, and the title alone is a mouthful: “Unified Hallucination Fuzzing for Multimodal Large Language Models.” Jane, I gotta say, just reading that title makes me think of two things at once — hallucinations and fuzzing. Those don’t usually go together.
Jane: Right, Tom. And that’s exactly why I’m so excited. We’ve talked about AI models making things up before, but this paper is trying to systematically *break* them to find out *when* and *why* they do it. It’s like instead of just asking a model questions, you’re actively trying to trick it into lying.
Tom: So it’s not just a new benchmark, it’s a whole new way of stress-testing. I love that. And the authors — this is a huge team. I’m seeing folks from National University of Singapore, Alibaba’s DAMO Academy, Zhejiang University, Berkeley, HKUST. This is a serious collaboration.
Jane: It really is. And having that kind of firepower behind it means they could build something pretty comprehensive. They’re not just looking at one type of mistake, like “is there a cat in the picture.” They’re looking at object-level, instruction-level, and knowledge-level hallucinations. That’s a much bigger net.
Tom: A bigger net, and a sharper one. The idea that they’re using “fuzzing” — that’s a term from software security, right? You throw random garbage at a program to see if it crashes. They’re doing the same thing to these multimodal models, but with *intelligent* garbage.
Jane: Exactly. And that’s what makes this feel like a real shift. Static benchmarks are getting saturated. Models are memorizing the answers. This paper is saying, “Okay, let’s stop giving them the same test and start actively probing for the weak spots.” I think this could change how we evaluate AI trustworthiness.
Tom: And that’s the big implication for me. If we can’t trust the evaluation, we can’t trust the model in a hospital or a courtroom. This feels like a step toward building that trust, by first finding all the ways it can break.
Jane: For sure. And I can’t wait to see what they actually found when they ran these tests. Let’s get into the summary next.
Summary: Tom: So, Jane, we’ve got the title, we’ve got the big names behind it. Let’s talk about what they actually did. The paper’s summary, “Unified Hallucination Fuzzing for Multimodal Large Language Models,” lays out a pretty bold plan.
Jane: It does. They built a new benchmark called UniHall, and it’s not just a pile of questions. It’s organized around this unified taxonomy — Object, Instruction, and Knowledge hallucinations. Each one has subtypes. For objects, it’s things like existence, attributes, relations. For instructions, it’s context, refusal, sycophancy.
Tom: Sycophancy. That’s the one where the model just agrees with you to be nice, even if you’re wrong. I love that they’re calling that out as a hallucination. It’s not just about facts, it’s about behavior.
Jane: Exactly. And then they have the knowledge dimension, which is about fabricating facts or citations. So they’ve got this really detailed map of all the ways a model can lie. But the benchmark is just the starting point.
Tom: Right, because they don’t just leave it static. They built this framework called SAMF — Self-Adaptive Multimodal Fuzzing. And this is the clever part. It takes the seed questions from UniHall and then *mutates* them. It changes the image, it changes the prompt, it adds distractions, it adds conflicting information.
Jane: And it does this adaptively. It’s not random. It learns which mutations are most likely to trip up a specific model. So you get a test that’s tailored to find the weaknesses of, say, GPT-five versus Gemini.
Tom: That’s the “self-adaptive” part. It’s like a coach who watches you play and then designs drills specifically to exploit your bad habits. And they have both a heuristic version and a reinforcement learning version of this fuzzer.
Jane: Right. And the RL version is much faster at finding the breaking points. They report it’s about sixteen times faster to train than the heuristic one. That’s a huge practical advantage.
Tom: So they built the map, they built the stress test, and then they ran a ton of models through it. And the results are, frankly, a little scary. Let’s get into the details of what they found.
Improvements: Tom: So we’re back with “Unified Hallucination Fuzzing for Multimodal Large Language Models,” and Jane, we’ve talked about the benchmark and the fuzzing framework. What’s the big improvement this paper is suggesting over the old way of doing things?
Jane: I think the biggest one is that they’re moving from a static test to a dynamic one. Old benchmarks, like POPE or HallusionBench, they’re like a final exam. You study for it, you pass it, but it doesn’t mean you know the material. This paper is proposing a pop quiz, every single time.
Tom: A pop quiz that’s written specifically to be hard for you. That’s the key improvement. They’re not just measuring accuracy on a fixed set of questions. They’re measuring robustness against a constantly evolving set of adversarial inputs.
Jane: And they’ve got the metrics to back it up. They introduce this whole suite of new scoring methods — GHR, BHR, SHR, GHS. They’re not just saying “right or wrong.” They’re trying to measure *how* wrong, and *what kind* of wrong.
Tom: Yeah, the Breakdown Hallucination Rate is interesting. It doesn’t just say “this answer is a hallucination.” It breaks the answer down into individual claims and checks each one against evidence. So you can see if the model got the object right but the relationship wrong.
Jane: That’s a huge improvement in diagnostic power. It’s like the difference between a doctor saying “you’re sick” and a doctor saying “you have a bacterial infection in your left lung, but your right lung is fine.” That level of detail is what you need to actually fix the problem.
Tom: And they also have this risk-aware aggregation. They don’t treat a hallucination about a car in a parking lot the same as a hallucination about a tumor in a medical scan. They weight the errors by how much harm they could cause.
Jane: That’s so important for real-world deployment. It’s not enough to just have a low error rate. You need to have a low error rate on the things that matter. This paper is pushing the field to think about that.
Tom: So the improvements are about being more thorough, more precise, and more aware of consequences. That’s a big step. Now let’s look at the actual first page and see how they set all this up.
First Page: Tom: Alright, Jane, we’ve been talking about the big ideas. Let’s get down to the nitty-gritty of the first page of “Unified Hallucination Fuzzing for Multimodal Large Language Models.” What’s the setup?
Jane: The first page is all about the problem. They start by saying that even the best models, like GPT-five and Gemini, still hallucinate. And that’s a huge problem for high-stakes stuff like healthcare and law. You can’t have a model confidently making things up in a courtroom.
Tom: And they make a really good point about why current benchmarks aren’t enough. They say these static benchmarks are getting saturated. Models are just memorizing the answers. They call it a “false sense of security.” I love that phrase.
Jane: It’s so true. You get a model that scores ninety-five percent on a benchmark, and you think it’s great. But then you put it in the real world, and it falls apart because the real world isn’t a benchmark. It’s messy and unpredictable.
Tom: So they’re arguing for a paradigm shift. They want to go from static measurement to dynamic, adversarial stress testing. And that’s what the whole paper is about. They even show a figure on the first page that lays out their taxonomy.
Jane: Right, the UniHall taxonomy. It’s a nice visual. It shows the three main branches — Object, Instruction, Knowledge — and then the subtypes under each one. It’s a clean way to organize the chaos of all the different ways models can be wrong.
Tom: And they also mention this “helpfulness-hallucination trade-off.” That’s a fascinating finding. They’re saying that models that are heavily trained to be helpful, through reinforcement learning, tend to be more sycophantic. They’ll make things up just to please the user.
Jane: That’s the alignment tax. You optimize for one thing, being helpful, and you accidentally make another thing worse, being truthful. It’s a real dilemma for the people building these models.
Tom: It really is. So the first page sets up the problem, introduces the taxonomy, and hints at these big findings. It’s a strong opening. I’m really curious to see how they actually tested all these models and what the numbers look like.
Jane: Me too. But for now, we’ve got a great picture of what this paper is trying to do. It’s a call to action for the whole field.
Conclusion: Tom: Well, Jane, we’ve spent some time with “Unified Hallucination Fuzzing for Multimodal Large Language Models,” and I think it’s fair to say this one’s a game-changer.
Jane: Absolutely. We started with the problem — static benchmarks are failing us. Then we saw their solution — a comprehensive taxonomy, a dynamic fuzzing framework, and a suite of new metrics. It’s a complete package.
Tom: And the findings are sobering. They showed that even the best models degrade significantly under fuzzing. The “reasoning-grounding dissociation” is a real thing. A model can be great at logic puzzles but terrible at sticking to the facts of an image.
Jane: And that alignment tax we talked about is a big deal. It means we can’t just blindly optimize for helpfulness. We need to be more careful about how we train these things.
Tom: For me, the biggest takeaway is that we need to stop trusting leaderboards. They’re not telling us the whole story. This paper gives us the tools to find the real weaknesses, the ones that matter in the real world.
Jane: Exactly. And that’s what makes this work so important. It’s not just about making better benchmarks. It’s about making safer, more reliable AI. The fact that they’re thinking about risk levels and real-world harm is a huge step in the right direction.
Tom: So we’re saying goodbye to this paper, but we’re taking its message with us. We need to be more skeptical, more thorough, and more creative in how we test our AI systems.
Jane: Well said, Tom. It’s been a great discussion. To all our listeners, thanks for tuning in. We’ll be back soon with the next paper to break down.
Tom: Take care, everyone. And remember, don’t trust everything a model tells you. Especially if it’s trying to be helpful.
Pengfei Zhou, Jiajun Song, Zhiwei Tang, Yixing Ma, Xiaopeng Peng, Donghui Si, Yuhang Xu, Huiqi Song, Yiyuan Miao, Yichen Qian, Weihua Chen, Wangbo Zhao, Bohan Zhuang, Jiasheng Tang, Yang You
National University of Singapore · DAMO Academy, Alibaba Group · Renmin University of China · University of California, Berkeley · Zhejiang University · Hupan Lab · Rochester Institute of Technology · The Hong Kong University of Science and Technology
cs.CL, cs.AI
Submitted: 2026-07-15
Updated: 2026-08-11
Comments: 47 pages, 17 figures
Code: https://github.com/LanceZPF/EvalHall
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 43/100
Key concepts
- Unified Hallucination Fuzzing
- This is a method of actively probing or 'stress-testing' AI models. Instead of using standard questions, the researchers apply intelligent, adversarial inputs to systematically break the models and discover exactly when and why they generate false information.
- UniHall Taxonomy
- This is a detailed classification system used in the paper to categorize how models fail. It organizes hallucinations into three main types: Object (errors related to image attributes), Instruction (errors related to context or refusal), and Knowledge (fabricating facts or citations).
- Self-Adaptive Multimodal Fuzzing (SAMF)
- This is the core framework that applies the fuzzing. It takes initial test questions and adaptively mutates them—changing images, adding distractions, or conflicting data—to specifically target and exploit a model's weaknesses.
- Helpfulness-Hallucination Trade-off
- This concept describes a dilemma where models trained to be highly helpful (often through reinforcement learning) become sycophantic. They tend to make things up just to agree with the user, sacrificing factual accuracy for perceived politeness.
Terminology
Summary
Affiliations: National University of Singapore, DAMO Academy Alibaba Group, Renmin University of China, Zhejiang University, Hupan Lab, University of California Berkeley, Rochester Institute of Technology, The Hong Kong University of Science and Technology
Published: arXiv:2608.07525v1 [cs.CL] 15 Jul 2026
The paper addresses the persistent challenge of hallucination in Multimodal Large Language Models (MLLMs), which severely limiting their reliability in high-stakes applications.
The authors argue that existing evaluations, predominantly based on static benchmarks, suffer from narrow taxonomical coverage and rapid performance saturation, failing to reflect model robustness in evolving real-world scenarios.
The paper states: "While numerous benchmarks have been proposed to quantify this phenomenon, the current evaluation paradigm becomes insufficient due to coverage limitations and performance saturation. Previous efforts mainly rely on static datasets with fixed inputs and predefined answers. As models evolve, they rapidly saturate these benchmarks, effectively memorizing the exam without acquiring true robustness. Such evaluations also fail to capture the vast long-tail distribution of real-world hallucinations and cannot systematically probe model vulnerabilities under adversarial or complex instruction mutations. Consequently, high benchmark scores often mask underlying brittleness, creating a false sense of security that necessitates a paradigm shift from static measurement to dynamic, adversarial stress testing."
The authors summarize their contributions as follows:
-
We introduce UniHall, a risk-aware benchmark built on a unified Object–Instruction–Knowledge hallucination taxonomy.
-
We propose SAMF, the first mutation-driven MLLM fuzzing framework with adaptive evolution and structured oracle verification.
-
We reveal an alignment tax through systematic stress testing, showing that helpfulness-oriented RL can reduce instruction-level faithfulness.
The paper states: "We first propose a comprehensive taxonomy that classifies hallucinations into three distinct dimensions: Object-level (e.g., fabricated existence, attributes), Instruction-level (e.g., failure to refuse, sycophancy), and Knowledge-level (e.g., fictitious facts)."
The taxonomy includes:
-
Object-level Hallucination (≥1000 instances): Existence Hallucination (200+), Attribute Hallucination (600+ with subtypes: Category, Number, Spatial, Size, State, Color/Other Attribute), Relation Hallucination (200+)
-
Instruction-level Hallucination (≥500 instances): Context Hallucination (200+), Failure-to-refuse Hallucination (200+), Sycophantic Hallucination (200+)
-
Knowledge-level Hallucination (≥500 instances): Factual Hallucination (200+), Reference Hallucination (200+), Detail Hallucination (200+)
According to Table 2: Total instances: 2,170. Question Formats: Yes-or-No (YON): 900, Open-ended VQA (free-form): 1,270. Avg G: 1.65, Avg N: 0.84. Prompt length: 18/1,455/122 (Min/Max/Avg). Prompt Length After fuzzing: 135/29,086/2,042.
The paper states: The dataset integrates widely used open-domain benchmarks (e.g., Hallu-PI and BEAF) with domain-specific resources.
The annotation workflow involves: "1. Source Allocation: 3 Meta-reviewers assign relevant raw source data (overprovisioned by 5-10× the target per category) to 10 annotators. 2. ID and Taxonomy Assignment: Annotators map each instance to a standardized instance id and annotate metadata field. 3. Seed Data Preparation: Each instance is classified as Yes/No (YON) or open-ended VQA. 4. Quality Control: Annotators cross-check image–text alignment, taxonomy consistency, and visual clarity."
The paper states: "SAMF functions as an automated red-teaming mechanism, employing evolutionary strategies including reinforcement learning to generate semantically preserved yet adversarial mutations. To enable scalable and reliable assessment of these dynamic inputs, SAMF incorporates a structured hallucination metric driven by complementary oracles, combining symbolic rules, detectors, and LLM verifiers in a coarse-to-fine pipeline."
Textual Mutations target "cognitive robustness and reasoning stability via (i) redundancy/interference control (e.g., irrelevant context expansion, constraint stacking), (ii) complexity traps (e.g., syntactic complication, explicitly negated distractors), and (iii) knowledge-conflict induction (e.g., authority-bias injection)."
Specific operators include:
-
(Easy) Irrelevant Context Expansion
-
(Common) Constraint Stacking
-
(Medium) Syntactic Complication
-
(Hard) Disturbance Exclusion
-
(Extreme) Authority Bias Injection
Visual Mutations stress "attention control and selective grounding via (i) background distraction injection, (ii) redundant multi-view copies with controlled image quality degradation, and (iii) multi-image overload with selective grounding constraints."
Specific operators include:
-
(Easy) Background Multimodal Interference Injection
-
(Common) Multi-image Copy Variant 1
-
(Medium) Multi-image Copy Variant 2
-
(Hard) Multi-image Copy Variant 3
-
(Extreme) Multi-image Composition with Selective Grounding
The paper states: "SAMF leverages all mutation operators for extensive and fine-grained robustness evaluation. For strategy discovery, we use a 20% pilot subset of the benchmark; on this dynamic pilot set, our basic implementation performs operator search via a heuristic algorithm to balance search efficiency with hallucination coverage improvements."
For RL-based fuzzing: "RL agents learn adaptive policies that optimize operator selection and mutation intensity from oracle feedback; compared to heuristic strategies, RL-based fuzzing offers better long-horizon exploration and parallel-training efficiency for model-specific weakness discovery. Empirically, both are effective, but RL achieves higher hallucination discovery efficiency and more stable cross-model coverage."
The RL approach uses a Multi-Armed Bandit formulation: We formulate policy search as a multi-armed bandit problem where each arm represents a candidate fuzzing policy. For each (s, q) group, we maintain a pool of K policies (default K = 1024).
Two bandit algorithms are supported: UCB and Epsilon-Greedy.
Training efficiency: "On a single node with 8 × NVIDIA H20 (96GB) GPUs, our RL-based fuzzing controller is trained in 37m28s (2248s) using 8-GPU data parallelism. In contrast, our heuristic fuzzing baseline is trained on a single GPU on one node and takes 10h31m4s (37864s). The significant reduction in training time of RL-based fuzzing represents a speedup of approximately 16.8×."
The paper proposes a multi-layered metric suite with (i) multi-stage deterministic logic, (ii) claim-level structural decomposition, and (iii) oracle-based verification grounded in evidence sets.
GHR addresses the failure of surface-level matching by employing asymmetric decision logic to detect unsupported but fluent responses.
The decision function is defined as:
-
1 if SimNLP(y, N) ≥ τ1
-
1 if JudgeGPT(y, N) = TRUE
-
0 if Simtok(y, G) ≥ τ3
-
0 if SimNLP(y, G) ≥ τ2
-
0 otherwise
BHR utilizes deterministic oracles φ to verify claims against evidence E = A, C, E, V (Answers, Constraints, External knowledge, Visual facts).
Claims are decomposed into atomic claims represented as tuples: ci = (typei, framei, spani), with typei ∈ Object, Instruction, Knowledge.
SHR generalizes BHR by replacing NLP components with GPT-based judges, including claim decomposition C(y) and oracle judgments φ(·).
Each instance receives a continuous score: SHR(y) = (1/C(y)) Σ s(ci), where s(ci) ∈ [0, 1] reflects the severity of hallucination inferred from the judge output.
GHS directly captures subjective severity through graded GPT-based scoring. Each response is scored on a discrete 0–10 scale and normalized to [0, 1].
Defined as: M risk = (Σ x∈D rx · M(x)) / (Σ x∈D rx)
where rx is the instance-level risk.
Models include: Qwen2.5-VL-7B, Qwen3-VL-8B, LLaVA-Next-Llama3-8B, LLaVA-v1.6 (Mistral-7B, Vicuna-7B, Vicuna-13B, 34B), GLM-4.1V-9B-Base, GLM-4.6V-Flash, InternVL3.5 (8B, 14B, 30B-A3B, 38B with variants), MiMo-VL-7B (with RL, RL+Thinking, SFT, SFT+Thinking variants), GPT-5 series (nano, mini, 5, 5.1, 5.2), Gemini-2.5 series (Flash, Pro), Gemini-3 series (Flash, Pro).
The Reasoning-Grounding Dissociation: "Our results challenge the assumption that stronger general-purpose reasoning yields better hallucination resistance. In the open-source domain, Qwen3-VL-8B, despite better performance on standard VQA benchmarks and reasoning tasks than Qwen2.5-VL-7B, consistently shows higher hallucination rates across the Object and Knowledge dimensions. Similarly, in the proprietary models, GPT-5.2 exhibits higher hallucination rates than its predecessors (GPT-5.0/5.1). This suggests that iterations emphasizing complex reasoning can inadvertently weaken factual grounding, especially as longer reasoning rationales may introduce more unsupported claims."
The Alignment Tax: "We observe that aggressive reinforcement learning (RL) alignment, while improving conversational fluency, can exacerbate hallucinations. For example, models such as Gemini-Pro often produce more comprehensive and verbose responses under RL incentives, increasing opportunities for unsupported details. To satisfy reward preferences for helpfulness, models may fabricate plausible information rather than answer concisely or admit uncertainty. We term this phenomenon the alignment tax: improving perceived helpfulness can come at the cost of factual faithfulness."
The Challenge of Knowledge Hallucination: Among the three dimensions, knowledge-level hallucinations show the highest variance and instability. Static benchmarks are ill-suited here because world facts change over time.
Metric Sensitivity and Verbosity Bias: "We observe a divergence between Generic Hallucination Rate (GHR) and structured metrics (SHR/BHR). Verbose models such as Gemini can be penalized by GHR because longer outputs raise the chance of triggering surface-level detectors. In contrast, SHR/BHR decompose responses into atomic claims before verification and suggest that Gemini is not substantially more prone to fabrication than GPT."
The paper states: All models show notable performance decreases from static to fuzzing evaluation. For example, Qwen2.5-VL-7B's overall accuracy drops from 0.702 to 0.635, while the generic hallucination rate rises from 0.298 to 0.365.
Specific results from Table 11:
-
Gemini3-Flash: GHR rises from 31.97% (Standard) to 51.27% (RL-based Fuzz)
-
Qwen2.5-VL-7B: GHR rises from 29.76% (Standard) to 42.65% (RL-based Fuzz)
-
GPT-5.2: GHR rises from 27.68% (Standard) to 38.74% (RL-based Fuzz)
Category-wise Robustness: "Object-level hallucinations are relatively well handled: models reach high accuracy (e.g., 0.910 and 0.830 for Qwen2.5-VL-7B) and low hallucination rates. By contrast, instruction-level challenges such as refusal and sycophancy degrade markedly under fuzzed inputs (e.g., the hallucination rate for Refuse exceeds 0.9)."
Failure Mechanisms: "First, overgeneralization: with visual noise or style transfer, models may generate false object claims associated with the style rather than the actual content. Second, context misalignment: complex instruction mutations can shift the model toward language priors, causing it to ignore the visual context and follow the prompt's semantic trajectory."
Sycophancy under Stress: "When prompts contain authoritative bias or misleading preconditions (e.g., Authority Bias Injection), models show increased sycophancy. In severe cases, instruction-level tasks such as Refusal degrade most under fuzzing, with hallucination scores exceeding 0.8. Instead of relying on visual evidence, models may agree with the prompt premise."
Table 4 shows: "GHS aligns best with human judgments, achieving the highest Accuracy/F1/Kappa and strongest Pearson/Spearman in both original and risk-weighted settings. This suggests that humans often assess hallucination severity through holistic subjective intuition, which is closer to graded judge-based scoring (GHS) than deterministic claim verification."
The paper concludes: "In this work, we addressed the limitations of static evaluation paradigms by introducing SAMF, a dynamic framework that stresses MLLMs. Grounded in the proposed comprehensive taxonomy and UniHall benchmark, our approach systematically challenges MLLMs across Object, Instruction, and Knowledge dimensions through semantically guided mutations. By integrating adaptive perturbations with an ensemble of oracle-based detectors, SAMF reveals latent vulnerabilities and robustness of current MLLMs. Ultimately, this work establishes a new framework for universal hallucination assessment, offering critical insights into the development of trustworthy multimodal systems."
The framework, code and benchmark are available at https://github.com/LanceZPF/EvalHall.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system, along with what the improved system can do:
Improvement: Implement a fuzzing engine that automatically mutates inputs (text and images) during evaluation, rather than relying on static benchmarks.
What the improved system can do:
-
Automatically generate adversarial test cases by applying operators like:
-
Textual: Irrelevant context expansion, constraint stacking, syntactic complication, disturbance exclusion, authority bias injection
-
Visual: Background distraction injection, multi-image copies with degradation, style transfer, selective grounding constraints
-
Use a reinforcement learning-based bandit controller (UCB or epsilon-greedy) to learn which mutation policies most effectively expose hallucinations for a specific model
-
Retrain the fuzzing controller in 37 minutes on 8 GPUs, enabling rapid adaptation to new model architectures
-
Achieve up to 16.8× faster policy search than heuristic hill-climbing methods
The improved AI system can:
-
Self-test for hallucinations using adaptive fuzzing without human intervention
-
Localize failures to specific subtypes (e.g.,
sycophancy under authority bias
) -
Quantify severity in a way that aligns with human judgment
-
Detect alignment tax from RL training
-
Avoid verbosity bias in evaluation
-
Verify knowledge against current, dynamic evidence
-
Signal when to refuse answering
-
Scale to new models rapidly (37-minute retraining of fuzzing policies)
Abstract
Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high-stakes applications. Existing evaluations, predominantly based on static benchmarks, suffer from narrow taxonomical coverage and rapid performance saturation, failing to reflect model robustness in evolving real-world scenarios. To bridge this gap, we present a systematic evaluation framework integrating a comprehensive benchmark with self-evolving stress testing. First, we introduce UniHall, a fine-grained dataset grounded in a unified taxonomy spanning Object, Instruction, and Knowledge dimensions. Second, to address benchmark saturation, we propose Self-Adaptive Multimodal Fuzzing (SAMF), a self-adaptive framework that employs evolutionary mutation strategies to explore the boundaries of model hallucinations. Crucially, to ensure reliable assessment of dynamic inputs, SAMF incorporates a structured metric suite driven by an ensemble of multi-modal oracles. Our extensive experiments reveal that state-of-the-art MLLMs exhibit significant performance degradation under fuzzing compared to conventional settings, exposing a dissociation between reasoning capabilities and factual grounding. Furthermore, we identify a helpfulness-hallucination trade-off, where reinforcement learning alignment inadvertently exacerbates sycophancy in instruction-following tasks. The framework, code and benchmark are available at https://github.com/LanceZPF/EvalHall.
Sources
- Qwen2.5-VL Technical Report
- Hallucination of Multimodal Large Language Models: A Survey
- Detecting and Evaluating Medical Hallucinations in Large Vision Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- A Survey on Hallucination in Large Vision-Language Models
- ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
- Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming
- MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models
- Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans
- Gemini: A Family of Highly Capable Multimodal Models
- AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- On-the-Fly Data Augmentation via Gradient-Guided and Sample-Aware Influence Estimation
- MedFrameQA: A Multi-Image Medical VQA Benchmark for Clinical Reasoning
- HallE-Control: Controlling Object Hallucination in Large Multimodal Models
- Poison as Cure: Visual Noise for Mitigating Object Hallucinations in LVMs
- Agent-as-a-Router: Agentic Model Routing for Coding Tasks
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering