DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts

arXiv:2412.10510 · cs.CV, cs.CL · Submitted 2024-12-13 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts".

Jane: The gist: DEFAME presents Dynamic Evidence-based FAct-checking with Multimodal Experts, a modular, zero-shot MLLM pipeline designed for open-domain,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So we're talking about DEFAME today. It's this new system for fact-checking that uses multimodal experts to handle claims involving both text and images. The authors say it’s a modular, zero-shot pipeline that dynamically picks the right tools and search depth to grab evidence.

Jane: Exactly. The main point of this paper is that it moves beyond just looking at text alone or just relying on some fixed knowledge base. They claim DEFAME performs end-to-end verification, meaning it actually checks the claim using both pictures and words, while generating a structured report for you.

Lu: What's really interesting is how they structure the process. It operates in six stages: planning actions, executing those actions with tools like Google Search or Reverse Image Search, summarizing the results, developing the fact-check based on that evidence, predicting a verdict and then justifying it all out.

Meng: That sounds complex for what it is trying to achieve. So how does this dynamic tool selection work in practice? Does it just pick tools randomly, or is there some kind of intelligence guiding which search depth to use?

Tom: Well, the planning stage lets the MLLM suggest a targeted action sequence first. Then you have the execution stage where it invokes tools like web search or image search based on what the plan suggested. It’s supposed to be adaptive, not just running one fixed process through everything.

Jane: And that leads into how they handle images specifically because they use tools like Google Vision API for reverse image searches. That lets the system actually analyze visual evidence in claims, which is a big step up from text-only systems.

Lalam: From my side, I see this as a way to improve cultural understanding. If we can build systems that look at images and text together reliably, it means we can process information in ways that are more nuanced than what simple language models or even older fact-checkers could handle on their own.

Tom: Right, and the results they show are pretty strong too. They’ve tested DEFAME on benchmarks like VERITE, AVERITEC, and MOCHEG. They claim it surpasses all previous methods on those tests for uni- and multimodal fact-checking.

Jane: They even mentioned an improvement of over ten percent accuracy on MOCHEG when they compared it to earlier work by Yao et al., two thousand twenty-three <ref:2412.10510#pg2>. That shows a tangible lift in how well it handles the verification task across different claim types.

Meng: But what about the limitations? The paper does flag a few things, including potential credibility issues with external evidence and system stability problems because of web scraping. That’s a practical hurdle for any real-world deployment I see right away.

Lu: They also point out that because it relies on an LLM for reasoning, there's always the risk of hallucinations creeping in, which they say needs more analysis in future work. It’s not perfect yet.

Tom: So, to wrap up this part of the DEFAME paper, it’s presented as a unified framework that handles multimodal claims natively and retrieves evidence on the fly. It sets a new standard for general fact-checking because it doesn't require task-specific training or tuning to work across different challenges.

Jane: The implication here is that for anyone dealing with information where pictures and text are mixed, this system offers a transparent way to get an answer, not just a black box output. It’s designed to mimic how a human fact-checker would actually work by giving you the steps.

Tom: So, what does this mean for us? It means we have a more robust tool for verifying complex information online than we did before, especially when visual context is involved. We'll keep an eye on those benchmark results as they settle.

Jane: And next time we talk about this paper, we’ll look at how these kinds of multimodal systems could actually be integrated into daily tools for people on the go. That’s what's coming up next.

Conclusion: Tom: So we're wrapping up on DEFAME today, focusing on what that title actually means for us in the real world.

Jane: It’s about taking those two big ideas—dynamic evidence and multimodal experts—and putting them together into one single system for fact-checking.

Lu: What they’ve done is create this modular pipeline that doesn't rely on a single way to check something. It just picks the right tools, like web search or image recognition, based on what it needs at that moment.

Meng: So you’re saying instead of running one giant script through everything, it plans what to do first and then executes those specific steps dynamically?

Tom: Exactly. It’s designed to look at a claim and figure out which piece of evidence—text or image—is missing, and then decide exactly which search tool is best for finding it.

Lalam: From my side, I see this as moving beyond just checking text. We can now verify claims that involve pictures in a way that feels much more like how a person would actually investigate something.

Jane: It really does give us a transparent process because you can see the steps they take to get to the final answer.

Tom: And when you look at the results, they’ve shown this system beating previous top methods on big fact-checking tests like VERITE and MOCHEG. That's a solid number showing it works better than what came before.

Lu: They even created a new test set, CLAIMREVIEW2024+, which is interesting because it checks claims that are newer than the models were trained on.

Meng: That’s important for making sure the system isn't just repeating old answers; it’s actually looking at current information.

Jane: What this means for us is that we have a more general fact-checking tool now, one that doesn't need to be specifically trained for every single type of claim.

Tom: And the challenge ahead is making sure this dynamic tool selection stays accurate and reliable when it's out there in the wild. That’s where we need to keep watching things.

Tobias Braun, Mark Rothermel, Marcus Rohrbach 1 Anna Rohrbach 1

Technical University of Darmstadt & hessian.AI

cs.CV, cs.CL

Submitted: 2024-12-13

Updated: 2026-10-03

Code: https://github.com/multimodal-ai-lab/DEFAME

Importance score: 88/100

The gist: The gist: DEFAME presents Dynamic Evidence-based FAct-checking with Multimodal Experts, a modular, zero-shot MLLM pipeline designed for open-domain, text-image claim verification that dynamically

Key concepts

DEFAME Framework
A six-stage process designed for end-to-end fact verification. It dynamically chooses the right tools (like search or image retrieval) and search depth needed to gather necessary textual and visual evidence, mimicking a human fact-checker.
Multimodal Experts
The system uses Large Multimodal Models (MLLMs) to handle claims that involve both text and images. These experts are prompted at several stages of the process to plan actions, summarize retrieved evidence, develop arguments, and predict verdicts based on the combined multimodal input.
Dynamic Evidence Retrieval
Instead of relying on fixed knowledge or text-only methods, DEFAME dynamically decides when and how to search for information. It invokes specific tools (like Google Search or Reverse Image Search) only when needed during the execution stage, making it flexible for open-domain claims.
Six Stages of Verification
The structured procedure DEFAME follows: Plan Actions, Execute Actions (using tools), Summarize Results, Develop the Fact-Check argument based on evidence, Predict a Verdict (looping back if information is insufficient), and finally Justify the Verdict with supporting links.

Terminology

Summary

The gist: DEFAME presents Dynamic Evidence-based FAct-checking with Multimodal Experts, a modular, zero-shot MLLM pipeline designed for open-domain, text-image claim verification that dynamically selects tools and search depth to extract and evaluate textual and visual evidence.

DEFAME Framework Overview

DEFAME operates in a six-stage process that dynamically selects the tools and search depth to extract and evaluate textual and visual evidence. Unlike prior approaches that are text-only, lack explainability, or rely solely on parametric knowledge, DEFAME performs end-to-end verification. It is designed with transparency in mind, imitating a human fact-checking process and returning a detailed fact-check report to the user. The framework is described as an end-to-end AFC framework that natively processes multimodal claims and evidence while retrieving the latter dynamically as needed.

The Six Stages of Verification

The procedure is decomposed into six manageable stages, five of which are subject to MLLM prompting. The stages are:

  1. Plan Actions. In this stage, the MLLM is prompted to suggest a targeted action sequence to retrieve missing information.

  2. Execute Actions. Given a set of actions, DEFAME invokes the corresponding tool. Tools include Web Search via Google Search, Image Search via Google Image Search, Reverse Image Search (RIS) using the Google Vision API, and Geolocation integrating GEOCLIP.

  3. Summarize Results. At this stage, the gathered evidence is integrated into the fact-checking report. The model generates an abstractive summary of key findings for each tool output.

  4. Develop the Fact-Check. Corresponding to Stage 4 in Moreno Gil et al. (2021), DEFAME brings claim and summarized evidence together. It directs the MLLM to discuss the claim’s veracity step-by-step based on the evidence.

  5. Predict a Verdict. Next, DEFAME classifies the claim into one of the benchmark-specific categories by prompting the MLLM to summarize key findings and select a verdict. If NEI (Not Enough Information) is returned, the system loops back to Stage 1 to retrieve additional evidence.

  6. Justify the Verdict. This final stage generates a concise summary that distills key findings and critical evidence, including hyperlinks.

Performance and Evaluation

DEFAME has established itself as the new general state-of-the-art fact-checking system for uni- and multimodal fact-checking. It surpasses all previous methods on the popular benchmarks VERITE, AVERITEC, and MOCHEG. On AVERITEC (Schlichtkrull et al., 2024b), accuracy improved from 65.6% to 70.5% on MOCHEG (Yao et al., 2023), a +10.6% improvement in accuracy was achieved and on VERITE (Papadopoulos et al., 2024b), True/False accuracy was enhanced by +25.9% Additionally, DEFAME contributes a new benchmark, CLAIMREVIEW2024+, featuring claims after the knowledge cutoff of GPT-4O to avoid data leakage. DEFAME drastically outperforms the GPT4O baselines on these “unseen” statements.

Key Contributions and Limitations

The main contribution of this paper is introducing DEFAME, a straightforward end-to-end AFC framework that unifies advancements in the field into one single system. It is able to natively process multimodal claims and evidence and retrieve evidence dynamically as needed. The framework's design allows it to operate across all benchmarks without task-specific tuning or training data, making it the most general AFC method as of now. However, limitations exist regarding the credibility of external evidence and system stability due to web scraping. Furthermore, hallucinations are a risk inherent to LLMs that must be analyzed more closely in future work.

Failure Analysis and Human Evaluation

A failure analysis on 119 VERITE and CLAIMREVIEW2024+ instances identified five common failure modes <ref:2412.

Improvements for AI systems

  1. Improve system generality by creating a dynamic, multistep RAG system that unifying advancements in AFC into one single system, enabling it to natively process multimodal claims and evidence.

  2. Enhance reasoning depth by decomposing the fact-checking process into six stages, where each MLLM call is guided by context from the previous stage, allowing for more intricate, multi-hop reasoning and evidence retrieval.

  3. Increase temporal robustness by implementing a constraint in Stage 2: To prevent temporal leakage, all web-based tools restrict search results to sources published before the claim’s release date (if known).

  4. Improve explainability by ensuring the final output includes a concise summary generated in Stage 6 that distills the key findings and critical evidence, including hyperlinks, making it a human-friendly report document.

  5. Mitigate reliance on static knowledge by employing dynamic retrieval instead of parametric knowledge; DEFAME does not rely on parametric knowledge but rather retrieves evidence dynamically through external tools.

Sources

Related papers