DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts

summary

Video file (mp4)

The gist

The gist: DEFAME presents Dynamic Evidence-based FAct-checking with Multimodal Experts, a modular, zero-shot MLLM pipeline designed for open-domain, text-image claim verification that dynamically

In short

DEFAME is a dynamic evidence-based fact-checking system that uses multimodal experts to verify claims across text and images. It operates in six stages, planning actions, executing tools like web search or image search, summarizing results, developing the fact-check argument, predicting a verdict, and justifying the final answer. This end-to-end framework dynamically selects evidence retrieval methods for open-domain verification.

Key concepts

DEFAME Framework
A six-stage process designed for end-to-end fact verification. It dynamically chooses the right tools (like search or image retrieval) and search depth needed to gather necessary textual and visual evidence, mimicking a human fact-checker.
Multimodal Experts
The system uses Large Multimodal Models (MLLMs) to handle claims that involve both text and images. These experts are prompted at several stages of the process to plan actions, summarize retrieved evidence, develop arguments, and predict verdicts based on the combined multimodal input.
Dynamic Evidence Retrieval
Instead of relying on fixed knowledge or text-only methods, DEFAME dynamically decides when and how to search for information. It invokes specific tools (like Google Search or Reverse Image Search) only when needed during the execution stage, making it flexible for open-domain claims.
Six Stages of Verification
The structured procedure DEFAME follows: Plan Actions, Execute Actions (using tools), Summarize Results, Develop the Fact-Check argument based on evidence, Predict a Verdict (looping back if information is insufficient), and finally Justify the Verdict with supporting links.

Terminology used across episodes

This episode discusses

The paper

DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts · Read on arXiv

Tobias Braun, Mark Rothermel, Marcus Rohrbach 1 Anna Rohrbach 1

Technical University of Darmstadt & hessian.AI

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts".

Jane: The gist: DEFAME presents Dynamic Evidence-based FAct-checking with Multimodal Experts, a modular, zero-shot MLLM pipeline designed for open-domain,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So we're talking about DEFAME today. It's this new system for fact-checking that uses multimodal experts to handle claims involving both text and images. The authors say it’s a modular, zero-shot pipeline that dynamically picks the right tools and search depth to grab evidence.

Jane: Exactly. The main point of this paper is that it moves beyond just looking at text alone or just relying on some fixed knowledge base. They claim DEFAME performs end-to-end verification, meaning it actually checks the claim using both pictures and words, while generating a structured report for you.

Lu: What's really interesting is how they structure the process. It operates in six stages: planning actions, executing those actions with tools like Google Search or Reverse Image Search, summarizing the results, developing the fact-check based on that evidence, predicting a verdict and then justifying it all out.

Meng: That sounds complex for what it is trying to achieve. So how does this dynamic tool selection work in practice? Does it just pick tools randomly, or is there some kind of intelligence guiding which search depth to use?

Tom: Well, the planning stage lets the MLLM suggest a targeted action sequence first. Then you have the execution stage where it invokes tools like web search or image search based on what the plan suggested. It’s supposed to be adaptive, not just running one fixed process through everything.

Jane: And that leads into how they handle images specifically because they use tools like Google Vision API for reverse image searches. That lets the system actually analyze visual evidence in claims, which is a big step up from text-only systems.

Lalam: From my side, I see this as a way to improve cultural understanding. If we can build systems that look at images and text together reliably, it means we can process information in ways that are more nuanced than what simple language models or even older fact-checkers could handle on their own.

Tom: Right, and the results they show are pretty strong too. They’ve tested DEFAME on benchmarks like VERITE, AVERITEC, and MOCHEG. They claim it surpasses all previous methods on those tests for uni- and multimodal fact-checking.

Jane: They even mentioned an improvement of over ten percent accuracy on MOCHEG when they compared it to earlier work by Yao et al., two thousand twenty-three <ref:2412.10510#pg2>. That shows a tangible lift in how well it handles the verification task across different claim types.

Meng: But what about the limitations? The paper does flag a few things, including potential credibility issues with external evidence and system stability problems because of web scraping. That’s a practical hurdle for any real-world deployment I see right away.

Lu: They also point out that because it relies on an LLM for reasoning, there's always the risk of hallucinations creeping in, which they say needs more analysis in future work. It’s not perfect yet.

Tom: So, to wrap up this part of the DEFAME paper, it’s presented as a unified framework that handles multimodal claims natively and retrieves evidence on the fly. It sets a new standard for general fact-checking because it doesn't require task-specific training or tuning to work across different challenges.

Jane: The implication here is that for anyone dealing with information where pictures and text are mixed, this system offers a transparent way to get an answer, not just a black box output. It’s designed to mimic how a human fact-checker would actually work by giving you the steps.

Tom: So, what does this mean for us? It means we have a more robust tool for verifying complex information online than we did before, especially when visual context is involved. We'll keep an eye on those benchmark results as they settle.

Jane: And next time we talk about this paper, we’ll look at how these kinds of multimodal systems could actually be integrated into daily tools for people on the go. That’s what's coming up next.

Conclusion: Tom: So we're wrapping up on DEFAME today, focusing on what that title actually means for us in the real world.

Jane: It’s about taking those two big ideas—dynamic evidence and multimodal experts—and putting them together into one single system for fact-checking.

Lu: What they’ve done is create this modular pipeline that doesn't rely on a single way to check something. It just picks the right tools, like web search or image recognition, based on what it needs at that moment.

Meng: So you’re saying instead of running one giant script through everything, it plans what to do first and then executes those specific steps dynamically?

Tom: Exactly. It’s designed to look at a claim and figure out which piece of evidence—text or image—is missing, and then decide exactly which search tool is best for finding it.

Lalam: From my side, I see this as moving beyond just checking text. We can now verify claims that involve pictures in a way that feels much more like how a person would actually investigate something.

Jane: It really does give us a transparent process because you can see the steps they take to get to the final answer.

Tom: And when you look at the results, they’ve shown this system beating previous top methods on big fact-checking tests like VERITE and MOCHEG. That's a solid number showing it works better than what came before.

Lu: They even created a new test set, CLAIMREVIEW2024+, which is interesting because it checks claims that are newer than the models were trained on.

Meng: That’s important for making sure the system isn't just repeating old answers; it’s actually looking at current information.

Jane: What this means for us is that we have a more general fact-checking tool now, one that doesn't need to be specifically trained for every single type of claim.

Tom: And the challenge ahead is making sure this dynamic tool selection stays accurate and reliable when it's out there in the wild. That’s where we need to keep watching things.

More episodes

← Home