DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts
summary
The gist
The gist: DEFAME presents Dynamic Evidence-based FAct-checking with Multimodal Experts, a modular, zero-shot MLLM pipeline designed for open-domain, text-image claim verification that dynamically
In short
DEFAME is a dynamic evidence-based fact-checking system that uses multimodal experts to verify claims across text and images. It operates in six stages, planning actions, executing tools like web search or image search, summarizing results, developing the fact-check argument, predicting a verdict, and justifying the final answer. This end-to-end framework dynamically selects evidence retrieval methods for open-domain verification.
Key concepts
- DEFAME Framework
- A six-stage process designed for end-to-end fact verification. It dynamically chooses the right tools (like search or image retrieval) and search depth needed to gather necessary textual and visual evidence, mimicking a human fact-checker.
- Multimodal Experts
- The system uses Large Multimodal Models (MLLMs) to handle claims that involve both text and images. These experts are prompted at several stages of the process to plan actions, summarize retrieved evidence, develop arguments, and predict verdicts based on the combined multimodal input.
- Dynamic Evidence Retrieval
- Instead of relying on fixed knowledge or text-only methods, DEFAME dynamically decides when and how to search for information. It invokes specific tools (like Google Search or Reverse Image Search) only when needed during the execution stage, making it flexible for open-domain claims.
- Six Stages of Verification
- The structured procedure DEFAME follows: Plan Actions, Execute Actions (using tools), Summarize Results, Develop the Fact-Check argument based on evidence, Predict a Verdict (looping back if information is insufficient), and finally Justify the Verdict with supporting links.
Terminology used across episodes
This episode discusses
- DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts · Paper Radio
- Multimodal Automated Fact-Checking: A Survey
- COSMOS: Catching Out-of-Context Misinformation with Self-Supervised Learning
- Can LLMs Improve Multimodal Fact-Checking by Asking Relevant Questions?
- Are Large Language Models Good Fact Checkers: A Preliminary Study
- Claim Verification in the Age of Large Language Models: A Survey
- AMMeBa: A Large-Scale Survey and Dataset of Media-Based Misinformation In-The-Wild
- Distributed Optimization by Network Flows with Spatio-Temporal Compression
- Multimodal Large Language Models to Support Real-World Fact-Checking
- LLaVA-OneVision: Easy Visual Task Transfer
- Re-Search for The Truth: Multi-round Retrieval-augmented Large Language Models are Strong Fake News Detectors
- Large Language Model Agent for Fake News Detection
- MMFakeBench: A Mixed-Source Multimodal Misinformation Detection Benchmark for LVLMs
- RED-DOT: Multimodal Fact-checking via Relevant Evidence Detection
- Similarity over Factuality: Are we making progress on multimodal out-of-context misinformation detection?
- Generative Large Language Models in Automated Fact-Checking: A Survey
- OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs
- Take It Easy: Label-Adaptive Self-Rationalization for Fact Verification and Explanation Generation
- Customized large language models can outperform Community Notes in correcting misinformation · Paper Radio
The paper
DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts · Read on arXiv
Tobias Braun, Mark Rothermel, Marcus Rohrbach 1 Anna Rohrbach 1
Technical University of Darmstadt & hessian.AI
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts".
Jane: The gist: DEFAME presents Dynamic Evidence-based FAct-checking with Multimodal Experts, a modular, zero-shot MLLM pipeline designed for open-domain,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So we're talking about DEFAME today. It's this new system for fact-checking that uses multimodal experts to handle claims involving both text and images. The authors say it’s a modular, zero-shot pipeline that dynamically picks the right tools and search depth to grab evidence.
Jane: Exactly. The main point of this paper is that it moves beyond just looking at text alone or just relying on some fixed knowledge base. They claim DEFAME performs end-to-end verification, meaning it actually checks the claim using both pictures and words, while generating a structured report for you.
Lu: What's really interesting is how they structure the process. It operates in six stages: planning actions, executing those actions with tools like Google Search or Reverse Image Search, summarizing the results, developing the fact-check based on that evidence, predicting a verdict and then justifying it all out.
Meng: That sounds complex for what it is trying to achieve. So how does this dynamic tool selection work in practice? Does it just pick tools randomly, or is there some kind of intelligence guiding which search depth to use?
Tom: Well, the planning stage lets the MLLM suggest a targeted action sequence first. Then you have the execution stage where it invokes tools like web search or image search based on what the plan suggested. It’s supposed to be adaptive, not just running one fixed process through everything.
Jane: And that leads into how they handle images specifically because they use tools like Google Vision API for reverse image searches. That lets the system actually analyze visual evidence in claims, which is a big step up from text-only systems.
Lalam: From my side, I see this as a way to improve cultural understanding. If we can build systems that look at images and text together reliably, it means we can process information in ways that are more nuanced than what simple language models or even older fact-checkers could handle on their own.
Tom: Right, and the results they show are pretty strong too. They’ve tested DEFAME on benchmarks like VERITE, AVERITEC, and MOCHEG. They claim it surpasses all previous methods on those tests for uni- and multimodal fact-checking.
Jane: They even mentioned an improvement of over ten percent accuracy on MOCHEG when they compared it to earlier work by Yao et al., two thousand twenty-three <ref:2412.10510#pg2>. That shows a tangible lift in how well it handles the verification task across different claim types.
Meng: But what about the limitations? The paper does flag a few things, including potential credibility issues with external evidence and system stability problems because of web scraping. That’s a practical hurdle for any real-world deployment I see right away.
Lu: They also point out that because it relies on an LLM for reasoning, there's always the risk of hallucinations creeping in, which they say needs more analysis in future work. It’s not perfect yet.
Tom: So, to wrap up this part of the DEFAME paper, it’s presented as a unified framework that handles multimodal claims natively and retrieves evidence on the fly. It sets a new standard for general fact-checking because it doesn't require task-specific training or tuning to work across different challenges.
Jane: The implication here is that for anyone dealing with information where pictures and text are mixed, this system offers a transparent way to get an answer, not just a black box output. It’s designed to mimic how a human fact-checker would actually work by giving you the steps.
Tom: So, what does this mean for us? It means we have a more robust tool for verifying complex information online than we did before, especially when visual context is involved. We'll keep an eye on those benchmark results as they settle.
Jane: And next time we talk about this paper, we’ll look at how these kinds of multimodal systems could actually be integrated into daily tools for people on the go. That’s what's coming up next.
Conclusion: Tom: So we're wrapping up on DEFAME today, focusing on what that title actually means for us in the real world.
Jane: It’s about taking those two big ideas—dynamic evidence and multimodal experts—and putting them together into one single system for fact-checking.
Lu: What they’ve done is create this modular pipeline that doesn't rely on a single way to check something. It just picks the right tools, like web search or image recognition, based on what it needs at that moment.
Meng: So you’re saying instead of running one giant script through everything, it plans what to do first and then executes those specific steps dynamically?
Tom: Exactly. It’s designed to look at a claim and figure out which piece of evidence—text or image—is missing, and then decide exactly which search tool is best for finding it.
Lalam: From my side, I see this as moving beyond just checking text. We can now verify claims that involve pictures in a way that feels much more like how a person would actually investigate something.
Jane: It really does give us a transparent process because you can see the steps they take to get to the final answer.
Tom: And when you look at the results, they’ve shown this system beating previous top methods on big fact-checking tests like VERITE and MOCHEG. That's a solid number showing it works better than what came before.
Lu: They even created a new test set, CLAIMREVIEW2024+, which is interesting because it checks claims that are newer than the models were trained on.
Meng: That’s important for making sure the system isn't just repeating old answers; it’s actually looking at current information.
Jane: What this means for us is that we have a more general fact-checking tool now, one that doesn't need to be specifically trained for every single type of claim.
Tom: And the challenge ahead is making sure this dynamic tool selection stays accurate and reliable when it's out there in the wild. That’s where we need to keep watching things.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization