FactReview: Evidence-Grounded Peer Review with Execution-Based Claim Verification
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "FactReview: Evidence-Grounded Peer Review with Execution-Based Claim Verification".
Jane: The paper was written by Ling Yue, Chaoqian Ouyang, Hang Xu, Ruijun Huang, Yuchen Liu et al. from Rensselaer Polytechnic Institute and Sun Yat-sen University and Southeast University and The Hong Kong University of Science and Technology.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that's been making the rounds on arXiv, and it's called "FactReview: Evidence-Grounded Peer Review with Execution-Based Claim Verification." Jane, I gotta say, just the title alone tells me these folks are trying to fix something that's been bugging me for years.
Jane: Oh absolutely, Tom. And I think the title really captures the two big ideas here. First, they want reviews to be grounded in evidence, not just vibes. Second, they actually want to run the code to check if the claims hold up. That's a huge deal.
Tom: Right, because right now, if you're an LLM reviewer, you read the paper and you write a review. But you never actually check whether the numbers in the paper are real. You're basically taking the author's word for it.
Jane: Exactly. And the authors of FactReview point out something really important. They say that most LLM-based reviewing systems take only the manuscript as input. So the reviewer never looks at related work, never checks the references, and never runs the code. It's like reviewing a recipe without ever tasting the dish.
Tom: That's a great analogy, Jane. And the title also mentions "execution-based claim verification." So they're not just talking about checking the literature. They actually execute the released code under a fixed repair budget to see if the empirical claims hold up.
Jane: And that's the part that gets me excited. Because we've all seen papers where the abstract says one thing, but when you try to run the code, it just doesn't work. FactReview wants to catch that automatically.
Tom: So the title is really a promise. It's saying, "We're going to ground our reviews in evidence, and we're going to verify claims by actually running the code." That's a bold promise.
Jane: It is. And the authors back it up with real experiments. They tested it on thirty-five machine learning papers and four hundred sixty-three benchmark major claims. That's a solid evaluation.
Tom: And they found that FactReview covers eighty-four percent of the claims. That's pretty impressive for a system that's doing all this extra work.
Jane: Yeah, and the title also hints at something else. They're not trying to replace human reviewers. They're trying to give them better tools. The paper explicitly says they avoid issuing accept-reject decisions.
Tom: Which is smart, because that's a human judgment call. But giving reviewers the evidence they need to make that call? That's where the value is.
Jane: So the title really sets up the whole paper. It's about grounding reviews in evidence, verifying claims by running code, and ultimately helping human reviewers do their jobs better.
Tom: And I think that's a conversation worth having. Because peer review is under so much pressure right now. Submission volumes are rising, and reviewers don't have time to check everything.
Jane: Right. And that's exactly what this paper is trying to address. So let's dig into the summary and see how they actually built this thing.
Tom: Sounds good. But before we do, let me just say, this paper has me genuinely excited. It feels like we're moving toward a future where reviews are more than just opinions. They're backed by evidence.
Jane: And that's a future I want to live in. So stick around, because next we're going to talk about how FactReview actually works.
Summary: Jane: So we're back, and we're still talking about "FactReview: Evidence-Grounded Peer Review with Execution-Based Claim Verification." Tom, let's break down what this system actually does, because the summary is pretty dense.
Tom: Yeah, let's do it. So the core idea is that FactReview reformulates automated peer review as evidence-grounded claim verification. Instead of generating a block of review prose, it decomposes the paper into review-relevant claims, grounds each claim in retrieved related work, and when code is available, verifies empirical claims by running the code.
Jane: And that's a really important shift. Because most LLM reviewers just read the manuscript and write a review. They don't check anything. FactReview actually builds evidence for each claim.
Tom: Right. And the evidence comes from multiple sources. There's manuscript evidence, which is what the paper itself says. There's literature evidence, which is what related work says. And there's execution evidence, which comes from running the code.
Jane: And they assign each claim one of four labels. Supported, Partially supported, In conflict, or Inconclusive. That's a really clean way to think about it.
Tom: It is. And the key thing is that each label comes with explicit links to the evidence that justifies it. So you can see exactly why a claim was labeled the way it was.
Jane: And that's huge for transparency. Because right now, if an LLM reviewer says "this paper is good," you have no idea why. With FactReview, you can trace every judgment back to its evidence.
Tom: Exactly. And the results are pretty impressive. They tested it on thirty-five ML papers and four hundred sixty-three benchmark claims. FactReview covers eighty-four percent of the claims.
Jane: And under an evidence-aware rubric, it scores four point eight six out of five in overall quality. That's zero point seven above DeepReview-v2 and one point five above matched OpenReview comments.
Tom: Those are big margins. And the paper also found that removing execution evidence changes seventeen percent of claim statuses. That's more than any other single evidence source.
Jane: So running the code actually matters. It's not just a nice-to-have. It changes the verdict on a significant number of claims.
Tom: Right. And that's the part that really sets FactReview apart. Because most LLM reviewers don't run code at all. They just read the paper and trust what it says.
Jane: And the paper also includes a reviewer-assistance study. They found that FactReview reduces mean review time by fifty-eight percent while raising benchmark claim coverage from eighty-seven percent to ninety-nine percent.
Tom: That's a huge improvement. And it makes sense. If you have a system that's already extracted the claims, checked the literature, and run the code, you don't have to do all that work yourself.
Jane: Exactly. And the authors make a really important point at the end. They argue that LLM reviewers should audit empirical claims, not make accept-reject decisions.
Tom: That's a philosophical stance, and I think it's the right one. Because the final decision should be a human judgment. But giving humans the evidence they need to make that judgment? That's where the value is.
Jane: So the summary really paints a picture of a system that's doing something fundamentally different from what came before. It's not just generating text. It's building evidence.
Tom: And that's why I'm so excited about this paper. It feels like a step toward making peer review more rigorous and more transparent.
Jane: Me too. And I think the implications are huge. So let's talk about the improvements this paper suggests, because there's a lot to unpack there.
Tom: Absolutely. And I think the biggest improvement is the shift from opinion to evidence. That's a fundamental change in how we think about automated reviewing.
Improvements: Jane: So we're back, still on "FactReview: Evidence-Grounded Peer Review with Execution-Based Claim Verification." Tom, we've talked about what the system does. Now let's talk about the improvements it suggests. And I think the biggest one is the shift from generating opinions to verifying claims.
Tom: Yeah, that's the fundamental change. And it's not just a tweak. It's a completely different way of thinking about automated reviewing. Instead of asking "what does the reviewer think about this paper?", you ask "what evidence supports or contradicts each claim?"
Jane: And that's a much more useful question. Because a review that says "this paper is good" is not very helpful. But a review that says "this claim is supported by the reproduced results, but this other claim is only partially supported because the ablation is missing" — that's actually actionable.
Tom: Exactly. And the paper suggests another improvement: linking every judgment to its evidence. So you can see exactly why a claim was labeled the way it was.
Jane: And that's a big deal for transparency. Because right now, if an LLM reviewer says something, you have to trust it. With FactReview, you can check the evidence yourself.
Tom: Right. And the paper also suggests that execution evidence is crucial. They found that removing execution evidence changes seventeen percent of claim statuses. That's the single largest source of change.
Jane: So running the code isn't optional. It's essential. And that's a big improvement over systems that just read the manuscript.
Tom: And there's another improvement I want to highlight. FactReview doesn't issue accept-reject decisions. It leaves that to human reviewers.
Jane: And I think that's really important. Because the system is not trying to replace human judgment. It's trying to support it. And that's a much more realistic and responsible framing.
Tom: Absolutely. And the paper also suggests improvements in how we evaluate LLM reviewers. They use an evidence-aware rubric that measures groundedness, specificity, and coverage.
Jane: That's a much better way to evaluate than just asking "does this review sound good?" Because a review can sound good and still be completely wrong.
Tom: Right. And the results show that FactReview outperforms other systems on all these dimensions. It scores four point nine seven on groundedness, four point nine four on specificity, and four point six six on coverage.
Jane: Those are strong numbers. And they suggest that grounding reviews in evidence actually makes them better, not just more transparent.
Tom: And there's one more improvement I want to mention. The paper includes a reviewer-assistance study. They found that FactReview reduces mean review time by fifty-eight percent while raising benchmark claim coverage from eighty-seven percent to ninety-nine percent.
Jane: That's a huge practical improvement. Because reviewers are under so much time pressure. If a system can help them do their job faster and better, that's a win.
Tom: And the paper also suggests that FactReview can help with reproducibility. By actually running the code, it can catch cases where the paper claims one thing but the code produces something different.
Jane: And that's a really important contribution. Because reproducibility is a huge problem in machine learning. And FactReview is trying to address it directly.
Tom: So the improvements are really about making reviews more rigorous, more transparent, and more useful. And I think that's exactly what the field needs right now.
Jane: Me too. And I think the implications go beyond just peer review. This approach could be used for any kind of scientific claim verification.
Tom: That's a great point. And it's one of the reasons I'm so excited about this paper. It's not just a tool for reviewers. It's a step toward a more evidence-based approach to science.
Jane: So let's wrap up and talk about the big picture. Because I think this paper has implications that go far beyond just reviewing papers.
Conclusion: Tom: So we've spent the whole show on "FactReview: Evidence-Grounded Peer Review with Execution-Based Claim Verification," and I think we can all agree this is a big deal. Jane, what's the takeaway for our listeners?
Jane: The takeaway is that automated reviewing doesn't have to be about generating opinions. It can be about building evidence. FactReview shows that by extracting claims, grounding them in literature, and actually running the code, you can produce reviews that are more rigorous, more transparent, and more useful.
Tom: And the numbers back it up. eighty-four percent claim coverage, four point eight six out of five on overall quality, and a fifty-eight percent reduction in review time. Those are real results.
Jane: And I think the most important thing is that FactReview doesn't try to replace human reviewers. It tries to support them. It gives them the evidence they need to make better decisions.
Tom: Exactly. And that's a philosophy I can get behind. Because peer review is too important to leave to machines. But machines can help us do it better.
Jane: And I think the implications go beyond peer review. This approach to evidence-grounded claim verification could be used in any field where you need to check whether claims hold up.
Tom: Absolutely. And I think that's what makes this paper so exciting. It's not just a tool. It's a vision for how AI can support science.
Jane: So let's say goodbye to "FactReview: Evidence-Grounded Peer Review with Execution-Based Claim Verification." It's been a great paper to discuss.
Tom: It really has. And I'm looking forward to seeing where this line of research goes. Because I think we're just scratching the surface.
Jane: Me too. And to our listeners, thanks for tuning in. We'll be back soon with another paper from arXiv.
Tom: Until then, keep asking questions and checking the evidence. That's what science is all about.
Jane: And that's a wrap. See you next time.
Tom: Take care, everyone.
Ling Yue, Chaoqian Ouyang, Hang Xu, Ruijun Huang, Yuchen Liu, Libin Zheng, Wei Liu, Shaowu Pan, Shimin Di, Min-Ling Zhang
Rensselaer Polytechnic Institute · Sun Yat-sen University · Southeast University · The Hong Kong University of Science and Technology
cs.AI, cs.LG
Submitted: 2026-08-16
Updated: 2026-08-18
Code: https://github.com/DEFENSE-SEU/FactReview
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 70/100
The gist: The paper introduces FactReview, a system that reformulates automated peer review as evidence-grounded claim verification rather than free-form review generation.
Key concepts
- FactReview
- This system reformulates automated peer review by breaking down a paper into claims. It then grounds these claims in three types of evidence: the manuscript itself, related literature, and the execution of the paper's provided code.
- Execution-Based Claim Verification
- This is when a system checks if empirical claims are true by actually running the associated code. It moves beyond just reading a paper to testing whether those claims hold up when executed under specific conditions.
- Evidence-Grounded Peer Review
- This approach ensures that every judgment made during review is explicitly tied to supporting data. It replaces subjective opinions with verifiable evidence, making reviews more transparent and rigorous for human reviewers.
Terminology
Summary
The paper introduces FactReview, a system that reformulates automated peer review as evidence-grounded claim verification rather than free-form review generation. The authors state: "LLM-based reviewing systems typically take only the manuscript as input, leaving literature- and code-based claims hard to verify. We present FactReview, a system that extracts review-relevant claims, grounds them in related work, and, when code is available, executes released artifacts under a fixed repair budget to audit empirical claims."
The motivation is that existing LLM reviewers "operate on the manuscript alone: claims are seldom verified against the broader literature, and even less often against the released code. The resulting reviews are sensitive to writing quality, prone to accepting unverified author claims, and opaque about the basis of each judgment. The authors argue that
the strongest test of these claims often requires repository inspection, environment reconstruction, and actual execution."
The system works as follows: "FactReview decomposes a paper into review-relevant claims, grounds each claim in retrieved related work, and, when a code repository is available, verifies empirical claims by running the code under a fixed repair budget. Each claim is then assigned one of four evidence labels (Supported, Partially supported, In conflict, or Inconclusive) with explicit links to the manuscript, literature, and execution evidence that justify it."
The method has several stages: document parsing with MinerU, claim-centered extraction using a schema-constrained LLM extractor, literature positioning and reference verification (including a module called RefCopilot), execution-based claim verification via a stateful loop of prepare, plan, run, judge, fix, and finalize steps
with a repair budget of K=3 rounds, and final claim assessment and review synthesis. The authors emphasize: FactReview does not issue accept-reject decisions or replace human judgment. It targets the evidence-heavy parts of reviewing.
Key experimental results across 35 ML papers and 463 benchmark claims
show:
-
FactReview
covers 84% of claims
and achieves4.86/5 in overall quality, 0.7 above DeepReview-v2 and 1.5 above matched OpenReview comments
under an evidence-aware rubric. -
Removing execution evidence changes 17% of claim statuses, more than any other single evidence source.
-
In a reviewer-assistance study,
FactReview reduces mean review time by 58% while raising benchmark claim coverage from 87% to 99%.
-
Run-Review-Fix
recovers 2 of the 9 papers that initially fail execution and raises the claim pass rate from 67.7% to 82.3%.
-
Pipeline cost is
357K tokens and 773 seconds per paper on average,
with estimated costroughly 0.5–0.7 per paper.
The authors conclude: FactReview is best understood as an audit layer for reviewers. It does not replace expert judgment, but makes the factual basis of that judgment easier to inspect.
They argue that LLM reviewers should audit empirical claims, not make accept-reject decisions.
The code is public at https://github.com/DEFENSE-SEU/FactReview.
Improvements for AI systems
Improvements to AI Systems Based on FactReview
- Add an execution-based claim verification module to LLM review systems.
-
The improved system will parse a submitted manuscript, extract review-relevant claims (e.g.,
improves accuracy by 5%
), and then, when code is available, run the released repository under a bounded repair budget (max 3 fix rounds) to check whether reported numbers match actual outputs. -
It will label each claim as Supported, Partially supported, In conflict, or Inconclusive, with explicit links to the manuscript, literature, and execution evidence.
-
This reduces false acceptance of unverified empirical claims and flags overclaims that manuscript-only reviewers miss.
- Integrate literature grounding and reference-integrity checks into the review pipeline.
-
The system will retrieve related work via Semantic Scholar and paper search, build a technical positioning table comparing the submission against nearby methods along task, mechanism, and evaluation axes.
-
It will also verify each bibliography entry against arXiv, Semantic Scholar, OpenReview, and OpenAlex, flagging hallucinated, withdrawn, or metadata-mismatched references.
-
This improves novelty assessment and prevents citation fabrication in generated reviews.
- Replace free-form review generation with evidence-linked, claim-level output.
-
The improved system will produce a concise review where every substantive judgment is tied to a specific claim and its evidence source (manuscript, literature, reference check, or execution).
-
It will not issue accept/reject decisions, leaving final judgment to human reviewers.
-
This increases groundedness and specificity scores (FactReview achieves 4.97/5 and 4.94/5 respectively, vs. 3.26 and 3.26 for the best direct LLM baseline).
- Add a reviewer-assistance mode that reduces review time and improves claim coverage.
-
The system will generate a report (and optional teaser figure) that human reviewers can use as an audit layer.
-
In the study, adding the report reduced mean review time from 50.6 to 31.6 minutes, and adding the teaser figure further reduced it to 21.3 minutes, while benchmark-claim coverage rose from 86.9% to 98.9%.
-
This allows reviewers to assess more claims in less time without sacrificing accuracy.
- Implement a conservative repair policy for code execution.
-
The system will only fix environment, dependency, path, or launch-script issues (e.g., installing missing packages, correcting file paths, adding missing command arguments).
-
It will never modify model architectures, loss functions, datasets, evaluation logic, or reported baselines.
-
If execution succeeds but outputs cannot be aligned to a claim, the claim is marked Inconclusive, not forced into a positive or negative verdict.
-
This prevents false positives from
repaired
code that no longer reflects the paper's actual method.
- Add an evidence-source ablation capability for debugging and transparency.
-
The system will allow users to remove individual evidence sources (execution, literature, retrieval, or manuscript-only) to see how claim statuses change.
-
FactReview shows that removing execution evidence changes 17.0% of claim statuses, more than any other single source, and removing all non-manuscript evidence changes 26.1%.
-
This helps users understand which evidence type is most critical for a given paper and where the system's confidence is weakest.
- Enable backend-model flexibility with a fixed workflow.
-
The system will run the same claim-verification pipeline across different LLM backends (e.g., Claude Opus 4.6, GPT-5.4, GPT-4.1).
-
In the benchmark, Claude Opus 4.6 achieved 83.3% execution success with the shortest runtime, while GPT-5.4 reached 75.0%.
-
This allows organizations to trade cost vs. reliability (e.g., using a cheaper model for initial screening and a stronger model for final verification).
- Add a Run-Review-Fix (RRF) loop for execution recovery.
-
The system will automatically diagnose failed runs, apply bounded fixes, and re-execute up to 3 rounds.
-
In the benchmark, RRF recovered 2 of 9 initially failing papers and raised the claim pass rate from 67.7% to 82.3%, with paper success increasing from 55% to 65%.
-
This reduces the number of Inconclusive verdicts caused by environment or launch issues rather than actual claim failure.
- Provide a structured execution trace for auditability.
-
The system will record every command, return code, log, intermediate output, metric, alignment decision, repair attempt, runtime, and token cost for each execution job.
-
This trace is included in the evidence report, allowing human reviewers to verify exactly what was run and how the claim-alignment decision was made.
-
This increases trust and reproducibility of the automated review process.
- Support multi-label failure diagnosis for code-available papers.
-
The system will categorize execution failures by provenance (environment, runtime, metric availability, missing baselines, import errors, alignment mismatches) and show a funnel of earliest blocking stages.
-
This helps developers and reviewers understand why a repository failed to produce claim-aligned evidence, rather than just that it failed.
-
In the benchmark, environment failures were most common (10 labels), followed by runtime (8), unavailable metrics (6), missing baselines (5), import errors (4), and alignment mismatches (3).
What the improved AI system can do:
-
Audit empirical claims in ML papers by running released code and comparing outputs to reported numbers, flagging overclaims and unsupported results.
-
Ground every review judgment in inspectable evidence (manuscript, literature, reference checks, execution), reducing reliance on model rhetoric.
-
Reduce reviewer workload by 58% (from 50.6 to 21.3 minutes per paper) while increasing claim coverage from 87% to 99%.
-
Detect citation hallucinations and metadata errors in bibliographies, preventing fabricated references from entering the literature.
-
Provide a transparent, claim-level verdict (Supported / Partially supported / In conflict / Inconclusive) with provenance links, instead of a single opaque score.
-
Recover from environment and launch failures via bounded repair, increasing the fraction of papers that yield execution evidence from 55% to 65%.
-
Adapt to different LLM backends without changing the workflow, allowing cost-performance trade-offs.
-
Serve as a reviewer-assistance tool, not a replacement, leaving accept/reject decisions to humans while making the factual basis of those decisions easier to inspect.
Sources
- OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs
- Stop Automating Peer Review Without Rigorous Evaluation
- MARG: Multi-Agent Review Generation for Scientific Papers
- Reviewer2: Optimizing Review Generation Through Prompt Generation
- ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code
- PaperQA: Retrieval-Augmented Generative Agent for Scientific Research
- PaperBench: Evaluating AI's Ability to Replicate AI Research
- Fact or Fiction: Verifying Scientific Claims
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale
- Can large language models provide useful feedback on research papers? A large-scale empirical analysis
- LLM-based Corroborating and Refuting Evidence Retrieval for Scientific Claim Verification
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- SF-RAG: Structure-Fidelity Retrieval-Augmented Generation for Academic Question Answering
- AutoReproduce: Automatic AI Experiment Reproduction with Paper Lineage
- When Your Reviewer is an LLM: Biases, Divergence, and Prompt Injection Risks in Peer Review
- A Step Toward Quantifying Independently Reproducible Machine Learning Research
- ReviewerToo: Should AI Join The Program Committee? A Look At The Future of Peer Review
- Large language models for automated scholarly paper review: A survey
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection