Are Large Language Models Reliable Reviewers? A Benchmark for Error Detection in Financial Documents
summary
The gist
This paper introduces FinED-Bench, the first publicly available benchmark for financial error detection, designed to evaluate the capabilities of Large Language Models (LLMs) in identifying errors
In short
The episode analyzes the paper 'Are Large Language Models Reliable Reviewers?' which tests LLMs' ability to detect errors in financial documents. Hosts conclude that current models are unreliable, especially with complex reasoning errors. However, fine-tuning shows promise for improving error detection.
Key concepts
- FinED-Bench
- A benchmark created by the paper's authors using over nine hundred real financial documents published after most LLM knowledge cutoffs. It was designed to test if models can reliably detect injected errors in financial texts.
- Reasoning Errors
- A difficult type of error where the text in one part of a document contradicts information found elsewhere, such as a discrepancy between the written text and an embedded table. Models struggle most with this category.
- F1 Score
- A metric used to evaluate model performance in error detection. A higher F1 score indicates better overall accuracy, meaning the model is successfully identifying a high percentage of actual errors.
- Fine-tuning
- The process of training an existing LLM (like Qwen3-14B) on a specific, targeted task—in this case, error detection. This specialized training significantly boosts performance without harming the model's general capabilities.
Terminology used across episodes
This episode discusses
- Are Large Language Models Reliable Reviewers? A Benchmark for Error Detection in Financial Documents · Paper Radio
- MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes
- Large Language Models for Mathematical Reasoning: Progresses and Challenges
- FinReflectKG -- MultiHop: Financial QA Benchmark for Reasoning with Knowledge Graph Evidence
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Large Language Model Agent in Financial Trading: A Survey
- Open-FinLLMs: Open Multimodal Large Language Models for Financial Applications
- FinanceBench: A New Benchmark for Financial Question Answering
- BizBench: A Quantitative Reasoning Benchmark for Business and Finance
- CREAK: A Dataset for Commonsense Reasoning over Entity Knowledge
- Plutus: Benchmarking Large Language Models in Low-Resource Greek Finance
- CFBenchmark: Chinese Financial Assistant Benchmark for Large Language Model
- WHEN FLUE MEETS FLANG: Benchmarks and Large Pre-trained Language Model for Financial Domain
- Fin-R1: A Large Language Model for Financial Reasoning through Reinforcement Learning
- UA-GEC: Grammatical Error Correction and Fluency Corpus for the Ukrainian Language
- KG-MuLQA: A Framework for KG-based Multi-Level QA Extraction and Long-Context LLM Evaluation
- Qwen3 Technical Report
- FinGPT: Instruction Tuning Benchmark for Open-Source Large Language Models in Financial Datasets
- BloombergGPT: A Large Language Model for Finance
- ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection
- FinRobot: An Open-Source AI Agent Platform for Financial Applications using Large Language Models
The paper
Are Large Language Models Reliable Reviewers? A Benchmark for Error Detection in Financial Documents · Read on arXiv
Ying He, Zhouhong Gu, Zhecheng Hu, Yubo Zhou, Hao Shen, Jiaqing Liang, Zhaoqian Dai, Shuguang Ma, Fei Yu, Yanghua Xiao, Zhixu Li
Fudan University · Ant Group · Renmin University of China
Ensuring the accuracy of financial documents is critical for economic analysis, regulatory compliance, and corporate decision-making. Several studies have shown that Large Language Models (LLMs) perform well in many financial tasks, such as stock price movements and financial analytics. However, a critical task remains unexplored: the ability of LLMs to identify errors in financial documents. In this paper, we introduce FinED-Bench, the first publicly Benchmark for Financial Error Detection across three levels of cognitive complexity. FinED-Bench covers nine real-world financial scenarios, and includes over 900 documents reported in 2025 that are unseen by existing language models. We detail the benchmark construction process and evaluate several advanced LLMs (e.g., GPT-4o, Qwen3-14B) on this tasks, which requires both financial domain knowledge and reasoning capabilities. Experimental results show that current LLMs still struggle with this task, especially in high-complexity cases. Besides, supervised fine-tuning can significantly improve the performance of weaker LLMs on this task. Our data and code are available at https://github.com/hedyHe/FinED-Bench.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Are Large Language Models Reliable Reviewers? A Benchmark for Error Detection in Financial Documents".
Jane: The paper was written by Ying He, Zhouhong Gu, Zhecheng Hu, Yubo Zhou, Hao Shen et al. from Fudan University and Ant Group and Renmin University of China.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the show, everybody. I'm Tom, and I've got Jane here with me, and we are looking at a brand new paper that just hit arXiv. It's called "Are Large Language Models Reliable Reviewers? A Benchmark for Error Detection in Financial Documents."
Jane: And Tom, I have to say, that title just gets me. Because we keep hearing about how great these models are at writing financial reports, predicting stock movements, doing all this fancy analysis. But this paper asks the question that nobody's really asked before — can they actually catch mistakes?
Tom: Exactly. And the answer, spoiler alert, is not really. But let's back up for a second. Why does this even matter? I mean, who cares if a model misses a typo in a quarterly report?
Jane: Oh, it's way bigger than a typo. Think about the London Whale scandal back in two thousand twelve. JPMorgan lost six billion dollars because of an Excel error. Not a market crash, not a bad trade — a spreadsheet mistake. So when we're talking about financial documents, errors aren't just embarrassing, they cost real money.
Tom: Six billion. That's wild. And the paper actually cites a Gartner survey showing that eighteen percent of financial practitioners make errors daily, and fifty-nine percent make them several times a month. So the problem is everywhere.
Jane: Right. And that's why the authors built this benchmark called FinED-Bench. They collected over nine hundred real financial documents published in two thousand twenty-five after the knowledge cutoff of most models, so there's no contamination. Then they injected errors and had experts verify them.
Tom: So they're testing whether LLMs can act like a second pair of eyes on financial documents. And the results show that even the best model, GPT-4o, only gets an F1 score of forty-eight point three four percent. That means it's missing more than half the errors.
Jane: And that's the best one. Some of the smaller models are down in the single digits. So the title of the paper — are LLMs reliable reviewers? — the answer is a pretty firm no, at least for now.
Tom: But here's what I find fascinating. They broke the errors into three levels. General knowledge errors, like saying February 30th. Financial domain errors, like mixing up price-to-earnings ratio with price-to-book ratio. And then reasoning errors, where the text in one paragraph contradicts a table later in the document.
Jane: And that third category is where models really struggle. GPT-4o drops from fifty-two percent F1 on general errors to thirty-eight percent on reasoning errors. That's a huge gap. Because catching those requires reading the whole document and connecting dots across paragraphs.
Tom: So the paper is basically saying, hey, these models are great at generating financial text, but they're not ready to audit it. And that's a really important distinction for anyone thinking about using AI in finance.
Jane: Exactly. And we're going to dig into how they built this benchmark and what the experiments actually showed. Because there's some good news buried in here too — fine-tuning helps a lot. But first, let's talk about what this benchmark actually contains.
Summary: Tom: So Jane, we've established that this paper, "Are Large Language Models Reliable Reviewers?" — the short answer is no. But let's get into the details of what FinED-Bench actually is, because it's not just a simple dataset.
Jane: Right. So the authors collected nine hundred seventy-three financial documents, and the average length is about three thousand eight hundred words. But here's the thing — they also built a hard subset with documents ranging from thirty-two thousand to one hundred twenty thousand words. Those are like full prospectuses and legal contracts.
Tom: And within those documents, they've got four thousand one hundred twenty-three annotated error instances. But they didn't just randomly sprinkle errors in. They used a semi-automated pipeline. First, they had domain experts define fifteen different error types across those three categories we mentioned.
Jane: Then they used GPT-4o to actually inject the errors into the documents. But they were smart about it. They segmented the documents first, because you can't just ask a model to inject errors into a one hundred thousand-word document in one go. It won't work well.
Tom: And then they had a two-stage filtering process. First, GPT-4o re-evaluates its own injected errors to make sure they're contextually appropriate. Then, five human experts manually verify everything. They even report a Fleiss' kappa of zero point eight nine, which means the annotators really agreed with each other.
Jane: That's a really high agreement score. It means the benchmark is reliable — the errors are actually errors, and the labels are consistent. And that matters because if the benchmark itself is sloppy, then the evaluation results don't mean anything.
Tom: So what did they find when they ran the models? We mentioned GPT-4o gets forty-eight point three four percent overall F1. But let's talk about the other models. Qwen3-14B gets forty-three point one five percent, Qwen3-8B gets about forty percent. Then you've got GPT-4o-mini at eighteen point nine percent, and some of the smaller open-source models are down in the single digits.
Jane: And here's something really interesting. They tested domain-specific financial models too. There's one called Dianjin-R1-7B and another called Fin-R1. You'd think those would do better, right? They're trained on financial data.
Tom: But they don't. Dianjin-R1-7B gets fifteen point eight one percent F1, and Fin-R1 gets a dismal three point one nine percent. That's actually worse than the general-purpose Qwen3-8B. So the paper makes this point that fine-tuning on financial tasks like question answering or text classification doesn't automatically make a model good at error detection.
Jane: That's because error detection is a different skill. It requires verification, cross-referencing, and skepticism. A model trained to answer questions is trained to be confident. A model trained to detect errors needs to be suspicious. Those are almost opposite mindsets.
Tom: That's a great way to put it. And the paper also shows that reasoning capabilities matter a lot. When they disabled the thinking mode on Qwen3-14B, its F1 dropped from forty-three point one five percent to twenty-seven point one nine percent. That's a massive drop. So the chain-of-thought reasoning is actually doing heavy lifting here.
Jane: So the summary is: current LLMs are not reliable financial reviewers, reasoning helps but isn't enough, and domain-specific fine-tuning doesn't automatically transfer to this task. That's the core finding. But next I want to talk about what they suggest doing about it, because they did find one thing that works.
Improvements: Tom: So we've covered the bad news — these models aren't great at catching financial errors. But the paper, "Are Large Language Models Reliable Reviewers?" — they didn't just stop at showing the problem. They actually tried to fix it. And Jane, I think this is the most exciting part.
Jane: Definitely. So the authors did supervised fine-tuning on Qwen3-14B. They built a training dataset using the same pipeline as the benchmark, but without the human verification step. So it's a larger, noisier dataset, but it's cheap to produce.
Tom: And the results are pretty dramatic. After fine-tuning, Qwen3-14B goes from forty-three point one five percent F1 to fifty-three point eight five percent F1. That's a ten point seven percent absolute improvement. And the gains are especially big on the hardest error types.
Jane: Let me give you a specific example. Conflicting expressions — where two parts of the document say opposite things — that goes from twenty-nine point two one percent to fifty-one point zero six percent. That's nearly doubling. And ambiguous expressions go from two point one three percent to twenty-six point four two percent. That's a twelve-fold improvement.
Tom: Twelve-fold. That's not a small bump. And even the reasoning errors like time contradictions improve from forty-six percent to sixty-seven percent. So fine-tuning on this specific task really does help.
Jane: But here's the question that I always worry about with fine-tuning — does the model forget everything else? If you train it to detect errors, does it stop being good at answering financial questions or doing translation?
Tom: That's exactly what they tested. They evaluated the fine-tuned model on five other financial tasks from a benchmark called CFLUE. Multiple choice questions, financial translation, text classification, relation extraction, and text generation.
Jane: And the results are reassuring. The model's performance stays basically the same. Multiple choice drops slightly from seventy-three point eight six percent to seventy-three point zero eight percent, but translation actually improves a little. Text classification improves from sixty-six point six seven percent to sixty-eight point eight nine percent. So the fine-tuning doesn't hurt generalization.
Tom: That's a really important finding, because it means you can specialize a model for error detection without sacrificing its other capabilities. And that's not always the case with fine-tuning.
Jane: Right. Sometimes you get catastrophic forgetting. But here, the task is similar enough to the model's existing abilities that it just adds a new skill without erasing old ones. And that's encouraging for practical deployment.
Tom: So the improvement strategy is clear. If you want a model that can actually help review financial documents, you need to fine-tune it specifically for error detection. And the data generation pipeline they built makes that feasible, because you don't need human annotation for every training example.
Jane: Exactly. The human verification is only for the benchmark itself. For training data, you can rely on the model-based filtering. And that means you can scale up the training data cheaply. That's a practical path forward.
Tom: So we've got the problem, and we've got a partial solution. But I want to zoom out and look at the actual first page of the paper, because there's some context there that really frames why this matters. Let's do that next.
First Page: Tom: So we're flipping back to the very first page of "Are Large Language Models Reliable Reviewers?" and Jane, there's a figure there that really sets the stage for everything we've been talking about.
Jane: Oh, the figure with the three error types. It shows a concrete example from an actual financial document. And it's perfect because it makes the abstract concepts tangible.
Tom: Right. So the first example is a general knowledge error — the document says something happened on an illegal date. Like February 30th. Any human immediately knows that's wrong. The second is a financial domain error — using the wrong financial term. And the third is a reasoning error where the text claims passenger volume grew ten point nine percent, but the table right there shows five point seven percent.
Jane: And that third one is the killer. Because neither of those numbers is individually wrong. The text says one thing, the table says another. You have to read both and compare them to catch the inconsistency. And that's exactly the kind of error that a model might miss if it's just processing paragraphs in isolation.
Tom: And the paper's introduction also makes this point about why financial error detection is uniquely hard. It's not like checking grammar in an essay. You need domain knowledge, you need to verify numbers, you need to reason across long documents.
Jane: And they also mention the "London Whale" scandal in the intro. Six billion dollars lost because of an Excel error. That's the stakes. And then they cite that Gartner survey — fifty-nine percent of financial practitioners make errors several times a month. So this isn't a rare problem.
Tom: The first page also sets up the structure of the benchmark. They're covering nine real-world financial scenarios — research reports, tender announcements, insurance contracts, legal documents, prospectuses, company bylaws. And the error taxonomy has those fifteen subtypes.
Jane: And one thing I appreciate about the first page is that they're very careful about data contamination. They collected documents published in two thousand twenty-five after the knowledge cutoff of the models they're testing. So the models haven't memorized these documents during training.
Tom: That's a really important methodological point. If you test a model on documents it's already seen, you're not testing its reasoning — you're testing its memory. By using fresh documents, they ensure the evaluation is fair.
Jane: And the first page also previews their key finding — that performance drops as document length increases. They show F1 scores declining from about forty percent on short documents to under seventeen percent on documents over fifty thousand words. That's a brutal drop-off.
Tom: So even the best models are basically useless on the longest documents. And that's a problem because those long documents — prospectuses, legal contracts — are exactly where errors are most costly.
Jane: Right. A typo in a two-page research report is annoying. A contradiction in a one hundred-page bond prospectus could cost investors millions. So the models are failing exactly where we need them most.
Tom: Okay, so we've covered the problem, the benchmark, the results, and the improvement strategy. Let's wrap this up and think about what it all means.
Conclusion: Tom: Alright, we've spent a good chunk of time on "Are Large Language Models Reliable Reviewers?" and I think it's time to pull it all together. Jane, what's the big picture here?
Jane: The big picture is that LLMs are not ready to be financial auditors. The best model gets forty-eight percent F1, which means it misses more than half the errors. And on long documents, it's even worse. But the paper gives us a roadmap for improvement — fine-tuning on error detection data can push that up to nearly fifty-four percent.
Tom: And that's a meaningful improvement. But it's still not good enough for real-world deployment. If you're a compliance officer and the model misses half the errors, you still have to read everything yourself. So the model becomes a helper, not a replacement.
Jane: And that's actually the right framing. The paper isn't saying AI is useless for financial document review. It's saying we need to be realistic about its capabilities and invest in task-specific training.
Tom: I also think the benchmark itself is a contribution. FinED-Bench gives researchers a standardized way to measure progress. So as models improve, we can track whether they're actually getting better at this task.
Jane: And the data pipeline is reusable. The semi-automated construction with error injection and model-based filtering means you can generate new training data as documents evolve. That's important because financial language changes.
Tom: So what's the takeaway for our listeners? If you're working in finance and thinking about using AI to review documents, the message is: be cautious, but don't give up. The models aren't there yet, but the direction is promising.
Jane: And if you're a researcher, this paper opens up a clear research direction. How do we improve reasoning on long documents? How do we make error detection more robust? Those are open questions.
Tom: Alright, we've covered a lot of ground. The paper gives us a benchmark, a sobering evaluation, and a path forward. That's a solid contribution.
Jane: Agreed. And with that, we'll say goodbye to "Are Large Language Models Reliable Reviewers?" and get ready to look at what's next on arXiv. Thanks for listening, everyone.
Tom: See you on the next episode.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language