Are Large Language Models Reliable Reviewers? A Benchmark for Error Detection in Financial Documents

arXiv:2608.12342 · cs.CL, cs.LG · Submitted 2026-06-03 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Are Large Language Models Reliable Reviewers? A Benchmark for Error Detection in Financial Documents".

Jane: The paper was written by Ying He, Zhouhong Gu, Zhecheng Hu, Yubo Zhou, Hao Shen et al. from Fudan University and Ant Group and Renmin University of China.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everybody. I'm Tom, and I've got Jane here with me, and we are looking at a brand new paper that just hit arXiv. It's called "Are Large Language Models Reliable Reviewers? A Benchmark for Error Detection in Financial Documents."

Jane: And Tom, I have to say, that title just gets me. Because we keep hearing about how great these models are at writing financial reports, predicting stock movements, doing all this fancy analysis. But this paper asks the question that nobody's really asked before — can they actually catch mistakes?

Tom: Exactly. And the answer, spoiler alert, is not really. But let's back up for a second. Why does this even matter? I mean, who cares if a model misses a typo in a quarterly report?

Jane: Oh, it's way bigger than a typo. Think about the London Whale scandal back in two thousand twelve. JPMorgan lost six billion dollars because of an Excel error. Not a market crash, not a bad trade — a spreadsheet mistake. So when we're talking about financial documents, errors aren't just embarrassing, they cost real money.

Tom: Six billion. That's wild. And the paper actually cites a Gartner survey showing that eighteen percent of financial practitioners make errors daily, and fifty-nine percent make them several times a month. So the problem is everywhere.

Jane: Right. And that's why the authors built this benchmark called FinED-Bench. They collected over nine hundred real financial documents published in two thousand twenty-five after the knowledge cutoff of most models, so there's no contamination. Then they injected errors and had experts verify them.

Tom: So they're testing whether LLMs can act like a second pair of eyes on financial documents. And the results show that even the best model, GPT-4o, only gets an F1 score of forty-eight point three four percent. That means it's missing more than half the errors.

Jane: And that's the best one. Some of the smaller models are down in the single digits. So the title of the paper — are LLMs reliable reviewers? — the answer is a pretty firm no, at least for now.

Tom: But here's what I find fascinating. They broke the errors into three levels. General knowledge errors, like saying February 30th. Financial domain errors, like mixing up price-to-earnings ratio with price-to-book ratio. And then reasoning errors, where the text in one paragraph contradicts a table later in the document.

Jane: And that third category is where models really struggle. GPT-4o drops from fifty-two percent F1 on general errors to thirty-eight percent on reasoning errors. That's a huge gap. Because catching those requires reading the whole document and connecting dots across paragraphs.

Tom: So the paper is basically saying, hey, these models are great at generating financial text, but they're not ready to audit it. And that's a really important distinction for anyone thinking about using AI in finance.

Jane: Exactly. And we're going to dig into how they built this benchmark and what the experiments actually showed. Because there's some good news buried in here too — fine-tuning helps a lot. But first, let's talk about what this benchmark actually contains.

Summary: Tom: So Jane, we've established that this paper, "Are Large Language Models Reliable Reviewers?" — the short answer is no. But let's get into the details of what FinED-Bench actually is, because it's not just a simple dataset.

Jane: Right. So the authors collected nine hundred seventy-three financial documents, and the average length is about three thousand eight hundred words. But here's the thing — they also built a hard subset with documents ranging from thirty-two thousand to one hundred twenty thousand words. Those are like full prospectuses and legal contracts.

Tom: And within those documents, they've got four thousand one hundred twenty-three annotated error instances. But they didn't just randomly sprinkle errors in. They used a semi-automated pipeline. First, they had domain experts define fifteen different error types across those three categories we mentioned.

Jane: Then they used GPT-4o to actually inject the errors into the documents. But they were smart about it. They segmented the documents first, because you can't just ask a model to inject errors into a one hundred thousand-word document in one go. It won't work well.

Tom: And then they had a two-stage filtering process. First, GPT-4o re-evaluates its own injected errors to make sure they're contextually appropriate. Then, five human experts manually verify everything. They even report a Fleiss' kappa of zero point eight nine, which means the annotators really agreed with each other.

Jane: That's a really high agreement score. It means the benchmark is reliable — the errors are actually errors, and the labels are consistent. And that matters because if the benchmark itself is sloppy, then the evaluation results don't mean anything.

Tom: So what did they find when they ran the models? We mentioned GPT-4o gets forty-eight point three four percent overall F1. But let's talk about the other models. Qwen3-14B gets forty-three point one five percent, Qwen3-8B gets about forty percent. Then you've got GPT-4o-mini at eighteen point nine percent, and some of the smaller open-source models are down in the single digits.

Jane: And here's something really interesting. They tested domain-specific financial models too. There's one called Dianjin-R1-7B and another called Fin-R1. You'd think those would do better, right? They're trained on financial data.

Tom: But they don't. Dianjin-R1-7B gets fifteen point eight one percent F1, and Fin-R1 gets a dismal three point one nine percent. That's actually worse than the general-purpose Qwen3-8B. So the paper makes this point that fine-tuning on financial tasks like question answering or text classification doesn't automatically make a model good at error detection.

Jane: That's because error detection is a different skill. It requires verification, cross-referencing, and skepticism. A model trained to answer questions is trained to be confident. A model trained to detect errors needs to be suspicious. Those are almost opposite mindsets.

Tom: That's a great way to put it. And the paper also shows that reasoning capabilities matter a lot. When they disabled the thinking mode on Qwen3-14B, its F1 dropped from forty-three point one five percent to twenty-seven point one nine percent. That's a massive drop. So the chain-of-thought reasoning is actually doing heavy lifting here.

Jane: So the summary is: current LLMs are not reliable financial reviewers, reasoning helps but isn't enough, and domain-specific fine-tuning doesn't automatically transfer to this task. That's the core finding. But next I want to talk about what they suggest doing about it, because they did find one thing that works.

Improvements: Tom: So we've covered the bad news — these models aren't great at catching financial errors. But the paper, "Are Large Language Models Reliable Reviewers?" — they didn't just stop at showing the problem. They actually tried to fix it. And Jane, I think this is the most exciting part.

Jane: Definitely. So the authors did supervised fine-tuning on Qwen3-14B. They built a training dataset using the same pipeline as the benchmark, but without the human verification step. So it's a larger, noisier dataset, but it's cheap to produce.

Tom: And the results are pretty dramatic. After fine-tuning, Qwen3-14B goes from forty-three point one five percent F1 to fifty-three point eight five percent F1. That's a ten point seven percent absolute improvement. And the gains are especially big on the hardest error types.

Jane: Let me give you a specific example. Conflicting expressions — where two parts of the document say opposite things — that goes from twenty-nine point two one percent to fifty-one point zero six percent. That's nearly doubling. And ambiguous expressions go from two point one three percent to twenty-six point four two percent. That's a twelve-fold improvement.

Tom: Twelve-fold. That's not a small bump. And even the reasoning errors like time contradictions improve from forty-six percent to sixty-seven percent. So fine-tuning on this specific task really does help.

Jane: But here's the question that I always worry about with fine-tuning — does the model forget everything else? If you train it to detect errors, does it stop being good at answering financial questions or doing translation?

Tom: That's exactly what they tested. They evaluated the fine-tuned model on five other financial tasks from a benchmark called CFLUE. Multiple choice questions, financial translation, text classification, relation extraction, and text generation.

Jane: And the results are reassuring. The model's performance stays basically the same. Multiple choice drops slightly from seventy-three point eight six percent to seventy-three point zero eight percent, but translation actually improves a little. Text classification improves from sixty-six point six seven percent to sixty-eight point eight nine percent. So the fine-tuning doesn't hurt generalization.

Tom: That's a really important finding, because it means you can specialize a model for error detection without sacrificing its other capabilities. And that's not always the case with fine-tuning.

Jane: Right. Sometimes you get catastrophic forgetting. But here, the task is similar enough to the model's existing abilities that it just adds a new skill without erasing old ones. And that's encouraging for practical deployment.

Tom: So the improvement strategy is clear. If you want a model that can actually help review financial documents, you need to fine-tune it specifically for error detection. And the data generation pipeline they built makes that feasible, because you don't need human annotation for every training example.

Jane: Exactly. The human verification is only for the benchmark itself. For training data, you can rely on the model-based filtering. And that means you can scale up the training data cheaply. That's a practical path forward.

Tom: So we've got the problem, and we've got a partial solution. But I want to zoom out and look at the actual first page of the paper, because there's some context there that really frames why this matters. Let's do that next.

First Page: Tom: So we're flipping back to the very first page of "Are Large Language Models Reliable Reviewers?" and Jane, there's a figure there that really sets the stage for everything we've been talking about.

Jane: Oh, the figure with the three error types. It shows a concrete example from an actual financial document. And it's perfect because it makes the abstract concepts tangible.

Tom: Right. So the first example is a general knowledge error — the document says something happened on an illegal date. Like February 30th. Any human immediately knows that's wrong. The second is a financial domain error — using the wrong financial term. And the third is a reasoning error where the text claims passenger volume grew ten point nine percent, but the table right there shows five point seven percent.

Jane: And that third one is the killer. Because neither of those numbers is individually wrong. The text says one thing, the table says another. You have to read both and compare them to catch the inconsistency. And that's exactly the kind of error that a model might miss if it's just processing paragraphs in isolation.

Tom: And the paper's introduction also makes this point about why financial error detection is uniquely hard. It's not like checking grammar in an essay. You need domain knowledge, you need to verify numbers, you need to reason across long documents.

Jane: And they also mention the "London Whale" scandal in the intro. Six billion dollars lost because of an Excel error. That's the stakes. And then they cite that Gartner survey — fifty-nine percent of financial practitioners make errors several times a month. So this isn't a rare problem.

Tom: The first page also sets up the structure of the benchmark. They're covering nine real-world financial scenarios — research reports, tender announcements, insurance contracts, legal documents, prospectuses, company bylaws. And the error taxonomy has those fifteen subtypes.

Jane: And one thing I appreciate about the first page is that they're very careful about data contamination. They collected documents published in two thousand twenty-five after the knowledge cutoff of the models they're testing. So the models haven't memorized these documents during training.

Tom: That's a really important methodological point. If you test a model on documents it's already seen, you're not testing its reasoning — you're testing its memory. By using fresh documents, they ensure the evaluation is fair.

Jane: And the first page also previews their key finding — that performance drops as document length increases. They show F1 scores declining from about forty percent on short documents to under seventeen percent on documents over fifty thousand words. That's a brutal drop-off.

Tom: So even the best models are basically useless on the longest documents. And that's a problem because those long documents — prospectuses, legal contracts — are exactly where errors are most costly.

Jane: Right. A typo in a two-page research report is annoying. A contradiction in a one hundred-page bond prospectus could cost investors millions. So the models are failing exactly where we need them most.

Tom: Okay, so we've covered the problem, the benchmark, the results, and the improvement strategy. Let's wrap this up and think about what it all means.

Conclusion: Tom: Alright, we've spent a good chunk of time on "Are Large Language Models Reliable Reviewers?" and I think it's time to pull it all together. Jane, what's the big picture here?

Jane: The big picture is that LLMs are not ready to be financial auditors. The best model gets forty-eight percent F1, which means it misses more than half the errors. And on long documents, it's even worse. But the paper gives us a roadmap for improvement — fine-tuning on error detection data can push that up to nearly fifty-four percent.

Tom: And that's a meaningful improvement. But it's still not good enough for real-world deployment. If you're a compliance officer and the model misses half the errors, you still have to read everything yourself. So the model becomes a helper, not a replacement.

Jane: And that's actually the right framing. The paper isn't saying AI is useless for financial document review. It's saying we need to be realistic about its capabilities and invest in task-specific training.

Tom: I also think the benchmark itself is a contribution. FinED-Bench gives researchers a standardized way to measure progress. So as models improve, we can track whether they're actually getting better at this task.

Jane: And the data pipeline is reusable. The semi-automated construction with error injection and model-based filtering means you can generate new training data as documents evolve. That's important because financial language changes.

Tom: So what's the takeaway for our listeners? If you're working in finance and thinking about using AI to review documents, the message is: be cautious, but don't give up. The models aren't there yet, but the direction is promising.

Jane: And if you're a researcher, this paper opens up a clear research direction. How do we improve reasoning on long documents? How do we make error detection more robust? Those are open questions.

Tom: Alright, we've covered a lot of ground. The paper gives us a benchmark, a sobering evaluation, and a path forward. That's a solid contribution.

Jane: Agreed. And with that, we'll say goodbye to "Are Large Language Models Reliable Reviewers?" and get ready to look at what's next on arXiv. Thanks for listening, everyone.

Tom: See you on the next episode.

Ying He, Zhouhong Gu, Zhecheng Hu, Yubo Zhou, Hao Shen, Jiaqing Liang, Zhaoqian Dai, Shuguang Ma, Fei Yu, Yanghua Xiao, Zhixu Li

Fudan University · Ant Group · Renmin University of China

cs.CL, cs.LG

Submitted: 2026-06-03

Updated: 2026-08-14

Code: https://github.com/hedyHe/FinED-Bench

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 63/100

The gist: This paper introduces FinED-Bench, the first publicly available benchmark for financial error detection, designed to evaluate the capabilities of Large Language Models (LLMs) in identifying errors

Key concepts

FinED-Bench
A benchmark created by the paper's authors using over nine hundred real financial documents published after most LLM knowledge cutoffs. It was designed to test if models can reliably detect injected errors in financial texts.
Reasoning Errors
A difficult type of error where the text in one part of a document contradicts information found elsewhere, such as a discrepancy between the written text and an embedded table. Models struggle most with this category.
F1 Score
A metric used to evaluate model performance in error detection. A higher F1 score indicates better overall accuracy, meaning the model is successfully identifying a high percentage of actual errors.
Fine-tuning
The process of training an existing LLM (like Qwen3-14B) on a specific, targeted task—in this case, error detection. This specialized training significantly boosts performance without harming the model's general capabilities.

Terminology

Summary

This paper introduces FinED-Bench, the first publicly available benchmark for financial error detection, designed to evaluate the capabilities of Large Language Models (LLMs) in identifying errors within financial documents. The benchmark addresses a critical gap, as previous work focused on grammatical errors or hallucination detection in short, general-domain texts, failing to assess the numerical and long-context reasoning required for professional financial error detection.

The benchmark is built on a two-tier error taxonomy with 3 categories and 15 subcategories, developed through collaboration between finance domain experts and LLMs. The three main categories are: (1) General Knowledge Errors (GKEs), which affect readability and surface structure and can be detected by non-experts; (2) Financial Domain Knowledge Errors (FKEs), which distort intended meaning and require domain-specific knowledge; and (3) Financial Reasoning Errors (FREs), which are logical inconsistencies requiring document-level comprehension and reasoning. The 15 subcategories include illegal time, redundant statements, value format errors, numerical missing, non-numerical attribute value missing, terminology misuse, incorrect legal reference, ambiguous expression, numerical unit error, omitted financial element, conflicting expression, time contradiction, numerical inconsistency, calculation error, and clause conflict.

FinED-Bench contains 973 realistic financial documents with an average length of 3,784.6 words, containing 4,123 annotated error instances. All errors were annotated and verified by experienced financial experts. The benchmark covers nine real-world financial scenarios, including stock research reports, industry research reports, tender announcements, insurance contracts, legal documents, regulatory documents, company bylaws, listing prospectuses, and bond prospectuses. To avoid data contamination, all documents were published after February 2025, postdating the knowledge cutoffs of evaluated models. Additionally, a challenging subset, FinED-Bench-Hard, was constructed with 24 documents ranging from 32K to 120K words containing 83 error instances.

The dataset construction process uses a semi-automated pipeline. Financial documents are collected and parsed, then segmented into semantically complete fragments. Error generation is performed using In-Context Learning (ICL) with GPT-4o, where error types vary by financial scene. Generated errors undergo a two-stage filtering process: model-based filtering where GPT-4o re-evaluates errors for contextual appropriateness, and manual verification by a team of five experts who remove errors that may introduce unintended errors and refine error seeds. The number of error instances per document is controlled to ensure realism.

Experimental results reveal that current LLMs struggle significantly with this task. The best-performing model, GPT-4o, achieves an overall F1 score of 48.34%, with performance dropping substantially across error categories: from 52.33% for general knowledge errors to 38.00% for financial reasoning errors. This performance degradation is consistent across all models, highlighting limitations in handling complex financial logic. Qwen3-14B drops from 51.41% to 29.10% across these categories, and Qwen3-8B from 47.60% to 27.96%. The paper states: "Financial reasoning errors are more challenging for LLMs because they require multi-step calculations and the integration of financial-specific regulations, whereas the other two types of errors rely primarily on commonsense reasoning and factual recall."

Document length severely impacts performance, with F1 scores declining from 40.16% to 16.66% as document length increases from 2.5K to 50.2K words. The highest F1 scores occur in Stock Research Reports (58.07%) and Industry Research Reports (59.68%), both very short. Performance drops to 44.61% for Tender Announcements, then further to 38.85% for Insurance Contracts and 29.17% for Legal Documents. The longest document types perform worst: Company Bylaws (16.66%), Listing Prospectuses (14.76%), and Bond Prospectuses (9.52%).

The paper also finds that reasoning capabilities significantly enhance LLM performance on error detection. The reasoning-ablated (no-thinking) variants of Qwen3-8B and Qwen3-14B consistently underperform their full counterparts across all error categories. Specifically, Qwen3-8B without reasoning achieves 24.93% F1 on General Knowledge Errors compared to 47.60% with reasoning, while Qwen3-14B shows a similar pattern (33.72% vs. 51.41%).

GPT-4o shows a high-recall but low-precision pattern, primarily due to its overly sensitive detection. The paper notes: "GPT-4o tends to identify a much larger set of candidate errors, which greatly improves its recall but at the cost of precision. As a no-thinking model, GPT-4o struggles to ensure internal inconsistency in its response, leading to numerous false positive and misclassified error types."

The performance of domain-specific LLMs is largely constrained by the capabilities of their base models. Dianjin-R1-7B, a financial-domain LLM fine-tuned from Qwen2.5-7B-Instruct, achieves a higher overall F1 score than its base model (15.81% vs. 9.85%). In contrast, Fin-R1 exhibits a decline in performance, with its overall F1 score dropping from 9.85% to 3.19%. Despite these changes, both remain substantially below the performance of recent general-purpose LLMs such as Qwen3-8B (39.99%) and Qwen3-14B (43.15%). The paper concludes that the foundational capabilities of base models largely determine the upper bound of performance achievable through domain-specific fine-tuning and that fine-tuning on related tasks, such as QA, text summarization, or classification, does not improve the error detection ability of models.

Supervised fine-tuning can significantly improve the performance of weaker LLMs on this task. After fine-tuning on financial error detection data, Qwen3-14B achieves a 10.70% improvement in overall F1 score (from 43.15% to 53.85%). The improvements are particularly notable for challenging error types: Conflicting Expression increases from 29.21% to 51.06%, and Ambiguous Expression rises from 2.13% to 26.42%. Fine-tuning also enhances performance on format-based errors, with Illegal Time improving from 70.53% to 80.00% and Numerical Missing increasing from 52.48% to 65.99%. Even complex reasoning errors show gains, with Time Contradiction improving from 46.07% to 67.32% and Clause Conflict rising from 25.53% to 29.36%. Importantly, fine-tuning preserves the model's generalization ability, as evaluated across five financial tasks from CFLUE, with the fine-tuned model maintaining comparable performance and in several cases showing slight improvements.

The paper also includes a human baseline study involving two sophomore students majoring in finance, who achieved an overall F1 score of 63.63%, indicating that even humans with relevant domain background find accurately identifying all errors in long financial documents highly challenging.

The limitations acknowledged in the paper include: (1) Diversity of Financial Documents - many financial documents such as balance sheets and income statements are not covered, as errors in these often originate from underlying data sources requiring extensive historical data review; (2) Multimodal Elements - real-world financial documents often contain visual elements such as seals and signatures that require multimodal capabilities beyond text-only models. The paper also notes ethical concerns regarding sensitive information in the benchmark, stating that significant effort was invested to replace real data with carefully crafted synthetic alternatives.

Improvements for AI systems

Based on the paper, here are specific improvements for AI systems:

1. Financial Error Detection Module

  • What to add: A dedicated error-detection layer that classifies text into three hierarchical levels: General Knowledge Errors (format, missing values), Financial Domain Knowledge Errors (terminology, legal references), and Financial Reasoning Errors (contradictions, calculation inconsistencies).

  • Improved capability: The AI can now scan lengthy financial documents (up to 120K words) and flag subtle errors like February 30, 2025 (illegal time), price-to-book ratio instead of price-to-earnings ratio (terminology misuse), or cross-paragraph numerical inconsistencies (e.g., 10.9% growth claimed in text vs. 5.7% in a table).

2. Long-Context Reasoning Enhancement

  • What to add: Implement a two-stage processing pipeline: (1) segment documents into semantically complete fragments, (2) use a reasoning-enabled model (e.g., Qwen3-14B with thinking mode) to perform cross-fragment verification.

  • Improved capability: The AI can now detect errors that span distant sections of a document, such as a clause conflict in Article 10 vs. Article 11 of a contract, or a time contradiction between a submission deadline and bid opening date.

3. Supervised Fine-Tuning for Financial Error Detection

  • What to add: Fine-tune a base LLM (e.g., Qwen3-14B) on a dataset of 8,697 error-containing fragments and 9,515 error-free fragments, each annotated with error type, error span, and reasoning process.

  • Improved capability: The fine-tuned model achieves a 10.70% improvement in overall F1 score (from 43.15% to 53.85%), with notable gains in challenging categories: Conflicting Expression (29.21% → 51.06%), Ambiguous Expression (2.13% → 26.42%), and Time Contradiction (46.07% → 67.32%).

4. Error Type-Specific Detection Strategies

  • What to add: Implement distinct detection strategies per error category:

  • For calculation errors: prompt the model to output the calculation formula, not just a yes/no judgment.

  • For reasoning errors: merge adjacent paragraphs that share overlapping numerical values or terminology.

  • For general knowledge errors: use whole-document prompting rather than chunking.

  • Improved capability: The AI can now reliably identify calculation errors (e.g., Annual demand: 285,995 units where the sum of subcomponents is 276,015) and detect errors that require multi-hop reasoning across document sections.

5. Position-Aware Error Detection

  • What to add: Implement a weighting mechanism that gives higher attention to the beginning and ending sections of documents, based on the finding that errors in the middle are most frequently missed.

  • Improved capability: The AI can now achieve more balanced error detection across the entire document, reducing the 30-40% performance gap observed between middle-section errors and beginning/ending errors.

6. Human-in-the-Loop Verification

  • What to add: Integrate a two-stage filtering process: (1) model-based verification where the AI re-evaluates its own error injections for contextual appropriateness, and (2) manual verification by financial experts to remove false positives and refine error definitions.

  • Improved capability: The AI system can now produce benchmark-quality annotations with high inter-annotator agreement (Fleiss' κ = 0.89), ensuring that flagged errors are genuinely incorrect and not just stylistic variations.

7. Domain-Specific Fine-Tuning Preservation

  • What to add: When fine-tuning for error detection, evaluate the model on five auxiliary financial tasks (MCQs, translation, text classification, relation extraction, text generation) to ensure no catastrophic forgetting.

  • Improved capability: The fine-tuned model maintains or slightly improves performance on other financial tasks (e.g., MCQ accuracy drops only 0.78%, while financial text classification improves by 2.22%), demonstrating that error detection capability can be added without degrading general financial competence.

Abstract

Ensuring the accuracy of financial documents is critical for economic analysis, regulatory compliance, and corporate decision-making. Several studies have shown that Large Language Models (LLMs) perform well in many financial tasks, such as stock price movements and financial analytics. However, a critical task remains unexplored: the ability of LLMs to identify errors in financial documents. In this paper, we introduce FinED-Bench, the first publicly Benchmark for Financial Error Detection across three levels of cognitive complexity. FinED-Bench covers nine real-world financial scenarios, and includes over 900 documents reported in 2025 that are unseen by existing language models. We detail the benchmark construction process and evaluate several advanced LLMs (e.g., GPT-4o, Qwen3-14B) on this tasks, which requires both financial domain knowledge and reasoning capabilities. Experimental results show that current LLMs still struggle with this task, especially in high-complexity cases. Besides, supervised fine-tuning can significantly improve the performance of weaker LLMs on this task. Our data and code are available at https://github.com/hedyHe/FinED-Bench.

Sources

Related papers