HC-RAG: Evidence-Centric Retrieval-Augmented Generation over Heterogeneous Financial Filings
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "HC-RAG: Evidence-Centric Retrieval-Augmented Generation over Heterogeneous Financial Filings".
Jane: The paper was written by Siyuan Chen, Huaye Tan, You Li and Jiajun Liang from Sun Yat-sen University and Central South University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Hey everyone, welcome back to the show. Today we’re cracking open a fresh one from arXiv called “HC-RAG: Evidence-Centric Retrieval-Augmented Generation over Heterogeneous Financial Filings.” Jane, I gotta say, the title alone tells me these folks are trying to fix something real.
Jane: Absolutely, Tom. And the authors—Siyuan Chen, Huaye Tan, You Li, Jiajun Liang—they’re from Sun Yat-sen University and Central South University. They’re not just throwing another chatbot at us. They’re looking at how you actually answer questions about financial documents, like annual reports.
Tom: Right, and that’s a big deal because these reports are massive. We’re talking ten-K filings with tens of thousands of tokens. If you just chop them up into little chunks and search, you lose the structure. You lose the sections, the tables, the context.
Jane: Exactly. And that’s what the “heterogeneous” part means. You’ve got text, you’ve got tables, you’ve got metadata about companies and fiscal years. A regular search engine treats all that the same way, and it just doesn’t work. The paper’s whole point is that evidence in finance is structured, and your retrieval should be too.
Tom: So they built a graph. A financial evidence graph. Documents connect to sections, sections connect to text units and table units, and then there are edges linking companies and years. It’s like a map of the filing, not just a pile of paper.
Jane: And that map lets you retrieve the way an analyst would. First find the right company, then the right section, then the specific table or paragraph. It’s step-by-step, not a wild guess across the whole corpus.
Tom: I love that. It’s like going to a library, but instead of wandering the aisles, you know the exact floor, the exact shelf, and the exact book. The implications here are huge for anyone who does financial research for a living.
Jane: Or for anyone building AI tools for analysts. This could change how we build question-answering systems for finance, and maybe even for legal or medical documents that have similar structure. We’ll dig into how they actually did it next.
Summary of the Paper: Tom: So Jane, we’ve got the title and the authors down. Now let’s talk about what this paper actually does. And I want to bring in Lu, our senior researcher, because I think she’ll appreciate the architecture here.
Jane: Good idea. Lu, the paper’s core idea is that financial QA isn’t just about finding similar text. It’s about finding verifiable evidence. Can you break that down for us?
Lu: Sure, Jane. The authors argue that if you ask “what was Apple’s current ratio in two thousand twenty-four” you don’t just need any text that mentions Apple. You need the exact balance sheet, the exact row, the exact value. So they built a three-level retrieval system: first documents, then sections, then evidence units like tables or paragraphs.
Tom: And that’s the “hierarchical” part of HC-RAG. It narrows the search space step by step. Instead of comparing the question to every chunk in the corpus, it first picks the right filings, then the right sections, then the right evidence. That’s way more efficient and way more accurate.
Jane: But here’s the clever part, Lu. They also align text and tables in the same embedding space. So a table about revenue and a paragraph about revenue are mapped close together. That way, when you search, you can find both, and the system can decide which one matters more.
Lu: Exactly. And they don’t just use a fixed mix of text and table. They classify the question into four intents: calculation, trend, fact, or comparison. A calculation question gets more weight on tables. A trend question gets more weight on narrative text. It’s query-aware routing.
Tom: And they built a whole new benchmark to test this, called Multi-Doc-two thousand twenty-five. It’s got over two thousand three hundred questions from one hundred seventy-nine real ten-K filings of S andP five hundred companies. And it’s designed to test cross-company and cross-year questions, which most benchmarks ignore.
Lu: That’s the part I find most exciting. Most financial QA datasets are single-document. You read one report and answer. But real analysts compare companies. They track trends across years. This benchmark finally tests that.
Jane: And the results? Tom, you want to share the numbers?
Tom: Oh, absolutely. On their benchmark, HC-RAG hits sixty point two F1, which beats GraphRAG by almost eleven points and TAPEX-RAG by six points. And on DocFinQA, a long-document benchmark, they beat RAPTOR by six point six F1 points. Those are big jumps.
Lu: And the evidence localization results are even more telling. Their table hit rate is eighty-nine point six percent, compared to sixty-six point eight percent for the best baseline. That’s a massive improvement in finding the right table, which is exactly what financial QA needs.
Jane: So the summary is: they built a graph, they aligned text and tables, they routed by intent, and they proved it works. Next, let’s talk about what this means for real-world systems.
Improvements Suggested by the Paper: Tom: Alright, so we know what HC-RAG does. But what does it actually improve in practice? And Meng, I want you in on this because you’re the engineer who has to make things run.
Meng: Happy to be here, Tom. And honestly, the first thing that jumps out at me is efficiency. The paper reports index construction at four point eight minutes per document, which is faster than GraphRAG’s seven point six minutes. And inference is three point one seconds per query, which is totally usable in a real product.
Jane: That’s a big deal, Meng. Because a lot of these fancy graph systems are research prototypes. They work in a lab but they’re too slow or too expensive to deploy. HC-RAG seems to hit a sweet spot.
Meng: Exactly. And the asymmetric design is smart. They use a heavy table-aware encoder offline to align text and tables. But online, they flatten tables into short strings with row headers, column headers, and values. So you get the benefit of table understanding without the cost of running a big model on every query.
Lu: And that’s not just an engineering trick, Meng. It’s a design philosophy. They’re saying: do the expensive work once, during indexing, and keep the online retrieval light. That’s how you scale to thousands of filings.
Tom: The paper also shows robustness to noise. They added irrelevant evidence to the context, and HC-RAG degraded much slower than flat retrieval systems. Because the hierarchy acts as a filter—noise has to pass through document and section gates before it reaches the answer.
Meng: That’s huge for real-world use. In practice, you don’t always have clean queries. People ask vague questions. The system needs to not panic and pull in garbage. HC-RAG’s structure keeps it grounded.
Jane: And the intent routing, Lu, that’s another improvement. Instead of treating every question the same, the system adapts. For a calculation question, it pulls tables. For a trend question, it pulls MD andA text. The paper shows the routing weight actually shifts, from zero point two eight for a numerical question to zero point seven one for a trend question.
Lu: That’s the kind of adaptivity that makes a system feel intelligent. It’s not just retrieving more context. It’s retrieving the right kind of context for the task. And that’s a lesson that goes beyond finance.
Meng: Yeah, I could see this applied to legal research or medical guidelines. Any domain where documents have structure and where evidence type matters. The framework is domain-agnostic even though the benchmark is finance-specific.
Tom: So the improvements are: better accuracy, better evidence localization, better efficiency, and better robustness. That’s a full package. Now, let’s wrap up and think about what this means for the future.
Conclusion: Tom: Alright, we’re at the end of our time with “HC-RAG: Evidence-Centric Retrieval-Augmented Generation over Heterogeneous Financial Filings.” Jane, give us the final summary.
Jane: Sure, Tom. This paper tackles a real problem: answering questions over long, structured financial filings. It builds a graph that preserves the document-section-unit hierarchy, aligns text and tables in one space, and routes evidence based on query intent. And it proves the approach works with a new benchmark, Multi-Doc-two thousand twenty-five.
Lu: And the results are compelling. They beat strong baselines on both answer quality and evidence localization. The table hit rate of nearly ninety percent is the kind of number that makes you sit up and pay attention.
Meng: From an engineering standpoint, the efficiency numbers are just as important. You can index a filing in under five minutes and answer a query in about three seconds. That’s deployable. That’s a product.
Tom: And the implications go beyond finance. The idea of evidence-centric retrieval, of respecting document structure and adapting to query intent, that’s a blueprint for any domain where answers need to be verifiable.
Jane: We should also mention the limitations. The paper admits it relies on regular SEC-style filings. Scanned PDFs or irregular layouts would be harder. And the intent classifier can make mistakes, which would affect routing.
Lu: But those are next steps, not dead ends. The authors suggest better document parsing, better table understanding, and extending the graph to other domains. This is a solid foundation.
Tom: So we say goodbye to HC-RAG. It’s a strong contribution to financial AI and to retrieval-augmented generation in general. We’ll be watching for follow-up work.
Jane: Thanks for listening, everyone. We’ve got another paper lined up next, so stay tuned. Until then, keep asking good questions.
Siyuan Chen, Huaye Tan, You Li, Jiajun Liang
Sun Yat-sen University · Central South University
cs.CL, cs.MM
Submitted: 2026-06-03
Comments: 16 pages, 5 figures
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 69/100
The gist: vD → vS → vU, where vD is a document node, vS is a section node, and vU is a text or table evidence-unit node.
Key concepts
- Heterogeneous Financial Filings
- These are large financial documents (like 10-K reports) containing different data types, including text paragraphs, structured tables, and company metadata. HC-RAG is designed to handle this mix of data simultaneously.
- Evidence-Centric Retrieval
- Instead of searching for similar text chunks, this method focuses on finding verifiable evidence by following the document's structure. It narrows the search step-by-step (document $\rightarrow$ section $\rightarrow$ unit) to ensure accuracy.
- Graph Structure
- The system models the filing as a map where documents connect to sections, and sections connect to specific text or table units. This preserves the natural hierarchy of the financial report.
- Intent Routing
- The system classifies a user's question into one of four intents (calculation, trend, fact, comparison). This classification determines whether the retrieval process should prioritize tables or narrative text.
Terminology
Summary
Summary
The paper introduces HC-RAG, a hierarchical cross-modal retrieval-augmented generation framework designed for evidence-centric financial question answering over heterogeneous financial filings, specifically SEC Form 10-K annual reports. The authors identify that financial QA requires more than retrieving semantically similar passages; it involves identifying relevant companies and fiscal years, locating standardized filing sections, collecting textual and tabular evidence, and checking answers against original documents. Existing RAG systems flatten long filings into unordered chunks, pay limited attention to the typed structure of financial reports, and use fixed text-table fusion strategies without considering query intent.
To address these limitations, HC-RAG organizes filings into a typed financial evidence graph
with documents, sections, text units, table units, and metadata nodes. The graph is formally defined as G = (V, E, X), where V denotes all nodes, E denotes relations between nodes, and X stores text, table content, and metadata. Nodes include document nodes, section nodes, text-unit nodes, table-unit nodes, and metadata nodes. Edges describe containment (documents to sections, sections to evidence units), sequential order, cross edges (related companies, fiscal years, sectors, financial metrics), and modal edges connecting narrative descriptions with related table evidence. Retrieval follows document-section-unit paths, formalized as P: vD → vS → vU, where vD is a document node, vS is a section node, and vU is a text or table evidence-unit node.
The retrieval process is hierarchical and occurs in three steps. At the document level, HC-RAG ranks filing nodes by combining semantic similarity and metadata matching: sD(q, vD) = sim(q, vD) + αD match(q, vD). At the section level, it ranks sections under selected documents: sS(q, vS) = sim(q, vS) + αS prior(q, vS), where prior represents simple section preference from the question. At the evidence-unit level, it retrieves text and table units from selected sections: sU(q, vU) = sim(q, vU) + βgraph(q, vU), where graph gives extra weight to evidence units on valid document-section-unit paths, matching company and fiscal year, or having useful text-table connections. The final evidence set is selected as Eq∗ = TopK(sU(q, vU)).
HC-RAG uses an asymmetric offline-online strategy for cross-modal alignment. In the offline stage, a financial text encoder (FinBERT) encodes textual evidence, while a table-aware encoder (TAPAS or TAPEX) encodes table evidence, considering row and column information. Text and table units from the same section are treated as related pairs, and their representations are mapped into the same retrieval space. The alignment goal is Lalign = 1 − sim(z t, z b), where z t and z b are representations of matched text-table pairs, with unrelated pairs used as negative examples. During online inference, table evidence is converted into compact textual form including row header, column header, value, and unit, allowing a single similarity calculation for both text chunks and flattened table evidence.
The framework incorporates intent-aware evidence routing. For each query q, an intent classifier predicts one of four semantic intents: calculation, trend, fact, or comparison, using pintent = softmax(W hq + b). A routing weight λ is calculated as λ = σ(g(hq, pintent)), controlling the relative importance of text versus table evidence. For each candidate evidence unit v, the retrieval score is adjusted: S(q, v) = λs(q, v) if v is a text unit, and S(q, v) = (1 − λ)s(q, v) if v is a table unit. Text and table candidates are then merged into one evidence list, and top evidence units are selected for answer generation.
The generation layer converts each selected evidence unit into a compact context item ci = (di, si, mi, ui), where di is the source filing, si is the section, mi is the modality type, and ui is the evidence content. The final context Cq = c1, c2,..., cB is formed with evidence budget B. Different prompting requirements are used for different intent types: calculation questions prompt for exact values and intermediate steps; trend questions prompt for trend direction, related periods, supporting numbers, and explanations; fact questions prompt for concise answers with source evidence; comparison questions prompt for aligning values across entities or fiscal periods. The final answer is generated as â = LLM(q, Cq).
The paper also introduces Multi-Doc-2025, a benchmark containing 2,327 expert-verified QA pairs from 179 SEC 10-K filings of 87 S&P 500 companies across 12 GICS sectors and fiscal years 2022–2024. The official split is primary-company-disjoint, with 1,600 training, 252 validation, and 475 test examples. The benchmark is organized into five subsets: single-document fact/calculation, single-document table reasoning, cross-year trend reasoning, cross-company comparison, and full-cross reasoning. Semantic intent labels include calculation, trend, fact, and comparison, while structural attributes include is cross doc, is cross year, and is hybrid modal. Evaluation metrics include EM, token-level F1, numerical execution accuracy, hallucination rate, and slice-wise metrics by intent, subset, and difficulty.
Experiments are conducted on FinQA, TAT-QA, DocFinQA, FinanceBench, and Multi-Doc-2025, comparing against baselines including BM25-RAG, DPR-RAG, Contriever-RAG, vanilla dense RAG, Self-RAG, RAPTOR, GraphRAG, TAPEX-RAG, and dataset-specific models. Results show HC-RAG achieves the strongest answer-level performance, with the largest gains in long-document and multi-document settings. On DocFinQA, HC-RAG improves over RAPTOR by 6.6 F1 points. On Multi-Doc-2025, it reaches 60.2 F1, outperforming GraphRAG by 10.9 F1 points and TAPEX-RAG by 6.1 F1 points. On FinQA, HC-RAG achieves 74.2 EM and 72.8 Exec Acc; on TAT-QA, 79.4 F1; on FinanceBench, 78.7 F1 with 11.3% hallucination rate.
Evidence localization analysis on Multi-Doc-2025 shows HC-RAG achieves the strongest section-level localization (Section Hit@5 of 64.21), fine-grained evidence recall (Evidence R@5 of 22.49, Evidence R@10 of 26.49), table retrieval (Table Hit@5 of 89.64), and cross-document recall (48.69), while GraphRAG obtains slightly higher Doc Hit@5 (50.10 vs. 45.26). The authors note this pattern supports the evidence-centric design: the main advantage is not merely finding a relevant filing, but navigating to the correct section, table, and cross-document evidence units.
Ablation studies on Multi-Doc-2025 show removing the three-level index causes the largest performance drop (F1 from 60.2 to 49.1, a drop of 11.1 points), demonstrating flat retrieval is insufficient for long financial filings. Removing L1 cross-document edges reduces Cross-Doc F1 from 58.7 to 40.9. Removing cross-modal alignment reduces Hybrid Modal F1 from 62.1 to 49.3. Removing TAPAS-based table structure and query-aware fusion also leads to lower results, with smaller drops.
Efficiency analysis shows HC-RAG's index construction cost is 4.8 min/doc, lower than GraphRAG's 7.6 min/doc and RAPTOR's 5.3 min/doc, while inference latency is 3.1 s/query, close to GraphRAG (3.4) and RAPTOR (2.9). Adaptive fusion analysis shows the routing weight λ changes with query semantics: a numerical current-ratio question receives λ of 0.28 and uses balance-sheet evidence, while a cloud-revenue trend question receives λ of 0.71 and relies on MD&A text. Cross-document diagnostics show HC-RAG reaches 48.7 F1 on 3+ document settings, exceeding GraphRAG by 10.1 points and RAPTOR by 13.5 points, and leads on YoY EM (57.2), cross-company EM (52.8), and industry-trend F1 (56.4).
The paper acknowledges limitations: HC-RAG relies on the relatively regular structure of SEC-style annual reports, may not work equally well on scanned reports, noisy PDF files, or irregular layouts; table extraction quality can affect results; intent classifier mistakes can lead to suboptimal text-table weighting; and Multi-Doc-2025 focuses mainly on SEC 10-K filings from S&P 500 companies. Future directions include more robust document parsing for visually complex PDFs, improved table understanding for complex headers and units, more flexible evidence graphs for updates, and application to other document-heavy domains such as legal, medical, and regulatory analysis.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement and the resulting capabilities of the improved AI system:
Improvement: Replace flat chunking with a typed financial evidence graph containing document, section, text-unit, table-unit, and metadata nodes, connected via containment, sequential, cross, and modal edges.
Resulting Capability: The system can navigate from company/year → filing → section → evidence unit, following the actual structure of financial reports rather than treating them as unordered text. This enables traceable retrieval paths and reduces lost in the middle
issues in long documents.
Improvement: Implement staged retrieval: first rank documents (combining semantic similarity with metadata matching for company/year/sector), then rank sections within selected documents (using section priors like calculation questions → financial statements
), then retrieve evidence units only from selected sections.
Improvement: Offline, use FinBERT for text and TAPAS/TAPEX for tables, aligning text-table pairs in a shared embedding space via contrastive learning. Online, flatten tables into compact strings (row header, column header, value, unit) for fast retrieval without running heavy table encoders per query.
Improvement: Train a classifier on FinBERT embeddings to predict query intent (4 classes), then compute a routing weight λ via a feed-forward layer that adjusts text vs. table evidence scores: text gets λ·s(q,v), tables get (1−λ)·s(q,v).
Improvement: Use different prompt templates per intent: calculation prompts require exact values and intermediate steps; trend prompts require direction, periods, supporting numbers; fact prompts require concise sourced answers; comparison prompts require aligned values across entities/periods.
Improvement: Add explicit edges connecting related companies, fiscal years, sectors, and financial metrics, enabling retrieval paths that span multiple filings.
-
Answer complex financial questions requiring evidence from multiple filings, fiscal years, and modalities — e.g., "Compare R&D intensity of Apple vs. Microsoft in FY2024
or
Explain the decline in Intel's gross margin." -
Provide verifiable, traceable answers — every answer maps to specific document → section → text/table unit paths, enabling auditors and analysts to check sources directly.
-
Handle long documents (100k+ tokens) without performance degradation — hierarchical retrieval prevents
lost in the middle
and maintains evidence precision. -
Adapt retrieval strategy per query type — automatically shifts between table-heavy (calculations) and text-heavy (trend explanations) evidence based on semantic intent.
-
Reduce hallucination in financial answers — by grounding generation in structured, intent-routed evidence with source paths.
-
Scale to large filing corpora — with 4.8 min/doc indexing and 3.1s/query latency, it processes 179 filings of 87 companies efficiently.
-
Evaluate both answer quality and evidence localization — providing metrics for document hit, section hit, table hit, and cross-document recall, not just final answer correctness.
-
Generalize to unseen companies — the primary-company-disjoint split ensures the system works on new entities, not just memorized training data.
Sources
- Retrieval-Augmented Generation for Large Language Models: A Survey
- BloombergGPT: A Large Language Model for Finance
- FinanceBench: A New Benchmark for Financial Question Answering
- TAPEX: Table Pre-training via Learning a Neural SQL Executor
- RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models
- FinBERT: Financial Sentiment Analysis with Pre-trained Language Models
- CFGPT: Chinese Financial Assistant with Large Language Model
- FinGPT: Instruction Tuning Benchmark for Open-Source Large Language Models in Financial Datasets
- Instruct-FinGPT: Financial Sentiment Analysis by Instruction Tuning of General-Purpose Large Language Models
- DocFinQA: A Long-Context Financial Reasoning Dataset
- DocVQA: A Dataset for VQA on Document Images
- Chain-of-Table: Evolving Tables in the Reasoning Chain for Table Understanding
- A Comprehensive Survey of Retrieval-Augmented Generation (RAG): Evolution, Current Landscape and Future Directions
- Representation Learning with Contrastive Predictive Coding
- Learning Transferable Visual Models From Natural Language Supervision
- Dense Passage Retrieval for Open-Domain Question Answering
- Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering