HC-RAG: Evidence-Centric Retrieval-Augmented Generation over Heterogeneous Financial Filings

summary

Video file (mp4)

The gist

vD → vS → vU, where vD is a document node, vS is a section node, and vU is a text or table evidence-unit node.

In short

The episode discusses 'HC-RAG,' a system for answering questions over complex financial filings. The hosts detail how HC-RAG uses a graph structure to preserve document hierarchy, aligns text and tables, and routes queries by intent. It achieves high accuracy and efficiency on specialized benchmarks.

Key concepts

Heterogeneous Financial Filings
These are large financial documents (like 10-K reports) containing different data types, including text paragraphs, structured tables, and company metadata. HC-RAG is designed to handle this mix of data simultaneously.
Evidence-Centric Retrieval
Instead of searching for similar text chunks, this method focuses on finding verifiable evidence by following the document's structure. It narrows the search step-by-step (document $\rightarrow$ section $\rightarrow$ unit) to ensure accuracy.
Graph Structure
The system models the filing as a map where documents connect to sections, and sections connect to specific text or table units. This preserves the natural hierarchy of the financial report.
Intent Routing
The system classifies a user's question into one of four intents (calculation, trend, fact, comparison). This classification determines whether the retrieval process should prioritize tables or narrative text.

Terminology used across episodes

This episode discusses

The paper

HC-RAG: Evidence-Centric Retrieval-Augmented Generation over Heterogeneous Financial Filings · Read on arXiv

Siyuan Chen, Huaye Tan, You Li, Jiajun Liang

Sun Yat-sen University · Central South University

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "HC-RAG: Evidence-Centric Retrieval-Augmented Generation over Heterogeneous Financial Filings".

Jane: The paper was written by Siyuan Chen, Huaye Tan, You Li and Jiajun Liang from Sun Yat-sen University and Central South University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Authors: Tom: Hey everyone, welcome back to the show. Today we’re cracking open a fresh one from arXiv called “HC-RAG: Evidence-Centric Retrieval-Augmented Generation over Heterogeneous Financial Filings.” Jane, I gotta say, the title alone tells me these folks are trying to fix something real.

Jane: Absolutely, Tom. And the authors—Siyuan Chen, Huaye Tan, You Li, Jiajun Liang—they’re from Sun Yat-sen University and Central South University. They’re not just throwing another chatbot at us. They’re looking at how you actually answer questions about financial documents, like annual reports.

Tom: Right, and that’s a big deal because these reports are massive. We’re talking ten-K filings with tens of thousands of tokens. If you just chop them up into little chunks and search, you lose the structure. You lose the sections, the tables, the context.

Jane: Exactly. And that’s what the “heterogeneous” part means. You’ve got text, you’ve got tables, you’ve got metadata about companies and fiscal years. A regular search engine treats all that the same way, and it just doesn’t work. The paper’s whole point is that evidence in finance is structured, and your retrieval should be too.

Tom: So they built a graph. A financial evidence graph. Documents connect to sections, sections connect to text units and table units, and then there are edges linking companies and years. It’s like a map of the filing, not just a pile of paper.

Jane: And that map lets you retrieve the way an analyst would. First find the right company, then the right section, then the specific table or paragraph. It’s step-by-step, not a wild guess across the whole corpus.

Tom: I love that. It’s like going to a library, but instead of wandering the aisles, you know the exact floor, the exact shelf, and the exact book. The implications here are huge for anyone who does financial research for a living.

Jane: Or for anyone building AI tools for analysts. This could change how we build question-answering systems for finance, and maybe even for legal or medical documents that have similar structure. We’ll dig into how they actually did it next.

Summary of the Paper: Tom: So Jane, we’ve got the title and the authors down. Now let’s talk about what this paper actually does. And I want to bring in Lu, our senior researcher, because I think she’ll appreciate the architecture here.

Jane: Good idea. Lu, the paper’s core idea is that financial QA isn’t just about finding similar text. It’s about finding verifiable evidence. Can you break that down for us?

Lu: Sure, Jane. The authors argue that if you ask “what was Apple’s current ratio in two thousand twenty-four” you don’t just need any text that mentions Apple. You need the exact balance sheet, the exact row, the exact value. So they built a three-level retrieval system: first documents, then sections, then evidence units like tables or paragraphs.

Tom: And that’s the “hierarchical” part of HC-RAG. It narrows the search space step by step. Instead of comparing the question to every chunk in the corpus, it first picks the right filings, then the right sections, then the right evidence. That’s way more efficient and way more accurate.

Jane: But here’s the clever part, Lu. They also align text and tables in the same embedding space. So a table about revenue and a paragraph about revenue are mapped close together. That way, when you search, you can find both, and the system can decide which one matters more.

Lu: Exactly. And they don’t just use a fixed mix of text and table. They classify the question into four intents: calculation, trend, fact, or comparison. A calculation question gets more weight on tables. A trend question gets more weight on narrative text. It’s query-aware routing.

Tom: And they built a whole new benchmark to test this, called Multi-Doc-two thousand twenty-five. It’s got over two thousand three hundred questions from one hundred seventy-nine real ten-K filings of S andP five hundred companies. And it’s designed to test cross-company and cross-year questions, which most benchmarks ignore.

Lu: That’s the part I find most exciting. Most financial QA datasets are single-document. You read one report and answer. But real analysts compare companies. They track trends across years. This benchmark finally tests that.

Jane: And the results? Tom, you want to share the numbers?

Tom: Oh, absolutely. On their benchmark, HC-RAG hits sixty point two F1, which beats GraphRAG by almost eleven points and TAPEX-RAG by six points. And on DocFinQA, a long-document benchmark, they beat RAPTOR by six point six F1 points. Those are big jumps.

Lu: And the evidence localization results are even more telling. Their table hit rate is eighty-nine point six percent, compared to sixty-six point eight percent for the best baseline. That’s a massive improvement in finding the right table, which is exactly what financial QA needs.

Jane: So the summary is: they built a graph, they aligned text and tables, they routed by intent, and they proved it works. Next, let’s talk about what this means for real-world systems.

Improvements Suggested by the Paper: Tom: Alright, so we know what HC-RAG does. But what does it actually improve in practice? And Meng, I want you in on this because you’re the engineer who has to make things run.

Meng: Happy to be here, Tom. And honestly, the first thing that jumps out at me is efficiency. The paper reports index construction at four point eight minutes per document, which is faster than GraphRAG’s seven point six minutes. And inference is three point one seconds per query, which is totally usable in a real product.

Jane: That’s a big deal, Meng. Because a lot of these fancy graph systems are research prototypes. They work in a lab but they’re too slow or too expensive to deploy. HC-RAG seems to hit a sweet spot.

Meng: Exactly. And the asymmetric design is smart. They use a heavy table-aware encoder offline to align text and tables. But online, they flatten tables into short strings with row headers, column headers, and values. So you get the benefit of table understanding without the cost of running a big model on every query.

Lu: And that’s not just an engineering trick, Meng. It’s a design philosophy. They’re saying: do the expensive work once, during indexing, and keep the online retrieval light. That’s how you scale to thousands of filings.

Tom: The paper also shows robustness to noise. They added irrelevant evidence to the context, and HC-RAG degraded much slower than flat retrieval systems. Because the hierarchy acts as a filter—noise has to pass through document and section gates before it reaches the answer.

Meng: That’s huge for real-world use. In practice, you don’t always have clean queries. People ask vague questions. The system needs to not panic and pull in garbage. HC-RAG’s structure keeps it grounded.

Jane: And the intent routing, Lu, that’s another improvement. Instead of treating every question the same, the system adapts. For a calculation question, it pulls tables. For a trend question, it pulls MD andA text. The paper shows the routing weight actually shifts, from zero point two eight for a numerical question to zero point seven one for a trend question.

Lu: That’s the kind of adaptivity that makes a system feel intelligent. It’s not just retrieving more context. It’s retrieving the right kind of context for the task. And that’s a lesson that goes beyond finance.

Meng: Yeah, I could see this applied to legal research or medical guidelines. Any domain where documents have structure and where evidence type matters. The framework is domain-agnostic even though the benchmark is finance-specific.

Tom: So the improvements are: better accuracy, better evidence localization, better efficiency, and better robustness. That’s a full package. Now, let’s wrap up and think about what this means for the future.

Conclusion: Tom: Alright, we’re at the end of our time with “HC-RAG: Evidence-Centric Retrieval-Augmented Generation over Heterogeneous Financial Filings.” Jane, give us the final summary.

Jane: Sure, Tom. This paper tackles a real problem: answering questions over long, structured financial filings. It builds a graph that preserves the document-section-unit hierarchy, aligns text and tables in one space, and routes evidence based on query intent. And it proves the approach works with a new benchmark, Multi-Doc-two thousand twenty-five.

Lu: And the results are compelling. They beat strong baselines on both answer quality and evidence localization. The table hit rate of nearly ninety percent is the kind of number that makes you sit up and pay attention.

Meng: From an engineering standpoint, the efficiency numbers are just as important. You can index a filing in under five minutes and answer a query in about three seconds. That’s deployable. That’s a product.

Tom: And the implications go beyond finance. The idea of evidence-centric retrieval, of respecting document structure and adapting to query intent, that’s a blueprint for any domain where answers need to be verifiable.

Jane: We should also mention the limitations. The paper admits it relies on regular SEC-style filings. Scanned PDFs or irregular layouts would be harder. And the intent classifier can make mistakes, which would affect routing.

Lu: But those are next steps, not dead ends. The authors suggest better document parsing, better table understanding, and extending the graph to other domains. This is a solid foundation.

Tom: So we say goodbye to HC-RAG. It’s a strong contribution to financial AI and to retrieval-augmented generation in general. We’ll be watching for follow-up work.

Jane: Thanks for listening, everyone. We’ve got another paper lined up next, so stay tuned. Until then, keep asking good questions.

More episodes

← Home