HiFi-KPI: A Dataset for Hierarchical KPI Extraction from Earnings Filings
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "HiFi-KPI: A Dataset for Hierarchical KPI Extraction from Earnings Filings".
Jane: The paper details a comprehensive approach for hierarchical KPI extraction from earnings filings, focusing heavily on standardized data representation and rigorous evaluation metrics suitable for generative language models.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Now that we’ve established the dataset and the setup, let’s go over exactly what the HiFi-KPI paper actually delivers in terms of its findings and what they found when they tested their approach. Basically, this research provides a massive library of financial data designed specifically to train AI to read earnings filings with extreme precision.
Jane: That’s right; the paper lays out how they gathered this dataset, focusing on creating standardized structures for KPIs so that AI models can learn the hierarchy inherent in financial reporting. It simplifies the complex task of turning messy text into clean, usable data points by defining exactly what each piece of information should look like.
Lu: What's really striking is their focus on creating a structure that isn't just flat key-value pairs; it’s designed to capture the relationships between different metrics, which opens up some wild possibilities for how AI can actually reason about company performance over time.
Meng: From an engineering standpoint, this structured data means we can build much more reliable downstream applications because the input data itself is validated against a very high standard of accuracy, reducing the kind of garbage-in/garbage-out problems we always see.
Lalam: The implication for AI culture is significant; when we provide models with this level of structured truth, it forces the development community to focus on building systems that understand hierarchical context rather than just pattern matching superficial keywords.
Tom: It sounds like the core takeaway is that this dataset and its proposed methodology don't just offer a bigger training set; they offer a blueprint for teaching AI to reason about financial data with proper structure from the start.
Jane: Exactly; it shifts the focus from just extracting numbers to understanding how those numbers fit into the broader narrative of a company's financial health, which is exactly what we need in complex analysis.
Lu: I think this moves us closer to systems where AI can handle nuanced queries that require traversing different levels of detail within a massive filing structure, something that was previously very difficult to achieve consistently.
Meng: If we can reliably extract these hierarchical relationships, it means forecasting tools won't be as easily misled by poorly formatted or ambiguous data in the real world.
Lalam: The vision here is that this work helps establish a new cultural expectation for how AI should be trained in specialized domains; it shows the value of structured data as a foundation for complex, reliable reasoning.
Tom: So, we've seen how they built the structure and why it matters for reliability in financial analysis. Now, we need to look at how these researchers plan to actually refine the models that use this information moving forward.
The paper's summary: Tom: We've covered the dataset itself, but now we need to talk about how these researchers are actually planning to improve the extraction process using HiFi-KPI; they’re not just handing us data, they’re giving us a roadmap for making AI smarter in this area.
Jane: They propose introducing two different sets of cleaned and unified taxonomies, which lets users choose how detailed they want to get with the extraction, directly addressing that issue of overly specific labels in the original iXBRL format.
Lu: That flexibility is really exciting because it means AI systems won't be locked into one rigid structure; they can adapt their extraction strategy based on whether they need broad categories or deep, granular detail for a specific task.
Meng: The paper also suggests setting up benchmarks for different types of AI, including text classification and sequence labeling, using HiFi-KPI-Lite to see how well various methods actually perform on this structured data.
Tom: And I'm really curious about the practical side—how do they suggest we handle real-world complexities like different currency units when extracting data, since page two showed examples in USD, EUR, and CAD?
Jane: They noted that mixing those currencies is a huge error because their values are different, so their system has to build in logic to manage the unit conversions correctly during the extraction phase.
Lu: The method for selecting granularity they mentioned is clever; it gives users a tool to navigate the complexity of the iXBRL structure without having to manually map every single tag themselves, which is a big win for usability.
Meng: If we can actually make AI systems robust enough to handle those currency ambiguities and choose the right level of detail automatically, then we can deploy these extraction tools into real financial pipelines with much higher confidence.
Lalam: For me, this focus on flexible structure and unit management is crucial because it shows a path toward building AI that isn't just brittle; it’s building AI that can handle the messy reality of global financial reporting.
Tom: So, they are focusing on making the extraction process itself more adaptable and less prone to structural errors by giving users control over granularity and unit handling. Where do we go from here with this structured knowledge?
The paper's improvements: Tom: So we’ve covered everything on HiFi-KPI, from the authors to the structural improvements they're suggesting for better AI extraction, and now we're wrapping up with what all this actually means for us.
Jane: It boils down to this: having a standardized, large-scale corpus like HiFi-KPI allows us to build much more robust AI systems for financial analysis because the data quality is significantly higher than what was previously available.
Lu: I think the biggest implication lies in how these structured extractions feed into larger systems; it moves AI from just finding numbers to understanding the actual hierarchical relationship between those numbers within a company's reporting structure.
Meng: Practically, if we can get this level of reliable extraction, it means downstream tools that rely on these metrics for forecasting or risk assessment will be much more trustworthy because the input data is less prone to structural errors.
Lalam: For me, the most impactful vision here is how this structured data capability can improve the culture of AI development by providing a gold standard for complex reasoning tasks in regulated industries like finance.
Tom: What a way to put it; we've seen how they tackled the complexity of iXBRL and proposed methods to make extraction more flexible, and we're ready to see what comes next, but for now, this HiFi-KPI dataset seems like a solid foundation.
Jane: I agree; this work sets a high bar for data quality in complex domains, and it’s clear that the effort put into standardizing this structure will have long-term benefits for every AI application in finance.
Lu: It opens up avenues for creating truly sophisticated AI reasoners that can operate across different levels of financial detail without getting lost or confused by the input data's inconsistencies.
Meng: We really need to see how this structured output translates into actual deployable systems; if we can move beyond just having a dataset to actually building tools on top of it, that’s where the real engineering value is.
Lalam: The vision here is that as AI becomes more prevalent in professional fields, this kind of standardized knowledge base will become essential for ensuring that AI reasoning remains accurate and reliable when dealing with critical information like financial statements.
Conclusion: Tom: So we’ve covered everything on HiFi-KPI: A Dataset for Hierarchical KPI Extraction from Earnings Filings, from the authors to the structural improvements they're suggesting for better AI extraction, and now we're wrapping up with what all this actually means for us.
Jane: It boils down to this: having a standardized, large-scale corpus like HiFi-KPI allows us to build much more robust AI systems for financial analysis because the data quality is significantly higher than what was previously available.
Lu: I think the biggest implication lies in how these structured extractions feed into larger systems; it moves AI from just finding numbers to understanding the actual hierarchical relationship between those numbers within a company's reporting structure.
Meng: Practically, if we can get this level of reliable extraction, it means downstream tools that rely on these metrics for forecasting or risk assessment will be much more trustworthy because the input data is less prone to structural errors.
Lalam: For me, the most impactful vision here is how this structured data capability can improve the culture of AI development by providing a gold standard for complex reasoning tasks in regulated industries like finance.
Tom: What a way to put it; we've seen how they tackled the complexity of iXBRL and proposed methods to make extraction more flexible. We're ready to see what comes next, but for now, this HiFi-KPI dataset seems like a solid foundation.
Jane: I agree; this work sets a high bar for data quality in complex domains, and it’s clear that the effort put into standardizing this structure will have long-term benefits for every AI application in finance.
Lu: It opens up avenues for creating truly sophisticated AI reasoners that can operate across different levels of financial detail without getting lost or confused by the input data's inconsistencies.
Meng: We really need to see how this structured output translates into actual deployable systems; if we can move beyond just having a dataset to actually building tools on top of it, that’s where the real engineering value is.
Lalam: The vision here is that as AI becomes more prevalent in professional fields, this kind of standardized knowledge base will become essential for ensuring that AI reasoning remains accurate and reliable when dealing with critical information like financial statements.
Tom: Fantastic summary; it really shows how they’ve laid the groundwork for next-generation financial AI. We're definitely feeling very optimistic about what comes next in this space.
Jie Zhou, Chunping Ma, Dingkun Long, Guangwei Xu, Ning Ding, Haoyu Zhang, Pengjun Xie, Gongshen Liu
Association for Computational Linguistics
cs.CL, cs.AI
Submitted: 2026-08-20
Updated: 2026-08-21
Code: https://github.com/aaunlp/HiFi-KPI
Importance score: 88/100
The gist: The paper details a comprehensive approach for hierarchical KPI extraction from earnings filings, focusing heavily on standardized data representation and rigorous evaluation metrics suitable for
Key concepts
- Hierarchical KPI Extraction
- This refers to a method for extracting Key Performance Indicators (KPIs) from earnings filings while preserving the relationships between different metrics. The paper focuses on creating a structure that captures how metrics fit into broader financial reporting narratives, moving beyond simple key-value pairs.
- Standardized Data Representation
- The research emphasizes creating standardized structures for KPIs. This simplifies the complex task of turning messy text from earnings filings into clean, usable data points by precisely defining what each piece of information should look like for AI models.
- Granularity and Unit Management
- The proposed improvements involve giving users control over the level of detail they want during extraction and building logic to manage different currency units (like USD, EUR, CAD). This flexibility helps AI systems adapt to various needs without being locked into a single rigid structure.
- Structured Knowledge Base for AI
- The ultimate goal is to establish a standardized corpus that acts as a gold standard for complex reasoning tasks in finance. This structured knowledge base allows AI to move beyond just finding numbers to understanding the actual hierarchical relationships between those numbers.
Terminology
Summary
The paper details a comprehensive approach for hierarchical KPI extraction from earnings filings, focusing heavily on standardized data representation and rigorous evaluation metrics suitable for generative language models.
The core task involves acting as an expert data extraction assistant
to read text and extract financial or entity-related information. For each entity found, the system must extract five specific fields:
-
value: The numerical representation of the data point. -
currency/unit: The currency or unit (e.g., USD, shares, EUR). -
label: The financial label (e.g., revenues, earnings, eps, ebit), orXBRL−OOSif it does not fit the other categories. -
start date for period: The beginning date of the period (if available). -
end date for period: The ending date of the period (if available).
The output must be strictly valid JSON, containing an array under the "entities" key.
For evaluating performance on datasets like HiFi-KPI Lite, the paper defines how standard classification metrics are adapted for generative LLM predictions because these predictions are unrestricted.
-
True Positive (TP): A correct prediction.
-
False Negative (FN): A misclassified prediction relative to the true label.
-
False Positive (FP): A misclassified prediction relative to the predicted label.
The calculation of metrics is defined as follows:
-
Micro F1: Computed
as in any standard classification task.
-
Macro F1: Calculated by taking
the average F1 score of only the ground truth labels, excluding labels that appear solely in the predicted set.
Furthermore, for visualization (specifically Figure 6), a cumulative sum is computed by iterating over the label distribution from the most frequent to the least frequent label in the test set,
allowing for a macro-average F1 score calculation for the top x included labels.
The scope of extractable information is highly formalized, utilizing a complex taxonomy of XBRL labels. The paper provides an extensive list of specialized financial labels, including:
-
us-gaap:NetIncomeLoss -
us-gaap:OperatingIncomeLoss -
Various detailed income loss components (e.g.,
us-gaap:IncomeLossFromContinuingOperationsBeforeIncomeTaxesMinorityInterestAndIncomeLossFromEquityMethodInvestments). -
Specific metrics related to shares and dividends (e.g.,
us-gaap:EarningsPerShareBasic,us-gaap:DividendsAndInterestPaid). -
Revenue categories (e.g.,
us-gaap:Revenues,us-gaap:UnregulatedOperatingRevenue).
These XBRL labels are mapped to a simplified set of expert labels
for consistent extraction, including broad categories like Earnings,
EBIT,
and Revenues.
Improvements for AI systems
Improvement: The current system prompt relies heavily on the LLM's inherent ability to follow complex instructions and generate perfect JSON. To mitigate hallucination and structural failure, the extraction pipeline must be upgraded from a pure text-to-JSON approach to one utilizing Pydantic schema enforcement or a similar structured output library (e.g., integrating a constrained decoding mechanism like those found in advanced model APIs).
What the Improved AI System Can Do:
-
Guaranteed Structural Integrity: The system will be incapable of producing malformed JSON, even when faced with ambiguous source text. It will strictly adhere to the defined schema (
label,start date for period, etc.). -
Automated Type Casting and Validation: It will automatically validate that the
valuefield is a numeric type (float/integer) and that date fields conform to strict ISO 8601 standards, minimizing data preprocessing errors often associated with natural language output. -
Enhanced Contextual Linking: By enforcing the schema at the architectural level, we can build post-processing layers that verify if the extracted dates (
start date for periodandend date for period) logically precede or overlap correctly, flagging temporal inconsistencies (e.g., a start date after an end date).
Sources
- Hierarchical Text Classification Using Contrastive Learning Informed Path Guided Hierarchy
- ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question Answering
- SEC-QA: A Systematic Evaluation Corpus for Financial QA
- EDGAR-CORPUS: Billions of Tokens Make The World Go Round
- BloombergGPT: A Large Language Model for Finance
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering