HiFi-KPI: A Dataset for Hierarchical KPI Extraction from Earnings Filings
summary
The gist
The paper details a comprehensive approach for hierarchical KPI extraction from earnings filings, focusing heavily on standardized data representation and rigorous evaluation metrics suitable for
In short
The episode discusses HiFi-KPI, a dataset for hierarchical KPI extraction from earnings filings. Hosts discuss how this structured data improves AI training by providing a standardized blueprint for financial data. They cover the paper's focus on capturing metric relationships, the proposed methods for flexible extraction with user control over detail and currency units, and the long-term impact on building more robust AI systems in finance.
Key concepts
- Hierarchical KPI Extraction
- This refers to a method for extracting Key Performance Indicators (KPIs) from earnings filings while preserving the relationships between different metrics. The paper focuses on creating a structure that captures how metrics fit into broader financial reporting narratives, moving beyond simple key-value pairs.
- Standardized Data Representation
- The research emphasizes creating standardized structures for KPIs. This simplifies the complex task of turning messy text from earnings filings into clean, usable data points by precisely defining what each piece of information should look like for AI models.
- Granularity and Unit Management
- The proposed improvements involve giving users control over the level of detail they want during extraction and building logic to manage different currency units (like USD, EUR, CAD). This flexibility helps AI systems adapt to various needs without being locked into a single rigid structure.
- Structured Knowledge Base for AI
- The ultimate goal is to establish a standardized corpus that acts as a gold standard for complex reasoning tasks in finance. This structured knowledge base allows AI to move beyond just finding numbers to understanding the actual hierarchical relationships between those numbers.
Terminology used across episodes
This episode discusses
- HiFi-KPI: A Dataset for Hierarchical KPI Extraction from Earnings Filings · Paper Radio
- Hierarchical Text Classification Using Contrastive Learning Informed Path Guided Hierarchy
- ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question Answering
- SEC-QA: A Systematic Evaluation Corpus for Financial QA
- EDGAR-CORPUS: Billions of Tokens Make The World Go Round
- BloombergGPT: A Large Language Model for Finance
The paper
HiFi-KPI: A Dataset for Hierarchical KPI Extraction from Earnings Filings · Read on arXiv
Jie Zhou, Chunping Ma, Dingkun Long, Guangwei Xu, Ning Ding, Haoyu Zhang, Pengjun Xie, Gongshen Liu
Association for Computational Linguistics
Accurate tagging of earnings reports can yield significant short-term returns for stakeholders. The machine-readable inline eXtensible Business Reporting Language (iXBRL) is mandated for public financial filings. Yet, its complex, fine-grained taxonomy limits the cross-company transferability of tagged Key Performance Indicators (KPIs). To address this, we introduce the Hierarchical Financial Key Performance Indicator (HiFi-KPI) dataset, a large-scale corpus of 1.65M paragraphs and 198k unique, hierarchically organized labels linked to iXBRL taxonomies. HiFi-KPI supports multiple tasks and we evaluate three: KPI classification, KPI extraction, and structured KPI extraction. For rapid evaluation, we also release HiFi-KPI-Lite, a manually curated 8K paragraph subset. Baselines on HiFi-KPI-Lite show that encoder-based models achieve over 0.906 macro-F1 on classification, while Large Language Models (LLMs) reach 0.440 F1 on structured extraction. Finally, a qualitative analysis reveals that extraction errors primarily relate to dates. We open-source all code and data at https://github.com/aaunlp/HiFi-KPI.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "HiFi-KPI: A Dataset for Hierarchical KPI Extraction from Earnings Filings".
Jane: The paper details a comprehensive approach for hierarchical KPI extraction from earnings filings, focusing heavily on standardized data representation and rigorous evaluation metrics suitable for generative language models.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Now that we’ve established the dataset and the setup, let’s go over exactly what the HiFi-KPI paper actually delivers in terms of its findings and what they found when they tested their approach. Basically, this research provides a massive library of financial data designed specifically to train AI to read earnings filings with extreme precision.
Jane: That’s right; the paper lays out how they gathered this dataset, focusing on creating standardized structures for KPIs so that AI models can learn the hierarchy inherent in financial reporting. It simplifies the complex task of turning messy text into clean, usable data points by defining exactly what each piece of information should look like.
Lu: What's really striking is their focus on creating a structure that isn't just flat key-value pairs; it’s designed to capture the relationships between different metrics, which opens up some wild possibilities for how AI can actually reason about company performance over time.
Meng: From an engineering standpoint, this structured data means we can build much more reliable downstream applications because the input data itself is validated against a very high standard of accuracy, reducing the kind of garbage-in/garbage-out problems we always see.
Lalam: The implication for AI culture is significant; when we provide models with this level of structured truth, it forces the development community to focus on building systems that understand hierarchical context rather than just pattern matching superficial keywords.
Tom: It sounds like the core takeaway is that this dataset and its proposed methodology don't just offer a bigger training set; they offer a blueprint for teaching AI to reason about financial data with proper structure from the start.
Jane: Exactly; it shifts the focus from just extracting numbers to understanding how those numbers fit into the broader narrative of a company's financial health, which is exactly what we need in complex analysis.
Lu: I think this moves us closer to systems where AI can handle nuanced queries that require traversing different levels of detail within a massive filing structure, something that was previously very difficult to achieve consistently.
Meng: If we can reliably extract these hierarchical relationships, it means forecasting tools won't be as easily misled by poorly formatted or ambiguous data in the real world.
Lalam: The vision here is that this work helps establish a new cultural expectation for how AI should be trained in specialized domains; it shows the value of structured data as a foundation for complex, reliable reasoning.
Tom: So, we've seen how they built the structure and why it matters for reliability in financial analysis. Now, we need to look at how these researchers plan to actually refine the models that use this information moving forward.
The paper's summary: Tom: We've covered the dataset itself, but now we need to talk about how these researchers are actually planning to improve the extraction process using HiFi-KPI; they’re not just handing us data, they’re giving us a roadmap for making AI smarter in this area.
Jane: They propose introducing two different sets of cleaned and unified taxonomies, which lets users choose how detailed they want to get with the extraction, directly addressing that issue of overly specific labels in the original iXBRL format.
Lu: That flexibility is really exciting because it means AI systems won't be locked into one rigid structure; they can adapt their extraction strategy based on whether they need broad categories or deep, granular detail for a specific task.
Meng: The paper also suggests setting up benchmarks for different types of AI, including text classification and sequence labeling, using HiFi-KPI-Lite to see how well various methods actually perform on this structured data.
Tom: And I'm really curious about the practical side—how do they suggest we handle real-world complexities like different currency units when extracting data, since page two showed examples in USD, EUR, and CAD?
Jane: They noted that mixing those currencies is a huge error because their values are different, so their system has to build in logic to manage the unit conversions correctly during the extraction phase.
Lu: The method for selecting granularity they mentioned is clever; it gives users a tool to navigate the complexity of the iXBRL structure without having to manually map every single tag themselves, which is a big win for usability.
Meng: If we can actually make AI systems robust enough to handle those currency ambiguities and choose the right level of detail automatically, then we can deploy these extraction tools into real financial pipelines with much higher confidence.
Lalam: For me, this focus on flexible structure and unit management is crucial because it shows a path toward building AI that isn't just brittle; it’s building AI that can handle the messy reality of global financial reporting.
Tom: So, they are focusing on making the extraction process itself more adaptable and less prone to structural errors by giving users control over granularity and unit handling. Where do we go from here with this structured knowledge?
The paper's improvements: Tom: So we’ve covered everything on HiFi-KPI, from the authors to the structural improvements they're suggesting for better AI extraction, and now we're wrapping up with what all this actually means for us.
Jane: It boils down to this: having a standardized, large-scale corpus like HiFi-KPI allows us to build much more robust AI systems for financial analysis because the data quality is significantly higher than what was previously available.
Lu: I think the biggest implication lies in how these structured extractions feed into larger systems; it moves AI from just finding numbers to understanding the actual hierarchical relationship between those numbers within a company's reporting structure.
Meng: Practically, if we can get this level of reliable extraction, it means downstream tools that rely on these metrics for forecasting or risk assessment will be much more trustworthy because the input data is less prone to structural errors.
Lalam: For me, the most impactful vision here is how this structured data capability can improve the culture of AI development by providing a gold standard for complex reasoning tasks in regulated industries like finance.
Tom: What a way to put it; we've seen how they tackled the complexity of iXBRL and proposed methods to make extraction more flexible, and we're ready to see what comes next, but for now, this HiFi-KPI dataset seems like a solid foundation.
Jane: I agree; this work sets a high bar for data quality in complex domains, and it’s clear that the effort put into standardizing this structure will have long-term benefits for every AI application in finance.
Lu: It opens up avenues for creating truly sophisticated AI reasoners that can operate across different levels of financial detail without getting lost or confused by the input data's inconsistencies.
Meng: We really need to see how this structured output translates into actual deployable systems; if we can move beyond just having a dataset to actually building tools on top of it, that’s where the real engineering value is.
Lalam: The vision here is that as AI becomes more prevalent in professional fields, this kind of standardized knowledge base will become essential for ensuring that AI reasoning remains accurate and reliable when dealing with critical information like financial statements.
Conclusion: Tom: So we’ve covered everything on HiFi-KPI: A Dataset for Hierarchical KPI Extraction from Earnings Filings, from the authors to the structural improvements they're suggesting for better AI extraction, and now we're wrapping up with what all this actually means for us.
Jane: It boils down to this: having a standardized, large-scale corpus like HiFi-KPI allows us to build much more robust AI systems for financial analysis because the data quality is significantly higher than what was previously available.
Lu: I think the biggest implication lies in how these structured extractions feed into larger systems; it moves AI from just finding numbers to understanding the actual hierarchical relationship between those numbers within a company's reporting structure.
Meng: Practically, if we can get this level of reliable extraction, it means downstream tools that rely on these metrics for forecasting or risk assessment will be much more trustworthy because the input data is less prone to structural errors.
Lalam: For me, the most impactful vision here is how this structured data capability can improve the culture of AI development by providing a gold standard for complex reasoning tasks in regulated industries like finance.
Tom: What a way to put it; we've seen how they tackled the complexity of iXBRL and proposed methods to make extraction more flexible. We're ready to see what comes next, but for now, this HiFi-KPI dataset seems like a solid foundation.
Jane: I agree; this work sets a high bar for data quality in complex domains, and it’s clear that the effort put into standardizing this structure will have long-term benefits for every AI application in finance.
Lu: It opens up avenues for creating truly sophisticated AI reasoners that can operate across different levels of financial detail without getting lost or confused by the input data's inconsistencies.
Meng: We really need to see how this structured output translates into actual deployable systems; if we can move beyond just having a dataset to actually building tools on top of it, that’s where the real engineering value is.
Lalam: The vision here is that as AI becomes more prevalent in professional fields, this kind of standardized knowledge base will become essential for ensuring that AI reasoning remains accurate and reliable when dealing with critical information like financial statements.
Tom: Fantastic summary; it really shows how they’ve laid the groundwork for next-generation financial AI. We're definitely feeling very optimistic about what comes next in this space.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization