KAHAN: Knowledge-Augmented Hierarchical Analysis and Narration for Financial Data Narration
cs.AI
Submitted: 2025-09-21
Updated: 2025-09-21
Comments: Accepted at EMNLP 2025 Findings
Journal ref: Findings of the Association for Computational Linguistics: EMNLP 2025, pages 25761-25785, Suzhou, China. Association for Computational Linguistics
DOI: 10.18653/v1/2025.findings-emnlp.1405
Code: https://github.com/yajingyang/kahan
License: http://creativecommons.org/licenses/by/4.0/
The gist: We propose KAHAN, a knowledge-augmented hierarchical framework that systematically extracts insights from raw tabular data at entity, pairwise, group, and system levels.
Terminology
Abstract
We propose KAHAN, a knowledge-augmented hierarchical framework that systematically extracts insights from raw tabular data at entity, pairwise, group, and system levels. KAHAN uniquely leverages LLMs as domain experts to drive the analysis. On DataTales financial reporting benchmark, KAHAN outperforms existing approaches by over 20% on narrative quality (GPT-4o), maintains 98.2% factuality, and demonstrates practical utility in human evaluation. Our results reveal that knowledge quality drives model performance through distillation, hierarchical analysis benefits vary with market complexity, and the framework transfers effectively to healthcare domains. The data and code are available at https://github.com/yajingyang/kahan.
Sources
- On the Opportunities and Risks of Foundation Models
- Can GPT models be Financial Analysts? An Evaluation of ChatGPT and GPT-4 on mock CFA Exams
- DataNarrative: Automated Data-Driven Storytelling with Visualizations and Texts
- Foresight: Recommending Visual Insights
- The Llama 3 Herd of Models
- DnA-Eval: Enhancing Large Language Model Evaluation through Decomposition and Aggregation
- Demonstration of InsightPilot: An LLM-Empowered Automated Data Exploration System
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
- WebGPT: Browser-assisted question-answering with human feedback
- GPT-4 Technical Report
- GPT-4o System Card
- Qwen2.5 Technical Report
- InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation
- Injecting Domain-Specific Knowledge into Large Language Models: A Comprehensive Survey
- Table Meets LLM: Can Large Language Models Understand Structured Table Data? A Benchmark and Empirical Study
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Challenges in Data-to-Document Generation
- Golden Touchstone: A Comprehensive Bilingual Benchmark for Evaluating Financial Large Language Models
- MDSF: Context-Aware Multi-Dimensional Data Storytelling Framework based on Large language Model
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection