DataTales: A Benchmark for Real-World Intelligent Data Narration
cs.AI
Submitted: 2024-10-23
Updated: 2025-08-23
Comments: Accepted at EMNLP 2024 (main conference, long paper)
Journal ref: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 10764-10788, Miami, Florida, USA. Association for Computational Linguistics
DOI: 10.18653/v1/2024.emnlp-main.601
Code: https://github.com/yajingyang/DataTales
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: We introduce DataTales, a novel benchmark designed to assess the proficiency of language models in data narration, a task crucial for transforming complex tabular data into accessible narratives.
Terminology
Abstract
We introduce DataTales, a novel benchmark designed to assess the proficiency of language models in data narration, a task crucial for transforming complex tabular data into accessible narratives. Existing benchmarks often fall short in capturing the requisite analytical complexity for practical applications. DataTales addresses this gap by offering 4.9k financial reports paired with corresponding market data, showcasing the demand for models to create clear narratives and analyze large datasets while understanding specialized terminology in the field. Our findings highlights the significant challenge that language models face in achieving the necessary precision and analytical depth for proficient data narration, suggesting promising avenues for future model development and evaluation methodologies.
Sources
- FinQA: A Dataset of Numerical Reasoning over Financial Data
- ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question Answering
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Foresight: Recommending Visual Insights
- Table-GPT: Table-tuned GPT for Diverse Table Tasks
- LLaMA: Open and Efficient Foundation Language Models
- ToTTo: A Controlled Table-To-Text Generation Dataset
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- LLMs May Perform MCQA by Selecting the Least Incorrect Option
- Challenges in Data-to-Document Generation
- OpenAgents: An Open Platform for Language Agents in the Wild
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection