Proper Dataset Valuation by Pointwise Mutual Information
cs.LG, cs.GT
Submitted: 2024-05-28
Updated: 2026-09-05
Comments: Accepted at EC 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Data plays a central role in advancements in modern artificial intelligence, with high-quality data emerging as a key driver of model performance.
Terminology
Abstract
Data plays a central role in advancements in modern artificial intelligence, with high-quality data emerging as a key driver of model performance. This has prompted the development of principled and effective data curation methods in recent years. However, existing methods largely rely on heuristics, and whether they are truly effective remains unclear. For instance, standard evaluation methods that assess a trained model's performance on specific benchmarks may incentivize assigning high scores to data that merely resembles the test set. This issue exemplifies Goodhart's law: when a measure becomes a target, it ceases to be a good measure. To address this issue, we propose an information-theoretic framework for evaluating data curation methods. We define dataset quality in terms of its informativeness about the true model parameters, formalized using the Blackwell ordering of informativeness. Under this ordering, Blackwell's theorem ensures that more informative data yields optimal models with lower expected loss on the true underlying distribution. To measure informativeness, we show that the Blackwell order can be determined by the Shannon mutual information between the curated data and the test data. To estimate this mutual information, we introduce a novel method that trains Bayesian models on embedded datasets and computes mutual information from the posteriors of model parameters. Experiments on real-world data demonstrate that our mutual information-based evaluation assigns appropriately lower scores to data curation strategies that reduce dataset informativeness, while traditional test score-based evaluation methods may favor data curation strategies that overfit to the test set but compromise the training data's informativeness.
Sources
- A Survey on Data Selection for Language Models
- Invariant Risk Minimization
- MINE: Mutual Information Neural Estimation
- Neural Approximate Sufficient Statistics for Implicit Models
- Data Filtering Networks
- MINDE: Mutual Information Neural Diffusion Estimation
- Metadata Conditioning Accelerates Language Model Pre-training
- DU-Shapley: A Shapley Value Proxy for Efficient Dataset Valuation
- The Llama 3 Herd of Models
- Textbooks Are All You Need
- Benchmarking Neural Network Robustness to Common Corruptions and Perturbations
- Efficient Task-Specific Data Valuation for Nearest Neighbor Algorithms
- OpenDataVal: a Unified Benchmark for Data Valuation
- LAVA: Data Valuation without Pre-Specified Learning Algorithms
- Beta Shapley: a Unified and Noise-reduced Data Valuation Framework for Machine Learning
- DataInf: Efficiently Estimating Data Influence in LoRA-tuned LLMs and Diffusion Models
- DataComp-LM: In search of the next generation of training sets for language models
- Textbooks Are All You Need II: phi-1.5 technical report
- DEMI: Discriminative Estimator of Mutual Information
- Flow Matching for Generative Modeling
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks