Fairshare Data Pricing via Data Valuation for Large Language Models
cs.GT, cs.CL
Submitted: 2025-01-31
Updated: 2025-11-19
Journal ref: Advances in Neural Information Processing Systems 38 (NeurIPS 2025)
DOI: 10.52202/085713-1119
License: http://creativecommons.org/licenses/by/4.0/
The gist: Training data is the backbone of large language models (LLMs), yet today's data markets often operate under exploitative pricing -- sourcing data from marginalized groups with little pay or
Terminology
Abstract
Training data is the backbone of large language models (LLMs), yet today's data markets often operate under exploitative pricing -- sourcing data from marginalized groups with little pay or recognition. This paper introduces a theoretical framework for LLM data markets, modeling the strategic interactions between buyers (LLM builders) and sellers (human annotators). We begin with theoretical and empirical analysis showing how exploitative pricing drives high-quality sellers out of the market, degrading data quality and long-term model performance. Then we introduce fairshare, a pricing mechanism grounded in data valuation that quantifies each data's contribution. It aligns incentives by sustaining seller participation and optimizing utility for both buyers and sellers. Theoretically, we show that fairshare yields mutually optimal outcomes: maximizing long-term buyer utility and seller profit while sustaining market participation. Empirically when training open-source LLMs on complex NLP tasks, including math problems, medical diagnosis, and physical reasoning, fairshare boosts seller earnings and ensures a stable supply of high-quality data, while improving buyers' performance-per-dollar and long-term welfare. Our findings offer a concrete path toward fair, transparent, and economically sustainable data markets for LLM.
Sources
- Galactica: A Large Language Model for Science
- DataInf: Efficiently Estimating Data Influence in LoRA-tuned LLMs and Diffusion Models
- A Survey on Data Markets
- Unlearning Traces the Influential Training Data of Language Models
- On Influence Functions, Classification Influence, Relative Influence, Memorization and Generalization
- Data Shapley in One Training Run
- On the Feasibility of In-Context Probing for Data Attribution
- Fine-Tuning is Fine, if Calibrated
- Mathematical Language Models: A Survey
- A Survey of Large Language Models in Medicine: Progress, Application, and Challenge
Related papers
- Exact Regret Frontiers and Externality Scheduling in Centralized Serial-Dictatorship Bandits
- In-Context Credit Assignment via the Core
- Breaking 1/epsilon Barrier in Quantum Zero-Sum Games: Generalizing Metric Subregularity for Spectraplexes
- Enhancing Affine Maximizer Auctions with Correlation-Aware Payment
- LLM Bidders Preserve the Mechanism-Level Orderings of Human Bidders
- Towards Performatively Stable Equilibria in Decision-Dependent Games for Arbitrary Data Distribution Maps