Do Large Language Models Capture the Diversity in their Training Data?
cs.CL, cs.AI, cs.LG
Submitted: 2026-09-02
Updated: 2026-09-02
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models are trained to model conditional distributions over text, yet it remains inadequately understood whether they capture the full diversity of plausible outputs present in their
Terminology
Abstract
Large language models are trained to model conditional distributions over text, yet it remains inadequately understood whether they capture the full diversity of plausible outputs present in their training data. We study this question through an information-theoretic lens by comparing the conditional entropy of model-generated outputs with that of the corresponding training data. Given paired input-output samples, we use conditional entropy and its matrix-based analogue based on von Neumann entropy to measure output variability beyond what is explained by the conditioning input, without requiring multiple reference outputs for the same prompt. Across LLM families with publicly available training data, including OLMo, Pythia, and GPT-Neo, we consistently find that model-generated outputs exhibit lower conditional entropy than their training data, across different model scales, sequence lengths, and decoding strategies. We observe a similar conditional diversity gap beyond language modeling, including class-conditioned ImageNet generators and text-conditioned models trained on MS-COCO. To address this gap, we propose a post-hoc correction mechanism that generates multiple outputs for each input and reweights them through a matrix-entropy projection, increasing conditional diversity while remaining close to the original model distribution. We prove the concavity of the matrix-based conditional entropy functional, which makes the resulting entropy-constrained projection a convex optimization problem, and develop a scalable mirror-descent algorithm for its implementation. Our results reveal a systematic conditional diversity gap between modern generative models and their training data, and provide an information-theoretic framework for measuring and mitigating this gap.
Sources
- Evaluating Text-to-Image and Text-to-Video Synthesis with a Conditional Fr'{e}chet Distance
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
- Towards A Better Metric for Text-to-Video Generation
- FETV: A Benchmark for Fine-Grained Evaluation of Open-Domain Text-to-Video Generation
- SPARKE: Scalable Prompt-Aware Diversity and Novelty Guidance in Diffusion Models via RKE Score
- Image Generation Diversity Issues and How to Tame Them
- The Price of Format: Diversity Collapse in LLMs
- Extracting Training Data from Diffusion Models
- Regularizing Neural Networks by Penalizing Confident Output Distributions
- The Impact of the Single-Label Assumption in Image Recognition Benchmarking
- SAR: Self-Supervised Anti-Distortion Representation for End-To-End Speech Model
- WildVis: Open Source Visualizer for Million-Scale Chat Logs in the Wild
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering