Is Convergence Inevitable? Tracing Output Homogeneity Back to Base Models
Alexandrine Fortier, Hazel Chen, Peter West
University of British Columbia
cs.CL
Submitted: 2026-08-13
Updated: 2026-08-17
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: The paper investigates whether output homogeneity in language models originates during pretraining or is introduced by the alignment process.
Terminology
Summary
The paper investigates whether output homogeneity in language models originates during pretraining or is introduced by the alignment process. The authors argue that output homogeneity is likely learned during the pretraining phase, and only revealed or magnified during the alignment process.
The study first examines semantic similarity across alignment stages (SFT, DPO, RLVR) using the I NFINITY-C HAT 100 dataset and models from the Tülu 3 and OLMo 3 suites. They find that output convergence already strongly manifests after the first alignment stage, the instruction-tuning phase—SFT,
with distributions heavily right-shifted (mean pairwise cosine similarity above 0.7) and DPO/RLVR only modestly increasing convergence. This suggests convergence precedes preference learning entirely.
To distinguish whether SFT induces or reveals convergence, the authors conduct controlled SFT experiments using metaphor generation as a probe on Llama-3.1 8B. In a Structure Injection
experiment, they add 30 metaphor instruction-response pairs (without studied vehicles) to the LIMA dataset. They find that exposure to an answer format during fine-tuning is enough to reveal a preferred vehicle for most metaphors,
indicating the base model already converges on pretraining-learned patterns.
In an Idea Injection
experiment, they inject single metaphors with vehicles of varying prior frequency (1st, 2nd, 6th most frequent, and absent food-related vehicles). Results show the capacity for a vehicle to be amplified is highly dependent on its pre-injection convergence level,
with the most frequent vehicle reaching 48-92% frequency, while all food-related vehicles failed to be learned.
They conclude that the model is unable to learn new vehicles if these are not already present in the initial distribution,
meaning memorization from SFT is difficult if it does not align with the latent converging patterns.
Overall, convergence can be revealed and amplified, but not introduced by the SFT data.
To confirm latent homogeneity, the authors directly probe base models (Llama 3.1 8B, OLMo 3 7B, Qwen3 8B) using three prompting conditions: basic completion, few-shot, and assistant-style persona. They find that through prompting alone, convergence can be surfaced in base models,
with the assistant condition producing the most homogeneous outputs. Dominant vehicles emerge clearly (e.g., time is a river
or time is a thief
), and intra-vehicle cosine similarity is high (0.69-0.91). They note that vehicles that dominate in the base model do not always dominate after alignment,
suggesting base models may encode multiple latent preferences.
The paper concludes that semantic convergence may arise naturally from the objectives underlying LM training, making it difficult to mitigate through post-alignment interventions alone.
Limitations include using metaphor generation as a single probe task, conducting SFT experiments on a single base model, operating at sample-level injection rather than global data shifts, and not directly analyzing pretraining corpora to identify the precise mechanism.
Improvements for AI systems
Improvements to AI Systems:
-
Pretraining-Aware Alignment Regularization: Modify the alignment loss (e.g., SFT, DPO) to explicitly penalize over-convergence toward dominant latent patterns identified during pretraining. The improved system can detect when fine-tuning amplifies pre-existing output homogeneity and apply a diversity-preserving regularizer, reducing mode collapse without sacrificing instruction-following accuracy.
-
Latent Preference Probing Before Fine-Tuning: Add a pre-alignment diagnostic step that probes the base model (via assistant-style prompts) to map its latent output distributions for given tasks. The improved system can then flag tasks where the base model already exhibits strong convergence, allowing developers to either accept the bias or inject targeted counter-examples before SFT—since the paper shows post-hoc injection fails for novel patterns.
-
Data Curation for Novel Pattern Injection: Use the paper’s finding that SFT cannot introduce vehicles absent from the pretraining distribution. The improved system can automatically filter or augment fine-tuning datasets to include only patterns that are already latent in the base model, while also generating synthetic pretraining-like data (e.g., via masked language modeling on diverse corpora) to expand the latent space before alignment—enabling genuinely new behaviors to be learned.
-
Adaptive Persona Control for Homogeneity Mitigation: Leverage the result that assistant-style prompting surfaces convergence more strongly. The improved system can dynamically switch between prompt styles (e.g., neutral vs. assistant) during inference based on a user’s need for creativity vs. consistency, and can also apply temperature or top-p adjustments calibrated to the measured intra-vehicle cosine similarity to break out of dominant modes when requested.
-
Post-Alignment Diversity Auditing: Build a tool that computes pairwise semantic similarity across generated outputs (as in the paper) and compares it to the base model’s baseline convergence. The improved system can automatically alert developers when alignment has magnified homogeneity beyond acceptable thresholds, and suggest targeted fine-tuning on diverse exemplars that match the base model’s latent preferences—rather than attempting to override them.
-
Multi-Model Latent Preference Ensembling: Since base models encode multiple latent preferences (e.g., different dominant vehicles), the improved system can train an ensemble of aligned models, each fine-tuned to amplify a different latent preference from the same base. At inference, a router selects the model whose output distribution best matches the user’s stylistic or semantic requirements, effectively converting latent diversity into controllable output diversity.
-
Pretraining Objective Modification for Diversity: The paper suggests convergence arises from pretraining objectives. The improved system can alter the next-token prediction loss to include a diversity-promoting term (e.g., a regularizer that penalizes over-reliance on high-frequency semantic clusters) during pretraining. This would reduce the initial convergence, making later alignment less likely to produce homogeneous outputs—addressing the root cause rather than symptoms.
Abstract
The lack of diversity in LM content is widely attributed to the alignment process, but how and where exactly in the pipeline this collapse begins is unknown. We argue that output homogeneity is likely learned during the pretraining phase, and only revealed or magnified during the alignment process. Specifically, we find that semantic convergence is observed from the first alignment stage--the instruction-tuning phase (SFT)--suggesting that homogeneity might already exist in the pre-alignment model. To investigate this, we conduct controlled SFT experiments examining how training data influences output convergence on specific input/output pairs. We find that convergence can be revealed and amplified, but not introduced by the SFT data, supporting its role as a catalyst rather than a cause. To further test whether homogeneity originates before alignment, we measure convergence in base models. We find that instruct-like collapse can be induced through prompting alone, even without alignment. Taken together, our results suggest that semantic convergence may arise naturally from the objectives underlying LM training, making it difficult to mitigate through post-alignment interventions alone.
Sources
- Picking on the Same Person: Does Algorithmic Monoculture lead to Outcome Homogenization?
- Language Models are Few-Shot Learners
- Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
- You Didn't Have to Say It like That: Subliminal Learning from Faithful Paraphrases
- The Llama 3 Herd of Models
- Benchmarking Linguistic Diversity of Large Language Models
- Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models
- LoRA: Low-Rank Adaptation of Large Language Models
- Olmo 3
- Standardizing the Measurement of Text Diversity: A Tool and a Comparative Analysis of Scores
- Understanding the Effects of RLHF on LLM Generalisation and Diversity
- Evaluating the Diversity and Quality of LLM Generated Content
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- The Shrinking Landscape of Linguistic Diversity in the Age of Large Language Models
- Predicting vs. Acting: A Trade-off Between World Modeling & Agent Modeling
- The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning
- Evaluating the Evaluation of Diversity in Natural Language Generation
- Theoretical Proof that Auto-regressive Language Models Collapse when Real-world Data is a Finite Set
- We're Different, We're the Same: Creative Homogeneity Across LLMs
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering