Diffract: Spectral View of LLM Domain Adaptation

arXiv:2608.10850 · cs.LG · Submitted 2026-08-11 · Read on arXiv

Risk AI Research Lab · Applied AI Institute

cs.LG

Submitted: 2026-08-11

Updated: 2026-09-11

Comments: Accepted at ICML 2026. Code: https://github.com/Risk-AI-Research/diffract

Code: https://github.com/Risk-AI-Research/diffract

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 100/100

The gist: The paper "Diffract: Spectral View of LLM Domain Adaptation" studies continual pre-training (CPT) as a mechanism for adapting general-purpose large language models to specialized domains:

Terminology

Summary

The paper Diffract: Spectral View of LLM Domain Adaptation studies continual pre-training (CPT) as a mechanism for adapting general-purpose large language models to specialized domains: mathematics, instruction, code, and natural text. Using singular value decomposition of weight matrices, the authors find that CPT leaves singular value spectra largely invariant, with adaptation driven mainly by changes in singular vectors. An analysis of attention-head projection matrices reveals strong, domain-dependent head heterogeneity, which is exploited to define a head importance criterion: up to 60% of head updates can be removed without measurable quality loss. Selectively rewinding low-importance heads to their pre-trained state improves benchmark accuracy by up to 4% versus the fully trained baseline. Finally, the authors identify domain connectivity—linear interpolation between CPT checkpoints yields smooth domain-quality interpolation without notable degradation on either domain—and release Diffract, an open-source toolkit for scalable spectral analysis of billion-parameter models.

The study conducts pre-train and continual pre-training experiments on 1B, 7B, and 13B language models with OLMo 2 architecture for several domains: math, instruction, code, and text. The authors analyze the CPT delta ∆W = W domain − W pre-train, characterize its sparsity, and assess the ability to interpolate between checkpoints adapted to different domains. They investigate the singular spectra of weight matrices for checkpoints along the pre-train trajectory and for weight increments after CPT, uncovering several novel phenomena.

Key contributions include:

  1. Investigation of the dynamics of weight matrix singular value spectra, identifying the development of complex spectral structure in attention head matrices along the pre-train stage, which cannot be described by heavy-tailed self-regularization theory. This is associated with an increase in quality on language tasks and faster domain adaptation on CPT. The spectra during CPT remain almost stable, and domain adaptation is primarily driven by singular vector changes localized near the peaks in the SVD spectra.

  2. Identification of head heterogeneity—the varying behavior of attention heads during CPT stages on different domains, which becomes more pronounced as the pre-train token budget increases. Some heads demonstrate significant changes regardless of CPT domain, while others change in a domain-specific manner. This enables a criterion for ordering heads according to their effect on CPT quality, achieving a quality increase of up to 4% for math CPT of the 7B model upon a partial head rewind. CPT deltas have redundancy in parameters, which increases with model size: up to 60% of attention heads or up to 50% of the smallest singular values in the CPT delta can be dropped without significant degradation of model quality.

  3. Identification of the ability to linearly interpolate between checkpoints after CPT on different domains without quality decrease, for which the term domain connectivity is coined; interpolated model quality improves with the increase in pre-train stage token budget.

  4. Implementation of an open-source tool for matrix spectra analysis—the Diffract package—to ensure reproducibility of findings and facilitate future research.

The experimental setup consists of two sequential stages. Stage 1 pre-trains models from scratch on mixtures sampled from DCLM, using token budgets ranging from 20B to 400B for the 1B model, and also uses 1B and 7B checkpoints pre-trained for 4T tokens, 13B checkpoint pre-trained for 5T tokens, and 32B checkpoint pre-trained for 6T tokens provided by OLMo team. Stage 2 involves continual pre-training starting from pre-trained checkpoints on several data mixtures: DolminoMath-only data (MATH), balanced DCLM+DolminoMath mixtures, DCLM-heavy replay mixtures (TEXT), instruction data from FLAN and Stack Exchange (INST), instruction data combined with DolminoMath, and source code from StarCoder (CODE). The learning rate is initialized to the final value used in Stage 1 and annealed to zero throughout this phase.

Evaluation uses the OLMES framework. Language accuracy is measured as the average performance across ARC-Easy, HellaSwag, and WinoGrande in a 5-shot setting. Math accuracy uses 8-shot evaluation on GSM8K and 4-shot evaluation on MATH-500 with exact-match accuracy. Instruction following and reading comprehension are assessed on DROP and SQuAD datasets. Code generation capabilities are measured using HumanEval.

Main results show that math performance plateaus at 20B CPT tokens for the 4T pre-train checkpoint, and models with a longer pre-train stage achieve better final math quality metrics after the CPT stage. Spectral evolution along the pre-train stage reveals non-monotonic behavior of Frobenius and spectral norms, reaching maximum values at around 100B tokens of pre-train. Effective rank demonstrates an inflection point around 100B tokens with substantial layer-level outliers emerging. At initialization, weight matrices follow the Marchenko-Pastur law; after 20B pre-train, a power law tail starts to form in the spectra of attention heads, in accordance with heavy-tailed self-regularization theory. However, at 100B and larger pre-train budgets, the complexity of attention heads spectra increases—appearance of outliers and multiple narrow peaks—deviating significantly from HTSR models. The spectral structure of MLP blocks stays close to the HTSR model with power law tail for any pre-train token budget.

During CPT, the singular spectra of model weight matrices remain largely unchanged, confirmed by transplanting singular spectra from W pre-train into W domain without affecting model quality. Vector agreement analysis shows that the most significantly changing vectors are associated with peaks in singular value spectra. Head heterogeneity emerges for longer 4T pre-train, where attention heads change differently during CPT—some heads exhibit substantial changes across all domains, while others change in a domain-specific manner. For shorter 20B pre-train, changes are more uniformly distributed across heads.

Head-wise rewind analysis assesses the importance of individual attention heads for CPT quality by ordering heads by scalar importance criteria and incrementally rewinding them to the pre-train state. A novel head ordering criterion is proposed: defining text CPT as a reference domain (since it is carried out on text data similar to the pre-train stage), heads are ordered by the amount of change during target domain CPT compared to the reference CPT. The metric based on Frobenius norms of CPT deltas, per the equation s l,h = scale[0,1] ∥W l,h pre-train − W l,h domain∥F − scale[0,1] ∥W l,h pre-train − W l,h reference∥F, outperforms not only standard spectral heuristics but also the greedy ranking strategy. This achieves a quality increase of up to +4% for math CPT of a 7B model upon rewinding around 15% of heads, and allows rewinding up to 60% of heads without significant quality drop.

SVD truncation of CPT delta reveals that for the 1B model, CPT deltas remain high-rank, and truncation tolerance does not improve monotonically with the pre-train token budget. Model scale has a substantial effect on truncation tolerance: for the 7B model one can remove up to 50% of singular values without a measurable drop in GSM8K accuracy, whereas for the 13B model up to 70% can be removed; pronounced degradation begins beyond 80% singular value removal for the 7B model and beyond 90% for the 13B model.

Domain connectivity is studied by forming interpolants W interp(ω) = (1 − ω) W dom1 + ω W dom2 for ω ∈ [0, 1], referred to as model soups. For the math domain, model soup quality lies below the chord connecting the endpoints (concave) at 20B pre-train, is approximately linear at 400B, becomes mildly convex at 4T for the 1B model, and is clearly convex at 4T for the 7B model, as well as at 5T for the 13B model. This trend indicates that interpolation quality improves with both the pre-train token budget and model size. However, linear interpolation underperforms CPT trained directly on dataset mixtures.

The paper discusses limitations: conclusions are supported by consistent evidence across OLMo 2 1B, 7B, 13B, and 32B models but should be interpreted as robustness within the OLMo family rather than universality across modern LLMs. Not all phenomena are studied uniformly at every scale, and analysis is restricted to the AdamW-based OLMo 2 training recipe. Benchmark coverage remains limited. Validation on other model families, broader benchmark coverage, larger-scale experiments beyond 32B, and optimizer ablations are identified as important directions for future work.

Improvements for AI systems

Improvements to AI Systems:

  1. Domain-Adaptive Head Rewinding for Specialized LLMs
  • Improvement: Implement a post-CPT pruning/rewinding mechanism that uses the proposed head importance criterion (based on Frobenius norm deltas relative to a reference domain) to selectively reset low-importance attention heads to their pre-trained state.

  • What the improved system can do: Achieve up to +4% accuracy on domain-specific tasks (e.g., math) with 15% of heads rewound, while maintaining performance when up to 60% of heads are reset—reducing overfitting and improving generalization without retraining.

  1. Spectral-Aware SVD Truncation for Efficient CPT Delta Compression
  • Improvement: Apply SVD truncation to CPT weight deltas, retaining only the top singular values/vectors (e.g., 50% for 7B, 70% for 13B models) before deployment.

  • What the improved system can do: Reduce model storage and memory footprint by up to 50–70% for domain-adapted weights, with negligible quality loss on benchmarks like GSM8K—enabling cheaper deployment of specialized models on edge devices.

  1. Domain-Connectivity-Based Model Soups for Multi-Domain Adaptation
  • Improvement: Use linear interpolation between CPT checkpoints from different domains (e.g., math and code) to create a single multi-domain model, leveraging the discovered domain connectivity property.

  • What the improved system can do: Seamlessly combine capabilities from multiple specialized domains without retraining, achieving smooth quality trade-offs (e.g., maintaining math and code accuracy simultaneously) and enabling rapid creation of versatile models for mixed-task deployments.

  1. Pre-Train Budget-Aware CPT Scheduling
  • Improvement: Use the finding that longer pre-training (e.g., 4T tokens) improves CPT efficiency and interpolation quality to dynamically allocate pre-training compute before domain adaptation.

  • What the improved system can do: Optimize resource allocation by prioritizing longer pre-training for domains requiring high final quality, reducing wasted CPT tokens (e.g., math plateaus at 20B CPT tokens for 4T pre-trained models) and improving overall model performance per compute unit.

  1. Head-Heterogeneity-Based Targeted Fine-Tuning
  • Improvement: Identify attention heads that change minimally across domains (domain-agnostic) and freeze them during CPT, while only updating domain-specific heads.

  • What the improved system can do: Reduce CPT training time and computational cost by up to 60% (by skipping updates to low-importance heads), while preserving or improving domain accuracy—enabling faster iteration on specialized models.

  1. Spectral-Structure-Aware Initialization for Faster Domain Adaptation
  • Improvement: Initialize CPT with weight matrices that already exhibit complex spectral structures (e.g., multiple peaks in attention head spectra) as seen in longer pre-trained models, rather than starting from simpler spectra.

  • What the improved system can do: Accelerate convergence during CPT, achieving target domain quality with fewer training tokens (e.g., reaching plateau earlier), thereby reducing energy and time costs for domain adaptation.

  1. Robustness to CPT Delta Sparsity for Parameter-Efficient Updates
  • Improvement: Leverage the finding that CPT deltas have redundant parameters (up to 50% of smallest singular values can be dropped) to implement sparse updates during training.

  • What the improved system can do: Train domain-adapted models with sparse gradient updates, reducing memory and communication overhead in distributed training, while maintaining benchmark performance—enabling larger effective batch sizes or longer contexts within the same hardware budget.

  1. Reference-Domain-Based Importance Scoring for Multi-Task Models
  • Improvement: Use a text-CPT checkpoint as a universal reference to rank head importance across all target domains (math, code, instruction), enabling a single importance criterion for mixed-domain adaptation.

  • What the improved system can do: Automatically prioritize updates to heads that matter most for each domain, allowing a single model to be sequentially adapted to multiple domains with minimal quality loss—simplifying multi-task LLM pipelines.

Abstract

We study continual pre-training (CPT) as a mechanism for adapting general-purpose large language models to specialized domains: mathematics, instruction, code, and natural text. Using singular value decomposition of weight matrices, we find that CPT leaves singular value spectra largely invariant, with adaptation driven mainly by changes in singular vectors. An analysis of attention-head projection matrices reveals strong, domain-dependent head heterogeneity, which we exploit to define a head importance criterion: up to 60% of head updates can be removed without measurable quality loss. Selectively rewinding low-importance heads to their pre-trained state improves benchmark accuracy by up to 4% versus the fully trained baseline. Finally, we identify domain connectivity - linear interpolation between CPT checkpoints yields smooth domain-quality interpolation without notable degradation on either domain - and release Diffract, an open-source toolkit for scalable spectral analysis of billion-parameter models.

Sources

Related papers