How Should LLMs Consume High-Quality Data? Optimal Data Scheduling via Quality-Aware Functional Scaling Laws

arXiv:2605.25698 · cs.LG, cs.AI · Submitted 2026-08-24 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "How Should LLMs Consume High-Quality Data? Optimal Data Scheduling via Quality-Aware Functional Scaling Laws".

Jane: The paper was written by Zhitao Zhu, Xili Wang, Shizhe Wu, Jiawei Fu and Xiaoqing Liu from Peking University and Meituan Company (Meituan).

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Summary: Tom: Okay, so if the title was about *how* to schedule data, this next part of the paper summary dives into the actual results, pointing out a really important correlation that they found.

Jane: What I took away is this finding about a negative correlation between two things: how much performance drops early on (the early-phase gap), and how much it ultimately improves by the end (the final gain).

Lu: That negative correlation is the core empirical finding, isn't it? It suggests an inverse relationship—if you suffer a big deficit during the stable phase, your final gains are likely to be smaller or even negative.

Meng: That makes intuitive sense when you think about noise. If the model is struggling with the small batch size early on, that struggle might be accumulating bad habits or instability that can't be overcome later.

Lalam: I found it fascinating how they linked this deficit to whether the task was bottlenecked by signal accumulation or dominated by noise throughout the extended stable phase.

Tom: So, let’s look at the examples they give: benchmarks like GSM8K, MATH, and HumanEval+—those relate to composition and code. These are the ones that cause little early deficit but achieve huge gains.

Jane: And those big gains—we're talking up to +four point two three improvements! It’s not just a minor tweak; it's a major performance boost attributable to this scheduling technique.

Lu: This strongly suggests that for compositional reasoning and algorithmic code generation, the *signal* is the bottleneck. The model desperately needs extra gradient steps to synthesize complex relationships from limited data.

Meng: That aligns with what we see in advanced systems: if you're trying to solve a tricky multi-step math problem, you need consistent exposure to those steps, not just a flood of general text that might distract you.

Lalam: The contrast with CMMLU and BCB is also so telling. These benchmarks, where the deficit was large initially, showed smaller or even negative final gains. It seems the noise cost outweighed any potential signal benefit.

Jane: So, to simplify that for our listeners: if the task relies heavily on synthesizing new information from a few core principles—like math—it can tolerate and even benefit from the slow, steady accumulation of data steps.

Tom: But if the task is more sensitive to general knowledge or diverse inputs, that prolonged low-batch-size phase might just introduce too much confusion or noise for the model to handle effectively.

Lu: It's a clear mechanistic distinction: signal accumulation bottleneck versus noise dominance. The paper essentially mapped out which types of learning are fundamentally different in their requirements.

Meng: If I were implementing this, I'd focus my resources on identifying that optimal transition point—the moment the model switches from needing basic signal to being overwhelmed by noise.

Lalam: This really refines our understanding of model failure modes; it’s not just about reaching a plateau, but understanding *why* the plateau was reached at that specific point.

Improvements: Tom: We’ve established the correlation, and we know which tasks benefit most. Now, the paper moves into suggesting concrete improvements based on this dynamic scheduling approach.

Jane: The main implication here is moving toward a truly *task-aware* data scheduling. It means treating the training process less like a single pipeline and more like an adaptive curriculum for the model.

Lu: The idea of estimating the early-phase gap from a small pilot run is revolutionary because it allows for real-time adaptation. You don't have to wait until you've trained for weeks to realize your data mix was wrong.

Meng: From an engineering standpoint, this suggests we need diagnostic tools that can quickly profile the model’s initial instability or stability across different domains before committing massive compute resources.

Lalam: And the proposed scheduling structure is incredibly elegant: prioritizing compositional data—math, code, chain-of-thought—during the stable phase to maximize signal accumulation.

Tom: Because those complex tasks are what really benefit from that steady, high-quality input during the middle stages of training, right? It’s like hitting that sweet spot of deep learning.

Jane: And

Paper discussion segment 3: Jane: It means the AI shouldn't just read everything equally; instead, it should prioritize certain types of learning at specific times during its training life cycle. Think about how a student studies for finals; they don't just read textbooks randomly until the last minute.

Tom: Exactly! The paper suggests that when an LLM is first getting stable, it should really focus on the hard stuff, like mathematical reasoning or complex code generation, because those are the skills that build up incrementally.

Lu: That makes me think about how we could model this process dynamically; instead of a fixed curriculum, the AI’s own performance metrics could dictate whether it needs a boost in compositional data or if it should switch to broad knowledge acquisition.

Meng: From an engineering standpoint, implementing that dynamic switch sounds computationally expensive, though; you'd need continuous, real-time monitoring of the model's internal state to decide which data stream to activate next.

Jane: So the model would essentially self-diagnose if it’s stronger in knowledge retrieval or in logical deduction at any given moment, right?

Tom: Right! And that’s where the quality comes in—it's not just about volume; it's about ensuring that when it does tackle those math problems, they are high-quality examples of diverse reasoning.

Lu: And I think this opens up a whole new paradigm for data curation, moving away from just collecting large text dumps and toward building highly structured, task-specific knowledge graphs that can be selectively exposed.

Meng: But how do you guarantee the quality of that data stream across billions of tokens? Curation is one thing; maintaining perfect fidelity at scale is something else entirely.

Lalam: What I find exciting about this isn't just the scheduling itself, but what it implies for culture—if we can teach models to learn in this highly focused way, AI could help us solve complex societal problems by simulating expertise across different human domains.

Tom: You mean like letting the AI simulate a deep dive into, say, urban planning theory before tackling a massive historical archive?

Jane: That’s right. It suggests that the best learning is targeted and layered, building up from core skills outwards toward general understanding.

Meng: If we could apply this to specialized fields—like medicine or law—we could potentially create AI assistants that graduate their knowledge, becoming much safer and more reliable over time.

Lu: That refinement process is incredible; it’s not just training a model, it's engineering an educational trajectory for intelligence itself.

Lalam: It elevates AI from being just a powerful search engine to being a true collaborator—a digital tutor that understands the optimal learning pace for human or machine users alike.

Tom: Wow, so we're moving beyond just scale and into genuine intelligence management, aren't we? This whole idea of optimizing the *learning process* is huge.

Conclusion: Tom: We've spent a lot of time diving into how this paper suggests we should manage high-quality data, and it’s clear that scheduling is far more complex than just adding better examples to training.

Jane: The core idea is that the LLM isn't static; it needs a dynamic strategy for feeding itself, changing its approach based on whether it's struggling with signal accumulation or being overwhelmed by noise.

Tom: And Lu has pointed out how this resolves the old conflict between curriculum learning and decay schedules—it’s a whole new framework for the functional scaling laws!

Lu: It fundamentally changes how we think about training dynamics, allowing us to predict optimal performance based on whether the bottleneck is signal acquisition or noise accumulation.

Meng: From a practical standpoint, it suggests we should stop thinking of data as just "data" and start seeing it as a resource that needs to be strategically deployed over time in a much more efficient way.

Lalam: It really shines how this could improve the culture of AI by making us more precise in our goals, allowing us to train models that are reliably focused on complex reasoning instead of just broad memorization.

Tom: We've seen how it works across different tasks, too—the math and code benchmarks were particularly responsive to this scheduling method.

Jane: So, as a final summary of the findings in "How Should LLMs Consume High-Quality Data? Optimal Data Scheduling via Quality-Aware Functional Scaling Laws," we're looking at a model that truly learns in its own optimal rhythm.

Meng: It’s a massive shift towards fine-grained control over how the AI processes its knowledge base.

Lu: This paper has provided us with the theoretical tools to manage AI training like never before, guiding our future work on optimizing these dynamic systems.

Lalam: I think this will be a foundational principle that allows us to build much more trustworthy and capable intelligent agents in the world.

Tom: Well, it sounds like we've reached a powerful conclusion; I think we might be ready to look at some new frontiers in AI architecture next time.

Zhitao Zhu, Xili Wang, Shizhe Wu, Jiawei Fu, Xiaoqing Liu

Peking University · Meituan Company (Meituan)

cs.LG, cs.AI

Submitted: 2026-08-24

Updated: 2026-08-25

Importance score: 87/100

The gist: This paper investigates how Large Language Models (LLMs) should consume scarce high-quality data during training to maximize performance.

Key concepts

Negative Correlation
This core finding suggests an inverse relationship between how much performance drops early in training (the early-phase gap) and the final improvement achieved by the end. If a model suffers a significant deficit initially, its ultimate gains are likely to be smaller or even negative.
Signal Accumulation vs. Noise Dominance
This mechanism determines if a task needs consistent, complex input to synthesize relationships (signal accumulation) or if it is overwhelmed by general data during prolonged training phases (noise dominance). This distinction dictates the optimal scheduling strategy.
Task-Aware Data Scheduling
The concept of treating the training process as an adaptive curriculum. Instead of a fixed pipeline, this approach prioritizes specific types of data, like compositional math or code, at certain times to maximize signal accumulation for complex learning.

Terminology

Summary

This paper investigates how Large Language Models (LLMs) should consume scarce high-quality data during training to maximize performance. It addresses a critical conflict in current practices where conventional learning rate decay occurs exactly when models encounter their best data, effectively reducing learning intensity at the most vital stage. By developing a theoretical framework for joint optimization of data quality and training dynamics, the authors provide guidance on how to schedule high-quality tokens alongside batch-size transitions to optimize model capabilities.

Theoretical Framework

The researchers develop a quality-aware functional scaling law that extends existing functional scaling laws by incorporating time-varying data quality. They model label noise as a two-component mixture where high-quality data corresponds to small noise variance and low-quality data corresponds to large variance. By formulating the joint optimization of the data-quality schedule and batch-size schedule as a variational problem, they derive asymptotic closed-form optimal schedules.

The theory reveals that the optimal strategy depends on whether the learning bottleneck is noise accumulation or signal acquisition, leading to two distinct regimes:

((

  1. The noise-limited regime (where signal acquisition is not the primary bottleneck): High-quality data should be used as a signal amplifier. By lowering the batch size during this phase, the model converts cleaner data into more optimization steps without amplifying noise.

  2. The signal-limited regime (where noise accumulation is the primary bottleneck): High-quality data should be used as a noise suppressor. In this case, late placement of clean data reduces terminal noise without sacrificing the accumulation of signal.

((

The Drop-Stable-Rampup Strategy

To translate these theoretical insights into practice, the authors propose a new midtraining recipe called Drop-Stable-Rampup (DSR). This strategy is designed to exploit both roles of high-quality data by coordinating the transition from noisy pretraining data to curated midtraining corpora. The DSR schedule consists of three specific phases:

((

  1. Drop: Cut the batch size at the quality transition to accumulate signal via more gradient steps under low-noise conditions.

  2. Stable: Hold the batch size at a minimum level for a period to maximize signal acquisition for signal-limited capabilities.

  3. Rampup: Linearly grow the batch size (or decay the learning rate) to suppress terminal noise and ensure convergence.

((

Empirical Results and Validation

The authors validated their theory using a 15B Mixture-of-Experts (MoE) model midtrained on 108B tokens. The experimental results demonstrate that DSR significantly outperforms conventional baselines like Cosine-decay and Warmup-Stable-Decay (WSD). On average, DSR improves accuracy by +1.70 over WSD and +2.98 over Cosine-decay across 14 benchmarks.

The paper highlights that the performance gains are particularly large on tasks requiring complex reasoning:

((

Mathematical reasoning:

  • GSM8K: +4.23 improvement over WSD

  • MATH: +2.80 improvement over WSD

Algorithmic code generation:

  • HumanEval+: +2.44 improvement over WSD

  • MultiPL-E: +2.66 improvement over WSD

((

The study also reveals a heterogeneous training dynamic across different benchmarks. Tasks with small early-phase gaps—meaning they do not suffer significantly from the increased noise of small batch sizes during the stable phase—such as mathematical reasoning and algorithmic code, achieve the largest final gains. Conversely, tasks more sensitive to noise, such as certain knowledge-based benchmarks, may see smaller benefits from this specific scheduling approach.

Improvements for AI systems

To improve an AI system based on this research, I would implement a midtraining scheduling protocol called the following:

The improved AI system will utilize a dynamic, three-phase training regime during its midtraining stage (the transition from massive noisy pretraining to high-quality curated data). Instead of using standard constant batch sizes or simple learning rate decay, the system will implement the following specific operational changes:

  1. Implement a Drop phase at the quality transition: Upon switching from general web-crawled data to high-quality curated corpora (e.g., textbooks, math proofs, expert code), immediately reduce the global batch size (e.g., from 4096 to 512 or 1024).

  2. Maintain a Stable phase: Hold this reduced batch size constant for a period determined by the target task's complexity (empirically 25% of the midtraining budget) to maximize the number of gradient updates performed on clean data.

  3. Execute a Ramp-up phase: Linearly increase the batch size (or decay the learning rate) toward the end of training to suppress accumulated noise and ensure final convergence.

The improved AI system will be able to:

  1. Achieve significantly higher performance in complex reasoning tasks, specifically demonstrating measurable improvements in mathematical reasoning (e.g., +4.23 on GSM8K, +2.80 on MATH) and algorithmic code generation (e.g., +2.44 on HumanEval+).

  2. Optimize its signal-to-noise ratio by converting high-quality data into additional gradient steps (signal amplification) rather than just using it for final convergence (noise suppression).

  3. Exhibit superior performance in compositional reasoning tasks where the benefits of increased update intensity outweigh the temporary increase in gradient noise caused by smaller batch sizes.

Sources

Related papers