Adaptive Reparametrized Time for Score-Based Diffusion Sampling

arXiv:2607.02137 · cs.LG, cs.AI, cs.SY, eess.SY, math.OC · Submitted 2026-07-02 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Adaptive Reparametrized Time for Score-Based Diffusion Sampling".

Tom: This paper introduces Adaptive Reparameterized Time (ART), a continuous-time control formulation designed to learn optimal timestep allocation for score-based diffusion sampling.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: So, focusing on the title itself—"Adaptive Reparametrized Time for Score-Based Diffusion Sampling"—it really tells us that the paper isn't just tweaking an existing process; it's proposing a fundamental shift in how we handle the time aspect of diffusion sampling. Jane, can you explain what "reparameterized time" means in simple terms?

Jane: Well, reparameterized time is essentially mapping physical diffusion time onto a new clock variable, which they call t, such that the relationship between the two times is defined by a function ψ(t), where tau = psi(t) and they set boundary conditions like psi(zero) = zero and psi(T) = T. This allows them to control how fast the sampling progresses using a local rate variable, which is their control, denoted by theta.

Lu: That mapping mechanism sounds very elegant; it’s like giving us a dial to tune the temporal progression of the entire reverse trajectory rather than just picking discrete points on a timeline.

Meng: So it's not about changing the underlying model or the solver, but about learning *how* to step through that existing process more intelligently based on where we are in time. That’s a significant difference for integration into current pipelines.

Lalam: This control over the progression suggests a level of intrinsic adaptability in the sampling process itself, which is something we need to explore if we want our models to be more robust across different generation scenarios.

The paper's summary: Tom: Moving on to what this paper actually achieves, they propose Adaptive Reparametrized Time (ART) as a continuous-time control formulation that learns the optimal time change by treating the speed of the clock as the control, and they use a leading-order Euler error surrogate to create a principled objective for allocating those timesteps along the sampling trajectory.

Jane: They do this by defining state dynamics on this new clock using an evolution equation involving theta and F T, while simultaneously imposing a time budget constraint where the integral of theta from zero to T equals T. This whole setup is designed to let the sampler "allocate resolution adaptively along the reverse trajectory, placing finer discretization where it is most beneficial."

Lu: I find that formulation of J theta particularly compelling; by quantifying the local geometric and model-induced stiffness of the probability-flow field with that term Q px, psi q, they’ve created a cost density for when to push or pull the clock speed.

Meng: That cost density idea is interesting because it directly ties computational effort to numerical sensitivity; if Q is large, pushing too fast will likely introduce significant error, so slowing down makes sense. I wonder how computationally intensive calculating that term Q actually is in practice for a high-dimensional model.

Lalam: This approach offers a principled objective function for schedule learning, which means we're not just guessing; we have a mathematical reason to optimize the allocation of our computational budget during inference.

The paper's improvements: Tom: The core improvement they highlight is that ART consistently improves over standard methods like Uniform, DPM, and hand-crafted EDM schedules when using matched evaluation budgets. They show empirical gains in Figure three across various benchmarks, including the largest budgets where the hand-designed EDM schedule is already performing strongly.

Jane: What’s really interesting about their findings is the transferability aspect they report; they demonstrate that a schedule trained on CIFAR–ten can transfer directly to different timestep counts, various datasets, and even different sampling pipelines without needing any retraining at all.

Lu: The claim of transfer across representation spaces—from pixel-space to latent-space—and across different backbone models like EDM2 is quite impressive; that suggests the learned control mechanism is truly capturing the essential dynamics of the reverse process.

Meng: If it transfers so easily, that really lowers the barrier for adopting new sampling techniques in production; we don't have to rebuild our whole pipeline just to switch from one schedule type to another.

Lalam: This generalization across datasets and pipelines means that we can develop a single, highly optimized schedule that works broadly across different generative tasks, which simplifies our deployment strategy significantly.

Conclusion: Tom: So, wrapping up the discussion on "Adaptive Reparametrized Time for Score-Based Diffusion Sampling," the main implication is that by treating time progression as a controllable rate governed by reinforcement learning via ART-RL, we can learn schedules that are far more efficient and accurate than what's available through fixed or ad hoc methods.

Jane: Exactly. The method allows us to adapt the resolution allocation dynamically along the reverse trajectory based on local numerical error sensitivities identified by the Euler error surrogate. It moves schedule learning from a static prescription to a data-driven, continuous control problem that seeks to maximize sample quality under a fixed time budget.

Lu: I think the connection proven in Theorem one between ART and ART-RL is very important; it validates that the complex high-dimensional deterministic control problem can indeed be solved effectively by approximating it with this randomized Gaussian policy formulation, which is a powerful mathematical insight.

Meng: From a practical standpoint, if we take their finding that they can transfer schedules across different pipelines and datasets, it means we can deploy this kind of learning once and benefit from its broad applicability across our entire generative ecosystem without needing bespoke tuning for every new model or dataset.

Lalam: This work has the potential to foster a more flexible AI culture where the sampling process itself becomes an adaptive component rather than a static configuration, allowing us to handle diverse generation needs with greater intrinsic efficiency.

Department of Applied Mathematics, The Hong Kong Polytechnic University · Department of Industrial Engineering and Operations Research, Columbia University

cs.LG, cs.AI, cs.SY, eess.SY, math.OC

Submitted: 2026-07-02

Updated: 2026-09-30

Comments: 37 pages, 14 figures, 8 tables

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 90/100

The gist: This paper introduces Adaptive Reparameterized Time (ART), a continuous-time control formulation designed to learn optimal timestep allocation for score-based diffusion sampling.

Key concepts

Reparameterized Time
This involves mapping physical diffusion time onto a new clock variable, t. This mapping is defined by a function psi(t) where tau equals psi(t), allowing control over the sampling speed using a local rate variable called theta.
ART (Adaptive Reparametrized Time)
ART is a continuous-time control formulation that learns the optimal time change by treating the clock speed as the control. It uses an evolution equation involving theta and F T, subject to a time budget constraint, to allocate resolution adaptively along the reverse trajectory.
Cost Density (Q px)
This term quantifies local geometric and model-induced stiffness of the probability-flow field. It serves as a cost density that dictates when to push or pull the clock speed, tying computational effort directly to numerical sensitivity.

Terminology

Summary

This paper introduces Adaptive Reparameterized Time (ART), a continuous-time control formulation designed to learn optimal timestep allocation for score-based diffusion sampling. It addresses the limitation of fixed or hand-crafted schedules by treating the speed of the sampling clock as a controllable variable, thereby reallocating computational effort along the reverse trajectory to maximize sample quality under a fixed total time budget. The methodology culminates in ART-RL, an auxiliary randomized formulation using Gaussian policies that transforms this high-dimensional deterministic control problem into a continuous-time reinforcement learning problem, providing a principled and data-driven method for schedule learning that demonstrates superior performance and broad generalizability across various diffusion settings.

Methodology: Adaptive Reparameterized Time (ART)

The core idea of ART is to introduce a reparameterized sampling clock, denoted by the mapping ψ: physical diffusion time τ to reparameterized time t, such that τ = ψ(t), with boundary conditions ψ(0) = 0 and ψ(T) = T. This allows the progression of physical diffusion time to be controlled by a local rate variable, the control θ. The state dynamics on this new clock are governed by:

  1. State evolution:

& x ptq “ θ ptq F T ψ ptq, ψ pt0q = pT, ψp0q = 0, ψpTq = T (Equation 4b).

  1. Time budget constraint:

& ∫ T 0 θ ptq dt = T (Equation 6).

This mechanism enables the sampler to allocate resolution adaptively along the reverse trajectory, placing finer discretization where it is most beneficial.

Control Objective: Euler Error Surrogate

To motivate the control objective, ART quantifies the leading-order one-step Euler error residual, denoted as E i. A second-order Taylor expansion of this residual around a fixed step size h i yields an expression that shows the local error is quadratic in θ i, modulated by the term Qpx, ψq evaluated along the trajectory (Equation 7). The coefficient Q captures local geometric and model-induced stiffness of the probability-flow field, meaning regions where Q is large are where aggressive time progression would amplify discretization error. This motivates defining a local cost density as Qpx, ψq θ ptq squared. The resulting objective functional J θ ps, y, ϕq incorporates this cost density along with the time budget constraint (6):

& J θ ps, y, ϕq = E T ∫ T s Qpx ptq, ψ ptqq θ ptqq dt - γ θ ∫ T dt + γ T x psq = y, ψ psq = ϕ (Equation 9).

Reinforcement Learning Formulation: ART-RL

Since the deterministic control problem (10) is numerically prohibitive in high dimensions, ART-RL introduces an auxiliary randomized formulation. This involves replacing the deterministic control θ with a stochastic policy that assigns time-warping rates. Specifically, it employs a Gaussian policy class whose variance depends on the local numerical sensitivity encoded by Q:

& π pλq(t, x, ψq) = N (µ pt, x, ψq, λ Qpx, ψq dt) (Equation 11).

The connection between the deterministic and randomized formulations is established via Theorem 1. The optimal value function V of the original problem satisfies a Hamilton–Jacobi–Bellman (HJB) equation (15), and the value function V pλq of the randomized problem satisfies an equivalent HJB equation (16). Crucially, Theorem 1 proves that the mean of the optimal ART-RL Gaussian policy is optimal for the latter, meaning µ˚pt, x, ψq is the optimal policy for the original ART control problem.

Actor-Critic Algorithm

The solution to ART-RL is found using an actor–critic scheme adapted to continuous time. The critic estimates the value function V pλq (20), which satisfies a linear PDE (21). The actor updates its mean by maximizing this value function, leading to the policy mean µ˚pt, x, ψq = V J x Fpx, ψq - γ 2Qpx, ψq. This is formalized in Theorem 3(ii), which shows that the optimal policy µˆ is µ˜ as defined in Theorem 2-(ii). The resulting update rules for the parameters are derived by solving moment conditions (26) using stochastic approximation, leading to iterative updates for the critic and actor parameters (27a, 27b). These theoretical results are discretized into a uniform time grid to yield an implementable algorithm (Algorithm 1), which involves an inner loop generating a trajectory and an outer loop updating the parameters.

Improvements for AI systems

As a fastidious researcher, I have analyzed the ART (Adaptive Reparameterized Time) framework presented in this paper. The core contribution is a principled, control-theoretic method for learning optimal timestep allocation in score-based diffusion sampling by treating time progression as a controllable rate governed by reinforcement learning (ART-RL).

Here are the specific improvements and capabilities that can be realized by implementing the ART system:


)

The improved AI system will be able to generate higher quality, more stable, and more efficient samples across various tasks compared to standard diffusion samplers. This is achieved by optimizing the computational budget (number of function evaluations) at every stage of generation.

Specific improvements include:

  1. [][]

  2. Higher Sample Quality with Fixed Budget: The system will consistently achieve superior sample quality (lower FID/LeNet-FID) compared to uniform and hand-crafted schedules across a fixed computational budget, even when using the same pretrained score model and solver. This means that for a set number of function evaluations (NFE), the resulting image or sample will be significantly sharper, more coherent, or more representative of the target distribution than what is achievable with baseline methods.

  3. Adaptive Resolution Allocation: The system will intelligently allocate computational effort to regions where numerical error is most sensitive (i.e., later stages of the reverse process) and slow down in regions where noise is already low (early stages). This leads to a more principled use of the budget, ensuring that every function evaluation contributes maximally to accuracy rather than being spread uniformly.

  4. Robustness Across Solvers and Pipelines: The learned schedule is designed to be architecture-agnostic and solver-agnostic. This means the ART system can be dropped into existing pipelines (like EDM or DPM) by simply replacing the timestep grid, without requiring retraining of the score model or modification of the numerical integrator.

  5. Cross-Domain Generalization: The learned schedule exhibits broad generalizability across different datasets (CIFAR-10, AFHQv2, FFHQ, ImageNet) and representation spaces (pixel-space vs. latent-space). This allows a single learned schedule to be applied successfully to new image generation tasks without needing task-specific retraining.

  6. Amortized Training Cost: The system enables a one-time training cost for learning the optimal schedule, which can then be distilled into a fixed, deterministic lookup table (a precomputed time grid). This means that during actual inference (sampling), there is no additional computational overhead beyond reading this precomputed list of timesteps.

  7. High-Resolution Synthesis: When applied to modern high-resolution pipelines (like EDM2 on ImageNet–512), the ART schedule maintains superior performance, demonstrating its utility in state-of-the-art generative settings where complex backbones and latent spaces are involved.

In summary, the improved AI system will transition from relying on static, heuristic time grids to employing a dynamic, data-driven control mechanism that optimizes sampling efficiency and fidelity in real-time according to the specific trajectory of the reverse diffusion process.

Abstract

We study timestep allocation for score-based diffusion sampling, where a learned reverse-time dynamics is discretized on a finite grid. Uniform and hand-crafted schedules are standard choices, but they rely on fixed prescriptions and can therefore be suboptimal. To address this limitation, we propose Adaptive Reparameterized Time (ART), a continuous-time control formulation that learns a time change by treating the speed of the sampling clock as the control, so that a uniform grid on the learned clock induces adaptive timesteps in the original diffusion time. Based on a leading-order Euler error surrogate, ART provides a principled objective for allocating timesteps along the sampling trajectory. To solve this deterministic control problem, we introduce ART-RL, an auxiliary randomized formulation with Gaussian policies that turns schedule learning into a continuous-time reinforcement learning problem. We prove that the randomized ART-RL formulation is equivalent to ART at the optimizer level, in the sense that its optimal Gaussian policy recovers the optimal ART time-warping rate through its mean. We further establish policy evaluation and policy improvement characterizations and derive trajectory-based moment identities that yield implementable actor--critic updates for learning the schedule. Across experiments ranging from controlled low-dimensional settings to image generation, ART-RL can be plugged into existing diffusion samplers by changing only the timestep grid, consistently improving sample quality over strong baseline schedules at matched budgets while leaving the rest of the sampling pipeline unchanged. The learned schedules also exhibit broad generalization, transferring without retraining across sampling budgets, datasets, solvers, pipelines, and representation spaces.

Sources

Related papers