BRACE: Taming Sharp Irregularities via Barycentric Rational Forecasting for Fast Diffusion Transformers Inference

arXiv:2608.07572 · cs.CV, cs.AI · Submitted 2026-08-04 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "BRACE: Taming Sharp Irregularities via Barycentric Rational Forecasting for Fast Diffusion Transformers Inference".

Jane: The paper was written by the authors from Sichuan University and School of Artificial Intelligence, Sichuan University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the arXiv radio hour, everyone. I'm Tom, and as always, I'm here with my co-host, Jane. Today we're looking at a paper with a pretty bold title: "BRACE: Taming Sharp Irregularities via Barycentric Rational Forecasting for Fast Diffusion Transformers Inference."

Jane: And Tom, I have to say, that title is a mouthful, but it's actually describing a really practical problem. We all love these diffusion models that generate images and videos, but they're painfully slow because they need dozens of steps to produce one result.

Tom: Right, and the paper is essentially trying to make those models faster without wrecking the quality. The authors are from Sichuan University, and they're tackling this by looking at how the internal "features" of the model change over those steps.

Jane: Exactly. The key idea is that these features are mostly smooth, but sometimes they jump around wildly. The old methods for predicting those jumps, based on polynomial math, tend to blow up. This paper, BRACE, uses a different kind of math called rational functions to handle those wild jumps.

Tom: So instead of trying to guess the direction of the jump using derivatives, which is like guessing a car's path by looking at its speedometer, BRACE looks at the actual positions of the car over the last few seconds and draws a smoother, more stable curve through them.

Jane: That's a great analogy, Tom. And the payoff is that they can skip computing whole layers of the network, skipping ahead several steps at a time, while keeping the image looking almost as good as the full, slow version.

Tom: And that's the big deal here. We're talking about potentially doubling or tripling the speed of image generation, maybe even more, which could make these tools actually usable in real-time applications.

Jane: It really could change how we interact with generative software. Instead of waiting for a prompt to render, you could get near-instant feedback.

Tom: I'm curious to see how they actually pulled this off mathematically. Let's dig into the summary next.

Summary: Jane: So, Tom, we've established that BRACE is about speed, but let's get into what the paper actually claims to achieve. The summary is pretty clear that they're setting a new standard for how fast you can push these models.

Tom: Yeah, and the numbers back it up. They tested it on a bunch of different models. On ImageNet with the DiT-XL/two model, they got a three point five six times speedup while actually lowering the FID score, which measures image quality, compared to the baseline.

Jane: And it's not just images. They also ran it on FLUX.one for text-to-image and HunyuanVideo for text-to-video. In every single case, they beat the existing acceleration tricks, whether that was simple caching or the fancier derivative-based forecasting methods.

Tom: The key phrase in the summary is "training-free." That means you don't have to retrain the massive model to get this speedup. You just plug this new forecasting module in during the inference phase, and it works.

Jane: That's a huge practical advantage. Retraining a model like FLUX or HunyuanVideo would cost a fortune in compute. Being able to just swap out the prediction logic is a massive win for anyone actually deploying these systems.

Tom: And the quality isn't just "acceptable." They're showing that BRACE actually preserves the fine details better than the other fast methods. The summary mentions they achieve the lowest LPIPS, which is a metric for perceptual similarity, meaning the images look closer to the original.

Jane: So they're not just cutting corners and hoping for the best. They're using a smarter mathematical foundation that genuinely handles the tricky parts of the generation process better.

Tom: It sounds like they've really found a sweet spot. But I'm wondering, what exactly is the "sharp irregularity" problem they're solving? Let's get into the improvements they're proposing.

Jane: Good lead-in, Tom. Let's talk about the core improvement next.

Improvements: Tom: So, Jane, the paper's main improvement is moving away from what they call "derivative-driven polynomial extrapolation." That's a technical way of saying the old methods use the slope of the feature curve to guess the future.

Jane: And the problem with that is if the curve has a sharp corner, like a hairpin turn on a race track, the slope is basically meaningless. The derivative is huge or undefined, and the prediction goes flying off the track.

Tom: Exactly. The paper shows this empirically. They plotted the feature trajectories and found they're globally smooth but locally very sharp. The old Taylor-series methods, like TaylorSeer, fail at these sharp points.

Jane: So BRACE's improvement is to use a "barycentric rational" forecast instead. Instead of using slopes, it uses the actual historical feature values and combines them in a clever weighted average.

Tom: And the weights aren't random. They're based on Chebyshev polynomials, which are famous in numerical analysis for being super stable. This "Adapted Chebyshev Weights" scheme is what lets them extrapolate far into the future without the prediction oscillating wildly.

Jane: The paper also introduces a "Local Sliding Window" that only keeps the last few feature states. This is important because it means the model isn't trying to predict based on ancient history that's no longer relevant.

Tom: Right, it's like trying to predict the weather. You look at the last few hours, not the weather from last month. This local focus, combined with the rational math, is what gives them the stability.

Jane: And the math checks out. They provide a proof showing their error is bounded by the step size and the first derivative, whereas the Taylor method's error grows with the step size to a high power. That's a fundamental difference in stability.

Tom: So they've built a more robust foundation. But how does that actually play out in the real world? Let's look at the first page of the paper to see the big picture.

Jane: Let's do it.

First Page: Tom: We're back, and we're looking at the first page of "BRACE: Taming Sharp Irregularities via Barycentric Rational Forecasting for Fast Diffusion Transformers Inference." The page sets the stage by highlighting the massive computational cost of diffusion transformers.

Jane: And it introduces that core observation we've been talking about: the feature trajectories are globally smooth but locally non-smooth. The paper even includes a figure showing these sharp irregularities across different layers of the network.

Tom: The figure on the right is the real kicker. It shows TaylorSeer, a leading method, completely diverging from the true feature path at one of these sharp inflection points. Meanwhile, BRACE tracks the actual manifold almost perfectly.

Jane: That visual really sells the story. It's not just a minor improvement; it's a fundamental difference in how the method handles stress. The old method breaks down, and the new one stays stable.

Tom: The page also introduces the authors' motivation: they looked at classical numerical analysis, where rational extrapolation is known to outperform polynomials for rapid transitions. They basically took a well-established mathematical tool and applied it to a modern machine learning problem.

Jane: And the results speak for themselves. The abstract mentions state-of-the-art quality-efficiency trade-offs across various architectures with negligible computational overhead. That means the forecasting itself is almost free.

Tom: So the first page sets up the problem, shows the failure of existing methods, and introduces the elegant solution. It's a very compelling opening.

Jane: It really is. And it makes you wonder, where does this go from here? Let's wrap up our thoughts.

Conclusion: Tom: Well, we've had a great time unpacking "BRACE: Taming Sharp Irregularities via Barycentric Rational Forecasting for Fast Diffusion Transformers Inference." Let's do a final recap for our listeners.

Jane: Absolutely. The paper tackles the slow inference problem in diffusion transformers by replacing unstable derivative-based forecasting with a stable barycentric rational approach. It uses a local sliding window of cached features and adapted Chebyshev weights to predict future states.

Tom: And the results are impressive across the board. They achieved significant speedups on DiT-XL/two FLUX.one-dev, and HunyuanVideo while maintaining or even improving generation quality compared to other acceleration methods.

Jane: The key takeaway is that by choosing the right mathematical foundation, you can tame the sharp irregularities that plague these models, making them both faster and more reliable.

Tom: And the implications are huge. This could make real-time, interactive image and video generation a reality, which would be transformative for creative tools, virtual environments, and even scientific visualization.

Jane: It's a clever piece of work that shows how classical mathematics can still offer fresh solutions to modern engineering challenges.

Tom: Couldn't agree more, Jane. We'll be sad to see this paper go, but we're excited to see what's next on the arXiv feed. Thanks for listening, everyone.

Jane: Goodbye for now, and happy generating.

Sichuan University · School of Artificial Intelligence, Sichuan University

cs.CV, cs.AI

Submitted: 2026-08-04

Updated: 2026-08-28

Comments: 10 pages, 6 figures, 5 tables. Project page: https://youngkinlon.github.io/BRACE-Taming-Sharp-Irregularities-via-Barycentric-Rational-Forecasting-for-Fast-DiT-Inference/

Code: https://github.com/black-forest-labs/flux

Project page: https://youngkinlon.github.io/BRACE-Taming-Sharp-Irregularities-via-Barycentric-Rational-Forecasting-for-Fast-DiT-Inference

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 63/100

Key concepts

Sharp Irregularities
These are sudden, wild jumps in the internal features of diffusion model layers. Old methods based on polynomial math fail when these sharp corners occur because the derivative becomes meaningless, leading to unstable predictions.
Barycentric Rational Forecasting
This is BRACE's core improvement. Instead of using slopes (derivatives) to guess future feature positions, it uses a rational function that combines actual historical feature values in a clever weighted average for more stable extrapolation.
Adapted Chebyshev Weights
These are specific weights used in the forecasting method. They are based on Chebyshev polynomials, which are known for being very stable in numerical analysis, helping the model extrapolate far into the future without wild oscillations.
Training-free Acceleration
BRACE provides speedup by plugging its new forecasting module into existing models during inference. This means the massive base model does not need to be retrained to gain this acceleration, offering a major practical advantage.

Terminology

Summary

Summary

This paper introduces BRACE (Barycentric Rational Forecasting with Chebyshev Enhancement), a training-free acceleration framework for Diffusion Transformers (DiTs). The core problem addressed is the prohibitive computational cost of iterative sampling in DiTs, which limits real-time deployment. Existing acceleration methods that use feature caching fall into two paradigms: cache-and-reuse (e.g., DeepCache, FORA), which degrades quality over long skip intervals, and cache-then-forecast (e.g., TaylorSeer, HiCache), which uses derivative-driven polynomial extrapolation. The paper identifies a fundamental flaw in the latter: derivative-driven polynomial paradigms... suffer from the inherent instability of polynomial extrapolation over large skip intervals.

The authors empirically analyze DiT feature trajectories and reveal a dual geometric nature: they are globally smooth yet frequently exhibit sharp irregularities and local non-smoothness. This is visualized in Figure 1(a) and aligns with phase transition phenomena in diffusion processes. They show that finite-difference gradient estimates become unreliable at sharp inflection points, and rigid polynomials lack the structural adaptability to model abrupt shifts, leading to severe extrapolation divergence outside the local window. Figure 1(b) demonstrates that TaylorSeer diverges at inflection points while BRACE tracks the actual manifold.

To address this, BRACE shifts the paradigm from derivative-driven polynomial extrapolation to feature-driven rational forecasting. Instead of estimating unstable local derivatives, it directly aggregates raw historical features cached in a sparse local sliding window. The method uses a barycentric rational function with adapted Chebyshev weights to ensure numerical stability and tame abrupt transitions where traditional derivative-driven methods fail.

The BRACE framework operates as follows: During full computation steps, it caches exact intermediate features and their timesteps into a FIFO Local Sliding Window of capacity C. When a skip interval is triggered, a Domain Mapping operation normalizes the cached timestamps into the canonical interval [-1, 1]. Then, the Barycentric Rational Extrapolation module synthesizes the predicted feature using the second barycentric form: Fpred = (Σ w j / (x(t pred) - x(τ j)) * F τ j) / (Σ w j / (x(t pred) - x(τ j))). The adapted Chebyshev weights incorporate an alternating sign (-1)(j+1) as a stabilizer and a boundary sensitivity coefficient γ. The paper notes that with ultra-sparse capacity (C ≤ 3), optimal Chebyshev nodes are mathematically identical to equidistant points, enabling simple fixed-interval sampling.

The paper provides a stability analysis showing that BRACE's extrapolation error is bounded by the interval k and the Lipschitz constant L, whereas Taylor methods incur truncation errors of O(s(m+1)F(m+1)∞), which can diverge during sharp transitions.

Experiments are conducted across three tasks: class-conditional image generation on ImageNet-1K with DiT-XL/2, text-to-image generation with FLUX.1-dev, and text-to-video generation with HunyuanVideo. Results show BRACE consistently achieves state-of-the-art quality-efficiency trade-offs. For DiT-XL/2, BRACE achieves the lowest FID and sFID across all skip intervals (e.g., FID 2.46 at 3.56× speedup vs. 2.51 for HiCache and 2.59 for TaylorSeer). For FLUX.1-dev, BRACE achieves the highest ImageReward, CLIP Score, and CycleReward, and the lowest LPIPS in nearly all settings. For HunyuanVideo, BRACE achieves the highest VBench score of 80.14 at N=6, closest to the unaccelerated 50-step baseline of 80.68.

Ablation studies show that window capacity C=3 provides optimal performance, and boundary sensitivity γ=0.7 is established as the default for perceptual quality. The paper also benchmarks the adapted Chebyshev weights against Uniform, Berrut, and Floater-Hormann schemes, finding the adapted weights achieve superior alignment with DiT feature dynamics.

Improvements for AI systems

Based on the BRACE paper, here are the specific improvements I can implement in an AI system, and what the improved system can do:


  • Current limitation: The system uses Taylor-series or Hermite-based polynomial extrapolation to predict future diffusion features, which diverges at sharp trajectory irregularities.

  • Improvement: Implement the second barycentric rational form with adapted Chebyshev weights for feature forecasting. This directly aggregates raw cached features rather than estimating unstable high-order derivatives.

  • Current limitation: The system caches features globally or with large windows, causing long-range feature corruption and unstable extrapolation.

  • Improvement: Maintain a fixed-capacity FIFO queue of state tuples (F τ, τ) with capacity C = 2 or 3. This confines the forecasting horizon to a sparse local window, ensuring local smoothness and bypassing global nonlinearities.

  • Current limitation: Raw diffusion timesteps are used directly, causing numerical instability in the barycentric formulation.

  • Improvement: Apply an affine mapping to project cached timestamps into the canonical interval [−1, 1] before extrapolation. This ensures scale-invariance across different schedulers and provides well-defined theoretical error bounds.

  • Current limitation: Standard uniform, Berrut, or Floater-Hormann weights lack the ability to capture asymmetric temporal dynamics of DiT features.

  • Improvement: Implement the alternating-sign weight scheme with a tunable boundary sensitivity coefficient γ (optimal value 0.7 for perceptual quality, 0.4 for video). This prevents numerical collapse and flexibly accommodates varying model dynamics.

  • Current limitation: The system has no runtime verification of extrapolation stability, risking silent quality degradation.

  • Improvement: Implement the theoretical error bound from Eq. (11): E(x C + s) ≤ (Σw i·L/λ)·k, where L is the Lipschitz constant of features, k is the skip interval, and λ is a structural constant. Monitor this bound during inference and fall back to full computation if it exceeds a threshold.

  • Achieve 3.56×–5.55× speedup with FID as low as 2.46 (vs. 2.29 for full 50-step DDIM), outperforming TaylorSeer (2.59), HiCache (2.51), and FORA (8.45) at matched acceleration.

  • Maintain structural integrity in text rendering: preserve double letters, correct spelling, and sharp typography even at 5.55× acceleration, where baselines produce merged/deformed characters.

  • Preserve fine-grained details (animal fur, object edges, backgrounds) without over-smoothing or ghosting artifacts.

  • Achieve 5.56× FLOPs reduction with VBench score of 80.14, recovering 99.3% of the unaccelerated 50-step baseline (80.68), while TaylorSeer (79.87) and HiCache (79.93) lag behind.

  • Preserve temporal consistency: maintain correct entity presence (e.g., both bird and cat), anatomical correctness (no extra horse legs), and coherent motion across frames.

  • Achieve 4.16×–5.55× speedup with ImageReward ≥ 1.0021 (exceeding the full 50-step baseline of 0.9898), CLIP Score ≥ 27.72, and CycleReward ≥ 0.87, all state-of-the-art.

  • Maintain perceptual fidelity: LPIPS as low as 0.3770 (vs. 0.5206 for FORA), SSIM up to 0.7598, and PSNR up to 19.37 dB.

  • Training-free acceleration: Plug-and-play integration with any DiT architecture without retraining or fine-tuning.

  • Robustness to sharp transitions: Handle phase-transition regions in the denoising trajectory where feature states change drastically, without overshooting or oscillation.

  • Adaptive stability: Automatically adjust boundary sensitivity γ based on task (0.7 for images, 0.4 for video) to balance reconstruction fidelity vs. perceptual quality.

  • Numerical stability: Guarantee bounded extrapolation error proportional to the skip interval k and Lipschitz constant L, avoiding the O(s(m+1)·F(m+1)∞) divergence of Taylor methods.

  • Computational overhead: The barycentric rational extrapolation adds negligible cost (microseconds per layer) compared to the DiT forward pass.

  • Memory footprint: Only C ≤ 3 cached feature maps are retained, requiring minimal additional memory.

  • Compatibility: Works with DDIM, DPM-Solver, and rectified flow samplers; applicable to DiT-XL/2, FLUX.1-dev, HunyuanVideo, and other transformer-based diffusion models.

Abstract

Diffusion Transformers (DiTs) have demonstrated exceptional performance in high-fidelity image and video generation. To alleviate their massive computational overhead, temporal feature caching has been proposed to bypass redundant computations. However, existing cache-then-forecast methods driven by derivative-based polynomials often cause severe quality degradation under high acceleration due to unstable long-step predictions. To address this bottleneck, we propose Barycentric Rational Forecasting with Chebyshev Enhancement (BRACE). Motivated by the observation that DiT feature trajectories are globally smooth yet frequently exhibit sharp irregularities and local non-smoothness, BRACE shifts the paradigm from derivative-driven polynomial extrapolation to feature-driven rational forecasting. Specifically, it maintains a local sliding window to cache sparse historical features and leverages adapted Chebyshev weights to formulate a barycentric rational function, directly aggregating these raw features to ensure numerical stability. Extensive experiments demonstrate that BRACE achieves state-of-the-art quality-efficiency trade-offs across various DiT architectures with negligible computational overhead.

Sources

Related papers