How Far Can 5,500 Hours of Driving Take You? A Scaling Law Analysis of Video Diffusion Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "How Far Can 5,500 Hours of Driving Take You? A Scaling Law Analysis of Video Diffusion Models".
Tom: We present a systematic scaling-law study of video diffusion models trained from scratch on driving data: a family of models from 1M to 9B parameters,
Jane: First, who's behind it and why it matters.
Paper discussion segment 1: Tom: So, focusing on that five thousand five hundred hours of driving video data mentioned in the paper, one of the most interesting points they highlight is that repeating the same clips over and over during training epochs doesn't hurt performance as much as we might think.
Jane: That’s a really neat observation because it means that the five thousand five hundred hours isn't yet a hard bottleneck for these diffusion models when you run enough training steps. It suggests we don't have to worry about exhausting the dataset too quickly just by repeating things over and over.
Lu: That’s wild because I was expecting the data constraint to be a major hurdle, but it seems that the specific way they are resampling noise and timesteps in flow matching allows them to effectively reuse the existing corpus without saturation. It’s like finding a loophole in how we think about data limitations.
Meng: That’s good news for practical implementation because it means we don't need massive, curated datasets right away; we can leverage what we have and focus on training efficiency instead of just just getting more hours.
Lalam: For me, this confirms that our initial investment in driving data collection is still valuable, but it just dictates the method of using that data—how we structure the learning process matters more than just hoarding hours.
Tom: So, if they are suggesting we can reuse the existing driving footage without immediate saturation by repeating it within training epochs, what does that actually imply about how we should structure our long-term training plans for these systems?
Jane: It implies we should be comfortable running longer training cycles on the available data before switching gears entirely. We can treat the corpus as more dynamic than a static set of fixed points.
Lu: That’s wild because it changes the whole paradigm; it suggests that data isn't just a resource to be consumed, but something that can be actively manipulated in novel ways during training.
Meng: I’m curious if these scaling laws suggest any immediate practical constraints on our current GPU allocation strategies for training new driving world models, especially when we're trying to fit these models onto our existing hardware.
Lalam: From my view, the implications here are huge because it gives us a data-driven decision-making process for resource allocation when building foundational driving AI systems that need to be reliable everywhere.
Paper discussion segment 2: Tom: Moving on to the core of "How Far Can five thousand five hundred Hours of Driving Take You? A Scaling Law Analysis of Video Diffusion Models," the paper really lays out a framework for understanding the trade-offs between model size, training time, and compute power.
Jane: Exactly! It’s not just about getting better results; it’s about understanding the physics of learning—how model size, training time, and compute power all balance each other out according to these scaling laws, which gives us a much better way to think about training budgets moving forward.
Lu: I'm just blown away by the mathematical consistency of these scaling laws; they hold up across different axes, and the extrapolation is really wild—it suggests that even with a fixed five thousand five hundred hours of driving footage, we have a very reliable prediction for what an enormous model could realistically achieve. That kind of predictive power is incredible.
Meng: From my side, it confirms that even with limited driving data, if we scale up to the largest possible model size, we still have a clear path to world-class generation quality on nuScenes benchmarks; it gives us confidence in scaling up our architecture decisions for the future.
Lalam: For me, this work provides the mathematical backbone for building truly robust driving AI systems; it shows us exactly how to build a foundation that will actually perform well in real-world scenarios and make driving AI more reliable and capable across all cultures.
Tom: That predictive power is huge! So, if we look at the summary of "How Far Can five thousand five hundred Hours of Driving Take You? A Scaling Law Analysis of Video Diffusion Models," what's the big picture they’re telling us about the overall performance ceiling?
Jane: The big picture is that the asymptotic validation loss follows consistent power laws across model size and training exposure, which answers a huge question for us: whether compute is better spent on longer training or on a larger model.
Lu: They found that loss improves much faster with training exposure than with model size in terms of getting closer to the minimum, but the asymptotic loss keeps falling as capacity grows, meaning larger models still get you lower final numbers.
Meng: That’s interesting for our practical needs; it means we can't just pick one variable and ignore the other when designing a training pipeline; we have to consider both model architecture and how long we plan to train.
Lalam: It shows us exactly how to build a foundation for driving AI that isn't just good in one area but reliable across all real-world scenarios by balancing these factors intelligently.
Paper discussion segment 3: Tom: Now let’s talk about the improvements suggested by "How Far Can five thousand five hundred Hours of Driving Take You? A Scaling Law Analysis of Video Diffusion Models." They aren't just reporting results; they are giving us concrete ways to optimize our training budgets.
Jane: Right, Tom! It really boils down to understanding the trade-offs between model size, training time, and compute budget to get the best performance out of a fixed amount of driving data. It’s a practical guide for budgeting our resources better.
Lu: I just keep thinking about how robust those scaling laws are across all three axes; that kind of mathematical consistency is what makes this paper so powerful for future research in generative AI, especially when we look at how these models might be structured beyond just diffusion architectures.
Meng: From an engineering standpoint, it confirms that even with limited data, if we scale up to the largest possible model size, we still have a clear path to world-class generation quality on nuScenes benchmarks; it gives us confidence in scaling up our architecture decisions for the future.
Lalam: It shows us exactly how to build a foundation for driving AI that's not just good in one area but reliable across all real-world scenarios, showing us that reliability comes from careful resource allocation, not just brute force.
Tom: Absolutely! Huge thanks to Jane for breaking down those concepts so clearly, and Lu, Meng, Lalam—your insights on the broader implications are fantastic. We'll keep this paper top of mind as we move forward.
Jane: It’s been a great chat, Tom; we really got a handle on how these scaling laws work without getting bogged down in too much math.
Lu: Speaking of building foundations, I think this scaling analysis opens up so many creative possibilities for how we might structure future generative architectures beyond just diffusion models.
Meng: I’m still focused on the practical side; I'm curious if these laws suggest any immediate constraints on our current GPU allocation strategies for training new driving world models?
Lalam: From my view, the implications here are huge because it gives us a data-driven decision-making process for resource allocation when building foundational driving AI systems that need to be reliable everywhere.
Conclusion: Tom: So we've just wrapped up our deep dive into "How Far Can five thousand five hundred Hours of Driving Take You? A Scaling Law Analysis of Video Diffusion Models," and what a ride it was! We’ve seen how they use scaling laws to map out the optimal path for training.
Jane: It really boiled down to understanding how model size, training time, and compute power all balance each other out according to those scaling laws, which gives us a much better way to think about training budgets moving forward.
Lu: I just keep thinking about how robust those scaling laws are across all three axes; that kind of mathematical consistency is what makes this paper so powerful for future research in generative AI.
Meng: From an engineering standpoint, it confirms that even with limited data, if we scale up to the largest possible model size, we still have a clear path to world-class generation quality on nuScenes benchmarks.
Lalam: It shows us exactly how to build a foundation for driving AI that's not just good in one area but reliable across all real-world scenarios.
Tom: Absolutely! Huge thanks to Jane for breaking down those concepts so clearly, and Lu, Meng, Lalam—your insights on the broader implications are fantastic. We’ll be back next time!
Jane: It’s been a great chat, Tom; we really got a handle on how these scaling laws work without getting bogged down in too much math.
Lu: Speaking of building foundations, I think this scaling analysis opens up so many creative possibilities for how we might structure future generative architectures beyond just diffusion models.
Meng: I’m still focused on the practical side; I'm curious if these laws suggest any immediate constraints on our current GPU allocation strategies for training new driving world models?
Lalam: From my view, the implications here are huge because it gives us a data-driven decision-making process for resource allocation when building foundational driving AI systems that need to be reliable everywhere.
Tom: We'll have to keep watching these scaling laws closely as we move into training more complex tasks, and Jane, you were spot on about the two-pronged approach for optimization.
Jane: Definitely, and next time we'll be looking at something totally different, so stay tuned!
Victor Besnier, Anh-Quan Cao, Elias Ramzi, Spyros Gidaris, Tuan-Hung Vu, Andrei Bursuc, Eloi Zablocki, Matthieu Cord
Sorbonne Université
cs.CV
Submitted: 2026-08-28
Updated: 2026-09-23
Code: https://github.com/valeoai/VATIX
Project page: https://valeoai.github.io/VATIX
Importance score: 92/100
The gist: We present a systematic scaling-law study of video diffusion models trained from scratch on driving data: a family of models from 1M to 9B parameters, trained at different exposures on up to 5,500
Key concepts
- Scaling Law Analysis
- This involves studying the relationship between model size, training time, and compute power to understand how performance scales. The paper uses these laws to predict what an enormous model can realistically achieve with a fixed amount of driving data.
- Training Exposure
- This refers to how much training data or time a model is exposed during the training process. The discussion noted that loss improves faster with training exposure than with model size, but larger models still yield lower final loss values.
- Asymptotic Validation Loss
- This is the long-term performance level of a model as it continues to train. The paper found that this loss follows consistent power laws across different model sizes and training exposures, helping determine whether compute should be spent on more training or a larger architecture.
Terminology
Summary
We present a systematic scaling-law study of video diffusion models trained from scratch on driving data: a family of models from 1M to 9B parameters, trained at different exposures on up to 5,500 hours of driving. Validation loss follows consistent power laws in both model size and training exposure, answering the questions that shape a training budget: whether compute is better spent on longer training or on a larger model, and whether more data is needed. Loss improves much faster with training exposure than with model size, making longer training the most effective way to improve a fixed model under limited compute. However, larger models continue to achieve lower asymptotic loss, so compute-optimal scaling still favors increasing model size when sufficient compute and data are available.
Guided by these laws, we train a 9B-parameter model, to our knowledge the largest video diffusion model trained from scratch on driving data: it sets a new open-source state of the art for driving video generation, as measured on nuScenes.
The study investigates how far a fixed dataset, 5,500 hours of driving video, can take a video diffusion model. To answer this, we need to know exactly where each unit of compute is best spent: is it better to build a larger model, or to train for more steps? Two differences separate our setting from these studies [on LLMs]: First, the optimization objective relies on denoising or flow matching rather than next-token prediction. Second, we operate in the data-constrained regime, unlike LLMs whose text corpora are practically unbounded. In Chinchilla-style scaling [13], growing the data axis means sourcing more unique tokens; here, the driving corpus is fixed at 5,500 hours: far larger than any public driving dataset, yet still small enough that repeated epochs over the same clips are unavoidable. Training exposure, the total number of samples seen, unique or repeated alike, thus becomes the primary scalable resource. We fit a simple asymptotic power law, L(x) = L0 + A x−α, to the validation loss along three complementary axes: model scaling (x:= N), training scaling (x:= D), and compute scaling (x:= C).
The study proceeds in two steps: first we fit the scaling laws, then we spend the budget where they point. The fits are accurate across more than 200 runs spanning 1.6M to 1.1B parameters, and three findings follow:
(1) The training-exposure exponent is far steeper than the model-size exponent, so longer training is the fastest way to improve a fixed model, while the asymptotic loss keeps falling with capacity, so compute-optimal budgets still favor larger models.
(2) The fixed corpus does not break the laws: within the fewer than five epochs we run, repeating the data behaves like fresh samples [20, 26], so the 5,500 hours are not yet the bottleneck. We hypothesize this is because flow matching resamples noise and timestep at every pass, unlike the fixed targets of autoregressive training.
(3) The laws extrapolate: a single 9B-parameter model, to our knowledge the largest open-source video diffusion model trained from scratch on driving data, reaches a validation loss predicted before training to within 3.6% of a law extrapolated 8× beyond the largest fitted model, the lowest loss among all our models.
The analysis is structured around three axes:
Model scaling (x=N):
We investigate how the validation loss scales with model size N under fixed training exposure (10M samples). The fitted law is:
L(N) = 0.0595 + 0.1041 · N−0.2125, where the fitted exponent is α N = 0.2125. Extrapolating to a 9B-parameter model trained with the same exposure (D ≈ 10 7 samples) predicts a validation loss of L(9B) ≈ 0.0746.
Training scaling (x=D):
We investigate how validation loss scales with training exposure D while keeping model size fixed. The fitted law is L(D) = L0 + AD−αD across model sizes, where the convergence exponent α D shows no systematic trend across scale (0.74 on average, spread 0.66–0.84). We estimate that training up to 1.5 × 10 7 samples captures most of the achievable gain: with the mean fitted coefficients of Fig. 3a (A ≈ 0.005, αD ≈ 0.74), the loss reduction attainable beyond this exposure is below 0.004.
Compute scaling (x=C):
We investigate scaling under a fixed compute training budget C (TFLOPs). The fitted law is:
L(C) = 0.0522 + 0.5955 · C−0.15443. Combining the predicted optimal compute allocation with the planned 9B architecture gives, for a model trained on approximately 10 7 samples (1.438 × 10 9 TFLOPs), an estimated validation loss of L(1.438 × 10 9) ≈ 0.0753.
Guided by the scaling laws, we train our 9B model and evaluate it against prominent open-source driving world models [7, 9]. The resulting model sets a new open-source state of the art for driving video generation on nuScenes. Validation loss at 9B scale is predicted to be 0.0753 before training, and after training up to 1.2 × 10 7 samples (approximately 2 epochs), the model reaches a validation loss of 0.0781, the lowest among all evaluated models, demonstrating continued improvements when scaling from 1B to 9B parameters. This closely matches the extrapolated prediction, with only a 3.6% relative error. The post-hoc fitted asymptotic loss (L 0 ≈ 0.0748) sits within 1% of the a-priori compute-law prediction of 0.0753, providing an independent confirmation of the extrapolated estimate.
The study concludes that across more than 200 training runs spanning 1.6M to 1.1B parameters, three findings emerge: (1) the training-exposure exponent is far steeper than the model-size exponent, so longer training is the fastest way to improve a fixed model, while compute-optimal budgets still favor larger models as capacity keeps lowering the asymptotic loss; (2) the fixed corpus does not break the laws: repeated data behaves like fresh samples, so the 5,500 hours are not yet the bottleneck; and (3) the laws extrapolate: a 9B-parameter model reaches a validation loss predicted before training to within 3.6% of a law extrapolated 8× beyond the largest fitted model. The main limitation is the scale gap between the models used for fitting and the flagship: intermediate scales (e.g., 3B–4B) would tighten extrapolated estimates.
The resulting model sets a new open-source state of the art for driving video generation, as measured on nuScenes. We further post-finetune it with pseudo-trajectory conditioning to enable action control and fair comparison with baselines, which all rely on trajectory conditioning. The 9B model reaches a loss comparable to the fitted asymptote of the 1.1B model (L 0=0.078, Fig. 3a), a level that the smaller 1.1B model would require more than an order of magnitude additional training exposure to approach, in contrast to the 9B model reaching this regime within 1.2 × 10 7 samples while remaining far from its own asymptote (Fig. 6). The post-training fine-tuning on nuScenes achieves the best performance on both image and video-based perceptual metrics compared to prior work like GEM and Vista, which use parameter-efficient adaptation.
The main limitation is the scale gap between the models used for fitting and the flagship: intermediate scales (e.g., 3B–4B) would tighten extrapolated estimates. The main limitation is that the scaling laws assume a uniform learning rate; since the 9B model required adjusting the learning rate during training for stability, a hyperparameter grid search could further tighten our predictions.
The study answers its title question: 5,500 hours of driving take a from-scratch video diffusion model to a new open-source state-of-the-art for driving video generation, as measured on nuScenes. The main limitation is the scale gap between the models used for fitting and the flagship: intermediate scales (e.g., 3B–4B) would tighten extrapolated estimates.
The authors hope these scaling laws provide a practical basis for training future driving world models. This work answers its title question: 5,500 hours of driving take a from-scratch video diffusion model to a new open-source state-of-the-art for driving video generation, as measured on nuScenes. Across more than 200 training runs spanning 1.6M to 1.1B parameters, three findings emerge: (1) the trainingexposure exponent is far steeper than the model-size exponent, so longer training is the fastest way to improve a fixed model, while compute-optimal budgets still favor larger models as capacity keeps lowering the asymptotic loss; (2) the fixed corpus does not break the laws: repeated data behaves like fresh samples, so the 5,500 hours are not yet the bottleneck; and (3) the laws extrapolate: a 9B-parameter model, to our knowledge the largest open-source video diffusion model trained from scratch on driving data, reaches a validation loss predicted before training to within 3.6% of a law extrapolated 8× beyond the largest fitted model, the lowest loss among all our models. The main limitation is the scale gap between the models used for fitting and the flagship: intermediate scales (e.g., 3B–4B) would tighten extrapolated estimates.
The resulting model sets a new open-source state of the art for driving video generation on nuScenes. We further post-finetune it with pseudo-trajectory conditioning to enable action control and fair comparison with baselines, which all rely on trajectory conditioning. The 9B model reaches a loss comparable to the fitted asymptote of the 1.1B model (L 0=0.078, Fig. 3a), a level that the smaller 1.1B model would require more than an order of magnitude additional training exposure to approach, in contrast to the 9B model reaching this regime within 1.2 × 10 7 samples while remaining far from its own asymptote (Fig. 6). The post-training fine-tuning on nuScenes achieves the best performance on both image and video-based perceptual metrics compared to prior work like GEM and Vista, which use parameter-efficient adaptation.
The study answers its title question: 5,500 hours of driving take a from-scratch video diffusion model to a new open-source state-of-the-art for driving video generation, as measured on nuScenes. Across more than 200 training runs spanning 1.6M to 1.1B parameters, three findings emerge: (1) the trainingexposure exponent is far steeper than the model-size exponent, so longer training is the fastest way to improve a fixed model, while compute-optimal budgets still favor larger models as capacity keeps lowering the asymptotic loss; (2) the fixed corpus does not break the laws: repeated data behaves like fresh samples, so the 5,500 hours are not yet the bottleneck; and (3) the laws extrapolate: a 9B-parameter model, to our knowledge the largest open-source video diffusion model trained from scratch on driving data, reaches a validation loss predicted before training to within 3.6% of a law extrapolated 8× beyond the largest fitted model, the lowest loss among all our models. The main limitation is the scale gap between the models used for fitting and the flagship: intermediate scales (e.g., 3B–4B) would tighten extrapolated estimates.
The study answers its title question: 5,500 hours of driving take a from-scratch video diffusion model to a new open-source state-of-the-art for driving video generation, as measured on nuScenes. Across more than 200 training runs spanning 1.6M to 1.1B parameters, three findings emerge: (1) the trainingexposure exponent is far steeper than the model-size exponent, so longer training is the fastest way to improve a fixed model, while compute-optimal budgets still favor larger models as capacity keeps lowering the
Improvements for AI systems
Here are the specific improvements that can be made to AI systems, based on the findings of this research:
-
The most effective way to improve a fixed video diffusion model under limited compute is by increasing the amount of training data (training exposure), rather than solely relying on increasing model size. The training-exposure exponent is significantly steeper than the model-size exponent, indicating that longer training steps yield faster improvements in loss reduction for a given budget.
-
For achieving lower asymptotic validation loss, computational resources should be allocated toward larger models, even under fixed data constraints, as larger capacities continue to drive the loss down asymptotically.
-
The fixed dataset size of 5,500 hours is not yet a bottleneck for diffusion models when training exposures are sufficient (i.e., more than a few epochs). The use of flow matching—which resamples noise and timesteps at every pass—allows repeated data to behave like fresh samples, suggesting that the model can effectively leverage the existing corpus without immediate saturation.
-
A compute-optimal training budget is achieved by jointly scaling model size and training exposure. The compute scaling law suggests that for a target large model (like 9B), an optimal training exposure of approximately 107 samples (about 2 epochs) is sufficient to reach a validation loss close to the theoretical minimum predicted by extrapolating the scaling laws.
-
The derived models, especially the trained 9B parameter model, set a new open-source state-of-the-art for driving video generation on benchmarks like nuScenes.
-
The resulting 9B model can generate high-fidelity driving videos (as measured by FID and FVD metrics) that exhibit superior temporal coherence and long-range motion consistency compared to smaller models, especially when augmented with trajectory conditioning.
This improved AI system can perform the following:
-
Generate highly realistic, physically consistent video sequences of autonomous driving scenarios from scratch using only 5,500 hours of driving footage.
-
Produce world models capable of predicting future driving scenes and actions by leveraging learned representations from this fixed corpus, making it robust for simulation and planning tasks.
-
Perform sophisticated trajectory control and scenario generation by incorporating pseudo-trajectories (pseudo-trajectories extracted via OccAny) directly into the generation process, allowing the AI to generate videos conditioned on specific maneuvers (e.g., hard braking or sharp turns).
-
Serve as a competitive, open-source foundation model for driving video generation that surpasses existing proprietary models in terms of perceptual quality and temporal fidelity across various driving modalities.
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models