Do AI weather models miss extremes?
summary
The gist
- Wind (Track A, 6–48 h): "Jua EPT-2.1 Europa leads all-conditions wind (+8.4%)" and holds "+8.1 to +8.7% wind skill at every lead from 6 to 48 h." At gale-force (>P95), "DWD ICON Global has the
This episode discusses
- Do AI weather models miss extremes? · Paper Radio
- MAUSAM: An Observations-focused assessment of Global AI Weather Prediction Models During the South Asian Monsoon
- WeatherReal: A Benchmark Based on In-Situ Observations for Evaluating Weather Models
- AIFS -- ECMWF's data-driven forecasting system
- Generative AI for fast and accurate statistical computation of fluids
- EPT-2 Technical Report
- Universal Diffusion-Based Probabilistic Downscaling
- FourCastNet: A Global Data-driven High-resolution Weather Model using Adaptive Fourier Neural Operators
- Fixing the Double Penalty in Data-Driven Weather Forecasting Through a Modified Spherical Harmonic Loss Function
The paper
Do AI weather models miss extremes? · Read on arXiv
Marvin Vincent Gabler, Roberto Molinaro, Niall Siegenheim, Henry Martin, Mark Frey, Niels Poulsen, Philipp Seitz, Olivier Lam
Jua AI AG
First-generation AI weather models are often reported to underperform at extremes, mostly in reanalysis-based evaluations of deterministic regression systems. We verify eleven physical and AI forecast systems against European synoptic, solar, and rain-gauge stations over ten months for 10 m wind, 2 m temperature, hourly shortwave accumulation, and hourly precipitation, scoring mean absolute error (MAE) against ECMWF IFS in ERA5 1991-2020 climatological regimes. Among these systems, AI models do not show a uniform relative-skill deficit in the tails. Jua EPT-2.1 Europa leads all-conditions wind (+8.4%), while Jua EPT-2 HRRR leads temperature overall (+12.1%) and in the heat regime (+19.6 +/- 2.2%). EPT-2.1 Europa and DWD ICON Global lead at gale-force wind. Jua EPT-2.1 Helios leads solar overall (+10.2 +/- 1.7%), in overcast conditions (+16.4 +/- 3.4%), and in the clear-sky tail (+24.8 +/- 5.4%). For precipitation, three Jua models gain 14-15% at moderate intensity and 9-11% at P75-P95; EPT-2 Reasoning remains ahead above P95 (+1.7 +/- 0.5%). Failures are model-specific: ECMWF AIFS loses 4.9 +/- 2.0% in the heat tail, while NOAA GFS loses 22.8 +/- 2.0% there. Every model, including numerical weather prediction systems, shows a shared conditional bias toward the centre of the observed distribution, with an inter-model spread several times smaller than the shared signal. Missing relative skill at extremes is therefore not a property of AI weather models as a class, but of particular AI and physical models.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Do AI weather models miss extremes?".
Jane: The paper was written by Marvin Vincent Gabler, Roberto Molinaro, Niall Siegenheim, Henry Martin, Mark Frey et al. from Jua AI AG.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that's been making waves in the weather forecasting world, and it's got a title that asks a pretty direct question: "Do AI weather models miss extremes?" That's the paper, evidence from ten months of verification against European weather stations.
Jane: And Tom, I have to say, that title is doing a lot of work. For years now, the knock on AI weather models has been that they're great at the average day, but when a storm really kicks up or a heatwave really bites, they go soft. They smooth things out. This paper is essentially putting that claim on trial.
Tom: Exactly. And the authors are from Jua AI, a company that builds these models, so you might expect some bias. But they're not just defending themselves. They brought in eleven different forecast systems, including the big physics-based ones from ECMWF, NOAA, and DWD, and they pitted them all against real weather station data across Europe.
Jane: Real station data, that's the key part. A lot of earlier studies checked AI models against reanalysis, which is basically a best-guess reconstruction of the weather. It's smooth, it's gridded, and it can hide problems. Here, they're comparing against actual thermometers, wind gauges, and rain buckets on the ground. That's a much tougher test.
Tom: Right, and the results are genuinely surprising. The headline is that AI models don't uniformly fail at extremes. Some do, but some beat the physics-based reference model, ECMWF's IFS, by double digits in the hottest temperatures and the windiest conditions.
Jane: And that's a big deal, because the conventional wisdom was that if you train a model to minimize average error, it will inevitably play it safe. It'll predict the middle of the range. This paper shows that's not a law of nature. It depends on how you build the model.
Tom: So we've got a title that asks a question, and the answer is turning out to be, "It depends on which AI model you're talking about." That's a much more interesting answer than a simple yes or no.
Jane: And it means the conversation about AI in weather has to shift. It's not about whether AI can handle extremes. It's about which architectures and training methods actually preserve the tails of the distribution.
Tom: Stay with us, because in the next segment we're going to dig into the actual numbers and see which models came out on top, and which ones really did fall flat.
Summary: Tom: Welcome back. We're still on "Do AI weather models miss extremes?" and Jane, we left off with the big picture. Now let's get into the specifics, because the numbers here are pretty wild.
Jane: They really are. So they ran this for ten months, from September two thousand twenty-five through June two thousand twenty-six and they looked at four things: wind speed, temperature, solar radiation, and precipitation. For each one, they sorted the observations into buckets based on a thirty-year climate baseline. So you've got your calm days, your gale-force days, your heatwaves, your heavy rain.
Tom: And the standout result for me is temperature. Jua's regional model, EPT-two HRRR, beat the ECMWF reference by about twelve percent overall, and in the heat regime, the top five percent of temperatures, it was nearly twenty percent better. That's not a small edge.
Jane: And it's not just the Jua models. The paper shows that some AI models are genuinely bad at extremes. ECMWF's own AI model, AIFS, lost about five percent in the heat tail. And NOAA's GFS, which is a physics-based model, lost nearly twenty-three percent in the heat. So the failures are spread across both camps.
Tom: That's the crucial finding. It's not AI versus physics. It's good models versus bad models, and the bad ones come from both families. The paper even shows that every single model, including the physics ones, has a tendency to under-forecast the most extreme events. They all drift toward the middle.
Jane: Right, that shared bias is fascinating. After they corrected for average bias, every system still over-forecast the calm conditions and under-forecast the gales. So there's a common failure mode. But the size of that failure varies hugely from model to model.
Tom: And for solar, the results are even more dramatic. Jua's solar-specialized model, Helios, was over twenty-four percent better than the reference in clear-sky conditions. But in the middle range, where clouds are doing their thing, it was actually worse than the other AI models. So it traded accuracy in the bulk for accuracy in the tails.
Jane: That trade-off is a really important nuance. You can't have a model that's great at everything. The paper shows that specialization matters. A model built for solar forecasting will beat a generalist on sunny days, but it might struggle when the sky is mixed.
Tom: And precipitation tells a similar story. The AI models were strong in moderate rain, gaining fourteen to fifteen percent, but at the very heaviest rainfall, the advantage shrank to almost nothing. Only one model, EPT-two Reasoning, stayed clearly ahead.
Jane: So the answer to the title question is nuanced. AI models don't miss extremes as a class. But they also haven't fully solved the hardest cases, like the most intense downpours.
Tom: And that's where we're going next. We're going to talk about what the paper suggests we should do about it, and what it means for the future of forecasting.
Improvements: Tom: Back on the show, still with "Do AI weather models miss extremes?" and Jane, we've established that the results are mixed. So what does this paper actually suggest we should do differently?
Jane: The biggest takeaway for me is that we need to stop treating "AI weather model" as if it's one thing. The paper compares regression models, which are trained to minimize average error, against generative models, which learn to produce realistic weather patterns. And the generative ones, like the Jua EPT-two HRRR and Europa, are the ones that hold up better in the extremes.
Tom: Right, and that points to a concrete improvement. If you want a model that handles heatwaves and gales, you should probably be looking at generative architectures, not the older regression approach. The paper even cites earlier work showing that regression training actively smooths out the sharp gradients that you need for extreme events.
Jane: And there's another improvement hiding in the methodology. They used a rolling bias correction, where they look at the previous four weeks of errors and subtract that bias from the next week's forecast. That's a simple trick, and it helped all models, but it didn't fix the tail problem. The bias correction only shifts the average, it doesn't change the shape of the errors.
Tom: So the improvement isn't just in the models themselves. It's in how we verify them. This paper uses station data, not reanalysis, and it uses climatological thresholds to define extremes. That's a much more honest test. And the authors are clear that this is the kind of evaluation we need more of.
Jane: And they also point out that the shared center-seeking bias, where every model under-forecasts the biggest events, is a property of the whole class. That means fixing it might require a fundamental change in how we train these systems, not just tweaking the loss function.
Tom: But here's the thing that excites me. The paper shows that some models already beat the physics-based reference in the tails. So the improvements are happening. We're not waiting for a breakthrough. The breakthrough is already here, it's just not uniform across all models.
Jane: Exactly. And that's why this paper is so valuable. It gives us a roadmap. It tells us which architectures are working, which ones are falling behind, and where the remaining weaknesses are. The heavy precipitation tail is still a problem for everyone, and that's where the next generation of models needs to focus.
Tom: So the improvement is partly architectural, partly in evaluation, and partly in just paying attention to the tails during training. And that brings us to the bigger question of what this means for the world, which we'll get to in our final segment.
Conclusion: Tom: And we're wrapping up our look at "Do AI weather models miss extremes?" and Jane, I think we've got a clear picture now.
Jane: We do. The paper's answer is that AI models don't uniformly miss extremes, but the failures are real and they're model-specific. Some AI models are excellent in the heat and wind tails, beating the physics-based reference by double digits. Others, including some physics models, fall well behind.
Tom: And the shared bias toward the center of the distribution is still there for every model. That's the thing that keeps me up at night. Even the best models we have today will under-forecast the most extreme events. They'll tell you it's going to be very hot, but not quite as hot as it actually gets.
Jane: But the paper also gives us a path forward. Generative models are showing real promise, and the verification methodology here, using station data and climatological thresholds, is something the whole field should adopt. If we can't measure the tails honestly, we can't improve them.
Tom: And that matters beyond just forecasting. Think about what this means for energy planning, for agriculture, for disaster preparedness. If we can get even a few percent better at predicting the hottest days and the windiest hours, that translates into real savings and real safety.
Jane: Absolutely. And the paper is honest about its limits. It's ten months, it's Europe, it's a specific set of models. But it's a huge step forward in how we evaluate these systems. We're no longer guessing based on reanalysis. We're checking against the ground truth.
Tom: So goodbye to "Do AI weather models miss extremes?" It's been a great discussion, and we're taking a lot of hope from it. The next generation of models is already showing that extremes aren't out of reach.
Jane: And we'll be back soon with the next paper. Until then, keep an eye on the sky, and maybe check what your favorite forecast model is doing on the hottest day of the year.
Tom: Thanks for listening, everyone. See you next time.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language