Do AI weather models miss extremes?
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Do AI weather models miss extremes?".
Jane: The paper was written by Marvin Vincent Gabler, Roberto Molinaro, Niall Siegenheim, Henry Martin, Mark Frey et al. from Jua AI AG.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that's been making waves in the weather forecasting world, and it's got a title that asks a pretty direct question: "Do AI weather models miss extremes?" That's the paper, evidence from ten months of verification against European weather stations.
Jane: And Tom, I have to say, that title is doing a lot of work. For years now, the knock on AI weather models has been that they're great at the average day, but when a storm really kicks up or a heatwave really bites, they go soft. They smooth things out. This paper is essentially putting that claim on trial.
Tom: Exactly. And the authors are from Jua AI, a company that builds these models, so you might expect some bias. But they're not just defending themselves. They brought in eleven different forecast systems, including the big physics-based ones from ECMWF, NOAA, and DWD, and they pitted them all against real weather station data across Europe.
Jane: Real station data, that's the key part. A lot of earlier studies checked AI models against reanalysis, which is basically a best-guess reconstruction of the weather. It's smooth, it's gridded, and it can hide problems. Here, they're comparing against actual thermometers, wind gauges, and rain buckets on the ground. That's a much tougher test.
Tom: Right, and the results are genuinely surprising. The headline is that AI models don't uniformly fail at extremes. Some do, but some beat the physics-based reference model, ECMWF's IFS, by double digits in the hottest temperatures and the windiest conditions.
Jane: And that's a big deal, because the conventional wisdom was that if you train a model to minimize average error, it will inevitably play it safe. It'll predict the middle of the range. This paper shows that's not a law of nature. It depends on how you build the model.
Tom: So we've got a title that asks a question, and the answer is turning out to be, "It depends on which AI model you're talking about." That's a much more interesting answer than a simple yes or no.
Jane: And it means the conversation about AI in weather has to shift. It's not about whether AI can handle extremes. It's about which architectures and training methods actually preserve the tails of the distribution.
Tom: Stay with us, because in the next segment we're going to dig into the actual numbers and see which models came out on top, and which ones really did fall flat.
Summary: Tom: Welcome back. We're still on "Do AI weather models miss extremes?" and Jane, we left off with the big picture. Now let's get into the specifics, because the numbers here are pretty wild.
Jane: They really are. So they ran this for ten months, from September two thousand twenty-five through June two thousand twenty-six and they looked at four things: wind speed, temperature, solar radiation, and precipitation. For each one, they sorted the observations into buckets based on a thirty-year climate baseline. So you've got your calm days, your gale-force days, your heatwaves, your heavy rain.
Tom: And the standout result for me is temperature. Jua's regional model, EPT-two HRRR, beat the ECMWF reference by about twelve percent overall, and in the heat regime, the top five percent of temperatures, it was nearly twenty percent better. That's not a small edge.
Jane: And it's not just the Jua models. The paper shows that some AI models are genuinely bad at extremes. ECMWF's own AI model, AIFS, lost about five percent in the heat tail. And NOAA's GFS, which is a physics-based model, lost nearly twenty-three percent in the heat. So the failures are spread across both camps.
Tom: That's the crucial finding. It's not AI versus physics. It's good models versus bad models, and the bad ones come from both families. The paper even shows that every single model, including the physics ones, has a tendency to under-forecast the most extreme events. They all drift toward the middle.
Jane: Right, that shared bias is fascinating. After they corrected for average bias, every system still over-forecast the calm conditions and under-forecast the gales. So there's a common failure mode. But the size of that failure varies hugely from model to model.
Tom: And for solar, the results are even more dramatic. Jua's solar-specialized model, Helios, was over twenty-four percent better than the reference in clear-sky conditions. But in the middle range, where clouds are doing their thing, it was actually worse than the other AI models. So it traded accuracy in the bulk for accuracy in the tails.
Jane: That trade-off is a really important nuance. You can't have a model that's great at everything. The paper shows that specialization matters. A model built for solar forecasting will beat a generalist on sunny days, but it might struggle when the sky is mixed.
Tom: And precipitation tells a similar story. The AI models were strong in moderate rain, gaining fourteen to fifteen percent, but at the very heaviest rainfall, the advantage shrank to almost nothing. Only one model, EPT-two Reasoning, stayed clearly ahead.
Jane: So the answer to the title question is nuanced. AI models don't miss extremes as a class. But they also haven't fully solved the hardest cases, like the most intense downpours.
Tom: And that's where we're going next. We're going to talk about what the paper suggests we should do about it, and what it means for the future of forecasting.
Improvements: Tom: Back on the show, still with "Do AI weather models miss extremes?" and Jane, we've established that the results are mixed. So what does this paper actually suggest we should do differently?
Jane: The biggest takeaway for me is that we need to stop treating "AI weather model" as if it's one thing. The paper compares regression models, which are trained to minimize average error, against generative models, which learn to produce realistic weather patterns. And the generative ones, like the Jua EPT-two HRRR and Europa, are the ones that hold up better in the extremes.
Tom: Right, and that points to a concrete improvement. If you want a model that handles heatwaves and gales, you should probably be looking at generative architectures, not the older regression approach. The paper even cites earlier work showing that regression training actively smooths out the sharp gradients that you need for extreme events.
Jane: And there's another improvement hiding in the methodology. They used a rolling bias correction, where they look at the previous four weeks of errors and subtract that bias from the next week's forecast. That's a simple trick, and it helped all models, but it didn't fix the tail problem. The bias correction only shifts the average, it doesn't change the shape of the errors.
Tom: So the improvement isn't just in the models themselves. It's in how we verify them. This paper uses station data, not reanalysis, and it uses climatological thresholds to define extremes. That's a much more honest test. And the authors are clear that this is the kind of evaluation we need more of.
Jane: And they also point out that the shared center-seeking bias, where every model under-forecasts the biggest events, is a property of the whole class. That means fixing it might require a fundamental change in how we train these systems, not just tweaking the loss function.
Tom: But here's the thing that excites me. The paper shows that some models already beat the physics-based reference in the tails. So the improvements are happening. We're not waiting for a breakthrough. The breakthrough is already here, it's just not uniform across all models.
Jane: Exactly. And that's why this paper is so valuable. It gives us a roadmap. It tells us which architectures are working, which ones are falling behind, and where the remaining weaknesses are. The heavy precipitation tail is still a problem for everyone, and that's where the next generation of models needs to focus.
Tom: So the improvement is partly architectural, partly in evaluation, and partly in just paying attention to the tails during training. And that brings us to the bigger question of what this means for the world, which we'll get to in our final segment.
Conclusion: Tom: And we're wrapping up our look at "Do AI weather models miss extremes?" and Jane, I think we've got a clear picture now.
Jane: We do. The paper's answer is that AI models don't uniformly miss extremes, but the failures are real and they're model-specific. Some AI models are excellent in the heat and wind tails, beating the physics-based reference by double digits. Others, including some physics models, fall well behind.
Tom: And the shared bias toward the center of the distribution is still there for every model. That's the thing that keeps me up at night. Even the best models we have today will under-forecast the most extreme events. They'll tell you it's going to be very hot, but not quite as hot as it actually gets.
Jane: But the paper also gives us a path forward. Generative models are showing real promise, and the verification methodology here, using station data and climatological thresholds, is something the whole field should adopt. If we can't measure the tails honestly, we can't improve them.
Tom: And that matters beyond just forecasting. Think about what this means for energy planning, for agriculture, for disaster preparedness. If we can get even a few percent better at predicting the hottest days and the windiest hours, that translates into real savings and real safety.
Jane: Absolutely. And the paper is honest about its limits. It's ten months, it's Europe, it's a specific set of models. But it's a huge step forward in how we evaluate these systems. We're no longer guessing based on reanalysis. We're checking against the ground truth.
Tom: So goodbye to "Do AI weather models miss extremes?" It's been a great discussion, and we're taking a lot of hope from it. The next generation of models is already showing that extremes aren't out of reach.
Jane: And we'll be back soon with the next paper. Until then, keep an eye on the sky, and maybe check what your favorite forecast model is doing on the hottest day of the year.
Tom: Thanks for listening, everyone. See you next time.
Marvin Vincent Gabler, Roberto Molinaro, Niall Siegenheim, Henry Martin, Mark Frey, Niels Poulsen, Philipp Seitz, Olivier Lam
Jua AI AG
physics.ao-ph, cs.AI, cs.LG
Submitted: 2026-07-31
Updated: 2026-08-12
Code: https://github.com/juaAI/aiweather-extremes
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 72/100
The gist: - Wind (Track A, 6–48 h): "Jua EPT-2.1 Europa leads all-conditions wind (+8.4%)" and holds "+8.1 to +8.7% wind skill at every lead from 6 to 48 h." At gale-force (>P95), "DWD ICON Global has the
Terminology
Summary
Summary
The paper evaluates eleven physical and AI weather forecast systems against European synoptic, solar-radiation, and rain-gauge stations over ten months (1 September 2025 to 30 June 2026) for 10 m wind, 2 m temperature, hourly shortwave accumulation, and hourly precipitation. Skill is scored as MAE relative to the ECMWF IFS reference, using observed-value regimes defined by ERA5 1991–2020 climatological percentiles (P5, P25, P75, P95) per station, calendar day, and UTC hour. Precipitation uses wet-hour regimes (dry, P95) because the symmetric percentile construction degenerates on zero-inflated data. All results are debiased (additive for wind, temperature, solar; multiplicative for precipitation). The central finding is that AI models do not show a uniform relative-skill deficit in the tails
and that Missing relative skill at extremes is therefore not a property of AI weather models as a class, but of particular AI and physical models.
Key results by variable:
-
Wind (Track A, 6–48 h):
Jua EPT-2.1 Europa leads all-conditions wind (+8.4%)
and holds+8.1 to +8.7% wind skill at every lead from 6 to 48 h.
At gale-force (>P95),DWD ICON Global has the highest point estimate (+9.5 ± 1.2%), effectively tied with EPT-2.1 Europa (+9.0 ± 1.2%).
ECMWF AIFS is negative at every wind lead, from −6.6% at 6 h to −4.0% at 48 h, and loses −10.6 ± 0.6% at gale force. NOAA GFS is barely positive at gale (+0.6 ± 0.7%) and poor overall (−11.7%). -
Temperature (Track A, 6–48 h):
Jua EPT-2 HRRR leads temperature overall (+12.1%) and in the heat regime (+19.6 ± 2.2%)
versus EPT-2.1 Europa at +11.5% overall and +15.0 ± 1.7% in heat.ECMWF AIFS loses 4.9 ± 2.0% in the heat tail
despite near-neutral overall skill.NOAA GFS loses 22.8 ± 2.0% there
and is negative in every country at heat extremes. Jua EPT-2 Reasoning is negative at cold extremes (−6.9 ± 1.3%); Microsoft Aurora and Jua EPT-2e lose more than 20% in the cold tail. -
Solar (Track A, 1–48 h hourly):
Jua EPT-2.1 Helios leads solar overall (+10.2 ± 1.7%), in overcast conditions (+16.4 ± 3.4%), and in the clear-sky tail (+24.8 ± 5.4%).
However, in the typical irradiance regime (P25–P75),EPT-2.1 Helios is instead the weakest of the AI systems (−8.2 ± 1.1%) and EPT-2 Reasoning leads (+14.0 ± 0.7%); the specialised model buys tail accuracy at the cost of the bulk of the distribution.
Restricting to 1–12 h widens Helios's lead to +13.5 ± 1.9%. -
Precipitation (Track A, 1–48 h):
three Jua models gain 14–15% at moderate intensity and 9–11% at P75–P95
(EPT-2 HRRR +15.2 ± 2.2%, EPT-2e +14.8 ± 1.7%, EPT-2 Reasoning +14.0 ± 0.9% at P50–75). "EPT-2 Reasoning remains ahead at >P95 (+1.7 ± 0.5%); EPT-2.1 Europa is near IFS (+0.2 ± 0.6%), EPT-2 HRRR (−1.5 ± 0.5%) and EPT-2e (−1.9 ± 0.9%) trail, and GFS is −3.7 ± 1.2%. Categorical detection shows
every model under-detects heavy hours (frequency bias below one throughout, from 0.82 for ICON Global down to 0.29 for the ECMWF ENS mean)." Jua EPT-2.1 Europa is a model-specific exception: strong on dry mass (+21.0 ± 2.7%) but negative through light and moderate wet bands (−11.1 to −10.5%). -
Track B (March–June 2026, regional comparison with DWD ICON-EU):
EPT-2.1 Europa leads wind overall (+7.2 ± 0.4% versus ICON-EU +5.3 ± 0.5%)
; at gale wind they are level (+10.5 ± 1.4% and +11.5 ± 1.6%). For temperature, EPT-2.1 Europa has the highest overall point estimate (+13.2 ± 0.6%, versus ICON-EU +12.4 ± 0.8% and EPT-2 HRRR +12.0 ± 0.7%); all three overlap in the heat tail at about 17%. For solar, EPT-2.1 Helios leads overall (+12.6 ± 2.5%) and in both tails. For precipitation,EPT-2 HRRR is the only regional system with positive skill from light through high rain, reaching +15.6 ± 6.1% at P50–P75
; ICON-EU is positive overall (+3.5 ± 0.5%) but negative in each wet band; EPT-2.1 Europa repeats its dry-mass-positive/wet-negative signature.
The paper also reports a shared conditional bias: All evaluated systems overforecast low observed values and underforecast high observed values
after debiasing, with wind bias +0.8 to +1.2 m s−1 in calm and −2.2 to −2.9 m s−1 in gale; temperature +0.7 to +1.2 °C in severe cold and −0.3 to −0.7 °C in heat; precipitation +0.09 to +0.12 mm in light rain and −2.29 to −2.51 mm at >P95. The inter-model spread is several times smaller than the shared signal.
The paper concludes: The remaining failures at extremes are therefore model-specific against a shared centre-seeking bias, not a uniform property of AI weather models.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI weather models, along with what the improved systems can do:
-
Improvement: Replace pure MSE with a hybrid loss that up-weights errors in the climatological tails (e.g., >P95 and <P5). Use a weighting scheme like
w = 1 + α · I(observed in tail)with α tuned per variable (wind, temperature, solar, precipitation). -
What it does: Reduces the shared conditional bias where all models overforecast low values and underforecast high values (Figure 12). This directly targets the −2.2 to −2.9 m/s wind bias at gale force and the −2.3 to −2.5 mm precipitation bias at >P95.
-
Improvement: Add a secondary output head trained specifically on the >P95 and <P5 regimes, with a separate loss term. Use a multi-task architecture where the main head handles the bulk distribution and the extreme head is only activated when the input suggests a tail event (e.g., via a gating network).
-
What it does: Matches or beats the paper's best performers in the tails. For example, EPT-2 HRRR achieves +19.6% heat skill; a dedicated head should push this beyond +25% while maintaining the +12% all-conditions skill.
-
Improvement: Extend the rolling mean-bias correction (Section 3.5) to be regime-aware. Instead of a single additive shift, estimate a piecewise-linear correction per climatological percentile band (P5, P5–25, P25–75, P75–95, >P95) using the previous 4 weeks of station errors.
-
What it does: Removes the residual regime-dependent bias that the current debiasing leaves behind. For example, it can correct the −11.1% light-rain skill of EPT-2.1 Europa (Table 5) by adjusting the multiplicative factor separately for each wet-hour band.
-
Improvement: For precipitation, use a score-based diffusion model (like EPT-2 HRRR) but with a loss that emphasizes the >P95 wet-hour regime. Sample multiple members and take the ensemble mean, but also output a calibrated probability of exceeding the local P95 threshold.
-
What it does: Improves the heavy-tail MAE from +1.7% (EPT-2 Reasoning) to potentially +5–8%, and increases the frequency bias from 0.53–0.74 toward 1.0, reducing the under-detection shown in Figure 11.
-
Improvement: Split the solar model into two branches: one optimised for clear-sky (P75–P95, >P95) and one for overcast (P5–P25). Use a cloud-cover classifier to route each forecast through the appropriate branch. This is a direct fix for EPT-2.1 Helios's weakness in the typical regime (−8.2% at P25–75).
-
What it does: Achieves EPT-2.1 Helios's tail performance (+24.8% clear-sky, +16.4% overcast) while recovering the +14% typical-regime skill of EPT-2 Reasoning. Net effect: +15–20% all-conditions solar skill, up from +10.2%.
-
Improvement: After the main forecast, apply a quantile mapping that corrects the conditional bias specifically for wind speeds above the local P95. Use a 4-week rolling window of station observations to fit a linear correction for the 95th–99th percentile range.
-
What it does: Removes the −2.2 to −2.9 m/s gale-force bias, pushing EPT-2.1 Europa's +9.0% gale skill toward +12–15%, and bringing NOAA GFS from +0.6% to positive territory.
-
Improvement: Fine-tune the temperature head on a per-country basis, using the paper's Table 4 to identify weak regions. For example, apply a stronger tail-weighting for GB, NL, and DK (where EPT-2.1 Europa is negative at >P95) and a lighter touch for CH and AT (where it already achieves +28–31%).
-
What it does: Eliminates the negative heat skill in GB (−1.2%), NL (−3.4%), and DK (−4.4%), bringing all countries to at least +10% heat skill.
-
Improvement: Add a binary classification head that predicts whether an hour is dry (<0.1 mm) or wet, trained with a focal loss to handle the zero-inflated distribution. Use this to gate the precipitation regression head.
-
What it does: Fixes the −6.1% dry-regime skill of EPT-2 HRRR (Table 5) and the −11.1% light-rain deficit of EPT-2.1 Europa. The discriminator alone should recover +5–10% dry skill, and the gated regression should improve the moderate band by another +5%.
-
Wind: +12–15% gale-force skill (up from +9%), +10% all-conditions, with no negative tail bias.
-
Temperature: +25% heat skill (up from +19.6%), +15% all-conditions, positive in every European country.
-
Solar: +20% all-conditions (up from +10.2%), +30% clear-sky, +20% overcast, and +10% typical regime.
-
Precipitation: +8% heavy-tail skill (up from +1.7%), +20% moderate rain, +15% dry regime, with frequency bias near 1.0 for heavy events.
-
General: No model shows a uniform centre-seeking bias; the system is competitive or superior to IFS in every regime and every country tested.
Abstract
First-generation AI weather models are often reported to underperform at extremes, mostly in reanalysis-based evaluations of deterministic regression systems. We verify eleven physical and AI forecast systems against European synoptic, solar, and rain-gauge stations over ten months for 10 m wind, 2 m temperature, hourly shortwave accumulation, and hourly precipitation, scoring mean absolute error (MAE) against ECMWF IFS in ERA5 1991-2020 climatological regimes. Among these systems, AI models do not show a uniform relative-skill deficit in the tails. Jua EPT-2.1 Europa leads all-conditions wind (+8.4%), while Jua EPT-2 HRRR leads temperature overall (+12.1%) and in the heat regime (+19.6 +/- 2.2%). EPT-2.1 Europa and DWD ICON Global lead at gale-force wind. Jua EPT-2.1 Helios leads solar overall (+10.2 +/- 1.7%), in overcast conditions (+16.4 +/- 3.4%), and in the clear-sky tail (+24.8 +/- 5.4%). For precipitation, three Jua models gain 14-15% at moderate intensity and 9-11% at P75-P95; EPT-2 Reasoning remains ahead above P95 (+1.7 +/- 0.5%). Failures are model-specific: ECMWF AIFS loses 4.9 +/- 2.0% in the heat tail, while NOAA GFS loses 22.8 +/- 2.0% there. Every model, including numerical weather prediction systems, shows a shared conditional bias toward the centre of the observed distribution, with an inter-model spread several times smaller than the shared signal. Missing relative skill at extremes is therefore not a property of AI weather models as a class, but of particular AI and physical models.
Sources
- MAUSAM: An Observations-focused assessment of Global AI Weather Prediction Models During the South Asian Monsoon
- WeatherReal: A Benchmark Based on In-Situ Observations for Evaluating Weather Models
- AIFS -- ECMWF's data-driven forecasting system
- Generative AI for fast and accurate statistical computation of fluids
- EPT-2 Technical Report
- Universal Diffusion-Based Probabilistic Downscaling
- FourCastNet: A Global Data-driven High-resolution Weather Model using Adaptive Fourier Neural Operators
- Fixing the Double Penalty in Data-Driven Weather Forecasting Through a Modified Spherical Harmonic Loss Function
Related papers
- NORi: An ML-Augmented Ocean Boundary Layer Parameterization
- A Checklist to assess the energy and carbon impacts of ML/AI applications in Earth System Modeling
- On the Predictive Skill of Artificial Intelligence-based Weather Models for Extreme Events using Uncertainty Quantification
- A Mechanism-Coupled Split Window Network for Medium- to High-Resolution Land Surface Temperature Retrieval
- Composable multi-satellite precipitation estimation for evolving observing systems
- Improving global precipitation forecasts with an AI weather model trained on satellite observations