High-Resolution Probabilistic Data-Driven Weather Modeling with a Stretched-Grid
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "High-Resolution Probabilistic Data-Driven Weather Modeling with a Stretched-Grid".
Jane: The paper was written by Even Marius Nordhagen, Håvard Homleid Haugen, Magnus Sikora Ingstad, Aram Farhad Shafiq Salihi, Thomas Nils Nipen et al. from Norwegian Meteorological Institute and ECMWF.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Welcome back, everyone. Today we're diving into a paper that's got the whole weather forecasting community buzzing — it's called "High-Resolution Probabilistic Data-Driven Weather Modeling with a Stretched Grid." Jane, I have to say, just the title alone tells you this is something special.
Jane: It really does, Tom. And the author list reads like a who's who of European weather research — folks from the Norwegian Meteorological Institute and ECMWF, which is the European Centre for Medium-Range Weather Forecasts. You've got names like Even Marius Nordhagen, Simon Lang, Peter Dueben — these are serious players in the data-driven weather space.
Tom: So for our listeners who might not be deep in the meteorology world, what's the big deal here? What does "stretched grid" even mean?
Jane: Great question. Imagine you're looking at a map of the whole Earth, but you've got a magnifying glass held over Scandinavia. That's essentially what they've done — they dedicate super fine detail, about two point five kilometers, to the Nordic region, while the rest of the globe runs at a coarser thirty-one kilometers. It's like having a high-resolution camera focused on your backyard while still seeing the whole neighborhood.
Lu: And that's the clever part, Jane. Normally, to get that kind of detail everywhere, you'd need massive computing power. But by stretching the grid, they concentrate all that computational muscle where it matters most. The model still sees the global weather patterns — the jet streams, the big pressure systems — but it can also resolve local effects like mountain winds and sea breezes.
Tom: So it's not just a global model pretending to be regional — it's actually doing both at the same time.
Lu: Exactly. And that's a huge step forward. Most data-driven weather models operate at a uniform resolution of around twenty-five to thirty kilometers. That's fine for seeing a storm system, but it can't tell you whether it's going to rain on your specific street corner.
Jane: And that's where the "probabilistic" part comes in. This isn't just giving you one forecast — it generates a whole ensemble of possible futures. We're talking about a model that can produce as many different weather scenarios as you want, each one slightly different, to capture the uncertainty.
Meng: Right, and from an engineering standpoint, that's actually pretty remarkable. Generating an ensemble usually means running the model multiple times, which multiplies the cost. But because this model is stochastic — it injects random noise internally — each ensemble member comes at nearly the same cost as a single deterministic run.
Tom: So you're telling me we're getting more information for basically the same price?
Meng: That's the promise, Tom. And when you're running forecasts four times a day operationally, like they're doing at the Norwegian Met Institute, that cost efficiency matters a lot.
Jane: And I love that they're already running it operationally — since October two thousand twenty-five apparently. This isn't just a research curiosity; it's already being used to make real forecasts that people depend on.
Tom: That's fantastic. So we've got a model that's high-resolution where it counts, probabilistic in nature, and already in production. What's not to love?
Jane: Well, let's not get ahead of ourselves. There's a lot of technical detail in how they actually trained this thing — and some interesting challenges they had to overcome. That's where the real story is.
Lu: And I think the most interesting part is how they handled the loss function — how they taught the model to produce forecasts that look like real weather, not just smoothed-out averages.
Tom: Okay, I'm hooked. Let's get into the nitty-gritty after this short break.
Summary and Key Findings: Jane: Welcome back. We're continuing our discussion of "High-Resolution Probabilistic Data-Driven Weather Modeling with a Stretched Grid." Tom, we left off talking about how they trained this model — and I think that's where things get really interesting.
Tom: Absolutely. So Jane, what did they actually do differently in the training process?
Jane: Well, the big innovation here is in the loss function — that's the mathematical formula that tells the model how wrong it is during training. Most data-driven weather models use something called mean squared error, which basically penalizes the model for every mistake equally. But that has a problem they call the "double penalty" issue.
Lu: Right, and that's a classic problem in weather forecasting. If a model predicts a rainstorm slightly to the east of where it actually happens, it gets penalized twice — once for predicting rain where there wasn't any, and once for missing it where it actually fell. To minimize that error, the model learns to just smooth everything out, making the forecast blurry and unhelpful.
Tom: So it's like a student who figures out that the safest answer on a test is always the average — never too high, never too low.
Lu: Exactly. And that's terrible for weather forecasting, especially for extreme events. You want a model that can actually predict heavy rain or strong winds, not just a watered-down version of the average weather.
Jane: So instead, they used something called the Continuous Ranked Probability Score, or CRPS. That's a proper scoring rule that evaluates the whole probability distribution, not just a single value. And crucially, they made it stochastic — the model generates multiple ensemble members during training, and the loss function rewards the ensemble for being both accurate and appropriately spread out.
Meng: But here's the catch, and I think this is the part that really shows their engineering skill. When they just used CRPS point-by-point, the forecasts looked terrible spatially. The fields were noisy — like static on a TV screen. Each individual grid point had good statistics, but the overall pattern was incoherent.
Tom: So the model was technically right at each point, but the picture as a whole looked like garbage?
Meng: That's exactly it. And that's where the "FFT" part comes in — Fast Fourier Transform. They added a second term to the loss function that evaluates the forecast in spectral space, meaning it looks at the spatial scales of the features. Are the rain bands the right width? Is the energy distributed correctly across different wavelengths?
Jane: And that made all the difference. When they compared their model — which they call Bris CRPS-FFT — to a version trained without the spectral term, the difference was night and day. The spectral version produced rain fields that actually looked like real weather, with coherent structures and realistic sharpness.
Lu: And the verification results are genuinely impressive. Against two hundred fifty-four weather stations across Norway, their model matched or beat the operational MEPS system — that's the state-of-the-art numerical weather prediction system used in the region. For temperature, they were up to fifteen percent better in the fair CRPS score.
Tom: Fifteen percent better than an operational NWP system? That's not just a small improvement — that's a leap.
Jane: It is. And they also showed that their single ensemble members look more like real weather than the deterministic model's output. The power spectra — essentially the fingerprint of spatial scales in the forecast — matched MEPS much more closely, especially for wavelengths longer than ten kilometers.
Meng: There's still some excess noise at the very smallest scales, which they acknowledge. But compared to the alternatives, it's a massive improvement in spatial coherence.
Tom: So we've got a model that's probabilistic, high-resolution, and produces spatially realistic forecasts. What's the catch? What are the limitations?
Jane: Well, there are a few, and I think that's what we should dig into next.
Improvements and Future Directions: Tom: Welcome back. We're still talking about "High-Resolution Probabilistic Data-Driven Weather Modeling with a Stretched Grid." Jane, you mentioned there are some limitations. Let's hear them.
Jane: Right. So the first big one is temporal resolution. This model produces forecasts every six hours. That's fine for seeing the general evolution of weather, but it's too coarse for many practical applications. Think about aviation — pilots need updates much more frequently than every six hours. Or severe thunderstorm warnings — those can develop and dissipate within a couple of hours.
Lu: And that's a real gap. The spatial resolution is two point five kilometers, which is genuinely kilometer-scale. But the temporal resolution doesn't match that. The authors themselves say they want to move to hourly forecasts in future work. That would be a game-changer for nowcasting — predicting weather in the next few hours.
Tom: So we've got high-res in space but low-res in time. What else?
Meng: The other issue is the ensemble spread. When they compared their ensemble to MEPS, they found that Bris CRPS-FFT has too little spread for most variables — meaning the ensemble members are too similar to each other. The model is overconfident about what's going to happen.
Jane: And why is that? I mean, they trained it to be probabilistic.
Meng: The likely reason is that they trained it against the control analysis — the single best estimate of the true state of the atmosphere. But in reality, that analysis itself is uncertain. The model learns to predict that one specific analysis, so it doesn't capture the full range of possible initial conditions.
Lu: That's a classic issue in data-driven weather prediction. The numerical weather prediction systems like MEPS initialize with a whole ensemble of perturbed analyses — they nudge the initial conditions slightly to represent uncertainty. The data-driven model doesn't do that yet, so it's starting from a single point instead of a spread of points.
Tom: So the fix would be to initialize with an ensemble of analyses?
Jane: Exactly. And the authors suggest that as a potential improvement. If you feed the model perturbed initial states, the ensemble spread should naturally increase to better match the true uncertainty.
Meng: There's also the small-scale noise issue we mentioned. Even with the spectral loss, there's still excess energy at the finest scales — wavelengths below about ten kilometers. That means the model is generating some small-scale features that aren't physically realistic. They're not sure yet whether that needs a loss function tweak or an architectural change.
Lu: And I'd add one more thing — the verification is focused on surface variables over Norway. That makes sense given the stretched grid is centered there, but it means we don't know how well the model performs in other regions or for upper-atmosphere variables. The global part of the model is running at thirty-one kilometers, which is coarser than some other global models.
Tom: So there's room to grow, but the foundation is solid. What does this mean for the future of weather forecasting?
Jane: I think this paper is a proof point that data-driven models can operate at kilometer scale with genuine probabilistic skill. That's something a lot of people doubted was possible. And the fact that it's already running operationally at MET Norway — four times a day — shows that this isn't just a lab experiment.
Tom: Let's bring in Lalam to get a broader perspective on where this could go.
Conclusion: Tom: We're wrapping up our discussion of "High-Resolution Probabilistic Data-Driven Weather Modeling with a Stretched Grid." Lalam, what's your take on the bigger picture here?
Lalam: I think the most impactful aspect is what this means for extreme weather warnings. National meteorological services have a core responsibility to issue timely warnings for high-impact events — flash floods, severe wind, heat waves. These events are short-lived and localized, which is exactly what this model is designed to capture. The combination of kilometer-scale resolution and probabilistic ensembles gives forecasters the tools they need to say not just "it might rain" but "there's a seventy percent chance of thirty millimeters of rain in this specific valley in the next six hours."
Jane: That's a really important point. And it connects to something the paper mentions — that this model could be a pathway for next-generation extreme weather prediction. The speed of data-driven models means you can run many more scenarios, explore more possibilities, and get answers faster.
Lu: And I'd add that the stretched-grid approach itself is a template for how to make data-driven weather prediction accessible to smaller countries and regions. You don't need a supercomputer the size of a warehouse to run this — the computational cost is a fraction of traditional NWP. A national meteorological service with modest computing resources could fine-tune this approach for their own region.
Meng: From an engineering standpoint, the fact that they've already got it running operationally is the real proof. It's not just a research paper — it's a deployed system. That's the hardest step, and they've done it.
Tom: So let's summarize what we've learned today. This paper presents a data-driven weather model that combines a stretched grid for kilometer-scale regional detail with a probabilistic ensemble approach. It uses a clever loss function that evaluates forecasts both point-by-point and in spectral space, which is what gives it spatially coherent fields. And it's already performing competitively with — and in some cases better than — the operational MEPS system.
Jane: And the key innovations are the spectral CRPS loss to fix the spatial coherence problem, and the stretched-grid architecture that makes high resolution computationally feasible. The limitations — six-hour temporal resolution, insufficient ensemble spread, and some residual small-scale noise — are all addressable in future work.
Tom: This is one of those papers that feels like a turning point. We're seeing data-driven weather prediction move from proof-of-concept to operational reality, and the implications for public safety, agriculture, aviation, and renewable energy are enormous.
Jane: Absolutely. And with the system already running at MET Norway since October two thousand twenty-five this isn't hypothetical — it's happening right now. We'll be watching closely to see how they address the temporal resolution and ensemble spread in their next iterations.
Tom: Well said. That's our discussion of "High-Resolution Probabilistic Data-Driven Weather Modeling with a Stretched Grid." Thanks to everyone who joined us — Lu, Meng, and of course Lalam. We'll be back soon with another paper. Until then, keep your eyes on the sky.
Jane: And check your forecasts — they might just be generated by a graph neural network. Take care, everyone.
Even Marius Nordhagen, Håvard Homleid Haugen, Magnus Sikora Ingstad, Aram Farhad Shafiq Salihi, Thomas Nils Nipen, Ivar Ambjørn Seierstad, Inger-Lise Frogner, Mariana Clare, Simon Lang, Matthew Chantry, Peter Dueben, Jørn Kristiansen
Norwegian Meteorological Institute · ECMWF
physics.ao-ph, cs.AI
Submitted: 2026-08-15
Updated: 2026-08-18
Comments: 14 pages, 8 figures
License: http://creativecommons.org/licenses/by-sa/4.0/
Importance score: 70/100
The gist: L(x, y) = Σ v Σ t w v [L g(x g, y g) + λ r L r(x r, y r)] where L g and L r denote the global and regional loss terms, x denotes the ensemble predictions, y the targets, v the variables, t the
Key concepts
- Stretched Grid
- This technique dedicates very fine detail, like two point five kilometers, to a specific region (e.g., Scandinavia), while the rest of the globe uses a coarser resolution of thirty-one kilometers. This concentrates computational power where high detail is most needed.
- Probabilistic Modeling
- Instead of one forecast, this model generates an ensemble of possible futures by injecting internal random noise. This allows it to produce multiple scenarios, capturing the uncertainty inherent in weather predictions.
- Loss Function (CRPS-FFT)
- The Continuous Ranked Probability Score (CRPS) is used as a loss function that evaluates the entire probability distribution. By adding a Fast Fourier Transform term, the model learns to produce spatially coherent forecasts with realistic structures, unlike models that only penalize point-by-point errors.
- Ensemble Spread
- This refers to how different the members of an ensemble forecast are from each other. The discussion notes that the current model has too little spread because it is trained against a single best estimate, and adding perturbed initial states could increase this spread to better reflect true uncertainty.
Terminology
Summary
Summary
The paper presents Bris CRPS-FFT, a probabilistic data-driven weather model capable of providing an ensemble of high spatial resolution realizations of 87 variables at arbitrary forecast length and ensemble size. The model uses a stretched grid, dedicating 2.5 km resolution to a region of interest (the Nordic region), and 31 km resolution elsewhere on the globe. Based on a stochastic encoder–decoder architecture, the model is trained using a loss function based on the Continuous Ranked Probability Score (CRPS) evaluated point-wise in real and spectral space. The spectral loss component is shown to be necessary to create fields that are spatially coherent.
The model is compared to high-resolution operational numerical weather prediction forecasts from the MetCoOp Ensemble Prediction System (MEPS), showing competitive forecasts when evaluated against observations from surface weather stations. The model produced fields that are more spatially coherent than mean squared error based models and CRPS based models without the spectral component in the loss.
Methodology
The model uses graph neural networks (GNNs), which have nodes representing the atmosphere for a specific spatial location and edges representing how nodes communicate information. The model follows the stretched-grid approach developed in prior work, where regional analyses from MEPS at 2.5 km resolution are used over the Nordic region, and ERA5 is used at 31 km resolution elsewhere on the globe. The two input datasets are passed through an encoder to a hidden mesh with lower resolution but greater feature count, then through a processor with 16 message-passing steps, and finally through a decoder to the same grid as the input datasets. This process generates a 6-hour forecast, repeated autoregressively for arbitrary forecast length. The model has 229 million trainable parameters, 284,000 hidden mesh nodes, and 1.3 million input/output grid points.
To represent stochastic processes, noise is injected into the latent space through a multilayer perceptron and conditional layernorms, following AIFS-CRPS. The noise injector takes in Gaussian noise with dimensions (4, Nmesh), allowing the model to shape the noise according to the fields, and also takes information about the forecast step.
Training objective
The training objective consists of a global and a regional component. The global component, evaluated on the 31 km horizontal resolution grid, is the CRPS computed point-wise in space. The regional component, evaluated at 2.5 km horizontal resolution, combines a point-wise CRPS loss with a spectral CRPS loss. Mathematically, the training objective is expressed as:
L(x, y) = Σ v Σ t w v [L g(x g, y g) + λ r L r(x r, y r)]
where L g and L r denote the global and regional loss terms, x denotes the ensemble predictions, y the targets, v the variables, t the rollout length, and w v the variable-specific scalings. The global and regional loss terms are respectively:
L g(x, y) = L point(x, y)
L r(x, y) = L point(x, y) + λ f L freq(x, y)
The point-wise terms follow the almost-fair CRPS loss, decomposed into a mean absolute error (MAE) and a variability term:
L point(x, y) = Σ p w p [(1-ε)/M Σ i x i - y - 1/(2M(M-1)) Σ i Σ j x i - x j]
where p denotes a spatial node, w p the corresponding node weight, M is the ensemble size, and ε determines the degree of CRPS fairness
(ε = 0 corresponds to fully fair CRPS, while ε = 1/M yields conventional CRPS).
For the spectral loss term, the two-dimensional fast Fourier transform (FFT) of each field is computed for both the target and all ensemble members, and then evaluated with the almost-fair CRPS loss:
L freq(x, y) = L point(FFT(x), FFT(y))
Frequencies beyond the Nyquist limit (k < 2π/(2 × 2.5 km)) are removed using a low-pass filter. The spectral term is weighted with a factor λ f relative to the point-wise term.
Training procedure
The model follows a stage-wise training procedure with transfer learning. In stage A, the model is pre-trained on ERA5 upscaled to an O96 grid (100 km horizontal resolution) for 150,000 iterations. In stage B, the model is further trained using ERA5 on its native N320 grid (31 km resolution), first with 1-step rollout for 50,000 iterations, then fine-tuned on 12-hour rollout with a smaller learning rate for 30,000 iterations. In stage C, the stretched grid is introduced, using IFS analysis globally and MEPS analysis over the Nordic domain. Stage C1 trains with regional weighting λ r = 1.0 and spectral loss weighting λ f = 0.1 for 10,000 iterations, and stage C2 performs 1,000 iterations at each of 12-hour, 18-hour and 24-hour rollout. The ensemble size is M = 2 and CRPS fairness
ε = 0.05/M = 0.025 for all runs. The AdamW optimizer is used with a cosine loss scheduler. The total cost of training the final model is around 25,000 GPU hours.
Datasets
For pre-training (stages A and B), the ERA5 reanalysis is used with training period 1979-01-01T00Z to 2021-12-31T18Z (43 years), keeping the year 2022 for validation. In total, 97 variables are input to the model and 87 variables are predicted, where 3 are diagnostic. The variables include 15 single level variables and 6 variables at 12 pressure levels (geopotential height, temperature, specific humidity, wind speed u/v/w components). Forcing fields include solar insolation, sine/cosine of Julian day, latitude, local time, longitude, land sea mask, and surface geopotential height.
For the stretched-grid formulation (stage C), the analysis of MEPS over the Nordic domain and IFS regridded to the ERA5 grid elsewhere are used. The training period is from 2020-02-05T00Z to 2022-05-31T18Z (2.3 years), keeping 2022-06-01T00Z to 2023-06-01T00Z for validation. The boundaries of MEPS are trimmed 125 km in all directions to avoid learning from data in the relaxation zone of the NWP analyses.
Results
The evaluation focuses on surface variables of high relevance to end users: air temperature at 2m, wind speed at 10m, precipitation (accumulated over 6 hours) and mean sea-level pressure. Forecasts were verified against measurements from 254 synoptic weather stations throughout Norway, with bilinear interpolation to observation points and temperature adjustment for altitude using a constant lapse rate of 6.5°C/km.
Single member evaluation
Comparing precipitation fields, MEPS produces banded precipitation structures reflecting the dynamics of rain showers and can represent sharp features such as localized heavy precipitation. Bris MSE smoothens out unpredictable events such as heavy rainfall, resulting in overly smooth field structures. Bris CRPS demonstrates improved ability to capture localized events but introduces substantial small-scale noise, reducing spatial coherence. Bris CRPS-FFT strikes a better balance, preserving sharp features while preserving more spatial correlation, although some residual small-scale noise is still present.
Quantile-quantile plots for 6-hour accumulated precipitation and 10m wind speed show that while Bris CRPS-FFT underestimates the frequency of larger precipitation events, it performs better than Bris MSE, highlighting the added value of the probabilistic approach. All models underestimate wind speed, a bias partially attributed to verification of gridded model outputs against point measurements. Bris CRPS-FFT exhibits a more realistic distribution than AIFS-CRPS, highlighting the potential advantage of higher spatial resolution. Bris CRPS-FFT shows notable improvement over Bris MSE in predicting higher wind speeds.
Power spectra analysis using the discrete cosine transform (DCT) shows that Bris CRPS-FFT closely aligns with MEPS for wavelengths longer than 10 km, with excess energy evident at the smallest scales. Bris MSE exhibits a significant lack of energy for wavelengths shorter than 300 km, while Bris CRPS shows excessive noise at wavelengths below 50 km. The power spectra remain largely consistent between lead times of +6h and +60h.
Probabilistic evaluation
Ten ensemble members are generated for the entire verification period (2022-06-01T00Z to 2023-05-31T18Z) and compared with MEPS (30 members) and AIFS-CRPS (10 members at 31 km resolution). Using fair CRPS (fCRPS) as a function of lead time:
-
For temperature, Bris CRPS-FFT shows clear improvement at all lead times compared to MEPS, with up to 15% lower fCRPS for some lead times. This is likely due in part to MEPS's inability to retain analysis increments forward in time.
-
For wind speed, Bris CRPS-FFT performs comparably to MEPS, with both high-resolution models outperforming AIFS-CRPS.
-
For precipitation, Bris CRPS-FFT performs slightly worse than MEPS at lead times up to 12 hours, but achieves similar performance at longer lead times. The difference for short lead times is likely due to MEPS being initialized from an ensemble of perturbed initial states, whereas Bris is initialized solely from the analysis of the MEPS control run.
-
For mean sea-level pressure, Bris CRPS-FFT and MEPS perform similarly, with AIFS-CRPS outperforming both.
Evaluating the ensemble spread-skill relationship, for all variables other than wind speed, Bris CRPS-FFT has less spread than MEPS. A likely explanation is that Bris CRPS-FFT has been trained against the control analysis of MEPS, which in reality is uncertain. Despite having too little spread, the RMSE of the ensemble mean for Bris CRPS-FFT is in general better than for MEPS, except for wind speed.
Brier skill score (BSS) as a function of thresholds shows that for precipitation, Bris CRPS-FFT and MEPS perform similarly, with MEPS slightly better for small precipitation amounts and Bris CRPS-FFT slightly better at high precipitation amounts. For wind speed, Bris CRPS-FFT performs worse for all thresholds.
Conclusion
The paper concludes that Bris CRPS-FFT provides forecasts with similar forecast skill as the state-of-the-art local area NWP ensemble prediction system MEPS. For 2m temperature, Bris CRPS-FFT has up to 15% lower fCRPS for some lead times compared to MEPS. The spatial characteristics of single ensemble members are closer to that of NWP models than fields from an MSE-based DDM. Incorporating terms into the loss function that assess the model's ability to represent different spatial scales was essential to produce coherent fields.
The current model operates at 6-hour temporal resolution, which is too coarse for many applications. Future work will increase the temporal resolution to hourly forecasts. Although spatial coherence of Bris CRPS-FFT is significantly better than Bris CRPS, further modifications to the loss function or the model architecture may be necessary to further improve the spatial structures in the fields, particularly the energy spectrum at the very finest scales. The stretched-grid design enables efficient operational deployment; the system has been running operationally at MET Norway four times a day since October 2025.
Improvements for AI systems
Based on the paper, here are specific improvements I can make to AI systems and what the improved system can do:
-
Improvement: Add a spectral-domain CRPS term to the loss function, computed via FFT on low-pass filtered Fourier coefficients, in addition to point-wise CRPS.
-
What it does: Prevents the model from producing spatially incoherent fields (excessive small-scale noise) that occur with point-wise CRPS alone, while still maintaining sharp, localized features that MSE-based models lose.
-
Improvement: Implement a graph neural network with variable resolution—2.5 km over a region of interest and 31 km globally—using a hidden mesh with refinement layers proportional to input resolution.
-
What it does: Enables kilometer-scale probabilistic forecasts without the computational cost of a fully global high-resolution grid, making operational deployment feasible (runs 4× daily at MET Norway).
-
Improvement: Inject Gaussian noise into the latent space via an MLP with conditional layer norms, where noise dimensions are (4, N mesh) and conditioned on forecast lead time.
-
What it does: Generates an arbitrary number of distinct ensemble members at inference time with minimal additional cost, capturing forecast uncertainty without needing perturbed initial conditions.
-
Improvement: Train in stages: (A) coarse ERA5 at 100 km, (B) fine ERA5 at 31 km with 12-hour rollout, (C) stretched-grid with IFS/MEPS data and progressive rollout lengths (12h→18h→24h).
-
What it does: Reduces total training cost to 25,000 GPU-hours while achieving competitive skill, and allows the model to learn general weather patterns from long climate records before specializing to regional high-resolution data.
-
Improvement: Use ε = 0.025 (fairness parameter) with M=2 ensemble members during training, and weight the spectral loss term (λ f = 0.1) relative to point-wise loss.
-
What it does: Balances unbiased gradient estimates (fair CRPS) with stable training, and ensures the model learns correct energy distribution across spatial scales (verified via DCT power spectra).
-
Generates 10+ ensemble members of 87 atmospheric variables (temperature, wind, precipitation, pressure, humidity, clouds) at 2.5 km resolution over the Nordics and 31 km globally, for arbitrary forecast lengths (tested up to 60 hours).
-
Produces spatially coherent precipitation fields that preserve sharp, localized heavy-rain features (unlike MSE models that over-smooth) while avoiding the noise artifacts of point-wise CRPS models.
-
Achieves competitive or better forecast skill than operational NWP (MEPS): up to 15% lower fCRPS for 2m temperature, comparable performance for wind speed and precipitation, and better representation of extreme wind speeds in QQ plots.
-
Provides reliable probabilistic guidance with spread-skill relationships close to ideal for most variables, enabling better extreme-weather warnings (e.g., heavy precipitation, high winds) than deterministic models.
-
Operates at 6-hour temporal resolution with a path to hourly forecasts, making it suitable for nowcasting and short-range warning systems that require rapid updates.
-
Runs efficiently in production—the stretched-grid design and stage-wise training make it computationally feasible for national meteorological services to run multiple times daily.
Abstract
We present a probabilistic data-driven weather model capable of providing an ensemble of high spatial resolution realizations of 87 variables at arbitrary forecast length and ensemble size. The model uses a stretched grid, dedicating 2.5 km resolution to a region of interest, and 31 km resolution elsewhere. Based on a stochastic encoder-decoder architecture, the model is trained using a loss function based on the Continuous Ranked Probability Score (CRPS) evaluated point-wise in real and spectral space. The spectral loss components is shown to be necessary to create fields that are spatially coherent. The model is compared to high-resolution operational numerical weather prediction forecasts from the MetCoOp Ensemble Prediction System (MEPS), showing competitive forecasts when evaluated against observations from surface weather stations. The model produced fields that are more spatially coherent than mean squared error based models and CRPS based models without the spectral component in the loss.
Sources
- FourCastNet: A Global Data-driven High-resolution Weather Model using Adaptive Fourier Neural Operators
- AIFS -- ECMWF's data-driven forecasting system
- Fixing the Double Penalty in Data-Driven Weather Forecasting Through a Modified Spherical Harmonic Loss Function
- ExtremeCast: Boosting Extreme Value Prediction for Global Weather Forecast
- AIFS-CRPS: Ensemble forecasting using a model trained with a loss function based on the Continuous Ranked Probability Score
- Skillful joint probabilistic weather forecasting from marginals
- Forecasting Global Weather with Graph Neural Networks
- FourCastNet 3: A geometric approach to probabilistic machine-learning weather forecasting at scale
- EnScale: Temporally-consistent multivariate generative downscaling via proper scoring rules
- Diffusion-LAM: Probabilistic Limited Area Weather Forecasting with Diffusion
- A multi-scale loss formulation for learning a probabilistic model with proper score optimisation
- Probabilistic Weather Forecasting with Hierarchical Graph Neural Networks
- Decoupled Weight Decay Regularization
- Building Machine Learning Limited Area Models: Kilometer-Scale Weather Forecasting in Realistic Settings
- A comparison of stretched-grid and limited-area modelling for data-driven regional weather forecasting
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
Related papers
- NORi: An ML-Augmented Ocean Boundary Layer Parameterization
- A Checklist to assess the energy and carbon impacts of ML/AI applications in Earth System Modeling
- On the Predictive Skill of Artificial Intelligence-based Weather Models for Extreme Events using Uncertainty Quantification
- A Mechanism-Coupled Split Window Network for Medium- to High-Resolution Land Surface Temperature Retrieval
- Composable multi-satellite precipitation estimation for evolving observing systems
- Improving global precipitation forecasts with an AI weather model trained on satellite observations