Improving global precipitation forecasts with an AI weather model trained on satellite observations
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Improving precipitation forecasts in an AI weather model using observational data".
Jane: The paper was written by the authors from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Paper discussion segment 1: Tom: So, starting with "Improving precipitation forecasts in an AI weather model using observational data," it’s helpful to remember that simply reading the title gives us a huge amount of information about the scope of this research. It immediately signals a shift in how meteorologists approach forecasting.
Jane: Exactly. The core idea here isn't just building a bigger computer; it’s integrating artificial intelligence into established physical science models, and critically, grounding that whole system in real-world observations that are constantly flowing in.
Lu: When we read the title, we see the key components: precipitation forecasts, AI weather models, and observational data. This combination suggests a powerful feedback loop where the model is never operating in a vacuum; it’s always tethered to what actually happened yesterday.
Tom: And that moves us beyond historical models that relied solely on general atmospheric equations. The authors are essentially arguing for a hybridized system—one that respects the physics but enhances it with machine learning pattern recognition based on massive amounts of varied data inputs.
Meng: From the authorial perspective, the title implies a dedication to practical utility. They aren't just theorizing about better math; they are focused on making an actual improvement to a measurable forecast product—precipitation amount and likelihood.
Jane: And what’s really interesting is how this framing elevates the stakes. It suggests that simply having powerful weather data isn't enough; you need the right *method* to process it, one that can handle the inherent messiness of global atmospheric measurements.
Lalam: The implication for global services is enormous because it suggests a pathway for improvement that might be more accessible than building an entirely new national supercomputer. It’s about smart integration rather than pure brute force computational power.
Tom: So, we've established the components—AI, physics, and observation—and the overarching goal is better precipitation forecasting. This sets us up perfectly to look at what the paper summarizes in its abstract next, which should give us a deeper dive into *how* this integration actually functions beneath the hood.
Paper discussion segment 2: ident: We've established that "Improving precipitation forecasts in an AI weather model using observational data" moves the conversation from single predictions to quantified probabilities. Now, let's focus on the specific improvements and deeper implications the paper suggests.
Tom: So, having looked at the title, we now turn our attention to what the paper’s summary tells us about its methodology. If we strip away some of the academic jargon, the summary essentially tells us that they are training an AI model not just on general atmospheric data, but on a highly specific, integrated set of observational records.
Jane: What I took away from reading the summary is that the model isn't treating all data points equally. It’s learning to assign different weights to different types of measurements—maybe prioritizing satellite imagery over ground station readings when predicting certain types of storm systems.
Meng: The summary implies a sophisticated data preprocessing step that is critical. It suggests that before the AI can even begin its predictive work, the messy, disparate streams of observational data must be normalized and cleaned up to give the model a coherent picture of what is happening across vast distances.
Lu: From an architectural standpoint, this deep dive into the summary shows us how crucial data heterogeneity is. The model has to handle everything from oceanic temperature readings to localized wind shear measurements, all while maintaining a consistent predictive output for precipitation.
Jane: And what's really powerful about the summary is that it doesn't just say "it works." It suggests *why* it works—by demonstrating a mechanism that allows the model to correlate complex, non-linear relationships between different physical variables that might have previously been overlooked by simpler models.
Lalam: This reinforces the idea of systemic intelligence. The AI is learning correlations that are too subtle or too numerous for a human scientist to manually program into the equations, making it an incredibly powerful tool for pattern recognition in environmental chaos.
Tom: So, we've moved from understanding the scope to understanding the core mechanism described in the summary. This leads us
Paper discussion segment 3: ident: We’ve established that "Improving precipitation forecasts in an AI weather model using observational data" shifts us from single predictions to quantified probabilities, and now let’s zero in on the specific improvements and deeper implications the paper suggests.
Tom: If we distill the core advancement here, it's about moving beyond general regional forecasts. The model suggests a level of hyper-localization that is almost unprecedented in public weather services today.
Jane: Exactly. It means the forecast output isn't just for "this county"—it can be tailored to your specific watershed, or even a particular micro-climate valley where conditions might vary dramatically over just a few square miles.
Lu: This level of detail changes the entire operational playbook for sectors like agriculture. Instead of getting a blanket rainfall estimate, the system could correlate predicted precipitation directly with soil type and specific crop needs—giving you an 'evapotranspiration risk index' alongside the rainfall number.
Meng: That’s key because it shifts the predictive output from pure atmospheric physics to resource management intelligence. It’s not just telling you *if* it will rain, but how that rain will affect a defined local system, like a specific reservoir or a fragile coastal ecosystem.
Lalam: Furthermore, the paper emphasizes the model's ability to process systemic learning in real time. It suggests a mechanism of self-correction that is vital for climate change scenarios—where we are dealing with unprecedented variability.
Tom: Think of it as an adaptive intelligence layer. When a major, historical event happens—like a massive, unusual storm pattern—the model doesn't just fail; it uses the observational data from that event to recalibrate its internal understanding of physics and probability for the next time around.
Jane: This continuous refinement builds what we might call 'predictive trust.' It means the system gets better not just with more data, but by successfully learning from its own failures and global environmental shifts.
Lu: The implication here is profound for governance: it provides a shared, objective intelligence platform that can help coordinate massive cross-border planning efforts, whether it’s managing transboundary river flow or planning for global food security.
Meng: Essentially, the tool becomes a critical piece of infrastructure itself—a dynamic knowledge source that elevates policy decisions by grounding them in hyper-specific, adaptive data.
Tom: So, while the technical achievements are staggering, the ultimate implication is giving decision-makers a powerful, resilient tool to plan for an inherently unpredictable future.
Jane: It’s a perfect blend of deep learning and global necessity. Having covered how AI fundamentally improves forecasting through adaptation and specificity, we now have a fascinating look at how these advanced computational models are being applied to fields far removed from weather...
Conclusion: Tom: So, to bring everything together, this research on "Improving precipitation forecasts in an AI weather model using observational data" fundamentally changes how we think about scientific forecasting—it moves it from a statement of fact to a quantified assessment of risk.
Jane: Exactly. The takeaway isn't just better numbers; it’s the ability to translate complex atmospheric physics into actionable probabilities that empower people on the ground, regardless of their scientific background.
Lu: For me, the most remarkable aspect remains that foundational self-correction mechanism—it gives the system a palpable sense of growing reliability over time. It's a leap in systemic trustworthiness inspired by "Improving precipitation forecasts in an AI weather model using observational data."
Meng: And from an operational standpoint, that means planning for global infrastructure or agriculture can become incredibly resilient because the model is built around handling uncertainty, not eliminating it.
Lalam: What really stands out about the potential of this work is its deployability; bringing this level of sophisticated forecasting capability to remote or under-resourced communities.
Jane: It’s a beautiful illustration of how computation can truly meet human necessity. We've seen that the ultimate goal is not just prediction, but intelligent decision support.
Tom: Absolutely. The shift from generalized scientific reporting to hyper-local, context-specific intelligence is arguably the biggest change here for humanity.
Lu: It elevates the entire field, showing how deep learning can complement—not replace—the established laws of physics in complex systems like weather.
Jane: It has been such an insightful discussion on these advanced systems today; thank you so much to everyone for synthesizing this material with us.
Tom: You bet, Jane. It’s a powerful example of AI serving real-world planning needs, and it summarizes the core message of "Improving precipitation forecasts in an AI weather model using observational data" perfectly.
Jane: Well, thank you again to our listeners for joining us on this deep dive into predictive modeling.
Tom: We'll be tackling something totally different next time—we’ve got a fascinating look at deep learning applications in historical linguistics...
physics.ao-ph, cs.LG
Submitted: 2026-09-02
Updated: 2026-09-23
Comments: 16 pages, 4 figures. Submitted to Science
Project page: https://gpm.nasa.gov/data/imerg
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 83/100
The gist: The paper presents a detailed comparison and evaluation of various models—specifically IFS, AIFS-CRPS, and Laxmi—for improving precipitation forecasts using observational data, covering
Key concepts
- AI Weather Models
- These are systems that integrate artificial intelligence into established physical science models. They use machine learning pattern recognition based on massive amounts of varied data inputs to enhance traditional forecasting methods and move beyond historical equations.
- Observational Data Integration
- The model is trained on a specific, integrated set of observational records. This involves sophisticated preprocessing to normalize and clean up disparate data streams—like satellite imagery and ground station readings—allowing the AI to assign different weights to various measurements for accurate prediction.
- Hyper-localization
- The research suggests the model can provide forecasts tailored to specific areas, such as a particular watershed or micro-climate valley. This allows for predictions beyond general regional estimates, offering detail relevant to specific local conditions.
- Systemic Learning and Self-Correction
- The AI has a mechanism for self-correction. When major events occur, the model uses that observational data to recalibrate its internal understanding of physics and probability. This continuous refinement builds predictive trust over time.
Terminology
Summary
The paper presents a detailed comparison and evaluation of various models—specifically IFS, AIFS-CRPS, and Laxmi—for improving precipitation forecasts using observational data, covering climatological comparisons, extreme event prediction for tropical cyclones in both India and the U.S., decision-oriented skill assessment across global landfalling systems, and regional probability density function analysis.
Climatology Comparison:
The study includes a comparison of mean precipitation climatology between 2000 and 2023 at a resolution of 0.25°. This comparison contrasts the ERA5 mean precipitation (Figure S9A) with the IMERG mean precipitation (Figure S9B). Furthermore, it analyzes the difference in precipitation between IMERG and ERA5 (Figure S9C), noting that red indicat[es] that IMERG precipitates more at that location.
Forecasting Skill and Extreme Event Prediction:
The performance of the models is rigorously evaluated using multiple metrics.
-
Brier Scores for Indian Cyclones: Average Brier scores are presented for the IFS ensemble, AIFS-CRPS, and Laxmi across ten landfalling tropical systems in India between February 8th, 2024 and September 30th, 2025 (Figure S10). These scores are analyzed as a function of increasing precipitation thresholds.
-
U.S. Tropical Cyclone Forecasting: The study provides an analysis of observed and forecasted extreme precipitation for five landfalling tropical cyclones in the U.S.—Beryl, Debby, Francine, Helene, and Milton—spanning 2024 and 2025 (Figure S11). This comparison involves observed IMERG accumulated precipitation alongside the model Probability of Exceedance (PoE) over 50-member ensembles at the 100 mm threshold for IFS, operational AIFS-CRPS, and Laxmi. The event-level Brier score is specifically reported for each forecast panel.
Decision-Oriented Forecast Skill Evaluation:
A comprehensive evaluation of decision-oriented forecast skill is performed across 46 globally landfalling tropical cyclones that made landfall between February 8, 2024 and September 30, 2025 (Figure S12). The models—IFS, AIFS-CRPS, and Laxmi—are compared across five precipitation thresholds (50, 100, 150, 200, and 250 mm day−1). The evaluation computes several categorical metrics:
-
Critical Success Index (CSI)
-
Probability of Detection (POD)
-
False Alarm Rate (FAR) at each threshold.
Additionally, the analysis includes the best forecast by Brier Score at the 100 mm threshold over the active region.
The calculation for categorical metrics is based on a 30% probability warning threshold, defined such that a grid cell is classified as a warning if at least 15 of 50 ensemble members exceed the precipitation target.
Regional Probability Density Functions:
The study analyzes regional precipitation probability density functions for 24-hour accumulated precipitation (Figure S13). These distributions are presented for both the tropics and the extratropics, utilizing resolutions of 0.25° and 1.0°. The comparison involves multiple models: IMERG, IFS, Laxmi, AIFS-MSE, and AIFS-CRPS. In this figure:
-
Solid lines distinguish
models trained directly on IMERG observations.
-
Dashed lines denote
models trained solely on ERA5 reanalysis data or the physical IFS model.
The results are shown for four distinct geographical domains: (A) the tropics at 0.25°, (B) the extratropics at 0.25°, (C) the tropics at 1.0°, and (D) the extratropics at 1.0°.
Improvements for AI systems
Improvements to AI Systems and Capabilities
Improvement: Develop a unified, multi-objective loss function that dynamically weights physical impact metrics (like minimizing False Alarm Rate, maximizing Critical Success Index) alongside traditional statistical measures (CRPS). This system should move beyond simple end-point loss functions.
Technical Specification: Implement a Hybrid Decision-Loss Function (L hybrid):
L hybrid = alpha times CRPS + beta times CSI - 1 + gamma / (FAR) + delta / (POD)
Where alpha, beta, gamma, and delta are dynamically tuned hyperparameters based on the operational risk profile (e.g., increasing gamma during high-stakes warnings).
Improved AI Capability: The resulting system can be trained to optimize directly for operational utility rather than just statistical fidelity. It can predict not only the location of high rainfall but also the optimal warning threshold (e.g., 100 mm vs 150 mm) that maximizes expected societal benefit, minimizing both false alarms and missed events simultaneously.
Abstract
Artificial intelligence weather prediction (AIWP) systems now surpass state-of-the-art physical models for medium-range weather forecasting. Current global AIWP models are trained almost exclusively using one reanalysis dataset, ERA5, but it has known biases, particularly for precipitation. Here we fine-tune a graph-transformer architecture with IMERG precipitation data at 0.25° resolution. The resulting model improves medium-range continuous ranked probability scores by up to 19%, while also demonstrating superior skill for tropical storms and drizzle events. Our model exceeds the Brier skill score of state-of-the-art operational models on extreme rainfall prediction by 57% globally; however, a physics-based operational model remains more reliable for the heaviest precipitation events. Our results demonstrate that incorporating observations-based precipitation data directly into training can substantially improve precipitation forecasts.
Related papers
- NORi: An ML-Augmented Ocean Boundary Layer Parameterization
- A Checklist to assess the energy and carbon impacts of ML/AI applications in Earth System Modeling
- On the Predictive Skill of Artificial Intelligence-based Weather Models for Extreme Events using Uncertainty Quantification
- A Mechanism-Coupled Split Window Network for Medium- to High-Resolution Land Surface Temperature Retrieval
- Composable multi-satellite precipitation estimation for evolving observing systems
- Atmospheric Predictability Beyond 30 Days with Machine Learning