Composable multi-satellite precipitation estimation for evolving observing systems
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Composable multi-satellite precipitation estimation for evolving observing systems".
Jane: The paper was written by Yunfan Yang, Haofei Sun, Xiuyu Sun, Wei Han, Xiaoze Xu et al. from State Key Laboratory of Atmospheric Boundary Layer Physics and Atmospheric Chemistry, Institute of Atmospheric Physics, Chinese Academy of Sciences and College of Earth and Planetary Sciences, University of Chinese Academy of Sciences and Shanghai Academy of Artificial Intelligence for Science and Key Laboratory of Numerical Modeling for Tropical Cyclone of the China Meteorological Administration, Shanghai Typhoon Institute and State Key Laboratory of Severe Weather, Chinese Academy of Meteorological Sciences and China Meteorological Administration Earth System Modeling and Prediction Centre and School of Atmospheric Physics, Nanjing University of Information Science and Technology and National Satellite Meteorological Center, China Meteorological Administration.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone! We're looking at a fascinating new paper titled 'Composable multi-satellite precipitation estimation for evolving observing systems.' It’s written by Yunfan Yang and a massive team of researchers from places like the Chinese Academy of Sciences and the Shanghai Academy of Artificial Intelligence for Science.
Jane: That title sounds quite intimidating, Tom, but I think it's actually quite beautiful once you break it down. "Composable" basically means you can snap different pieces together like Lego bricks, which is perfect for when we have different satellites providing different bits of information.
Tom: Exactly, Jane! And the "evolving observing systems" part is the real genius here because our satellite constellations aren't static; new ones launch all the time.
Lu: It’s a brilliant way to think about AI research! Instead of building one giant, rigid model that breaks every time a new sensor comes online, they've built something that grows with our technology. I can see this being applied to almost any type of remote sensing data in the future.
Meng: I like the sound of that from a deployment standpoint. Usually, when you add a new data source, you're looking at a massive retraining headache for your engineers, but this "plug-and-play" approach seems to sidestep that entirely.
Lalam: It really does change the landscape of how we interact with planetary data. By making these systems adaptable, we're essentially creating a more resilient digital twin of our world, which helps cultures in high-risk areas feel much more secure about their environments.
Jane: That sense of security is huge, isn't it? If the model can just "plug in" a new satellite without needing months of retraining, we get faster updates for people who need them most.
Tom: So, how does this actually work under the hood? Let's move into what they actually built with this framework.
Summary: Tom: So, the team developed this framework called PRISMA, which stands for Precipitation Inference from Satellite Modalities via generAtive modeling. It’s a latent generative framework that uses something called "rectified flow" to estimate rainfall.
Jane: I'll try to make that a bit less technical for our listeners! Instead of trying to predict every single raindrop in a massive, high-resolution image, they compress the data into a smaller, "latent" space first. It's like looking at a simplified sketch of a landscape before you decide where to paint all the tiny details.
Meng: That's a much more efficient way to handle it. If you tried to run heavy diffusion models on raw, high-resolution satellite imagery, your compute costs would absolutely skyrocket and your latency would be terrible.
Lu: And because they use this latent space, they can train an "unconditional prior" first. They basically teach the AI what rain looks like in general using the IMERG data, and then they just add specific "branches" for different satellites like AGRI or GMI.
Jane: It’s like teaching a child what a dog looks like in general, and then showing them a specific Golden Retriever or a Poodle so they can recognize those exact breeds.
Lalam: That kind of modular intelligence is what will drive the next wave of environmental awareness. We're moving away from static software and toward living systems that learn from every new piece of evidence we provide.
Meng: I'm curious about the actual training process, though. Did they have to retrain everything when they added the microwave data?
Tom: Actually, no, and that’s the best part! They keep the main "backbone" frozen and only train these small, instrument-specific branches.
Jane: Which brings us to how much better this actually performs compared to what we're using now.
Improvements: Tom: The results are pretty staggering, especially when you look at the precision. They saw the Critical Success Index improve by up to forty point three percent when they added microwave observations into the mix.
Jane: And it wasn't just about finding where it's raining; it was about seeing through the clouds. The infrared sensors see the clouds, but they can be a bit "smudgy" regarding the actual rain, whereas the microwave sensors can actually penetrate those layers to see the heavy stuff.
Meng: I noticed they mentioned that for typhoon cases, like Typhoon Krosa, this method actually restored the structure of the eyewall and those spiral rainbands. That’s a massive leap over previous models that just gave you a blurry blob of rain.
Lu: It's incredible to think about the resolution of detail they're getting! By combining these modalities, they aren't just guessing; they are reconstructing the physical reality of the storm using multiple perspectives at once.
Jane: They even validated it against actual ground stations in China, not just other satellite products. That makes the results much more trustworthy because it's checking against real rain falling on real soil.
Meng: I was looking at their inference time, too. They managed to keep the average time at about thirty-seven seconds per member on an A100 GPU, which is fast enough to be actually useful for operational forecasting.
Lalam: That speed is what turns a research paper into a life-saving tool. When you can provide accurate, structured data about a typhoon's core in under a minute, you're giving emergency responders the gift of time.
Tom: It really is the difference between seeing a shadow and seeing the object itself. Let's wrap this all up and see what we can take away from it.
Conclusion: Jane: We've covered a lot of ground today, from "composable" architectures to the way PRISMA uses latent space to reconstruct intense typhoons. It seems like a major step toward more flexible and accurate weather monitoring.
Tom: This paper, 'Composable multi-satellite precipitation estimation for evolving observing systems,' really sets a new standard for how we can integrate heterogeneous data without breaking the model.
Lu: I'm already thinking about how this could be scaled to soil moisture or even sea surface temperatures. The potential for a universal "Earth observation foundation model" is right there in front of us!
Meng: From my side, the modularity is the real winner. If a company wants to build a custom weather service, they can just plug in their own proprietary sensor data and go.
Lalam: And ultimately, this helps us build a more empathetic relationship with our planet's cycles. We're getting better at listening to what the Earth is telling us through these satellites.
Jane: Well said, Lalam! Thanks for joining us, everyone.
Tom: We'll see you next time for the next big paper on arXiv! Goodbye!
State Key Laboratory of Atmospheric Boundary Layer Physics and Atmospheric Chemistry, Institute of Atmospheric Physics, Chinese Academy of Sciences · College of Earth and Planetary Sciences, University of Chinese Academy of Sciences · Shanghai Academy of Artificial Intelligence for Science · Key Laboratory of Numerical Modeling for Tropical Cyclone of the China Meteorological Administration, Shanghai Typhoon Institute · State Key Laboratory of Severe Weather, Chinese Academy of Meteorological Sciences · China Meteorological Administration Earth System Modeling and Prediction Centre · School of Atmospheric Physics, Nanjing University of Information Science and Technology · National Satellite Meteorological Center, China Meteorological Administration
physics.ao-ph, cs.AI
Submitted: 2026-05-14
Updated: 2026-09-15
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 84/100
The gist: This paper introduces PRISMA (Precipitation Inference from Satellite Modalities via generAtive modeling), a "plug-and-play latent generative framework" for multi-sensor precipitation estimation.
Key concepts
- PRISMA
- Precipitation Inference from Satellite Modalities via generAtive modeling is a framework that uses rectified flow to estimate rainfall. It compresses satellite data into a smaller 'latent' space, making it more efficient to process than using raw, high-resolution imagery.
- Composable Architecture
- A modular design that allows new sensors or satellites to be added as 'plug-and-play' branches. This prevents the need for massive retraining of the entire model whenever new data sources become available, allowing the system to evolve alongside improving technology.
- Latent Space
- A simplified, compressed representation of complex data. Instead of processing massive, high-resolution satellite images directly—which would be computationally expensive—the AI uses this smaller space to make estimations more efficiently and reduce latency.
Terminology
Summary
This paper introduces PRISMA (Precipitation Inference from Satellite Modalities via generAtive modeling), a plug-and-play latent generative framework
for multi-sensor precipitation estimation. It addresses the limitations of existing methods that are either computationally inefficient or lack the flexibility to incorporate new sensors without retraining, providing a scalable solution for effective disaster monitoring.
The PRISMA Architecture
PRISMA formulates precipitation estimation as a conditional generative modeling problem in latent space,
where the underlying precipitation state is represented by a learned prior and constrained by heterogeneous observations. By performing generation in a compressed latent space rather than on raw fields, the model avoids the computational burden of allocating capacity to non-precipitating regions while still capturing sharp gradients and rare extremes.
The framework is guided by three core design principles:
-
Generative modeling in a
compressed latent space
to reduce training instability and computational cost. -
The use of
instrument-specific tokenizer[s]
adapted from pretrained visual foundation models to create compact, physically informed representations. -
A modular architecture where each satellite source is incorporated through an
independent conditioning branch
that can be attached or removed without modifying the core generative model.
Two-Stage Training and Inference
The framework is operationalized through a two-stage training procedure
that decouples the learned precipitation prior from the availability of specific satellite observations. In Stage 1, a precipitation tokenizer and a DiT-based Rectified Flow backbone are trained on IMERG data to learn an unconditional precipitation prior in latent space.
In Stage 2, for each satellite source, a dedicated conditioning branch is trained to provide observation-guided signals that steer the frozen generative backbone.
At inference, arbitrary subsets of these trained branches are composed on the fly
through spatially weighted feature injection. This allows for sensor-agnostic multi-source conditioning,
meaning that integrating a new sensor only requires training an additional tokenizer–conditioning pair, thereby preserving the generative prior while enabling flexible... conditioning.
Experimental Results and Validation
The framework was evaluated using geostationary infrared observations from the FY-4B AGRI and passive microwave observations from the GPM GMI. Within microwave swaths, PRISMA improves the Critical Success Index by up to 40.3% and reduces root-mean-square error by 22.6% relative to infrared-only estimation. In typhoon case studies, microwave conditioning restores eyewall and spiral rainband structures,
reducing storm-core mean absolute error by up to 42.3%.
The model's efficacy was confirmed through several validation methods:
-
Quantitative comparison against the IMERG Final product to assess satellite-product consistency.
-
Independent rain-gauge validation
across China using the national automatic weather station (AWS) network, which confirmed consistent gains in real-world precipitation detection. -
Analysis of extreme events, such as the
historically rare North China rainstorm of July 2025,
to test robustness.
Improvements for AI systems
1. Adaptive Spatially-Weighted Fusion Module
-
Improvement: Replace the current heuristic priority-based weighting (e.g., microwave > infrared) with a lightweight, meta-learning weight predictor. This module would take sensor metadata (zenith angle, signal-to-noise ratio, and instrument calibration age) and local latent feature variance as inputs to dynamically output the spatial weight maps w(m).
-
Capability: The AI system can autonomously resolve sensor conflicts in real-time. For example, it could automatically down-weight microwave observations in high-zenith angle regions where noise is high and shift priority to infrared, or prioritize microwave signals specifically when detecting the sharp gradients of a typhoon eyewall.
2. Cross-Modal Geophysical Latent Alignment (CGLA)
-
Improvement: Integrate a contrastive learning objective (such as InfoNCE loss) during the tokenizer fine-tuning stage. This forces the latent representations of heterogeneous sensors (IR, MW, Radar) observing the same geophysical phenomenon to converge toward a shared, modality-agnostic manifold.
-
Capability: This enables
near-zero-shot sensor integration.
When a new satellite instrument is launched, the system can incorporate its data with significantly less training data because the model already understands how to map that instrument's unique physics into the existing generative backbone's latent space.
3. Ground-Truth Informed Residual Learning (GTIRL)
-
Improvement: Transition the Stage 2 conditioning objective from direct precipitation estimation to a residual learning task. The conditional branches would be trained to predict the correction term x required to transform the unconditional prior (or a single-sensor estimate) into a state that matches independent ground-gauge observations.
-
Capability: The system can break the
target-product dependency
loop. Instead of merely mimicking the biases, smoothing, and errors inherent in gridded products like IMERG, the AI produces precipitation estimates that are physically anchored to actual ground-truth measurements.
4. Heteroscedastic Uncertainty-Aware Denoising
-
Improvement: Implement a feedback mechanism where the model's predicted local variance sigma 2(z) in the latent space modulates the denoising schedule sigma'(t) and the number of sampling steps per spatial region.
-
Capability: The AI performs
intelligent computational resource allocation.
It can deploy high-fidelity, multi-step denoising cycles on complex, high-variance convective storm cores while simultaneously using ultra-fast, low-step inference for stable, non-precipitating atmospheric regions, maximizing both accuracy and operational throughput.
5. Multi-Task Geophysical Foundation Backbone
-
Improvement: Expand the DiT backbone from a single precipitation prior to a multi-task pre-training regime using a massive multi-modal dataset (e.g., simultaneous training on soil moisture, sea surface temperature, and precipitable water).
-
Capability: The framework evolves into a
Universal Geophysical State Estimator.
By simply swapping the instrument-specific tokenizers and the final decoder, the same core generative engine can perform any Earth system retrieval task with high physical consistency across different atmospheric variables.
Sources
- Flow Matching for Generative Modeling
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
- Cosmos World Foundation Model Platform for Physical AI
- Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
Related papers
- NORi: An ML-Augmented Ocean Boundary Layer Parameterization
- A Checklist to assess the energy and carbon impacts of ML/AI applications in Earth System Modeling
- On the Predictive Skill of Artificial Intelligence-based Weather Models for Extreme Events using Uncertainty Quantification
- A Mechanism-Coupled Split Window Network for Medium- to High-Resolution Land Surface Temperature Retrieval
- Improving global precipitation forecasts with an AI weather model trained on satellite observations
- Atmospheric Predictability Beyond 30 Days with Machine Learning