Probabilistic Emulation of the Community Radiative Transfer Model Using Machine Learning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Probabilistic Emulation of the Community Radiative Transfer Model Using Machine Learning".
Jane: The paper was written by Lucas Howard, Aneesh C. Subramanian, Gregory Thompson, Benjamin Johnson and Thomas Auligne from University of Colorado, Boulder and University Center for Atmospheric Research, Joint Center for Satellite Data Assimilation.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper called "Probabilistic Emulation of the Community Radiative Transfer Model Using Machine Learning." Jane, I have to say, that title is a mouthful, but the idea behind it is actually pretty straightforward once you unpack it.
Jane: Absolutely, Tom. So the Community Radiative Transfer Model, or CRTM, is basically the tool that weather forecasters use to simulate what satellites see. When a satellite looks down at Earth, it's measuring radiation, and CRTM helps translate our weather model's predictions into what that radiation should look like. It's the bridge between the model and the observations.
Tom: Right, and that bridge is essential for data assimilation, which is how we feed satellite observations into weather forecasts to improve them. But here's the catch — CRTM is slow. Really slow, compared to how much data we have. The paper says for some sensors, less than one percent of available observations actually get used.
Jane: That's wild. We're throwing away ninety-nine percent of the data because we can't process it fast enough. And that's where machine learning comes in. The idea is to train a neural network to mimic what CRTM does, but do it way faster. The title says "probabilistic" because it doesn't just predict a single brightness temperature — it also predicts how confident it is in that prediction.
Tom: And that's a big deal. If the emulator knows when it's uncertain, forecasters can decide to trust it or fall back on the full CRTM. It's like having a quick calculator that tells you when it's guessing versus when it's certain.
Jane: Exactly. The authors are from the University of Colorado and the Joint Center for Satellite Data Assimilation, so these are the folks actually building the operational tools. They're not just doing this for fun — they want to fix a real bottleneck in weather forecasting.
Tom: And that's what excites me. This isn't just a cool AI trick. It's a practical fix for a problem that's limiting how good our forecasts can be. More data used means better forecasts, and better forecasts save lives and money.
Jane: Well said, Tom. And the fact that they're targeting the GOES satellites, which watch the Western Hemisphere constantly, means this could have immediate impact on severe weather monitoring.
Tom: Alright, so we've got the big picture. Next up, we're going to look at what the paper actually did and how they built this thing.
Summary: Jane: So Tom, we're back with "Probabilistic Emulation of the Community Radiative Transfer Model Using Machine Learning." Let's talk about what the team actually did. They built a neural network that takes the same inputs as CRTM — things like temperature, water vapor, cloud properties, ozone, all across one hundred twenty-seven pressure levels in the atmosphere.
Tom: one hundred twenty-seven levels, that's a lot of data. Plus surface variables like soil temperature and snow cover, and even sensor angles. The paper lists over a thousand input variables total. That's a massive input space.
Jane: Right, and the network has three hidden layers with five hundred twelve nodes each. It's a dense, fully connected network. Nothing fancy architecturally, but that's intentional. They want speed, and a simple network is fast.
Tom: So what does it output? It predicts brightness temperatures for ten channels on the GOES Advanced Baseline Imager, channels seven through sixteen. And on top of that, it predicts an error standard deviation for each channel. That's the probabilistic part.
Jane: And they trained it on data generated by CRTM itself. They used the GFS weather model to create realistic atmospheric states, then ran CRTM to get the "true" brightness temperatures. The neural network learns to match those outputs.
Tom: So it's learning to imitate CRTM, not to predict real observations directly. That's an important distinction. The emulator is only as good as CRTM, but if CRTM is good enough for operational use, then a fast emulator that matches it is also good enough.
Jane: Exactly. And the results are impressive. The average error across all channels is about zero point three Kelvin. For clear sky conditions, nine out of ten infrared channels have errors under zero point one Kelvin. That's really precise.
Tom: And they also checked whether the predicted error is reliable. If the network says the error is small, is it actually small? And for the most part, yes. The calibration is good for the vast majority of cases.
Jane: They also used explainable AI to check that the network learned real physics, not just memorized the training data. We'll get into that more later, but for now, the summary is: they built a fast, accurate, probabilistic emulator of CRTM.
Tom: And that's a solid foundation. But the real question is, how fast is it? And can it actually replace CRTM in an operational setting? That's what we're going to dig into next.
Improvements: Tom: Alright Jane, so we've covered what the paper did. Now let's talk about what this paper suggests as improvements over the current state of things. The big one is speed. They benchmarked the neural network against CRTM on a single CPU, and the ML model was about five times faster.
Jane: Five times faster, and that's without any optimization. The paper notes that CRTM is also slower for cloudy conditions, so the speed advantage would be even bigger in those cases. And in an operational setting, you could use GPUs to make the neural network even faster.
Tom: Right, so the improvement isn't just marginal. It's a step change in computational efficiency. And that's what opens the door to using more of the available satellite data.
Jane: But there's another improvement that I think is even more important. The probabilistic output. Previous ML emulators of CRTM were deterministic — they just gave a point estimate. This one predicts the error too. That means you can set a threshold. If the predicted error is low, you trust the emulator. If it's high, you fall back on CRTM.
Tom: That's a really smart workflow. It's not about replacing CRTM entirely. It's about using the emulator where it's reliable and keeping CRTM as a safety net. The paper calls this out explicitly as a potential operational strategy.
Jane: And the error predictions are generally reliable. The calibration curves show that for most cases, the predicted error matches the actual error well. It only breaks down for very large, rare errors, and even then the relationship is still monotonic — bigger predicted errors still mean bigger actual errors.
Tom: So even in the tail, you can rank-order your confidence. That's useful.
Jane: Another improvement is the explainability. They used SHAP values to verify that the network is learning real physics. For example, the water vapor channels show sensitivity to water vapor at the expected altitudes. The ozone channel is sensitive to ozone. And channel seven which has a reflected solar component, is heavily influenced by solar zenith angle.
Tom: That's the kind of validation that gives you confidence the model will generalize. It's not just fitting noise. It's actually capturing the underlying radiative transfer physics.
Jane: So the improvements are: speed, probabilistic error prediction, and physics-based validation. That's a strong package. But the question is, what does this mean for actual weather forecasting? Let's talk about that.
First Page: Jane: So Tom, let's go back to the very beginning of "Probabilistic Emulation of the Community Radiative Transfer Model Using Machine Learning." The first page sets the stage with a really important problem. Weather forecast quality has improved a lot over the decades, and a big reason is satellite data assimilation.
Tom: Right, and the paper makes the point that we're not using most of that data. For some sensors, less than one percent of observations get assimilated. The computational cost of the observation operator — that's the tool that maps model states to what the satellite sees — is the bottleneck.
Jane: And it's going to get worse. The paper mentions that hyperspectral instruments are coming, and they have way more channels than current sensors. So the computational challenge is going to become more acute.
Tom: That's a sobering thought. We're about to get even more data, and we're already struggling to use what we have.
Jane: But the first page also sets up the solution. Machine learning has been used to create fast observation operators before, but those were deterministic. This paper's contribution is the probabilistic aspect, which is crucial for data assimilation because you need to know the error characteristics of your observation operator.
Tom: And that's what makes this paper stand out. It's not just "here's a faster model." It's "here's a faster model that also tells you when it's uncertain." That's what makes it usable in a real data assimilation system.
Jane: The authors also frame this as a practical path forward. They're not proposing a complete replacement for CRTM. They're proposing a hybrid approach where the ML emulator handles the bulk of the data, and CRTM handles the cases where the emulator is uncertain.
Tom: So the first page really lays out the motivation clearly. We have a problem — too much data, not enough compute. And we have a potential solution — probabilistic ML emulators. The rest of the paper is about proving that solution works.
Jane: And it does work, at least in this test case. The errors are small, the speedup is real, and the physics checks out. So the question is, what's next? How do we get from this research to operational use?
Tom: That's exactly what we should talk about. Let's wrap this up with the big picture.
Conclusion: Tom: Alright, let's bring it home. We've been talking about "Probabilistic Emulation of the Community Radiative Transfer Model Using Machine Learning." Jane, what's the one-sentence summary?
Jane: It's a neural network that mimics CRTM for the GOES ABI instrument, runs about five times faster, predicts its own error, and checks out on the physics. That's a big step toward using way more satellite data in weather forecasts.
Tom: And the implications are huge. Better forecasts mean better severe weather warnings, better flight planning, better agriculture decisions. All of that starts with using the data we already have more effectively.
Jane: And this paper shows a clear path. The probabilistic output is the key innovation. It's what makes the emulator trustworthy enough to use operationally. You can set a threshold, use the fast model where it's confident, and fall back on CRTM where it's not.
Tom: There's still work to do, of course. The paper notes that the error predictions get less reliable for very large errors, and the Jacobians diverge from CRTM at high altitudes where there's little water vapor. But those are refinements, not roadblocks.
Jane: And the explainable AI results give us confidence. The model is learning real physics, not just memorizing patterns. That means it's more likely to generalize to new conditions.
Tom: So we're saying goodbye to this paper, but the ideas in it are going to stick around. The combination of speed, probabilistic uncertainty, and physics validation is a template for future work in this area.
Jane: Absolutely. And as more instruments come online, this approach is going to become even more valuable. We're going to need fast, reliable emulators for all of them.
Tom: Well said, Jane. That's a wrap on "Probabilistic Emulation of the Community Radiative Transfer Model Using Machine Learning." Thanks for listening, and we'll see you next time with a fresh paper to dig into.
Jane: Take care, everyone.
Lucas Howard, Aneesh C. Subramanian, Gregory Thompson, Benjamin Johnson, Thomas Auligne
University of Colorado, Boulder · University Center for Atmospheric Research, Joint Center for Satellite Data Assimilation
physics.ao-ph, cs.LG, stat.ML
Submitted: 2025-04-22
Updated: 2026-08-11
Comments: 26 pages, 9 figures, 1 table
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 62/100
The gist: The paper presents a machine learning-based probabilistic emulator of the Community Radiative Transfer Model (CRTM), applied to the GOES Advanced Baseline Imager (ABI).
Key concepts
- Community Radiative Transfer Model (CRTM)
- This is a tool used by weather forecasters to simulate what satellites observe. It translates predictions from a weather model into the radiation that should be measured by sensors, acting as a bridge between the model and observations.
- Machine Learning Emulator
- A neural network trained to mimic the CRTM's function. This emulator is designed to run much faster than the original CRTM, allowing for quicker processing of satellite data assimilation while still providing outputs that match the full model.
- Probabilistic Output
- The emulator does not just predict a single value; it also predicts the confidence or error associated with that prediction. This allows users to set thresholds: trust the fast emulator when its predicted error is low, and use CRTM as a safety net when uncertainty is high.
- Data Assimilation Bottleneck
- The computational cost of the observation operator—the tool that maps model states to satellite observations—is too slow compared to the volume of available data. This bottleneck limits how much useful satellite data can be incorporated into weather forecasts.
Terminology
Summary
The paper presents a machine learning-based probabilistic emulator of the Community Radiative Transfer Model (CRTM), applied to the GOES Advanced Baseline Imager (ABI). The motivation is that the computational cost of the observation operator, used to map between model forecast variables and direct observations, is a significant bottleneck preventing the exploitation of a higher volume of available data
and that for some sensors, less than 1% of available observations are assimilated.
The authors note that in NOAA's operational forecast system, only 0.02% of ABI observations are assimilated largely due to the constraints imposed by CRTM's computational efficiency.
The neural network (NN) emulator takes identical input variables as CRTM, which for this application includes nine atmospheric variables at each of the 127 pressure levels of the GFS, 16 surface variables, and seven metadata variables
— totaling 1,166 input variables. The network outputs predicted brightness temperatures for channels 7-16 and predicted error standard deviations for each channel (with the error referring to disagreement with CRTM predictions).
The architecture consists of 3 hidden fully connected layers with 512 nodes in each hidden layer
using the SWISH activation function. The output layer for brightness temperatures uses sigmoid activation with outputs scaled to between 0 and 1 using tunable parameters β min = 180 K and β max = 355 K. The output layer for predicted standard deviations uses a softmax-based activation with a minimum standard deviation δ = 0.001 K.
Training data was generated using the Global Forecast System Finite Volume Cubed Sphere (GFS FV3), with ABI data subsampled to a uniform grid with 64 km spacing. 30 days of simulated scans at 6-hour intervals were generated for both the GOES-16 and GOES-17 platforms,
totaling 151 simulated scans for February 15, 2022–March 15, 2022. The dataset was split into training, validation, and test sets in an approximate ratio of 80/10/10. The Adam optimizer was used with a batch size of 32, maximum of 200 epochs, gradient clipping at 0.5, and adaptive learning rate reduction. The loss function was the continuous rank probability score (CRPS), which will be higher for imprecise predictions (where the spread is large) or inaccurate (where the difference between the predicted mean and target is large).
Results show that RMSE of the predicted brightness temperature is 0.3 K averaged across all channels. For clear sky conditions, the RMSE is less than 0.1 K for 9 out of 10 infrared channels.
Channel 7, which has a significantly reflected short-wave IR component,
has notably larger error than other channels. Cloudy conditions also produce larger errors consistent with the more complicated scattering dynamics involved.
The normalized RMSE (errors divided by predicted standard deviation) is close to but larger than 1 for all channels and conditions, indicating that the NN slightly underpredicts the true errors.
Calibration curves show that "for smaller errors, the predictions are very well calibrated, while for larger errors the NN tends to underestimate the true error systematically. However, the general trend is monotonic – larger predicted error standard deviations tend to have larger errors. The predictions are
well calibrated for a wide range of true errors and only become notably separated from the ideal curve for values roughly an order of magnitude greater than the RMSE."
Computational benchmarking shows that the ML model is faster than CRTM by approximately a factor of 5
based on running both models on a single scan 100 times on a single CPU. CRTM is also substantially slower for cloudy conditions, meaning that the speed-up provided by the ML model would be even higher for these data points.
Jacobian comparisons for water vapor channels show that for lower altitudes, the mean and spread match relatively well. Above 200 hPA the ML model Jacobian diverges significantly from CRTM,
likely due to the lack of variability in stratospheric water vapor variability.
Explainable AI using Shapely Additive Explanations (SHAP) was applied to 1000 clear-sky and 1000 cloudy pixels. Results confirm the NN learned relevant physics: The impact of water vapor at different pressure levels is as expected in the water vapor channels
with peaks at appropriate altitudes for upper-level (channel 8), mid-level (channel 9), and lower-level (channel 10) water vapor channels. Ozone greatly impacts channel 12 brightness temperature (the ozone channel), again consistent with expectations.
For channel 7, solar zenith angle has a very large impact on both predicted brightness temperature and predicted error in both clear-sky and cloudy conditions
— consistent with the reflected short-wave IR component of that channel.
The authors conclude that "the NN we have trained is faster (roughly an order of magnitude faster than CRTM based on initial tests), approximately as accurate as other NN CRTM emulators that have been built, and generates reliable probabilistic error predictions. They envision two applications for the probabilistic component:
first, to use a threshold predicted error above which NN results will not be used, and second, as input into the observation error term. The XAI results
provide evidence that the NN is generating its predictions based on physically reasonable patterns and can be relied on to produce physically reasonable predictions on out-of-sample data presented to it in the future."
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems and what the improved system can do:
1. Probabilistic Output Layer for Uncertainty-Aware Prediction
-
Improvement: Replace deterministic point-output layers with a dual-head architecture that predicts both the mean and the standard deviation of the target variable (brightness temperature). Use a loss function like CRPS (Continuous Ranked Probability Score) instead of MSE.
-
What the improved AI system can do: It will output a full probability distribution for each prediction, not just a single value. This allows downstream users (e.g., data assimilation systems) to know how confident the model is in each prediction, enabling selective use of predictions only when the predicted error is low (e.g., thresholding at 0.5 K). This is critical for operational use where unreliable predictions can degrade forecast quality.
2. Error Calibration via Predicted Standard Deviation
-
Improvement: Train the model to explicitly predict the error standard deviation (σ) with respect to the target (CRTM). Use a calibration curve (predicted σ vs. actual error) as a post-training diagnostic and fine-tune the output activation (e.g., softplus with a minimum δ) to ensure the predicted σ is neither over- nor under-confident.
-
What the improved AI system can do: It will provide reliable uncertainty estimates—meaning that when the model says
the error is ±0.2 K,
the actual error is indeed within that range 68% of the time. This enables safe integration into data assimilation: the system can automatically discard observations where the predicted σ exceeds a safety threshold (e.g., 0.5 K), reducing the risk of assimilating bad data.
3. Physics-Consistent Feature Attribution (SHAP) for Trust and Debugging
-
Improvement: Integrate SHAP (Shapley Additive Explanations) as a standard evaluation step after training, and use it to verify that the model’s internal reasoning aligns with known physics (e.g., water vapor channels should be most sensitive to water vapor at the correct pressure levels; ozone channel to ozone; solar zenith angle only for the shortwave IR channel).
-
What the improved AI system can do: It will be auditable—operators can verify that the model is not relying on spurious correlations or memorizing the training set. This increases trust for deployment on out-of-sample data (e.g., different seasons, geographic regions, or sensor geometries). If SHAP reveals unexpected patterns, the model can be retrained or regularized before operational use.
4. Jacobian Consistency Check for Adjoint-Based Data Assimilation
-
Improvement: After training, compute the Jacobian (derivative of output with respect to each input variable, e.g., water vapor at each pressure level) and compare it to the physics-based model (CRTM). Use this as a validation metric, and if the Jacobian diverges in regions of low input variability (e.g., stratospheric water vapor), add targeted training data or a regularization term to enforce smoothness.
-
What the improved AI system can do: It will produce stable and accurate gradients, which are essential for variational data assimilation (e.g., 4D-Var) that requires the adjoint of the observation operator. Without this, the assimilation system could diverge or produce incorrect analyses.
5. Computational Efficiency with Cloud-Aware Speedup
-
Improvement: Optimize the neural network architecture (e.g., use dense layers with 512 nodes, 3 hidden layers, SWISH activation) and benchmark it against CRTM on a per-scan basis, specifically measuring speedup for cloudy vs. clear-sky conditions. Use this to report a realistic speedup factor (e.g., 5x on CPU, potentially higher with GPU).
-
What the improved AI system can do: It will process a full satellite scan (e.g., GOES-16 ABI) in under 1 second on a single CPU, compared to 5 seconds for CRTM. This enables real-time assimilation of a much larger fraction of available observations (e.g., from 50%), directly improving forecast skill without additional hardware costs.
6. Channel-Specific Error Reporting and Selective Use
-
Improvement: Report RMSE and normalized RMSE per channel (e.g., channel 7 has higher error due to reflected solar IR). Use this to set channel-specific acceptance thresholds for the probabilistic output.
-
What the improved AI system can do: It will automatically skip channels where the model is known to be less accurate (e.g., channel 7) or where the predicted σ is high, and only assimilate high-confidence channels. This maximizes the information gain while minimizing the risk of corrupting the analysis with poor-quality predictions.
7. Training with Early Stopping and Adaptive Learning Rate
-
Improvement: Implement early stopping (e.g., stop if validation loss does not improve for 10 epochs) and adaptive learning rate reduction (e.g., reduce by factor of 5 if validation loss plateaus for 3 epochs). Use L2 regularization (λ=1e-8) to prevent overfitting.
-
What the improved AI system can do: It will train faster and more reliably, avoiding overfitting to the training data. This ensures the model generalizes well to unseen conditions (e.g., different seasons, storm systems, or sensor angles), which is critical for operational robustness.
8. Data Augmentation for Rare Conditions
-
Improvement: In the training data generation, explicitly include a wider range of atmospheric states (e.g., more stratospheric water vapor variability, extreme cloud conditions) to prevent the model from being under-constrained in low-variability regions (as seen in the Jacobian divergence above 200 hPa).
-
What the improved AI system can do: It will produce physically consistent predictions even in rare but important conditions (e.g., volcanic ash, severe storms, polar stratospheric clouds), reducing the risk of large errors in extreme weather events.
Summary of What the Improved AI System Can Do:
-
Predict brightness temperatures with a mean error <0.3 K (and <0.1 K for most clear-sky IR channels) while also providing a reliable uncertainty estimate for each prediction.
-
Be used as a drop-in replacement for CRTM in data assimilation, with a 5–10x speedup, enabling assimilation of 10–100x more satellite observations.
-
Be trusted for operational use because its internal reasoning is verified against physics (via SHAP and Jacobian checks), and it can automatically filter out low-confidence predictions.
-
Provide a clear path to real-time, all-sky radiance assimilation, which is currently a major bottleneck in numerical weather prediction.
If you need, I can also provide the specific code changes (e.g., loss function, output layer, training loop) to implement these improvements.
Abstract
The continuous improvement in weather forecast skill over the past several decades is largely due to the increasing quantity of available satellite observations and their assimilation into operational forecast systems. Assimilating these observations requires observation operators in the form of radiative transfer models. Significant efforts have been dedicated to enhancing the computational efficiency of these models. Computational cost remains a bottleneck, and a large fraction of available data goes unused for assimilation. To address this, we used machine learning to build an efficient neural network based probabilistic emulator of the Community Radiative Transfer Model (CRTM), applied to the GOES Advanced Baseline Imager. The trained NN emulator predicts brightness temperatures output by CRTM and the corresponding error with respect to CRTM. RMSE of the predicted brightness temperature is 0.3 K averaged across all channels. For clear sky conditions, the RMSE is less than 0.1 K for 9 out of 10 infrared channels. The error predictions are generally reliable across a wide range of conditions. Explainable AI methods demonstrate that the trained emulator reproduces the relevant physics, increasing confidence that the model will perform well when presented with new data.
Related papers
- NORi: An ML-Augmented Ocean Boundary Layer Parameterization
- A Checklist to assess the energy and carbon impacts of ML/AI applications in Earth System Modeling
- On the Predictive Skill of Artificial Intelligence-based Weather Models for Extreme Events using Uncertainty Quantification
- A Mechanism-Coupled Split Window Network for Medium- to High-Resolution Land Surface Temperature Retrieval
- Composable multi-satellite precipitation estimation for evolving observing systems
- Improving global precipitation forecasts with an AI weather model trained on satellite observations