An adaptive and evolvable deep reinforcement learning framework for weather prediction
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "An adaptive and evolvable deep reinforcement learning framework for weather prediction".
Jane: The paper was written by Qiang Wu, Han Li and Jianping Huang from Collaborative Innovation Center for Western Ecological Safety, Lanzhou University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the arXiv Radio Hour, folks. I’m Tom, and with me as always is the wonderful Jane. Jane, we’ve got a paper today that I think is going to get us both pretty excited — it’s called “An adaptive and evolvable deep reinforcement learning framework for weather prediction.”
Jane: Tom, I’m already hooked just by the title. You know, weather forecasting is one of those things where we’ve seen this explosion of AI models — Pangu, GraphCast, all these names — and everyone’s always asking which one is the best. But this paper flips that question on its head.
Tom: Exactly. Instead of asking which model wins, they ask how do we get them all to work together. And the title says it all — adaptive and evolvable. The system learns to adapt in real time, and it can evolve over time as new models come out.
Jane: Right, and that’s such a smart framing. Because the reality is, no single model is best at everything. One might be great at temperature, another at wind, another at longer lead times. So why force a winner when you can have a team?
Tom: And the team metaphor is actually pretty apt here. The paper introduces something called FTAE-Weather — Feitian Adaptive Ensemble Weather. It’s like a coach that watches all the players, sees who’s performing well in the current conditions, and decides who gets the ball.
Jane: I love that. And the “evolvable” part means the coach can also cut players who aren’t contributing and bring in new talent as it becomes available. So the system doesn’t just get frozen in time — it keeps improving as the field of AI weather forecasting improves.
Tom: Which is a huge deal. Think about how fast these models are being released. A framework that can just absorb new architectures without retraining everything from scratch? That’s the kind of practical thinking that actually moves the field forward.
Jane: And the results back it up. They’re reporting RMSE reductions anywhere from seventeen percent to over seventy-eight percent compared to the best individual model, depending on the variable and lead time. That’s not a small improvement.
Tom: Yeah, and they’re doing it with fewer than zero point zero one percent additional parameters. That’s almost nothing. So you’re getting massive gains without building yet another massive model.
Jane: I think that’s the real story here. We don’t need another foundation model. We need better ways to coordinate the ones we already have. And that’s exactly what this paper is proposing.
Tom: And that’s just the title and the big idea. Wait until we get into how the reinforcement learning actually works — the Weight-Agent, the Evolve-Agent, the whole architecture. That’s coming up next.
Summary: Tom: So Jane, we’ve set the stage with the title. Now let’s get into the actual summary of “An adaptive and evolvable deep reinforcement learning framework for weather prediction.” What are the authors actually proposing here?
Jane: So the core idea is that they’ve built a framework with two main agents. The first is the Weight-Agent, which is the tactical layer. It looks at the current atmospheric state and the predictions from all the expert models, and it decides — in real time — how much to trust each model for each variable and each lead time.
Tom: And that’s where the deep reinforcement learning comes in. The agent isn’t trained with a fixed rule. It learns through trial and error, getting rewarded when its weight assignments lead to better forecasts.
Jane: Right. And the reward function is actually pretty clever. It’s not just about minimizing error. It also rewards the system for beating the best individual model, and for ranking well among all the experts. So the agent is constantly pushed to find combinations that are genuinely better than any single model.
Tom: Then there’s the second agent — the Evolve-Agent. This one works on a slower timescale. It watches which models are consistently getting high weights, and which ones are being ignored. Then it prunes the underperformers and brings in new models from the reserve pool.
Jane: So it’s like a self-improving system. The tactical agent figures out the best combination right now, and the strategic agent makes sure the pool of available models stays fresh and relevant.
Tom: And they’ve got this asynchronous caching mechanism too. That’s a practical detail that matters a lot. Different models run at different speeds, and if you had to wait for the slowest one every time, training would be painfully slow. Caching the predictions decouples the learning from the inference speed.
Jane: That’s such an engineer’s touch, and I mean that as a compliment. It shows they were thinking about real-world deployment, not just academic benchmarks.
Tom: The paper also introduces two fusion strategies. FTAE-Unify uses one set of weights for all variables, which is simpler. FTAE-Refine allows different weights for different variables, plus a bias correction term. And then there are cascade variants that use a short-horizon ensemble recursively for longer lead times.
Jane: And that cascade approach turns out to be the real winner, especially at longer lead times. At two hundred forty hours and three hundred sixty hours, the cascade version is the best across all ten variables they evaluated. That’s a two-week forecast, which is genuinely hard territory.
Tom: The summary really paints a picture of a system that gets better as the field gets better. Every new AI weather model released — regardless of architecture — can be absorbed into this framework and strengthen the collective forecast.
Jane: And that’s the part that gets me excited. We’re not just talking about a one-time improvement. This is a framework that compounds. The more diverse the models, the more opportunity for the Weight-Agent to find complementary strengths.
Tom: Alright, so we’ve got the big picture. But how does this actually play out across different lead times? That’s where the results get really interesting, and we’ll dig into that next.
Improvements: Tom: Welcome back. We’re still on “An adaptive and evolvable deep reinforcement learning framework for weather prediction,” and Jane, I want to talk about what this paper actually improves over what already exists.
Jane: Good place to start, Tom. Because there are already ensemble methods out there. People have been averaging multiple weather models for years. But the key improvement here is that the weighting is dynamic and state-dependent.
Tom: Meaning what, exactly?
Jane: Meaning traditional ensembles apply the same combination rule regardless of what the atmosphere is doing. This framework reads the current atmospheric state and adjusts the weights accordingly. So if there’s a tropical cyclone forming, it might favor one model. If there’s midlatitude storm development, it might favor another.
Tom: And that’s a fundamental shift. The paper even points out that mixture-of-experts approaches in weather forecasting have relied on fixed gating rather than state-conditioned weighting. This is the first time, as far as I can tell, that reinforcement learning is being used to make those gating decisions dynamically.
Jane: Another improvement is the Evolve-Agent. Nobody else is doing that. Most ensemble methods are static — you pick your models, you combine them, you’re done. This framework actively curates the model pool over time, pruning models that aren’t contributing and adding new ones as they’re released.
Tom: And that addresses a real pain point. The field is moving so fast. A model that was state-of-the-art two years ago might be obsolete now. This system doesn’t require manual intervention to keep up — it does it automatically.
Jane: There’s also the parameter efficiency angle. They’re adding fewer than zero point zero one percent extra parameters. That’s remarkable because most improvements in this field come from scaling up — bigger models, more training data, more compute. This is the opposite. It’s a lightweight coordination layer on top of existing models.
Tom: And the improvements are measurable. At one hundred twenty hours, FTAE-Refine beats the best baseline model on most variables. At one hundred sixty-eight hours, the cascade version wins on all ten variables. At two hundred forty and three hundred sixty hours, it’s not even close — the cascade approach dominates.
Jane: The paper also shows that the benefit of fusion is lead-time dependent. At twenty-four hours, the best individual model — Aurora — is still the winner. There’s no benefit to fusion when the models haven’t had time to diverge. But as the lead time grows, the error structures become more differentiated, and that’s when the fusion really shines.
Tom: So it’s not claiming to be better everywhere. It’s claiming to be better where it matters most — at the longer lead times where forecasting gets really hard.
Jane: Exactly. And that’s an honest and useful result. It tells you when to use this framework and when you don’t need it.
Tom: I also appreciate that they’re thinking about operational deployment. The asynchronous caching, the lightweight footprint — those are the kind of details that make a framework actually usable in practice, not just in a research paper.
Jane: Right. And that leads us to the actual first page of the paper, where they lay out the motivation and the broader vision. Let’s take a look at that.
First Page: Tom: So Jane, we’ve covered the title, the summary, the improvements. Now let’s actually walk through the first page of “An adaptive and evolvable deep reinforcement learning framework for weather prediction.” What do the authors say right out of the gate?
Jane: The first page sets up the problem beautifully. They start by talking about how numerical weather prediction — the traditional physics-based approach — is hitting diminishing returns. Higher resolution costs more and more compute but delivers less and less improvement.
Tom: And then the machine learning revolution comes along and changes the game. Models trained on ERA5 reanalysis data are now rivaling or even beating the physics-based models. But here’s the catch — they disagree with each other on which aspects of the atmosphere they predict best.
Jane: That’s the paradox of plenty, as they call it. Pangu-Weather is great at upper-air circulation. GraphCast is great at near-surface variables. FourCastNet is great at extended integrations. But no single model wins everywhere.
Tom: And that’s the problem they’re solving. Not building a better model, but building a better coordinator. They say it right there — the missing ingredient is an agent that can read the atmosphere and decide, in real time, which model to trust for each variable, location, and lead time.
Jane: And that’s exactly what deep reinforcement learning provides. The agent observes the atmospheric state, selects combination weights as actions, and refines its policy from delayed verification signals. It’s learning from the outcome of its decisions.
Tom: The first page also introduces the two-agent structure — the Weight-Agent for tactical decisions and the Evolve-Agent for strategic curation. And it frames the whole thing as turning model diversity from a coordination challenge into a compounding scientific advantage.
Jane: That last point is really the thesis of the paper. The proliferation of AI weather models is often seen as a burden — too many options, no clear winner. This paper says no, that diversity is an asset. If you can coordinate it properly, every new model makes the collective forecast stronger.
Tom: And they back that up with the numbers. RMSE reductions from seventeen point two percent to seventy-eight point three percent over the best individual model across ten atmospheric variables. That’s not incremental — that’s substantial.
Jane: The first page also makes a broader philosophical point. Progress in AI weather forecasting shouldn’t just be measured by individual model performance. It should also be measured by the collective capability that emerges when diverse architectures are systematically coordinated.
Tom: That’s a reframing of the whole field. Instead of winner-take-all competition, it’s cooperation. Instead of building bigger and bigger models, it’s building smarter ways to use what we already have.
Jane: And I think that message resonates beyond weather forecasting. Any field that’s seeing a proliferation of specialized AI models — medical diagnosis, financial prediction, climate modeling — could benefit from this kind of coordination framework.
Tom: Alright, so we’ve covered a lot of ground. Let’s bring it all together in our final segment.
Conclusion: Tom: So we’ve spent this whole episode on “An adaptive and evolvable deep reinforcement learning framework for weather prediction,” and Jane, I think it’s time to wrap it up.
Jane: Agreed, Tom. Let’s recap what makes this paper special. The core idea is that instead of building yet another AI weather model, the authors built a framework that coordinates existing models. A Weight-Agent uses deep reinforcement learning to assign dynamic, state-dependent weights to each model for each variable and lead time.
Tom: And the Evolve-Agent keeps the model pool fresh by pruning underperformers and adding new models as they’re released. So the system improves automatically as the field progresses.
Jane: The results speak for themselves. At lead times from seventy-two to three hundred sixty hours, the framework consistently outperforms the best individual model. At two hundred forty and three hundred sixty hours, the cascade version wins on every single variable they tested. And all of this comes with fewer than zero point zero one percent additional parameters.
Tom: The key insight is that model diversity is an asset, not a burden. When you have models with different strengths — some better at temperature, some at wind, some at longer ranges — the opportunity is to combine them intelligently. That’s what this framework does.
Jane: And it does it in a way that’s practical. The asynchronous caching means training doesn’t get bottlenecked by the slowest model. The lightweight footprint means it can run alongside operational pipelines without significant additional cost.
Tom: There are limitations, of course. If all the models share a systematic bias, the ensemble inherits it. And the current evaluation is based on deterministic metrics — RMSE and ACC — not probabilistic verification. But those are directions for future work, not flaws in the approach.
Jane: I think the broader implication is what excites me most. This framework could be applied beyond weather. Any field with a growing number of specialized AI models could benefit from this kind of coordination. It’s a template for turning fragmentation into strength.
Tom: And that’s a powerful message to end on. “An adaptive and evolvable deep reinforcement learning framework for weather prediction” — a paper that says we don’t need to keep building bigger models. We need to get smarter about using the ones we have.
Jane: Well said, Tom. That’s a wrap on this one. Thanks for joining us, listeners. We’ll be back next time with another paper from the arXiv. Until then, keep looking up.
Tom: And keep listening. See you all soon.
Qiang Wu, Han Li, Jianping Huang
Collaborative Innovation Center for Western Ecological Safety, Lanzhou University
physics.ao-ph, cs.LG
Submitted: 2026-07-07
Updated: 2026-08-12
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 63/100
The gist: (i) an open library of pretrained forecasting models (expert model zoo), (ii) a tactical Weight-Agent that generates situation-dependent fusion weights via deep reinforcement learning, and (iii) a
Terminology
Summary
Summary
The paper introduces Feitian Adaptive Ensemble Weather (FTAE-Weather), a lightweight framework that reframes weather forecasting as a coordination problem rather than a winner-take-all contest. Instead of training yet another foundation model, FTAE-Weather learns to orchestrate existing pretrained forecasters via deep reinforcement learning. The framework has three components: (i) an open library of pretrained forecasting models (expert model zoo), (ii) a tactical Weight-Agent that generates situation-dependent fusion weights via deep reinforcement learning, and (iii) a strategic Evolve-Agent that curates the active model pool over time.
The expert model zoo maintains a library of pretrained models drawn from diverse architectural families: Transformer-based models (Pangu-Weather, Aurora, FuXi), graph-based models (GraphCast, AIFS), diffusion-based probabilistic models (GenCast), and neural-operator or spectral-operator models (FourCastNet3, SFNO). Each model is treated as a black-box forecaster through a common input–output interface, so the eight pretrained models can be evaluated and combined without modifying or retraining their internal parameters.
The Weight-Agent produces fusion weights conditioned on the expert predictions, the atmospheric state, and the forecast lead time. Weight generation is framed as a reinforcement learning problem: the agent observes the expert forecasts and the ERA5 verification state, selects weight configurations as actions, receives a reward based on forecast skill, and updates its policy to maximise expected return. A convolutional encoder captures multiscale atmospheric patterns and inter-model discrepancies; a Softmax layer enforces the simplex constraint so that weights sum to one at every grid point. An asynchronous caching mechanism stores precomputed model outputs, decoupling policy updates from the differing inference speeds of constituent models.
FTAE-Weather supports two fusion granularities. FTAE-Unify applies one shared set of scalar weights across all atmospheric variables, with the bias correction term fixed at zero, producing a strict convex combination. FTAE-Refine allows different atmospheric variables to receive distinct weight allocations, with a residual correction term compensating for systematic biases remaining in the weighted combination. The paper also introduces cascade variants: FTAE-Cascade trains a 6-hour base ensemble and freezes its weights, then produces extended forecasts by iteratively applying the frozen 6-hour ensemble, with the output at each step feeding back as initial conditions for the next.
The Evolve-Agent updates the active model pool on a slower timescale. Its signal comes directly from the Weight-Agent: models that consistently receive high weights demonstrate practical value, while those assigned low weights contribute little. After each evaluation epoch, the Evolve-Agent ranks models by their fitness score and applies a prune-and-inject protocol that removes the lowest-ranked models and replaces them with candidates from the reserve pool. This feedback loop ties tactical weight allocation to strategic model curation, allowing the system to retire outdated models and incorporate newly released architectures without manual intervention.
The evaluation uses ERA5 reanalysis at 0.25° resolution, with data split chronologically: 1980–2017 for training, 2018–2019 for validation, and 2020–present for testing. Forecast skill is measured by latitude-weighted RMSE and ACC. The results reveal a clear lead-time dependence in the benefits of multi-model fusion. At the short lead time of 24 hours, the best individual model, Aurora, remains dominant, indicating that a single high-performing model still benefits from strong initial-condition constraints and short-range stability. By 72 hours, inter-model complementarity begins to emerge: FTAE-Cascade shows advantages for wind and humidity variables, while lower errors in geopotential height and temperature variables are still mainly achieved by the best baseline models. After 120 hours, the benefits of fusion become more evident, with FTAE-Cascade outperforming the baseline models for most variables. At the 168-hour lead time, the advantage of cascade fusion further increases, with FTAE-Cascade achieving the lowest RMSE across all evaluated variables. When the lead time is further extended to 240 and 360 hours, the errors of the baseline models continue to increase, whereas FTAE-Cascade maintains the lowest RMSE for all variables, indicating more stable error control in medium- to long-range forecasting.
The paper reports that FTAE-Weather reduces RMSE by from 17.2% to 78.3% over the best individual model in 10 atmospheric variables and outperforms conventional ensemble baselines across lead times from 72 to 360 hours, while adding fewer than 0.01% extra parameters. The authors conclude that multi-model fusion does not necessarily outperform the best individual model at all forecast lead times; its primary advantage is concentrated in medium- to long-range prediction. As the forecast horizon increases, different models exhibit increasingly distinct error-growth characteristics and systematic biases across variables, and the cascade fusion strategy can better exploit this complementarity, thereby suppressing accumulated errors in multi-variable forecasts and improving the overall stability of long-range prediction.
The discussion highlights why adaptive weighting outperforms static combination: traditional ensemble methods assign fixed or slowly updated weights that cannot respond to the atmospheric state at forecast time, whereas FTAE-Weather generates weights conditioned on the current analysis and the full set of model forecasts. The authors note limitations, including that the framework's skill ceiling is bounded by the collective capability of the model pool, that the evaluation uses deterministic metrics and could be extended to probabilistic verification and extreme-event prediction, and that the Evolve-Agent currently ranks models by a single fitness score. The broader implication is that architectural diversity in AI weather forecasting need not be a coordination problem—every new specialist model can be absorbed into an FTAE ensemble and improve the collective forecast at negligible parameter cost, turning model diversity from a coordination challenge into a compounding scientific advantage.
Improvements for AI systems
Based on the FTAE-Weather paper, here are the specific improvements I can implement in an AI system, along with the resulting capabilities:
-
Implementation: Add a reinforcement-learning-based Weight-Agent that dynamically assigns fusion weights to multiple pretrained forecasters based on the current atmospheric state, rather than using static or simple-averaging ensembles.
-
Specifics: Use a lightweight CNN encoder (≈512-dim latent) to process the initial state and expert predictions; output dense weight fields (721×1440×69) via Softmax normalization; train with an off-policy actor–critic (TD3-style) using a composite reward combining skill (quantile-transformed sMAPE), rank, and gain-over-best terms.
-
Implementation: Replace single-step fusion with a recursive cascade: train a 6-hour base ensemble, freeze its weights, then iteratively apply it for longer lead times (up to 360h). This preserves temporal consistency and reduces accumulated error growth.
-
Specifics: For each 6-hour step, feed the ensemble output back as initial conditions; this outperforms direct fusion at all lead times ≥120h, achieving up to 78.3% RMSE reduction over the best individual model.
-
Implementation: Add a per-variable residual correction term (b t) predicted by a small MLP from the latent state, trained with L2 loss against the verification residual. This corrects systematic biases that uniform weighting cannot address.
-
Specifics: Only enable this in the
Refine
mode; forUnify
mode, set b t=0 to avoid under-constrained optimization. -
Implementation: Add an Evolve-Agent that periodically (every evaluation epoch, e.g., months) ranks active models by accumulated weight preference (fitness score F i) and prunes the bottom 2–3, injecting new architectures from a reserve pool.
-
Specifics: Use the Weight-Agent's revealed preferences (variable-averaged weights over time) as the fitness signal—no separate neural policy needed; this keeps the pool current without manual curation.
-
Implementation: Cache all expert model outputs during training so that policy updates are decoupled from the slowest constituent model's inference time. This makes training cost independent of model heterogeneity.
-
Specifics: Store (state, action, reward, next state) tuples in a replay buffer; sample mini-batches uniformly to break temporal correlations.
-
Outperform any single model: Achieves lower RMSE than the best individual forecaster (Pangu-Weather, GraphCast, Aurora, FuXi, etc.) across 10 atmospheric variables (z500, z850, t850, t2m, u10m, v10m, t500, q500, u500, v500) at lead times from 72h to 360h.
-
Adapt to forecast horizon: Automatically switches from trusting a single model (Aurora at 24h) to cascade fusion (at ≥120h) as model error structures diverge, exploiting complementarity.
-
Handle regime changes: Generates state-dependent weights, so it can favor one architecture during tropical convection and another during midlatitude baroclinic development—something static ensembles cannot do.
-
Minimal overhead: Adds fewer than 0.01% extra parameters compared to a typical foundation model (e.g., 5M vs. 50B+), making it feasible to run alongside operational pipelines.
-
Model-agnostic integration: Treats each expert as a black-box through a common input–output interface; new models can be absorbed without retraining the fusion layer.
-
Graceful degradation: If a model becomes outdated or fails, the Evolve-Agent prunes it and injects a replacement, maintaining forecast quality without manual intervention.
-
Probabilistic extension: Can incorporate diffusion-based models (e.g., GenCast) by using ensemble means for deterministic fusion, with a path toward full probabilistic weighting (CRPS-based rewards) for calibrated uncertainty.
-
Benchmarking tool: Provides a data-driven
second opinion
for operational centers, highlighting when standard fixed-weight combinations deviate from optimal regime-dependent weighting. -
At 120h: FTAE-Cascade achieves lowest RMSE for 9/10 variables; improvement over best baseline is particularly clear for geopotential height, wind, and humidity.
-
At 168h–360h: Cascade fusion achieves lowest RMSE for all 10 variables, with error reduction growing as lead time extends (e.g., 17.2% to 78.3% improvement over best individual model).
-
At 24h: The system correctly identifies Aurora as the best model and assigns it near-total weight, avoiding unnecessary fusion overhead.
This system effectively converts a fragmented ecosystem of specialist weather models into a single, self-improving prediction platform that compounds scientific progress rather than requiring a winner-take-all choice.
Abstract
No single AI weather model excels at all variables, pressure levels, and lead times. Rather than building yet another architecture, we reframe the forecasting problem as one of coordination. Here we present Feitian Adaptive Ensemble Weather (FTAE-Weather), a lightweight framework that learns, through deep reinforcement learning, when and where to trust each member of an open pool of pretrained forecasters. A tactical Weight-Agent reads the current atmospheric state and assigns variable- and horizon-specific fusion weights, while a strategic Evolve-Agent periodically prunes underperforming models and absorbs newly released ones. Asynchronous prediction caching keeps training cost independent of the slowest constituent model. Adding fewer than 0.01 percent extra parameters, FTAE-Weather reduces RMSE by from 17.2 percent to 78.3 percent over the best individual model in 10 atmospheric variables and outperforms conventional ensemble baselines across lead times from 72 to 360 hours. The framework thus converts a growing, fragmented inventory of specialist models into a single prediction system that strengthens as the field of AI weather forecasting releases new architectures-turning model diversity from a coordination challenge into a compounding scientific advantage.
Sources
- FengWu: Pushing the Skillful Global Medium-range Weather Forecast beyond 10 Days Lead
- FourCastNet 3: A geometric approach to probabilistic machine-learning weather forecasting at scale
- AIFS -- ECMWF's data-driven forecasting system
- Forecasting Global Weather with Graph Neural Networks
- Kilometer-Scale Convection Allowing Model Emulation using Generative Diffusion Modeling
- MoWE : A Mixture of Weather Experts
- Proximal Policy Optimization Algorithms
Related papers
- NORi: An ML-Augmented Ocean Boundary Layer Parameterization
- A Checklist to assess the energy and carbon impacts of ML/AI applications in Earth System Modeling
- On the Predictive Skill of Artificial Intelligence-based Weather Models for Extreme Events using Uncertainty Quantification
- A Mechanism-Coupled Split Window Network for Medium- to High-Resolution Land Surface Temperature Retrieval
- Composable multi-satellite precipitation estimation for evolving observing systems
- Improving global precipitation forecasts with an AI weather model trained on satellite observations