Replacing Tunable Parameters in Weather and Climate Models with State-Dependent Functions using Reinforcement Learning
Pritthijit Nath, Sebastian Schemm, Henry Moss, Peter Haynes, Emily Shuckburgh, Mark J. Webb
University of Cambridge · Lancaster University · Met Office Hadley Centre
cs.LG, physics.ao-ph
Submitted: 2026-08-14
Updated: 2026-08-18
Comments: 79 pages, 24 figures
Journal ref: Journal of Advances in Modeling Earth Systems (JAMES) 18 (8), e2026MS005745
DOI: 10.1029/2026MS005745
Code: https://github.com/p3jitnath/climate-rl
Project page: https://brentyi.github.io/tyro
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 87/100
Terminology
Summary
Summary
This paper investigates the use of reinforcement learning (RL) to learn state-dependent functions that replace static, tunable parameters in weather and climate model parametrisations. The authors state: "This study presents a framework that learns components of parametrisation schemes online as a function of the evolving model state using reinforcement learning (RL) and evaluates policy-driven parameter updates across idealised testbeds spanning a simple climate bias correction (SCBC), a radiative-convective equilibrium (RCE), and a zonal mean energy balance model (EBM) with single-agent and federated multi-agent settings."
The core motivation is that "Weather and climate models rely on parametrisations to represent unresolved sub-grid processes. Traditional schemes rely on fixed coefficients that are weakly constrained and tuned offline, contributing to persistent biases that limit their ability to adapt to underlying physics. The authors propose that
RL offers a state-aware alternative to static, hand-tuned parametrisations, where policies (action selection rules) set parameters as a function of the model state at each timestep and the learning task is framed as a sequence of decisions optimised for long-horizon performance in the coupled model during training."
The methodology frames the climate model as a Markov Decision Process (MDP): the evolving numerical model state is treated as the state of the decision process, parameter updates are treated as actions, and the reward quantifies how desirable the resulting model behaviour is over time.
The authors evaluate nine RL algorithms: REINFORCE, DPG, TRPO, PPO, AVG, DDPG, TD3, SAC, and TQC. They state: "Across nine RL algorithms, Truncated Quantile Critics (TQC), Deep Deterministic Policy Gradient (DDPG), and Twin Delayed DDPG (TD3) achieved the highest skill and stable convergence, with performance assessed against a static baseline using area-weighted RMSE, temperature and pressure-level diagnostics."
The experiments are conducted across a hierarchy of idealised testbeds. For the SCBC environment, the authors report: "DDPG learns to maintain a stable heating profile centred around-0.2, locking the model output precisely to the target with no variability across seeds. TD3 produces a similar mean profile but with larger, persistent variance in the heating increments. TQC, in contrast, adapts its heating actions dynamically while maintaining the observed temperature throughout the episode. They conclude that
DDPG and TQC demonstrate strong skill in bias correction and heating control."
For the RCE environment, the authors find: All RL algorithms show great skill in reducing errors at 100 hPa and 200 hPa. These levels correspond to the tropopause and mid- to upper troposphere, regions that are highly sensitive to model parametrisations.
They also note that DPG emerges as the most reliable algorithm, consistently appearing in the top-3 across all seeds and configurations
for RCE, with TQC and DDPG also performing strongly.
For the EBM environments, the authors report: TQC emerges as the most robust and consistently reliable, ranking highest in all eight configurations.
They further state: "RL assisted runs using the best performing seeds reduce areaWRMSE across most latitudes and substantially shrink these biases, with the most consistent gains in ebm-v1, where learning A and B per latitude provides sufficient flexibility to partially compensate for the structural errors over Antarctica."
The paper also explores federated multi-agent RL: "Federated reinforcement learning (FedRL) enabled regional regime-aware control by decomposing global space into zonal agents with periodic policy aggregation. In ebm-v2/3 with DDPG, the 6-agent configuration with frequent updates (fed05) achieved low zonal RMSE in tropical and mid-latitudes, surpassing both ebm-v1 and climlab. The authors note that
federated coordination with well-chosen hyperparameters enables scalable and regime aware learning, outperforming both static baselines and globally trained single-agent RL."
The authors emphasise the physical interpretability of the learnt policies: "A key advantage of the proposed RL framework is the physical interpretability of learnt policies, particularly in geophysical environments where actions correspond to parameters within the parametrisation schemes. In the single-agent ebm-v1 experiments, agents demonstrated latitude-dependent corrections such as at the poles, where the RL agents increased radiative coefficients A enhancing OLR-driven cooling and alleviating persistent warm biases."
The paper also tests robustness under imposed forcing perturbations: "As a secondary evaluation, we tested whether the learnt control policies remained physically reasonable under imposed climate-change-like perturbations, using a +4 K sea-surface-temperature perturbation in the RCE and a 4 W m−2 forcing in the EBM. The authors report:
In the RCE experiment, the perturbed profiles remain within a physically plausible thermodynamic envelope, with no evidence of grossly unrealistic lapse-rate behaviour after warming is imposed. However, they caution:
The learnt control may act like a climatological nudging term, pulling the perturbed simulation back towards the reference state rather than allowing the forced response to emerge. This ensuing weak response highlights a limitation of the present simple configuration."
The authors conclude: "This work shows that RL, when deployed in scalable and federated forms, can support the design of numerical-model parametrisation components that respond to local state and physical constraints in controlled idealised settings, with future work needed to assess robustness, uncertainty, and behaviour under altered forcing regimes relevant to weather and climate applications. They also state:
DDPG emerged as one of the most consistently strong performers across the test cases and therefore provides a sensible starting point for follow-up experiments where the core methodology is readily extensible to more complex systems, such as the current Unified Model (UM) and the forthcoming Momentum system at the UK Met Office."
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems and the resulting capabilities:
Improvement: Replace static, offline-tuned parameters in physics-based models with a reinforcement learning policy network that adjusts parameters in real-time based on the model's evolving state.
Capabilities:
-
Dynamically modulate radiative coefficients (A, B) in energy balance models to correct meridional temperature biases during simulation, rather than relying on fixed values
-
Adjust convective lapse rates per vertical level in radiative-convective equilibrium models to reduce temperature profile errors at 100 hPa and 200 hPa
-
Apply state-dependent heating corrections that maintain target temperatures with minimal drift (e.g., sustaining a −0.2 heating increment to lock temperature at 321.75 K)
Bottom line: The improved AI system can learn state-dependent, physically interpretable parameter control policies for numerical models, scale them via federated multi-agent learning, select appropriate algorithms based on environment characteristics, and integrate with existing HPC infrastructure—all while maintaining stability and physical realism under perturbed conditions.
Abstract
Weather and climate models rely on parametrisations to represent unresolved sub-grid processes. Traditional schemes rely on fixed coefficients that are weakly constrained and tuned offline, contributing to persistent biases that limit their ability to adapt to underlying physics. This study presents a framework that learns components of parametrisation schemes online as a function of the evolving model state using reinforcement learning (RL) and evaluates policy-driven parameter updates across idealised testbeds spanning a simple climate bias correction (SCBC), a radiative-convective equilibrium (RCE), and a zonal mean energy balance model (EBM) with single-agent and federated multi-agent settings. Across nine RL algorithms, Truncated Quantile Critics (TQC), Deep Deterministic Policy Gradient (DDPG), and Twin Delayed DDPG (TD3) achieved the highest skill and stable convergence, with performance assessed against a static baseline using area-weighted RMSE, temperature and pressure-level diagnostics. For the EBM, single-agent RL outperformed static parameter tuning with the strongest gains in tropical and mid-latitude bands, while federated RL on multi-agent setups enabled specialised control and faster convergence, with a six-agent DDPG configuration using frequent aggregation yielding the lowest area-weighted RMSE across the tropics and mid-latitudes. The learnt corrections were also physically meaningful as agents modulated EBM radiative parameters to reduce meridional biases, adjusted RCE lapse rates to match vertical temperature errors, and stabilised heating increments to limit drift. Overall, results show that RL can learn skilful state-dependent parametrisation components in idealised settings, offering a scalable pathway for online learning within numerical models and a starting point for evaluation in weather and climate models.
Sources
- Climate-Invariant Machine Learning
- Flower: A Friendly Federated Learning Research Framework
- Machine Learning Climate Model Dynamics: Offline versus Online Performance
- OpenAI Gym
- Addressing Function Approximation Error in Actor-Critic Methods
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Stable Machine-Learning Parameterization of Subgrid Processes in a Comprehensive Atmospheric Model Learned From Embedded Convection-Permitting Simulations
- BenchRL-QAS: Benchmarking reinforcement learning algorithms for quantum architecture search
- Deep Reinforcement Learning for Autonomous Driving: A Survey
- The JASMIN super-data-cluster
- Tune: A Research Platform for Distributed Model Selection and Training
- Continuous control with deep reinforcement learning
- Communication-Efficient Learning of Deep Networks from Decentralized Data
- Playing Atari with Deep Reinforcement Learning
- Reinforcement Learning Increases Wind Farm Power Production by Enabling Closed-Loop Collaborative Control
- Training language models to follow instructions with human feedback
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
- Proximal Policy Optimization Algorithms
- HITTER: A HumanoId Table TEnnis Robot via Hierarchical Planning and Learning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks