Dynamics Models for Offline Hyperparameter Selection in Real-World RL

arXiv:2608.11349 · cs.LG, cs.AI · Submitted 2026-08-11 · Read on arXiv

Jordan Coblin, Han Wang, Martha White, Adam White

cs.LG, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-13

Comments: Accepted to the 2026 Reinforcement Learning Conference (RLC)

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: A key obstacle to deploying reinforcement learning in real-world systems is hyperparameter selection, particularly when simulators are unavailable and online experimentation is costly.

Terminology

Summary

A key obstacle to deploying reinforcement learning in real-world systems is hyperparameter selection, particularly when simulators are unavailable and online experimentation is costly. Prior work has proposed calibration models trained on offline data to approximate environment dynamics and enable offline hyperparameter selection, but these methods have so far been evaluated only in simple simulated settings. In this paper, we present the first application of calibration models in a real-world industrial setting: a municipal water treatment plant. We evaluate several calibration model approaches, including a k-nearest neighbors model with a Laplacian distance metric, on high-dimensional, non-stationary sensor data for nexting prediction tasks. Our results show that these models can generate realistic long-horizon rollouts and recover meaningful hyperparameter sensitivity trends. We further examine how calibration models scale to year-long datasets, how they support the selection of fine-tuning learning rates for pre-trained agents, and how robust they are under distribution shift. Overall, our findings provide a proof of concept for using offline dynamics models to support RL deployment in real-world environments, while highlighting important practical challenges for future work.

Improvements for AI systems

Improvements to AI systems:

  1. Offline Hyperparameter Selection via Learned Dynamics Models
  • The AI system can now select RL hyperparameters (e.g., learning rates, discount factors) without needing a simulator or live experiments, by training a calibration model on historical sensor data.

  • Specifically, a k-nearest neighbors model with a Laplacian distance metric can generate realistic long-horizon rollouts from high-dimensional, non-stationary data, enabling the system to rank hyperparameter configurations by predicted performance before deployment.

  1. Robust Fine-Tuning of Pre-Trained Agents
  • The system can use calibration models to predict the effect of fine-tuning learning rates on a pre-trained agent’s performance, allowing it to choose an optimal rate that avoids catastrophic forgetting or slow convergence—even when the target environment’s data distribution shifts over time.
  1. Distribution Shift Detection and Adaptation
  • By comparing calibration model rollout accuracy against live sensor data, the system can detect when the environment has drifted from the training distribution. This triggers a re-calibration or triggers a switch to a more conservative policy, improving reliability in non-stationary industrial settings.
  1. Scalable Long-Horizon Prediction on Year-Long Datasets
  • The system can process and model year-long, high-dimensional time-series data efficiently, using the calibration model to generate multi-step predictions that remain realistic over long horizons—enabling proactive control decisions (e.g., chemical dosing adjustments) weeks in advance.
  1. Practical Deployment in Safety-Critical Environments
  • The improved system can be deployed in real-world industrial processes (e.g., water treatment) where online experimentation is costly or dangerous. It provides a safe, offline method to evaluate policy changes, reducing the risk of performance degradation during live trials.

What the improved AI system can do:

  • Pre-deploy RL agents in a municipal water treatment plant by simulating thousands of hyperparameter combinations offline, selecting the best one with high confidence.

  • Automatically fine-tune a pre-trained agent’s learning rate every few weeks, using only historical sensor data, to maintain performance as water quality and flow patterns change.

  • Issue early warnings when the calibration model’s prediction error exceeds a threshold, signaling a distribution shift that requires human intervention or model retraining.

  • Generate realistic, long-term forecasts of sensor readings (e.g., turbidity, pH, chlorine levels) to support planning and anomaly detection, without needing a physics-based simulator.

Abstract

A key obstacle to deploying reinforcement learning in real-world systems is hyperparameter selection, particularly when simulators are unavailable and online experimentation is costly. Prior work has proposed calibration models trained on offline data to approximate environment dynamics and enable offline hyperparameter selection, but these methods have so far been evaluated only in simple simulated settings. In this paper, we present the first application of calibration models in a real-world industrial setting: a municipal water treatment plant. We evaluate several calibration model approaches, including a k-nearest neighbors model with a Laplacian distance metric, on high-dimensional, non-stationary sensor data for nexting prediction tasks. Our results show that these models can generate realistic long-horizon rollouts and recover meaningful hyperparameter sensitivity trends. We further examine how calibration models scale to year-long datasets, how they support the selection of fine-tuning learning rates for pre-trained agents, and how robust they are under distribution shift. Overall, our findings provide a proof of concept for using offline dynamics models to support RL deployment in real-world environments, while highlighting important practical challenges for future work.

Sources

Related papers