Reward Observability and the Limits of Offline Checkpoint Selection in RSSM World Models

arXiv:2607.01736 · cs.LG, cs.AI, cs.SY, eess.SY · Submitted 2026-07-02 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Reward Observability and the Limits of Offline Checkpoint Selection in RSSM World Models".

Jane: The gist The Composite Reward Observability Fraction (CROF) is a single-number offline checkpointselection score that combines reward/observation subspace-alignment score with three structural scores,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So we’ve covered how this paper, "Reward Observability and the Limits of Offline Checkpoint Selection in RSSM World Models," uses these structural diagnostics to pick better world models from training runs. It really boils down to using the Reward Observability Fraction, or ROF, as a key indicator for performance.

Jane: Right. They found that by combining ROF with three other structural checks into this Composite Reward Observability Fraction, CROF, you get a single number that helps you select the best model checkpoint for doing things like Model Predictive Control or training an Actor-Critic policy.

Lu: And the implication is pretty clear—that focusing only on validation loss or prediction error doesn't guarantee good closed-loop performance for the agent.

Meng: So, if you're building an AI system, this means you need to look beyond just the immediate training metrics and check if your model has learned a structurally sound representation of the environment’s constraints.

Lalam: It suggests that making these structural properties explicit in the selection process helps ensure that what we deploy is actually capable of performing well when it encounters novel situations.

Tom: Exactly. The authors show that using CROF leads to a model-based A2C policy that beats a standard model-free baseline by about twenty-four point five return points while using about sixty-five times fewer real-environment interactions, which is a big deal for sample efficiency.

Jane: It means we can get much better results with less expensive, real-world testing because we’re selecting models that are structurally more aligned with what the task actually requires.

Lu: This work points toward training methods that actively encourage this structural alignment rather than just hoping the model learns it by chance during training.

Meng: It gives us a direction for improving how we design these world models from scratch, focusing on those geometric properties of the reward and observation spaces.

Lalam: Ultimately, CROF is a metric that tracks both planning performance and model-based A2C performance, which is what the authors argue makes it a solid tool for this kind of selection task.

Conclusion: Tom: So we've seen how they used this Composite Reward Observability Fraction, or CROF, to pick the best world model from training runs for planning tasks like Model Predictive Control.

Jane: It’s a single score that combines reward alignment with three structural checks to help select which model checkpoint is actually good for the agent.

Lu: What this means is that just looking at how well a model performs during training isn't enough to know if it will work when it tries to control something in the real world.

Meng: It’s about making sure the model has learned a structure that actually helps it navigate, not just memorized some specific training data points.

Lalam: The paper shows that this selection method leads to an AI policy that beats a baseline by about twenty-five return points while using much fewer real-world interactions.

Tom: So what’s the big picture here for us listening right now? It suggests we need better ways to check the inner workings of these complex world models before we commit to deploying them.

Jane: Exactly, it pushes us toward using more rigorous structural diagnostics instead of just relying on validation losses alone.

Lu: This moves the focus from just fitting data points to understanding how the model understands the physics and reward structure itself.

Meng: From an engineering standpoint, this gives us a way to filter out those overfit models that look good on paper but fail in practice when you actually try to control something.

Lalam: It hints that improving these structural alignment metrics could fundamentally change how we train world models for complex tasks across different domains.

cs.LG, cs.AI, cs.SY, eess.SY

Submitted: 2026-07-02

Updated: 2026-10-08

Comments: Preprint, 22 pages (17 main text + 5 pages appendix), 4 figures, 9 tables. Video: https://youtu.be/BlRAI2hdKDs , Code: https://github.com/nsmoly/LunarLander_RSSM and https://github.com/jonstraveladventures/rof-reproduction

Code: https://github.com/nsmoly/LunarLander_RSSM

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 86/100

The gist: The gist The Composite Reward Observability Fraction (CROF) is a single-number offline checkpointselection score that combines reward/observation subspace-alignment score with three structural

Key concepts

Composite Reward Observability Fraction (CROF)
A single score used to select the best world models. It merges a reward/observation alignment score with three structural scores, aiming to find models whose internal structure is best suited for predicting rewards.
Reward Observability Fraction (ROF)
Measures how much of the reward gradient's energy lies in the observable subspace. It quantifies the dependence of a reward predictor on what can be directly observed, which is key for selecting good models.
World Model Architecture
The RSSM family learns latent dynamics for planning and imagination. This specific work adapts RSSM to LunarLander-v3, using a hidden state and stochastic latent state to model physics and actions.
Model Predictive Control (MPC) via CEM
A planning technique where the system iteratively refines a distribution over action sequences. It uses the world model's imagination to sample candidate actions over a future horizon to find the best control strategy.

Terminology

Summary

The gist The Composite Reward Observability Fraction (CROF) is a single-number offline checkpointselection score that combines reward/observation subspace-alignment score with three structural scores, and it selects world models that train a model-based A2C policy that beats a fairly evaluated model-free A2C baseline by ∼24.5 return points while using ∼65× fewer real-environment interactions.

World Model Architecture

The RSSM family learns latent dynamics for planning and imagination-based actor–critic training, and this work adapts an RSSM [4, 6] to LunarLander-v3. The world model uses an observation of p=8-dimensional (6 continuous physics dimensions plus 2 binary leg-contact indicators) with K=4 discrete actions. The core maintains a deterministic hidden state ht ∈ R dh (dh=256, singlelayer GRU) and a stochastic latent zt ∈ R dz (dz=16), giving a full latent state st = [ht; zt] ∈ R 272. The total loss is a weighted sum: L = wrLrecon + wρLrew + wdLdone + wKL LKL, with weights wr=1.0, wρ=1.2, wd=0.5, wKL=1.0 (KL free-bits floor βKL=0.5). The decoder uses a shared 2-layer MLP backbone feeding four heads: an MSE-regression physics head (R 6), a BCE contact head (R squared logits, sigmoid at inference), a BCE done head (R 1), and a separate 3-layer reward MLP (R 272→256→256→1) that operates directly on [h; z] rather than the shared backbone.

Planning and Policy Training

The method employs two primary downstream tasks: Model Predictive Control (MPC) via Cross-Entropy Method (CEM) and an Actor-Critic policy trained in latent imagination. CEM iteratively refines a per-step categorical distribution over a planning horizon of H=25 steps, sampling N=384 candidate action sequences to accumulate discounted reward PHk=1 γk−1 rˆt+k (γ=0.97). The Actor-Critic policy is trained entirely within the latent imagination of the world model, where the actor and critic are compact 3-layer MLPs that read the latent state [h; z]. Each training step samples a real observation–action sequence, warms up the RSSM posterior for 5 steps on real data, then imagines 15 further steps by rolling out the prior with actions sampled from the actor.

Structural Diagnostics and Metrics

The study defines a suite of 40 structural validation-time metrics drawn from optimal-control theory. These include Jacobian-based controllability/observability analysis (fixed and timevarying linearizations) and three reward/observation subspace-alignment scores. The Reward Observability Fraction (ROF) measures the fraction of the reward gradient’s energy in the observable subspace, which is a key metric in this regime. The Composite Reward Observability Fraction (CROF) combines ROF with three structural regularizers into a single-number offline checkpointselection score. The CROF-selected world model trains a model-based A2C policy that beats a fairly evaluated model-free A2C baseline by ∼24.5 return points while using ∼65× fewer real-environment interactions.

Predictive Performance and Selection

The analysis shows that the strongest single predictor is the Reward Observability Fraction (ROF), which measures the reward predictor’s dependence on the observable subspace. The CROF-selected world model trains a model-based A2C policy that beats a fairly evaluated model-free A2C baseline by ∼24.5 return points while using ∼65× fewer real-environment interactions, and the same world model also drives a strong zero-shot CEM-MPC policy. The smoothed MPC oracle (WM 310) yields a comparable WM-AC peak of +209.6, so on this task CROF matches an oracle pick at zero real-environment-selection cost.

Conclusion

The CROF raw pick (WM 280) yields a model-based A2C with peak mean return +217.5 (79/100 perfect landings) — ∼24.5 points above the model-free A2C baseline (+193.0) at ∼65× fewer real-environment interactions. The CROF raw WM-AC pick (WM 280, identified offline) reaches peak mean +217.48 and requires ∼65.3× fewer real environment interactions than MF-AC. The CROF raw pick (WM 280) wins the best-by-mean A2C check point selection, showing that CROF’s auxiliary terms are doing real work.

Scope and Limitations

The mechanistic argument for ROF leans on the specific structure of LunarLander’s reward (per-step shaping plus terminal events gated by simulator-internal flags), so we expect ROF to carry less signal on tasks whose reward is fully Markovian in (ot, at). The Reacher control confirms this directionally (MLP R2 = 0.9998, ROF Spearman ρs = +0.143), and we present CROF as a metric targeting a specific, but practically common, structural condition where rewards are non-Markovian or shaped rather than as a domain-agnostic metric. The paper also shows that overfitting is structural: multi-step prediction RMSE keeps improving long after the rewards/observations/controllable subspaces have drifted out of alignment. The final actor checkpoint still requires real-environment evaluation, which is an obvious gap in the current pipeline.

Future Work

Natural extensions include cross-environment validation (continuous actions, higherdimensional observations, varied reward structures), using a ROF-like term as an auxiliary training regularizer, and developing an analogous offline metric for A2C-policy checkpoint selection. The paper also notes that the fixed linearization is the stronger MPC predictor. The authors conclude that CROF tracks both CEM-MPC and model-based A2C performance, while standard criteria (validation loss, multi-step RMSE, empirical sensitivities) instead pick deeply overfit checkpoints where MPC has poor closed-loop performance. The paper is presented as arXiv:2607.01736v2 [cs.LG] 4 Jul 2026.

Acknowledgments

We thank Larry Jackel, Urs Muller, Alexander Popov, and Alexey Kamenev for their careful reviews and helpful suggestions. The paper is a study of how to predict the downstream closed-loop performance of a learned latent world model from validation-time diagnostics alone.

References

[1] Rishabh Agarwal, Marlos C. Machado, Pablo Samuel Castro, and Marc G. Bellemare. Contrastive behavioral similarity embeddings for generalization in reinforcement learning. In International Conference on Learning Representations (ICLR), 2021.

[2] Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine. Deep reinforcement learning in a handful of trials using probabilistic dynamics models. In Advances in Neural Information Processing Systems (NeurIPS), volume 31, 2018.

[3] Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi. Dream to control: Learning behaviors by latent imagination. In International Conference on Learning Representations (ICLR), 2020.

[4] Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap. Mastering diverse domains through world models. arXiv preprint arXiv:2301.04104, 2023.

[5] Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson. Learning latent dynamics for planning from pixels. In International Conference on Machine Learning (ICML), pages 2555–2565. PMLR, 2019.

[6] Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap. Mastering diverse domains through world models. arXiv preprint arXiv:2301.04104, 2023.

[7] Rudolf E. Kalman. Mathematical description of linear dynamical systems. Journal of the Society for Industrial and Applied Mathematics, Series A: Control, 1(2):152–192, 1963.

[8] Nathan Lambert, Brandon Amos, Omry Yadan, and Roberto Calandra. Objective mismatch in model-based reinforcement learning.

Improvements for AI systems

  1. Improved checkpoint selection using Composite Reward Observability Fraction (CROF). The CROF score combines ROF with three structural regularizers into a single-number offline checkpoint selection score, which is shown to be the strongest single predictor of MPC performance. This allows the system to select world-model checkpoints that beat a fairly evaluated model-free A2C baseline by ∼24.5 return points while using ∼65× fewer real-environment interactions.

  2. Enhanced closed-loop policy training via CROF-selected models. The paper demonstrates that the CROF raw WM-AC pick (WM 280) reaches peak mean +217.48 and requires ∼65.3× fewer real-environment interactions than MF-AC, confirming that using the structural metric guides the selection of models for model-based A2C training.

  3. Better handling of non-Markovian rewards in planning by utilizing Reward Observability Fraction (ROF). ROF measures the fraction of the reward gradient that lies in the observable subspace, which is crucial because prediction accuracy (even multi-step) captures statistical fit, while planning quality depends on the structural alignment between reward-relevant, controllable, and observable subspaces.

  4. More robust control-theoretic analysis for latent dynamics. The system can now utilize Jacobian-based controllability/observability analysis (fixed and timevarying linearizations) to assess the latent space structure, which identifies dimensions that are simultaneously controllable and observable, providing a geometric property that determines planning quality independent of training losses.

  5. Optimized policy imagination through structural diagnostics. By using structural regularizers like penalizes low controllability rank, penalizes low observability rank, and penalizes high open-loop observation RMSE, the system ensures that the learned latent dynamics support closed-loop planning, preventing selection of checkpoints where the latent is barely organized.

Sources

Related papers