When Does Reward Teach State? A Hidden-Automaton Instrument and a Group-Language Warning Signal
cs.LG
Submitted: 2026-07-11
Updated: 2026-09-08
Comments: 22 pages, 9 figures, 10 tables (12-page main text). Ancillary file hidden-automata-rl-code.zip holds reproduction code and the per-run data behind every table and figure. v4: held-out numbers from the 5-seed uniform-budget rerun (86/103), observation-only control, 5-seed intervention, recovery-threshold and minimality robustness, experimental protocol in the main text, claim scoping
License: http://creativecommons.org/licenses/by/4.0/
The gist: Does a reinforcement-learning agent that earns reward learn its task's hidden state? We study this question with hidden finite automata that the agent partially controls.
Terminology
Abstract
Does a reinforcement-learning agent that earns reward learn its task's hidden state? We study this question with hidden finite automata that the agent partially controls. Because each automaton is known, we can normalize reward by the best achievable return and probe the network for the true state at every step. Together the two measurements separate failures that reward alone conflates. An agent can encode too little of a state its network could hold, or encode the state and still control poorly. Weak on-policy RL matches random play while the state probe stays at chance. State learning depends on the optimizer, the training budget, and the task's structure. Permutation automata provide a warning before training: no input symbol maps two distinct states to the same successor. On a stratified held-out set, 86 of 103 permutation automata fail the state probe, and the classification is stable across probe read-outs and recovery thresholds. Most of these failures come with weak reward. High reward without the state occurs but is rare. Non-permutation automata can also fail. Oracle-normalized reward alone therefore does not establish that the task's state was learned.
Sources
- Goal Misgeneralization in Deep Reinforcement Learning
- Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals
- Transformers Learn Shortcuts to Automata
- What Formal Languages Can Transformers Express? A Survey
- Language Models Need Inductive Biases to Count Inductively
- A Formal Framework for Understanding Length Generalization in Transformers
- Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
- Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
- Emergent Linear Representations in World Models of Self-Supervised Sequence Models
- Designing and Interpreting Probes with Control Tasks
- Understanding intermediate layers using linear classifier probes
- DeepMDP: Learning Continuous Latent Space Models for Representation Learning
- Learning Invariant Representations for Reinforcement Learning without Reconstruction
- Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs
- Reinforcement Learning with Unsupervised Auxiliary Tasks
- Neural Networks and the Chomsky Hierarchy
- On the Ability and Limitations of Transformers to Recognize Formal Languages
- On the Practical Computational Power of Finite Precision RNNs for Language Recognition
- Representing Formal Languages: A Comparison Between Finite Automata and Recurrent Neural Networks
- Learning Reward Machines: A Study in Partially Observable Reinforcement Learning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks