Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL
cs.LG, cs.AI, cs.CL
Submitted: 2026-09-17
Updated: 2026-09-17
Comments: 29 pages, 9 figures, 11 tables
Code: https://github.com/TheMikeMerrill/personal-site.gitroot@03738d59a47f:
License: http://creativecommons.org/licenses/by/4.0/
The gist: Agent trajectories record what an agent does and what happens next.
Terminology
Abstract
Agent trajectories record what an agent does and what happens next. Yet standard supervised fine-tuning (SFT) applies loss only to agent-authored action tokens, using environment observations as context but not as prediction targets. We ask whether this convention provides the best initialization for subsequent reinforcement learning. We introduce ActObs, which also supervises the observation tokens already present in each trajectory. Although deployed agents never generate observations, learning to predict them encourages the policy to model action consequences without adding data, parameters, sequence tokens, or forward passes. The methods perform similarly after SFT but diverge after GRPO. On Qwen3-4B, GRPO from ActObs achieves higher pass@k at every evaluated sampling budget than its action-only counterpart on Terminal-Bench 2.0. On Qwen3-8B, it trades some pass@1 reliability for higher pass@k (+3.4 pp at pass@16) and solves more distinct tasks. The advantage extends to cross-domain code editing on aider-polyglot (+4.2 pp at pass@1 at 4B), whose tasks are unseen during SFT and RL. ActObs retains more entropy during RL while requiring less policy movement, leaving the final policy closer to its SFT initialization. Our analysis traces this difference to SFT: action and observation gradients rapidly become orthogonal, while action-only training leaves a large residual observation gradient and degrades environment prediction below the base model. Joint supervision prevents this one-sided specialization, preserving consequence prediction and preparing the policy for downstream exploration.
Sources
- Beyond Next-Observation Prediction: Agent-Authored World Modeling for Sequential Decision Making
- FireAct: Toward Language Agent Fine-tuning
- Evaluating Large Language Models Trained on Code
- Why Do LLM Agents Fail in Exploring New Environments? A World-Modeling Perspective
- Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Endless Terminals: Scaling RL Environments for Terminal Agents
- World Modelling Improves Language Model Agents
- Tmax: A simple recipe for terminal agents
- TAPO: Transition-Aware Policy Optimization for LLM Agents
- When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff
- Policy and World Modeling Co-Training for Language Agents
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- Offline Exploration-Aware Fine-Tuning for Long-Chain Mathematical Reasoning
- On Data Engineering for Scaling LLM Terminal Capabilities
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- ECHO: Terminal Agents Learn World Models for Free
- Qwen3 Technical Report
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks