Unifying Policy Learning and State Prediction through Spatial Language Modeling

arXiv:2610.12172 · cs.RO, cs.AI · Submitted 2026-10-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Unifying Policy Learning and State Prediction through Spatial Language Modeling".

Dev: The gist Learning how actions change scene geometry can provide complementary supervision for goal-directed manipulation through Spatial Language Modeling.

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: So to recap this paper on "Unifying Policy Learning and State Prediction through Spatial Language Modeling," the core thesis is that learning how actions change scene geometry provides complementary supervision for goal-directed manipulation tasks.

Dev: They introduce Spatial Language Modeling, which uses a shared vocabulary of discrete coordinates and semantic tokens to represent scene contours, goals, action targets, and future states.

Rosa: The model organizes these elements into a task-specific grammar that allows one autoregressive Transformer to learn both generating actions and predicting the state conditioned on those actions through a common next-token objective.

Taro: It matters because this shared prediction interface lets transition supervision update the same parameters used for action generation, which is really useful for incorporating non-expert interaction data into policy learning.

Dev: They train it from scratch using random-play transition pretraining, where recorded action coordinates condition subsequent state predictions while excluding those action coordinates from the prediction loss.

Rosa: Then they use expert demonstration training to supervise both the action targets and selected subsequent states with losses that update the same parameters, making the geometric effects of interaction an explicit learning target.

Taro: The way they structure this turns action generation and state prediction into sequence-completion tasks that share a vocabulary, which is a very clean way to formulate this problem for an autoregressive Transformer.

Dev: They demonstrate closed-loop control in both simulation and on a real robot, showing that the model can decode executable action targets during runtime and update its history with new observations.

Rosa: The results show competitive simulation performance and higher task success rates than existing real-robot policy baselines for Push-T.

Taro: So what this means for us is that we have a way to use geometric consequences of actions as supervision, which is something we haven't fully integrated into the same learning framework before.

Dev: It provides a mechanism to connect transition supervision with goal-directed policy learning by making both targets accessible to the same model.

Rosa: The paper really lays out how this spatial representation makes sense for tying geometric change directly into the learned policy behavior during manipulation tasks.

Conclusion: Dev: Thinking about "Unifying Policy Learning and State Prediction through Spatial Language Modeling," the authors are trying to bridge the gap between learning *how* to move an object and learning *what* the resulting scene geometry will look like.

Rosa: They accomplish this by using a grammar-structured spatial representation where every coordinate in that space has a defined role, whether it's just where the pusher is, or where the goal is located.

Dev: The implication for us right now is that we can use these geometric consequences as an explicit learning target instead of relying solely on expert demonstrations to show us the right sequence of actions.

Rosa: For someone just listening to this, it means we’re building a system that understands the physical consequences of its movements in real-time, which should make it more robust when things don't go exactly as planned.

Dev: They showed this works on Push-T in simulation and on a real robot, so the potential is there for applying this to other manipulation problems where understanding the geometry change is critical.

Rosa: The authors are showing that connecting transition supervision with policy learning through this spatial modeling framework is a viable path forward for complex robotic control.

Minye Wu, Zehao Wang, Tinne Tuytelaars

KU Leuven

cs.RO, cs.AI

Submitted: 2026-10-08

Updated: 2026-10-08

The gist: The gist Learning how actions change scene geometry can provide complementary supervision for goal-directed manipulation through Spatial Language Modeling.

Key concepts

Spatial Language Modeling
This is a representation where the entire scene—including contours, goals, targets, and future states—is organized into a structured sequence using discrete coordinates. A shared vocabulary ensures that a single coordinate can represent the same physical location whether it's part of an observed object or a predicted future state.
Task-Specific Grammar
This is the rule set that organizes the spatial elements into Spatial Language Sentences (SLS). It dictates how scene observations, goals, and action targets must be ordered and identified. This grammar allows the model to learn the relationships between different parts of the scene during training.
Randomplay Pretraining
The model is first trained on random interactions where recorded pusher motions condition predictions of subsequent states. Crucially, these action coordinates are excluded from the prediction loss, allowing the model to learn how actions change scene geometry without being penalized for predicting those specific action locations.
Closed-Loop Control
During actual control, the model receives a history of past states and actions. It generates only executable action targets based on this history. New observations refresh the history, allowing the policy to use its learned transition knowledge to generate immediate actions while maintaining context.

Terminology

Summary

The gist Learning how actions change scene geometry can provide complementary supervision for goal-directed manipulation through Spatial Language Modeling.

Spatial Language Modeling

The core concept introduced is Spatial Language Modeling, which represents scene contours, goals, action targets, and future states with a shared vocabulary of discrete coordinates and semantic tokens A task-specific grammar organizes these elements into spatial sequences allowing one autoregressive Transformer to learn action generation and action-conditioned state prediction through a common next-token objective. This representation uses discrete workspace coordinates to provide an interface at the level of explicit geometry, where a coordinate retains the same physical reference whether it appears in an observed contour, an executable action target, or a predicted future state. Semantic tokens identify whether coordinates belong to the pusher, the manipulated object, the goal, or a movement command.

Training Scheme

The model is trained from scratch using two sources of interaction data Randomplay pretraining first trains the model to predict scene transitions conditioned on recorded pusher motions where action coordinates condition subsequent state predictions and are excluded from the prediction loss. Expert-demonstration training then supervises both action targets and selected subsequent states, with these losses updating the same parameters, making the geometric effects of interaction an explicit learning target alongside goal-directed action generation. This training scheme connects transition information from random interactions with action supervision from expert demonstrations.

Closed-Loop Control

During closed-loop control, an observed-history SLS prefix describes recent scene states and pusher actions The model completes this prefix with consecutive action blocks, which are parsed into target coordinates and executed by the robot. Future-state blocks are excluded from control decoding, and new observations refresh the history for the next completion. The policy can therefore use a model trained with transition supervision while generating only actions at execution time.

Evaluation and Results

The approach was evaluated on Push-T in simulation and on a real robot where it achieved competitive simulation performance and higher task success and target coverage than the evaluated real-robot policy baselines. In simulation, the model matched the best Max of 0.95 achieved by DP-C and DP-T, while reaching the highest Avg of 0.93 compared with 0.91 and 0.79. On the real robot, our full method achieved a success rate of 0.80, compared with 0.45 for DP-Coord, 0.50 for DP-Tag, and 0.65 for ACT. The model also supports a separate prediction setting where it generates successive scene states given an initial scene and a supplied action trajectory.

Contributions

The main contributions include introducing a grammar-structured spatial representation that expresses observed geometry, goals, action targets, and future states in a shared coordinate vocabulary. Furthermore, the work developed a training scheme that combines action-conditioned transition pretraining on random interactions with joint action and state supervision on expert demonstrations. Finally, the method demonstrates closed-loop Push-T control in simulation and on a real robot, reporting results that support the proposed formulation.

Limitations

The method has several limitations including reliance on calibrated AprilTag localization and known object geometry for the real-robot system. Additionally, the grammar is manually designed, and evaluation focuses only on planar Push-T, meaning generalization to other object geometries remains to be established. Errors accumulate during autoregressive state rollouts, which limits prediction over longer horizons. The model's performance is also dependent on the grid resolution chosen for discretization.

Conclusion

Spatial Language Modeling expresses scene geometry, goals, action targets, and future states as grammar-structured coordinate sequences One autoregressive model learns action generation and action-conditioned state prediction through a shared next-token objective. Training combines transition pretraining on random interactions with joint action and state supervision on expert demonstrations. During closed-loop execution, the model generates action targets directly and updates its context using new observations. Experiments demonstrate effective control in simulated and real-world Push-T, while rollouts conditioned on supplied actions capture successive geometric scene changes. This supports shared spatial modeling as a practical way to connect action-conditioned transition supervision with goal-directed policy learning.

--- Page 1 ---

Unifying Policy Learning and State Prediction through Spatial Language Modeling Minye Wu Zehao Wang Tinne Tuytelaars KU Leuven Abstract: Learning how actions change scene geometry can provide complementary supervision for goal-directed manipulation We introduce Spatial Language Modeling, which represents scene contours, goals, action targets, and future states with a shared vocabulary of discrete coordinates and semantic tokens. A task-specific grammar organizes these elements into spatial sequences, allowing one autoregressive Transformer to learn action generation and action-conditioned state prediction through a common next-token objective. We train the model from scratch using random-play transition pretraining followed by joint action and state training on expert demonstrations. During pretraining, recorded action coordinates condition subsequent state predictions and are excluded from the prediction loss. During control, the model decodes only executable action targets and updates its history with newly observed states. We evaluate the approach on Push-T in simulation and on a real robot. The model achieves competitive simulation performance and higher task success and target coverage than the evaluated real-robot policy baselines. Training ablations show improved control with joint action and state sequences, with further gains from random-play pretraining. Given supplied action trajectories, the same model also predicts successive scene states, capturing the geometric effects of pushing. Figure 1: Spatial Language Modeling for Push-T. Scene geometry, goals, pusher actions, and optional future states are represented as structured coordinate-token sequences. The model generates actions for closed-loop execution and learns state transitions through future-state supervision.

--- Page 2 ---

1 Introduction Robot manipulation requires selecting actions that transform a scene toward a goal In the Push-T task [1], a planar pusher moves a T-shaped block to a target pose. The resulting motion depends on both object geometry and contact. Demonstrations of these interactions provide two complementary training signals: expert actions indicate how to act, while subsequent observations reveal the geometric changes that follow. Policies such as Action Chunking Transformers, Diffusion Policy, and vision-language-action models learn actions from demonstrations [2, 1, 3, 4]. World models use future predictions for rollout generation and planning [5, 6], while recent work also jointly models actions and future observations [7]. Predicting the geometric consequences of actions can provide complementary supervision for policy learning and enable the use of non-expert interaction data. A shared prediction interface allows this transition supervision to update the same model parameters used for action generation. We therefore seek a spatial representation that makes both targets accessible to the same autoregressive model. Discrete workspace coordinates provide such an interface at the level of explicit geometry. We divide a bounded planar workspace into grid cells and assign each cell a coordinate token All possible grid locations form the coordinate vocabulary Qcoord. Object contours, target regions, pusher positions, and action trajectories can then be represented in the same spatial vocabulary. A coordinate retains the same physical reference whether it appears in an observed contour, an executable action target, or a predicted future state. Semantic tokens identify whether coordinates belong to the pusher, the manipulated object, the goal, or a movement command. Geometry and actions thus retain their distinct roles within a shared spatial reference frame.

--- Page 3 ---

We develop this representation as Spatial Language Modeling A task-specific grammar organizes scene observations, goals, action targets, and optional subsequent states into a Spatial Language Sentence(SLS). The grammar specifies how these elements are ordered and identified, while the model learns their coordinate content. A serializer converts observed geometry and recorded actions into spatial sentences, and a parser maps generated action blocks back to pusher targets. This formulation turns action generation and action-conditioned state prediction into sequence-completion tasks that share an output vocabulary and a next-token training objective. We train an autoregressive Transformer from scratch using two sources of interaction data Randomplay pretraining first trains the model to predict scene transitions conditioned on recorded pusher motions. The action coordinates remain in the input sequence, but their prediction losses are masked; subsequent states supply targets describing the resulting geometry. Expert-demonstration training then supervises both action targets and selected subsequent states. These losses update the same parameters, making the geometric effects of interaction an explicit learning target alongside goal-directed action generation.

Improvements for AI systems

  1. Improved action generation through shared spatial modeling: The system can learn action generation and action-conditioned state prediction through a common next-token objective by training on both randomplay pretraining and expert demonstrations, allowing it to connect transition information from random interactions with action supervision from expert demonstrations.

  2. Enhanced closed-loop control: The model can perform real-time manipulation by using an observed-history SLS prefix to generate actions, where the system does so by completing the prefix with consecutive action blocks, which are parsed into target coordinates and executed by the robot.

  3. Complementary geometric supervision: The model gains a crucial understanding of physical consequences during training because it can predict scene changes: Given supplied action trajectories, the same model also predicts successive scene states, capturing the geometric effects of pushing.

  4. Improved policy robustness via pretraining: By incorporating transition supervision, the system improves performance by leveraging random-play pretraining which provides additional state-transition supervision conditioned on recorded pusher actions, leading to a success rate increase from 0.65 to 0.80 in real-robot evaluations.

Sources

Related papers