Social-WM: Safety-Aware Latent World Models for Robot Social Navigation

arXiv:2609.40177 · cs.RO · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "Social-WM: Safety-Aware Latent World Models for Robot Social Navigation".

Rosa: Safe social navigation requires a robot to anticipate not only the future consequences of its actions,

Dev: First, who's behind it and why it matters.

Paper summary: Rosa: Well team, we're here to discuss "Social-WM: Safety-Aware Latent World Models for Robot Social Navigation." The core thesis seems to be that safe social navigation demands a robot not just predict the future consequences of its actions, but also check if those nominal actions are actually possible given the physical and social setup.

Dev: I agree with Rosa; it sounds like they're focusing on that crucial gap between what a robot *wants* to do and what it *can* physically do in a crowded environment. It claims they learn these safety issues by looking at the actual observed future after every command, which is interesting because it bypasses needing explicit pedestrian tracking or online reinforcement learning.

Taro: I'm keen on that part about action-conditioned future prediction where the target is the actual observed future; that suggests the model learns safety consequences directly from real-world transitions, not just from simulated environments. It addresses what happens when things go wrong in unpredictable social situations.

Rosa: Exactly, and this seems particularly relevant because it moves beyond just predicting a trajectory to understanding action realizability in a dynamic setting. It claims their framework helps the robot understand that a nominal forward action might be executable in open space but needs to be constrained when approaching someone or something else.

Dev: That distinction between nominal and realizable actions is what really interests me from an engineering standpoint, because it means they can distinguish between an action that looks fine on paper and one that would actually lead to a collision or a social faux pas. How does this discrepancy manifest in the system's loop rate?

Taro: The way they introduce the realizable inverse dynamics objective seems like a smart way to ground those latent transitions in physically achievable motion, making it concrete for planning purposes. It suggests that the model learns what actions are actually possible by looking at how robot poses transition from one state to another.

Rosa: And then at deployment, they use this latent world model with a generative CVAE to propose candidate actions, and then they evaluate them by measuring the discrepancy between the nominal prediction and this newly learned realizable action estimate. That closed-loop propose–imagine–evaluate–select planning cycle is pretty neat for integrating safety constraints.

Dev: The efficiency aspect is also something I'm watching; if the planning operates entirely in latent space and only takes about fifty-one milliseconds per step, that’s fast enough to be practical, especially compared to some of those heavier methods we've seen before. But Rosa, how long can this system reliably operate outside of a highly controlled lab environment before these safety assumptions start breaking down?

Taro: That brings up the question of world misbehavior; if the real world presents something completely unexpected that the training data didn't cover, how robust is this system in handling those novel scenarios where it might need to adapt its understanding of what's realizable?

Paper summary: Rosa: The paper suggests they achieved competitive navigation success even when transferring this model zero-shot from Social-HMthree dee to Social-MPthree dee, which implies a certain level of generalization in terms of social navigation goals. However, they do flag a limitation: the overall success rate only improves modestly because many remaining failures are timeouts.

Dev: Timeouts are always an issue when you're dealing with real-time constraints; so if the planning process stalls due to complexity rather than a safety failure, that impacts our reliability metrics significantly. I wonder if that makes us worry about the latency in those long-horizon route selections they mentioned as unchanged.

Taro: That points toward where future work needs to focus, and I think it means we should expect more research into combining this local safety reasoning with things like adaptive subgoal selection or spatial memory for longer routes, because the current model seems focused on short-horizon safety and action realizability.

Rosa: So, in simple terms, this Social-WM framework is a system that learns what actions are socially safe by comparing what it thinks it should do against what it can actually execute under physical limitations. It's about making the robot smarter about constraints rather than just faster at following commands.

Dev: And from my side, the real promise is that we get a mechanism to identify inadmissible candidate actions before they even leave the planner, which should help us manage failure modes proactively during high-speed execution.

Taro: The implication for the broader world is that robots can navigate social spaces with a much more nuanced understanding of physical constraints and social etiquette, moving beyond just following pre-programmed paths to actually behaving in a way that respects surrounding actors.

Rosa: Exactly, so we're looking at systems capable of proposing actions and then imagining them through the lens of safety constraints, which opens up new possibilities for collaborative robots in less structured environments.

Dev: I hope the latency remains tight during deployment; if this framework can truly handle real-time demands without excessive processing time, it could move from research to practical applications much faster than we've seen before.

Taro: The system’s ability to reason about action realizability suggests that future autonomy will require models that explicitly model physical and social limitations in their planning, not just abstract goal attainment.

Rosa: That really puts the focus on how robots interact with people, moving toward a more constrained, yet safe, form of social navigation.

Dev: So we've covered the summary of Social-WM: Safety-Aware Latent World Models for Robot Social Navigation. Now that we've seen what they did, let's talk about what this framework actually means for deployment and the robot’s behavior in the real world.

Conclusion: Rosa: So, to wrap up this discussion on Social-WM, we've seen how this system learns safety by comparing what it expects to happen versus what actually happens when a robot tries a command in a social setting.

Dev: Yeah, and the real focus there was on keeping that loop rate tight; I mean if the latency gets too high, all that predictive power doesn't matter in real-time operations.

Taro: From an autonomy standpoint, it’s fascinating how they frame safety as this measurable discrepancy between a nominal plan and a realizable one, which gives us a concrete signal to trust or reject an action.

Rosa: Exactly, so the whole point of Social-WM is building that awareness into the latent model itself so the robot knows when to pull back from an ambitious move.

Dev: And I'm still thinking about how it handles those timeouts we discussed; if it can't plan fast enough, does that safety signal become unreliable under time pressure?

Taro: That’s a huge question for me; if the world misbehaves unexpectedly, does this learned discrepancy hold up when the environment violates its training assumptions?

Rosa: It seems like they aimed to show a framework that could operate in varied social contexts, suggesting it might be more robust than systems tuned for just one specific lab setup.

Dev: I wonder how long this kind of model can reliably function outside of a highly controlled simulation before we see significant degradation in performance?

Taro: That points directly to the need for better long-horizon planning and adapting to unseen environmental behaviors, which is definitely where the next big challenge lies for autonomy research.

Lehigh University

cs.RO

Submitted: 2026-09-30

Updated: 2026-10-01

Comments: 9 pages, 5 figures

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 82/100

The gist: Safe social navigation requires a robot to anticipate not only the future consequences of its actions, but also whether a nominal action can actually be executed under surrounding physical and social

Key concepts

Nominal Action vs. Realizable Action
A nominal action is the command the robot intends to issue, while a realizable action is what the robot can actually execute given its physical environment and social context. The model learns to find the difference between these two, which signals potential safety issues when heading toward obstacles.
Realizable Inverse Dynamics
This component predicts the actual motion that occurred based on the robot's observed poses, rather than just following a command. By comparing this predicted actual motion with what was commanded, the model learns to associate latent state transitions with physically achievable movements.
Safety-Relevant Signal
The discrepancy calculated between nominal and realizable actions serves as a safety signal. A large discrepancy indicates that the intended action is likely blocked or inadmissible due to physical or social constraints, helping the robot avoid unsafe situations before acting.
Latent World Model Planning
This framework uses a compressed representation of the environment (latent state) to plan actions. It predicts future latent states based on history and nominal actions, enabling fast planning that evaluates potential futures in a simplified, efficient space.

Terminology

Summary

Safe social navigation requires a robot to anticipate not only the future consequences of its actions, but also whether a nominal action can actually be executed under surrounding physical and social constraints.

The gist

Social-WM learns these safety-relevant consequences directly through action-conditioned future prediction, where the target is the actual observed future following each command.

How it works

Social-WM introduces an efficient latent world-model planning framework trained from egocentric RGB video sequences to jointly learn future dynamics and action realizability. The core mechanism involves distinguishing between a nominal action and the realizable action, which reflects physical and social safety constraints. A key observation is that a nominal forward action may be fully executed in free space, but needs to be constrained when heading towards a pedestrian or obstacle.

The framework operates through several integrated components:

  1. Visual Representation: A frozen DINOv2-small encoder is used to encode each RGB observation into spatial patch tokens, which are then pooled into latent representations, denoted as the latent state vector.

  2. Action-conditioned Future Prediction: A transformer predictor, Fϕ, predicts the next latent feature based on the history and nominal actions: zˆt+1 = Fϕ(zt−T +1:t, at−T +1:t). Crucially, although conditioned on nominal actions, the target is the actual observed future, allowing the model to learn safety consequences from real-world transitions.

  3. Realizable Inverse Dynamics: To explicitly ground the dynamics in physically achievable motion, a realizable inverse-dynamics objective is introduced. This model predicts the action actually executed, denoted as aˆ˜t = Iψ zt, mlt, where it is supervised using the realizable action a˜t recovered from consecutive robot poses rather than the issued command.

Planning and Safety Evaluation

At deployment, candidate actions are generated using a lightweight generative model (a CVAE) conditioned on latent and action history: a(k)t:t+H−1 = Gθ zt−T +1:t, at−T +1:t, ϵ(k). These candidates are then rolled forward through the frozen world model to produce an imagined latent trajectory. The safety signal is derived by measuring the discrepancy between the nominal and predicted realizable actions: c(k)safe = (1/H) Στ=t to H-1 τ a(k)τ − Iψ zˆ(k)τ, zˆ(k)τ+1 − zˆ(k)τ. A small residual indicates consistency, while a large residual indicates that the candidate is likely to be blocked or otherwise inadmissible.

Candidates satisfying "c(k)safe < τ" form the admissible set. The final action is selected by minimizing a goal-conditioned cost function, where the cost depends on whether it is a position goal (closest distance to pg) or an image goal (similarity between imagined latent and goal latent). This closed-loop process allows for propose–imagine–evaluate–select planning using both goal progress and action realizability.

Key Contributions

The paper makes three main contributions:

  1. Formulating the discrepancy between nominal and realizable actions as a safety-relevant signal for learning action-conditioned latent dynamics in social navigation.

  2. Introducing realizable inverse dynamics, which associates latent transitions with actually achievable motion and identifies admissible candidate actions from imagined futures before execution.

  3. Developing an efficient goal-independent latent worldmodel planner that supports both position- and image-goal social navigation and generalizes from SocialHM3D to SocialMP3D.

Experimental Results

On the benchmark Social-HM3D, Social-WM achieved a success rate (SR) of 63.77% and reduced human collisions (HColl) to 21.67%, a 44.6% relative reduction over NavThinker. Under zero-shot transfer to Social-MP3D, the model maintained competitive performance, achieving an SR of 54.15% and HColl of 23.34%. Furthermore, transitionlevel diagnostics showed that blocked transitions produced much larger nominal–realizable residuals than free transitions (AUROC of 0.961), confirming that the learned discrepancy is strongly associated with constrained execution. The model's efficiency is maintained because planning operates in latent space, requiring only 51 ms per step, which is faster than Diffusion Policy and LeWM-style MPC.

Limitations

Despite substantial improvements in safety, the paper notes that the overall success rate improves only modestly because many remaining failures are timeouts. The current world model primarily reasons about short-horizon safety and action realizability, meaning long-horizon route selection remains unchanged, suggesting future work should focus on combining local safety reasoning with route-level replanning, spatial memory, and adaptive subgoal selection.

Improvements for AI systems

Here are the specific improvements that can be made to existing AI systems by implementing the Social-WM framework, along with a detailed description of what these improved systems will be able to do:


  1. The core improvement is the introduction of an explicit safety signal derived from the discrepancy between a nominal action and its realizable counterpart, achieved through a combination of action-conditioned latent prediction and realizable inverse dynamics.

  2. The improved AI system (Social-WM) can perform closed-loop, real-time planning by integrating two distinct foresight mechanisms:

Ease of Planning and Safety Guarantee: Unlike systems that only predict the nominal future or rely on external tracking (like pedestrian state estimation), Social-WM generates candidate action sequences in latent space and immediately evaluates their physical/social feasibility. This allows the robot to select actions based on a combined cost function of goal progress and guaranteed realizability.

  1. Enhanced Safety Constraint Handling: The system explicitly learns that certain nominal commands are physically or socially impossible (e.g., attempting to move into a pedestrian's space). It can actively reject unsafe action chunks before execution, leading to significantly reduced human collisions (e.g., reducing H-Coll by over 40% in benchmarks) and increased personal-space compliance (PSC).

  2. Generalization Across Navigation Goals: The latent world model is designed to be goal-independent, supporting both position-goal navigation (reaching a specific coordinate) and image-goal navigation (reaching a specific visual scene), without requiring explicit pedestrian tracking or privileged human state information during deployment.

  3. Robustness via Offline Training and Zero-Shot Transfer: The system can be trained entirely from egocentric RGB video sequences offline, learning the complex dynamics of social interaction. It demonstrates strong generalization capabilities, maintaining competitive performance on unseen environments (zero-shot transfer), without requiring online reinforcement learning or explicit tracking modules for pedestrians.

  4. Multi-Modal Goal Guidance Capability: The planner can adapt to different forms of goal specification—either a final target position or an image goal—by adjusting the latent cost function accordingly, allowing for flexible task execution in social settings.

  5. Improved Latent Representation Quality: By leveraging frozen, high-quality visual encoders (like DINOv2 spatial tokens), the system ensures that the learned latent space preserves crucial geometric information (pose decodability) and supports effective goal ranking, leading to more accurate planning decisions based on the imagined future.

Abstract

Safe social navigation requires a robot to anticipate not only the future consequences of its actions, but also whether a nominal action can actually be executed under surrounding physical and social constraints. We present Social-WM, an efficient latent world-model planning framework trained from egocentric RGB video sequences. Our key observation is that social-navigation experience contains a systematic discrepancy between the nominal action and the realizable action: a nominal forward action may be fully executed in free space, but needs to be constrained when heading towards a pedestrian or obstacle. Social-WM learns these safety-relevant consequences directly through action-conditioned future prediction, where the target is the actual observed future following each command. We further introduce a realizable inverse-dynamics objective that associates observed latent transitions with the action actually realized rather than the nominal one. At deployment, candidate actions are imagined through the latent world model, and the inverse dynamics model estimates their realizability; nominal--realizable discrepancy then provides a safety signal before execution. The learned dynamics and realizability model remain goal-independent and support both position- and image-goal navigation. On Social-HM3D, Social-WM achieves 63.77% success while reducing human collisions to 21.67%, and maintains strong performance under zero-shot transfer to Social-MP3D, without explicit pedestrian tracking, privileged human state, or online reinforcement learning.

Sources

Related papers