Rethinking Legibility in Social Robot Hallway Navigation: Impact of Intent Representation and Human Distraction

arXiv:2609.40158 · cs.RO, cs.HC · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Rethinking Legibility in Social Robot Hallway Navigation".

Dev: Legibility in social robot navigation is crucial for ensuring human safety and smooth coordination in dynamic, constrained environments where human attention can be divided.

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: Welcome everyone to our show today as we look at a really interesting paper on arXiv titled "Rethinking Legibility in Social Robot Hallway Navigation: Impact of Intent Representation and Human Distraction." This research dives into how robots should move in crowded spaces, especially when people are paying attention to other things. We'll discuss what the authors found regarding intent representation and how human distraction affects that legibility.

Dev: I’m ready for it, Rosa. Given my background in control systems, I’m particularly interested in whether these proposed intent representations translate into actually stable and low-latency movement when we put them into a Model Predictive Control framework, which is what this paper seems to be using.

Taro: From an autonomy standpoint, I want to hear about what happens when the environment gets messy; if the robot’s legibility cues don't hold up when pedestrians are distracted or the situation changes unexpectedly, how resilient is that motion strategy?

Rosa: Exactly, Taro. So basically, this paper looks at a major gap in existing research where we often only think about static observers and not dynamic ones in crowded hallways. The core thesis here is that how a robot communicates what it’s going to do—its intent—matters a lot more than just where it’s trying to go when humans are actually interacting with it.

Dev: So, the paper claims that moving beyond simple destination-based cues towards interaction-level coordination can be more effective in crowded settings, even when people aren't looking directly at the robot. That sounds like something we need to test on our hardware loop rates.

Taro: I’m interested in the specific representations they tested; are we talking about abstract concepts, or do they have concrete ways to encode that interaction-level coordination into the robot's actual trajectory planning? If it’s too abstract, it won't work when things go wrong.

Rosa: The authors compared several different ways to encode intent, ranging from simple goal-based legibility to more complex ideas like Social Momentum, and they found some really interesting trade-offs in hallway navigation. They specifically looked at how these different representations perform under conditions where human attention is divided during the interaction.

Dev: So, if I understand correctly, the paper suggests that some forms of intent representation are more robust than others when we can't rely on a person being perfectly focused on us? That has implications for our failure modes when we encounter unexpected human behavior or distraction.

Paper summary: Taro: If the paper shows that dynamic adaptation in intent—like selecting a passing side based on predicted human choice—works better than a fixed intent, that tells us we need more sophisticated real-time decision-making in our autonomy stacks to handle unpredictable social dynamics.

Rosa: That’s the big picture there. The whole point is to see how these different ways of encoding intent shape both how good the navigation actually is and what people feel when they observe it, especially when their attention is divided. It sets up a real challenge for designing robots that are not just safe, but also socially smooth in busy environments.

Dev: So the paper essentially argues that for social robot navigation in constrained settings like hallways, we need to focus on interaction-level cues rather than just destination-based ones, and that these cues should adapt based on what the human is actually doing or paying attention to.

Taro: And it suggests that even when those attention cues aren't perfectly clear subjectively, the objective measures of coordination still show a benefit from having a legible strategy in place. That persistence under distraction is something I think is important for real-world deployment because in reality, people are almost always distracted.

Rosa: Precisely, Taro. The authors’ conclusion emphasizes that adaptive strategies reinforce the legibility effect and that this coordination benefit continues even when subjective human impressions become less sensitive to the differences between strategies. This suggests we should build systems that can dynamically adjust their intent signaling based on observed human context during movement.

Dev: From a control engineering view, if the paper confirms that dynamic selection of passing sides, like in DPL or SM, yields better Human Average Acceleration metrics objectively, then our MPC framework needs to be able to incorporate those real-time predictions about human choice into its cost function for trajectory generation.

Taro: If the system needs to dynamically adapt its intent based on predicted human behavior during the interaction, that means our planning loop has to become much more tightly coupled with perception of the social context, not just the immediate geometric constraints of the hallway.

Rosa: So when we look at these results in "Rethinking Legibility in Social Robot Hallway Navigation: Impact of Intent Representation and Human Distraction," it’s clear that moving from simple destination-based planning to interaction-level intent representation is key for navigating crowded spaces safely and smoothly.

Dev: And the finding that adaptive strategies outperform fixed ones, even when people are distracted, points directly toward building more robust online estimation and adaptation capabilities into our navigation algorithms to handle those real-world human distractions effectively.

Paper summary: Taro: I think the biggest implication is that we need to design autonomy where the robot doesn't just follow a pre-set path but actively tries to maintain a legible social presence by constantly adjusting how it signals its intent based on what it senses about the human partner.

Rosa: That really puts things into perspective, Taro. The paper suggests that for these robots, legibility isn't just about moving smoothly in a vacuum; it’s about managing the perception of coordination in a messy hallway where attention shifts around.

Dev: I think we should be looking at how to implement those dynamic representations within our MPC framework to ensure low latency and reliable execution, even when the human input data is noisy due to distraction.

Taro: If we can figure out how to reliably estimate the human’s momentary focus or intent through sensory input, then that adaptive legibility becomes a powerful tool for achieving safer and more natural social interactions in dynamic environments.

Rosa: It sounds like this paper really pushes us toward integrating social context directly into the robot's motion planning decisions, which is exactly what we need to consider when we move these systems out of the controlled lab environment and into actual public spaces.

Dev: I’ll be checking how long these control loops can maintain that level of adaptability under sustained real-world operational stress; that’s where I see the biggest engineering hurdle for us right now.

Taro: That's a fair point, Dev. The future work they suggest about automatically balancing functional efficiency against legibility using an attention parameter lambda is precisely the kind of adaptation we need to explore for field deployment.

Rosa: So when we talk about the title "Rethinking Legibility in Social Robot Hallway Navigation: Impact of Intent Representation and Human Distraction," it really captures that this isn't just a technical tweak, but a fundamental rethinking of how robots should communicate their intentions in complex social settings.

Dev: It seems like the paper points us toward developing systems where intent is not just a static output from the planner, but something actively negotiated based on real-time interaction feedback and environmental awareness.

Taro: That means we’re not just building better path planners; we're building robots that are better at understanding and responding to the social state of their environment in real time.

Rosa: And that’s what makes this paper so compelling for all of us, showing how subtle shifts in intent encoding can lead to measurable improvements in both objective coordination and how people actually perceive the robot's behavior during a navigation task.

Conclusion: Rosa: So, to wrap up this discussion on "Rethinking Legibility in Social Robot Hallway Navigation: Impact of Intent Representation and Human Distraction," we've seen how changing how a robot signals its plan really affects both its performance and how people react when they're distracted.

Dev: Yeah, it’s clear the core focus here is moving beyond just where the robot is going to making sure that intention is actually legible to people interacting with it in busy hallways.

Taro: I think what stuck with me was how they showed that even when humans are distracted, those legibility cues still help resolve conflicts between robots or between robots and people.

Rosa: Exactly, Taro, and the authors found that using dynamic intent representations, like adapting the passing side based on predictions of human choice, works better than fixed ones.

Dev: From a control standpoint, if those adaptive strategies are producing lower Human Average Acceleration objectively in controlled tests while maintaining reasonable loop rates for MPC to handle them, that’s solid data.

Taro: But I wonder how resilient those systems really are when the world gets unexpectedly chaotic; does this hold up when the human partner suddenly changes their behavior drastically?

Rosa: That’s a fair question, Taro, and the authors themselves pointed out that while objective measures showed benefits under distraction, subjective impressions became less sensitive to algorithmic differences.

Dev: That means we need to figure out if those objective gains translate into reliable performance outside of controlled lab conditions where we can strictly script the interactions.

Taro: If we can transfer these concepts to real-world deployment, it suggests that robots will be much better at navigating crowded public spaces without needing a perfectly focused human partner.

Rosa: Exactly, and this work opens up a lot of ideas for how we design social robots that are not just efficient but also socially intuitive in messy environments.

Dev: We need to think about how to actually bake that dynamic adaptation into the MPC cost function so the system can handle those real-time adjustments smoothly without introducing latency issues.

Taro: So, the next step seems to be developing systems where robots can estimate human attention levels dynamically and adjust their legibility signals accordingly.

Rosa: That sounds like a really exciting direction for future research, Dev; we should definitely keep an eye on how those online estimation models evolve.

PRANAV GOYAL, ANDREW STRATTON, CHRISTOFOROS MAVROGIANNIS

University of Michigan at Ann Arbor

cs.RO, cs.HC

Submitted: 2026-09-30

Updated: 2026-09-30

Comments: 24 pages, 6 figures

Code: https://github.com/fluentrobotics/Legible_MPPI

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 88/100

The gist: Legibility in social robot navigation is crucial for ensuring human safety and smooth coordination in dynamic, constrained environments where human attention can be divided.

Key concepts

Intent Representation
This refers to how a robot communicates its plan or goal to a human. The study compared different methods, such as using the robot's final destination (goal-based) versus cues about which way it intends to pass (passing-side). The paper found that interaction-level cues are better than just stating a fixed destination.
Legible Motion
This is the quality of a robot's movement that makes its intentions clear to humans. The researchers tested several types, including goal-based and dynamic passing side legibility. They found that certain forms of legible motion lead to smoother human movements and better coordination during navigation.
Divided Attention
This occurs when a person is trying to do two things at once, like navigating a hallway while simultaneously listening to instructions. The study examined how distraction affects whether humans can correctly interpret the robot's signals. They found that even when distracted, legible motion still helps resolve conflicts.
Dynamic Adaptation
This involves a robot changing its behavior in real-time based on what it observes from the human. Dynamic Passing Side Legibility was superior because it automatically adjusted which way to pass based on predictions of human choice, showing that intent should evolve during interaction.

Terminology

Summary

Legibility in social robot navigation is crucial for ensuring human safety and smooth coordination in dynamic, constrained environments where human attention can be divided. This work investigates how the choice of intent representation and the level of human attention shape navigation performance and human impressions in hallway settings.

The gist: Legible motion benefits from interaction-level intent representations that support coordination, with some effects persisting even when human attention is divided.

Research Questions

The authors address two central questions motivated by gaps in existing literature:

  1. How does intent representation shape the effectiveness of legible motion in constrained navigation settings? This explores whether destination-based formulations scale to realistic environments or if interaction-level coordination cues are more effective when pedestrians focus on conflict resolution rather than final destinations.

  2. How does human attention influence the impact of legible motion during interaction? This examines how divided attention affects the ability of humans to register communicative signals from robots, suggesting that even under distraction, legibility can shape behavioral outcomes.

Methodology and Framework

The research was conducted using a shared Model Predictive Control (MPC) framework to embed alternative formulations of legible motion. The study involved two controlled user studies in a hallway setting:

  1. Study 1 (N = 45) investigated the role of intent representation, comparing Goal-based legibility (GL), Passing-side legibility (PL), Dynamic passing side legibility (DPL), and Social Momentum (SM) against a Non-legible baseline (NL).

  2. Study 2 (N = 45) examined the effect of pedestrian attention, treating distraction as a central experimental variable to see how it shapes the influence of legible motion on behavior and perception.

Intent Representation Formulations

The study compared several ways to encode robot intent:

- Goal-based legibility (GL): Modeled intent as the robot’s intended destination, which was found to be associated with higher reported workload and less smooth human motion.

- Passing-side legibility (PL): Created artificial subgoals to the left and right of endpoints, conveying the intended side of passage.

- Dynamic passing side legibility (DPL): Dynamically adapted at run time by selecting the passing side associated with the lower predicted probability of being chosen by a human, based on CV predictions.

- Social Momentum (SM): Reasoned about passing directly in the joint interaction space, rewarding trajectories that reinforce the currently emerging passing side using angular momentum.

Impact of Human Attention and Distraction

Study 2 introduced a distraction task where participants simultaneously completed a verbal comprehension task while navigating. The analysis showed:

  1. Subjective ratings of competence and discomfort became less sensitive to algorithmic differences under distraction, as users provided more uniform evaluations.

  2. Objective measures of coordination, such as Human Average Acceleration (Human AA), continued to favor legible strategies (DPL, SM) over the non-legible baseline (NL).

  3. The persistence of coordination benefits suggests that legible motion continued to shape how conflicts were resolved even when its effects were less apparent in participants’ subjective impressions.

Key Findings and Conclusions

The findings demonstrated that:

  1. Passing-side formulations with dynamic adaptation (DPL, SM) produced the most favorable outcomes, consistently outperforming the non-legible baseline (NL) and goal-based legibility (GL). They were rated most competent and elicited lower human average acceleration.

  2. Goal-based legibility (GL) was associated with higher workload and less fluent trajectories, making it unsuitable for close human–robot navigation.

  3. Adaptation strengthens legible motion: dynamic strategies outperformed fixed intent legibility, suggesting that the robot should use pedestrian motion to guide how it selects and updates its passing intent, reinforcing the reciprocal nature of navigation.

  4. Legible motion benefits persist under divided attention; while subjective ratings were less sensitive, objective measures showed that legible motion continued to promote smoother coordination.

Design Implications

The results suggest that SRN systems should prioritize interaction-level representations of intent over destination inference in constrained environments and incorporate adaptive mechanisms that allow expressed intent to evolve during interaction. Treating human attention as an explicit contextual variable is important for designing behaviors that balance efficiency with legibility. The work concludes that the benefit of legible motion depends critically on how intent is represented and adapted to the human, reinforcing that interaction-level legibility shapes coordination behavior, not only impressions.

Limitations and Future Work

The studies were conducted in a controlled laboratory setting with scripted tasks. Future work will focus on transferring systems to field deployment, exploring different encounter geometries and passing norms independently of legibility formulations, and developing systems that automatically adapt the attention parameter (lambda) to dynamically balance functional efficiency and legibility. The research also suggests that the benefits of legible motion may persist even when pedestrians are inattentive, motivating future work on online estimation and adaptation of this factor.

Improvements for AI systems

Based on the provided scientific paper, here are specific improvements that can be made to Artificial Intelligence systems, categorized by the capabilities they would gain:


)1. Improved Social Navigation and Conflict Resolution in Dynamic Environments:

The core improvement lies in shifting robot navigation from purely destination-based planning (Goal-based Legibility, GL) to interaction-level coordination planning.

AI Systems can be improved to:

  • Implement a Model Predictive Control (MPC) framework that incorporates a specific legibility cost function, as defined by Equation (4):

The legibility term motivates the selection of actions maximizing of P (g = g∗ s r:t+1, M).

  • Replace simple goal-vector targeting with representations that prioritize passing-side or interaction-centric intent. This means the AI system should learn to predict and signal its intended side of passage rather than just its final destination.

  • Employ a dynamic adaptation mechanism (like Dynamic Passing Side Legibility, DPL, or Social Momentum, SM) where the robot's intent representation is updated in real-time based on predicted human behavior. For example, if the AI predicts a human will prefer passing on the left based on their trajectory prediction (Equation 3), it should dynamically adjust its own trajectory to signal an intention to pass on that side.

)2. Enhanced Robustness Under Cognitive Load:

The AI system should be designed with explicit awareness of human attentional states, moving beyond the assumption of fully attentive observers.

AI Systems can be improved to:

  • Integrate a variable parameter, modeled by a distraction factor (equivalent to the parameter λ in Equation 5), into its belief update mechanism for human intent.

  • When the system detects high cognitive load (e.g., via external sensors or internal task monitoring), it should automatically increase its legibility signaling intensity. This means when humans are distracted, the AI system should prioritize vivid motion cues that allow rapid disambiguation of immediate intent, even if subjective ratings might not immediately reflect this improvement.

)3. Optimized Trajectory Smoothness and Human Comfort:

The AI's optimization objective function must explicitly balance functional safety/efficiency with legibility to ensure smooth physical interaction.

AI Systems can be improved to:

  • Incorporate a composite cost function (Equation 3):

J = a F J F + a L J L, comprising: a functional cost (J F) that captures human safety, efficiency, and obstacle avoidance as a weighted sum; a legibility cost (J L) that promotes early and confident communication of the robot’s intent.

  • The AI should be trained to minimize trajectory jerk and acceleration (Human AA), ensuring motion is not only safe but also physically comfortable for humans. This directly addresses the finding that non-legible motion leads to more abrupt human movements under low legibility conditions.

)4. Data-Driven Strategy Selection:

The AI system should be able to select the most appropriate legibility strategy based on the current environment and interaction context.

AI Systems can be improved to:

  • Develop a meta-policy that evaluates different intent representations (GL, PL, DPL, SM) in real-time using predicted human motion.

  • The AI should automatically switch between strategies: use GL when the environment is simple or human attention is high; switch to adaptive passing-side methods (DPL/SM) in complex hallway settings or when it detects signs of potential conflict.

)5. Objective Performance Monitoring for Safety Guarantees:

The system needs internal metrics that track coordination rather than just external perception.

AI Systems can be improved to:

  • Track objective measures like Conflict-Resolution Time and Robot Responsibility during interactions. This allows the AI to learn which motion strategies are objectively best at resolving conflicts, independent of immediate subjective user ratings. For instance, it should prioritize maneuvers that lead to faster conflict resolution and a larger share of avoidance for itself.

Abstract

We focus on legible robot motion generation in social navigation settings. Legibility in human-robot interaction (HRI) is often described as the property of robot motion that enables an observer to confidently infer the robot's intent. While mature frameworks exist for generating legible motion in front of static observers, social robot navigation presents a new challenge: the robot must clearly convey its intent while ensuring human safety in dynamic pedestrian environments where human attention is often divided. With the goal of enabling robots to generate legible motion in dynamic and constrained spaces, we investigate how the choice of representation and the level of human attention shape navigation performance and human impressions. Focusing on the ubiquitous and demanding scenario of hallway navigation, we conduct two controlled user studies involving alternative legibility formulations implemented within a shared model predictive control framework. Study 1 (N = 45) investigates the role of intent representation, showing that passing-side legibility, particularly when adaptively updated, leads to smoother human motion and is perceived as more competent and less mentally and physically demanding than destination-based and non-legible baselines. Study 2 (N = 45) examines the effect of pedestrian attention, demonstrating that legible motion allows for smooth human motion even under distraction, even if this is not consistently reflected in subjective ratings. Together, these findings suggest that effective legible motion in social robot navigation benefits from interaction-level intent representations that support coordination, with some effects persisting even when human attention is divided. Code is available at https://github.com/fluentrobotics/Legible MPPI.

Sources

Related papers