Probabilistic Plan Legibility with Off-the-shelf Planners
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "Probabilistic Plan Legibility with Off-the-shelf Planners".
Rosa: Legible planning is addressed by proposing a method to generate plans that best disambiguate their goals from other candidates from an observer’s perspective,
Dev: First, who's behind it and why it matters.
Title and authors: Rosa: So, we're looking at "Probabilistic Plan Legibility with Off-the-shelf Planners," and it sounds like the title itself suggests a focus on making plans understandable from an outside viewpoint. I was reading about Michele Persiani and Thomas Hellstrom being the authors, which is interesting because they are working on something that connects planning to how humans perceive those plans.
Dev: Yeah, the title hints at using probability to figure out which plan is the clearest one for someone looking at it from a distance. It’s about legibility in a probabilistic setting, which I think means they aren't just looking for any plan; they're trying to find the one that maximizes clarity based on what we can observe.
Taro: From an autonomy standpoint, if plans are meant to be understood by someone else—like a human collaborator—then making that understanding probabilistic is key because real-world observation is never perfect. It allows the system to account for uncertainty in how the observer interprets the plan.
Rosa: Exactly, Taro; it’s not about finding one single "best" plan universally, but rather a family of plans where we quantify how well each one helps distinguish its intended goal from all the other possibilities in a set. It seems like they are tackling the problem of implicit communication between an AI and a human.
Dev: And that distinction between finding the mathematically perfect plan versus finding one that is practically legible for a collaborator really matters for deployment, Rosa; we need to know if this works reliably in noisy environments where observation is limited.
Taro: I wonder how robust this probabilistic definition holds up when the task space itself has many competing goals, especially when the constraints of the PDDL domain make it impossible to generate a perfectly legible plan at all.
Rosa: That’s a big question, Taro; the authors actually mention that in some cases it might be impossible to generate a sufficiently legible plan due to constraints in the task space, which means we have to deal with those failures head-on.
Dev: So, if we can't get perfect legibility sometimes, how does the algorithm handle those instances where generating a clear plan just isn't feasible within the given planning rules?
Taro: Well, they introduce a concept called n-legibility to address that, which looks at legibility across different prefixes of the plan. It suggests that even if the whole thing isn't clear, maybe each small step helps build up some discernible pattern for the observer.
Rosa: That makes sense; so it’s not an all-or-nothing situation regarding plan clarity; there are partial successes we can measure using this n-legibility metric. It seems like they are building a way to quantify that partial success.
The paper's summary: Dev: Moving on to the actual summary of "Probabilistic Plan Legibility with Off-the-shelf Planners," the core idea is proposing a method for legible planning in arbitrary PDDL domains without needing to build custom planners from scratch. They extend earlier work on legibility into classical planning and introduce a probabilistic way to define what it means for a plan to be legible based on observations and prior beliefs about the task goals.
Rosa: It’s important that they emphasize that this method can work across various PDDL domains because they don't require constructing ad-hoc planners, which is a huge practical win for us; we can use existing tools like Fast-Downward and just plug in their algorithm to get results quickly.
Taro: The second crucial part of the summary is how they connect the planner’s task space to the observer’s task space through a second-order theory of mind function, which acts as a transformation T that helps us estimate how an observer will interpret our actions.
Dev: That theory of mind connection is where it gets deep; it allows the planner to essentially model what the observer believes about itself or the plan, shifting the legibility computation into the observer's perspective model, which they call ' O’s model of R.
Rosa: So, in simple terms, they are saying we need a mental bridge—a second-order theory of mind—to translate our internal planning steps into a language that makes sense to the human collaborator trying to follow us. It’s about modeling the inference process itself rather than just the execution sequence.
Taro: That seems like it solves a major problem in human-robot teaming because we're not just outputting actions; we are outputting intentions that are structured around how someone else thinks, which is a necessary step for implicit communication.
Dev: And the paper lays out an algorithm using these concepts to actually produce those plans, starting by finding candidate plans and then transforming them through the theory of mind function before selecting the one that maximizes legibility while balancing it against plan length.
Rosa: So, they’ve got a concrete procedure here: generate candidates, map them to the observer's view using T, and then select the plan that gets us closest to legibility while keeping a leash on how long the resulting plan is. It sounds like a very structured approach for practical implementation.
Taro: The mention of n-legibility as a weighted average across prefixes shows they’ve thought about how to measure progress incrementally, which is useful for monitoring the planner during long planning horizons.
The paper's improvements: Rosa: Now let's talk about the specific improvements they suggest in "Probabilistic Plan Legibility with Off-the-shelf Planners." They propose using a procedure that starts by finding a diverse set of candidate plans, then transforming those instances through the theory of mind function T to get the observer's perspective, and finally selecting a plan pi that maximizes legibility while applying a regularization term involving gamma.
Dev: The key improvement I see is this regularization factor gamma; they explicitly introduce it to balance the trade-off between maximizing legibility and keeping the plan length manageable, which addresses the finding that legibility is often inversely correlated with efficiency.
Taro: That balancing act seems really important because if we only maximized legibility without gamma, we could end up with plans that are three or even six times longer than the optimal ones for the same task, as some of their empirical findings suggest in domains like blocks-world or logistics.
Rosa: It’s a huge practical improvement because it acknowledges that in real-world scenarios, you can't always afford to be perfectly legible if it means your robot takes way too long to execute the plan. This factor allows us to tune that trade-off based on the situation we’re in.
Dev: The paper also shows improvements in performance by using a generalized measure called-legibility, which is defined as a weighted average of legibility across all prefixes of the plan, allowing for a more nuanced evaluation than just looking at the final plan alone.
Taro: And this structure allows them to apply this framework to any PDDL domain using off-the-shelf planners, which broadens the applicability significantly beyond just one specific type of robotic task. It makes it more generalizable across different problem types.
Rosa: So, the improvements boil down to creating a systematic way—using that theory of mind connection and that regularization term—to generate plans that are both understandable and reasonably efficient for deployment in diverse robotic settings.
Conclusion: Dev: To wrap up on this paper, the main implication is that we can now produce plans using existing PDDL planners by integrating probabilistic goal recognition with a second-order theory of mind. This allows the planner to generate plans that are understandable to collaborators implicitly, which is really important for human-robot teaming scenarios where there’s no explicit communication channel.
Rosa: It confirms that legibility is achievable, but it also firmly establishes that it’s inherently a trade-off with plan cost; we can't just get high legibility without making the plans longer, and this relationship depends heavily on the specific domain and the theory of mind model used.
Taro: From my view, this work moves us closer to autonomous agents that can implicitly communicate their intent by producing these legible plans, which is a step toward more intuitive interaction in complex environments where explicit signaling isn't possible.
Dev: I agree with Taro; the empirical findings show that for goal-specific actions, dropping part of the observations often helps legibility more than in domains with universally applicable actions. Plus, the regularization factor gamma is clearly a necessary tool to keep those plans from ballooning in length unnecessarily.
Rosa: So, to summarize "Probabilistic Plan Legibility with Off-the-shelf Planners," we have a framework that uses probabilistic goal recognition and theory of mind to generate plans understandable by an external observer, provided we manage the efficiency trade-off correctly.
Taro: It’s definitely a solid contribution because it provides a methodology for generating plans that prioritize interpretability without needing entirely new planning tools for every specific problem.
Dev: Yeah, it gives us a usable toolset right now by leveraging off-the-shelf PDDL planners and making the legibility calculation more robust through the probabilistic framework.
Umea University
cs.RO, cs.AI
Submitted: 2026-09-04
Updated: 2026-09-04
Comments: Accepted at the 9th ICAPS Workshop on Planning and Robotics. ICAPS 2021
Code: https://github.com/pucrs-automated-planning/goal-planrecognition-dataset
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 80/100
The gist: Legible planning is addressed by proposing a method to generate plans that best disambiguate their goals from other candidates from an observer’s perspective, which matters because it provides a
Key concepts
- Probabilistic Formulation of Legibility
- Legibility is defined by how well a plan's observations distinguish its true goal from other possible goals. Mathematically, it measures the similarity between the distribution of observations resulting from a plan and the prior beliefs about possible goals. A highly legible plan produces observations that clearly point to one specific goal.
- Theory of Mind for Connecting Task Spaces
- This is a second-order reasoning mechanism that allows the planner to model how an observer interprets its actions. It acts as a function transforming the planner's plan instance into an instance viewed from the observer's perspective, enabling the planner to estimate what the human collaborator believes about its intentions.
- Legibility Gain (Lgain) and Cost Gain (Cgain)
- These performance measures quantify the trade-off between clarity and efficiency. Lgain measures how much better a plan is at distinguishing its goal compared to other generated plans, while Cgain compares the length of the most legible plan to an optimal, short plan. The paper shows that increasing legibility often requires accepting a longer plan.
- Regularization Factor ($\gamma$)
- This parameter is used in the final planning step to balance two competing objectives: maximizing legibility and minimizing plan cost (length). Because achieving perfect legibility can lead to excessively long plans, $\gamma$ is introduced to penalize longer plans, ensuring the generated plans remain practical for execution.
Terminology
Summary
Legible planning is addressed by proposing a method to generate plans that best disambiguate their goals from other candidates from an observer’s perspective, which matters because it provides a mechanism for implicit communication in human-robot teaming scenarios by allowing autonomous agents to produce plans understandable to collaborators.
Probabilistic Formulation of Legibility
The paper adopts a probabilistic approach where plan legibility is based on the observations that plans lead to and prior beliefs over the task. At any moment, there is a set of possible goals, denoted as G = g, gˆ 0, g1, …, gn. To be legible, a plan must produce observations that make its true goal gˆ easily discernible from the other goals in G. This is mathematically defined by: legibility(π, g) = H(Pg(G), P(Gπ)P(Π)), where H is a similarity function of two probability distributions. The most legible plan, πlegible, is then defined as: πlegible = argmax π∈Πgˆ legibility(π, gˆ).
Theory of Mind for Connecting Task Spaces
A crucial contribution is the requirement for a second-order theory of mind to connect the planner’s and the observer’s task spaces. This theory of mind acts as a function T which transforms the plan instances utilized by the planner (Ξpl) into an instance from the observer's perspective (Ξ′pl): T = Tpl ◦ Tobs: Ξpl → Ξobs → Ξ′pl. This second-order reasoning allows the planner to estimate how the observer interprets its actions, effectively allowing it to access the observer’s beliefs about itself. The paper reformulates legibility as being computed inside Ξ′pl, which is O’s model of R.
Algorithm for Producing Legible Plans
The paper proposes an algorithm to produce legible plans using off-the-shelf PDDL planners. This procedure involves several steps:
-
Πgˆ ← DIVERSE-PLAN(Ξ, g, k ˆ) (Finding a set of candidate plans)
-
Ξ′ ← T(Ξ) (Transforming the planning instance using the theory of mind T)
-
gˆ′ ← T(ˆg) and G′ ← T(G) (Transforming the goal and its set into the observer's perspective)
-
π ← argmax π∈Πgˆ ˆ T(π)–legibility − γπ (Finding the plan that maximizes legibility while regularizing towards cheap plans)
Performance Measures and Trade-offs
The performance of the proposed method is evaluated using two main measures: Legibility Gain (Lgain) and Cost Gain (Cgain). Lgain corresponds to the ratio between the legibility of the optimal plan π and that of the most legible plan among k generated plans: Lgain = nˆ-legibility(πlegible) / nˆ-legibility(π). Cgain indicates how costly (in plan length) the legible plan is compared with the optimal plan: Cgain = πlegible / π. The results generally show that plan legibility is a trade-off with plan efficiency, however, not all planning domains allows to increase legibility in the same way and a regularizing factor to balance legibility and efficiency was proved necessary.
Empirical Findings
Statistical evaluation over several PDDL domains revealed that planning domains with goal-specific actions will benefit more from dropping part of the observations rather than domains with universally applicable actions.
Furthermore, the results demonstrated that the regularization factor γ plays an important role,
as in some cases, without regularization legible plans could reach a length up to 3-6 times higher than the optimal plans for the same instances
in certain domains like blocks-world or logistics. The findings also show that legibility is positively correlated with cost, indicating that it was always possible to increase legibility in exchange of making plans more lengthy.
Conclusion
The paper successfully proposed a procedure to compute legible plans using off-the-shelf PDDL planners by integrating probabilistic goal recognition with a second-order theory of mind. This approach allows the planner to know how its actions are perceived by the observer, facilitating implicit communication in human-robot teaming contexts without requiring external input like scene labeling. The statistical tests confirm that while legibility is achievable, it is inherently a trade-off with plan cost, and this relationship is domain and theory of mind dependent.
The gist: Legible planning requires a second order theory of mind to model how the observer infers the planner’s goal, leading to an algorithm that produces plans maximizing legibility while regularizing for plan efficiency.
Table 1: Average instance measures over the tested planning domains.
Improvements for AI systems
As a fastidious and diligent AI researcher, I have analyzed the provided paper, Probabilistic Plan Legibility with Off-the-shelf Planners,
by Persiani and Hellstrom. The core contribution is a framework for generating plans that are understandable from an external observer's perspective—a concept termed legible planning
—by modeling the observer's task space using a second-order theory of mind.
Here are the specific improvements to AI systems and what the resulting improved systems can achieve, based on this research:
)1. Implementation of a Second-Order Theory of Mind (Theory of Mind - ToM) for Planning:
The system will be augmented with a mechanism to estimate the observer's mental model of the planner. Instead of assuming shared task models, the system will explicitly model how an observer (e.g., a human collaborator or another agent) infers the planner's goals based on their own beliefs.
)2. Probabilistic Goal Recognition Integration:
The planning algorithm will incorporate a probabilistic goal recognition model, specifically utilizing cross-entropy and Boltzmann distributions to quantify the likelihood that a set of observations leads to a specific goal distribution. This allows the system to move beyond simple optimality (cost minimization) toward finding plans whose resulting observation sequences are most likely to lead an observer to correctly identify the true objective.
)3. Legibility-Aware Planning Algorithm:
A new planning procedure, based on Algorithm 1, will be implemented using off-the-shelf PDDL planners. This procedure will iteratively generate diverse plans and evaluate them using a generalized measure of n̂-legibility (a weighted average of prefix legibilities). Crucially, this algorithm includes a regularization factor (γ) that explicitly balances plan legibility against plan efficiency (cost), allowing the system to choose between highly legible but costly plans and optimal but opaque ones.
)4. Domain-Agnostic Planning for Legibility:
The framework is designed to be applied to arbitrary PDDL domains without requiring the construction of ad-hoc planners. The improved system can integrate existing, well-known planning tools (like Fast-Downward) by transforming their internal representations via the Theory of Mind function (T), ensuring that legibility is achievable across a wide variety of planning problems.
The resulting improved AI systems will be capable of:
-
A robot or autonomous agent can produce plans that are not only optimal in terms of execution cost but are also
explainable
orintuitive
to a human collaborator (e.g., a commander). This is achieved by ensuring the plan's sequence of actions strongly suggests the true goal among several competing possibilities. -
In Human-Robot Teaming (HRT) scenarios, the robot can implicitly communicate its intent by producing legible plans that guide human decision-making, even when there is no explicit communication channel between them. The system will prioritize goals that are
best discriminated
by the observer's inferred model of the world. -
The system will be able to adapt its planning strategy dynamically based on the required level of legibility versus efficiency, using the regularization factor (γ) to tune this trade-off for different operational contexts (e.g., prioritizing high legibility during critical safety maneuvers).
-
The system can operate effectively in environments where the observer has a potentially inaccurate or incomplete mental model of the agent's task, as it explicitly models this discrepancy through its second-order theory of mind, enabling better performance even when model reconciliation is not possible.
Abstract
Legible planning is the creation of plans that best disambiguate their goals from a set of other candidates from an observer's perspective. In this paper we propose a method for legible planning for arbitrary PDDL domains, by extending previous research on legibility to classical planning without requiring to construct ad-hoc planners. We also discuss how the observer perspective may be estimated through a second order theory of mind that connects the planner's and the observer's task spaces. Our solution can for example be deployed in human-robot teaming scenarios, where an autonomous robot in a team can implicitly communicate its goal by producing legible plans. We present benchmark results on several PDDL planning domains. Our results generally show that plan legibility is a trade-off with plan efficiency, however, not all planning domains allows to increase legibility in the same way and a regularizing factor to balance legibility and efficiency was proved necessary.
Sources
- Plan Explanations as Model Reconciliation: Moving Beyond Explanation as Soliloquy
- Explainable Planning
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving