PO-PDDL: Learning Symbolic POMDPs from Visual Demonstrations for Robot Planning Under Uncertainty
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "PO-PDDL: Learning Symbolic POMDPs from Visual Demonstrations for Robot Planning Under Uncertainty".
Jane: The paper was written by N/A from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Paper discussion segment 3: Tom: Now that we’ve absorbed the core concepts from the summary, let's pivot and look at the advanced enhancements that "PO-PDDL: Learning Symbolic POMDPs from Visual Demonstrations for Robot Planning Under Uncertainty" suggests. These improvements are where the real revolutionary potential lies.
Jane: If the previous segments described what it *is*, this segment describes how to make it robust enough to actually survive in a chaotic, unpredictable world—which is everything outside of a perfect research lab.
Lu: One critical enhancement focuses on online adaptation, meaning the system shouldn't need a massive retraining effort every time its operating conditions change slightly.
Meng: This addresses what they call 'model drift.' It’s that slow, insidious slide toward inaccuracy that plagues complex AI systems over long periods of use, and the paper proposes continuous correction.
Lalam: Furthermore, they introduce targeted data collection, which is a massive efficiency gain for researchers. Instead of filming hours of random footage to find one specific failure mode, the system directs its own learning efforts.
Jane: The genius here is self-diagnosis: the system doesn't wait for a human to point out its weakness; it pinpoints exactly where its current model assumptions are breaking down during operation.
Tom: It gives the robot meta-cognition—it becomes smart about its own limitations. If it knows, for instance, that grasping highly reflective metal is consistently difficult, it knows precisely that and can ask the demonstrator to provide more data on just that specific material.
Lu: This takes us into generalized failure handling. Instead of simply reporting a generalized "Error," the system is prompted to generate a full diagnostic report.
Meng: That report must explain *why* its initial assumptions failed in the first place, giving us root-cause analysis rather than just an exit code.
Jane: It’s about formal hypothesis testing being built directly into the planning process itself. The robot doesn't just fail; it fails gracefully and analytically.
Tom: For example, it could articulate: "My plan failed because I assumed gravity was constant, but my internal sensor readings suggest a significant variable force is at play."
Lalam: This level of transparent reasoning is the hallmark of true generalization. It allows us to teach abstract concepts, like maintaining stability against varying loads or uneven terrain.
Jane: So, these improvements aren't just quick patches; they are building a self-correcting, continuously learning cognitive layer over the entire architecture.
Tom: If we can build systems that are both capable of complex reasoning and completely transparent about their own limitations, that fundamentally changes our relationship with trust in AI.
Tom: This makes us consider the broader implications for how we interact with these machines, which brings us to a discussion on another major form of intelligence: Large Language Models.
Conclusion: Tom: So, after discussing the comprehensive enhancements and underlying mechanisms of "PO-PDDL: Learning Symbolic POMDPs from Visual Demonstrations for Robot Planning Under Uncertainty," we are left with a truly profound vision of robotic intelligence.
Jane: It forces us to rethink what 'learning' means entirely. We used to think learning was simply data accumulation, but this framework shows that true learning is about formalizing the underlying rules—the grammar—of a task.
Lu: That shift from data collection to rule extraction is massive. It implies that knowledge can be distilled and generalized far more effectively than relying solely on sheer volume of sensory input.
Meng: Exactly. The critical take-away here is the move away from black-box perception and toward verifiable, structured reasoning. This ability to map ambiguity into a probabilistic state is what unlocks generalization in a way pure end-to-end models simply cannot achieve today.
Lalam: And I believe the ultimate implication for engineers building these systems is accountability. By providing this traceable, symbolic logic layer over the raw perception data, we give operators a level of understanding they haven't had before—they know *why* the robot chose its path.
Jane: It’s more than just making robots perform tasks; it’s about building systems that can logically reason about their own uncertainties and limitations while performing them. That self-
Paper discussion segment 3: Tom: We’ve established how this framework learns basic rules from watching demonstrations; now we need to talk about making it actually work outside a lab.
Jane: The central idea for improvement is adaptability. Lab tests are clean environments, but the real world is messy—things change all the time.
Lu: One critical enhancement addresses online adaptation. This means the robot shouldn't need to stop and be fully retrained every single time something minor shifts, like humidity changing a surface’s friction coefficient.
Meng: That continuous refinement tackles what we call model drift. It’s the slow slide toward inaccuracy that plagues many complex systems over time; the robot must learn and adjust right in the moment without human intervention.
Lalam: Furthermore, they propose targeted data collection, which is a huge efficiency boost for researchers and engineers alike. The system doesn't waste time gathering random footage; it self-diagnoses its weakest points.
Jane: This means the human demonstrator’s role changes completely. Instead of providing general examples, they are now directing expertise precisely where the robot’s current model is failing—maybe only focusing on reflective objects, for instance.
Tom: This leads us to generalized failure handling. The enhancements suggest moving past just saying "Error" when something breaks down.
Lu: Instead, the system needs to provide a full diagnostic report explaining *why* its assumptions failed in the first place. It’s about formal hypothesis testing built right into the planning process itself.
Meng: The robot can articulate: "My plan failed because I assumed gravity was constant, but my internal model suggests a significant variable force is at play." This moves us beyond simply knowing *what* went wrong to understanding *why* the physics changed.
Jane: This level of transparent reasoning is what we mean by true generalization. It allows us to teach abstract concepts, like maintaining stability against varying loads, which is far more powerful than teaching it how to push a specific box across a specific floor.
Tom: So, these improvements aren't just quick fixes; they’re creating a self-correcting, continuously learning cognitive layer over the entire system architecture. The biggest shift here is accountability: if the robot can explain *why* it failed, that fundamentally changes how we trust it and how we build with it.
Jane: This ability to map ambiguity into structured, probabilistic states is what unlocks true general intelligence in physical tasks—it gives us a roadmap for understanding the limits of machine knowledge.
Tom: It really solidifies that blending deep learning perception with classical symbolic AI methods isn't just an option; it's a necessity for achieving robust intelligence in physical systems, allowing them to reason about their own inherent uncertainties.
Tom: Now that we’ve seen how much work goes into making robots reason about uncertainty, let’s pivot and look at another form of complex intelligence: Large Language Models.
Conclusion: Tom: So, to wrap up our deep dive into "PO-PDDL: Learning Symbolic POMDPs from Visual Demonstrations for Robot Planning Under Uncertainty," it’s clear this framework represents a significant leap toward truly robust physical intelligence.
Jane: We've seen that by combining the power of visual observation with the structure of symbolic planning, these robots gain an unprecedented level of cognitive awareness.
Lu: What remains most impactful is how the system shifts from merely executing tasks to actively questioning its own assumptions, treating its knowledge base as a continuous hypothesis that must be validated against reality.
Meng: That ability to diagnose failure—to understand *why* the plan failed because an underlying assumption was violated—is fundamentally what separates this approach from older, purely reactive machine learning models.
Lalam: It means that human oversight doesn't become a safety net; it becomes highly specialized diagnostic expertise, focusing precisely on the boundaries of the robot's current understanding.
Jane: This self-awareness is key; it elevates the system from a sophisticated tool into an autonomous reasoning agent capable of handling genuine novelty.
Tom: Ultimately, this entire process solidifies that for physical systems to achieve general intelligence, they cannot rely solely on raw perception data; they must be anchored by verifiable, symbolic logic.
Jane: So while "PO-PDDL: Learning Symbolic POMDPs from Visual Demonstrations for Robot Planning Under Uncertainty" provides a detailed roadmap for achieving this level of structured reasoning, it’s only one chapter in the grand story of AI.
Tom: With that final thought on structured planning, we have truly wrapped up our deep dive into this remarkable paper. And speaking of entirely different forms of intelligence, next up, we are going to pivot completely gears and look at how Large Language Models are changing the landscape of natural language understanding in a totally different way.
N/A
cs.RO, cs.AI
Submitted: 2026-08-20
Updated: 2026-08-21
Project page: https://po-pddl.github.io
Importance score: 81/100
The gist: The paper introduces PO-PDDL (Partial Observable PDDL), a framework designed for "Learning Symbolic POMDPs from Visual Demonstrations for Robot Planning Under Uncertainty." The core contribution lies
Key concepts
- Online Adaptation
- This enhancement allows the system to adjust its operation without needing massive retraining efforts when environmental conditions change slightly. It ensures the robot remains accurate and robust in unpredictable real-world settings.
- Model Drift
- This refers to the gradual decline in accuracy that complex AI systems experience over long periods of use. The paper addresses this by proposing continuous correction mechanisms, allowing the robot to self-adjust in real time.
- Targeted Data Collection
- A major efficiency gain where the system directs its own learning efforts. Instead of gathering random footage, it self-diagnoses its weakest points and requests specific data needed to fix known failure modes.
- Generalized Failure Handling
- Moving beyond simply reporting an 'Error,' this process requires the robot to generate a full diagnostic report. It must explain *why* its initial assumptions failed, providing root-cause analysis.
Terminology
Summary
The paper introduces PO-PDDL (Partial Observable PDDL), a framework designed for Learning Symbolic POMDPs from Visual Demonstrations for Robot Planning Under Uncertainty.
The core contribution lies in creating a structured, symbolic representation that allows for robust planning under partial observability.
Comparison and Limitations of Prior Work (pomdp coder):*
The text details several limitations observed in the pomdp coder* model. Regarding transition effects, while pomdp coder* can express stochastic programs,
its learned transition model is described as weakly structured.
Its preconditions are deemed incomplete because they must be enumerated as Python conditions,
and its effects lack an explicit stochastic effect-mode structure for modeling action failures. Furthermore, the implementation of helper functions, such as maybe reveal object on open (Listing 13: Information reveal encoded as transition), demonstrates a flaw where an observation event into a physical state transition, corrupting the latent state rather than updating the belief.
The observation model also exhibits significant issues. Although pomdp coder* has an explicit interface returning symbolic payloads (positive, negative, and metadata) (Listing 14: Learned observation payload), its space is overly broad and redundant.
It mixes fully observable facts with partially observable facts,
including gripper holding state, drawer open/closed state, visible object locations, object properties, and support-clear observations. The required observation space should instead focus on task-relevant partially observable predicates.
This redundancy is illustrated by the fact that the learned observation function emits observations for many ordinary visible facts (e.g., holding state and drawer state) even if they are not necessary for belief updates (Listing 15: Learned observation model). Overall, pomdp coder* is criticized because it learns as a loose Python program and does not use a two-stage learning pipeline that reconstructs a reliable fully observable domain before identifying truly partially observable predicates,
making it unreliable for long-horizon belief-space planning.
The PO-PDDL Approach and Key Capabilities:
In contrast, the PO-PDDL representation is presented as superior. A major advantage highlighted is the ability to perform Online Adaptation of Operator Effect Probabilities. The structured PO-PDDL representation allow[s] operator effect probabilities to be updated during online execution.
Because each outcome can be mapped to a symbolic effect mode, the planner can recalibrate transition probabilities from real feedback without relearning the domain structure.
This capability is demonstrated qualitatively using an example on block in drawer. When simulating test-time skill degradation (e.g., making right-drawer grasping unreliable), the model initially prefers the right drawer based on learned demonstrations. However, After repeated failures, the estimated success probability of right-drawer grasping decreases, and the planner switches to opening the left drawer and grasping the block from there.
This adaptation also serves as a diagnostic tool: a skill whose estimated success probability drops repeatedly indicates a mismatch between the learned model and the current execution environment,
allowing researchers to prioritize targeted data collection or fine-tuning.
Planning Implementation Details:
The planning process utilizes several hyperparameters, including:
-
lambda: The
Empirical–uniform prior mixing weight for initial belief generation,
set to 0.5. -
T: The
Temperature for all LLMs and VLMs,
set to 0.0. -
k: The number of scenarios in Belief Tree Search, set to 500.
-
ds: Maximum search depth in Belief Tree Search, set to 50.
-
dr: Rollout policy execution depth, set to 20.
Improvements for AI systems
As a fastidious and diligent AI researcher where mistakes are prohibitively costly, my analysis of the pomdp coder* approach reveals fundamental structural weaknesses that render it unreliable for long-horizon, safety-critical planning. The primary issue is that the model treats structured symbolic reasoning (PDDL/POMDPs) as a loose, unstructured Python program output, leading to brittle state representations and flawed belief updates.
My proposed improvements focus on enforcing strict modularity and explicit structural bindings across the entire pipeline: State to Action to Observation to Belief Update.
Here are the specific improvements I recommend for building a robust POMDP planning system:
The current method's failure to model stochastic effects explicitly is unacceptable, as it cannot distinguish between an intended action and its possible outcomes (e.g., success, failure due to obstruction).
Improvement: The system must adopt a formal Stochastic Effect Mode Structure.
- Mechanism: For every action A, the transition function T(s' s, A) must be defined not merely by deterministic updates, but as a weighted distribution over distinct effect modes.
T(s' s, A) = sum omega in P(omega s, A) times T omega(s' s)
- Specific Implementation: Each action must be decomposed into mandatory outcomes:
-
Success Mode (Eff success): The intended outcome (e.g., object moved, container opened). This mode must explicitly update the state variables (
object location,support occupancy). -
Failure Mode (Eff failure): The most common failure (e.g., obstruction, grip slippage). This mode must define the resulting state change (e.g., no change to location, or only a minor perturbation) and assign a non-zero probability P(omega failure s, A).
-
Partial/Warning Mode (Eff warning): For complex actions (like opening a drawer), this mode might model partial success (e.g., the drawer opens, but nothing is revealed).
- Benefit: This forces the model to learn probabilities of outcomes, not just state changes, making it robust to real-world execution failures and allowing for explicit risk assessment during planning.
The current practice of using observation events (like opening a drawer) to directly rewrite the latent physical state (object location) is a critical bug that corrupts the belief state.
The current observation model is redundant, mixing high-level facts (gripper state) with low-level visibility heuristics. This overloads the belief update process and obscures the true partial observability targets.
The ability to adapt probabilities online (as shown in Figure 6) is valuable, but it must be formalized for actionable planning.
The resulting system will be a Structured Belief-Space Planner with the following superior capabilities:
-
Robust Execution Planning: It plans over explicit probability distributions of outcomes, allowing it to select actions that maximize expected utility even when failure is likely (e.g., choosing a path with 0.7 times 0.9 success rate over one with 1.0 but high obstruction risk).
-
Guaranteed State Integrity: It maintains a mathematically clean separation between the physical, symbolic state and the noisy, observed evidence, eliminating procedural bugs that corrupt the belief state.
-
Efficient Learning & Adaptation: It uses targeted domain shift detection to identify which specific skill or predicate is failing (e.g.,
drawer grasping
) rather than generalizing failure across all skills, dramatically improving data efficiency and reducing the cost of retraining. -
Formal Reasoning: By enforcing explicit stochastic effect modes, it moves beyond simple state-transition
Sources
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- Large-Language-Model-Guided State Estimation for Partially Observable Task and Motion Planning
- Robots That Ask For Help: Uncertainty Alignment for Large Language Model Planners
- LBAP: Improved Uncertainty Alignment of LLM Planners using Bayesian Inference
- LLMs for Robotic Object Disambiguation
- VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents
- Tru-POMDP: Task Planning Under Uncertainty via Tree of Hypotheses and Open-Ended POMDPs
- End-to-end PDDL Planning with Hardcoded and Dynamic Agents
- Unifying Deep Predicate Invention with Pre-trained Foundation Models
- One Demo Is All It Takes: Planning Domain Derivation with LLMs from A Single Demonstration
- Generating Symbolic World Models via Test-time Scaling of Large Language Models
- Iterative Formalization and Planning in Partially Observable Environments
- Seeing is Believing: Belief-Space Planning with Foundation Models as Uncertainty Estimators
- PlanU: Large Language Model Reasoning through Planning under Uncertainty
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving