SkillWrapper: Generative Predicate Invention for Task-level Robot Planning

arXiv:2511.18203 · cs.RO · Submitted 2025-11-22 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "SkillWrapper: Generative Predicate Invention for Task-level Robot Planning".

Rosa: Detailed Research Summary: SkillWrapper - Generative Predicate Invention for Task-level Robot Planning This research introduces SkillWrapper,

Dev: First, who's behind it and why it matters.

Title and authors: Rosa: So, we're looking at the paper "SkillWrapper: Generative Predicate Invention for Task-level Robot Planning," and the title itself really tells you what this is about—it’s about generating these symbolic representations for robot planning. I think it suggests they are trying to bridge that gap between raw visual data and actual task execution by creating these abstract rules that a planner can use.

Dev: It sounds like they’re focusing on how to make the system think in terms of skills and objects, rather than just reacting to pixels. The authors seem to be tackling the challenge of learning high-level symbolic representations from low-level sensor data, which is a big hurdle in autonomous systems right now.

Taro: I'm curious if this approach scales well beyond the specific environment they used for their initial experiments, Rosa? If it relies heavily on visual input for predicate invention, how robust will that be when things look different in the real world?

Rosa: That’s a fair question, Taro. The paper mentions testing on various setups like simulation and even single-arm manipulation with an RGB-D camera and bimanual setups with KUKA arms. I'm wondering if the generalization holds up outside of those controlled lab settings for extended periods.

Dev: From my end, I'm thinking about the execution side of things; if the inference time for generating these predicates gets too high, we lose our loop rate, which is critical for real-time control. The paper doesn't really detail how they manage that latency when the foundation model is doing most of the heavy reasoning.

Taro: That brings up a point about robustness; what happens when the world misbehaves in a way that wasn't represented in their training data? Does SkillWrapper have a mechanism for handling unexpected failures, or does it just get stuck because its learned predicates don't cover the new situation?

The paper's summary: Rosa: So, to summarize what this paper is actually proposing with "SkillWrapper: Generative Predicate Invention for Task-level Robot Planning," it essentially introduces a method that uses a foundation model to actively collect data and then invent new, human-interpretable predicates. This process is designed to create symbolic models for task planning that are provably sound and complete with respect to the data they see.

Dev: It sounds like the core mechanism involves an active exploration loop where the AI proposes sequences of actions, observes outcomes, and then uses those observations to propose new predicates that describe what happened. The paper emphasizes that this invention is driven by observing preconditions in successful transitions within the dataset.

Taro: So they're not just learning a fixed set of rules; they are dynamically discovering the necessary abstract concepts as it interacts with the environment, which is an interesting direction for autonomy research.

Rosa: Exactly, Taro. The authors build a formal theory of this generative predicate invention specifically to give guarantees about whether the learned model will actually work when plugged into a classical planner. They focus on proving that for every successful transition in their dataset, there must be an abstract action with matching preconditions and effects.

Dev: That formality is key; it moves the work beyond just being a clever heuristic and gives us some kind of theoretical assurance about correctness. I'm interested in how they handle the operator learning part, since that’s where they try to distill multiple skill instances into a single usable operator.

The paper's improvements: Rosa: The paper outlines several specific improvements for this approach, and one of the main ones is moving from reactive control to proactive task composition. This means the system can use those learned symbolic predicates to generate high-level plans rather than just reacting moment by moment.

Dev: That proactive planning capability sounds ambitious; it implies the robot can look ahead and compose a sequence of actions based on these abstract symbols, which is much more complex than simple reactive loops we see in other work, like InCoM.

Taro: I think the ability to generate predicates from visual comparisons of input images is a big step. It means the system can invent concepts purely based on what it sees, which gives it a lot of flexibility in understanding novel object states without needing explicit programming for every single possibility.

Rosa: And there's also the focus on active data collection through skill sequence proposals. The idea is that by intentionally designing those sequences to sometimes violate preconditions or explore new skills, the system forces itself to find the missing symbolic knowledge it needs to be complete.

Dev: From a control engineer's point of view, if they can generate these predicates reliably, it should lead to much more predictable failure modes. Instead of unpredictable errors arising from state space issues, we might see failures traceable back to an incorrect predicate definition.

Conclusion: Rosa: So, wrapping up the "SkillWrapper: Generative Predicate Invention for Task-level Robot Planning" paper, it really boils down to a method that uses foundation models to generate symbolic predicates for robot planning while providing formal guarantees of soundness and completeness. This is significant because it shows a path toward building robots that can reason at a level beyond just low-level control loops.

Dev: I agree, the theoretical framework they provide around empirical correctness and asymptotic convergence gives us confidence that this isn't just an experimental curiosity; it has some mathematical backing for its performance when applied to the task-level planning problem.

Taro: From an autonomy research standpoint, I think the most impactful aspect is how it allows for novel skill discovery when the current model fails, so the robot can adapt its internal logic dynamically without manual reprogramming of preconditions.

Rosa: That dynamic adaptation is what really excites me—the idea of a system that learns its own logic through interaction with the environment in this structured way. We should definitely keep an eye on how this translates into real-world deployment timelines, Rosa?

Dev: And I’m keeping my eyes peeled for those latency metrics; if they can manage the inference speed while maintaining that formal guarantee of correctness, then we might actually see something practical sooner than expected.

Taro: I just think the ability to create generalized operators that handle multi-type objects conservatively is a really smart way to ensure that when we generalize across different domains, we don't end up with messy, overly permissive rules.

Rosa: Well, this paper presents "SkillWrapper: Generative Predicate Invention for Task-level Robot Planning" as a solid step forward in how we formalize the learning of symbolic models for complex robot tasks. We’ll have to see if these theoretical guarantees hold up when we put them through the wringer in actual field testing.

Dev: Indeed, it's a complex system, and I'm looking forward to seeing the next iterations focusing on making that loop rate as tight as possible under real-world conditions.

Ziyi Yang, Benned Hedegaard, Ahmed Jaafar, Yichen Wei, Skye Thompson, Shreyas Raman, Haotian Fu, Stefanie Tellex, George Konidaris, David Paulius

Brown University

cs.RO

Submitted: 2025-11-22

Updated: 2026-09-28

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 90/100

The gist: This research introduces SkillWrapper, a novel method designed to learn provably sound and complete symbolic models for robot planning by leveraging large Foundation Models (FMs) to actively collect

Key concepts

Generative Predicate Invention
This is the core process where the AI system proposes new high-level rules or conditions (predicates) based on visual observations. It learns to create these abstract concepts by comparing different visual states, allowing the robot to reason about task success using human-understandable language rather than just raw pixels.
Theoretical Guarantees
The research provides formal mathematical proofs that the learned symbolic model is correct and complete relative to the observed data. This means there are rigorous bounds on how much data is needed for the system to be reliable, ensuring that any plan derived from it will actually work in practice.
Foundation Models (FMs)
Large pre-trained models are used as the engine for learning complex representations, especially for inventing new predicates. These FMs analyze visual input and generate reasoning about object states or spatial relations, providing the semantic grounding necessary to create abstract planning knowledge.

Terminology

Summary

This research introduces SkillWrapper, a novel method designed to learn provably sound and complete symbolic models for robot planning by leveraging large Foundation Models (FMs) to actively collect data and generate human-interpretable, plannable representations using only RGB image observations. SkillWrapper is positioned as the first work to use an off-the-shelf foundation model to automatically learn these symbolic representations with theoretical guarantees of correctness.

SkillWrapper operates on a three-step process: active exploration for data collection, predicate invention based on contrastive examples, and operator learning using the accumulated data and invented predicates. Crucially, the paper establishes a formal theory of generative predicate invention specifically tailored for task-level planning to rigorously characterize the conditions under which a learned model is provably sound and complete with respect to downstream planning.

The theoretical guarantees are robust:

  1. Empirical Correctness (Theorem 1): The analysis proves that the learned model M is correct (sound and complete) with respect to observed data D. Specifically, for every successful transition s i, omega i, s'i in the dataset D, there must exist an abstract action a such that the precondition of the current state (s i) is met by the abstract action (Pre a), and applying that action results in a state (s'i) whose effect ((s i, s'i)) matches the expected effect (Eff a).

  2. Asymptotic Convergence (Theorem 2): The sample complexity analysis provides a bound on the necessary number of transitions required to achieve a desired completeness level, showing that this sample size is polynomial in terms of the model capacity. Furthermore, Theorem 2 demonstrates asymptotic convergence to a correct model with respect to the underlying data distribution (mu), providing strong theoretical assurance for generalization.

The paper details how predicate invention is driven:

  • Condition 1: Predicate invention is primarily triggered by observing preconditions within successful transitions.

  • Operator Learning: Operators are learned by clustering observed transitions based on their lifted abstract changes, allowing the system to learn a single operator across multiple skill instantiations.

The resulting learned predicates and operators are explicitly shown to be human-interpretable, as evidenced by illustrative examples in domains like Burger and Franka.

SkillWrapper utilizes an FM to perform the heavy lifting of representation learning, specifically for predicate invention, while maintaining a structured pipeline:

  1. Skill Sequence Proposal: The system proposes skill sequences using a System Prompt and a Skill Sequence Proposal Prompt. This ensures that the generated sequences occasionally violate preconditions and utilize at least one unexplored skill pair in each sequence, driving active exploration.

  2. Predicate Invention: This is the core inventive step, where the agent is tasked with proposing a single new high-level predicate. This task relies on visual comparison of two input images taken before an execution, demanding that invented predicates be grounded purely in visual state and accurately describe object states or spatial relations pertinent to task success.

  3. Predicate Evaluation: A rigorous two-stage evaluation process is employed: the FM first generates a response in any format, followed by a second stage where it provides a summary of the initial output.

The paper highlights that SkillWrapper draws inspiration from established fields such as model learning, abstraction learning, and Task and Motion Planning (TAMP), aiming to generate human-interpretable abstractions using pre-trained FMs.

SkillWrapper has undergone extensive empirical testing across diverse setups:

  • Simulation: Experiments in the simulated Robotouille kitchen environment demonstrate its capability with a robot possessing five skills (Pick, Place, Cut, Cook, Stack).

  • Single-Arm Manipulation: Tested on a Franka Emika Research 3 arm equipped with an RGB-D camera.

  • Bimanual Manipulation: Tested on a setup involving two horizontally mounted KUKA LBR iiwa 7 R800 manipulators.

Key Empirical Findings:

  • Superiority over Baselines: SkillWrapper significantly outperforms baselines that lack access to privileged knowledge and even surpasses the performance of System Predicates on hard problems.

  • Data Efficiency: In real-world experiments, SkillWrapper achieves competitive performance with Expert Operators while requiring only a handful of exploratory interactions, suggesting an excellent trade-off between data efficiency and model correctness.

  • FM Comparison (Qualitative): A qualitative comparison between GPT-5 and Qwen3 on predicate invention shows that GPT-5 is substantially more reliable in reasoning over contrastive pairs of transitions for predicate invention, indicating its superior capability in this specific reasoning task.

Despite its theoretical rigor, the work acknowledges several limitations:

Improvements for AI systems

Here are specific improvements for AI systems based on the SkillWrapper paper, detailing what these improved systems can achieve:


  1. Enhanced Reasoning for Long-Horizon Task Planning:

Based on SkillWrapper's core contribution (learning provably sound and complete symbolic models), an improved AI system can transition from reactive, low-level control to proactive, high-level task composition.

  1. Skill Composition and Abstraction: The system will be able to take a set of raw sensory inputs (RGB images) and autonomously generate human-interpretable symbolic predicates (e.g., the robot's gripper is visually not holding any object). This allows the AI to create an abstract transition model that operates on these predicates rather than low-level state vectors, enabling reasoning independent of the specific low-level state space.

  2. Provably Sound and Complete Planning: Unlike prior generative methods that lack formal guarantees, this system will possess a learned task-level model where the preconditions and effects of skills are formally characterized. This ensures that any plan generated by a downstream classical planner is guaranteed to be sound (will lead to a real outcome) and complete (will find a solution if one exists), dramatically reducing the risk of executing physically impossible or logically unsound sequences.

  3. Active, Data-Driven Model Refinement: The system will employ an active learning loop where it intelligently proposes skill sequences based on learned heuristics like Balance (prioritizing successful executions) and Coverage (ensuring observation of all inter-skill dependencies). This means the AI doesn't just randomly explore; it systematically seeks out the data most likely to reveal missing symbolic knowledge, leading to a more efficient learning curve.

  4. Novel Skill Discovery: The system can autonomously invent new predicates when its current model fails to explain observed transitions (e.g., distinguishing between a successful and failed execution of the same skill). This capability allows the robot to adapt its internal logic dynamically as it encounters novel environmental constraints or object interactions without requiring manual reprogramming of preconditions.

  5. Generalizable Operator Learning: The system will learn generalized operators that handle multi-type objects by conservatively assigning them to the lowest level of the type hierarchy consistent with the data, preventing over-generalization. This allows a single operator definition (e.g., Pick) to be used across diverse object categories (e.g., picking up a bottle vs. picking up an apple) while maintaining compositional structure and avoiding spurious preconditions that hinder generalization in different environments or domains.

  6. Scalable Sample Complexity Management: The system's theoretical analysis provides concrete bounds for required data, demonstrating how sample complexity scales with the model capacity (maximum number of predicates/operators). This allows researchers to predict the necessary data size needed to achieve a target level of completeness and error tolerance, making it feasible to design experiments for complex tasks.

  7. Robustness Against Domain Shifts: By relying on semantic grounding provided by Foundation Models rather than raw visual features, the system's generalization capability is enhanced. It can be prompted with new object types (via text descriptions) and correctly identify referents in novel environments, only failing when the visual appearance fundamentally contradicts the learned semantics (e.g., confusing a small plate for a saucer), leading to more robust performance across different physical settings.

Abstract

Generalizing from individual skill executions to long-horizon tasks is a core challenge in building autonomous robots. A promising direction is learning high-level, symbolic representations of low-level robot skills, enabling abstract reasoning independent of the low-level state space. Recent advances in foundation models have made it possible to generate symbolic predicates that operate on raw sensory inputs-a process we call generative predicate invention-to facilitate downstream representation learning. However, prior work learns these abstractions using heuristic or ad-hoc procedures, leaving unclear which formal properties they ought to satisfy, and how these properties can guide representation learning. We address these questions by characterizing conditions under which learned representations support sound and complete task-level planning, and using them to guide the design of SkillWrapper, a system that autonomously learns symbolic representations of black-box skills without predefined tasks, predicates, or operators. Our approach leverages foundation models to actively collect robot data and learn human-interpretable, plannable representations directly from RGB observations. Our extensive empirical evaluation in simulation and on real robots shows that SkillWrapper learns abstract representations that enable robots to compose black-box skills to solve unseen, long-horizon tasks in the real world.

Sources

Related papers