OrbitTAMP: Grounding Language Models for Task and Motion Planning in Spacecraft Rendezvous

summary

Video file (mp4)

The gist

Spacecraft rendezvous and proximity operations (RPO) are currently planned through an expertise-intensive process that creates a bottleneck for scalable operations, and this paper introduces a

In short

The paper introduces a hierarchical framework to ground Large Language Model (LLM) reasoning in spacecraft task-and-motion planning (TAMP). It structures planning into semantic grounding, decision completion, and physical verification stages. This system uses reusable behaviors and structured mission representations to convert vague language intent into physically realizable spacecraft motion plans.

Key concepts

Hierarchical Planning Formulation
This is an agentic architecture that scaffolds LLMs with structured modules for planning and verification. It systematically completes operator underspecification by using a graph of reusable behaviors and domain-specific planning modules to turn incomplete intent into physically possible spacecraft motion.
MissionConfig
A structured representation used to hold all mission information. It includes fields like preferences (P), hard constraints (H), and sequences for behaviors, durations, and waypoints. This structure allows operator-specified data to be preserved throughout the entire planning pipeline.
Behavior–Waypoint Graph G(D, E)
This graph represents continuous waypoint domains and admissible behavior transitions rather than discrete states. It serves as a reusable library of spacecraft behaviors that guides the completion of mission decisions by defining what actions are possible in a given orbital domain.
Intent Parsing
The initial stage where a pretrained LLM maps natural language commands to a partial mission specification (M0). This step is crucial for ensuring only stated intent is extracted, while explicitly marking any unsupported fields as unspecified, setting the stage for downstream completion.

Terminology used across episodes

This episode discusses

The paper

OrbitTAMP: Grounding Language Models for Task and Motion Planning in Spacecraft Rendezvous · Read on arXiv

Yuji Takubo, Daniele Gammelli, Marco Pavone, Simone D’Amico

Stanford University · Italian Institute of Artificial Intelligence (AI4I) · NVIDIA Research

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "OrbitTAMP: Grounding Language Models for Task and Motion Planning in Spacecraft Rendezvous".

Dev: Spacecraft rendezvous and proximity operations (RPO) are currently planned through an expertise-intensive process that creates a bottleneck for scalable operations,

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: So we're looking at this paper, "OrbitTAMP: Grounding Language Models for Task and Motion Planning in Spacecraft Rendezvous," which tackles how to use large language models for spacecraft operations. Basically, the authors are addressing the bottleneck where engineers have to translate high-level goals into safe trajectories, and they propose a hierarchical framework to make that process more structured.

Dev: That sounds like it could really help with scaling up these kinds of missions because right now, the process seems too dependent on individual expertise for every single planning step. What the paper claims is that this framework grounds the LLM's reasoning in orbital dynamics and operational constraints rather than just letting it generate plans without any physical checks.

Taro: I'm interested in how they handle the complexity of real-world situations, especially when things go wrong; I want to know what happens when the world misbehaves during these kinds of operations.

Rosa: The core thesis seems to be that by breaking down the planning into stages—semantic grounding, completing decisions, and then physical verification—they can ensure that only stated intent is extracted from the natural language command while systematically filling in the blanks with reusable behaviors and domain-specific modules. This preserves operator input throughout the entire pipeline, which is something I think is really crucial for human oversight.

Dev: From an engineering standpoint, I see this hierarchical approach as a way to manage complexity by separating what the LLM understands semantically from what actually needs to be physically executed; it moves away from just hoping the LLM spits out a valid sequence of maneuvers. The paper suggests they use a "hierarchical planning formulation in which operator underspecification is preserved and systematically completed".

Taro: Preserving the underspecification while completing it by downstream modules is interesting, because when you're dealing with autonomous decision-making, you need a robust way to handle those gaps without completely losing the original intent of what was requested. How does this structure handle situations where the initial natural language command is vague?

Rosa: The paper introduces a MissionConfig structure that explicitly marks unresolved decisions with "⊥ for completion by downstream modules," which is a neat way to keep track of what's missing while still letting the system move forward. This means the intent parsing stage just sets up M0, and subsequent modules fill in the gaps based on behavioral graphs and constraints.

Dev: That structure sounds like it provides a clear audit trail, which is something we need when we're dealing with spacecraft operations where failure modes are so critical. I’m thinking about the loop rate here; if this framework adds too many layers of processing, could the latency become an issue for real-time response?

Paper summary: Taro: The paper does touch on the verification stage, which involves trajectory optimization to enforce dynamics and constraints defined in H, suggesting that this final layer is where the physical realization happens. I wonder what kind of replanning capability this provides when an unexpected event occurs mid-maneuver.

Rosa: The paper demonstrates that the full process involves intent parsing mapping to M0, behavior sequencing to M1, waypoint generation to M2, and finally trajectory optimization to get the final state and control trajectories. It’s a very structured way of moving from a vague idea to an executable plan.

Dev: I'm also looking at the model evaluation part, where they test different LLM backends like Qwen3 point 5-9B and GPT-five point six Terra, and they found that GPT-five point six Terra achieves "ninety-eight percent exact M0 recovery on all three splits," while smaller models show lower exact recovery. That comparison is important for understanding the practical requirements for the underlying LLM component.

Taro: If a frontier model like GPT-five point six Terra shows that it can recover the stated intent with high accuracy, does that imply we need to rely on very powerful models just to get us started, or can these compact open-weight models actually be sufficient if they are properly scaffolded?

Rosa: The results suggest that the structured scaffolding substantially improves performance over oneshot LLM-based planning when both architectures are intent-aligned, which points toward the scaffolding being a major factor in making these systems usable. We also see that test-time computation can be strategically allocated, showing verifier-guided revision improving intent parsing accuracy from seventy-five percent to eighty-eight percent at Nrev = two.

Dev: That allocation of compute is telling; it means you can tailor the processing intensity based on where the uncertainty is highest, like using a verifier to refine the semantics before diving deep into behavior planning search to improve trajectory quality among candidates. This seems like a smart way to manage computational load for real-time applications.

Taro: It’s encouraging that they show how refining the intent parsing and then doing a separate planning search for trajectory quality are distinct steps, which suggests we can optimize each module independently for better reliability in failure scenarios. I'm curious about the limitations they mention; what is it that this framework simply cannot do when deployed outside of a highly controlled lab environment?

Rosa: The paper does acknowledge that the system relies on an external graph of reusable spacecraft behaviors and domain-specific planning modules, meaning its success is heavily tied to how well those reusable components are defined upfront. It doesn't necessarily imply it works perfectly in an unstructured, completely novel operational scenario without that prior structure.

Dev: So the main limitation seems to be the dependency on the quality and completeness of that pre-defined behavior graph and constraints, rather than a pure failure of the LLM's reasoning itself, which is a helpful distinction for us designing hardware interfaces. We have to ensure those domain-specific modules are rigorously validated against orbital dynamics.

Paper summary: Taro: Given what we've seen about the potential for these language-driven agents to interface with spacecraft operations, I think the implication is that we can start moving toward systems where mission planning isn't entirely hardcoded, but where a highly structured, verifiable AI agent handles the translation of human goals into executable physics.

Rosa: Exactly; this framework provides an auditable human–spacecraft interface and supports scalable language-driven planning for future distributed space systems. It suggests that we can build interfaces that are both intuitive for humans and robust enough to handle the physical realities of orbital mechanics.

Dev: For control engineers like myself, the implication is that if we can use this structured approach, we gain a layer of formal verification before we even get to the trajectory optimization stage, which helps manage those failure modes you mentioned earlier. We get better stability because the input to the optimizer is already constrained by behavior sequences and durations.

Taro: From an autonomy perspective, this means we move closer to systems where autonomous decision-making isn't just about following discrete logic but can incorporate semantic reasoning about complex task sequencing. It opens up possibilities for more adaptable agents in unpredictable environments.

Rosa: It really seems like OrbitTAMP gives us a way to bridge the gap between high-level, natural language intent and the low-level, physically realizable maneuvers required for spacecraft rendezvous. The structure is what makes it scalable.

Dev: Scalability in this context means we can deploy these planning agents to handle a wider variety of mission profiles without needing a completely new set of manual planning rules for every single one. That reduces the burden on mission control engineers significantly.

Taro: It’s exciting because it moves us past systems that are just executing pre-programmed sequences and toward agents that can reason about the task itself in a more generalized way. That level of abstraction is where the real autonomy is going to come from.

Rosa: So, looking at this OrbitTAMP paper, it’s clear that by layering semantic understanding on top of structured physical constraints, we build something that's both intuitive for mission planners and rigorous enough for actual spacecraft control.

Dev: And from a loop rate perspective, the structure allows us to isolate where the computation happens so we can optimize the latency in those critical decision points without having to re-solve the entire complex trajectory problem every time.

Taro: I just think what this paper shows is a path toward more capable autonomous systems where the planning component itself is intelligent enough to handle ambiguity by using a hierarchy of tools, rather than relying on brittle, single-step reasoning.

Rosa: It’s definitely a foundation for what we could call an auditable human–spacecraft interface that scales beyond the current manual planning bottlenecks in rendezvous operations.

Conclusion: Rosa: So we've just been diving deep into how this paper tackles spacecraft rendezvous planning using language models, and now we’re wrapping up with some final thoughts on what OrbitTAMP actually means for us.

Dev: It really boils down to taking those complex, high-level mission instructions from humans and giving them a structured path that the AI can actually follow in space, which is pretty smart thinking for loop rate management.

Taro: I’m still thinking about the autonomy side; if this framework works as described, does it mean we can trust an AI to handle unexpected maneuvers during a proximity operation without constant human intervention?

Rosa: Exactly, Taro, and the authors are pointing toward a scalable foundation for distributed space systems because they've managed to externalize that domain knowledge into something structured.

Dev: I agree with Rosa on the scalability point; having that hierarchy means we can isolate where the computation is heavy and optimize those latency bottlenecks without having to redo the whole trajectory solve every time.

Taro: But I’m still cautious about deployment outside of a perfect simulation; how robust is this framework when things go completely off-script in a real operational environment?

Rosa: That's the million-dollar question, isn't it? The paper shows that by grounding the LLM in orbital dynamics and using those reusable behaviors, we build an auditable interface that’s much safer than relying on pure prompt compliance.

Dev: I think the implication is a significant step toward moving away from purely pre-programmed sequences toward agents that can reason about the task itself, which opens up possibilities for more adaptable autonomous systems.

Taro: That sounds promising, but we still need to see how well those domain-specific modules handle truly novel failure modes that weren't explicitly in their training data.

Rosa: Well, the core finding is that this hierarchical architecture successfully separates semantic grounding from physical verification, which gives us a powerful tool for designing human-spacecraft interfaces that are both intuitive and rigorous.

Dev: It’s a huge step toward reducing the burden on mission control engineers by providing an AI planner that can handle a wider variety of mission profiles without needing entirely new manual planning rules for every single one.

Taro: So the big picture is building agents that can reason about task sequencing in unpredictable environments, rather than just executing discrete logic steps.

Rosa: Precisely, and this work lays a solid foundation for future language-driven planning in distributed space systems because it’s systematic and verifiable.

Dev: We've shown how test-time computation can be strategically allocated to refine intent parsing and then perform downstream planning search for better trajectory quality among candidates, which is a practical win.

Taro: It really shows that the system isn't just guessing; it’s systematically completing unspecified decisions using the established constraints.

Rosa: And that systematic approach is what makes this framework so valuable for building scalable and auditable language-driven planning for those complex rendezvous operations we've been discussing.

More episodes

← Home