OrbitTAMP: Grounding Language Models for Task and Motion Planning in Spacecraft Rendezvous

arXiv:2610.01093 · cs.RO, cs.AI, math.OC · Submitted 2026-10-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "OrbitTAMP: Grounding Language Models for Task and Motion Planning in Spacecraft Rendezvous".

Dev: Spacecraft rendezvous and proximity operations (RPO) are currently planned through an expertise-intensive process that creates a bottleneck for scalable operations,

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: So we're looking at this paper, "OrbitTAMP: Grounding Language Models for Task and Motion Planning in Spacecraft Rendezvous," which tackles how to use large language models for spacecraft operations. Basically, the authors are addressing the bottleneck where engineers have to translate high-level goals into safe trajectories, and they propose a hierarchical framework to make that process more structured.

Dev: That sounds like it could really help with scaling up these kinds of missions because right now, the process seems too dependent on individual expertise for every single planning step. What the paper claims is that this framework grounds the LLM's reasoning in orbital dynamics and operational constraints rather than just letting it generate plans without any physical checks.

Taro: I'm interested in how they handle the complexity of real-world situations, especially when things go wrong; I want to know what happens when the world misbehaves during these kinds of operations.

Rosa: The core thesis seems to be that by breaking down the planning into stages—semantic grounding, completing decisions, and then physical verification—they can ensure that only stated intent is extracted from the natural language command while systematically filling in the blanks with reusable behaviors and domain-specific modules. This preserves operator input throughout the entire pipeline, which is something I think is really crucial for human oversight.

Dev: From an engineering standpoint, I see this hierarchical approach as a way to manage complexity by separating what the LLM understands semantically from what actually needs to be physically executed; it moves away from just hoping the LLM spits out a valid sequence of maneuvers. The paper suggests they use a "hierarchical planning formulation in which operator underspecification is preserved and systematically completed".

Taro: Preserving the underspecification while completing it by downstream modules is interesting, because when you're dealing with autonomous decision-making, you need a robust way to handle those gaps without completely losing the original intent of what was requested. How does this structure handle situations where the initial natural language command is vague?

Rosa: The paper introduces a MissionConfig structure that explicitly marks unresolved decisions with "⊥ for completion by downstream modules," which is a neat way to keep track of what's missing while still letting the system move forward. This means the intent parsing stage just sets up M0, and subsequent modules fill in the gaps based on behavioral graphs and constraints.

Dev: That structure sounds like it provides a clear audit trail, which is something we need when we're dealing with spacecraft operations where failure modes are so critical. I’m thinking about the loop rate here; if this framework adds too many layers of processing, could the latency become an issue for real-time response?

Paper summary: Taro: The paper does touch on the verification stage, which involves trajectory optimization to enforce dynamics and constraints defined in H, suggesting that this final layer is where the physical realization happens. I wonder what kind of replanning capability this provides when an unexpected event occurs mid-maneuver.

Rosa: The paper demonstrates that the full process involves intent parsing mapping to M0, behavior sequencing to M1, waypoint generation to M2, and finally trajectory optimization to get the final state and control trajectories. It’s a very structured way of moving from a vague idea to an executable plan.

Dev: I'm also looking at the model evaluation part, where they test different LLM backends like Qwen3 point 5-9B and GPT-five point six Terra, and they found that GPT-five point six Terra achieves "ninety-eight percent exact M0 recovery on all three splits," while smaller models show lower exact recovery. That comparison is important for understanding the practical requirements for the underlying LLM component.

Taro: If a frontier model like GPT-five point six Terra shows that it can recover the stated intent with high accuracy, does that imply we need to rely on very powerful models just to get us started, or can these compact open-weight models actually be sufficient if they are properly scaffolded?

Rosa: The results suggest that the structured scaffolding substantially improves performance over oneshot LLM-based planning when both architectures are intent-aligned, which points toward the scaffolding being a major factor in making these systems usable. We also see that test-time computation can be strategically allocated, showing verifier-guided revision improving intent parsing accuracy from seventy-five percent to eighty-eight percent at Nrev = two.

Dev: That allocation of compute is telling; it means you can tailor the processing intensity based on where the uncertainty is highest, like using a verifier to refine the semantics before diving deep into behavior planning search to improve trajectory quality among candidates. This seems like a smart way to manage computational load for real-time applications.

Taro: It’s encouraging that they show how refining the intent parsing and then doing a separate planning search for trajectory quality are distinct steps, which suggests we can optimize each module independently for better reliability in failure scenarios. I'm curious about the limitations they mention; what is it that this framework simply cannot do when deployed outside of a highly controlled lab environment?

Rosa: The paper does acknowledge that the system relies on an external graph of reusable spacecraft behaviors and domain-specific planning modules, meaning its success is heavily tied to how well those reusable components are defined upfront. It doesn't necessarily imply it works perfectly in an unstructured, completely novel operational scenario without that prior structure.

Dev: So the main limitation seems to be the dependency on the quality and completeness of that pre-defined behavior graph and constraints, rather than a pure failure of the LLM's reasoning itself, which is a helpful distinction for us designing hardware interfaces. We have to ensure those domain-specific modules are rigorously validated against orbital dynamics.

Paper summary: Taro: Given what we've seen about the potential for these language-driven agents to interface with spacecraft operations, I think the implication is that we can start moving toward systems where mission planning isn't entirely hardcoded, but where a highly structured, verifiable AI agent handles the translation of human goals into executable physics.

Rosa: Exactly; this framework provides an auditable human–spacecraft interface and supports scalable language-driven planning for future distributed space systems. It suggests that we can build interfaces that are both intuitive for humans and robust enough to handle the physical realities of orbital mechanics.

Dev: For control engineers like myself, the implication is that if we can use this structured approach, we gain a layer of formal verification before we even get to the trajectory optimization stage, which helps manage those failure modes you mentioned earlier. We get better stability because the input to the optimizer is already constrained by behavior sequences and durations.

Taro: From an autonomy perspective, this means we move closer to systems where autonomous decision-making isn't just about following discrete logic but can incorporate semantic reasoning about complex task sequencing. It opens up possibilities for more adaptable agents in unpredictable environments.

Rosa: It really seems like OrbitTAMP gives us a way to bridge the gap between high-level, natural language intent and the low-level, physically realizable maneuvers required for spacecraft rendezvous. The structure is what makes it scalable.

Dev: Scalability in this context means we can deploy these planning agents to handle a wider variety of mission profiles without needing a completely new set of manual planning rules for every single one. That reduces the burden on mission control engineers significantly.

Taro: It’s exciting because it moves us past systems that are just executing pre-programmed sequences and toward agents that can reason about the task itself in a more generalized way. That level of abstraction is where the real autonomy is going to come from.

Rosa: So, looking at this OrbitTAMP paper, it’s clear that by layering semantic understanding on top of structured physical constraints, we build something that's both intuitive for mission planners and rigorous enough for actual spacecraft control.

Dev: And from a loop rate perspective, the structure allows us to isolate where the computation happens so we can optimize the latency in those critical decision points without having to re-solve the entire complex trajectory problem every time.

Taro: I just think what this paper shows is a path toward more capable autonomous systems where the planning component itself is intelligent enough to handle ambiguity by using a hierarchy of tools, rather than relying on brittle, single-step reasoning.

Rosa: It’s definitely a foundation for what we could call an auditable human–spacecraft interface that scales beyond the current manual planning bottlenecks in rendezvous operations.

Conclusion: Rosa: So we've just been diving deep into how this paper tackles spacecraft rendezvous planning using language models, and now we’re wrapping up with some final thoughts on what OrbitTAMP actually means for us.

Dev: It really boils down to taking those complex, high-level mission instructions from humans and giving them a structured path that the AI can actually follow in space, which is pretty smart thinking for loop rate management.

Taro: I’m still thinking about the autonomy side; if this framework works as described, does it mean we can trust an AI to handle unexpected maneuvers during a proximity operation without constant human intervention?

Rosa: Exactly, Taro, and the authors are pointing toward a scalable foundation for distributed space systems because they've managed to externalize that domain knowledge into something structured.

Dev: I agree with Rosa on the scalability point; having that hierarchy means we can isolate where the computation is heavy and optimize those latency bottlenecks without having to redo the whole trajectory solve every time.

Taro: But I’m still cautious about deployment outside of a perfect simulation; how robust is this framework when things go completely off-script in a real operational environment?

Rosa: That's the million-dollar question, isn't it? The paper shows that by grounding the LLM in orbital dynamics and using those reusable behaviors, we build an auditable interface that’s much safer than relying on pure prompt compliance.

Dev: I think the implication is a significant step toward moving away from purely pre-programmed sequences toward agents that can reason about the task itself, which opens up possibilities for more adaptable autonomous systems.

Taro: That sounds promising, but we still need to see how well those domain-specific modules handle truly novel failure modes that weren't explicitly in their training data.

Rosa: Well, the core finding is that this hierarchical architecture successfully separates semantic grounding from physical verification, which gives us a powerful tool for designing human-spacecraft interfaces that are both intuitive and rigorous.

Dev: It’s a huge step toward reducing the burden on mission control engineers by providing an AI planner that can handle a wider variety of mission profiles without needing entirely new manual planning rules for every single one.

Taro: So the big picture is building agents that can reason about task sequencing in unpredictable environments, rather than just executing discrete logic steps.

Rosa: Precisely, and this work lays a solid foundation for future language-driven planning in distributed space systems because it’s systematic and verifiable.

Dev: We've shown how test-time computation can be strategically allocated to refine intent parsing and then perform downstream planning search for better trajectory quality among candidates, which is a practical win.

Taro: It really shows that the system isn't just guessing; it’s systematically completing unspecified decisions using the established constraints.

Rosa: And that systematic approach is what makes this framework so valuable for building scalable and auditable language-driven planning for those complex rendezvous operations we've been discussing.

Yuji Takubo, Daniele Gammelli, Marco Pavone, Simone D’Amico

Stanford University · Italian Institute of Artificial Intelligence (AI4I) · NVIDIA Research

cs.RO, cs.AI, math.OC

Submitted: 2026-10-01

Updated: 2026-10-01

Comments: 20 pages, 8 figures

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 91/100

The gist: Spacecraft rendezvous and proximity operations (RPO) are currently planned through an expertise-intensive process that creates a bottleneck for scalable operations, and this paper introduces a

Key concepts

Hierarchical Planning Formulation
This is an agentic architecture that scaffolds LLMs with structured modules for planning and verification. It systematically completes operator underspecification by using a graph of reusable behaviors and domain-specific planning modules to turn incomplete intent into physically possible spacecraft motion.
MissionConfig
A structured representation used to hold all mission information. It includes fields like preferences (P), hard constraints (H), and sequences for behaviors, durations, and waypoints. This structure allows operator-specified data to be preserved throughout the entire planning pipeline.
Behavior–Waypoint Graph G(D, E)
This graph represents continuous waypoint domains and admissible behavior transitions rather than discrete states. It serves as a reusable library of spacecraft behaviors that guides the completion of mission decisions by defining what actions are possible in a given orbital domain.
Intent Parsing
The initial stage where a pretrained LLM maps natural language commands to a partial mission specification (M0). This step is crucial for ensuring only stated intent is extracted, while explicitly marking any unsupported fields as unspecified, setting the stage for downstream completion.

Terminology

Summary

Spacecraft rendezvous and proximity operations (RPO) are currently planned through an expertise-intensive process that creates a bottleneck for scalable operations, and this paper introduces a hierarchical framework to ground Large Language Model (LLM) reasoning in orbital dynamics and operational constraints. The gist is: A hierarchical framework for spacecraft task-and-motion planning (TAMP) grounds LLM reasoning in a graph of reusable behaviors and domain-specific planning modules to convert incomplete intent into admissible and physically realizable spacecraft motion.

Framework Overview

The proposed architecture organizes the planning process into three stages: semantic grounding, completion of unspecified mission decisions, and physical trajectory verification. This hierarchy is illustrated in Fig. 1, where the intent parsing initializes a partial mission specification M0; behavior sequencing produces M1; waypoint generation produces M2; and trajectory optimization generates the final state and control trajectories. The central contribution is a hierarchical planning formulation in which operator underspecification is preserved and systematically completed, realized as an agentic architecture that scaffolds pretrained, off-the-shelf LLMs with structured, spacecraft-specific planning and verification modules. This design ensures that stated intent is preserved by construction rather than by prompt compliance while unstated decisions are resolved by the behavior graph under orbital dynamics and safety constraints.

Common Planning Representation

To make the planning problem tractable, two complementary abstractions are introduced: a graphbased representation of reusable spacecraft behaviors and a structured representation of the mission specification called MissionConfig. The Behavior–Waypoint Graph G(D, E) represents continuous waypoint domains and admissible behavior transitions rather than discretized states or motion primitives. The MissionConfig structure includes fields such as P (preferences), H (hard constraints), b1:M (behavior sequence), d1:M (duration sequence), T (total flight time), and w1:M (waypoint sequence). This representation allows operator-specified information to be preserved throughout the pipeline, marking unresolved decisions with ⊥ for completion by downstream modules.

Hierarchical Planning Stages

The framework proceeds through four main stages, each grounding LLM reasoning at a different level:

  1. Intent Parsing: A pretrained LLM maps the natural-language command to a partial mission specification, M0, using a curated prompt. This stage ensures that only stated intent is extracted and that unsupported fields are unspecified.

  2. Behavior-Duration Sequence Planning: This module uses the behavior graph to complete the behavior sequence and phase durations, yielding M1. It approximates downstream evaluation using an evaluator Q to estimate feasibility without requiring a full trajectory optimization solve for every candidate, ranking candidates by a deterministic lexicographic map ρP induced by operator preferences P.

  3. Waypoint Generation: This layer instantiates the waypoint sequence for the selected behavior-duration sequence (b⋆1:M, d⋆1:M), completing the mission specification to M2. It uses a policy πw to propose continuous relative-state waypoints within admissible domains D⋆1:M, reconciling them with operator input without overwriting specified information.

  4. Trajectory Optimization: The final module takes the completed MissionConfig (M2) and computes nominal open-loop state and control trajectories by solving a discrete-time optimal control problem subject to dynamics, waypoint, and safety constraints defined in H.

Case Study Results

The architecture was demonstrated in an on-orbit inspection case study involving two or three phases. The results showed that the structured scaffolding substantially improves performance over oneshot LLM-based planning. Specifically, the hierarchical architecture is shown to be preferred over direct generation baselines when both architectures are intent-aligned, suggesting that explicit candidate enumeration and trajectory-aware evaluation provide a mechanism for selecting among multiple admissible alternatives according to feasibility and operator preferences.

Model Capability and Test-Time Compute

The study evaluated three LLM backends: compact open-weight models (Qwen3.5-9B, Ministral3-8B) and a frontier model (GPT-5.6 Terra). GPT-5.6 Terra achieves 98% exact M0 recovery on all three splits, whereas compact models show lower exact recovery and greater sensitivity to structural shifts like the unseen radio-communication register. Furthermore, test-time computation can be strategically allocated: verifier-guided revision improves intent parsing accuracy from 75% to 88% at Nrev = 2, while behavior–planning search improves trajectory quality among candidates. The results indicate that verifier-guided semantic refinement and downstream planning search act on largely distinct aspects of the pipeline.

Conclusion

The paper establishes a scalable and auditable foundation for language-driven agentic planning of spacecraft RPO by externalizing domain knowledge through a structured mission representation. The hierarchical architecture successfully separates semantic grounding, completion of unspecified mission decisions, and physical trajectory verification while preserving operator-provided information throughout the planning process. This framework provides an "auditable human–spacecraft interface and supports scalable language-driven planning for future distributed space systems.

Improvements for AI systems

As a fastidious researcher, I have analyzed the OrbitTAMP framework presented in this paper. The core innovation lies in systematically grounding Large Language Model (LLM) reasoning within a hierarchical planning architecture that respects spacecraft-specific physical constraints and behavioral semantics.

Here are the specific, actionable improvements to AI systems based on this research:


  1. The fundamental improvement is the transition from direct LLM generation to a grounded, multi-stage TAMP pipeline.

  2. The system can now reliably translate high-level human intent into executable spacecraft trajectories by decomposing the problem into three distinct, verifiable stages: Semantic Grounding, Decision Completion, and Physical Verification.

Specific Capabilities of the Improved AI System:

The improved system can perform the following tasks with high fidelity and safety guarantees:

  1. Reliable Intent Recovery from Natural Language Commands:

  2. Systemic Resolution of Unspecified Mission Decisions (Behavior Sequencing, Timing, Waypoints):

  3. Generation of Dynamically Feasible and Safe Trajectories:

Detailed breakdown of the improvements based on the paper's methodology:

  1. A natural language command (e.g., approach slowly from the-V bar, circumnavigate for two orbits, and then depart to a 50m separation) will be parsed by an LLM to generate a partial mission specification (M0).

  2. The system uses a pre-defined graph of reusable spacecraft behaviors and domain-specific planning modules to systematically complete M0 into M1 by resolving:

  3. The correct sequence of maneuvers (behavior sequence, e.g., approach, circumnavigate, depart), the corresponding phase durations (timing), and the required waypoint specifications (M2).

  4. The final stage employs trajectory optimization to convert the completed specification (M2) into a state-control trajectory that is guaranteed to be dynamically feasible and safe with respect to orbital dynamics and operational constraints.

This architecture ensures that:

  • Operator intent is preserved throughout the pipeline via the MissionConfig structure (M0, M1, M2).

  • The system avoids hallucinating unstated mission requirements by forcing them through structured planning modules.

  • The final output is not just a sequence of numbers, but a verified trajectory that adheres to safety constraints (like keep-out zones) and operational feasibility.

Sources

Related papers