Masked Generative Motion Planning with Geometry-Guided Token Search
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "Masked Generative Motion Planning with Geometry-Guided Token Search".
Dev: The gist The Masked Generative Motion Planning (MGMP) introduces a method that extends learned trajectory priors from efficient parallel generation to structural repair,
Rosa: First, who's behind it and why it matters.
Paper summary: Dev: So, wrapping up on Masked Generative Motion Planning with Geometry-Guided Token Search, the authors are showing that their method achieves high success rates on challenging tasks and generalizes across unseen layouts and environments.
Taro: The implications for autonomy is that this learned trajectory prior can serve not just as a generator for initial motions but also as a structured proposal space for adapting to new geometric constraints during operation.
Rosa: The main point is moving the repair mechanism from local continuous refinement to a guided search over discrete alternatives, which they call GGTS. This allows for route-level restructuring beyond what local deformation can achieve.
Dev: It’s important to remember that this discrete search provides the main capability for basin escape, while the final polishing stage is only used to resolve residual local infeasibilities. The method is sensitive to initialization and local planning basins, which the authors noted as a limitation.
Taro: So for future work, they're looking at learning reusable structure across different planning problems, which suggests this token-based approach might be adaptable beyond just motion planning scenarios.
Rosa: That’s the gist of what they did here: using discrete search to restructure trajectories guided by geometry to achieve structural repair instead of relying on continuous optimization for the whole job.
Conclusion: Rosa: So, we’re wrapping up on Masked Generative Motion Planning with Geometry-Guided Token Search. The title itself hints at this whole shift from just generating paths to actually fixing them structurally using geometry as a guide.
Dev: Yeah, it’s about taking those learned motion ideas and using them not just to make something that *looks* right initially, but to actively search for the right route when things go wrong in real-time.
Taro: What I find interesting is how they framed test-time repair as a search over discrete tokens instead of just letting continuous optimization wiggle around the existing path. It turns it into a targeted quest.
Rosa: Exactly. They’re using scene geometry to tell the AI exactly *where* to look for errors and which alternative motion ideas are worth checking. It's not blindly trying every possible change in space anymore; it's focusing on what matters geometrically.
Dev: From an engineering side, that means instead of running a huge, continuous optimization routine every time a collision pops up, the AI just swaps out specific pieces of the trajectory based on those geometric scores they calculate. That should keep the loop rate manageable and hopefully faster.
Taro: And that’s what makes it powerful for autonomy—it’s giving the system a way to make large structural changes, like finding a completely different route around a new obstacle layout, without having to start the whole generation process over again.
Rosa: It really moves the capability beyond just local adjustments, which is where those older methods usually hit a wall. They show this approach works even on things like single- and dual-arm planning and real robot tasks.
Dev: The results they showed on things like the Ring Maze—ninety-six percent success—that’s a solid number when you think about how much more robust it is than just local refinement alone. It shows the discrete search part handles the big stuff, and only a little continuous polishing is needed for the fine details.
Taro: It suggests that learned priors aren't just good starting points; they can be turned into a structured library of valid alternatives you can search through when the environment changes unexpectedly. That opens up new ways to handle complex planning scenarios.
Rosa: So, it’s about making those learned trajectory ideas work as a dynamic toolbox for structural repair during operation, rather than just a static plan generator. That’s what this paper is really pushing for.
Lipeng Zhuang, Yingdong Ru, Shiyu Fan, Edmond S. L. Ho Gerardo Aragon Camarasa Paul Henderson
University of Glasgow
cs.RO, cs.AI
Submitted: 2026-10-07
Updated: 2026-10-07
The gist: The gist The Masked Generative Motion Planning (MGMP) introduces a method that extends learned trajectory priors from efficient parallel generation to structural repair, enabling route-level
Key concepts
- Masked Generative Transformer (MGT)
- This model learns compact motion tokens from expert trajectories. It predicts masked codes in parallel, allowing for the generation of diverse trajectory candidates efficiently during the initial planning phase. These discrete tokens form a reusable space for subsequent structural repair.
- Geometry-Guided Token Search (GGTS)
- GGTS uses scene geometry to guide two key decisions: identifying where to edit (by selecting tokens causing violations) and determining what new tokens to fill in that region. This targeted search avoids random changes, focusing on making large, coordinated structural edits.
- Discrete Trajectory Tokens
- Instead of working with continuous trajectories, MGMP maps expert paths into discrete codes or tokens. These tokens represent fundamental motion elements. By searching and replacing these discrete tokens, the method can restructure the entire motion sequence rather than just slightly deforming it.
- Test-Time Refinement Process
- This iterative process refines a trajectory by making decisions round by round: which positions to re-mask and what new codes to insert. Candidates are screened using geometric scores and verified against exact objectives before being committed, ensuring only improving edits are accepted.
Terminology
Summary
The gist The Masked Generative Motion Planning (MGMP) introduces a method that extends learned trajectory priors from efficient parallel generation to structural repair, enabling route-level restructuring beyond local trajectory deformation
How it works
MGMP reframes motion planning by formulating it as masked generation over discrete trajectory tokens, reusing the same token space for test-time structural repair. A trajectory tokenizer maps expert trajectories into discrete codes and reconstructs continuous trajectories, while a Masked Generative Transformer (MGT) predicts masked codes conditioned on the start state, goal state, and environment.
Geometry-Guided Token Search (GGTS)
To make the search effective, Geometry-Guided Token Search (GGTS) uses scene geometry to guide two decisions: where to edit and what to fill. Geometric attribution determines where to edit by selecting the token positions most responsible for the current violation while preserving unaffected parts of the trajectory. To decide what to fill, the MGT proposes alternative motion tokens within the selected repair region, which are first screened by a geometric score and then verified through exact trajectory evaluation before an improving edit is finalized. This targeted, coarse-to-exact search avoids unguided perturbations and exhaustive codebook scans, inducing large, coordinated changes in motion structure.
Test-Time Refinement Process
The refinement process involves iterative rounds where each round makes two decisions: which positions to re-mask and what codes to put there. The first MGT forward pass defines the token distribution pθ(zi c) at each position, and the Gauss-Southwell rule is used to select regions requiring refinement based on the magnitude of the gradient. Following this, candidate proposals are screened using a first-order approximation of the geometric objective J to rank replacements, and shortlisted candidates are verified using hard-token trajectories against the exact objective J. The best replacement is committed only if it improves the current solution.
Evaluation and Generalization
MGMP is validated on diverse tasks and environments, showing generalization across unseen layouts, additional obstacles and unseen geometries, single- and dual-arm planning, and real-world Baxter tasks. Diagnostic benchmarks like Ring Maze require substantial route changes to escape local planning basins, where MGMP achieves 96% success. The method generalizes to unseen obstacle layouts and geometries by utilizing DeepSDF for differentiable collision costs during refinement.
Performance Metrics
MGMP achieves 96% success on Ring Maze and 82% repair success on Controlled Route Invalidation, exceeding the strongest external baselines by 23 and 25 percentage points, respectively. In real-world experiments with a Baxter robot, MGMP performs consistently well across shelf placing, button pressing, and drawing tasks. The method demonstrates that GGTS provides most of the basin-escape capability while the final polishing stage resolves residual local infeasibilities.
Ablation Study Insights
The ablation study isolates the main design choices in GGTS, showing that geometry-only remasking provides a higher success rate than confidence-only remasking for where to edit. Furthermore, the analysis indicates that discrete search performs the main structural repair while continuous optimization resolves residual local infeasibilities. The use of k=2 for candidate verification and a 2x2 refinement schedule achieved the best overall performance in several settings.
Conclusion
MGMP reframes generative motion planning by keeping the learned trajectory prior active beyond initial generation, while GGTS turns prior-supported motion alternatives into a geometry-guided, targeted search space. Rather than relying on local continuous refinement to discover new routes implicitly, MGMP directly restructures trajectories through discrete token replacement, using continuous optimization only for residual correction. Controlled basin-escape and route-invalidation experiments show that discrete search provides the main repair capability while broader evaluations demonstrate generalization across diverse planning settings. More broadly, MGMP shows how learned trajectory priors can serve not only as generators, but also as structured proposal spaces for geometry-guided test-time adaptation.
--- Page 1 ---
MASKED GENERATIVE MOTION PLANNING WITH GEOMETRY-GUIDED TOKEN SEARCH Lipeng Zhuang Yingdong Ru Shiyu Fan Edmond S. L. Ho Gerardo Aragon Camarasa Paul Henderson University of Glasgow, Glasgow, United Kingdom ABSTRACT Generative motion planners typically use learned trajectory priors for initial generation, while leaving test-time repair to local continuous refinement. We introduce Masked Generative Motion Planning (MGMP), which extends the learned prior from efficient parallel generation to structural repair. A masked generative transformer generates discrete trajectory candidates in parallel, and Geometry-Guided Token Search (GGTS) uses scene geometry to target where to edit and which priorsupported alternatives to evaluate. This turns refinement into an efficient search over discrete motion alternatives, enabling route-level restructuring beyond local trajectory deformation. MGMP achieves 96% success on Ring Maze and 82% repair success on Controlled Route Invalidation on Kuka, exceeding the strongest external baselines by 23 and 25 percentage points, respectively. It further generalizes to unseen layouts, additional obstacles, unseen geometries, single- and dual-arm planning, and real-world Baxter tasks. 1 INTRODUCTION Motion planning in robotics aims to find a feasible trajectory between a start and goal configuration while satisfying geometric constraints (LaValle, 2006). Trajectory-optimization methods directly refine trajectories against task-specific objectives (Ratliff et al., 2009; Schulman et al., 2014), but their performance is typically sensitive to initialization in cluttered and non-convex environments (Pan et al., 2014). Generative planners mitigate this sensitivity by learning multimodal trajectory priors and producing diverse candidate motions that can subsequently be refined to satisfy collision constraints. Generative motion planning therefore involves two coupled inference problems: efficient initial trajectory generation and effective test-time geometric refinement. Existing planners make trade-offs on both fronts. For initial generation, diffusion-based methods rely on iterative denoising in a high-dimensional continuous space (Janner et al., 2022; Carvalho et al., 2023), leading to high inference cost; fewer denoising steps improve efficiency but can reduce sample quality. For test-time refinement, generated trajectories are typically refined through continuous guidance or trajectory optimization (Carvalho et al., 2023; Luo et al., 2024; Yang et al., 2025). Weak guidance may leave collisions unresolved, while stronger guidance can move trajectories away from the learned prior (Yang et al., 2025). Such continuous refinement is effective when feasibility can be recovered through local adjustments; however, in harder cases such as near-enclosing obstacles, reaching a feasible alternative may require many iterations, remain sensitive to initialization and guidance strength, and still fail to escape unfavorable local basins. We therefore ask: How can a learned trajectory prior directly expose qualitatively different routes as candidates for structural repair, rather than expecting continuous refinement to reach them? We introduce Masked Generative Motion Planning (MGMP), which addresses this question by formulating motion planning as masked generation over discrete trajectory tokens, reusing the same token space for test-time structural repair. A trajectory tokenizer learns compact motion tokens from feasible trajectories. For initial generation, a Masked Generative Transformer (MGT) (Chang et al., 2022; 2023) predicts token sequences in parallel, producing high-quality trajectory candidates.
--- Page 2 ---
Figure 1: Overview of MGMP. We replace continuous trajectory generation and refinement with discrete masked generation and Geometry-Guided Token Search, enabling efficient planning and structural route repair. We evaluate it across challenging simulated and real-world settings, including basin escape, unseen geometries, additional obstacles, dual-arm planning, and Baxter tasks. in a single forward pass and reducing the need for costly resampling and post-hoc repair during online planning. Importantly, the same representation allows us to cast repair as a guided search over prior-supported motion alternatives. To make this search effective, we introduce Geometry-Guided Token Search (GGTS), which uses scene geometry to guide two decisions: where to edit and what to fill. Geometric attribution determines where to edit by selecting the token positions most responsible for the current violation while preserving unaffected parts of the trajectory. To decide what to fill, the MGT proposes alternative motion tokens within the selected repair region, which are first screened by a geometric score and then verified through exact trajectory evaluation before an improving edit is finalized. This targeted, coarse-to-exact search avoids unguided perturbations and exhaustive codebook scans; instead, a small number of targeted discrete replacements induce large, coordinated changes in motion structure, enabling exploration of alternative routes beyond the current planning basin.
--- Page 3 ---
Figure 2: Overview of GGTS. We use geometric attribution to select edit positions, first-order scores to screen token replacements, and sequential verification to accept improving candidates, followed by optional continuous polish.
Improvements for AI systems
-
textbfMasked Generative Motion Planning (MGMP) for Structural Repair: Enables route-level restructuring beyond local deformation by
formulating motion planning as masked generation over discrete trajectory tokens.
This allows the system to achieveroute-level restructuring beyond local trajectory deformation
by leveraging a shared token space for both initial generation and test-time repair. -
textbfGeometry-Guided Token Search (GGTS) for Targeted Repair: Turns refinement into
an efficient search over discrete motion alternatives, enabling route-level restructuring.
It uses scene geometry to determinewhere to edit and what to fill,
moving beyond unguided perturbations by employing a process where it selects token positions based on the gradient of the geometric objective J. -
textbfEnhanced Robustness via Multi-Stage Refinement: The system separates structural repair from initial generation quality by using
discrete token replacement provides structural repair, while continuous optimization performs only local correction.
This allows forbasin escape
through discrete search, as evidenced by achieving96% success on Ring Maze
and repairing trajectories even when the initial generation is poor. -
textbfGuaranteed Repair Isolation: Controlled Route Invalidation allows researchers to isolate repair capability by initializing all methods from the same trajectory and testing
whether structural repair capability from initial-generation quality.
This proves that MGMP can achieve82% repair success on Controlled Route Invalidation,
demonstrating its ability to fix trajectories that requiresubstantial route changes.
-
textbfReal-World Task Generalization: MGMP demonstrates generalization to complex, unseen scenarios, including
unseen layouts, additional obstacles, unseen geometries
and real-world Baxter tasks. This is achieved by using techniques likeDeepSDF
for collision costs during refinement and adapting the MGT architecture via AdaLN conditioning for different robot arms (single-, dual-arm planning).
Sources
- Motion Planning in Compressed Representation Spaces
- ShapeNet: An Information-Rich 3D Model Repository
- Path Planning for Masked Diffusion Model Sampling
- Riemannian Motion Policies
- SafeFlowMatcher: Safe and Fast Planning using Flow Matching with Control Barrier Functions
- Masked Generative Policy for Robotic Control
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving