Flow Motion Policy: Manipulator Motion Planning with Flow Matching Models

arXiv:2604.07084 · cs.RO, cs.AI · Submitted 2026-08-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Flow Motion Policy: Manipulator Motion Planning with Flow Matching Models".

Jane: The paper was written by Davood Soleymanzadeh, Xiao Liang and Minghui Zheng from Texas A&M University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everybody. Today we're digging into a fresh arXiv paper called "Flow Motion Policy: Manipulator Motion Planning with Flow Matching Models." Jane, I've got to say, the title alone got me excited — it's combining two of my favorite things: robot arms and generative models.

Jane: Same here, Tom. And I love that the authors — Davood Soleymanzadeh, Xiao Liang, and Minghui Zheng from Texas A andM — are tackling a problem that's been bugging roboticists for decades. Getting a robot arm to move from point A to point B without hitting anything is way harder than it sounds.

Tom: Right, and the clever part is how they're using flow matching. For our listeners who haven't heard of it, flow matching is like teaching a model to gradually turn random noise into a meaningful trajectory, kind of like how you'd slowly sculpt a block of clay into a statue.

Jane: That's a great way to put it. And what's really cool here is that instead of just predicting one path, this model can generate many different paths for the same problem. That's huge because motion planning is inherently multi-modal — there are often several equally good ways to get from start to goal.

Tom: Exactly. And the team at Texas A andM is showing that this stochastic approach lets them do something called best-of-N sampling at inference time. You generate a bunch of candidate paths, check which ones are collision-free, and pick the first good one you find.

Jane: It's like brainstorming multiple routes to work in the morning and then checking traffic before you commit. The paper shows this dramatically improves success rates compared to deterministic planners that just commit to one path and hope for the best.

Tom: And the implications go beyond just this paper. This is part of a bigger movement toward open-loop neural planners that don't need a privileged collision checker running during planning. That's a game-changer for real-world deployment where computation time matters.

Jane: I love that they're benchmarking against both classical planners like Bi-RRT and BIT*, and also against neural approaches like MPNets and PerFACT. It's a thorough comparison that really shows where this method stands.

Tom: And where it stands is pretty impressive. We'll get into the numbers in a bit, but let me just tease this — the planning times are orders of magnitude faster than sampling-based methods. That's the kind of result that makes engineers sit up and pay attention.

Jane: I'm curious about how they actually built this thing. The architecture — using PointNet++ for point clouds and a transformer encoder — sounds like a modern recipe for success. Let's dig into that next.

Summary: Tom: So we're back, still talking about "Flow Motion Policy: Manipulator Motion Planning with Flow Matching Models." Jane, you wanted to dig into the architecture — and honestly, it's a beautiful piece of engineering.

Jane: It really is. The model takes in the robot's current configuration, the goal configuration, and point clouds of both the robot and the workspace. Then it embeds all of that information and feeds it into a transformer encoder. The transformer basically learns to pay attention to the parts of the scene that matter for planning.

Tom: And then the flow head takes over. It's a relatively small MLP that predicts the velocity field — the direction the trajectory should move at each step of the flow process. The whole model is only about one point four million parameters, which is remarkably lightweight.

Jane: That's a fraction of what some other neural planners use. The paper compares it to Neural MP, which has around twenty million parameters. So Flow Motion Policy is not just faster at inference — it's also much cheaper to train and deploy.

Tom: The training process is elegant too. They use cuRobo, which is a fast optimization-based planner, to generate training data. Then they train the flow model to match the distribution of those expert trajectories. It's supervised learning, but instead of predicting one path, the model learns the whole distribution.

Jane: And that distribution is what enables the best-of-N sampling we mentioned earlier. At inference time, they sample a batch of candidate paths, run collision checking in parallel on the GPU, and pick the first collision-free one. The whole thing happens in fractions of a second.

Tom: The numbers in the paper are striking. With one hundred samples, they hit a ninety-six point seven five percent success rate on the Bins task, compared to forty-eight percent with just one sample. That's a massive jump from simply leveraging the stochastic nature of the model.

Jane: And the planning times are still tiny — under a second in most cases. Compare that to Bi-RRT which takes over two seconds on the same tasks, and you start to see why this matters for real-world robotics.

Tom: I also appreciate that they didn't just compare against sampling-based planners. They benchmarked against other neural methods too — MPNets, SIMPNet, GAIDE, PerFACT. Flow Motion Policy holds its own or beats them across the board, especially when you enable the best-of-N sampling.

Jane: The ablation studies are really thorough as well. They tested different policy head architectures — MLP, U-Net, Transformer, DiT — and showed that the simple MLP head is not only faster but often just as good or better. That's a nice reminder that bigger isn't always better.

Tom: And they compared flow matching against diffusion-based policies too. Flow matching wins on speed because it needs fewer integration steps — around twenty Euler steps versus one hundred denoising steps for diffusion. That's a huge practical advantage.

Jane: So the summary is: a lightweight model, trained on expert data, that can generate diverse candidate paths and pick the best one at inference time. It's fast, it's accurate, and it doesn't need a collision checker during planning. What's not to love?

Tom: I'm dying to know how this holds up in the real world. The paper has a real-world deployment section — let's talk about that.

Improvements: Tom: Alright, we're back with "Flow Motion Policy: Manipulator Motion Planning with Flow Matching Models." Jane, I promised we'd talk about the real-world deployment, and honestly, that's where this paper really shines.

Jane: It does. They set up a UR5e robot arm with an Intel RealSense camera and tested the policy in three real-world environments: bins, articulated shelves, and regular shelves. And the results are honestly remarkable.

Tom: With best-of-N sampling using one hundred paths, they hit a one hundred percent success rate on both the bins and articulated shelves tasks. That's perfect execution across all trials. The base policy without sampling only got fifty percent and thirty percent respectively.

Jane: The shelves environment was tougher — they got sixty percent there. The paper attributes that to the training data not having enough similar scenarios. It's a good reminder that even the best neural planner is only as good as its training distribution.

Tom: But here's what I find really impressive — the improvement from inference-time optimization. The base policy got thirty-three point four percent overall, and with best-of-N sampling, that jumped to eighty-six point seven percent. That's not a small tweak; that's the difference between unusable and production-ready.

Jane: And the beauty is that this improvement comes without retraining. You just sample more paths at inference time. It's like giving the model more chances to get it right, and since the flow model captures the multi-modality of the problem, those chances are actually diverse.

Meng: I've been listening in, and I have to ask — how does the collision checking work in parallel? The paper mentions using cuRobo's batched collision utilities, but what's the actual overhead?

Tom: Great question, Meng. The key is that all the candidate paths are checked simultaneously on the GPU. So instead of checking one path at a time, you check a hundred at once. The cost scales sublinearly because the GPU is designed for this kind of parallel work.

Jane: And that's why the planning time stays under a second even with one hundred samples. The generation is fast because the model is small, and the checking is fast because it's batched. It's a really elegant combination.

Lu: I'd add that this approach also sidesteps one of the biggest bottlenecks in classical planning — collision checking can account for up to ninety percent of computation time in sampling-based planners. By moving collision checking to the end and doing it in parallel, you avoid the sequential cost entirely.

Meng: So the real-world deployment shows this isn't just a simulation toy. It works on actual hardware with noisy point clouds from a real camera. That's the kind of evidence that makes me want to try this in our own lab.

Jane: And that's the exciting part — this isn't just a paper that looks good on paper. It's a paper that works in practice. The improvements they're proposing — stochastic generative policies with inference-time optimization — are immediately applicable to real robotic systems.

Tom: Absolutely. And it opens up so many questions about what's next. Can this scale to more complex tasks? Can it handle dynamic environments? We'll touch on those in our conclusion.

Conclusion: Tom: And we're wrapping up our discussion of "Flow Motion Policy: Manipulator Motion Planning with Flow Matching Models." Jane, what's your final takeaway?

Jane: My takeaway is that this paper represents a real shift in how we think about neural motion planning. Instead of trying to predict the single perfect path, Flow Motion Policy learns a distribution over feasible paths and then uses that distribution to sample multiple candidates at inference time. That's a fundamentally more robust approach.

Tom: And the numbers back it up — ninety-six point seven five percent success on the Bins task, eighty-six point seven percent overall in real-world deployment, all with planning times under a second. Compared to classical planners that take seconds and neural planners that are deterministic, this is a clear step forward.

Jane: I also love that they showed the importance of the generative formulation itself. They compared flow matching against Gaussian Mixture Models and diffusion models, and flow matching came out ahead in both speed and success rate. That's a strong argument for flow-based approaches in robotics.

Tom: And let's not forget the practical implications. This is a lightweight model — one point four million parameters — that runs on a single GPU and works with a single camera setup. That's accessible to a lot of robotics labs and companies, not just the big players.

Jane: The limitations are honest too — the shelves environment showed that generalization is still a challenge, and the paper mentions that static environments are an assumption. But those are natural next steps, not deal-breakers.

Lu: I'd add that the combination of flow matching with best-of-N sampling could extend beyond motion planning. Any sequential decision-making problem with multi-modal solutions — like manipulation planning or even navigation — could benefit from this pattern.

Meng: And from an engineering standpoint, the fact that you can get this performance without a privileged collision checker during planning is huge. It simplifies the system architecture and makes deployment much easier.

Tom: So we'll say goodbye to Flow Motion Policy and the team at Texas A andM — great work, genuinely exciting results. Next up on the show, we've got another paper that caught our eye, and I think it's going to spark some great conversation.

Jane: Thanks for listening, everyone. We'll see you on the next episode.

Tom: Take care, and keep planning those paths.

Davood Soleymanzadeh, Xiao Liang, Minghui Zheng

Texas A&M University

cs.RO, cs.AI

Submitted: 2026-08-15

Updated: 2026-08-18

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 58/100

The gist: Motion planning for manipulators aims to compute a valid, collision-free path connecting an initial configuration to a desired goal configuration.

Key concepts

Flow Matching
A technique used to teach a model to gradually turn random noise into a meaningful trajectory, similar to sculpting clay. It enables the model to learn a distribution over feasible paths instead of just one single path.
Best-of-N Sampling
At inference time, the model generates multiple candidate paths. The system then checks these candidates for collisions and selects the first collision-free path found, which significantly improves success rates compared to deterministic planners.
Open-Loop Neural Planners
This is a movement toward neural planners that do not require a privileged collision checker during the planning phase. This is important for real-world deployment where computation time must be minimized.
Flow Matching vs. Diffusion Models
Flow matching was compared against diffusion models, showing it wins on speed because it requires fewer integration steps—around twenty Euler steps versus one hundred denoising steps for diffusion.

Terminology

Summary

Summary

The paper introduces Flow Motion Policy, an open-loop, end-to-end neural motion planner for robotic manipulators that leverages the stochastic generative formulation of flow matching methods to capture the inherent multi-modality of planning datasets. The framework enables efficient inference-time best-of-N sampling, generating multiple end-to-end candidate paths, evaluating their collision status after planning, and executing the first collision-free solution.

Problem Context: Motion planning for manipulators aims to compute a valid, collision-free path connecting an initial configuration to a desired goal configuration. Sampling-based planners (SBPs) construct a tree in configuration space but rely on a privileged collision checker during planning, which accounts for up to 90% of computation time, making parallelism and inference-time optimization challenging. Trajectory optimization methods iteratively refine an initial trajectory but also require a privileged collision checker for collision status and gradient information. End-to-end neural motion planners learn a direct mapping from planning problem information (start configuration, goal configuration, environment observations) to a feasible path without relying on a privileged collision checker during planning. However, most existing approaches are deterministic policies that produce a single path for a given planning problem and cannot exploit best-of-N sampling at inference time.

Method: Flow Motion Policy models a conditional distribution over joint configuration increments over a planning horizon H: pθ(δqt:t+H−1qt, qgoal, P), where qt is the current timestep configuration, qgoal is the goal configuration, and P is the environment observation representation. The path is generated autoregressively by predicting configuration increments. The framework learns a flow-based policy under the flow matching framework to model the distribution of planning data p0(x0), where x0 = δqt:t+H−1. This is achieved by constructing a Continuous Normalizing Flow with an optimal transport probability path: xτ = (1 − τ)x0 + τϵ, τ ∈ [0, 1], which linearly interpolates between clean data x0 at τ = 0 and Gaussian noise ϵ ∼ N(0, I) at τ = 1. The model learns an estimator vθ(xτ, τ, O) conditioned on planning observation by regressing it toward the conditional vector field, with the loss function: L = ET(τ), p0(x0), pτ(xτx0) ∥vθ(xτ, τ, O) − uτ(xτx0)∥2.

Network Architecture: The architecture embeds current and goal configurations via shared multi-layer perceptrons (MLPs): ht = MLP(qt), hgoal = MLP(qgoal). Robot and scene point clouds are down-sampled and embedded with set abstraction layers from PointNet++. These embeddings (ht, hgoal, hr, hw) are augmented with learnable token embeddings and fed into a transformer encoder. The flow head is instantiated as an MLP that operates on learnable action tokens and incorporates planning information from the transformer encoder output, predicting the continuous-time vector field vt used to generate actions through flow-based sampling.

Inference: During inference, samples are obtained by integrating the learned vector field vθ(xτ, τ, O) from τ = 1 to τ = 0 to recover x̂0 ∼ p0. Flow Motion Policy is evaluated autoregressively in an open-loop manner, rolled out for a fixed number of steps. Planning success is defined as the end-effector reaching within a predefined threshold of the planning goal pose configuration. For best-of-N sampling, the framework samples N candidate paths and evaluates each trajectory using a collision checker, selecting the first collision-free path: q* = arg min qi C(qi), where the cost function C(qi) = ΣTt=1 I d(qit, P) < δsafe, with d(qit, P) denoting the minimum signed distance between the robot at configuration qit and the environment representation P, I(.) an indicator function, and δsafe the safety distance threshold.

Training: The authors leverage an LLM-based workspace generator to create a large number of motion planning environments and use cuRobo, a fast optimization-based motion planner, for planning data collection to train Flow Motion Policy.

Evaluation: The framework is benchmarked against Bi-Directional RRT (Bi-RRT), Batch Informed Trees (BIT*), Motion Planning Networks (MPNets), Spatial-informed Motion Planning Network (SIMPNet), GAIDE, and PerFACT across held-out planning tasks (TableTop, Box, Bins, Shelf Task I, Task II, Task III). Metrics include planning time and success rate.

Key Results:

  • Flow Motion Policy with inference-time optimization achieves performance comparable to classical sampling-based motion planning benchmarks while requiring significantly lower planning time. For example, on TableTop, FMP-100 achieves 84% success in 0.58±0.44 seconds, while Bi-RRT achieves 88% in 2.85±0.97 seconds.

  • The base policy without inference-time optimization is an order of magnitude faster than sampling-based planners but has lower success rates. FMP-1 achieves 48% success on TableTop in 0.16±0.01 seconds.

  • Flow Motion Policy outperforms neural informed samplers (MPNets, SIMPNet, GAIDE) in both success rate and planning time.

  • The base policy performs comparably to PerFACT while requiring lower planning time due to a smaller network size (1.4M parameters vs. 4.15M for PerFACT).

  • Flow Motion Policy consistently outperforms the adapted Neural MP baseline across all held-out planning tasks, attributed to the flow matching formulation providing a more expressive generative modeling framework than Gaussian Mixture Models and the substantially lighter model size (1.4M vs. 20M parameters).

Ablation Studies:

  • Policy head architecture: Flow Motion Policy with an MLP head achieves the fastest planning time among all evaluated policy head architectures (U-Net, Transformer, DiT-Block). Diffusion Motion Policy requires significantly longer inference time due to iterative denoising with 100 steps.

  • Inference-time optimization: Increasing the number of sampled trajectories improves planning success rate across all policy head architectures but increases planning time. More complex heads attain improved success rates but with higher computational overhead per trajectory.

  • Diffusion timesteps: Decreasing the number of diffusion timesteps reduces average planning time but at the expense of average planning success rate.

  • Inference flow steps: Increasing the number of Euler solver steps during inference does not lead to significant improvement in success rate across all policy head architectures, while average planning time increases with the number of solver steps.

Real-World Deployment: Flow Motion Policy is evaluated on a UR5e robotic manipulator in real-world environments using a calibrated Intel RealSense D435i RGB-D camera with AprilTag markers for calibration. Across real-world tasks (Bins, Articulated, Shelves), Flow Motion Policy with best-of-N sampling achieves a substantially higher success rate compared to the base policy without inference-time optimization (86.7% vs. 33.4%). In the bins environment, the policy with inference-time optimization solves all planning instances (10/10), while the base policy achieves only 50% success. In articulated environments, the policy with inference-time optimization achieves 100% success, while the base policy drops to 30%. In shelves environments, both variants exhibit reduced performance due to limited representation of similar scenarios in the training dataset.

Limitations: The current formulation assumes static environments and does not explicitly account for dynamic obstacles. The generated paths may lack smoothness or local optimality and require post-processing before execution on real robots.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system, along with the resulting capabilities:

1. Replace deterministic motion planning heads with a flow-matching generative head

  • The current system outputs a single deterministic path. I will replace the output layer with a conditional flow-matching model that learns a distribution over feasible joint-configuration increments (δq) given current state, goal, and point-cloud observations.

  • This enables sampling multiple candidate paths from the learned distribution rather than a single prediction.

2. Add inference-time best-of-N sampling with parallel collision checking

  • I will implement a batch-sampling mechanism that generates N candidate trajectories (e.g., N=100) in parallel.

  • I will add a lightweight, GPU-parallel collision checker (using signed distance functions) that evaluates all N paths simultaneously and selects the first collision-free trajectory.

  • This converts the planner from a single-shot predictor to an anytime, multi-hypothesis planner.

3. Implement an autoregressive action-chunking rollout with flow integration

  • I will replace the single-step prediction with an autoregressive loop that predicts H-step action chunks (e.g., H=16) using Euler integration of the learned vector field (20 steps).

  • This allows the system to plan longer horizons while maintaining temporal consistency, and it enables re-planning at each chunk boundary.

4. Use a transformer encoder with tokenized multi-modal embeddings

  • I will encode the robot point cloud, workspace point cloud, current configuration, and goal configuration as separate tokens (using PointNet++ for point clouds and MLPs for configurations).

  • I will add learnable token-type embeddings and feed these into a transformer encoder to condition the flow head on fused spatial and kinematic context.

5. Replace GMM or diffusion heads with a lightweight MLP flow head

  • I will use a small MLP (1.4M parameters) as the flow head instead of larger U-Net, transformer, or DiT heads (2.3M–3.2M parameters) or a 20M-parameter RNN-GMM.

  • This reduces inference latency by 3–5× while maintaining comparable or better success rates due to the flow-matching formulation.

6. Add training with optimal-transport conditional flow matching

  • I will train the model using the flow-matching loss: L = vθ(x t, t, O) − (ε − x0)2, where x t = (1−t)x0 + tε, with t U[0,1].

  • This provides a stable, fast-converging training objective compared to diffusion denoising, and it directly optimizes for the vector field used at inference.

1. Generate multiple diverse, collision-free motion plans in real time

  • Given a start configuration, goal configuration, and a point cloud of the environment, the system produces 100 candidate paths in 0.6 seconds (vs. 2–6 seconds for sampling-based planners) and selects the first collision-free one.

  • It achieves 84–97% success rates on tabletop, box, bin, and shelf tasks, matching or exceeding Bi-RRT and BIT* while being 4–10× faster.

2. Operate as both a direct planner and an anytime planner

  • With N=1, it acts as a fast open-loop planner (0.16s) for time-critical applications.

  • With N=100, it acts as a robust planner that improves success rates by 30–50 percentage points (e.g., from 48% to 84% on tabletop tasks) with only a modest increase in latency.

3. Generalize to unseen environments and real-world deployment

  • The system works with raw RGB-D point clouds (e.g., from an Intel RealSense D435i) and does not require a privileged collision checker during path generation.

  • In real-world tests on a UR5e manipulator, it achieves 86.7% success with best-of-N sampling (vs. 33.4% without), across bins, articulated shelves, and shelf tasks.

4. Provide faster inference than diffusion-based policies

  • The flow-matching head requires only 20 Euler steps (vs. 100 denoising steps for diffusion), reducing planning time by 2–5× (e.g., 0.58s vs. 2.87s for DiT-based diffusion with N=100).

  • It also uses 1.4M parameters vs. 3.2M for DiT, enabling deployment on edge GPUs.

5. Support parallel batch planning for multi-robot or multi-goal scenarios

  • The batched sampling and collision-checking design allows the system to plan for multiple robots or multiple goals simultaneously, which is not possible with sequential sampling-based planners.

6. Enable safe execution by selecting the first collision-free path

  • The system evaluates all candidate paths for collisions and executes the first valid one, reducing the risk of executing a colliding trajectory—a critical safety improvement over deterministic neural planners.

7. Adapt to new tasks with minimal fine-tuning

  • Because the flow-matching formulation captures multi-modality, the system can be fine-tuned on new environment types (e.g., shelves, bins) with only a small dataset, while retaining the ability to sample diverse paths.

Sources

Related papers