FLASH: Efficient Visuomotor Policy via Sparse Sampling
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "FLASH: Efficient Visuomotor Policy via Sparse Sampling".
Rosa: Fast Legendre-polynomial Action policy via Sparse History-anchored flow (FLASH Policy) is a generative visuomotor policy that represents trajectories as continuous Legendre polynomial coefficients to enable single-step inference and precise torque…
Dev: First, who's behind it and why it matters.
Paper summary: Rosa: So we're looking at this paper titled "FLASH: Efficient Visuomotor Policy via Sparse Sampling," which claims it replaces discrete action chunk generation with a continuous Legendre polynomial trajectory representation. My main question is, Rosa and Dev, what exactly does this mean for the overall policy structure?
Dev: It means they're moving away from generating separate actions and instead treating the entire trajectory as a continuous function defined by coefficients of a Legendre polynomial, which lets them do single-step inference and precise torque control. This seems to tackle the latency issues inherent in iterative denoising methods.
Taro: From an autonomy standpoint, I'm interested in how this handles unexpected situations; if the world misbehaves during execution, does this continuous representation allow for smoother recovery than chunk-based policies?
Rosa: That’s a good point, Taro; the paper suggests that because the trajectory is represented as a continuous polynomial function of time variable s from zero to one, you can compute desired velocity feed-forward signals directly through simple differentiation of those coefficients. This analytic computation of velocity seems much more robust than relying on discrete steps.
Dev: And what makes this approach so fast, Rosa? The summary mentions replacing iterative denoising with a flow matching mechanism that initiates generation from history polynomial coefficients instead of uninformative Gaussian noise to shorten the transport distance and enable accurate single-step inference. That sounds like a big win for real-time systems.
Taro: If the flow matching starts from the history, it implies that we aren't starting from scratch every time; we are leveraging what happened before to guide what happens next, which should definitely make adaptation faster when dynamics shift unexpectedly.
Rosa: Exactly, Taro; they use a sparse temporal sampling strategy where expert trajectories are fitted to these polynomial coefficients at a significantly long temporal stride, allowing one inference to cover an extended action horizon without increasing the model scale. This is how they keep the model size manageable while achieving that extended prediction capability.
Dev: The efficiency gains sound substantial, especially since they report that per-episode inference time is thirty-one point four zero milliseconds and per-call latency hits zero point three two milliseconds, which is significantly faster than diffusion policies and prior flow matching policies mentioned in the paper "FLASH: Efficient Visuomotor Policy via Sparse Sampling."
Paper summary: Taro: That speed metric is what really gets my attention; if we can achieve that level of inference speed, it opens up possibilities for much more reactive autonomy where the robot can respond to environmental changes almost instantaneously. What about the actual control loop rate on a physical system?
Rosa: That's where I want to get specific; while they focus heavily on the computational speed of generating coefficients, we need to see how this translates to the physical hardware. The paper describes using cross-horizon kinematic continuity constraints and a closed-form KKT correction (Mellinger and Kumar, two thousand eleven) to ensure C1 continuity at transition points between polynomial chunks <ref:2605.15492#pg0>.
Dev: That constraint handling is crucial for smoothness; it’s not just about speed, but about making sure the resulting continuous trajectory is physically viable and smooth enough for the torque controller to handle without excessive jitter or instability.
Taro: If those constraints enforce C1 continuity, it suggests that even when we're using sparse sampling for training, the model maintains a high level of kinematic fidelity across time steps, which is vital when navigating complex environments where small errors compound quickly.
Rosa: Precisely; they also added "fit padding steps" to pull constraint anchors into a stable interior of the fitting window, which I see as another layer in ensuring those continuity constraints are met reliably during deployment. So, we have this representation that's fast and smooth across time steps.
Dev: It sounds like the architecture is designed to be highly efficient at inference by leveraging that continuous representation and history-anchored flow matching, minimizing the computational work needed for every decision point. This really speaks to the engineer in us about how much overhead we can cut out of the pipeline.
Taro: I think the core implication here is shifting policy learning from discrete action chunks to a continuous functional space, which fundamentally alters how we model motion planning and control altogether when dealing with complex, dynamic environments where precise trajectory following is essential.
Rosa: And what about its real-world viability? My primary concern is whether this works outside the controlled lab setting for extended periods; I want to know if the learned policy remains robust when faced with novel physical interactions or sensor noise that isn't perfectly modeled in training.
Dev: The paper itself doesn't detail long-term operational robustness, but it does point out a limitation in its objective function: they use a combined loss of Flow-Matching Loss and Polynomial Consistency Loss, and the parameters for the least-squares solver and KKT correction are precomputed and fixed during training. This suggests that adapting the solver itself might be difficult if we want to fine-tune it on a completely different physical setup.
Paper summary: Taro: That limitation is important; if the underlying mathematical framework relies so heavily on precomputed solvers, we might face challenges when deploying this system in a domain with significantly different dynamics than what was captured in the expert demonstrations.
Rosa: So, to summarize this paper "FLASH: Efficient Visuomotor Policy via Sparse Sampling," it presents a method that uses continuous Legendre polynomial coefficients for action representation and history-anchored flow matching to achieve single-step inference speeds up to one hundred seventy-five times faster than diffusion policies.
Dev: That speed difference is significant, especially when considering the reported success rates, which are state-of-the-art across simulated and real tasks, achieving at least ninety-two percent success rates with a single inference.
Taro: The implication for autonomy is that we can move toward systems that exhibit much lower latency in their decision-making loops, potentially allowing for high-frequency interactions with the physical world without the control system lagging behind the environment's dynamics.
Rosa: And if we look at its title and authors, Jiaqi Bai, Jindou Jia, Yuxuan Hu, Gen Li, Xiangyu Chen, Tuo An, Kuangji Zuo... it shows a strong collaboration between theoretical formulation and practical implementation in this area of policy learning.
Dev: The way they handle the continuity constraints using KKT corrections and fit padding steps is a clever engineering solution to keep that mathematical representation physically sound during execution.
Taro: I think the paper suggests that by focusing on representing the action in a functional space rather than discrete points, we are building a more naturally continuous model of robot motion, which is what truly matters for complex manipulation tasks.
Rosa: So, for the listeners tuning in right now, we're talking about FLASH: Efficient Visuomotor Policy via Sparse Sampling, a method that uses polynomial representations and flow matching to drastically cut down inference time while maintaining high accuracy.
Dev: It's certainly a paper that shows how mathematical representation can directly translate into tangible performance metrics like speed and tracking error reduction compared to older methods.
Taro: The impact could be seen in applications requiring high-speed interaction, like agile manipulation or even drone control, where minimizing the time between sensing and acting is paramount for safe operation.
Conclusion: Rosa: So we’ve just been diving into FLASH: Efficient Visuomotor Policy via Sparse Sampling, which tackles how to make robot policies way faster using continuous mathematical representations of motion instead of discrete chunks.
Dev: Exactly, and I'm really focused on the technical side—how they manage that inference speed and what it means for our actual control loop performance.
Taro: From an autonomy angle, I keep thinking about how this continuous flow matches how a real system would move when things get messy or unexpected.
Rosa: That’s the core of my question, Taro; I want to know if this policy is reliable enough to handle the unpredictable nature of real-world manipulation outside of a perfectly controlled lab setting, and for how long can we trust it?
Dev: I worry about failure modes; if there's a sudden change in dynamics or sensor noise during operation, can this model recover smoothly without causing jerky movements or instability in the torque control loop?
Taro: If the underlying mechanism allows for that kind of smooth, continuous motion modeling, it suggests a much more resilient behavior when confronted with environmental shifts compared to systems based on fixed action sequences.
Rosa: I’m hoping this paper shows some evidence that this mathematical foundation translates into actual robust physical performance, not just theoretical speed gains in simulation.
Dev: We need to look closely at the constraint handling they used, like those KKT corrections and fit padding steps, because those are what ensure the resulting trajectory is actually physically sound for a controller to execute.
Taro: That continuity management is critical; if the model maintains C1 continuity between different polynomial segments, it should prevent that kind of jarring transition we see in older methods.
Rosa: It seems like this work is positioning itself as a way to bridge the gap between theoretical continuous dynamics and practical, high-speed robotic execution.
Dev: I'm really interested in the performance claims they made regarding latency; if we can get that kind of low per-call latency down, it changes how quickly a robot can react to unexpected physical events.
Taro: That rapid reaction time is exactly what we need for complex autonomy, but I still want to probe what happens when the world throws something completely novel at the policy.
Rosa: It’s exciting because they’re showing us a path toward policies that are both highly efficient computationally and kinematically smooth in their outputs.
Dev: Indeed, this paper is pushing the boundary on efficiency while trying to keep those real-world execution concerns—like loop rate and stability—in mind.
Taro: So, we’re looking at a framework that moves away from discrete steps toward a continuous mathematical description of movement for better autonomy.
Rosa: Right, and it's certainly got some strong backing with the success rates they reported across various tasks.
Jiaqi Bai, Jindou Jia, Yuxuan Hu, Gen Li, Xiangyu Chen, Tuo An, Kuangji Zuo
Nanyang Technological University
cs.RO, cs.CV
Submitted: 2026-05-15
Updated: 2026-10-05
Comments: Accepted at NeurIPS 2026. Code: https://github.com/NTUMARS/FLASH-Policy
Project page: https://b1ue-jay.github.io/FLASH
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 87/100
The gist: Fast Legendre-polynomial Action policy via Sparse History-anchored flow (FLASH Policy) is a generative visuomotor policy that represents trajectories as continuous Legendre polynomial coefficients to
Key concepts
- Polynomial Representation
- Actions are modeled as continuous functions of time using Legendre polynomials. This mathematical structure ensures the trajectory is spatially smooth and allows for easy calculation of desired velocities through simple differentiation, which is crucial for accurate control.
- Sparse Temporal Sampling
- Expert trajectories are fitted to polynomial coefficients at very long time intervals. This strategy enables a single inference step to cover an extended action horizon without needing a massive model, making the policy computationally efficient.
- History-Anchored Flow Matching
- The generation process starts not from random noise, but from the polynomial coefficients of previous actions (history). This anchors the flow matching process to known past states, which shortens the transport distance and enables highly accurate single-step inference.
Terminology
Summary
Fast Legendre-polynomial Action policy via Sparse History-anchored flow (FLASH Policy) is a generative visuomotor policy that represents trajectories as continuous Legendre polynomial coefficients to enable single-step inference and precise torque control, achieving state-of-the-art performance in efficiency and accuracy.
How it works
The core innovation of FLASH involves replacing discrete action-chunk generation with a continuous Legendre polynomial trajectory representation, which decouples the network’s output dimensionality from the prediction horizon. This is achieved through two primary mechanisms:
- Sparse temporal sampling strategy: Expert trajectories are fitted to polynomial coefficients at a
significantly long temporal stride,
allowing a single inference to cover an extended action horizon without increasing model scale. The optimal polynomial coefficients are obtained via ordinary least squares (OLS):
C⋆ = (S⊤S)−1 S⊤Q, where Q is the stacked matrix of observed sparse actions and S is the Legendre basis matrix evaluated at corresponding discrete timesteps.
- History-anchored flow matching: Instead of initiating generation from uninformative Gaussian noise, FLASH
initiates the flow matching process from history polynomial coefficients,
which shortens the transport distance and enablesaccurate single-step inference.
This is achieved by defining a generative flow directly between history and future polynomials, conditioned on observations.
Key Technical Components
The policy integrates several sophisticated components to ensure smoothness, continuity, and efficiency:
)&Polynomial Representation:
The action trajectory is parameterized as a continuous function of a normalized time variable s ∈ [0, 1] using Legendre polynomials: a(s) = X K j=0 cj Φj (s), where cj are the coefficients. This representation offers analytic computation of desired velocity via simple differentiation
and guarantees the spatial smoothness of the trajectory.
)&Cross-horizon Continuity and Two-Tier Extension:
To alleviate discontinuities between adjacent polynomial chunks, cross-horizon kinematic continuity constraints are introduced to guarantee C1 continuity (position and velocity) at the transition points.
This is enforced using a closed-form KKT correction (Mellinger and Kumar, 2011) to project the unconstrained coefficients onto the constraint manifold. Furthermore, fit padding steps
are appended to pull constraint anchors into a stable interior of the fitting window.
Learning Objectives and Training
The model is trained using a combined objective function:
)&Flow-Matching Loss (LFM):
This loss minimizes the transport velocity along every intermediate point of the polynomial flow: LFM = Eτ∼U[0,1], Ch, C1 − (C1 − Ch)∥2.
)&Polynomial Consistency Loss (Lcons):
To bridge the gap between continuous flow and discrete inference, a consistency term is added to enforce transport in a single evaluation of fθ at τ = 0: Lcons = ECh, C1 − (Ch + fθ(Ch, 0, e))∥2.
)&Total Objective:
The final training loss is Ltotal = LFM + λcons Lcons. All parameters related to the least-squares solver and KKT correction are precomputed and remain fixed during training.
Performance Results
Extensive experiments on five simulated and two real-world manipulation tasks demonstrate superior performance:
-
Success Rates: FLASH achieves
state-of-the-art success rates (≥ 92% across all tasks)
with a single inference (NFE=1). It outperforms the homologous FLASH-G by an average of 17.6 absolute percentage points. -
Inference Speed: The per-episode inference time is
31.40 ms,
which isup to 175× faster than diffusion policies and 18× faster than prior flow matching policies.
Per-call latency is exceptionally low at 0.32 ms, making it16× faster
than the median of the baselines. -
Controller Tracking Error: FLASH achieves a
5× to 7× reduction in controller tracking error compared to discrete-action baselines,
with mean and peak errors substantially lower than FM-DiT (e.g., 4.70× lower mean error).
Mechanistic Breakdown of Speed Advantages
The systematic speed advantage is attributed to two independent contributions:
(i) Sparse sampling training: This strategy allows a single forwardly generated polynomial to cover 4× the control duration of discrete action policies,
significantly lowering inference cost.
(ii) Flow matching initiated from history polynomials: This mechanism, combined with sparse sampling, surpasses FLASH-G by 1.7× at 0.32 ms per call, proving the independent contribution of the history-anchored flow mechanism to inference speed.
Improvements for AI systems
As a fastidious researcher, I have analyzed the FLASH (Fast Legendre-polynomial Action policy via Sparse History-anchored flow) paper. The core innovation lies in decoupling inference frequency from control frequency using continuous polynomial representations, history-anchored flow matching, and analytic velocity feedforward.
Based on this research, here are the specific improvements you can make to existing AI systems and what the resulting system can achieve:
The FLASH framework fundamentally improves visuomotor policy learning by addressing the curse of dimensionality
in high-frequency control while maintaining long-horizon capability. The resulting improved AI system will possess the following capabilities:
-
A single, ultra-fast inference step capable of generating a continuous, smooth trajectory over a significantly extended action horizon (up to 4× physical time span compared to discrete policies) in milliseconds (e.g., 31.40 ms per episode).
-
Achieving state-of-the-art success rates across complex manipulation tasks (≥92% across five simulated and two real-world tasks), surpassing current diffusion and flow matching baselines by up to 17.6 percentage points when utilizing history priors instead of noise priors.
-
Substantially reducing controller tracking error (5× to 7× reduction compared to discrete-action baselines) due to the inherent spatial smoothness of Legendre polynomial representations and the direct use of analytic velocity feedforward signals for torque control, eliminating reliance on noisy numerical differentiation.
-
Demonstrating significantly faster training convergence (up to 4× faster than ACT), enabling rapid acquisition of robust manipulation skills with less computational budget.
-
Enabling dynamic execution speed modulation: the system can be executed at any desired playback speed (e.g., 0.5× to 1.33× acceleration) post-hoc by adjusting the evaluation stride, allowing for online acceleration of simple sub-tasks or high-precision deceleration during critical maneuvers without requiring retraining or additional computation.
-
Maintaining high execution repeatability: The system exhibits extremely low inter-rollout standard deviations (e.g., ±0.004 deg for 7 joints), ensuring reliable and precise physical control in real-world deployment environments where discrete policies often fail due to jitter.
Specifically, the improved AI system can perform the following actions:
Feature Specific Capability of Improved System
:---:---
Precision Control Execute high-precision insertion tasks (e.g., millimeter-level tolerance) with superior tracking accuracy (0.274 deg MAE), directly using analytic velocity feedforward to supply the torque controller, ensuring fine motor control is not degraded by numerical approximations.
Efficiency & Latency Achieve real-time robotic control where the inference time is low enough (31.4 ms) to satisfy stringent safety and responsiveness requirements in physical hardware, unlike iterative denoising methods which are too slow for real-time loops.
Robust Skill Acquisition Learn complex, multi-step manipulation skills rapidly by leveraging sparse temporal sampling to fit a compact polynomial representation of the trajectory, leading to faster training convergence than standard action-based or noise-based generative models.
Long-Horizon Planning Plan and execute trajectories that span substantially longer physical durations (up to 4× the horizon) in a single inference, which is critical for tasks requiring complex sequencing or long movements without increasing model complexity.
Adaptive Execution Dynamically adjust the execution speed of a learned trajectory on-the-fly (post-hoc modulation) to adapt to task dynamics—slowing down precisely during high-risk maneuvers while accelerating through free space—without needing new training data.
Sources
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning
- Mean Flows for One-step Generative Modeling
- Action-to-Action Flow Matching
- ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations
- Progressive Distillation for Fast Sampling of Diffusion Models
- Crowd-FM: Learned Optimal Selection of Conditional Flow Matching-generated Trajectories for Crowd Navigation
- Denoising Diffusion Implicit Models
- Score-Based Generative Modeling through Stochastic Differential Equations
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
- BEAST: Efficient Tokenization of B-Splines Encoded Action Sequences for Imitation Learning
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving