SplineWAM: Adaptive Action Horizons for World Action Models via B-Spline Representations

arXiv:2609.39873 · cs.RO · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "SplineWAM: Adaptive Action Horizons for World Action Models via B-Spline Representations".

Dev: World action models (WAMs) are large embodied policies that jointly predict future video and actions,

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: So we're looking at this paper today called "SplineWAM: Adaptive Action Horizons for World Action Models via B-Spline Representations," and what it claims is that they've replaced those fixed action chunks with something much more flexible.

Dev: I agree, Rosa, the core idea seems to be adapting the representation itself so that the temporal resolution of the prediction actually matches how complex the motion is happening in real time.

Taro: From an autonomy standpoint, this sounds promising because it directly addresses a fundamental limitation where uniform chunking fails to capture both long free-space movements and fast contact manipulations.

Rosa: Exactly, Taro, they introduce this adaptive action representation using cubic B-splines that use knot times and control points to define the trajectory.

Dev: That means instead of a rigid grid of actions, you get a set of parameters—the knots and control points—which are fitted adaptively so the prediction is dense where motion is hard to approximate.

Taro: I see how that tackles the issue we discussed earlier; it lets one parameter budget decode chunks that have different temporal resolutions and durations depending on what the robot is actually doing.

Rosa: Precisely, and they use these knot times explicitly to align video supervision, favoring a knot-grid alignment strategy to share the nonuniform temporal structure of the action data.

Dev: That adaptive fitting is clever because it means the execution span and even how long we wait between policy calls are determined by the prediction itself, which is a significant departure from fixed chunking.

Taro: That leads me to think about what happens when things go wrong in the real world; does this system have a way to handle unexpected misbehavior during that adaptive decoding?

Rosa: They address that with something called JP-RTC, which stands for Jacobian-Pullback Real-Time Chunking, which imposes continuity on the decoded actions rather than just relying on the spline parameters.

Dev: The mechanism behind JP-RTC involves correcting those parameters through the decoder’s local Jacobian by solving a regularized least-squares problem where the update minimizes a specific cost function.

Paper summary: Taro: That sounds like they are actively trying to maintain coherence with what has already been executed, which is crucial for real-time deployment on robots.

Rosa: The results they show on simulated suites LIBERO-Plus and RoboCasa are quite compelling, showing improvements in success rates and reductions in the number of policy calls per episode.

Dev: I noticed they report specific figures for those results; for instance, success rate gains of eight point two and four point four points over a fixed action chunking WAM on the two suites they tested.

Taro: And those call reductions, which are reported as twenty-two percent and twenty-six percent fewer policy calls per episode on LIBERO-Plus and RoboCasa, suggest a tangible efficiency gain in deployment scenarios.

Rosa: It really shows that by making the representation adaptive, you get better performance metrics without necessarily increasing the raw computational load in a fixed way.

Dev: But we have to consider how this translates to actual robot hardware; I'm wondering about the loop rate and latency when implementing this kind of complex spline decoding on an embodied model.

Taro: That’s a fair concern, Dev, because while the representation is adaptive, the underlying neural network still has to process that complex spline structure within its inference cycle.

Rosa: The paper notes that for real-robot tasks under asynchronous execution, SplineWAM can decode one point two to one point six times as much executed motion per call on the physical robot compared to the baseline chunking method.

Dev: That factor of one point two to one point six is significant because it means we are getting more useful motion information out of each inference cycle, which directly impacts latency management and throughput in a deployed system.

Taro: This really speaks to the idea that for tasks involving both long free-space motions and sudden contact, this method provides a better temporal resolution profile than uniform sampling allows.

Rosa: The paper points out that this representation is best suited for tasks with that mix of motion, where the continuous smooth parameterization helps stretch the horizon only where it actually needs to be stretched.

Paper summary: Dev: However, they also laid out some limitations; they mentioned that precision at contact tasks can be limited because a cubic spline trades exactness for compression wherever the tolerance allows it.

Taro: That limitation is important because when you're doing fine manipulation, losing accuracy right when you need it most could lead to failures during those critical moments.

Rosa: And they also noted that as the action space gets larger, the achievable compression falls, and they are using a single knot vector for all action dimensions which limits compression further.

Dev: So while it's efficient for certain scenarios, we can't expect perfect accuracy across every single dimension if the dimensionality of the task increases substantially.

Taro: That points toward future work where maybe multi-dimensional or hierarchical spline representations could be explored to maintain high fidelity in very complex action spaces.

Rosa: Exactly, and thinking about how this impacts real-world robotics, the biggest implication is that we can deploy policies on robots that are much more efficient in terms of computation while still handling varied motion complexity effectively.

Dev: For me, the practical implication is making those asynchronous deployments feasible because JP-RTC makes the spline representation usable for real robots in a way that was challenging with naive chunking.

Taro: I think the broader impact on autonomy is that it allows for more robust and less brittle action execution policies when facing unpredictable environmental interactions.

Rosa: So, to wrap up this discussion on SplineWAM: the core contribution is using adaptively fitted cubic B-splines to create action representations where the temporal resolution scales with motion complexity, leading to fewer policy calls and higher success rates in simulations.

Dev: And for deployment, JP-RTC helps bridge the gap between that adaptive representation and real-time execution by ensuring parameter continuity during decoding.

Taro: It suggests a way forward for embodied AI policies to be more efficient in terms of computational budget while still being robust to the dynamic nature of physical tasks.

Conclusion: Rosa: So we've been digging into this work called "SplineWAM: Adaptive Action Horizons for World Action Models via B-Spline Representations," and what we really need to focus on now is what this whole concept actually means for our work.

Dev: I agree, Rosa; the core idea of replacing fixed action chunks with these adaptive cubic B-splines is a significant structural change that needs careful consideration regarding performance and stability.

Taro: From an autonomy standpoint, I'm interested in how this flexible representation handles situations where the environment throws something unexpected at it, like sudden contact or a long period of free movement.

Rosa: Exactly, Taro; the authors are showing how they can tailor the temporal resolution of their predictions to match the motion complexity on demand rather than using a single uniform grid.

Dev: That variability is interesting for my loop rate concerns; if we’re decoding chunks with wildly different resolutions, it complicates things immensely for maintaining a consistent execution timeline and managing latency.

Taro: But that adaptability is what makes it potentially powerful; if the system can compress the long free-space movements while keeping high detail during contact, that could lead to much more efficient policy calls overall.

Rosa: That's the big picture, Taro; it suggests a new way to structure how a world action model processes motion—one that respects the physical reality of what the robot is doing moment by moment.

Dev: I'm still thinking about deployment outside of controlled simulations; does this adaptive fitting mechanism have any known failure modes when dealing with real-world sensor noise or unexpected dynamics?

Taro: The authors do mention some limitations, like how precision drops at contact tasks if the spline has to compress too much, which brings up my point about handling misbehavior; it seems there are trade-offs between compression and exactness.

Rosa: So, the main point is that SplineWAM offers a representation where temporal resolution scales with motion complexity, which could lead to better efficiency in real applications.

Dev: It’s definitely an interesting paper because it tackles the fundamental problem of uniform chunking being inefficient for dynamic tasks while introducing a method to maintain continuity through JP-RTC.

Taro: We should keep watching how they address those compression limits when dealing with higher-dimensional action spaces, as that seems like a current bottleneck for this approach.

Jun Guo, Xiaoshen Han, Qiwei Li, Nan Sun, Peiyan Li, Heyun Wang, Hang Lai

Tsinghua University · Shanghai Jiao Tong University · Xiaomi Robotics

cs.RO

Submitted: 2026-09-30

Updated: 2026-09-30

Comments: Project website: https://splinewam.github.io/

Project page: https://splinewam.github.io

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 89/100

The gist: World action models (WAMs) are large embodied policies that jointly predict future video and actions, and SplineWAM introduces an adaptive action representation using cubic B-splines to allow these

Key concepts

Action Chunking
Traditional methods divide actions into uniform time segments. This is inefficient because free movement needs long segments, while contact requires very short ones, leading to wasted computation during slow motion.
Cubic B-spline Representation
Instead of fixed points, the action is represented by a set of knot times and control points. These parameters are fitted adaptively so that the prediction's temporal resolution matches how complex the actual movement is, allowing for smoother, more efficient modeling.
Knot-grid Alignment
This strategy aligns video supervision with the spline representation instead of a uniform grid. It ensures that video data is supervised most densely during complex action phases where the motion structure changes rapidly.

Terminology

Summary

World action models (WAMs) are large embodied policies that jointly predict future video and actions, and SplineWAM introduces an adaptive action representation using cubic B-splines to allow these models to handle varying motion complexities efficiently. The gist: SplineWAM replaces the fixed-length action chunk of a WAM with a fixed-size window of cubic B-spline knot times and control points, fitted adaptively so the temporal resolution follows the complexity of the motion, allowing one parameter budget to decode chunks of varying temporal resolution and duration.

The Problem Addressed

Current World Action Models (WAMs) typically rely on action chunking, which samples actions on a uniform temporal grid. This fixed representation is suboptimal because manipulation is not temporally uniform; free-space motion requires a long horizon, while contact-rich manipulation demands fine temporal resolution that a fixed grid cannot provide. A fixed action chunk spends more inference time during free-space motion and cannot replan promptly enough once contact begins, leading to redundant calls over free space and poor throughput when served in the cloud.

The SplineWAM Representation

SplineWAM replaces the raw action chunk with a fixed-size window of cubic B-spline parameters. The representation is defined by a set of knot times and control points, denoted as the vector z = [U, C], where U is the knot vector and C collects the control points. The policy predicts a fixed number of such pairs. The crucial innovation lies in Adaptive fitting, where knots are placed densely where the trajectory is hard to approximate, meaning the temporal resolution of the prediction follows the complexity of the motion. This allows one parameter budget to decode chunks of varying temporal resolution and duration, enabling a variable execution span and interval until the next policy call.

Video-Action Temporal Alignment

The representation carries explicit knot times, which are used to align video supervision. Instead of a uniform grid, SplineWAM uses two alignment strategies: Time-window alignment or Knot-grid alignment. Knot-grid alignment is preferred because it allows the video supervision to share the nonuniform temporal structure of the action supervision, meaning frames are supervised most densely where the action trajectory is complex. During training, video targets are sampled at these fitted knot times rather than on a uniform grid.

Asynchronous Execution with JP-RTC

For asynchronous deployment on real robots, SplineWAM introduces Jacobian-Pullback Real-Time Chunking (JP-RTC). This method imposes chunk continuity on the decoded actions rather than the spline parameters alone. JP-RTC corrects the parameters through the decoder so that the executed prefix agrees with the actions already committed. The update is derived by solving a regularized least-squares problem: δ⋆ minimises∥JSδ−eS∥22 + λ∥δ∥22, where JS is the decoder's Jacobian, and λ selects a correction within an underdetermined system.

Key Results and Deployment

On simulated suites LIBERO-Plus and RoboCasa, SplineWAM improves success rate by 8.2 and 4.4 points over an action chunking WAM while cutting policy calls per episode by 22% and 26%. On three bimanual real-robot tasks under asynchronous execution, it leads or matches the baseline while decoding 1.2 to 1.6 times as much executed motion per call. The method is best suited for tasks with a mix of free-space and contact-rich motion, where the continuous smooth parameterization provides efficiency gains by stretching the horizon only where motion is compressible. JP-RTC makes the spline representation usable asynchronously, recovering performance lost when naive chunking is applied in parameter space.

Limitations

The paper notes that precision at contact tasks can be limited because a cubic spline trades exactness for compression wherever the tolerance allows it, and accuracy collapses as soon as knots resolving contact are removed. Furthermore, JP-RTC still leaves a gap to synchronous execution on both simulated suites, which is larger on LIBERO-Plus than on RoboCasa. The fitting tolerance is a single scalar chosen per benchmark and held fixed across all tasks. Additionally, the compression achieved is set by the hardest dimension to approximate; as the action space grows, achievable compression falls. The representation uses a single knot vector for all action dimensions, which limits compression achievable when dealing with higher-dimensional actions.

Implementation Details

The architecture involves training a video expert and an action expert jointly under a shared attention mask. The training objective is defined on the spline parameters: The action loss is defined on spline parameters, not on decoded actions. The fixed-size local windows are constructed using boundary support to ensure the curve remains evaluable across the interval. A row stores one absolute knot time and one control point, [t, c], which are mapped to a common range before reaching the model. The constrained prefix for JP-RTC covers S = 8 control steps.

Improvements for AI systems

Here are specific improvements to AI systems derived from the SplineWAM research, detailing what these improved systems can achieve:


  1. The core improvement is replacing fixed-length action chunks with a flexible, motion-dependent representation using a cubic B-spline window of knot times and control points.

  2. This allows the model to predict an action trajectory that dynamically adjusts its temporal resolution based on the complexity of the required motion (e.g., predicting fine, high-resolution actions during grasping/insertion and long, low-resolution actions during free-space travel).

  3. The system can achieve significantly higher performance in complex manipulation tasks compared to traditional action chunking models:

  4. By aligning video supervision directly to the fitted knot times (rather than a uniform grid), the model concentrates training effort on difficult, high-complexity motion segments, leading to an improvement in success rate (up to 8.2 points on LIBERO-Plus) and better generalization across different visual conditions.

  5. The system can operate efficiently in resource-constrained or asynchronous deployment environments:

  6. By introducing Jacobian-Pullback Real-Time Chunking (JP-RTC), the model can execute actions asynchronously while maintaining continuity with previously committed actions, reducing the number of policy calls per episode by 22% to 26%. This directly increases throughput when serving models in a shared cloud infrastructure.

  7. The system can achieve higher action fidelity and efficiency during real-world execution:

  8. On real-robot tasks, SplineWAM can decode 1.2 to 1.6 times as much executed motion per policy call compared to fixed-chunk baselines, meaning the robot executes a longer span of relevant motion before needing a new decision, leading to better task progress (e.g., exceeding baseline on four out of six accuracy measures in real-world tests).

  9. The system exhibits superior adaptability during execution:

  10. Because the knot times define both the executed span and the interval until the next call, it inherently manages trade-offs: it replans promptly when contact begins (shortening chunks) while extending its horizon during smooth free-space motion (lengthening chunks).

  11. The system provides a more robust inference mechanism across various task types:

  12. It maintains high accuracy on tasks requiring both long-horizon planning (like mobile manipulation) and high-precision, contact-rich manipulation (like inserting a plug into a socket) by dynamically tuning the compression based on the local trajectory characteristics.

  13. The system's deployment is more robust to environmental perturbations:

  14. By using adaptive fitting based on an error tolerance, the model learns representations that remain effective even when facing variations in background textures, camera viewpoints, or lighting conditions (as shown by superior performance margins in the ablation studies).

Sources

Related papers