ActionSplice: In-Flight Action Editing for Interactive World Models
cs.CV, cs.LG
Submitted: 2026-09-08
Updated: 2026-09-26
Comments: 15 pages, 6 figures. Project page: https://pardistaghavi.github.io/actionsplice-website/
Project page: https://pardistaghavi.github.io/actionsplice-website
License: http://creativecommons.org/licenses/by/4.0/
The gist: Chunk-autoregressive video world models typically condition each generated chunk on one action.
Terminology
Abstract
Chunk-autoregressive video world models typically condition each generated chunk on one action. An action received during sampling must therefore wait for the next chunk, condition future solver evaluations on a state produced under the previous action, or trigger rollback that repeats completed evaluations. We introduce ActionSplice, an inference framework that formulates this problem as Counterfactual State Transport (CST). A lightweight corrector transports the interrupted backbone-native representation toward the matched state induced by the revised action at the same solver step. The world model and sampler remain frozen, and sampling resumes without replaying completed evaluations. The retargeting variant CST* R updates the entire active chunk, while the temporal-splicing variant CST* T preserves a temporal prefix and updates only the suffix. Across minWM-Wan Action2V and HY-WM1.5, CST* R reduces rollback-relative LPIPS by 61.5% and 75.9% relative to direct condition swapping. CST* T reduces suffix LPIPS by 56.1% and 77.5%, respectively, while providing 2.73 times and 1.69 times pixel-ready speedups over waiting. Under the HY-WorldPlay protocol, CST R obtains a PSNR of 25.66 dB, an SSIM of 0.6902, and an LPIPS of 0.1337 against the original rollout.
Sources
- Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation
- Matrix-game 2.0: An open-source, real-time, and streaming interactive world model
- Diffusion ReRoll: Revisable Denoising for Robotic Sequential Prediction
- Beyond Few-Step Inference: Accelerating Video Diffusion Transformer Model Serving with Inter-Request Caching Reuse
- Light Interaction: Training-Free Inference Acceleration for Interactive Video World Models
- LongCat-Video Technical Report
- SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations
- BiWM: Advancing Open-Source Interactive Video World Models with Bidirectional Autoregression
- WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling
- Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models
- Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
- Echo-Forcing: A Scene Memory Framework for Interactive Long Video Generation
- Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation
- Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
- LongLive: Real-time Interactive Long Video Generation
- Anchor Forcing: Anchor Memory and Tri-Region RoPE for Interactive Streaming Video Diffusion
- X-Cache: Cross-Chunk Block Caching for Few-Step Autoregressive World Models Inference
- minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models
- Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation
- C$^3$ache: Accelerating World Action Models with Cross Inference Chunk Cache
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models