FACT: Failure-Aware Causal Training for World-Action Models

arXiv:2608.10232 · cs.RO, cs.AI, cs.LG · Submitted 2026-08-10 · Read on arXiv

Quanquan Peng, Yutong Liang, Rui Yan, Nicklas Hansen, Xiaolong Wang

University of California San Diego

cs.RO, cs.AI, cs.LG

Submitted: 2026-08-10

Updated: 2026-08-12

Project page: https://fact-wam.github.io

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 95/100

The gist: FACT: Failure-Aware Causal Training for World-Action Models introduces a causal World-Action Model (WAM) that predicts future video and task progress conditioned on the executed action.

Terminology

Summary

FACT: Failure-Aware Causal Training for World-Action Models introduces a causal World-Action Model (WAM) that predicts future video and task progress conditioned on the executed action. The paper states: "We introduce FACT, a causal World-Action Model that predicts future video and task progress conditioned on the executed action. This action-conditioned interface allows failure rollouts to supervise action consequences, turning bad actions into valid future targets rather than being discarded."

The method separates action imitation from future prediction using a teacher-forcing mask: We introduce a teacher-forced action-conditioned mask that separates action generation from future prediction, allowing failed actions to supervise future and value learning without undermining policy. The architecture uses a shared causal diffusion transformer with token order P (observation prefix), A (noisy predicted action), G (clean teacher-forced action), V (value), and I (future video). World-side predictions condition on G instead of A, so failure rollouts mask the action imitation loss but keep future-video and value supervision active.

Training uses flow-matching denoising losses. For success demonstrations, all three losses (action, value, future video) are supervised. For failure rollouts, the action loss is masked: failure rollouts mask Lact but keep value and future-video supervision with a lowered progress target—failures teach consequences, not behavior. The value target for failures is clipped: vt(agt t:t+H) = clip(pt+H − λfail 1fail(t + H), 0, 1), with λfail = 1.

Experiments are conducted on RoboTwin simulation (50 tasks) and real-world bimanual manipulation (five seen tasks, three unseen variants). In simulation, video co-training improves FACT from 81.8% to 85.6% average success, and failure co-training further improves it to 87.5%, bringing it close to Motus (87.8%) while running roughly 3× faster. On real-world seen tasks, FACT outperforms Motus (82% vs. 64%); failure-aware training raises this from 82% to 89%, and optional scoring further reaches 92%. On unseen variants, failure-aware training raises success from 67% to 77% and optional scoring to 82%.

Ablations show: removing video co-training reduces real-world seen-task success from 82% to 58%; removing the causal mask drops success from 82% to 77%; adding failure rollouts without masking the action loss drops success to 63%. Failure data reduces future hallucination: under the same bad-action condition, the success-only model still predicts a successful grasp, while failure-aware co-training predicts the observed failed outcome. Quantitatively, failure-aware co-training improves PSNR on failure-rollout futures from 19.51 to 25.92 while leaving success-rollout futures nearly unchanged (26.12 to 26.08). Failure-data scaling shows monotonic improvement from 32.7% to 57.3% as failure-rollout fraction increases from 0% to 100%. Value-guided candidate scoring with N=4 candidates improves task completion, and value traces reflect action outcomes (dropping at missed grasps and recovering after re-grasping).

The paper concludes: "We presented FACT, a causal World-Action Model that reverses the usual WAM order by generating actions before predicting future video and task-progress value. A teacher-forcing mask makes the clean executed action the condition for all world-side predictions, allowing failure rollouts to supervise future and value prediction while their action imitation loss is disabled."

Improvements for AI systems

Improvements to AI Systems:

  1. Action-Conditioned Future Prediction with Failure Supervision: Implement a causal World-Action Model where the executed action (clean, teacher-forced) conditions all future video and value predictions, while the predicted action is generated separately. This allows the system to learn from failed rollouts by masking only the action-imitation loss, not the future-video or value losses. The improved system can predict accurate outcomes of both successful and failed actions, reducing hallucination of success in failure scenarios (e.g., PSNR improves from 19.51 to 25.92 on failure futures).

  2. Teacher-Forced Action Mask for Robust Policy Learning: Use a token-order mask (P, A, G, V, I) where world-side predictions condition on the clean action G, not the noisy predicted action A. This separates action generation from future prediction, preventing bad actions from corrupting policy learning while still leveraging their consequences. The improved system maintains policy stability (success drops only from 82% to 77% without the mask, vs. 63% when failure actions are imitated directly).

  3. Failure-Aware Value Target Clipping: Set value targets for failure rollouts as clip(p t+H − λ fail * 1 fail, 0, 1) with λ fail=1, lowering the progress target for failures. This teaches the system that failed actions lead to reduced task progress, enabling value traces that accurately reflect action outcomes (e.g., value drops at missed grasps, recovers after re-grasping). The improved system can guide action selection by scoring candidate actions based on predicted value, improving task completion (e.g., from 89% to 92% with N=4 candidate scoring on real-world tasks).

  4. Video Co-Training for Generalization: Jointly train future-video prediction alongside action and value losses, even without failure data. This improves success rates significantly (e.g., real-world seen tasks from 58% to 82% with video co-training) by providing richer visual supervision that enhances world understanding and action consequence modeling.

  5. Failure-Data Scaling for Robustness: Increase the fraction of failure rollouts in training data (from 0% to 100%) to monotonically improve success (from 32.7% to 57.3% in simulation). The improved system becomes increasingly robust to unseen or novel task variants by learning from diverse failure modes, raising real-world unseen-task success from 67% to 77% (and 82% with scoring).

  6. Efficient Causal Diffusion Transformer for Real-Time Control: Use a shared causal diffusion transformer with flow-matching denoising, enabling faster inference (≈3× faster than Motus) while achieving comparable or better performance. The improved system can run in real-time for robotic manipulation, predicting future video and task progress conditioned on actions, suitable for closed-loop control in dynamic environments.

What the Improved AI System Can Do:

  • Predict future video frames and task-progress values conditioned on executed actions, accurately reflecting both successes and failures.

  • Learn from failed demonstrations without imitating bad actions, improving robustness and generalization to unseen tasks.

  • Use value-guided candidate scoring (e.g., N=4) to select optimal actions, enhancing task success rates (up to 92% on real-world seen tasks, 82% on unseen variants).

  • Operate in real-time for bimanual manipulation, with reduced hallucination of success in failure conditions and stable policy learning even with noisy action predictions.

Abstract

Recent world-action models (WAMs) show that co-training policies with future prediction can provide physical priors for action generation. Building on the future-prediction ability of video models, many WAMs generate future videos and recover actions with inverse-dynamics models, or use these predicted videos as goal conditions for action generation. In both cases, the world model is trained mostly on successful demonstrations and has little reason to predict the consequences of bad actions. We introduce FACT, a causal World-Action Model that predicts future video and task progress conditioned on the executed action. This action-conditioned interface allows failure rollouts to supervise action consequences, turning bad actions into valid future targets rather than being discarded. Failure-aware training makes the progress predictor aware of both successful and failed action outcomes, which can optionally be used to score sampled action candidates at inference. Extensive experiments on simulation and real-world bimanual manipulation tasks show that FACT outperforms many existing baselines, improves as failure data are incorporated into training, and reduces success-biased future hallucination under bad actions. See more details at https://fact-wam.github.io/

Sources

Related papers