ContactGuard: Pre-Contact Execution Monitoring with Action-Conditioned Latent World Models

arXiv:2608.13438 · cs.RO, cs.AI, cs.CV · Submitted 2026-08-13 · Read on arXiv

Gehan Zheng, Matthew Johnson-Roberson, Weiming Zhi

Vanderbilt University · The University of Sydney · Australian Centre for Robotics, The University of Sydney

cs.RO, cs.AI, cs.CV

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: 14 pages, 5 figures, 8 tables

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 100/100

The gist: ContactGuard is a pre-contact execution monitoring system for chunked visuomotor policies in contact-rich robot manipulation tasks.

Terminology

Summary

ContactGuard is a pre-contact execution monitoring system for chunked visuomotor policies in contact-rich robot manipulation tasks. The paper addresses the problem that contact-rich manipulation failures are often detected only after the robot has committed to contact, which is especially limiting in wrist-camera setups where a poor approach may already push, miss, slip, or disturb the object before conventional detectors react.

ContactGuard is built on the Joint-Embedding Predictive Architecture (JEPA) principle, where a predictor is trained to forecast the embedding of a target signal rather than reconstruct it in pixel space. The system consists of:

  1. A latent world model trained from unlabelled robot trajectories using next-latent supervision. The world model uses a shared ViT-Tiny encoder that processes multi-view camera observations independently, mean-pools the per-view embeddings, and passes them through a learned linear projection to produce a single latent. An action-conditioned causal predictor (four AdaLN-zero-conditioned Transformer blocks followed by a two-layer prediction MLP) predicts future latents under planned actions.

  2. A lightweight failure probe (logistic regression) trained from a small labelled set of pre-contact clips. The probe maps the predicted post-contact latent to failure likelihood.

The training objective is next-latent regression with SIGReg regularisation: L = (1/L) Σ ẑi+C − zi+C2 + λL reg.

At deployment, ContactGuard runs alongside an existing chunked visuomotor policy (such as ACT) as a policy-decoupled predictive verifier. The deployed policy is treated as a black-box proposer. The monitor:

  • Scans the upcoming action chunk for an imminent contact event (in this paper, gripper closure detected via open-to-close transition)

  • Anchors prediction k pre = 15 frames before planned gripper closure (0.5 seconds at 30 Hz)

  • Rolls the frozen latent world model forward for K = 30 steps under the planned actions

  • Applies the frozen probe to the predicted post-contact latent ẑt+K (reaching 0.5 seconds after closure)

  • Aborts execution if P(fail) > τ, where τ is selected per-task on validation split

Experiments use a 14-DoF AgileX Piper dual-arm robot with three synchronized RGB cameras across four grasp settings: cup and box grasps in pick-and-place, pencil grasping in pencil-and-notebook, and towel grasping in towel-fold. The world model is trained on unlabelled real-robot trajectories from ACT rollouts and human teleoperation. Each labelled set contains roughly 250 grasp attempts.

Prediction quality (Table 1): ContactGuard outperforms both Direct-linear (which predicts failure directly from anchor latent and planned action) and LeWM (single-view world model) baselines. For example, on the Cup task, ContactGuard achieves 0.992 AUC and 0.940 balanced accuracy versus 0.661/0.580 for Direct-linear and 0.928/0.780 for LeWM.

External detector comparison (Table 2): ContactGuard achieves the highest AUC on all four tasks compared to FAIL-Detect, RND, and SAFE. On larger offline pools, ContactGuard achieves AUC of 0.982±.003 (Cup), 0.984±.005 (Box), 0.992±.001 (Pencil), and 0.978±.004 (Towel).

Information source analysis (Table 3): The paper demonstrates the signal comes from the imagined consequence of the specific proposed action. Replacing ẑt+K with current latent zt reduces AUC. Corrupting actions—shuffling action chunks across episodes collapses AUC to near chance (0.535, 0.493, 0.483, 0.493 on Cup/Box/Pencil/Towel), while zero and mean actions also remain below correctly aligned rollout.

Counterfactual action-swap test: Holding observation fixed while replacing only the planned action chunk with one from a failed attempt increases predicted failure probability by +0.25/+0.33/+0.56/+0.65 on Cup/Box/Pencil/Towel respectively, confirming the monitor responds to the pending action rather than static visual risk.

Proprioceptive ablation (Table 5): Counter to intuition, dropping robot state improves test AUC on all four tasks (e.g., Cup: 0.660 with state vs 0.920 image-only), suggesting naive state fusion acts as a domain-specific shortcut in the small-data regime.

Runtime (Table 4): At the deployed horizon (K=30, 1.00 second), full encode-rollout-probe pass takes 19.18 ms on RTX 5090, fitting comfortably within the pre-closure slack.

The paper's contributions are: (1) formulating pre-contact execution monitoring for chunked visuomotor policies, (2) introducing ContactGuard as an action-conditioned latent world-model monitor, (3) showing predicted future latents provide failure information beyond current observations and raw planned actions, and (4) demonstrating live real-robot pre-contact aborts without candidate action search, pixel-level video prediction, or modification of the underlying visuomotor policy.

ContactGuard prevents failures by abstaining but does not recover from them or complete the task after an abort. It targets imminent contact events whose outcome is determined within the next action chunk; extending to longer-horizon skills would require hierarchical or repeated event-level monitoring.

Improvements for AI systems

Improvements to AI systems based on ContactGuard:

  1. Action-conditioned pre-contact failure prediction for robotic manipulation

The improved AI system can predict whether a planned action chunk will cause a failure before physical contact occurs, by rolling a latent world model forward under the proposed actions and classifying the imagined post-contact latent. This enables pre-emptive aborting of unsafe grasps, preventing object displacement, slipping, or missed picks that conventional post-contact detectors miss.

  1. Policy-decoupled safety monitoring for black-box visuomotor policies

The improved system can act as an external verifier that wraps any existing chunked visuomotor policy (e.g., ACT, diffusion policies) without modifying its weights or requiring access to internal representations. It treats the policy as a black-box proposer, scans upcoming action chunks for imminent contact events (e.g., gripper closure), and aborts execution when predicted failure likelihood exceeds a task-specific threshold—enabling safe deployment of pre-trained policies in new contact-rich settings.

  1. Latent-space world modeling for action consequence simulation

The improved system can forecast high-dimensional future observations in a compact latent space (via JEPA-style next-latent prediction with SIGReg regularization), avoiding expensive pixel-space video prediction. This allows real-time (19 ms) rollouts of 30 future steps under planned actions, making it feasible for closed-loop monitoring on embedded or wrist-camera robot platforms.

  1. Data-efficient failure detection with minimal labels

The improved system can learn failure prediction from only 250 labeled grasp attempts per task, by combining an unlabeled-pretrained latent world model (trained on raw trajectories) with a lightweight logistic-regression probe. This reduces annotation cost and enables rapid adaptation to new manipulation tasks or object geometries.

  1. Counterfactual action sensitivity for risk assessment

The improved system can distinguish between static environmental risk and action-induced risk by holding observations fixed and swapping only the planned action chunk. This enables the AI to answer would this specific action cause failure? rather than is this scene risky?, supporting more precise decision-making in dynamic or cluttered environments.

  1. Proprioceptive-free perception for robust monitoring

The improved system can achieve higher failure-prediction accuracy by excluding robot joint-state inputs when training the world model, as naive state fusion introduces domain-specific shortcuts in small-data regimes. This leads to more generalizable monitoring across different robot configurations or after hardware changes.

  1. Real-time pre-contact abort with no action search

The improved system can abort unsafe actions within the pre-contact slack (0.5 seconds before gripper closure) without performing candidate action search, pixel-level video prediction, or replanning. This enables safe execution of high-speed manipulation skills where reaction time is critical.

  1. Transferable event-level monitoring for hierarchical skills

The improved system can be extended to longer-horizon tasks by decomposing them into repeated event-level monitors (e.g., each grasp, each insertion), enabling staged safety checks throughout a multi-step manipulation pipeline rather than only at the final outcome.

Abstract

Contact-rich manipulation failures are often detected only after the robot has committed to contact. This is especially limiting in wrist-camera setups: close gripper--object views help observe contact, but a poor approach may already push, miss, slip, or disturb the object before conventional detectors react. We introduce ContactGuard, a pre-contact execution monitor for chunked visuomotor policies. Given the policy's planned action chunk, ContactGuard predicts its short-horizon consequence in latent visual space and aborts if the predicted future latent indicates likely failure. Its latent world model is trained from unlabelled robot trajectories to predict compact multi-view visual embeddings under planned actions, avoiding pixel-level video prediction. A lightweight failure probe is then trained from a small labelled set of pre-contact clips. At deployment, ContactGuard anchors prediction before an imminent contact event, rolls the model forward under the policy's own actions, and verifies the predicted post-contact latent. Across real-world contact-rich manipulation tasks, ContactGuard predicts failure more accurately than direct and corrupted-action ablations, and transfers to live robot as a pre-contact abort signal without modifying the underlying policy.

Sources

Related papers