Compact Visuotactile World Models for Lifting: Prediction, Reward Alignment, and Force Constraints

arXiv:2609.09597 · cs.RO, cs.AI · Submitted 2026-09-09 · Read on arXiv

cs.RO, cs.AI

Submitted: 2026-09-09

Updated: 2026-09-10

Comments: 8 pages, 2 figures. Code and tabulated results included as ancillary material

License: http://creativecommons.org/publicdomain/zero/1.0/

The gist: Accurate tactile forecasts need not improve force-constrained control.

Terminology

Abstract

Accurate tactile forecasts need not improve force-constrained control. We study a 652,157-parameter action-conditioned visuotactile world model with matched behavior cloning, policy learning in imagination, independent reactive implicit Q-learning, and model-assisted force feedback. A fixed protocol executes 34 policies on 120 fresh MuJoCo environments spanning geometry and physical-parameter shifts, plus 324 independently replayed action branches on 12 additional ID environments. Visuotactile dynamics reduce force action-effect MAE from 0.413 N for persistence to 0.338 N. Model-assisted feedback raises ID force-budgeted success from 73.3% to 93.3%, with paired difference +20.0 [+6.7,+33.4] percentage points (95% CI), with the difference occurring during scripted lowering. Its pooled difference is +3.9 [-4.5,+11.7] points. Imagined RL achieves 11.9% pooled joint success versus 25.0% for reactive IQL. An empirical tactile-residual stress test adds 330 executions. The evidence concerns rigid-box lifting after a common approach, without physical-robot transfer or a closed-loop safety guarantee.

Related papers