3D Point World Models: Point Completion Enables More Accurate Dynamics Learning

arXiv:2607.00148 · cs.RO, cs.CV · Submitted 2026-06-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "3D Point World Models".

Dev: The gist:

Rosa: First, who's behind it and why it matters.

Paper summary: Dev: So, wrapping up the discussion on "three dee Point World Models: Point Completion Enables More Accurate Dynamics Learning," the authors are really showing that learning predictive models of the world can enable robotic control through planning, and they achieve this by building a task-agnostic world model operating in three dee space <ref:2607.00148#pg1,3D Point World Models: Point Completion Enables More Accurate Dynamics Learning>.

Rosa: The implication here is that by explicitly completing those point clouds first, you get a much better foundation for dynamics learning, which leads to more accurate geometric reasoning and planning performance across different robot embodiments.

Taro: It suggests that this method is robust enough to handle complex behaviors like pick-and-place and can adapt to novel combinations of tasks, even though they still acknowledge the limitations around SAM3 failures and data distribution mismatches.

Dev: Ultimately, the work on three deePWM demonstrates that world models based on explicit point cloud completion lead to improved rollout quality, better predicted geometry, and higher downstream planning performance across simulated and real tasks <ref:2607.00148#pg1>.

Rosa: So, for anyone listening who’s thinking about how robots can improvise solutions on new tasks without needing task-specific fine-tuning every time, this paper suggests a path forward by integrating perception components more tightly with the dynamics model.

Conclusion: Rosa: So, we’re looking at this paper, "three dee Point World Models: Point Completion Enables More Accurate Dynamics Learning," and what it really boils down to is they built a whole world model that works purely in three dimensions by first making sure their point clouds are actually complete before they try to learn how things move.

Dev: That’s right. They took something partial—a robot's view of the scene—and they used some clever steps, like segmenting objects and then filling in the missing geometry, to get a full three dee picture before feeding it into their dynamics engine.

Taro: What I find interesting is that this lets them do long-horizon rollouts reliably. Usually, when you have partial data or geometry errors accumulating over time, the predictions drift pretty fast and you lose track of where things actually are in the world.

Rosa: Exactly. The paper shows that by using this completed three dee scene for learning, they can get much more accurate predictions over long sequences of actions than what most other models can manage when dealing with incomplete input.

Dev: And from an engineering standpoint, it’s important because they’re training the model to predict per-point velocities based on those complete snapshots, which keeps the physics consistent throughout the simulation.

Taro: It also suggests that this isn't just about getting better rolls in a lab setting; they show it can handle more complex stuff like pick-and-place and even adapt when they’re thrown a completely new task combination.

Rosa: So, in the end, this work suggests that if you want robots to actually plan for long distances or handle tricky real-world situations, focusing on getting a solid three dee representation first is a pretty crucial step.

Dev: It definitely points toward needing tighter integration between perception—getting those point clouds right—and the dynamics learning itself.

Taro: And that leads us to the question of how much latency you can tolerate before this whole process starts breaking down in a real-time system.

Oregon State University

cs.RO, cs.CV

Submitted: 2026-06-30

Updated: 2026-10-07

Comments: 21 Pages

Code: https://github.com/pvskand/3dpwm

Project page: https://3dpwm.github.io

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 86/100

The gist: The gist: 3D Point World Models (3DPWM) are a task-agnostic world model that operates entirely in 3D space by first completing partial point clouds and then learning action-conditioned dynamics in

Key concepts

Point Cloud Completion
This process takes incomplete 3D data, like a partial point cloud from a camera, and reconstructs it into a full 3D scene. The system uses segmentation and specialized completion techniques to fill in missing geometry, ensuring the model learns dynamics on complete scenes.
Action-Conditioned Dynamics
This refers to learning how the robot's state changes based on an action taken. The model predicts the next observation by considering both the current scene geometry and the specific movement (action) applied by the robot, allowing for realistic simulation of robot behavior.
Point Transformer V3 (PTV3)
This is a deep learning architecture used to encode 3D points into a compact, meaningful latent representation. By using PTV3, the model can efficiently capture the spatial information of every point in the scene, which is then used for predicting future states.
Model Predictive Control (MPC)
This is a planning framework that uses the learned world model to predict future outcomes based on a sequence of actions. It allows the system to plan optimal control sequences by minimizing a cost function defined over the predicted point cloud states.

Terminology

Summary

The gist: 3D Point World Models (3DPWM) are a task-agnostic world model that operates entirely in 3D space by first completing partial point clouds and then learning action-conditioned dynamics in this completed 3D scene, enabling reliable long-horizon rollouts and more accurate cost evaluation for model-based planning

Problem Addressed

Large video-based dynamics models lack explicit 3D spatial structure and suffer from geometrically inconsistent long-term rollouts with compounding errors Emerging 3D dynamics models based on partial point clouds improve geometric consistency but remain sensitive to occlusions and accumulated prediction drift Prior work in world model-based robot control has predominantly learned task-specific world models from 2D images, which typically train both transition and reward functions that must be fine-tuned when adapting to new tasks Models trained on complete point clouds in simulation struggle in real-world deployment where complete geometry is unavailable

Proposed Solution: 3DPWM Architecture

The proposed system, 3DPWM, is a task-agnostic world model that operates entirely in 3D space by first completing partial point clouds and then learning action-conditioned dynamics in this completed 3D scene The system takes in robot proprioception (with joint angles q) and a single-view partial point cloud as the observation O at every step Given this partial observation, it first completes the point cloud to obtain the complete observation of the visible scene Ocomplete Then, given an action a ∈ A, it trains a transition function T: Ocomplete × A → Ocomplete that predicts the next observation

Data Processing and Dynamics Modeling

The process of obtaining complete point clouds involves several steps:

  1. Computing the robot’s mesh configuration and sampling points from it to obtain the point cloud of the robot Probot t = FK(qt, URDF)

  2. Segmenting each visible object in the scene using SAM3 [44] with manual prompts and backprojecting the depth map to recover partial object point clouds

  3. Applying point cloud completion [28] to produce completed object point clouds, Pobj t

The final observation is a combination of the robot end-effector points and complete object point clouds; Pt = Probot t ∪ Sobj Pobj t

For the 3D Point World Model, each point is featurized with its current spatial location as well as both current and previous timestep velocities – Ft = [xt, yt, zt, vt, vt−1] This observation representation (Pt, Ft) is encoded using Point Transformer V3 (PTV3) [46] to obtain a point-based latent representation zt ∈ RM×d Action conditioning is induced by encoding the action at with a small MLP to get zact t = MLP(at) and performing cross-attention where action-relevant points can attend to action commands The next timestep point cloud can be obtained by Pˆt+1 = Pt + vˆt+1 (Eq. 1)

Training Objective and Planning

The training objective is to train the dynamics model using a Huber loss [48] between the predicted and ground-truth per-point velocities This is done by training based on multi-step rollouts – iteratively applying our model to produce the following H time steps induced by an action sequence at, at+H Planning with 3DPWM employs a model predictive control framework (MPC) using a sampling based MPPI planner [49] The cost function is defined in the observation space of the complete point cloud Pt, similar to how a reward function for a downstream task is written with a simulator

Key Contributions and Results

The main contributions include:

We propose 3DPWM, a system comprising point cloud completion and a learned 3D world model

We show that 3DPWM is capable of accurate long-horizon rollouts in the point cloud space across different robot embodiments that facilitates better closed and open-loop planning than baselines

We also demonstrate 3DPWM ability to adapt to novel unseen combination of tasks

Finally, we show effective sim-to-real transfer of the system on tabletop manipulation tasks where 3DPWM performs 2.5× better than baselines

Experiments across different robotic embodiments and tabletop manipulation benchmarks demonstrate that 3DPWM achieves significantly more reliable long-horizon rollouts (100-300+ steps), supports both open-loop and closed-loop planning, and enables successful sim-to-real transfer Across all tasks, 3DPWM achieves significantly lower Chamfer distance on actionconditioned rollouts (average rollout lengths reported in Tab. 1)

Robustness and Adaptation

In contrast, our approach supports more complex manipulation behaviors, including pick-and-place

We also demonstrate 3DPWM ability to adapt to novel unseen combination of tasks

The system shows adaptation to unseen long-horizon tasks: on short-horizon tasks, 3DPWM achieves 75% on MugCleanup and 60% on Coffee, with the latter’s lower performance likely due to precise placement requirements The system also shows robustness to camera viewpoints and lighting conditions, maintaining higher performance when compared with ParticleFormer+

Limitations

Failures in any stage can propagate through the pipeline, leading to cascading errors Common sources of failure include SAM3 failures where object segmentation often failed to detect objects that were partially occluded or held by the gripper Point completion failures arose due to a mismatch between the training partial point cloud distribution and the real-world partial point cloud distribution, as the data generation pipeline did not include occlusions caused by other objects in the scene Latency is introduced at each stage, with sequential rollouts being the main bottleneck for planning

Conclusion

In this work, we demonstrate that world models based on explicit point cloud completion lead to improved rollout quality 3DPWM achieves better predicted geometry, reward correlation, and downstream planning performance across simulated and real tasks, including adaptation to new longhorizon tasks Our analysis suggests tighter integration of perception components and reduced planning latency are critical directions for further improvement

Acknowledgements

The authors would like to thank the anonymous reviewers for their insightful feedback that helped improve the quality of the work SP would like to thank Alejo for his help with Franka arm during the sim-to-real experiments, Wesley for numerous brainstorming sessions during the early stages of the project and with his help on the Point Completion codebase, Prof. Cindy Grimm for allowing us to use her lab’s Franka arm, and DMV & ViRL labmates for their feedback on the draft

Author Contributions

Skand led the project by implementing the data generation pipeline for point completion and dynamics model training, trained dynamics model, and wrote the planner Hung Nguyen led the point completion training and helped with sim-to-real experiments in integrating SAM3 with point completion module Chanho Kim helped with writing of the paper, participated in discussions throughout the project and provided inputs on methodology Li Fuxin helped with writing of the paper and provided guidance throughout the project Stefan Lee helped with writing of the paper and provided guidance throughout the project

References

[1] D. Ha and J. Schmidhuber World models arXiv preprint arXiv:1803.

Improvements for AI systems

  1. Improved world model fidelity through explicit geometry recovery: 3DPWM enables reliable long-horizon rollouts by first completing partial point clouds, which addresses prior work's issue of suffering from compounding rollout errors in occluded scenes by ensuring the dynamics model operates on complete geometry.

  2. Enhanced planning capability via action-conditioned dynamics: The system supports more complex manipulation behaviors, including pick-and-place, by learning action-conditioned dynamics in this completed 3D scene, which facilitates more accurate cost evaluation for modelbased planning compared to models trained on partial observations.

  3. Improved generalization to novel tasks: 3DPWM demonstrates the ability to adapt to novel unseen combination of tasks because it leverages complete point clouds offer better training regime for point clouds based world models, allowing adaptation beyond the initial training distribution.

  4. Robust sim-to-real transfer: The integration of a completion module allows for effective sim-to-real transfer on tabletop manipulation tasks where 3DPWM performs 2.5× better than baselines, bridging the gap between simulation and physical reality by recovering complete geometry from partial views at deployment time.

  5. Improved planning robustness via action conditioning: The use of cross-attention with a learnable register token is found to be important for performance because it ensures that action-relevant points (i.e., corresponding to the robot) can attend to action commands, leading to better closed-loop planning success rates.

Abstract

Learning predictive models of the world enables robotic control through planning, potentially allowing robots to improvise solutions on new tasks. However, large video-based dynamics models lack explicit 3D spatial structure and suffer from geometrically inconsistent long-term rollouts with compounding errors. Emerging 3D dynamics models based on partial point clouds improve geometric consistency but remain sensitive to occlusions and accumulated prediction drift. To address these challenges, we present 3D Point World Models (3DPWM) - a task-agnostic world model that operates entirely in 3D space by first completing partial point clouds and then learning action-conditioned dynamics in this completed 3D scene. By operating on completed geometry, 3DPWM enables reliable long-horizon rollouts and more accurate cost evaluation for model-based planning while supporting adaptation to new tasks. Experiments across different robotic embodiments and tabletop manipulation benchmarks demonstrate that 3DPWM achieves significantly more reliable long-horizon rollouts (100-300+ steps), supports both open-loop and closed-loop planning, and enables successful sim-to-real transfer.

Sources

Related papers