DefWAM: 3D Point Flow Prediction Learns Effective Goal-Conditioned Deformable Object Manipulation
Abstract
Manipulating deformable objects toward desired 3D configurations is a fundamental capability for robots operating in unstructured physical environments. However, goal-conditioned deformable object manipulation in 3D requires reasoning over high-dimensional shape variation, complex dynamics, and complex gripper-object interactions. We propose DefWAM, a 3D World Action Modeling approach that learns goal-conditioned 3D point-flow prediction as a unified objective for modeling both robot behavior and deformable scene evolution. By representing the robot and scene as point clouds in a unified 3D space, and actions as 3D point flows, DefWAM predicts goal-conditioned joint robot and scene flow and models deformable dynamics directly in particle space. This joint 3D point-flow objective enables our approach to learn when and how to use different contact modes, such as grasping, releasing, and re-grasping, to drive deformable objects toward diverse target shapes. Our experiments show that DefWAM outperforms action-only 3D goal-conditioned imitation learning baselines, particularly under large variation in initial and goal configurations. Furthermore, DefWAM's predicted scene evolution enables inference-time steering to avoid collisions by implicitly simulating possible robot-scene futures.