ResearchPod Summary
Vision-Language-Action (VLA) policies often struggle to generalize to novel objects that differ in appearance or geometry from their training data. While collecting more human-led demonstrations is the standard solution, it is prohibitively expensive and time-consuming. This paper asks: can we synthesize high-quality, physically plausible training data for novel objects by repurposing existing successful demonstrations without additional data collection?
The authors introduce Pose6DAug, a framework that performs "failure-driven" data augmentation. Instead of using 2D video editing—which often fails to maintain multi-view consistency or physical realism—Pose6DAug operates in 3D space.
Pose6DAug significantly improves the generalization of VLA policies. In experiments on the RoboCasa benchmark, fine-tuning a VLA model with Pose6DAug-augmented data resulted in a 16.5% relative improvement in success rates on novel objects compared to state-of-the-art baselines. The authors demonstrate that by maintaining 3D structural coherence, their method avoids the artifacts common in 2D-only generative editing, leading to more reliable robot behavior in pick-and-place tasks.
This work provides a scalable path for robot learning by maximizing the utility of existing demonstration data. By bridging the gap between visual generative models and 3D kinematic constraints, Pose6DAug offers a practical way to teach robots to handle a wider variety of objects without requiring constant human intervention or expensive new data collection.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.