Tianshan Zhang, Yijia Duan, Yanjun Li, Zeyu Zhang, Hao Tang
5 min
Dexterous manipulation of articulated objects—such as opening doors or drawers—is challenging because the object's motion is not directly controlled but must emerge from sustained physical contact between the hand and the object. Existing reinforcement learning approaches often overfit to nominal dynamics, leading to failure when contact loads (like damping) change. The authors propose DragMesh-2, a framework that treats articulated manipulation as a contact-driven task. To improve robustness without expensive tactile or force sensors, they introduce PICA (Physically Informed Contact-Aware) training. PICA injects physical signals into the policy learning process, including contact maintenance, detachment risk, and action-boundary regularization, while using a temporal encoder to predict recent contact responses.
DragMesh-2 significantly outperforms state-of-the-art baselines in robustness to changing contact loads. While standard policies often achieve high success under nominal conditions, they frequently collapse when damping increases. PICA-trained policies maintain higher success rates across varying damping conditions (x1, x2, and x4 multipliers). The authors demonstrate that nominal success can be misleading, as policies trained for longer durations often achieve high success by entering a saturated, low-robustness regime. By explicitly incorporating physical proxies and temporal contact-response modeling, PICA shifts learning toward stable, contact-conditioned interaction rather than relying on dynamics shortcuts.
This work bridges the gap between object-centric motion generation and realistic, hand-driven dexterous interaction. By providing a pure-geometry dataset and a systematic evaluation protocol that reports OOD (out-of-distribution) robustness and action-saturation diagnostics, the authors offer a reproducible foundation for future research in humanoid and loco-manipulation. The framework demonstrates that it is possible to achieve physically plausible, robust interaction by leveraging observable physical signals and temporal history, reducing the reliance on specialized hardware like tactile sensors.
Dexterous interaction with articulated objects is important for household, assistive, and humanoid manipulation, where multi-finger hands can provide compliant contact patterns beyond parallel-jaw grasping. However, articulated-object manipulation differs from static-object manipulation: the target part cannot be directly actuated, and its motion must emerge through sustained physical hand--handle contact. This makes the transition from object-centric articulated generation to hand-driven dexterous hand--object interaction non-trivial, since geometric trajectory replay or open-loop execution does not model the contact dynamics required to move the articulated part. Moreover, policies trained only for task completion under fixed dynamics can overfit nominal contact loads, especially without tactile or force feedback, and may degrade when the contact load changes. To address these challenges, we present DragMesh-2, a contact-driven framework for dexterous interaction with articulated objects that extends articulated interaction from object-centric generation to hand-driven dexterous hand--object interaction, where articulated motion must arise through physical contact. We further propose PICA, a physically informed contact-aware training mechanism that injects physical signals into policy learning without tactile or force feedback, improving robustness and task success under changing contact loads. Finally, we conduct systematic evaluation across multiple damping conditions and articulated-object categories to study robustness under contact-load variation, and provide a pure-geometry dexterous interaction resource to support future loco-manipulation and humanoid hand--object interaction research. Across seven GAPartNet objects, DragMesh-2 achieves stronger robustness under contact-load variation than the compared methods while maintaining high task success across damping conditions.
Alex: So the goal shifts from "get it open" to "maintain a stable physical connection while getting it open." That's a meaningful difference.
Sam: It is. By treating the physical response of the object as the primary signal — rather than just tracking whether the task was completed — the robot stops trying to work around the physics and starts working with them.
Alex: I want to push on something, though. If the robot is learning from its own history, why doesn't it just get better the longer it trains? Why does the training design matter so much?
Sam: That's a common assumption, and the paper addresses it directly. Simply running the training loop for longer doesn't help if the conditions never change. If you only ever practice on one type of door, you just get better at that specific door. You're still memorizing a path — it's just a more refined one. The robot never learns to interpret feedback from an unfamiliar situation.
Alex: So they have to actively introduce uncertainty during training?
Sam: Exactly. They use a technique called dynamics randomization. Imagine practicing to open a door, but every time you try, someone secretly changes how heavy or stiff it feels. Because the conditions keep shifting, you can't rely on muscle memory. You're forced to pay attention to what the door is actually doing in the moment. That's what dynamics randomization does for the robot — it makes shortcuts useless, so the robot has to develop genuine adaptability.
Alex: It's like practicing a sport against different opponents instead of just hitting a ball against the same wall over and over.
Sam: That's a precise analogy. And it's worth being clear about what this approach still can't do. Because the robot doesn't have actual force sensors on its fingers, it's inferring the contact state indirectly — from motion history, not from direct touch. That works well under moderate conditions, but it has limits.
Alex: So under really demanding conditions — an extremely heavy door, say — does the system start to break down?
Sam: It does. When the resistance becomes extreme, the robot's indirect inferences aren't enough to compensate, and performance drops. The authors acknowledge this openly. They note that the main bottleneck is the absence of real force-sensing hardware, and they point to that as the natural direction for future work — either adding physical sensors or developing more sophisticated ways to track how an object moves in response to being pulled.
Alex: So this is a meaningful step toward robots that can handle the messiness of the real world, but it's not a complete replacement for physical sensing yet.
Sam: That's a fair reading of the paper. What the research demonstrates is that by designing the training process to include physical signals — tracking grip error, predicting detachment risk — you can get considerably closer to robust, adaptable manipulation. The key insight is that the training process itself has to respect the physics of the world, not just reward the robot for getting the task done by any means available.
Alex: That's a useful distinction to end on. Thanks for walking through the mechanics, Sam. And thank you for listening to ResearchPod.