Zeyu Shen, Haoxiang You, Yilang Liu, Zhicheng Zheng, Lihan Zha, Kashu Yamazaki, Mingtong Zhang, Suning Huang, Jiankai Sun, Qianzhong Chen, Lucy He, Kaiyuan Liu, Haoran Chang, Katerina Fragkiadaki, Dhruv Shah, Mac Schwager, Peter Henderson, Ian Abraham, Canwen Xu
5 min
This paper investigates the potential of coding agents—large language models that write and execute code—to serve as both solvers for complex, long-horizon robotics tasks and as automated teachers for training generalist robot policies. The authors introduce EMBODIEDSWE-BENCH, a simulation benchmark featuring 28 challenging tasks, such as IKEA furniture assembly, tool packing, and food cutting, which require precise physical interaction and long execution horizons.
The researchers evaluate frontier coding agents on their ability to solve these tasks by iteratively writing, executing, and debugging control programs. To overcome the limitations of individual task-specific solutions, they introduce EMBODIEDSWE-GEN, a hierarchical data engine. This engine takes a single verified coding-agent solution and expands it into a large, diverse dataset by systematically varying scenes, strategies, task phases, physical dynamics, and visual appearances. These generated trajectories are then used to finetune Vision-Language-Action (VLA) models.
Frontier coding agents can solve a significant portion of the benchmark tasks, with the best models achieving high success rates. The authors demonstrate that these programmatic solutions provide high-quality supervision for training VLA models, with performance scaling predictably as the number of generated demonstrations increases. Furthermore, the agent-aided diversification pipeline improves the VLA's ability to generalize to held-out task variations compared to standard domain randomization. Finally, the authors show that a VLA finetuned exclusively on these simulated, agent-generated demonstrations can successfully complete a four-stage lamp-disassembly task on a real robot.
This work provides a scalable path for generating high-quality robotics data without the need for expensive human teleoperation or manual demonstration collection. By leveraging the reasoning capabilities of coding agents to explore and solve complex tasks, researchers can bootstrap generalist robot policies that are more robust and capable of handling long-horizon, dexterous manipulation.
We study coding agents for long-horizon, dexterous robotics and ask whether their solutions can provide scalable supervision for learning general robot policies. To test this, we develop EMBODIEDSWE-BENCH, a simulation benchmark for coding agents spanning contact-rich manipulation, deformable objects, and long-horizon tasks requiring up to half an hour of continuous interaction. We find that frontier coding agents can solve complex long-horizon tasks and transfer prior solutions across both tasks and embodiments. We also design supporting tools that help agents more effectively solve these tasks. However, the resulting solutions require substantial iterative interaction and are typically specialized to individual task instances. We therefore introduce EMBODIEDSWE-GEN, which expands a single solution from coding agent into large diverse trajectories for training a VLA. VLA performance improves with more generated demonstrations, and agent-aided diversification improves generalization to held-out task variations. We also show that a VLA finetuned solely on coding-agent-generated simulation demonstrations completes a long-horizon task on real robot. Together, our framework uses coding agents to solve complex robotics tasks and turn verified solutions into scalable supervision for robot policies.
Sam: That's the case they're making. Though I'd hold onto "long-horizon" for a second, because that's doing a lot of work in what makes this hard. In a simple pick-and-place, one mistake doesn't sink the outcome. In something like motherboard assembly, the robot has to hold state accurately over minutes — drift by a few millimetres on the first bolt, and the whole assembly fails downstream.
Alex: So the coding agent isn't succeeding by being more precise. It's succeeding by structuring the problem differently?
Sam: Exactly that. Because the agent's program has access to the simulator's internal state, it can write explicit checks into the control loop — verify the bolt is actually seated before moving to the next step. It's writing its own procedure, with built-in verification, rather than hoping a single trained policy gets it right in one pass. [[RP_SECTION:simulator-fidelity-challenges|Simulator Fidelity Challenges]]
Alex: There's a clear performance gap in their results though — tool-packing does well, something delicate like the bulb-unscrewing task does worse. Is that a simulator fidelity problem?
Sam: That's exactly where it breaks down. Unscrewing a bulb depends on the contact physics of thread release — a subtle interaction the simulator has to model correctly for the agent's checks to mean anything. If the contact model is off, the agent's verification logic is checking against a physics that isn't quite the real one. They use domain randomization — varying friction, mass, texture — to bridge that gap, but randomization only covers dynamics you thought to randomize. If the simulator's contact geometry for that specific bulb is wrong in a way they didn't anticipate, the policy trained on it will struggle in the lab. It's a fairly standard sim-to-real problem, just showing up in a new place.
Alex: And going back to the grading question — you said a program has to be functioning, not just a lucky action sequence. Does that fully close off reward hacking, or can the agent still find a loophole in how the simulator reports state?
Sam: It narrows it, doesn't close it. They add an LLM judge on top of the automated grading specifically to catch cases where the agent found a bug in the simulator rather than a genuine solution. But it's an arms race — the agent is optimizing for the shortest path to a passing score, and sometimes that path is an exploit rather than competence. The ceiling on the agent's apparent intelligence, in a real sense, is the robustness of the harness grading it. [[RP_SECTION:closing-the-feedback-loop|Closing the Feedback Loop]]
Alex: So the natural next step is closing that loop — having the agent's real-world failures feed back into correcting the simulator, instead of treating the simulator as fixed ground truth.
Sam: That's the direction the authors point to, yes. Right now the simulator is upstream and fixed; the more interesting version has hardware failures updating the simulator's parameters over time.
Alex: If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Sam: Thanks for listening.