Marius Memmel, Andrew Wagenmaker, Chuning Zhu, Patrick Yin, Dieter Fox, Abhishek Gupta
5 min
Robotic manipulation often suffers from the sim-to-real gap, where policies trained in simulation fail when deployed on physical hardware due to discrepancies in physical parameters like mass, friction, or articulation. The authors ask: can we autonomously refine a simulator using minimal real-world data to enable successful zero-shot transfer of task-specific policies?
To address this, the authors propose Active Exploration for System Identification (ASID). The framework operates in three stages:
The authors demonstrate that ASID effectively identifies unknown physical parameters across several challenging tasks, including rod balancing, sphere manipulation, and laptop articulation. By using Fisher information to guide exploration, the robot collects significantly more informative data than random exploration or standard mutual-information-based baselines. Consequently, the refined simulator allows for successful zero-shot transfer of policies that would otherwise fail under standard domain randomization or uninformed training.
ASID provides a principled, sample-efficient way to bridge the sim-to-real gap. By decoupling the exploration phase from the task-solving phase, it allows robots to learn about their environment's physics autonomously. This reduces the need for extensive human-designed simulation assets and enables robots to adapt to new, dynamic environments with minimal real-world interaction.
Model-free control strategies such as reinforcement learning have shown the ability to learn control strategies without requiring an accurate model or simulator of the world. While this is appealing due to the lack of modeling requirements, such methods can be sample inefficient, making them impractical in many real-world domains. On the other hand, model-based control techniques leveraging accurate simulators can circumvent these challenges and use a large amount of cheap simulation data to learn controllers that can effectively transfer to the real world. The challenge with such model-based techniques is the requirement for an extremely accurate simulation, requiring both the specification of appropriate simulation assets and physical parameters. This requires considerable human effort to design for every environment being considered. In this work, we propose a learning system that can leverage a small amount of real-world data to autonomously refine a simulation model and then plan an accurate control strategy that can be deployed in the real world. Our approach critically relies on utilizing an initial (possibly inaccurate) simulator to design effective exploration policies that, when deployed in the real world, collect high-quality data. We demonstrate the efficacy of this paradigm in identifying articulation, mass, and other physical parameters in several challenging robotic manipulation tasks, and illustrate that only a small amount of real-world data can allow for effective sim-to-real transfer. Project website at https://weirdlabuw.github.io/asid
Alex: [thoughtful] Notice that the exploration policy never has to solve the task.
Sam: [confirming, precise] That's the separation of concerns. Exploration only has to be informative, and task performance comes afterward in the calibrated simulator. It also tells you where the method can fail. If the exploration doesn't excite the parameters, the information stays low and the update is biased. So the bottleneck is the quality of the exploration more than the quantity of data.
Alex: [analytical, processing] And the evidence? A single-episode claim is the kind of thing I'd want to see stress-tested.
Sam: [measured, acknowledging the point] I'm working from the framework's description rather than the full tables, so I can't give you a clean read on effect size or how much data it saves. Any referee would ask how the advantage changes as the number of unknown parameters grows, and whether four tasks is enough to say it generalizes.
Alex: [curious, leaning in] What about when the real world just doesn't behave like the model?
Sam: [measured, grounded] That's the core limitation. The method assumes the true dynamics belong to a known parametric family. If the real physics deviates fundamentally, with unmodeled contact dynamics or complex fluid interactions for instance, identification will fail. The simulator can't represent what it's trying to match.
Alex: [reflective] So targeted exploration can't fix a structural error. It finds the best fit inside a restricted space, and if the truth lies outside that space, the answer is misleading.
Sam: [nodding, precise] Yes. It's a parametric identification tool, not a universal learner. The trade is efficiency and zero-shot transfer for problem classes you can already parameterize, in exchange for no coverage of genuinely novel physical behavior. Non-parametric dynamics models or latent representations would be a natural direction, though that goes beyond what this framework does.
Alex: [reflective] I'd put it this way. Instead of fighting the sim-to-real gap with brute-force randomization, you measure the gap carefully before you jump, as long as your model class can describe it.
Sam: [professional] That's the takeaway, and it also marks the boundary of the method: its value depends on how well the model class matches the physics you actually face.
Alex: [steady] If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Sam: [warm] Thanks for listening.