Jianshu Zhang, Ce Zhang, Xiyuan Yang, Chenwei Xu, Haoran Lu, Yijiang Li, Yaqi Xie, Katia P. Sycara, Han Liu
5 min
Abstract
Manipulation behaviors vary widely across objects and scenes, but they share a small set of reusable skills, and planning with these skills helps embodied agents generalize to new tasks. Yet an agent can only plan with skills it knows. Recovering skills from observed experience, the inverse of planning, builds this knowledge over time and yields skill data for training future agents. Vision-Language Models (VLMs) describe individual manipulation events well, but can they organize a stream of events into reusable skills? We formulate this problem as Streaming Embodied Skill Discovery (SESD): a model watches videos in sequence and maintains a persistent skill library that shapes its later decisions. To systematically measure this ability, we introduce Video2Skill, a benchmark that covers robot tabletop manipulation and human kitchen activity and tests three core capabilities: (i) locating manipulation events in time, (ii) grouping events of the same transformation, and (iii) deciding when to reuse an existing skill or create a new one. Across 19 open-source VLMs, many models group events at near-chance level, and scale does not consistently help. Their errors depend on how perception and library updates are coupled: joint models merge distinct transformations into one skill, while models that update the library from text descriptions duplicate recurring ones. Supervised fine-tuning, including our counterfactual library-state rebalancing (CLaRe), improves grouping but exposes a deeper bottleneck: trained models consolidate familiar skills yet rarely expand the library. Their libraries stall below half the reference size, and transformations unseen in training are located in time but almost never given a new skill. Recognizing when existing skills are insufficient thus emerges as the central challenge.
Alex: What about the factorized models? You said they fail the other way.
Sam: They improve on recurrence grouping but fragment heavily. They may detect that a transformation occurred, yet fail to retrieve the right existing schema, so they mint a new label for a recurring event.
Alex: So neither paradigm bridges recognizing a visual event and correctly indexing it in an evolving library. Did CLARE help on that?
Sam: Only in a setting-specific way. On the HD-EPIC domain it improved creation recall but slightly reduced overall partition agreement. It makes the model more sensitive to change without making it better at the create-or-reuse decision itself.
Alex: The ablation trajectories seem the most telling evidence here. If you front-load all the novel transformations early, does the library end up bigger?
Sam: No, and that's the key observation. Whether new transformations are front-loaded or spread out, the final library size is essentially the same. Even with oracle-history supervision, the model rarely creates a schema for an unseen transformation.
Alex: So it's not a capacity issue. It's a decision bias. The model forces an observation into an existing, slightly wrong bucket rather than create a new one.
Sam: That's the paper's reading. The supervision teaches the boundaries of skills the model already knows, but not the boundary of its own ignorance. The authors describe the behavior as pattern matching against internal state rather than true inverse planning. I'd treat that as an interpretation consistent with the data, not something the experiments isolate directly.
Alex: What would a fix look like, then?
Sam: The authors argue that current metrics reward how well a model organizes what it sees, but don't penalize failing to expand the repertoire. Their suggestion is to evaluate temporal localization and the quality of the abstraction jointly.
Alex: Which makes sense. Finding the event isn't enough. The model has to know whether it has seen something like it before.
Sam: And at present, the evidence suggests it mostly assumes it has. If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Alex: Thanks for listening.