Manipulation behaviors vary widely across objects and scenes, but they share a small set of reusable skills, and planning with these skills helps embodied agents generalize to new tasks. Yet an agent can only plan with skills it knows. Recovering skills from observed experience, the inverse of planning, builds this knowledge over time and yields skill data for training future agents. Vision-Language Models (VLMs) describe individual manipulation events well, but can they organize a stream of events into reusable skills? We formulate this problem as Streaming Embodied Skill Discovery (SESD): a model watches videos in sequence and maintains a persistent skill library that shapes its later decisions. To systematically measure this ability, we introduce Video2Skill, a benchmark that covers robot tabletop manipulation and human kitchen activity and tests three core capabilities: (i) locating manipulation events in time, (ii) grouping events of the same transformation, and (iii) deciding when to reuse an existing skill or create a new one. Across 19 open-source VLMs, many models group events at near-chance level, and scale does not consistently help. Their errors depend on how perception and library updates are coupled: joint models merge distinct transformations into one skill, while models that update the library from text descriptions duplicate recurring ones. Supervised fine-tuning, including our counterfactual library-state rebalancing (CLaRe), improves grouping but exposes a deeper bottleneck: trained models consolidate familiar skills yet rarely expand the library. Their libraries stall below half the reference size, and transformations unseen in training are located in time but almost never given a new skill. Recognizing when existing skills are insufficient thus emerges as the central challenge.
Sam: When vision-language models build a skill library from streaming video, they fail in two opposite ways. They either merge distinct transformations into one broad skill, or split a single recurring skill into several duplicates. That's the pattern reported in a benchmark paper called Video2Skill.
Alex: So the bottleneck isn't recognizing the action. It's keeping track of what the model has already learned across a long video?
Sam: Yes. The paper frames it as an inverse planning problem. At each new observation, the model has to decide whether it warrants a new schema or reuses one already in the library. It's a bit like a dynamic dictionary: is this a new word, or a synonym for one I know?
Alex: How does the benchmark force that decision?
Sam: They call the setup Streaming Embodied Skill Discovery. The model keeps a persistent library as the video streams in. Scoring uses the Adjusted Rand Index, which checks whether the model's clustering of events matches ground truth regardless of what it names each skill. So label wording doesn't matter, only the grouping.
Alex: Does scaling help? And is the failure structural or just a training data issue?
Sam: The authors point to structure, specifically how perception and library updates are coupled. Unified models tend to overmerge. Factorized ones, where perception and library update are separate, tend to fragment.
Alex: That sounds like it compounds. If the model mislabels an early event, the library state it conditions on afterwards is already wrong.
Sam: That's the intuition behind their training fix, Counterfactual Library-state Rebalancing, or CLARE. They edit the library state in training examples, for instance removing a schema the video requires, while holding the visual observations fixed. The model has to recover the right library state from the visual evidence alone.
Alex: Did it fix the stalling? I recall the abstract saying libraries still stall at about half the reference size, even with fine-tuning.
Sam: That's the limitation that matters most. Supervision improves grouping, but the models rarely expand the library for transformations unseen in training. Reuse recall stays near ninety-eight percent, while creation recall barely reaches forty.
Alex: So supervision is buying consistency, not necessarily accuracy. And the Adjusted Rand Index would reward exactly that.
Sam: Right. Better grouping of familiar events raises partition agreement, while the model still doesn't trigger a new schema for a genuinely novel one. The metric can improve while the thing you care about stays flat.
Alex: What about the factorized models? You said they fail the other way.
Sam: They improve on recurrence grouping but fragment heavily. They may detect that a transformation occurred, yet fail to retrieve the right existing schema, so they mint a new label for a recurring event.
Alex: So neither paradigm bridges recognizing a visual event and correctly indexing it in an evolving library. Did CLARE help on that?
Sam: Only in a setting-specific way. On the HD-EPIC domain it improved creation recall but slightly reduced overall partition agreement. It makes the model more sensitive to change without making it better at the create-or-reuse decision itself.
Alex: The ablation trajectories seem the most telling evidence here. If you front-load all the novel transformations early, does the library end up bigger?
Sam: No, and that's the key observation. Whether new transformations are front-loaded or spread out, the final library size is essentially the same. Even with oracle-history supervision, the model rarely creates a schema for an unseen transformation.
Alex: So it's not a capacity issue. It's a decision bias. The model forces an observation into an existing, slightly wrong bucket rather than create a new one.
Sam: That's the paper's reading. The supervision teaches the boundaries of skills the model already knows, but not the boundary of its own ignorance. The authors describe the behavior as pattern matching against internal state rather than true inverse planning. I'd treat that as an interpretation consistent with the data, not something the experiments isolate directly.
Alex: What would a fix look like, then?
Sam: The authors argue that current metrics reward how well a model organizes what it sees, but don't penalize failing to expand the repertoire. Their suggestion is to evaluate temporal localization and the quality of the abstraction jointly.
Alex: Which makes sense. Finding the event isn't enough. The model has to know whether it has seen something like it before.
Sam: And at present, the evidence suggests it mostly assumes it has. If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Alex: Thanks for listening.