Jianshu Zhang, Keliang Wu, Chengxuan Qian, Xiyuan Yang, Ce Zhang, Ariel Tian, Anbang Liu, Haoran Lu, Han Liu
5 min
Abstract
Embodied agents now take on ever longer tasks. For long tasks, knowing only whether a task finally succeeds or fails says little; the steps along the way matter. Progress Reward Models (PRMs) score how far a task has come at every step, and serve as dense rewards, verifiers and monitors. Yet in long tasks the current frame alone often cannot tell how far the task has come, because progress depends on what happened before. We call this problem context-dependent progress estimation. Existing benchmarks on progress estimation mostly focus on short tasks whose progress can be read from the current observation, and whether PRMs can estimate progress when context is needed remains underexplored. We therefore build ContextProgress-Bench, with 24 manipulation tasks for 120 episodes. The benchmark covers three settings: (i) State Recall, where information needed for progress appeared earlier but is not in the current frame; (ii) Sequence Tracking, where steps follow a fixed order, so progress requires knowing which steps are done and which comes next; and (iii) Recurrence Disambiguation, where look-alike frames sit at very different progress. We then run a paired diagnosis: each PRM keeps the same input format in both runs, and in one run its instruction integrates the right context. Even PRMs that read the entire history get lost in estimating progress, yet with the right context the same five models cut their progress error by 77-82%. Embodied PRMs are thus not incapable of progress estimation, but lost without the right context. We therefore propose ProgressCompass, an autonomous agentic loop that reorients an existing PRM and uses current general-purpose VLMs to supply the context the PRM needs. Wrapped in the loop, the same frozen PRM cuts its progress error by 63% and raises its rank agreement by 76%. With such a compass, PRMs estimate progress far better on longer, more complex tasks.
Sam: A narrow verifier still has an operating point. If it's too permissive it skips steps, and if it's too strict it stalls.
Alex: They treat the false-accept rate as a probability term in the error bound. The containment comes from the design. The Navigator never retroactively rewrites the past, so a premature advance is a local error, not a collapse of the whole episode.
Sam: That's a different failure profile from a monolithic model guessing from one noisy frame. What did the negative samples show?
Alex: They included early-stop and mismatched-instruction cases. The score stayed tied to the subtask being verified rather than drifting with the visuals. That supports the idea that the system checks transitions against expected state, not just pixels. I'd read it as supporting evidence, though. The load-bearing evidence is the diagnostic gap and the theorem.
Sam: Then the 77 to 82 percent drop deserves a caveat. That's with correct context supplied. A referee would ask how closely a real Orienter approximates it.
Alex: That's the right question. The supplied-context condition works as a ceiling on what retrieval can buy. The paper says ProgressCompass closes most of the performance gap, but "most" is not all. And the loop has a cost. They parallelized it and cut wall-clock time by over 65 percent, which keeps the overhead manageable.
Sam: Manageable for evaluation, perhaps. For real-time robotics it's still an agentic loop.
Alex: I'd read it as an offline tool for now. For real-time control, you'd want to internalize the context awareness in the model itself.
Sam: Distilling the Orienter-Verifier logic into a single-pass model would keep the context benefit without the loop.
Alex: That's the natural next step, and the paper provides the pieces for it: a diagnostic framework that explains why the gap exists, and a benchmark to measure whether future architectures close it. Whether the same decoupling carries over to long-form reasoning is open. It would depend on defining clear state transitions in abstract domains.
Sam: So the takeaway is that progress estimation behaves like a retrieval problem, and a frozen backbone can score well when it's given the right context at the right time.
Alex: Yes. It shifts the question from how large the model should be to where the context comes from.
Sam: If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Alex: Thanks for listening.