Embodied agents now take on ever longer tasks. For long tasks, knowing only whether a task finally succeeds or fails says little; the steps along the way matter. Progress Reward Models (PRMs) score how far a task has come at every step, and serve as dense rewards, verifiers and monitors. Yet in long tasks the current frame alone often cannot tell how far the task has come, because progress depends on what happened before. We call this problem context-dependent progress estimation. Existing benchmarks on progress estimation mostly focus on short tasks whose progress can be read from the current observation, and whether PRMs can estimate progress when context is needed remains underexplored. We therefore build ContextProgress-Bench, with 24 manipulation tasks for 120 episodes. The benchmark covers three settings: (i) State Recall, where information needed for progress appeared earlier but is not in the current frame; (ii) Sequence Tracking, where steps follow a fixed order, so progress requires knowing which steps are done and which comes next; and (iii) Recurrence Disambiguation, where look-alike frames sit at very different progress. We then run a paired diagnosis: each PRM keeps the same input format in both runs, and in one run its instruction integrates the right context. Even PRMs that read the entire history get lost in estimating progress, yet with the right context the same five models cut their progress error by 77-82%. Embodied PRMs are thus not incapable of progress estimation, but lost without the right context. We therefore propose ProgressCompass, an autonomous agentic loop that reorients an existing PRM and uses current general-purpose VLMs to supply the context the PRM needs. Wrapped in the loop, the same frozen PRM cuts its progress error by 63% and raises its rank agreement by 76%. With such a compass, PRMs estimate progress far better on longer, more complex tasks.
Alex: Current progress reward models aren't blind to task progress. They're lost without historical context. A recent diagnostic study suggests the bottleneck is information retrieval, not model capacity.
Sam: That's a strong thing to say about frozen models. What's the evidence that capacity isn't the limit?
Alex: A paired diagnosis. For each model, they ran the same query twice, once with raw history and once with interpreted context injected. When the correct context was supplied, progress estimation error dropped by 77 to 82 percent. The models could score the task, but they lacked the facts to anchor the score. Raw history alone wasn't enough. They needed explicit context.
Sam: What does that failure look like in a real manipulation task?
Alex: Imagine a robot moving three blocks to a mat. If it's holding a block, it can't know its progress without knowing which blocks were already moved. The current frame doesn't encode that.
Sam: And if the robot repeats a motion, it can't tell the first repetition from the second just by looking at the block.
Alex: That's the recurrence disambiguation problem. The authors built ContextProgress-Bench to isolate three demands: state recall, sequence tracking, and recurrence disambiguation. Each one forces the model to link history to the current frame.
Sam: Is this only an empirical pattern, or do they argue it's structural?
Alex: They offer a theorem. It breaks the error into three floors, for state, sequence, and recurrence, that no current-frame estimator can get under, whatever its capacity. Raw observations are ambiguous, so even an infinite-capacity model reading them would hit those floors. Scaling alone can't remove them. The full proof is in the appendix.
Sam: So the method has to supply the missing information from outside the frame. How does ProgressCompass do that?
Alex: It decouples retrieval from scoring. A vision-language model acts as an Orienter, reading frames to provide the missing context. The frozen reward model then does the scoring.
Sam: Like an open-book exam. The student knows the physics, but the textbook tells them which chapter they're in. Where does the loop go wrong, though? An Orienter could feed the scorer a confident but wrong context.
Alex: That's what the Verifier is for, and the authors keep its role narrow. It only checks for specific, expected transitions, not global progress. The Navigator then updates the global state only after the Verifier confirms a transition. That's how they guard against phase drift, where the system gets ahead of itself.
Sam: A narrow verifier still has an operating point. If it's too permissive it skips steps, and if it's too strict it stalls.
Alex: They treat the false-accept rate as a probability term in the error bound. The containment comes from the design. The Navigator never retroactively rewrites the past, so a premature advance is a local error, not a collapse of the whole episode.
Sam: That's a different failure profile from a monolithic model guessing from one noisy frame. What did the negative samples show?
Alex: They included early-stop and mismatched-instruction cases. The score stayed tied to the subtask being verified rather than drifting with the visuals. That supports the idea that the system checks transitions against expected state, not just pixels. I'd read it as supporting evidence, though. The load-bearing evidence is the diagnostic gap and the theorem.
Sam: Then the 77 to 82 percent drop deserves a caveat. That's with correct context supplied. A referee would ask how closely a real Orienter approximates it.
Alex: That's the right question. The supplied-context condition works as a ceiling on what retrieval can buy. The paper says ProgressCompass closes most of the performance gap, but "most" is not all. And the loop has a cost. They parallelized it and cut wall-clock time by over 65 percent, which keeps the overhead manageable.
Sam: Manageable for evaluation, perhaps. For real-time robotics it's still an agentic loop.
Alex: I'd read it as an offline tool for now. For real-time control, you'd want to internalize the context awareness in the model itself.
Sam: Distilling the Orienter-Verifier logic into a single-pass model would keep the context benefit without the loop.
Alex: That's the natural next step, and the paper provides the pieces for it: a diagnostic framework that explains why the gap exists, and a benchmark to measure whether future architectures close it. Whether the same decoupling carries over to long-form reasoning is open. It would depend on defining clear state transitions in abstract domains.
Sam: So the takeaway is that progress estimation behaves like a retrieval problem, and a frozen backbone can score well when it's given the right context at the right time.
Alex: Yes. It shifts the question from how large the model should be to where the context comes from.
Sam: If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Alex: Thanks for listening.