A safe action is not necessarily a viable one. A frozen vision-language-action (VLA) policy can favor a locally admissible move that leaves no policy-supported route to safe task completion. We call this the feasibility-likelihood gap: likelihood ranks the next move, while feasibility depends on the futures it leaves open. To bring those futures into the decision, we derive the exact next-block marginal of the history-conditioned policy-environment trajectory law restricted to safe task completion. The derivation reveals a candidate-dependent feasible-future mass: its support records whether safe completion remains possible under the frozen continuation process, while its magnitude measures how much weighted safe-completion mass remains. Since exact evaluation is impractical online, we develop a selective finite-candidate approximation and establish conditions for recovering the best retained viable candidate. Our alarm-triggered, training-free reranker VICS-G lowers mean cumulative safety cost by 1.9%-57.5% across six Safety-CHORES settings while remaining within 2.5 percentage points of policy sampling in success and 0.82 steps in mean episode length. Our approach offers a promising and practical path toward safer task completion, grounded in an exact policy-relative target yet requiring neither policy retraining nor online rollouts.
Sam: A vision-language-action policy can pick a move that looks safe and likely but leaves the robot no way to finish the task. In a preprint from Tu Nguyen and colleagues, this is called the feasibility-likelihood gap. The robot keeps retrying a blocked path because the policy rates it the most probable move, even though the path can't reach the goal.
Alex: So the failure isn't recklessness. The policy is being careful locally, and that carefulness is what traps it.
Sam: Yes. Immediate safety is satisfied, but no viable future remains. The authors' fix is to rerank candidate actions by what they call feasible-future mass: how much of the probability over future continuations ends in safe completion. Across their benchmarks, safety costs dropped by up to fifty-seven percent. That's the headline number, and I'd read it as a best case.
Alex: It's like a GPS that checks only the next turn and never asks whether the road ahead is washed out. But how do you compute that without simulating every candidate?
Sam: That is the technical core. They derive an exact next-block marginal by marginalizing over all future trajectories, which can't be run online. So they build a selective reranker, VICS-G. A gate decides whether the current action needs a second look. If it does, the decoder scores each candidate on three things: policy likelihood, local robustness, and the new continuation term.
Alex: A gate is an approximation, though. What if it discards the only viable path?
Sam: That is a real failure mode. The authors derive conditions for recovery. As long as the gate doesn't reject every viable option, the reranker can still find the best one. They trade perfect coverage for a training-free, rollout-free system, and the guarantee is conditional on that.
Alex: If the physical world doesn't match what the continuation estimate assumes, doesn't the feasible-future idea break?
Sam: That is why they test a variant that uses execution feedback. The decoder sees the history of commanded versus realized transitions and adjusts the next decision. If the robot commands a move and hits an obstacle, the decoder knows the intended action didn't happen. In their tests, this variant gave higher success rates at lower observed cost. Those are the supporting experiments, though. The core claim rests on the reranking itself.
Alex: And the cost? Does the extra processing slow the robot down?
Sam: The authors acknowledge that scaling these estimates is computationally expensive. The gate is the mitigation. The fast frozen policy runs by default, and the expensive check fires only when the gate flags a possible dead end. They also explore VICS-R, which spends a limited simulation budget on checking alternatives. Even a modest lookahead shifts outcomes in the harder environments.
Alex: And the policy weights are never touched?
Sam: Correct. It is entirely training-free, with the decoder acting as a wrapper around a frozen policy. The paper presents it as modular, but I'd want to see it tried across more policy families before treating that as established.
Alex: So can it discover safer strategies, or is it only filtering?
Sam: Filtering and reranking only. It can't expand capability, because the model can only choose among the candidates it proposes. What changes is the selection criterion, which prunes the safe dead ends. The model acts on the best available future rather than the most likely immediate move.
Alex: So the intelligence sits in evaluating consequences, not generating actions. Then performance is bounded by the safety monitor. What happens when the monitor is wrong?
Sam: That is the fundamental limitation. If the monitor labels a dangerous state as safe, the reranker will steer the robot into it. The authors are explicit that there are no formal safety guarantees. It is a heuristic improvement, not a provable safeguard, and that is where a careful referee would push hardest.
Alex: And the monitor's calibration sets the trade-off. Too conservative and the robot stalls, too loose and you get violations. Can that be tuned without retraining?
Sam: The sensitivity analyses address that. Adjusting the continuation weight shifts the balance toward task completion or toward safety, and you can do it at deployment time. That is far cheaper than retraining the policy. The caveat is that the setting still depends on how good the monitor is.
Alex: So the contribution is a cheap, tunable evaluation layer on top of a frozen policy, with the monitor as its ceiling.
Sam: Yes. It changes which actions get chosen, not what the policy can do, and its value tracks the quality of the safety signal feeding it. If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Alex: Thanks for listening.