ResearchPod Summary
Vision-Language Navigation (VLN) agents often use on-policy exploration to improve robustness by exposing the agent to diverse states. However, when an agent deviates from the expert path, the original instruction no longer matches the visual observations, creating a semantic supervision gap. The authors introduce Phi-Nav, a framework that bridges this gap by treating exploratory trajectories as new learning opportunities. The approach uses a three-stage cycle: the agent samples a trajectory, a hindsight speaker synthesizes a path-level instruction grounded in the actual visual observations, and the agent performs a second imitation pass using this new trajectory-instruction pair as an expert demonstration.
To ensure the synthesized instructions are reliable and consistent with training data, Phi-Nav employs two primary mechanisms. First, it uses expert-in-context learning, where the hindsight speaker is provided with an example of an expert path and instruction to maintain stylistic and structural consistency. Second, it implements a trajectory-instruction alignment weighting module. This module evaluates the semantic fidelity of the generated instructions using both global and landmark-aware frame-word similarities. By adaptively weighting the hindsight supervision, the framework ensures that only semantically faithful instructions influence the policy, effectively filtering out potential hallucinations from the vision-language model.
Evaluations on the R2R-CE and RxR-CE benchmarks demonstrate that Phi-Nav consistently improves navigation success rates and path-length-weighted success rates compared to standard on-policy training methods. Notably, the framework achieves competitive performance while requiring fewer expert demonstrations, highlighting its superior sample efficiency. By transforming semantically unlabeled exploratory movements into dense training signals, Phi-Nav provides a scalable solution for training embodied agents in complex environments with limited human-annotated data.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.