Yoojin Oh, Jeongsol Kim, Yeonwoo Seo, Jangho Park, Seonghyun Jin, Sunwoo Park, Youngmin Kim, Youngjun Jun, Kyumin Choi, Jong Chul Ye
9 min
Abstract
Vision-Language-Action (VLA) foundation models have scaled rapidly to enhance manipulation performance and generalizability, but this scaling incurs high computational costs that render real-world deployment increasingly challenging. Existing approaches typically mitigate this issue by designing smaller architectures or reducing the iterative denoising steps in flow-based policies. In this work, we propose FastOPD, a foundation-to-lightweight VLA framework that enables the practical deployment of large-scale VLAs through efficient on-policy distillation. Specifically, FastOPD adapts a flow map for single-state teacher supervision and combines it with a self-consistency objective to construct a compact student that learns the teacher dynamics. Furthermore, we theoretically demonstrate that minimizing this objective allows the distilled student to recover a distribution on par with that induced by an ideal few-step teacher model. We evaluate FastOPD across diverse foundation policies in simulation and real-world experiments. On LIBERO, FastOPD retains 84% of the performance of $π_{0.5}$ with only two inference steps, reducing inference latency by 78.1% while outperforming existing few-step distillation baselines in average success rate. With LingBot-VLA as the teacher, FastOPD improves the single-step success rate over the base student by 15.9 percentage points on RoboTwin 2.0. We further demonstrate its applicability to a World Action Model (WAM) and deploy a compact student distilled from MolmoAct2 on a real robot.
Alex: It needs another objective. The student learns a flow map, which is a jump between stages of action generation. At the sampled checkpoint, it matches the teacher's velocity: the local direction and rate of change in the evolving action sample.
Sam: So that velocity is not the robot arm's physical speed. It's how the candidate action changes during generation.
Alex: Yes. Then self-consistency connects that local guidance to longer jumps. The target for a direct jump averages the student's local direction at the starting point and at an estimated midpoint. Those local directions are the same kind of quantities the teacher supervises.
Sam: The teacher anchors the local directions, and the consistency objective makes the longer jump respect them. What does that look like in an actual experiment?
Alex: On LIBERO, a single-arm manipulation benchmark, the teacher is pi zero point five. The student is SmolVLA with an added layer that tells its action generator where a jump starts and ends. During distillation, the vision-language backbone stays frozen; training updates the action expert and projection layers.
Sam: Is the consistency target just something intuitively sensible, or do the authors establish a connection to the teacher?
Alex: They establish a conditional connection. Their proof bounds the difference between student and teacher-derived shortcut maps using the training loss. It requires a sufficiently regular teacher velocity field and assumptions that supervised states adequately cover the states used in training and sampling.
Sam: And does that establish recovery of the original teacher's action distribution, or a particular approximation to it?
Alex: A particular approximation. The theorem concerns an ideal few-step shortcut constructed from the teacher's velocity field. Under the assumptions, reducing the objective to zero makes the student's terminal distribution approach that shortcut distribution on the same sampling schedule. It does not guarantee the original teacher's full task performance.
Sam: Let's separate the theory from the observed performance. What is the strongest practical comparison?
Alex: On LIBERO, FastOPD's best average success rate is about eighty-two percent. The larger teacher achieves about ninety-eight percent. The student uses only two generation steps for that result, and beats the initial student and all tested few-step distillation baselines in average success.
Sam: That's useful compression, but a real performance gap. What does the speed improvement mean under their measurement setup?
Alex: The reported inference latency falls by about seventy-eight percent relative to the teacher's usual sampling setup. They measure on a single RTX 3090, using batch size one. Their timing includes model inference but excludes preprocessing and post-processing.
Sam: Could the training savings come at the expense of how quickly the student learns?
Alex: Their comparison suggests otherwise. FastOPD reaches a similar single-step success level to conventional on-policy distillation about six times faster in wall-clock training time. That supports the central efficiency claim: fewer teacher evaluations can still provide useful learning.
Sam: LIBERO is one benchmark. Does the advantage survive more complicated manipulation or a different kind of teacher?
Alex: They also test RoboTwin, which evaluates bimanual manipulation in clean and randomized scenes. With LingBot-VLA as teacher, single-step success improves by about sixteen percentage points over the base student. But at longer sampling schedules, the distilled student performs slightly below the base student.
Sam: So the advantage is strongest where the original small model struggles: generating actions with very few steps. It's not an across-the-board improvement.
Alex: That's the authors' interpretation too. They also distill Fast-WAM, a World Action Model built on a video-generation model. It improves the same student's few-step performance, suggesting the method can transfer knowledge across substantially different teacher architectures.
Sam: Which experiment shows that both parts of the objective are necessary, rather than one part carrying the result?
Alex: The loss-component ablation is especially informative. Self-consistency alone collapses to poor task performance. Teacher matching alone improves longer-step generation but struggles with the larger jumps needed for few-step generation. Combining them produces the strong LIBERO result.
Sam: That also limits the shortcut story. You cannot simply skip teacher checks and hope the student keeps itself coherent.
Alex: The local teacher anchor matters. A separate weight sweep shows that weakening it too much causes performance to collapse. Consistency helps propagate useful guidance, but it does not supply that guidance by itself.
Sam: What happens on a physical robot, where a successful simulation policy is not the whole deployment story?
Alex: They evaluate picking up a ball and placing it on a plate, using the right arm of a bimanual YAM robot. A student distilled from MolmoAct2 reaches fifty percent success. It improves over the base student and completes successful episodes faster than the base student's longer-step version.
Sam: Half the trials succeeding is evidence of transfer, not evidence of reliable deployment. How broad is that physical evaluation?
Alex: Each policy gets fifty evaluation episodes, but only that ball-and-plate task is evaluated. Completion time is averaged over successful episodes only. And the simulation results all come from a single training run, so they do not establish robustness across training seeds.
Sam: Given those boundaries, who should read the full document, and where should they begin?
Alex: Researchers deploying robot foundation policies should start with the simulation comparisons and the loss-component ablation. Check the few-step gains against the remaining teacher-performance gap. For a physical deployment, read the real-world execution settings; for the theory, inspect the ideal-shortcut definition and coverage assumptions.
Sam: And what should everyone else carry away?
Alex: A small robot model can learn to take bigger computational steps, but local teacher guidance and consistency have to work together. This study shows a promising way to buy speed—not a way to erase the cost in capability.
Sam: Thanks for listening.