Wei Wu, Fangjing Wang, Fan Lu, He Sun, Shi Liu, Yunnan Wang, Yibin Yan, Yong Wang, Shuailei Ma, Xinyang Wang, Yibin Liu, Shuai Yang, Tianxiang Zhou, Kejia Zhang, Lei Zhou, Cheng Su, Nan Xue, Bin Tan, Han Zhang, Youchao Zhang, Fei Liao, Xing Zhu, Yujun Shen, Kecheng Zheng
5 min
LingBot-VLA 2.0 is a vision-language-action (VLA) model designed to bridge the gap between laboratory-based robot learning and real-world deployment. The authors argue that current VLA models often struggle with the heterogeneity of real-world robotics, including diverse robot configurations, complex action spaces (such as mobile bases and dexterous hands), and the need for temporal reasoning in dynamic environments. To address these, the researchers introduce a system that scales data, expands control interfaces, and incorporates predictive modeling.
The researchers focus on three functional improvements:
Evaluations on the GM-100 benchmark and long-horizon mobile manipulation tasks demonstrate that these modifications lead to stronger cross-embodiment capabilities. By jointly addressing generalization, action-space coverage, and predictive dynamics, the authors provide a pragmatic framework for moving VLA models from controlled laboratory settings toward practical, real-world robotic applications.
Despite recent progress of VLA foundation models, the disparity between laboratory conditions and real-world applications continues to impede their practical implementation. To bridge this gap, we present LingBot-VLA 2.0, which advances LingBot-VLA through improvements in three functional domains. (1) Generalization across tasks and embodiments. Compared to the previous version, we revamp the data processing pipeline and curate around 60,000 hours of data for pretraining, including 50,000 hours of robot trajectories spanning 20 robot configurations and 10,000 hours of egocentric human videos. (2) Expanded action space in addition to dual-arm hardware platforms. In particular, our system accommodates degrees of freedom for the heads, waists, mobile bases, and dexterous hands, thereby empowering the robots to tackle more complex tasks in practical scenarios. (3) Predictive dynamics modeling for improved temporal reasoning. Specifically, we formulate future prediction as a proxy task, facilitated by a video representation model for semantic priors and a depth estimation model for geometric cues. Evaluations on the GM-100 benchmark, conducted in a generalist setting, validate the beneficial impact of these proposed modifications. Furthermore, benefiting from the expanded pretraining data that covers whole-body degrees of freedom, LingBot-VLA-2.0 demonstrates strong cross-embodiment long-horizon mobile manipulation capability across the two robotic platforms.
Alex: Precisely. The researchers call this "dual-query distillation" — two streams of questions running in parallel. One anchored in the present, one reaching into the near future. Together, they give the robot a much richer model of what's happening and what's likely to happen next.
Sam: That does sound computationally heavy, though. How do they keep it fast enough to actually control a robot in real time?
Alex: They use what's called a "Mixture-of-Experts" architecture. Rather than routing every decision through one large, slow brain, the system divides the work. Different specialist modules — "experts" — handle different types of movement. The system figures out which expert is most relevant for a given task and activates that one.
Sam: So it's a divide-and-conquer approach. Instead of one generalist doing everything slowly, you have a team of specialists, and you call on the right one for the job.
Alex: That's a good way to put it. And it also solves the problem of handling many different robot bodies. There's a shared layer that stores universal movement principles — things that apply to any robot — and then separate specialist layers that handle the quirks of each specific body. One robot might have wheels; another might have a flexible waist. The shared layer handles what's common; the specialists handle what's unique.
Sam: So the model is learning a kind of common language for movement, and then each robot gets its own dialect on top of that?
Alex: That's a useful way to think about it. And that separation is what allows the system to generalize — to take what it learned on one robot and apply it, at least partially, to a robot it's never controlled before.
Sam: And all of this is trained on a substantial amount of data, I assume?
Alex: About 60,000 hours in total. Roughly 50,000 hours of robot movement data drawn from around 20 different robot types, and another 10,000 hours of human activity video — footage of people doing everyday tasks, which helps the model understand how humans interact with objects in the world.
Sam: It's a meaningful step. Not just making one robot better at one task, but trying to build something that can transfer across bodies, environments, and situations.
Alex: That's the ambition. The persistent gap between lab performance and real-world performance has been one of the central challenges in robotics for years. What this paper proposes is a more unified approach to learning — one that accounts for physical variety, temporal prediction, and computational efficiency at the same time. Whether it fully closes that gap remains to be seen, but the direction is clear.
Sam: Thanks for walking through that, Alex. It's a good reminder of how much engineering sits behind even a simple-looking robot task.
Alex: Thanks for listening to ResearchPod.