Native Video-Action Pretraining for Generalizable Robot Control by Prof Yinghao Xu
Abstract
Robots need three capabilities to operate in the physical world: spatial understanding, predictive dynamics, and real-time action. In this talk, the speaker will present three pieces of the LingBot embodied intelligence stack: LingBot-Map, LingBot-VA, and LingBot-VA2. LingBot-Map addresses streaming 3D reconstruction, building compact and temporally consistent spatial memory from continuous visual input. LingBot-VA then moves from perception to control by introducing causal video-action modeling, where a robot policy predicts future visual states together with the actions that cause them. Building on this direction, LingBot-VA2 develops a native video-action foundation model designed for embodied control, with semantic visual-action tokenization, causal pretraining, sparse MoE scaling, human-robot co-training, and asynchronous closed-loop inference. Together, these systems illustrate a path from reconstructing the world, to predicting how it changes, to acting in it efficiently and robustly. The speaker will discuss the key technical ideas, real-world demonstrations, and open challenges toward scalable embodied foundation models.
About the Speaker
Yinghao Xu is an Assistant Professor in the Department of Computer Science and Engineering at the Hong Kong University of Science and Technology (HKUST). Previously, he was a Staff Research Scientist at RobbyAnt, working on world models and embodied AI. Before that, he was a postdoctoral researcher at Stanford University. His research lies at the intersection of 3D computer vision, generative AI, and embodied AI, with a recent focus on building world models that unify 3D reconstruction, world simulation, and embodied action. He was the recipient of the Yunfan Award at WAIC 2024 and was nominated for the Snap Fellowship in 2022.