Research / arXiv / Jul 16, 2026
AeroAct: Action-Centered World-Action Models for Language-Conditioned Quadrotor Flight
AeroAct predicts smooth quadrotor trajectory-action chunks from egocentric visual history, proprioception, and language-conditioned goals.
AeroAct addresses language-conditioned quadrotor flight under rapidly changing first-person views. The model must connect a semantic goal with visual history and produce smooth control references that a physical aircraft can execute.
Its action-centered world-action model adapts a pretrained video diffusion Transformer to predict local trajectory-action chunks from egocentric visual history, proprioception, and language.
During training, future first-person frames provide dense supervision about the visual consequences of motion. During deployment, the system directly decodes actions from the learned representation.
The authors also introduce a handheld collection device that couples camera observations with motion estimates to reproduce flight-like egocentric trajectories. Simulation and real-world experiments report benefits from temporal visual context in target tracking and object search.
The cited arXiv record provides the complete model, data-generation pipeline, flight experiments, and reported results.
