Research / arXiv / Jun 10, 2026
EgoEngine: From Egocentric Human Videos to High-Fidelity Dexterous Robot Demonstrations
EgoEngine converts egocentric human manipulation video into a robot-view observation sequence and a task-aligned executable action trajectory.
EgoEngine addresses the visual and action gaps between human egocentric demonstrations and robot execution. It begins with first-person RGB video that preserves the manipulated object, hand motion, and surrounding task context.
The framework produces two coordinated outputs. One is a robot observation video that replaces the human embodiment while retaining scene context and temporal alignment. The other is a task-aligned robot action trajectory generated under feasibility constraints.
The paper evaluates the conversion pipeline in simulation and on real robots. Its reported experiments demonstrate zero-shot visuomotor dexterous policy learning from egocentric human videos in the tested tasks.
The capture lesson is practical: useful source video must preserve the complete action, stable object identity, visible state changes, and enough temporal context for observation and trajectory conversion to remain aligned.
The cited arXiv publication contains the full architecture, conversion procedure, evaluation design, and project link.
