Research / arXiv / May 24, 2026
HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos
HumanEgo turns minutes of human egocentric manipulation video into entity-level hand-object representations for zero-shot transfer to robot policies.
HumanEgo studies how short first-person human demonstrations can become useful robot-learning data. The framework starts from egocentric manipulation video, where the camera naturally follows the hands, active object, and changing task state.
Its core representation lifts each demonstration into entity-level hand-object interaction. A flow-matching policy then uses dense auxiliary objectives to extract more supervision from each trajectory and bridge visual and kinematic differences between people and robots.
The authors report 92.5 percent average success across four real-world tasks using 30 minutes of human video per task, and 75 percent with 15 minutes. They also report a 41 percent improvement over matched-time robot teleoperation in their evaluation.
For egocentric data teams, the paper highlights the value of complete manipulation sequences, visible object-state changes, consistent framing, and varied demonstrations. These capture qualities determine how much interaction structure a later learning pipeline can recover.
HumanEgo is released as an open-source framework. The cited arXiv record links the project resources and provides the complete method, experimental setup, metrics, and author analysis.
