Back to Blogs

Research / arXiv / Jul 17, 2026

Exo2EgoPose: Leveraging Exocentric Demonstrations for Vision-Language guided Egocentric 3D Hand Pose Forecasting

Exo2EgoPose uses exocentric demonstrations to guide vision-language prediction of future 3D hand poses from dynamic egocentric observations.

Exo2EgoPose: Leveraging Exocentric Demonstrations for Vision-Language guided Egocentric 3D Hand Pose Forecasting related egocentric vision research visual
TINTELE GLOBAL CO., LIMITED editorial context image

Exo2EgoPose introduces vision-language-guided egocentric 3D hand-pose forecasting. The task predicts future hand poses from first-person visual observations, a language instruction, and current pose states.

Dynamic motion and partial views make egocentric forecasting difficult. The framework uses stable, wider exocentric demonstrations as guidance for spatial context and temporal development.

A dual-level exocentric reconstruction module learns video-level and chunk-level representations. A global-to-local modulation module then uses those reconstructed features to refine the egocentric prediction.

The paper evaluates the method on AssemblyHands, Ego-Exo4D, and the authors' EgoMe-pose benchmark, reporting substantial improvements over the evaluated methods.

The cited arXiv record contains the full architecture, benchmark construction, experimental comparisons, and updated paper versions.

egocentric vision3D hand posevision-language learningexocentric guidance
Explore more EGO field guidesDiscuss a camera data project