Collect
Record human demonstrations with dual egocentric views and dynamics-guided feedback.
Learn from human demonstrations alone.
Walk, carry, and interact over long distances.
1 Beihang University 2 Shanghai Innovation Institute 3 Sichuan University
4 Alibaba Group
* Work done during an internship at Alibaba Token Hub (ATH), Alibaba Group.
01 / The gap
Long-range loco-manipulation connects walking, reaching, and interacting across a scene. People demonstrate these skills naturally, but differences in body proportions and controller response make their motions difficult for a humanoid to realize.
EgoAlign bridges this gap by turning egocentric human demonstrations into robot-compatible action and state supervision, enabling task learning without physical-robot demonstrations.
02 / The method
We capture egocentric views, body motion, and hand commands, then adapt the motion and reconstruct corresponding robot states. Using only these human task demonstrations, we fine-tune a VLA for long-range, closed-loop humanoid execution.
Record human demonstrations with dual egocentric views and dynamics-guided feedback.
Combine kinematic scale alignment with controller-in-the-loop refinement.
Pair pre-action robot states with same-tick motion tokens from the final rollout.
03 / Demonstration collection
A PICO headset and five trackers capture body motion, while two GoPro cameras record downward and level views. Corresponding human and robot cameras have matched ground-relative heights.
Real-time SONIC–MuJoCo feedback lets demonstrators see the robot’s realized motion and adjust their movements accordingly.
The collection example below shows six human demonstrations in the time used to collect one teleoperated demonstration.
04 / Scale alignment
EgoAlign preserves global travel references while adapting upper-body interaction geometry. Kinematic alignment and controller-in-the-loop refinement correct hand-position errors; during deployment, visual feedback guides the robot’s walking toward its goal.
Final causal replay constructs robot-state and action supervision for the adapted references. At deployment, the VLA predicts motion tokens, while SONIC decodes them using live robot state history.
05 / Physical deployment
With no physical-robot task demonstrations, the fine-tuned VLA deploys zero-shot on a Unitree G1. Through SONIC’s continuous whole-body interface, it performs long-range object relocation, position-generalized navigation, and foot interaction.
The robot approaches a pedal bin and opens it with its foot. Further demonstrations show navigation to different bin locations and responses to a moving target.
Across an approximately 10-m task, the robot approaches a basket, picks it up, transports it, and places it on a table—all learned from human demonstrations. The videos also show heading adjustments toward a relocated table.
Bridging the human–humanoid gap
By aligning egocentric data with the target robot, EgoAlign enables human-only task training for zero-shot, long-range loco-manipulation.
Reference
@misc{jiang2026egoalign,
title = {{EgoAlign}: Bridging the Human-Humanoid Gap for Long-Range Loco-Manipulation},
author = {Yiming Jiang and Jin Chen and Chongyang Xu and
Yilun Chen and Aimin Hao and Yisheng He},
year = {2026},
eprint = {2609.38046},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2609.38046}
}