Tongyi Lab logoEgoAlign

Learn from human demonstrations alone.
Walk, carry, and interact over long distances.

EgoAlign

Bridging the Human–Humanoid Gap
for Long-Range Loco-Manipulation

Yiming Jiang1,4,* Jin Chen2,4 Chongyang Xu3,4 Yilun Chen4,† Aimin Hao1,‡ Yisheng He4,†,‡

1 Beihang University   2 Shanghai Innovation Institute   3 Sichuan University
4 Alibaba Group

* Work done during an internship at Alibaba Token Hub (ATH), Alibaba Group.

† Co-project leaders · ‡ Co-corresponding authors.

Human-only task trainingZero-shot physical deployment

01 / The gap

Beyond stationary tasks.
Learn to move and interact.

Long-range loco-manipulation connects walking, reaching, and interacting across a scene. People demonstrate these skills naturally, but differences in body proportions and controller response make their motions difficult for a humanoid to realize.

EgoAlign bridges this gap by turning egocentric human demonstrations into robot-compatible action and state supervision, enabling task learning without physical-robot demonstrations.

02 / The method

From human demonstrations
to robot-compatible data.

We capture egocentric views, body motion, and hand commands, then adapt the motion and reconstruct corresponding robot states. Using only these human task demonstrations, we fine-tune a VLA for long-range, closed-loop humanoid execution.

01

Collect

Record human demonstrations with dual egocentric views and dynamics-guided feedback.

02

Adapt

Combine kinematic scale alignment with controller-in-the-loop refinement.

03

Reconstruct

Pair pre-action robot states with same-tick motion tokens from the final rollout.

03 / Demonstration collection

Capture human skill.
See the robot’s response.

A PICO headset and five trackers capture body motion, while two GoPro cameras record downward and level views. Corresponding human and robot cameras have matched ground-relative heights.

Body motion, hand commands, and complementary egocentric views.

Feedback while collecting

Real-time SONIC–MuJoCo feedback lets demonstrators see the robot’s realized motion and adjust their movements accordingly.

More demonstrations, less on-site effort

The collection example below shows six human demonstrations in the time used to collect one teleoperated demonstration.

A collection example; the paper reports on-site timing statistics separately.

04 / Scale alignment

Align the motion.
Account for the controller.

EgoAlign preserves global travel references while adapting upper-body interaction geometry. Kinematic alignment and controller-in-the-loop refinement correct hand-position errors; during deployment, visual feedback guides the robot’s walking toward its goal.

NoAlign and aligned motion during object pickup.

Final causal replay constructs robot-state and action supervision for the adapted references. At deployment, the VLA predicts motion tokens, while SONIC decodes them using live robot state history.

05 / Physical deployment

Human-only task training.
Long-range robot behavior.

With no physical-robot task demonstrations, the fine-tuned VLA deploys zero-shot on a Unitree G1. Through SONIC’s continuous whole-body interface, it performs long-range object relocation, position-generalized navigation, and foot interaction.

Navigation & Foot-Based Interaction

The robot approaches a pedal bin and opens it with its foot. Further demonstrations show navigation to different bin locations and responses to a moving target.

Complete behavior shown in video; navigation and foot interaction are evaluated separately in the paper.

Long-Range Object Relocation

Across an approximately 10-m task, the robot approaches a basket, picks it up, transports it, and places it on a table—all learned from human demonstrations. The videos also show heading adjustments toward a relocated table.

Bridging the human–humanoid gap

Train on human demonstrations.
Go the distance on a humanoid.

By aligning egocentric data with the target robot, EgoAlign enables human-only task training for zero-shot, long-range loco-manipulation.

Reference

Citation

@misc{jiang2026egoalign,
  title = {{EgoAlign}: Bridging the Human-Humanoid Gap for Long-Range Loco-Manipulation},
  author = {Yiming Jiang and Jin Chen and Chongyang Xu and
            Yilun Chen and Aimin Hao and Yisheng He},
  year = {2026},
  eprint = {2609.38046},
  archivePrefix = {arXiv},
  primaryClass = {cs.RO},
  url = {https://arxiv.org/abs/2609.38046}
}