InfiniHand˙

STREAMING · WORLD-SPACE · EGOCENTRIC

InfiniHand: Streaming World-Space
Hand Motion Estimation
from Egocentric Video

Kerui Ren1,2Kaiwen Song1,3Weiguang Zhao1,4Yuxi Wang5Yufei Liu2Bo Dai6Haoyu Guo1Chunhua Shen7,1Mulin Yu1Tao Lu1Junting Dong1
1Shanghai Artificial Intelligence Laboratory2Shanghai Jiao Tong University3University of Science and Technology of China4University of Liverpool5Nanyang Technological University6The University of Hong Kong7Zhejiang University

TL;DR A unified streaming model turns egocentric video into accurate hand motion in world coordinates, with large-scale egocentric pretraining and fast inference.

01 / INTERACTIVE DEMO

See the hands. Follow the motion.

Seven scenes. Four methods. Drag the three dividers to compare.

ARCTIC · Box

Loading scene…
World-space motion

Loading 3D mesh…

Camera-space reconstructionDrag ↔ to compare
OursViDiHandHaWoRWiLoR
↔
↔
↔
0:00 / 0:00

Left: world-space hand motion—drag to rotate, scroll to zoom, and switch between Ours and HaWoR. Right: synchronized camera-space results—drag the three dividers to compare Ours, ViDiHand, HaWoR, and WiLoR. Use the shared timeline to play or scrub both views.

02 / MOTIVATION

What do we need for an
egocentric video annotation pipeline?

  1. 01

    More accurate
    reconstruction

    A streaming 3D foundation model supplies geometric priors that help locate hands in space. Combining this context with hand-centered appearance supports detailed reconstruction even under occlusion.

    Geometry-aware hand motion
  2. 02

    More diverse
    egocentric data

    We curate approximately 5,000 hours of public egocentric video and clean the annotations into a large-scale training corpus. Two-stage training builds accurate hand geometry and improves generalization to in-the-wild interactions.

    5,000 hours · two-stage training
  3. 03

    Faster
    inference

    One streaming framework jointly estimates hand locations, hand motion, and camera trajectories. Reusing features and overlapping backbone and head execution delivers 11.19 FPS—more than twice HaWoR’s throughput.

    11.19 FPS · 2.04× HaWoR
03 / METHOD

From local articulation to world-space motion.

InfiniHand pipeline with hand localization, MANO reconstruction, streaming camera estimation, and sparse bundle adjustment.
InfiniHand overview. Stage I learns hand localization and camera-space reconstruction. Stage II jointly estimates hands and camera trajectories with streaming geometric memory; sparse refinement supports long-sequence consistency.
04 / QUALITATIVE RESULTS

Detailed hands, across diverse interactions.

Camera-space comparisons with baselines on in-domain and in-the-wild egocentric videos.
Camera-space reconstruction. InfiniHand recovers hand articulation under substantial occlusion and generalizes to in-the-wild scenes.
Aligned world-space hand trajectories compared with baselines and ground truth.
World-space hand motion. Aligned qualitative comparisons show the articulation and spatial arrangement of successive hand poses.
05 / CITATION

Citation

BibTeX
@misc{ren2026infinihand,
  title  = {InfiniHand: Streaming World-Space Hand Motion Estimation from Egocentric Video},
  author = {Ren, Kerui and Song, Kaiwen and Zhao, Weiguang and Wang, Yuxi and
            Liu, Yufei and Dai, Bo and Guo, Haoyu and Shen, Chunhua and
            Yu, Mulin and Lu, Tao and Dong, Junting},
  year   = {2026},
  eprint = {2609.35743},
  archivePrefix = {arXiv},
  url    = {https://arxiv.org/abs/2609.35743}
}