HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark

arXiv:2608.13555 · cs.RO, cs.AI, cs.CV · Submitted 2026-08-13 · Read on arXiv

Dairu Liu, Zekun Qi, Jiayu Zeng, Ruixi Yu, Yu Guan, Yintianrun Zhang, Xuchuan Chen, Sikai Liang, Zekai Li, Chenghuai Lin, Xinqiang Yu, Wenyao Zhang, He Wang, Li Yi

Nankai University · Tsinghua University · Galbot · Shanghai Jiao Tong University · Peking University · Shanghai Qi Zhi Institute

cs.RO, cs.AI, cs.CV

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: Accepted to ECCV 2026

Code: https://github.com/GalaxyGeneralRobotics/HumanTracker

Project page: https://dairuliu.github.io/humantracker

License: http://creativecommons.org/licenses/by-sa/4.0/

Importance score: 95/100

The gist: HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark Summary This paper introduces HumanTracker, a large-scale benchmark and a preference-aligned metric for evaluating

Terminology

Summary

HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark

Summary

This paper introduces HumanTracker, a large-scale benchmark and a preference-aligned metric for evaluating humanoid motion tracking. The work addresses two major issues in the field: (1) the misalignment between traditional kinematic metrics and human perception of motion quality, and (2) the limited scope and diversity of existing evaluation datasets.

Problem Statement

The paper identifies that Motion tracking looks easy to measure: one can compare the rollout pose to the reference pose and report kinematic errors such as mean joint error and key point error. However, "we find a clear gap between these numbers and what people see in videos. A rollout may achieve low kinematic error but still look bad, especially under contacts. Feet may slide on the ground, and contacts may break and reattach at the wrong time. The authors argue that By treating tracking as a per-frame pose matching problem and averaging errors over joints and time, these metrics fail to capture contact, support, balance, and the accumulation of errors in closed-loop control."

Additionally, the paper notes that the most commonly used evaluation suite for humanoid tracking is still an AMASS test set with only 140 sequences, which lacks diversity and under-represents the long tail of human movement, including challenging contact transitions, asymmetric balancing, and complex recoveries.

HumanTracker Benchmark

The benchmark contains approximately 153 hours of optical motion trajectories from 24 professional performers, including dance teachers, fitness coaches, tennis coaches and full-time motion-capture actors. The data is organized into four motion families: Daily (89 hours, 9.7k clips), Highly Dynamic (11 hours, 2.7k clips), Interaction (48 hours, 10.9k clips), and Ground (5 hours, 1.6k clips). The paper states that "The four motion families are defined by the failure regimes that they expose. Daily contains walking, turning and routine gestures, and therefore measures steady-state stability and residual drift under comparatively regular support. Highly Dynamic contains jumps, kicks, acrobatics and fast dance footwork, for which impacts and rapid support switching amplify phase and timing errors. Interaction contains human body motions associated with actions involving objects or the surrounding environment. Ground covers kneeling, sitting, rolling and recovery, where a low centre of mass and multiple simultaneous contacts make the controller sensitive to contact geometry and friction."

Each clip includes a top-level motion-family label, a natural-language description, a fitted SMPL sequence and a robot-space reference trajectory in qpos format. The dataset is split 9:1 into disjoint training and test partitions, with the family distribution preserved and duplicate motions kept within one partition.

HumanScore Metric

HumanScore is a preference-aligned metric trained on 12K motion pairs containing 24K motions. The preference data was collected from six doctoral researchers specializing in humanoid robotics who annotated 6,000 original trajectory pairs, which were bilaterally mirrored... yielding 12,000 preference records.

The reward model processes a 539-dimensional vector per frame, containing 70 dimensions for the current reference state and 469 dimensions for simulated state, control, measured contact dynamics, root motion and current keypoint kinematics. The model uses a Transformer encoder with masked mean pooling to form a trajectory representation, and is trained with the Bradley–Terry loss for strict preferences and a symmetric loss for Similar pairs.

Evaluation Results

The paper evaluates four state-of-the-art trackers: GMT, TWIST2, SONIC, and Humanoid-GPT. Results show that "Humanoid-GPT is the strongest overall tracker, leading most comparisons and all three metrics on Daily and Highly Dynamic. SONIC is the closest competitor. It achieves the highest completion rate on Interaction and the highest HumanScore on Ground."

Preference Alignment

Table 4 shows that HumanScore agrees with human preferences more consistently than any individual analytic diagnostic, achieving an Align Rate of 0.9083, compared to MPJPE (0.8049), MPJVE (0.8404), KPT Position MAE (0.8405), Foot Contact Accuracy (0.7882), Avg Joint Accel (0.6933), and Avg Joint Jerk (0.7232).

Sensitivity Analysis

The paper finds that removing measured contact features degrades performance most clearly on Ground, and that Longer context improves alignment by revealing sliding, jitter, drift and recovery that isolated poses cannot capture. Adding future reference information performs slightly worse than the baseline.

Limitations

The authors acknowledge that "HumanScore is trained from HumanTracker training motions and rollouts produced by four trackers, so its motion-disjoint test measures generalization to unseen motions, not to an entirely unseen robot, simulator or controller family. They also note that Its 539-dimensional input also includes privileged simulator state and contact quantities that are available for benchmarking but may not be observable on hardware. Furthermore, Directly optimizing a learned score can create behaviours that exploit model imperfections; using it as a reinforcement-learning reward will require explicit regularization and an independent human evaluation."

Improvements for AI systems

Improvements to AI Systems:

  1. Contact-Aware Motion Quality Assessment
  • Integrate HumanScore’s preference-aligned reward model into humanoid control policies as a dense reward signal, replacing or augmenting kinematic error terms.

  • The improved AI system can generate motions that prioritize realistic foot contact, balance recovery, and smooth transitions—reducing visible artifacts like foot sliding and premature contact breaks during locomotion, jumping, and interaction tasks.

  1. Long-Horizon Closed-Loop Error Correction
  • Use the benchmark’s 153-hour diverse motion dataset (including Ground and Interaction families) to train a trajectory-level predictive model that anticipates drift and instability over multi-step horizons.

  • The improved AI system can proactively adjust control commands before errors accumulate, enabling robust performance on asymmetric balancing, sit-to-stand, and recovery from perturbations—beyond what per-frame pose matching achieves.

  1. Preference-Aligned Reinforcement Learning Regularization
  • Combine HumanScore with explicit regularization (e.g., KL penalty against reference trajectories, contact-consistency constraints) to prevent reward hacking.

  • The improved AI system can optimize for human-perceived quality without exploiting model blind spots, yielding motions that are both numerically accurate and visually natural across unseen motion families.

  1. Cross-Modal Motion Description and Retrieval
  • Leverage the natural-language descriptions paired with SMPL and robot-space trajectories to train a vision-language-action model for motion generation from text prompts.

  • The improved AI system can translate high-level instructions (e.g., “perform a dynamic kick with a quick recovery”) into physically feasible, contact-correct humanoid motions, expanding applicability in animation, teleoperation, and human-robot collaboration.

  1. Sim-to-Real Transfer with Privileged-State Distillation
  • Train a teacher policy using the benchmark’s privileged simulator states (539-dim vector) and then distill it into a student policy that relies only on observable keypoints and proprioception.

  • The improved AI system can deploy on physical hardware while retaining HumanScore-aligned motion quality, as it learns to infer contact and balance from vision and joint sensors alone.

  1. Adaptive Motion Family Recognition
  • Train a classifier on the four motion families (Daily, Highly Dynamic, Interaction, Ground) to dynamically switch control strategies or reward weights during deployment.

  • The improved AI system can automatically adjust its behavior—e.g., prioritizing impact absorption for jumps, contact geometry for ground work, or steady-state stability for walking—leading to superior performance across diverse, unstructured tasks.

Abstract

Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in videos. Kinematic errors average per-frame pose differences but miss the physical artifacts that matter most, particularly unstable support and incorrect contacts such as foot skating and mistimed touch-downs. Meanwhile, widely used test suites are small and lack the diversity needed to stress contact-rich, long-horizon behaviors. We introduce HumanTracker to make humanoid tracking evaluation both perceptually aligned and scalable. The HumanTracker benchmark contains approximately 153 hours of optical motion trajectories from multiple professional performers, organized into four motion families with text labels for fine-grained diagnosis. We further propose HumanScore, a preference-aligned metric trained on 12K motion pairs containing 24K motions. Across representative state-of-the-art trackers, HumanScore better predicts human preferences and reveals contact and stability failures that kinematic metrics often miss.

Sources

Related papers