HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark
Dairu Liu, Zekun Qi, Jiayu Zeng, Ruixi Yu, Yu Guan, Yintianrun Zhang, Xuchuan Chen, Sikai Liang, Zekai Li, Chenghuai Lin, Xinqiang Yu, Wenyao Zhang, He Wang, Li Yi
Nankai University · Tsinghua University · Galbot · Shanghai Jiao Tong University · Peking University · Shanghai Qi Zhi Institute
cs.RO, cs.AI, cs.CV
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: Accepted to ECCV 2026
Code: https://github.com/GalaxyGeneralRobotics/HumanTracker
Project page: https://dairuliu.github.io/humantracker
License: http://creativecommons.org/licenses/by-sa/4.0/
Importance score: 95/100
The gist: HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark Summary This paper introduces HumanTracker, a large-scale benchmark and a preference-aligned metric for evaluating
Terminology
Summary
HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark
Summary
This paper introduces HumanTracker, a large-scale benchmark and a preference-aligned metric for evaluating humanoid motion tracking. The work addresses two major issues in the field: (1) the misalignment between traditional kinematic metrics and human perception of motion quality, and (2) the limited scope and diversity of existing evaluation datasets.
Problem Statement
The paper identifies that Motion tracking looks easy to measure: one can compare the rollout pose to the reference pose and report kinematic errors such as mean joint error and key point error.
However, "we find a clear gap between these numbers and what people see in videos. A rollout may achieve low kinematic error but still look bad, especially under contacts. Feet may slide on the ground, and contacts may break and reattach at the wrong time. The authors argue that
By treating tracking as a per-frame pose matching problem and averaging errors over joints and time, these metrics fail to capture contact, support, balance, and the accumulation of errors in closed-loop control."
Additionally, the paper notes that the most commonly used evaluation suite for humanoid tracking is still an AMASS test set with only 140 sequences,
which lacks diversity and under-represents the long tail of human movement, including challenging contact transitions, asymmetric balancing, and complex recoveries.
HumanTracker Benchmark
The benchmark contains approximately 153 hours of optical motion trajectories from 24 professional performers,
including dance teachers, fitness coaches, tennis coaches and full-time motion-capture actors.
The data is organized into four motion families: Daily (89 hours, 9.7k clips), Highly Dynamic (11 hours, 2.7k clips), Interaction (48 hours, 10.9k clips), and Ground (5 hours, 1.6k clips). The paper states that "The four motion families are defined by the failure regimes that they expose. Daily contains walking, turning and routine gestures, and therefore measures steady-state stability and residual drift under comparatively regular support. Highly Dynamic contains jumps, kicks, acrobatics and fast dance footwork, for which impacts and rapid support switching amplify phase and timing errors. Interaction contains human body motions associated with actions involving objects or the surrounding environment. Ground covers kneeling, sitting, rolling and recovery, where a low centre of mass and multiple simultaneous contacts make the controller sensitive to contact geometry and friction."
Each clip includes a top-level motion-family label, a natural-language description, a fitted SMPL sequence and a robot-space reference trajectory in qpos format.
The dataset is split 9:1 into disjoint training and test partitions, with the family distribution preserved and duplicate motions kept within one partition.
HumanScore Metric
HumanScore is a preference-aligned metric trained on 12K motion pairs containing 24K motions.
The preference data was collected from six doctoral researchers specializing in humanoid robotics
who annotated 6,000 original trajectory pairs,
which were bilaterally mirrored... yielding 12,000 preference records.
The reward model processes a 539-dimensional vector
per frame, containing 70 dimensions for the current reference state and 469 dimensions for simulated state, control, measured contact dynamics, root motion and current keypoint kinematics.
The model uses a Transformer encoder
with masked mean pooling
to form a trajectory representation, and is trained with the Bradley–Terry loss
for strict preferences and a symmetric loss for Similar pairs.
Evaluation Results
The paper evaluates four state-of-the-art trackers: GMT, TWIST2, SONIC, and Humanoid-GPT. Results show that "Humanoid-GPT is the strongest overall tracker, leading most comparisons and all three metrics on Daily and Highly Dynamic. SONIC is the closest competitor. It achieves the highest completion rate on Interaction and the highest HumanScore on Ground."
Preference Alignment
Table 4 shows that HumanScore agrees with human preferences more consistently than any individual analytic diagnostic,
achieving an Align Rate of 0.9083, compared to MPJPE (0.8049), MPJVE (0.8404), KPT Position MAE (0.8405), Foot Contact Accuracy (0.7882), Avg Joint Accel (0.6933), and Avg Joint Jerk (0.7232).
Sensitivity Analysis
The paper finds that removing measured contact features degrades performance most clearly on Ground,
and that Longer context improves alignment by revealing sliding, jitter, drift and recovery that isolated poses cannot capture.
Adding future reference information performs slightly worse than the baseline.
Limitations
The authors acknowledge that "HumanScore is trained from HumanTracker training motions and rollouts produced by four trackers, so its motion-disjoint test measures generalization to unseen motions, not to an entirely unseen robot, simulator or controller family. They also note that
Its 539-dimensional input also includes privileged simulator state and contact quantities that are available for benchmarking but may not be observable on hardware. Furthermore,
Directly optimizing a learned score can create behaviours that exploit model imperfections; using it as a reinforcement-learning reward will require explicit regularization and an independent human evaluation."
Improvements for AI systems
Improvements to AI Systems:
- Contact-Aware Motion Quality Assessment
-
Integrate HumanScore’s preference-aligned reward model into humanoid control policies as a dense reward signal, replacing or augmenting kinematic error terms.
-
The improved AI system can generate motions that prioritize realistic foot contact, balance recovery, and smooth transitions—reducing visible artifacts like foot sliding and premature contact breaks during locomotion, jumping, and interaction tasks.
- Long-Horizon Closed-Loop Error Correction
-
Use the benchmark’s 153-hour diverse motion dataset (including Ground and Interaction families) to train a trajectory-level predictive model that anticipates drift and instability over multi-step horizons.
-
The improved AI system can proactively adjust control commands before errors accumulate, enabling robust performance on asymmetric balancing, sit-to-stand, and recovery from perturbations—beyond what per-frame pose matching achieves.
- Preference-Aligned Reinforcement Learning Regularization
-
Combine HumanScore with explicit regularization (e.g., KL penalty against reference trajectories, contact-consistency constraints) to prevent reward hacking.
-
The improved AI system can optimize for human-perceived quality without exploiting model blind spots, yielding motions that are both numerically accurate and visually natural across unseen motion families.
- Cross-Modal Motion Description and Retrieval
-
Leverage the natural-language descriptions paired with SMPL and robot-space trajectories to train a vision-language-action model for motion generation from text prompts.
-
The improved AI system can translate high-level instructions (e.g., “perform a dynamic kick with a quick recovery”) into physically feasible, contact-correct humanoid motions, expanding applicability in animation, teleoperation, and human-robot collaboration.
- Sim-to-Real Transfer with Privileged-State Distillation
-
Train a teacher policy using the benchmark’s privileged simulator states (539-dim vector) and then distill it into a student policy that relies only on observable keypoints and proprioception.
-
The improved AI system can deploy on physical hardware while retaining HumanScore-aligned motion quality, as it learns to infer contact and balance from vision and joint sensors alone.
- Adaptive Motion Family Recognition
-
Train a classifier on the four motion families (Daily, Highly Dynamic, Interaction, Ground) to dynamically switch control strategies or reward weights during deployment.
-
The improved AI system can automatically adjust its behavior—e.g., prioritizing impact absorption for jumps, contact geometry for ground work, or steady-state stability for walking—leading to superior performance across diverse, unstructured tasks.
Abstract
Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in videos. Kinematic errors average per-frame pose differences but miss the physical artifacts that matter most, particularly unstable support and incorrect contacts such as foot skating and mistimed touch-downs. Meanwhile, widely used test suites are small and lack the diversity needed to stress contact-rich, long-horizon behaviors. We introduce HumanTracker to make humanoid tracking evaluation both perceptually aligned and scalable. The HumanTracker benchmark contains approximately 153 hours of optical motion trajectories from multiple professional performers, organized into four motion families with text labels for fine-grained diagnosis. We further propose HumanScore, a preference-aligned metric trained on 12K motion pairs containing 24K motions. Across representative state-of-the-art trackers, HumanScore better predicts human preferences and reveals contact and stability failures that kinematic metrics often miss.
Sources
- A Cross-Dataset Study for Text-based 3D Human Motion Retrieval
- Switch-JustDance: Benchmarking Whole Body Motion Tracking Controllers Using a Commercial Console Game
- PHUMA: Physically Reliable Humanoid Locomotion Dataset
- GMT: General Motion Tracking for Humanoid Whole-Body Control
- RoboReward: General-Purpose Vision-Language Reward Models for Robotics
- OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning
- CLONE: Closed-Loop Whole-Body Humanoid Teleoperation for Long-Horizon Tasks
- OmniTrack: General Motion Tracking via Physics-Consistent Reference
- SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control
- Robo-Dopamine: General Process Reward Modeling for High-Precision Robotic Manipulation
- KungfuBot: Physics-Based Humanoid Whole-Body Control for Learning Highly-Dynamic Skills
- Collision-Free Humanoid Traversal in Cluttered Indoor Scenes
- OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction
- UniTracker: Learning Universal Whole-Body Motion Tracker for Humanoid Robots
- TWIST2: Scalable, Portable, and Holistic Humanoid Data Collection System
- Motion-X++: A Large-Scale Multimodal 3D Whole-body Human Motion Dataset
- Learning Athletic Humanoid Tennis Skills from Imperfect Human Motion Data
- ResMimic: From General Motion Tracking to Humanoid Whole-body Loco-Manipulation via Residual Learning
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving