YOCO: You Only Calibrate Once! Fast Mocap Calibration for Dexterous Teleoperation
summary
The gist
The gist: YOCO presents a fast few-shot, fine-tuning-free calibration framework that corrects biased hand-pose streams from a small set of paired raw and target poses to improve downstream dexterous
In short
YOCO introduces a fast calibration framework that corrects biases in low-cost hand-pose streams using few-shot learning. It conditions a network on paired raw and target poses to predict LoRA updates for a frozen hand estimation module. This method improves downstream dexterous teleoperation performance significantly, achieving high success rates by inferring session-specific noise from minimal calibration data.
Key concepts
- Few-shot Calibration
- Instead of training a separate model for every user or session, YOCO uses a small set of paired raw and target poses to quickly learn how to correct biases. It treats the calibration process as generating 'low-rank' updates that specialize a main estimation module, allowing it to adapt rapidly without extensive retraining.
- Hypernetwork
- A Hypernetwork is a network whose weights are generated by another network. In YOCO, the calibration HyperNetwork takes the paired poses as input and outputs these low-rank LoRA matrices. This allows the system to dynamically generate session-specific correction parameters based on the observed calibration examples.
- LoRA Updates
- LoRA (Low-Rank Adaptation) updates are small, efficient adjustments made to a frozen neural network. YOCO predicts these updates from the calibration set, injecting them into the main hand estimation module. This allows for session-specific tuning of the model's behavior without needing to fine-tune or optimize the entire large model.
- MANO Hand Estimation Module
- This is a frozen component of YOCO that maps 3D hand keypoints to MANO pose parameters. The calibration process modifies this module using LoRA updates, rather than retraining it from scratch. This separation allows the system to focus solely on learning the bias correction while keeping the core shape estimation stable.
Terminology used across episodes
This episode discusses
- YOCO: You Only Calibrate Once! Fast Mocap Calibration for Dexterous Teleoperation · Paper Radio
- FSGlove: An Inertial-Based Hand Tracking System with Shape-Aware Calibration
- Transformer IMU Calibrator: Dynamic On-body IMU Calibration for Inertial Motion Capture
- Human-Exoskeleton Kinematic Calibration to Improve Hand Tracking for Dexterous Teleoperation
- MediaPipe Hands: On-device Real-time Hand Tracking
- InterHand2.6M: A Dataset and Baseline for 3D Interacting Hand Pose Estimation from a Single RGB Image
- Learning Visuotactile Skills with Two Multifingered Hands
- OPEN TEACH: A Versatile Teleoperation System for Robotic Manipulation
- Learning to Transfer Human Hand Skills for Robot Manipulations
- Geometric Retargeting: A Principled, Ultrafast Neural Hand Retargeting Algorithm
- RAPID Hand: A Robust, Affordable, Perception-Integrated, Dexterous Manipulation Platform for Generalist Robot Autonomy
- DEXOP: A Device for Robotic Transfer of Dexterous Human Manipulation
- Hyper-GoalNet: Goal-Conditioned Manipulation Policy Learning with HyperNetworks
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- Scalable Diffusion Models with Transformers
The paper
YOCO: You Only Calibrate Once! Fast Mocap Calibration for Dexterous Teleoperation · Read on arXiv
Yu Zhang, Yunqi Li, Yushi Du, Yi Ma, Yanchao Yang
The University of Hong Kong
Dexterous teleoperation requires reliable human-hand state estimations. However, common low-cost motion-capture gloves and markerless trackers often exhibit biases that vary across users, glove fit, and recording sessions, degrading retargeting and demonstration quality. We present YOCO, a fast few-shot, fine-tuning-free calibration framework that corrects biased hand-pose streams from a small set of paired raw and target poses. Instead of optimizing a separate model for every operator or session, YOCO conditions a calibration HyperNet on the paired examples and predicts LoRA-style updates for a frozen MANO hand-estimation module, turning per-user calibration into a lightweight feed-forward adaptation step while preserving the geometric prior of MANO and the efficiency of a compact estimator. We train YOCO with synthetic drift augmentations on InterHand2.6M and evaluate on augmented InterHand sequences, offline real glove data, and dexterous teleoperation tasks. Across these settings, YOCO improves calibration efficiency, hand-state estimation quality and teleoperation performance compared with uncalibrated input and standard calibration baselines.
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "YOCO: You Only Calibrate Once! Fast Mocap Calibration for Dexterous Teleoperation".
Rosa: The gist: YOCO presents a fast few-shot, fine-tuning-free calibration framework that corrects biased hand-pose streams from a small set of paired raw and target poses to improve downstream dexterous teleoperation.
Dev: First, who's behind it and why it matters.
Paper summary: Rosa: So we're looking at YOCO: You Only Calibrate Once! Fast Mocap Calibration for Dexterous Teleoperation, and the main thing here is tackling those biases in low-cost hand trackers that change from user to user and even from one recording session to another.
Dev: Yeah, it claims they have a fast few-shot calibration framework that corrects these biased hand-pose streams using just a small set of paired raw and target poses to improve dexterity in teleoperation.
Taro: So, the core idea seems to be moving away from traditional methods that learn glovesignal-to-pose mappings or correct IMU alignment, and instead focusing on calibrating the derived three dee hand-keypoint interface after a tracker has already produced its pose estimate <ref:2610.11657#pg1>.
Rosa: Right, so they formulate this few-shot calibration by conditioning a HyperNet on those paired examples to predict LoRA updates for a frozen MANO handestimation module. It's designed to be session-specific without needing per-session gradient optimization.
Dev: That sounds efficient, but I want to know how it actually runs in terms of the loop rate and latency since we're talking about real-time control systems here.
Taro: The formulation uses a Target network that maps the twenty-one three dee hand keypoints to MANO pose parameters, and its parameters are trained once on clean data and then kept fixed <ref:2610.11657#pg3>. Then the Hypernetwork takes the calibration set as input and outputs a set of low-rank LoRA matrices that get injected into that frozen target network.
Rosa: So, the learning objective is to minimize the expected joint-position error of this resulting calibrated estimator, encouraging it to infer mocap bias from those paired examples and produce updates that correct unseen noisy poses drawn from the same distribution.
Dev: They use synthetic biases in the MANO pose space—things like per-finger zero offsets, perjoint residual offsets, and global or finger gains—to train the hypernetwork to compensate for realistic mocap bias. That sounds like a good way to test robustness outside of just clean data.
Taro: The evaluation shows that YOCO substantially reduces hand-pose error on biased mocap streams relative to uncalibrated input, reducing the PA-MPJPE from sixteen point five five mm down to ten point four five mm on the augmented InterHand2 point 6M evaluation set, with the largest improvement seen on fingertip-sensitive metrics like Tip-PA dropping from thirty-seven point one five mm to twenty-three point two nine mm.
Rosa: And it's not just about accuracy; they claim YOCO outperforms optimization- and fine-tuning-based baselines, and that this advantage stays even when they limit the LoRA capacity across matched setups.
Dev: I see, so the speed aspect is important because deployment latency matters a lot in teleoperation. The paper mentions that on CPU, YOCO requires thirteen point one six seconds for calibration before real deployment, compared to maybe thirty-four to over two hundred seconds for gradient fine-tuning baselines.
Paper summary: Taro: That difference in time is significant because it means the calibration itself isn't taking a huge chunk of time on the machine that's running the robot control loop. Also, they found that YOCO achieves an overall success rate of ninety percent in teleoperation tasks, compared to sixty-five percent for LoRA-Finetune and thirty-five percent for the uncalibrated baseline.
Rosa: So, what does this mean for someone who just listens to the show? It suggests that even though we use low-cost trackers that are inherently biased, we can get a much more reliable hand state estimation by doing this one-time calibration using just a few paired poses.
Dev: It means we don't have to constantly re-optimize the entire model for every single user or every session; instead, this HyperNet generates updates on the fly based on that small calibration set. That cuts down on computation significantly.
Taro: The authors also pointed out a limitation: they focus heavily on pose calibration and leave the MANO shape estimation to a separate stage, and while the assumption is that bias is stationary across sessions, high-frequency drift on less stable hardware might need re-calibration or a temporal extension of the HyperNet.
Rosa: Exactly, so it's a strong tool for correcting session-stationary bias with minimal upfront work, but we still have to consider those cases where the tracking gets unstable over time.
Dev: We also see that the gains are consistent across different users in their dataset rather than being driven by just one specific operator. That suggests this method offers a more general improvement in reliability for the whole class of low-cost glove data they tested.
Taro: The results on contact-sensitive tasks were particularly telling, like Earbud Pinch Grasping rising from a one/ten success rate with LoRA-Finetune to an eight/ten success rate with YOCO, and Tissue Box Opening going from eight/ten to a full ten/ten.
Rosa: So the paper, YOCO: You Only Calibrate Once! Fast Mocap Calibration for Dexterous Teleoperation, presents a method where you use a small set of paired poses to condition a HyperNet that predicts LoRA updates for a frozen MANO estimator to correct hand pose biases quickly and robustly.
Dev: It's about coupling that sensor-independent pose interface with amortized, session-conditioned weight generation without needing per-session gradient optimization, which should be very appealing for deployment on real hardware.
Taro: The implication is that even with imperfect low-cost tracking, you can achieve much better dexterity in teleoperation tasks when you apply this calibration to correct the errors in the hand state estimation.
Rosa: We'll keep an eye on how it holds up outside of that single tensile glove evaluation, but for now, YOCO looks like a solid way to get those measurable gains on downstream manipulation quality and success rates.
Conclusion: Rosa: So we've looked at how YOCO uses just a few paired poses to correct those hand pose biases from low-cost trackers to make teleoperation better for dexterous tasks.
Dev: Yeah, it's about this hypernetwork generating updates for a frozen model instead of retraining the whole thing every time you get new data.
Rosa: The title itself, "You Only Calibrate Once!", really captures that idea of doing the heavy lifting upfront and then using it efficiently for everything else in a session.
Dev: And the authors are focused on making this calibration method fast enough to actually be useful in a real-time control loop, which is where my brain immediately goes.
Rosa: They're showing how this approach beats optimization and fine-tuning methods when it comes to accuracy on these specific contact tasks, like pinching things or opening boxes.
Dev: The numbers they give are pretty compelling, especially that jump in success rates from about sixty-five percent for other methods to ninety percent with YOCO.
Rosa: It means if you're building a system using cheaper tracking hardware, this calibration step is something you can do once and then rely on for much better performance down the line.
Dev: But we have to keep an eye on those caveats they mentioned regarding high-frequency drift, because that could be a problem if the hardware isn't super stable.
Rosa: Exactly, it's a solid tool for getting immediate gains in manipulation quality with less setup time, but we still have to figure out how long it holds up in the real world without constant re-calibration.
Dev: So what this means for you listening right now is that if you're working on a robotic system where hand tracking is noisy, there might be a way to get better results without spending all your time optimizing and retuning models for every single new recording session.
More episodes
- 2610.10846-Cross-Embodiment Robot Foundation World Models with Latent Actions
- 2610.10601-Teaching a Robot Dog New Tricks: Diverse Quadruped Skills via Combined Reinforcement and Imitation Learning with Adversarial Task Selection
- 2610.10637-TacHair: Tactile Contact-Distribution Guided Online Correction for Robotic Hair Stroking and Perception
- 2610.10646-Masked Generative Motion Planning with Geometry-Guided Token Search
- 2610.10812-Skill-SLM: Agent Skill-driven Small Language Models for Reliable Robot Operation
- 2610.10801-Same Action, Different Outcome: Variability in Dynamic Cloth Manipulation
- 2610.10810-Diagnosing and Recovering from Observation-Space Shift at Long-Horizon Skill Seams
- 2610.10748-TAPNAV: Humanoid Navigation through Tactile Active Perception
- 2610.10855-OmniHOI: Dexterous Hand-Object Interaction from Monocular Human Video
- 2610.11003-ActiveReg: Information-Driven Active Regional Probing for Partial-to-Full Bone Registration