HumanoidTTT: Test-Time Capability Reuse for Efficient Humanoid Control

arXiv:2610.00198 · cs.RO · Submitted 2026-09-18 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "HumanoidTTT: Test-Time Capability Reuse for Efficient Humanoid Control".

Dev: HumanoidTTT introduces a framework for test-time capability reuse in continual humanoid control, addressing the challenges of reliable motion reuse under changing robot states and managing validated capabilities within finite storage.

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: We're looking at the paper "HumanoidTTT: Test-Time Capability Reuse for Efficient Humanoid Control," and the authors are Jingtai Yang, Yining Wu, Yanjun Li, Zeyu Zhang, and Hao Tang from Peking University.

Dev: The title itself really tells us they’re focused on capability reuse at test time, which implies a system that has to decide quickly whether to use something it already knows or generate something new.

Taro: It sounds like the core idea is managing how the robot uses its past experiences when it needs to perform a task in the moment.

Rosa: Right, and what this paper seems to be doing is introducing a way for validated motions to be directly reused only if they are applicable given the robot's current state, which is a key distinction.

Dev: I see that the authors are tackling two main challenges: figuring out when a motion can actually replace fresh generation and how to decide which capabilities are worth keeping in the store.

Taro: That split into a read-side applicability problem and a write-side retention problem seems like a smart way to break down such a complex continual learning challenge.

The paper's summary: Rosa: So, HumanoidTTT proposes this framework for test-time capability reuse in continual humanoid control, aiming to make motion reuse reliable even when the robot's state is changing.

Dev: In simple terms, the system authorizes direct reuse of validated complete motions only when the current robot state satisfies specific entry certificates, which means it bypasses generating a fresh motion if that condition is met.

Taro: That conditional reuse based on an applicability certificate sounds like a safety measure to prevent blindly playing back old actions that might not be safe now.

Rosa: Precisely, and on the other hand, the paper also introduces Test-Time Capability Consolidation, which adaptively decides which qualified capabilities should persist in a bounded Full-Motion Store based on how useful they are observed during subsequent deployment reuse.

Dev: So, it’s not just about reusing motions; it's also about learning and pruning the stored capabilities to keep the store efficient as new things come in and old things get less useful.

Taro: That adaptive retention mechanism based on utility feedback sounds like a necessary step for any system trying to manage finite memory while still learning effectively.

The paper's improvements: Rosa: One of the main improvements they highlight is the selective full-motion reuse, where a stored motion only replaces fresh generation when the robot's entry state permits it, which is a big step for reliability.

Dev: That direct reuse path means that instead of going through the Frozen Motion Generator and then qualification checks, accepted motions go straight to execution at a much faster speed.

Taro: The paper also mentions that they separate frequent lightweight management from compute-intensive motion generation by using a heterogeneous CPU–GPU execution path, which sounds like smart engineering for real-time performance.

Rosa: That’s right; the CPU handles the checks and store lookups, keeping the critical GPU path free for when a fresh motion is actually needed.

Dev: And regarding retention, they use an online reinforcement learning process called DoubleDQN to decide whether to skip or replace an existing capability when the store is full, using subsequent deployment reuse utility as its reward signal.

Taro: It’s interesting how they keep the core motion generation and acceptance criteria fixed while only letting the memory management part adapt through that online learning loop.

Conclusion: Rosa: So, to wrap up on "HumanoidTTT: Test-Time Capability Reuse for Efficient Humanoid Control," this paper introduces selective reuse based on state applicability and an adaptive consolidation policy for the store.

Dev: The implication is a significant speedup, showing a sixteen point four times faster end-to-end deployment compared to fresh generation, which is quite substantial for any robot system.

Taro: I think the real impact here is in making continual capability reuse practical by tying the reuse decision directly to current physical feasibility and memory constraints.

Rosa: And we see strong results, including zero unsafe accepts and a notable reduction in median preparation latency down to about twenty-nine point five milliseconds for accepted hits.

Dev: The system manages the trade-off between having a large store of knowledge and keeping that store relevant under deployment demand, which is crucial for long-term robot operation.

Taro: Overall, this work systematically evaluates reliability and efficiency in managing these validated capabilities under finite capacity, setting a solid foundation for future research in this area.

School of Computer Science, Peking University

cs.RO

Submitted: 2026-09-18

Updated: 2026-09-18

Project page: https://aigeeksgroup.github.io/HumanoidTTT

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 89/100

The gist: HumanoidTTT introduces a framework for test-time capability reuse in continual humanoid control, addressing the challenges of reliable motion reuse under changing robot states and managing validated

Key concepts

Selective Full-Motion Reuse
This mechanism lets the system reuse a stored motion only if the robot's current state is in the set of states where that motion is guaranteed to work. If it works, the system skips generating a new motion and uses the stored one directly.
Test-Time Capability Consolidation
When storage is full, this process decides which existing capabilities to keep and which new ones to discard. It learns this by observing how often previously stored motions are actually reused during deployment, optimizing the limited storage space.
Entry Applicability Set
This is a set of robot states that are certified as safe or valid entry points for a specific motion. A motion can only be reliably reused if the robot's current state belongs to this pre-defined set of admissible entry conditions.

Terminology

Summary

HumanoidTTT introduces a framework for test-time capability reuse in continual humanoid control, addressing the challenges of reliable motion reuse under changing robot states and managing validated capabilities within finite storage. The core contribution is enabling direct reuse of qualified motions only when the current state satisfies entry certificates, while adaptively retaining useful capabilities through deployment feedback.

The gist: HumanoidTTT enables reliable and efficient reuse of validated complete motions while adaptively retaining useful capabilities through online consolidation.

Selective Full-Motion Reuse

This mechanism authorizes the direct reuse of validated complete motions only from certified applicable entry states, allowing accepted reuse to bypass fresh generation. Each stored capability consists of a validated complete motion and an applicability certificate defining its Entry Applicability set, which is the set of admissible robot entry states from which the motion can be reliably reused. For each recurring request, a matched capability is directly reused only when the current entry state falls within this set, allowing the stored complete motion to bypass the Frozen Motion Generator and proceed directly to execution; otherwise, it falls back to fresh motion generation followed by execution-based qualification.

Test-Time Capability Consolidation

This process adapts which qualified capabilities persist in a bounded Full-Motion Store using subsequent deployment reuse as feedback. When newly qualified capabilities compete for limited storage, the consolidation policy decides whether to skip the new capability or replace an existing one, and adapts its retention decisions using subsequent reuse utility observed during deployment. The policy is invoked exactly when a qualified miss creates competition for a full Store, expressed by the condition where a qualified request arrives at a full Store. The online network selects either SKIP or REPLACE(j) from the current Store-conditioned state, and the reward is defined as execution-successful Store reuses/observed requests.

System Components and Logic

The framework separates continual capability reuse into two design principles: a read-side applicability problem and a write-side retention problem. The read path involves checking for an exact capability/signature match against the stored capabilities, requiring fixed physical checks and candidate-relative certificate membership, defined by the formula: Accept(i) = Match(qt, di) ∧ Phys(st, Ht) ∧ x(i)t ∈ Ai. If a match is found and the physical criteria are met, an accepted motion is retrieved for direct execution without motion generation.

Kernel-Aware Real-Time Deployment

HumanoidTTT utilizes a heterogeneous CPU–GPU execution path to separate frequent lightweight capability management from compute-intensive motion generation. For accepted reuse requests, Full-Motion Store lookup, Entry Applicability evaluation, complete-motion retrieval, and Store writeback remain on the CPU without activating the motion generator. Capacity-constrained test-time consolidation also runs in CPU FP32, where the compact DoubleDQN is invoked only when a qualified capability competes for a full Store. This event-driven execution keeps online adaptation off the critical GPU generation path and avoids making test-time learning a fixed cost for every request.

Experimental Results

Experiments demonstrate zero unsafe accepts and a 16.4× end-to-end speedup over fresh generation, while online consolidation improves avoided generator calls by 13.2 per 200 requests over its frozen counterpart. HumanoidTTT reduces median preparation latency from 5057.996 ms to 29.514 ms on accepted reuse hits, corresponding to a 171.4× speedup, and achieves a 16.4× end-to-end speedup over fresh generation. Under a capacity-10 Full-Motion Store, online consolidation avoids 13.2 additional generator calls per 200-request stream compared with the same pretrained retention policy with online updates disabled.

Contributions

The main contributions are:

  1. Introducing Selective Full-Motion Reuse, which turns execution-qualified complete motions into persistent capabilities and enables reliable direct reuse across admissible entry states, bypassing fresh motion generation on accepted reuse hits.

  2. Introducing Test-Time Capability Consolidation, which learns which qualified capabilities should persist in a bounded Full-Motion Store based on subsequent reuse utility observed during deployment, while keeping motion generation and acceptance criteria fixed and freezing each applicability certificate after acquisition.

  3. Systematically evaluating HumanoidTTT in terms of full-motion reuse reliability, accepted-hit efficiency, end-to-end inference efficiency, and finite-capacity capability retention.

Limitations

HumanoidTTT currently freezes each Entry Applicability certificate after capability acquisition, leaving persistent changes in robot dynamics, environmental conditions, or admissible entry regions outside the current adaptation scope. Future work will extend Entry Applicability to update from deployment experience and evaluate the resulting capability reuse and finite-capacity consolidation over longer and more diverse real-world deployments.

Improvements for AI systems

Here are specific improvements for existing AI systems based on the HumanoidTTT framework, detailing what these improved systems can achieve:


  1. The ability to execute previously validated full-motion sequences reliably across changing robot states without requiring a complete motion regeneration pipeline.

  2. The capability to serve recurring motion requests with a significant reduction in end-to-end latency (up to 16.4× faster than fresh generation) by bypassing the Motion Generator entirely on accepted reuse hits, leading to near-instantaneous execution of known, safe motions (median preparation latency drop from 5058ms to 29.5ms).

  3. The implementation of a dynamic, data-driven memory management system (Test-Time Capability Consolidation) that learns which previously executed motions are most useful for future requests and adaptively retains them in a bounded store, optimizing storage utilization while maximizing the potential for fast reuse.

  4. The creation of a robust safety mechanism that guarantees zero unsafe accepts during reuse, ensuring that only motions whose execution is certified to be safe from the current robot entry state are played back, effectively eliminating the risk of blindly replaying potentially unsafe previous behaviors.

  5. The enhancement of real-time inference efficiency by employing a heterogeneous CPU–GPU execution path: lightweight capability management and safety checks run on the CPU (FP32), while only compute-intensive motion generation is dispatched to the GPU (via ONNX Runtime CUDA) when a fresh motion is required, drastically reducing critical-path computation for recurrent requests.

  6. The development of an adaptive, frozen certificate system where the entry applicability set for each stored motion is fixed after acquisition but can be leveraged by subsequent deployment experience to improve future retention decisions via online reinforcement learning (Double-DQN), allowing the system to learn which capabilities are worth keeping without risking the correctness of the motion generation or acceptance criteria themselves.

Abstract

Recent advances in motion generation and whole-body tracking have enabled humanoid robots to execute increasingly diverse motions, yet the same motion capabilities may be requested repeatedly during continual deployment. Reliable reuse is challenging because intervening motions can change the robot's entry state, making previously successful motions unsafe to replay blindly. Meanwhile, validated capabilities accumulate during deployment, while bounded storage requires deciding which ones are worth retaining. To address these challenges, we present HumanoidTTT, a framework for test-time capability reuse in continual humanoid control. Specifically, we introduce Selective Full-Motion Reuse, which authorizes direct reuse of validated complete motions only from certified applicable entry states, allowing accepted reuse to bypass fresh generation. We further introduce Test-Time Capability Consolidation, which adapts which qualified capabilities persist in a bounded Full-Motion Store using subsequent deployment reuse as feedback. Experiments demonstrate zero unsafe accepts and a 16.4 times end-to-end speedup over fresh generation, while online consolidation improves avoided generator calls by 13.2 per 200 requests over its frozen counterpart. Overall, HumanoidTTT enables reliable and efficient reuse of validated motion capabilities while adaptively retaining useful capabilities throughout continual deployment. Code: https://github.com/AIGeeksGroup/HumanoidTTT. Website: https://aigeeksgroup.github.io/HumanoidTTT.

Sources

Related papers