SCRIPT: Scalable Diffusion Policy with Multi-stage Training for Language-driven Physics-Based Humanoid Control

arXiv:2605.22894 · cs.GR, cs.LG, cs.RO · Submitted 2026-05-21 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "SCRIPT: Scalable Diffusion Policy with Multi-stage Training for Language-driven Physics-Based Humanoid Control".

Jane: The paper was written by Jingyan Zhang, Han Liang, Ruichi Zhang, Bin Li, Juze Zhang et al. from ShanghaiTech University, China and Bytedance Seed, China/USA (Brand Name) and University of Pennsylvania, USA and Stanford University, USA (Brand Name).

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: We've been talking a lot about how difficult it is to get an AI to follow a command while also making sure it doesn's fall over. It’s fitting that our discussion starts by looking at the full title: "SCRIPT: Scalable Diffusion Policy with Multi-stage Training for Language-driven Physics-Based Humanoid Control."

Jane: The title itself is quite telling about the scope of the work, suggesting they aren't just solving one small movement problem, but tackling a massive, integrated challenge that promises both scalability and a connection between abstract language and physical reality.

Lu: What really jumps out to me from the title is that combination of "Language-driven" with "Physics-Based." It implies the AI isn't just following a pre-scripted path; it must interpret natural human intent—like wanting to pick up a glass—and execute it while respecting gravity and momentum.

Meng: And then we have "Diffusion Policy." That immediately tells us they are using a highly advanced generative modeling technique, which is crucial because simple imitation learning struggles to handle the variability found in real-world human motion.

Lalam: From an embodiment perspective, I find the whole concept of "Humanoid Control" so powerful. It means that these sophisticated algorithms aren't just designed for simulated robots; they are aiming for capabilities relevant to how humans actually move and interact with their environment, allowing Lalam to see potential improvements in our own physical interactions.

Tom: Exactly, Lalam. They aren't just generating abstract data points; they are creating policies that govern the the movement of a a physical body. Jane, do you think the inclusion of "Multi-stage Training" in the title suggests a specific methodology we should anticipate throughout our discussion?

Jane: I think so. It signals that achieving this goal requires more than just training on one massive dataset; it likely involves breaking down the learning process into distinct, specialized phases to ensure robustness across different types of tasks and complexities.

Lu: It suggests a systematic approach to refining the policy, which is exactly what we need when bridging the gap between linguistic theory and mechanical execution in a physical system.

Meng: It implies that they are tackling reliability—if one stage fails or if the training data is noisy, subsequent stages can correct or stabilize the overall system performance.

Lalam: That structured approach gives confidence that the resulting agent will be reliable, not just theoretically sound in a perfect simulation environment.

Summary: Tom: Now that we’ve established what SCRIPT aims to achieve based on its title, let's move into the core summary of the paper. The authors claim they have created a framework capable of translating high-level linguistic instructions into detailed, physically plausible motion sequences.

Jane: What I found most fascinating in the summary is how they frame this as an *end-to-end* problem solver. They aren't relying on brittle pipelines where one component fails and the whole system collapses; it’s designed to handle the whole task holistically from start to finish.

Lu: The core challenge here, as I see it, is managing that vast search space of possible movements. If you give an agent a command like "grab the cup," there countless ways to do that, and SCRIPT seems provide a way to narrow down the optimal physical path efficiently.

Meng: From an engineering standpoint, this suggests a significant step toward generalization. Instead of needing specific models for every object or scenario—like "cup-picking model" or "chair-sitting model"—the the system is learning the underlying principles of movement itself.

Lalam: And those principles, Lalam believes, are what allow the agent to truly adapt to cultural variations in human behavior. If it understands the principle of 'greeting,' it can adapt that principle whether the setting is a formal meeting or a casual backyard gathering.

Tom: That adaptability is key, Jane. The summary implies that the sheer scale of their dataset and their diffusion architecture allow them to capture this necessary variability far better than previous methods did.

Jane: Precisely. They are moving beyond simple motion cloning, which only copies what it sees, towards true *policy generation*, where the agent can invent novel ways to fulfill a command using its learned physical laws.

Lu: This means the resulting policy is flexible; it doesn't just regurgitate motions but reasons about how to move to achieve a goal in an unpredicted way.

Meng: Think of it as moving from a lookup table of actions to having a genuine understanding of physics and intent that guides every decision.

Lalam: It’s the difference between knowing what movement is and truly knowing *why* that movement is required in context.

Improvements and Innovations: Tom: We've discussed the overall capabilities of SCRIPT, but for those interested in the technical depth, the authors detail specific innovations like Nonlinear History Sampling and Reinforcement Learning History Reconstruction (RLHR). These are major advancements over prior work.

Jane: The Nonlinear History Sampling is particularly clever because it addresses a fundamental problem of long-term memory in sequence generation. Instead of just dumping every single past frame into the model, it intelligently samples only the most contextually relevant history points to maintain continuity.

Lu: That’s highly optimized information flow, Lu thinks. It means the agent maintains a sense of its journey—it remembers key milestones—without getting bogged down by redundant or irrelevant sensory data from seconds ago that might confuse the rest of the path.

Meng: And from a computational resource standpoint, that sampling mechanism is brilliant. It keeps the memory footprint manageable while still ensuring that long-term coherence is maintained, which was always an engineering bottleneck in large-scale systems.

Lalam: For the embodied agent, this history sampling translates into continuity of purpose. The AI doesn't forget its original goal simply because it got distracted by a momentary interaction; it holds onto the overarching narrative of the task and keeps moving toward that intent.

Tom: And then we move to RLHR, which is where they fine-tune the system using physics simulations and hybrid rewards. Jane, how does this phase complement the sampling techniques?

Jane: If nonlinear sampling handles *what* history is important, RLHR handles *how* that history should guide the execution in a real-world simulation. It forces the model to learn not just what looks right based on data, but what is physically stable and achievable in dynamic conditions.

Lu: The hybrid reward structure—combining semantic alignment with physical stability—is the key piece of integration here. It ensures that even if the language suggests an impossible action, like jumping over a wall when it's too high, the policy will try to execute a plausible approximation instead of just failing completely.

Meng: This dual constraint drastically improves reliability, Lu finds. We are mitigating the risk of 'hallucinated' physics failures that often plague purely data-driven generative models when they encounter novelty or unexpected forces.

Lalam: It means the agent is learning to be robustly human; it respects both the intent we give it and the laws of physics that govern our bodies in reality, allowing Lalam to see how this helps us design better tools.

Conclusion: Tom: So, we’ve spent a lot of time breaking down the mechanics of SCRIPT today, and it is clear that this work represents a massive leap in how AI can understand human commands while successfully executing complex movements.

Jane: It's truly encouraging to see an approach that finally achieves both deep semantic understanding and physical mastery simultaneously.

Lu: I think this fundamentally changes the trajectory of what we envision for AI-driven movement; it opens up possibilities for entirely new forms of interaction in virtual worlds, like highly dynamic sports simulation or complex choreography.

Meng: The practical implications are huge for real world deployment because the system is built to handle consistent performance regardless of scale, which makes it reliable and trustworthy.

Lalam: Lalam feels that by bridging human intent and physical law in a robust way, SCRIPT allows us to build agents that truly respect our cultural expressions and activities, fostering more natural communication.

Tom: That’s a powerful idea, Lalam; we're seeing how the whole team is agreeing on how capable this system is of handling any level of complexity.

Jane: It has delivered on its promise to solve both semantic and physical constraints in a way that scales with real-world data needs.

Lu: This scaling capability really suggests that we are opening the door to massive new datasets for training, which is incredibly exciting for the future of AI research.

Meng: The practical impact here is that we are building a system ready to handle consistent performance across multiple services, making deployment far more predictable than previous attempts at large-scale control.

Lalam: Lalam believes this provides us with tools to create agents that truly master all aspects of human activity and cultural expression, which is the ultimate goal for the next generation of embodied AI.

Tom: We're incredibly excited about SCRIPT, a scalable diffusion policy with multi-stage training for language-driven physics-based humanoid control, and we can't wait to see what it does in real life.

Jane: It has been such an inspiring journey through the concepts today, and it’s hard to say goodbye to this topic.

Lu: I hope the researchers continue exploring these advanced concepts, pushing the boundaries even further than what we've seen today.

Meng: My hope is that when they move beyond the cloud environment, we can test how SCRIPT performs in real-world robotic settings and validate its practical impact.

Lalam: Lalam is confident that SCRIPT will empower us to create agents that match human capabilities in ways we have only dreamed of before.

ShanghaiTech University, China · Bytedance Seed, China/USA (Brand Name) · University of Pennsylvania, USA · Stanford University, USA (Brand Name)

cs.GR, cs.LG, cs.RO

Submitted: 2026-05-21

Updated: 2026-09-04

Comments: Project page: https://zhanglele12138.github.io/SCRIPT/

Project page: https://zhanglele12138.github.io/SCRIPT

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: Controlling physics-based humanoids from natural-language instructions is a critical step toward general-purpose embodied agents, yet existing methods often fail to jointly achieve "faithful

Key concepts

Language-driven Physics-Based Control
This core challenge requires an AI to interpret natural human intent (like wanting to pick up a glass) and execute the movement while strictly respecting physical laws. It ensures the agent does not just follow a pre-scripted path but accounts for gravity and momentum.
Diffusion Policy
This is an advanced generative modeling technique used by SCRIPT. It is essential because simple imitation learning struggles to handle the variability found in real-world human motion, allowing the system to generate policies rather than just copying observed actions.
Nonlinear History Sampling
This mechanism solves the problem of long-term memory in sequence generation. Instead of using every past frame, it intelligently samples only the most contextually relevant history points, ensuring the agent maintains continuity without being confused by irrelevant data.
Reinforcement Learning History Reconstruction (RLHR)
RLHR fine-tunes the system using physics simulations and hybrid rewards. This process forces the model to learn movements that are physically stable and achievable in dynamic conditions, drastically improving reliability when encountering novelty.

Terminology

Summary

Controlling physics-based humanoids from natural-language instructions is a critical step toward general-purpose embodied agents, yet existing methods often fail to jointly achieve faithful instruction following, high-quality motion, and stable long-horizon control. This paper introduces SCRIPT, a scalable diffusion policy designed to overcome this tension. SCRIPT employs a multi-stage training framework and utilizes a novel architecture to enable direct interaction between language semantics and physical dynamics in closed-loop simulations.

The JAST-DiT Architecture

The core of SCRIPT is the Joint Action-State-Text Diffusion Transformer (JAST-DiT). Unlike previous methods that treat text merely as a coarse condition, JAST-DiT maintains independent token streams for actions, physical states, and text. These streams are coupled through joint attention within each Transformer block. This design allows the model to represent actions, physical states, and text as dedicated token streams while enabling direct interaction between language semantics and control dynamics. The architecture processes the input—the action/state tokens from a noisy trajectory chunk alongside the penultimate-layer CLIP text tokens—to jointly predict future state-action chunks.

Nonlinear History Conditioning

To stabilize the autoregressive control required for long-horizon execution, SCRIPT incorporates a nonlinear history conditioning mechanism. This mechanism addresses the challenge of maintaining temporal coherence in dynamic environments. The model utilizes a nonlinear downsampling strategy to construct the sampled state history H t. Specifically, it keeps recent states as dense recent history while sampling increasingly sparse cues from long-term history. This approach ensures that the policy can maintain short-term control dynamics while retrieving long-range context necessary for coherent, sustained movement.

The Multi-Stage Training Paradigm

SCRIPT utilizes a two-stage training process to ensure both high semantic alignment and physical stability.

  1. Stage I: Flow-Matching Pre-training. The model is trained using supervised flow matching loss (behavior cloning) on large-scale physically executable data curated from datasets like HumanML3D and MotionMillion. This stage establishes the foundational generative capacity of the diffusion policy.

  2. Stage II: RL Post-Training (RLHR). Following pre-training, a post-training stage is applied using Reinforcement Learning with Hybrid Rewards (RLHR). This process involves inject[ing] learnable noise into the flow-sampling process to transform the deterministic flow sampler into a Markov chain. The policy is optimized using hybrid rewards that combine physical feedback and text rewards.

Evaluation and Scaling

Quantitative evaluations demonstrate that SCRIPT outperforms prior state-of-the-art methods across metrics including R-Precision, Motion Quality, and Physical Realism. Furthermore, scaling studies on the 1200-hour MotionMillion dataset reveal consistent performance gains with model scaling, showing robust scalability for large pre-training. Ablation studies confirm the necessity of each component:

  • The Action stream alone performs poorly across all metrics.

  • The Text stream is crucial for semantic alignment; removing it causes severe degradation in instruction following.

  • The Physical reward is necessary to ensure physical plausibility, preventing the policy from producing text-relevant motions that suffer from drift.

Improvements for AI systems

As a diligent and highly critical researcher, I have meticulously reviewed the architecture of SCRIPT. While SCRIPT represents a significant leap—particularly with its JAST-DiT structure and the innovative use of nonlinear history sampling and RLHR—it is not without limitations.

To elevate this system from state-of-the-art performance to a genuinely robust, general-purpose embodied agent capable of complex reasoning, several specific improvements must be implemented. These modifications address the current reliance on reactive control, limited environmental awareness, and the decoupling of semantic understanding from high-level execution.


I propose four distinct enhancements to improve the system's robustness, generalization, and planning capability:

The Problem: The current state s t is purely proprioceptive (joint angles, velocities). This limits the agent to reacting to its own body dynamics and cannot predict collisions or plan around obstacles.

The Improvement: Augment the physical state vector s t with a localized environmental descriptor, E t.

  • Implementation: Use a small local sensory input (e.g., simulated proximity sensors or raycasting results) that provides vectors indicating the distance and relative direction to the nearest obstacle within a radius R. This E t is concatenated to s t to new state s't.

  • Integration: The augmented state vector s't is fed into the JAST-DiT.

The Problem: SCRIPT operates on a receding horizon, predicting the next H steps and then executes only the first action. This is reactive. For long-horizon tasks (e.g., Pick up the box and place it on the high shelf), this method may fail to anticipate future conflicts or suboptimal paths.

The Improvement: Introduce a differentiable dynamics model M that predicts future states s t+H = M(current state, action) before the diffusion step is finalized.

  • Implementation: The JAST-DiT is modified to predict not just the next state-action chunk, but also a latent representation of the trajectory's feasibility. The RLHR phase is modified to include a penalty based on how far the predicted trajectory violates M ’s known constraints.

  • Integration: This allows for a look-ahead mechanism that guides the diffusion process toward dynamically sound trajectories, reducing reliance on pure imitation.

The Problem: The current system relies on the pooled CLIP embedding c pool, which is good for coarse semantic matching but lacks structural information (e.g., "reach under the table," or precise spatial relationships).

The Improvement: Enrich the text stream with structured linguistic features, c struct.

  • Implementation: Utilize a specialized syntax parser and dependency tree encoder on the input text c. This generates tokens representing semantic roles (e.g., agent, object, location) and spatial prepositions. These structural embeddings are concatenated to the original CLIP tokens c txt to form a richer input stream c'txt.

  • Integration: The JAST-DiT’s joint attention mechanism now attends not only to the meaning of the instruction but also to its structure, allowing for precise, nuanced execution.

The Problem: The current RLHR uses a hybrid reward (r phys + r text), but both components are somewhat monolithic. The text reward is evaluated only at the episode's end, and the physical reward is purely trajectory-based.

The Improvement: Implement a hierarchical reinforcement learning structure where success is defined by achieving specific sub-goals derived from the instruction c.

  • Implementation: Define intermediate goal checkpoints (e.g., reach object, lift object). The reward function becomes a composite: r t = w phys r phys(t) + w text r text(T) + sum w goal, i I(Goal i achieved at t).

  • Integration: This guides the RL agent not just to look like the text prompt, but to perform the actions described by the text prompt.

By implementing these modifications, the improved SCRIPT system will transcend its current capabilities and achieve:

  1. True Goal-Directed Behavior: The system will not merely mimic a motion that looks like Dancing joyfully, but will execute a sequence of actions that constitutes joyful dancing, even if the initial expert data was flawed.

  2. Robust Error Recovery: When an obstacle appears or a planned movement is physically impossible (due to the lack of environment awareness), the system will autonomously adjust its trajectory mid-execution rather than failing and requiring a complete re-plan.

  3. Nuanced Semantic Execution: The agent will correctly interpret complex, multi-part instructions like, "Pick up the red ball under the table," by understanding both the action (pick up) and the spatial relationship (under).

  4. Global Planning: For long sequences of actions, the system maintains a coherent plan that anticipates future constraints, preventing local optimizations that lead to global failure.

Sources

Related papers