You've Got a Golden Ticket: Improving Generative Robot Policies With A Single Noise Vector

arXiv:2603.15757 · cs.RO, cs.AI · Submitted 2026-03-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "You've Got a Golden Ticket".

Dev: A novel method for improving generative robot policies involves swapping repeated sampling from a Gaussian prior with a single,

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: So wrapping up what we've seen in "You've Got a Golden Ticket: Improving Generative Robot Policies With A Single Noise Vector," the authors are presenting a method that swaps standard Gaussian noise sampling for a single constant initial noise input, or golden ticket, to enhance the performance of pretrained diffusion or flow matching policies.

Dev: That boils down to finding an optimal vector that boosts performance without updating any model weights or training any new networks, which is a huge win for deployment speed.

Taro: I see the implication as suggesting that we don't necessarily need complex training pipelines to fine-tune generative models; sometimes, a simple input modification can unlock latent capabilities already present in the pretrained structure.

Rosa: The authors are emphasizing that this approach is applicable to all diffusion or flow matching policies and many vision language models, which broadens the scope of where this technique could be useful in robotics and beyond.

Dev: From an engineering viewpoint, the search methods they propose like Random Search and CEM give us concrete tools to find these tickets by maximizing expected rewards through Monte-Carlo policy evaluation.

Taro: The potential impact is significant because it offers a low-overhead way to improve autonomy performance, suggesting that we can achieve better task success rates with less intensive model retraining effort.

Rosa: It seems the authors are pointing toward finding a fixed initial noise vector, which they call the golden ticket, as a key component for improving generative robot policies.

Dev: The search process itself is framed as an optimization problem aimed at maximizing cumulative discounted expected rewards on the downstream task, which gives us a clear objective function to pursue.

Taro: Ultimately, this work suggests that there’s an empirical basis for the lottery ticket hypothesis in robotics, showing that certain noise vectors are inherently better suited for specific robot control outcomes.

Rosa: It really makes us wonder how frequently these tickets appear naturally in the noise space and what underlying geometric properties dictate their effectiveness across different tasks.

Conclusion: Rosa: So, we've been looking at how this paper tackles improving generative robot policies by swapping repeated Gaussian sampling for a single "golden ticket" noise vector.

Dev: Yeah, and the core idea is that this constant initial input can significantly boost policy performance without needing to retrain the model weights at all.

Taro: I think the authors are really pushing the lottery ticket hypothesis here, suggesting there's an inherent structure in random noise that we can exploit for better control.

Rosa: Exactly, and when you look at the results they present, it seems these tickets outperform standard Gaussian noise on a significant chunk of tasks, which is what got my attention.

Dev: From an engineering standpoint, I'm curious about how stable this approach is in real-world deployment; does it hold up outside of the controlled simulation environment for extended periods?

Taro: That's a critical question because if we can get these tickets to generalize across different environments and unexpected scenarios, that would have a huge impact on autonomy.

Rosa: And I want to explore those implications further, especially regarding how this technique could affect the long-term viability of deploying complex generative models in physical systems.

Robotics and AI Institute (RAI) · Arizona State University · Northeastern University

cs.RO, cs.AI

Submitted: 2026-03-16

Updated: 2026-10-03

Comments: 34 pages, 14 figures, 12 tables

Code: https://github.com/irom-princeton/dppo

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 91/100

The gist: A novel method for improving generative robot policies involves swapping repeated sampling from a Gaussian prior with a single, well-chosen constant initial noise vector called a "golden ticket" to

Key concepts

Golden Ticket
A single, well-chosen constant initial noise vector used instead of sampling from the standard Gaussian prior distribution when initializing a generative robot policy. The hypothesis suggests this fixed input contains special properties that naturally lead to better performance on specific tasks.
Lottery Ticket Hypothesis
The idea that random noise vectors, like those from a Gaussian prior, contain 'winning tickets'—specific initial states that are inherently better for certain outcomes. In robotics, this means finding a fixed noise vector that naturally steers the policy toward higher expected rewards on a specific task.
Monte-Carlo Policy Evaluation
A method used to estimate the cumulative discounted expected rewards obtained from running a robot policy over many episodes. This estimation is crucial for searching for the optimal golden ticket by providing a quantifiable measure of how well different noise vectors perform on a given task.
Cross-Entropy Method (CEM)
An optimization algorithm used to search for the best golden ticket. It iteratively refines a distribution over noise vectors by selecting the top performers and updating the distribution parameters, aiming to converge on a noise vector that maximizes expected task rewards.

Terminology

Summary

A novel method for improving generative robot policies involves swapping repeated sampling from a Gaussian prior with a single, well-chosen constant initial noise vector called a golden ticket to enhance performance on downstream tasks. This approach is significant because it allows for policy improvement without updating model weights or training additional networks, making it applicable to any diffusion or flow matching policy and offering substantial gains across various benchmarks.

The Gist

Performance of a pretrained, frozen diffusion or flow matching policy can be improved by simply replacing the sampling of initial noise from the prior distribution with a well-chosen, constant initial noise input, which we call a golden ticket.

Policy Improvement Hypothesis and Motivation

The work is motivated by the lottery ticket hypothesis, which suggests that random Gaussian noise images contain special pixel blocks (winning tickets) that naturally tend to be denoised into specific content independently. In the context of robot control policies, this hypothesis is framed as: "For a pretrained diffusion or flow matching policy and a downstream task, there exists a fixed initial noise vector w∗ that, used in place of sampling from the Gaussian prior, increases the policy’s expected return on held-out task evaluations." This approach avoids the computational and model design challenges associated with existing methods like updating weights or training additional networks.

The Golden Ticket Search Problem

Finding an optimal golden ticket is framed as a search problem where candidate initial noise vectors, called lottery tickets, are optimized to maximize cumulative discounted expected rewards on the downstream task. The objective is formally defined as:

"w∗ = argmax w 1 M P M m=1 P t>0 γ t R(s(m)t, πw(s(m)t))"

This search is conducted using Monte-Carlo policy evaluation to estimate episodic rewards obtained from policy rollouts. The proposed search algorithms include:

  1. Random Search (RS): Samples N initial noise vectors and evaluates each in M search environments.

  2. Cross-Entropy Method (CEM): Iteratively refines a Gaussian sampling distribution over noise vectors by selecting the top-K elites based on mean return and updating the distribution parameters.

  3. Zeroth-Order Search (ZOS): Performs a local random walk in noise space anchored at a pivot, advancing the pivot to better candidates based on their mean return.

Empirical Performance and Multi-Task Extension

Experiments demonstrate that golden tickets outperform Gaussian noise in 46 out of 51 tasks, with at least matching it in 49. The gains are substantial:

In franka sim cube picking, the base policy averages 38.5% success and the best golden tickets average 96%.

The approach extends naturally to multi-task settings: a single ticket can boost performance across several tasks when searched jointly, and tickets optimized for one task may transfer to unseen tasks within the same suite. For instance, joint multi-task search yields tickets that lift the multi-task average on SimplerEnv by 14% across 7 tasks.

Search Method Analysis and Comparison

The paper compares RS, CEM, and ZOS against state-of-the-art latent steering methods like DSRL. Wall-clock time comparisons show that searching for golden tickets is dramatically cheaper than DSRL in wall-clock time to reach a fixed env-step budget. Furthermore, CEM+racing (a variant of CEM incorporating sequential halving) shows a sample-efficiency gain over plain CEM while maintaining similar performance to DSRL on tasks like Lift and Can. The search methods are also analyzed for subspace ablation; for instance, restricting the proposal distribution to lower-dimensional subspaces (like CEM-gripper) can capture fast, cheap gains but may lack the degrees of freedom to fully fix policy failure modes.

Hardware and Real-World Utility

The method shows substantial gains on hardware. For example:

For cup pushing, the base policy was evaluated over 10 episodes. We searched 10 tickets, each for 5 episodes, and then the best performing ticket (Ticket 2) and evaluated an additional 10 episodes.

This resulted in a 60% improvement in success rate for the cup pushing task. Similarly, on real Franka hardware, searching for a few tickets can drive success rates from 80% to 98% with relatively small search budgets (e.g., 150 episodes total). The results show that golden tickets can generalize within certain spatial locations, and the search process allows steering the policy to miss with extreme reliability.

Conclusion on Effectiveness

The lottery ticket hypothesis is empirically supported, showing that golden tickets often exist and outperform Gaussian noise. They are competitive with state-of-the-art latent steering methods trained with RL on certain tasks (e.g., Can).

Improvements for AI systems

Here are the specific improvements that can be made to AI systems based on this research:

  • Improve out-of-the-box performance of pretrained generative robot policies (diffusion or flow matching) by utilizing a single, optimized constant initial noise vector, termed a golden ticket, instead of repeatedly sampling from the prior distribution.

  • Enable policy improvement without updating original model weights or training additional networks.

  • Develop and deploy simple search algorithms (Random Search, Cross-Entropy Method (CEM), Zeroth-Order Search (ZOS)) to efficiently find these golden tickets using Monte Carlo policy evaluation on task rewards.

  • Enhance performance in downstream tasks by leveraging the discovered golden tickets, achieving absolute success rate improvements up to 55% for some simulated tasks and 28% for real-world tasks within a limited search budget (e.g., <150 episodes per task).

  • Improve multi-task Vision-Language-Action (VLA) models by using a single golden ticket that can boost the average performance across multiple related tasks (e.g., improving average performance by 14% across 7 tasks using one ticket).

  • Facilitate transfer learning between different related robot manipulation tasks, as a golden ticket optimized for one task can also improve performance in others within the same VLA policy suite.

  • Create a deployable system that requires no additional infrastructure or models beyond the pretrained policy and an environment capable of calculating sparse task rewards.

This improved AI system can perform:

  1. Pick objects from various spatial locations with high reliability (e.g., achieving 96% success rate on cube picking in simulation).

  2. Assemble complex robotic components (like pistons) with significantly higher success rates (up to 88%) compared to base policies.

  3. Execute fine manipulation tasks like threading, drawer cleanup, and box cleanup with improved accuracy and speed.

  4. Perform complex point-cloud based actions, such as picking bananas or pushing cups, with near-perfect success rates (e.g., 100% success on cup pushing).

  5. Adapt a single foundational VLA model to perform a suite of related manipulation tasks with an average performance boost of 14%.

  6. Improve the generalization and robustness of robot policies across different observation modalities (RGB, point-cloud, state) and embodiments (single-arm vs. bimanual).

Sources

Related papers