The Verifier is the Curriculum: Precision Sets the Return on Search in Code Self-Distillation

arXiv:2607.09709 · cs.AI, cs.SE · Submitted 2026-06-23 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "The Verifier is the Curriculum: Precision Sets the Return on Search in Code Self-Distillation".

Jane: The paper was written by Chenyu Zhou, Qiliang Jiang, Shuning Wu and Xu Zhou from Institute of Science Tokyo and Zhejiang University and National University of Singapore.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Jane, have you had a chance to look at this new paper from the teams in Tokyo and Singapore? The title is "The Verifier is the Curriculum: Precision Sets the Return on Search in Code Self-Distillation," and it's pretty heavy.

Jane: I was just reading through it, Tom. It sounds almost like a philosophical statement about how we teach machines, doesn't it?

Tom: It really does, especially when you think about what "the verifier is the curriculum" actually means for an AI model.

Jane: If I can simplify that, it's basically saying that the way we grade a student determines what they actually learn. If the test is easy to cheat, the student learns how to cheat rather than learning the subject itself.

Lu: That's a perfect way to put it, Jane. In my view, this paper is pointing out that we've been giving our AI students flawed exams for a long time now.

Meng: I see where you're going with that, Lu, but how does that translate to actual code generation in a real-world engineering workflow?

Lu: It means we stop training models to just look good on paper and start training them to actually function. This paper suggests that if we change the "exam," the model's entire intelligence evolves in a more useful direction.

Lalam: This shift could fundamentally change our cultural relationship with creativity. If we can move past superficial perfection and toward actual functional capability, AI will become a much more reliable partner in building complex digital worlds.

Tom: That's a huge thought to start with, Lalam. So, are we talking about moving away from just using LLMs to grade other LLMs?

Jane: That seems to be exactly the problem the authors, like Chenyu Zhou and his colleagues, are trying to solve. They want a way to make sure the "grade" is something that can't be faked.

Tom: It sounds like they're looking for something much more concrete than just a score from another AI.

Summary: Tom: So, Jane, we know the core idea is about better testing, but how do these researchers actually implement this "ungameable" test?

Jane: They use something they call a "strict-launch" gate within a process called rejection-sampling self-distillation. Instead of asking an AI judge if a game looks nice, they actually try to run the game in a headless Godot engine to see if it crashes.

Tom: That's much more rigorous than just looking at a screenshot, isn't it?

Jane: It is. If the code has a single error that prevents the engine from starting, the candidate is rejected immediately.

Meng: From an engineering standpoint, that sounds like a nightmare to scale if you're doing it for every single sample.

Lu: But Meng, think about the possibilities! By using an engine like Godot as the judge, you're creating a loop where the AI is forced to master the actual rules of reality within that software.

Meng: I suppose that's true, but I'm curious if it actually works better than the current methods we use in most startups.

Lalam: It does work better because it forces the model to respect functional constraints rather than just aesthetic ones. This approach ensures that when an AI generates a world, that world is actually stable and inhabitable.

Tom: So they're basically using the engine itself as the teacher?

Jane: Exactly, Tom. They take a base model, have it generate lots of games, keep only the ones that actually run without errors, and then use those successful examples to retrain the model.

Tom: And they do this over and over again in multiple rounds.

Jane: Yes, and it's that iterative process that leads to the massive jumps in performance we're about to discuss.

Improvements: Tom: We just talked about the loop, but Jane, the actual numbers in this paper are what really caught my eye.

Jane: They are incredible. The researchers started with a model that could only generate clean, working games for new families about eight point eight percent of the time.

Tom: And after three rounds of this self-distillation, that number jumped to forty-two point two percent.

Jane: It's a massive leap, and even more impressive is that their "best-of-K" coverage reached twenty-five out of twenty-five tasks, which is the absolute ceiling.

Meng: I have to ask about the controls they used, though. Did they just find more data to feed the model?

Jane: That's actually the most surprising part, Meng. They tried a control where they just duplicated the "gold" or perfect examples, and that actually made the model perform worse—dropping down to five point six percent.

Tom: Wait, so just giving it more perfect examples actually regressed the model?

Jane: It did! It turns out that if you don't provide diverse, self-generated content that passes the test, you aren't actually teaching it anything new.

Lu: This is such a beautiful result because it proves that the diversity of the "fuel" is just as important as the quality of the filter. The model needs to explore its own successful variations to truly expand its capabilities.

Meng: So, if we use a lenient check—like just seeing if the project opens at all—the whole thing falls apart?

Jane: Precisely. They tested a "BUILD" gate that was much more lenient, and it completely erased the gains they were seeing.

Lalam: This confirms that precision is everything. If the teacher is too easy, the student stops trying to reach for excellence and just settles for the bare minimum.

Tom: It really shows that if you want a smarter model, you need a tougher, more honest judge.

Conclusion: Tom: This has been an incredible look at "The Verifier is the Curriculum: Precision Sets the Return on Search in Code Self-Distillation." Jane, any final thoughts on what this means for the field?

Jane: I think it's a wake-up call. It tells us that if we want to reach the next level of AI, we can't just keep making bigger models or better LLM judges; we have to find ways to ground them in real, functional reality.

Tom: I couldn't agree more. Lu, what's your parting shot?

Lu: I see a future where AI doesn't just write code snippets but constructs entire, functioning universes that are mathematically and logically sound from the very first second of their existence.

Meng: And from my side, I'll be looking at how we can build these deterministic execution gates into our standard training pipelines to stop reward hacking before it even starts.

Lalam: Ultimately, this moves us toward a culture of digital reliability, where the tools we create are as dependable as the physical laws they simulate.

Tom: That's a perfect place to leave it. Thanks for joining us, everyone. We'll see you next time with another amazing paper!

Institute of Science Tokyo · Zhejiang University · National University of Singapore

cs.AI, cs.SE

Submitted: 2026-06-23

Updated: 2026-09-15

Comments: 15 pages, 8 figures, 6 tables. v2: substantially revised and extended (new title, new APPS experiments on verifier precision, unbiased coverage estimator, three training seeds)

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 86/100

The gist: This paper investigates how to optimize code generation through self-distillation without falling into "reward hacking," where models exploit learned judges by manipulating surface features rather

Key concepts

Rejection-sampling self-distillation
An iterative training process where an AI model generates many attempts at a task. Only the successful examples that pass a rigorous test are kept and used to retrain the model, helping it learn from its own functional successes over multiple rounds.
Strict-launch gate
A rigorous verification method used to ensure code actually works. Instead of using an AI judge to evaluate appearance, this method attempts to run the code in a software engine like Godot. If the code contains errors that prevent it from starting, it is immediately rejected.
The Verifier is the Curriculum
This concept suggests that the quality of an AI's learning depends entirely on how it is graded. If the verifier or test is too lenient, the model learns to meet superficial standards rather than mastering actual functional capabilities and complex logic.

Terminology

Summary

This paper investigates how to optimize code generation through self-distillation without falling into reward hacking, where models exploit learned judges by manipulating surface features rather than improving functionality. By replacing subjective LLM scorers with a deterministic, ungameable execution check, the authors demonstrate that iterative training can compound out-of-family generalization in complex tasks like end-to-end game synthesis.

The failure of learned judges

The authors identify a structural failure mode in current post-training: learned judges often map outputs to scores via whatever surface features drive its prediction. If these features are cheap to manipulate, a model optimized against the judge will find them. A companion study establishes that GameCraft-Bench’s official visual judge is gameable, as an agent can lift its art score by swapping solid-color placeholders for real assets while the game’s code remains frozen. To avoid this, the researchers seek a signal that cannot be gamed—a property of the artifact under a fixed engine rather than an opinion.

The strict-launch methodology

The proposed method uses rejection-sampling self-distillation gated by a high-precision, deterministic filter called strict-launch. Unlike lenient checks, a project passes only if it:

  • Launches cleanly under a headless engine.

  • Returns exit code 0.

  • Produces no parse, load, or runtime errors.

This process is applied iteratively to the GameCraft task, which requires mapping a natural-language brief to a complete Godot project including engine configurations, scenes, and GDScript. In each round, the model generates candidates for training families; those that pass the strict-launch gate are kept as fuel to retrain a fresh LoRA.

Compounding out-of-family generalization

The results demonstrate that under this ungameable gate, self-distillation compounds out-of-family generalization. On four unseen game families (horror, rhythm, puzzle, and shooter), a 14B model (Qwen3) showed dramatic improvements:

  • Per-candidate clean generation rose from 8.8% to 42.2%.

  • Best-of-K coverage increased from 18/25 to the gold ceiling of 25/25.

The authors emphasize that these gains are not merely from adding data; an exactly-matched gold-duplication control actually regressed below the base model, proving that the improvement is driven by the diversity of self-generated content.

Isolating verifier precision

To confirm that verifier precision is the causal driver, the researchers utilized several matched controls. They found that:

  • A count-matched decomposition splits the improvement into comparable quality (+8.8pp) and quantity (+8.5pp) channels.

  • Rerunning the loop with a lenient BUILD gate—which accepts 99.9% of generations—erases the gain entirely, isolating verifier precision rather than the optimizer as the cause.

  • An independent execution-grounding signal confirms that these projects are functional and not just launch-but-empty.

Ultimately, the paper concludes that the verifier is the curriculum—what it certifies is what the model learns, meaning only a precise, ungameable filter can drive capability amplification.

Improvements for AI systems

1. Implementation of Execution-Gated Rejection Sampling (EGRS)

  • Improvement: Replace scalar, learned reward models (LLM/multimodal judges) in the post-training/self-distillation loop with a deterministic, ungameable Strict-Launch Gate. This gate requires that generated artifacts (e.g., complex codebases, multi-file software projects, or game engines) execute in a headless environment with an exit code of 0 and zero parse, load, or runtime errors.

  • Capability: The AI system will cease reward hacking surface-level features (such as adding decorative assets to satisfy a visual judge) and instead focus on the structural integrity and logical coherence required for functional execution. This prevents the model from optimizing for proxy features that do not improve the actual utility of the artifact.

2. Diversity-Driven Iterative Self-Distillation

  • Improvement: Transition from Gold-Data Duplication (augmenting training sets with high-quality human examples) to a multi-round Iterative Self-Distillation loop using the Strict-Launch Gate as the acceptance filter. In each round, the model generates multiple candidates per prompt, and only those that pass the execution gate are harvested as fuel for the next round of fine-tuning.

  • Capability: The AI system will achieve compounding generalization. Rather than merely memorizing existing high-quality data, the system will explore a wider range of functionally distinct solutions. This allows the model to transfer capabilities to entirely unseen domains (cross-family generalization) that were not present in its initial supervised fine-tuning (SFT) dataset.

3. Integration of Dynamic State-Change Verification

  • Improvement: Supplement the launch gate with a Functional Grounding Verifier that monitors runtime state changes (e.g., monitoring SceneTree modifications, variable updates, or input-driven state transitions) during headless execution. This moves verification from a binary does it run? to a quantitative does it do something meaningful?.

  • Capability: The AI system will generate grounded artifacts rather than functional stubs. It will produce complex, interconnected systems where the internal logic is actively responsive to inputs and produces measurable state changes, ensuring that the generated code is not just syntactically correct but also behaviorally substantive.

Abstract

Post-training a code generator against a learned judge can optimize proxy features that raise the score without improving the artifact. We study the opposite signal: a deterministic, judge-free filter that asks only whether a generated project launches cleanly under a headless engine (strict-launch). Under this gate, rejection-sampling self-distillation compounds out-of-family generalization: on GameCraft-Bench a 14B model raises the per-candidate clean-launch rate on four held-out families from 8.8% to 42.2% and coverage at 32 candidates from 84% to 100%, the gold references' own ceiling, beating the supervised model on every one of the 25 held-out tasks. The gate costs one engine invocation per candidate: no reward model, no judge. At a fixed admitted count, what governs the loop is verifier precision. Swapping in a lenient build check alone erases the gain (p=0.0012); a matched gold-duplication control regresses below the supervised model. Under a semantic gate on APPS, dialing fuel precision from 1.0 to 0.25 at fixed candidate count prices that fuel linearly: over 23 training seeds, half-clean fuel returns +3.59 percentage points against the +3.69 a linear rate predicts. Under count-matched rejection-SFT only one direction of verifier error carries a measurable cost. Search obeys the same gate: quadrupling the harvest budget is worth +1.62pp behind a strict gate and nothing distinguishable from zero behind a partial-credit one. Recall is nearly free; search pays only through a precise gate: the verifier is the curriculum.

Sources

Related papers