Dreaming in Code for Curriculum Learning in Open-Ended Worlds
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Dreaming in Code for Curriculum Learning in Open-Ended Worlds".
Jane: The paper was written by Konstantinos Mitsides, Maxence Faldor and Antoine Cully from Department of Computing, Imperial College London.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: You know, Jane, I was looking at this title earlier and it just stuck with me.
Jane: Are you talking about "Dreaming in Code for Curriculum Learning in Open-Ended Worlds"?
Tom: That's the one, and it sounds like something straight out of a sci-fi novel.
Jane: It does, but when you break it down, it's actually a very clever way to describe how we train intelligent agents.
Tom: Right, because instead of just giving them a set of rules, we're letting them "dream" up their own training scenarios through code.
Jane: And the authors, Mitsides, Faldor, and Cully from Imperial College London, are suggesting that this is the way to handle environments that never really end.
Lu: It's much more than just a catchy title, though.
Tom: What do you mean by that, Lu?
Lu: I think the beauty is in the idea of "open-endedness," where the world evolves alongside the intelligence trying to master it.
Jane: That sounds like a constant dance between the creator and the learner.
Meng: It sounds like a massive engineering headache to me, honestly.
Tom: Why is that, Meng?
Meng: Because if the world is always changing, you're constantly chasing a moving target in your testing.
Jane: But isn't that the whole point of the research, to stop the agent from hitting a plateau?
Meng: It is, but you have to make sure the code the AI is dreaming up actually functions without crashing the whole simulation.
Lalam: I see it as a shift in how we perceive digital evolution.
Tom: How so, Lalam?
Lalam: We're moving away from static, human-made lessons toward a culture of self-generated experience.
Jane: That would mean the AI is essentially writing its own textbook as it learns.
Tom: It's a pretty wild concept to start with, but I'm curious to see how they actually pull it off.
Jane: We'll find out when we look at how they're actually building these worlds in the next segment.
Summary: Tom: So, Jane, we've established the vibe, but how does this "dreaming" process actually work in practice?
Jane: They use a framework called DiCode, where a foundation model acts like an architect.
Tom: An architect that writes in Python, right?
Jane: Exactly, the model synthesizes actual, executable code that tells the game engine how to set up a new level.
Meng: And they're doing this within a benchmark called Craftax, which is a really complex, procedurally generated world.
Tom: I heard Craftax is pretty intense with all its different biomes and mechanics.
Meng: It's very demanding, and that's why the closed-loop part of DiCode is so important.
Jane: Could you explain that loop for us, Meng?
Meng: Sure, the agent plays the generated levels, and then its performance metrics are fed back to the foundation model.
Tom: So the model sees where the agent is struggling and adjusts the code to fix that?
Meng: Precisely, it's a constant cycle of generation, testing, and refinement.
Lu: It reminds me of how a human teacher observes a student.
Jane: Like noticing a student can do multiplication but struggles with division, so you give them more division practice?
Lu: Yes, but the AI is doing it by literally rewriting the physics or the resource availability in the game.
Tom: That's a massive leap from just changing a few numbers in a spreadsheet.
Lalam: It suggests a future where the boundary between the creator and the learner becomes completely blurred.
Jane: It's almost like the environment itself has become a living part of the learning process.
Tom: I'm really interested to see if this actually produces better results than the old ways of doing things.
Jane: Well, the results are actually quite startling, so let's talk about those next.
Improvements: Tom: Jane, I was looking at the results section, and the numbers are pretty staggering.
Jane: They really are, especially that sixteen percent improvement in mean return over the strongest baseline.
Tom: But it's not just a small incremental gain, is it?
Jane: No, because they achieved success on tasks that were completely impossible for previous methods.
Meng: I was looking at the combat data, and it's quite impressive from a technical standpoint.
Tom: Which tasks are you talking about, Meng?
Meng: The late-game combat, like facing a Gnome Warrior or a Gnome Archer.
Jane: And the other methods had a zero percent success rate there?
Meng: Yep, they just hit a wall, while DiCode actually managed to navigate those challenges.
Lu: It's because the model started acting like a real teacher.
Tom: What do you mean by that, Lu?
Lu: The paper mentions that the model would remove "scaffolding" once the agent got good.
Jane: So, if the agent was being given extra resources to survive, the model would eventually take those away to increase the difficulty?
Lu: Exactly, it pushes the agent right to the edge of its ability.
Tom: That sounds like it would prevent the agent from getting lazy or just coasting.
Meng: It also keeps the training data fresh, so you don't get stuck in a loop of easy tasks.
Lalam: This shows that true mastery comes from facing the right kind of struggle.
Jane: It's a very biological way of looking at machine learning.
Tom: It really makes you wonder how much more we can achieve if we let the AI design its own challenges.
Jane: We're definitely going to have more to think about as we wrap this up.
Conclusion: Tom: Well, we've covered a lot of ground with "Dreaming in Code for Curriculum Learning in Open-Ended Worlds."
Jane: It really feels like a turning point for how we think about training in complex environments.
Tom: From the architecture at Imperial College London to the actual success in Craftax, it's a complete package.
Jane: It's definitely moving us closer to agents that can learn almost anything if given the right path.
Lu: I think we're looking at the beginning of a new era of generative intelligence.
Meng: I just hope we can keep the computational costs under control as these models get bigger.
Lalam: Regardless of the cost, the cultural impact of creating self-evolving digital minds will be profound.
Tom: Thanks for joining us, everyone.
Jane: See you next time!
Department of Computing, Imperial College London
cs.LG, cs.AI, cs.CL
Submitted: 2026-02-09
Updated: 2026-09-12
Comments: ICML 2026. Project page: https://konstantinosmitsides.github.io/dreaming-in-code
Code: https://github.com/konstantinosmitsides/dreaming-in-code
Project page: https://konstantinosmitsides.github.io/dreaming-in-code
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 83/100
The gist: This paper introduces Dreaming in Code (DiCode), a Unsupervised Environment Design (UED) framework designed to address the performance plateaus common in open-ended learning.
Key concepts
- DiCode
- DiCode is a framework where a foundation model acts as an architect by writing executable Python code. This code instructs a game engine to set up new training levels. The system uses a closed loop, feeding the agent's performance back to the model to refine future challenges.
- Craftax
- Craftax is a complex, procedurally generated benchmark world used to test intelligent agents. It features various biomes and mechanics, providing a demanding environment that requires agents to master diverse tasks, such as late-game combat against Gnome Warriors or Gnome Archers, to demonstrate true proficiency.
- Scaffolding
- Scaffolding refers to the extra resources or easier conditions provided by the model to help an agent learn. As the agent improves, the model removes these supports and increases difficulty, pushing the learner toward mastery by ensuring they face constant, appropriate challenges.
Terminology
Summary
This paper introduces Dreaming in Code (DiCode), a Unsupervised Environment Design (UED) framework designed to address the performance plateaus common in open-ended learning. By utilizing foundation models to synthesize executable environment code, DiCode provides a mechanism for scaffolding learning toward increasing competence,
ensuring agents receive a continual stream of experiences that remain both novel and learnable.
The Core Framework
DiCode operates by allowing a foundation model to dream
new environment instances through the synthesis of executable generation logic. Unlike traditional UED methods that tune low-dimensional parameters, DiCode represents environments as executable programs that can be programmatically specified and composed. This allows for structurally evolving environments
that introduce long-horizon dependencies. The framework is instantiated in Craftax, where the foundation model materializes code-level variations of the world to bridge competence gaps. Crucially, by utilizing a fixed world engine rather than learning a world model, DiCode ensures all generated experiences adhere to valid physics and consistent mechanics.
The Generation Cycle
The framework functions through an interleaved process consisting of training and a generation cycle. The generation cycle follows four sequential steps:
-
Selection of a parent level from an archive based on its learnability score.
-
Generation of a natural language description for the new level, conditioned on the parent and the agent’s current competence.
-
Synthesis of an executable Python program based on that description using a foundation model (Qwen3-235B).
-
Validation through a compilation check and short agent trajectory execution to filter out errors.
This process creates a closed-loop curriculum
where the agent's evolving skill set continuously guides the generative process, effectively maintaining the agent in a zone of proximal development.
Experimental Results
Empirical evaluations on the Craftax benchmark demonstrate that DiCode enables agents to acquire complex, long-horizon skills that are otherwise unattainable. The researchers observed several key performance gains:
-
A 16% improvement in mean return over the strongest baseline.
-
Non-zero success rates on late-game combat tasks (such as defeating Gnome Warriors) where prior methods
effectively collapse to zero.
-
Superiority in mastering
instrumental milestones,
such as crafting iron armor, which are critical for sustaining long-term progress.
Qualitative analysis reveals that the foundation model spontaneously develops teacher-like
strategies, such as removing resource scaffolding to increase difficulty as the agent improves.
Importance of Closed-Loop Feedback
To confirm that the success was driven by the curriculum rather than just generative capability, an ablation study was conducted using an open-loop variant (DiCode-OL). In this version, the feedback loop is removed, and the model generates tasks without access to parent level descriptions or agent performance profiles. The results showed a substantial degradation in final performance,
with DiCode-OL achieving a 15% reduction in score compared to the full DiCode framework. This confirms that generating executable environments alone... is insufficient
and that the gains are derived from the closed-loop curriculum steering generation toward the agent's learnability frontier.
Improvements for AI systems
1. Architectural Shift from Parameter-Based UED to Programmatic Environment Synthesis
-
Improvement: Replace traditional Unsupervised Environment Design (UED) methods that optimize low-dimensional, continuous parameters with a framework that utilizes Foundation Models (FMs) to synthesize executable code representing the environment's transition dynamics (T) and initial state distribution (rho).
-
Capability: The improved AI system can master complex, long-horizon dependencies and hierarchical skills by training in environments where the underlying game logic and interaction rules are structurally evolved through code, rather than just being perturbed by noise or parameter shifts.
2. Closed-Loop Semantic Feedback Integration (Competence-Gap Bridging)
-
Improvement: Implement a feedback loop that conditions the generative process on a multi-dimensional
performance profile
(success rates of specific sub-skills) and the semantic description of aparent level.
-
Capability: The system can autonomously identify precise competence gaps—such as an agent's inability to transition from resource gathering to crafting—and synthesize targeted
stepping-stone
environments. These environments provide specific scaffolding (e.g., pre-loading prerequisite items in the initial state or reducing environmental pressure) to bridge the gap between current capability and the target task.
3. Hierarchical Logic and Objective Evolution
-
Improvement: Move beyond simple reward shaping to a system where FMs programmatically redefine high-level task semantics, including goal definitions (g) and termination conditions.
-
Capability: The improved AI can evolve through complex task hierarchies (e.g., transitioning from
basic survival
tocomplex combat
todeep exploration
) by training on environments that progressively layer new logical dependencies onto mastered behaviors.
4. Asynchronous Generative Curriculum Orchestration
-
Improvement: Decouple the heavy inference latency of large Foundation Models from the Reinforcement Learning (RL) optimization loop through an asynchronous generation pipeline.
-
Capability: The system can maintain high-throughput, continuous training in complex, open-ended worlds without being bottlenecked by the computational cost of generating new, code-based environments, ensuring a steady stream of
Goldilocks
difficulty experiences.
5. Programmatic Robustness and Generalization Testing
-
Improvement: Utilize the FM to generate
mutated
versions of successful levels by systematically removing scaffolding or increasing environmental complexity via code-level modifications (e.g., changing mob spawn logic or resource scarcity). -
Capability: The system can force an agent to generalize learned behaviors from highly scaffolded
easy
environments to realistic, unassisted environments, preventing the agent from overfitting to specific environmental layouts orcheating
through provided resources.
Abstract
Open-ended learning frames intelligence as emerging from continual interaction with an ever-expanding space of environments. While recent advances have utilized foundation models to programmatically generate diverse environments, these approaches often focus on discovering isolated behaviors rather than orchestrating sustained progression. In complex open-ended worlds, the large combinatorial space of possible challenges makes it difficult for agents to discover sequences of experiences that remain consistently learnable. To address this, we propose Dreaming in Code (DiCode), an unsupervised environment design (UED) framework in which foundation models (large language models) synthesize executable environment code to scaffold learning toward increasing competence. In DiCode, "dreaming" takes the form of materializing code-level variations of the world. We instantiate DiCode in Craftax, a challenging open-ended reinforcement learning benchmark characterized by rich mechanics and long-horizon progression. Empirically, DiCode enables agents to acquire long-horizon skills, achieving a 17% improvement in mean return over the strongest baseline and non-zero success on late-game combat tasks where prior methods fail. Our results suggest that code-level environment design provides a practical mechanism for curriculum control, enabling the construction of intermediate environments that bridge competence gaps in open-ended worlds.
Sources
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment Design
- OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code
- Mastering Diverse Domains through World Models
- Replay-Guided Adversarial Environment Design
- Curriculum Learning for Reinforcement Learning Domains: A Framework and Survey
- General Intelligence Requires Rethinking Exploration
- Multi-task curriculum learning in a complex, visual, hard-exploration domain: Minecraft
- Stabilizing Transformers for Reinforcement Learning
- MaestroMotif: Skill Design from Artificial Intelligence Feedback
- Evolving Curricula with Regret-Based Environment Design
- The NetHack Learning Environment
- Teacher algorithms for curriculum learning of Deep RL in continuously parameterized environments
- Eurekaverse: Environment Curriculum Generation via Large Language Models
- No Regrets: Investigating and Improving Regret Approximations for Curriculum Discovery
- Teacher-Student Curriculum Learning
- Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement Learning
- Kinetix: Investigating the Training of General Agents through Open-Ended Physics-Based Control Tasks
- OMNI: Open-endedness via Models of human Notions of Interestingness
- Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks