Dreaming in Code for Curriculum Learning in Open-Ended Worlds
summary
The gist
This paper introduces Dreaming in Code (DiCode), a Unsupervised Environment Design (UED) framework designed to address the performance plateaus common in open-ended learning.
In short
This episode discusses the paper "Dreaming in Code for Curriculum Learning in Open-Ended Worlds." It explores the DiCode framework, where a foundation model writes Python code to design training levels within the Craftax benchmark. By adjusting difficulty through a closed-loop system, DiCode significantly outperforms previous methods in complex environments.
Key concepts
- DiCode
- DiCode is a framework where a foundation model acts as an architect by writing executable Python code. This code instructs a game engine to set up new training levels. The system uses a closed loop, feeding the agent's performance back to the model to refine future challenges.
- Craftax
- Craftax is a complex, procedurally generated benchmark world used to test intelligent agents. It features various biomes and mechanics, providing a demanding environment that requires agents to master diverse tasks, such as late-game combat against Gnome Warriors or Gnome Archers, to demonstrate true proficiency.
- Scaffolding
- Scaffolding refers to the extra resources or easier conditions provided by the model to help an agent learn. As the agent improves, the model removes these supports and increases difficulty, pushing the learner toward mastery by ensuring they face constant, appropriate challenges.
Terminology used across episodes
This episode discusses
- Dreaming in Code for Curriculum Learning in Open-Ended Worlds · Paper Radio
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment Design
- OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code
- Mastering Diverse Domains through World Models
- Replay-Guided Adversarial Environment Design
- Curriculum Learning for Reinforcement Learning Domains: A Framework and Survey
- General Intelligence Requires Rethinking Exploration
- Multi-task curriculum learning in a complex, visual, hard-exploration domain: Minecraft
- Stabilizing Transformers for Reinforcement Learning
- MaestroMotif: Skill Design from Artificial Intelligence Feedback
- Evolving Curricula with Regret-Based Environment Design
- The NetHack Learning Environment
- Teacher algorithms for curriculum learning of Deep RL in continuously parameterized environments
- Eurekaverse: Environment Curriculum Generation via Large Language Models
- No Regrets: Investigating and Improving Regret Approximations for Curriculum Discovery
- Teacher-Student Curriculum Learning
- Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement Learning
- Kinetix: Investigating the Training of General Agents through Open-Ended Physics-Based Control Tasks
- OMNI: Open-endedness via Models of human Notions of Interestingness
- Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World
The paper
Dreaming in Code for Curriculum Learning in Open-Ended Worlds · Read on arXiv
Department of Computing, Imperial College London
Open-ended learning frames intelligence as emerging from continual interaction with an ever-expanding space of environments. While recent advances have utilized foundation models to programmatically generate diverse environments, these approaches often focus on discovering isolated behaviors rather than orchestrating sustained progression. In complex open-ended worlds, the large combinatorial space of possible challenges makes it difficult for agents to discover sequences of experiences that remain consistently learnable. To address this, we propose Dreaming in Code (DiCode), an unsupervised environment design (UED) framework in which foundation models (large language models) synthesize executable environment code to scaffold learning toward increasing competence. In DiCode, "dreaming" takes the form of materializing code-level variations of the world. We instantiate DiCode in Craftax, a challenging open-ended reinforcement learning benchmark characterized by rich mechanics and long-horizon progression. Empirically, DiCode enables agents to acquire long-horizon skills, achieving a 17% improvement in mean return over the strongest baseline and non-zero success on late-game combat tasks where prior methods fail. Our results suggest that code-level environment design provides a practical mechanism for curriculum control, enabling the construction of intermediate environments that bridge competence gaps in open-ended worlds.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Dreaming in Code for Curriculum Learning in Open-Ended Worlds".
Jane: The paper was written by Konstantinos Mitsides, Maxence Faldor and Antoine Cully from Department of Computing, Imperial College London.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: You know, Jane, I was looking at this title earlier and it just stuck with me.
Jane: Are you talking about "Dreaming in Code for Curriculum Learning in Open-Ended Worlds"?
Tom: That's the one, and it sounds like something straight out of a sci-fi novel.
Jane: It does, but when you break it down, it's actually a very clever way to describe how we train intelligent agents.
Tom: Right, because instead of just giving them a set of rules, we're letting them "dream" up their own training scenarios through code.
Jane: And the authors, Mitsides, Faldor, and Cully from Imperial College London, are suggesting that this is the way to handle environments that never really end.
Lu: It's much more than just a catchy title, though.
Tom: What do you mean by that, Lu?
Lu: I think the beauty is in the idea of "open-endedness," where the world evolves alongside the intelligence trying to master it.
Jane: That sounds like a constant dance between the creator and the learner.
Meng: It sounds like a massive engineering headache to me, honestly.
Tom: Why is that, Meng?
Meng: Because if the world is always changing, you're constantly chasing a moving target in your testing.
Jane: But isn't that the whole point of the research, to stop the agent from hitting a plateau?
Meng: It is, but you have to make sure the code the AI is dreaming up actually functions without crashing the whole simulation.
Lalam: I see it as a shift in how we perceive digital evolution.
Tom: How so, Lalam?
Lalam: We're moving away from static, human-made lessons toward a culture of self-generated experience.
Jane: That would mean the AI is essentially writing its own textbook as it learns.
Tom: It's a pretty wild concept to start with, but I'm curious to see how they actually pull it off.
Jane: We'll find out when we look at how they're actually building these worlds in the next segment.
Summary: Tom: So, Jane, we've established the vibe, but how does this "dreaming" process actually work in practice?
Jane: They use a framework called DiCode, where a foundation model acts like an architect.
Tom: An architect that writes in Python, right?
Jane: Exactly, the model synthesizes actual, executable code that tells the game engine how to set up a new level.
Meng: And they're doing this within a benchmark called Craftax, which is a really complex, procedurally generated world.
Tom: I heard Craftax is pretty intense with all its different biomes and mechanics.
Meng: It's very demanding, and that's why the closed-loop part of DiCode is so important.
Jane: Could you explain that loop for us, Meng?
Meng: Sure, the agent plays the generated levels, and then its performance metrics are fed back to the foundation model.
Tom: So the model sees where the agent is struggling and adjusts the code to fix that?
Meng: Precisely, it's a constant cycle of generation, testing, and refinement.
Lu: It reminds me of how a human teacher observes a student.
Jane: Like noticing a student can do multiplication but struggles with division, so you give them more division practice?
Lu: Yes, but the AI is doing it by literally rewriting the physics or the resource availability in the game.
Tom: That's a massive leap from just changing a few numbers in a spreadsheet.
Lalam: It suggests a future where the boundary between the creator and the learner becomes completely blurred.
Jane: It's almost like the environment itself has become a living part of the learning process.
Tom: I'm really interested to see if this actually produces better results than the old ways of doing things.
Jane: Well, the results are actually quite startling, so let's talk about those next.
Improvements: Tom: Jane, I was looking at the results section, and the numbers are pretty staggering.
Jane: They really are, especially that sixteen percent improvement in mean return over the strongest baseline.
Tom: But it's not just a small incremental gain, is it?
Jane: No, because they achieved success on tasks that were completely impossible for previous methods.
Meng: I was looking at the combat data, and it's quite impressive from a technical standpoint.
Tom: Which tasks are you talking about, Meng?
Meng: The late-game combat, like facing a Gnome Warrior or a Gnome Archer.
Jane: And the other methods had a zero percent success rate there?
Meng: Yep, they just hit a wall, while DiCode actually managed to navigate those challenges.
Lu: It's because the model started acting like a real teacher.
Tom: What do you mean by that, Lu?
Lu: The paper mentions that the model would remove "scaffolding" once the agent got good.
Jane: So, if the agent was being given extra resources to survive, the model would eventually take those away to increase the difficulty?
Lu: Exactly, it pushes the agent right to the edge of its ability.
Tom: That sounds like it would prevent the agent from getting lazy or just coasting.
Meng: It also keeps the training data fresh, so you don't get stuck in a loop of easy tasks.
Lalam: This shows that true mastery comes from facing the right kind of struggle.
Jane: It's a very biological way of looking at machine learning.
Tom: It really makes you wonder how much more we can achieve if we let the AI design its own challenges.
Jane: We're definitely going to have more to think about as we wrap this up.
Conclusion: Tom: Well, we've covered a lot of ground with "Dreaming in Code for Curriculum Learning in Open-Ended Worlds."
Jane: It really feels like a turning point for how we think about training in complex environments.
Tom: From the architecture at Imperial College London to the actual success in Craftax, it's a complete package.
Jane: It's definitely moving us closer to agents that can learn almost anything if given the right path.
Lu: I think we're looking at the beginning of a new era of generative intelligence.
Meng: I just hope we can keep the computational costs under control as these models get bigger.
Lalam: Regardless of the cost, the cultural impact of creating self-evolving digital minds will be profound.
Tom: Thanks for joining us, everyone.
Jane: See you next time!
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language