PhantomEnvironments: Training LLM Agents in Fictional Worlds
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "PhantomEnvironments: Training LLM Agents in Fictional Worlds".
Jane: Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, wrapping up our discussion on "PhantomEnvironments: Training LLM Agents in Fictional Worlds," the authors are essentially arguing that we can train LLMs into strong search agents using zero-marginal-cost, rule-generated synthetic environments instead of relying on costly human curation or potentially hallucinated data.
Jane: The title itself really captures the essence of their work; they're not just training an agent, they're showing how to train one within these fictional worlds and what that actually means for the future.
Lu: It implies a pathway where we can design specific, structured training scenarios purely through logic, allowing us to target exact capabilities we want the agents to develop.
Meng: From a practical standpoint, this suggests that we don't always need massive datasets of real-world interactions if we can create highly targeted synthetic data that teaches core reasoning skills efficiently.
Lalam: The paper’s implication is that agentic search skill can be learned procedurally through RL interaction with these simple, verifiable environments, which is a powerful way to build robust capabilities into the model itself.
Tom: It boils down to proving that we can generate effective training material for complex AI tasks cheaply and reliably, leading to agents that show solid transferability to actual search benchmarks.
Jane: If these models can learn this kind of scalable search skill from fictional worlds, it means we might see a new way to bootstrap the development of powerful reasoning agents across many different domains.
Conclusion: Tom: So we've seen how these LLMs learn search skills inside these fictional setups, but let's talk about what that title means for us as listeners and thinkers out there.
Jane: I think the name "PhantomEnvironments" is really clever because it suggests these worlds are almost invisible or ghostly, yet they guide the agent's learning process in a very real way.
Lu: From a creative standpoint, imagining entire fictional universes built purely from rules opens up possibilities for designing training scenarios we can't even conceive of right now.
Meng: Practically speaking, the authors are showing how you can create high-quality training data without needing massive amounts of real-world interaction or expensive human labeling.
Lalam: I see this as a huge step toward making agentic skill acquisition more scalable because it proves that structured interaction in synthetic settings is an effective way to build procedural abilities.
Tom: That’s a big shift, Lalam; moving away from just feeding them raw data and instead training them on these meticulously crafted logic puzzles.
Jane: It really boils down to the authors showing us that we don't have to wait for the real world to teach an agent everything it needs to know about complex search tasks.
Lu: Think about the sheer breadth of knowledge you could inject into these fictional rules; every new world is a completely novel training ground waiting to be explored.
Meng: If this methodology works consistently across different types of problems, then the real impact is that we could rapidly deploy capable agents for many specialized tasks where real-world data is scarce.
Lalam: The most impactful vision I see here is that we can start building a culture where foundational agentic reasoning skills are learned through these synthetic environments before they ever encounter messy, unpredictable human data.
Tom: That sounds like an exciting path forward for how we develop the next generation of helpful AI systems.
Jane: It gives us a tangible framework for how to test and refine these models systematically without needing constant access to the internet or expensive feedback loops.
Lu: We need to keep thinking about what kind of complex interactions those fictional worlds might allow the agents to develop that we haven't even considered yet.
Meng: I’m curious if this rule-based generation method can actually scale up efficiently when we move from simple Q andA games to more intricate, multi-step operational tasks.
Anmol Kabra, Swathi Saravana Selvam, Albert Gong, Chao Wan, Christian Belardi, Dongyoung Go
Cornell University · Stanford University
cs.LG, cs.AI, cs.CL
Submitted: 2026-09-30
Updated: 2026-09-30
Code: https://github.com/kilian-group/phantom-envs
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 90/100
The gist: Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply.
Key concepts
- PhantomEnvironments Construction
- These are multi-turn RL environments built from fictional worlds using codified rules instead of LLMs. They consist of entities with programmed relationships and template-generated articles. Because they are rule-based, they have zero generation cost and are verifiable by running parallel Prolog queries to confirm ground-truth answers.
- Search Scaling
- This is an emergent skill where the trained agents learn to allocate their search budget in a way that scales linearly with question difficulty. Instead of spending the same amount of effort regardless of how hard a question is, the model learns to adjust its search strategy based on complexity, which is a key skill learned from interacting with these synthetic environments.
- Complexity Axes Analysis
- The study examined three ways complexity can be structured: linear hops over entities, comparisons of attributes, and constraints. The findings reveal that 'linear hops' are the most important axis for transferring skills to real-world tasks, suggesting that learning to navigate sequential entity relationships is more valuable than learning attribute comparisons or filtering candidates.
Terminology
Summary
Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply. The gist: LLMs can instead be trained into capable search agents using synthetic environments generated entirely by rules, whose generation requires no LLM and has zero marginal cost.
PhantomEnvironments Construction
The core innovation is the creation of PhantomEnvironments,
multi-turn RL environments derived from fictional worlds and template-generated multi-hop questions. These environments are built entirely from codified rules, requiring no humans or LLMs in the pipeline, which results in them being verifiable by construction and incur zero generation cost.
The worlds consist of fictional entities connected through programmed relationships (like family or friendship), and articles about these individuals are generated using template-based formats. Questions are generated with context-free grammars, ensuring every question-answer pair is fully verifiable by executing a parallel Prolog query that returns ground-truth answers.
Training Methodology
The paper follows the RL fine-tuning recipe of Search-R1, instructing models to use XML tags like for queries and for final answers. The optimization uses Group Relative Policy Optimization (GRPO), where the model is rewarded only on the final answer, and environment outputs within tags are masked out of the objective function. Training involves approximately 55K questions in these synthetic environments, which are compared against a random subsample of 55K questions from the NaturalQuestions-HotpotQA corpus to assess transfer to real-world data.
Key Findings on Transfer and Scaling
The agents trained in PhantomEnvironments demonstrate significant and consistent transfer to real-world multi-hop search benchmarks. Performance improves by roughly 1.7× on older benchmarks (like HotpotQA) and 2.2× on newer, harder ones (like SynthWorlds-RM). Crucially, the paper shows that Qwen2.5 models learn to allocate their search budget roughly linearly with question difficulty—an emergent property termed search scaling
—suggesting this is a skill learned from environment interaction alone. Furthermore, agents trained in these synthetic environments are robust to larger-scale deployments with noisy search results and are effective at closing the Knowledge Advantage (KA) gap when evaluation targets fall beyond LLM training cutoffs.
Complexity Axes Analysis
The study analyzed three orthogonal axes of complexity: linear hops over entities, comparisons of attributes, and constraints that filter candidates. The results show that Hops is the strongest axis that drives real-world transfer,
as linear hops are at least as effective as the other two axes combined. Specifically, comparison questions target matching real-world categories and significantly outperform constraint questions, which are shown to hurt
transfer because they reward a verbatim-query shortcut rather than decomposition.
Generalization and Robustness
The agents trained on fixed fictional universes generalize well to unseen environments of the same size and even 10× larger universes, confirming that they have learned generalizable agentic search skill rather than memorizing facts.
The analysis further reveals that learning agentic search through RL largely teaches a procedural skill, which can be learned complement to the LLMs’ memorized parametric knowledge. The paper concludes that rule-generated synthetic environments are a zeromarginal-cost and effective source of agent training data.
Limitations and Future Directions
The limitations noted include the lag on in-domain benchmarks where knowledge memorization helps, and the necessity to design environments to avoid reward shortcuts. Future work is suggested for environment composition, investigating if one can read off a recipe of synthetic axes
that closes specific capability gaps. The authors also note that compute budgets limit training larger models and open-web benchmarks like BrowseComp-Plus.
AI Use Statement
In this work, we used generative AI tools for code implementation and experimentation. We have not used generative AI tools for generating synthetic datasets (the paper is about LLM-free synthetic data), or interpreting results; the rest of the required disclosure tasks are not applicable to this work. Additionally, we used generative AI tools for brainstorming, editing for readability, and plotting. We have reviewed all AI-assisted work by verifying code and confirming all literature. We take responsibility for the final content of this work, including text, claims or artifacts produced with the aid of generative AI.
Reproducibility Statement
We use open-source LLMs, training code, and evaluation benchmarks, and our training experiments are across multiple seeds. All details are in the appendix. We report standard errors in all tables and plots and note significance. We have open-sourced our code at github.com/kilian-group/phantom-envs.
Acknowledgements
Authors acknowledge help from AI models in various project phases. DG is supported by Empire AI Postdoctoral Fellowship.
Improvements for AI systems
Here are specific improvements for AI systems based on the PHANTOMENVIRONMENTS: TRAINING LLM AGENTS IN FICTIONAL WORLDS
paper, categorized by capability.
) 1. Develop Zero-Marginal-Cost, Verifiable Synthetic RL Training Environments.
The system can generate entire multi-turn reinforcement learning trajectories from fictional worlds using only codified rules (e.g., Prolog queries), requiring no LLM generation or human curation for the environment itself.
- Enhance Agentic Search Skills via Rule-Generated Transfer Learning.
The improved AI system can be fine-tuned using Reinforcement Learning (RL) on these synthetic environments to master core agentic search skills, specifically:
-
Decomposing complex multi-hop questions into sub-questions.
-
Formulating targeted, focused search queries (tags).
-
Retrieving information from a knowledge corpus.
-
Composing knowledge across multiple turns of interaction.
- Achieve Robust Transfer to Real Multi-Hop Benchmarks.
By training on rule-generated synthetic data, the resulting LLM agents can achieve significant performance gains (up to 7.1×) on real-world multi-hop search benchmarks (e.g., SynthWorlds, FRAMES). This means the agent learns generalized search strategy rather than memorizing specific facts from a fixed training corpus.
- Implement Emergent Search Scaling Behavior.
The system can be designed such that the model learns to allocate its search budget roughly linearly with question difficulty, demonstrating search scaling.
This allows for predictable performance gains as questions become more complex, regardless of the underlying model size (emergent in capable LLMs like Qwen2.5).
- Optimize Environment Complexity for Targeted Skill Acquisition.
The system can dynamically compose training environments along specific complexity axes:
-
Focus on pure linear hops to drive general transfer.
-
Use comparison questions to improve performance on real-world comparison tasks (e.g., matching attributes).
-
Avoid constraint questions that reward verbatim query shortcuts, which actively hinder agentic decomposition.
- Ensure Robustness to Noisy and Large-Scale Deployment.
The agents trained in these synthetic environments are robust to deployment challenges, including:
-
Overlapping or near-duplicate documents in a large pooled search index (simulating noisy search).
-
Out-of-domain evaluation benchmarks (where knowledge memorization is less useful than learned skills), as the training decoupled the agent from specific temporal data snapshots.
Abstract
Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply. Existing approaches rely on costly human-curated data or on LLM-generated environments that risk hallucinations and benchmark contamination. We show that LLMs can instead be trained into capable search agents using synthetic environments generated entirely by rules, whose generation requires no LLM and has zero marginal cost. We build PhantomEnvironments, multi-turn RL environments from fictional worlds, where agents must search a corpus of templated articles to answer multi-hop questions. Despite sharing no facts with the real world, these strikingly simple environments yield agents that transfer to real-world multi-hop search benchmarks, often outperforming real-world training data on newer benchmarks. Trained agents generalize to unseen fictional universes, and Qwen models learn to scale their search budget roughly linearly with question difficulty, suggesting emergent search scaling from environment interaction alone. Ablating environment complexity reveals that hop count drives transfer more than constraints or comparisons: even the simplest rule-generated environments are a surprisingly effective, free resource for training generalizable LLM agents.
Sources
- KARL: Knowledge Agents via Reinforcement Learning
- Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
- BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent
- Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
- Endless Terminals: Scaling RL Environments for Terminal Agents
- Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use
- PhantomWiki: On-Demand Datasets for Reasoning and Retrieval Evaluation
- The Llama 3 Herd of Models
- SynthWorlds: Controlled Parallel Worlds for Disentangling Reasoning and Knowledge in Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Measuring Massive Multitask Language Understanding
- An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Learning from Synthetic Data Improves Multi-hop Reasoning
- DeepDive: Advancing Deep Search Agents with Knowledge Graphs and Multi-Turn RL
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
- REASONING GYM: Reasoning Environments for Reinforcement Learning with Verifiable Rewards
- ZeroSearch: Incentivize the Search Capability of LLMs without Searching
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks