Recursive Agent Optimization

arXiv:2605.06639 · cs.LG, cs.AI, cs.CL, cs.MA · Submitted 2026-05-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Recursive Agent Optimization".

Tom: Recursive Agent Optimization (RAO) introduces a reinforcement learning approach for training recursive agents, which are models capable of spawning and delegating sub-tasks to new instances of themselves.

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So, wrapping up our discussion on "Recursive Agent Optimization," we’ve seen how this reinforcement learning approach trains agents to utilize recursion dynamically through a shared policy across an execution tree <ref:2605.06639#pg2>.

Jane: The title itself suggests that the focus is on optimizing the process of recursive agent training, specifically teaching models when and how to delegate sub-tasks <ref:2605.06639#pg1>.

Lu: It connects this work to hierarchical reinforcement learning because it uses natural language for generating subtasks instead of relying on fixed action abstractions, which is a significant connection <ref:2605.06639#pg1>.

Meng: From an engineering viewpoint, the implication is that we can build systems that solve problems by decomposing them recursively into manageable steps, which reduces the complexity burden on any single component <ref:2605.06639#pg0>.

Lalam: For me, Lalam, this advance suggests we are moving toward AI systems that exhibit a more sophisticated internal reasoning structure that can handle extremely long or intricate goals efficiently <ref:2605.06639#pg2>.

Tom: It really shows how training can be structured to induce an implicit curriculum over simpler subproblems, which is a neat way to learn <ref:2605.06639#pg1>.

Jane: The overall implication is that we're seeing methods that enable agents to tackle problems substantially harder than those they were initially trained on through recursive delegation <ref:2605.06639#pg0>.

Conclusion: Tom: So, we've been diving deep into this paper called "Recursive Agent Optimization," and now we’re getting to the closing thoughts from Tom and Jane about what this all means for the field.

Jane: Yeah, I think it’s important to start by looking at the title itself, "Recursive Agent Optimization," because it really captures the core idea of how these agents are being trained.

Lu: That title is spot on because it points directly to the mechanism: optimizing agents that can repeat themselves in a structured way, which is a really interesting concept for scaling up reasoning.

Meng: From an engineering standpoint, I see "optimization" as the key phrase here; it suggests they aren't just building these recursive systems from scratch, but they’re learning how to tune them effectively through this RL approach.

Lalam: And what this means simply is that we’re moving toward AI that can solve really complex problems by breaking them down and having the agent delegate work back to itself intelligently.

Tom: Exactly, Lalam, so it's about teaching the AI how to manage its own workload in a nested fashion, which is pretty wild when you think about the scale of tasks it can handle.

Jane: And looking at who wrote this paper, those authors have clearly put a lot of thought into how to design this system for real-world application rather than just theoretical exercises.

Lu: I've read the background, and their approach to defining that reward structure is what makes this paper stand out; it’s not just a simple "do the task" reward.

Meng: That reward design part seems crucial for me because if they can train the agents on quality of delegation instead of just speed, then we might see much more reliable results when deploying these recursive systems.

Lalam: I think the biggest implication here is that we could build AI systems that don't just process information linearly but can explore massive search spaces by constantly spawning specialized sub-agents.

Tom: It sounds like this research suggests a future where AI doesn't just answer questions in one go, but actually creates a miniature team of specialized thinkers to tackle the overall goal.

Jane: And we should keep an eye on how they’ve framed the results concerning context windows and task difficulty; that tells us how far this concept can take us before we hit new limitations.

Apurva Gandhi, Satyaki Chakraborty, Xiangjun Wang, Aviral Kumar

Carnegie Mellon University & Amazon AGI Labs

cs.LG, cs.AI, cs.CL, cs.MA

Submitted: 2026-05-07

Updated: 2026-10-02

Project page: https://apga.github.io/RAO

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 94/100

The gist: Recursive Agent Optimization (RAO) introduces a reinforcement learning approach for training recursive agents, which are models capable of spawning and delegating sub-tasks to new instances of

Key concepts

Recursive Agents
These are models that can solve complex problems by breaking them down into smaller pieces and assigning those pieces as new tasks to copies of themselves. They use a mechanism, like an async function call, to launch these sub-tasks and wait for their results.
RAO Reward Design
The reward system combines two parts: rewarding the agent for solving its own task successfully and rewarding it when its delegated sub-tasks are completed well. This encourages the agent to learn effective delegation strategies rather than just solving one big problem.
Policy Optimization Objective
The objective function trains a shared policy across all tasks—the main goal and all generated sub-goals. This structure creates an implicit curriculum, where learning simpler subproblems helps the model learn more complex ones effectively.

Terminology

Summary

Recursive Agent Optimization (RAO) introduces a reinforcement learning approach for training recursive agents, which are models capable of spawning and delegating sub-tasks to new instances of themselves. This method is significant because it teaches agents how to effectively utilize recursive inference—a capability that allows them to scale to longer contexts, generalize better, and achieve reduced wall-clock time compared to single-agent systems.

The gist

RAO trains a single LLM policy across all nodes of a dynamically generated execution tree, teaching the model both how to solve assigned tasks and how to generate useful delegated subtasks for spawned copies of itself.

Recursive Agent Inference and Implementation

Recursive agents are implemented as an extension of an agent that interleaves natural-language reasoning with code execution in a Python REPL. To support recursion, the action space is extended by exposing an asynchronous function, such as async launch subagent(goal,...) → Any2, which launches a new instance of the same policy on a delegated sub-task and returns its output to the parent. This allows parents to decide when and how to delegate—sequentially or concurrently using standard Python libraries like asyncio—and how to aggregate results. The system imposes explicit limits on recursion depth and environment steps for bounded computation, with implementations resembling Recursive Language Models (RLMs) but demonstrating positive results for recursion depths greater than one.

Recursive Agent Optimization (RAO) Reward Design

RAO trains all nodes in the recursively-generated execution tree jointly by combining a local reward defined at each node with policy optimization over the tree. The local node reward is defined as:

R(X, τX) = s̃(X, τX) success(X) / proxy + λ · 1/C(X) ∑ c∈C(X) s˜(c, τc) delegation bonus.

The first term rewards the agent for solving its own assigned task according to the available supervisory signal (s̃), while the second term rewards it when sub-tasks are successfully completed by its children, using the success rate of immediate children rather than raw counts to align with quality of delegation. The parameter λ controls the strength of this delegation bonus, which is most useful when an initial policy under-utilizes delegation.

Policy Optimization Objective

The objective function J(θ) optimizes a shared policy over both root tasks and policy-generated descendants: J(θ) = ∑ d=0 EX∼Dd(θ) h EτX∼πθ (·X) R(X, τX). This structure is effective because the same parameters are trained across a hierarchy of related tasks, which induces a structured intermediate supervision and an implicit curriculum over simpler subproblems. To compute the gradient, advantages A(τ(g)) are calculated by comparing each node’s local reward to a leave-one-out baseline computed from root rollout rewards (Eq. 3). Depth-level inverse-frequency weighting is used to mitigate the effect of deeper parts of the tree dominating learning by downweighting trajectories from depths that appear more frequently in the batch.

Experimental Results and Benefits

RAO training yields several benefits across benchmarks:

  1. It allows recursive agents to solve tasks that exceed the base model’s context window, even when training is limited to smaller windows (e.g., 8K in TEXTCRAFT-SYNTH).

  2. It enables generalization to substantially harder tasks than those seen during training by leveraging recursive delegation (up to 10 levels deep).

  3. When subproblems can be solved independently, recursion can reduce wall-clock execution time relative to non-recursive baselines (up to 2.5×).

  4. RAO improves training efficiency over single-agent training by exploiting recursive decomposition during learning.

The paper further notes that RAO learns task-appropriate delegation strategies rather than applying recursion uniformly, as evidenced by the fact that successful rollouts on OOLONG-REAL often have a maximum depth of 1, and on DEEPDIVE, the average maximum depth is 2.9. The results suggest that inference-time scaffolds should be trained to be used effectively rather than merely wrapped around a fixed model.

Related Work and Future Directions

RAO is positioned as the first work to jointly study recursive agent training with recursion beyond depth 1, end-to-end RL, asynchronous subagent execution, joint training of root agents and subagents, and evaluation on complex agentic applications. It connects RAO to hierarchical reinforcement learning through its use of natural language for generating subtasks rather than fixed action abstractions. Future work is suggested in training generalist recursive agents that transfer delegation strategies across domains and designing surrogate sampling procedures when full recursive rollouts are too costly.

Improvements for AI systems

Here are the specific improvements and capabilities derived from Recursive Agent Optimization (RAO) for AI systems:

  1. ​Extending Effective Working Memory and Context Handling: The system can solve tasks with horizons far beyond the model’s native context window (e.g., 256K tokens in OOLONG-REAL).

  2. ​Improved Generalization to Harder Problems: By leveraging recursive delegation (up to 10 levels deep), the agent can tackle significantly more difficult tasks than those it was trained on, as recursion naturally generates a curriculum of progressively harder subproblems.

  3. ​Enhanced Training Efficiency: RAO exploits recursive decomposition during learning, leading to faster training times compared to single-agent systems by providing dense process rewards and a self-induced curriculum over simpler subproblems.

  4. ​Adaptive Test-Time Compute Allocation: The agent learns when and how much to delegate based on task complexity (e.g., using crafting depth heuristics), allowing it to allocate test-time compute efficiently—using parallelism for independent subtasks and sequential execution for dependent ones, leading to up to 2.5× faster wall-clock time on complex tasks.

  5. ​Better Task Decomposition and Planning: The model is trained not just to solve the root task, but also how to formulate useful, concrete sub-tasks (e.g., Find X, Verify Y), enabling sophisticated planning into sequential or parallel execution trees rather than relying on a fixed orchestration scheme.

  6. ​Self-Organizing Agent Behavior: The system learns a divide-and-conquer strategy, allowing it to adapt its recursion depth dynamically to the difficulty of the problem, moving beyond uniform delegation toward task-appropriate decomposition (e.g., using depth heuristics).

  7. ​Robustness in Long Context Analysis: For very long documents (like D&D transcripts), the agent learns a reliable chunking strategy, avoiding simple heuristic failures like regex/string matching and instead using subagents to process context chunks sequentially or in parallel, ensuring accurate analysis of the entire input.

In essence, the improved AI system transforms from a single-pass solver into a sophisticated, self-organizing reasoning engine capable of tackling complex real-world problems that demand long-horizon planning and adaptive resource management.

Abstract

We introduce Recursive Agent Optimization (RAO), a reinforcement learning approach for training recursive agents: agents that can spawn and delegate sub-tasks to new instantiations of themselves recursively. Recursive agents implement an inference-time scaling algorithm that naturally allows agents to scale to longer contexts and generalize to more difficult problems via divide-and-conquer. RAO provides a method to train models to best take advantage of such recursive inference, teaching agents when and how to delegate and communicate. We find that recursive agents trained in this way enjoy better training efficiency, can scale to tasks that go beyond the model's context window, generalize to tasks much harder than the ones the agent was trained on, and can enjoy reduced wall-clock time compared to single-agent systems.

Sources

Related papers