RLMOpt: Adaptive Prompt Optimization via Recursive Language Models

arXiv:2608.10471 · cs.AI · Submitted 2026-08-11 · Read on arXiv

Subhash Bangalore Satheesha, Nirvik Pande, Deepthi Duddempudi, Bharath Dandala

Autonomize AI · Carnegie Mellon University

cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 75/100

The gist: RLMOpt is a prompt optimizer that makes the search policy itself language-model-driven through a recursive language model (RLM).

Terminology

Summary

RLMOpt is a prompt optimizer that makes the search policy itself language-model-driven through a recursive language model (RLM). The RLM agent operates over a tool-based environment, inspecting task information, analyzing failures, generating candidates, allocating evaluation budget, and deciding when to stop. A deterministic harness complements the agent by enforcing objective scoring, Pareto-based selection, and regression constraints.

The paper evaluates RLMOpt across four benchmarks spanning structured clinical information extraction (Chia), multi-hop question answering (HotpotQA), verifiable instruction following (IFBench-2025), and multi-turn tool-calling agents (BFCL). In a matched comparison at a single seed, RLMOpt obtains the best held-out score on all four benchmarks and leads the four-task mean (0.610 against 0.589 for GEPA). Repeating each benchmark across seeds yields 11 matched benchmark–seed comparisons, in which RLMOpt outperforms GEPA in 9 cases. Across all 11 runs, it never produced a prompt that underperformed its seed, whereas GEPA fell below its starting point twice. It is also more efficient, achieving these results with fewer search rollouts while producing prompts that are 27–79% the size of those produced by GEPA.

The results further show that optimization gains are determined primarily by the headroom available in the seed prompt, rather than by the search budget. Efficient optimization therefore depends on reaching the available headroom reliably and with minimal search.

The paper's contributions are:

  1. Introducing RLMOpt, a prompt optimizer whose outer search procedure is controlled by an RLM agent rather than a fixed optimization algorithm. The agent performs adaptive exploration through a tool interface, while a deterministic harness enforces evaluation and selection constraints.

  2. Developing a harness-controlled optimization framework for reliable adaptive search under limited evaluation data, including per-field scoring, Pareto-based selection, and regression constraints that prevent candidate updates from sacrificing existing capabilities.

  3. Demonstrating improved optimization efficiency across four benchmarks, where RLMOpt achieves the strongest held-out performance while requiring fewer downstream evaluations than existing prompt optimizers.

  4. Characterizing when prompt optimization provides value by analyzing the relationship between initial prompt quality and achievable improvement, showing that optimizer gains are determined by the remaining performance headroom of the underlying model.

  5. Extending adaptive prompt optimization to multi-component agent systems, where the optimized object includes tool-use behavior and interaction trajectories rather than only single-response prompts.

In the head-to-head comparison, RLMOpt achieves the best held-out score on all four benchmarks: Chia (0.568 vs 0.562 for GEPA), HotpotQA (0.727 vs 0.702), IFBench-25 (0.460 vs 0.440), and BFCL-mt (0.686 vs 0.653). The four-task mean is 0.610 for RLMOpt against 0.589 for GEPA. The margins over GEPA exceed one paired standard error on BFCL-mt (1.83) and HotpotQA (1.04), and sit inside it on Chia and IFBench-25.

In the multi-seed comparison, RLMOpt leads GEPA on three of the four benchmark means; Chia is the exception, where GEPA is ahead by 0.008. RLMOpt is also the more stable method: its HotpotQA standard deviation (0.019) is a third of GEPA's (0.056). Across all 11 runs RLMOpt never returns a prompt below its seed, while GEPA does twice, once on HotpotQA and once on IFBench-25.

Regarding compute efficiency, RLMOpt reaches these scores at lower compute than GEPA: fewer downstream rollouts on every benchmark, and on three of the four fewer tokens and less wall-clock time as well. The widest gap is on the agentic task, 1,854s against 5,344s. Smaller budgets recover much of the score: on HotpotQA a B=200 run reaches 0.717 against the B=500 run's 0.727, at a quarter of GEPA's rollouts.

RLMOpt's prompts are smaller than GEPA's while scoring higher, at 27–79% of GEPA-light's size on every benchmark. The largest difference is on BFCL multi-turn, where the winning instruction is 5,419 characters against 19,818.

The analysis shows that the key factor determining optimization gains is the amount of prompt-accessible headroom left by the seed. All four headline benchmarks fall in the headroom regime where the task LM can exploit substantial remaining headroom. Chia starts at 0.435 and improves by +0.133 to 0.568, while the agentic BFCL task improves from 0.602 to 0.686 (+0.084). HotpotQA and IFBench-2025 begin closer to the task LM's prompting ceiling and consequently show smaller gains of +0.027 and +0.030, respectively. In the ceiling regime, when the seed is already near the task LM's prompting ceiling, additional optimization provides little opportunity for improvement, and RLMOpt's no-regression floor returns the seed prompt rather than accepting a noisy candidate with a lower score.

The paper also discusses limitations, including that RLMOpt differs from baselines along several dimensions simultaneously, the stopping policy remains a source of variance, and the evaluation scope covers four benchmarks and two task LMs. The observed margins over GEPA are generally comparable to the per-run standard error, so the evidence for an advantage comes from consistency across benchmarks and seeds rather than from large margins on individual runs.

Improvements for AI systems

Improvements to AI systems:

  1. Adaptive, language-model-driven search policies: Replace fixed optimization algorithms (e.g., genetic or Bayesian) with a recursive language model agent that dynamically inspects task failures, generates candidate prompts, allocates evaluation budget, and decides when to stop—enabling more efficient and context-aware prompt optimization.

  2. Regression-constrained optimization: Enforce a no-regression floor during candidate selection, ensuring that any new prompt never underperforms the seed. This prevents catastrophic forgetting of existing capabilities during optimization, making the system safer for deployment.

  3. Pareto-based multi-objective selection: Use per-field scoring and Pareto dominance to select candidates that improve multiple dimensions (e.g., accuracy, instruction adherence, tool-use behavior) simultaneously, rather than optimizing a single aggregate metric.

  4. Headroom-aware optimization: Detect whether a seed prompt is in the headroom regime (large potential gains) or ceiling regime (near model's prompting ceiling). In the ceiling regime, the system automatically stops searching and returns the seed, saving compute and avoiding noisy degradation.

  5. Compact prompt generation: Produce prompts that are 27–79% smaller than baseline optimizers while scoring higher, reducing token costs and latency in downstream inference—critical for agentic and multi-turn systems.

  6. Multi-component agent optimization: Extend optimization beyond single-response prompts to include tool-use behavior, interaction trajectories, and multi-turn decision policies, enabling end-to-end improvement of complex agent systems.

  7. Compute-efficient search: Achieve comparable or better scores with fewer downstream rollouts (e.g., 1,854s vs 5,344s on agentic tasks) by using adaptive budget allocation and early stopping based on headroom analysis.

What the improved AI system can do:

  • Optimize its own instructions for new tasks with minimal human input, adapting search strategy in real time based on observed failures and remaining headroom.

  • Guarantee non-degradation of performance during optimization, making it safe for iterative self-improvement in production.

  • Handle complex, multi-step agentic tasks (e.g., tool calling, multi-turn dialogue) by optimizing not just the prompt but the entire interaction policy.

  • Operate under limited evaluation budgets, recovering most of the performance gain with 40% of the search budget (e.g., HotpotQA: 0.717 vs 0.727 at B=200 vs B=500).

  • Generate concise, high-performing prompts that reduce inference cost and latency, especially beneficial for large-scale or real-time applications.

  • Automatically identify when further optimization is futile (ceiling regime) and stop, saving compute and avoiding overfitting to noisy evaluation data.

Abstract

Prompt optimizers automate the search for prompts that improve language-model performance, but existing methods rely on a predefined optimization procedure: the algorithm determines which candidates to explore and how the search progresses, while the language model generates or refines prompt proposals. We introduce RLMOpt, a prompt optimizer that makes the search policy itself language-model-driven through a recursive language model (RLM). The RLM agent operates over a tool-based environment, inspecting task information, analyzing failures, generating candidates, allocating evaluation budget, and deciding when to stop. A deterministic harness complements the agent by enforcing objective scoring, Pareto-based selection, and regression constraints. We evaluate RLMOpt across four benchmarks spanning structured clinical information extraction (Chia), multi-hop question answering (HotpotQA), verifiable instruction following (IFBench-2025), and multi-turn tool-calling agents (BFCL). In a matched comparison at a single seed, RLMOpt obtains the best held-out score on all four benchmarks and leads the four-task mean (0.610 against 0.589 for GEPA). Repeating each benchmark across seeds yields 11 matched benchmark-seed comparisons, in which RLMOpt outperforms GEPA in 9 cases. Across all 11 runs, it never produced a prompt that underperformed its seed, whereas GEPA fell below its starting point twice. It is also more efficient, achieving these results with fewer search rollouts while producing prompts that are 27-79% the size of those produced by GEPA. Our results further show that optimization gains are determined primarily by the headroom available in the seed prompt, rather than by the search budget. Efficient optimization therefore depends on reaching the available headroom reliably and with minimal search

Sources

Related papers