Dynamic Optimizations of LLM Ensembles with Two-Stage Reinforcement Learning Agents

arXiv:2502.04492 · cs.CL · Submitted 2025-02-06 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Dynamic Optimizations of LLM Ensembles with Two-Stage Reinforcement Learning Agents".

Jane: Dynamic Optimizations of LLM Ensembles with Two-Stage Reinforcement Learning Agents introduces RL-Focal,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, Jane, we're looking at this paper today, "Dynamic Optimizations of LLM Ensembles with Two-Stage Reinforcement Learning Agents." The main idea seems to be proposing a two-stage reinforcement learning agent framework called RL-Focal designed to dynamically optimize how we build and use ensembles of Large Language Models for different tasks. It claims this approach can solve the instability we see when using static ensemble learners on models that are task agnostic.

Jane: That makes sense, Tom; it sounds like they're building a system that learns which ensemble strategies work best for specific problems while also handling disagreements between those models. The abstract says it addresses two key challenges: dynamically routing problems to different ensemble sets and fusing conflicting outputs from those selected ensembles regardless of which set is chosen.

Lu: It’s interesting how they tackle the selection problem first, by having a Decider RL-agent that learns to pick a small ensemble size m i out of N LLMs based on maximizing both error diversity and reasoning performance for task i. That sounds like a smart way to minimize redundant inferences right from the start.

Meng: From an engineering standpoint, minimizing inference redundancy is always a big deal because we're dealing with limited compute resources, so I’m curious how this selection process translates into actual resource savings in a live production environment.

Lalam: I think what excites me most is the Fusion RL-agent; it learns to resolve those reasoning conflicts within the selected ensemble, which means we get a better final answer even if the models disagree on something specific. This could really improve how our AI assists users by providing more coherent and robust outputs based on what's available.

Tom: Exactly, Lalam, and Jane pointed out that this framework is formulated as a decentralized partially observable Markov Decision Process, which is quite sophisticated for handling these sequential decisions across the two agents. The Decider agent’s reward function specifically includes a size-penalization constant alpha that pushes it to reduce the pool size when possible.

Jane: And then we have this focal diversity metric, lambda focal(E), which is designed to quantify error correlation among the component models in an ensemble E by averaging the focal negative correlation metric rho focal(M i; E). This seems like a clever way to capture how much error those models are sharing.

Paper summary: Lu: The paper formalizes this using an ambiguity decomposition that shows the quadratic error of the ensemble estimator is guaranteed to be less than or equal to the average quadratic estimators of its component estimators, which is formalized as (y - ens) squared = X i w(i - y) squared - X i w(i - ens), and that the greater the ambiguity, meaning higher error correlation between individual estimators, the lower overall error the ensemble may result.

Meng: That decomposition is helpful for understanding why we need diversity; it shows that high correlation leads to a larger covariance term in the bias-variance-covariance decomposition, which we want to keep low. But how do the agents actually learn this correlation measure, rho focal ?

Lalam: Well, the learning process involves both agents being trained using Multi-Agent Proximal Policy Optimization with a Centralized Critic Net to stabilize optimization by creating a bridge between the two agents with global information. That shared critic seems important for keeping the two stages aligned.

Tom: Right, and then we look at their loss function L RLFocal(theta) = one over n sum k=zero n r(k) t(theta i) t, clip, which incorporates clipped policy updates to ensure stability. This clipping helps keep the learning process from taking too big of a step in policy space.

Jane: The paper also mentions that the Central Critic observes a "global state s t =

o(one) t+one o(two) t: " and estimates the value V phi to stabilize optimization, which acts as this bridge between the Decider Agent and the Fusion Agent. This seems like a robust way to manage their interaction.

Lu: The authors are quite thorough in showing how this two-stage process works iteratively based on task-adaptive rewards and policy updates, suggesting a deep dive into adaptive control for ensemble management. Their approach moves beyond just picking the best model statically.

Meng: So, if I'm thinking about practical impact, this sounds like we could potentially deploy an AI system that automatically adjusts its own inference pipeline based on the complexity of the query in real-time, which would drastically cut down on wasted compute time.

Lalam: That’s a huge cultural shift for how we use these tools; instead of running a fixed setup, the system evolves with what it's doing, making our AI interaction much more efficient and responsive to the user's needs.

Paper summary: Tom: Moving on to the conclusion of this paper, the title "Dynamic Optimizations of LLM Ensembles with Two-Stage Reinforcement Learning Agents" really captures that dynamic nature—it’s about optimizing ensembles in a way that adapts over time using reinforcement learning. The authors are showing how RL can guide both the initial selection and subsequent fusion steps in a very structured, iterative manner.

Jane: And what I find most significant is their focus on moving away from static methods, which have been common in post-inference ensemble learning, by introducing these adaptive agents to handle task variations effectively. It moves the field toward systems that can truly respond to changing environments.

Lu: The implication here is that we might see a future where AI applications don't rely on one fixed ensemble setup but rather deploy a system capable of intelligently reconfiguring its component models for every new request, which is an exciting prospect for creative use cases.

Meng: From my perspective, the practical implication is that this framework could be used to manage complex AI workflows where different tasks require fundamentally different model compositions, leading to much better resource allocation across the board.

Lalam: And for me, it means a more reliable and nuanced experience from our AI products because it’s constantly tuning its internal structure to handle ambiguity and conflict in the outputs.

Tom: So, just wrapping up these points about "Dynamic Optimizations of LLM Ensembles with Two-Stage Reinforcement Learning Agents," this paper shows how RL can be applied not just to training models, but to orchestrating complex ensembles for dynamic performance. It’s a framework that learns to select and then fuse the right models for the job on the fly.

Jane: That really brings us back around to the core idea: using two distinct reinforcement learning agents, one for selection and one for fusion, to create an adaptive system that handles model diversity and reasoning conflicts intelligently.

Lu: It sets a new direction for how we think about ensemble methods, pushing it into the domain of dynamic control where adaptation is built directly into the decision-making process.

Meng: I’m looking forward to seeing how this translates from theoretical framework to something that runs reliably on actual production hardware, though I'll be watching their implementation details closely.

Lalam: I think the potential for our AI culture is huge here; having an AI that dynamically optimizes its own architecture for better task performance feels like a major step toward truly intelligent, self-tuning digital systems.

Conclusion: Tom: So we've been diving deep into RL-Focal, and now we're getting to the wrap-up section of this paper on "Dynamic Optimizations of LLM Ensembles with Two-Stage Reinforcement Learning Agents."

Jane: It’s been fascinating tracing how the Decider and Fusion agents work together to manage those ensemble selections and outputs.

Tom: Exactly, Jane, they’ve shown how these two distinct reinforcement learning agents can dynamically optimize which LLMs to use and then figure out the best way to combine their results based on the task at hand.

Lu: The authors really nail that idea of creating a system that adapts its entire ensemble strategy in real-time rather than sticking to one fixed setup.

Meng: From my side, it’s interesting how this translates into a system that can adjust its resource usage based on the query complexity, which is something I need to see more of in practice.

Lalam: For me, the most impactful vision here is seeing AI systems that don't just process data but actually tune their own internal structure for better task performance and cultural relevance.

Tom: That’s the core idea—moving from static methods to an adaptive orchestrator that learns how to select and fuse models dynamically based on what the specific query demands.

Jane: It really boils down to using reinforcement learning not just for training models, but for controlling the entire ensemble process itself.

Lu: The implications are huge because it opens up possibilities for creating AI applications that aren't locked into a single architecture but can evolve their component selection strategy as they encounter new kinds of problems.

Meng: I wonder if this level of dynamic adaptation could lead to much more efficient and tailored resource allocation across different AI workflows.

Lalam: I think the cultural impact is significant because it suggests a future where AI is inherently self-optimizing, constantly refining its approach to deliver better and more nuanced results for everyone.

Selim Furkan Tekin, Fatih Ilhan, Gaowen Liu, Ramana Rao Kompella

Georgia Institute of Technology · Cisco Systems

cs.CL

Submitted: 2025-02-06

Updated: 2026-09-29

Code: https://github.com/sftekin/rl-focal

Importance score: 80/100

The gist: Dynamic Optimizations of LLM Ensembles with Two-Stage Reinforcement Learning Agents introduces RL-Focal, a two-stage reinforcement learning agent framework designed to dynamically optimize ensembles

Key concepts

Decider RL-agent
This first agent is responsible for selecting which LLMs from a larger pool should be included in the ensemble. It learns to choose a small, effective subset of models by balancing two goals: maximizing error diversity and achieving high reasoning performance on the task.
Fusion RL-agent
This second agent takes the outputs from the selected ensemble and resolves any conflicting answers between those models. It learns to make a final, robust decision by choosing the output choice with the highest probability among all selected models.
Focal Diversity Metric ($\lambda_{focal}(E)$)
This metric measures how diverse the errors are among different LLMs within a chosen ensemble. It quantifies the 'error correlation' between models, helping the Decider agent select a set of models that are as varied as possible to improve overall performance.
DEC-POMDP
The framework is modeled as a Decentralized Partially Observable Markov Decision Process. This means both agents make decisions based on their limited, partial view of the current state, requiring them to learn complex coordination strategies to achieve optimal outcomes.

Terminology

Summary

Dynamic Optimizations of LLM Ensembles with Two-Stage Reinforcement Learning Agents introduces RL-Focal, a two-stage reinforcement learning agent framework designed to dynamically optimize ensembles of Large Language Models by learning which ensemble strategies are most effective for given downstream tasks. This research is significant because it addresses the instability inherent in applying static ensemble learners to task-agnostic LLMs by developing an adaptive system capable of (i) dynamically routing diverse problems to different ensemble sets and (ii) fusing conflicting outputs from those selected ensembles regardless of which set is chosen.

The gist

RL-Focal is a two-stage RL agent framework that routes and ensembles LLMs, where the Decider RL-agent learns to select an ensemble of small size (mi) among N LLMs by maximizing both error-diversity and reasoning-performance, while the Fusion RL-agent learns to resolve reasoning conflicts from different LLMs within the selected ensemble.

How it works

The framework is formulated as a decentralized partially observable Markov Decision Process (DEC-POMDP) with two agents: the Decider Agent and the Fusion Agent. The process involves iterative learning based on task-adaptive rewards and policy updates.

  1. The Decider RL-agent is responsible for model selection to minimize redundant inferences. Its state observation includes et,∥et∥1, σ1,..., σK, where et represents the current model pool and σ are diversity metrics calculated based on historical data within a window size T. The agent's action is a binary vector a(1)t ∈ (a1,..., aN), indicating whether each model should be included in the pool. The reward function R is defined as:

R(at, ot, y) = 1 if yˆfusion = y, −1 − α ·∥Et∥1 N otherwise. The size-penalization constant α forces the Decider Agent to decrease pool size.

  1. The Fusion RL-agent learns to resolve reasoning conflicts from the outputs of the selected ensemble. Its observation o(2)t includes rt−1,..., rt−T, p1,..., pm, where pi = Mi(xt) is the probability vector that model Mi assigned to choices for input xt. The agent's action is to make the final decision at = ˆyfusion by choosing the choice with the maximum probability: at = arg maxa∈A πθ2(a o(2)t).

Key Mechanism: Focal Diversity Metric

To enable effective ensemble pruning, RL-Focal introduces diversity metrics. The focal negative correlation metric ρ focal is used to quantify error diversity among component models concerning each model within the ensemble. The focal diversity metric λ focal(E) is the average of these scores across all models in the ensemble E:

λ focal(E) = 1/N Σ Mi∈E ρ focal(Mi; E). This metric is designed to capture error correlation among multiple LLMs. The objective for the Decider Agent is defined as maximizing this diversity: max E∈E λ focal(E).

Learning and Optimization

Both agents are trained using Multi-Agent Proximal Policy Optimization (MAPPO) with a Centralized Critic Net to approximate the value of the current state. The loss function LRLFocal incorporates clipped policy updates to ensure stability:

LRLFocal(θ) = 1/n Σ Xn k=0 min rˆ(k)t(θi)Aˆt, cliprˆ(k)t(θi), 1 − ε, 1 + ε Aˆt. The Central Critic observes the global state st = [o(1)t+1, o(2)t] and estimates the value Vϕ to stabilize optimization by creating a bridge between two agents with global information.

Performance and Cost Efficiency

Extensive evaluations on five benchmarks show that RL-Focal achieves a performance improvement of 8.48% with an ensemble of small size compared to the best individual LLM in a pool and offers stronger robustness. The framework is noted for being significantly less number of parameters, having no supervised training, and exhibiting less inference time latency compared to other methods like LLM-Blender or TOPLA, demonstrating superior cost-efficiency. Furthermore, the theoretical analysis proves that maximizing focal diversity leads to a low (or even negative) error covariance between member models, thereby reducing the overall ensemble error by minimizing the term (1 − 1/N)Covar in the bias-variance-covariance decomposition.

Online Adaptation

The framework supports online adaptation through periodic updates. The Decider and Fusion agents are updated every nupdate queries to maintain stability and adaptability to evolving query tasks, enabling RL-Focal to "timely adapt to a changing environment.

Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed Dynamic Optimizations of LLM Ensembles with Two-Stage Reinforcement Learning Agents (RL-Focal). The core innovation lies in moving ensemble learning from static, pre-defined strategies to a dynamic, task-adaptive reinforcement learning framework.

Here are the specific improvements that can be made to AI systems using this paper, and what these improved systems can achieve:


) Dynamic Optimization of LLM Ensembles with Two-Stage Reinforcement Learning Agents (RL-Focal) Improvements:

  1. Dynamic Ensemble Composition (Decider Agent): The system will dynamically select the optimal subset size and specific models from a large pool of N LLMs based on the current query's task characteristics.

  2. Task-Adaptive Model Selection via Focal Diversity Metric: Instead of relying on simple performance metrics, the Decider Agent will actively prune ensemble combinations to maximize focal diversity (low error correlation) among the selected models, ensuring that the chosen set possesses complementary strengths rather than redundant capabilities.

  3. Conflict Resolution Fusion (Fusion Agent): The system will employ a second RL agent that learns to intelligently resolve reasoning conflicts and synthesize potentially conflicting outputs from the dynamically selected small ensemble, resulting in a single, high-quality final decision.

  4. Adaptive Learning Framework: The entire system operates as a meta-learning framework where both agents iteratively learn from environmental feedback (task-adaptive rewards) and previous experimental knowledge to adapt their policies online, allowing it to evolve with changing query distributions or domain shifts without requiring full retraining on every new task.

) Capabilities of the Improved AI System:

  1. Highly Robust and Adaptive Reasoning: The system can tackle complex, multi-faceted reasoning tasks (e.g., open-ended math problems like GSM8K) by dynamically selecting the most appropriate combination of specialized models (e.g., one for factual recall, one for complex symbolic manipulation).

  2. Cost-Efficient Inference: By employing a small ensemble size and dynamically pruning unnecessary models, the system can achieve performance gains (up to 8.48% improvement) while operating at a significantly lower computational cost compared to large, fixed ensembles or traditional supervised methods.

  3. Task-Specific Alignment and Safety: The Decider Agent can be fine-tuned for specific alignment objectives (Helpfulness, Safety, Truthfulness), allowing the system to dynamically select models optimized for a query's required ethical or factual constraints (as demonstrated in Table 6).

  4. Autonomous Decision Making Under Uncertainty: In scenarios where multiple base models disagree significantly (e.g., complex inference tasks like the Murder Mysteries dataset), the Fusion Agent provides a high-confidence final answer by resolving these conflicts, leading to superior accuracy compared to majority voting or single-model outputs (as shown in Table 10).

  5. Continuous Self-Improvement: Because the framework is RL-based, it can continuously update its model selection and fusion policies online based on real-world performance data, enabling the AI system to autonomously refine its ensemble strategy over time without manual intervention.

Sources

Related papers