MoRSE: Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts

arXiv:2608.09251 · cs.MA, cs.AI, cs.CL, cs.LG · Submitted 2026-08-10 · Read on arXiv

Peiwen Li, Shiyang Zhang, Yangtian Zhang, Sizhuang He, David van Dijk, Rex Ying

Yale University

cs.MA, cs.AI, cs.CL, cs.LG

Submitted: 2026-08-10

Updated: 2026-08-11

Comments: 25 pages, 8 figures, 9 tables

Code: https://github.com/lpwpower/MoRSE

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: MoRSE: Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts Abstract Large language model-based multi-agent systems have recently shown strong potential for complex, long-horizon

Terminology

Summary

MoRSE: Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts

Abstract

Large language model-based multi-agent systems have recently shown strong potential for complex, long-horizon tasks. However, existing methods mainly rely on coarse prompt-level differentiation without parameter adaptation for diverse subtasks, resulting in insufficient inter-agent heterogeneity and limited specialized capability that bottleneck performance on tasks with complex requirements. To address this, we introduce a Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts (MoRSE) that distinguishes agents with (role, subtask)-conditional specialization at both the task structure and parameter levels. To make agents' responsibility explicit at the task structure level, we formulate a task-oriented multi-agent system that decomposes each task into a dependency-aware Directed Acyclic Graph of subtasks and assigns each agent a specific (role, subtask), introducing task-level specialization across collaborating agents. Additionally, to address the diverse role and subtask parameter adaptation demands, we propose a dynamic Mixture of (role, subtask) LoRA Experts module with a prototype-based semantic router for subtasks, augmenting agents with parameter-level specialization on a shared LLM substrate cost-effectively. Then, to co-optimize experts and router stably under sparse task rewards, we further propose a hierarchical group-relative policy optimization with two-layer credit assignment that isolates expert updates from the cross-route variance introduced by routing decisions, disentangling expert quality from routing quality. Experiments on code-generation benchmarks across three backbones demonstrate the effectiveness of our approach, with improvements in both whole-task and step-wise performance, and the gains from trained specialization generalize across held-out task categories and domains.

Introduction

LLM-based multi-agent systems (MAS) have emerged as an effective paradigm for addressing complex, long-horizon tasks, by orchestrating multiple role-specialized agents into structured multi-stage pipelines. This paradigm has demonstrated strong performance across diverse domains such as collaborative software development with simulated engineering teams and multi-step reasoning through inter-agent debate and critique, with recent work further pushing collaboration to populations of hundreds or even thousands of agents. Yet simply adding more agents to a shared backbone yields diminishing returns; the real bottleneck is how distinctly each agent contributes to the joint solution.

Existing MAS methods pursue such distinctness mainly through prompt-level role descriptions over frozen base models; but without explicit subtask assignment, agents tend to produce redundant, overlapping responsibilities and correlated outputs that saturate performance quickly as the system scales, leaving inter-agent heterogeneity insufficient at the task structure level. A growing line of work moves beyond prompts toward parameter-level adaptation, including role-conditioned adaptation in a single-agent setting, and multi-agent reinforcement learning (RL) frameworks like AT-GRPO and MARTI that extend RL optimization to multi-agent settings. Yet these methods adapt mainly at the predefined and limited role level; the diverse subtasks within a single complex task impose distinct demands that a single shared base model cannot satisfy, leaving inter-agent heterogeneity insufficient at the parameter level and bottlenecking performance on tasks with complex requirements.

To address these two levels of insufficiency, the paper studies sufficient agent heterogeneity on complex open-ended multi-agent collaboration tasks from both the task structure and parameter levels. MoRSE explicitly manages subtasks within the whole task and learns role- and subtask-specific parameter specialization across collaborating agents cost-effectively on a shared backbone via a stable optimization scheme. Here, an agent's role denotes its reusable function drawn from a small discrete set (e.g., merging upstream context or executing a step), whereas its subtask denotes the instance-specific unit of work described in natural language; the two axes impose complementary specialization demands and jointly define each agent's responsibility.

Specifically, the paper introduces a task-oriented MAS framework that exposes each agent's (role, subtask) responsibility, reducing task-structure-level redundancy and providing a substrate for fine-grained parameter-level specialization. Inspired by graph-based planning, the model decomposes each task instance into a dependency-aware Directed Acyclic Graph (DAG) of subtasks and assigns each agent a specific (role, subtask). Building on this framework, the paper further proposes a dynamic mixture of (role, subtask) LoRA experts that enhances inter-agent heterogeneity at the parameter level. It attaches role and subtask experts selected per agent call by a learned prototype-based semantic router. Jointly optimizing the experts and router under sparse, open-ended task rewards, however, can entangle expert quality with routing quality. Finally, the paper proposes a hierarchical group-relative policy optimization with two-layer credit assignment that splits the task-level reward into a within-route advantage updating experts within a routed combination and a cross-route advantage updating the router across combinations. The paper shows analytically that it reduces gradient variance for expert updates and empirically that it stabilizes co-optimization, reliably realizing the parameter-level heterogeneity gains.

Experiments on complex code-generation benchmarks across multiple LLM backbones demonstrate that MoRSE achieves the best held-out test scores over all baselines in both whole-task and step-wise performance. On held-out task categories, MoRSE further surpasses both its untrained variant and a standard fixed-LoRA fine-tuning baseline, indicating that the learned role-subtask experts generalize better under distribution shift.

The contributions are summarized as follows:

  • A task-oriented MAS framework with dependency-aware DAG decomposition and per-agent (role, subtask) assignment as substrate to reduce redundancy at the task-structure level.

  • A dynamic mixture of (role, subtask)-conditioned LoRA experts on a shared backbone with a prototype-based semantic router for effective parameter-level specialization.

  • A hierarchical group-relative policy optimization with two-layer credit assignment that disentangles expert quality from routing quality to stabilize co-optimization.

Preliminaries

Empirical Motivation

Two preliminary studies on SRDD with Qwen3-4B are conducted. The first measures the redundancy of role-prompted MAS to confirm the heterogeneity deficit is real, and to ground the adoption of subtask-based decomposition as the framework substrate. The second probes the limitation of parameter-efficient fine-tuning on a single base LM in satisfying the heterogeneous learning demands of diverse roles and subtasks, directly motivating the design of Mixture of Role-Subtask LoRA Experts.

Redundancy in role-prompted, graph-structured MAS. For each SRDD task instance involving n collaborating agents, heterogeneity is measured by the mean pairwise cosine similarity of node-level prompts and outputs under a frozen text encoder. Both MacNet variants collapse the per-sample distribution toward Simoutput ≈ 1 (median shift ∆out-prompt = +0.14 for both, indicating that the review stage does not relieve the redundancy). In contrast, MoRSEbase spreads the per-sample outputs apart (∆ = −0.15, lowering median Simoutput from 0.75 to 0.60). However, even with subtask decomposition, the median Simoutput still sits around 0.60, confirming that structural decomposition alone does not close the heterogeneity deficit—motivating the parameter-level adaptation.

Role-subtask mismatch in the base LM. The paper asks whether different roles and subtasks require mismatched rank-ρ adaptations of the base LM that no single shared LoRA can simultaneously provide under its rank-ρ bottleneck. Adapting subspace-based representation analyses to per-(role, subtask) groups, the paper takes execute and merge as two example roles: real role labels yield substantially larger pairwise rank-ρ subspace divergence (1 − ∥Va⊤ Vb∥2F/ρ, ρ=8) than a label-shuffled binary baseline across the deeper 8 transformer layers (gap ∆≈0.40). Fixing the role to execute and varying subtasks, the paper further clusters SRDD subtask descriptions into K=8 groups and applies the same diagnostic across the same deeper layers: real cluster labels yield divergence ∼0.75 vs. ∼0.45 under label-shuffled baselines, with σ-bands non-overlapping at every layer and visually stable cluster-pair structure across layers. Together, these results reveal specialized adaptation needs along two axes—roles (persistent, process-level) and subtasks (fine-grained, semantic)—directly motivating a dynamic mixture of LoRA experts with a router for (role, subtask) specialization on a shared base model.

Problem Formulation

A complex task instance T admits decomposition into V interdependent subtasks vi Vi=1 with natural-language descriptions si Vi=1. Each subtask vi is handled by a collaborating agent ai with role ri drawn from a small discrete role set R. Agent ai produces an artifact yi conditioned on its (role, subtask) pair (ri, si) and the upstream artifacts yj j∈Pa(i), where Pa(i) ⊆ 1,..., V indexes the subtasks that vi depends on.

All agents share a frozen base language model with parameters θ0, and (role, subtask)-conditional heterogeneity is realised via parameter-efficient adaptation θi = θ0 + ∆θi(ri, si) with dim ∆θi ≪ dim θ0, so that yi ∼ p(· ri, si, yj j∈Pa(i); θi). Given a reward R(yi Vi=1) defined over the joint outputs, all learnable parameters Θ of the chosen ∆θ parameterisation are jointly optimised to maximise expected reward over a task distribution D.

Methodology

The framework is instantiated with three tightly coupled components realising (role, subtask)-conditional specialization at both the task-structure and parameter levels.

Task-Oriented Multi-Agent System (ToMAS)

To address the task-structure-level redundancy, each task instance is explicitly decomposed into subtasks structured by dependencies, inspired by graph-based planning, and the resulting DAG is executed with a step-level rule-based verifier at every node. This produces two signals required by the downstream modules: per-node (ri, si) labels that condition parameter specialization, and step-level rewards ui i∈V that supply the training signal for that specialization via per-step credit assignment.

DAG-based task decomposition and execution. Given a task instance T, a planner dynamically produces a per-instance dependency-aware DAG G = (V, E), where each node vi ∈ V carries a subtask description si and each directed edge (vi, vj) ∈ E indicates that subtask j depends on the artifact yi produced by subtask i. The DAG is executed in topological order: at each node vi, a Merger first aggregates the upstream artifacts yj j∈Pa(i) into a coherent context Ci ∼ p(· r, T, yj j∈Pa(i); θir), r=merge (which reduces to yj when Pa(i) = 1 and to ∅ when Pa(i) = ∅); the Executor then produces the current artifact yi ∼ p(· r, T, si, Ci; θir), r=execute. Each agent call carries its own role r, which selects a role-conditioned LoRA composition. Training proceeds sequentially to expose per-subtask credit signals, while at inference the DAG also allows independent branches to run in parallel for efficiency.

Rule-based verifier for reward signals. A task-completion-only reward is too sparse to provide sufficient supervision for fine-grained per-subtask credit assignment; therefore a rule-based verifier is applied at every DAG node to produce a step reward ui on each output artifact yi, which also blocks yi from downstream propagation upon failing task-specific validity criteria (e.g., output format, basic correctness), preventing error compounding across the DAG. These step rewards ui i∈V feed downstream policy optimization.

Prototype-Routed Mixture of Role-Subtask LoRA Experts (MoLE)

Motivated by the role-subtask architecture mismatch under parameter-efficient fine-tuning, the paper addresses it via learnable specialization on the shared frozen backbone θ0. A dynamic mixture of dual-factorized LoRA experts is employed—a role expert pool Φr activated per role, and a subtask expert pool Φs from which a subset is selected per subtask by a lightweight prototype-based semantic router—keeping training and inference cost close to a single-model system.

Mixture of dual-factorized LoRA experts. The expert pool is partitioned into role experts Φr (one per role r ∈ R) and subtask experts Φs (shared across subtasks), all attached as LoRA pairs to the frozen backbone θ0. For each agent call at node vi with role ri, the active set Ei combines the role expert e(ri) with a subtask-expert subset Es,i: Ei = e(ri) ∪ Es,i, e(ri) ∈ Φr, Es,i ⊂ Φs. This active set yields the LoRA-composed projection at each adapted layer l: W(i)l = Wl + (α/ρ) Σk∈Ei wk · Bl,k Al,k, where (Al,k, Bl,k) is the rank-ρ LoRA pair of expert k, α is a scaling, and uniform aggregation weights wk = 1/Ei are used. The subtask-expert subset Es,i is selected by the prototype router, except at Merger join nodes (Pa(i) > 1), where it is inherited from upstream Executors as Es,i(merge) = ∪j∈Pa(i) Es,j(execute), exposing the merge role to the same subtask-specific adaptation directions that produced the upstream artifacts.

Dynamic prototype-based semantic router. A router πψ scores subtask experts against per-expert learnable prototypes, inducing a soft semantic partition over the open-ended subtask space rather than a hard cluster-to-expert assignment. The subtask description si is encoded via a frozen embedder, then a trainable projection hη is applied to produce hη(si) ∈ RD in the same latent space as the per-expert prototypes P ∈ RKs×D (Ks = Φs): zk = (hη(si)⊤pk)/(∥hη(si)∥2∥pk∥2), πψ(k si) = softmax(z)k, where pk ∈ RD is the k-th subtask expert's prototype, zk is its cosine similarity to the projected subtask embedding, and πψ(k si) is the probability of routing si to expert k; top-K is sampled from πψ(· si) stochastically during training and greedily at inference.

Trainable parameters. The trainable components are the LoRA expert matrices (Al,k, Bl,k) for all experts in Φ:= Φr ∪ Φs, and the router parameters ψ = (η, P) comprising the embedding-projection parameters η and the prototypes P. Backbone θ0 remains frozen throughout training. Together Φ and ψ instantiate the abstract learnable parameters Θ = Φ ∪ ψ.

Hierarchical GRPO with Two-Layer Credit Assignment (HGRPO)

Co-optimizing Φ and πψ under step rewards ui i∈V entangles their updates under a single shared advantage, exposing each to uncontrolled reward variation and inflating gradient variance. This is addressed with a two-layer credit assignment that uses separate within-route and cross-route advantages to update experts and the router: the within-route conditional baseline strictly reduces the LoRA-expert gradient variance compared to standard GRPO (a single shared advantage for both expert and router updates), while the cross-route formulation supplies the router with a well-scaled per-route signal.

Two-layer credit assignment. At each node vi during training, B routes are sampled from πψ(· si) (each fixing an expert combination Ei(b)), M candidates yi(b,m) are generated per route, and step rewards ui(b,m) are obtained from the verifier. Two advantages are formed with conditional baselines: Awithin(b,m) = (ui(b,m) − µ(b))/(σ(b) + ϵ), A(b)cross = (µ(b) − µ̄)/(σC + ϵ), where µ(b), σ(b) are the within-route mean and standard deviation, µ̄ = (1/B) Σb µ(b) is the across-route baseline, and σC is the standard deviation of µ(b) Bb=1. Following GRPO, Awithin updates the LoRA experts under fixed Ei(b) and Across updates the router via the route log-likelihood: LLoRA = −Σb,m Awithin(b,m) · log p(yi(b,m) T, ri, si, Ci; Ei(b)), Lrouter = −απ Σb A(b)cross · log πψ(Es,i(b) si), with απ balancing router and LoRA gradients. The within-route baseline µ(b) removes the cross-route component σC2 from the LoRA-expert gradient variance; the cross-route signal Across supplies the router with a well-scaled per-route gradient.

Variance reduction analysis. Decompose the reward variance into within- and cross-route components, σW2:= Eb[Varm(u(b,m) b)] (within-route), σC2:= Varb[µ(b)] (cross-route), with µ(b):= Em[u(b,m) b] the mean reward of route b (so Var(u) = σW2 + σC2).

Proposition 3.1 (HGRPO LoRA-expert variance reduction). HGRPO and standard GRPO share the same expected gradient; at finite group sizes M, B, the HGRPO LoRA-expert gradient has strictly lower variance, Var[gH LoRA] = Var[gLoRA] − σC2 · E∥∇Φ log p∥2. This isolates the LoRA expert from cross-route variance it cannot influence, attributing reward signals to expert quality rather than routing decisions; the construction applies the classical conditional-baseline variance reduction in a two-layer form to the router–expert credit structure. Proposition 3.1 is stated for the unnormalized advantages, while the implemented estimator further standardizes each advantage by the per-route scale σ(b) + ϵ. A short bridge argument shows that the variance reduction persists for this normalized estimator in a reweighted form (Corollary G.1).

In conclusion, the three modules jointly instantiate the abstract framework: ToMAS supplies the decomposition vi, Pa(i) and step rewards ui; MoLE realises the (ri, si)-conditional adaptation ∆θi on the shared backbone θ0; and HGRPO optimises the learnable parameters Θ = Φ∪ψ with provable variance reduction on the LoRA-expert gradient.

Experiments

Datasets. Two code-generation benchmarks with complex task requirements are used: SRDD (1,200 software-requirement descriptions across 40 application categories) and SciCode (80 scientific computing problems spanning 5 scientific domains, decomposed into 338 subproblems with executable unit tests at the step and problem level).

Baselines. Three multi-agent baselines spanning role-based, graph-structured, and search-based designs are compared — ChatChain, MacNet, and AFlow — together with a single-agent reference (Base model) that runs each backbone alone. All methods share the same Qwen3-4B-Instruct, Llama-3.1-8B-Instruct, and Gemma-4-31B-IT backbones, prompts, and decoding configuration, with each method's native revision mechanism capped at the same maximum number of attempts.

Evaluation metrics. On SRDD, Exec (↑) (the fraction of generated codebases that compile and run end-to-end) and ECI (↑) (a [0, 1]-valued composite over completeness, executability, and consistency, summarised as the per-sample mean (ECI Mean) and product (ECI Product)) are reported. On SciCode, Step Pass (↑) (micro step-level pass rate across all subproblems), Mean Step Pass (↑) (per-problem step accuracy averaged across problems), and Problem Pass (↑) (fraction of problems for which all steps pass) are reported.

Main Results

Table 1 reports two comparisons. The All set scores the inference performance of methods on the full benchmark without training; the Test split is held out from MoRSE's training for direct comparison against baselines, where MoRSE attains the column-best score on every reported test cell.

SciCode provides both process-level performance (Step Pass, Mean Step Pass) and end-to-end success (Problem Pass); SRDD complements this with overall quality aligning with the task requirements (ECI: completeness and task-description alignment given executability). On the All set, the untrained MoRSEbase achieves the best overall performance comparing with all baselines, showing that the role-subtask DAG decomposition framework adds value without any training. On the held-out Test split, the trained MoRSE further lifts MoRSEbase on all metrics, confirming the value of specialized parameter-level adaptation.

As a side observation, prior MAS without task-specific adaptation can lag below the single-agent reference, and the gap widens on stronger backbones whose base is near-saturating. Also, the post-training uplift is smaller on the strongest backbone (Gemma-4-31B), where the base model is closer to saturation and, with the LoRA configuration fixed across backbones, the trainable fraction shrinks with model width (0.170% on Qwen3-4B, 0.109% on Llama-3.1-8B, 0.048% on Gemma-4-31B), making the adaptation relatively less expressive. The structure and the training remain complementary there, as the untrained framework sometimes falls slightly below the single agent on the held-out split while training recovers the gap and matches or exceeds it on every metric.

To conclude, on both process and end-to-end aspects, the role-subtask DAG framework improves inference performance at the task-structure level, and specialized adaptation provides a further substantial uplift at the parameter level.

Deeper Analysis

Ablation study. Table 1 shows the standalone effect of the role-subtask DAG framework via MoRSEbase. Table 2 ablates the remaining two core modules: (i) the MoLE architecture and (ii) HGRPO training.

C vs. D (fixed MoLE, swap standard GRPO → HGRPO) isolates the training: HGRPO substantially recovers and extends the gain, directly verifying the hierarchical-credit design. B vs. C (under standard GRPO) reveals the converse: plugging in MoLE without hierarchical credit is fragile, indicating that experts and router can be challenging to jointly optimize under a standard optimization method. The training dynamics behind this fragility, where the standard-GRPO surrogate destabilizes while HGRPO remains controlled across epochs, are shown in Appendix H. A vs. D then confirms the joint effect with substantial gains across all metrics; the gain is further validated by B vs. D, where MoRSE surpasses a non-routing fixed-LoRA baseline trained with standard GRPO whose rank is matched to MoRSE's activated parameters per call.

Together, MoLE and HGRPO are tightly coupled and both are required to deliver the full improvement.

Parameter-level heterogeneity gain from MoLE. Figure 4 shows that MoLE further reduces residual node-output redundancy at the parameter level: trained dual role×subtask LoRA experts shift the per-sample Simoutput distribution toward more diversity, with the majority of samples landing below the diagonal, letting each node speak in its own voice rather than echoing its neighbours.

Out-of-distribution (OOD) generalization. A central premise of MoLE is that role-subtask decomposition exposes potentially generalizable units—different tasks may share recurring subtask skills, so LoRA experts learned on training categories have the potential to compose for unseen ones. This is tested on category-level held-outs: SRDD withholds 8 of 40 categories; SciCode trains on Physics+Math+Material Science and tests on Chemistry+Biology.

Figure 5 contrasts four configurations on the OOD test: prior role-based DAG MAS MacNet, our untrained MoRSEbase (A), a standard non-routing LoRA with GRPO fine-tuning baseline (B = MoRSE w/o MoLE), and the full trained MoRSE (D). Two findings emerge across both backbones and both datasets. (i) Training transfers to OOD: trained MoRSE consistently improves over the untrained MoRSEbase, indicating that subtask-level skills learned on one category help held-out ones. (ii) Dynamic role-subtask routing matters at OOD: MoRSE also surpasses the standard fine-tuning baseline, indicating that dynamic expert routing generalizes to held-out subtasks better than a fixed LoRA.

Related Work

LLM-based multi-agent systems. Existing MAS coordinate cooperating agents via prompting protocols, recently with dependency-aware DAGs; ToMAS instead exposes per-node (role, subtask) structure as the substrate for downstream parameter-level adaptation in MoLE.

Mixture of LoRA experts. Mixture-of-LoRA-Experts combines parameter-efficient adapters with sparse gating, and recent GRPO variants apply policy-gradient RL to MoE/LoRA-MoE under a single shared advantage; HGRPO instead performs a bi-level credit decomposition with within-route and cross-route conditional baselines, sharing the same expected gradient but with strictly lower variance at finite group sizes.

RL for multi-agent LLM systems. Recent multi-agent LLM RL operates on static topologies and parameterizes only the role axis under a single shared advantage; a parallel line extends GRPO along the rollout-time axis. MoRSE differs along three orthogonal axes—per-instance dynamic DAGs, dual role-subtask LoRA factorization on a shared backbone, and structural-axis credit decomposition via HGRPO's bi-level baselines—and is composable with rollout-time hierarchy.

Conclusion

MoRSE, a task-oriented multi-agent system, addresses the heterogeneity bottleneck of role-prompted MAS via (role, subtask)-conditional specialization at both task-structure and parameter levels. ToMAS decomposes each task into a dependency-aware DAG with per-agent (role, subtask) labels for task-structure-level specialization; MoLE attaches a dynamic mixture of role and subtask LoRA experts with a prototype-based semantic router on a shared backbone, addressing the diverse role and subtask demands that a single shared base model cannot satisfy; HGRPO stably co-optimizes experts and router via two-layer credit assignment that disentangles expert quality from routing quality. Experiments across three diverse backbones show notable whole-task and step-wise improvements, with better generalization across held-out task categories and domains.

Improvements for AI systems

Improvements to AI Systems Based on MoRSE

  1. Dynamic Role-Subtask Specialization: Replace monolithic fine-tuning with a dual-factorized LoRA expert pool—one expert per persistent role (e.g., execute, merge) and a shared pool of subtask experts selected by a prototype-based semantic router. This allows a single frozen backbone to adapt its parameters per agent call based on both the agent's function and the specific natural-language subtask, enabling fine-grained specialization without retraining the full model.

  2. Task-Structure-Level Heterogeneity via DAG Decomposition: Automatically decompose complex tasks into dependency-aware Directed Acyclic Graphs (DAGs) of subtasks, assigning each agent a unique (role, subtask) pair. This reduces output redundancy and overlapping responsibilities among agents, ensuring each agent contributes distinct, non-correlated artifacts—improving whole-task performance on long-horizon, multi-step problems.

  3. Stable Co-Optimization of Experts and Router via Hierarchical Credit Assignment: Use a two-layer advantage structure (within-route and cross-route) during reinforcement learning to separate expert-quality updates from routing-quality updates. This provably reduces gradient variance for expert updates, preventing the instability that arises when jointly optimizing a router and expert pool under sparse rewards—enabling reliable convergence in multi-agent RL.

  4. Composable Subtask-Level Generalization: Train subtask experts on a subset of task categories, then compose them for held-out categories via the semantic router. Because the router maps unseen subtask descriptions to learned prototypes, the system transfers specialized skills across domains (e.g., from physics to biology) without retraining, outperforming fixed-LoRA baselines under distribution shift.

  5. Step-Level Verifier with Error Blocking: Integrate a rule-based verifier at each DAG node to generate per-subtask rewards and block failed artifacts from propagating downstream. This prevents error compounding across long chains, provides dense training signals for credit assignment, and improves step-wise pass rates on benchmarks with executable unit tests.

  6. Cost-Effective Parameter Adaptation on Shared Backbone: Attach only LoRA adapters (rank-8) to a frozen base model, with active parameters per call matching a single LoRA's size. This achieves multi-agent heterogeneity at a fraction of the memory and compute cost of full fine-tuning, scaling to larger backbones (e.g., 31B parameters) with minimal trainable overhead.

What the Improved AI System Can Do:

  • Solve complex, multi-step code-generation and scientific-computing tasks with higher end-to-end success rates and step-level accuracy than prompt-only or fixed-LoRA multi-agent systems.

  • Adapt to new task categories or domains without retraining, by recombining learned subtask experts via semantic routing.

  • Maintain stable training under sparse, open-ended rewards, avoiding the collapse or oscillation seen with standard GRPO when optimizing routers and experts jointly.

  • Operate efficiently on resource-constrained settings, since only small adapters are trained on a frozen backbone, while still achieving specialized behavior across diverse roles and subtasks.

Sources

Related papers