MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale

arXiv:2608.10333 · cs.LG · Submitted 2026-08-11 · Read on arXiv

Yuhang Yao, Zeyu Wang, Wanyi Chen, Tongyun Yang, Yuhang Han, Jie Xiao, Chengke Bao, Tianyi Zhao, Lynn Ai, Eric Yang, Tianyu Shi

Gradient · Soochow University · Independent Researcher · Shanghai Jiao Tong University · Carnegie Mellon University · University of California, Los Angeles

cs.LG

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: Preliminary version in CAIS RL-Eval

Code: https://github.com/yh-yao/MERA-Evolve

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 75/100

The gist: MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale MERA is a verifier-backed multi-cycle protocol for improving small language models in agentic systems, treating a

Terminology

Summary

MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale

MERA is a verifier-backed multi-cycle protocol for improving small language models in agentic systems, treating a single model invocation as the unit of adaptation. The paper states: MERA instead improves the small model itself, using a single model invocation as the unit of adaptation. The method addresses the tension in deployed agent systems where traces needed for adaptation are naturally produced online, but directly changing the serving policy from raw traces is risky.

The runtime loop uses an input-only router that observes only the serialized prompt for the current invocation and selects among a strong model, a cheap model, and optionally a specialized student. A skill selector may dispatch stable templates for recurring local structure, and a verifier checks outputs with fallback to a stronger model on failure. The update loop operates on canonicalized step slices containing the prompt, local context, tool schemas, generated output, verifier result, retry count, fallback metadata, and any skill assignment.

Each update round follows a dependency-respecting schedule: "the LLM adapter is trained and queried with the SkillBook procedure prepended to its prompt, so the SkillBook update must precede LLM adaptation within a cycle (Skill → LLM); training the two in parallel would fit the adapter on a stale procedure. The router is placed last so its labels reflect the current-cycle skill state and small-model outcomes. This yields the canonical Skill → LLM → Router schedule."

Updates are admitted only through joint replay: SkillBook, router, and verifier evidence define easy, hard, and uncertain regions... A new router, skill state, or adapter is promoted only if replay preserves quality while reducing cost or fallback risk. The paper emphasizes that replay itself is only a pre-deployment gate: a production deployment still requires shadow or canary validation, drift monitoring, and rollback.

The main experimental results on code generation use Qwen2.5-Coder-1.5B as the small model and GPT-5.6 Luna as the teacher/fallback model, with 546 training tasks and 582 held-out HumanEval+MBPP tasks. Table 1 reports: Multi-cycle fine-tuning lifts the SLM from 28.7% to 44.2% (SFT) and 49.7% (SFT+GRPO). Matched GRPO exceeds SFT by 5.5–6.7 points (95% paired-t intervals exclude zero).

For deployed policies with verifier-backed fallback, Table 2 shows: With verifier fallback, MERA matches near-Luna quality at 60.8% cost, while the RouteLLM- and FrugalGPT-style baselines remain near always-Luna cost. The paper notes: The system win is the evolved SLM plus verifier fallback; the learned pre-router is weak, and most quality preservation comes from verification rather than routing alone.

On TAU-2, a multi-turn tool-use benchmark, Table 3 shows: The trained Qwen3.5-2B improves from 14/35 to 18/35 and can reach the performance of the unadapted Qwen3.5-4B endpoint. However, the paper cautions: The paired test has seven wins, three losses, and 25 ties (one-sided McNemar p = 0.171875), so the comparison is underpowered and supports feasibility on tool use rather than broad cross-domain generality.

A finance break-even analysis connects adaptation cost to serving savings. Table 4 shows that for 500 training rows, the break-even is 8.68 days nominally, or 12.47 days with a conservative 70% savings realization factor.

The paper's discussion states: "The strongest evidence is the three-seed SLM lift under matched SFT and SFT+GRPO. Verifier fallback turns that improvement into a near-Luna cost–quality point at 60.8% cost, while query/response routers stay near always-Luna cost."

Key limitations acknowledged in the paper include: the quality of both routing labels and student admission decisions depends on verifier coverage, the framework may realize its gains gradually and may leave a substantial fraction of difficult long-horizon reasoning on the strongest model for a long time, replay cannot perfectly capture distribution shift induced by changing the runtime policy itself, and the learned pre-router remains weak in our rebuilt study: most deployed quality preservation comes from executable verification and fallback rather than accurate upfront routing.

The conclusion states: "MERA uses shared executable traces primarily to improve the small model, with SkillBook, routing, and verifier-backed admission as supporting machinery. On held-out HumanEval+MBPP, four-cycle adaptation raises Qwen2.5-Coder-1.5B from 28.7% to 49.7% direct pass, and verifier-backed deployment retains 88.3% pass at 60.8% of always-Luna cost. On TAU-2, an adapted Qwen3.5-2B improves from 14/35 to 18/35 and matches an unadapted 4B endpoint, though the comparison is underpowered."

Improvements for AI systems

Improvements to AI systems:

  1. Self-Improving Small Models via Multi-Cycle Adaptation – AI systems can now iteratively fine-tune their own small language models (SLMs) using online execution traces, without needing human-labeled data. The system improves its direct pass rate from 28.7% to 49.7% on held-out code generation tasks over four cycles, using a verifier to gate each update.

  2. Cost-Aware Verifier-Backed Fallback – AI systems can deploy a cheap SLM with a verifier that checks outputs and falls back to a strong teacher model only on failure. This achieves 88.3% of teacher quality at 60.8% of the cost, compared to naive routing baselines that stay near always-teacher cost. The system learns when to trust its own cheap model and when to escalate.

  3. Dependency-Respecting Update Scheduling – AI systems can now safely update multiple components (skill templates, LLM adapter, router) in a canonical order (Skill → LLM → Router) to avoid training on stale data. This prevents performance degradation from parallel updates and ensures each component's labels reflect the current state of others.

  4. Replay-Based Admission Control – AI systems can pre-validate any proposed update (new router, skill, or adapter) by replaying historical traces. A new component is promoted only if it preserves quality while reducing cost or fallback risk. This acts as a safety gate before any live deployment, reducing the risk of regressions.

  5. SkillBook Procedure Prepend – AI systems can learn and reuse stable templates for recurring local structure (e.g., common code patterns, tool-use sequences) by prepending a learned SkillBook to the model's prompt. This improves adaptation efficiency and enables the system to generalize across tasks with shared structure.

  6. Online Trace Canonicalization – AI systems can convert raw execution traces into structured, canonical step slices (prompt, context, tool schemas, output, verifier result, retry count, fallback metadata, skill assignment). This enables clean, reproducible training data extraction from live agent interactions, making continuous improvement feasible in production.

  7. Verifier-Coverage-Aware Routing – AI systems can now explicitly account for verifier limitations. The system learns to route only when verifier coverage is reliable, and it acknowledges that quality preservation in deployment comes primarily from verification and fallback, not upfront routing. This prevents over-reliance on weak pre-routers.

  8. Feasibility for Multi-Turn Tool Use – AI systems can adapt small models for complex tool-use benchmarks (e.g., TAU-2) using the same protocol, improving from 14/35 to 18/35 and matching an unadapted 4B model's performance. This shows the method extends beyond code generation to general agentic tasks, though with weaker statistical power.

  9. Break-Even Cost Modeling – AI systems can now estimate the economic viability of adaptation upfront. With 500 training rows, the break-even point is 8.68 days (or 12.47 days with conservative savings), allowing operators to decide whether to run the adaptation loop based on expected serving volume and cost.

  10. Shadow/Canary Deployment Gate – AI systems can combine replay-based pre-deployment validation with shadow or canary testing, drift monitoring, and rollback. This creates a full safety pipeline for continuous model evolution in production, reducing the risk of silent quality degradation.

What the improved AI system can do:

  • Continuously improve its own small model on live tasks without human labels, reaching near-teacher quality at 60% cost.

  • Safely evolve its policies in production by validating updates via replay and canary testing before full rollout.

  • Learn reusable skill templates for recurring patterns, reducing token usage and improving accuracy on similar future tasks.

  • Automatically decide when to use its cheap model versus a strong teacher, based on verifier feedback, to optimize cost-quality trade-offs.

  • Adapt to new domains (e.g., multi-turn tool use) with minimal data, though with caution about statistical power.

  • Provide operators with clear cost-benefit metrics to decide when adaptation is worthwhile.

Abstract

LLM agents execute heterogeneous sequences of model calls within a single task: some invocations require careful reasoning, while others are structured steps such as formatting or tool-argument construction. Prior routing methods exploit this asymmetry by assigning easy invocations to a cheaper small model and difficult ones to a large model. Such policies reduce inference cost, but they leave the small model's capability unchanged, so attainable savings remain bounded by the work the student can already solve. MERA instead improves the small model itself, using a single model invocation as the unit of adaptation. In each cycle, MERA replays failed student invocations to obtain execution-verified teacher demonstrations, distills recurring procedures into an iteratively updated SkillBook, and fine-tunes a student LoRA adapter via supervised learning and optional GRPO. Routing serves as supporting machinery for deployment: the improved student is served behind a cost-calibrated router with verifier-backed fallback, and a candidate SkillBook, adapter, or router is admitted only when joint replay preserves task quality. Empirically, four-cycle adaptation raises Qwen2.5-Coder-1.5B from 28.7% to 49.7% pass on held-out HumanEval+MBPP. Under verifier-backed fallback, the deployed policy retains 88.3% pass at 60.8% of always-Luna cost. On TAU-2, a fine-tuned Qwen3.5-2B improves from 14/35 to 18/35 and matches an unadapted 4B model. These results indicate that verifier-backed multi-cycle adaptation can increase small-model capability, rather than only routing around a fixed student.

Sources

Related papers