The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World
Ashritha Gonuguntla
cs.LG, cs.CL
Submitted: 2026-08-08
Updated: 2026-08-11
Comments: 8 pages, 3 figures. Accepted at the Conference on Language Modeling 2026. Code: https://github.com/AshrithaG/replay-gap Data: https://huggingface.co/datasets/ashritha0907/replay-gap-trajectories
Code: https://github.com/AshrithaG/replay-gap
License: http://creativecommons.org/licenses/by/4.0/
The gist: LLM routers promise efficiency by matching each request to the cheapest adequate model, and are increasingly applied per step inside multi-step agents.
Terminology
Abstract
LLM routers promise efficiency by matching each request to the cheapest adequate model, and are increasingly applied per step inside multi-step agents. Yet agentic routers are evaluated like single-turn routers: by replaying logged trajectories and substituting another model's recorded outputs, assuming the rest of the trajectory is unaffected. We test this assumption with branching rollouts: we fork live SWE-bench agent trajectories at controlled points, rebuild the environment, continue each fork with a different model, and compare against same-model control forks that isolate sampling and replay noise. Across six paired runs (900 rollouts), swaps exceed their matched control floors by +0.25 to +0.66 normalized edit distance (multiplicity-corrected CIs exclude zero), rewriting 61-94% of post-fork actions; 74-77% of early swaps diverge at the first post-fork action, versus 6-35% of controls, leaving only 3% of replayed states valid. Divergence decreases with fork depth in both directions. All five outcome flips we observe occur in swap arms, upgrades rescuing unsolved instances and a downgrade losing the sole solve, and zero occur across 359 control forks. Scoring these same swaps with a log-stitching replay evaluator, replay mispredicts every success-relevant outcome call and predicts patches with 0.00-0.11 similarity to reality. Auditing the noise floor, temperature-0 "determinism" is configuration-dependent: FP8-served controls diverge on over 90% of forks while AWQ-served ones remain near-identical; and under tight budgets the stronger model more often exhausts its steps without submitting. Replay-based benchmarks score the wrong world for agentic routing; we release our harness and all trajectories.
Sources
- Switchcraft: AI Model Router for Agentic Tool Calling
- AutoMix: Automatically Mixing Language Models
- FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
- RouterBench: A Benchmark for Multi-LLM Routing System
- RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs
- Universal Model Routing for Efficient LLM Inference
- RouteJudge: An Open Platform for Reproducible and Preference-Aware LLM Routing
- RouterArena: An Open Platform for Comprehensive Comparison of LLM Routers
- ODAR: Principled Adaptive Routing for LLM Reasoning via Active Inference
- RouteLLM: Learning to Route LLMs with Preference Data
- Route to Reason: Adaptive Routing for LLM and Reasoning Strategy Selection
- Step-level Optimization for Efficient Computer-use Agents
- Qwen3 Technical Report
- TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks