Benchmarking LLM-Guided Control-Plane Policies for Backend Fault Isolation in HAProxy

arXiv:2608.10532 · cs.NI, cs.LG · Submitted 2026-08-11 · Read on arXiv

Aman Chauhan, Vishnu Pendyala

cs.NI, cs.LG

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: 43 pages, 6 figures, 15 tables. Submitted to Journal of Network and Computer Applications (Elsevier)

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: This paper investigates whether a Large Language Model (LLM) can replace the static routing policy in a load balancer (HAProxy) to automatically isolate faulty backend servers that are degraded

Terminology

Summary

This paper investigates whether a Large Language Model (LLM) can replace the static routing policy in a load balancer (HAProxy) to automatically isolate faulty backend servers that are degraded (returning HTTP 500s) but not down. The authors frame this as an intent-assurance problem, where the LLM acts as a closed-loop controller that reads server-side telemetry every 10 seconds and takes actions (adjusting weights, draining servers) through guardrailed API calls.

The study uses a reproducible benchmark with a persistent structural fault: every third backend has a 5% error rate baked into its configuration. The authors sweep 15 open-weight models across five families (Qwen3, Qwen3.5, Granite 4.0, Gemma 4, GPT-OSS), spanning 0.35B to 35B total parameters, with reasoning modes toggled, fleet sizes of 3 to 9 backends, and two load-balancing algorithms, totaling 240 runs.

Key findings:

  1. Capability threshold near 3B active parameters: "We find a capability threshold near 3B active parameters. Below it, LLM policies are typically unreliable and sometimes worse than no policy; above it, every model, regardless of architecture, saturates near an 88% reduction in client-perceived 5xx errors over the static baseline. The threshold is approximate: Gemma 4 E2B clears it with 2B active parameters, while the dense 3B Granite 4.0 Micro does not."

  2. Sub-threshold failure modes: Below the threshold, models split into two regimes: The conservative (non-drainer) models drain rarely but correctly (e.g., Qwen3 0.6B with precision 0.92 but recall 0.34), and The trigger-happy (over-drainer) models catch the faults but spray collateral onto healthy slots (e.g., Qwen3.5 2B with precision 0.71, recall 0.94). The worst configuration regressed the client 5xx rate by −667% relative to doing nothing.

  3. Instability even among competent models: End-state correctness is only 0.304 across all runs, against an ever-drained exact-match of 0.665: 83 of the 149 runs that at some point localized the fault correctly (55.7%) had reverted that decision by the final step. The churn lands on faulty slots: the macro faulty-slot flip rate is 4.12 versus a healthy-slot flip rate of 0.066, a 62× gap.

  4. Latency cost of control: Draining concentrates load onto surviving servers, inflating tail latency 2.6 to 2.8 times. Specifically, the supra-threshold arm’s p95 and p99 (40.4 and 49.0 ms) are roughly 2.6× and 2.8× the uncontrolled baseline. Mean latency rises by 4.7 ms, while median latency is barely affected.

  5. Scaling behavior: Effectiveness degrades gracefully rather than catastrophically as the fleet grows, from 89–93% at N=3 to roughly 77% at N=9. The degradation is a recall problem: Recall trends down from 0.893 at N=3 to 0.732 at N=9... Precision moves in the opposite direction, from 0.877 to 0.961. Time-to-resolution climbs from 59–79s at N=3 to 240–289s at N=9.

  6. Reasoning effort does not help: Enabling thinking multiplies median completion tokens per step from 157 to 1,546 (∼10×). This causes the control loop to miss deadlines: "thinking-on slips to a 12 s median cadence with 72.7% of runs falling behind real time and only ∼50 of 60 steps completed, and GPT-OSS-high slips to an 18 s cadence with every run falling behind and a median of only 27.5 of 60 steps completed. The result is that median 5xx reduction is 86.6% for nothink configurations versus 86.3% for think configurations," and GPT-OSS high effort drops to 74.0%.

  7. Cost-effectiveness: The cheapest non-reasoning configurations avoid 6–13× more failed requests per token than their reasoning-on counterparts (e.g. GPT-OSS low at 1,214 versus high at 93).

Practical recommendation: The efficient operating point is a supra-threshold model in its cheapest non-reasoning mode, wrapped inside deterministic guardrails. The guardrails (capping at one drain per interval, ±50 weight deltas, three actions per cycle) are essential: Without the one-drain-per-cycle cap, those plans would quarantine multiple backends at once, precisely the catastrophic over-reaction that produced the worst regressions.

Limitations: The study uses a single fault class, one load balancer, one backend stack, fleets up to nine backends, a single seed, a fixed system prompt, and a fixed 10-second poll interval. The authors note that the token counts are hardware-agnostic, but the cadence and steps-completed figures reflect this GPU’s throughput.

Improvements for AI systems

Improvements to AI Systems:

  1. Add a capability-gated deployment guard: Before enabling an LLM as a closed-loop controller, run a lightweight calibration test (e.g., 10 simulated fault scenarios) to verify the model has ≥3B active parameters or demonstrably clears the capability threshold. If below threshold, fall back to static policy or a rule-based drainer, preventing catastrophic regressions (−667% 5xx errors).

  2. Implement decision-stability enforcement: Add a hysteresis mechanism that requires the same drain/weight action to be proposed in two consecutive control cycles before executing it. This reduces the observed 62× flip-rate gap between faulty and healthy slots and prevents the 55.7% of runs that correctly localized faults from reverting to wrong states.

  3. Add a latency-aware action scheduler: For reasoning-enabled models, cap thinking time per step (e.g., 5 seconds) and pre-compute fallback actions if the model exceeds the 10-second poll deadline. This prevents the 72.7% of runs falling behind real time and the 50% completion rate seen with thinking-on configurations, while preserving the negligible 0.3% 5xx reduction difference.

  4. Introduce collateral-damage penalties in the reward function: Train or fine-tune the LLM policy with an explicit penalty for draining healthy backends (precision < 0.9). This addresses the trigger-happy regime (e.g., Qwen3.5 2B with 0.71 precision) where over-draining sprays collateral onto healthy slots, and aligns with the guardrail of one-drain-per-cycle.

  5. Build a fleet-size adaptive controller: Dynamically adjust the control loop’s aggressiveness based on backend count. For N≥7, increase the allowed drain/weight-change frequency and add a recall-boosting mechanism (e.g., proactive health checks) since recall drops from 0.893 (N=3) to 0.732 (N=9) and time-to-resolution triples.

  6. Add a tail-latency guardrail: Monitor p99 latency after each action; if it exceeds 2.5× the uncontrolled baseline, automatically revert the last weight change or drain. This prevents the 2.8× p99 inflation observed, which is not captured by median or mean metrics.

  7. Create a cost-aware reasoning toggle: Automatically disable reasoning mode when the request rate exceeds a threshold (e.g., >1 action per 10 seconds) or when the model’s token budget per action exceeds a limit. This preserves the 6–13× better failed-request-per-token efficiency of non-reasoning modes while maintaining the 86.6% 5xx reduction.

What the Improved AI System Can Do:

  • Reliably isolate faulty backends with ≥88% reduction in client-perceived 5xx errors, but only when the model is verified capable, avoiding worse-than-baseline behavior.

  • Maintain correct fault localization decisions over time, reducing end-state reversion from 70% to near-zero via stability enforcement.

  • Operate in real-time even with reasoning models, completing 60/60 control steps within the 10-second cadence.

  • Achieve precision >0.9 on drain actions, minimizing collateral damage to healthy servers.

  • Scale gracefully to fleets of 9+ backends, keeping recall above 0.85 and time-to-resolution under 120 seconds.

  • Keep p99 tail latency within 1.5× of baseline, preventing load concentration from degrading user experience.

  • Maximize cost-efficiency, avoiding 6–13× more failed requests per token than reasoning-heavy alternatives.

Sources

Related papers