Measuring the Serving Stack Instead of the Model: Hidden Confounds in Local Tool-Use Evaluation
cs.CL, cs.AI, cs.SE
Submitted: 2026-09-22
Updated: 2026-09-22
Comments: 9 pages, 4 figures, 3 tables. Accepted at the 2nd Workshop for Research on Agent Language Models (REALM) @ EMNLP 2026
Code: https://github.com/LijuanTang94/serving-confound-repo
License: http://creativecommons.org/licenses/by/4.0/
The gist: A coding agent must emit a valid tool call--a parseable invocation of a tool in the provided schema--before the harness can execute its chosen action.
Terminology
Abstract
A coding agent must emit a valid tool call--a parseable invocation of a tool in the provided schema--before the harness can execute its chosen action. We study how local serving stacks affect this protocol step and show that measured outcomes can depend on the serving layer rather than model behavior alone. In Ollama, the default tools= request is gated per model by a static template flag: some models are accepted and return calls as text, some return native tool calls, while Phi-3 and Gemma-3 are rejected before inference. In our harness, rejection and retry exhaustion are not preserved as structured failure metadata, so downstream analysis can misclassify them as model non-calls and naively report 0% fidelity. Adding a text tool list while retaining the native channel recovers much of the measured fidelity for accepted models, whereas a uniform text protocol reduces fidelity for Llama-3.2, which has native tool-call support. Cross-stack probes on Ollama, llama.cpp, vLLM, and SGLang show different handling of the same request. Constrained decoding removes parse failures but can induce non-termination, and turn-pooled versus per-instance estimates differ by up to about 55 points. We conclude with a checklist for treating serving behavior as part of the evaluation protocol.
Sources
- Why Do Multi-Agent LLM Systems Fail?
- Evaluating Large Language Models Trained on Code
- Don't Adapt Small Language Models for Tools; Adapt Tool Schemas to the Models
- ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
- The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility
- Gorilla: Large Language Model Connected with Massive APIs
- Efficient Guided Generation for Large Language Models
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows
- LOCA-bench: Benchmarking Language Agents Under Controllable and Extreme Context Growth
- Don't Fine-Tune, Decode: Syntax Error-Free Tool Use via Constrained Decoding
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering