QuoteBench: How Matched Scores Can Hide Command-Path Failures
Shangao Li, Yao Zhang, Volker Tresp, Yuanyuan Yang
Stony Brook University · LMU Munich · Munich Center for Machine Learning
cs.AI, cs.SE
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: 29 pages, 5 figures. Project page: https://quotebench.lsamc.website/
Code: https://github.com/anthropics/claude-code
Project page: https://quotebench.lsamc.website
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: The paper addresses a critical blind spot in evaluating LLM coding agents that issue Bash commands.
Terminology
Summary
The paper addresses a critical blind spot in evaluating LLM coding agents that issue Bash commands. The authors argue that "LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish command-generation errors from failures introduced after generation. They note that
Bash quoting failures can corrupt literals, break routine agent actions, and trigger repair loops, and that even simple tasks such as
writing exact bytes, passing a literal argument, editing JSON, or invoking a remote-like wrapper must preserve quotes, dollar signs, backticks, newlines, glob characters, and expansion timing."
The key insight is that agent success does not reveal whether the first command preserved its payload, and command-generation success does not show whether it survives deployment.
The paper introduces QuoteBench to isolate this failure mode, which current benchmarks leave entangled with other capabilities
— broad coding benchmarks combine command construction with planning, repository navigation, and recovery,
while command-generation benchmarks score emitted programs under fixed transport.
QuoteBench consists of 56 one-shot Bash tasks: 14 operation families, each with a benign control and three hazardous payload variants.
The families cover literal file content, hostile filenames, regular expressions and globbing, heredocs, argument and environment passing, JSON and Git state, and two local simulations of a second shell parser.
The hazardous tiers hold the operation fixed while adding quotes, expansion characters, multiline data, leading dashes, or parser-boundary conflicts.
Each task provides a fixture, an instruction, and a final-state validator.
Validators check exact file bytes, received argument vectors, parsed JSON, directory state, or Git history
and score only the resulting state, so any semantically correct implementation receives credit.
The authors emphasize that Exit codes cannot substitute: across the failing executions, 23.4–47.0% exit zero while leaving the wrong final state.
The task families were selected from a pre-release mechanism survey of 86 de-identified incidents in author-owned agent sessions and 412 screened public reports,
supporting coverage of repeatedly observed command-construction mechanisms
rather than prevalence estimation.
The paper's central methodological contribution is a crossed design for mechanism identification
that independently varies generation contract
and execution transport.
The generation contract tells the model how to express an action,
while the execution transport determines how that action reaches a shell.
Three contracts are evaluated:
-
Raw contract:
the model emits one Bash program, executed verbatim as the script argument to bash-c
-
Native contract:
the model fills a provider shell-tool call
with the command field extracted and executedon the same raw path
-
Disclosed-boundary contract:
the model is told that its reply R will be interpolated into bash-c
R""
Two transports are used:
-
Raw transport: direct execution
-
Nested transport: adds
the double-quoted parser boundary found in remote, container, and CI command paths
The four cells are denoted RR (raw reply, raw transport), RN (raw reply, nested transport), NR (disclosed-boundary reply, raw transport), and NN (disclosed-boundary reply, nested transport). The key decomposition is: YN N − YRR = (YRN − YRR) + (YN N − YRN),
where Fixed-reply transport damage compares YRN with YRR
and Contract-conditioned compensation compares YN N with YRN.
The nested transport is validated as realistic: replaying each stored raw reply through a real ssh localhost
R remote command reproduces the nested damage to the decimal for seven of eight configurations and within one task for the eighth.
At their best observed settings,
three models pass all 56 tasks, with the remaining scores range from 14.3% to 98.2%.
The paper notes that "Raw generation itself is close to saturated at the frontier: the six frontier configurations pass 91.1–100% of tasks on the direct path, so raw scores carry almost no discriminative signal. The entire signal lives on the nested side."
The native-tool campaign shows provider-native shell-tool success ranges from 85.7% to 98.0%, compared with 95.4–99.3% on the raw path,
with native-minus-raw changes ranging from +2.6 to −10.0 points.
The central finding: Moving a fixed raw-generated reply from raw to nested transport costs every same-window configuration 55.4–73.2 points.
This damage is not confined to adversarial payloads: the 14 benign control tasks alone lose 28.6–57.1 points, because models emit double-quote-active characters even for ordinary commands.
The damage is fully attributable to the added parser: Reparsing preserves 123 of the 415 direct-path successes, corresponding to configuration-level retention of 25.0–35.4%. The remaining 292 become failures.
The damage is completely reversible with correct handling: "Escaping the reply at the interpolation point (bash-c ⟨quoted input⟩) reproduces the raw-path outcome exactly for all 448 public pairs. Replaying the reply as a temporary script does the same for all 448 public and 126 private-v1 pairs."
The paper's headline example: "GPT-5.6-sol's matched gap is only −3.6 points, even though fixed-reply transport loses 64.3 points and the realized contract-conditioned contrast restores 60.7. The near-zero matched change is therefore the sum of two large opposing components."
This masking is systematic: At the reported rungs GPT-5.6-sol meets this reading (−3.6 = −64.3 + 60.7), and ten of the 30 rung-level crossovers meet the same cut, all at ladder tops.
Six of eight same-window configurations show 30.4 to 60.7 points of realized compensation,
all with family-bootstrap intervals excluding zero.
The two non-compensating configurations are Qwen3.5-27B (0.0) and Gemini-3.1-Flash-Lite (−5.4).
The compensation is genuine behavioral adaptation: "The same disclosed-boundary replies that recover the nested path lose 28.6–64.3 points when replayed on the raw path (N R versus RR): the six compensating models rewrote their commands for the declared boundary and pay for it where the boundary is absent."
Compensation concentrates where the hazard is explicit
: "Payload-quoting families such as json-write (+50.0) and sed-replace (+46.9) recover about half their damage, but implicit hazards remain difficult: find-glob (−12.5), grep-count (+15.6), and hostile-filenames (+18.8) stay broken even under disclosure."
The adaptation is conditioned on the declared grammar rather than applying a fixed defense
: in a grammar crossover, GPT-5.6-sol passes 53/56 of its single-disclosed replies on the single-quote wrapper but only 10/56 on the double-quote one, and its diagonal (matched) advantage over the anti-diagonal is +80.4 points.
The paper finds that Effort improves matched success for some models but not others, and the same effort label corresponds to different token budgets across models.
The nested-replay pass rate stays between 23.2% and 33.9%, moves by at most 5.4 points within any one ladder
across effort rungs, while compensation rises substantially (e.g., from +10.7 at low to +64.3 at max for Opus-4.8
).
The paper warns: "An unset effort field maps to different parts of each provider's ladder. Across the seven same-window configurations with both measurements, Opus-4.8's unset arm resembles xhigh, Opus-5's resembles medium, and Gemini-3.1-Pro's falls below its entire ladder."
The RR and N N orderings agree only partially: their Kendall rank correlation is 0.57 (task-cluster bootstrap 95% interval [0.32, 0.82], excluding perfect agreement).
The one unambiguous reversal is GPT-5.6-sol versus Gemini-3.5-Flash (behind by one task under RR, ahead by eighteen under N N).
The paper provides a concrete example: selecting by raw success picks GPT-5.5 (56/56 raw), which reaches 50/56 on the nested path, whereas the path-aware pick reaches 51/56. The regret is small at the saturated frontier, but the reversed top rank.
The mechanism persists across effort settings, repeated draws, userlands (GNU versus BSD coreutils environments), and held-out payloads.
On private-v2 payloads, Both models pass 92.9% of the private tasks on the direct raw path, then lose 73.8 and 76.2 points when the same replies cross the added parser.
Three additional draws preserve negative damage and positive compensation for both models.
The decomposition should extend beyond shell
: Replaying each stored raw reply through a naive JSON string embedding causes losses from 51.8 to 66.1 points,
while a correct round-trip serializer costs exactly zero.
The paper makes three contributions:
-
A final-state benchmark of command-path reliability
— 56 exact-state tasks from 14 operation families -
A crossed design for mechanism identification
— independently varying generation contract and execution transport with fixed-reply replay -
Robustness and a measured, not novel, fix
— the transport loss persists across settings, and the fixes (correct escaping, temporary script) are trivial;precisely because the fixes are trivial, the contribution is the measurement, not the repair
The authors conclude: "Evaluations of command-issuing agents should report the model configuration, generation contract, execution path, operating point, and final-state validator rather than treat a matched score as an intrinsic model property. They emphasize that
The command interface is part of the evaluated system, not neutral plumbing."
The paper acknowledges: QuoteBench isolates one mechanism: one-shot Bash generation under quotation and interpolation hazards.
Its 14 constructed families support mechanism attribution,
but results characterize this benchmark and stored-reply portability, not deployment prevalence.
The effort-ladder rungs rely on a single stored generation per task,
and effort labels are not comparable compute budgets.
The native-tool campaign is observational,
and other shells and multi-turn recovery remain open.
Improvements for AI systems
Based on this paper, here are the specific improvements I can make to AI systems and what the improved systems can do:
1. Add parser-boundary awareness to command generation
-
When generating Bash commands, explicitly model the transport path (direct execution vs. nested interpolation through
bash-c "R", SSH, containers, CI systems) -
Detect when generated output contains characters that will be reinterpreted at a parser boundary (double quotes,
, backticks, backslashes, newlines) -
Automatically escape or restructure commands when the declared execution path involves interpolation
2. Implement transport-conditioned generation
-
Accept a
transport specification
as input (raw, nested, JSON-embedded, remote) and generate commands that survive that specific path -
For nested transport, produce replies that are safe when interpolated into
bash-c "R"— either by avoiding double-quote-active characters or by pre-escaping them -
For JSON-embedded transport, use proper round-trip serialization rather than naive string embedding
3. Add final-state validation to command execution
-
After executing a command, verify the actual resulting state (file bytes, argument vectors, parsed JSON, directory structure, Git history) rather than relying on exit codes alone
-
Detect when exit code is zero but the final state is wrong (the paper shows 23.4–47.0% of failures exit zero)
-
Trigger automatic retry with corrected quoting when state validation fails
4. Separate command generation from transport in evaluation and execution
-
Maintain two distinct pipelines: one for command construction and one for transport-safe delivery
-
When a command fails at the transport layer, distinguish this from generation errors and apply transport-specific fixes (escaping, script files) rather than regenerating the command
5. Add quote-hazard awareness to planning
-
Before executing any command involving literals, filenames, regex, globs, heredocs, JSON, or Git state, check for characters that could be corrupted by quoting (
",',,`, ``,*,?,[,], newlines, leading-) -
For payload-quoting tasks (JSON writes, sed replacements), apply explicit escaping; for implicit hazards (find-glob, grep-count, hostile filenames), add defensive quoting even when not obvious
6. Implement contract-conditioned compensation
-
When the system is told its output will be interpolated into a shell command, automatically rewrite commands to be safe under that declared grammar
-
Learn which hazard families benefit from disclosure (json-write, sed-replace) and which remain broken (find-glob, grep-count, hostile-filenames) — apply additional safeguards for the latter
1. Survive nested execution paths without degradation
-
Maintain 91–100% task success on direct paths AND retain that performance when commands are routed through SSH, containers, or CI systems with double-quoted interpolation
-
Avoid the 55–73 point drop that current systems experience when moving from raw to nested transport
2. Pass exact-state validation on first attempt
-
Write exact file bytes, pass literal arguments, edit JSON correctly, and invoke remote-like wrappers while preserving quotes, dollar signs, backticks, newlines, glob characters, and expansion timing
-
Achieve correct final state even when the command contains hostile filenames, multiline data, leading dashes, or parser-boundary conflicts
3. Detect and recover from transport-induced failures
-
When a command fails due to parser re-interpretation (not generation error), identify the specific corrupted characters and apply targeted escaping or switch to a script-file execution method
-
Distinguish between
command was wrong
andcommand was right but transport mangled it
— and apply the appropriate fix
4. Report transport-aware performance metrics
-
Provide separate scores for command generation quality and transport survival
-
Report matched scores alongside their decomposition (fixed-reply transport damage + contract-conditioned compensation) so users understand what a single number actually means
5. Adapt to declared execution contracts
-
When told "your reply will be interpolated into
bash-c "R"", generate commands that are safe under double-quote interpolation -
When told
your reply will be JSON-embedded
, use proper round-trip serialization -
When no contract is declared, default to the most conservative transport-safe output
6. Maintain performance across effort levels and environments
-
Keep nested-path success stable across low/medium/high effort settings (avoid the current 23–34% floor)
-
Preserve command reliability across GNU and BSD coreutils environments and across held-out payload variants
7. Avoid deployment-configuration ranking reversals
- Ensure that a model ranked highly under raw transport is also ranked highly under nested transport — or explicitly report both rankings so users can make path-aware selections
Abstract
LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish command-generation errors from failures introduced after generation. QuoteBench measures this boundary with exact final-state validation on 56 one-shot tasks from 14 incident-derived families, crossing the generation contract with the execution transport around one deliberately unescaped added parser. Escaping at the interpolation point reproduces each replayed reply's raw-path outcome, so any recovery under a disclosed boundary must come from the model changing its generation. Across eight same-window configurations, replaying the same reply through the added parser lowers success by 55.4 to 73.2 percentage points; disclosure recovers 30.4 to 60.7 points for six configurations, and zero or slightly negative for the other two. Raw generation is nearly saturated at the frontier; boundary adaptation is what still separates models. GPT-5.6-sol's matched gap of-3.6 points hides-64.3 points of damage and +60.7 points of compensation. The deployment configuration reorders models: one reversal among 26 comparable pairs is unambiguous and four more sit on single-task margins. Evaluations of command-issuing agents should report the model configuration, generation contract, execution path, operating point, and final-state validator rather than treat a matched score as an intrinsic model property.
Sources
- NeurIPS 2020 NLC2CMD Competition: Translating Natural Language to Bash Commands
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks
- EnvBench: A Benchmark for Automated Environment Setup
- NL2Bash: A Corpus and Semantic Parser for Natural Language Interface to the Linux Operating System
- AgentBench: Evaluating LLMs as Agents
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
- Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- BashCoder-R1: Towards Robust and Explainable Bash Code Generation with Robustness-Aware Group Relative Policy Optimization
- CARE: Pre-Execution Command Verification for Shell-Executing LLM Agents
- Stop Comparing LLM Agents Without Disclosing the Harness
- WebArena: A Realistic Web Environment for Building Autonomous Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection