The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance

arXiv:2608.11694 · cs.CL, cs.AI · Submitted 2026-08-12 · Read on arXiv

Shailja Thakur, Sungeun An, Chad DeLuca, Hima Patel

IBM Research India · IBM Research Almaden

cs.CL, cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

Code: https://github.com/IBM/BenchDrift

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

Terminology

Summary

Paper: arXiv:2608.11694v1 [cs.CL], 12 Aug 2026


The paper addresses a fundamental flaw in how LLM benchmark scores are reported: A benchmark score comes from a single phrasing of each problem. That single phrasing is treated as if it stood for the whole space of ways the same problem could be asked, but it does not. The authors demonstrate that rephrasing a problem while keeping its meaning and answer fixed routinely flips a model's answer in both directions, so some failures become successes and some successes become failures. They call this phenomenon drift.

The authors introduce BenchDrift, a framework that generates meaning-preserving variations of benchmark problems along four axes, namely linguistic, referential, pragmatic, and structural, and measures how often, and why, correctness flips under each.

  1. Linguistic — transformations change wording and style, such as rephrasing or switching between active and passive voice

  2. Referential — transformations change which entities, units, or values are named, such as swapping a name or converting a unit

  3. Pragmatic — transformations change tone, framing, and context, such as adding a persona or recasting the problem as a direct question

  4. Structural — transformations change how the problem is organised, such as reordering information or adding text irrelevant to the answer

  5. Generator — Proposes candidate rephrasings of the problem

  6. Validator — Checks each candidate against the original ground-truth answer and discards any that no longer have it

  7. Target model — The model under evaluation. It answers the original problem and every candidate that survives validation

  8. Judge — Scores those answers against the ground truth, accepting '12', 'twelve', and 'a dozen' as one answer

  • Positive drift: The model fails the original phrasing and succeeds on the variation: a capability the original wording was hiding (C(q) = 0, C(q′) = 1)

  • Negative drift: The model succeeds on the original and fails the variation: apparent success that rested on a wording choice (C(q) = 1, C(q′) = 0)

Rates are computed as fractions of the same N problems: Pos. = (1/N) q: C(q) = 0, ∃q′, C(q′) = 1 and Neg. = (1/N) q: C(q) = 1, ∃q′, C(q′) = 0.

Best = Rep. + Pos., Worst = Rep. − Neg. The reported score sits inside an interval whose upper margin is the positive drift rate and whose lower margin is the negative rate.

  • Benchmarks: GSM8K, MMLU, and MATH-Hard, with 500 problems sampled per benchmark

  • Models: Eight open-weight models spanning 7B to 34B parameters: Qwen3-8B, Mistral-7B, Phi-4, Granite-3.3-8B, GPT-OSS-20B, Granite-34B-Code, Llama-3.1-8B, and Qwen2.5-7B

  • Main configuration: Mistral-Large-Instruct-2411 as generator, GPT-OSS-120B as validator, Llama-3.3-70B-Instruct as judge

  • Judging: Target models answer with chain-of-thought prompting at temperature 0, which removes sampling noise

  • Variations: 20.2 variations remain per problem on average, ranging from 3 to 92

Averaged across all 24 model-benchmark pairs, the gap between worst-case accuracy (correct under every phrasing) and best-case accuracy (correct under at least one phrasing) is 74.7 percentage points.

Concrete example: Phi-4 on GSM8K makes this concrete. Its reported score is 93.4%. That score drops to 38.2% in the worst case and rises to 99.2% in the best case, a 61-point window with the reported number sitting near the top.

"The asymmetry between the two directions — negative minus positive drift — tracks baseline accuracy almost linearly across our 24 pairs (Pearson r = 0.98, slope 1.08). Above a 60% baseline, negative drift exceeds positive drift in all twelve pairs."

"A weak model has little to lose and something to gain, so rephrasing surfaces knowledge the original wording failed to elicit (Mistral-7B on MATH-Hard: 37.6% positive drift against 9.0% negative). A strong model has much to lose, so the same operation mostly costs it correct answers."

Critical implication: The best models on a benchmark are therefore the ones whose scores depend most on the wording they happened to be given.

"Swapping every pipeline role shifts average negative drift by at most 11.9 percentage points and average positive drift by at most 2.9, both within the per-cell confidence intervals of Table 2. The ranking of which models are most fragile is preserved too (ρ ≥ 0.98 against the main configuration)."

"The axes that break the most answers are the ones that leave a problem's content untouched: persona framing and paragraph reordering change no number and no relation, yet cost more correct answers than entity substitution, which rewrites the tokens the answer depends on. This indicates that the damage tracks surface form, not comprehension."

"The negative-drift range alone spans more than fourfold: interrogative expansion, which breaks a problem into a series of smaller explicit questions, costs 48.9% of correct answers, while adding whitespace costs only 11.7%."

Models agree on fragility: "computing each transformation's negative drift separately for every model and correlating the rankings gives a mean pairwise Spearman ρ of 0.77 across the 28 model pairs. Models differ in how much they drift overall, but largely agree on which transformations cost the most correct answers."

"Negative drift is lowest when length is roughly preserved (19.6%) and rises in both directions — to 28.9% when the problem gets shorter and 34.7% when it more than doubles in length. Shortening a problem causes about as much negative drift as lengthening it, which rules out drift as an artifact of problems simply becoming easier."

"Negative drift falls as confidence rises, from 38.1% in the low bin to 18.5% in the high bin, but does not approach zero — nearly one in five answers the model was most confident about is still lost to a meaning-preserving rephrasing. The pattern holds within every model tested, from 13% to 26% in the high bin. Therefore, a model's own confidence cannot be used to tell which of its correct answers will survive rephrasing."

With the full budget, 67.2% of problems drift. Five variations recover 74% of those, eight recover 88%, and ten recover 93%; a single variation finds only 30%.

"At a cost close to DSPy's and far below GEPA's, BenchDrift returns a 73.3-point range (26.7% to 100%) that neither optimiser reports, because both search for a single best-performing prompt and discard everything else once they find it."

  1. "A meaning-preserving variation framework across four axes, with every variation validated against the original ground truth, and a bidirectional drift measurement over a common denominator that attributes each flip to the transformation that produced it"

  2. Two tests that rule out the leading alternative explanations for drift: that rephrasing merely simplifies problems, and that it only disturbs answers the model was unsure of

  3. Reliability checks that swap the generator, validator, and judge models

  4. Open release of the generated and validated variations used in our experiments

  5. Open release of the BenchDrift pipeline and full variation taxonomy

"A benchmark score computed from one phrasing is not a stable estimate of what a model can do. Across eight models and three benchmarks, accuracy under different valid phrasings of the same problems spans, on average, 74.7 percentage points, and models more often lose correct answers to rephrasing than they gain new ones. We find that drift is systematic. It concentrates in specific transformations, holds up under reliability checks that swap out the entire pipeline, and BenchDrift attributes each flip to the transformation and axis responsible."

  • No human validation of the judge or validator — both are LLMs, and agreement between them doesn't guarantee agreement with humans

  • Three benchmarks, all closed-form — drift on open-ended tasks is not covered

  • Generator and judge model families overlap with target models — shared family-level biases cannot be fully ruled out

  • Axis membership is inferred, not recorded — the mapping is deterministic but self-defined

  • Cost — running BenchDrift is more expensive than single-phrasing evaluation, though Section 4.7 shows most drift is detectable with fewer variations

  • Scope — BenchDrift measures stability, not prompt optimization; specific rephrasings don't transfer across problems

Improvements for AI systems

Based on this paper, here are specific improvements to AI systems and what the improved systems can do:

Improvement: Modify the model's inference layer to generate multiple internal rephrasings of each input problem (using the four axes), answer each, and report a confidence interval rather than a single score.

Capability: The system can now tell users I am 93% confident on this phrasing, but only 61% confident across 20 valid rephrasings — flagging problems where its answer depends on wording. It can also output a stability score alongside each answer, letting users know whether the result is robust or fragile.

Improvement: Add a post-processing module that, before finalizing an answer, generates 5–10 meaning-preserving variations (especially pragmatic and structural axes, which cause the most damage) and checks if the answer flips.

Capability: The system can automatically reject or flag answers that are not stable across rephrasings, reducing silent failures. For high-stakes applications (medical, legal, financial), it can escalate unstable answers to a human or request clarification rather than confidently returning a wording-dependent result.

Improvement: During fine-tuning, augment the training set with the four-axis variations (linguistic, referential, pragmatic, structural) and add a loss term that penalizes answer flips across variations of the same problem.

Capability: The model learns to encode the underlying problem semantics rather than surface form, improving its worst-case accuracy. This directly addresses the finding that strong models lose up to 34.7% of correct answers when problems are lengthened or shortened — training on length variations would harden against this.

Improvement: Implement a runtime switch: for problems where the model's baseline confidence is below 60% (where positive drift dominates), generate multiple variations and take the majority answer; for problems above 60% confidence (where negative drift dominates), stick with the original answer but verify with a single structural variation.

Capability: The system dynamically allocates compute to maximize robustness. Weak models gain new correct answers (up to 37.6% positive drift), while strong models avoid losing correct answers (up to 48.9% negative drift on interrogative expansion) — improving overall accuracy without doubling inference cost.

Improvement: Modify evaluation pipelines to report not just a single accuracy number but a three-value output: worst-case, reported, and best-case accuracy, along with the drift rates per axis.

Capability: Model comparison becomes honest — a model with 90% reported but 50% worst-case is clearly less reliable than one with 85% reported and 75% worst-case. This prevents over-reliance on benchmark scores that the paper shows can be inflated by up to 61 points (Phi-4 on GSM8K: 93.4% reported vs. 38.2% worst-case).

Improvement: Train the model to map answers to canonical forms (e.g., "12", twelve, a dozen → "12") before scoring, and to internally rephrase the problem into a canonical template before answering.

Capability: The system becomes less sensitive to referential variations (unit conversions, name swaps) and structural variations (reordering), reducing the 11.7–48.9% negative drift caused by these transformations. This is especially useful for multi-turn dialogue systems where users naturally rephrase questions.

Improvement: Use the finding that even high-confidence answers (top confidence bin) still lose 18.5% of correct answers to rephrasing — so the system should not trust its own confidence. Instead, for tasks above a criticality threshold, automatically run a mini-BenchDrift (3–5 variations) and require agreement across all variations before providing a final answer.

Capability: In medical diagnosis, legal advice, or code generation, the system will refuse to answer or will explicitly state This answer is not stable across phrasings when drift is detected, preventing confidently wrong outputs.

Improvement: Instead of searching for a single best prompt (like DSPy or GEPA), the system maintains a portfolio of validated variations per problem type and selects the one that maximizes the worst-case accuracy across all variations, not the average.

Capability: The system avoids the trap where a prompt optimizes for one phrasing but fails on others. It can report a guaranteed performance floor, not just a peak, making it safer for deployment in unpredictable user-input environments.

Abstract

A benchmark score comes from a single phrasing of each problem. That single phrasing is treated as if it stood for the whole space of ways the same problem could be asked, but it does not. We show that rephrasing a problem while keeping its meaning and answer fixed routinely flips a model's answer in both directions, so some failures become successes and some successes become failures. We call this drift. BenchDrift generates meaning-preserving variations of benchmark problems along four axes, namely linguistic, referential, pragmatic, and structural, and measures how often, and why, correctness flips under each. Across eight models and three benchmarks (GSM8K, MMLU, MATH-Hard), we observe that drift is large in both directions. Two findings stand out. First, phrasing sensitivity does not fade as models get better. Instead, it changes sign. Weak models gain more from rephrasing than they lose, while strong models lose far more than they gain. We find that the best models on a benchmark are therefore the ones whose scores depend most on the wording they happened to be given. Second, the models largely agree on which rephrasings cost the most correct answers even though they differ in how much they drift, so fragility belongs to the rephrasing and not to the model. Furthermore, rephrasing breaks answers a model was confident about, whether the problem is made shorter or longer. Code and Data: https://github.com/IBM/BenchDrift/tree/demo-ui

Sources

Related papers