Accuracy and Order Sensitivity Diverge Under Label-Free Strategies

arXiv:2608.11947 · cs.CL, cs.AI · Submitted 2026-08-12 · Read on arXiv

Karl Hanna, Chen Feng

Queen's University Belfast

cs.CL, cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: 20 pages. Code available at https://github.com/cotenthusiast/choicebench

Code: https://github.com/cotenthusiast/choicebench

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

Terminology

Summary

Summary

This paper investigates whether preventing a model from seeing option labels while committing to an answer removes positional influence and improves performance on multiple-choice question (MCQ) benchmarks. The authors evaluate two label-free strategies across six models (GPT-4.1 mini, Gemini 2.5 Flash, Llama 3.1 8B Instant via Groq API, Qwen 2.5 7B Instruct Turbo via Together AI API, Qwen 2.5 7B Instruct local, and Llama 3.1 8B Instruct local) and two benchmarks (MMLU and ARC-Challenge).

The first strategy, two-stage prompting, first generates a free-text answer without showing the options, then in Stage 2 matches that answer to the most closely corresponding option. The second strategy, independent hypothesis scoring, prompts the model k times, each time presenting the question with a single option and asking for a score, then selects the highest-scoring option. This strategy is positionally unbiased by construction, since each call is isolated and there are no positions to influence the response.

The key finding is that "Neither strategy reliably improves accuracy. Two-stage prompting reduces it for nearly all model–benchmark pairs, while independent hypothesis scoring is more mixed: it reduces accuracy for most models, but improves Llama-local, with the gain concentrated on ARC-Challenge." Specifically, two-stage reduces end-to-end accuracy in 11 of 12 model–benchmark pairs, while independent hypothesis reduces it in 8 of 11 valid pairs (Gemini on ARC-Challenge was excluded due to provider-side API failures).

The authors decompose two-stage prompting along two factors: whether options are hidden or visible in Stage 1, and whether Stage 2 uses an LLM call or embedding-based semantic matching. The decomposition shows that "when the Stage 2 LLM call is replaced by semantic matching with options hidden, performance drops sharply because the matcher cannot reliably map the free-text answer to an option. However, when paired with a Stage 1 where the options are visible, much of the loss is recovered, which indicates that the bottleneck is hiding the options in Stage 1 rather than the matching step." The only configuration that consistently matches baseline is visible options in Stage 1 paired with an LLM matcher.

For measuring positional effects, the authors use flip rate (a direct per-question measure of order sensitivity) and recall standard deviation (RStd, measuring how unevenly the model performs across answer positions). Two-stage prompting does not reliably improve either measure: some model–benchmark pairs become more stable or balanced, while others become less so. RStd increases in 5 of 12 pairs and decreases in 7, while flip rate increases in 7 of 12 pairs and decreases in 5. The two metrics agree in direction in 8 of 12 pairs, with Qwen-local being the clearest exception where they move in opposite directions.

A notable finding is that lower positional sensitivity does not necessarily coincide with higher accuracy; for GPT-4.1 mini on MMLU, flip rate roughly halves while accuracy falls. Specifically, GPT-4.1 mini's flip rate falls from 21.6 to 11.8 on MMLU while accuracy falls from 81.8 to 80.1. This demonstrates the decoupling thesis: reducing positional effects, even eliminating them entirely, does not always translate into accuracy gains.

Independent hypothesis removes positional influence by construction, yet accuracy still does not consistently improve. Only Llama-local sees a meaningful increase, rising 14.4 pp on ARC-Challenge (from 58.0 to 72.4). However, the authors note that on ARC, three unrelated methods converge to similar gains for Llama-local (cyclic at 72.9, independent hypothesis at 72.4, Visible+LLM at 70.4), suggesting Llama-local's unusually high recoverable performance rather than anything specific to removing positional information. On MMLU, the independent hypothesis gain for Llama-local is modest (roughly 1.0–2.4 pp) and sensitive to the random tie-break seed.

Cyclic permutation improves accuracy in 10 of 12 pairs (5 of 6 on MMLU and 5 of 6 on ARC). The authors note that cyclic is the best-performing model-agnostic method among those evaluated. PriDe, which requires logprob access, can substantially reduce accuracy even when log-probabilities are available, with Llama-local dropping from 50.2 to 44.2 on MMLU and from 58.0 to 48.8 on ARC-Challenge.

The paper also reports that semantic matching drives flip rate to exactly zero for every model on both benchmarks, confirming its position-blindness by construction, but at a substantial cost to accuracy. For example, semantic matching with hidden options yields accuracies as low as 32.0 for Llama-local on MMLU.

The authors conclude: These results show that eliminating positional influence does not reliably improve accuracy, while two-stage prompting does not reliably eliminate that influence in the first place. They also note that reducing positional effects and improving accuracy come apart across the strategies evaluated, including one that removes positional influence by construction and still does not consistently improve accuracy.

Improvements for AI systems

Improvements to AI systems:

  1. Add a positional-influence vs. accuracy diagnostic mode: Before deploying a model on multiple-choice tasks, run a calibration check that measures flip rate and RStd alongside accuracy. If flip rate is high but accuracy is also high (as with GPT-4.1 mini on MMLU), do not attempt to reduce positional bias via label-free prompting, as this will likely degrade accuracy. Instead, keep the standard option-present format.

  2. Implement a selective two-stage fallback only when semantic matching is reliable: For models that show poor free-text-to-option mapping (e.g., Llama-local on MMLU with accuracy dropping to 32.0), disable two-stage prompting entirely. Instead, use visible-option Stage 1 with an LLM-based matcher, which is the only configuration that consistently matches baseline accuracy. This prevents the sharp accuracy loss caused by hiding options.

  3. Use cyclic permutation as the default debiasing strategy when accuracy is the priority: Since cyclic permutation improves accuracy in 10 of 12 model–benchmark pairs and is model-agnostic, integrate it as a built-in option in MCQ evaluation pipelines. For production systems, apply cyclic permutation to the option order before inference, and aggregate predictions across permutations to reduce order sensitivity without sacrificing accuracy.

  4. Add a guardrail against independent hypothesis scoring for models with low base accuracy: For models like Llama-local on ARC-Challenge where independent scoring yields a large gain (58.0 → 72.4), enable this strategy only when a pre-check shows the model has high recoverable performance (e.g., cyclic permutation also shows gains >10 pp). For models where gains are marginal or seed-sensitive (e.g., Llama-local on MMLU), avoid this method to prevent unstable results.

  5. Build a hybrid selector that chooses between methods based on model–benchmark characteristics: Train a lightweight meta-classifier (using features like model family, parameter count, API vs. local, and baseline flip rate) to predict which strategy (standard, cyclic, two-stage, or independent scoring) maximizes accuracy for a given model–benchmark pair. This selector would use the paper’s empirical results to avoid known failure cases (e.g., two-stage for GPT-4.1 mini, independent scoring for Qwen 2.5 7B) and exploit known successes (e.g., cyclic for most models, independent scoring for Llama-local on ARC).

  6. Introduce a position-blindness verification step in evaluation harnesses: Before reporting accuracy improvements from any debiasing method, automatically compute flip rate and RStd. If a method claims to remove positional influence (e.g., semantic matching) but accuracy drops below a threshold (e.g., >10 pp loss), flag it as unsuitable for deployment. This prevents over-reliance on position-blind methods that sacrifice correctness.

What the improved AI system can do:

  • Automatically choose the optimal MCQ answering strategy per model and benchmark, avoiding accuracy loss from naive debiasing.

  • Detect when reducing positional bias will not help (or will hurt) accuracy, and refrain from applying such changes.

  • Provide a safety net: if a debiasing method fails (e.g., semantic matching yields low accuracy), the system falls back to the best-performing configuration (visible options + LLM matcher).

  • Offer a transparent report of positional sensitivity metrics (flip rate, RStd) alongside accuracy, so users can make informed decisions about deployment.

  • Achieve higher average accuracy across diverse models and benchmarks by leveraging cyclic permutation as a safe default, while selectively applying independent scoring only where it is proven beneficial.

Abstract

Multiple-choice benchmarks are widely used to evaluate large language models, but MCQ scores conflate knowledge with sensitivity to option order, which makes them unreliable measures of model knowledge. In this paper, we test whether preventing a model from seeing option labels while committing to an answer removes positional influence and, in turn, improves performance. We evaluate two different strategies for mitigating bias. The first uses a generation-then-matching approach, and the second scores options in isolation, which is positionally unbiased by construction. Neither reliably improves accuracy. A complete decomposition shows that the bottleneck is withholding options, not the matching step. The only configuration that consistently matches the baseline is the one that shows the model all options paired with an LLM matcher. However, eliminating positional influence entirely still does not reliably yield accuracy gains, while cyclic permutation often improves them. For two-stage prompting, an aggregate measure of recall imbalance and a direct per-question measure of order sensitivity both fail to show reliable debiasing.

Sources

Related papers