Choices Speak Louder than Questions

arXiv:2502.18798 · cs.CL, cs.AI · Submitted 2025-02-26 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Choices Speak Louder than Questions".

Jane: The evaluation of Large Language Models (LLMs) using Multiple-Choice Question Answering (MCQA) is unreliable because model decisions are often more influenced by superficial characteristics of answer options than by genuine…

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, to recap where we are, we've heard that this paper introduces choice sensitivity and a new scoring method called NPSQ to check if models are truly understanding questions or just picking answers based on superficial clues. What is the central thesis they are pushing?

Jane: The main argument is that if a model frequently picks the correct answer without actually understanding the question, then its accuracy score doesn't really reflect its comprehension abilities. This paper examines how much LLMs depend on answer choices instead of properly grasping what’s being asked in multiple-choice benchmarks.

Lu: They systematically analyze how overall performance can be attributed to the information in the answer choices by comparing model performance with and without the questions input, which is a key part of their systematic analysis section.

Meng: So, they aren't just saying models are bad at understanding; they are showing *how* much of that apparent success is actually driven by those distractors versus the prompt itself.

Lalam: It’s about moving beyond just looking at the final answer score and getting a clearer picture of whether the model has done the work of reading and reasoning through the question.

Tom: And they claim their proposed NPSQ method allows for a more robust and interpretable assessment, which is important because it means we can trust our evaluations more when we deploy these models.

Jane: Precisely, they demonstrate that traditional MCQA evaluation metrics are often highly sensitive to superficial features of the answer choices, but the new approach helps isolate the impact of the question itself.

Conclusion: Tom: Looking at "Choices Speak Louder than Questions," it seems like the authors are really calling attention to the gap between how we test models today and what true comprehension looks like in practice. Who were the researchers behind this work?

Jane: The paper was written by Gyeongje Cho, Yeonkyoung So, and Jaejin Lee from Seoul National University, and they presented their findings under review at a conference called ICLR two thousand twenty-six.

Lu: The implications are big because if models rely too much on these superficial answer choices, it suggests that the training methods might be rewarding the model for memorizing option patterns instead of learning the underlying logic.

Meng: From a practical impact view, this means we need to develop testing pipelines that actively try to break models by using deliberately irrelevant options to see if they actually understand the prompt.

Lalam: For me, this research suggests that focusing on methods like NPSQ helps us build an AI culture where we prioritize genuine reasoning over simple pattern matching in our applications.

Tom: So, in simple terms, the whole point of "Choices Speak Louder than Questions" is to give us a better tool—NPSQ—to measure if an AI is actually learning the concept or just memorizing patterns tied to specific answer formats.

Gyeongje Cho, Yeonkyoung So, Jaejin Lee

Graduate School of Data Science, Seoul National University

cs.CL, cs.AI

Submitted: 2025-02-26

Updated: 2026-01-12

Importance score: 76/100

The gist: The evaluation of Large Language Models (LLMs) using Multiple-Choice Question Answering (MCQA) is unreliable because model decisions are often more influenced by superficial characteristics of answer

Key concepts

Choice Sensitivity
This measures how much a model’s prediction relies on the provided answer options instead of understanding the actual question. High sensitivity means the model is easily swayed by distractors rather than focusing on what the question is asking.
Choice-Driven Component
This part of a score calculates how much a decision is influenced solely by the answer choices, ignoring any input from the original question. It's found to be zero when no question is present, helping researchers isolate choice influence.
Normalized Probability Shift by the Question (NPSQ)
NPSQ is a new metric that isolates the impact of the question from the answer choices. It calculates how much more likely a correct answer is given the question compared to just considering the choices alone, providing a cleaner measure of true understanding.

Terminology

Summary

The evaluation of Large Language Models (LLMs) using Multiple-Choice Question Answering (MCQA) is unreliable because model decisions are often more influenced by superficial characteristics of answer options than by genuine comprehension of the question. This paper addresses this issue by introducing a new scoring method, Normalized Probability Shift by the Question (NPSQ), to isolate the impact of the question from that of the answer choices, providing a more robust and interpretable assessment of model understanding.

Choice Sensitivity Definition

Choice sensitivity is formally defined as the extent to which a model’s predictions are predominantly influenced by the provided answer choices rather than by its understanding of the question itself. This phenomenon occurs when model predictions are predominantly influenced by the provided answer choices rather than by the actual questions. To quantify this, researchers analyze two components of a score: choice-driven and question-driven. The choice-driven component is determined by calculating the score with the question replaced by an empty string, while the question-driven component captures the additional contribution from the question itself calculated by subtracting the choice-driven component from the overall score.

Quantifying Choice Sensitivity

The relative contribution of these components to a decision is analyzed through two differences:

  1. The choice-driven difference, denoted as ∆choice, which captures how much more the model prefers x1 over x2 based solely on the answer choices provided.

  2. The question-driven difference, denoted as ∆question, which assesses how much the presence of the question influences this preference.

If ∆choice > ∆question, it indicates that the model’s preference for x1 over x2 is more strongly driven by differences in the answer choices than by any influence from the question itself, signifying a choice-sensitive decision.

Normalized Probability Shift by the Question (NPSQ)

The paper introduces NPSQ as a new evaluation method designed to isolate the impact of the question from that of the answer choices. It is defined as:

NPSQ(Q, C, x) = log P(x Q, C) − log P(x C).

This metric normalizes the probability shift by dividing the negative log-probability of the choice x in the absence of Q. A key property is that if Q is not present, NPSQ will always equal zero for all choices, meaning the choice-driven component of NPSQ is always zero, thus isolating its determination to the relationship between the question and the choices rather than being influenced soley by the choice information alone.

Experimental Findings on Choice Sensitivity

Experiments across various models (Qwen 2.5, Llama 3.1, Mistral) and input formats (cloze, symbols, hybrid) revealed systematic differences in sensitivity:

  1. Approximately 20–60% of the answer choices by language models are primarily influenced by the choices themselves. Choice sensitivity ranges from approximately 0.2 to 0.4 for symbols and hybrid formats, and around 0.5 to 0.6 for cloze format.

  2. The symbols and hybrid formats consistently exhibit lower choice sensitivity compared to the cloze format, suggesting that incorporating answer choice information in the prompt may help reduce the model’s reliance on spurious patterns.

  3. Normalization by token length fails to mitigate choice sensitivity, as choice sensitivity does not decrease after applying length normalization in some cases.

Robustness to Adversarial Choices

The study tested model responses when original distractors were replaced with carefully crafted adversarial choice[s], which is an intentionally irrelevant and implausible option that does not mislead a human examinee. The results showed that while standard metrics like accuracy (acc) and length-normalized accuracy (acc norm) are significantly impacted by adversarial choices, the NPSQ metric remains more stable. For instance, in the ARC-Challenge dataset for Llama3.1-8B-Instruct, acc npsq showed only 10.13% of predictions affected, whereas acc and acc norm saw greater shifts. This demonstrates that NPSQ provides a more robust and reliable measure of a model’s true understanding of the question.

Impact of Model Characteristics

The analysis highlighted several factors influencing choice sensitivity:

((a) Choice sensitivity across model sizes on the Qwen2.5 series (instruction-tuned): Larger models tend to exhibit lower choice sensitivity, particularly in the cloze format.

((b) Impact of the number of few-shot examples on choice sensitivity in Llama3.1-8B-Instruct: Increasing the number of few-shot examples does not consistently reduce choice sensitivity and often increases it in the symbols and hybrid formats.

**((c) Instruction tuning: Instruction-tuned models generally demonstrate lower choice sensitivity compared to their base versions, indicating that instruction tuning "helps to minimize spurious, choice-driven behavior.

Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed the core contributions of this paper, CHOICES SPEAK LOUDER THAN QUESTIONS, focusing on the definition of choice sensitivity and the introduction of Normalized Probability Shift by the Question (NPSQ).

The primary improvement is not in training a new model architecture, but in developing a superior, more reliable evaluation framework for existing Large Language Models (LLMs).

Here are the specific improvements and what they enable:


) Improved AI System Capabilities: Robust and Fair LLM Evaluation Frameworks.

The proposed system allows researchers to move beyond superficial performance metrics (like raw log-likelihood or length-normalized likelihood) that are easily gamed by answer choices, enabling a truly reliable assessment of a model's internal reasoning capabilities.

Specific Improvements:

  1. [System Improvement: Introduction of the NPSQ Metric]

  2. [System Capability 1: Isolating Question Comprehension]

  3. [System Capability 2: Identifying Choice Sensitivity Artifacts]

  4. [System Capability 3: Mitigating Adversarial Choice Manipulation]

  5. The system incorporates the novel metric, the Normalized Probability Shift by the Question (NPSQ), defined as:

NPSQ(Q, C, x) = log P(x Q, C) − log P(x C).

  1. [System Capability 1: Isolating Question Comprehension]

The improved system can precisely quantify the question-driven component of a model's decision. By calculating the NPSQ, researchers can determine how much the presence of the question (Q) genuinely shifts a model's probability towards or away from a specific answer choice (x), independent of that choice’s inherent popularity or superficial structure.

This enables developers to distinguish between:

  • Answers selected because they are inherently plausible options (high choice-driven component).

  • Answers selected because the model successfully integrated the question's constraints into its reasoning (high question-driven shift).

  1. [System Capability 2: Identifying Choice Sensitivity Artifacts]

The system automatically calculates a quantitative Choice Sensitivity score across various input formats (cloze, symbols, hybrid) and model scales. This score identifies which specific model-choice combinations are most vulnerable to superficial cues in the answer options.

This allows for targeted model refinement; if a specific format (e.g., cloze) shows high sensitivity, engineers know to prioritize prompt design or instruction tuning specifically for that format's evaluation pipeline, rather than applying a one-size-fits-all fix.

  1. [System Capability 3: Mitigating Adversarial Choice Manipulation]

The system is designed to be robust against adversarial choices—intentionally irrelevant or implausible distractors—by using the NPSQ metric instead of traditional accuracy (acc) or length-normalized accuracy (acc norm). As demonstrated in Figure 8, NPSQ maintains performance stability when an adversarial choice is introduced, whereas raw metrics suffer significant drops.

This capability ensures that models are evaluated based on their true comprehension of the question's logic, rather than their tendency to select the most statistically probable or superficially appealing option.

In summary, this research provides a diagnostic tool for LLM evaluation. The improved AI system can reliably answer: Does this model understand the question, or is it just picking the best-looking answer?

Sources

Related papers