ReasonLab: A Controlled and Auditable Evaluation of Prompting Techniques for Multiple-Choice QA
cs.CL, cs.AI
Submitted: 2026-05-07
Updated: 2026-09-22
License: http://creativecommons.org/licenses/by/4.0/
The gist: Probing the capabilities of Large Language Models (LLMs) and building robust solutions for Multiple-Choice Question Answering (MCQA) remain central challenges in natural language understanding.
Terminology
Abstract
Probing the capabilities of Large Language Models (LLMs) and building robust solutions for Multiple-Choice Question Answering (MCQA) remain central challenges in natural language understanding. Furthermore, the rapid proliferation of LLMs has created the implicit assumption that more sophisticated prompting techniques yield better performance. Several studies claim such gains, but report them under differing models, prompt wordings and answer-extraction rules, so the gains cannot be attributed to the technique alone. We address this gap with ReasonLab, an evaluation framework in which the prompting technique is a first-class experimental variable alongside the model and the dataset, and which retains every generation for inspection. Using ReasonLab we conduct a controlled study of 8 prompting techniques across 10 MCQA datasets, 27 model configurations and 480,927 evaluations at temperature 0. We find that the prompting technique is a minor determinant of accuracy: on configurations without a reasoning budget the reasoning triggers improve on direct prompting by only 3.92 to 4.69 pp and are indistinguishable from one another, and on configurations with reasoning enabled no technique differs by more than 0.51 pp. Self-Generate is the only technique with a consistent effect, a reduction of 2.95 pp. We further investigate three phenomena: (1) the comparison of models on a common set of datasets, where model size does not predict accuracy, (2) the trade-offs across thinking budgets, where enabling reasoning is worth up to 12.74 pp whereas an eightfold budget increase adds only 0.48 to 2.10 pp, and (3) the variation in dataset difficulty, with 60% of benchmarks below 70% accuracy and a 43.9 pp spread from easiest to hardest. These results suggest that, for MCQA, the prompting technique is a minor lever compared with enabling model reasoning, and that substantial headroom remains.
Sources
- On the Measure of Intelligence
- FailureSensorIQ: A Multi-Choice QA Dataset for Understanding Sensor Relationships and Failure Modes
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Measuring Massive Multitask Language Understanding
- Time-MQA: Time Series Multi-Task Question Answering with Context Enhancement
- Prompt Repetition Improves Non-Reasoning LLMs
- Holistic Evaluation of Language Models
- SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
- Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
- s1: Simple test-time scaling
- Qwen3 Technical Report
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
- Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
- Generate rather than Retrieve: Large Language Models are Strong Context Generators
- Scaling Physical Reasoning with the PHYSICS Dataset
- PromptBench: A Unified Library for Evaluation of Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering