Check The Scoreboard: An Analysis of Scoring Schemes on Multiple-Choice Evaluation
cs.CL
Submitted: 2026-08-30
Updated: 2026-08-30
Comments: EMNLP 2026
Code: https://github.com/nbalepur/mcqa-scoring
License: http://creativecommons.org/licenses/by/4.0/
The gist: Multiple-choice question answering (MCQA) benchmarks in NLP use number-right scoring (accuracy), but in educational testing, the scoring scheme, the combination of the response mode models follow and
Terminology
Abstract
Multiple-choice question answering (MCQA) benchmarks in NLP use number-right scoring (accuracy), but in educational testing, the scoring scheme, the combination of the response mode models follow and the rule for grading responses, is a key design choice that dictates which abilities to reward. We examine how alternatives to number right change what MCQA measures with six education-inspired schemes that assess abilities beyond accuracy: distractor elimination, abstention, confidence calibration, and self-correction. On LLM benchmarks, these schemes: 1) shift rankings of 31 LLMs beyond rephrased number right prompts; 2) better predict the LLMs users prefer in LLM Arena; and 3) reveal distinct model capabilities, like that GPT-5 rarely abstains and readily self-corrects, while weaker open-weight models often abstain and hesitate to eliminate choices. Given the benefits of alternative scoring schemes, we discuss ways to extend them to tasks beyond MCQA.
Sources
- gpt-oss-120b & gpt-oss-20b Model Card
- AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
- BeHonest: Benchmarking Honesty in Large Language Models
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena
- Olmo 3
- OpenAI GPT-5 System Card
- Kimi K2: Open Agentic Intelligence
- Qwen3 Technical Report
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering