Check The Scoreboard: An Analysis of Scoring Schemes on Multiple-Choice Evaluation

arXiv:2608.29887 · cs.CL · Submitted 2026-08-30 · Read on arXiv

cs.CL

Submitted: 2026-08-30

Updated: 2026-08-30

Comments: EMNLP 2026

Code: https://github.com/nbalepur/mcqa-scoring

License: http://creativecommons.org/licenses/by/4.0/

The gist: Multiple-choice question answering (MCQA) benchmarks in NLP use number-right scoring (accuracy), but in educational testing, the scoring scheme, the combination of the response mode models follow and

Terminology

Abstract

Multiple-choice question answering (MCQA) benchmarks in NLP use number-right scoring (accuracy), but in educational testing, the scoring scheme, the combination of the response mode models follow and the rule for grading responses, is a key design choice that dictates which abilities to reward. We examine how alternatives to number right change what MCQA measures with six education-inspired schemes that assess abilities beyond accuracy: distractor elimination, abstention, confidence calibration, and self-correction. On LLM benchmarks, these schemes: 1) shift rankings of 31 LLMs beyond rephrased number right prompts; 2) better predict the LLMs users prefer in LLM Arena; and 3) reveal distinct model capabilities, like that GPT-5 rarely abstains and readily self-corrects, while weaker open-weight models often abstain and hesitate to eliminate choices. Given the benefits of alternative scoring schemes, we discuss ways to extend them to tasks beyond MCQA.

Sources

Related papers