Scoring the Wrong Question: Readout Failures in Constrained-Option Evaluation
cs.CL, cs.AI, cs.LG
Submitted: 2026-09-27
Updated: 2026-09-27
Code: https://github.com/EleutherAI/lm-evaluation-harness
Terminology
Sources
- Improving LLM First-Token Predictions in Multiple-Choice Question Answering via Output Prefilling
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- SummEval: Re-evaluating Summarization Evaluation
- Measuring Massive Multitask Language Understanding
- TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
- Language Models (Mostly) Know What They Know
- Attention Deficits in Language Models: Causal Explanations for Procedural Hallucinations
- G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
- Leveraging Large Language Models for Multiple Choice Question Answering
- Perturbation CheckLists for Evaluating NLG Evaluation Metrics
- Mind the Gap: A Closer Look at Tokenization for Multiple-Choice Question Answering with LLMs
- Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences
- PerspectiveGap: A Benchmark for Multi-Agent Orchestration Prompting
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- Look at the Text: Instruction-Tuned Language Models are More Robust Multiple Choice Selectors than You Think
- "My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models
- MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
- Large Language Models Are Not Robust Multiple Choice Selectors
- Beyond Yes and No: Improving Zero-Shot LLM Rankers via Scoring Fine-Grained Relevance Labels
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering