Does CoT-Pass@k Really Check the CoT? A Multilingual Mathematical Audit
cs.CL
Submitted: 2026-09-26
Updated: 2026-09-26
Code: https://github.com/ttasalti/evalhub
Terminology
Sources
- Measuring what Matters: Construct Validity in Large Language Model Benchmarks
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Evaluating Large Language Models Trained on Code
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Beyond Pass@k: Breadth-Depth Metrics for Reasoning Boundaries
- Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish
- The Last Word Often Wins: A Format Confound in Chain-of-Thought Corruption Studies
- Gemma 4 Technical Report
- Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
- Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning?
- A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains
- Measuring Faithfulness in Chain-of-Thought Reasoning
- Let's Verify Step by Step
- Are Your LLMs Capable of Stable Reasoning?
- Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
- C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning
- Qwen2.5 Technical Report
- Qwen3 Technical Report
- Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering