Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
cs.SE, cs.AI, cs.LG
Submitted: 2026-04-14
Updated: 2026-08-27
Code: https://github.com/MrLYG/CodeRQ-Bench
Terminology
Sources
- Evaluating Large Language Models Trained on Code
- What Defines Good Reasoning in LLMs? Dissecting Reasoning Steps with Multi-Aspect Evaluation
- ClassEval: A Manually-Crafted Benchmark for Evaluating LLMs on Class-level Code Generation
- Competitive Programming with Large Reasoning Models
- ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning
- Measuring Faithfulness in Chain-of-Thought Reasoning
- CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution
- Mitigating Hallucinations in Large Language Models via Causal Reasoning
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models
- CodeMind: Evaluating Large Language Models for Code Reasoning
- CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation
- Code Summarization Beyond Function Level
- OpenAI o1 System Card
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness
- Code Llama: Open Foundation Models for Code
- CoRe: Benchmarking LLMs Code Reasoning Capabilities through Static Analysis Tasks
- Simple BERT Models for Relation Extraction and Semantic Role Labeling
- On the Hardness of Faithful Chain-of-Thought Reasoning in Large Language Models
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties