From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench
cs.SE, cs.AI, cs.CL
Submitted: 2026-08-27
Updated: 2026-08-27
Code: https://github.com/DeepSoftwareAnalytics/MCR-bench
Terminology
Sources
- Sea Change in Software Development: Economic and Productivity Analysis of the AI-Powered Developer Lifecycle
- CodeFuse-CR-Bench: A Comprehensiveness-aware Benchmark for End-to-End Code Review Evaluation in Python Projects
- Retrieval-Augmented Code Review Comment Generation
- Deep Assessment of Code Review Generation Approaches: Beyond Lexical Similarity
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- SecureReviewer: Enhancing Large Language Models for Secure Code Review through Secure-aware Fine-tuning
- StarCoder 2 and The Stack v2: The Next Generation
- ReleaseEval: A Benchmark for Evaluating Language Models in Automated Release Note Generation
- When More Retrieval Hurts: Retrieval-Augmented Code Review Generation
- Kimi K2: Open Agentic Intelligence
- Towards an Understanding of Context Utilization in Code Intelligence
- RepoReasoner: Evaluating Repository-Level Code Reasoning Ability of Long-Context Language Models
- Qwen3 Technical Report
- Large Language Models are Qualified Benchmark Builders: Rebuilding Pre-Training Datasets for Advancing Code Intelligence Tasks
- SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment Generation
- Sphinx: Benchmarking and Modeling for LLM-Driven Pull Request Review
- AACR-Bench: Evaluating Automatic Code Review with Holistic Repository-Level Context
- HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation
- A Survey of Large Language Models for Code: Evolution, Benchmarking, and Future Trends
- MemoryBank: Enhancing Large Language Models with Long-Term Memory
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties