What is the Difference Between Me and You? Benchmarking the Quality Gap Between Human-Written and AI-Generated Code
cs.SE, cs.AI
Submitted: 2026-09-11
Updated: 2026-09-21
Comments: Preprint. This manuscript is currently under peer review
Code: https://github.com/github/codeql
Project page: https://pmd.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Program Synthesis with Large Language Models
- Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
- Evaluating Large Language Models Trained on Code
- FullStack Bench: Evaluating LLMs as Full Stack Coders
- DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
- Measuring Coding Challenge Competence With APPS
- Qwen2.5-Coder Technical Report
- CodeSearchNet Challenge: Evaluating the State of Semantic Code Search
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- Infini-gram: Scaling Unbounded n-gram Language Models to a Trillion Tokens
- Debt Behind the AI Boom: A Large-Scale Empirical Study of AI-Generated Code in the Wild
- A Large-Scale Comprehensive Measurement of AI-Generated Code in Real-World Repositories
- Investigating The Smells of LLM Generated Code
- Assessing the Quality and Security of AI-Generated Code: A Quantitative Analysis
- An Empirical Study on Automatically Detecting AI-Generated Source Code: How Far Are We?
- Show Me Your Code! Kill Code Poisoning: A Lightweight Method Based on Code Naturalness
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation
- Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models
- A Survey of Large Language Models for Code: Evolution, Benchmarking, and Future Trends
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties