When Models Edit Too Much: On the Fidelity of Minimal Code Edits
cs.SE, cs.AI, cs.CL
Submitted: 2026-09-03
Updated: 2026-09-03
Code: https://github.com/nreHieW/over-editing
Terminology
Sources
- Program Synthesis with Large Language Models
- Can It Edit? Evaluating the Ability of Large Language Models to Follow Code Editing Instructions
- Evaluating Large Language Models Trained on Code
- EDIT-Bench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits
- SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
- Less is More: Adaptive Program Repair with Bug Localization and Preference Learning
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- DeepSeek-V3 Technical Report
- Edit, But Verify: An Empirical Audit of Instructed Code-Editing Benchmarks
- CoreCodeBench: Decoupling Code Intelligence via Fine-Grained Repository-Level Tasks
- Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
- CodeEditorBench: Evaluating Code Editing Capability of Large Language Models
- SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories?
- Qwen2.5-Coder Technical Report
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- QiMeng-PRepair: Precise Code Repair via Edit-Aware Reward Optimization
- TACO: Topics in Algorithmic COde generation dataset
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties