Does Fault Localization Beat a Fresh Attempt? A Placebo-Controlled Study of Test-Guided Code Repair
cs.SE, cs.AI, cs.LG
Submitted: 2026-09-01
Updated: 2026-09-01
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Loc2Repair: A Framework for Evaluating the Impact of File-Level Issue Localization in Repo-Level LLM Repair
- RGFL: Reasoning Guided Fault Localization for Automated Program Repair Using Large Language Models
- Falsification, Not Exposure: An Internally Preregistered Placebo-Controlled Decomposition of Self-Repair Feedback in Frozen Small Code Models
- Form, Not Content? A Preregistered, Placebo-Controlled Evaluation of Learned Error-Conditioned Self-Repair Through Prompts and Weights in Frozen Small Code Models
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- Automated Repair of Programs from Large Language Models
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation
- Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems
- Characterizing Multi-Hunk Patches: Divergence, Proximity, and LLM Repair Challenges
- Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?
- On the Role of Fault Localization Context for LLM-Based Program Repair
- SHERLOC: Structured Diagnostic Localization for Code Repair Agents
- A Study on the Impact of Fault localization Granularity for Repository-Scale Code Repair Tasks
- Try Again, Don't Look Back: Blind Resampling Outperforms Self-Repair in Small Code Models
- Aligning the Objective of LLM-based Program Repair
- Practical Program Repair in the Era of Large Pre-trained Language Models
- RepairLLaMA: Efficient Representations and Fine-Tuned Adapters for Program Repair
- Automated Repair of C Programs Using Large Language Models
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties