Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?

arXiv:2604.17338 · cs.SE, cs.CL · Submitted 2026-04-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Precise Debugging Benchmark".

Tom: The gist The PRECISE DEBUGGING BENCHMARKING (PDB) framework introduces two novel metrics, edit-level precision and bug-level recall,

Jane: First, who's behind it and why it matters.

Title and authors: Tom: So we're diving into this new paper, "Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?" and it's asking a really direct question about how well these AI models are actually fixing code.

Jane: Exactly. It suggests that right now, when we test these models, they often just generate a whole new correct solution instead of making the small, targeted changes needed to fix the specific bug they were asked about.

Lu: This framework introduces two metrics: edit-level precision and bug-level recall, which are designed to reward models for making exactly the right edits instead of just rewriting everything.

Meng: So we're moving away from just checking if the final code runs, and instead measuring how many necessary modifications were actually applied to get there.

Tom: Right. The paper sets up this whole system called PDB, which is designed to be general enough to take any coding dataset and turn it into a debugging challenge with this new precision focus.

Jane: And the core idea behind that is that they generate bugs by taking verified atomic bugs and putting them together into bigger programs, but keeping them independent so they don't mess with each other.

Lu: They define precision as measuring how many necessary edits are made compared to the total set of ground-truth edits, which is decomposed into different groups of lines.

Meng: And recall is about how many distinct bugs out of the total number of bugs in the program actually get resolved by the model's proposed fixes.

Tom: The results show that frontier models behave very differently when you look at this edit-level evaluation compared to just standard unit tests, which is quite telling.

Jane: Apparently, some models like GPT-five point one-Codex and DeepSeek-V3 point 2-Thinking score high on unit test pass rates but their edit precision is relatively low, under forty-five percent.

Lu: But then you have Qwen3-Coder-480B, which might have a lower unit test score around seventy percent, yet its precision jumps up to sixty-six percent.

Meng: That suggests that focusing on precise fixes matters more than just getting the code to pass a simple test suite.

Tom: And they also found that using iterative or agentic debugging strategies doesn't actually improve these precision or recall scores much at all, which is a big hint about where we need to focus our next steps.

Jane: It seems like these post-training pipelines for coding models really need to change if we want them to be truly precise debuggers.

Lu: The paper points out that current evaluation methods are mostly based on questions and answers from places like Stack Overflow, which introduces a lot of contamination into the data they use.

Meng: So the framework itself is designed to be dataset-agnostic, meaning it can take any existing code collection and turn it into this structured benchmark for rigorous testing.

Tom: The authors also show four different types of debugging strategies that emerge from these tests, ranging from models that pass with precision to those that are just pass-oriented regeneration machines.

Jane: It seems like the real value is in distinguishing between a model that's being truly targeted and one just trying to generate a big solution quickly.

Lu: The conclusion is pretty straightforward: we need edit-level evaluation because it helps us separate targeted fixing from superficial pass-driven rewriting.

Meng: They are pushing for this framework to become both an evaluation tool and a way to give reward signals that focus on finding faults and making minimal edits.

Tom: It sounds like the big picture here is that we're moving toward training coding models with feedback specifically focused on fault localization and how small the required fixes should be.

Jane: Before we wrap up this discussion, Lu, what's your take on how this shifts the possibilities for AI development?

Lu: I think this framework gives us a way to create a much tighter feedback loop where models learn to be surgical rather than just broad generators of code.

Meng: From an engineering standpoint, it means we can start training these models with reward signals that directly penalize unnecessary changes while still ensuring functional correctness is maintained.

Lalam: I see this as really important for how we shape the culture around building these tools; if they learn precision first, the resulting code will be much cleaner and easier to maintain over time.

Tom: So, we're looking at a path where models focus on fault localization and edit minimality through this precise debugging benchmark. That's what we have from "Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?".

Jane: It really underscores the gap between what these models are doing now and what real-world debugging requires.

Lu: We'll keep an eye on how developers use this PDB framework to test and improve their own coding models moving forward.

Meng: And we need to see if this approach can translate into better, more efficient AI assistants in the next generation of tools.

Lalam: It’s about making the output of these systems not just correct, but intelligently corrected.

The paper's summary: Tom: So we’re talking about this Precise Debugging Benchmark paper now and what all that summary means for us in practice.

Jane: Basically, they’re saying that right now, when we test these models, we’re mostly just checking if the final code runs or if it passes a simple unit test.

Lu: But this work introduces two new ways to measure performance: edit-level precision and bug-level recall. It’s about measuring how targeted the fixes actually are versus just throwing in whole new solutions.

Meng: So they’re trying to figure out if an AI is actually fixing the specific error it was given, or if it's just regenerating a whole new program because the old one failed.

Tom: Exactly. The paper shows that there’s a big difference between models that look good on basic tests and models that actually make surgical changes to fix problems.

Jane: They tested this on frontier models, and some of them, like GPT-five point one-Codex, can pass those basic tests really well but their edit precision is quite low.

Lu: On the flip side, other models show high precision even if their unit test scores aren't as high—which is a huge point because it shows that making the right change matters more than just getting a passing score.

Meng: The engineers are seeing this too: iterative or agentic debugging strategies don't really boost these precision or recall numbers much, which tells us we need to rethink how we train these models post-training.

Tom: So what does this mean for the next generation of coding AI? It suggests that the focus needs to shift from just functional correctness to fault localization and making minimal, targeted edits.

Jane: It points toward a future where we can reward models not just for getting the right answer, but for being surgical about how they get there.

Lu: The PDB framework itself is designed to be general, so it can take any coding dataset and turn it into this structured debugging challenge for rigorous testing.

Meng: If we use these precision metrics, we can actually train models to be more efficient because they’ll learn to avoid unnecessary changes during the creation process.

Tom: It sounds like the goal is creating a tighter feedback loop where models focus on minimal edits and precise fault localization instead of just broad code generation.

Jane: It really forces us to ask what kind of reward signals we should be feeding these coding models if we want them to become truly reliable debugging partners.

Lu: We need this framework because it helps separate targeted debugging behavior from superficial pass-driven rewriting, which is a big problem right now.

The paper's improvements: Tom: So we’re looking at what this paper suggests we should actually do with all this PDB framework now that we know how it works.

Jane: The main idea is that this evaluation system isn't just for showing off; it’s meant to change the way we train and improve these coding AI models.

Lu: They are proposing using these edit-level precision and bug-level recall scores as direct reward signals during training. It’s about teaching the AI to be surgical, not just broad.

Meng: That means we can test if making precise edits actually helps the model avoid over-editing when it's trying to get a functional result. That’s a practical engineering goal right there.

Tom: Exactly. If we use these signals, we test whether precision-aware objectives can reduce those massive amounts of unnecessary code rewriting while still keeping the program working correctly.

Jane: It moves the focus away from just getting a final correct output toward optimizing the *path* to that correct output through minimal changes.

Lu: The framework itself is set up to be dataset-agnostic, meaning it can plug into any existing coding data and turn it into this structured debugging test automatically.

Meng: That opens up a lot of possibilities for creating self-improving coding models that prioritize fault localization over just generating massive blocks of code.

Tom: It’s about building a system where the AI learns that making a small, accurate fix is more valuable than writing ten lines of new code to solve the same issue.

Jane: This shifts the entire culture around training these models toward rewarding efficiency and accuracy in their modifications rather than just output size or superficial success.

Lu: It’s about closing that loop where the AI gets feedback specifically focused on how small and necessary its edits should be, which is a big step toward better reasoning.

Tom: So, what does this mean for the future of coding assistance? It means we could see models that are not just fast at writing code, but also incredibly efficient at fixing it when things go wrong.

Jane: It suggests that the next big step isn't just bigger models, but smarter training methods that care deeply about the quality and precision of every single line change they propose.

Lu: This is how we move toward systems that are truly helpful because they are precise, not just lucky with their regeneration.

Conclusion: Tom: Alright folks, so we’re wrapping up our look at this Precise Debugging Benchmark paper and what all this means for how we build coding models.

Jane: Basically, they’ve shown us that just checking if code runs isn't enough anymore; we need to measure exactly how precise the fixes are when an AI is debugging a program.

Lu: The big implication is that we can start training AI models with rewards that actually push them toward making small, targeted edits instead of just blindly rewriting everything.

Meng: From an engineering standpoint, this gives us a way to test if precision-aware objectives can actually make the model avoid over-editing while still hitting the functional correctness targets.

Lalam: I think this is huge for culture because it sets a new standard for what we expect from AI assistants; it moves us past just generating code to actually fixing it intelligently.

Tom: It’s about demanding surgical accuracy instead of just broad attempts at success. This whole Precise Debugging Benchmark paper really forces us to rethink our post-training pipelines for these coding tools.

Jane: So, the authors are suggesting we adopt this framework as both a way to evaluate models and a structure for training them in a better way.

Lu: They’re providing the infrastructure so we can rigorously test whether these precise metrics actually translate into better real-world debugging skills.

Meng: The limitation they point out is that currently, the data generation might be focused on Python programs, so we need to see how quickly this framework adapts to other languages.

Tom: True. And it’s worth remembering that these precision and recall metrics are more accurate than just unit tests, but they still might miss some very subtle fixes that look different but do the same thing.

Jane: It’s a good reminder that we aren't quite there yet with truly precise debugging, which is exactly what this work highlights.

Lu: We’ve seen how different strategies perform in PDB-SINGLE and PDB-WILD, showing us four distinct ways models can behave when they’re trying to fix bugs.

Meng: It shows that the current iterative methods aren't really improving these precision scores much, which tells us we need a real overhaul of how we handle model training.

Tom: That’s the core message of the Precise Debugging Benchmark paper: we need to focus on fault localization and edit minimality for better coding AI.

Jane: It really underscores the gap between what these models can do now and what robust, precise debugging actually requires.

Lu: We'll keep watching how developers use this PDB framework to test their models and try to bring that precision into the next generation of tools.

Wang Bill Zhu, Miaosen Chai, Shangshang Wang, Yejia Liu, Song Bian, Honghua Dong

University of Southern California · Microsoft

cs.SE, cs.CL

Submitted: 2026-04-19

Updated: 2026-10-05

Comments: NeurIPS 2026 Evaluations and Datasets

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 91/100

The gist: The gist The PRECISE DEBUGGING BENCHMARKING (PDB) framework introduces two novel metrics, edit-level precision and bug-level recall, to evaluate how far frontier LLMs are from precise debugging by

Key concepts

Edit-Level Precision
This metric measures how close a model's suggested fix is to the ground truth. It calculates the ratio of correctly applied edits to total edits made. A high score means the model makes very few unnecessary modifications, focusing only on fixing what is broken.
Bug-Level Recall
Recall measures how many of the actual bugs present in a program were successfully resolved by the model's changes. This focuses on whether the core issues are addressed, regardless of how many lines of code were touched. It assesses if the model correctly identifies and fixes all intended defects.
PDB Framework
This is a general system that turns any coding dataset into a debugging benchmark. It automatically synthesizes verified atomic bugs and combines them into complex programs to create realistic testing scenarios for LLMs, allowing for precise evaluation of debugging skills.

Terminology

Summary

The gist The PRECISE DEBUGGING BENCHMARKING (PDB) framework introduces two novel metrics, edit-level precision and bug-level recall, to evaluate how far frontier LLMs are from precise debugging by measuring how many necessary edits are made and how many bugs are resolved.

How it works

The PDB framework is a general, dataset-agnostic pipeline that converts any coding dataset into a debugging benchmark with precision-aware evaluation. PDB generates buggy programs by synthesizing verified atomic bugs and composing them into multibug programs.

The framework involves two main steps:

  1. Synthesizing verified atomic bugs to produce ground-truth edit scripts.

  2. Composing these bugs into multi-bug programs while preserving bug independence (i.e., avoiding compounding interactions).

Metrics and Evaluation

PDB evaluates model patches using novel edit-level precision and bug-level recall, explicitly rewarding targeted fixes and penalizing unnecessary modifications.

The metrics are defined as follows:

(1) Precision:

precision = 1/Ê X k i=1 FU (Cˆ i) · Ei, where Egt ∈ ECb is the set of ground-truth edits, which can be decomposed as Egt = E1∪E2∪· · ·∪Ek, where E1,..., Ek are contiguous and non-overlapping.

(2) Recall:

recall = 1/k X k i=1 FU (Cˆ i), where k is the number of bugs.

The paper notes that precision functions as an edit-level metric by averaging over the edits Ê, while recall is a bug-level metric averaged over the k bugs.

Experimental Findings

Experiments on PDB-SINGLE reveal behaviors that unit tests fail to capture. Frontier models exhibit strikingly different rankings under editlevel evaluation. For instance, GPT-5.1-Codex and DeepSeek-V3.2-Thinking achieve high unit-test pass rates (>76%) but low edit precision (≤45%), whereas Qwen3-Coder-480B attains comparatively lower unit-test pass rates (70%) yet substantially higher precision (66%)76%) but low edit precision (≤45%), while Qwen3-Coder-480B attains comparatively lower unit test pass rates (70%) yet substantially higher precision (66%)>.

Furthermore, the paper shows that iterative and agentic debugging strategies do not substantially improve precision or recall. This highlights the need to rethink post-training pipelines for coding models.

Analysis of Debugging Strategies

The results from PDB-SINGLE and PDB-WILD demonstrate four different types of model debugging strategies. These include:

  1. Pass with precision: Models like Claude-Sonnet-4.5 and Gemini-2.5-Pro debug correctly (>75%) with the highest precision (>71%) and recall (>81%)75%) with the highest precision (>71%) and recall (>81%)>.

  2. Weak but precise: Models like Qwen3-Coder-480B show moderately high precision (66%) and recall (77%), despite only achieving a 70% unit score.

  3. Weak, imprecise, but identifying: Models like Kimi-K2-Instruct reliably identify buggy regions but struggle to produce correct and precise fixes, with precision below 57%.

  4. Pass-oriented: Models such as DeepSeek-V3.2 exhibit substantially lower precision (≤48%) with recall below unit test scores, indicating a regeneration-heavy strategy that relies on broad rewrites.

Conclusion and Implications

The findings demonstrate the necessity of edit-level evaluation for distinguishing targeted debugging behavior from superficial pass-driven regeneration. The paper concludes that current systems remain far from achieving precise, edit-aware debugging, highlighting a fundamental limitation in current post-training pipelines for coding LLMs. The PDB framework is suggested as both an evaluation benchmark and infrastructure for closing the training loop in self-improving coding models by providing reward signals focused on fault localization and edit minimality.

Limitations

A limitation noted is that current prompts and data generation procedures target Python programs, which may limit immediate applicability to other programming languages. Additionally, the work focuses on benchmark construction and evaluation rather than training or posttraining models using PDB signals. Finally, while precision and recall metrics are more accurate than unit-test-only evaluation, they may still fail to capture certain correct but semantically equivalent fixes.

Data Generation Details

The PDB generation pipeline involves synthesizing atomic bugs from existing coding datasets and composing them into multi-bug programs. This process utilizes Orthogonal Defect Classification (ODC) categories to guide bug injection, ensuring diversity across defect types. The evaluation uses a tolerance parameter epsilon for precision evaluation to allow up to Ei + ϵ edited lines per bug. The final benchmarks, PDB-SINGLE and PDB-WILD, consist of 5,751 and 484 examples respectively<ref:4,PDB-SINGLE and PDB-WILD evaluation benchmarks with the PDB framework, from BigCodeBench [Zhuo et al., 2024], LiveCodeBench [Jain et al., 2024], and SWE-smith [Yang et al., 2025a]>.

Improvements for AI systems

  1. Precise evaluation of debugging capabilities allows for targeted post-training pipeline adjustments by providing edit-level precision and bug-level recall metrics, explicitly rewarding targeted fixes and penalizing unnecessary modifications. This enables developers to focus on fault localization and minimal, intent-preserving edits rather than relying solely on functional correctness.

  2. Development of a general framework like PRECISE DEBUGGING BENCHMARKING (PDB) allows for the creation of dataset-agnostic plug-and-play framework that converts existing coding datasets into debugging benchmarks, making it possible to rigorously evaluate LLM behavior independently of code generation.

  3. Training models using PDB signals can be used to test whether precision-aware objectives can reduce over-editing while preserving functional correctness, suggesting a pathway for self-improving coding models that focus on fault localization and edit minimality.

Abstract

Unlike code completion, debugging requires localizing faults and applying targeted edits. We observe that frontier LLMs often regenerate correct but over-edited solutions during debugging. To evaluate how far LLMs are from precise debugging, we introduce the Precise Debugging Benchmark (PDB) framework, which automatically converts any coding dataset into a debugging benchmark with precision-aware evaluation. PDB generates buggy programs by synthesizing verified atomic bugs and composing them into multi-bug programs. We define two novel metrics, edit-level precision and bug-level recall, which measure how many necessary edits are made and how many bugs are resolved. We release two evaluation benchmarks: PDB-Single-Hard on single-line bugs, and PDB-Multi on multi-line bugs. Experiments show that frontier models, such as GPT-5.1-Codex and DeepSeek-V3.2-Thinking, achieve unit-test pass rates above 76% but exhibit precision below 45%, even when explicitly instructed to perform minimal debugging. Finally, we show that iterative and agentic debugging strategies do not substantially improve precision or recall, highlighting the need to rethink post-training pipelines for coding models.

Sources

Related papers