Self-Bootstrapping Automated Program Repair: Using LLMs to Generate and Evaluate Synthetic Training Data for Bug Repair

summary

Video file (mp4)

The gist

This paper presents a novel methodology for enhancing Automated Program Repair (APR) through synthetic data generation utilizing Large Language Models (LLMs).

In short

The episode discusses a paper on 'Self-Bootstrapping Automated Program Repair' where LLMs generate and evaluate their own synthetic training data for bug fixing. The authors found that filtering synthetic data by quality significantly improved the model's success rate by about forty-seven percent compared to using real-world data alone, demonstrating that high-quality filtering is more important than sheer quantity.

Key concepts

Self-Bootstrapping
This concept means using the same AI models that will eventually fix bugs to generate the practice examples they learn from. The AI acts as both the teacher and the student, creating its own training material.
Synthetic Training Data
This refers to artificial data, like buggy code examples and their correct fixes, created by large language models instead of relying solely on scarce human-curated real-world datasets. This addresses the problem of insufficient bug data.
Quality Filtering
The authors found that filtering synthetic samples based on multiple model evaluations—scoring them on correctness, security, and quality—was crucial. Samples scoring above a certain threshold were kept, which improved performance more than simply using all generated data.

Terminology used across episodes

This episode discusses

The paper

Self-Bootstrapping Automated Program Repair: Using LLMs to Generate and Evaluate Synthetic Training Data for Bug Repair · Read on arXiv

David de-Fitero-Dominguez, Antonio Garcia-Cabot, Eva Garcia-Lopez

University of Alcalá

This paper presents a novel methodology for enhancing Automated Program Repair (APR) through synthetic data generation utilizing Large Language Models (LLMs). Current APR systems are constrained by the limited availability of high-quality training data encompassing diverse bug types across multiple programming languages. The proposed approach addresses this limitation through a two-phase process: a synthetic sample generation followed by a rigorous quality assessment. Multiple state-of-the-art LLMs were employed to generate approximately 30,000 paired examples of buggy and fixed code across 12 programming languages and 13 bug categories. Subsequently, these samples underwent cross-model evaluation against five criteria: correctness, code quality, security, performance, and completeness. Experimental evaluation on the VulRepair test set dataset showed statistically significant improvements in Perfect Prediction rates, with the quality-filtered synthetic dataset achieving 17.18% (Top@1) and 23.00% (Top@5) compared to the baseline's 11.68% and 18.88% respectively, representing a 47% relative improvement in Top@1 and 22% in Top@5. The methodology was validated through rigorous statistical testing, including ANOVA and post-hoc Tukey's Honest Significant Difference analysis. Furthermore, the best-performing configurations surpassed existing systems despite using a less computationally intensive decoding strategy. This research establishes a self-bootstrapping paradigm in which LLMs generate and evaluate their own training data, suggesting promising directions for addressing data scarcity in similar software engineering tasks and advancing the development of robust, adaptable tools for automated code maintenance.

DOI: 10.1016/j.eswa.2026.132154

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Self-Bootstrapping Automated Program Repair: Using LLMs to Generate and Evaluate Synthetic Training Data for Bug Repair".

Jane: The paper was written by David de-Fitero-Dominguez, Antonio Garcia-Cabot and Eva Garcia-Lopez from University of Alcalá.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Authors: Tom: Welcome back to the show, everyone. We’ve got a fascinating paper on the table today, and Jane, I have to say, the title alone got me hooked: “Self-Bootstrapping Automated Program Repair: Using LLMs to Generate and Evaluate Synthetic Training Data for Bug Repair.”

Jane: It’s a mouthful, Tom, but it’s really about a clever idea. The authors—David de-Fitero-Domínguez, Antonio García-Cabot, and Eva García-López from the University of Alcalá in Spain—they’re tackling a huge problem: there just isn’t enough good data to train AI to fix bugs in code automatically.

Tom: Right, and they’re saying, why wait for humans to collect thousands of real bugs and fixes? Why not have the AI create its own training examples?

Jane: Exactly. And that’s the “self-bootstrapping” part. The same kind of models that will eventually fix the bugs are used to generate the examples they learn from.

Tom: So it’s like a student writing their own practice problems and then grading them before the exam. That’s wild.

Jane: It is, and the authors show it works. They generated about thirty thousand examples across twelve programming languages and thirteen bug types. Then they had six different large language models grade each example on things like correctness, security, and code quality.

Tom: And the grading matters, right? Because if you just throw all that synthetic data at a model, it might learn bad habits.

Jane: Precisely. They found that filtering for high-quality samples made a huge difference. Their best model improved its bug-fix success rate by about forty-seven percent compared to training only on real-world data.

Tom: Forty-seven percent is a big deal. And they’re not just fixing simple typos—they’re tackling security vulnerabilities in C and C++ code, which is notoriously tricky.

Jane: That’s what makes this paper so exciting. It suggests we can break the data bottleneck that’s been holding back automated repair for years.

Tom: So, Jane, is the takeaway that we don’t need human-curated datasets anymore?

Jane: Not quite. They still used a real dataset as a baseline, but the synthetic data gave the model a big boost on top of it. It’s a complement, not a replacement, at least for now.

Tom: Got it. So the title really captures the essence—the models are both the teachers and the students here. I can’t wait to dig into how they actually built this pipeline.

Jane: Me too. Next up, we’ll walk through their methodology step by step.

Paper Summary: Tom: So, Jane, we’ve set the stage. Now let’s get into the nitty-gritty of how this “self-bootstrapping” actually works in practice.

Jane: The paper lays out a two-phase process. First, they prompt six different large language models—think of them as six different experts—to generate pairs of buggy code and the corrected version.

Tom: And they didn’t just stick to one language. They covered Python, Java, C++, Go, Ruby, Rust, Swift, Kotlin, PHP, C#, JavaScript, and even Pascal.

Jane: That’s right. Each model was asked to produce five thousand examples, with a random bug type and language assigned each time. They ended up with just under thirty thousand valid samples.

Tom: But generating the data is only half the story. The clever part is the evaluation phase. Each of those six models then graded every single sample.

Jane: Yes, and that’s the cross-model evaluation. Each sample gets scored by all six models on five criteria: correctness, code quality, security, performance, and completeness. They use a weighted average, with correctness counting the most.

Tom: So you’re getting a consensus. If one model thinks a fix is great but another thinks it’s sloppy, you catch that.

Jane: Exactly. They found that models disagree quite a bit. Some were stricter graders than others. But by averaging those scores, they could filter out the weak examples.

Tom: And the filtering threshold? They kept samples scoring above eight point five out of ten. That left them with about twenty thousand high-quality examples.

Jane: Right, about seventy percent of the original pool. And then they fine-tuned a smaller model—Qwen two point five Coder 7B—on different combinations of this synthetic data and the real-world VulRepair dataset.

Tom: The results are pretty striking. The model trained with the filtered synthetic data beat the baseline by a wide margin. And crucially, the unfiltered data performed worse than the filtered data.

Jane: That’s the key insight. More data isn’t automatically better. The quality filter is what made the difference.

Tom: It’s like studying for a test with a stack of practice problems, but half of them have wrong answers. You’d be better off with fewer, correct ones.

Jane: Precisely. And they ran the experiments fifty times to make sure the results weren’t just luck. They even used statistical tests—ANOVA and Tukey’s HSD—to confirm the improvements were real.

Tom: So this isn’t a fluke. The methodology is sound.

Jane: It is. And what’s really exciting is that their best model outperformed existing systems like VulMaster and VulRepair, even though those systems used a more computationally expensive search strategy.

Tom: That’s a huge win for efficiency. You’re getting better results with less computing power.

Jane: Exactly. The paper shows that investing in high-quality training data can pay off more than investing in fancier decoding algorithms.

Tom: So, what does this mean for the future of automated program repair? That’s what we’ll dig into next.

Improvements Suggested: Tom: Alright, Jane, we’ve covered the method and the results. But what does this paper actually change for the field? What are the improvements it’s suggesting?

Jane: The biggest improvement is the self-bootstrapping paradigm itself. Instead of relying on scarce, expensive, human-curated datasets, we can use LLMs to generate and evaluate their own training data.

Tom: And that’s not just for bug repair. The authors argue this could apply to other software engineering tasks where paired examples are hard to come by—like code translation or refactoring.

Jane: Right. And they did a really interesting comparison. They tested whether code-specialized models, like CodeLlama or DeepSeek-Coder, would generate better synthetic data than general-purpose models.

Tom: And what did they find?

Jane: Surprisingly, the general-purpose models won. The code-specialized ones were more likely to refuse generating security-related bugs, which hurt their output.

Tom: That’s fascinating. So the models that are supposedly better at code were actually worse at this task because of safety training.

Jane: Exactly. It’s a reminder that instruction-following and willingness to explore are just as important as raw coding ability.

Tom: So what’s the practical improvement for someone actually building an automated repair tool?

Jane: The paper shows you can get better results with a smaller model and fewer decoding candidates. Their model used just five candidates and still beat systems that used fifty.

Tom: That means faster, cheaper repairs. That’s a real win for deployment in real-world settings.

Jane: And they also did an ablation study, tweaking the quality threshold and the weighting of the evaluation criteria. They found the approach is pretty robust—moderate thresholds all work well.

Tom: So you don’t have to fine-tune the parameters obsessively to get the benefit.

Jane: Right. As long as you’re filtering for quality, you’ll see improvement. That makes the methodology much easier to adopt.

Tom: I also noticed they mention using syntax validators as a future direction. That could catch the few invalid examples that slip through.

Jane: Yes, they found about ninety-seven percent of the filtered samples were syntactically valid, but adding a compiler check could push that higher.

Tom: So the improvements are about both data quality and practical efficiency. This feels like a blueprint for other tasks facing data scarcity.

Jane: Absolutely. The authors have shown a path to creating large, diverse, high-quality datasets without manual curation.

Tom: I’m curious to hear what Lu and Meng think about the practical implications. Let’s bring them in.

Conclusion: Tom: We’ve had a great run with “Self-Bootstrapping Automated Program Repair: Using LLMs to Generate and Evaluate Synthetic Training Data for Bug Repair.” Jane, let’s wrap this up.

Jane: It’s a paper that shows us a new way to think about training data. Instead of being a bottleneck, data generation can be part of the model’s own capabilities.

Tom: The key numbers are hard to forget: a forty-seven percent relative improvement in Top@one repair success, and beating existing systems with far less computational effort.

Jane: And the core lesson—quality filtering beats raw quantity. The unfiltered data actually hurt performance compared to the filtered set.

Tom: The authors also gave us a rigorous statistical framework, which is rare in this field. They didn’t just report numbers; they proved the differences weren’t random.

Jane: That’s a standard we should all expect going forward. It makes the results much more trustworthy.

Tom: And the future work is exciting. They mention scaling to larger models, exploring more bug types, and even using compilers as additional quality checks.

Jane: There’s also the potential to apply this self-bootstrapping idea to other tasks—code translation, documentation, refactoring. The sky’s the limit.

Tom: So, as we say goodbye to this paper, I think the takeaway is that the models can be their own teachers. And that’s a powerful idea.

Jane: It really is. Thanks for joining us, everyone. We’ll be back with the next paper soon.

Tom: Until then, keep coding, keep fixing, and keep listening.

More episodes

← Home