Self-Bootstrapping Automated Program Repair: Using LLMs to Generate and Evaluate Synthetic Training Data for Bug Repair
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Self-Bootstrapping Automated Program Repair: Using LLMs to Generate and Evaluate Synthetic Training Data for Bug Repair".
Jane: The paper was written by David de-Fitero-Dominguez, Antonio Garcia-Cabot and Eva Garcia-Lopez from University of Alcalá.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Welcome back to the show, everyone. We’ve got a fascinating paper on the table today, and Jane, I have to say, the title alone got me hooked: “Self-Bootstrapping Automated Program Repair: Using LLMs to Generate and Evaluate Synthetic Training Data for Bug Repair.”
Jane: It’s a mouthful, Tom, but it’s really about a clever idea. The authors—David de-Fitero-Domínguez, Antonio García-Cabot, and Eva García-López from the University of Alcalá in Spain—they’re tackling a huge problem: there just isn’t enough good data to train AI to fix bugs in code automatically.
Tom: Right, and they’re saying, why wait for humans to collect thousands of real bugs and fixes? Why not have the AI create its own training examples?
Jane: Exactly. And that’s the “self-bootstrapping” part. The same kind of models that will eventually fix the bugs are used to generate the examples they learn from.
Tom: So it’s like a student writing their own practice problems and then grading them before the exam. That’s wild.
Jane: It is, and the authors show it works. They generated about thirty thousand examples across twelve programming languages and thirteen bug types. Then they had six different large language models grade each example on things like correctness, security, and code quality.
Tom: And the grading matters, right? Because if you just throw all that synthetic data at a model, it might learn bad habits.
Jane: Precisely. They found that filtering for high-quality samples made a huge difference. Their best model improved its bug-fix success rate by about forty-seven percent compared to training only on real-world data.
Tom: Forty-seven percent is a big deal. And they’re not just fixing simple typos—they’re tackling security vulnerabilities in C and C++ code, which is notoriously tricky.
Jane: That’s what makes this paper so exciting. It suggests we can break the data bottleneck that’s been holding back automated repair for years.
Tom: So, Jane, is the takeaway that we don’t need human-curated datasets anymore?
Jane: Not quite. They still used a real dataset as a baseline, but the synthetic data gave the model a big boost on top of it. It’s a complement, not a replacement, at least for now.
Tom: Got it. So the title really captures the essence—the models are both the teachers and the students here. I can’t wait to dig into how they actually built this pipeline.
Jane: Me too. Next up, we’ll walk through their methodology step by step.
Paper Summary: Tom: So, Jane, we’ve set the stage. Now let’s get into the nitty-gritty of how this “self-bootstrapping” actually works in practice.
Jane: The paper lays out a two-phase process. First, they prompt six different large language models—think of them as six different experts—to generate pairs of buggy code and the corrected version.
Tom: And they didn’t just stick to one language. They covered Python, Java, C++, Go, Ruby, Rust, Swift, Kotlin, PHP, C#, JavaScript, and even Pascal.
Jane: That’s right. Each model was asked to produce five thousand examples, with a random bug type and language assigned each time. They ended up with just under thirty thousand valid samples.
Tom: But generating the data is only half the story. The clever part is the evaluation phase. Each of those six models then graded every single sample.
Jane: Yes, and that’s the cross-model evaluation. Each sample gets scored by all six models on five criteria: correctness, code quality, security, performance, and completeness. They use a weighted average, with correctness counting the most.
Tom: So you’re getting a consensus. If one model thinks a fix is great but another thinks it’s sloppy, you catch that.
Jane: Exactly. They found that models disagree quite a bit. Some were stricter graders than others. But by averaging those scores, they could filter out the weak examples.
Tom: And the filtering threshold? They kept samples scoring above eight point five out of ten. That left them with about twenty thousand high-quality examples.
Jane: Right, about seventy percent of the original pool. And then they fine-tuned a smaller model—Qwen two point five Coder 7B—on different combinations of this synthetic data and the real-world VulRepair dataset.
Tom: The results are pretty striking. The model trained with the filtered synthetic data beat the baseline by a wide margin. And crucially, the unfiltered data performed worse than the filtered data.
Jane: That’s the key insight. More data isn’t automatically better. The quality filter is what made the difference.
Tom: It’s like studying for a test with a stack of practice problems, but half of them have wrong answers. You’d be better off with fewer, correct ones.
Jane: Precisely. And they ran the experiments fifty times to make sure the results weren’t just luck. They even used statistical tests—ANOVA and Tukey’s HSD—to confirm the improvements were real.
Tom: So this isn’t a fluke. The methodology is sound.
Jane: It is. And what’s really exciting is that their best model outperformed existing systems like VulMaster and VulRepair, even though those systems used a more computationally expensive search strategy.
Tom: That’s a huge win for efficiency. You’re getting better results with less computing power.
Jane: Exactly. The paper shows that investing in high-quality training data can pay off more than investing in fancier decoding algorithms.
Tom: So, what does this mean for the future of automated program repair? That’s what we’ll dig into next.
Improvements Suggested: Tom: Alright, Jane, we’ve covered the method and the results. But what does this paper actually change for the field? What are the improvements it’s suggesting?
Jane: The biggest improvement is the self-bootstrapping paradigm itself. Instead of relying on scarce, expensive, human-curated datasets, we can use LLMs to generate and evaluate their own training data.
Tom: And that’s not just for bug repair. The authors argue this could apply to other software engineering tasks where paired examples are hard to come by—like code translation or refactoring.
Jane: Right. And they did a really interesting comparison. They tested whether code-specialized models, like CodeLlama or DeepSeek-Coder, would generate better synthetic data than general-purpose models.
Tom: And what did they find?
Jane: Surprisingly, the general-purpose models won. The code-specialized ones were more likely to refuse generating security-related bugs, which hurt their output.
Tom: That’s fascinating. So the models that are supposedly better at code were actually worse at this task because of safety training.
Jane: Exactly. It’s a reminder that instruction-following and willingness to explore are just as important as raw coding ability.
Tom: So what’s the practical improvement for someone actually building an automated repair tool?
Jane: The paper shows you can get better results with a smaller model and fewer decoding candidates. Their model used just five candidates and still beat systems that used fifty.
Tom: That means faster, cheaper repairs. That’s a real win for deployment in real-world settings.
Jane: And they also did an ablation study, tweaking the quality threshold and the weighting of the evaluation criteria. They found the approach is pretty robust—moderate thresholds all work well.
Tom: So you don’t have to fine-tune the parameters obsessively to get the benefit.
Jane: Right. As long as you’re filtering for quality, you’ll see improvement. That makes the methodology much easier to adopt.
Tom: I also noticed they mention using syntax validators as a future direction. That could catch the few invalid examples that slip through.
Jane: Yes, they found about ninety-seven percent of the filtered samples were syntactically valid, but adding a compiler check could push that higher.
Tom: So the improvements are about both data quality and practical efficiency. This feels like a blueprint for other tasks facing data scarcity.
Jane: Absolutely. The authors have shown a path to creating large, diverse, high-quality datasets without manual curation.
Tom: I’m curious to hear what Lu and Meng think about the practical implications. Let’s bring them in.
Conclusion: Tom: We’ve had a great run with “Self-Bootstrapping Automated Program Repair: Using LLMs to Generate and Evaluate Synthetic Training Data for Bug Repair.” Jane, let’s wrap this up.
Jane: It’s a paper that shows us a new way to think about training data. Instead of being a bottleneck, data generation can be part of the model’s own capabilities.
Tom: The key numbers are hard to forget: a forty-seven percent relative improvement in Top@one repair success, and beating existing systems with far less computational effort.
Jane: And the core lesson—quality filtering beats raw quantity. The unfiltered data actually hurt performance compared to the filtered set.
Tom: The authors also gave us a rigorous statistical framework, which is rare in this field. They didn’t just report numbers; they proved the differences weren’t random.
Jane: That’s a standard we should all expect going forward. It makes the results much more trustworthy.
Tom: And the future work is exciting. They mention scaling to larger models, exploring more bug types, and even using compilers as additional quality checks.
Jane: There’s also the potential to apply this self-bootstrapping idea to other tasks—code translation, documentation, refactoring. The sky’s the limit.
Tom: So, as we say goodbye to this paper, I think the takeaway is that the models can be their own teachers. And that’s a powerful idea.
Jane: It really is. Thanks for joining us, everyone. We’ll be back with the next paper soon.
Tom: Until then, keep coding, keep fixing, and keep listening.
David de-Fitero-Dominguez, Antonio Garcia-Cabot, Eva Garcia-Lopez
University of Alcalá
cs.SE, cs.AI
Submitted: 2026-08-16
Updated: 2026-08-18
Journal ref: Expert Systems with Applications 319 (2026)
DOI: 10.1016/j.eswa.2026.132154
Code: https://github.com/sahil280114/codealpaca
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 68/100
The gist: This paper presents a novel methodology for enhancing Automated Program Repair (APR) through synthetic data generation utilizing Large Language Models (LLMs).
Key concepts
- Self-Bootstrapping
- This concept means using the same AI models that will eventually fix bugs to generate the practice examples they learn from. The AI acts as both the teacher and the student, creating its own training material.
- Synthetic Training Data
- This refers to artificial data, like buggy code examples and their correct fixes, created by large language models instead of relying solely on scarce human-curated real-world datasets. This addresses the problem of insufficient bug data.
- Quality Filtering
- The authors found that filtering synthetic samples based on multiple model evaluations—scoring them on correctness, security, and quality—was crucial. Samples scoring above a certain threshold were kept, which improved performance more than simply using all generated data.
Terminology
Summary
This paper presents a novel methodology for enhancing Automated Program Repair (APR) through synthetic data generation utilizing Large Language Models (LLMs). Current APR systems are constrained by the limited availability of high-quality training data encompassing diverse bug types across multiple programming languages. The proposed approach addresses this limitation through a two-phase process: a synthetic sample generation followed by a rigorous quality assessment. Multiple state-of-the-art LLMs were employed to generate approximately 30,000 paired examples of buggy and fixed code across 12 programming languages and 13 bug categories. Subsequently, these samples underwent cross-model evaluation against five criteria: correctness, code quality, security, performance, and completeness. Experimental evaluation on the VulRepair test set dataset showed statistically significant improvements in Perfect Prediction rates, with the quality-filtered synthetic dataset achieving 17.18% (Top@1) and 23.00% (Top@5) compared to the baseline's 11.68% and 18.88% respectively, representing a 47% relative improvement in Top@1 and 22% in Top@5. The methodology was validated through rigorous statistical testing, including ANOVA and post-hoc Tukey's Honest Significant Difference analysis. Furthermore, the best-performing configurations surpassed existing systems despite using a less computationally intensive decoding strategy. This research establishes a self-bootstrapping paradigm in which LLMs generate and evaluate their own training data, suggesting promising directions for addressing data scarcity in similar software engineering tasks and advancing the development of robust, adaptable tools for automated code maintenance.
The main contributions of this work are: (1) A novel methodology for generating and evaluating synthetic data for APR using multiple LLMs, with cross-model evaluation to ensure quality; (2) Empirical evidence demonstrating that quality-filtered synthetic data outperforms both unfiltered data and manually collected training samples; and (3) A rigorous statistical validation framework that confirms the significance of results across multiple experimental configurations, establishing more robust evaluation standards for APR research.
The methodology uses LLMs in a dual role, first as generators of buggy/fixed code pairs, then as evaluators of those same examples, creating a cycle where models both create and assess training data. The generation process integrates multiple frontier models, including NVIDIA's Llama-3.1-Nemotron-70B-Instruct-HF, Meta's Meta-Llama-3.1-70B-Instruct, Qwen2.5-72B-Instruct, Mistral-Small-Instruct-2409, Gemma-2-27B-it, and Qwen2.5-32B-Instruct. The generation pipeline randomly samples from 12 programming languages (Python, Java, C++, Go, Ruby, Rust, Swift, Kotlin, PHP, C#, JavaScript, and Pascal) and 13 bug types (including off-by-one errors, infinite loops, resource leaks, concurrency issues, improper error handling, memory management errors, and security vulnerabilities). Each model generates 5,000 samples, creating a representative dataset across multiple programming contexts.
For quality evaluation, each generated sample is evaluated across multiple dimensions on a 0-10 scale, with the same frontier LLMs used for sample generation serving as evaluators. Each of the six models evaluates the entire dataset, regardless of which model originally generated each sample, meaning each sample receives six independent quality assessments. The evaluation process compares the buggy code with its fixed version to generate a diff
showing changed, added, or removed lines. Individual scores are combined into a single quality score using weighted averages, prioritizing correctness (0.3) as the fundamental requirement, with code quality and length (0.2 each), and security, performance, and completeness (0.1 each). A quality threshold of 8.5 is applied to determine which samples to include in the final dataset.
After the generation phase, the dataset contained 29,646 unique synthetic samples with an average score of 8.58 and standard deviation of 0.86. After applying the quality threshold of 8.5, the dataset was reduced to 20,832 samples (70% retention), with the average quality score increasing to 8.94 and standard deviation reduced to 0.26. Llama-3.1-Nemotron-70B-Instruct-HF contributed the most samples (21.07%), followed by Qwen2.5-72B-Instruct (18.93%) and Meta-Llama-3.1-70B-Instruct (18.75%).
For fine-tuning experiments, the Qwen 2.5 Coder 7B model was selected. Five experimental configurations were evaluated: (Baseline – vulrep) using VulRepair training set only; (1 - vulrep synt 85) Baseline plus filtered synthetic dataset with threshold 8.5; (2 - vulrep commitpack) Baseline plus CommitPackFT; (3 - vulrep synt 85 commitpack) Baseline plus filtered synthetic plus CommitPackFT; and (4 - vulrep synt full) Baseline plus unfiltered synthetic dataset. LoRA was used for efficient fine-tuning with three epochs, learning rate of 3e-4 using a cosine learning rate scheduler, batch size of 256, and maximum sequence length of 4,096 tokens.
Results showed that all configurations using synthetic data outperformed the baseline in both Top@1 and Top@5 settings. The quality-filtered synthetic dataset (vulrep synt 85) achieved substantially better results than the unfiltered version (vulrep synt full), with an improvement of approximately 2 percentage points in both evaluation settings. The combination of filtered synthetic data with CommitPackFT (vulrep synt 85 commitpack) yielded the highest performance in the Top@1 setting (17.26%), while the filtered synthetic data alone (vulrep synt 85) performed best in the Top@5 scenario (23.00%).
Statistical validation confirmed the significance of these differences. All configurations passed the Shapiro-Wilk test for normality and Levene's test for homogeneity of variances. ANOVA results revealed significant differences between configurations in both Top@1 (F=1735.92, p<0.0001) and Top@5 (F=646.27, p<0.0001) settings. Post-hoc Tukey's HSD tests confirmed statistically significant differences between most configuration pairs, with the comparison between vulrep synt 85 and vulrep synt 85 commitpack in Top@1 being the only non-significant pair (p=0.806).
Ablation studies investigated sensitivity to filtering parameters. For threshold analysis, values from 7.0 to 9.0 were tested, finding that moderate thresholds (7.5-8.5) yielded optimal performance, with 7.5 achieving the highest mean PP% (19.58%). The threshold of 8.5 used in main experiments (19.08%) proved near optimal. Higher thresholds (9.0) showed significant performance degradation (17.30%). For weight analysis, the original balanced weighting achieved the best performance (19.08%), followed closely by quality-focused weights (18.78%). All weighted configurations significantly outperformed the baseline without synthetic data.
A comparative study with code-specialized LLMs found that general-purpose models with strong code capabilities are better suited for synthetic data generation than code-specialized models. Code LLMs showed substantial variation in format compliance, with some models achieving near-perfect validity rates (Devstral: 99.96%, DeepSeek-Coder-V2: 99.94%) while others showed significant compliance issues (CodeLlama-70b: 27.58%, Qwen2.5-Coder-32B: 46.18%), primarily due to refusals to generate security-related vulnerabilities. Downstream evaluation confirmed that general-purpose synthetic data yielded better results, with general-purpose data significantly outperforming code LLM data in Top@1 (11.18% vs 10.51%, p=0.045).
The paper acknowledges several threats to validity, including potential correlated errors among evaluator models sharing similar training paradigms, the expert-judgment-based weighting scheme, the pragmatic rather than systematic choice of the 8.5 quality threshold, exclusive evaluation on the VulRepair test set focusing on C/C++ vulnerabilities, and the PP% metric's potential undervaluation of functionally equivalent but syntactically different fixes.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system, along with what the improved system can do:
-
Implementation: Add a two-phase module where the AI generates paired buggy/fixed code examples, then evaluates them using the same or different LLMs across 5 criteria (correctness, code quality, security, performance, completeness) on a 0–10 scale.
-
Specifics: Use a weighted score (correctness 0.3, code quality 0.2, length 0.2, security 0.1, performance 0.1, completeness 0.1) with a quality threshold of 8.5 to filter training data.
-
Implementation: Instead of relying on a single evaluator, use 6 diverse LLMs (e.g., Llama-3.1-70B, Qwen2.5-72B, Mistral-Small, Gemma-2-27B) to score each generated sample. Aggregate scores to reduce individual model bias.
-
Specifics: Use temperature 0.2 for evaluation (vs. 0.7 for generation) to ensure consistent scoring. Enforce structured JSON output using the Outlines library.
-
Implementation: When fine-tuning a code model (e.g., Qwen 2.5 Coder 7B), augment the real-world training set with the filtered synthetic dataset (threshold ≥8.5) rather than using unfiltered data.
-
Specifics: This yields a 47% relative improvement in Top@1 Perfect Prediction (17.18% vs. 11.68%) and 22% in Top@5 (23.00% vs. 18.88%) compared to training on real data alone.
-
Implementation: Modify the model's output format to generate patches as diffs (line ranges with modified lines only) rather than full corrected code. Use special tokens (
,, ``) to structure the output. -
Specifics: This improves efficiency and precision, reducing token generation while maintaining repair accuracy.
-
Implementation: Add a post-training evaluation module that runs each configuration 50 times and applies ANOVA + Tukey's HSD test to verify that performance differences are statistically significant (p<0.0001) before deployment.
-
Specifics: This prevents deploying models with improvements that are due to random chance.
-
Implementation: Before finalizing the system, run ablation studies on quality thresholds (7.0–9.0) and evaluation criteria weights to confirm the chosen parameters are near-optimal and robust.
-
Specifics: Moderate thresholds (7.5–8.5) yield optimal performance; extreme filtering (9.0) degrades results by 2 percentage points.
-
Generate diverse, realistic bug-fix pairs across 12 programming languages (Python, Java, C++, Go, Ruby, Rust, Swift, Kotlin, PHP, C#, JavaScript, Pascal) and 13 bug categories (off-by-one, concurrency, resource leaks, SQL injection, etc.), without requiring existing datasets.
-
Self-assess its own training data using a consensus of multiple LLMs, filtering out low-quality examples (e.g., incorrect fixes, poor security) before training.
-
Achieve higher repair accuracy than systems trained on real-world data alone, even surpassing state-of-the-art systems like VulMaster (20.0% → 23.0% Top@5) while using only 5 candidates instead of 50.
-
Overcome limited training data by generating high-quality synthetic examples on demand, particularly for rare languages (e.g., Rust, Pascal) or complex bug types (e.g., concurrency, memory leaks) where real-world data is scarce.
-
Maintain performance even when real-world training data is small, as the synthetic data supplements rather than replaces it.
-
Provide multi-dimensional quality scores (correctness, security, performance, completeness) for every generated fix, enabling automated triage of which patches to deploy.
-
Detect and reject low-quality samples before they contaminate the training set, reducing the risk of the model learning incorrect repair patterns.
-
Extend the methodology to other paired-code tasks like code translation, refactoring, or documentation generation, where high-quality paired examples are also scarce.
-
Reduce reliance on exhaustive inference strategies (e.g., beam search with 50 candidates) by improving training data quality, cutting computational cost during deployment.
-
Provide statistically rigorous evidence of improvement, with ANOVA and Tukey's HSD tests confirming that observed gains are not due to chance.
-
Enable reproducible experiments with 50 repeated runs and standard deviation reporting, setting a higher bar for empirical validation in APR research.
Abstract
This paper presents a novel methodology for enhancing Automated Program Repair (APR) through synthetic data generation utilizing Large Language Models (LLMs). Current APR systems are constrained by the limited availability of high-quality training data encompassing diverse bug types across multiple programming languages. The proposed approach addresses this limitation through a two-phase process: a synthetic sample generation followed by a rigorous quality assessment. Multiple state-of-the-art LLMs were employed to generate approximately 30,000 paired examples of buggy and fixed code across 12 programming languages and 13 bug categories. Subsequently, these samples underwent cross-model evaluation against five criteria: correctness, code quality, security, performance, and completeness. Experimental evaluation on the VulRepair test set dataset showed statistically significant improvements in Perfect Prediction rates, with the quality-filtered synthetic dataset achieving 17.18% (Top@1) and 23.00% (Top@5) compared to the baseline's 11.68% and 18.88% respectively, representing a 47% relative improvement in Top@1 and 22% in Top@5. The methodology was validated through rigorous statistical testing, including ANOVA and post-hoc Tukey's Honest Significant Difference analysis. Furthermore, the best-performing configurations surpassed existing systems despite using a less computationally intensive decoding strategy. This research establishes a self-bootstrapping paradigm in which LLMs generate and evaluate their own training data, suggesting promising directions for addressing data scarcity in similar software engineering tasks and advancing the development of robust, adaptable tools for automated code maintenance.
Sources
- Program Synthesis with Large Language Models
- A Survey on Evaluating Large Language Models in Code Generation Tasks
- Evaluating Large Language Models Trained on Code
- DeepSeek-V3 Technical Report
- DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
- Fixing Rust Compilation Errors using LLMs
- The Llama 3 Herd of Models
- CodeBERT: A Pre-Trained Model for Programming and Natural Languages
- LoRA: Low-Rank Adaptation of Large Language Models
- Execution-based Evaluation for Data Science Code Generation Models
- A Survey on Automated Program Repair Techniques
- Qwen2.5-Coder Technical Report
- LLM-Powered Code Vulnerability Repair with Reinforcement Learning and Semantic Reward
- Evaluating Language Models as Synthetic Data Generators
- ContrastRepair: Enhancing Conversation-Based Automated Program Repair via Contrastive Test Case Pairs
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- Data Generation Using Large Language Models for Text Classification: An Empirical Case Study
- Best Practices and Lessons Learned on Synthetic Data
- CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation
- OctoPack: Instruction Tuning Code Large Language Models
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties