Does Fixing Break Security? An Empirical Study of Security Degradation in Iterative LLM-Driven Infrastructure-as-Code Repair

arXiv:2608.13404 · cs.SE, cs.CR · Submitted 2026-08-14 · Read on arXiv

Benjamin Agyekum, Fabio Santos

Colorado State University

cs.SE, cs.CR

Submitted: 2026-08-14

Updated: 2026-08-17

Comments: 20 pages, 3 figures, 5 tables. Accepted at the 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). To appear in LIPIcs Vol. 394

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

The gist: Background and Motivation: Iterative feedback loops have become the dominant paradigm for improving LLM-generated Infrastructure-as-Code (IaC): validators such as Checkov and terraform validate feed

Terminology

Summary

Background and Motivation: Iterative feedback loops have become the dominant paradigm for improving LLM-generated Infrastructure-as-Code (IaC): validators such as Checkov and terraform validate feed error signals back to the model for successive repair attempts. Prior work reports cumulative-best metrics, which are monotonically non-decreasing by construction, so the raw per-iteration security trajectory has never been examined in the IaC domain. The paper notes that "prior work on iterative IaC repair universally reports cumulative-best metrics: the highest compliance achieved up to iteration k. These are monotonically non-decreasing by construction and may mask important dynamics in the repair trajectory. Across 5,968 scenario timelines, the cumulative-best CIS pass rate rises from 73% to 83%, yet the raw trajectory dips at iteration 5 (82.6% vs. the 83.4% peak at iteration 4) as previously-passing checks fail."

Research Questions: The study addresses four research questions: RQ1: Does iterative LLM repair introduce security regressions in IaC? RQ2: What types of security controls are most vulnerable to regression? RQ3: How do prompting strategy, model, and temperature affect regression? RQ4: Is there a security-correctness trade-off in iterative IaC repair?

Methodology: The study analyzed 5,968 scenario timelines from the IaC-Eval benchmark (458 Terraform generation scenarios) across 15 configurations (six RAG and nine non-RAG configurations, three temperatures each: 0.1, 0.4, and 0.7), yielding 4,440 iteration transitions with Checkov data on both sides. The configurations comprise four prompting strategies (Zero-Shot, Few-Shot, Chain-of-Thought, and RAG) using two models (Gemini 2.0 Flash and Mistral Large Latest), with up to 5 repair iterations per scenario. The study tracked 30 individual CIS check IDs mapped to CIS AWS Foundations Benchmark controls across six security categories (encryption, access control, logging, networking, data protection, and other). Two detection modes were employed: standard (inclusive) detection counts any check that passes at iteration i and fails at iteration i+1, while strict detection requires a check to be exclusively passed (not in failed) in iteration i and exclusively failed (not in passed) in iteration i+1, excluding ambiguous multi-resource cases.

Key Findings:

RQ1 (Regression prevalence): Under standard detection, 13.8% of scenarios (823 of 5,968, 95% CI: [12.9%, 14.7%]) exhibit at least one security regression event, and 24.8% of transitions (1,103 of 4,440, CI: [23.6%, 26.1%]) contain at least one regression, producing 2,639 total regression events. Under strict detection, 194 scenarios (3.3%, CI: [2.8%, 3.7%]) exhibit regression, with 282 total events across 231 transitions (5.2% of transitions, CI: [4.6%, 5.9%]). Approximately 76% of standard-mode-flagged scenarios are not flagged in strict mode, consistent with multi-resource ambiguity rather than exclusive check failures. The paper notes this is lower than general code (37.6% in Shukla et al.), but still substantial.

RQ2 (Vulnerable controls): Regression events are highly concentrated rather than spread evenly (chi-squared goodness-of-fit test: χ2 = 1445, p < 0.001, Cramér's V = 0.33). In standard mode, access control dominates (38.5%), driven by three IAM-related checks: CKV AWS 356 (298 events), CKV AWS 111 (297), and CKV AWS 109 (291). Strikingly, the category ranking reverses in strict mode: networking (42.6%) and encryption (21.6%) dominate, while access control drops to 9.2%. The three most frequently regressing checks (886 events, 33.6% of all standard-mode regressions) all govern IAM privilege boundaries.

RQ3 (Configuration factors): In standard mode, the model effect is dramatic: "Gemini's scenario regression rates range from 1.9% to 4.7% across temperatures, while Mistral ranges from 32.6% to 41.2%. Mistral scenarios are over 17 times more likely to regress (OR = 17.29, p < 0.001). However, this gap vanishes entirely in strict mode: neither RAG+Gemini nor RAG+Mistral produces a single strict-mode scenario regression (0 of 1,297 and 0 of 929 respectively). The RAG effect reverses between modes: In standard mode, RAG has a higher regression rate... OR = 1.67, p < 0.001. In strict mode, however, RAG records zero scenario regressions, versus 4–6% for the non-RAG strategies." Temperature has a statistically significant but practically negligible effect (Cramér's V = 0.049).

RQ4 (Security-correctness trade-off): Transitions with regressions show significantly more code modification than those without: the average number of lines changed between consecutive iterations is 140.1 vs. 53.8, a 2.6× difference. Check volatility is the single strongest signal in our study: in strict mode, 11.55 vs. 2.38 (4.9×), with Cohen's d = 1.49. "In transitions immediately following a syntax fix... These show a 21% higher regression rate than transitions between two already-valid iterations (28.1% vs. 23.3%, OR = 1.29, p < 0.001). Counter-intuitively, scenarios that always had valid syntax ('clean' timelines) show a higher regression rate than those that struggled with syntax errors: clean timelines regress in 20.8% of cases versus 12.1% for syntax-struggled timelines" (OR = 1.90, p < 0.001).

Root Cause Taxonomy: Resource restructuring (79.0%) is the dominant cause of regressions, followed by configuration drift (15.5%), argument removal (3.6% standard, 8.2% strict), and unclassified (1.9%). In strict mode, the distribution shifts to 68.4% restructuring, 19.9% drift, 8.2% argument removal, and 3.5% unclassified.

Temporal Patterns: Regression events peak in early-to-middle transitions (0→1: 769 events, 1→2: 711, 2→3: 616), with the per-transition rate peaking at 2→3 (29.5%) and 1→2 (26.6%). "Of 2,639 standard-mode regressions, 967 (36.6%) self-correct in a subsequent iteration... On average, self-correction takes 1.2 iterations after the regression occurs (median: exactly 1), with 80.4% correcting within a single step. Self-correction is at least as common in strict mode (44.0%). 28.5% of all scenarios (1,698 of 5,968) exhibit oscillating checks," with the three most oscillation-prone checks all IAM-related.

Optimal Stopping Point: "Iteration 3 offers the best trade-off. We anchor this recommendation on the pass-rate trajectory... the pass rate reaches 83.1% at iteration 3, within 0.3pp of the 83.4% maximum at iteration 4, and then decreases to 82.6% at iteration 5."

Conclusions: The paper presents six findings: (1) Regression is real: 13.8% of scenarios regress in standard mode, 3.3% in strict mode. (2) Detection mode changes conclusions: access control dominates standard mode while networking and encryption are the genuine strict-mode risks, the 17× Mistral-vs-Gemini gap vanishes, and RAG reverses from worst to best. (3) Resource restructuring is the dominant root cause (79.0%). (4) Code churn and check volatility are indicators of regressions, with strict-mode volatility the strongest signal (d = 1.49). (5) Self-correction is common but unstable: 36.6% of standard-mode regressions self-correct, yet 28.5% of scenarios oscillate. (6) Iteration 3 is the optimal stopping point, balancing an 83.1% pass rate against manageable regression risk. The paper concludes: "Iterative IaC repair does introduce security regressions, but most apparent regressions are multi-resource measurement artifacts. The conservative, defensible rate is approximately 3.3% of scenarios. Our findings motivate security-aware feedback-loop design and provide actionable iteration-budget guidance."

Improvements for AI systems

Improvements to AI Systems:

  1. Add regression-aware stopping criteria to iterative repair loops. Instead of blindly running a fixed number of repair iterations, the AI system should monitor per-iteration check status and halt at iteration 3 (the empirically optimal point) when pass rate peaks at 83.1%, avoiding the regression dip at iteration 5. This prevents unnecessary code churn that correlates with regressions (140.1 vs. 53.8 lines changed).

  2. Implement a dual-mode regression detector with ambiguity resolution. The AI system should distinguish between exclusive check failures (strict mode) and multi-resource ambiguities (standard mode). When a check transitions from pass to fail, the system must verify whether the failure is exclusive to a single resource or confounded by multiple resources, reducing false-positive regression alerts by 76%.

  3. Prioritize security checks by true regression risk, not frequency. The system should re-weight its validation focus: in strict mode, networking (42.6%) and encryption (21.6%) checks are the genuine regression risks, not access control (9.2%) which dominates standard-mode artifacts. Specifically, the system should allocate more repair budget to CKV AWS networking/encryption checks and treat IAM privilege-boundary checks (CKV AWS 356, 111, 109) as lower-priority for regression monitoring.

  4. Add a resource-restructuring guardrail. Since 79.0% of regressions stem from resource restructuring (e.g., splitting, merging, or renaming Terraform resources), the AI system should detect when a repair iteration restructures resources and automatically re-run a full security validation on all affected resources, not just the modified lines. This prevents silent security regressions during large refactors.

  5. Implement a self-correction tracker with oscillation detection. The system should track each check's pass/fail history across iterations. If a check oscillates (pass→fail→pass), the system should flag it as unstable and either (a) freeze the last-passing configuration or (b) apply a targeted fix, rather than continuing to modify code. This addresses the 28.5% of scenarios with oscillating checks and the 36.6% self-correction rate that is unreliable.

  6. Add a code-churn and check-volatility early-warning system. Before applying a repair, the system should estimate the expected lines-changed and check-volatility. If the proposed edit exceeds a threshold (e.g., >140 lines changed or >11.5 check volatility), the system should either simplify the edit or split it into smaller, independently validated steps, reducing regression likelihood by up to 4.9×.

  7. Configure model and prompting strategy based on regression risk profile. The system should use RAG-based prompting for strict-mode reliability (0% scenario regressions) and avoid high-temperature settings for models like Mistral (32.6–41.2% regression rate). For non-RAG configurations, the system should prefer Gemini over Mistral (1.9–4.7% vs. 32.6–41.2% regression rate) when strict security guarantees are needed.

  8. Add a post-syntax-fix regression checkpoint. Since transitions immediately after syntax fixes show 21% higher regression rates, the system should perform an extra validation pass specifically after resolving syntax errors, before proceeding to the next repair iteration. This catches regressions introduced during the syntax-correction step.

  9. Implement a regression-aware iteration budget optimizer. The system should dynamically allocate iteration budgets per scenario: stop at iteration 3 for most cases, but extend to iteration 4 only if the pass rate is still improving and no regression events have occurred in the last two transitions. This balances the 83.4% peak pass rate at iteration 4 against the regression risk at iteration 5.

  10. Add a root-cause classification module for regression feedback. When a regression is detected, the system should automatically classify it into one of four categories (resource restructuring, configuration drift, argument removal, unclassified) and apply category-specific remediation: for restructuring, re-validate all resources; for drift, restore the previous configuration; for argument removal, re-add the missing argument with a comment explaining its security necessity.

Abstract

Background: Iterative feedback loops are the dominant paradigm for improving LLM-generated Infrastructure-as-Code (IaC): validators such as Checkov and terraform validate feed error signals back for successive repair attempts. Prior work reports cumulative-best metrics, which are non-decreasing by construction, so the raw per-iteration security trajectory has never been examined for IaC. Aims: We study security regression (a previously-passing CIS Benchmark check that fails after a repair iteration) to determine whether and how often iterative LLM repair degrades security while fixing other issues. Method: We analyze 5,968 scenario timelines from the IaC-Eval benchmark, each one scenario run through one configuration for up to 5 repair iterations. The 15 configurations (six model-specific RAG, nine model-aggregated non-RAG, three temperatures each) yield 4,440 iteration transitions with Checkov data on both sides. We track 30 individual CIS check IDs and classify root causes from code diffs, under two detection modes: standard (inclusive) and strict (exclusive check failures only). Results: Under standard detection, 13.8% of scenarios (24.8% of transitions) exhibit at least one regression. Under strict detection the rate falls to 3.3% of scenarios (5.2% of transitions), indicating most apparent regressions are multi-resource measurement artifacts. Resource restructuring (79.0%) is the dominant root cause. Regression transitions show 2.6x more code churn (Cohen's d=0.90) and 4.9x higher strict-mode check volatility (d=1.49). Of standard-mode regressions, 36.6% self-correct within an average of 1.2 iterations; iteration 3 is the optimal stopping point. Conclusions: Iterative IaC repair does introduce security regressions, but the conservative, defensible rate is about 3.3% of scenarios. Our findings motivate security-aware feedback-loop design and actionable iteration-budget guidance.

Sources

Related papers