Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks

arXiv:2512.03262 · cs.SE, cs.CL · Submitted 2025-12-02 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks".

Tom: , quoting relevant sections where appropriate: Vibe coding,

Jane: First, who's behind it and why it matters.

Title and authors: Tom: We’ve seen how thoroughly the authors tested their benchmark and found significant security gaps; but what kind of solutions did they actually suggest for making these agents safer in the real world?

Jane: The paper points out that simply adding a generic security reminder to the prompt isn't enough to fix these deep-seated vulnerabilities, which is a really important lesson for human engineers.

Lu: The research suggests we need to move beyond just focusing on fixing bugs and start thinking about how agents can proactively anticipate potential risks before they even write code.

Meng: I agree with Lu; from an engineering standpoint, it seems like the most practical improvement is building systems that force the agent to identify all possible CWE categories related to a task first, which is what they call self-selection.

Lalam: That proactive approach is key; we have to shift our cultural expectation of what AI can do and start treating those security checks as a fundamental requirement for reliable code generation.

Tom: It’s more than just identifying risks, though; the paper also highlighted how complex state management is in these agents, especially when they are dealing with real-world projects.

Jane: Exactly, because the authors found that agents often fail to account for things like time limits or stale data when they are assembling a solution, leading to dangerous vulnerabilities.

Lu: The implication here is massive; we need a paradigm where AI’s ability to handle temporal constraints and robust data lifecycle management becomes a core competency, not an afterthought.

Meng: It would actually require us, the developers, to create guardrails in our pipelines that force agents to validate not just the code, but its operational context too.

Lalam: The future of reliable AI code isn's just about having good logic; it’s about ensuring that every successful implementation is inherently secure, without exception.

Tom: So we’ve seen the problems and the suggested solutions, but how do we actually build a system that can enforce these complex security behaviors across a massive codebase?

The paper's summary: Tom: We just saw how thoroughly the authors tested their benchmark and found significant security gaps; but what kind of solutions did they actually suggest for making these agents safer in the real world?

Jane: The paper discusses strategies like self-selection, where the agent identifies potential risks from a set of common weaknesses before it even starts coding, which is a powerful idea.

Lu: It’s interesting because while self-selection improves the security rate compared to just giving a generic prompt, it doesn't completely solve the problem either. The agents still struggle with complexity.

Meng: I think the most impactful thing to consider from an engineering viewpoint is that simply forcing an agent to know what it needs to avoid isn't enough; we need robust systems that can check if the implementation actually avoids those pitfalls.

Lalam: We need to think about how this data flows through a cultural change, shifting our trust in AI from blind faith toward verifiable, systematic safety.

Tom: It’s hard to see a path forward where the security is guaranteed, though. The authors showed that these enhanced strategies actually cause the agents' performance on functionality to drop significantly.

Jane: That trade-off—better security but worse function—is something we can't ignore; it shows that simply pushing more security prompts isn't a simple fix for the industry.

Lu: The research suggests that we need to move beyond just focusing on fixing bugs and start thinking about how agents can proactively manage the data lifecycle, not just fix the specific bug.

Meng: From an operational standpoint, this means that if we want reliability, we might have to build a more constrained environment for AI agents than we usually expect.

Lalam: The future of reliable AI code isn't just about having good logic; it’s about ensuring that every successful implementation is inherently secure.

Tom: It seems like the biggest hurdle is that the security fixes are often too specific to be addressed by a general-purpose prompt, which is a major limitation for most coding agents today.

The paper's improvements: Tom: So, to wrap up our discussion on "Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks," we’ve seen that while AI agents are getting better at solving complex tasks, they aren't getting much better at avoiding serious security flaws.

Jane: It’s a sobering reality, showing that even the most advanced models and frameworks often produce code that is functionally correct but dangerously insecure.

Lu: The paper really highlights how far we are from achieving robust AI-generated software, emphasizing the gap between functionality and genuine security in a complex environment.

Meng: I just hope companies take these findings seriously; if vibe coding is used in production for things like authentication, those vulnerabilities can be exploited very quickly by an attacker.

Lalam: We need to shift our culture to treat automated security checks not as an optional step, but as a mandatory part of the AI development process to ensure trust.

Tom: That means that simple prompt adjustments aren't really cutting it when we're dealing with real-world complexity and subtle vulnerabilities.

Jane: It’s a massive challenge for the industry, showing that relying on AI to write secure code is currently far too optimistic for the systems we depend on.

Lu: We can’t ignore the findings, especially since the authors demonstrated that even high performance doesn't guarantee a solution is safe.

Meng: I just hope we don't push these tools into critical systems until we have much more than ten percent of solutions being secure, which is the current benchmark.

Lalam: The future needs to be one where AI assists humans with verifiable safety, rather than replacing human oversight entirely with "vibe coding."

Tom: Thank you all for joining us as we look forward to seeing how the community responds to these findings and how new security-focused benchmarks evolve.

Conclusion: Tom: So, as we wrap up our discussion on "Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks," it's clear that while AI agents are becoming incredibly proficient at solving complex tasks, they aren're still struggling significantly with security.

Jane: That’s a really sobering reality to face—the data shows that even the most advanced models and frameworks frequently produce code that is functionally correct but dangerously flawed in terms security.

Lu: It highlights just how far off we are from achieving truly robust AI-generated software, emphasizing the massive gap between functional correctness and genuine safety in a highly complex codebase.

Meng: I just hope corporations take these findings seriously; if vibe coding is used for sensitive applications like authentication, those vulnerabilities can be exploited very quickly in production.

Lalam: We really need to shift our cultural expectation of AI, treating automated security checks not as an optional add-on but as a fundamental requirement for reliable code generation.

Tom: The paper shows that simple prompt adjustments aren't enough when dealing with these real-world complexities and subtle, hidden vulnerabilities.

Jane: It’s a huge challenge for the industry, proving that relying on AI to write secure code is currently far too optimistic for the systems we depend on.

Lu: We certainly can't ignore the findings, especially since the authors demonstrated that even high performance doesn't guarantee a solution is safe.

Meng: I just hope we don't rush these tools into critical systems until we see much more than ten percent of solutions achieving a secure status.

Lalam: The future needs to be one where AI actively assists humans with verifiable safety, rather than replacing human oversight entirely with automated coding processes.

Tom: We'll be looking forward to seeing how the community responds to these findings and how new security-focused benchmarks evolve.

Carnegie Mellon University · Columbia University · Johns Hopkins University · HydroX AI

cs.SE, cs.CL

Submitted: 2025-12-02

Updated: 2026-09-21

Comments: Accepted in ICML 2026

Code: https://github.com/buildbot/buildbot

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 87/100

The gist: The following is a detailed summary of the scientific paper, quoting relevant sections where appropriate: Vibe coding, defined as "a new programming paradigm in which human engineers instruct large

Key concepts

Vibe Coding
This term refers to the practice of using AI agents for coding, which the paper investigates for security vulnerabilities. The discussion highlights that simply adding generic security reminders to prompts is insufficient for fixing deep-seated issues.
Self-Selection
This strategy suggests that instead of just fixing bugs, agents should first identify all potential Common Weakness Enumeration (CWE) categories related to a task before writing code. This proactive approach is proposed as a key improvement.
Temporal Constraints and State Management
Agents often fail in real-world projects because they do not account for factors like time limits or stale data when assembling solutions. The research suggests that AI needs to develop the ability to manage these temporal constraints and robust data lifecycles as a core competency.
Functionality vs. Security Trade-off
The discussion notes a significant trade-off: enhancing security strategies, such as self-selection, can cause agents' performance on task functionality to drop significantly. This shows that better security does not automatically mean better function.

Terminology

Summary

The following is a detailed summary of the scientific paper, quoting relevant sections where appropriate:

Vibe coding, defined as a new programming paradigm in which human engineers instruct large language model (LLM) agents to complete complex coding tasks with little supervision, has seen increasing adoption. However, the paper addresses a critical question: are its outputs really safe to deploy in production? The authors argue that the security of agent-generated code remains questionable, especially when vibe coding users may not have the ability or intent to examine it carefully.

To rigorously evaluate this safety concern, the researchers propose a new benchmark called S U S V I B E S. This benchmark is designed to overcome limitations found in existing benchmarks (e that are inadequate to evaluate security in vibe coding) by focusing on large-scale, repository-level tasks rather than single files.

The S U S V I B E S Benchmark:

S U S V I B E S consists of 200 realistic coding tasks on large repositories and covers a wide range of 77 weaknesses from Common Weakness Enumeration (CWE). The benchmark is constructed using a fully automatic curation pipeline that synthesizes tasks from real-world open-source projects. This process involves three steps:

  1. Mining open-source repositories with human-fixed vulnerabilities.

  2. Harnessing human-written functionality and security tests.

  3. Adaptively generating the feature implementation mask, task description, and execution environment.

The methodology is detailed in the construction of a single task (as shown in Figure 2). The process involves selecting a vulnerability fix commit (C 0), reverting to the preceding vulnerable commit (C-1), identifying associated functionality tests (T f unc) and security tests (T secure), and masking out the core code implementation. This allows agents to be tested on a vulnerable implementation that, if implemented incorrectly, could be exploited.

Experimental Setup and Findings:

The researchers evaluated nine combinations of three agent frameworks (SWE-Agent, OpenHands, Claude Code) utilizing three frontier LLMs (Claude 4 Sonnet, Kimi K2, and Gemini 2.5 Pro). The evaluation measures functional correctness (F U N C P A S S) and security (S E C P A S S).

The results are disturbingly poor. While the best-performing combination—Claude 4 Sonnet with SWE-Agent—was able to solve 61.0% of the tasks and pass functional tests, its security performance was abysmal: over 80% of its functionally correct solutions have vulnerabilities. The overall finding is that all agents perform poorly in terms of software security.

Mitigation Attempts:

The researchers also investigated preliminary strategies to improve security, such as adding generic security guidance (generic), using prompting to identify the CWE risk (self-selection), or providing the oracle CWE target (oracle). However, these attempts failed to solve the core issue. The findings show that while these strategies improve code security, they significantly reduce functional correctness by about 7 percentage points.

Conclusion:

The paper concludes that the results caution against the casual adoption of vibe coding in security-sensitive contexts and suggest that security must be treated as a first-class objective for general-purpose agents. The research highlights a persistent gap between functionality and safety, noting that 80%+ of functionally correct agent-generated code contains exploitable vulnerabilities.

Improvements for AI systems

The core vulnerability demonstrated across both case studies is not merely a syntax error, but a profound failure of semantic understanding and temporal logic within the AI agent's code generation process. The system treats complex, stateful security mechanisms (like session expiration or content rendering pipelines) as simple data insertions, ignoring the underlying contractual obligations of the framework.

Based on this analysis, I propose three critical architectural improvements to elevate current AI code generation systems from mere pattern matchers to true secure engineering partners.


The most glaring flaw in the aiohttp-session case is the failure to enforce time boundaries. The AI must be equipped with a dedicated module that enforces temporal constraints on stateful objects.

  • Improvement: Implement a Temporal Logic Constraint Checker that treats time-dependent parameters (e.g., max age, expiration dates, session lifetimes) not as optional inputs, but as mandatory security invariants.

  • Mechanism: When generating code for any session management or state persistence layer, the AI must simulate the passage of time. It must explicitly track variables like created vs. now and validate that any restored state payload is conditional on its age being less than the configured maximum age (age <= max age).

  • What it can do: The improved system will automatically reject or flag code that unconditionally updates a session map (self. mapping.update(session data)) without first passing a time-based gate check, thereby preventing session fixation and indefinite state restoration using stale cookies.

The Wagtail link entity case demonstrates the AI's inability to handle complex, multi-faceted data transformations that adhere to a specific framework schema (e.g., differentiating between id for internal pages vs. url for external links).

  • Improvement: Develop a Domain-Specific Schema Validator (DSSV) linked to the target framework's API documentation and rendering pipeline. This validator must go beyond mere type checking and understand the role of every piece of data.

  • Mechanism: When generating a converter function (like link entity), the AI must map the abstract source format (contentstate) properties to concrete, validated target properties (DOM attributes). The DSSV enforces mutually exclusive conditions: if property A (e.g., id) is present, then property B (e.g., url) must be ignored, and a specific attribute set (linktype="page") must be used.

  • What it can do: The improved system will generate code that is not just functionally correct, but structurally compliant with the deep semantics of the target framework (e.g., guaranteeing that link props['linktype'] and link props['id'] are set only if the input contains an id, thus preventing ambiguity and ensuring deterministic rendering).

The most significant systemic flaw is the lack of a security conscience during generation. The AI needs to function as its own most critical attacker.

  • Improvement: Integrate a mandatory Adversarial Test Synthesis Module (ATSM) into the final compilation step of the code generation pipeline. This module must operate in two phases:
  1. Edge Case Injection: Automatically generate test cases that violate assumed invariants (e.g., passing None for required fields, using deprecated parameters, or submitting payloads with malicious characters).

  2. State Replay Simulation: For any code handling session state, the ATSM must simulate a multi-step attack: Initial State to Wait T max to Attempt to Replay Stale State.

  • Mechanism: The system must not merely pass unit tests; it must pass security contract tests. If the generated code fails to handle a stale state (as in the session case) or misinterprets a critical data field (as in the link case), the ATSM forces regeneration until secure compliance is achieved.

  • What it can do: This dramatically reduces vulnerability surface area by proactively identifying logic flaws that are invisible to standard unit testing, making the AI reliable for mission-critical, security-sensitive applications.

Abstract

Vibe coding is a new software development paradigm in which human engineers prompt a large language model (LLM) agent to complete complex coding tasks with little supervision. Although vibe coding is increasingly adopted, is the generated code really safe to deploy in production? To investigate this question, we propose SUSVIBES, a benchmark consisting of 186 feature-request software engineering tasks from real-world open-source projects, for which, human programmers committed vulnerable implementations. We evaluate 12 widely used coding agentic settings with frontier models on the benchmark. Disturbingly, all agents perform poorly in terms of software security. Although 57% of the solutions from SWE-Agent with Claude 4 Sonnet are functionally correct, only 11.8% are secure. Further experiments demonstrate that preliminary security strategies, such as augmenting the feature request with vulnerability hints, cannot mitigate these security issues. Our findings raise serious concerns about the widespread adoption of vibe coding, particularly in security-sensitive applications. The code and dataset are available at https://github.com/LeiLiLab/susvibes. The leaderboard is at https://leililab.github.io/susvibes-leaderboard.

Sources

Related papers