TitanCA: Lessons from Orchestrating LLM Agents to Discover 100+ CVEs

arXiv:2604.17860 · cs.CR · Submitted 2026-08-24 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "TitanCA: Lessons from Orchestrating LLM Agents to Discover 100+ CVEs".

Jane: The paper was written by N/A (Author list for primary paper not present in excerpt) from National Research Foundation, Singapore and Monash University and Singapore Management University and University of Groningen and IEEE Security & Privacy Magazine and Centre of Research on Intelligent Software Engineering (RISE), Singapore Management University and University of Computer Studies and Government Technology Agency of Singapore and Imperial College London and Tsinghua University and Southern University of Science and Technology (SUSTech) and The University of Sydney.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Paper discussion segment 1: Jane: Now that we know the scope, let's look at how the authors summarize TitanCA’s process in "TitanCA: Lessons from Orchestrating LLM Agents to Discover one hundred plus CVEs." The summary outlines a powerful sequence of steps that are designed to filter out noise and focus only on genuine risks.

Tom: They don't just rely on one pass over the code; they use a multi-layered approach, which is what makes the system so robust against common AI pitfalls like hallucination or superficial pattern matching.

Lu: It sounds like they are building a sort of confidence ladder—each module has to validate the findings of the previous one before moving forward.

Meng: The description suggests that this process is intentionally redundant, which is smart because security analysis cannot afford to miss anything due to a single point of failure or bias.

Lalam: It’s less about finding *any* bug and more about proving, step-by-step, that a bug exists and can be exploited under specific conditions.

Jane: They emphasize that the initial search is intentionally broad, which is crucial because if you narrow the scope too early, you risk missing entire classes of vulnerabilities.

Tom: And then they use structured reasoning to guide the subsequent modules, effectively forcing the AI to write a logical justification for every claim it makes about a bug.

Meng: That’s what I found most compelling—it turns the black box of an LLM into a transparent, auditable chain of thought that security engineers can actually follow.

Lu: It's not just saying "this is vulnerable"; it's saying, "The logic flow from line A to line B creates condition C, which is exploitable because of feature D."

Lalam: This meticulous level of documentation in the process itself elevates the tool from a mere detector to a research assistant that teaches you *why* something is wrong.

Tom: We've covered the overall structure, but next we need to understand how this system gets even better—how it adapts to environments outside of its initial training data.

Jane: That brings us perfectly into the discussion of architectural improvements suggested by the paper.

Paper discussion segment 2: Tom: Moving into the improvements section of "TitanCA: Lessons from Orchestrating LLM Agents to Discover one hundred plus CVEs," the focus shifts heavily onto adaptability, which is frankly where most large-scale enterprise tools struggle. The core concept here is bridging the gap between generalized AI knowledge and highly specific company codebases.

Jane: They introduce the Adapter module, often called PairVul in the paper, which acts as a crucial bridge for domain-specific challenges. This means it can handle unique coding styles or internal frameworks that a model trained on public GitHub repositories would never know about.

Meng: That’s absolutely critical for deployment because if the tool is only good on textbook examples, it’s useless in the messy reality of proprietary enterprise code.

Lu: The key mechanism they highlight is the feedback loop—the system doesn't just wait for a massive retraining dataset; it actively learns from its own mistakes in a real-world environment.

Lalam: This moves the technology from being a static snapshot of knowledge to being a living, breathing component that evolves alongside the organization's software development practices.

Tom: It’s about continuous improvement through operational data, which is significantly more valuable and accessible than trying to source massive new public datasets.

Jane: The paper details how they use false positives generated during initial testing as valuable training material, which is a very pragmatic engineering insight.

Meng: You are essentially teaching the model to recognize what *not* to flag in your specific context, which drastically reduces alert fatigue for the developers who have to triage these findings.

Lu: This concept of 'distribution shift' learning suggests that TitanCA is designed not just for high accuracy in ideal conditions, but for resilience in messy, real-world conditions.

Lalam: It’s a demonstration of how sophisticated systems need to incorporate human operational feedback into their core learning loop to remain useful over time.

Tom: So we understand the architecture and how it adapts; next up, we're going to talk about the final piece—how this

Paper discussion segment 3: Tom: We’ve seen how TitanCA: Lessons from Orchestrating LLM Agents to Discover one hundred plus CVE works in its core design, but now we need to talk about the practical improvements and implications this architecture suggests for wider adoption.

Jane: It's not just a cool concept, Tom, it’s a massive operational shift because the authors highlight that simply having many agents isn't enough; cost-aware pipeline ordering is just as important.

Meng: That makes a lot of sense from an engineering standpoint; if you put the super expensive multi-agent deliberation at the start of the process, your server costs would absolutely skyrocket.

Lu: The implication there is that we are moving toward a system that can actually scale, Meng, by prioritizing those cheap filtering steps first to handle the sheer volume of codebases.

Lalam: I think this cost-awareness speaks to how AI will integrate into the daily life of developers, Lu; it suggests tools that respect their time and their budget.

Tom: It’s about making the right choices in sequencing, Jane, not just having more agents running at once.

Jane: And another big improvement they discuss is how they treat the false positives—the F0 point three metric approach is really telling us something important about workflow.

Meng: That focus on precision is huge because it directly addresses developer burnout; if the tool gives a trustworthy warning, it’s a genuine security concern, not just noise.

Lu: It changes the cultural expectation of what 'automated testing' means; we are moving away from a hope for perfect recall toward certainty in practical deployment.

Lalam: When the system is designed to provide high-confidence findings instead of overwhelming lists, the developer’s relationship with that AI shifts into a partnership.

Tom: That partnership relies on the structured reasoning they built into R2Vul, which forces a check on how we use LLMs.

Jane: Right, it makes sure the explanation is logically sound, not just sounding plausible; it grounds the reasoning in actual code execution paths.

Meng: It’s a way of saying that if you can't explain *why* the AI thinks something is vulnerable, it isn't useful for practical remediation.

Lu: The creative potential here lies in using this explainability to build trust across a massive, distributed workforce that doesn's working on one monolithic codebase.

Lalam: It allows the organization to adopt a culture of proactive security, knowing their AI assistant is reliable enough to act as an early warning system.

Tom: We’ve seen how the structure improves reliability, but we also need to talk about how this system adapts and learns from its own errors.

Conclusion: Tom: : So, wrapping up our deep dive today, it’s clear that TitanCA represents a significant leap forward in how we approach automated vulnerability discovery.

Jane: : It moves beyond simple pattern matching and builds an entire, highly structured reasoning pipeline that prioritizes depth and accuracy over sheer speed.

Meng: : For me, the most impressive takeaway is the emphasis on adaptability; it’s not just a static tool but a system designed to learn within a specific organizational context.

Lu: : Exactly. The combination of multi-agent debate and domain adaptation means this technology can mature alongside the software development process itself.

Lalam: : It really underscores that solving complex security problems requires mimicking the best parts of human collaboration—multiple viewpoints leading to one strong consensus.

Tom: : And when you consider the measurable results they achieved, finding over one hundred CVEs, it lends incredible weight to the potential of this methodology.

Jane: : It’s a powerful reminder that LLMs, when properly orchestrated and constrained by rigorous processes, are immensely valuable tools for security engineers.

Meng: : I think the industry needs to pay close attention not just to the detection rate, but to the reliability metrics they introduced throughout their process.

Lu: : Definitely; that focus on precision over recall in certain stages is what makes this approach practically viable for real-world operations.

Lalam: : It’s a blueprint for building trust into automated systems, which is arguably harder than building the system itself.

Tom: : We’ve covered so much ground today regarding "TitanCA: Lessons from Orchestrating LLM Agents to Discover one hundred plus CVEs."

Jane: : I feel like we only scratched the surface of what this technology can achieve across different types of codebases.

Tom: : And while we wrap up our discussion on TitanCA, I'm really looking forward to diving into our next topic, where we’ll be looking at novel approaches in threat modeling...

N/A (Author list for primary paper not present in excerpt)

National Research Foundation, Singapore · Monash University · Singapore Management University · University of Groningen · IEEE Security & Privacy Magazine · Centre of Research on Intelligent Software Engineering (RISE), Singapore Management University · University of Computer Studies · Government Technology Agency of Singapore · Imperial College London · Tsinghua University · Southern University of Science and Technology (SUSTech) · The University of Sydney

cs.CR

Submitted: 2026-08-24

Updated: 2026-08-25

Project page: https://titancaproject.github.io/cves

Importance score: 92/100

The gist: TitanCA: Lessons from Orchestrating LLM Agents to Discover 100+ CVEs Introduction and Motivation Software vulnerabilities remain a persistent threat, but traditional Static Application Security

Key concepts

Structured Reasoning
The system uses structured reasoning to guide its modules, forcing the AI to write a logical justification for every claim about a bug. This turns the LLM's output into an auditable chain of thought that security engineers can follow, ensuring it moves beyond just stating that a bug exists.
Domain Adaptation (Adapter Module)
The Adapter module (PairVul) acts as a crucial bridge between generalized AI knowledge and specific company codebases. It handles unique coding styles or internal frameworks not found in public datasets, making the tool effective for proprietary enterprise code.
Continuous Learning
TitanCA incorporates a feedback loop to learn from its own mistakes in real-world environments. It uses false positives generated during initial testing as valuable training material, allowing the system to evolve and adapt to an organization's specific software development practices.
Cost-Aware Pipeline Ordering
This involves prioritizing cheaper filtering steps first in the sequence of agents. This prevents expensive multi-agent deliberation from skyrocketing server costs, allowing the system to scale and integrate practically into a developer workflow.

Terminology

Summary

TitanCA: Lessons from Orchestrating LLM Agents to Discover 100+ CVEs

Introduction and Motivation

Software vulnerabilities remain a persistent threat, but traditional Static Application Security Testing (SAST) tools suffer from high false-positive rates. This leads developers to ignore security warnings, leaving many vulnerabilities untriaged. While recent advances in large language models (LLMs) offer new avenues for automated code understanding—models trained on billions of lines of source code can reason about semantics and identify patterns rule-based analyzers miss—applying LLMs presents challenges such as hallucination, the need for high precision over recall, and a lack of awareness regarding specific deployment contexts.

TitanCA Solution and Scope

TitanCA addresses these issues by orchestrating multiple LLM-powered agents into a unified vulnerability discovery pipeline. Instead of relying on a single model, the system composes agents into a four-module pipeline where each refines the output of the previous stage, progressively filtering false positives and increasing confidence in findings.

This project is a collaboration between Singapore Management University (SMU) and GovTech Singapore. The work described corresponds to Phase 1 of the TitanCA project, which concluded in January 2026. As of March 2026, TitanCA analyzed code from over 127,000 GitHub repositories, identified 203 zero-day vulnerabilities (all subsequently remediated), and published 118 CVE identifiers as a direct result of the approach.

TitanCA Architecture: The Four Modules

The system implements a just-in-time (JIT) vulnerability-discovery approach that analyzes code as it is committed. The following modules are designed to complement one another, creating a layered defense that prioritizes precision while maintaining recall:

  1. Module 1: The Matcher (VulCoCo)

This module performs vulnerable clone detection. VulCoCo generates embeddings of candidate functions and compares them against a curated database of known vulnerable functions, drawing sources from the NVD and the TitanVul dataset, which contains approximately 40,000 vulnerability samples. It uses similarity search over a vector database to identify candidates and then employs an LLM as a semantic validator. This initial scan maximizes recall because the similarity matches are highly likely to be genuine vulnerabilities. If a candidate function is flagged as vulnerable, the remaining modules are skipped; otherwise, it flows to the next stage.

  1. Module 2: The Filter (R2Vul)

identifies true vulnerabilities among candidates that did not match known patterns but may contain novel weaknesses. R2Vul is trained using structured reasoning distillation and contrastive supervision with RLAIF (Reinforcement Learning from AI Feedback) to distinguish between grounded reasoning (where the logical chain matches the label) and misleading reasoning. The model is trained to produce an explicit chain of justification for its judgment. At inference time, a lightweight calibration step compares conditional log-likelihood ratios, converting the log-odds margin into a confidence score. This calibration reduces the false positive rate from 28% to 20% while preserving over 77% recall. If the function is classified as safe, prediction ends; otherwise, it flows to Module 3.

  1. Module 3: The Inspector (VulTrial)

This module introduces multi-agent deliberation using a mock-courtroom metaphor. It involves four specialized agents: a Security Researcher (who presents the case), a Code Author (who defends the code), a Moderator (who acts as the judge and distills the exchange into an impartial summary), and a Review Board (which functions as the jury to issue the final verdict). This adversarial structure forces consideration of both sides, helping surface subtle vulnerabilities that a single model might miss.

  1. Module 4: The Adapter (PairVul)

This module addresses domain-specific adaptation, crucial for deployment in new organizational contexts where coding conventions and vulnerability patterns may differ from the training data. PairVul analyzes false positives generated during initial deployment to identify recurring error patterns. It then finetunes the detection models using these newly labeled examples, creating a feedback loop that improves precision within a specific environment without requiring expensive manual annotation of new training data.

Data Foundation and Infrastructure

The TitanCA project invested heavily in data engineering, constructing a dataset of over 342,000 likely vulnerability samples. To address the label noise often found in public datasets, the team proposed CleanVul [5], which applies systematic cleaning procedures combining automated heuristics with LLM-assisted verification. The operational infrastructure is substantial: approximately 500 terabytes of data engineering workloads support the monitoring of over 127,000 GitHub repositories.

Supporting the Open Source Community and Results

The system's ability to detect major issues is demonstrated by its severity profile: 35% are rated critical severity, 95% are at least medium severity, and 91% exhibit low attack complexity. Approximately half of the discovered vulnerabilities fall within the CWE Top-25 Most Dangerous Software Weaknesses. The most frequently identified types include CWE-787 (out-of-bounds write), CWE125 (out-of bounds read), and CWE-190 (integer overflow or wraparound).

Key Lessons Learned

The project distilled several key architectural and operational insights:

  • Orchestration over Monolithic Models: The most important insight is that a pipeline of specialized, collaborating agents outperforms a single model. This division of labor mirrors how human security teams operate.

  • Cost-aware Pipeline Ordering: Restructuring the pipeline so that cheap filtering runs first and expensive deliberation runs later reduced our per-function cost substantially.

  • Precision is Critical: In production environments, false positives are operationally damaging. The architecture is designed to maximize the F0.3 metric (which weights precision three times more than recall), arguing that relying on the standard F1 score is misleading.

  • Structured Reasoning Improves Reliability: R2Vul trains LLMs to produce grounded reasoning chains, which significantly improves the reliability and interpretability of model outputs.

  • Multi-agent Debate Catches What Single Models Miss: The mock courtroom design of VulTrial addresses confirmation bias, forcing an adversarial structure that produces more robust judgments than any single model alone.

  • Adaptation is Essential for Deployment: PairVul provides a mechanism to handle domain shift, ensuring the system can be applied effectively in different target organizations.

Improvements for AI systems

The current state of AI code generation is insufficient for mission-critical or high-stakes applications because it lacks verifiable security guarantees and fails to reason about complex operational vulnerabilities. Based on the principles outlined in SecureAgentBench, I propose a shift from simple generative models to a robust, multi-agent, reflective system architecture.

Here are the specific improvements required for an AI system to achieve enterprise-grade secure code generation:


Improvement: The AI must transition from a monolithic Large Language Model (LLM) to a coordinating ensemble of specialized, interacting agents that mimic the roles of a development team: the Planner, the Generator, and the Critic/Verifier.

How it Works:

  1. Planner Agent (Security Requirements): This agent ingests not just functional requirements but also non-functional security policies (e.g., Must handle user input via parameterized queries, Must adhere to least-privilege principles). It generates a high-level, secure architectural design before writing any code.

  2. Generator Agent (Code Draft): This agent writes the initial code draft, constrained by the Planner's secure design blueprint.

  3. Critic/Verifier Agent (Security Review & Iteration): This is the most critical addition. It does not merely check syntax; it performs deep security analysis on the Generator's output, flagging potential vulnerabilities (e.g., race conditions, improper input sanitization) and forcing the Generator to iterate until remediation is achieved.

What the Improved AI System Can Do:

  • It can proactively identify and remediate design flaws (e.g., recommending using a message queue instead of direct database calls for inter-service communication) rather than just fixing coding errors.

  • It provides a verifiable audit trail, showing why the code was changed (e.g., The Critic flagged potential XSS; the Generator updated the input sanitization function accordingly).

Sources

Related papers