LLM-Assisted Dynamic Threat Analysis for Attacker-Reachable Software Weaknesses in Autonomous Vehicles
Md Wasiul Haque, Sagar Dasgupta, Mizanur Rahman, Md Rayhanur Rahman
The University of Alabama
cs.SE, cs.CR, cs.LG
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: 17 pages, 8 figures, 8 tables
Code: https://github.com/google/oss-fuzz-gen
Project page: https://autowarefoundation.github.io/autoware-documentation/main/home
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: LLM-Assisted Dynamic Threat Analysis for Attacker-Reachable Software Weaknesses in Autonomous Vehicles Objectives: As autonomous vehicles reach public roads, their software becomes safety-critical.
Terminology
Summary
LLM-Assisted Dynamic Threat Analysis for Attacker-Reachable Software Weaknesses in Autonomous Vehicles
Objectives: As autonomous vehicles reach public roads, their software becomes safety-critical. A defect reachable from an attacker input can change how a vehicle steers or brakes. Static analysis flags many candidate sites, but confirming one is reachable and exploitable needs executable artifacts whose manual construction is the bottleneck. The paper asks whether large language models (LLMs) can automate this for the open-source stack Autoware.
Methods: The authors perform a compiler-precise static analysis of Autoware (185 packages), recovering 1,375 decision rules, 2,274 validation checks, and 482 input-to-safety-output flows, from which they derive a weakness taxonomy and a sample of 740 reachable sites. For each site, 2 local open-weight LLMs, a no-static-context ablation, and a naive-template baseline generate artifacts. All 3,700 sets are compiled against the real build under sanitizers, repaired via a compiler in-the-loop stage, and fuzzed when they compile.
Findings: The principal result is a build-integration failure taxonomy that locates the binding constraint one stage before the fuzzer: 80% of first-shot compile failures arise from dependency-wiring rather than program logic. The reasoning model compiled 64% of its harnesses first, against 6% for the code-specialized model, and repair reached full object-compileability for the reasoning model only by stubbing the real target; consequently, under half of its harnesses reached the fuzzer and all 37 crashes came from a stub. Within budget, no candidate weakness was dynamically confirmed, reflecting this build-integration barrier rather than a benign attack surface.
Novelty: The first end-to-end feasibility study of LLM-assisted dynamic analysis on a full AV stack, pairing repository-scale static candidate identification with LLM-generated confirmation and pinpointing where the automation fails.
Practical Applications: Scalable software-safety assurance is a prerequisite for automated driving. Unaided LLM-assisted dynamic analysis is not yet a trustworthy assurance stage; effort is better spent on build integration than on prompt design, while the static analysis already guides safety review.
Detailed Results:
-
Static Analysis: The pipeline analyzes 185 packages and 1,857 source files. It recovers 786 ROS interfaces (248 subscriptions, 500 publishers, 38 services), 1,816 parameter surfaces, and 2,749 compiler-precise condition sites. Safety-relevant decision logic is concentrated in planning (1,161) and control (203). Validation is uneven: 1,789 range/threshold guards, 391 container checks, 43 null checks, 22 timeout/staleness checks, and 6 numeric-validity checks. The heuristic analysis recovers 482 package-level paths.
-
Target Selection: From the 2,749 condition sites, a safety-relevance score (0–11) is assigned. Sites scoring ≥8 are P1, 5–7 are P2. All P1 and P2 sites are retained, yielding 740 targets: 214 in P1 and 526 in P2, spanning 34 packages and 107 source files.
-
Build Integration: No weakness was confirmed across the 5 conditions and 740 targets. Of the 2,960 LLM-generated harnesses, 2,259 failed to compile initially. 1,436 omitted required ROS or Autoware headers and 381 included target.cpp files through invalid paths, together accounting for 1,817 of 2,259 failures. The remainder comprised API mismatches (225), signature mismatches (110), syntax errors (81), and 26 other errors.
-
Model Choice and Repair: The reasoning model (gpt-oss:20b) compiled 473 of 740 harnesses with static context (63.9%), whereas the code-specialized model (codestral:22b) compiled only 46 (6%) and 3 (under 1%) across its two conditions. Removing static context reduced the reasoning model's rate from 63.9% to 24.2%. Repair raised the reasoning model's object compileability to 100% in both conditions after a mean of 1.2 rounds with context and 1.3 without. However, only 289 of 740 context-enabled harnesses and 315 of 740 no-context harnesses linked and reached the fuzzer. The no-context condition linked more harnesses because it produced simpler self-contained stubs.
-
Confirmation Analysis: After repair, 615 LLM harnesses linked and fuzzed for the full budget without a crash: 24 and 10 from the code-specialized model (with and without static context), and 271 and 310 from the reasoning model (with and without context). These runs provide only weak disconfirmation because most surviving harnesses exercised local stubs rather than the intended Autoware implementation. The remaining 2,308 condition–target pairs never reached the fuzzer. All 37 crashes occurred in generated stub code rather than in Autoware.
-
Case Studies: Four cases illustrate the results. First, a P1 control-gate harness compiled and fuzzed cleanly only after replacing real interfaces with no-op stubs. Second, the code-specialized model rarely produced compilable harnesses due to invalid target paths, missing headers, and leaked Markdown syntax. Third, static context increased the reasoning model's first-shot compileability from 24.2% to 63.9%. Fourth, a velocity-smoother crash occurred inside a model-reimplemented helper with no Autoware frame on the stack.
Conclusions: The results support a qualified negative. The static analysis identifies a broad safety-relevant attack surface, but the dynamic stage shows that the main obstacle is not fuzzing, but faithful integration with the real build. Compiler-in-the-loop repair raises object compileability to 100% for the stronger model, but largely through stub convergence. As a result, only 652 of 2,960 harnesses reach the fuzzer, and all 37 crashes occur in generated stub code rather than in Autoware. The repair loop fails because it optimizes for satisfying the compiler rather than preserving the connection to the real target. Future automation should focus on dependency resolution, native linking, and verified execution of the intended target rather than on prompt refinement or compile-only repair.
Improvements for AI systems
Improvements to AI Systems:
-
Build-Integration-Aware Code Generation: Enhance LLMs to generate code that directly resolves ROS/Autoware dependency wiring (headers, package paths, API signatures) by training on repository-level build graphs and compile-error patterns, not just code syntax. The improved system can produce harnesses that compile against the real target on the first attempt, reducing the 80% dependency-wiring failure rate.
-
Compiler-in-the-Loop Repair with Target-Fidelity Objective: Replace the current repair loop (which optimizes for compiler satisfaction) with a dual-objective repair that maximizes both compileability and semantic linkage to the intended target (e.g., by checking that the generated harness calls the real Autoware functions, not stubs). The improved system can repair harnesses while preserving the connection to the actual implementation, preventing stub convergence and enabling genuine dynamic confirmation.
-
Static-Context-Aware Prompting for Safety-Critical Code: Use the recovered static analysis (1,375 decision rules, 2,274 validation checks, 482 flows) as structured context in prompts, but also include build-system metadata (CMakeLists, package.xml, include paths). The improved system can generate harnesses with 64% first-shot compileability (vs. 6% without context) and maintain that advantage through repair, leading to more harnesses reaching the fuzzer.
-
Dependency-Resolution Module: Integrate a pre-processing step that maps each target site to its required headers, libraries, and ROS message types, then injects these into the LLM prompt or post-processes the output. The improved system can eliminate the 1,436 missing-header and 381 invalid-path errors, increasing the number of fuzzable harnesses from 652 to potentially over 2,000.
-
Stub-Detection and Fidelity Scoring: Add a runtime or static check that flags when a generated harness substitutes a stub for a real Autoware function (e.g., by comparing symbol tables or call graphs). The improved system can reject or repair such harnesses, ensuring that fuzzing exercises the actual target code, thereby providing meaningful disconfirmation or confirmation of weaknesses.
-
Targeted Fuzzing with Build-Integration Verification: After linking, verify that the harness's execution path includes the intended Autoware source file (e.g., via coverage-guided checks or stack-trace validation). The improved system can distinguish crashes in real code from crashes in stubs, enabling accurate vulnerability confirmation and avoiding false positives like the 37 stub-only crashes.
-
Adaptive Model Selection Based on Task Phase: Use the reasoning model (gpt-oss:20b) for initial harness generation and repair (due to its 64% compileability), but switch to a code-specialized model for dependency-wiring subtasks (e.g., generating correct include paths and API calls) where the reasoning model fails. The improved system can combine strengths, achieving higher first-shot compileability and lower repair rounds.
-
Repository-Scale Static-to-Dynamic Feedback Loop: Feed the dynamic results (e.g., which sites fail to compile, which stubs are generated) back into the static analysis to refine the weakness taxonomy and target selection. The improved system can prioritize sites that are more likely to be dynamically confirmable, focusing effort on feasible targets and improving overall assurance efficiency.
What the Improved AI System Can Do:
-
Generate build-integrated test harnesses for 740 safety-critical sites in Autoware with >80% first-shot compileability (vs. 6–64% currently).
-
Repair harnesses while preserving real-target linkage, ensuring that fuzzing exercises actual Autoware code, not stubs.
-
Dynamically confirm or disconfirm attacker-reachable weaknesses in planning and control modules within a 24-hour budget, scaling to full AV stacks.
-
Provide a trustworthy assurance stage for autonomous vehicle software, reducing manual effort in vulnerability confirmation by orders of magnitude.
Abstract
Autonomous vehicles depend on large safety-critical software stacks, where weaknesses reachable from adversarial inputs may affect steering, braking, or other control decisions. Static analysis can identify candidate sites, but dynamically confirming exploitability requires executable test artifacts that are difficult to construct manually. We investigate whether large language models (LLMs) can automate this process for Autoware, an open-source autonomous-driving stack. We perform compiler-precise static analysis across 185 packages, identifying 1,375 decision rules, 2,274 validation checks, and 482 input-to-safety-output flows, from which we derive a weakness taxonomy and sample 740 reachable sites. Two local open-weight LLMs, a no-static-context ablation, and a naive-template baseline generate 3,700 artifact sets, which are compiled against the real build under sanitizers, repaired through compiler-in-the-loop feedback, and fuzzed when executable. The main result is a build-integration failure taxonomy showing that 80% of first-shot compilation failures arise from dependency wiring rather than program logic. The reasoning model compiled 64% of harnesses on the first attempt, compared with 6% for the code-specialized model. Repair achieved full object-compileability for the reasoning model only through extensive stubbing; fewer than half of its harnesses reached the fuzzer, and all 37 observed crashes originated in stubbed code rather than Autoware. No candidate weakness was dynamically confirmed within budget. These results show that build integration, not candidate generation or fuzzing, is the primary barrier to reliable LLM-assisted dynamic analysis of full autonomous-vehicle software stacks.
Sources
- Security Vulnerabilities in Software Supply Chain for Autonomous Vehicles
- LLM Agents can Autonomously Exploit One-day Vulnerabilities
- Waymo's Safety Methodologies and Safety Readiness Determinations
- On a Formal Model of Safe and Scalable Self-driving Cars
- Open-Source Autonomous Driving Software Platforms: Comparison of Autoware and Apollo
- Towards an open standard for assessing the severity of robot security vulnerabilities, the Robot Vulnerability Scoring System (RVSS)
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties