SyzHarness: Patch-Based Kernel Bug Reproduction with LLM-Synthesized Fuzzing Harnesses
cs.CR, cs.AI
Submitted: 2026-09-20
Updated: 2026-10-01
Code: https://github.com/google/oss-fuzz-gen
Project page: https://google.github.io/security-research/kernelctf/rules
License: http://creativecommons.org/licenses/by/4.0/
The gist: Automated kernel vulnerability reproduction is essential for bug triage, patch validation, and regression testing, but still lacks an effective and efficient solution.
Terminology
Abstract
Automated kernel vulnerability reproduction is essential for bug triage, patch validation, and regression testing, but still lacks an effective and efficient solution. The core challenge is twofold: a reproducer must first recover the trigger scaffold needed to reach the vulnerable state and determine the precise concrete values that actually trigger the bug. Existing directed fuzzing approaches are ineffective at recovering the necessary trigger scaffold, while LLM- only generation is brittle because it struggles with concrete-value discovery and runtime nondeterminism. We design SyzHarness, a framework that combines LLM reasoning with coverage-guided fuzzing for patch-based Linux kernel vulnerability reproduction. Given a patch, SyzHarness uses an LLM agent grounded by code navigation tools to synthesize a parameterized fuzzing harness that fixes the prerequisite setup logic while exposing only uncertain, bug- critical input parameters to be mutated by Syzkaller. SyzHarness then translates this harness into a Syzkaller- compatible interface and iteratively refines it using hierarchical reachability feedback. We evaluate SyzHarness on multiple datasets of triggerable real-world Linux kernel vulnerabilities. On 100 KernelCTF cases, SyzHarness achieves a 78% bug reproduction success rate. On the SyzDirect benchmark, SyzHarness achieves a 73% bug reproduction success rate, substantially outperforming prior directed greybox fuzzing. On 50 recent, known-triggerable syzbot bugs fixed after March 2026, SyzHarness reproduces 40/50 (80%) using only the fix commits as input.
Sources
- LibLMFuzz: LLM-Augmented Fuzz Target Generation for Black-box Libraries
- GPTrace: Effective Crash Deduplication Using LLM Embeddings
- kAgent: An execution-guided crash resolution agent for the Linux kernel
- PILOT: Command-line Interface Fuzzing via Path-Guided, Iterative Large Language Model Prompting
- Cottontail: Large Language Model-Driven Concolic Execution for Highly Structured Test Input Generation
- From CVE Entries to Verifiable Exploits: An Automated Multi-Agent Framework for Reproducing CVEs
- Attention Distance: A Novel Metric for Directed Fuzzing with Large Language Models
- CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
- CKGFuzzer: LLM-Based Fuzz Driver Generation Enhanced By Code Knowledge Graph
- Directed Greybox Fuzzing via Large Language Model
- Understanding Gaps in LLM Pipelines Towards Scalable Fuzzing Harness Generation: An Empirical Study and Enhancement
- PBFuzz: Agentic Directed Fuzzing for PoV Generation
- LLAMAFUZZ: Large Language Model Enhanced Greybox Fuzzing
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs