Towards LLM-Enhanced Android Taint Analysis
cs.SE, cs.CR
Submitted: 2026-08-25
Updated: 2026-08-25
Code: https://github.com/google-gemini/gemini-cli
License: http://creativecommons.org/licenses/by/4.0/
The gist: Taint analysis is a fundamental technique for detecting sensitive data leaks in Android apps.
Terminology
Abstract
Taint analysis is a fundamental technique for detecting sensitive data leaks in Android apps. However, traditional static tools, such as FlowDroid, still face well-known challenges due to the complexity of accurately modeling the Android framework. In this paper, we investigate whether off-the-shelf Large Language Models (LLMs) can effectively reason about taint flows in Android apps. Our preliminary approach relies on an agentic interaction strategy, enabling the LLM to iteratively explore code and reason about data flows. We conduct an initial evaluation on the DroidBench benchmark against FlowDroid, where our approach outperforms the baseline: Gemini-3 Flash achieves an F1-score of 0.96, compared to 0.55 for FlowDroid. In particular, we observe improvements in challenging categories such as inter-component communication (0.95 vs. 0.17), implicit flows (0.94 vs. 0.00), and reflection (1.00 vs. 0.50), where FlowDroid typically struggles. On a small set of real-world apps, the LLM-based approach also identifies additional potential data leaks not reported by FlowDroid. These preliminary findings suggest that LLM reasoning may effectively complement traditional static taint analysis, motivating future research on hybrid LLM-enhanced taint analysis pipelines.
Sources
- Multi-Agent Taint Specification Extraction for Vulnerability Detection
- Understanding the Effectiveness of Large Language Models in Detecting Security Vulnerabilities
- IRIS: LLM-Assisted Static Analysis for Detecting Security Vulnerabilities
- Exploring Code Analysis: Zero-Shot Insights on Syntax and Semantics with LLMs
- LLbezpeky: Leveraging Large Language Models for Vulnerability Detection
- LAMD: Context-driven Android Malware Detection and Classification with LLMs
- RepairAgent: An Autonomous, LLM-Based Agent for Program Repair
- Automatic Code Summarization via ChatGPT: How Far Are We?
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties