Semi-Automated Detection of Gaps in LLM Security Knowledge
cs.CR, cs.AI, cs.HC
Submitted: 2026-07-20
Updated: 2026-09-22
Comments: v3: fixed typos in abstract metadata; no changes to the paper
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Large language models (LLMs) are increasingly used for a range of software, hardware and human-centered security tasks.
Terminology
Abstract
Large language models (LLMs) are increasingly used for a range of software, hardware and human-centered security tasks. Consequently, LLM performance on security tasks is an active area of measurement and research, often with a focus on identifying areas in which LLM security "knowledge" may be insufficient. Popular strategies for identifying LLM security knowledge gaps include building corpora of challenge questions or task benchmarks, strategies that require substantial manual work and security expertise to design and execute. We introduce a partially-automated method for assessing LLM knowledge of a security area. The method uses authoritative information from Consumer Protection Agencies (CPAs) to identify instability in LLM responses that can be indicative of knowledge gaps. We demonstrate the method for 2 security topics, identity theft and impostor scams, and 5 LLMs in 2 leading LLM families, Gemini and GPT, using publicly available information about identity theft and impostor scams from 6 CPAs. The method distinguishes between models that have and don't have sufficient knowledge to accurately identify the security topics in text narratives.
Sources
- Language Models (Mostly) Know What They Know
- Security and Privacy Prompts in the Wild: What Users Ask LLMs and How LLMs Respond
- HelpBench: Assessing the Ability of LLMs to Provide Privacy, Safety, and Security Advice
- Assessing the Software Security Comprehension of Large Language Models
- How Can ChatGPT Support Human Security Testers to Help Mitigate Supply Chain Attacks?
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs