PiMRef: Deducing Ever-evolving Spear-phishing Emails with Knowledge Base Invariants
cs.CR, cs.AI
Submitted: 2025-07-21
Updated: 2026-09-07
Code: https://github.com/code-philia/PhishEmail
Project page: https://sites.google.com/view/pimref/home
License: http://creativecommons.org/licenses/by/4.0/
The gist: Phishing email is a critical step in the cybercrime kill chain due to the high reachability of victims' email accounts and the low cost of launching phishing campaigns.
Terminology
Abstract
Phishing email is a critical step in the cybercrime kill chain due to the high reachability of victims' email accounts and the low cost of launching phishing campaigns. This ever-evolving nature of phishing emails makes traditional rule-based and feature-engineering-based phishing email detectors fight an uphill battle in the cat-and-mouse game of defense and attack. In this work, we show that, large language models (LLMs) can be effectively exploited to generate profile-grounded spear-phishing, compromising major paradigms of phishing email detectors. To defend against such LLM-based spear-phishing attacks, we propose PiMRef, the first reference-based solution to detect ever-evolving phishing emails using knowledge-based invariants, targeting the identity-impersonation attacks that characterize spear-phishing. Our rationale lies in the fact that convincing phishing emails often include ``disprovable claims'', which contradict certain real-world facts. Technically, given an email, PiMRef (i) discovers the claimed identity of the sender, (ii) verifies the sender's email domain against a dynamically expandable knowledge base, and (iii) infers call-to-action instructions that encourage next-step engagement. Compared to existing baselines such as D-Fence, HelpHed, and ChatSpamDetector, PimRef reduces the false-positive rate to 0.81% while maintaining a recall of 90.7%-93.1% on conventional phishing benchmarks such as Nazario and PhishPot. On SpearMail, our newly constructed benchmark of 14,672 LLM-generated spear-phishing emails targeting 681 public profiles, PimRef reaches a recall of 86.4% without incurring additional false positives. Our code is publicly available at https://github.com/code-philia/PhishEmail.
Sources
- That Ain't You: Detecting Spearphishing Emails Before They Are Sent
- Characterizing Robocalls with Multiple Vantage Points
- CATBERT: Context-Aware Tiny BERT for Detecting Social Engineering Emails
- A Modular and Adaptive System for Business Email Compromise Detection
- ChatSpamDetector: Leveraging Large Language Models for Effective Phishing Email Detection
- An Empirical Analysis of SMS Scam Detection Systems
- ScamChatBot: An End-to-End Analysis of Fake Account Recovery on Social Media via Chatbots
- An Explorative Study of Pig Butchering Scams
- Social Turing Tests: Crowdsourcing Sybil Detection
- SENet: Visual Detection of Online Social Engineering Attack Campaigns
- Personalized Security Indicators to Detect Application Phishing Attacks in Mobile Platforms
- Prompted Contextual Vectors for Spear-Phishing Detection
- URL Inspection Tasks: Helping Users Detect Phishing Links in Emails
- When LLMs Go Online: The Emerging Threat of Web-Enabled LLMs
- Lateral Phishing With Large Language Models: A Large Organization Comparative Study
- Teach LLMs to Phish: Stealing Private Information from Language Models
- KnowPhish: Large Language Models Meet Multimodal Knowledge Graphs for Enhancing Reference-Based Phishing Detection
- VisualPhishNet: Zero-Day Phishing Website Detection by Visual Similarity
- Instruction Tuning for Large Language Models: A Survey
- Mistral 7B
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs