APOLLO: A GPT-based tool to detect phishing emails and generate explanations that warn users
cs.HC, cs.CR
Submitted: 2024-10-10
Updated: 2024-10-10
Journal ref: Proceedings of the ACM on Human-Computer Interaction (2025), Volume 9, Issue 4
DOI: 10.1145/3733049
License: http://creativecommons.org/licenses/by/4.0/
The gist: Phishing is one of the most prolific cybercriminal activities, with attacks becoming increasingly sophisticated.
Terminology
Abstract
Phishing is one of the most prolific cybercriminal activities, with attacks becoming increasingly sophisticated. It is, therefore, imperative to explore novel technologies to improve user protection across both technical and human dimensions. Large Language Models (LLMs) offer significant promise for text processing in various domains, but their use for defense against phishing attacks still remains scarcely explored. In this paper, we present APOLLO, a tool based on OpenAI's GPT-4o to detect phishing emails and generate explanation messages to users about why a specific email is dangerous, thus improving their decision-making capabilities. We have evaluated the performance of APOLLO in classifying phishing emails; the results show that the LLM models have exemplary capabilities in classifying phishing emails (97 percent accuracy in the case of GPT-4o) and that this performance can be further improved by integrating data from third-party services, resulting in a near-perfect classification rate (99 percent accuracy). To assess the perception of the explanations generated by this tool, we also conducted a study with 20 participants, comparing four different explanations presented as phishing warnings. We compared the LLM-generated explanations to four baselines: a manually crafted warning, and warnings from Chrome, Firefox, and Edge browsers. The results show that not only the LLM-generated explanations were perceived as high quality, but also that they can be more understandable, interesting, and trustworthy than the baselines. These findings suggest that using LLMs as a defense against phishing is a very promising approach, with APOLLO representing a proof of concept in this research direction.
Sources
- On the Opportunities and Risks of Foundation Models
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- Emerging Phishing Trends and Effectiveness of the Anti-Phishing Landing Page
- Phishing, Personality Traits and Facebook
- Devising and Detecting Phishing: Large Language Models vs. Smaller Human Models
- Detecting Phishing Sites Using ChatGPT
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- AI Transparency in the Age of LLMs: A Human-Centered Research Roadmap
- Code Generation with AlphaCodium: From Prompt Engineering to Flow Engineering
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
- Attention Is All You Need
- Mixture-of-Agents Enhances Large Language Model Capabilities
- Prompt Engineering a Prompt Engineer
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support