Measuring and Evaluating the Performance of Generative AI Models for Scam Detection
Cem Topcuoglu, Seyed Ali Akhavani, Harel Berger, Sadia Afroz, Michalis Pachilakis, Vibhor Sehgal, Leyla Bilge, Engin Kirda
cs.CR
Submitted: 2026-07-19
Code: https://github.com/cemtopcuoglu/genai
License: http://creativecommons.org/licenses/by/4.0/
The gist: Online scams continue to cause substantial financial and personal harm.
Terminology
Abstract
Online scams continue to cause substantial financial and personal harm. As a result, detection systems based on Large Language Models (LLMs) have been integrated into security products ranging from email gateways and browser extensions to fraud-monitoring dashboards. As this adoption accelerates, a common belief has taken hold: that these models are broadly suitable for scam detection. In this work, we investigate whether LLMs, with their strong capabilities in understanding intent, context, and reasoning, can effectively detect scams across diverse scenarios without task-specific fine-tuning. We curate and release a unique benchmark dataset of real-world scams spanning multiple formats and topics. We evaluate nine LLMs of varying sizes and architectures, examining their performance under different prompting strategies and comparing them to a fine-tuned BERT-based classifier. Our results show that while larger LLMs generally outperform smaller ones, effective prompting substantially boosts the performance of smaller models. Moreover, LLMs are better at generalizing to unseen scams compared to fine-tuned models, suggesting that pre-trained knowledge contributes meaningfully to scam detection. We release our dataset and evaluation framework to facilitate future research in robust scam detection using language models.
Sources
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- Feature Engineering vs BERT on Twitter Data
- Comparing BERT against traditional machine learning text classification
- Incorporating BERT into Parallel Sequence Decoding with Adapters
- CMRxRecon: An open cardiac MRI dataset for the competition of accelerated image reconstruction
- From ML to LLM: Evaluating the Robustness of Phishing Webpage Detection Models against Adversarial Attacks
- Spam-T5: Benchmarking Large Language Models for Few-Shot Email Spam Detection
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity
- Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation
- LLaMA: Open and Efficient Foundation Language Models
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs