Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows
Keyur Gabani
cs.CR, cs.AI, cs.LG
Submitted: 2026-07-12
Comments: 22 pages, 3 figures, 6 tables. Ancillary files include the evidence matrix, search note, and numeric claim check
License: http://creativecommons.org/licenses/by/4.0/
The gist: LLMs are now proposed for fraud detection, scam investigation, content moderation, and other trust-and-safety workflows.
Terminology
Abstract
LLMs are now proposed for fraud detection, scam investigation, content moderation, and other trust-and-safety workflows. Much of the public literature still evaluates them as models, with less attention to their behavior as components in operational pipelines. This creates a practical evidence question: what would justify placing an LLM inside a live workflow with latency, cost, escalation, human-review, and adversarial-risk constraints? We address this question through a fraud-first survey of deployment evidence. We code 49 operationally relevant sources on LLM use in fraud detection, investigation support, content moderation, and cross-cutting robustness (18 fraud, 14 moderation, 17 cross-cutting), supplemented by 15 contextual references that establish the survey boundaries. These sources include systems, benchmarks, frameworks, and deployment-relevant surveys, not 49 production deployments. The main finding is an evidence imbalance. Fraud supplies the largest task-specific portion of the coded corpus. The moderation papers, however, include more explicit public evidence on latency, cost, governance, and fairness. Among the 18 fraud and investigation sources, none report clean per-decision latency, per-decision dollar cost, or calibration evidence; most report offline task performance, retrieval gains, or case-study accuracy instead. The survey contributes a role-and-evidence organizing frame, FORTE, for locating LLMs as classifiers, retrieval interfaces, explanation generators, reviewer assistants, agents, feature extractors, or escalation components. It also contributes a minimum deployment-evidence checklist covering latency budget, cost per decision, decision threshold, explanation integrity, and adversarial pressure. The resulting agenda identifies studies needed to support deployment claims for LLM-based fraud and trust-and-safety work.
Sources
- Understanding Structured Financial Data with LLMs: A Case Study on Fraud Detection
- LLM-Enhanced Self-Evolving Reinforcement Learning for Multi-Step E-Commerce Payment Fraud Risk Detection
- Advanced Real-Time Fraud Detection Using RAG-Based LLMs
- LLM-Assisted Authentication and Fraud Detection
- AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts
- SLM-Mod: Small Language Models Surpass LLMs at Content Moderation
- CASE: An Agentic AI Framework for Enhancing Scam Intelligence in Digital Payments
- EXPLICATE: Enhancing Phishing Detection through Explainable AI and LLM-Powered Interpretability
- Co-Investigator AI: The Rise of Agentic AI for Smarter, Trustworthy AML Compliance Narratives
- FAA Framework: A Large Language Model-Based Approach for Credit Card Fraud Investigations
- Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
- PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models
- Selective Conformal Risk Control
- Enhancing the Interpretability of SHAP Values Using Large Language Models
- Safeguarding Large Language Models: A Survey
- Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety
- A Survey on Large Language Model (LLM) Security and Privacy: The Good, the Bad, and the Ugly
- Year-over-Year Developments in Financial Fraud Detection via Deep Learning: A Systematic Literature Review
- FlexGuard: Continuous Risk Scoring for Strictness-Adaptive LLM Content Moderation
- Measuring What LLMs Think They Do: SHAP Faithfulness and Deployability on Financial Tabular Classification
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs