Open-Jev Judgments on CallScreenBench: Calibrated One-Pass Scam Screening with a Small Language Model
cs.CL
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: 13 pages, 6 figures, 5 tables
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Screening a phone call for fraud needs a trustworthy probability after every caller turn, in milliseconds.
Terminology
Abstract
Screening a phone call for fraud needs a trustworthy probability after every caller turn, in milliseconds. Jev-style typed decisions promise exactly that: declared options go in, one calibrated probability per option comes out of a single forward pass, with no generated text. We test an open implementation of this readout, JevLite, on scam-call screening: Qwen3-4B is LoRA-tuned so that the temperature-scaled softmax over two answer-label logits is P(scam). On 41 held-out CallScreenBench scenarios (577 per-turn decisions) a three-seed ensemble reaches AUROC.974 with calibration error.052, non-inferior to an LLM judge (MiniMax-M3) at a pre-registered.02 margin, with no false alarms on legitimate calls, decisions 1.14 turns earlier under the same hang-up rule, and 64.5 ms per decision on one consumer GPU, 4.9x lower than the same backbone fine-tuned to generate its answer. The gain is in the readout and calibration, not accuracy: a fine-tuned ModernBERT encoder is not significantly worse, the recipe was selected with test-set exposure, and all callers are synthetic. We claim no architectural novelty; the contribution is the application and an evaluation reporting calibration, false alarms and decision timing alongside AUROC.
Sources
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Voice-Enabled AI Agents can Perform Common Scams
- Language Models (Mostly) Know What They Know
- CallScreenBench: Benchmarking Small Language Models as Phone Secretaries
- Incremental Risk Assessment of Progressive Elder Financial Scams via Instruction-Tuned Small Language Models
- "It Warned Me Just at the Right Moment": Exploring LLM-based Real-time Detection of Phone Scams
- Anatomy of a Scam Call: What 10,000 real scam and spam calls reveal about how phone scammers operate
- The Machines Are Calling: Measuring Automated and Synthetic Voices in Unwanted Inbound Calls
- Combating Phone Scams with LLM-based Detection: Where Do We Stand?
- Qwen3 Technical Report
- ShieldGemma: Generative AI Content Moderation Based on Gemma
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering