Open-Jev Judgments on CallScreenBench: Calibrated One-Pass Scam Screening with a Small Language Model

arXiv:2609.23959 · cs.CL · Submitted 2026-09-21 · Read on arXiv

cs.CL

Submitted: 2026-09-21

Updated: 2026-09-21

Comments: 13 pages, 6 figures, 5 tables

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

The gist: Screening a phone call for fraud needs a trustworthy probability after every caller turn, in milliseconds.

Terminology

Abstract

Screening a phone call for fraud needs a trustworthy probability after every caller turn, in milliseconds. Jev-style typed decisions promise exactly that: declared options go in, one calibrated probability per option comes out of a single forward pass, with no generated text. We test an open implementation of this readout, JevLite, on scam-call screening: Qwen3-4B is LoRA-tuned so that the temperature-scaled softmax over two answer-label logits is P(scam). On 41 held-out CallScreenBench scenarios (577 per-turn decisions) a three-seed ensemble reaches AUROC.974 with calibration error.052, non-inferior to an LLM judge (MiniMax-M3) at a pre-registered.02 margin, with no false alarms on legitimate calls, decisions 1.14 turns earlier under the same hang-up rule, and 64.5 ms per decision on one consumer GPU, 4.9x lower than the same backbone fine-tuned to generate its answer. The gain is in the readout and calibration, not accuracy: a fine-tuned ModernBERT encoder is not significantly worse, the recipe was selected with test-set exposure, and all callers are synthetic. We claim no architectural novelty; the contribution is the application and an evaluation reporting calibration, false alarms and decision timing alongside AUROC.

Sources

Related papers