CallScreenBench: Benchmarking Small Language Models as Phone Secretaries
Jiaqi Gan, Haoyuan Tang, Jamey Z. Liang, Siying Chen, Ankit Raj, Kidus Zewde, Yuchen Zhou, Yuxin Zhang, Simiao Ren
cs.CR, cs.AI
Submitted: 2026-08-21
Updated: 2026-08-25
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Language models small enough to run on a handset, quantized to a few bits, are increasingly capable of acting for their user, making on-device task automation newly plausible.
Terminology
Abstract
Language models small enough to run on a handset, quantized to a few bits, are increasingly capable of acting for their user, making on-device task automation newly plausible. One such task is answering the phone. A phone secretary takes an unknown inbound call on its owner's behalf. Unlike the agents evaluated by most benchmarks, it has no task to complete and no cooperative user: the caller holds the goal, may be an adversary, and must be judged from the opening turn with no oracle. What matters is not task success, but whether the owner would endorse how their proxy handled the call. We present CallScreenBench, which scores this setting on five quality dimensions. Each dimension is printed beside the counter-metric that bills it and is never averaged into a single number. We also report a guardedness profile for a toolless proxy that holds no credentials and calls no tools. Across six on-device models (0.6-4B parameters, 4-bit quantization), quality scales with capability, but triage does not. The appearance that it does is an artifact of measurement. Scripted degenerate agents supply the missing floors: after correcting for them, the number of model pairs whose triage performance separates falls from 11 of 15 to zero at the preregistered operating point. An agent that simply hangs up and echoes the caller also scores perfect message fidelity. We report which of our own metrics these floors defeat and declare no pass/fail threshold.
Sources
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics
- RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- VoiceBench: Benchmarking LLM-Based Voice Assistants
- Can Large Language Models Be an Alternative to Human Evaluations?
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Moshi: a speech-text foundation model for real-time dialogue
- Towards Human-Like Interactive Speech Recognition With Agentic Correction and Semantic Evaluation
- MULTI-Bench: A Multi-Turn Interactive Benchmark for Assessing Emotional Intelligence ability of Spoken Dialogue Models
- Semantic Distance: A New Metric for ASR Performance Analysis Towards Spoken Language Understanding
- LLMs Get Lost In Multi-Turn Conversation
- ACUTE-EVAL: Improved Dialogue Evaluation with Optimized Questions and Multi-turn Comparisons
- Voice Activity Projection: Self-supervised Learning of Turn-taking Events
- Holistic Evaluation of Language Models
- Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency
- Gemma 3 Technical Report
- Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
- The Llama 3 Herd of Models
- VocalBench-DF: A Benchmark for Evaluating Speech LLM Robustness to Disfluency
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs