When to Call an LLM: A Confidence-Gated Hybrid for Cost-Effective Emotion Recognition in Conversational AI
cs.AI
Submitted: 2026-09-16
Updated: 2026-09-16
License: http://creativecommons.org/licenses/by/4.0/
The gist: Emotion recognition in conversation (ERC) is a production capability behind agent-assist prompts, escalation routing, and post-call analytics in contact-center-as-a-service (CCaaS) platforms, where
Terminology
Abstract
Emotion recognition in conversation (ERC) is a production capability behind agent-assist prompts, escalation routing, and post-call analytics in contact-center-as-a-service (CCaaS) platforms, where cost and latency constraints matter as much as accuracy. We report a systems-level comparison of three deployment options for dialogue-contextual ERC: a low-cost stacked ensemble (sentence embeddings, windowed context, RandomForest/XGBoost/logistic-regression stacking), off-the-shelf LLM prompting (GPT-4o-mini; zero-shot, few-shot, chain-of-thought), and a confidence-gated hybrid that escalates only the ensemble's least-confident predictions to the LLM - modeled on IVA-to-human-agent escalation policies used in production contact centers. On IEMOCAP, the ensemble significantly outperforms every LLM configuration (0.595 vs. 0.460-0.536 weighted F1, p < 0.0001) at a fraction of the cost and sub-10ms latency; on MELD and CMU-MOSI the ranking reverses, showing neither pure system is a safe default. The confidence-gated hybrid resolves this by Pareto-dominating both pure systems on all three datasets (0.620, 0.643, 0.824 weighted F1) while routing the majority of traffic through the near-zero-cost ensemble, translating to roughly 10-85 per million utterances versus 99-170 for an LLM-only pipeline. The escalation policy is not an opaque cost/accuracy dial: escalated turns disproportionately follow an emotion or sentiment shift, giving operators an interpretable, auditable routing signal, and the ensemble's confidence is well-calibrated and safely under- rather than over-confident. The pattern holds across three datasets and two LLM providers. Confidence-gated cascading is established in general ML systems; our contribution is showing it transfers cleanly to dialogue-contextual ERC, yielding a concrete deployment recipe for CCaaS and conversational-AI platforms deciding how to allocate LLM spend.
Sources
- FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
- AutoMix: Automatically Mixing Language Models
- Measuring the Robustness of NLP Models to Domain Shifts
- Learning to Route LLMs with Confidence Tokens
- A Unified Approach to Routing and Cascading for LLMs
- A Mixture-of-Experts Model for Multimodal Emotion Recognition in Conversations
- Cascaded Language Models for Cost-effective Human-AI Decision-Making
- Calibrated Selective Classification
- Tryage: Real-time, intelligent Routing of User Prompts to Large Language Models
- Machine Learning with a Reject Option: A survey
- EmoBERTa: Speaker-Aware Emotion Recognition in Conversation with RoBERTa
- LLMs Get Lost In Multi-Turn Conversation
- Enhancing Granular Sentiment Classification with Chain-of-Thought Prompting in Large Language Models
- OmniVox: Zero-Shot Emotion Recognition with Omni-LLMs
- RouteLLM: Learning to Route LLMs with Preference Data
- Gatekeeper: Improving Model Cascades Through Confidence Tuning
- Mixture-of-Agents Enhances Large Language Model Capabilities
- MOSI: Multimodal Corpus of Sentiment Intensity and Subjectivity Analysis in Online Opinion Videos
- EcoAssistant: Using LLM Assistant More Affordably and Accurately
- Reassessing the Role of Chain-of-Thought in Sentiment Analysis: Insights and Limitations
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection