EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models
Junyu Wang, Siyuan Zhang, Peiyuan Jiang, Jian Zong, Jingyu Zhang, Tianrui Wang, Yuqin Lin, Zhenghui Chen, Shuqing Xie, Ziyang Ma, Meng Ge, Xiaobao Wang, Longbiao Wang, Jianwu Dang
Tianjin University · Fuzhou University · Shanghai Jiaotong University · Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences
cs.CL
Submitted: 2026-08-10
Updated: 2026-08-11
Comments: Accepted at ACM Multimedia 2026 (MM '26)
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models Abstract Despite significant advances in instruction-following and auditory
Terminology
Summary
EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models
Abstract
Despite significant advances in instruction-following and auditory comprehension, the evaluation of Emotional Intelligence (EI) in Spoken Language Models (SLMs) remains confined to rudimentary paralinguistic perception, lacking a systematic, theory-driven cognitive framework. We introduce EmoSBench, the first comprehensive EI evaluation benchmark for SLMs constructed upon the four-branch theoretical model, covering Perceiving, Understanding, Using, and Managing Emotion across ten sub-tasks. Preliminary assessments on EmoSBench reveal a substantial gap: even leading proprietary models like GPT-4o-Audio achieve only 52.6%, significantly trailing human baselines. To bridge this gap, we develop EmoS, a specialized evaluator model optimized via Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO). To facilitate its effective training, we curate EmoDialogue, a bilingual dataset providing necessary fine-grained supervision through response pairs with rigorously defined EI gradations. Concurrently, we introduce a reward mechanism integrating a Steep Exponential Accuracy Reward (SEAR) and a Rationale Fidelity Reward (RFR) to enforce precise ordinal scoring and valid reasoning. Experiments demonstrate that EmoS reaches 83.8% accuracy, approaching human-level performance. Furthermore, evaluations on authentic, unconstrained spoken interactions validate its robust real-world generalization, establishing a foundational framework for advancing emotionally intelligent dialogue systems.
1. Introduction
Current Spoken Language Models (SLMs) have evolved into the core engines of modern conversational systems, achieving sophisticated auditory comprehension, context awareness, and instruction-following. Early evaluation frameworks primarily assessed the IQ of SLMs by focusing on acoustic perception and textual reasoning. Although subsequent dialogue benchmarks successfully introduced tasks requiring contextual interaction, their core focus remains confined to semantic understanding and logical deduction, typically treating the speech modality merely as an auxiliary input for text processing.
However, human spoken dialogue is intrinsically rich in paralinguistic cues, such as emotion, prosody, speed, volume, and rhythm, which significantly modulate semantic interpretation. For instance, the phrase I am aware of this schedule
conveys distinctly different meanings when articulated with a tone of disappointment versus relief. Recognizing this, recent works like WavReward and ParaS2S have begun to address this by incorporating the perception of paralinguistic features into preliminary assessments of model Emotional Intelligence (EI).
Nevertheless, genuine EI, as rigorously defined in psychological theory, transcends mere acoustic perception. Fundamentally, EI represents a complex cognitive ability to reason about emotions and utilize emotional information to enhance thought and problem-solving. This conceptualization is formalized in the influential four-branch model proposed by Mayer and Salovey, which characterizes EI as a hierarchical progression—starting from the basic sensory recognition of affective signals and ascending to the higher-order regulation of emotions to achieve social or personal goals. Drawing upon this theoretical foundation, our work posits that EI in SLMs comprises four hierarchical skills: Perceiving (decoding emotional and non-verbal cues), Understanding (comprehending emotional transitions across turns), Using (harnessing emotions for decision-making), and Managing (regulating emotions to achieve goals). Based on this, we define ten sub-task scenarios to systematically simulate the multifaceted challenges of high-EI dialogue in the real world.
Preliminary evaluations reveal a significant performance disparity: even the most advanced contemporary models, including leading open-source SLMs such as Qwen3-Omni and Kimi-Audio, and closed-source counterparts such as GPT-4o-Audio and Gemini 2.5 Pro, perform considerably below the human baseline on this comprehensive EI task. This finding highlights a critical methodological limitation: the prevailing practice of relying on proprietary models (e.g., GPT-4o-Audio) as evaluators for generating high-EI dialogue inadvertently sets an artificially low performance ceiling. This approach masks the intrinsic deficiencies of current systems and underscores the urgent need for specialized EI evaluation models. However, the development of such specialized evaluators is currently obstructed by a fundamental resource gap, as existing datasets primarily focus on coarse-grained emotion classification and paralinguistic perception, thereby lacking the fine-grained, theoretically aligned annotations necessary to train models with valid EI reasoning capabilities.
To bridge this data scarcity and enable the construction of reliable evaluators, we introduce EmoDialogue, a large-scale dataset comprising both single-turn and multi-turn samples. Each instance consists of user speech paired with four candidate responses, rigorously graded from 1 to 4 based on their EI quality. Leveraging this, we develop EmoS, a specialized reward model for spoken EI evaluation, employing a sequential two-stage training paradigm. Specifically, EmoS is first initialized through Supervised Fine-Tuning (SFT) to establish a robust foundational distribution of EI. Subsequently, by integrating Group Relative Policy Optimization (GRPO) with multi-reward reinforcement learning (RL) techniques, we effectively align EmoS's predictive scoring and analytical rationales with our theoretical guidelines.
The contributions are threefold:
-
We propose EmoSBench, a novel, theoretically grounded benchmark that evaluates EI across four key dimensions (Perceiving, Understanding, Using, Managing) through ten sub-tasks, offering a comprehensive and realistic reflection of high-EI interactions.
-
We introduce EmoDialogue, the first high-quality dataset featuring responses with clear, theoretically defined scoring gradients, emphasizing the quality of emotional interaction over simple emotion classification.
-
We develop EmoS, the first specialized reward model for spoken dialogue EI evaluation, alleviating the performance ceilings and scoring biases inherent in relying on proprietary models like GPT-4o-Audio as judges.
2. Related Work
2.1 Spoken Language Models
The evolution of large language models (LLMs) and large-scale audio-text paired data has transitioned spoken language models (SLMs) from traditional modular pipelines (ASR-LLM-TTS) toward unified architectures. These advanced SLMs have acquired strong capabilities in auditory understanding, open-domain dialogue, and multi-turn instruction-following. Early advances in this field primarily focused on enhancing cross-modal interaction. For instance, Qwen2-Audio improved fluent speech interaction and robust instruction-following by scaling up pretraining data and applying Direct Preference Optimization (DPO) to the audio-text alignment process. Building on this, AudioReasoner introduced Chain-of-Thought (CoT) reasoning techniques into the audio modality, significantly augmenting the model's capacity for complex logical deduction based on spoken input. However, concerning speech synthesis, these models did not operate in a strictly end-to-end manner; the decoupling of semantic generation from acoustic synthesis often created an information bottleneck, leading to significant loss of paralinguistic nuances during response generation.
Recent research has pivoted toward end-to-end SLMs to achieve unified modeling. GLM-4-Voice utilizes a low-frame-rate tokenizer to discretize continuous speech signals, enabling the model to learn bidirectional dependencies between speech and text more efficiently. Similarly, Step-Audio 2 integrates the generation of discrete audio tokens directly into the language modeling process, thereby enhancing its responsiveness to paralinguistic cues and improving the naturalness of interactions. Furthermore, to address the latency challenges inherent in real-time dialogue, Qwen2.5-Omni introduces a Thinker-Talker architecture, which decouples reasoning from streaming generation, allowing for the simultaneous output of text and speech with minimal latency. Kimi-Audio adopts a dual-head generation mechanism combined with a chunk-wise flow-matching detokenizer to balance low latency and high-fidelity expressive speech synthesis. Most recently, Qwen3-Omni significantly scales the architecture up to 30 billion parameters, leveraging optimized training objectives to further push the performance ceiling of multi-modal alignment and complex reasoning. Despite substantial architectural progress in achieving fluency, low latency, and expressiveness, no existing model has systematically addressed the complex challenges of EI in spoken dialogue. The assessment and optimization of EI for speech remains a largely unexplored frontier. To the best of our knowledge, EmoS represents the first work dedicated to evaluate the EI capabilities of SLMs within a theoretically grounded framework.
2.2 Benchmark for Spoken Language Models
Evaluation paradigms for SLMs have shifted from early tasks focused on discrete recognition towards a more holistic assessment of speech-to-speech interaction. Early benchmarks like AIR-Bench, SD-Eval, and VoxDialogue primarily utilize text-based metrics or LLMs to gauge semantic alignment and paralinguistic cues. Other frameworks target specific capabilities, such as VoxEval for knowledge comprehension and VoiceBench for instruction-following. However, a common limitation among these frameworks is their heavy reliance on transcription or a narrow focus on basic acoustic information perception.
To address the information loss inherent in text-based metrics, recent work has pivoted toward direct, audio-centric evaluation. URO-Bench serves as a speech-to-speech benchmark covering multi-turn dialogue and paralinguistic assessment, and WavReward further incorporates implicit conversational scenarios to directly evaluate the authenticity of acoustic interaction. Despite these advances, existing benchmarks tend to conflate EI with rudimentary acoustic perception. While a few methods have begun to explore the Understanding dimension, the Using and Managing dimensions remain notably understudied. In contrast, EmoSBench constitutes the first benchmark that systematically addresses all four hierarchical dimensions of EI, extending the evaluation scope from mere perception to the complex cognitive processes involved in using and managing emotions.
3. Methodology
3.1 EmoSBench
To address the limitations of prevailing approaches that equate Emotional Intelligence (EI) with simple paralinguistic perception, we establish a systematic evaluation framework anchored in the influential four-branch model. This framework is designed to comprehensively assess the advanced emotional capabilities required by Spoken Language Models (SLMs) for complex interpersonal interactions. We categorize EI into four hierarchical branches encompassing ten specific sub-task scenarios.
Perceiving Emotion: This dimension evaluates the foundational ability to decode emotional information and non-verbal cues from single-turn speech. It comprises two sub-tasks: (1) Basic Acoustic Information Perception, which tests the accuracy in identifying explicit emotional and paralinguistic features; and (2) Implicit Attitude Analysis, which assesses the model's capacity to capture subtle cues from faint acoustic hints or textual implications to generate empathetic responses.
Understanding Emotion: This dimension assesses the model's capacity to comprehend the causes and dynamic evolution of emotions in multi-turn contexts. This is examined through: (3) Emotional State Tracking, which requires the model to identify emotional transitions across dialogue turns; and (4) Emotional Causation Analysis, which evaluates the ability to connect the current complex emotional state to previously mentioned historical triggers.
Using Emotion: This dimension examines the proactive harnessing of emotions to facilitate cognitive processing and decision-making. We propose: (5) Emotion-Cognition Matching, investigating the understanding of how emotional states facilitate cognitive tasks; and (6) Emotion-Driven Plan Adjustment, requiring the model to detect acoustic-semantic conflicts (e.g., hesitant tone vs. committed text) and guide users to reconsider or optimize decisions rather than offer blind agreement.
Managing Emotion: As the pinnacle of spoken EI, this dimension evaluates the strategic capability to regulate emotions to achieve specific communicative goals. It encompasses four sub-tasks: (7) Social Strategy Execution, coordinating text and tone to achieve challenging social objectives like constructive criticism; (8) Proactive Mitigation and Emotional Buffering, employing textual and acoustic means to soften the impact of negative information; (9) Conflict Resolution and De-escalation, utilizing an empathy-first strategy in high-pressure scenarios; and (10) Value Alignment and Safety Response, maintaining dialogue while upholding principles against harmful or biased inputs.
3.2 EmoDialogue
To facilitate the training and rigorous evaluation of EI in spoken dialogue systems, we construct EmoDialogue, a comprehensive dataset comprising over 70,000 pairs of user inputs and model responses in both Chinese and English. Given the absence of speech dialogue data with fine-grained annotations for multidimensional EI, we predominantly employ a synthesis strategy based on advanced LLMs to generate training data.
Data Synthesis and Annotation. We leverage the reasoning capabilities of the DeepSeek-R1 LLM to simulate realistic user-model interactions. For every user input, the system generates four distinct candidate responses, hierarchically graded on a four-point ordinal scale to reflect varying levels of EI. Generally, Score 1 denotes responses characterized by logical incoherence or mechanical repetition, rendering them unusable; Score 2 represents factually correct but emotionally detached or blunt responses, indicating a deficiency in empathy; Score 3 corresponds to responses that are fundamentally empathetic, socially appropriate, and helpful, meeting the standard for general conversation; and Score 4 is reserved specifically for excellent responses that demonstrate high-EI capabilities, such as optimal emotion utilization or proactive emotional guidance. Concurrently, the LLM produces detailed metadata, including the intended emotional tone, paralinguistic descriptions, and a rationale for each assigned score based on sub-task criteria.
Dialogue Audio Synthesis. To obtain high-fidelity and expressive speech synthesis, we utilize the Doubao TTS API with advanced emotional control, ensuring acoustic diversity by explicitly controlling parameters such as emotion category, speech rate, and volume. Furthermore, to enhance realism, non-verbal paralinguistic cues like sighs, laughter, or hesitation pauses are strategically injected into the audio stream.
Quality Control. A rigorous multi-stage quality control pipeline is implemented to ensure the acoustic quality and consistency of the dataset. First, all synthesized audio is automatically transcribed using the Whisper-Large-V3 model, and samples with a Word Error Rate (WER) or Character Error Rate (CER) exceeding 5% are filtered out. Subsequently, a manual review is conducted to perform secondary verification on the remaining samples, checking for audio naturalness, accuracy of emotional expression, and consistency with annotated descriptions.
Benchmark Test Set Construction. To ensure reliable and unbiased evaluation, we curate EmoSBench, a core evaluation benchmark consists of approximately 4,000 audio dialogue samples. Five expert annotators select samples from each sub-task based on the principles of score discriminability, topic diversity, and high inter-annotator agreement.
Real-World Evaluation Set. While synthetic datasets provide large-scale supervision signals, authentic spoken interactions typically involve more complex acoustic conditions and spontaneous paralinguistic features. To validate the generalization capability of our model in unconstrained real-world scenarios, we construct an additional test set derived from YouTube videos. Specifically, we collect diverse conversational segments and subject them to independent evaluation by four expert annotators based on our established EI rubrics. To ensure data reliability, we enforce a strict consensus mechanism, retaining a sample only if at least three out of the four annotators agree on the same score. Ultimately, we retain 531 EI evaluation pairs to validate the model's robustness and generalization beyond synthesized speech.
3.3 EmoS
Our task is formally defined as a scalar regression problem within a generative framework: given a spoken dialogue context x, the model pitheta is tasked with generating a reasoning process followed by a predicted EI score ypred ∈ 1, 2, 3, 4 for the final response. Performance is measured by comparing ypred against the ground truth label ygt. To optimize performance, we adopt a two-stage training paradigm: an initial Supervised Fine-Tuning (SFT) phase to establish a foundational distribution of the EmoSBench, followed by Reinforcement Learning (RL) to achieve high-precision alignment. Consistent with recent findings in multimodal alignment and reasoning tasks, we observe that RL significantly outperforms SFT in handling scenarios that require precise scalar feedback and complex decision boundaries. Consequently, we adopt Group Relative Policy Optimization (GRPO) as our primary training paradigm. Unlike traditional Proximal Policy Optimization (PPO), which relies on a computationally expensive critic network to estimate the baseline, GRPO derives the baseline directly from group scores. This design reduces memory overhead and stabilizes training, making it particularly suitable for optimizing large-scale SLMs.
Reward Engineering. Conventional RL approaches often employ a binary reward signal. However, EI evaluation constitutes a meticulous ordinal assessment where the magnitude of the prediction error is critically important. For instance, misclassifying a Score 1
(low EI) response as Score 4
(high EI) represents a more serious cognitive failure compared to a Score 2
misclassification. A simple binary reward fails to distinguish these discrepancies. Moreover, linear penalty mechanisms often lead to optimization laziness, where models converge on safe, approximate predictions rather than precise assessments. To enforce strict precision, we posit that the utility of a prediction decays exponentially with respect to the error magnitude. Consequently, we design a Steep Exponential Accuracy Reward:
Ra(Δ) = alpha · e(−lambdaΔ), if Δ < 3; 0, otherwise
where Δ = ypred − ygt denotes the absolute difference between the predicted score and the ground-truth label, alpha = 2.0 is a scaling factor, and lambda = 2.0 is a distinctiveness coefficient. This function ensures that the reward drops precipitously from the maximum value (Δ = 0) to a negligible level for even minor deviations (Δ = 1), thereby penalizing vague guesses and steering the model toward exactitude.
To ensure that high-accuracy predictions are anchored in valid deductive logic rather than spurious correlations, we introduce a Rationale Fidelity Reward (Rr) utilizing Qwen3-8B as an automated critic. Specifically, we implement an adaptive reasoning weighting mechanism via a dynamic coefficient omega(Δ) that modulates the contribution of the fidelity reward based on the prediction error:
omega(Δ) = 1.0, if Δ = 0; 0.1, otherwise
This mechanism ensures that while the model is primarily reinforced for correct answers (omega = 1.0), it continues to receive residual supervision (omega = 0.1) for its internal reasoning even during incorrect predictions. This prevents the collapse of the inferential learning signal in early training stages and creates a robust optimization gradient toward the correct solution. The total reward R is thus formulated as:
R = Ra(Δ) + omega(Δ) · Rr + Rf
where Rf denotes the penalty for format non-compliance, ensuring adherence to the required CoT structure.
Optimization via GRPO. Following DeepSeek-R1, we optimize the policy pitheta by sampling a group of outputs O1, O2,..., OG for each query q. For each output Oi, we compute the total reward Ri and its advantage Ai by normalizing within the group:
Ai = (Ri − mean(R1,..., RG)) / std(R1,..., RG)
The policy is then updated by maximizing the GRPO objective, which incorporates a clipping mechanism to prevent excessively large updates and a KL-divergence penalty for stability relative to the reference model piref.
4. Experiment
4.1 Experiment Setup
Datasets and Metrics. We utilize the training split of the EmoDialogue dataset, comprising approximately 70,000 bilingual (English and Chinese) dialogue evaluation pairs, while reserving 4,000 samples for the independent test set. To ensure rigorous assessment, we adopt exact matching accuracy as the sole metric, where a prediction is only considered correct when the predicted label strictly matches the ground truth label.
Baselines. To comprehensively benchmark the Emotional Intelligence (EI) capabilities of existing models, we evaluate a diverse array of Spoken Language Models (SLMs), categorized into open-source and proprietary models. For open-source SLMs, we evaluate Qwen2-Audio-Instruct, Freeze-Omni, GLM-4-Voice, LLama-Omni 2, Step-Audio 2, Kimi-Audio, Qwen3-Omni-Instruct, and our backbone Qwen2.5-Omni-7B. We also evaluate Audio-Flamingo3; because it does not support Chinese, we report its results only on the English benchmark. For closed-source models, we assess GPT-4o-Audio and Gemini 2.5 Pro.
Implementation Details. We select Qwen2.5-Omni-7B as the foundation model for EmoS. All training is conducted on 4 NVIDIA A800 (80GB) GPUs with an effective batch size of 16 (per-device batch size 1, gradient accumulation 4). We employ a two-stage training strategy comprising Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL). In the SFT stage, the model is trained for one epoch on the EmoDialogue dataset with a learning rate of 1×10−5. In the RL stage, the model is optimized for 2,000 steps using the Group Relative Policy Optimization (GRPO) paradigm with a learning rate of 1 × 10−6. We set the sampling temperature to 1.0, generate G = 8 responses per query to estimate the group-relative baseline, and set the KL coefficient beta to 0.02.
4.2 Results and Analysis
Comparison with State-of-the-Art Models. The results demonstrate that EmoS achieves an average accuracy of 83.8%, approaching the human baseline and significantly surpassing all existing models. Most open-source spoken language models (SLMs) exhibit severe limitations, with average accuracies predominantly stagnating below 40%. Models such as Qwen2-Audio (27.4%) and Freeze-Omni (26.1%) yield results barely indistinguishable from random guessing (26.4%), rendering them practically unusable for this domain. While Qwen3-Omni manages around 50%, its deployment is heavily constrained by the computational overhead of its 30B parameters. Our backbone model, Qwen2.5-Omni, attains a modest 37.2%. It shows fundamental capability in tasks that rely primarily on language understanding and reasoning (e.g., 41.8% in EST), while its performance is already deficient in basic acoustic perception (33.5% in BAP) and drops markedly in tasks requiring psychological and social cognition, such as Emotion-Driven Plan Adjustment (27.8%). This performance structure suggests that while standard multi-modal pre-training endows models with general instruction-following ability, it does not equip them with the nuanced cognitive reasoning essential for complex EI scenarios, thus forming a clear bottleneck for handling higher-order empathetic and contextual tasks.
Crucially, EmoS markedly outperforms leading proprietary commercial models, including GPT-4o-Audio (52.6%) and Gemini 2.5 Pro (54.0%). Although these commercial models exhibit stronger general capabilities than open-source counterparts, particularly in auditory perception, they still lag behind EmoS by approximately 29.8%. This performance disparity validates our hypothesis that relying on general-purpose SOTA models as judges sets an artificially low performance ceiling. The dominance of EmoS confirms that specialized training with theoretically grounded data is indispensable for mastering the strategic and empathetic aspects of high-EI spoken dialogue.
Ablation Study. Direct SFT on the EmoDialogue dataset provides a substantial boost, raising accuracy from 37.2% to 57.0%, which confirms the high quality and pedagogical value of our dataset. However, replacing SFT with GRPO further elevates performance to 63.0%, indicating that the exploration mechanism inherent in RL allows the model to discover more optimal reasoning paths and align better with complex decision boundaries than traditional supervised learning. The incremental gains further validate the effectiveness of our specialized reward engineering. Introducing the Steep Exponential Accuracy Reward (SEAR) improves accuracy to 67.5%, confirming that penalizing near-miss errors with a steep gradient effectively combats optimization laziness and enforces exact ordinal alignment. Furthermore, the Rationale Fidelity Reward (RFR) pushes performance to 70.1%, ensuring that correct scores are derived from valid psychological reasoning rather than spurious correlations. When combined, the (GRPO + SEAR + RFR) configuration achieves 73.2%. Most notably, the finalized two-stage paradigm achieves the peak performance of 83.8%. This demonstrates a powerful synergistic effect: SFT establishes a robust foundational distribution of EI knowledge, while the subsequent RL phase, guided by our specialized rewards, refines the model's capabilities toward high-precision evaluation and deductive reasoning.
Evaluation on Real-World Scenarios. While EmoS achieves an accuracy of 83.8% on EmoSBench, closely approaching human performance, we consider that this peak accuracy might partially stem from the inherent domain alignment between the synthesized training and testing data. To verify whether the model has learned transferable emotional reasoning rather than merely adapting to synthesized speech patterns, evaluating its performance on out-of-domain data is necessary.
The results on the real-world YouTube dataset show that the inherent complexity of authentic speech (e.g., background noise, spontaneous disfluencies) causes noticeable performance degradation across all models. Leading commercial APIs, including GPT-4o-Audio (36.0%) and Gemini 2.5 Pro (43.3%), struggle to maintain prior performance levels, underscoring the difficulty of processing naturalistic spoken dialogue. Despite these challenging conditions, EmoS sustains an overall accuracy of 62.9%. Although the absolute performance gap narrows compared to the synthetic benchmark, EmoS still outperforms the strongest baseline, Gemini 2.5 Pro, by a margin of 19.6%. This sustained superiority on authentic data suggests that, despite the potential advantages on synthetic benchmarks, our training pipeline effectively enhances the model's generalization to real-world scenarios. It demonstrates that EmoS maintains a clear performance advantage over existing models when handling complex, unconstrained spoken interactions.
5. Conclusion
In this paper, we address the critical absence of comprehensive Emotional Intelligence (EI) assessment in spoken dialogue systems. We propose EmoSBench, the first theoretically grounded benchmark designed to evaluate EI beyond superficial acoustic perception, systematically covering four hierarchical dimensions of Perceiving, Understanding, Using, and Managing Emotion. Experiments on this benchmark reveal a profound gap between current general-purpose models and human-level emotional reasoning. To bridge this deficiency, we curate the EmoDialogue dataset to train EmoS, a specialized evaluator optimized via Supervised Fine-Tuning and Group Relative Policy Optimization (GRPO). By integrating a Steep Exponential Accuracy Reward (SEAR) and Rationale Fidelity Reward (RFR), EmoS achieves precise scoring and valid reasoning, significantly outperforming existing baselines. Furthermore, evaluations on authentic, unconstrained spoken interactions validate its robust generalization capabilities in real-world scenarios. Ultimately, our work establishes a foundational framework for spoken EI evaluation. In future work, we plan to continuously expand the diversity of our dataset to encompass a wider array of real-world acoustic conditions, thereby pushing the boundaries of the model's generalization capabilities. Building upon this, we aim to leverage EmoS as a reliable reward model to facilitate the alignment of Spoken Language Models (SLMs), fostering the development of genuinely emotionally intelligent agents.
Improvements for AI systems
Improvements to AI Systems Based on This Paper:
-
Emotionally Intelligent Speech Evaluation Module: Integrate EmoS as a specialized reward model into SLM training pipelines to replace generic LLM judges (e.g., GPT-4o-Audio) that cap performance. This enables systems to score responses on a 1–4 ordinal EI scale with 83.8% accuracy, approaching human baselines, and provides valid Chain-of-Thought rationales for each score.
-
Four-Branch EI Reasoning Capability: Equip SLMs with hierarchical emotional cognition—Perceiving (acoustic/paralinguistic decoding), Understanding (emotional state tracking and causation across turns), Using (emotion-cognition matching and plan adjustment), and Managing (social strategy, mitigation, conflict resolution, value alignment). This allows systems to handle ten distinct sub-tasks, from detecting implicit attitudes to de-escalating high-pressure conflicts.
-
Steep Exponential Accuracy Reward (SEAR) for Precise Scoring: Implement a reward function where utility decays exponentially with prediction error (Δ), penalizing near-miss scores (e.g., predicting 3 instead of 4) far more heavily than binary or linear rewards. This forces the model to avoid vague guesses and achieve exact ordinal alignment, improving scoring precision by 10% over standard RL.
-
Rationale Fidelity Reward (RFR) with Adaptive Weighting: Use an automated critic (e.g., Qwen3-8B) to verify that correct scores are derived from psychologically valid reasoning, not spurious correlations. The dynamic coefficient ω(Δ) gives full weight (1.0) to rationale fidelity when predictions are correct and residual weight (0.1) when incorrect, preventing reasoning collapse during training and ensuring explainable, trustworthy evaluations.
-
Two-Stage Training Paradigm (SFT + GRPO): Combine Supervised Fine-Tuning on a large bilingual dataset (EmoDialogue, 70k pairs) with Group Relative Policy Optimization using multi-reward RL. This yields a synergistic boost from 57.0% (SFT only) to 83.8% (final), enabling the model to explore optimal reasoning paths and align with complex decision boundaries beyond what supervised learning alone achieves.
-
Fine-Grained EI-Graded Dialogue Dataset (EmoDialogue): Train systems on responses with four ordinal EI levels (1=incoherent, 2=blunt, 3=empathetic, 4=high-EI) with detailed rationales, emotional tones, and paralinguistic annotations. This provides supervision for nuanced emotional interaction quality, not just emotion classification, enabling models to distinguish between
factually correct but emotionally detached
andproactively emotionally intelligent
responses. -
Real-World Generalization via Acoustic Robustness: Apply the training pipeline to handle unconstrained speech with background noise, disfluencies, and spontaneous paralinguistic cues (validated on YouTube-derived data). The system maintains 62.9% accuracy in real-world conditions, outperforming GPT-4o-Audio (36.0%) and Gemini 2.5 Pro (43.3%) by 19.6%+, enabling deployment in noisy, authentic conversational settings.
-
Bilingual Emotional Intelligence Support: Leverage the bilingual (Chinese and English) training data to enable EI evaluation and reasoning across languages, allowing systems to assess and generate emotionally intelligent responses for diverse global user bases without language-specific retraining.
What the Improved AI System Can Do:
-
Evaluate spoken dialogue responses on a 4-point EI scale with near-human accuracy, providing transparent reasoning for each score.
-
Generate high-EI responses that perceive subtle acoustic cues (e.g., sighs, hesitation), understand emotional transitions, use emotions to guide decisions, and manage conflicts strategically.
-
Detect acoustic-semantic conflicts (e.g., hesitant tone vs. committed text) and proactively guide users to reconsider decisions.
-
De-escalate high-pressure scenarios using empathy-first strategies and uphold safety values against harmful inputs.
-
Operate reliably in real-world, noisy environments with spontaneous speech, outperforming commercial APIs by significant margins.
-
Serve as a reward model for aligning future SLMs, enabling the development of emotionally intelligent agents that reason about and regulate emotions in dialogue.
Abstract
Despite significant advances in instruction-following and auditory comprehension, the evaluation of Emotional Intelligence (EI) in Spoken Language Models (SLMs) remains confined to rudimentary paralinguistic perception, lacking a systematic, theory-driven cognitive framework. We introduce EmoSBench, the first comprehensive EI evaluation benchmark for SLMs constructed upon the four-branch theoretical model, covering Perceiving, Understanding, Using, and Managing Emotion across ten sub-tasks. Preliminary assessments on EmoSBench reveal a substantial gap: even leading proprietary models like GPT-4o-Audio achieve only 52.6%, significantly trailing human baselines. To bridge this gap, we develop EmoS, a specialized evaluator model optimized via Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO). To facilitate its effective training, we curate EmoDialogue, a bilingual dataset providing necessary fine-grained supervision through response pairs with rigorously defined EI gradations. Concurrently, we introduce a reward mechanism integrating a Steep Exponential Accuracy Reward (SEAR) and a Rationale Fidelity Reward (RFR) to enforce precise ordinal scoring and valid reasoning. Experiments demonstrate that EmoS reaches 83.8% accuracy, approaching human-level performance. Furthermore, evaluations on authentic, unconstrained spoken interactions validate its robust real-world generalization, establishing a foundational framework for advancing emotionally intelligent dialogue systems.
Sources
- SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
- VoiceBench: Benchmarking LLM-Based Voice Assistants
- Qwen2-Audio Technical Report
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models
- Moshi: a speech-text foundation model for real-time dialogue
- Kimi-Audio Technical Report
- LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis
- OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- EmoBench-M: Benchmarking Emotional Intelligence for Multimodal Large Language Models
- WavReward: Spoken Dialogue Models With Generalist Reward Evaluators
- Baichuan-Omni-1.5 Technical Report
- MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
- Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
- Step-Audio 2 Technical Report
- Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models
- Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
- Qwen2.5-Omni Technical Report
- Qwen3-Omni Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering