Reading Emotions in the Token Space: Discriminative Adaptation of SpeechLLMs for Emotion Recognition
cs.CL, cs.AI, cs.SD, eess.AS
Submitted: 2026-09-17
Updated: 2026-09-17
Code: https://github.com/emo-box/EmoBox
Project page: https://emo-box.github.io/leaderboard1.html
License: http://creativecommons.org/licenses/by/4.0/
The gist: SpeechLLMs have shown strong potential for emotion recognition, yet they read the predicted emotion off a generative decoder not suited for classification: it can emit labels outside the target set
Terminology
Abstract
SpeechLLMs have shown strong potential for emotion recognition, yet they read the predicted emotion off a generative decoder not suited for classification: it can emit labels outside the target set and favors frequent classes. We propose a discriminative adaptation that reads the final prompt token's hidden state through a classification head, producing a label in one forward pass without modifying the backbone. Because this readout starts from the hidden state the model would otherwise decode, it gives a controlled comparison of generative and discriminative inference in an otherwise identical speechLLM. We keep the head a single linear layer, trading little accuracy for interpretability: each emotion becomes one direction in the LLM output token space, revealing associated tokens. On IEMOCAP, across two speechLLM architectures, it improves Macro F1 and removes hallucinations, with largest gains on realistic ASR transcripts. Our analysis reveals that these emotion directions encode indirect associations mirroring biases in web-scale text.
Sources
- A Simple Method to Enhance Pre-trained Language Models with Speech Tokens for Classification
- An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
- Label Supervised LLaMA Finetuning
- Beyond Generative Decoding: Discriminative Hidden-State Readout from a Native Omni-Modal LLM for Multimodal Sentiment Analysis
- EmoQ: Speech Emotion Recognition via Speech-Aware Q-Former and Large Language Model
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering