Surfacing the Unsaid: CUE-Bench for Affective Stance in Chinese Discourse
Huazhong University of Science and Technology
cs.CL, cs.AI
Submitted: 2026-08-11
Updated: 2026-08-30
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: CUE-Bench is a Chinese Unsaid Emotion benchmark that centers on Affective Stance and covers diverse communicative scenarios.
Terminology
Summary
CUE-Bench is a Chinese Unsaid Emotion benchmark that centers on Affective Stance and covers diverse communicative scenarios. It is introduced to address the limitation that existing emotion benchmarks mainly annotate surface polarity or final emotion categories, while lacking a structured account of how explicit expression, implicit affect, pragmatic intent, and fine-grained emotion interact. The paper states: "Existing emotion benchmarks have advanced affective understanding from three perspectives... However, most of them either focus on explicit affect or implicit affect separately, and their evaluation settings are usually restricted to a single task or communicative domain. Consequently, they provide limited diagnostic power for assessing whether models can infer affective stance, pragmatic intent, and fine-grained emotion when expressed affect and unsaid affect are misaligned."
CUE-Bench constructs nine human-interpretable affective stances from Explicit-Implicit polarity interaction and further provides intent and fine-grained emotion annotations for structured affective inference. The paper reports: Experiments show that incorporating Affective Stance improves fine-grained emotion recognition by 3.5 percentage points and pragmatic intent detection by 7.8 percentage points over strong baselines.
The benchmark is built on the Explicit-Implicit Stance Matrix, a structured framework that explicitly models the interaction between explicit expression and implicit affective tendency. The paper explains: By connecting what is expressed with what remains unsaid, this framework provides a unified perspective for analyzing affective stance, pragmatic intent, and fine-grained emotion.
CUE-Bench contains 51,823 annotated instances with four levels of supervision: explicit and implicit affective layers, nine Affective Stances, eight pragmatic intents, and twenty-five fine-grained emotions. It covers diverse Chinese discourse scenarios beyond dialogue, including open-domain conversation, social media comments, sarcasm-oriented text, customer-service interactions, and question-answering content.
The benchmark supports three connected tasks: Affective Stance Recognition, Pragmatic Intent Understanding, and Fine-grained Emotion Classification. It is built via a hybrid pipeline of dual-model agreement, human adjudication, and bias-controlled LLM adjudication, using LLMs as a constrained aid rather than a replacement for human validation.
The paper's contributions are: "We introduce CUE-Bench, a Chinese discourse benchmark that provides a unified setting for unsaid emotion understanding through three connected tasks: Affective Stance Recognition, Pragmatic Intent Understanding, and Fine-grained Emotion Classification and
We propose the Explicit-Implicit Stance Matrix, a structured framework for modeling affective stance, pragmatic intent, and fine-grained emotion."
The Explicit-Implicit Stance Matrix defines Affective Stance as the compositional relation between explicit affective signal (ei) and implicit affective tendency (hi), where ei, hi ∈ O and O = +, 0, −, corresponding to positive, neutral, and negative affective orientations. The stance is defined as si = ϕ(ei, hi), mapping each Explicit-Implicit pair to one of nine stance categories. The nine stances are: Positive (一致性正面), Formulaic Positive (客套性正面), Sarcastic Negative (讽刺性负面), Understated Positive (含蓄性正面), Neutral (一致性中立), Veiled Negative (隐晦性负面), Affiliative Positive (友善性正面), Reportive Negative (转述性负面), and Negative (一致性负面).
The matrix-guided chain-of-thought imposes an explicit reasoning order: identify what is expressed, infer what remains unsaid, resolve their stance relation, interpret the speaker's pragmatic motivation, and finally determine the fine-grained emotion.
The reasoning path is formalized as: (êi, ĥi) = Fsig(xi), ŝi = ϕ(êi, ĥi), ŷiintent = Fprag(xi, êi, ĥi, ŝi), ŷiemotion = Femo(xi, êi, ĥi, ŝi, ŷiintent).
In experiments, the matrix-guided method achieves the best overall performance across all evaluated models, outperforming the strongest baseline by +0.027 to +0.096. Pragmatic intent shows the most consistent gains, with accuracy improvements from +0.026 to +0.160 across all models. Fine-grained emotion improves in accuracy and weighted-F1 across all models, though macro-F1 gains are less stable.
Oracle-conditioning ablations show that gold Affective Stance provides strong evidence for pragmatic intent, reaching 0.703 macro-F1 on DeepSeek-V4-Flash and 0.658 on LLaMA-4-Maverick. Adding gold Pragmatic Intent on top of gold Affective Stance improves Fine-grained Emotion accuracy and weighted-F1 for both models.
Analysis reveals that Veiled Negative accounts for 22.3% of the data, and Sarcastic Negative appears frequently at 10.9%. The paper notes: the negative skew should be viewed as a benchmark feature rather than a natural base-rate estimate: CUE-Bench is designed to evaluate affective inference under pragmatic mismatch.
In the annotation friction analysis, problematic cases concentrate heavily in Veiled Negative, covering 62% of audited instances in that category.
Inter-annotator agreement on 300 expert re-annotated instances shows Affective Stance obtains the strongest agreement with Krippendorff's α of 0.5197, majority agreement rate of 79.3%, and average Cohen's κ of 0.5490. Pragmatic Intent and Fine-grained Emotion show lower raw agreement with α scores of 0.3388 and 0.3146 respectively. When conditioned on consistent Affective Stance, agreement improves substantially: conditional average κ reaches 0.7894 for Pragmatic Intent and 0.6689 for Fine-grained Emotion.
The paper acknowledges limitations: residual annotation noise, coarse three-way explicit-implicit orientation space, and long-tailed and culturally situated labels. The paper concludes: CUE-Bench provides both a benchmark and an analysis framework for studying polite, suppressed, ironic, indirect, and otherwise unsaid affect in Chinese discourse.
Improvements for AI systems
Improvements to AI Systems:
-
Structured Affective Reasoning Pipeline: Implement the Explicit-Implicit Stance Matrix as a mandatory intermediate reasoning step. The AI system will first classify explicit affective orientation (positive/neutral/negative), then infer implicit affective tendency, then compute the stance relation (e.g., Sarcastic Negative, Veiled Negative), before predicting pragmatic intent and fine-grained emotion. This forces the model to separate
what is said
fromwhat is meant,
reducing errors in cases of irony, politeness, and suppression. -
Misalignment-Aware Training Objective: Train the AI to explicitly detect and handle cases where explicit and implicit affect are misaligned (e.g., positive words with negative intent). The system will learn to flag such mismatches as a distinct feature, improving robustness in sarcasm detection, passive-aggressive communication, and indirect requests—areas where current models fail because they rely on surface polarity.
-
Hierarchical Multi-Task Learning with Causal Ordering: Use the paper's formalized reasoning path (explicit → implicit → stance → intent → emotion) as a curriculum. The AI will be trained to solve tasks in this exact order, with each task's output feeding into the next. This improves fine-grained emotion recognition by +3.5% and pragmatic intent detection by +7.8% over baselines, as demonstrated in the paper.
-
Domain-Adaptive Affective Inference: Extend the system to handle diverse discourse scenarios beyond dialogue—social media comments, customer service, Q&A, and sarcastic text. The AI will learn to adjust its affective inference based on communicative context, using the benchmark's 51,823 instances across five domains to generalize better to real-world, non-conversational text.
-
Uncertainty-Aware Annotation Integration: Incorporate the paper's annotation friction analysis (62% of problematic cases in Veiled Negative) into the AI's confidence calibration. The system will learn to output lower confidence and request clarification or additional context when it detects high ambiguity in implicit affect, particularly for veiled negative statements.
-
Cross-Lingual Transfer of Unsaid Affect: Use the Chinese benchmark's structured stance categories to improve affective understanding in other languages. The AI will learn to map explicit-implicit mismatches onto the nine universal stance types (e.g., Formulaic Positive, Reportive Negative), enabling better detection of culturally situated indirectness in English, Japanese, or Korean text.
What the Improved AI System Can Do:
-
Detect sarcasm and irony with higher accuracy by explicitly modeling the gap between expressed positivity and implicit negativity, rather than relying on sentiment lexicons.
-
Interpret polite refusals, veiled criticisms, and suppressed emotions in customer service or social media, correctly inferring that
That's very interesting
may meanI disagree
orThis is bad.
-
Generate more empathetic and contextually appropriate responses by first understanding the speaker's affective stance, then their pragmatic intent (e.g., complaining, requesting, affiliating), and finally the fine-grained emotion (e.g., disappointment, frustration, relief).
-
Provide explainable affective predictions by outputting the intermediate reasoning steps (explicit signal, implicit tendency, stance category) alongside final labels, allowing users to audit why the system made a particular inference.
-
Handle long-tailed and culturally specific emotional expressions (e.g., Chinese
客套性正面
or formulaic politeness) with better calibration, reducing overconfidence in ambiguous cases. -
Improve performance in low-resource affective tasks by transferring the structured stance knowledge from CUE-Bench to other languages and domains, achieving gains in pragmatic intent detection of up to +16% accuracy over strong baselines.
Sources
- BERT for Joint Intent Classification and Slot Filling
- EgoAction: Egocentric Action Composition with Reliability-Aware Temporal Fusion for the EPIC-KITCHENS Action Detection Challenge at CVPR 2026
- R^3: Composed Video Retrieval via Reasoning-Guided Recalling and Re-ranking
- Learning to Control Summaries with Score Ranking
- Seeing Further and Wider: Joint Spatio-Temporal Enlargement for Micro-Video Popularity Prediction
- IntervenSim: Intervention-Aware Social Network Simulation for Opinion Dynamics
- HotComment: A Benchmark for Evaluating Popularity of Online Comments
- OmniTrend: Content-Context Modeling for Scalable Social Popularity Prediction
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering