ViTOED: A Dataset for Target-Oriented Emotion Detection on Vietnamese Social Media Texts

arXiv:2608.12776 · cs.CL · Submitted 2026-08-13 · Read on arXiv

University of Information Technology · Vietnam National University

cs.CL

Submitted: 2026-08-13

Updated: 2026-09-17

Comments: Accepted for publication at 2026 International Conference on Multimedia Analysis and Pattern Recognition (MAPR 2026)

Code: https://github.com/doccano/doccano

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

The gist: This paper introduces ViTOED, a novel dataset for target-oriented emotion detection in Vietnamese social media texts.

Terminology

Summary

This paper introduces ViTOED, a novel dataset for target-oriented emotion detection in Vietnamese social media texts. The ViTOED comprises 10,985 user comments and 21,244 manually annotated opinion quadruples (source, target, expression, polarity) that follow strict guidelines. The dataset reveals Vietnamese-specific phenomena, such as implicit sources and targets and vocabulary ambiguities, enabling deeper analysis of user emotions toward entities. We propose a baseline using structured sentiment graphs and evaluate various Vietnamese pre-trained language models. The empirical results highlight challenges in span detection and relation extraction and indicate substantial room for model improvement in Vietnamese Target-Oriented Emotion Detection tasks.

The paper's main contributions are: (1) introducing ViTOED, a new human-annotated Vietnamese dataset for emotion detection by target, constructed under a strict annotation process; (2) analyzing the dataset to identify a phenomenon in Vietnamese regarding the difference between the target, source, and expression in expressing emotion in a sentence; and (3) evaluating various Vietnamese pre-trained language models using the Structured Sentiment Analysis Graph proposed by [4] on the dataset to investigate the performance of machine learning models for target-oriented emotion tasks in Vietnamese.

The dataset was created using data from social networks, including 6,000 sentences from the UIT-VSMEC and 5,010 comments collected raw from social media platforms. Annotators were given guidelines and annotated a set of 210 sample sentences to evaluate agreement. The annotation process involved four rounds, with inter-annotator agreement improving from 0.17 to 0.32 for Source, 0.31 to 0.61 for Target, 0.34 to 0.77 for Expression, and 0.15 to 0.47 for Polarity. The low IAA for Source is because many sentences are difficult to detect as containing a source, while the moderate agreement for Polarity depends on consistency across the quadruples.

The dataset was divided into train, dev, and test subsets with a ratio of 7:1:2. The distributions of source, target, and expression across the splits are similar, with sentiment leaning slightly toward the negative class. The analysis shows that opinions missing the Source (T-E type) account for 52.10% of the total opinions, which is the most common type in social media comments. The top words for Source include tao (I), "t (abbreviated I), ta (I), e (younger), and Em (younger). For Target, the popular words are mày (You), mấy (any), thằng (he), tao (me), and ta (me). For Expression, positive words include yêu (love), thương (love), đẹp (beauty), vui (happy), and cảm ơn (thanks); neutral words include đi (go), ko (not), mua (buy), bảo (said), and coi (see); negative words include sợ (scare), đéo (f*ck), buồn (sad), chửi (scold), and ghét" (hate).

The baseline model employs the Structured Sentiment Graph proposed by [4], using tokenized input passed through word, pos-tag, and character embedding layers, then through an LSTM layer to construct character representation vectors. These are concatenated with BERT-contextualized vectors and passed into a BiLSTM. The sentiment graph is created using two parsing graph representations: head-first (first token as root) and head-final (final token as root). The models evaluated include multilingual models (mBERT, XLM-R, mT5, mBART) and monolingual Vietnamese models (PhoBERT, ViSoBERT, CafeBERT, ViT5, BARTPho).

The experimental results show that monolingual pre-trained language models, such as ViSoBERT and PhoBERT, achieve better performance than multilingual models on this task. The head-first approach is efficient for extracting entities such as source and target, while the head-final approach is suitable for detecting expressions. mT5 and CafeBERT achieved the highest results for the target-oriented emotion detection task, with mT5 good at extracting source entities and CafeBERT good at extracting target entities. ViT5 is robust at exploiting the expression, parsing, and sentiment graphs, while ViSoBERT shows high performance in detecting the targeted expression.

Error analysis identified three main challenges: (1) difficulty in recognizing source, target, and expression entity spans due to the complexity of the sentence vocabulary; (2) misprediction between the root and edges, making it challenging to detect the correct target emotion; and (3) confusion between source and target due to Vietnamese morphology, affecting model performance in correctly identifying the source and target in a sentence. The model also fails to predict missing entities or correctly align entities, leading to label misalignment and decreased accuracy.

Improvements for AI systems

Improvements to AI Systems:

  1. Enhance Vietnamese-specific emotion detection models by integrating ViTOED’s annotated quadruples (source, target, expression, polarity) into training pipelines. This allows models to learn implicit sources/targets (e.g., dropped pronouns like tao or mày) and handle vocabulary ambiguities (e.g., "đi" as neutral vs. imperative) more accurately, improving F1 scores for span detection and relation extraction in Vietnamese social media text.

  2. Develop a hybrid parsing strategy that combines head-first and head-final graph representations dynamically. The improved system can switch between these based on the entity type (e.g., head-first for source/target extraction, head-final for expression detection), reducing root-edge mispredictions and improving overall structured sentiment graph accuracy.

  3. Implement a source-target disambiguation module using morphological and syntactic features from Vietnamese (e.g., pronoun drop, word-order flexibility). This module can pre-classify whether a token is more likely a source or target, reducing confusion errors that currently lower model performance.

  4. Add a missing-entity prediction layer to the baseline model. This layer learns to identify when a source or target is implicit (52.1% of cases) and generates a placeholder or infers the entity from context, improving quadruple completeness and reducing label misalignment.

  5. Fine-tune monolingual Vietnamese PLMs (e.g., ViSoBERT, PhoBERT) with a multi-task objective that jointly optimizes span detection, relation extraction, and polarity classification. The improved system can leverage shared representations to boost performance on all subtasks, especially for rare or ambiguous expressions.

  6. Incorporate an error-aware training mechanism that uses the paper’s error analysis (e.g., difficulty in recognizing complex spans, root-edge confusion) to generate adversarial examples or reweight training samples. This makes the model more robust to Vietnamese-specific challenges like mixed dialectal terms (e.g., đéo) or abbreviated pronouns ("t").

What the Improved AI System Can Do:

  • Accurately detect target-oriented emotions in Vietnamese social media comments, including implicit sources/targets and ambiguous expressions, with higher precision and recall than current baselines.

  • Generate structured sentiment graphs (source, target, expression, polarity) that are complete and correctly aligned, even for sentences with missing entities or complex vocabulary.

  • Adaptively choose the best parsing strategy per entity type, reducing errors in root-edge assignments and improving downstream tasks like opinion summarization or hate speech detection.

  • Distinguish between source and target in morphologically complex Vietnamese sentences, enabling more nuanced analysis of who is expressing emotion toward whom.

  • Perform robustly across both formal and informal Vietnamese text, including abbreviations, slang, and dialectal variations, making it suitable for real-world social media monitoring and sentiment analysis applications.

Abstract

This paper introduces ViTOED, a novel dataset for target-oriented emotion detection in Vietnamese social media texts. The ViTOED comprises 10,985 user comments and 21,244 manually annotated opinion quadruples (source, target, expression, polarity) that follow strict guidelines. The dataset reveals Vietnamese-specific phenomena, such as implicit sources and targets and vocabulary ambiguities, enabling deeper analysis of user emotions toward entities. We propose a baseline using structured sentiment graphs and evaluate various Vietnamese pre-trained language models. The empirical results highlight challenges in span detection and relation extraction and indicate substantial room for model improvement in Vietnamese Target-Oriented Emotion Detection tasks.

Sources

Related papers