"You are an expert annotator": Automatic Best-Worst-Scaling Annotations for Emotion Intensity Modeling
cs.CL
Submitted: 2024-03-26
Updated: 2026-09-16
Comments: Published at NAACL 2024: https://aclanthology.org/2024.naacl-long.439/
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Labeling corpora constitutes a bottleneck to create models for new tasks or domains.
Terminology
Abstract
Labeling corpora constitutes a bottleneck to create models for new tasks or domains. Large language models mitigate the issue with automatic corpus labeling methods, particularly for categorical annotations. Some NLP tasks such as emotion intensity prediction, however, require text regression, but there is no work on automating annotations for continuous label assignments. Regression is considered more challenging than classification: The fact that humans perform worse when tasked to choose values from a rating scale lead to comparative annotation methods, including best-worst scaling. This raises the question if large language model-based annotation methods show similar patterns, namely that they perform worse on rating scale annotation tasks than on comparative annotation tasks. To study this, we automate emotion intensity predictions and compare direct rating scale predictions, pairwise comparisons and best-worst scaling. We find that the latter shows the highest reliability. A transformer regressor fine-tuned on these data performs nearly on par with a model trained on the original manual annotations.
Sources
- Towards Ethics by Design in Online Abusive Content Detection
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Using Natural Language Explanations to Rescale Human Judgments
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering