Do We Still Need Humans in the Loop? Human vs. LLM Annotation in Active Learning for TikTok Hate Speech Detection
cs.CL, cs.AI
Submitted: 2026-04-15
Updated: 2026-08-31
License: http://creativecommons.org/licenses/by/4.0/
The gist: Annotating data remains a costly bottleneck for supervised NLP.
Terminology
Abstract
Annotating data remains a costly bottleneck for supervised NLP. Active learning (AL) reduces the number of human labels needed by selecting only the most informative instances, while instruction-tuned LLMs attack the same bottleneck from the other side, making labels cheap enough to annotate entire corpora. This raises two questions: can LLM labels replace human labels within the AL loop, and does AL remain necessary when entire corpora can be cheaply labeled? We investigate both by training supervised hate speech classifiers on a new dataset of 278K German political TikTok comments, comparing human and LLM annotation under matched conditions. LLM annotation at scale outperforms human-supervised classifiers at roughly one-tenth the cost, for both a closed-source (GPT-5.2) and an open-weight (Qwen3.5-122B-A10B) LLM, and the advantage is robust under soft-label evaluation. It hinges on the annotation interface: only a two-question decomposition mirroring the human annotation task unlocks it. AL provides no reliable advantage over random sampling in our prefiltered pool. Error structure depends on the LLM: only GPT-5.2 matches the human FP/FN balance, while other variants over-flag border-control and economic-competition discourse. Humans remain essential as evaluators; for training labels, the question shifts to which LLM, which interface, and what shape of pool.
Sources
- A Dataset for the Detection of Dehumanizing Language
- Beyond the Explicit: A Bilingual Dataset for Dehumanization Detection in Social Media
- ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks
- BERTopic: Neural topic modeling with a class-based TF-IDF procedure
- Bayesian Active Learning for Classification and Preference Learning
- Can Reasoning Help Large Language Models Capture Human Annotator Disagreement?
- Automated Annotation with Generative AI Requires Validation
- A Survey of Active Learning for Text Classification using Deep Neural Networks
- Measuring the Reliability of Hate Speech Annotations: The Case of the European Refugee Crisis
- The Llama 3 Herd of Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering