Evaluating Decision Models for Text Annotation in Computational Social Science
cs.CL, cs.CY
Submitted: 2026-09-21
Updated: 2026-09-23
Comments: 47 pages, 7 figures
Code: https://github.com/hazemibrahim97/decision-models-css
License: http://creativecommons.org/licenses/by/4.0/
The gist: Computational social science increasingly relies on large language models for text annotation, and the validity of published findings now rests on the labels generated by such models.
Terminology
Abstract
Computational social science increasingly relies on large language models for text annotation, and the validity of published findings now rests on the labels generated by such models. Decision models, a new model class built for categorical question answering, answer typed questions with a choice, a probability distribution over the label set, and a confidence score rather than free text, at a small fraction of frontier inference prices. Whether their answers are accurate, and whether that stated confidence can be trusted on social science constructs, are unknown. Here, we mirror the evaluation of Ziems et al. (2024) on 18 computational social science classification tasks (7,977 items), comparing the first commercial decision model and two open-weight counterparts against 19 frontier and open-weight language models under the same zero-shot protocol. The decision model trails the per-task best LLM on 14 of 15 evaluation tasks, with a median deficit of 11.6 macro-F1 points, at a median 44 times lower measured cost. Its confidence is better calibrated than the verbalized confidence of 16 of the 19 LLMs, yet three frontier models show lower median calibration error (0.157 against 0.066). While items above 0.9 confidence are typically labeled accurately (median accuracy 0.815), on one task, empathy in peer-support dialogues, the model reports high confidence while performing near chance. Nonetheless, our results suggest that decision models are useful as a first step in the annotation pipeline: routing low-confidence items to an LLM matches or exceeds the LLM alone at a quarter to half of its cost.
Sources
- ChatGPT-4 Outperforms Experts and Crowd Workers in Annotating Political Twitter Messages with Zero-Shot Learning
- Automated Annotation with Generative AI Requires Validation
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
- RLCD: Reinforcement Learning from Contrastive Distillation for Language Model Alignment
- Language Models are Few-Shot Learners
- Neutralizing the Narrative: AI-Powered Debiasing of Online News Articles
- Large Language Models are often politically extreme, usually ideologically inconsistent, and persuasive even in informational contexts
- Deep reinforcement learning from human preferences
- Training language models to follow instructions with human feedback
- Constitutional AI: Harmlessness from AI Feedback
- Qwen2.5 Technical Report
- Selective Classification for Deep Neural Networks
- On Calibration of Modern Neural Networks
- Teaching Models to Express Their Uncertainty in Words
- Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
- Consistent Estimators for Learning to Defer to an Expert
- TempoWiC: An Evaluation Benchmark for Detecting Meaning Shift in Social Media
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering