Introducing HALC: A general pipeline for the systematic and reliable construction of prompts for automated coding with LLMs in the computational social sciences
cs.CL, cs.AI, cs.LG
Submitted: 2025-07-29
Updated: 2026-09-08
Comments: 62 pages, 7 figures and 15 tables. Published in Communication Methods and Measures (Open Access)
Journal ref: Communication Methods and Measures, 20(3), 225-254 (2026)
DOI: 10.1080/19312458.2026.2693637
Code: https://github.com/ollama/ollama
Project page: https://datamod2023.github.io/pdf/datamod_Lambert.pdf
License: http://creativecommons.org/licenses/by/4.0/
The gist: LLMs are seeing widespread use for task automation, including automated coding in the social sciences.
Terminology
Abstract
LLMs are seeing widespread use for task automation, including automated coding in the social sciences. However, even though researchers have proposed different prompting strategies, their effectiveness varies across LLMs and tasks. Often trial and error practices are still widespread. Our study aims to fill this gap and evaluate how LLMs can be used in a systematic and transparent way to produce reliable codings in content analyses. We propose HALC-a general pipeline that allows for the systematic and reliable construction of prompts for any given coding task and model. We develop this pipeline based on current literature and findings of a prestudy investigating consistency and influencing factors of LLM codings. We also apply HALC on two other datasets covering different thematic contexts, document types, languages, and coding units to test its applicability. Based on more than three million LLM requests, our results demonstrate that the pipeline is capable of identifying prompts for reliable codings in different settings. We also discuss shortcomings and further potential for development.
Sources
- LLM-Assisted Content Analysis: Using Large Language Models to Support Deductive Coding
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Mistral 7B
- DeepSeek-V3 Technical Report
- Automated Annotation with Generative AI Requires Validation
- Is Temperature the Creativity Parameter of Large Language Models?
- Testing the Reliability of ChatGPT for Text Annotation and Classification: A Cautionary Remark
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Evaluation is all you need. Prompting Generative Large Language Models for Annotation Tasks in the Social Sciences. A Primer using Open Models
- A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT
- Chain of Draft: Thinking Faster by Writing Less
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering