Knowing Your Uncertainty -- On the application of LLM in social sciences

arXiv:2512.05461 · cs.CY, cs.AI, cs.HC · Submitted 2025-12-05 · Read on arXiv

cs.CY, cs.AI, cs.HC

Submitted: 2025-12-05

Updated: 2026-09-06

Comments: 56 pages, 13 figures

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

The gist: Large language models (LLMs) are rapidly being integrated into computational social science research, yet their blackboxed training and designed stochastic elements in inference pose unique

Terminology

Abstract

Large language models (LLMs) are rapidly being integrated into computational social science research, yet their blackboxed training and designed stochastic elements in inference pose unique challenges for scientific inquiry. This article argues that applying LLMs to social scientific tasks requires explicit assessment of uncertainty -- an expectation long established in both quantitative methodology in the social sciences and machine learning. We introduce a unified framework for evaluating LLM uncertainty based on Hill numbers, a family of diversity measures. By transforming existing uncertainty quantification (UQ) metrics into Hill numbers, the framework provides a common and intuitive scale for interpreting variation in LLM outputs while accommodating different notions of semantic similarity and different sensitivities to output distributions. We show how it might help the application of LLMs in social sciences through four empirical applications.

Sources

Related papers