Multi-Perspective LLM Annotations for Valid Analyses in Subjective Tasks
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Multi-Perspective LLM Annotations for Valid Analyses in Subjective Tasks".
Tom: Large language models are increasingly used to annotate texts, but their outputs reflect some human perspectives better than others.
Jane: First, who's behind it and why it matters.
Paper summary: Jane: So, wrapping up our discussion on "Multi-Perspective LLM Annotations for Valid Analyses in Subjective Tasks," we've seen how the PERSPECTIVEDRIVEN INFERENCE framework works by shifting the focus from a single ground truth to estimating a vector of group-specific quantities.
Tom: That’s right, and it really comes down to the idea that correcting LLM annotation error is important when you are dealing with a distribution of perspectives rather than just one objective truth.
Lu: The authors’ main contribution is that they developed an adaptive sampling strategy that smartly concentrates human annotation effort on groups where the LLM proxies are least accurate, which is a targeted subgroup-improvement method.
Meng: It’s interesting how they show that this approach can maintain coverage above ninety percent across different age groups for both politeness and offensiveness ratings, unlike some LLM-only methods which fall short.
Lalam: This research really suggests that for subjective tasks, we should be aiming to capture the variation in viewpoints across different demographics rather than just trying to find one unified score.
Tom: Exactly, and the authors show that this method can achieve targeted improvements for harder-to-model demographic groups, like Age fifty plus, demonstrating that error-driven upsampling is useful.
Jane: So in simple terms, it’s about using the LLM's errors to guide human input so we can get a more faithful representation of how people actually rate things across different perspectives.
Lu: The implication for the field is that this approach could be very useful in RLHF or LLM-as-judge evaluations where we need to understand those nuanced group differences in language use.
Meng: It seems like a practical tool for making our data collection and evaluation processes more robust when dealing with subjective human judgments that have inherent group differences.
Lalam: This paper shows that modeling the distribution of opinions, rather than just a single opinion, is where the real statistical value lies for understanding language usage.
Tom: We’ve covered how PERSPECTIVEDRIVEN INFERENCE estimates this distribution using adaptive sampling based on predicted LLM error across different groups.
Jane: And the overall message is that by leveraging these LLM annotations strategically, we can conduct more robust statistical inference about human perspectives than relying on simpler methods.
Conclusion: Tom: So, we've been deep into how this work uses AI to guide human input to get better results on subjective tasks, and now we're coming to the finish line with some thoughts on what this whole piece actually means for us.
Jane: It’s been fascinating tracking how they moved away from just looking for one single answer and started thinking about the whole landscape of opinions across different groups.
Lu: I think it really highlights how we can use statistical methods to make sense of messy, subjective human language, which is a huge area for creative AI exploration.
Meng: From an engineering standpoint, the idea of targeting where the AI is weakest in its judgment seems like a very smart way to manage our annotation budget efficiently.
Lalam: And I think this work has big implications for how we build systems that understand and respect diverse human viewpoints in a meaningful way.
Tom: Exactly, Lu, it’s not just about getting a better score; it’s about making sure the data we use to judge things actually reflects the real variety of people's opinions.
Jane: They did this by formalizing the problem as estimating a vector of group-specific quantities instead of just one number.
Lu: That shift from scalar to vector is where things get really interesting; it opens up whole new avenues for how we model social dynamics in text.
Meng: It means we can be much more strategic about where we invest our time in annotating, focusing on the areas that give us the most valuable information per annotation effort.
Lalam: And for culture, this suggests that AI systems could be trained on annotations that represent a much richer and more nuanced picture of how different people actually talk to each other.
Tom: It’s really about moving past simple majority votes to capture the actual spectrum of human disagreement when it comes to things like politeness or offensiveness.
Jane: So, in simple terms, this paper shows us a smarter way to use AI-generated labels so that we can get a much more accurate and representative picture of how different people feel about things.
Lu: And the authors’ approach with adaptive sampling really shows how these kinds of error signals can be used to guide human effort in a very practical way.
Meng: It’s a proof of concept, but it shows that this kind of targeted correction based on where the AI is struggling has real potential for improving downstream applications.
Lalam: This moves us toward building AI that doesn't just predict what one person might say, but understands the entire spectrum of how different communities view language.
Tom: It’s a really powerful way to use these models to create data that actually tells a deeper story about human interaction.
Johns Hopkins University · University of Washington
cs.CL
Submitted: 2026-03-22
Updated: 2026-09-30
Journal ref: AACL 2026
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 92/100
The gist: Large language models are increasingly used to annotate texts, but their outputs reflect some human perspectives better than others.
Key concepts
- estimate $\Theta^*$
- This is the main goal: estimating a vector of expected annotation values, one for each demographic group. Instead of finding one overall answer, this method seeks to understand how the annotation value differs across all defined groups.
- LLM error as $(H\hat{i} - H_i)^2$
- This measures how wrong an LLM's annotation ($H\hat{i}$) is compared to the true human judgment ($H_i$). This squared error serves as a signal to identify which demographic groups are causing the LLM the most trouble.
- Adaptive sampling strategy
- This is a dynamic process that decides where to spend human annotation time. It prioritizes annotating groups where the predicted LLM error is high, ensuring that limited human resources are focused on improving estimates for those challenging demographic subgroups.
Terminology
Summary
Large language models are increasingly used to annotate texts, but their outputs reflect some human perspectives better than others. This work introduces PERSPECTIVEDRIVEN INFERENCE, a method that treats the distribution of annotations across groups as the quantity of interest and estimates it using a small human annotation budget through an adaptive sampling strategy.
The gist
This framework formalizes multi-perspective corpus inference as estimation of a vector estimand (one entry per demographic group) rather than a single scalar, and develops an adaptive sampling strategy that concentrates human annotation effort on groups where LLM annotations are predicted to be less reliable, showing targeted improvements for harder-to-model demographic groups relative to uniform sampling baselines while maintaining coverage.
Background and Motivation
Annotators’ background shows considerable heterogeneity in preferences on subjective tasks, and collapsing annotator disagreement into a majority vote discards variation that matters for understanding how people use and perceive language. While methods like persona prompting exist to simulate group perspectives, they do not always improve LLM performance; they can exaggerate group differences or misrepresent minority viewpoints. The goal of this research is not to train LLMs to make more accurate predictions, but to leverage their annotations to conduct robust downstream statistical inference that faithfully retains annotator disagreement.
Framework and Problem Setup
The problem is set on a corpus of texts where each instance is annotated by multiple annotators, each from a particular demographic group. The objective is to estimate a vector of group-specific quantities, denoted as estimate Θ∗ = (Θ∗g1,..., Θ∗gk),
where each θ∗gk equals the expected annotation value within demographic group gk. This estimation is performed using LLM annotations (Hˆi) as a cheap but potentially inaccurate proxy for human judgments, defined by the LLM error as (Hˆi − Hi)2.
Perspective-Driven Inference Mechanism
The core of PERSPECTIVEDRIVEN INFERENCE involves an adaptive loop:
-
Collect initial human annotations via
burn-in uniform sampling.
-
Enter an
adaptive loop
where the human annotation probability πi is set proportional to the predicted LLM squared error from annotator demographics, specifically:πi ∝ err di(di), normalized so that Pn i=1 πi = nhuman.
-
The function err di is approximated by training a black-box classifier (XGBoost) on feature–label pairs of
(dj,(Hˆj−Hj)2)
to predict the squared LLM error from the annotator’s demographic profile.
Evaluation and Key Findings
The framework was evaluated on politeness and offensiveness rating tasks using GPT-5.2 across eight different LLMs and three prompting strategies (zero-shot, few-shot, persona). The key findings include:
-
PDI maintains
coverage above 90% across all three age groups
for both politeness and offensiveness, unlike LLM-only methods which often fall below 90%. -
PDI achieves the
lowest delta for Age 50+
in both tasks (e.g., 11.23% delta for politeness), outperforming PPI (13.63%) and all three LLM-only variants, demonstrating thaterror-driven upsampling
is useful. -
The adaptive allocation confirms the utility of the error signal: PDI allocates
a 33% increase
in annotations to Age 50+ compared to uniform sampling (PPI). -
Persona prompting is found to be unreliable, sometimes shifting predictions
systematically in the wrong direction,
and often fails precisely for the groups it is meant to represent.
Limitations and Ethical Considerations
The researchers note several limitations:
-
The stratification by gender, age, and education are independent but correlate in practice; modeling intersectional subgroups would require larger budgets.
-
The error predictor relies only on demographic features, potentially leaving performance on the table if richer textual signals were included.
-
Adaptive upsampling concentrates effort where LLMs perform worst, which often means marginalized groups, raising concerns about the labor borne by these annotators and requiring explicit compensation tradeoffs for participants.
-
The framework is a
proof of concept,
and its results should be interpreted in light of how demographic attributes were collected, as they arecoarse proxies for perspective.
Conclusion
PERSPECTIVEDRIVEN INFERENCE demonstrates that correcting LLM annotation error can be important when the target estimand is a distribution over perspectives rather than a single ground truth. It shows that PDI can be viewed as a targeted subgroup-improvement method (rather than a universally superior alternative to uniform PPI)
by directing human effort where LLM performance disparities are large and groups are skewed, while maintaining validity and coverage. This approach is task-agnostic and applicable to settings like RLHF or LLM-as-judge evaluation.
Improvements for AI systems
Here are specific, high-impact improvements to AI systems based on the PERSPECTIVEDRIVEN INFERENCE (PDI) framework described in the paper:
-
Improve Subjective Task Annotation Reliability for Demographic Subgroups
-
System Capability: Targeted Annotation Budget Allocation
-
The improved system can dynamically allocate a limited human annotation budget across different demographic groups based on the predicted reliability of existing Large Language Model (LLM) annotations for that specific group.
-
Mechanism: Perspective-Driven Inference (PDI)
-
The system utilizes an adaptive sampling strategy. It trains a black-box error predictor (e.g., using XGBoost) on annotator demographic features to estimate the squared LLM error, specifically targeting groups where the LLM proxy is least accurate.
-
System Capability: Precision Improvement for Hard-to-Model Groups
-
The system concentrates expensive human annotation effort on underrepresented or poorly modeled perspectives (e.g., older age strata like Age 50+ in politeness ratings, as shown in the results), leading to significantly improved accuracy (lower Delta) for those groups compared to uniform sampling baselines.
-
System Capability: Valid Downstream Inference
-
The system produces group-level estimates of the quantity of interest (e.g., mean politeness rating per demographic group, rather than a single aggregate), allowing for statistically valid inference even when no single
ground truth
exists. -
System Capability: Robust Confidence Intervals
-
The system provides provably valid confidence intervals for these group-level estimates using Inverse Probability Weighting (IPW) correction and bootstrap methods, ensuring the reported uncertainty reflects the true distribution of perspectives.
-
System Capability: Task Agnosticism
-
The framework is task-agnostic and can be applied to any subjective annotation task where annotator identity shapes labels (e.g., politeness, offensiveness, theory of mind modeling), provided a demographic axis for stratification can be defined.
-
System Capability: Detecting LLM Bias vs. Sampling Variance
-
The framework explicitly distinguishes between error driven by poor model performance (which PDI corrects) and variance caused by insufficient human data, allowing researchers to diagnose the root cause of annotation inaccuracy.
-
System Capability: Enhanced Model Selection Guidance
-
By analyzing the performance across multiple LLM variants (zero-shot vs. few-shot vs. persona prompting), the system can guide developers toward prompt engineering strategies (e.g., favoring few-shot examples over persona conditioning for certain tasks) that yield better initial proxies.
-
System Capability: Identifying Systemic Disagreement
-
The framework, when combined with disagreement analysis tools (like
Fightin' Words
or LLooM), can identify high-level thematic disagreements between demographic groups, revealing exactly what aspects of text cause annotator divergence across perspectives.
Abstract
Large language models are increasingly used to annotate texts, but their outputs reflect some human perspectives better than others. Existing methods for correcting LLM annotation error assume a single ground truth. However, this assumption fails in subjective tasks where disagreement across demographic groups is meaningful. Here we introduce Perspective-Driven Inference, a method that treats the distribution of annotations across groups as the quantity of interest, and estimates it using a small human annotation budget. We contribute an adaptive sampling strategy that concentrates human annotation effort on groups where LLM proxies are least accurate. We evaluate on politeness and offensiveness rating tasks, showing targeted improvements for harder-to-model demographic groups relative to uniform sampling baselines, while maintaining coverage.
Sources
- PPI++: Efficient Prediction-Powered Inference
- SoMi-ToM: Evaluating Multi-Perspective Theory of Mind in Embodied Social Interactions
- This human study did not involve human subjects: Validating LLM simulations as behavioral evidence
- Valid Survey Simulations with Limited Human Data: The Roles of Prompting, Fine-Tuning, and Rectification
- What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data
- Modeling Annotator Disagreement with Demographic-Aware Experts and Synthetic Perspectives
- Beyond Consensus: Perspectivist Modeling and Evaluation of Annotator Disagreement in NLP
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering