How Much Human Label Variation Does Formal Semantic Structure Explain? Group-Level Effects and Item-Level Ceilings in NLI

arXiv:2607.15870 · cs.CL · Submitted 2026-07-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "How Much Human Label Variation Does Formal Semantic Structure Explain?".

Tom: Human label variation in Natural Language Inference (NLI) is increasingly viewed as signal rather than noise, prompting researchers to investigate what structures within language—specifically formal semantics—account for this observed disagreement.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: Alright, let's get into the specifics of who wrote this and what exactly they are looking at in "How Much Human Label Variation Does Formal Semantic Structure Explain?: Group-Level Effects and Item-Level Ceilings in NLI." The title itself points directly to the two main things they measured: group-level effects and item-level ceilings.

Jane: That’s right, Tom. The authors are looking at how much formal semantic structure explains human label variation across a massive dataset called ChaosNLI, which consists of three thousand one hundred thirteen items from SNLI and MNLI that were specifically chosen because they had low original agreement—exactly three out of five majorities.

Lu: That sample selection is important because it means the effects they see are "within-sample associations that may be attenuated relative to unrestricted NLI data," which adds a layer of necessary caution to their findings.

Meng: Attenuated effects mean we need to be careful about how much weight we put on these results when applying them to larger, less controlled datasets, like the general internet data we use for pre-training.

Lalam: That's a good point; it tells us that while the structure matters here in this specific context, we can't assume the same level of influence when we generalize those findings widely.

The paper's summary: Tom: Now for the summary of what they found. Basically, they established three key bounds regarding this formal structure. They found a robust group-level boundary where hypotheses whose monotonicity profile is not purely upward show reliably higher label entropy, with Cliff’s delta being −zero point two eight four.

Jane: So, in simpler terms, Tom, they discovered that if the way a hypothesis behaves structurally isn't strictly increasing in a certain way—for instance, not purely upward monotone—then human disagreement on that item is reliably higher than we would expect from random chance.

Lu: That group-level boundary survives tests against operator presence reduction and length reductions, but they also found that the majority margin mirrors this entropy, showing upward hypotheses have higher margins than downward ones with a delta of +zero point two zero one.

Meng: So, the structure itself has a measurable impact on disagreement within a group of related items, which is significant because it’s not just about whether an item is hard or easy to classify overall.

Lalam: It suggests that we can start looking for these structural patterns in our data to predict areas where human judgment might be more volatile, which could help us target specific types of reasoning tasks.

The paper's improvements: Tom: Moving on to the second major finding, they established an item-level ceiling. They found that these formal profiles only explain about three point three to three point six percent of the entropy variance and reach a median-split AUC of zero point six zero six for a logistic classifier.

Jane: That means, Tom, even though we found that structural patterns exist at the group level, using those profiles alone isn't enough to predict whether a specific individual item will attract high disagreement; it’s too weak for that task.

Lu: It suggests that operator and monotonicity profiles don't identify which items will attract high disagreement at all, which is a crucial limitation if we want to use these features for direct item-level prediction.

Meng: That tells me we shouldn't rely on these features alone if our goal is to build a system that can predict the most difficult items in a set; it points toward needing other, more complex features.

Lalam: It sets a realistic expectation for us; we can't just assume the formal structure will perfectly tell us which item is going to cause trouble, so we need to keep that in mind when designing prediction pipelines.

Conclusion: Tom: So, wrapping things up on "How Much Human Label Variation Does Formal Semantic Structure Explain?: Group-Level Effects and Item-Level Ceilings in NLI," the main conclusion is that there’s a group-level boundary where hypotheses not being purely upward monotone show higher label entropy, but this structural information doesn't change what annotators disagree about compositionally.

Jane: Exactly. They checked the composition of disagreement across this boundary using three different contrasts—error share, explanation-type shares, and pragmatic/world knowledge shares—and they all returned null results, meaning the mixture of error types doesn't differ significantly on either side of that formal divide.

Lu: It reinforces that the structure modifies how much disagreement there is in a group but doesn't fundamentally alter what kind of disagreement is happening across that boundary.

Meng: So, the implication for practical engineering is that formal structure has a measured, bounded weight in explaining disagreement, but it’s not the sole factor driving uncertainty. We need to focus on other item properties if we're trying to build robust classifiers based on this work.

Lalam: For our culture here at the startup, this means we can use these findings to design more nuanced error detection and explanation generation systems that recognize when disagreement is driven by structural shifts versus just random noise or annotation variability.

Tom: Well said, Lalam. It sounds like the big picture is that formal semantics plays a measured role in explaining some of the human variation, but it doesn't capture the entire story of why people disagree on inference tasks. We'll keep an eye on future work following this study.

University of Bremen

cs.CL

Submitted: 2026-07-17

Updated: 2026-10-08

Code: https://github.com/oudeis01/nli-hlv-structure

Importance score: 83/100

The gist: Human label variation in Natural Language Inference (NLI) is increasingly viewed as signal rather than noise, prompting researchers to investigate what structures within language—specifically

Key concepts

Group-Level Boundary
This is a structural division within the dataset based on how well a hypothesis's semantic profile (specifically its monotonicity) aligns with expected patterns. Hypotheses that are not strictly increasing in this profile are reliably associated with higher disagreement levels among human annotators.
Item-Level Ceiling
This concept defines the limit of predictive power for formal structure when trying to predict high disagreement for a single sentence. The study found that semantic features alone only explain a small fraction (3.3 to 3.6%) of the variation in disagreement, meaning they are insufficient for predicting which specific items will cause high disagreement.
Composition Invariance
This refers to whether the *type* of disagreement—for instance, whether it's due to error share or pragmatic inference—changes as you cross the identified group-level boundary. The study found that this composition does not significantly differ between the two sides of this boundary, suggesting formal structure affects disagreement in a consistent way.
Monotonicity Profile
This is a measurement derived from dependency parses that describes how consistently certain linguistic features (like quantifiers or negation) change across different hypotheses. A 'not purely upward' profile serves as the primary indicator for identifying the group-level boundary where higher disagreement occurs.

Terminology

Summary

Human label variation in Natural Language Inference (NLI) is increasingly viewed as signal rather than noise, prompting researchers to investigate what structures within language—specifically formal semantics—account for this observed disagreement. This paper directly addresses this question by measuring how much formal semantic structure explains human label variation across a large-scale dataset, ChaosNLI. The findings establish three key bounds: a robust group-level association, an item-level ceiling that is too weak to predict high disagreement, and the invariance of disagreement composition across the identified boundary.

Data and Measurement

The study utilizes the 3,113 SNLI and MNLI items from ChaosNLI (ChaosNLI-S/M), which were intentionally selected for low original agreement (exactly 3-out-of-5 majorities). This range restriction means the effects observed are within-sample associations that may be attenuated relative to unrestricted NLI data. The core measurement involves a rulebased operator tagger applied over dependency parses, which assigns profiles for four operator families: syntactic negation, quantifiers, negative polarity items with a licensing flag, and monotonicity triggers. This tagger was validated against the MED benchmark (0.883 agreement at the edit site) and yielded a sentence-level summary agreement of 0.807.

Group-Level Boundary (Contribution 1)

The primary finding is the existence of a group-level boundary. Hypotheses whose monotonicity profile is not purely upward show reliably higher label entropy (Cliff’s δ = −0.284). This effect survives tests against operator presence reduction, length reductions, and parse depth confounds. The analysis also found that the majority margin mirrors this entropy: upward hypotheses have higher margins than downward (δ = +0.201). However, the boundary is not reducible to simple features; regressions on operator counts and complexity covariates showed only modest effects (e.g., a coefficient of 0.036 bits for marked hypothesis in M0).

Item-Level Ceiling (Contribution 2)

The paper establishes an item-level ceiling, demonstrating that formal structure alone is insufficient to predict high disagreement at the item level. The four-class partition explains only 3.3 to 3.6 percent of entropy variance and reaches a median-split AUC of 0.606 for a logistic classifier. This suggests that operator and monotonicity profiles do not identify which items will attract high disagreement, and applications requiring item-level predictions should not rely on these features alone due to this low ceiling.

Composition Invariance (Contribution 3)

The study investigates whether the composition of disagreement differs across the group-level boundary using a 498-item overlap sample. Three high-powered contrasts—C1 (error share), C2 (explanation-type shares), and C3 (pragmatic/world knowledge shares)—all returned null results. Specifically, for error share, the contrast yielded δ = −0.076, and for explanation categories, δ = −0.029. This indicates that the composition of disagreement does not detectably differ across the formal boundary, as the same mixture of error, logical-structure conflict, and pragmatic inference appears on both sides.

Robustness and Limitations

The analysis was rigorously controlled through five preregistered checks, including a zero-inflation check (C-R1) which confirmed that composition invariance was not an artifact of zero inflation. The study is conditioned by several limitations: the sample is range-restricted toward high disagreement, task specificity is limited to NLI over English sentences, and the bounded outcome variable (entropy) restricts OLS coefficients to reporting direction and relative size only. Crucially, the paper explicitly states it makes no causal claims, treating all results as observed associations. The final conclusion frames formal structure as having a measured, bounded weight in explaining disagreement, leaving the larger share of variation to other item properties and annotator-side variables.

Conclusion

The research concludes that while a robust group-level boundary exists—hypotheses that are not purely upward monotone show reliably higher label entropy—this structure does not change what annotators disagree about compositionally. Formal semantic structure contributes to disagreement by a small amount, but its predictive power at the item level is low, suggesting it belongs in the inventory of disagreement sources with a measured weight. The paper suggests that future questions regarding unrestricted samples or lexical versus compositional tagging remain open empirical avenues.


Key Enumerations from the Paper:

  1. Three bounds emerge: a group-level boundary, an item-level ceiling, and composition invariance across the boundary.

  2. The group-level boundary is defined by hypotheses whose monotonicity profile is not purely upward monotone, showing higher label entropy (Cliff’s δ = −0.284).

  3. The item-level ceiling explains only "3.

Improvements for AI systems

Based on the scientific paper How Much Human Label Variation Does Formal Semantic Structure Explain?, here are the specific improvements for AI systems and what those improved systems can achieve:


)Improved AI System Capabilities:

The core improvement is moving from treating all human label variation as noise to quantifying which formal semantic structures (monotonicity, negation, quantification) actually contribute to it. This allows for more targeted model development and error analysis.

  1. Confidence-Weighted Disagreement Detection (Group-Level):

A system can use the derived group-level boundary (Cliff’s δ = −0.284) to flag items exhibiting reliably higher label entropy than expected under a purely upward monotone distribution. It would specifically target hypotheses whose monotonicity profile is not purely upward monotone, indicating areas where formal structure is driving human disagreement, rather than random noise or simple operator presence.

  1. Item-Level Disagreement Prioritization (Ceiling-Aware):

A system can be trained to recognize the low item-level ceiling (R2 = 0.033, AUC = 0.606) and use this knowledge to set realistic expectations for automated prediction models. Crucially, it can identify items that are too hard or highly ambiguous for current formal semantic structure alone to resolve (i.e., items falling below the median-split AUC of 0.606), preventing the system from wasting resources trying to predict high-disagreement items based solely on surface-level formal features.

  1. Invariant Disagreement Analysis (Compositional Understanding):

The system can be designed to assess whether disagreement is shifting in kind across formal boundaries using validated error shares and explanation-type shares. If the system observes that the composition of disagreement (e.g., shifts from factual knowledge conflict to logical structure conflict) remains invariant across a formal boundary, it confirms that the structure merely changes the magnitude of disagreement, not its fundamental nature. This allows for more robust classification schemes where disagreements are categorized consistently regardless of whether they are in a monotone or non-monotone context.

  1. Robustness Against Surface Features (Defense Verification):

The system can be explicitly trained to recognize and disregard superficial features like simple operator presence, sentence length, or parse depth as weak predictors of actual high disagreement (as demonstrated by the null results in Section 4.4). This prevents the AI from over-relying on easily measurable but ultimately insignificant linguistic proxies for true inferential uncertainty.

)What improved AI Systems Can Do:

  1. Refined NLI Classification and Uncertainty Estimation:

An improved system could be used to perform Natural Language Inference (NLI) with a more nuanced understanding of human ambiguity. Instead of a single probability score, it could output an uncertainty metric that is informed by formal semantics, allowing downstream applications (like complex reasoning systems or legal document analysis) to distinguish between high disagreement due to semantic complexity and high disagreement due to annotator error/noise.

  1. Targeted Data Curation for Training:

For supervised learning tasks where the target is high-disagreement examples, the system can use these formal structure profiles as features. This allows researchers to curate training sets specifically focusing on items that exhibit high entropy driven by non-upward monotonicity, leading to more effective training data for models designed to handle complex reasoning tasks (e.g., multi-hop reasoning).

  1. Improved Error Detection and Explanation Generation:

When an NLI model predicts a low confidence score, the system can use the knowledge from Section 6 (Composition Invariance) to generate more precise explanations. If the disagreement is invariant across formal boundaries, the system knows that its failure mode is not due to misinterpreting a specific formal structure (like negation vs. quantification), but rather due to other item-specific factors or annotator variability, leading to better debugging of the model's failures.

  1. Calibrated Confidence in Prediction Thresholds:

By understanding the item-level ceiling (AUC 0.606), an AI system can be calibrated to know when a prediction is likely unreliable due to inherent semantic complexity versus genuine difficulty. This prevents the system from applying high confidence scores to items that are fundamentally hard for current formal structure-based models to resolve, leading to more trustworthy and less overconfident outputs in real-world applications.

Sources

Related papers