Uncertainty-Aware Probabilistic Constrained Clustering from Entangled Pairwise Supervision
Shaojie Zhang, Ke Chen
Department of Computer Science, The University of Manchester
cs.LG, cs.AI
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: 40 pages, including appendices
Code: https://github.com/zalandoresearch/fashion-mnist
Project page: https://www.cs.toronto.edu/~kriz/cifar.html
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 95/100
Terminology
Summary
The paper addresses a realistic setting in pairwise constrained clustering where supervision is not idealized hard binary labels but rather real-valued probabilistic relations that entangle multiple sources of uncertainty. The authors state: "Pairwise constrained clustering typically relies on hard must-link/cannot-link labels, whereas realistic pairwise supervision may be real-valued and entangle intrinsic ambiguity, expert judgment, and stochastic corruption." They formalize this as uncertainty-aware probabilistic constrained clustering (UPCC).
The motivation comes from practical scenarios: when comparing two clinical cases, a recorded observation may reflect ambiguity in the cases themselves, systematic clinician judgment bias, and documentation or data-entry errors.
Existing deep constrained clustering (DCC) methods mainly target hard, expert-agnostic constraints, treating soft labels mostly numerically rather than semantically.
The paper specifies an observation model distinguishing three factors:
-
Aleatoric relation: The canonical target defined as
the aleatoric co-membership relation R⋆ab:= Pr[xa and xb belong to the same cluster under S X] ∈ [0, 1]
where S is a latent random partition. -
Epistemic judgment: Expert-conditioned judgment built on the aleatoric relation, modeled via a logistic-normal density:
pjud(y R⋆ab, e, ϕab) = (1/(y(1−y))) φ(logit(y); logitε(R⋆ab) + me(ϕab), τ2e(ϕab))
where me(ϕab) is a structured mean effect and τe(ϕab) is a centered residual perturbation. -
Stochastic corruption:
ye,ab ce,ab = 0 ∼ pjud(· R⋆ab, e, ϕab), ye,ab ce,ab = 1 ∼ Uniform(0, 1)
with corruption probability πe.
The paper establishes conditional identifiability of the canonical aleatoric relation through Theorem 3.5, which states that if the induced conditional laws of the observed response coincide under Ξ and Ξ′ for every observable triple (a, b, e), then rθ(xa, xb) = rθ′(xa, xb) for all (a, b) ∈ Pobs.
This requires Assumption 3.1 (centering on observable support) and Assumption 3.4 (structural separability).
The paper introduces ProbPair, a deep constraint embedding approach that links probabilistic pairwise relations to angular representation learning. The key components are:
-
Probabilistic readout:
ŷaibi:= σ((cos(zai, zbi) − m)/T) ∈ (0, 1)
where m and T are learnable parameters. -
ProbPair loss:
LPP = −(1/C) Σi [yi log ŷaibi + (1−yi) log(1−ŷaibi)]
-
Reconstruction regularizer:
Lrec = (1/X) Σⱼ ‖xj − x̂j‖22
with the overall objectiveL = LPP + λrec Lrec
Unlike hard-constraint objectives, Eq. (7) directly accommodates non-binary relation strengths. As a result, the induced geometry need not collapse to a purely binary regime, but can preserve both confident and uncertain pairwise structure.
The paper develops the Estimator–Corrector–Integrator ProbPair (ECI-PP) framework with three roles:
Estimator: To enable the subsequent learning of structured observational effects from entangled supervision, ECI-PP first constructs a model-side reference that is less coupled to the raw observations.
It uses K-fold cross-fitting: We partition C into K disjoint folds C(k) Kk=1 and instantiate K non-shared Estimators
where each is trained on C C(k) and produces out-of-fold beliefs. The initialization uses an information-based decisiveness score
κi based on KL divergence.
Corrector: Given the Estimators' out-of-fold beliefs, the Corrector aims to learn structured expert-specific corrections in the observed supervision.
Each expert-specific Corrector is parameterized by a network qνe on Ie:= i: ei = e with input [ŷoofi; ϕi]
and trained with the loss: L(e)cor = (1/Ie) Σi∈Ie [ρ(∆i − ∆̂i) + λcor∆̂i2]
where ρ is the Huber loss.
Integrator: The Integrator, parameterized as Int = (fω, gω′, m, T), is the learner that integrates the refined supervision into a clustering-oriented embedding.
It uses corrected relations yicor, screening weights wi:= (1 − gapi)γ, and Bayesian-confidence supervision yiBC:= (yi + ni yicor)/(1 + ni)
with ni:= n0 wi.
The framework operates iteratively: "ECI-PP operates by passing supervision through the Estimator–Corrector–Integrator pipeline to produce, via Eq. (6), Integrator-side relations yiint, and then feeding (yiint, wi) back to the Estimators for iterative refinement."
Main comparison: Under lv0.01 single-expert and multi3 multi-expert settings with corruption probability 0.3 and 9k constraints, ECI-PP is strongest, ranking first in 41 out of 48 test entries (2 expert regimes×8 datasets×3 metrics) and within the top two in 46/48.
Under single-expert supervision, ECI-PP outperforms the strong non-ProbPair baseline, SpherePair, in most entries (4–10% absolute NMI margins).
Under multi-expert supervision, the NMI margin over SpherePair widens to 6–32%.
Robustness: Across varying expert quality, corruption probability, and multi-expert configurations, "the stability advantage of ECI-PP is broadly visible across datasets, particularly on FMNIST, ImageNet10, and STL10, and is most pronounced under increasing corruption, where most baselines degrade sharply while ECI-PP remains comparatively stable."
Held-out diagnostics: The paper examines internal signals on sample-disjoint held-out constraints, showing that relation Brier scores decrease, the post-correction residual discrepancy drops, and the screening signal aligns with corruption above chance.
Sensitivity and ablation: "The default setting is robust over broad ranges: larger K gives only mild gains relative to its cross-fitting cost, information-based warm-up is mildly helpful, small Correctors with moderate regularization are sufficient, and the screening and boundary parameters (γ, n0, ξ) are generally insensitive."
The paper's main contributions are: "(i) We formalize UPCC through a heterogeneous observation model and analyze identifiability of the canonical aleatoric target on observable pairs. (ii) We introduce ProbPair, a deep constraint embedding approach linking probabilistic pairwise relations to angular embeddings. (iii) We develop ECI-PP, a practical framework for aleatoric relation learning under entangled supervision. (iv) We empirically show across diverse benchmarks that ECI-PP outperforms state-of-the-art DCC methods and is robust under fallible, heterogeneous, and corrupted supervision."
The paper acknowledges several qualifications: "The identifiability analysis is population-level and conditional; fully factor-specific finite-sample estimation would require stronger assumptions or additional observations beyond the current surrogate residual structure. Accordingly, the Corrector and reliability signal are practical surrogates rather than estimators of individual latent factors. Additionally,
our experiments use controlled expert simulations, leaving real-world annotator-provided constraints as a natural next step."
Improvements for AI systems
Improvement 1: Uncertainty-Aware Constraint Integration for Clustering Systems
-
What to implement: Modify deep clustering algorithms to accept probabilistic pairwise relations (e.g., 0.7 must-link) instead of binary labels, using a ProbPair-style loss that maps relation strength to angular embedding similarity via a learnable temperature and margin.
-
Capability gained: The AI system can learn from noisy, human-judged similarities (e.g., medical case comparisons) without discarding ambiguous or low-confidence annotations, preserving graded structure that binary constraints would collapse. This yields more accurate cluster boundaries when supervision is inherently uncertain.
Improvement 2: Entangled Supervision Decomposition via Estimator–Corrector–Integrator Pipeline
-
What to implement: Add a three-stage module to any representation learner: (a) an Estimator using cross-fitted out-of-fold predictions to build a reference belief; (b) a Corrector network that learns expert-specific biases and noise from residuals; (c) an Integrator that fuses raw and corrected beliefs with confidence-weighted screening (e.g.,
wi = (1 - gap) γ) to train the final embedding. -
Capability gained: The system can separate aleatoric (inherent data ambiguity) from epistemic (expert bias) and stochastic (corruption) sources in supervision, enabling robust learning even when multiple annotators disagree or when labels are randomly flipped. It can self-correct its own training signal, improving performance under heterogeneous, fallible human input.
Improvement 3: Identifiability-Guided Model Design for Latent Relation Recovery
-
What to implement: Enforce the paper’s identifiability conditions (centering on observable support, structural separability) in the loss function by regularizing the model’s predicted relation function to be invariant to expert-specific perturbations and corruption, e.g., adding a penalty that minimizes the discrepancy between predicted and corrected relations across experts.
-
Capability gained: The AI system can recover the true underlying co-membership probability (not just a proxy) from entangled supervision, making its clustering decisions more interpretable and reliable in high-stakes domains like clinical diagnosis or fraud detection, where knowing why two items are grouped is as important as the grouping itself.
Improvement 4: Corruption-Resilient Active Learning and Data Screening
-
What to implement: Use the paper’s screening signal (based on the gap between raw and corrected beliefs) to automatically filter or down-weight corrupted supervision instances during training, and to prioritize new data collection from experts where the gap is largest.
-
Capability gained: The system becomes robust to adversarial or accidental label corruption (e.g., up to 30% random noise) without manual data cleaning. It can also actively query annotators for the most informative, least-certain relations, reducing annotation cost while maintaining clustering quality.
Improvement 5: Cross-Fitted Uncertainty Estimation for Semi-Supervised Learning
-
What to implement: Integrate K-fold cross-fitting of belief estimators (as in ECI-PP) into any semi-supervised or self-training pipeline, using out-of-fold predictions to compute per-sample uncertainty and to weight pseudo-labels during iterative refinement.
-
Capability gained: The AI system can avoid overfitting to its own noisy predictions during self-training, leading to more stable and accurate performance on unlabeled data, especially when the labeled set is small or contains conflicting annotations. This is directly applicable to image, text, and tabular clustering tasks.
Improvement 6: Angular Embedding with Probabilistic Readout for Metric Learning
-
What to implement: Replace standard cosine-similarity-based metric learning losses with the ProbPair formulation:
ŷ = σ((cos(z a, z b) - m)/T), trained with binary cross-entropy on probabilistic targets, plus a reconstruction regularizer to preserve input structure. -
Capability gained: The system produces embeddings where similarity scores are calibrated probabilities (not just arbitrary distances), enabling downstream tasks like nearest-neighbor retrieval, outlier detection, or hierarchical clustering to make probabilistic statements (e.g., “85% confidence these two images belong to the same class”), which is useful for human-in-the-loop decision support.
Abstract
Pairwise constrained clustering typically relies on hard must-link/cannot-link labels, whereas realistic pairwise supervision may be real-valued and entangle intrinsic ambiguity, expert judgment, and stochastic corruption. Existing deep constrained clustering (DCC) methods mainly target hard, expert-agnostic constraints, treating soft labels mostly numerically rather than semantically. We formalize this setting as uncertainty-aware probabilistic constrained clustering (UPCC), defining a canonical aleatoric target through a heterogeneous observation process and analyzing its conditional identifiability. We introduce ProbPair, an angular pairwise objective for probabilistic relations, and build ECI-PP, an estimator--corrector--integrator framework that refines imperfect supervision via belief estimation, correction, and reliability-aware integration. Across challenging probabilistic supervision settings, experiments on diverse benchmarks show that ECI-PP outperforms state-of-the-art DCC methods and remains robust with a shared default configuration.
Sources
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks