Dual-Primal Graph VAEs for Noisy Label Aggregation
Patrick Stinson, Nikolaus Kriegeskorte
Zuckerman Institute · Columbia University
cs.LG
Submitted: 2026-08-11
Updated: 2026-08-13
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 77/100
The gist: Dual-Primal Graph VAEs for Noisy Label Aggregation proposes a graph VAE architecture for inferring ground-truth labels from noisy crowdsourced annotations.
Terminology
Summary
Dual-Primal Graph VAEs for Noisy Label Aggregation proposes a graph VAE architecture for inferring ground-truth labels from noisy crowdsourced annotations. The model treats ground-truth labels as latent variables and uses a GAT-based decoder on the crowdsourcing graph (primal) and a GAT-based encoder on its dual. The paper states: "We propose overcoming the shortcomings of current latent variable models and representation-based models by integrating representation learning in a graph autoencoder architecture. Specifically, we use a graph attention network (GAT) to specify a generative model (decoder) for the edges given task latent encodings. The encoder model is a GAT on the dual of the graph in which edge representations are transformed to latent encoding distributions."
The generative model outputs edge distributions given task latent variables (ground-truth labels and task-specific parameters), with worker and task embeddings updated via multi-head attention message passing. The inference model transforms edge embeddings (initialized from observed labels) into distributions over ground-truth labels and task parameters, using separate attention mechanisms for task-linked and worker-linked neighbors. Optimization uses the ELBO with a tempered majority vote prior for labels and a standard normal prior for task parameters, with Reinmax for discrete gradient estimation.
Key results on crowdsourcing benchmarks (CF, face, dog, senti, prod, adult, RTE, LabelMe) show the model achieves the best average rank (2.0) compared to MV, iBCC, EBCC, GOVERN, LAA variants, and CrowdFM. For example, on CF it achieves.898 ±.002 vs.880 for MV, and on face.658 ±.003 vs.637 for MV. Ablations show the GAT architecture is crucial (C=0 degrades performance most), while resampling, dropout, and non-label latent variables (Dr > 0) all contribute positively.
The paper also demonstrates generalization by augmenting the crowdsourcing graph with DNN representations trained on noisy labels (e.g., IDNT, TAIDTM). On simulated MNIST and CIFAR10 labelings with low/mid/high noise, and real datasets CIFAR10-N and LabelMe, combining noisy labels with DNN representations (x̂) substantially boosts accuracy over either source alone. For instance, on CIFAR10 high noise, DPGVAE + x̂IDNT achieves.987 ±.002 vs.979 ±.002 for IDNT alone, and on LabelMe, DPGVAE + x̂TAIDTM achieves.962 ±.002 vs.909 ±.002 for TAIDTM alone. The paper notes: "by using the augmented graph and combining the information in the noisy labelings with the image information encoded by the DNNs (referred to as x̂), we achieve a much higher performance over just using the trained DNN... or just using our method with the noisy labelings."
The paper also addresses posterior collapse (Proposition 1 gives a condition on embedding dimensionality) and observes automatic disentanglement of label and task latents, with mutual information traces showing transient entanglement that resolves during training.
Improvements for AI systems
Improvements to AI Systems:
-
Robust Label Aggregation in Weakly Supervised Learning: Integrate the dual-primal GAT VAE as a front-end for any supervised model trained on noisy labels. The system can infer clean ground-truth labels from multiple annotators, then train downstream classifiers on these denoised labels, improving accuracy in settings where expert labels are unavailable (e.g., medical imaging, social media moderation).
-
Unified Multimodal Crowdsourcing with Learned Representations: Extend the augmented graph approach to fuse heterogeneous data sources—e.g., raw text, images, or audio embeddings from pretrained DNNs—with noisy human labels. The improved system can jointly infer labels and refine task embeddings, enabling semi-supervised learning where only a few clean labels exist but many noisy annotations and unlabeled raw data are available.
-
Adaptive Active Learning for Annotation Efficiency: Use the model’s uncertainty over ground-truth labels (from the posterior distribution) to select the most informative tasks for additional annotation. The system can query workers only on tasks where label entropy is high, reducing annotation cost while maintaining accuracy, especially in low-budget crowdsourcing campaigns.
-
Disentangled Latent Factor Discovery in Complex Datasets: Leverage the automatic disentanglement of label and task-specific latents to separate semantic content (e.g., object class) from nuisance factors (e.g., style, viewpoint). This enables controllable generation, domain adaptation, and transfer learning where the label latent is reused across tasks while task latents are discarded or fine-tuned.
-
Posterior-Collapse-Resistant Variational Training: Apply the condition from Proposition 1 (on embedding dimensionality) to other VAE-based systems. The improved AI can avoid degenerate latent representations in text, graph, or image generation, maintaining diverse and informative latents even with strong decoders, improving sample quality and interpretability.
-
Graph-Structured Noise Correction for Social and Biological Networks: Generalize the dual-primal architecture to any relational data with noisy node attributes or edge labels (e.g., protein interaction networks, social trust networks). The system can denoise edge labels and infer latent node properties, enabling more reliable link prediction and community detection.
-
Reinmax-Based Discrete Optimization in Hybrid Models: Replace Gumbel-Softmax or straight-through estimators with Reinmax for any discrete latent variable model (e.g., discrete VAE, mixture-of-experts routing). The improved system achieves lower-variance gradient estimates, faster convergence, and better final performance on tasks requiring discrete decisions (e.g., neural architecture search, discrete codebook learning).
Abstract
Inferring the ground-truth from noisy crowdsourced labels is an important theoretical and practical problem. Neural network-based methods offer an alternative to classical Bayesian models which require specifying a family of generative models used for inference. However, current models either still rely on fairly simple generative models for inference or require pseudo-labels or synthetic data to train the aggregate classifier. We propose a graph VAE architecture in which the decoder and encoder use GAT-based message passing on the adjacency graph of a crowdsourced dataset and its dual, respectively. The ground-truth labels are treated as latent variables, enabling unsupervised representation learning without needing to train a separate classifier. We show our model achieves state of the art performance on crowdsourcing benchmarks. We then demonstrate the generality of our approach by showing how the original crowdsourcing graph can be augmented to incorporate side information such as representations from neural network classifiers trained on the noisy labels to substantially boost their classification performance at test time.
Sources
- Deep Convolutional Networks on Graph-Structured Data
- Subsampling Graphs with GNN Performance Guarantees
- Auto-Encoding Variational Bayes
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks