Rethinking Vacuity for OOD Detection in Evidential Deep Learning
cs.AI
Submitted: 2026-05-07
Updated: 2026-08-28
Code: https://github.com/mcnamacl/Vacuity_Analysis
License: http://creativecommons.org/licenses/by/4.0/
The gist: Vacuity, or Uncertainty Mass (UM), is commonly used as a metric to evaluate Out-of-Distribution (OOD) detection in Evidential Deep Learning (EDL).
Terminology
Abstract
Vacuity, or Uncertainty Mass (UM), is commonly used as a metric to evaluate Out-of-Distribution (OOD) detection in Evidential Deep Learning (EDL). It generally involves dividing the number of classes (K) by the total strength of belief (S) of the model's predictions, where S is derived from summing the Dirichlet parameters. As such, UM is sensitive to the cardinality of K. As a result, when comparing In Distribution (ID) and OOD results, it is important that K ID and K OOD are equal; something that is not always ensured in practice. We provide an empirical demonstration of how results for AUROC and AUPR can substantially differ when class cardinality between ID and OOD differs by 1, with AUROC differing by as much as 0.346 and AUPR by 0.634 for standard EDL, and AUROC by 0.427 and AUPR by 0.745 for IB-EDL (both from Implementation B, Llama3-8B, ARC-E). Our findings isolate an evaluation artefact: when K differs between ID and OOD, AUROC/AUPR can be artificially inflated without any change in model predictions. We further discuss the evaluation of EDL over causal language models using Multiple-Choice Question-Answer (MCQA) datasets and argue for clearer definitions of ID and OOD in this context. Our primary contribution is an empirical and theoretical demonstration that vacuity-based OOD detection in EDL-fine-tuned LLMs is highly sensitive to uncontrolled differences in evaluated class cardinality.
Sources
- Evidential Deep Learning to Quantify Classification Uncertainty
- Calibrating LLMs with Information-Theoretic Evidential Deep Learning
- Evidential Transformation Network: Turning Pretrained Models into Evidential Models for Post-hoc Uncertainty Estimation
- LoRA: Low-Rank Adaptation of Large Language Models
- Revisiting Essential and Nonessential Settings of Evidential Deep Learning
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Uncertainty Estimation by Fisher Information-based Evidential Deep Learning
- Contrastive Training for Improved Out-of-Distribution Detection
- Character-level Convolutional Networks for Text Classification
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection