Unified Multimodal Uncertain Inference
cs.CV, cs.LG
Submitted: 2026-04-09
Updated: 2026-09-22
Comments: Update CI and modality training exps
Code: https://github.com/adoptedirelia/UMUI
License: http://creativecommons.org/licenses/by/4.0/
The gist: We introduce Unified Multimodal Uncertain Inference (UMUI), a multimodal inference task spanning text, audio, and video, where models must produce calibrated probability estimates of hypotheses
Terminology
Abstract
We introduce Unified Multimodal Uncertain Inference (UMUI), a multimodal inference task spanning text, audio, and video, where models must produce calibrated probability estimates of hypotheses conditioned on a premise in any modality or combination. While uncertain inference has been explored in text, extension to other modalities has been limited to single-modality binary entailment judgments, leaving no framework for fine-grained probabilistic reasoning in or across other modalities. To address this, we curate a human-annotated evaluation set with scalar probability judgments across audio, visual, and audiovisual settings, and additionally evaluate on existing text and audio benchmarks. We introduce CLUE (Calibrated Latent Uncertainty Estimation), which combines self-consistent teacher calibration and distribution-based confidence probing to produce calibrated predictions. We demonstrate that our 3B-parameter model achieves equivalent or stronger performance than zero-shot baselines up to 32B parameters across all modalities.
Sources
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models
- Qwen2-Audio Technical Report
- Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
- Language Models (Mostly) Know What They Know
- Audio Entailment: Assessing Deductive Reasoning for Audio Understanding
- MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval
- Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
- Always Tell Me The Odds: Fine-grained Conditional Probability Estimation
- LoRA ensembles for large language model fine-tuning
- WikiVideo: Article Generation from Multiple Videos
- Seeing Through the MiRAGE: Evaluating Multimodal Retrieval Augmented Generation
- Conformal Thinking: Risk Control for Reasoning on a Compute Budget
- Visual Entailment: A Novel Task for Fine-Grained Image Understanding
- Qwen2.5-Omni Technical Report
- RANKVIDEO: Reasoning Reranking for Text-to-Video Retrieval
- Qwen3 Technical Report
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models