ENCP: Episode-Normalized Conformal Prediction for Vision-and-Language Navigation
cs.LG, cs.AI, cs.RO
Submitted: 2026-09-15
Updated: 2026-09-15
Comments: 8 pages, 5 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: Uncertainty estimation for Vision-Language-Navigation (VLN) models is a critical task since it can help identify ambiguous and unreliable predictions, enabling agents to make safer navigation
Terminology
Abstract
Uncertainty estimation for Vision-Language-Navigation (VLN) models is a critical task since it can help identify ambiguous and unreliable predictions, enabling agents to make safer navigation decisions. As one of the most advanced uncertainty estimation frameworks, conformal prediction (CP) offers a promising approach for uncertainty estimation in VLN. However, given that VLN agent requires a sequence of steps, standard calibration in conformal prediction fails to provide coverage guarantee it promises over a dependent, variable-length VLN episode. To this end, we propose Episode-Normalized Conformal Prediction (ENCP), which rescales a nonconformity score by the policy's residual confidence and calibrates one maximum score per episode. Under exchangeable calibration and test episodes, this construction covers the ground truth at every step with probability at least 1 - α, while allowing dependence among steps within an episode. Across four VLN policies and three nonconformity scores on R2R and REVERIE dataset, ENCP meets all reported empirical step-coverage targets on the seen-to-unseen evaluation. These results demonstrate that ENCP can provide model-agnostic uncertainty estimates, which might be useful for determining when a VLN agent should defer to a more capable predictor, including human assistance.
Sources
- Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments
- Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models
- Concrete Problems in AI Safety
- Unsolved Problems in ML Safety
- Safe Planning in Dynamic Environments using Conformal Prediction
- History Aware Multimodal Transformer for Vision-and-Language Navigation
- Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language Navigation
- Self-Monitoring Navigation Agent via Auxiliary Progress Estimation
- To Ask or Not to Ask? Detecting Absence of Information in Vision and Language Navigation
- On Calibration of Modern Neural Networks
- Can You Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift
- Robots That Ask For Help: Uncertainty Alignment for Large Language Model Planners
- Introspective Planning: Aligning Robots' Uncertainty with Inherent Task Ambiguity
- Matterport3D: Learning from RGB-D Data in Indoor Environments
- Beyond the Nav-Graph: Vision-and-Language Navigation in Continuous Environments
- Learning to Navigate Unseen Environments: Back Translation with Environmental Dropout
- Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning
- Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles
- Selective Classification for Deep Neural Networks
- Classification with Valid and Adaptive Coverage
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks