Verifiable abstention makes AI leak diagnosis accountable in urban water distribution networks
cs.AI
Submitted: 2026-08-19
Updated: 2026-09-01
Comments: 45 pages, 5 main figures, 1 main table, 5 supplementary figures, 15 supplementary tables. Code and data availability described in the paper
License: http://creativecommons.org/licenses/by/4.0/
The gist: Leak localization is usually evaluated as forced-choice prediction, although sparse hydraulic observations may not justify excavation.
Terminology
Abstract
Leak localization is usually evaluated as forced-choice prediction, although sparse hydraulic observations may not justify excavation. Here, we quantify a pressure-information limit and use it to recast localization as selective, evidence-gated decision-making. A physics-grounded executor falsifies competing leak, demand, sensor and valve hypotheses in a hydraulic twin. Deterministic code computes every number and every acceptance predicate; an independent large language model auditor may add a rejection but never overturn a failed check. Forced retrieval placed only 95 of 300 leaks in the correct zone. Across 550 mixed events, the gate acted on 223 (214 correct); on a third-party 33-leak benchmark, all four accepted events were correct. In a replay of 194 audited City D repairs, the pressure tier authorized five excavation recommendations, three matching the repaired district, while the district-inflow tier returned the correct district for 85 events. Observability limits with machine-checkable abstention enable auditable utility intervention.
Sources
- Overcoming Common Flaws in the Evaluation of Selective Classification Systems
- Are Uncertainty Quantification Capabilities of Evidential Deep Learning a Mirage?
- Verifiably Robust Conformal Prediction
- Controlling Counterfactual Harm in Decision Support Systems Based on Prediction Sets
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Debating with More Persuasive LLMs Leads to More Truthful Answers
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection