Verifiable abstention makes AI leak diagnosis accountable in urban water distribution networks

arXiv:2608.18836 · cs.AI · Submitted 2026-08-19 · Read on arXiv

cs.AI

Submitted: 2026-08-19

Updated: 2026-09-01

Comments: 45 pages, 5 main figures, 1 main table, 5 supplementary figures, 15 supplementary tables. Code and data availability described in the paper

License: http://creativecommons.org/licenses/by/4.0/

The gist: Leak localization is usually evaluated as forced-choice prediction, although sparse hydraulic observations may not justify excavation.

Terminology

Abstract

Leak localization is usually evaluated as forced-choice prediction, although sparse hydraulic observations may not justify excavation. Here, we quantify a pressure-information limit and use it to recast localization as selective, evidence-gated decision-making. A physics-grounded executor falsifies competing leak, demand, sensor and valve hypotheses in a hydraulic twin. Deterministic code computes every number and every acceptance predicate; an independent large language model auditor may add a rejection but never overturn a failed check. Forced retrieval placed only 95 of 300 leaks in the correct zone. Across 550 mixed events, the gate acted on 223 (214 correct); on a third-party 33-leak benchmark, all four accepted events were correct. In a replay of 194 audited City D repairs, the pressure tier authorized five excavation recommendations, three matching the repaired district, while the district-inflow tier returned the correct district for 85 events. Observability limits with machine-checkable abstention enable auditable utility intervention.

Sources

Related papers