WATCH: Adaptive Monitoring for AI Deployments via Weighted-Conformal Martingales

arXiv:2505.04608 · cs.LG, cs.AI, stat.ML · Submitted 2025-05-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "WATCH: Adaptive Monitoring for AI Deployments via Weighted-Conformal Martingales".

Jane: The paper was written by Drew Prinster, Xing Han, Anqi Liu and Suchi Saria from Johns Hopkins University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Authors: Tom: Welcome back to the show, everybody. Today we are digging into a paper that has a real mouthful of a title: “WATCH: Adaptive Monitoring for AI Deployments via Weighted-Conformal Martingales.” Jane, I’m going to need you to break that down for me, because I got lost at “weighted-conformal.”

Jane: Happy to, Tom. So, the core problem is that when you put an AI model into the real world, like in a hospital or a self-driving car, the data it sees is going to change over time. This paper is about building a “watchdog” for that AI. The watchdog’s job is to tell you, “Hey, something’s different, and you need to retrain me,” or “Don’t worry, I’ve got this.”

Tom: So it’s like a check-engine light for your neural network?

Jane: Exactly. But the clever part is in the “weighted” part. A normal watchdog might scream “Problem!” the moment the data shifts even a little bit, even if the shift is totally harmless. This paper’s method, WATCH, is smart enough to say, “This shift is fine, I can adapt to it,” and only raise the alarm when the shift is actually dangerous.

Tom: That’s a huge deal. I mean, think about a sepsis prediction model in a hospital. If a new batch of patients comes in that’s slightly younger, that’s probably fine. But if a new bacterial strain shows up that’s deadly to kids, you need to know *now*. This paper is basically giving us a way to tell those two situations apart.

Jane: Right. And the authors, Drew Prinster, Xing Han, Anqi Liu, and Suchi Saria from Johns Hopkins, they’re not just hand-waving this. They’ve built a whole new mathematical framework for it. They call them Weighted-Conformal Test Martingales, or WCTMs. It sounds scary, but it’s just a way to keep score of the evidence that something has gone wrong.

Tom: Keeping score of evidence. I like that. So, instead of just a binary “alarm on” or “alarm off,” it’s more like a running tally of how suspicious the new data is looking?

Jane: Precisely. And that running tally has some really nice mathematical guarantees. It can tell you, “The chance that I’m raising a false alarm is less than one in a hundred,” which is exactly what you need when you’re asking doctors to trust the system.

Tom: So we’ve got a smart watchdog that can adapt to small changes and scream about big ones. What’s the catch?

Jane: The catch is that it needs to be able to tell the difference between a “small” change and a “big” one. And that’s what the rest of the paper is all about. They have this whole system for diagnosing *why* the alarm went off.

Tom: Diagnosis. Okay, now you’ve really got my attention. Let’s get into the meat of how this thing actually works in the next segment.

Paper Summary: Tom: So, Jane, we’ve established that WATCH is a smart alarm system. But how does it actually decide what’s a benign shift versus a harmful one? Let’s get into the summary of the paper.

Jane: The key is that they run two different monitors at the same time. One monitor, the main one, is watching the overall performance of the AI. The other monitor, they call it the X-CTM, is only watching the input data, the “X” variables.

Tom: So, one monitor is looking at the whole picture, and the other is just looking at the inputs?

Jane: Exactly. Let’s go back to the sepsis example. The input X is things like patient age, heart rate, lab results. The output Y is the risk of sepsis. A “concept shift” is when the relationship between X and Y changes. Like, a new strain of bacteria that makes a normal heart rate much more dangerous than it used to be. The X-CTM wouldn’t see anything wrong, because the inputs look normal. But the main WCTM would see the predictions are suddenly wrong and raise the alarm.

Tom: And that tells you it’s a concept shift, not a data problem.

Jane: Right. But what if the hospital suddenly starts admitting a lot of elderly patients? That’s a shift in the input X. The X-CTM would see that and say, “Hey, the inputs are changing.” Now, if that shift is mild, the main WCTM can adapt to it. It reweights its calculations to account for the new patient population, and it doesn’t raise the alarm.

Tom: So it’s like the AI is saying, “Oh, you’re showing me older patients now? No problem, let me just recalibrate my expectations.”

Jane: Exactly. But if the shift is extreme—like, suddenly everyone is over one hundred years old, which the model has never seen—the main WCTM can’t adapt. The reweighting fails, and it raises the alarm. And because the X-CTM also raised the alarm, you know the problem is with the input data, not the underlying relationship.

Tom: So you get both the alarm and the diagnosis. That’s incredibly practical. It’s not just saying “something is wrong,” it’s saying “the input data has shifted so much that I can’t cope.”

Jane: And that’s the real contribution here. They’ve built a monitoring system that doesn’t just detect a problem, it helps you understand the nature of the problem so you can fix it. This isn’t just theoretical, either. They tested it on real-world datasets like medical expenditure surveys and bike-sharing data.

Tom: Real-world data, that’s what I like to hear. So it’s not just a math exercise. They actually showed it working. I’m curious about the improvements they claim over the old methods. Let’s get into that.

Improvements: Tom: So, Jane, we’ve got this system that can adapt and diagnose. But what was wrong with the old methods? Why do we need WATCH in the first place?

Jane: The old methods, the standard Conformal Test Martingales, or CTMs, are like a smoke detector that’s set to go off if there’s any smoke at all. They’re testing for a very strict assumption: that the data is completely exchangeable, meaning the order doesn’t matter and the distribution never changes.

Tom: So any change at all, even a harmless one, sets them off?

Jane: Exactly. In the paper, they show this happening. They create a benign shift in the data, like a shift towards younger patients, and the standard CTM raises a false alarm. It’s crying wolf. And if you have a system that cries wolf all the time, doctors and engineers are going to start ignoring it.

Tom: That’s a huge problem. If the alarm is always going off, it becomes useless.

Jane: Right. The improvement WATCH brings is that it tests a different, more realistic null hypothesis. It says, “I’m going to assume that the relationship between inputs and outputs stays the same, and that the input distribution can shift, but only in a way that I can adapt to.” So it’s not testing for *any* change, it’s testing for a *harmful* change.

Tom: So it’s a more specific test, which means fewer false alarms.

Jane: And it’s not just about avoiding false alarms. They also show that WATCH is faster at detecting the truly harmful shifts. They compared it to other methods that try to track the model’s risk directly, and WATCH detected the harmful concept shifts much faster. In some cases, it was three times faster.

Tom: Three times faster? That’s a massive difference when you’re talking about patient safety.

Jane: It is. And there’s another improvement I really like. They added a “penalty” for when the model becomes unhelpful. If the AI is so confused that it just predicts the entire range of possible outcomes, that’s not a safe state either. WATCH treats that as a problem, which makes the system more robust.

Tom: So it’s not just about being right, it’s about being useful. If the model is so uncertain it’s useless, that’s a failure mode too. That’s a really thoughtful touch. I can see why this is getting attention. Let’s bring in the rest of the team to get their take on the bigger picture.

Conclusion: Tom: Okay, let’s bring in Lu and Meng to get their final thoughts on “WATCH: Adaptive Monitoring for AI Deployments via Weighted-Conformal Martingales.”

Lu: I’m really excited about the theoretical foundation here. The idea of generalizing conformal test martingales to test any null hypothesis you can define is powerful. It’s not just for monitoring; you could use this for any kind of sequential hypothesis testing problem. It opens up a whole new toolbox.

Meng: From an engineering standpoint, the speed is what gets me. The paper shows their method is O(t) in time complexity, meaning it scales linearly with the data. The older changepoint detection methods were O(t2), which gets brutally slow with lots of data. This is something we could actually deploy in a production system without it becoming a bottleneck.

Jane: And that’s the real takeaway, isn’t it? This isn’t just a neat math trick. It’s a practical tool that could make AI deployments in high-stakes fields like healthcare and autonomous driving much safer.

Tom: Absolutely. It’s a way to give us confidence that our AI systems are still working as intended, and to tell us exactly what’s wrong when they aren’t. It’s a huge step towards responsible AI.

Lu: And the future work is just as exciting. They mention extending this to monitor foundation models and AI agents. That’s where the real-world impact is going to explode.

Meng: For now, though, it’s a solid, well-tested framework for keeping an eye on the models we have today.

Tom: Well said, everyone. We’ve covered the problem, the solution, and the improvements. We’ll be keeping an eye on this one. Thanks for listening, and we’ll see you for the next paper.

Johns Hopkins University

cs.LG, cs.AI, stat.ML

Submitted: 2025-05-07

Updated: 2026-09-11

Comments: The International Conference on Machine Learning (ICML), 2025

Code: https://github.com/aaronhan223/watch

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 72/100

Key concepts

Adaptive Monitoring
A system that monitors an AI model's performance in real-world settings. Unlike simple alarms, adaptive monitoring can distinguish between harmless data shifts and genuinely dangerous changes, allowing the system to adjust without raising false alarms.
Concept Shift
A change in the relationship between the input data (X) and the model's output (Y). For example, a new bacterial strain might make a normal heart rate much more dangerous than previously thought, changing the underlying rules of prediction.
Weighted-Conformal Martingales (WCTMs)
A new mathematical framework used by the authors. It functions as a running tally or scorekeeper of evidence to determine how suspicious new data is, providing mathematical guarantees about the chance of raising a false alarm.

Terminology

Summary

Summary

This paper introduces Weighted-Conformal Test Martingales (WCTMs), a novel generalization of standard conformal test martingales (CTMs), and proposes a practical framework called WATCH (Weighted Adaptive Testing for Changepoint Hypotheses) for the continual, post-deployment monitoring of AI/ML systems.

The paper argues that responsible AI deployment requires not only proof of system reliability but also continual monitoring to detect and address unsafe behavior. Existing methods, such as standard CTMs, are limited because they are restricted to monitoring exchangeability or IID assumptions, do not allow for online adaptation to shifts, and cannot diagnose the cause of degradation. The proposed WCTMs address these limitations by laying a theoretical foundation for online monitoring of any unexpected changepoints in the data distribution while controlling false alarms.

Theoretical Contributions:

The paper's main theoretical contribution is the introduction of weighted-conformal p-values and the construction of WCTMs from sequences of these p-values. A generalized weighted-conformal p-value is defined as:

[

p n+1 = sum i=1 n+1 i [1v i > v n+1 + u n+1 1v i = v n+1]

]

where is an arbitrary weight vector summing to one. The paper presents Theorem 3.1, which proves the independence and exact validity of online WCP p-values under a null hypothesis H 0: if the assumptions hold, then P 1, P 2, are IID uniform on [0,1]. This theorem generalizes the proof for standard CTMs by weighting permutations according to their likelihood rather than assuming exchangeability.

Based on this theorem, the paper proves Proposition 3.2, which establishes that WCTMs achieve anytime-valid false-alarm control via Ville's inequality:

[

P H 0 (t: t / 0 at least c) at most 1/c

]

The paper also presents Proposition 3.3, which shows that a Shiryaev-Roberts procedure applied to WCTMs controls the average run length (ARL) under the null hypothesis, with E H 0[tau 1] at least c.

Main Practical Testing Objective:

The paper focuses on testing the null hypothesis H 0(cs), which assumes that the conditional label distribution YX remains invariant while the marginal input distribution X may shift according to an estimated density-ratio function (x). This hypothesis is formalized as:

[

H 0(cs):=

(X i, Y i) about F YX times F X, & i in [n + t ad - 1]

(X n+t, Y n+t) about F YX times G X, & t at least t ad

]

Observing an extreme value of the weighted-conformal p-value conveys evidence against this hypothesis, indicating a concept shift in YX or an inaccurate density-ratio adaptation.

WATCH Framework Implementation:

The WATCH framework implements WCTMs that continuously adapt to mild covariate shifts while raising alarms for extreme covariate shifts or concept shifts. Key implementation details include:

  1. Online Adaptation with Dynamic Initialization: A secondary standard CTM (X-CTM) monitors only for changepoints in the marginal X distribution using a nearest-neighbor nonconformity score. When this X-CTM exceeds a pre-determined adaptation threshold at time t ad, it triggers the adaptation phase of the main WCTM, which begins estimating density-ratio weights(t)(x) online.

  2. Root-Cause Analysis: The parallel implementation of both the primary WCTM and the secondary X-CTM enables diagnosis. If both detect changepoints, the shift is diagnosed as an extreme covariate shift; if only the WCTM detects a changepoint, it is diagnosed as a concept shift in YX.

Experimental Results:

The paper conducts comprehensive experiments on real-world tabular datasets (MEPS, Bike Sharing, Superconductivity) and image datasets (MNIST-C, CIFAR-10-C).

  1. Adaptation to Benign Shifts (Section 4.1): On tabular data with benign covariate shifts, WCTMs avoid unnecessary alarms across all datasets and both anytime-valid and scheduled monitoring criteria, while standard CTMs raise unnecessary alarms. WCTMs maintain target coverage and improve prediction informativeness by decreasing interval widths.

  2. From Mild to Extreme Covariate Shifts (Section 4.2): On CIFAR-10-C with varying corruption levels, WCTMs adapt to mild shifts without triggering alarms but raise alarms for severe shifts. For example, with level-1 brightness corruption, CTM quickly raises an unnecessary alarm while WCTM avoids it; at corruption level 5, WCTM does raise the alarm.

  3. Detection Speed (Section 4.3): Table 1 reports average detection delay (ADD) results. Among anytime-valid methods, WCTMs and CTMs achieve comparable ADD (e.g., 115.8 vs 114.8 on MEPS), but are over three times faster than the sequential testing methods from Podkopaev and Ramdas (2021b) (PR-ST, 536.6). Among stagewise methods, WCTM-based Shiryaev-Roberts procedures are comparable to CTM-based ones but significantly faster than PR-CD methods. The paper notes that WCTM and SR-WCTM methods have O(t) time complexity, whereas PR-CD-online has O(t2) complexity.

  4. Root-Cause Analysis (Section 4.4): Figure 4 demonstrates that the combination of WCTM and X-CTM can diagnose different shift types: benign covariate shifts (X-CTM detects, WCTM adapts without alarm), extreme covariate shifts (both detect), and concept shifts (only WCTM detects).

The paper concludes that WATCH achieves three main goals: (1) adaptation to benign shifts to avoid unnecessary alarms and improve prediction utility; (2) fast detection of harmful shifts; and (3) root-cause analysis to identify whether degradation is due to harmful covariate shift or fundamental concept shift. Future directions include developing WCTM algorithms for other nonparametric null hypotheses, exploring connections to conditional permutation tests, analyzing efficiency with respect to alternative hypotheses, and extending implementations to healthcare and foundation models.

Improvements for AI systems

Based on the paper, here are specific improvements I can implement in an AI system, along with the resulting capabilities:

  • Implementation: Add a post-deployment monitoring layer using Weighted-Conformal Test Martingales (WCTMs) that continuously evaluates the sequence of weighted-conformal p-values from the deployed model's predictions.

  • Specifics:

  • Compute p-values using Eq. (17) with online density-ratio estimation via a probabilistic classifier (e.g., MLP) that distinguishes source vs. target distributions.

  • Use the composite jumper betting strategy (Eq. 7) to construct the WCTM.

  • Set alarm threshold c (e.g., c=100) for anytime-valid control (Prop. 3.2).

  • Implementation: Dynamically reweight calibration scores using estimated density ratios (x) when a secondary X-CTM (monitoring only input distribution) exceeds an adaptation threshold.

  • Specifics:

  • Maintain a secondary CTM using nearest-neighbor nonconformity scores on X only.

  • When X-CTM crosses threshold tau ad (e.g., 10), switch main WCTM to weighted mode, fixing calibration set at [n+ tau ad-1].

  • Update density-ratio weights online using the probabilistic classifier.

  • Implementation: Raise an alarm when the WCTM wealth M t/M 0 exceeds a user-defined threshold, indicating concept shift or unanticipated covariate shift.

  • Specifics:

  • Use Shiryaev-Roberts procedure (Eq. 15) for scheduled monitoring with average run length control (Prop. 3.3).

  • For faster detection, use the anytime-valid criterion (Prop. 3.2) with threshold c=100.

  • Implementation: Diagnose alarm cause by comparing WCTM and X-CTM paths:

  • If both WCTM and X-CTM alarm → extreme covariate shift.

  • If only WCTM alarms → concept shift in YX.

  • If only X-CTM alarms → benign covariate shift (adapt, no retraining needed).

  • Implementation: When the weighted-conformal p-value would produce a trivial prediction set (e.g., n+1 at least alpha), use an anticonservative p-value (set u n+1=0) to accelerate alarm for harmful over-coverage.

  1. Continuously monitor deployed models with statistical guarantees: false alarm rate ≤ 1/c (anytime-valid) or ARL ≥ c (scheduled).

  2. Adapt to mild covariate shifts (e.g., demographic changes) without unnecessary retraining, maintaining target coverage (e.g., 90%) while improving prediction set sharpness.

  3. Rapidly detect harmful shifts (concept shifts or extreme covariate shifts) with detection delay up to 3x faster than existing methods (e.g., Podkopaev & Ramdas, 2021).

  4. Diagnose the root cause of performance degradation, distinguishing between:

  • Benign covariate shift (adapt, no alarm)

  • Extreme covariate shift (alarm, retrain with new data)

  • Concept shift (alarm, retrain with updated labels)

  1. Handle both regression and classification tasks with flexible nonconformity scores (e.g., absolute residual, one-minus-softmax).

  2. Operate in real-time with O(t) computational complexity per timestep, suitable for high-frequency monitoring (e.g., healthcare early warning systems).

  3. Provide anytime-valid evidence (e-process) that can be used for human-in-the-loop decision-making, not just binary alarms.

  4. Integrate with existing conformal prediction pipelines without requiring retraining of the base model, only adding a lightweight monitoring layer.

Sources

Related papers