Deployment of AI-Assisted Interventions: Capacity Constraints and Noisy Compliance

arXiv:2604.14370 · stat.ME, cs.LG · Submitted 2026-04-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Deployment of AI-Assisted Interventions".

Jane: AI tools are increasingly used to guide targeted interventions in healthcare, education, and recruiting, but standard practice fails when service capacity is limited and compliance with algorithmic suggestions is noisy.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: Welcome back everyone! Today we’re talking about this fascinating paper from arXiv, "Deployment of AI-Assisted Interventions: Capacity Constraints and Noisy Compliance." It tackles a really practical problem where AI is used to guide people toward getting help in areas like healthcare or recruiting.

Jane: It sounds like the core issue they’re exploring is that when you use an algorithm to flag who needs intervention, it doesn't always work out perfectly in the real world, especially when there are limits on how much service can actually be provided.

Lu: This paper really zeroes in on how standard practice of just picking the most accurate prediction score might fail if you don't account for capacity limitations and how people actually respond to those suggestions. It claims that relying only on predictive accuracy for setting intervention thresholds leads to outcomes that aren't optimal because it misses important competing forces at play.

Meng: Competing forces, right? So it’s not just about making the best prediction; it’s about managing the system itself, which makes sense from a practical standpoint.

Lalam: From my perspective as a model, I see this paper highlighting how we need to look beyond just maximizing raw accuracy metrics when designing systems that interact with real users under resource constraints.

Tom: Exactly, and the authors lay out two key issues they are trying to solve: making sure all available capacity gets filled versus making sure the most important individuals actually get served even when there’s competition. They argue that a good threshold needs to balance those two things carefully.

Jane: That balancing act is crucial because if you focus too much on filling every slot, you might end up serving people who don't need it as much, which isn't efficient for the system.

Lu: The model they use formalizes this by looking at how individual request behavior changes based on the intervention nudge, which they capture with a Bernoulli random variable dependent on p zero + P s i. This is how you account for the probabilistic nature of human response to an AI suggestion.

Meng: So when demand goes up, and the capacity is tight, that nudge can cause some people to request service while others might skip it entirely, and that's where the noise comes in.

Lalam: That noise in compliance is a big factor; patients might ignore reminders or candidates might decline outreach, which means the system isn't just acting on a perfect prediction.

Paper summary: Tom: Right, and they show that prediction-based thresholds, which only look at metrics like the True Positive Rate or Positive Predictive Value, fall short because they completely ignore these operational constraints. They assume better predictions automatically equal better system performance.

Jane: That's a very common pitfall in deploying AI tools; assuming the model’s internal score is the only thing that matters without checking if it fits within the real-world service limits.

Lu: The paper characterizes this optimal threshold as being determined by the minimum of two different thresholds: one that tries to account for cannibalization and another one that accounts for underutilization.

Meng: Cannibalization sounds like when a high-value request gets pushed out by too many lower-value requests when capacity is scarce, which is a real concern for resource allocation.

Lalam: And the paper shows that under scarce capacity conditions, the optimal policy flags fewer individuals to reduce total requests while concentrating service on those with higher scores.

Tom: So, if capacity is tight compared to demand—when M/N is small—the system should lean toward flagging fewer people because of that crowding effect. But if capacity is abundant, things flip and it prioritizes filling the slots up.

Jane: It’s a dynamic setting, isn't it? The optimal action changes depending on whether the system is stretched thin or has plenty of room to operate.

Lu: Precisely, and they define this behavior by two regimes based on the capacity ratio M/N; when M/N is less than or equal to p zero the optimal threshold becomes tau* A score, focusing on which requesters get served.

Meng: So if we're in that high-demand, low-capacity situation, the system shifts its focus entirely to selecting the individuals who truly warrant service based on their predicted scores. That’s a very grounded engineering decision.

Lalam: And conversely, when capacity is abundant, M/N is greater than p zero + P, and the optimal threshold switches to tau c, which focuses more on ensuring full utilization of those available slots.

Tom: And this brings us to how we evaluate the models, because standard metrics like AUC don't capture this trade-off, leading to what they call suboptimality.

Jane: They propose Operational AUC as a better way to judge algorithm selection because it weights performance over the actual capacity distribution rather than just looking at one single metric across the board.

Lu: The definition of OpAUC includes a term related to p zero and P, which means it explicitly factors in the baseline request rate and how much the intervention nudges affect those rates, all while considering the capacity distribution mu(rho).

Paper summary: Meng: That makes sense for practical application; we need a metric that reflects how well an AI tool performs when it's actually running on a constrained system, not just in a perfect test environment.

Lalam: And the case study results really support this idea; they found that an algorithm with lower overall AUC could actually show higher system efficacy when deployed at its optimal threshold across different capacity conditions.

Tom: That’s a powerful finding, because it means we shouldn't just pick the highest AUC score in isolation; we need to consider how that algorithm behaves across the whole operating range of capacities.

Jane: So, when you look at the paper "Deployment of AI-Assisted Interventions: Capacity Constraints and Noisy Compliance," what does this all boil down to in simple terms for our listeners?

Lu: Essentially, it shows that designing these AI interventions isn't just about building a brilliant predictor; it’s about designing the entire deployment strategy to manage the tension between serving high-value individuals and making sure you don't waste service slots when capacity is limited.

Meng: From an engineering standpoint, this means we stop looking at a single prediction accuracy number and start looking at how that prediction interacts with the system's physical limits, like the available service slots.

Lalam: If we apply this thinking to our culture, it suggests that AI tools should be implemented not just for their predictive power, but in a way that accounts for how people actually respond under pressure and resource scarcity.

Tom: It shifts the focus from just developing a model to figuring out the best way to implement it given real operational constraints, which is something our listeners need to hear about.

Jane: So, we’re moving away from simply optimizing for a single number and towards designing systems that are robust across different operational scenarios, whether capacity is high or low.

Lu: This framework requires practitioners to understand their capacity distribution and behavioral parameters more deeply than just the model's raw output.

Meng: It’s about shifting our focus from algorithm development to implementation strategy, which is a huge change for how we approach deploying these tools in practice.

Lalam: The implication for the future is that we need to build systems that are inherently aware of these constraints, perhaps by incorporating this understanding into how the model's output is translated into an action.

Tom: That’s a big thought, and it really gives listeners a practical lens through which to view the utility of these AI tools in their own fields.

Conclusion: Tom: So, we've seen how this paper tackles the real challenge of putting AI tools into practice when there are limits on what you can actually deliver.

Jane: Exactly, and they focus on how that threshold needs to balance making sure everyone gets help with managing competition for limited resources in a noisy environment.

Lu: It’s really clever how they frame the problem as balancing two competing pressures, which is something I find fascinating from a theoretical standpoint.

Meng: From an engineering side, the idea of needing two different operational thresholds based on the capacity ratio makes perfect sense when you're trying to design a robust system.

Lalam: And for me, it points toward a future where AI tools are designed not just for their peak performance in isolation, but for their actual resilience under real-world conditions.

Tom: That’s what I want to emphasize—it’s about the deployment strategy as much as the algorithm itself that matters here.

Jane: They’re suggesting that simply picking the best prediction score isn't enough; you have to adjust your approach based on how busy the system is right now.

Lu: The concept of shifting between two different behaviors depending on whether capacity is scarce or abundant gives us a lot to explore in terms of adaptive AI architectures.

Meng: It means we don't just optimize for one scenario; we build for the whole range of conditions, which is a much more practical way to approach system design.

Lalam: This has huge implications because it suggests that the next generation of AI tools needs to be inherently aware of these operational constraints from the very start.

Tom: And that really makes you wonder what kinds of real-world systems we can build if we start thinking about this kind of nuanced resource management.

Jane: It moves the conversation away from just achieving high accuracy numbers and toward building systems that function reliably in complex, constrained settings.

Lu: It opens up a whole new avenue for creative thinking on how AI agents can dynamically adjust their behavior based on those capacity signals.

Meng: So, this isn't just academic theory; it’s about making sure the software actually performs well when it’s running in a real environment where things get messy.

Lalam: And I see that improving our tools this way means we can create systems that are much more responsive and fair to the individuals they serve.

Tom: It really paints a picture of what’s next for deploying these kinds of AI interventions in critical fields like healthcare or recruiting.

Columbia Business School · Columbia University

stat.ME, cs.LG

Submitted: 2026-04-15

Updated: 2026-09-27

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 90/100

The gist: AI tools are increasingly used to guide targeted interventions in healthcare, education, and recruiting, but standard practice fails when service capacity is limited and compliance with algorithmic

Key concepts

Underutilization
This occurs when there is plenty of service capacity available compared to the actual demand for interventions. If the system flags too few individuals, valuable service slots go unused, wasting potential resources. The optimal threshold must prevent this by ensuring capacity is filled when it's abundant.
Cannibalization
This happens when service capacity is scarce relative to demand. If the system flags too many individuals, lower-value requests might crowd out or displace the higher-value individuals that were flagged first. The threshold must balance serving high-value needs against ensuring all available slots are used.
Prediction-based Thresholds
These are thresholds set using only metrics like True Positive Rate (TPR) or Positive Predictive Value (PPV). They focus purely on the algorithm's accuracy in predicting who needs help, completely ignoring real-world operational limits like service capacity or the actual behavior of requesters.
Operational AUC (OpAUC)
This is a proposed evaluation metric that replaces standard metrics like AUC when operational constraints exist. OpAUC measures performance by weighting results over the distribution of available capacity, providing a more principled way to select algorithms for real-world deployment.

Terminology

Summary

AI tools are increasingly used to guide targeted interventions in healthcare, education, and recruiting, but standard practice fails when service capacity is limited and compliance with algorithmic suggestions is noisy. This work demonstrates that relying solely on predictive accuracy for setting intervention thresholds leads to suboptimal system performance because it ignores the competing forces of capacity underutilization and cannibalization.

The gist

The optimal threshold must balance two effects: ensuring all capacity is filled (utilization) and ensuring high-value individuals are served despite competition between requests (cannibalization).

Model Components and Behavior

The model formalizes AI-assisted interventions where algorithms generate predicted scores, which are then used to set a quantile threshold for flagging individuals for service. Individual request behavior is modeled as a Bernoulli random variable dependent on the nudging indicator: di ∼ Bernoulli(p0 + ∆P si). This captures the probabilistic response to an intervention, where p0 is the baseline probability and ∆P is the lift due to nudging. The system operates under capacity constraints, where service slots are allocated uniformly at random among requesters when demand exceeds capacity.

Competing Forces in Threshold Selection

The paper identifies two competing forces that any deployment must account for:

  1. Underutilization: Occurs when capacity is abundant relative to demand, and flagging too few individuals leads to wasted capacity.

  2. Cannibalization: Occurs when capacity is scarce relative to demand, and if too many individuals are flagged, lower-value requests crowd out higher-value flagged ones.

The optimal threshold must balance these forces. The paper characterizes this as the minimum of two interpretable thresholds, one that accounts for cannibalization and one that accounts for underutilization.

Suboptimality of Natural Thresholding Policies

Two common natural policies are shown to be suboptimal:

  1. Prediction-based Thresholds: These depend only on predictive metrics such as True Positive Rate (TPR; sensitivity) or Positive Predictive Value (PPV), ignoring operational constraints. They fail because they ignore operational constraints entirely.

  2. Capacity-matching Thresholds: These adjust outreach to match expected requests to available capacity, ensuring full utilization but ignoring the cannibalization effect when demand is high.

Optimal Threshold Characterization

The optimal threshold is characterized by Theorem 1 as the minimum of the score-optimal threshold and the capacity-matching threshold:

τ∗A M/N p0, ∆P = min τ∗A score(p0, ∆P), τc M/N p0, ∆P

The behavior of this optimal threshold is determined by two regimes based on the capacity ratio M/N:

  1. When capacity is scarce compared to baseline requests (M/N ≤ p0), the optimal threshold is τ∗A score, focusing on which requesters are served.

  2. When capacity is abundant (M/N ≥ p0 + ∆P), the optimal threshold is τc, focusing on how many requesters are served to ensure full utilization.

Algorithm Selection Metric

Standard evaluation metrics like AUC, which weight all thresholds equally, are shown to be misaligned with operational performance. The paper proposes the Operational AUC (OpAUC) as a principled replacement:

OpAUC (Aµ, p0, ∆P) = Z ρ · p0 + ∆P TPRA (qA(τ∗A(ρ p0,∆P))) / (p0 + ∆P (1 − τ∗A(ρ p0,∆P))) dµ(ρ)

OpAUC aligns algorithm selection with realized system efficacy by weighting performance over the capacity distribution. In a case study, an algorithm with lower overall AUC can achieve higher system efficacy when deployed at its optimal threshold across varying capacity conditions.

Case Study Findings

In a sepsis early warning system simulation, prediction-based thresholds incurred efficacy losses of up to 40% in the worst case. Furthermore, for algorithm selection under limited capacity (M ∈ [10, 30]), XGBoost outperformed Epic because OpAUC was higher for XGBoost when the operating capacity was distributed as µ ∼ Uniform(0.05, 0.15). This demonstrates that algorithm selection should be guided by performance over the capacity-induced operating range.

Practical Implications

The framework shifts focus from algorithm development to implementation:

  1. Practitioners can improve efficacy by adjusting thresholds based on capacity and behavioral parameters, without needing to retrain models.

  2. Algorithm selection should use OpAUC instead of AUC when operational constraints are present, as OpAUC provides a principled basis for navigating this tradeoff.

Future Work

Future work suggests incorporating heterogeneity in baseline request probability and behavioral responses across the population. The framework is designed to be immediately actionable by requiring only additional knowledge of the operating capacity distribution and behavioral parameters, which are typically estimable from operational data.

Improvements for AI systems

Based on the provided research, here are specific, actionable improvements for deploying AI-assisted interventions in capacity-constrained environments:


The core improvement is shifting the focus from maximizing predictive accuracy (AUC) to maximizing System Efficacy by optimizing both the service threshold and algorithm selection based on operational realities.

Here are the specific improvements categorized by system component:

  1. // Threshold Selection Improvement: Implement a Two-Point Optimal Threshold Policy

A standard practice is to choose a single threshold based solely on predictive accuracy (e.g., maximizing AUC). The paper proves this is suboptimal because it ignores capacity constraints and behavioral responses.

  • Use the derived optimal threshold, which is the minimum of two thresholds:

  • The Score-Optimal Threshold (to maximize efficacy per served request, accounting for cannibalization).

  • The Capacity-Matching Threshold (to ensure full utilization of scarce service slots).

  1. // Algorithm Selection Improvement: Replace AUC with Operational AUC (OpAUC)

When multiple predictive models are available, selecting the one with the highest overall AUC is misleading because it may perform poorly at the specific thresholds a system will actually deploy.

  • Use the Operational AUC (OpAUC) as the selection metric. This metric weights each algorithm's performance over its own optimal threshold across all expected capacity ratios.

  • This ensures that when selecting an algorithm, you choose one that maximizes the realized system efficacy under the constraints of your specific deployment environment (capacity distribution and behavioral parameters).

  1. // Operational Deployment Strategy: Dynamic Threshold Adjustment Based on System State

The optimal threshold is not static; it must adapt based on the current operational regime (capacity ratio, baseline demand).

  • Implement a switching mechanism where the system dynamically selects between the Score-Optimal Threshold and the Capacity-Matching Threshold based on real-time capacity indicators (e.g., current patient load relative to shift capacity).

  • This ensures that in periods of high competition (low capacity), you prioritize maximizing value per slot; in periods of abundance, you prioritize filling all available slots.

  1. // Provider Behavior Modeling: Account for Prioritization

Standard models assume random allocation when requests exceed capacity. In reality, providers often prioritize high-scoring requests (e.g., reserved slots).

  • Model the provider as a mixture mechanism where they allocate slots based on a prioritization weight parameter (β1).

  • The framework allows you to select thresholds that are robust across different prioritization levels, providing a principled hedge against uncertainty regarding how much the provider trusts and uses the algorithm.

The improved AI system can achieve:

  1. // Enhanced Efficacy in Sepsis Early Warning Systems:

Instead of simply identifying the most accurate model (highest AUC), this system will select the model that maximizes expected patient survival/treatment success given current hospital capacity constraints and nurse workload. This leads to a demonstrated improvement in true positive case identification (up to 40% relative gain) compared to using standard prediction-based thresholds.

  1. // Optimized Resource Allocation:

The system will ensure that scarce resources (like coordination nurse time or inpatient beds) are not wasted due to underutilization, while simultaneously preventing the cannibalization effect where lower-value requests crowd out high-value ones when capacity is tight. The system will allocate service slots in a way that balances filling every slot with serving the highest-value individuals possible.

  1. // Robust Algorithm Deployment:

When multiple AI models are used (e.g., Epic vs. XGBoost), this system will select the model that is best suited for the specific operational context—not just the one with the best general accuracy score—thereby maximizing real-world impact rather than theoretical performance metrics.

  1. // Actionable Implementation:

Practitioners can achieve substantial, measurable gains simply by adjusting thresholds based on known operational parameters (like capacity ratios and expected patient baseline request rates), without needing to retrain complex models from scratch.

Abstract

AI tools increasingly drive targeted interventions in various service settings, including healthcare, education, and public services. Algorithms score individuals, trigger outreach to those above a threshold (e.g., high-risk or high-value), and encourage them to request service; then providers deliver service to those who request. Much of the work in this area has focused on improving predictive accuracy, implicitly assuming that better predictions lead to better outcomes. We show that predictions are only one component of a larger service system: when service capacity is limited and behavioral responses to outreach are probabilistic, system efficacy depends on operational forces that predictive accuracy does not capture. In such settings, the optimal score threshold must balance two effects: ensuring all capacity is filled (utilization) and, when capacity is constrained, preventing low-value requests from crowding out high-value ones (cannibalization). We characterize the optimal threshold and prove that thresholds based solely on predictive accuracy are generally suboptimal. Further, algorithm selection metrics such as AUC can be misaligned with operational performance: they weight all thresholds equally, while optimal deployment uses a subset of thresholds that depends on both capacity and compliance behavior. We introduce a new metric, Operational AUC (OpAUC), and show that it identifies the efficacy-optimal algorithm. Finally, we conduct a case study on sepsis early warning data that illustrates the magnitude of the improvement available from better threshold selection and shows that a predictor with lower AUC can achieve higher system efficacy under optimal deployment.

Sources

Related papers