Selective Prediction Reduces the Negative Effects of Automation Bias Overall but Increases False Negatives

arXiv:2508.07617 · cs.HC, cs.AI · Submitted 2026-08-11 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Selective Prediction Reduces the Negative Effects of Automation Bias Overall but Increases False Negatives".

Jane: The paper was written by Sarah Jabbour, David Fouhey, Nikola Banovic, Stephanie D. Shepard, Ella Kazerooni et al. from University of Michigan and New York University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title and Authors: Tom: Welcome back to the show, everybody. Today we're digging into a paper with a title that really makes you stop and think: "Selective Prediction Reduces the Negative Effects of Automation Bias Overall but Increases False Negatives." Jane, that title is a whole story in itself, isn't it?

Jane: It really is, Tom. It's basically saying, "We found a fix, but the fix has its own problem." And that's exactly the kind of nuance we need more of in AI research. The team is from the University of Michigan and NYU, with Sarah Jabbour leading the charge, and they ran a study with two hundred fifty-nine clinicians to test how AI tools change the way doctors treat patients with acute respiratory failure.

Tom: Right, and the setup is clever. They had doctors make treatment decisions with no AI, with AI that sometimes made mistakes, and with AI that was allowed to say, "I'm not sure, you decide on your own." That last one is the selective prediction part — the AI hides its prediction when it thinks showing it might hurt.

Jane: And the big finding is that when the AI just gave its prediction, even when it was wrong, doctors over-relied on it and their accuracy dropped from sixty-two percent down to fifty percent. That's automation bias in action. But when the AI withheld those bad predictions, accuracy came back up to fifty-five percent — so it helped, but it didn't fully restore the baseline.

Tom: But here's the twist that the title is pointing at. Even though overall accuracy improved, the type of errors changed. Doctors started making more false negatives — meaning they were more likely to undertreat patients who actually needed treatment. That's a serious clinical concern, because missing a treatment can be worse than giving an unnecessary one.

Jane: Exactly. And that's the kind of finding that should make us pause before we just assume that letting the AI abstain is a neutral, safe default. The paper is really challenging that assumption, and I think that's why it's getting so much attention.

Tom: So we're not just talking about a better algorithm here. We're talking about how the very act of telling a human "the AI is staying quiet on this one" changes their behavior. That's a human factors question, not just a machine learning question.

Jane: And it's a question that has real stakes. If we're going to deploy these systems in hospitals, we need to know not just whether they're accurate, but how they change the decisions people actually make. This paper is a big step toward that.

Tom: I'm already curious about how they actually ran this study and what the doctors were seeing on their screens. Let's get into the details in the next segment.

Summary and Methodology: Tom: So we're back with "Selective Prediction Reduces the Negative Effects of Automation Bias Overall but Increases False Negatives." Jane, let's talk about how they actually pulled this off, because it's not a typical lab experiment.

Jane: Right. They built a vignette-based study, which means they gave clinicians realistic patient cases — complete with chest X-rays, lab results, vital signs, and medical history — and asked them to figure out what was causing each patient's shortness of breath. The three possible culprits were pneumonia, heart failure, and COPD, and patients could have any combination of them.

Tom: And that multilabel setup is important, because it's not just a yes-or-no decision. The AI had to make three separate predictions for each patient, and selective prediction could hide some of those predictions while showing others. That's much closer to how these tools would actually work in practice.

Jane: Exactly. The doctors first did three cases with no AI at all to get a baseline. Then they were randomized into two groups: one saw all the AI predictions, and the other saw the AI's predictions only when they were accurate. When the AI was wrong, it showed a message saying "the model defers to you" instead of giving a score.

Tom: And they made sure the participants actually understood that deferral message. There was a knowledge check — you had to confirm that "the model may or may not believe the patient has the condition, but is choosing not to provide its input." So they were really testing the effect of abstention itself, not confusion about what abstention meant.

Jane: That's a key design choice. They wanted to isolate the behavioral response to being told "the AI is staying quiet," separate from any misunderstanding. And even with that understanding, the false negative rate still went up significantly — from thirty-one percent at baseline to forty-two percent when the AI abstained.

Tom: That's a huge jump. And it's not just a statistical blip — they used mixed-effects models to account for the fact that the same doctors were making multiple decisions and the same patient cases were being used across participants. So the finding is pretty robust.

Jane: And they also looked at subgroups. The effect was stronger for advanced practice practitioners than for physicians, and it was particularly pronounced for clinicians who said they had never used AI in their practice before. Those AI-naive clinicians saw their false negative rate jump by thirteen percentage points when the AI abstained.

Tom: So the people who are least familiar with AI are the ones most likely to be thrown off when it stays quiet. That makes intuitive sense — if you don't have experience with how these tools behave, you might interpret "I'm not sure" as "this is probably not a problem."

Jane: Right. And that's exactly the kind of insight you can't get from a simulation. You have to put real clinicians in front of real cases and watch what they do.

Tom: Okay, so we know the problem. What does this mean for how we should actually build and deploy these systems? That's what I want to dig into next.

Improvements and Implications: Tom: Welcome back. We're still on "Selective Prediction Reduces the Negative Effects of Automation Bias Overall but Increases False Negatives." Jane, we've established that selective prediction helps overall but shifts errors toward undertreatment. What does the paper suggest we do about that?

Jane: Well, the paper doesn't offer a simple fix, and I think that's actually the point. The authors argue that selective prediction should not be treated as a default safety mechanism. Instead, you have to evaluate it in each specific context and think carefully about which types of errors are more acceptable.

Tom: And in the context of treating acute respiratory failure, undertreatment is especially dangerous. The paper points out that delayed antibiotics for severe infections can lead to worse outcomes. So a system that nudges doctors toward undertreatment is trading one harm for another.

Jane: Exactly. And that's why the authors are calling for more research into why clinicians respond the way they do when the AI abstains. They have a couple of hypotheses. One is that clinicians might interpret "the model defers to you" as the model saying "there's no disease here," even when they understand it's abstaining.

Tom: That's the availability bias idea — the conditions the AI does show predictions for become more salient in the doctor's mind, so they focus on those and forget to consider the ones the AI stayed quiet about. The paper mentions that as a possible mechanism.

Jane: Right. And another implication is that training matters. The authors suggest that formal training on what selective prediction means could mitigate some of these effects. But they also note that even with a knowledge check in the study, the false negative effect still appeared. So training alone might not be enough.

Tom: So what's the practical path forward? Do we just abandon selective prediction?

Jane: No, I don't think so. The paper shows it does reduce the harm from automation bias — accuracy went from fifty percent back up to fifty-five percent. It's just that we need to be more thoughtful about when and how we use it. Maybe the AI should abstain less often, or maybe it should abstain only when the cost of a false positive is higher than the cost of a false negative.

Tom: That's a really important design question. And the paper also highlights that AI-experienced clinicians didn't show the same false negative effect. So there's hope that with more exposure, doctors can learn to calibrate their response to abstention.

Jane: Exactly. But we can't just assume that will happen on its own. We need to study it, measure it, and design for it. This paper is a great example of why we need to evaluate AI systems in the context of real workflows, not just in isolation.

Tom: Alright, let's wrap this up in the final segment.

Conclusion: Tom: We're wrapping up our discussion of "Selective Prediction Reduces the Negative Effects of Automation Bias Overall but Increases False Negatives." Jane, give us the final takeaway for our listeners.

Jane: Sure. This paper is a reminder that AI systems don't exist in a vacuum. When you put a tool in front of a human decision-maker, the way it behaves — including when it stays quiet — changes their behavior. Selective prediction can reduce the harm of automation bias, but it also shifts errors toward undertreatment, which can be just as dangerous in a clinical setting.

Tom: And that's a finding that should matter to anyone building AI for high-stakes decisions, not just in healthcare. The paper's core message is that we need to evaluate these systems in real workflows, with real users, and pay attention not just to accuracy but to the types of errors people make.

Jane: Right. And the authors are careful to note that their study is a vignette-based survey, so it's not the same as a real hospital setting. But it's an important first step in identifying potential harms before we deploy these systems widely.

Tom: It also opens up a lot of questions for future work — how to train clinicians to respond to abstention, how to design selective prediction mechanisms that account for the cost of different errors, and how to make sure the people who are least familiar with AI aren't the ones most harmed by it.

Jane: Definitely. This is one of those papers that doesn't give you a clean answer, but it gives you a much better understanding of the problem. And that's exactly what we need more of.

Tom: Well said. Thanks for joining us, and we'll be back soon with another paper. Until then, keep questioning the tools you use.

Jane: And remember — when an AI says "I'm not sure," that's not the same as saying "you're fine." Take care, everyone.

Sarah Jabbour, David Fouhey, Nikola Banovic, Stephanie D. Shepard, Ella Kazerooni, Michael W. Sjoding, Jenna Wiens

University of Michigan · New York University

cs.HC, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: 14 pages, 10 figures, 5 tables. v2: Revised results and analysis; title updated. Previously circulated as "On the Limits of Selective AI Prediction: A Case Study in Clinical Decision Making."

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 71/100

The gist: This paper investigates whether selective prediction—an AI design strategy where the system withholds potentially harmful predictions and explicitly defers to the human—causes humans to behave as

Terminology

Summary

This paper investigates whether selective prediction—an AI design strategy where the system withholds potentially harmful predictions and explicitly defers to the human—causes humans to behave as they would in a workflow with no AI involvement at all. The authors test this assumption in a multilabel clinical decision-making setting, specifically treating the causes of acute respiratory failure (ARF) in hospitalized patients.

Study Design: The authors conducted a vignette-based study of 259 clinicians (125 randomized to Clinician + AI, 134 randomized to Clinician + Selective Prediction) who were tasked with treating hospitalized patients with ARF. ARF has three common causes—pneumonia, heart failure, and COPD—and patients can have any combination, creating a multilabel decision-making setting. Clinicians first completed three patient cases without AI input (Clinician Alone baseline), then were randomized to receive AI predictions either with or without selective prediction for six additional cases. Three of these cases contained inaccurate predictions (where the selective prediction mechanism identified potentially harmful predictions), and three contained only accurate predictions. The selective prediction mechanism withheld all incorrect predictions and a subset of correct predictions (approximately 11% of withheld predictions were correct).

Key Findings:

  1. Inaccurate AI hurts clinician accuracy: In Clinician Alone, treatment accuracy was 62% (95% CI: [47-75]). In Clinician + AI, treatment accuracy dropped significantly to 50% [36-65] (p-value < 0.001), suggesting overreliance on inaccurate predictions. This drop was driven by increases in both false positives (from 40% [18-67] to 55% [28-79], p-value < 0.05) and false negatives (from 31% [22-42] to 41% [31-53], p-value < 0.05).

  2. Selective prediction partly recovers accuracy but changes error types: Compared to Clinician + AI, treatment accuracy was higher under selective prediction (55% [40-69] vs. 50% [36-65]). Accuracy remained slightly lower than Clinician Alone (55% vs. 62%), though this difference was not statistically significant. Under Clinician + Selective Prediction, clinicians' false positive rate returned to baseline levels (41% [19-68] vs. 40% [18-67]). However, their false negative rate remained significantly elevated compared to Clinician Alone (42% [31-53] vs. 31% [22-42]; p-value < 0.05), indicating clinicians are more likely to undertreat patients who need intervention when the AI abstains.

Exploratory Analyses:

  • By medical training length: Effects of AI were stronger for Advanced Practice Practitioners (APPs) than physicians. In Clinician + AI, APP accuracy decreased by 18 percentage points relative to Clinician Alone, compared to a 10 pp decrease for physicians. Under Clinician + Selective Prediction, APP accuracy recovered by 13 pp to within 5 pp of baseline, whereas physician accuracy recovered by only 2 pp. False negative rates were elevated for both groups (10 pp higher for APPs, 11 pp higher for physicians).

  • By prior AI exposure: Effects of selective prediction were stronger on AI-naive clinicians. In Clinician + AI, accuracy decreased by 12 pp for AI-naive clinicians versus 9 pp for AI-experienced clinicians. Under Clinician + Selective Prediction, accuracy recovered by only 3 pp for AI-naive clinicians but recovered fully by 9 pp for AI-experienced clinicians. AI-naive clinicians' false negative rates were elevated by 13 pp relative to Clinician Alone, whereas they were only elevated by 2 pp for AI-experienced clinicians.

  • Perceived case difficulty: Selective prediction increased perceived case difficulty. At baseline, average perceived difficulty was 1.17 (95% CI: [1.11-1.23]). Clinician + AI increased this to 1.20 (1.13-1.28), and Clinician + Selective Prediction increased it further to 1.27 (1.20-1.35).

Conclusions: The core assumption underlying selective prediction evaluation—that when AI abstains, humans behave as they would without AI—does not fully hold. The authors state: When clinicians are informed that the model has withheld its prediction, abstention appears to influence decision-making, rather than functioning as a neutral absence of information. They note that AI abstention shifted clinicians towards undertreatment by increasing false negative treatment errors, which is clinically important because undertreatment (false negatives) may actually pose more harm than overtreatment (false positives). The authors conclude that selective prediction should not be treated as a default safety mechanism and that the effects of selective prediction should be evaluated in each end-use context, with careful attention to which types of errors are most acceptable.

Improvements for AI systems

Based on the paper's findings, here are the specific improvements I can make to AI systems, along with what the improved system can do:


Improvement: Modify the selective prediction mechanism to account for the asymmetric error costs revealed in the study. The paper shows that AI abstention increases human false negative rates (undertreatment) by 11 percentage points (31% → 42%) while restoring false positive rates to baseline.

What the improved system does: When deciding whether to abstain, the system weighs the cost of a potential false negative (e.g., undertreating pneumonia) against the cost of a false positive (e.g., unnecessary antibiotics). For conditions where undertreatment is more harmful (e.g., sepsis, heart failure), the system errs toward showing its prediction even if uncertain, rather than abstaining. For conditions where overtreatment is more harmful (e.g., unnecessary surgery), it abstains more aggressively.

Improvement: Replace the neutral the model defers to you message with context-aware messaging that counteracts the observed bias toward undertreatment.

Improvement: Differentiate selective prediction behavior based on the user's prior AI exposure and training level, as the paper shows AI-naive clinicians and APPs are more negatively affected.

Improvement: Add a lightweight decision-support step after the user makes their independent decision on an abstained condition.

Improvement: Replace the oracle-based abstention mechanism (which withholds all incorrect predictions plus 11% correct ones) with a threshold calibrated using the empirical human error rates from this study.

Improvement: The paper shows selective prediction increases perceived case difficulty (1.17 → 1.27 on a 4-point scale). The system should account for this cognitive load.

Improvement: For conditions where the model abstains and the clinician's independent decision is a no treatment (negative), the system triggers a secondary check.

The improved AI system:

  • Reduces undertreatment by counteracting the false negative bias induced by abstention

  • Adapts to user expertise to minimize harm for vulnerable groups (AI-naive, APPs)

  • Balances error types based on clinical context rather than treating all errors equally

  • Maintains the benefits of selective prediction (reduced automation bias, restored false positive rates)

  • Provides transparent, actionable abstention that supports rather than biases human judgment

Abstract

AI has the potential to augment human decision making. However, even high-performing models can produce inaccurate predictions when deployed. These inaccuracies, combined with automation bias, where humans overrely on AI predictions, can result in worse decisions. Selective prediction, in which potentially unreliable model predictions are hidden from users, has been proposed as a solution. This approach assumes that when AI abstains and informs the user so, humans make decisions as they would without AI involvement. To test this assumption, we study the effects of selective prediction on human decisions in a clinical context. We conducted a user study of 259 clinicians tasked with diagnosing and treating hospitalized patients. We compared their baseline performance without any AI involvement to their AI-assisted accuracy with and without selective prediction. Our findings indicate that selective prediction mitigates the negative effects of inaccurate AI in terms of decision accuracy. Compared to no AI assistance, clinician accuracy declined when shown inaccurate AI predictions (66% [95% CI: 56%-75%] vs. 56% [95% CI: 46%-66%]), but recovered under selective prediction (64% [95% CI: 54%-73%]). However, while selective prediction nearly maintains overall accuracy, our results suggest that it alters patterns of mistakes: when informed the AI abstains, clinicians underdiagnose (18% increase in missed diagnoses) and undertreat (35% increase in missed treatments) compared to no AI input at all. Our findings underscore the importance of empirically validating assumptions about how humans engage with AI within human-AI systems.

Sources

Related papers