Harmonizing Safety and Speed: A Human-Algorithm Approach to Enhance the FDA's Medical Device Clearance Policy

arXiv:2407.11823 · cs.LG, cs.HC, math.OC, stat.ML · Submitted 2026-08-12 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Harmonizing Safety and Speed: A Human-Algorithm Approach to Enhance the FDA's Medical Device Clearance Policy".

Jane: The paper was written by Mohammad Zhalechian, Soroush Saghafian and Omar Robles from Indiana University and Harvard University and Emerging Health Consulting.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back, everyone. Today we're looking at a paper that's been making waves in the regulatory world, and it's called "Harmonizing Safety and Speed: A Human-Algorithm Approach to Enhance the FDA's Medical Device Clearance Policy." Jane, I gotta say, that title alone got me hooked.

Jane: Oh, absolutely, Tom. And I love the word "harmonizing" because it gets right at the tension here. The FDA has to be safe, but it also has to be fast. This paper is trying to do both at the same time.

Tom: Right, and it's not just about speed for speed's sake. It's about the fact that the current system, the five hundred ten(k) pathway, clears thousands of devices a year. But a lot of them get recalled later. The paper opens with a pretty stark stat: the current recall rate is about ten percent.

Jane: Ten percent is huge when you think about how many devices we're talking about. And the authors aren't just pointing at the problem. They're building a tool to fix it. They've assembled this massive dataset of over thirty-one thousand submissions.

Tom: And that's where it gets really interesting. They're not just throwing data at a wall. They're using machine learning to predict which devices are likely to be recalled, based only on information available at the time of submission. That's the key, right? You can't use hindsight.

Jane: Exactly. You have to put yourself in the FDA's shoes on day one. What do you know about this device? What do you know about its predicates, the devices it's claiming to be similar to? That's the information they're feeding into their model.

Tom: And the model works. They got an area under the curve score of zero point seven eight, which for anyone listening who isn't a stats nerd, that means it's pretty good at telling the risky devices from the safe ones.

Jane: It's a strong signal, for sure. But the really clever part, and I think this is what makes the paper special, is that they don't just let the algorithm make the final call. They're designing a policy where the algorithm flags the easy cases, and the hard cases go to human experts.

Tom: A human-algorithm approach. That's the core of the title. And it makes so much sense. You don't want a machine rejecting a brand new idea just because it looks a little different from what's come before. You want a human to look at that.

Jane: Right, and you don't want a human spending hours on a device that's almost identical to one that's been safe for years. The algorithm can clear that in a second. So it's about using each tool where it's strongest.

Tom: And the results they project are pretty wild. They're talking about a thirty-two point nine percent improvement in the recall rate and a forty point five percent reduction in the FDA's workload. That's a massive win on both fronts.

Jane: It really is. And it's not just about saving money, though they estimate about.7 billion in annual savings. It's about not having unsafe devices on the market in the first place. It's about preventing those recalls from ever happening.

Tom: So we've got the big picture. But I'm dying to know how they actually built this thing. How did they get the data on the predicates? That seems like a nightmare.

Jane: Oh, it was a nightmare. They had to build a text mining algorithm just to find out which devices were used as predicates. The FDA doesn't link them in a structured way. That's a story for the next segment, for sure.

Summary: Tom: So we're back with "Harmonizing Safety and Speed: A Human-Algorithm Approach to Enhance the FDA's Medical Device Clearance Policy." And Jane, you were just about to tell me how they got the data.

Jane: Right, the data. So, the FDA has all these summary documents for each five hundred ten(k) submission. But the predicate devices, the ones the new device is compared to, are just mentioned in the text. They're not in a neat little database field.

Tom: So they had to read every single one?

Jane: Essentially, yes. They built a text mining algorithm to scan all those PDFs, find the predicate numbers, and link them to the recall data. It was a huge engineering effort just to get to the starting line.

Tom: And that effort paid off. They found that the number of recalls for the predicate devices is a huge predictor of whether the new device will be recalled. It sounds obvious when you say it, but nobody had proven it with data at this scale.

Jane: It's not just the number, though. It's the timing. They created this "weighted recall score" that gives more weight to recent recalls. And that was a really strong predictor. The idea being, if a predicate was just recalled, the manufacturer might not have had time to fix the underlying problem.

Tom: So a device that leans on a recently-recalled predicate is a big red flag. That makes a lot of sense. And they also found that the age of the predicate matters, but in a nuanced way.

Jane: Exactly. Newer predicates are actually riskier, which is counterintuitive. If a device has only been on the market for a year, it hasn't been tested by time. The long-term problems haven't shown up yet. Older predicates, with a long track record, are safer.

Tom: So there's a learning curve effect. The device needs time on the market to prove itself. And the model picks up on all of this.

Jane: It does. And it's not just about the predicates. The model also looks at the applicant device itself. Things like the medical specialty, the product code, even the country of origin. All of that feeds into the risk score.

Tom: And they found some interesting stuff there, too. For example, devices that are life-sustaining or supporting are more likely to be recalled. That makes sense, because they're under more scrutiny.

Jane: Right, they're higher risk, so they're watched more closely. The model is really good at picking up on these complex interactions.

Tom: Now, here's the part I keep coming back to. They didn't just build a model and say, "Trust the machine." They built a policy around it. Can you walk me through how that works?

Jane: Sure. So the model gives every device a risk score from zero to one. The policy sets two thresholds: a low one and a high one. If a device scores below the low threshold, it's an automatic accept. If it scores above the high threshold, it's an automatic reject.

Tom: And everything in between?

Jane: That's the "defer" zone. Those are the hard cases. The algorithm isn't confident enough to make a call, so it sends them to the FDA's human committees for a deep dive.

Tom: So the algorithm handles the easy stuff, and the humans handle the gray areas. That's the human-algorithm approach in action.

Jane: Exactly. And they optimized those thresholds to balance three things: catching unsafe devices, not rejecting safe ones, and not overloading the FDA with work. It's a delicate balancing act.

Tom: And they have a whole section on the math behind finding those thresholds. It's a non-convex optimization problem, which is a fancy way of saying it's hard to solve.

Jane: It is hard, but they came up with a clever nested search algorithm to find a near-optimal solution. They even proved some structural properties about the problem that make it easier to solve in certain cases.

Tom: So they've got the model, they've got the policy, and they've got the math to back it up. But the big question is, does it actually work in the real world? That's what we need to talk about next.

Improvements: Tom: We're back with "Harmonizing Safety and Speed: A Human-Algorithm Approach to Enhance the FDA's Medical Device Clearance Policy." And Jane, we just talked about how they built the thing. Now we need to talk about what happens when you actually run it.

Jane: Right. And the results are pretty compelling. They tested their policy on a hold-out set of data, and compared it to what the FDA actually did. The improvement in the recall rate was thirty-two point nine percent.

Tom: That's a huge number. And they also cut the workload by forty point five percent. So they're catching more bad devices and doing less work at the same time. How is that possible?

Jane: Because they're not doing the same work. The FDA currently reviews every single device with the same level of scrutiny. This policy lets the algorithm handle the clear-cut cases, so the human reviewers can focus their energy on the ambiguous ones.

Tom: So it's about focusing effort where it matters most. And they were conservative in their evaluation, too. They assumed the human committees wouldn't get any better at their jobs, even with the extra information the algorithm provides.

Jane: That's the conservative part. In reality, the committees would have the risk score in front of them. They'd know why the algorithm was uncertain. That information alone could help them make better decisions. So the real-world improvement could be even larger.

Tom: And they didn't just stop at recalls. They looked at the cost. They estimated the replacement cost of the recalled devices they would have caught, using Medicare claims data. That's where the.7 billion in annual savings comes from.

Jane: And that's just the replacement cost. It doesn't include the cost of injuries, the lawsuits, the pain and suffering. The true cost of a recall is much higher than the price of the device itself.

Tom: So this policy could save billions of dollars and, more importantly, prevent patient harm. But it's not a silver bullet. There are some trade-offs.

Jane: There are. The policy rejects more devices than the FDA currently does. Some of those rejected devices would have been safe. The paper calls those "false rejections."

Tom: But they argue that's not as bad as it sounds, because for five hundred ten(k) devices, there's almost always a similar device already on the market. So rejecting one doesn't leave patients without options.

Jane: That's a good point. The clinical harm is minimal. But it does impose a cost on the manufacturer. They have to go back to the drawing board, maybe resubmit. That's a real burden.

Tom: And there's also the question of gaming the system. If manufacturers know the algorithm's rules, they might try to game it. They might pick predicates that look safe on paper, even if they're not the most similar.

Jane: That's a real concern. The paper suggests keeping some details of the model and the thresholds confidential to make that harder. And they also suggest periodically recalibrating the policy to adapt to new tactics.

Tom: So it's not a set-it-and-forget-it solution. It needs to be maintained and updated. But the framework is solid.

Jane: I think so. And the authors are careful to point out that this is a decision support tool, not a replacement for the FDA. The final call still rests with the human reviewers.

Tom: It's a tool to help them do their jobs better. And that's a really exciting vision for how AI can be used in regulation. It's not about replacing people; it's about empowering them.

Jane: Exactly. And that's a message that goes way beyond medical devices. This framework could be used for any kind of regulatory review, from drugs to financial products.

Tom: So we've got the results, the cost savings, and the caveats. But what does this actually mean for the future? What's the big picture here?

Conclusion: Tom: Alright, we're wrapping up our discussion on "Harmonizing Safety and Speed: A Human-Algorithm Approach to Enhance the FDA's Medical Device Clearance Policy." Jane, it's been a fascinating ride.

Jane: It really has, Tom. We started with a problem: the FDA's five hundred ten(k) pathway clears devices that later get recalled. And we ended with a solution that uses machine learning to flag risky devices and human experts to make the final call.

Tom: The key numbers are hard to ignore. A thirty-two point nine percent improvement in recall rate, a forty point five percent reduction in workload, and an estimated.7 billion in annual savings. Those are the kind of results that get people's attention.

Jane: And the approach is so elegant. It doesn't ask the algorithm to do everything. It just asks it to do what it's good at, which is sorting the obvious cases. The hard cases still get the human touch.

Tom: That's the "harmonizing" part of the title. It's about finding the right balance between safety and speed, between automation and human judgment.

Jane: And I think that's the biggest takeaway for me. This isn't just a paper about medical devices. It's a blueprint for how to use AI in high-stakes decision-making. It shows that the best results come from collaboration, not replacement.

Tom: Absolutely. And the authors did their homework. They built a massive dataset, they developed a robust model, and they created a practical policy. It's a complete package.

Jane: It is. And while there are still questions to answer, like how to handle manufacturers who try to game the system, the foundation is solid. This is a real step forward.

Tom: So with that, we'll say goodbye to "Harmonizing Safety and Speed." It's a paper that could genuinely change how the FDA operates and how we think about AI in regulation.

Jane: And it sets the stage for our next paper, which I'm told is just as exciting. So stick around, everyone. Thanks for listening.

Tom: See you in a bit.

Mohammad Zhalechian, Soroush Saghafian, Omar Robles

Indiana University · Harvard University · Emerging Health Consulting

cs.LG, cs.HC, math.OC, stat.ML

Submitted: 2026-08-12

Updated: 2026-08-14

Comments: Accepted for publication in Management Science

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 69/100

The gist: - "Weighted Recall Score accounts for the timing of predicate recall events.

Key concepts

Five hundred ten(k) pathway
This is the current FDA system used to clear thousands of medical devices annually. The paper notes that while it is fast, a significant number of these cleared devices are later recalled, creating safety concerns.
Human-Algorithm Approach
This policy design uses AI to handle straightforward cases and flag risky ones. Complex or ambiguous cases are then sent to human experts for final review, balancing speed with necessary human judgment.
Predicate Devices
These are existing devices that a new device is claimed to be similar to when seeking clearance. The paper's model uses information about these predicates—such as their recall history and age—to predict the risk of the new device.

Terminology

Summary

Summary

The paper develops a combined human-algorithm approach to assist the United States Food and Drug Administration (FDA) in improving its 510(k) medical device clearance process by reducing recall risk and regulatory workload. The 510(k) pathway allows manufacturers to gain medical device approval by demonstrating substantial equivalence to a legally marketed device (a predicate device). The paper notes that the inherent ambiguity of this regulatory procedure has been associated with high recall among many devices cleared through this pathway, raising significant safety concerns.

The authors first develop machine learning methods to estimate the risk of recall of 510(k) medical devices based on the information available at the time of submission. They then propose a data-driven clearance policy that recommends acceptance, rejection, or deferral to FDA's committees for in-depth evaluation. The paper conducts an empirical study using a unique dataset of over 31,000 submissions assembled based on data sources from the FDA and Centers for Medicare and Medicaid Service (CMS).

Data and Methods

The dataset includes 510(k) applicant device submission data for the years 2008 to 2020 and FDA recall data for the years 2008 to 2021. The authors developed a text mining algorithm to identify the predicate devices listed by manufacturers for each cleared 510(k) applicant device using summary documents. The primary outcome is a binary recall event indicating whether the applicant device had at least one recall between its FDA clearance date and the end of the study period. The authors constructed 17 predictors for each applicant device, including continuous measures, binary indicator flags, and multi-level categorical variables.

Several machine learning models were trained and evaluated, including regularized logistic regression (Lasso and Ridge), decision tree, random forest, and gradient boosting. The gradient boosting model was selected, achieving a cross-validation Area Under the Curve (AUC) score of 0.78 and an out-of-sample AUC of 0.76. The paper notes that all models except the decision tree model attain relatively similar performance in terms of the CV-AUC metric, ranging from 0.77 to 0.78.

Key Findings on Predictors

Using SHAP analysis, the authors identified the most important variables for predicting recall risk. They found that the number of recalls reported for predicate devices, along with variables relevant to the age of the predicates, hold significant predictive power for a recall event. Specifically:

  • Weighted Recall Score accounts for the timing of predicate recall events. We observe that a higher value of this variable is associated with a higher recall risk.

  • Num. of Class 2 Recalls is highly predictive of the recall risk.

  • These observations not only indicate that the number of predicate recall events matters, but also highlight the importance of their timing, an aspect that has been overlooked in the literature.

  • Regarding predicate age: a higher value of Predicate Newest Age is associated with a lower recall risk and high values of Predicate Oldest Age can lead to a low recall risk.

  • Altogether, these findings highlight a learning curve effect. That is, if a predicate device has been on the market only briefly, latent problems may remain undiscovered that may lead to a higher risk of future recalls.

Proposed Policy

The paper proposes a data-driven clearance policy with two main components: (1) an ML predictor to estimate recall risk, and (2) an optimization approach that determines whether an applicant device can be accepted/rejected or should be deferred to an FDA-assigned committee. The policy uses two optimized thresholds (low threshold l and high threshold h):

  • If predicted risk is below l, the policy recommends accepting the device.

  • If predicted risk exceeds h, the policy recommends rejecting the device.

  • If predicted risk falls between l and h, the device is deferred for human expert evaluation.

The optimization model minimizes the weighted sum of rates of acceptance of unsafe devices and rejection of safe devices, subject to constraints on the rates of rejection of unsafe devices, acceptance of safe devices, and the FDA's workload. The authors developed a nested search algorithm to solve the optimization problem when the workload constraint is binding, and derived structural properties including a closed-form solution for relaxed versions of the problem.

Results

Compared to the FDA's current practice, which has a recall rate of 10.3% and a normalized workload measure of 100%, a conservative evaluation of the proposed policy shows a 32.9% improvement in the recall rate and a 40.5% reduction in the workload. The representative policy uses thresholds l = 0.060 and h = 0.177, resulting in a 12.1% rejection rate.

The paper also estimates potential cost savings: "Our analyses further suggest annual cost savings of approximately 1.7 billion for the healthcare system driven by avoided replacement costs, which is equivalent to 1.1% of the entire United States annual medical device expenditure." The cost analysis was based on Medicare allowed amounts from CMS administrative claims data for 2013-2020, with the authors determining the medical specialty for over 99% of the 1,351 unique HCPCS codes.

Managerial Insights

The paper concludes that there is a need for a combined human-algorithm approach, where devices with a mid-range predicted risk of recall (non-easy cases) are deferred to human experts for further evaluation. The findings address FDA concerns about the 510(k) process, including the importance of predicate recall history and age. The paper notes that best practices should take advantage of the fact that the number of recall events for predicates, particularly recent recalls, is a significant predictor of recall risk for an applicant device. The proposed policy is described as having four specific benefits: transparent structure, explainability, appropriate metric selection, and building upon human-in-the-loop decision-making frameworks.

Improvements for AI systems

Based on the paper, here are specific improvements for AI systems:

Improvement: Implement a three-tier decision system (accept/reject/defer) instead of binary classification. The AI identifies clear-cut cases for automated decisions and flags ambiguous ones for human review.

Capability: The system can now:

  • Automatically approve low-risk submissions (below threshold l)

  • Automatically reject high-risk submissions (above threshold h)

  • Defer mid-risk cases to human experts with supplementary risk information

  • Achieve 32.9% improvement in recall rate while reducing workload by 40.5%

Improvement: Add a workload constraint to the optimization objective, balancing decision accuracy against human review capacity.

Improvement: Incorporate predicate device characteristics and recall history as primary predictive features, particularly:

  • Weighted Recall Score (timing-weighted recall events)

  • Number of Class 2 recalls

  • Predicate age metrics (newest, median, oldest)

Improvement: Use the nested search algorithm (Algorithm 1) to find near-optimal decision thresholds when workload constraints are binding.

Improvement: Calibrate model performance and decision thresholds by medical specialty, as AUC varies from 0.65 to 0.85 across specialties.

Improvement: Account for time-based censoring in recall prediction by using survival analysis techniques (Cox proportional hazards with Lasso/Ridge penalties) as robustness checks.

Improvement: Integrate cost estimation into the decision framework, using Medicare allowed amounts by medical specialty.

Improvement: Use SHAP analysis to identify and communicate the top 15 predictors of recall risk.

Related papers