Regulatory Approval Is Not Enough: Gaps in Trustworthy AI Reporting in FDA-Cleared Medical Devices

summary

Video file (mp4)

This episode discusses

The paper

Regulatory Approval Is Not Enough: Gaps in Trustworthy AI Reporting in FDA-Cleared Medical Devices · Read on arXiv

Ahmed M Salih, Oliver Díaz, Alejandro Guzman, Noah Marquez Vara, Fotios Avgoustidis, Rituraj Singh, Saman Barakat, Zahra Raisi-Estabragh, Karim Lekadir

University of Leicester · British Heart Foundation Leicester Centre of Research Excellence · Universitat de Barcelona · Chalmers University of Technology · Universidad de Sevilla · Queen Mary University of London · Institució Catalana de Recerca i Estudis Avançats

Background: AI/ML-enabled medical devices are increasingly deployed in healthcare under evolving regulatory frameworks. As these systems become more integrated into clinical decision-making, there is growing expectation that they demonstrate key dimensions of trustworthy AI to support clinician, patient, and public trust. Whether publicly available regulatory documentation provides sufficient evidence to independently assess the trustworthiness of cleared AI systems remains unclear. Methods: We analysed FDA AI/ML-enabled medical device summary reports published between 2021 and 2025. Reports underwent automated keyword screening followed by multi-stage manual consensus review to identify documented evidence for the six FUTURE-AI principles: Fairness, Universality, Traceability, Usability, Robustness, and Explainability. Descriptive, temporal, and clinical-domain analyses were performed. Multivariable logistic regression assessed whether year of clearance or clinical domain predicted higher reporting transparency, defined as evidence reported for three or more principles. Results: Of 1,105 FDA summary reports screened, 519 were included. Trustworthy AI reporting was limited and uneven. Nearly one quarter (24.7%) provided no evidence for any principle, and none documented evidence across all six. Robustness was most frequently reported (57.6%), while Traceability (8.3%) and Explainability (3.5%) were the most pronounced gaps. Neither year of clearance (OR 1.02, 95% CI 0.88-1.19) nor clinical domain (OR 0.73, 95% CI 0.46-1.15) predicted higher reporting transparency. Interpretation: Substantial, persistent trustworthy AI reporting gaps exist in FDA documentation. Regulatory approval alone should not be considered a proxy for trustworthiness. Standardised, audit-ready reporting across the AI lifecycle is needed to support independent assessment and responsible adoption of healthcare AI.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Regulatory Approval Is Not Enough: Gaps in Trustworthy AI Reporting in FDA-Cleared Medical Devices".

Jane: The paper was written by Ahmed M Salih, Oliver Díaz, Alejandro Guzman, Noah Marquez Vara, Fotios Avgoustidis et al. from University of Leicester and British Heart Foundation Leicester Centre of Research Excellence and Universitat de Barcelona and Chalmers University of Technology and Universidad de Sevilla and Queen Mary University of London and Institució Catalana de Recerca i Estudis Avançats.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that's going to make a lot of people in healthcare stop and think. It's called "Regulatory Approval Is Not Enough: Gaps in Trustworthy AI Reporting in FDA-Cleared Medical Devices."

Jane: And Tom, I have to say, that title is doing a lot of work. It's basically saying that just because the FDA clears a medical device that uses AI, that doesn't mean we actually know enough about it to trust it. That's a pretty bold claim.

Tom: It is, and it comes from a big international team — researchers from the University of Leicester, the University of Barcelona, Chalmers, Queen Mary University of London. They went through over a thousand FDA summary reports for AI-enabled medical devices cleared between two thousand twenty-one and two thousand twenty-five.

Jane: And what they found is honestly a bit concerning. They used this framework called FUTURE-AI, which lays out six principles that trustworthy AI in healthcare should follow: Fairness, Universality, Traceability, Usability, Robustness, and Explainability.

Tom: Right, and the paper's title is basically their conclusion in a nutshell. They found that regulatory approval doesn't tell you whether a device is actually trustworthy in all these dimensions. It just tells you it met some minimum bar for safety and effectiveness.

Jane: Exactly. And I think that's the key insight here. We've been treating FDA clearance as this gold standard, like if it's approved, it must be good. But this paper is saying, hold on, let's look at what's actually being disclosed publicly about these devices.

Tom: And spoiler alert, it's not a lot. Nearly a quarter of the reports they looked at had no evidence for any of the six trustworthy AI principles. Not one. And no report covered all six.

Jane: That's wild. So we're clearing devices that could be making decisions about patient care, and the public documentation doesn't even mention fairness or explainability in most cases. That's a transparency problem.

Tom: It really is. And the authors are careful to say this doesn't mean the devices are bad. It means we can't independently verify how trustworthy they are from the public record. And that's a problem for clinicians who have to decide whether to use these tools.

Jane: So the title is really a warning. It's saying, don't confuse regulatory approval with a stamp of trustworthiness. They're different things, and we need to treat them that way.

Tom: And we're going to dig into exactly what they found in the data next. Because the numbers tell a pretty stark story about which principles are being reported and which are being ignored.

Jane: I'm curious to see how that breaks down. Let's get into it.

Summary: Tom: So we're back with "Regulatory Approval Is Not Enough: Gaps in Trustworthy AI Reporting in FDA-Cleared Medical Devices." Last time we talked about the big picture — the title is a warning about confusing approval with trustworthiness. Now let's get into the actual findings.

Jane: And the findings are pretty stark. They started with one thousand one hundred five FDA summary reports, and after screening, they ended up with five hundred nineteen that they analyzed in depth. The team looked for evidence related to those six FUTURE-AI principles I mentioned: Fairness, Universality, Traceability, Usability, Robustness, and Explainability.

Tom: And the distribution is really uneven. Robustness was the most commonly reported principle — about fifty-eight percent of reports had some evidence for it. That makes sense, because regulators have always cared about whether a device performs reliably.

Jane: But then you look at the other end of the spectrum. Traceability was only in about eight percent of reports, and Explainability was in just three point five percent. That's almost nothing. So we have devices that can make predictions, but the public documentation rarely explains how those predictions are made or how the system can be audited.

Tom: And it's not like things are improving much over time. The paper looked at trends from two thousand twenty-one to two thousand twenty-five and while Fairness reporting did go up from about thirteen percent to thirty-four percent, other principles stayed flat or even declined. Usability actually dropped from thirty-six percent to sixteen percent.

Jane: That's a really important point, Tom. The authors ran a statistical analysis to see if year of clearance predicted higher reporting transparency — meaning evidence for three or more principles. And it didn't. The odds ratio was basically one and the p-value was zero point seven one. So no meaningful improvement over time.

Tom: Right, so all this talk about trustworthy AI in policy circles hasn't translated into better public reporting. And the same goes for clinical domain. Radiology devices weren't significantly better than other specialties. The gap is systemic.

Jane: And I think that's the most striking part. This isn't one bad manufacturer or one lazy specialty. It's a structural issue across the entire FDA clearance process for AI devices.

Tom: The authors also looked at what kind of evidence was being reported when principles were addressed. For robustness, it was mostly repeatability and reproducibility testing. For universality, external validation across multiple sites. But for fairness, it was usually just subgroup performance comparisons — not advanced bias mitigation.

Jane: And explainability was the worst. Only eighteen devices out of five hundred nineteen disclosed any explainability-related methods, and even then it was often just saliency maps or highlighted regions of interest. Nothing about how clinicians actually understand the model's reasoning.

Tom: So the summary is this: we're approving AI medical devices with very limited public evidence about their fairness, their explainability, their traceability, and their usability. And that's a problem for anyone who has to decide whether to trust these systems.

Jane: And it's a problem for patients too, because they're the ones ultimately affected by these decisions. But the paper doesn't just stop at identifying the problem. It has some concrete recommendations, and that's what we should talk about next.

Tom: Good point. Let's look at what they're proposing to fix this.

Improvements: Tom: We're still with "Regulatory Approval Is Not Enough: Gaps in Trustworthy AI Reporting in FDA-Cleared Medical Devices." We've covered the findings — the uneven reporting, the lack of improvement over time. Now let's talk about what the authors think should change.

Jane: And they have four main recommendations. The first one is pretty straightforward: they want comprehensive reporting across all six trustworthy AI principles, not just the ones that are easiest to measure. Right now, robustness gets reported because it's easy to quantify. Explainability doesn't, because it's harder.

Tom: Right, and that's a real imbalance. The paper argues that regulatory summaries should encourage balanced reporting across all dimensions, so clinicians can see the full picture of a device's strengths and limitations.

Jane: The second recommendation is about standardization. They're proposing a dedicated Trustworthy AI section in every FDA summary report, structured around the six FUTURE-AI principles. Think of it like the CONSORT-AI or TRIPOD-AI guidelines that already exist for clinical trials and prediction models.

Tom: That would make it so much easier to compare devices. Right now, every manufacturer uses different terminology, so it's hard to know if one device's "generalizability" is the same as another's "universality." A standardized template would fix that.

Jane: The third recommendation is about human-centered evaluation. The paper found that usability reporting often just references general human factors standards like IEC sixty-two thousand three hundred sixty-six but doesn't actually evaluate how clinicians interact with the AI. They want more evidence about workflow integration, clinician trust, and whether people actually understand the model outputs.

Tom: And that connects to the fourth recommendation, which is about the AI lifecycle. The authors make a really important point here: AI devices aren't static. They can change over time as they encounter new data, new patient populations, new clinical settings. So trustworthiness isn't something you establish once at approval.

Jane: Exactly. A model that's fair and robust when it's cleared might not stay that way after deployment. The paper argues that manufacturers should describe post-deployment monitoring strategies — tracking performance, detecting data drift, monitoring subgroup fairness, and having mitigation plans when things degrade.

Tom: And they've got a nice table in the paper that maps the FUTURE-AI principles to existing regulatory frameworks like the GMLP guidelines, the EU AI Act, and the US Executive Order. So it's not like these are pie-in-the-sky ideas. They align with what regulators are already saying.

Jane: Right, the principles are already there in policy documents. What's missing is the operationalization. The paper is essentially saying, you've told us these things matter, now let's actually see the evidence in the public record.

Tom: And I think that's the core contribution here. It's not just a critique. It's a roadmap for how to make regulatory reporting actually useful for clinicians, hospitals, and patients who need to make decisions about AI.

Jane: And it's a reminder that transparency isn't just a nice-to-have. It's what enables trust, and trust is what enables adoption. We'll wrap up with our final thoughts next.

Conclusion: Tom: Alright, we're wrapping up our discussion of "Regulatory Approval Is Not Enough: Gaps in Trustworthy AI Reporting in FDA-Cleared Medical Devices." And I think we should end by thinking about what this paper actually changes.

Jane: For me, the biggest takeaway is that we need to stop treating FDA approval as the end of the conversation. It's the beginning. The paper analyzed five hundred nineteen FDA summary reports and found that no single report covered all six trustworthy AI principles, and a quarter covered none at all. That's a transparency gap that affects real decisions.

Tom: And the authors were careful to say this doesn't mean the devices are unsafe. It means the public can't independently verify their trustworthiness. And that's a problem for clinicians who need to decide whether to use these tools, and for patients who are affected by those decisions.

Jane: The recommendations are really actionable too. Standardized reporting templates, balanced coverage across all principles, more human-centered evaluation, and lifecycle monitoring. These aren't radical ideas. They're practical steps that regulators and manufacturers could take.

Tom: And the paper's title really captures the message. Regulatory approval is necessary, but it's not sufficient. We need more than a stamp of approval. We need evidence that these systems are fair, explainable, traceable, usable, and robust — not just at approval, but throughout their entire lifecycle.

Jane: It's a sobering paper, but I think it's also an optimistic one. It shows that the tools for assessing trustworthy AI exist. The FUTURE-AI framework gives us a common language. What's missing is the will to require this kind of reporting.

Tom: And that's something we can change. Regulators can update their requirements. Manufacturers can voluntarily disclose more. Clinicians can demand better information. Patients can ask questions.

Jane: So as we say goodbye to this paper, I think the message is clear: trust in AI healthcare isn't automatic. It has to be earned through transparency. And this paper gives us a roadmap for how to do that.

Tom: Thanks for joining us, everyone. We'll be back next time with another paper from the arXiv. Until then, keep asking questions about the technology that's shaping healthcare.

Jane: Take care, everyone.

More episodes

← Home