Regulatory Approval Is Not Enough: Gaps in Trustworthy AI Reporting in FDA-Cleared Medical Devices
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Regulatory Approval Is Not Enough: Gaps in Trustworthy AI Reporting in FDA-Cleared Medical Devices".
Jane: The paper was written by Ahmed M Salih, Oliver Díaz, Alejandro Guzman, Noah Marquez Vara, Fotios Avgoustidis et al. from University of Leicester and British Heart Foundation Leicester Centre of Research Excellence and Universitat de Barcelona and Chalmers University of Technology and Universidad de Sevilla and Queen Mary University of London and Institució Catalana de Recerca i Estudis Avançats.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that's going to make a lot of people in healthcare stop and think. It's called "Regulatory Approval Is Not Enough: Gaps in Trustworthy AI Reporting in FDA-Cleared Medical Devices."
Jane: And Tom, I have to say, that title is doing a lot of work. It's basically saying that just because the FDA clears a medical device that uses AI, that doesn't mean we actually know enough about it to trust it. That's a pretty bold claim.
Tom: It is, and it comes from a big international team — researchers from the University of Leicester, the University of Barcelona, Chalmers, Queen Mary University of London. They went through over a thousand FDA summary reports for AI-enabled medical devices cleared between two thousand twenty-one and two thousand twenty-five.
Jane: And what they found is honestly a bit concerning. They used this framework called FUTURE-AI, which lays out six principles that trustworthy AI in healthcare should follow: Fairness, Universality, Traceability, Usability, Robustness, and Explainability.
Tom: Right, and the paper's title is basically their conclusion in a nutshell. They found that regulatory approval doesn't tell you whether a device is actually trustworthy in all these dimensions. It just tells you it met some minimum bar for safety and effectiveness.
Jane: Exactly. And I think that's the key insight here. We've been treating FDA clearance as this gold standard, like if it's approved, it must be good. But this paper is saying, hold on, let's look at what's actually being disclosed publicly about these devices.
Tom: And spoiler alert, it's not a lot. Nearly a quarter of the reports they looked at had no evidence for any of the six trustworthy AI principles. Not one. And no report covered all six.
Jane: That's wild. So we're clearing devices that could be making decisions about patient care, and the public documentation doesn't even mention fairness or explainability in most cases. That's a transparency problem.
Tom: It really is. And the authors are careful to say this doesn't mean the devices are bad. It means we can't independently verify how trustworthy they are from the public record. And that's a problem for clinicians who have to decide whether to use these tools.
Jane: So the title is really a warning. It's saying, don't confuse regulatory approval with a stamp of trustworthiness. They're different things, and we need to treat them that way.
Tom: And we're going to dig into exactly what they found in the data next. Because the numbers tell a pretty stark story about which principles are being reported and which are being ignored.
Jane: I'm curious to see how that breaks down. Let's get into it.
Summary: Tom: So we're back with "Regulatory Approval Is Not Enough: Gaps in Trustworthy AI Reporting in FDA-Cleared Medical Devices." Last time we talked about the big picture — the title is a warning about confusing approval with trustworthiness. Now let's get into the actual findings.
Jane: And the findings are pretty stark. They started with one thousand one hundred five FDA summary reports, and after screening, they ended up with five hundred nineteen that they analyzed in depth. The team looked for evidence related to those six FUTURE-AI principles I mentioned: Fairness, Universality, Traceability, Usability, Robustness, and Explainability.
Tom: And the distribution is really uneven. Robustness was the most commonly reported principle — about fifty-eight percent of reports had some evidence for it. That makes sense, because regulators have always cared about whether a device performs reliably.
Jane: But then you look at the other end of the spectrum. Traceability was only in about eight percent of reports, and Explainability was in just three point five percent. That's almost nothing. So we have devices that can make predictions, but the public documentation rarely explains how those predictions are made or how the system can be audited.
Tom: And it's not like things are improving much over time. The paper looked at trends from two thousand twenty-one to two thousand twenty-five and while Fairness reporting did go up from about thirteen percent to thirty-four percent, other principles stayed flat or even declined. Usability actually dropped from thirty-six percent to sixteen percent.
Jane: That's a really important point, Tom. The authors ran a statistical analysis to see if year of clearance predicted higher reporting transparency — meaning evidence for three or more principles. And it didn't. The odds ratio was basically one and the p-value was zero point seven one. So no meaningful improvement over time.
Tom: Right, so all this talk about trustworthy AI in policy circles hasn't translated into better public reporting. And the same goes for clinical domain. Radiology devices weren't significantly better than other specialties. The gap is systemic.
Jane: And I think that's the most striking part. This isn't one bad manufacturer or one lazy specialty. It's a structural issue across the entire FDA clearance process for AI devices.
Tom: The authors also looked at what kind of evidence was being reported when principles were addressed. For robustness, it was mostly repeatability and reproducibility testing. For universality, external validation across multiple sites. But for fairness, it was usually just subgroup performance comparisons — not advanced bias mitigation.
Jane: And explainability was the worst. Only eighteen devices out of five hundred nineteen disclosed any explainability-related methods, and even then it was often just saliency maps or highlighted regions of interest. Nothing about how clinicians actually understand the model's reasoning.
Tom: So the summary is this: we're approving AI medical devices with very limited public evidence about their fairness, their explainability, their traceability, and their usability. And that's a problem for anyone who has to decide whether to trust these systems.
Jane: And it's a problem for patients too, because they're the ones ultimately affected by these decisions. But the paper doesn't just stop at identifying the problem. It has some concrete recommendations, and that's what we should talk about next.
Tom: Good point. Let's look at what they're proposing to fix this.
Improvements: Tom: We're still with "Regulatory Approval Is Not Enough: Gaps in Trustworthy AI Reporting in FDA-Cleared Medical Devices." We've covered the findings — the uneven reporting, the lack of improvement over time. Now let's talk about what the authors think should change.
Jane: And they have four main recommendations. The first one is pretty straightforward: they want comprehensive reporting across all six trustworthy AI principles, not just the ones that are easiest to measure. Right now, robustness gets reported because it's easy to quantify. Explainability doesn't, because it's harder.
Tom: Right, and that's a real imbalance. The paper argues that regulatory summaries should encourage balanced reporting across all dimensions, so clinicians can see the full picture of a device's strengths and limitations.
Jane: The second recommendation is about standardization. They're proposing a dedicated Trustworthy AI section in every FDA summary report, structured around the six FUTURE-AI principles. Think of it like the CONSORT-AI or TRIPOD-AI guidelines that already exist for clinical trials and prediction models.
Tom: That would make it so much easier to compare devices. Right now, every manufacturer uses different terminology, so it's hard to know if one device's "generalizability" is the same as another's "universality." A standardized template would fix that.
Jane: The third recommendation is about human-centered evaluation. The paper found that usability reporting often just references general human factors standards like IEC sixty-two thousand three hundred sixty-six but doesn't actually evaluate how clinicians interact with the AI. They want more evidence about workflow integration, clinician trust, and whether people actually understand the model outputs.
Tom: And that connects to the fourth recommendation, which is about the AI lifecycle. The authors make a really important point here: AI devices aren't static. They can change over time as they encounter new data, new patient populations, new clinical settings. So trustworthiness isn't something you establish once at approval.
Jane: Exactly. A model that's fair and robust when it's cleared might not stay that way after deployment. The paper argues that manufacturers should describe post-deployment monitoring strategies — tracking performance, detecting data drift, monitoring subgroup fairness, and having mitigation plans when things degrade.
Tom: And they've got a nice table in the paper that maps the FUTURE-AI principles to existing regulatory frameworks like the GMLP guidelines, the EU AI Act, and the US Executive Order. So it's not like these are pie-in-the-sky ideas. They align with what regulators are already saying.
Jane: Right, the principles are already there in policy documents. What's missing is the operationalization. The paper is essentially saying, you've told us these things matter, now let's actually see the evidence in the public record.
Tom: And I think that's the core contribution here. It's not just a critique. It's a roadmap for how to make regulatory reporting actually useful for clinicians, hospitals, and patients who need to make decisions about AI.
Jane: And it's a reminder that transparency isn't just a nice-to-have. It's what enables trust, and trust is what enables adoption. We'll wrap up with our final thoughts next.
Conclusion: Tom: Alright, we're wrapping up our discussion of "Regulatory Approval Is Not Enough: Gaps in Trustworthy AI Reporting in FDA-Cleared Medical Devices." And I think we should end by thinking about what this paper actually changes.
Jane: For me, the biggest takeaway is that we need to stop treating FDA approval as the end of the conversation. It's the beginning. The paper analyzed five hundred nineteen FDA summary reports and found that no single report covered all six trustworthy AI principles, and a quarter covered none at all. That's a transparency gap that affects real decisions.
Tom: And the authors were careful to say this doesn't mean the devices are unsafe. It means the public can't independently verify their trustworthiness. And that's a problem for clinicians who need to decide whether to use these tools, and for patients who are affected by those decisions.
Jane: The recommendations are really actionable too. Standardized reporting templates, balanced coverage across all principles, more human-centered evaluation, and lifecycle monitoring. These aren't radical ideas. They're practical steps that regulators and manufacturers could take.
Tom: And the paper's title really captures the message. Regulatory approval is necessary, but it's not sufficient. We need more than a stamp of approval. We need evidence that these systems are fair, explainable, traceable, usable, and robust — not just at approval, but throughout their entire lifecycle.
Jane: It's a sobering paper, but I think it's also an optimistic one. It shows that the tools for assessing trustworthy AI exist. The FUTURE-AI framework gives us a common language. What's missing is the will to require this kind of reporting.
Tom: And that's something we can change. Regulators can update their requirements. Manufacturers can voluntarily disclose more. Clinicians can demand better information. Patients can ask questions.
Jane: So as we say goodbye to this paper, I think the message is clear: trust in AI healthcare isn't automatic. It has to be earned through transparency. And this paper gives us a roadmap for how to do that.
Tom: Thanks for joining us, everyone. We'll be back next time with another paper from the arXiv. Until then, keep asking questions about the technology that's shaping healthcare.
Jane: Take care, everyone.
Ahmed M Salih, Oliver Díaz, Alejandro Guzman, Noah Marquez Vara, Fotios Avgoustidis, Rituraj Singh, Saman Barakat, Zahra Raisi-Estabragh, Karim Lekadir
University of Leicester · British Heart Foundation Leicester Centre of Research Excellence · Universitat de Barcelona · Chalmers University of Technology · Universidad de Sevilla · Queen Mary University of London · Institució Catalana de Recerca i Estudis Avançats
cs.CY, cs.LG
Submitted: 2026-07-04
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 68/100
Terminology
Summary
Summary
This paper presents the first systematic assessment of trustworthy AI reporting in FDA-cleared medical device summary reports, using the FUTURE-AI framework as a structured, multi-dimensional lens. The study analyzed 1,105 FDA AI/ML-enabled medical device summary reports published between 2021 and 2025, of which 519 were included after automated keyword screening and multi-stage manual consensus review by seven researchers. The analysis evaluated documented evidence related to six FUTURE-AI principles: Fairness, Universality, Traceability, Usability, Robustness, and Explainability.
The study found that trustworthy AI reporting was limited and uneven. Nearly one quarter of reports (24.7%, 128/519) provided no evidence for any principle, and no report documented evidence across all six principles. Most reports provided evidence for only one (29.1%, 151/519) or two principles (21.6%, 112/519). Reporting across three principles was observed in only 17.3% of reports (90/519), while evidence spanning four (6.7%, 35/519) or five principles (0.6%, 3/519) was rare.
Robustness was the most frequently documented principle, with evidence identified in 57.6% of reports (299/519). Lower reporting rates were observed for Universality (33.5%, 174/519), Usability (26.2%, 136/519), and Fairness (25.0%, 130/519). Traceability and Explainability represented the most pronounced reporting gaps, with evidence disclosed in only 8.3% (43/519) and 3.5% of reports (18/519), respectively. Explainability was the least reported dimension, with only five submissions explicitly using the term explainability
and 18 devices disclosing explainability-related methods, most commonly through generic visualizations, saliency maps, or highlighted regions of interest.
Temporal trend analyses revealed that Robustness remained the most consistently reported principle throughout the study period, ranging between 50.3% and 67.5%. Fairness showed an upward trajectory, rising from 12.8% in 2021 to 33.5% in 2025, while Universality peaked at 45.5% in 2024 before declining sharply. Usability exhibited a steady decline from 36.2% to 16.2%. Traceability remained persistently low and volatile, and Explainability never exceeded 5.5% and dropped to 0.0% in 2023.
Multivariable logistic regression using High Transparency
(defined as reporting evidence for ≥ 3 trustworthy AI principles) as the dependent variable showed that neither year of FDA clearance (OR 1.02, 95% CI 0.88–1.19, p=0.71) nor clinical domain (OR 0.73, 95% CI 0.46–1.15, p=0.17) were significantly associated with higher reporting transparency. The study concluded that the identified transparency gap is systemic rather than confined to specific clinical specialties or time periods.
The paper highlights that current reporting practices prioritize model performance, technical validation, and documentation, while standardized evaluation and reporting approaches for fairness, explainability, and other human-centered trustworthy AI dimensions remain limited. The authors emphasize that regulatory approval alone should not be considered a proxy for trustworthiness, and they propose four recommendations: (1) comprehensive reporting across all trustworthy AI principles, (2) standardized trustworthy AI reporting templates and terminology, (3) greater emphasis on human-centred evaluation, and (4) assessing trustworthiness throughout the AI lifecycle, including post-deployment monitoring strategies.
The study acknowledges limitations, including that the analysis was restricted to publicly available FDA summary reports and did not include confidential regulatory submissions or internal manufacturer documentation, and that the binary assessment framework may not fully capture differences in the depth or quality of reported evidence.
Improvements for AI systems
Based on the paper's findings, here are specific improvements I can implement in AI systems for medical devices:
-
Current gap: 24.7% of devices report zero trustworthy AI principles; none report all six
-
Implementation: Build a structured reporting template that forces documentation across all six FUTURE-AI dimensions (Fairness, Universality, Traceability, Usability, Robustness, Explainability) before regulatory submission
-
What the improved system can do: Automatically flag incomplete submissions and generate compliance scores per principle
-
Current gap: Only 3.5% of devices report explainability evidence; 0% in neurology and ophthalmology
-
Implementation: Integrate model-agnostic explainability methods (SHAP, LIME, saliency maps) directly into the AI pipeline, with automatic generation of feature importance reports and visualization outputs
-
What the improved system can do: Produce interpretable outputs for every prediction, including highlighted regions of interest and confidence scores, with standardized documentation for regulatory review
-
Current gap: Only 8.3% of devices document traceability
-
Implementation: Build immutable logging of data provenance, model versions, training configurations, and validation splits; implement automatic version control and audit trail generation
-
What the improved system can do: Provide complete lineage from raw data to final prediction, enabling regulators and clinicians to trace any output back to its source data and model version
-
Current gap: Fairness reporting is low (25%) and rarely includes bias metrics
-
Implementation: Add automated subgroup performance analysis across demographic variables (age, sex, race, ethnicity), with statistical tests for performance disparities and automatic bias alerts
-
What the improved system can do: Continuously monitor fairness metrics post-deployment, flagging when subgroup performance diverges beyond predefined thresholds
-
Current gap: No evidence of temporal stability assessment or monitoring strategies in most reports
-
Implementation: Build continuous monitoring for data drift, population shift, and performance degradation using statistical process control methods
-
What the improved system can do: Automatically detect when model inputs or performance deviate from validation conditions, triggering alerts for recalibration or human oversight
-
Current gap: Usability reporting declining (36.2% to 16.2%); often limited to generic standards compliance
-
Implementation: Integrate clinician workflow simulation and decision-support effectiveness testing into the development cycle
-
What the improved system can do: Generate evidence of how clinicians interact with AI outputs, including time-to-decision, override rates, and trust calibration metrics
-
Current gap: Heterogeneous terminology across manufacturers
-
Implementation: Build an automated report generator that maps internal validation evidence to standardized FUTURE-AI terminology and formats
-
What the improved system can do: Produce consistent, comparable regulatory summaries that allow clinicians to directly compare trustworthiness across competing devices
-
Current gap: No device provides evidence across all six principles; trustworthiness treated as static
-
Implementation: Create a composite trustworthiness score that updates dynamically based on ongoing monitoring data
-
What the improved system can do: Provide a real-time, transparent trustworthiness indicator that reflects current performance across all six dimensions, not just approval-time validation
These improvements would transform AI medical devices from approved but opaque
systems into continuously auditable, explainable, and fairness-aware tools that clinicians can independently evaluate and trust.
Abstract
Background: AI/ML-enabled medical devices are increasingly deployed in healthcare under evolving regulatory frameworks. As these systems become more integrated into clinical decision-making, there is growing expectation that they demonstrate key dimensions of trustworthy AI to support clinician, patient, and public trust. Whether publicly available regulatory documentation provides sufficient evidence to independently assess the trustworthiness of cleared AI systems remains unclear. Methods: We analysed FDA AI/ML-enabled medical device summary reports published between 2021 and 2025. Reports underwent automated keyword screening followed by multi-stage manual consensus review to identify documented evidence for the six FUTURE-AI principles: Fairness, Universality, Traceability, Usability, Robustness, and Explainability. Descriptive, temporal, and clinical-domain analyses were performed. Multivariable logistic regression assessed whether year of clearance or clinical domain predicted higher reporting transparency, defined as evidence reported for three or more principles. Results: Of 1,105 FDA summary reports screened, 519 were included. Trustworthy AI reporting was limited and uneven. Nearly one quarter (24.7%) provided no evidence for any principle, and none documented evidence across all six. Robustness was most frequently reported (57.6%), while Traceability (8.3%) and Explainability (3.5%) were the most pronounced gaps. Neither year of clearance (OR 1.02, 95% CI 0.88-1.19) nor clinical domain (OR 0.73, 95% CI 0.46-1.15) predicted higher reporting transparency. Interpretation: Substantial, persistent trustworthy AI reporting gaps exist in FDA documentation. Regulatory approval alone should not be considered a proxy for trustworthiness. Standardised, audit-ready reporting across the AI lifecycle is needed to support independent assessment and responsible adoption of healthcare AI.
Related papers
- Reasoning Enhances Robustness to Prompt Injection in LLM-Based Consensus
- Generative AI Purpose-built for Social and Mental Health: A Real-World Pilot
- PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks
- What is an intelligent system?
- AI University: An LLM-Powered Learning Assistant for Engineering---A Finite Element Method Case Study
- Generative AI Use in Entrepreneurship: An Integrative Review and an Empowerment-Entrapment Framework