From Safety Documentation to Safety Knowledge Support: An Evidence-Grounded LLM Framework for Medical Devices
Tuhinangshu Gangopadhyay, Rasmus Adler, Peter Liggesmeyer, Jan Reich
Fraunhofer Institute for Experimental Software Engineering · RPTU Kaiserslautern-Landau
cs.SE, cs.AI
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: ISSRE 2026, AISQ, 8 pages
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: The paper argues that the central research problem for using large language models (LLMs) in medical-device safety analysis is not safety-text generation, but source-linked safety-knowledge support.
Terminology
Summary
The paper argues that the central research problem for using large language models (LLMs) in medical-device safety analysis is not safety-text generation, but source-linked safety-knowledge support. The authors propose an evidence-grounded framework that connects device artifacts, controlled knowledge storage and retrieval, method-specific generation of candidate safety items, critique and uncertainty checks, and recorded expert review. The framework prepares, links, checks, and updates candidate safety artifacts for expert decision-making; it does not decide whether a device is safe and does not provide regulatory approval.
The paper makes four contributions: (1) it summarizes recent patterns in LLM-assisted safety engineering and classifies them by safety task and dominant LLM method; (2) it identifies gaps that limit current approaches in regulated medical-device development; (3) it proposes a framework for evidence-grounded, traceable, reviewable, and updatable safety-knowledge support; and (4) it defines an evaluation strategy based on non-public or newly built medical-device case studies, expert reference analyses, and safety-specific metrics.
The paper reviews 90 papers published between 2023 and 2026 and identifies five main patterns: prompt engineering feasibility studies, retrieval-based grounding (RAG), knowledge-graph-based structured safety knowledge, multi-step and multi-agent pipelines, and LLMs used as critics (especially for assurance cases). The authors identify seven research gaps: weak domain grounding, limited hazard coverage, limited traceability, unsupported claims and weak scoring, unclear human role, weak evaluation, and limited lifecycle support.
The proposed framework consists of several process components: evidence intake, safety-knowledge storage and retrieval, task setup, method-specific generation, critique and uncertainty checks, and expert review and lifecycle updates. The central artifact is the source-linked safety item,
which contains fields such as item ID, item type, device link, source evidence, generated statement, rationale, assumptions, uncertainty flags, reviewer decision, reviewer justification, and change history.
The evaluation strategy compares three configurations (prompt-only generation, retrieval-grounded generation, and the full framework) using metrics including coverage, correctness, relevance, duplicate rate, source support, unsupported-claim rate, traceability, review effort, and review usefulness. The authors note that the full framework would be considered useful if it increases the share of relevant and source-supported safety items, reduces unsupported claims and duplicate items, improves traceability, and reduces expert search and consistency-checking effort, compared with the baselines.
The paper concludes that the intended industry benefit is strongest where manufacturers must maintain risk-management files across many device variants, software versions, and post-market feedback loops. The authors plan to implement and evaluate the framework in the MedSafe project, expecting a support system for testing whether medical-device safety analysis can become more efficient, reviewable, traceable, and updatable while preserving human responsibility.
Improvements for AI systems
Improvements to AI systems:
-
Source-linked generation with mandatory evidence fields: The AI system must generate safety items as structured objects with explicit fields for source evidence, device link, assumptions, and uncertainty flags—not free-text output. It can automatically reject or flag any generated claim lacking a verifiable source link.
-
Two-stage critique-and-uncertainty pipeline: The AI system runs a separate critic module after generation to check for unsupported claims, duplicates, and logical inconsistencies, and assigns uncertainty scores. It can then either revise the item or mark it for mandatory human review, reducing blind acceptance of AI output.
-
Lifecycle-aware memory with change history: The AI system maintains a persistent, versioned knowledge store of safety items, with each item tracking its full change history (who changed what, when, and why). It can automatically detect when new device artifacts or post-market feedback invalidate or require updates to existing safety items.
-
Task-specific generation methods instead of one-size-fits-all prompting: The AI system selects generation strategies based on the safety task type (e.g., hazard identification vs. assurance-case argumentation), using retrieval-augmented generation for evidence-heavy tasks and structured knowledge-graph traversal for relationship-heavy tasks. It can switch methods dynamically based on the input artifact.
-
Human-review-optimized output formatting: The AI system pre-organizes generated items into review-ready tables with pre-filled reviewer decision fields, justifications, and traceability links. It can highlight which items need expert attention (high uncertainty, unsupported claims) and which are low-risk, reducing expert search and consistency-checking effort.
-
Evaluation-driven self-improvement loop: The AI system can be benchmarked against the paper’s metrics (coverage, source support, unsupported-claim rate, duplicate rate, traceability) on internal case studies. It can then fine-tune its retrieval, generation, and critique components to maximize source-supported relevance while minimizing duplicates and unsupported claims.
What the improved AI system can do:
-
Produce safety-analysis artifacts (hazard lists, risk-control measures, assurance arguments) where every claim is traceable to a specific device artifact, standard, or post-market report, with automatic flagging of unsupported statements.
-
Maintain a continuously updated risk-management file across multiple device variants and software versions, automatically propagating changes from new evidence to affected safety items.
-
Act as a decision-support tool for human experts: it drafts, critiques, and revises candidate safety items, but never makes final safety decisions—it explicitly outputs uncertainty levels and review requests.
-
Reduce expert workload by pre-filtering duplicates, grouping related items, and suggesting source evidence for each candidate, while preserving full auditability for regulatory submissions.
Abstract
Medical devices are becoming more software-intensive, connected, and AI-enabled. Their development requires risk-management evidence aligned with ISO 14971 and, for software, IEC 62304. This evidence must be kept consistent across requirements, design decisions, software changes, verification results, complaints, and post-market data. These tasks are costly and depend on scarce safety and domain experts. Large language models (LLMs) may reduce parts of this effort because medical-device safety work is highly document-based. However, current LLM-based safety-engineering studies often address isolated methods, rely on generic prompting or public examples, and provide limited support for source links, traceability, uncertainty handling, lifecycle updates, and recorded expert review. This limits their use in regulated medical-device development. This paper argues that the central research problem is not safety-text generation, but source-linked safety-knowledge support. We propose an evidence-grounded framework that connects device artifacts, controlled knowledge storage and retrieval, method-specific generation of candidate safety items, critique and uncertainty checks, and recorded expert review. The framework prepares, links, checks, and updates candidate safety artifacts for expert decision-making. It does not decide whether a device is safe and does not provide regulatory approval. We also outline an evaluation strategy using non-public or newly built medical-device case studies and expert reference analyses to assess coverage, correctness, relevance, traceability, duplicate rate, unsupported claims, and review effort.
Sources
- Supporting Risk Management for Medical Devices via the Riskman Ontology and Shapes (Preprint)
- An LLM-Integrated Framework for Completion, Management, and Tracing of STPA
- Evaluating the Effectiveness of GPT-4 Turbo in Creating Defeaters for Assurance Cases
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties