MedDeID enables locally governed clinical-text de-identification from real or synthetic training data

arXiv:2609.10049 · cs.CL, cs.LG · Submitted 2026-09-09 · Read on arXiv

cs.CL, cs.LG

Submitted: 2026-09-09

Updated: 2026-09-09

Comments: 64 pages total: 32-page main manuscript with 4 figures, followed by 32-page Supplementary Information

License: http://creativecommons.org/licenses/by/4.0/

The gist: Clinical notes contain personally identifiable information (PII), restricting reuse for research and medical AI, especially when data cannot leave an institution.

Terminology

Abstract

Clinical notes contain personally identifiable information (PII), restricting reuse for research and medical AI, especially when data cannot leave an institution. We developed MedDeID, an on-premises framework combining in-house annotation and synthetic-note generation with model training, inference, pseudonymisation and evaluation. On an independently annotated, adjudicated 300-note Dutch hospital benchmark, a hospital-trained compact transformer detected 98.9% of identifying text while redacting 0.24% of text outside annotated identifiers; a synthetic-only counterpart detected 96.1%. On 100 primary-care notes, the synthetic-trained model achieved higher recall than the hospital-trained model (90.3% versus 87.0%) and greater robustness to identifier-format perturbations. An English instantiation trained without real text detected 99.7% and 98.9% of annotated identifier characters on two external synthetic benchmarks. These results demonstrate transfer of the workflow to another language, but not clinical English performance. MedDeID provides a route to locally governed de-identification using real or synthetic training data.

Related papers