🛰️ Daily AI Frontier
‹ back to 2026-09-10

MedDeID enables locally governed clinical-text de-identification from real or synthetic training data

Research Medical/Healthcare AI

Ranking

Overall 79
Content 95
Popularity 42

Observed public metrics from 1 member.

Representative image for MedDeID enables locally governed clinical-text de-identification from real or synthetic training data

Merged summary

TL;DR - MedDeID is an on-premises framework for training and deploying clinical-text de-identification models using real or synthetic notes. Its synthetic-only models achieved strong identifier recall, offering institutions a path to reuse sensitive clinical data without transferring it externally.

  • A compact transformer trained on Dutch hospital data detected 98.9% of identifying text while redacting 0.24% of text outside annotated identifiers.
  • The Dutch synthetic-only model detected 96.1% of identifiers and outperformed the hospital-trained model on primary-care notes in recall (90.3% versus 87.0%).
  • Synthetic training improved robustness to changes in identifier formatting.
  • An English synthetic-only implementation reached 99.7% and 98.9% character-level detection on two synthetic benchmarks, but clinical English performance was not established.

Sources (1)

MedDeID enables locally governed clinical-text de-identification from real or synthetic training data

arXiv cs.CL Stig Hellemans, Tom Stroobants, Elyne Scheurwegs, Pieter Meysman, Philippe G. Jorens, Kris Laukens 2026-09-09 arXiv:2609.10049
Public signals Hugging Face upvotes 0
Providers: Hugging Face · Upvotes 0 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:21:30.987425 UTC

TL;DR - MedDeID is an on-premises framework for training and deploying clinical-text de-identification models using real or synthetic notes. Its synthetic-only models achieved strong identifier recall, offering institutions a path to reuse sensitive clinical data without transferring it externally.

  • A compact transformer trained on Dutch hospital data detected 98.9% of identifying text while redacting 0.24% of text outside annotated identifiers.
  • The Dutch synthetic-only model detected 96.1% of identifiers and outperformed the hospital-trained model on primary-care notes in recall (90.3% versus 87.0%).
  • Synthetic training improved robustness to changes in identifier formatting.
  • An English synthetic-only implementation reached 99.7% and 98.9% character-level detection on two synthetic benchmarks, but clinical English performance was not established.
item →