Memorisation bias in medical AI
TL;DR - This paper identifies “memorisation bias,” where medical AI predictions on a patient’s future records are altered because the model previously saw that patient’s anonymized historical data. The effect can distort clinical accuracy and persist for decades, creating risks when training-data contributors later return for care.
- Memorisation bias appears across multiple data modalities and model architectures.
- For new conditions absent from a patient’s historical training records, diagnostic sensitivity decreased.
- For unchanged health states, both sensitivity and specificity were artificially inflated.
- Current de-identification practices hinder identifying returning contributors, suggesting training and deployment protocols may need revision.