Making Clinical Language Models Auditable: Concept-Guided Fine-Tuning for Robust Prediction
Ranking
Overall
78
Content
95
Popularity
37
Observed public metrics from 1 member.
Merged summary
TL;DR - CAST is an SAE-based fine-tuning framework that makes clinical text classifiers more robust and auditable by identifying and suppressing features tied to note artifacts rather than patient state. On MIMIC-IV mortality prediction, it improves over corresponding fine-tuned encoder baselines while providing concept-level audit trails.
- Sparse autoencoders expose human-auditable features from intermediate Transformer activations.
- An LLM-assisted pipeline with ICD-10 retrieval constraints labels SAE latents as clinical concepts or artifacts.
- Verified artifact latents are suppressed through residual subtraction during fine-tuning.
- Post-hoc attributions show which clinical concepts supported each prediction and which artifacts were suppressed.
Sources (1)
Making Clinical Language Models Auditable: Concept-Guided Fine-Tuning for Robust Prediction
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - CAST is an SAE-based fine-tuning framework that makes clinical text classifiers more robust and auditable by identifying and suppressing features tied to note artifacts rather than patient state. On MIMIC-IV mortality prediction, it improves over corresponding fine-tuned encoder baselines while providing concept-level audit trails.
- Sparse autoencoders expose human-auditable features from intermediate Transformer activations.
- An LLM-assisted pipeline with ICD-10 retrieval constraints labels SAE latents as clinical concepts or artifacts.
- Verified artifact latents are suppressed through residual subtraction during fine-tuning.
- Post-hoc attributions show which clinical concepts supported each prediction and which artifacts were suppressed.