🛰️ Daily AI Frontier
‹ back to 2026-08-27

Making Clinical Language Models Auditable: Concept-Guided Fine-Tuning for Robust Prediction

Research Medical/Healthcare AI

Ranking

Overall 78
Content 95
Popularity 37

Observed public metrics from 1 member.

Merged summary

TL;DR - CAST is an SAE-based fine-tuning framework that makes clinical text classifiers more robust and auditable by identifying and suppressing features tied to note artifacts rather than patient state. On MIMIC-IV mortality prediction, it improves over corresponding fine-tuned encoder baselines while providing concept-level audit trails.

  • Sparse autoencoders expose human-auditable features from intermediate Transformer activations.
  • An LLM-assisted pipeline with ICD-10 retrieval constraints labels SAE latents as clinical concepts or artifacts.
  • Verified artifact latents are suppressed through residual subtraction during fine-tuning.
  • Post-hoc attributions show which clinical concepts supported each prediction and which artifacts were suppressed.

Sources (1)

Making Clinical Language Models Auditable: Concept-Guided Fine-Tuning for Robust Prediction

arXiv cs.CL Jin Mu, Guanhua Chen 2026-08-27 arXiv:2608.27397
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-24 14:32:37.850560 UTC

TL;DR - CAST is an SAE-based fine-tuning framework that makes clinical text classifiers more robust and auditable by identifying and suppressing features tied to note artifacts rather than patient state. On MIMIC-IV mortality prediction, it improves over corresponding fine-tuned encoder baselines while providing concept-level audit trails.

  • Sparse autoencoders expose human-auditable features from intermediate Transformer activations.
  • An LLM-assisted pipeline with ICD-10 retrieval constraints labels SAE latents as clinical concepts or artifacts.
  • Verified artifact latents are suppressed through residual subtraction during fine-tuning.
  • Post-hoc attributions show which clinical concepts supported each prediction and which artifacts were suppressed.
item →