Right Diagnoses, Decorative Reasoning:A Perturbation Audit of Medical Chain-of-Thought
TL;DR - A perturbation audit of 14 LLMs finds that medical chain-of-thought rationales are often disconnected from the models’ diagnoses. This challenges the use of visible reasoning as evidence that a model reached an answer through clinically faithful reasoning.
- The audit uses 30 clinically motivated operators, including severity reversals, negation flips, demographic swaps, and evidence ablations, across four medical QA benchmarks.
- On clinically meaningful destructive edits, the panel-wide Chain-Decoupling Rate was 72.9%: models often neither updated the rationale nor changed the answer.
- Corrupting the chain did not affect accuracy, and removing chain-of-thought prompting did not reduce accuracy.
- Two board-certified clinicians reviewed 197 perturbed questions and found that 98.5% retained defensible gold answers, supporting the validity of the perturbation analysis.