🛰️ Daily AI Frontier
‹ back to 2026-08-26

Right Diagnoses, Decorative Reasoning:A Perturbation Audit of Medical Chain-of-Thought

arXiv cs.AI Medical/Healthcare AI Mengzhu Xu, Jifan Gao, Xia Jiang, Yaoxin Wu, Xi Long 2026-08-25

TL;DR - A perturbation audit of 14 LLMs finds that medical chain-of-thought rationales are often disconnected from the models’ diagnoses. This challenges the use of visible reasoning as evidence that a model reached an answer through clinically faithful reasoning.

  • The audit uses 30 clinically motivated operators, including severity reversals, negation flips, demographic swaps, and evidence ablations, across four medical QA benchmarks.
  • On clinically meaningful destructive edits, the panel-wide Chain-Decoupling Rate was 72.9%: models often neither updated the rationale nor changed the answer.
  • Corrupting the chain did not affect accuracy, and removing chain-of-thought prompting did not reduce accuracy.
  • Two board-certified clinicians reviewed 197 perturbed questions and found that 98.5% retained defensible gold answers, supporting the validity of the perturbation analysis.

view merged work →