🛰️ Daily AI Frontier
‹ back to 2026-08-26

Right Diagnoses, Decorative Reasoning:A Perturbation Audit of Medical Chain-of-Thought

Research Medical/Healthcare AI

Ranking

Overall 80
Content 100
Popularity 34

Observed public metrics from 1 member.

Merged summary

TL;DR - A perturbation audit of 14 LLMs finds that medical chain-of-thought rationales are often disconnected from the models’ diagnoses. This challenges the use of visible reasoning as evidence that a model reached an answer through clinically faithful reasoning.

  • The audit uses 30 clinically motivated operators, including severity reversals, negation flips, demographic swaps, and evidence ablations, across four medical QA benchmarks.
  • On clinically meaningful destructive edits, the panel-wide Chain-Decoupling Rate was 72.9%: models often neither updated the rationale nor changed the answer.
  • Corrupting the chain did not affect accuracy, and removing chain-of-thought prompting did not reduce accuracy.
  • Two board-certified clinicians reviewed 197 perturbed questions and found that 98.5% retained defensible gold answers, supporting the validity of the perturbation analysis.

Sources (1)

Right Diagnoses, Decorative Reasoning:A Perturbation Audit of Medical Chain-of-Thought

arXiv cs.AI Mengzhu Xu, Jifan Gao, Xia Jiang, Yaoxin Wu, Xi Long 2026-08-25 arXiv:2608.24790
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-14 14:17:12.187916 UTC

TL;DR - A perturbation audit of 14 LLMs finds that medical chain-of-thought rationales are often disconnected from the models’ diagnoses. This challenges the use of visible reasoning as evidence that a model reached an answer through clinically faithful reasoning.

  • The audit uses 30 clinically motivated operators, including severity reversals, negation flips, demographic swaps, and evidence ablations, across four medical QA benchmarks.
  • On clinically meaningful destructive edits, the panel-wide Chain-Decoupling Rate was 72.9%: models often neither updated the rationale nor changed the answer.
  • Corrupting the chain did not affect accuracy, and removing chain-of-thought prompting did not reduce accuracy.
  • Two board-certified clinicians reviewed 197 perturbed questions and found that 98.5% retained defensible gold answers, supporting the validity of the perturbation analysis.
item →