🛰️ Daily AI Frontier
‹ back to 2026-08-24

Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization

arXiv cs.AI Medical/Healthcare AI Praphul Singh, Shanu Kumar, Akshat Agarwal 2026-08-21
Representative image for Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization

TL;DR - This paper audits the weight changes between general-purpose LLMs and their medical-specialized counterparts, finding that decoder updates closely reproduce benchmark gains but do not yield a simple component-level explanation of specialization.

  • Examines aligned Gemma-3-to-MedGemma and Qwen2.5-to-HuatuoGPT-o1 checkpoint pairs.
  • Full decoder deltas strongly reconstruct medical benchmark movement, with endpoint-normalized retention of 0.974 and 1.183.
  • MLP layers are the strongest broad component family in both pairs, but controls and rollback tests prevent uniquely attributing the gains to them.
  • Conclusions are limited to text-only multiple-choice benchmarks and do not establish clinical validity or circuit-level mechanisms.

view merged work →