🛰️ Daily AI Frontier
‹ back to 2026-09-22

Decomposing Error and Style in Automated Clinical Coding

arXiv cs.CL Medical/Healthcare AI Han-Chin Shing, Jack Moriarty, Ryan Ware, Afton Marchbanks, Carlyn Canvasser, Stefanie Higgins, Harsh Gupta, Fang Wang, Joseph Paul Cohen 2026-09-21

TL;DR - Automated clinical coding evaluations often mistake systematic differences in coder or site practices for model errors. Explicitly conditioning models on a 10-dimensional “coding style” representation substantially improves ICD coding scores and narrows performance gaps between prompting methods.

  • Two coding teams agreed on 73% of codes across 110 ACI-Bench encounters; independent clinical auditing raised agreement only to 77%.
  • The study models coding as conditioned on both the clinical note and a coder- or site-specific policy governing what and how much to code.
  • Across five datasets, matching the model to the data’s coding style increased ICD F1 by up to 26 points, while extreme style mismatch reduced it by up to 21 points.
  • Four prompt-based methods initially scoring 39–49 F1 converged to 52–56 F1 when supplied with style information, with all reported differences statistically significant.

view merged work →