🛰️ Daily AI Frontier
‹ back to 2026-09-22

Decomposing Error and Style in Automated Clinical Coding

Research Medical/Healthcare AI

Ranking

Overall 82
Content 95
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Merged summary

TL;DR - Automated clinical coding evaluations often mistake systematic differences in coder or site practices for model errors. Explicitly conditioning models on a 10-dimensional “coding style” representation substantially improves ICD coding scores and narrows performance gaps between prompting methods.

  • Two coding teams agreed on 73% of codes across 110 ACI-Bench encounters; independent clinical auditing raised agreement only to 77%.
  • The study models coding as conditioned on both the clinical note and a coder- or site-specific policy governing what and how much to code.
  • Across five datasets, matching the model to the data’s coding style increased ICD F1 by up to 26 points, while extreme style mismatch reduced it by up to 21 points.
  • Four prompt-based methods initially scoring 39–49 F1 converged to 52–56 F1 when supplied with style information, with all reported differences statistically significant.

Sources (1)

Decomposing Error and Style in Automated Clinical Coding

arXiv cs.CL Han-Chin Shing, Jack Moriarty, Ryan Ware, Afton Marchbanks, Carlyn Canvasser, Stefanie Higgins, Harsh Gupta, Fang Wang, Joseph Paul Cohen 2026-09-21 arXiv:2609.24877
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-26 14:14:41.372127 UTC

TL;DR - Automated clinical coding evaluations often mistake systematic differences in coder or site practices for model errors. Explicitly conditioning models on a 10-dimensional “coding style” representation substantially improves ICD coding scores and narrows performance gaps between prompting methods.

  • Two coding teams agreed on 73% of codes across 110 ACI-Bench encounters; independent clinical auditing raised agreement only to 77%.
  • The study models coding as conditioned on both the clinical note and a coder- or site-specific policy governing what and how much to code.
  • Across five datasets, matching the model to the data’s coding style increased ICD F1 by up to 26 points, while extreme style mismatch reduced it by up to 21 points.
  • Four prompt-based methods initially scoring 39–49 F1 converged to 52–56 F1 when supplied with style information, with all reported differences statistically significant.
item →