Decomposing Error and Style in Automated Clinical Coding
TL;DR - Automated clinical coding evaluations often mistake systematic differences in coder or site practices for model errors. Explicitly conditioning models on a 10-dimensional “coding style” representation substantially improves ICD coding scores and narrows performance gaps between prompting methods.
- Two coding teams agreed on 73% of codes across 110 ACI-Bench encounters; independent clinical auditing raised agreement only to 77%.
- The study models coding as conditioned on both the clinical note and a coder- or site-specific policy governing what and how much to code.
- Across five datasets, matching the model to the data’s coding style increased ICD F1 by up to 26 points, while extreme style mismatch reduced it by up to 21 points.
- Four prompt-based methods initially scoring 39–49 F1 converged to 52–56 F1 when supplied with style information, with all reported differences statistically significant.