Grounding Healthcare LLMs in a Causal Knowledge Graph: Framework, Metrics, and a Cardiovascular Pilot
TL;DR - This paper introduces a causal-knowledge-graph framework for evaluating whether healthcare LLMs ground intervention recommendations in mechanisms, harms, evidence, and uncertainty. A cardiovascular pilot shows that raw answer accuracy can conceal weak causal and evidential reasoning.
- The framework preserves provenance using assertions as graph nodes with stable identifiers and extracts scenario-specific subgraphs.
- Four conditions compare ungrounded, knowledge-graph, causal-graph, and integrated grounding.
- Integrated grounding achieved the best causal-edge F1 (0.838), adverse-effect F1 (0.833), evidence accuracy (0.738), and unsupported-claim rate (0.114).
- The ungrounded condition had the highest intervention accuracy (0.948) despite no measurable causal or evidential grounding.