Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering
Ranking
Observed public metrics from 1 member.
Merged summary
TL;DR - An arXiv preprint presenting nMAS, a multi-agent LLM pipeline that automates evidence-linked, rubric-grounded feature engineering from EHR data for heart-failure phenotyping. It matters because feature engineering consumes 39–45% of data scientists' workload, and this approach adds auditability and provenance that rule-based or plain-LLM methods lack.
- Evaluated on 500 dummy patient records spanning nine EHR source tables; produced 132 structured features and 70 rubric-scored aggregated features, checked for structural integrity, rubric compliance, and provenance, with a restricted LLM performing the audit.
- Adding aggregated features raised held-out AUROC from 0.895 → 0.963 (HFrEF) and 0.870 → 0.910 (HFpEF).
- An independent LLM rubric assessment of evidence support and methodological soundness scored the features at 81.5% of maximum points.
- Limitations stated by the authors: single-institution cohort and dummy records only; external validation still needed.
Sources (1)
Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering
TL;DR - An arXiv preprint presenting nMAS, a multi-agent LLM pipeline that automates evidence-linked, rubric-grounded feature engineering from EHR data for heart-failure phenotyping. It matters because feature engineering consumes 39–45% of data scientists' workload, and this approach adds auditability and provenance that rule-based or plain-LLM methods lack.
- Evaluated on 500 dummy patient records spanning nine EHR source tables; produced 132 structured features and 70 rubric-scored aggregated features, checked for structural integrity, rubric compliance, and provenance, with a restricted LLM performing the audit.
- Adding aggregated features raised held-out AUROC from 0.895 → 0.963 (HFrEF) and 0.870 → 0.910 (HFpEF).
- An independent LLM rubric assessment of evidence support and methodological soundness scored the features at 81.5% of maximum points.
- Limitations stated by the authors: single-institution cohort and dummy records only; external validation still needed.