Does Splitting a Triage Decision Across Agents Hide Bias or Help Catch It? A Multi-Agent Simulation Study of LLM-Based Resource Allocation Under Audit Capacity Constraints
Ranking
Observed public metrics from 1 member.
Merged summary
TL;DR - An arXiv simulation study tests whether splitting a life-or-death triage decision across a role-differentiated multi-agent LLM pipeline (assessment, allocation, independent audit) reduces demographic bias versus a single agent, and finds it does not — what actually matters is whether the auditor has capacity to review cases at all. It matters because oversight layers are widely assumed to catch bias, but here they only help if audit coverage holds under load.
- Setup: synthetic disaster-triage simulator with paired cases identical except one demographic attribute; 192 episodes / 2,304 resolved case pairs on GPT-4o-mini, single-agent control vs. nine-agent pipeline, across three independently varied pressure dimensions.
- Pipeline structure did not change bias incidence: 6.9% vs. 6.1% biased outcomes (p = 0.498).
- Audit capacity drove detection: 30.0% of biased outcomes went entirely undetected overall, rising to 43.8% with an overloaded auditor and falling to 18.4% when not overloaded.
- The loss came from coverage, not judgment quality — review coverage collapsed 100.0% → 65.6% under load (p < 0.001) while judgment on reviewed cases held (81.6% vs. 85.7%, p = 1.000, direction reversed); risk-ordered audit queues recovered coverage to 91.7% (p = 0.028) at the same capacity. Authors note limits: one model, modest samples, no adversarial replication.
Sources (1)
Does Splitting a Triage Decision Across Agents Hide Bias or Help Catch It? A Multi-Agent Simulation Study of LLM-Based Resource Allocation Under Audit Capacity Constraints
TL;DR - An arXiv simulation study tests whether splitting a life-or-death triage decision across a role-differentiated multi-agent LLM pipeline (assessment, allocation, independent audit) reduces demographic bias versus a single agent, and finds it does not — what actually matters is whether the auditor has capacity to review cases at all. It matters because oversight layers are widely assumed to catch bias, but here they only help if audit coverage holds under load.
- Setup: synthetic disaster-triage simulator with paired cases identical except one demographic attribute; 192 episodes / 2,304 resolved case pairs on GPT-4o-mini, single-agent control vs. nine-agent pipeline, across three independently varied pressure dimensions.
- Pipeline structure did not change bias incidence: 6.9% vs. 6.1% biased outcomes (p = 0.498).
- Audit capacity drove detection: 30.0% of biased outcomes went entirely undetected overall, rising to 43.8% with an overloaded auditor and falling to 18.4% when not overloaded.
- The loss came from coverage, not judgment quality — review coverage collapsed 100.0% → 65.6% under load (p < 0.001) while judgment on reviewed cases held (81.6% vs. 85.7%, p = 1.000, direction reversed); risk-ordered audit queues recovered coverage to 91.7% (p = 0.028) at the same capacity. Authors note limits: one model, modest samples, no adversarial replication.