Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay
TL;DR - An executed-replay audit in ALFWorld finds that common step-level credit signals for training LLM agents identify causally important actions no better than chance. This challenges correctness-based credit evaluations and shows that training comparisons must control for effective sample size.
- Causal contribution was sparse: only 30.5% of measurable decision points affected outcomes.
- LLM judges, outcome-conditioned log-probability ratios, and policy confidence failed to recover causally pivotal steps above chance.
- Implicit credit primarily tracked policy fluency, while outcome conditioning added essentially no causal information.
- Across seven training arms, none reliably beat the untrained policy; apparent differences were explained by training dose rather than credit quality.