🛰️ Daily AI Frontier
‹ back to 2026-07-24

The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and What Actually Works

Research LLM Agents

Merged summary

TL;DR - Dense prediction rewards can collapse GRPO-trained LLM agents because groupwise standard-deviation normalization amplifies reward variance in all-fail groups. Delivering the same signal through an auxiliary loss instead improved ALFWorld performance by roughly 20 points.

  • Across Qwen3-1.7B/4B/8B, prediction accuracy reached 1.0 while task success fell to 0% and episodes hit the maximum horizon.
  • Removing only GRPO’s standard-deviation normalization restored performance to baseline parity.
  • The authors propose that dense signals are safe when their within-group variance declines as the model masters them.
  • Results are currently single-seed; replication and group-size controls are preregistered and ongoing.

Sources (1)

The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and What Actually Works

arXiv cs.LG Yu Wang 2026-07-23 arXiv:2607.21273 doi:10.5281/zenodo.21505228

TL;DR - Dense prediction rewards can collapse GRPO-trained LLM agents because groupwise standard-deviation normalization amplifies reward variance in all-fail groups. Delivering the same signal through an auxiliary loss instead improved ALFWorld performance by roughly 20 points.

  • Across Qwen3-1.7B/4B/8B, prediction accuracy reached 1.0 while task success fell to 0% and episodes hit the maximum horizon.
  • Removing only GRPO’s standard-deviation normalization restored performance to baseline parity.
  • The authors propose that dense signals are safe when their within-group variance declines as the model masters them.
  • Results are currently single-seed; replication and group-size controls are preregistered and ongoing.
item →