One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL
Ranking
Overall
80
Content
100
Popularity
34
Observed public metrics from 1 member.
Merged summary
TL;DR - This paper identifies “simulator collapse,” where agent policies overfit to a single mode-collapsed LLM user simulator. Diversifying simulator behavior improves generalization to unseen simulators and real users.
- Verbalized Sampling broadens simulator responses at inference time, improving held-out success by up to 9%.
- Co-Training jointly trains policies against multiple simulators, increasing gains to 14%.
- Both methods preserve policy diversity across three multi-turn benchmarks.
- The authors release SCOPE, an open-source framework for population co-training in multi-agent RL.
Sources (1)
One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - This paper identifies “simulator collapse,” where agent policies overfit to a single mode-collapsed LLM user simulator. Diversifying simulator behavior improves generalization to unseen simulators and real users.
- Verbalized Sampling broadens simulator responses at inference time, improving held-out success by up to 9%.
- Co-Training jointly trains policies against multiple simulators, increasing gains to 14%.
- Both methods preserve policy diversity across three multi-turn benchmarks.
- The authors release SCOPE, an open-source framework for population co-training in multi-agent RL.