🛰️ Daily AI Frontier
‹ back to 2026-08-13

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL

Research LLM Agents

Ranking

Overall 80
Content 100
Popularity 34

Observed public metrics from 1 member.

Representative image for One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL

Merged summary

TL;DR - This paper identifies “simulator collapse,” where agent policies overfit to a single mode-collapsed LLM user simulator. Diversifying simulator behavior improves generalization to unseen simulators and real users.

  • Verbalized Sampling broadens simulator responses at inference time, improving held-out success by up to 9%.
  • Co-Training jointly trains policies against multiple simulators, increasing gains to 14%.
  • Both methods preserve policy diversity across three multi-turn benchmarks.
  • The authors release SCOPE, an open-source framework for population co-training in multi-agent RL.

Sources (1)

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL

arXiv cs.CL Simon Yu, Nicholas Tomlin, Marwa Abdulhai, Ximing Lu, Derek Chong, Abe Hou, Dilara Soylu, Sergey Levine, Christopher D. Manning, Weiyan Shi 2026-08-12 arXiv:2608.12253
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-31 14:23:00.629065 UTC

TL;DR - This paper identifies “simulator collapse,” where agent policies overfit to a single mode-collapsed LLM user simulator. Diversifying simulator behavior improves generalization to unseen simulators and real users.

  • Verbalized Sampling broadens simulator responses at inference time, improving held-out success by up to 9%.
  • Co-Training jointly trains policies against multiple simulators, increasing gains to 14%.
  • Both methods preserve policy diversity across three multi-turn benchmarks.
  • The authors release SCOPE, an open-source framework for population co-training in multi-agent RL.
item →