Testing Interchangeability in LLM Agent Teams
TL;DR - This study tests whether role-matched agents can be exchanged between established LLM teams without harming performance. Swaps barely affect task scores but increase communication per unit of progress by 16–63%, showing that learned team conventions make agents less interchangeable than outcomes alone suggest.
- Eight independently formed teams retained private notebooks over ten formation episodes before agents were swapped and evaluated on held-out tasks.
- In Hanabi, a swapped agent was costlier than an inexperienced agent, consistent with interference from conventions learned with its previous partner.
- In Collab-Overcooked, replacing the agenda-setting agent caused most additional communication to come from the teammate who remained.
- Swap penalties tracked how far independently formed teams diverged: greedy decoding reduced both, while longer team histories increased both.