Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing
Ranking
Overall
79
Content
85
Popularity
66
Observed public metrics from 1 member.
Merged summary
TL;DR - This study finds that sparse MoE routing selects relevant experts with substantially overlapping representation subspaces, rather than geometrically complementary ones. This “coherent overlap” matters because geometric similarity alone does not imply expert redundancy or justify pruning.
- Across six MoE architectures, actual routes explain token representations better than matched alternatives despite substantial expert-subspace overlap.
- In 39 factorial comparisons, selected experts outperformed the strongest unselected rivals, while the actual prefix consistently narrowed that advantage.
- Adding later experts improved next-token prediction in 24 of 39 frozen-route comparisons; the remaining results were inconclusive.
- Controlled training favored Top-2 routing over Top-1 across all three seeds.
Sources (1)
Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing
Public signals
Hugging Face upvotes 5
TL;DR - This study finds that sparse MoE routing selects relevant experts with substantially overlapping representation subspaces, rather than geometrically complementary ones. This “coherent overlap” matters because geometric similarity alone does not imply expert redundancy or justify pruning.
- Across six MoE architectures, actual routes explain token representations better than matched alternatives despite substantial expert-subspace overlap.
- In 39 factorial comparisons, selected experts outperformed the strongest unselected rivals, while the actual prefix consistently narrowed that advantage.
- Adding later experts improved next-token prediction in 24 of 39 frozen-route comparisons; the remaining results were inconclusive.
- Controlled training favored Top-2 routing over Top-1 across all three seeds.