🛰️ Daily AI Frontier
‹ back to 2026-08-28

Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms

arXiv cs.CL LLMs & Foundation Models Siye Wu, Kai Yang, Yuchen Cai, Xin Xu, Peng-Yuan Wang, Jiaxuan Wang, Jiashun Liu, Jiafei Lyu, Yangkun Chen, Saiyong Yang, Yanghua Xiao 2026-08-27

TL;DR - This paper compares three paradigms for consolidating domain-specific RLVR capabilities into one language model: expert task-vector merging, mixed-domain RL, and multi-teacher on-policy distillation. Their average performance is similar, but substantial benchmark-level differences make the best choice dependent on cost, data balance, and capability-preservation priorities.

  • Average performance differs by at most 1.4 points across paradigms, while individual benchmark gaps reach 8.6 points.
  • Cross-domain relationships reflected in task-vector geometry help explain domain-level performance variation.
  • Mix RL is sensitive to domain proportions, MOPD is bounded by teacher performance, and Merge compresses expert updates into a single model update.
  • All three improve single-sample accuracy without measurable solution-coverage gains or degradation of held-out capabilities.

view merged work →