Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms
Ranking
Observed public metrics from 1 member.
Merged summary
TL;DR - This paper compares three paradigms for consolidating domain-specific RLVR capabilities into one language model: expert task-vector merging, mixed-domain RL, and multi-teacher on-policy distillation. Their average performance is similar, but substantial benchmark-level differences make the best choice dependent on cost, data balance, and capability-preservation priorities.
- Average performance differs by at most 1.4 points across paradigms, while individual benchmark gaps reach 8.6 points.
- Cross-domain relationships reflected in task-vector geometry help explain domain-level performance variation.
- Mix RL is sensitive to domain proportions, MOPD is bounded by teacher performance, and Merge compresses expert updates into a single model update.
- All three improve single-sample accuracy without measurable solution-coverage gains or degradation of held-out capabilities.
Sources (1)
Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms
TL;DR - This paper compares three paradigms for consolidating domain-specific RLVR capabilities into one language model: expert task-vector merging, mixed-domain RL, and multi-teacher on-policy distillation. Their average performance is similar, but substantial benchmark-level differences make the best choice dependent on cost, data balance, and capability-preservation priorities.
- Average performance differs by at most 1.4 points across paradigms, while individual benchmark gaps reach 8.6 points.
- Cross-domain relationships reflected in task-vector geometry help explain domain-level performance variation.
- Mix RL is sensitive to domain proportions, MOPD is bounded by teacher performance, and Merge compresses expert updates into a single model update.
- All three improve single-sample accuracy without measurable solution-coverage gains or degradation of held-out capabilities.