When Model Merging Rivals Joint Multi-Task Reinforcement Learning: A Task-Vector Geometry Analysis
TL;DR - On the AppWorld agent benchmark, merging separately trained Qwen3-8B RL specialists matched joint multi-task reinforcement learning. Near-orthogonal task vectors explain why sophisticated merging methods performed similarly to simple averaging.
- TIES and RAM+ merges were statistically indistinguishable from joint RL on task-goal completion.
- Specialist task vectors had low cosine similarity (0.06–0.10) despite roughly 65% parameter-support overlap.
- Task-vector direction and support were decoupled, causing support- and sign-based methods to approximate uniform averaging.
- Calibration against random-init and same-run baselines indicated the shared direction reflected learning rather than low-rank parameterization.