AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
TL;DR - Strong models can build inference-time harnesses that transfer capabilities to weaker models without parameter updates. Across four Theory-of-Mind benchmarks, these harnesses nearly doubled average target-model performance from 0.49 to 0.91.
- Harnesses were iteratively refined using 5% of benchmark data, then evaluated on the full test set.
- Gains primarily came from deterministic code, benchmark-specific routing, and strict output formatting—not deeper reasoning or broader sampling.
- More builder-model reasoning consistently improved harness quality.
- Weaker target models benefited most, while platform effects were comparatively modest.