AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
Ranking
Overall
88
Content
95
Popularity
73
Observed public metrics from 1 member.
Merged summary
TL;DR - Strong models can build inference-time harnesses that transfer capabilities to weaker models without parameter updates. Across four Theory-of-Mind benchmarks, these harnesses nearly doubled average target-model performance from 0.49 to 0.91.
- Harnesses were iteratively refined using 5% of benchmark data, then evaluated on the full test set.
- Gains primarily came from deterministic code, benchmark-specific routing, and strict output formatting—not deeper reasoning or broader sampling.
- More builder-model reasoning consistently improved harness quality.
- Weaker target models benefited most, while platform effects were comparatively modest.
Sources (1)
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
Public signals
Hugging Face upvotes 115 · Semantic Scholar citations 1 · Semantic Scholar influential citations 0
TL;DR - Strong models can build inference-time harnesses that transfer capabilities to weaker models without parameter updates. Across four Theory-of-Mind benchmarks, these harnesses nearly doubled average target-model performance from 0.49 to 0.91.
- Harnesses were iteratively refined using 5% of benchmark data, then evaluated on the full test set.
- Gains primarily came from deterministic code, benchmark-specific routing, and strict output formatting—not deeper reasoning or broader sampling.
- More builder-model reasoning consistently improved harness quality.
- Weaker target models benefited most, while platform effects were comparatively modest.