What is Missing from AI Post-Training AI: An Empirical Analysis
TL;DR - An empirical study finds that LLM agents can execute post-training workflows effectively but rarely reconsider the training strategy chosen at the outset. This limits AI-for-AI systems because they optimize locally rather than adapting their high-level approach as evidence accumulates.
- Agents consistently locked in an initial strategy and spent subsequent budgets on local adjustments.
- Experience-driven scaffolding improved execution by 12.6 points on GSM8K and 40.8 points on HumanEval, but did not induce strategy changes.
- Human guidance redirected initial choices, yet agents reverted to local adjustment loops once training began.
- Additional inference compute helped on easier tasks but produced almost no gain on the hardest task, suggesting the missing capability is spontaneous strategy reevaluation.