$Φ$-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?
TL;DR - Φ-Bench evaluates whether LLMs can engineer and optimize the infrastructure that powers language models, moving beyond isolated coding tasks to realistic, long-horizon systems work. It highlights remaining barriers to autonomous AI infrastructure optimization.
- Tasks are derived from frontier optimization problems and grounded in real-world code repositories.
- Coverage ranges from kernel-level function completion to extended implementation and end-to-end system optimization.
- The benchmark spans multiple layers of the LLM infrastructure stack and varying levels of complexity.
- Experiments with frontier LLMs characterize current capabilities and limitations, though the abstract provides no specific performance results.