StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
TL;DR - StartupBench evaluates general-purpose agents on end-to-end workflows derived from AI startup products with demonstrated market adoption. The strongest evaluated model completes only about 30% of tasks, indicating a substantial gap between partial progress and reliable real-world delivery.
- Tasks span diverse professional domains and require complete, deliverable-oriented outputs.
- Fine-grained rubrics assess the complex requirements of each workflow under a unified agent harness.
- Agents often make meaningful partial progress but fail to complete workflows successfully.
- Complex instruction following and domain-specific expertise are identified as major failure sources.