ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders
Merged summary
TL;DR - ICAE-Bench evaluates coding agents as interactive project builders, starting from ambiguous product requirements rather than fully specified tasks. It measures whether agents can clarify requirements, use tools, debug, and reconstruct working repository-level software.
- Tasks derive controlled ambiguity from real open-source repositories with executable behavior.
- An automated User Agent reveals grounded hidden constraints without adding requirements or leaking implementation details.
- Standardized black-box tests assess functional correctness alongside semantic, API, and structural fidelity.
- Diagnostics also evaluate design quality and the quality of agent-user interaction.
Sources (1)
ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders
TL;DR - ICAE-Bench evaluates coding agents as interactive project builders, starting from ambiguous product requirements rather than fully specified tasks. It measures whether agents can clarify requirements, use tools, debug, and reconstruct working repository-level software.
- Tasks derive controlled ambiguity from real open-source repositories with executable behavior.
- An automated User Agent reveals grounded hidden constraints without adding requirements or leaking implementation details.
- Standardized black-box tests assess functional correctness alongside semantic, API, and structural fidelity.
- Diagnostics also evaluate design quality and the quality of agent-user interaction.