🛰️ Daily AI Frontier
‹ back to 2026-09-24

WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents

Research LLM Agents

Ranking

Overall 86
Content 100
Popularity 54

Observed public metrics from 1 member.

Merged summary

TL;DR - WhatWorkedBench evaluates whether AI research agents can predict how code-component changes affect experimental outcomes after limited experimentation. Results show that Gaussian processes and code-equivalence information substantially improve effect recovery.

  • The benchmark spans 36 tasks, 30 data sources, eight workflow types, and 1,248 configurations.
  • Agents inspect code, choose measurements, and predict scores across all component configurations using response surfaces.
  • Gaussian-process fitting improved effect recovery from 0.632 to 0.698 in the original Flash cohort and from 0.621 to 0.720 in an additional cohort.
  • Encoding behaviorally equivalent configurations raised Gaussian-process recovery from 0.248 to 0.462 on workflows with six binary options.

Sources (1)

WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents

arXiv cs.AI Jingjie Ning, Xueqi Li, Yibo Kong, Dongting Li 2026-09-23 arXiv:2609.27490
Public signals Hugging Face upvotes 8
Providers: Hugging Face · Upvotes 8 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:16:29.605485 UTC

TL;DR - WhatWorkedBench evaluates whether AI research agents can predict how code-component changes affect experimental outcomes after limited experimentation. Results show that Gaussian processes and code-equivalence information substantially improve effect recovery.

  • The benchmark spans 36 tasks, 30 data sources, eight workflow types, and 1,248 configurations.
  • Agents inspect code, choose measurements, and predict scores across all component configurations using response surfaces.
  • Gaussian-process fitting improved effect recovery from 0.632 to 0.698 in the original Flash cohort and from 0.621 to 0.720 in an additional cohort.
  • Encoding behaviorally equivalent configurations raised Gaussian-process recovery from 0.248 to 0.462 on workflows with six binary options.
item →