Signal or Noise? A Benchmark Study of Agent Skills in Web Development
Ranking
Observed public metrics from 1 member.
Merged summary
TL;DR - WebDev-Skills-Bench evaluates reusable coding-agent skills across 50 web projects and finds that injecting matched skills often hurts performance while sharply increasing token costs. The results suggest skill injection should be routed and audited for each skill-project-model combination rather than treated as universally beneficial.
- Across four models, target-skill injection reduced mean Pass@2 by 1.3%–4.2% and increased token costs by 72%–394%.
- Skills improved results in only 17%–36% of skill-project pairs, with weak ranking transfer between models.
- Length-matched irrelevant controls distinguished prompt-length distraction from cases where skill content itself misled the model.
- Helpful skills favored anti-pattern rules over example-heavy content, while losses were concentrated on easier, earlier tasks.
Sources (1)
Signal or Noise? A Benchmark Study of Agent Skills in Web Development
TL;DR - WebDev-Skills-Bench evaluates reusable coding-agent skills across 50 web projects and finds that injecting matched skills often hurts performance while sharply increasing token costs. The results suggest skill injection should be routed and audited for each skill-project-model combination rather than treated as universally beneficial.
- Across four models, target-skill injection reduced mean Pass@2 by 1.3%–4.2% and increased token costs by 72%–394%.
- Skills improved results in only 17%–36% of skill-project pairs, with weak ranking transfer between models.
- Length-matched irrelevant controls distinguished prompt-length distraction from cases where skill content itself misled the model.
- Helpful skills favored anti-pattern rules over example-heavy content, while losses were concentrated on easier, earlier tasks.