🛰️ Daily AI Frontier
‹ back to 2026-08-13

Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents

Research LLM Agents

Ranking

Overall 90
Content 100
Popularity 67

Observed public metrics from 1 member.

Merged summary

TL;DR - This paper analyzes how reusable skills can cause LLM-agent failures and higher execution costs. Across two benchmarks, it identifies 307 skill-induced regressions and introduces SkillTriage for evidence-based failure attribution.

  • The study finds 125 functional failures and 182 efficiency regressions.
  • Seemingly relevant skills can cause agents to implement requirements incorrectly or omit them.
  • Efficiency regressions are not attributable to prompt length alone.
  • Excessive verification and heavy implementation pipelines account for 67 and 30 excessive-procedure cases, respectively.

Sources (1)

Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents

arXiv cs.AI Gen Dong, Yanjie Gao, Liqun Li, Tianyin Xu, Yu Hua, Fan Yang 2026-08-12 arXiv:2608.11888
Public signals Hugging Face upvotes 3 · Semantic Scholar citations 1 · Semantic Scholar influential citations 0
Providers: Hugging Face · Upvotes 3 OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 1 · Influential citations 0 X · N/A Fetched 2026-09-12 14:28:12.942080 UTC

TL;DR - This paper analyzes how reusable skills can cause LLM-agent failures and higher execution costs. Across two benchmarks, it identifies 307 skill-induced regressions and introduces SkillTriage for evidence-based failure attribution.

  • The study finds 125 functional failures and 182 efficiency regressions.
  • Seemingly relevant skills can cause agents to implement requirements incorrectly or omit them.
  • Efficiency regressions are not attributable to prompt length alone.
  • Excessive verification and heavy implementation pipelines account for 67 and 30 excessive-procedure cases, respectively.
item →