🛰️ Daily AI Frontier
‹ back to 2026-08-06

Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?

Research LLM Agents

Ranking

Overall 71
Content 85
Popularity 38

Observed public metrics from 1 member.

Representative image for Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?

Merged summary

TL;DR - Skill-Use is a benchmark testing whether LLM agents can autonomously recognize, retrieve, and follow "skills" (structured procedure documents) under progressive disclosure; it finds current agents unreliable, with the best configuration scoring only 0.613 SU.

  • Decomposes skill use into three facets: Trigger (invokes the relevant skill), Compliance (follows the prescribed procedure), and Boundary (avoids forbidden operations); the combined SU score credits execution only after triggering.
  • Benchmark scope: 79 real skills paired with 177 executable tasks across nine domains, grounded in real files, run in isolated Docker sandboxes and graded by a trajectory-based rubric.
  • Evaluated eight LLMs under two agent harnesses; triggering and procedural compliance emerged as independent bottlenecks rather than a single failure mode.
  • Both absolute scores and model rankings shifted with the harness, implying skill use is a harness-conditioned capability, not a fixed model property — so harness choice must be reported in agent evaluations.

Sources (1)

Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?

arXiv cs.CL Jinyi Han, Yuanjian Xu, Ying Liao, Xinyi Wang, Zishang Jiang, Zixiang Di, Fanyang Lu, Zhichao Hu, Yanghua Xiao 2026-08-05 arXiv:2608.04828
Public signals Hugging Face upvotes 1
Providers: Hugging Face · Upvotes 1 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-03 14:31:56.780025 UTC

TL;DR - Skill-Use is a benchmark testing whether LLM agents can autonomously recognize, retrieve, and follow "skills" (structured procedure documents) under progressive disclosure; it finds current agents unreliable, with the best configuration scoring only 0.613 SU.

  • Decomposes skill use into three facets: Trigger (invokes the relevant skill), Compliance (follows the prescribed procedure), and Boundary (avoids forbidden operations); the combined SU score credits execution only after triggering.
  • Benchmark scope: 79 real skills paired with 177 executable tasks across nine domains, grounded in real files, run in isolated Docker sandboxes and graded by a trajectory-based rubric.
  • Evaluated eight LLMs under two agent harnesses; triggering and procedural compliance emerged as independent bottlenecks rather than a single failure mode.
  • Both absolute scores and model rankings shifted with the harness, implying skill use is a harness-conditioned capability, not a fixed model property — so harness choice must be reported in agent evaluations.
item →