🛰️ Daily AI Frontier
‹ back to 2026-08-22

MaliciousSkillBench: A Comprehensive Benchmark for Malicious Agent Skill Detection

Research LLM Agents

Ranking

Overall 83
Content 100
Popularity 45

Observed public metrics from 1 member.

Representative image for MaliciousSkillBench: A Comprehensive Benchmark for Malicious Agent Skill Detection

Merged summary

TL;DR - MaliciousSkillBench is a benchmark for detecting malicious reusable skill packages in LLM agents. Its evaluations reveal that detectors performing well on random splits generalize poorly to unseen sources and often over-flag benign skills.

  • The benchmark consolidates 13 public sources, normalizing 8,414 raw malicious records into 7,539 unique identities across 4,588 structural families.
  • Its primary dataset contains 9,740 skills: 7,505 malicious and 2,235 benign, with 11 harmonized attack categories covering 4,983 malicious identities.
  • Learned text detectors reach 0.882–0.932 Macro-F1 on random splits but only 0.653–0.665 under source-disjoint evaluation.
  • The strongest TF-IDF SVM retains 95.6% malicious recall on held-out sources but incurs a 62.4% benign false-positive rate; off-the-shelf scanners reduce false positives only by sharply sacrificing recall.

Sources (1)

MaliciousSkillBench: A Comprehensive Benchmark for Malicious Agent Skill Detection

arXiv cs.CR Yue Wang, Yi Liu, Gelei Deng, Ying Zhang, Yuekang Li, Zhenyu Chen, Leo Zhang 2026-08-20 arXiv:2608.19901
Public signals Hugging Face upvotes 0
Providers: Hugging Face · Upvotes 0 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-21 14:33:37.289322 UTC

TL;DR - MaliciousSkillBench is a benchmark for detecting malicious reusable skill packages in LLM agents. Its evaluations reveal that detectors performing well on random splits generalize poorly to unseen sources and often over-flag benign skills.

  • The benchmark consolidates 13 public sources, normalizing 8,414 raw malicious records into 7,539 unique identities across 4,588 structural families.
  • Its primary dataset contains 9,740 skills: 7,505 malicious and 2,235 benign, with 11 harmonized attack categories covering 4,983 malicious identities.
  • Learned text detectors reach 0.882–0.932 Macro-F1 on random splits but only 0.653–0.665 under source-disjoint evaluation.
  • The strongest TF-IDF SVM retains 95.6% malicious recall on held-out sources but incurs a 62.4% benign false-positive rate; off-the-shelf scanners reduce false positives only by sharply sacrificing recall.
item →