🛰️ Daily AI Frontier
‹ back to 2026-08-04

SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation

arXiv cs.AI LLM Agents Zelin Tan, Yiqun Zhang, Hao Li, Zhiyao Cui, Hejia Geng, Shao Zhang, Hangfan Zhang, Yang Chen, Xiaosong Wang, Lilong Wang, Zhenfei Yin, Shuyue Hu, Chen Zhang, Lei Bai 2026-08-03
Representative image for SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation

TL;DR - SKT is a verified synthetic-data pipeline that generates skill-grounded tasks and executable trajectories from large pools of agent skills, so LLM agents can be fine-tuned to actually identify, apply, and coordinate reusable skills rather than merely being handed them.

  • Pipeline selects single- and multi-skill configurations, synthesizes tasks with rule-based plus agent-based verification and feedback-guided repair, and keeps only successful trajectories that substantively exercise every required skill.
  • From 2,000 public skills it produced 4,000 task packages and 27,164 verified trajectories; a disjoint test pool yields SkillEval, a held-out executable benchmark for skill use.
  • Supervised fine-tuning on SKT trajectories consistently improved skill-use performance across multiple models, benchmarks, and agent harnesses.
  • Ablations indicate gains hinge on high-quality verified supervision, transfer across agent interfaces rather than one harness, and scale with broader skill coverage.

view merged work →