SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation
Ranking
Overall
76
Content
80
Popularity
68
Observed public metrics from 1 member.
Merged summary
TL;DR - SKT is a verified synthetic-data pipeline that generates skill-grounded tasks and executable trajectories from large pools of agent skills, so LLM agents can be fine-tuned to actually identify, apply, and coordinate reusable skills rather than merely being handed them.
- Pipeline selects single- and multi-skill configurations, synthesizes tasks with rule-based plus agent-based verification and feedback-guided repair, and keeps only successful trajectories that substantively exercise every required skill.
- From 2,000 public skills it produced 4,000 task packages and 27,164 verified trajectories; a disjoint test pool yields SkillEval, a held-out executable benchmark for skill use.
- Supervised fine-tuning on SKT trajectories consistently improved skill-use performance across multiple models, benchmarks, and agent harnesses.
- Ablations indicate gains hinge on high-quality verified supervision, transfer across agent interfaces rather than one harness, and scale with broader skill coverage.
Sources (1)
SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation
Public signals
Hugging Face upvotes 31
TL;DR - SKT is a verified synthetic-data pipeline that generates skill-grounded tasks and executable trajectories from large pools of agent skills, so LLM agents can be fine-tuned to actually identify, apply, and coordinate reusable skills rather than merely being handed them.
- Pipeline selects single- and multi-skill configurations, synthesizes tasks with rule-based plus agent-based verification and feedback-guided repair, and keeps only successful trajectories that substantively exercise every required skill.
- From 2,000 public skills it produced 4,000 task packages and 27,164 verified trajectories; a disjoint test pool yields SkillEval, a held-out executable benchmark for skill use.
- Supervised fine-tuning on SKT trajectories consistently improved skill-use performance across multiple models, benchmarks, and agent harnesses.
- Ablations indicate gains hinge on high-quality verified supervision, transfer across agent interfaces rather than one harness, and scale with broader skill coverage.