SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation
- Published
- Source
- arXiv
- Paper number
- 796
- Field
- AI / General
- arXiv ID
- 2608.02287
Key points
- From 2,000 public skills, it created 4,000 task packages and 27,164 verified execution trajectories, and it also built the SkillEval benchmark from a separate skill pool that does not overlap with the training data using the same pipeline.
- As the skill budget used for training increases from 0 to 100, 500, 1000, and 2000, the Qwen3.5-9B SkillEval score rises monotonically to 55.24, 59.8, 67.4, 70.8, and 72.48.
- The benefit transfers across execution environments: a model trained on OpenCode trajectories improves from 4.94 to 9.29 on SkillsBench in DeepAgents, and training on same-environment trajectories raises it to 13.62, meaning it retains 50.1 percent of that gain.
- A single checkpoint trained on mixed trajectories from two environments beats the original models in all 8 paired comparisons between benchmarks and execution environments, with gains ranging from 2.64 to 18.81 points. SkillEval rises from 55.24 to 74.05 on OpenCode and from 51.62 to 69.96 on DeepAgents.
- The mixed-trained model stays within an average of 2.71 points of the environment-specific model and even surpasses it in 3 of the 8 comparisons, which means separate models per environment are less necessary.
- When the 77 SkillsBench tasks are grouped by the number of required skills, one-skill tasks improve from 9.54 to 16.94, two-skill tasks from 3.95 to 20.72, and tasks with three or more skills from 4.04 to 12.08, so the gain is not limited to single-skill tasks.
Paper links
External research summaries. These are not HDATF publications or measured product results.