OpenThoughts-Agent: Data Recipes for Agentic Models
- Published
- Source
- arXiv
- Paper number
- 486
- Field
- AI / General
- arXiv ID
- 2606.24855
Key points
- The authors run more than 100 controlled ablation studies on a six-stage SFT data pipeline to quantify the effect of each stage.
- Instruction selection is the most important factor in the pipeline, which shows that the strongest model is not necessarily the best teacher.
- Fine-tuning Qwen3-32B on 100K examples yields 54.0 percent on SWE-Bench Verified and 26.2 percent on Terminal-Bench 2.0.
- The average score is 44.8 percent, which is 3.9 points better than the previous best open-data agent model, Nemotron-Terminal-32B at 40.9 percent.
- It shows scaling behavior that beats competing open datasets across all training-set sizes.
- It confirms that an 8B model trained with SFT and RL in two stages outperforms both pure SFT and pure RL.
Paper links
External research summaries. These are not HDATF publications or measured product results.