AgenticQwen: Training Small Agentic Language Models with Dual Data Flywheels for Industrial-Scale Tool Use
- Published
- Source
- arXiv
- Paper number
- 161
- Field
- LLMs / Agents / Efficiency
- arXiv ID
- 2604.21590
Key points
- Industrial deployment of LLM-based agents is constrained by the high compute cost and latency associated with very large or proprietary foundation models.
- There is a lack of small language models that are specialized and optimized for robust multi-step agent capabilities and tool interaction.
- Static synthetic datasets used for agent reinforcement learning often cause data homogeneity and rapid saturation of the learning signal, limiting further gains.
- The paper trains small AgenticQwen models, including 8B and 30B-A3B variants, with a multi-round reinforcement learning approach that combines Reasoning RL and Agentic RL.
- It introduces a dual data flywheel mechanism, composed of a Reasoning Data Flywheel and an Agentic Data Flywheel, to iteratively generate increasingly complex and diverse training data.
- A strong LLM, Qwen3-235B, is used to create a fully simulated training environment that models both the user and the tools, removing dependence on external proprietary APIs for data synthesis and simulation.
Paper links
External research summaries. These are not HDATF publications or measured product results.