AgenticQwen: Training Small Agentic Language Models with Dual Data Flywheels for Industrial-Scale Tool Use

Published
Source
arXiv
Paper number
161
Field
LLMs / Agents / Efficiency
arXiv ID
2604.21590

Key points

  • Industrial deployment of LLM-based agents is constrained by the high compute cost and latency associated with very large or proprietary foundation models.
  • There is a lack of small language models that are specialized and optimized for robust multi-step agent capabilities and tool interaction.
  • Static synthetic datasets used for agent reinforcement learning often cause data homogeneity and rapid saturation of the learning signal, limiting further gains.
  • The paper trains small AgenticQwen models, including 8B and 30B-A3B variants, with a multi-round reinforcement learning approach that combines Reasoning RL and Agentic RL.
  • It introduces a dual data flywheel mechanism, composed of a Reasoning Data Flywheel and an Agentic Data Flywheel, to iteratively generate increasingly complex and diverse training data.
  • A strong LLM, Qwen3-235B, is used to create a fully simulated training environment that models both the user and the tools, removing dependence on external proprietary APIs for data synthesis and simulation.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)