Learning to Orchestrate Agents in Natural Language with the Conductor

Published
Source
arXiv
Paper number
171
Field
Agents / Multi-Agent / Orchestration
arXiv ID
2512.04388

Key points

  • Manually designing the optimal prompt engineering strategy and multi-agent workflow for complex tasks is time-consuming and difficult.
  • Because no single model is universally optimal, it is difficult to make effective use of the specialized strengths of different LLMs.
  • Existing multi-agent coordination methods often depend on human-designed scaffolds, predefined routing, and fixed tool-use patterns, which limits flexibility and adaptability.
  • Conductor is a 7B LLM trained end to end with reinforcement learning, using the GRPO algorithm, to generate agent workflows in natural language that specify communication patterns with worker LLMs and subtasks.
  • Conductor orchestrates a diverse pool of seven strong worker LLMs, both commercial and open source, by dynamically decomposing problems and assigning target subtasks with tailored instructions.
  • The framework allows flexible communication topologies, including recursive self-calls, enabling dynamic refinement of coordination strategies and adaptive resource allocation at test time.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)