Learning to Orchestrate Agents in Natural Language with the Conductor
- Published
- Source
- arXiv
- Paper number
- 171
- Field
- Agents / Multi-Agent / Orchestration
- arXiv ID
- 2512.04388
Key points
- Manually designing the optimal prompt engineering strategy and multi-agent workflow for complex tasks is time-consuming and difficult.
- Because no single model is universally optimal, it is difficult to make effective use of the specialized strengths of different LLMs.
- Existing multi-agent coordination methods often depend on human-designed scaffolds, predefined routing, and fixed tool-use patterns, which limits flexibility and adaptability.
- Conductor is a 7B LLM trained end to end with reinforcement learning, using the GRPO algorithm, to generate agent workflows in natural language that specify communication patterns with worker LLMs and subtasks.
- Conductor orchestrates a diverse pool of seven strong worker LLMs, both commercial and open source, by dynamically decomposing problems and assigning target subtasks with tailored instructions.
- The framework allows flexible communication topologies, including recursive self-calls, enabling dynamic refinement of coordination strategies and adaptive resource allocation at test time.
Paper links
External research summaries. These are not HDATF publications or measured product results.