Latent Agents: A Post-Training Procedure for Internalized Multi-Agent Debate

Published
Source
arXiv
Paper number
166
Field
Agents / Distillation / Efficiency
arXiv ID
2604.24881

Key points

  • Explicit multi-agent debate improves LLM accuracy and reduces hallucinations, but it is computationally expensive because it requires multiple model calls and verbose outputs.
  • Existing methods for internalizing LLM reasoning mainly focus on explicit reasoning steps in a single agent and do not handle the complex interactions of multi-agent debate.
  • Suppressing undesirable LLM behavior often leads to an alignment tax, which is a loss of general model capability caused by coarse control mechanisms.
  • We developed IMAD, a two-stage fine-tuning pipeline consisting of supervised fine-tuning (SFT) on the full multi-agent debate trajectory followed by reinforcement learning (RL) with dynamic rewards.
  • The RL stage uses a dynamic reward function that rewards correct answers while progressively reducing incentives for explicit debate language and tightening output length limits, forcing potential internalization.
  • We apply activation steering to identified agent subspaces within the internalized model to control specific behaviors, including suppression of malicious traits.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)