Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent

Published
Source
arXiv
Paper number
521
Field
LLMs / NLP
arXiv ID
2606.30616

Key points

  • A 35B MoE shows a performance path based on expanding agent horizon rather than parameter count.
  • Knowledge-Action Graph (KAG): a long-horizon trajectory infrastructure that stores evidence, actions, observations, and verification results as connected objects.
  • A three-stage training pipeline: all-domain SFT, domain-specific teacher RL, and domain-routed multi-teacher on-policy distillation.
  • It generates agent trajectories averaging 45K tokens and integrates six heterogeneous domains.
  • It outperforms 1T models on SEAL-0 (56.4), IFBench (80.6), and FrontierScience-Olympiad (79.0).
  • Salient Vocabulary Alignment reduces conflicts between reasoning patterns across domains.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)