Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent
- Published
- Source
- arXiv
- Paper number
- 521
- Field
- LLMs / NLP
- arXiv ID
- 2606.30616
Key points
- A 35B MoE shows a performance path based on expanding agent horizon rather than parameter count.
- Knowledge-Action Graph (KAG): a long-horizon trajectory infrastructure that stores evidence, actions, observations, and verification results as connected objects.
- A three-stage training pipeline: all-domain SFT, domain-specific teacher RL, and domain-routed multi-teacher on-policy distillation.
- It generates agent trajectories averaging 45K tokens and integrates six heterogeneous domains.
- It outperforms 1T models on SEAL-0 (56.4), IFBench (80.6), and FrontierScience-Olympiad (79.0).
- Salient Vocabulary Alignment reduces conflicts between reasoning patterns across domains.
Paper links
External research summaries. These are not HDATF publications or measured product results.