Motif 3: Technical Report
- Published
- Source
- arXiv
- Paper number
- 862
- Field
- AI / General
- arXiv ID
- 2608.09119
Key points
- It uses an ultra-fine MoE architecture that selects only 8 experts per token out of 384, increasing capacity while reducing compute.
- Grouped Differential Latent Attention, or GDLA, achieves both noise suppression and KV cache compression.
- It is pretrained on 12.5 trillion tokens and then uses multi-teacher on-policy distillation, or MOPD, to unify diverse capabilities into one model.
- It supports a 256K context length and stabilizes large-scale training with MXFP8 precision.
- Among open-weight models, it shows competitive performance on agent tasks, math reasoning, and scientific knowledge.
Paper links
External research summaries. These are not HDATF publications or measured product results.