Motif 3: Technical Report

Published
Source
arXiv
Paper number
862
Field
AI / General
arXiv ID
2608.09119

Key points

  • It uses an ultra-fine MoE architecture that selects only 8 experts per token out of 384, increasing capacity while reducing compute.
  • Grouped Differential Latent Attention, or GDLA, achieves both noise suppression and KV cache compression.
  • It is pretrained on 12.5 trillion tokens and then uses multi-teacher on-policy distillation, or MOPD, to unify diverse capabilities into one model.
  • It supports a 256K context length and stabilizes large-scale training with MXFP8 precision.
  • Among open-weight models, it shows competitive performance on agent tasks, math reasoning, and scientific knowledge.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)