Mach-Mind-4-Flash Technical Report

Published
Source
arXiv
Paper number
594
Field
Machine Learning
arXiv ID
2607.09375

Key points

  • A 35B MoE model with 3B active parameters reaches 100B-plus-level performance through post-training alone, without scaling pretraining compute.
  • MOPD, or Multi-Teacher On-Policy Distillation, trains domain-specific RL experts independently and then merges them with routed reverse-KL to remove the seesaw effect of mixed-reward RL.
  • HMPO is a single-stage token-efficient optimization method that compresses reasoning chains by 19 to 46 percent while keeping accuracy loss below 0.7 percentage points.
  • It reports AIME'26 at 92.70, BFCL-v4 at 75.80, and ClawBench at 84.20, showing that 3B active parameters can compete with 1T-scale models.
  • SonicMoE GEMM kernels and segmented shared-expert fusion speed training up by 17 percent end to end.
  • It also reshapes the Pareto frontier for token efficiency by sharply reducing token usage at similar accuracy.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)