Mach-Mind-4-Flash Technical Report
- Published
- Source
- arXiv
- Paper number
- 594
- Field
- Machine Learning
- arXiv ID
- 2607.09375
Key points
- A 35B MoE model with 3B active parameters reaches 100B-plus-level performance through post-training alone, without scaling pretraining compute.
- MOPD, or Multi-Teacher On-Policy Distillation, trains domain-specific RL experts independently and then merges them with routed reverse-KL to remove the seesaw effect of mixed-reward RL.
- HMPO is a single-stage token-efficient optimization method that compresses reasoning chains by 19 to 46 percent while keeping accuracy loss below 0.7 percentage points.
- It reports AIME'26 at 92.70, BFCL-v4 at 75.80, and ClawBench at 84.20, showing that 3B active parameters can compete with 1T-scale models.
- SonicMoE GEMM kernels and segmented shared-expert fusion speed training up by 17 percent end to end.
- It also reshapes the Pareto frontier for token efficiency by sharply reducing token usage at similar accuracy.
Paper links
External research summaries. These are not HDATF publications or measured product results.