Mellum2 Technical Report
- Published
- Source
- arXiv
- Paper number
- 279
- Field
- LLMs / NLP
- arXiv ID
- 2605.31268
Key points
- This report introduces Mellum 2, an open-weight 12B-parameter MoE language model with 2.5B active parameters per token.
- Across code generation, math and reasoning, tool use, knowledge, and safety benchmarks, Mellum 2 is competitive with open-weight baselines in the 4B to 14B range while operating at the per-token compute cost of a 2.5B dense model.
- The authors release the base, instruct, and thinking checkpoints under the Apache 2.0 license together with this report, which describes the underlying architectural choices, data pipelines, and training recipe.
Paper links
External research summaries. These are not HDATF publications or measured product results.