Solar Open 2 Technical Report

Published
Source
arXiv
Paper number
698
Field
LLMs / NLP
arXiv ID
2607.20062

Key points

  • A 250B-A15B MoE structure scales up from Solar Open 1 (102B) while routing only 2.3% of the parameters.
  • Hybrid attention, 75% linear and 25% softmax, plus NoPE, supports a 1M-token context.
  • The Korean tokenizer is 24% more efficient than global models, meaning it represents the same Korean text with fewer tokens.
  • It ranks first among comparable open models and commercial APIs on average Korean benchmarks, and is nearly on par with DeepSeek-V4-Pro (1.6T) on Ko-GDPval.
  • It integrates 12 domain experts into a single model through multi-teacher on-policy distillation (MOPD).

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)