OLMo 2

Published
Source
arXiv
Paper number
003
Field
LLMs / Open Models
arXiv ID
2501.00656

Key points

  • The release includes weights, data, code, recipes, logs, and intermediate checkpoints, making it useful for reproducible research and downstream adaptation.
  • In the training recipe, Dolmino Mix 1124 and late-stage curriculum learning improve downstream capability without relying only on more compute.
  • The instruction model OLMo 2-Instruct builds on the Tülu 3 recipe and extends it with reinforcement learning that uses verifiable rewards in the final stage.
  • As a result, the base model remains more transparent while staying competitive with open-weight systems in the Llama 3.1, Qwen 2.5, and Gemma 2 families.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)