Base Models Can Reason By Taking a Cue From Training Data

Published
Source
arXiv
Paper number
1166
Field
Machine Learning
arXiv ID
2610.06851

Key points

  • Fixing only the first two token cues raised Olmo-3-7B's MATH-500 accuracy from 42% to 78% (the RL version reaches 75%).
  • They showed that RL does not so much teach new abilities as increase the probability of existing reasoning cues appearing.
  • By causally editing the training data, they made 'chicken' work as a reasoning cue just like 'okay', proving the cue effect comes from the data.
  • In a safety case study, different cues elicited different refusal and compliance behaviors, with implications for safety evaluation.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)