Base Models Can Reason By Taking a Cue From Training Data
- Published
- Source
- arXiv
- Paper number
- 1166
- Field
- Machine Learning
- arXiv ID
- 2610.06851
Key points
- Fixing only the first two token cues raised Olmo-3-7B's MATH-500 accuracy from 42% to 78% (the RL version reaches 75%).
- They showed that RL does not so much teach new abilities as increase the probability of existing reasoning cues appearing.
- By causally editing the training data, they made 'chicken' work as a reasoning cue just like 'okay', proving the cue effect comes from the data.
- In a safety case study, different cues elicited different refusal and compliance behaviors, with implications for safety evaluation.
Paper links
External research summaries. These are not HDATF publications or measured product results.