OneReason Technical Report
- Published
- Source
- arXiv
- Paper number
- 361
- Field
- Information Retrieval
- arXiv ID
- 2606.06260
Key points
- The paper analyzes the cause of the problem in generative recommendation, where thinking mode fails to outperform non-thinking mode, along two axes: perception and cognition.
- In pretraining, it obtains strong item-level token perception through an itemic tokenizer and four-granularity training.
- It uses a three-level cognition-enhanced chain-of-thought, or CoT, SFT scheme that goes from R0, which stands for perception, to R1, derivation, to R2, evolution, and finally R3, recommendation.
- The authors use a specialize-then-unify RL strategy that first specializes by task and then unifies the models to strengthen thinking ability.
- On the Kuaishou production multi-domain benchmark, the thinking mode consistently outperforms the non-thinking mode.
- They find a transfer effect in which CoT supervision also improves non-thinking inference.
- OneReason-8B and 0.8B models are planned for open sourcing.
Paper links
External research summaries. These are not HDATF publications or measured product results.