OneReason Technical Report

Published
Source
arXiv
Paper number
361
Field
Information Retrieval
arXiv ID
2606.06260

Key points

  • The paper analyzes the cause of the problem in generative recommendation, where thinking mode fails to outperform non-thinking mode, along two axes: perception and cognition.
  • In pretraining, it obtains strong item-level token perception through an itemic tokenizer and four-granularity training.
  • It uses a three-level cognition-enhanced chain-of-thought, or CoT, SFT scheme that goes from R0, which stands for perception, to R1, derivation, to R2, evolution, and finally R3, recommendation.
  • The authors use a specialize-then-unify RL strategy that first specializes by task and then unifies the models to strengthen thinking ability.
  • On the Kuaishou production multi-domain benchmark, the thinking mode consistently outperforms the non-thinking mode.
  • They find a transfer effect in which CoT supervision also improves non-thinking inference.
  • OneReason-8B and 0.8B models are planned for open sourcing.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)