LaRec: Unleashing LLM-based Latent Reasoning for Generative Recommendation

Published
Source
arXiv
Paper number
741
Field
Information Retrieval
arXiv ID
2607.24617

Key points

  • It greatly reduces response latency by reasoning in latent space instead of through long text reasoning.
  • During latent pretraining, step-level alignment and process direction alignment provide rich supervisory signals to intermediate reasoning steps.
  • A user-specific Gaussian mixture distribution is used to sample the starting point of reasoning within the relevant latent space.
  • GRPO is combined with a sparse-hit reward and a dense semantic-similarity reward to align reasoning with the recommendation objective.
  • The latency of latent reasoning is close to no-reasoning and much lower than chain-of-thought.
  • An online A/B test on a real advertising platform records improvements such as a 2.93% conversion rate.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)