RecGPT-V3 Technical Report

Published
Source
arXiv
Paper number
661
Field
Information Retrieval
arXiv ID
2607.15591

Key points

  • The Memory Hub manages user behavior histories as memory units compressed by 94.5%, so the system does not have to reprocess the full history on every request.
  • It solves the tag-item information bottleneck with a hybrid multimodal model in which the LLM directly infers Semantic IDs.
  • It compresses 3,000-token reasoning into 15 tokens using latent CoT tokens, a 200x reduction, while keeping the content recoverable when needed.
  • In live A/B tests, it improved IPV by 1.28%, CTR by 1.00%, and GMV by 3.97%, while cutting GPU cost by 52.4%.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)