RecGPT-V3 Technical Report
- Published
- Source
- arXiv
- Paper number
- 661
- Field
- Information Retrieval
- arXiv ID
- 2607.15591
Key points
- The Memory Hub manages user behavior histories as memory units compressed by 94.5%, so the system does not have to reprocess the full history on every request.
- It solves the tag-item information bottleneck with a hybrid multimodal model in which the LLM directly infers Semantic IDs.
- It compresses 3,000-token reasoning into 15 tokens using latent CoT tokens, a 200x reduction, while keeping the content recoverable when needed.
- In live A/B tests, it improved IPV by 1.28%, CTR by 1.00%, and GMV by 3.97%, while cutting GPU cost by 52.4%.
Paper links
External research summaries. These are not HDATF publications or measured product results.