Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation
- Published
- Source
- arXiv
- Paper number
- 811
- Field
- Information Retrieval
- arXiv ID
- 2608.02738
Key points
- It uses Behavioral Multi-Token Prediction (BMTP) to filter spurious transitions across session boundaries, yielding cleaner pretraining data.
- It keeps the knowledge encoder and task learner as separate parameter sets so that updates and task optimization do not interfere with each other.
- It shows a 4% to 12% improvement over the baseline on eight public benchmarks.
- It was deployed in Shopee production and recorded a 1.75% increase in GMV and a 1.53% increase in ad revenue.
- Over 90 consecutive days in streaming settings, baseline methods stagnated while KGD maintained its advantage.
Paper links
External research summaries. These are not HDATF publications or measured product results.