Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation
AuthorsZixuan Wang, Yuhong Chen, Yuxuan Zhu, Guidong Lei, Zhiluohan Guo, Yu Zhao, Kun Wang, Bangyang Hong, Kangle Wu, Yabo Ni, Anxiang Zeng, Cong Fu, Hui Li
This work keeps recommender systems up to date by separating reusable behavioral knowledge from task-specific adaptation, delivering strong benchmark and real-world business gains.
Key results
Improvement reached on some metrics across eight Amazon-2023 Reviews benchmarks.
Days over which KGD sustained its advantage on a streaming production evaluation.
Increase measured in the Shopee Homepage Search A/B test.
Increase measured in the Shopee Homepage Search A/B test.
Approximate sample count in the industrial recommendation stream.
Milliseconds of serving latency on A30 GPUs after deployment.
What the paper found
Knowledge–Geometry Decoupling, or KGD, addresses two failures in streaming recommendation: next-token prediction treats every adjacent behavior as meaningful, and shared parameters force pretrained knowledge and task-specific geometry into conflict. Its Behavioral Multi-Token Prediction, BMTP, filters future-item supervision through collaborative similarity from LightGCN and semantic similarity from Qwen3-Embedding, while its transfer mechanism uses read-only cross-attention plus an Anchored Calibration Residual, or ACR, that adds task-specific low-rank directions orthogonal to the pretrained embedding. This lets the encoder refresh on new behavior without task gradients overwriting it, while the task learner independently reshapes the representation for retrieval or ranking. Using ManCAR on eight Amazon-2023 Reviews benchmarks, KGD improves strong pretrain-transfer baselines by 4–12%, reaching 12% on some metrics. On a 90-day production stream, it maintains gains as frozen and entangled baselines degrade. In Shopee Homepage Search, deployed with the OneRank ranking backbone, KGD processes approximately 13B samples and lifted GMV per user by 1.75% and advertising revenue per user by 1.53% in a live A/B test. Training takes about two hours per daily update on NVIDIA A100 GPUs, while serving latency remains 120 ms on A30 GPUs, demonstrating refreshable transfer without a latency penalty.
Original abstract
Industrial recommenders increasingly adopt the pretrain-then-transfer paradigm, yet behavioral distribution drift raises two questions: what to learn from behavior sequences, and how to transfer the learned knowledge while the pretrained model is continually refreshed. To resolve them, we propose Knowledge-Geometry Decoupling (KGD). For what to learn, conventional next-token prediction treats adjacency as dependency and may encode spurious transitions across unrelated sessions. We introduce Behavioral Multi-Token Prediction (BMTP) to retain only collaboratively or semantically related future items as supervision, yielding cleaner and more transferable behavioral knowledge. For how to transfer, pretrained knowledge and task-specific geometry impose conflicting optimization demands on shared parameters. To handle it, KGD assigns them to separate parameter sets: a refreshable encoder owns behavioral knowledge, while a task learner reads contextualized encoder states through read-only cross-attention and writes task-specific geometry through Anchored Calibration Residual (ACR) orthogonal to the pretrained embedding. The decoupled ownership enables continual knowledge refresh without task-gradient interference or invalidating downstream adaptation. KGD improves over strong pretrain-transfer baselines by 4-12% on eight public benchmarks and sustains its advantage over a 90-day production stream where baselines show no gains. KGD has been fully deployed in Shopee. In a live A/B test on Shopee Homepage Search, it increases GMV per user by 1.75% and advertising revenue by 1.53%, demonstrating its high practical value. We provide the core implementation of KGD at https://github.com/FuCongResearchSquad/KGD4REC.
Read the original paperMore in Continual Learning
Browse all 24 papers →ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience
Haodong Lu, Dong Gong
ASCENT lets deployed LLM agents learn from verified successes on the fly by converting hindsight about their own trajectories into lasting weight updates.
From Knowledge Access to Source Learning: Developing Source-Specific Competence
Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang
SourceLearn helps LLM agents progressively build reusable expertise about trusted information sources instead of repeatedly starting from scratch.
Local Support Learning
Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes
Local Support Learning helps large language models learn new skills without overwriting what they already know by activating updates only where they are locally needed.