Unifying Generative Recall and Multi-Objective Ranking in a Single Decoder-Only Sequence
AuthorsRuochen Yang, Shuang Wen, Pengbo Xu, Yusheng Huang, Jiangxia Cao, Shuang Yang, Zhaojie Liu, Jiawei Sheng, Tingwen Liu
Resources
UniR² turns recommendation recall and ranking into two coordinated tasks within one decoder-only Transformer, aiming to improve consistency and efficiency in production systems.
Key results
Scale of the Kuaishou live-streaming evaluation data
UniR2 improvement over the strongest recall baseline
Offline ranking improvement for the click-through-rate objective
End-to-end reduction from cached context reuse and unified serving
Online A/B-test improvement on 5% of traffic
What the paper found
Researchers at Kuaishou Technology and the Institute of Information Engineering, Chinese Academy of Sciences, propose UniR2, a single decoder-only Transformer that replaces the conventional recall-then-ranking cascade. UniR2 places user context, a hierarchical semantic-ID trajectory, and item features into one heterogeneous sequence, allowing the generated recall trajectory to serve as a representation bridge for multi-objective ranking. Its Dual-Query Prefix-Causal Attention gives SID generation causal access to user history while allowing ranking tokens to attend bidirectionally to the user profile, full SID trajectory, and item features. Shared attention weights provide representation reuse, while separate feed-forward paths, stop-gradient boundaries, and ranking-only LoRA preserve optimization isolation; an MMoE head predicts click, long-view, and gift objectives. On Kuaishou live-streaming data covering 400M users and 3M authors, UniR2 improved recall HR@64 by 4.54% over the strongest baseline and raised CTR UAUC by 1.45%. Cached key-value states eliminate repeated user-context encoding, reducing end-to-end inference time by 54.29%. In a two-week online A/B test on 5% of traffic, Kuaishou App play volume increased by 1.177%, demonstrating that unified modeling can improve retrieval quality, ranking accuracy, and serving efficiency simultaneously.
Original abstract
Modern industrial recommendation systems typically separate recall and ranking into two independent stages. Although this cascade supports corpus-level retrieval and fine-grained multi-objective scoring, it causes objective inconsistency, information loss at the candidate hand-off, and redundant user-side context computation. Meanwhile, the generative recall and ranking scaling share a common Transformer-based modeling philosophy, where architectural consistency creates a natural opportunity for unified integration. However, direct sharing remains challenging since the two tasks require different information visibility and optimization methods. Therefore, we propose \textbf{UniR$^2$}, a \textbf{Uni}fied decoder-only Transformer that unifies Generative \textbf{R}ecall and Multi-Objective \textbf{R}anking within a single heterogeneous sequence comprising user context, SID trajectory, and item features. Within this sequence, the generated trajectory serves as a representation bridge between recall and ranking, where Dual-Query Prefix-Causal Attention provides task-specific visibility. The two tasks share the base attention weights but retain separate optimization boundaries, with ranking-side LoRA preserving ranking adaptability without disrupting the generative backbone. Extensive offline experiments on large-scale industrial data demonstrate the effectiveness and efficiency of UniR$^2$ for both recall and ranking. Long-term online A/B tests on Kuaishou platform further show consistent positive gains, validating the practicality of unified model in large-scale recommendation systems.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.