NTH

Unifying Generative Recall and Multi-Objective Ranking in a Single Decoder-Only Sequence

AuthorsRuochen Yang, Shuang Wen, Pengbo Xu, Yusheng Huang, Jiangxia Cao, Shuang Yang, Zhaojie Liu, Jiawei Sheng, Tingwen Liu

August 6, 2026 2 min read
Watch on YouTube
The one-line take

UniR² turns recommendation recall and ranking into two coordinated tasks within one decoder-only Transformer, aiming to improve consistency and efficiency in production systems.

Key results

400M
Evaluation users

Scale of the Kuaishou live-streaming evaluation data

4.54%
Recall HR@64 improvement

UniR2 improvement over the strongest recall baseline

1.45%
CTR UAUC improvement

Offline ranking improvement for the click-through-rate objective

54.29%
Inference-time reduction

End-to-end reduction from cached context reuse and unified serving

1.177%
Kuaishou App play-volume gain

Online A/B-test improvement on 5% of traffic

What the paper found

Researchers at Kuaishou Technology and the Institute of Information Engineering, Chinese Academy of Sciences, propose UniR2, a single decoder-only Transformer that replaces the conventional recall-then-ranking cascade. UniR2 places user context, a hierarchical semantic-ID trajectory, and item features into one heterogeneous sequence, allowing the generated recall trajectory to serve as a representation bridge for multi-objective ranking. Its Dual-Query Prefix-Causal Attention gives SID generation causal access to user history while allowing ranking tokens to attend bidirectionally to the user profile, full SID trajectory, and item features. Shared attention weights provide representation reuse, while separate feed-forward paths, stop-gradient boundaries, and ranking-only LoRA preserve optimization isolation; an MMoE head predicts click, long-view, and gift objectives. On Kuaishou live-streaming data covering 400M users and 3M authors, UniR2 improved recall HR@64 by 4.54% over the strongest baseline and raised CTR UAUC by 1.45%. Cached key-value states eliminate repeated user-context encoding, reducing end-to-end inference time by 54.29%. In a two-week online A/B test on 5% of traffic, Kuaishou App play volume increased by 1.177%, demonstrating that unified modeling can improve retrieval quality, ranking accuracy, and serving efficiency simultaneously.

Original abstract

Modern industrial recommendation systems typically separate recall and ranking into two independent stages. Although this cascade supports corpus-level retrieval and fine-grained multi-objective scoring, it causes objective inconsistency, information loss at the candidate hand-off, and redundant user-side context computation. Meanwhile, the generative recall and ranking scaling share a common Transformer-based modeling philosophy, where architectural consistency creates a natural opportunity for unified integration. However, direct sharing remains challenging since the two tasks require different information visibility and optimization methods. Therefore, we propose \textbf{UniR$^2$}, a \textbf{Uni}fied decoder-only Transformer that unifies Generative \textbf{R}ecall and Multi-Objective \textbf{R}anking within a single heterogeneous sequence comprising user context, SID trajectory, and item features. Within this sequence, the generated trajectory serves as a representation bridge between recall and ranking, where Dual-Query Prefix-Causal Attention provides task-specific visibility. The two tasks share the base attention weights but retain separate optimization boundaries, with ranking-side LoRA preserving ranking adaptability without disrupting the generative backbone. Extensive offline experiments on large-scale industrial data demonstrate the effectiveness and efficiency of UniR$^2$ for both recall and ranking. Long-term online A/B tests on Kuaishou platform further show consistent positive gains, validating the practicality of unified model in large-scale recommendation systems.

Read the original paper

More in Transformers

Browse all 42 papers →
03Transformer

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Zehao Jin, Ruixuan Deng, Junran Wang

A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.

Read analysis