TGR: Advancing Industrial Recommendation from Generative-Paradigm Ranking toward Unified Generation and Reasoning
AuthorsTGR Team, Lei Cheng, Haonan Hu, Beibei Kong, Yudong Li, Zang Li, Yunsheng Pang, Hongyang Su, Jianchao Tu, Yunlong Wang, Bing Wen, Junzhang Zhu, Shaojie Zhu, Chengxiang Zhuo
Resources
Tencent’s TGR framework replaces fragmented recommendation pipelines with generative ranking, slate generation, and offline reasoning, delivering measurable gains at massive production scale.
Key results
Sample scale used for industrial offline evaluation
Training speedup versus HSTU
Maximum improvement over OneRec across two industrial scenarios
Speedup over matched OneRec-style decoding
Relative improvement for cold-start new users
What the paper found
TGR, Tencent’s industrial recommendation framework, moves beyond the conventional retrieval–ranking cascade through three layers: generative-paradigm ranking, end-to-end generation, and amortized reasoning. Its CCFormer ranker combines unified feature tokenization, field-separated cross-attention, subspace token mixing, and hierarchical sequence compression, preserving per-item multi-task scores while scaling to long histories. On a 4B-sample production dataset, CCFormer trains 2.21x faster than HSTU and uses roughly half the GFLOPs. For end-to-end recommendation, BARGE repairs semantic-ID generation with item context-aware attention, hierarchical path reranking, and dual-path decoding, improving Hit@5 over OneRec by up to 16.9% without increasing parameters, beam width, or serving latency. HiGR instead generates an entire slate using PCRQ-VAE semantic IDs, a coarse-to-fine Hierarchical Slate Decoder, and ORPO listwise alignment for ranking fidelity, user interest, and diversity; it delivers a 5x inference speedup and below-50-ms P99 latency while reducing GPU demand by about 60%. TGR-Reason adds LLM-derived priors without online chain-of-thought: a LatentRec-trained Think model based on Qwen3-1.7B exports reason tokens offline, which Direct Reasoning Injection feeds into the generator. This raises cold-start new-user Hit@1 by 477.8% and improves online Effective Consumption Rate by 1.75%, showing that reasoning can be amortized into production recommendation rather than executed for every request.
Original abstract
Industrial recommender systems typically rely on cascaded retrieval, pre-ranking, ranking, and reranking stages, whose separately optimized models limit scaling, fragment decision making, and lack semantic knowledge and reasoning. We present TGR (Tencent Generative Recommendation), an industrial framework that advances recommendation toward the generative paradigm along three coupled directions. TGR-GenRank upgrades ranking through CCFormer, which combines unified feature tokenization, a scalable Transformer backbone, feature-field separated cross attention, subspace token mixing, and hierarchical sequence compression while retaining per-item multi-task outputs. TGR-GenRec explores end-to-end generation under two paradigms: BARGE bridges item-boundary loss and semantic drift in hierarchical semantic-ID generation through item context-aware attention, hierarchical path reranking, and orthogonal dual-path decoding; HiGR performs whole-slate generation with prefix-structured semantic IDs, coarse-to-fine decoding, and listwise multi-objective alignment. TGR-Reason injects offline-generated semantic-ID reason tokens into online decoding, providing reasoning without request-time rollout. TGR is deployed across Tencent production surfaces serving hundreds of millions of users. CCFormer delivers significant gains in five A/B-tested scenarios and is fully launched in two, including +3.57% CTR and +1.71% advertising revenue. BARGE improves Hit@5 by 10.2-16.9% and yields +0.60% CTR and +1.70% reading time after full rollout. HiGR improves offline slate quality by 15.9-21.3% with a 5x inference speedup and achieves up to +1.22% watch time and +1.73% video views. TGR-Reason raises cold-start new-user Hit@1 by 477.8% and delivers +1.75% effective consumption and +13.09% new-user exposure-to-conversion online.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.