NTH

TGR: Advancing Industrial Recommendation from Generative-Paradigm Ranking toward Unified Generation and Reasoning

AuthorsTGR Team, Lei Cheng, Haonan Hu, Beibei Kong, Yudong Li, Zang Li, Yunsheng Pang, Hongyang Su, Jianchao Tu, Yunlong Wang, Bing Wen, Junzhang Zhu, Shaojie Zhu, Chengxiang Zhuo

September 10, 2026 2 min read
Watch on YouTube
The one-line take

Tencent’s TGR framework replaces fragmented recommendation pipelines with generative ranking, slate generation, and offline reasoning, delivering measurable gains at massive production scale.

Key results

4B
CCFormer production dataset

Sample scale used for industrial offline evaluation

2.21
CCFormer training speedup

Training speedup versus HSTU

16.9%
BARGE Hit@5 improvement

Maximum improvement over OneRec across two industrial scenarios

5
HiGR inference speedup

Speedup over matched OneRec-style decoding

477.8%
TGR-Reason cold-start Hit@1 lift

Relative improvement for cold-start new users

What the paper found

TGR, Tencent’s industrial recommendation framework, moves beyond the conventional retrieval–ranking cascade through three layers: generative-paradigm ranking, end-to-end generation, and amortized reasoning. Its CCFormer ranker combines unified feature tokenization, field-separated cross-attention, subspace token mixing, and hierarchical sequence compression, preserving per-item multi-task scores while scaling to long histories. On a 4B-sample production dataset, CCFormer trains 2.21x faster than HSTU and uses roughly half the GFLOPs. For end-to-end recommendation, BARGE repairs semantic-ID generation with item context-aware attention, hierarchical path reranking, and dual-path decoding, improving Hit@5 over OneRec by up to 16.9% without increasing parameters, beam width, or serving latency. HiGR instead generates an entire slate using PCRQ-VAE semantic IDs, a coarse-to-fine Hierarchical Slate Decoder, and ORPO listwise alignment for ranking fidelity, user interest, and diversity; it delivers a 5x inference speedup and below-50-ms P99 latency while reducing GPU demand by about 60%. TGR-Reason adds LLM-derived priors without online chain-of-thought: a LatentRec-trained Think model based on Qwen3-1.7B exports reason tokens offline, which Direct Reasoning Injection feeds into the generator. This raises cold-start new-user Hit@1 by 477.8% and improves online Effective Consumption Rate by 1.75%, showing that reasoning can be amortized into production recommendation rather than executed for every request.

Original abstract

Industrial recommender systems typically rely on cascaded retrieval, pre-ranking, ranking, and reranking stages, whose separately optimized models limit scaling, fragment decision making, and lack semantic knowledge and reasoning. We present TGR (Tencent Generative Recommendation), an industrial framework that advances recommendation toward the generative paradigm along three coupled directions. TGR-GenRank upgrades ranking through CCFormer, which combines unified feature tokenization, a scalable Transformer backbone, feature-field separated cross attention, subspace token mixing, and hierarchical sequence compression while retaining per-item multi-task outputs. TGR-GenRec explores end-to-end generation under two paradigms: BARGE bridges item-boundary loss and semantic drift in hierarchical semantic-ID generation through item context-aware attention, hierarchical path reranking, and orthogonal dual-path decoding; HiGR performs whole-slate generation with prefix-structured semantic IDs, coarse-to-fine decoding, and listwise multi-objective alignment. TGR-Reason injects offline-generated semantic-ID reason tokens into online decoding, providing reasoning without request-time rollout. TGR is deployed across Tencent production surfaces serving hundreds of millions of users. CCFormer delivers significant gains in five A/B-tested scenarios and is fully launched in two, including +3.57% CTR and +1.71% advertising revenue. BARGE improves Hit@5 by 10.2-16.9% and yields +0.60% CTR and +1.70% reading time after full rollout. HiGR improves offline slate quality by 15.9-21.3% with a 5x inference speedup and achieves up to +1.22% watch time and +1.73% video views. TGR-Reason raises cold-start new-user Hit@1 by 477.8% and delivers +1.75% effective consumption and +13.09% new-user exposure-to-conversion online.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis