NTH

From Language to Behavior: Scaling Sequence Transformers for Industrial Recommendation Ranking with Rec-Native Designs

AuthorsJie Chen, Xiangqian Yu, Yanchao Lian, Tan Lu, Run Yang, Zhengchun Shang, Xing Wang, Cheng Chen, Ke Hu, Qiang Li, Tianjiu Yin, Xiaobing Liu

September 9, 2026 2 min read
Watch on YouTube
The one-line take

ReST adapts scalable Transformers for industrial recommendation by reusing user-history computation across candidates while improving ranking quality under strict latency constraints.

Key results

1.31%
Online AUC improvement

Improvement in the one-week TikTok Shop Ads A/B test.

11.93%
Advertiser Value lift

Revenue-related online metric improvement in the A/B test.

50
P99 latency budget

Milliseconds, met by the deployed ReST configuration.

5.8
Shared-prefix training throughput

Times higher end-to-end training throughput with user-level prefix reuse.

20
Serving cost reduction

Up to times lower sequence-model inference cost through shared-prefix serving.

0.8135
ML-20M ReST AUC

Best reported AUC for ReST on the ML-20M benchmark.

What the paper found

This paper introduces ReST, a recommendation-native Transformer framework from ByteDance for scaling user-behavior modeling in industrial ranking, where histories are noisy, timestamps are irregular, supervision is sparse, and one history must score many candidates under strict latency limits. ReST combines Dual-Gated Attention, which filters unreliable behavior before and after aggregation, RoPE plus multi-granularity Rotary Temporal Embedding, and Stabilized Residual Normalization for deeper scaling. Its asymmetric design uses a heavy sequence encoder that runs once per user and a lightweight cross decoder with projection-free key/value attention and token-specific parameters for candidate-specific scoring. Training-only sequence CVR and sequence/non-sequence alignment objectives counter sequence starvation caused by strong DLRM features, while user-level and serving-time shared-prefix reuse reduce redundant computation. Unlike LLaMA-style Transformers and HSTU, ReST continues improving with longer histories, greater depth, and wider hidden states; on ML-20M it reaches an AUC of 0.8135. In a one-week TikTok Shop Ads A/B test, the deployed system improved online AUC by 1.31% and Advertiser Value by 11.93% within a 50 ms P99 budget. Shared-prefix training increased throughput by 5.8, and serving reduced sequence-model inference cost by up to 20 times.

Original abstract

Scaling Transformers has driven large gains in language modeling, but transplanting this to behavior-sequence modeling in production ranking is challenging: recommendation differs in signal quality, where behavior sequences are noisy, temporally irregular, and sparsely supervised, and in computation asymmetry, where each request scores many candidates against one shared user history under tight latency budgets. We propose ReST, a recommendation-native Transformer scaling framework. For signal quality, it introduces a sequence encoder with dual-gated attention, rotary positional and temporal embedding, stabilized residual normalization, and training-only auxiliary objectives. For computation asymmetry, it factorizes ranking into a heavy reusable encoder and a lightweight cross decoder with projection-free KV attention and token-specific parameterization, coupling user-level shared-prefix training with shared-prefix serving for compute-once, decode-many-times ranking. Across industrial and public benchmarks, ReST achieves higher accuracy and scales more consistently along sequence length, depth, and width, where LLM-style Transformer blocks saturate. A one-week online A/B test on a production advertising platform improves online AUC by 1.31% and lifts a core revenue metric by 11.93% within a 50 ms P99 budget; ReST has since been fully deployed in production, showing that behavior-sequence scaling remains a promising, under-exploited axis for production ranking.

Read the original paper

More in Transformers

Browse all 42 papers →
03Transformer

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Zehao Jin, Ruixuan Deng, Junran Wang

A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.

Read analysis