OneModel: A Unified Foundation for Platform-Scale Multi-Scenario Ranking
AuthorsYinqi Zhang, Peiyu Hu, Yuntian Tang, Siying Gu, Jiahao Liang, Longxin Kou, Haiqing Hu, Shuman Zhuang, Yubin Xu, Chenggen Sun, Bin Ye, Donghui Xu, Zhaoyu Liu, Jiang Rong, Yuting Jia, Zhaokai Luo, Leilei Ma, Yiying Xie, Yao Hu
Resources
OneModel unifies user behavior across feeds, ads, and commerce into one production-scale ranking system that improves engagement and business metrics.
Key results
Relative online A/B improvement
Relative online A/B improvement
Relative online A/B improvement
OneModel production latency after optimization
OneModel capacity in the cost-effectiveness comparison
What the paper found
OneModel is a unified foundation model for final ranking across Xiaohongshu’s organic recommendation, advertising, and merchant services, replacing fragmented scenario-specific systems with a shared user representation. It converts heterogeneous behaviors into timestamped event sequences, uses an action-oriented GenRank-style Transformer with 500-event context, and combines long-term attention pooling with the latest state to capture both stable preferences and immediate intent. Scenario-aware Information Modulation, or SAIM, gates feed-forward channels by business stream, while a multi-objective loss jointly trains next-item prediction and stream-specific ranking, with gradient isolation and selective backpropagation improving optimization stability. For deployment, cached incremental user states, feature decomposition, prefetching, shared user-tower computation, and graph-level inference optimization let the model scale from 173M to 230M dense parameters while reducing latency from 270ms to 90ms. Production A/B tests report 0.33% higher Explore Feed time spent, 1.25% higher engagement, 3.43% higher advertising value, and 8.18% higher advertising CTR; Merchant Recommendation also raises DGMV by 1.1867% and GPM by 2.1585%. The results position OneModel as a practical multi-stream alternative to separate ranking stacks and a complementary industrial direction to systems such as HSTU and GenRank, rather than a general-purpose language model like ChatGPT or Claude.
Original abstract
Platform-scale recommender systems often span multiple business streams such as organic recommendation, advertising, and merchant services, where user behaviors form a continuous cross-stream trajectory. Maintaining separate ranking systems fragments user representations and increases engineering cost. We propose \textbf{OneModel}, a unified framework for multi-stream final ranking. OneModel maps heterogeneous behaviors into shared event sequences, learns long-context user representations with an action-oriented backbone, and introduces \emph{Scenario-aware Information Modulation} to balance cross-stream transfer and stream-specific specialization. For production deployment, OneModel further adopts stratified user representation, multi-objective training, and optimized online serving with feature decomposition, user feature prefetching, shared user-tower computation, and graph-level inference optimization. We deploy OneModel in production at \emph{Xiaohongshu}, where it delivers consistent offline gains over strong baselines and scales favorably with context length and model capacity. Online A/B tests improve Time Spent by \textbf{+0.33\%} and Engagement by \textbf{+1.25\%} in Explore Feed, lift advertising value by \textbf{+3.43\%} and CTR by \textbf{+8.18\%} in Feed Advertising, and raise DGMV by \textbf{+1.1867\%} and GPM by \textbf{+2.1585\%} in Merchant Recommendation, validating unified multi-stream ranking as an effective production foundation.
Read the original paperMore in Foundation Models
Browse all 47 papers →How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi
AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.
TabFM: A Zero-Shot Foundation Model for Tabular Data
Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das
TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.
When Do Biological Reasoning Models Use Their Biological Inputs?
Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik
The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.