MetaStrategy: Generative Ranking with Executable LLM Strategies
AuthorsChengyu Lai, Jiuning Lin, Zhibo Xiao, Xiaodong Zhu, Ruiquan Lan, Bin Zhang, Zihong Huang, Wendong Zhang, Chuxin Chen, Yinjiang Cai, Shuai Zhong, Lingqing Zhang, Dimin Wang, Jialin Zhu, Han Zhu
Resources
MetaStrategy uses a compact LLM to generate safe, executable ranking policies that improve real-world recommendation outcomes without adding online latency.
Key results
Logged requests used for production-path replay training and evaluation
Deployable Student produced by routed on-policy distillation
Complementary Teachers distilled into the deployed Student
Offline marginal utility contributed by the routed OPD Student
Share of replay calls in which the Student Generator was selected
Treatment-side gain in the seven-day Taobao A/B test
What the paper found
MetaStrategy, deployed on Alibaba’s Taobao Homepage Guess You Like feed, reframes generative ranking as executable strategy generation rather than item or slate generation. A request-conditioned LLM emits a typed JSON bundle covering objective weights, card-type and category preferences, experience constraints, and top-CTR policies; deterministic validation and compilation translate it into parameters for an isolated Generator that competes with roughly ten incumbent Generators under a shared list-level Generator-Evaluator architecture. Training uses production-path replay over 65,536 logged requests, combining selection, relative-rank, and baseline-lift rewards with a self-competitive curriculum that turns frequent compiled strategies into frozen competitors. Evaluator-routed reward-augmented on-policy distillation then transfers complementary behavior from 4B Teachers into a deployable 0.8B Student. Offline, the Student achieves 98.03% validity, 16.24% selection, and 0.73% incremental GE lift, outperforming larger direct-RL variants on marginal contribution. Nearline, diff-triggered generation keeps LLM inference outside synchronous ranking, producing no observable response-time increase. In a seven-day randomized A/B test, MetaStrategy wins 27.93% of treatment-side GE calls and raises exposure PV by 1.49%, click PV by 2.11%, item-detail page views by 3.12%, and transaction amount by 2.83%.
Original abstract
Industrial recommender systems rank heterogeneous content under coupled user, business, commercial, and experience objectives. Existing generative ranking methods typically construct item sequences directly, making them difficult to integrate with mature predictive models, operational rules, and field-level guardrails. We present MetaStrategy, a framework that instead generates a structured, executable ranking strategy. Conditioned on request context, a large language model (LLM) policy emits a typed JSON bundle controlling objective weights, content and category preferences, experience constraints, and position policies. A deterministic validator and compiler instantiate an isolated Generator that competes atomically with incumbents under the list-level Evaluator of the Generator-Evaluator (GE) architecture. We train the policy in a production-path replay environment that re-executes logged requests through the current re-ranking stack without user exposure. The method combines selection, relative-rank, and baseline-lift rewards, a self-competitive curriculum that feeds frequent strategies back as competitors, and Evaluator-routed reward-augmented on-policy distillation that transfers complementary 4B-parameter Teachers into a compact 0.8B-parameter Student. We deploy MetaStrategy in Taobao Homepage Guess You Like through diff-triggered nearline generation; LLM inference remains outside synchronous ranking, with no observable increase in response time (RT). In a seven-day user-randomized online A/B test, MetaStrategy wins 27.93% of treatment-side GE calls and significantly improves click page views (click PV) by 2.11%, item-detail page views (IPV) by 3.12%, and transaction amount by 2.83%.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.