NTH

MetaStrategy: Generative Ranking with Executable LLM Strategies

AuthorsChengyu Lai, Jiuning Lin, Zhibo Xiao, Xiaodong Zhu, Ruiquan Lan, Bin Zhang, Zihong Huang, Wendong Zhang, Chuxin Chen, Yinjiang Cai, Shuai Zhong, Lingqing Zhang, Dimin Wang, Jialin Zhu, Han Zhu

August 14, 2026 2 min read
Watch on YouTube
The one-line take

MetaStrategy uses a compact LLM to generate safe, executable ranking policies that improve real-world recommendation outcomes without adding online latency.

Key results

65,536
Replay requests

Logged requests used for production-path replay training and evaluation

0.8B
Student model size

Deployable Student produced by routed on-policy distillation

4B
Teacher model size

Complementary Teachers distilled into the deployed Student

0.73%
Incremental GE lift

Offline marginal utility contributed by the routed OPD Student

16.24%
GE selection rate

Share of replay calls in which the Student Generator was selected

2.11%
Online click PV improvement

Treatment-side gain in the seven-day Taobao A/B test

What the paper found

MetaStrategy, deployed on Alibaba’s Taobao Homepage Guess You Like feed, reframes generative ranking as executable strategy generation rather than item or slate generation. A request-conditioned LLM emits a typed JSON bundle covering objective weights, card-type and category preferences, experience constraints, and top-CTR policies; deterministic validation and compilation translate it into parameters for an isolated Generator that competes with roughly ten incumbent Generators under a shared list-level Generator-Evaluator architecture. Training uses production-path replay over 65,536 logged requests, combining selection, relative-rank, and baseline-lift rewards with a self-competitive curriculum that turns frequent compiled strategies into frozen competitors. Evaluator-routed reward-augmented on-policy distillation then transfers complementary behavior from 4B Teachers into a deployable 0.8B Student. Offline, the Student achieves 98.03% validity, 16.24% selection, and 0.73% incremental GE lift, outperforming larger direct-RL variants on marginal contribution. Nearline, diff-triggered generation keeps LLM inference outside synchronous ranking, producing no observable response-time increase. In a seven-day randomized A/B test, MetaStrategy wins 27.93% of treatment-side GE calls and raises exposure PV by 1.49%, click PV by 2.11%, item-detail page views by 3.12%, and transaction amount by 2.83%.

Original abstract

Industrial recommender systems rank heterogeneous content under coupled user, business, commercial, and experience objectives. Existing generative ranking methods typically construct item sequences directly, making them difficult to integrate with mature predictive models, operational rules, and field-level guardrails. We present MetaStrategy, a framework that instead generates a structured, executable ranking strategy. Conditioned on request context, a large language model (LLM) policy emits a typed JSON bundle controlling objective weights, content and category preferences, experience constraints, and position policies. A deterministic validator and compiler instantiate an isolated Generator that competes atomically with incumbents under the list-level Evaluator of the Generator-Evaluator (GE) architecture. We train the policy in a production-path replay environment that re-executes logged requests through the current re-ranking stack without user exposure. The method combines selection, relative-rank, and baseline-lift rewards, a self-competitive curriculum that feeds frequent strategies back as competitors, and Evaluator-routed reward-augmented on-policy distillation that transfers complementary 4B-parameter Teachers into a compact 0.8B-parameter Student. We deploy MetaStrategy in Taobao Homepage Guess You Like through diff-triggered nearline generation; LLM inference remains outside synchronous ranking, with no observable increase in response time (RT). In a seven-day user-randomized online A/B test, MetaStrategy wins 27.93% of treatment-side GE calls and significantly improves click page views (click PV) by 2.11%, item-detail page views (IPV) by 3.12%, and transaction amount by 2.83%.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis