NTH

Co-Evolving LLM Evaluators and Policies via DynamicRubric

AuthorsBeining Wang, Weihang Su, Hongtao Tian, Hao Kong, Tao Yang, Ting Yao, Qingyi Pan, Yueyue Wu, Qingyao Ai, Min Zhang, Yiqun Liu

July 31, 2026 2 min read
Watch on YouTube
The one-line take

DynamicRubric keeps LLM evaluators useful as policies improve by generating adaptive rubrics tailored to each set of competing responses.

Key results

8B
DR-Generator backbone

Qwen3-8B initializes the dynamic rubric generator and policy.

24K
Evaluator training examples

Filtered Nectar examples used to optimize DR-Generator.

36K
Policy training prompts

Filtered prompts retained for DR-Policy optimization.

57.4
JudgeBench accuracy

DR-Generator-8B evaluator score, compared with 46.4 for zero-shot Qwen3-32B.

67.1
AlpacaEval2 win rate

Qwen3-8B policy optimized with DynamicRubric supervision.

21.0
ArenaHardv2.0 win rate

DynamicRubric-optimized Qwen3-8B policy performance.

What the paper found

Researchers at Tsinghua University and Tencent’s WeChat propose DynamicRubric, a co-evolution framework for preventing evaluator score gaps from collapsing as an LLM policy improves. Their probability-allocation analysis shows that the score difference between two candidate responses exactly determines the local gain from shifting probability mass toward the preferred response. DynamicRubric therefore samples responses from the current policy, has a DR-Generator produce weighted binary rubric items tailored to that response set, and uses a frozen DR-Verifier to aggregate yes/no judgments into comparable scores. The generator is trained with a weighted Bernoulli-variance discriminability objective plus a Bradley–Terry-style anchor-ranking objective, while the policy is optimized with GRPO. Using Qwen3-8B, the system trains on 24K evaluator examples and 36K policy prompts; its DR-Generator reaches 57.4 JudgeBench accuracy, outperforming zero-shot Qwen3-32B at 46.4 despite using one-quarter the nominal model scale. DynamicRubric policy supervision raises Qwen3-8B to 67.1 on AlpacaEval2, 21.0 on ArenaHardv2.0, 30.7 on WildBench, and 64.3 on WritingBench, exceeding static rubric supervision from a 235B generator and reward models up to 70B. Benefits also transfer to mathematical, scientific-QA, and coding benchmarks, remain robust with Llama-3.1-8B-Instruct, and have been deployed across all WeChat Search AI-answering traffic, serving tens of millions of requests per day.

Original abstract

Post-training with evaluator feedback on policy-induced samples serves as a major mechanism for improving large language models. As policies improve, these sampled responses become close in quality. These close candidates create a bottleneck for policy optimization: collapsed relative evaluator score gaps yield weak or misleading policy supervision. We theoretically characterize why these gaps matter through a probability allocation view, showing that the directional gain of shifting probability mass from one response to another is exactly the evaluator score gap between them. This identifies relative score gaps as the policy optimization signals that guide updates. Motivated by this view, we propose DynamicRubric, a response-set-conditioned evaluator--policy co-evolution framework that generates weighted binary rubric items for each candidate set and aggregates the resulting judgments into response-level scores. In our experiments with 8B backbones, DynamicRubric improves evaluator performance and provides stronger policy supervision than baselines using a 70B reward model or a 235B static rubric generator. DynamicRubric-optimized policies also show gains on verifiable reasoning and coding tasks. A DynamicRubric-optimized model is fully deployed in WeChat Search's AI answering scenario, where it serves all online traffic across tens of millions of requests per day and improves key online metrics. These results suggest a principle for evaluator-guided post-training: evaluators should evolve with the policies they supervise.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis