NTH

DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space

AuthorsJiangwang Chen, Zixin Song, Junlin Liu, Shuaiyu Zhou, Haiyan Wu, Haihan Shi, Chenxi Zhou, Hanqing Li, Xiao Yang, Da Zhu, Guanjun Jiang, Hai Wan, Xibin Zhao

August 1, 2026 2 min read
Watch on YouTube
The one-line take

DecoEvo helps language models improve by evolving both their problem-solving instructions and the rubrics used to expose overlooked weaknesses.

Key results

15
Backbone-benchmark combinations won

DecoEvo achieved the highest mean score across all five benchmarks and three backbones.

5.0%
GPT-4o average relative gain

Relative improvement over SkillOpt on the five-benchmark average.

64.8
GPT-4o five-benchmark average

DecoEvo average score, compared with 61.7 for SkillOpt.

45.3
GPT-4o HealthBench score

DecoEvo score on HealthBench, compared with 41.7 for SkillOpt.

56.9
HealthBench rubric-alignment F1

DecoEvo’s GPT-4o criterion-level F1, compared with 44.6 for SkillOpt.

2.4–2.9
Verification ablation loss

Point loss on HealthBench when held-out generator verification was removed.

What the paper found

DecoEvo, from researchers at Tsinghua University, Peking University, the University of Chinese Academy of Sciences, and Alibaba’s Qwen Business Unit, adapts frozen black-box LLMs by co-evolving two inspectable text artifacts: a solver skill and a rubric-generator skill. Its key innovation is score decoupling: solver revisions use criterion-level rubric feedback, while generator revisions are accepted through task-conditioned structural audits for omitted requirements and rubric-blind near-tie comparisons for missed distinctions, rather than aggregate solver scores that could reward easier rubrics. Evaluated on HealthBench, WritingBench, ResearchQA, and two cross-benchmark transfers, LLMEval-Med and EQ-Bench Creative Writing v3, DecoEvo achieved the best result in all 15 benchmark–backbone combinations using OpenAI GPT-4o and Alibaba Qwen3-4B and Qwen3-8B. On GPT-4o, its five-benchmark average was 64.8 versus 61.7 for SkillOpt, a 5.0% relative gain; on HealthBench it reached 45.3 versus SkillOpt’s 41.7. Rubric alignment also improved: GPT-4o’s HealthBench F1 rose to 56.9 from 44.6. Ablations show that removing held-out generator verification caused the largest loss, 2.4–2.9 points, supporting persistent, audited rubric evolution over score-coupled co-adaptation.

Original abstract

Text-space optimization adapts large language models (LLMs) by editing external natural-language artifacts rather than model weights, so the optimized artifacts remain inspectable and the model can be treated as a black box. However, most existing text-space methods keep evaluation fixed. On open-ended tasks, this can become a bottleneck: once the solver improves on the criteria a rubric measures, omitted dimensions remain invisible to the optimization signal. Simply evolving the rubric is also unreliable when updates are selected by the current solver's score, because apparent progress can come from making the rubric easier to satisfy. We introduce DecoEvo (Decoupled Co-Evolution), which co-evolves a solver skill and a rubric-generator skill under decoupled objectives without using gold rubrics during optimization. The solver skill is updated using criterion-level feedback, while the rubric-generator skill is revised through complementary audits of requirement coverage and response discrimination that are independent of aggregate solver score. This separation focuses generator updates on newly exposed solver weaknesses, reducing repeated emphasis on criteria the solver already satisfies. Under each benchmark's official evaluation, DecoEvo outperforms all compared methods across five benchmarks and three LLM backbones, yielding 2.8--5.0\% relative gains over SkillOpt in the five-benchmark average.

Read the original paper

More in Optimization

Browse all 36 papers →