DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space
AuthorsJiangwang Chen, Zixin Song, Junlin Liu, Shuaiyu Zhou, Haiyan Wu, Haihan Shi, Chenxi Zhou, Hanqing Li, Xiao Yang, Da Zhu, Guanjun Jiang, Hai Wan, Xibin Zhao
Resources
DecoEvo helps language models improve by evolving both their problem-solving instructions and the rubrics used to expose overlooked weaknesses.
Key results
DecoEvo achieved the highest mean score across all five benchmarks and three backbones.
Relative improvement over SkillOpt on the five-benchmark average.
DecoEvo average score, compared with 61.7 for SkillOpt.
DecoEvo score on HealthBench, compared with 41.7 for SkillOpt.
DecoEvo’s GPT-4o criterion-level F1, compared with 44.6 for SkillOpt.
Point loss on HealthBench when held-out generator verification was removed.
What the paper found
DecoEvo, from researchers at Tsinghua University, Peking University, the University of Chinese Academy of Sciences, and Alibaba’s Qwen Business Unit, adapts frozen black-box LLMs by co-evolving two inspectable text artifacts: a solver skill and a rubric-generator skill. Its key innovation is score decoupling: solver revisions use criterion-level rubric feedback, while generator revisions are accepted through task-conditioned structural audits for omitted requirements and rubric-blind near-tie comparisons for missed distinctions, rather than aggregate solver scores that could reward easier rubrics. Evaluated on HealthBench, WritingBench, ResearchQA, and two cross-benchmark transfers, LLMEval-Med and EQ-Bench Creative Writing v3, DecoEvo achieved the best result in all 15 benchmark–backbone combinations using OpenAI GPT-4o and Alibaba Qwen3-4B and Qwen3-8B. On GPT-4o, its five-benchmark average was 64.8 versus 61.7 for SkillOpt, a 5.0% relative gain; on HealthBench it reached 45.3 versus SkillOpt’s 41.7. Rubric alignment also improved: GPT-4o’s HealthBench F1 rose to 56.9 from 44.6. Ablations show that removing held-out generator verification caused the largest loss, 2.4–2.9 points, supporting persistent, audited rubric evolution over score-coupled co-adaptation.
Original abstract
Text-space optimization adapts large language models (LLMs) by editing external natural-language artifacts rather than model weights, so the optimized artifacts remain inspectable and the model can be treated as a black box. However, most existing text-space methods keep evaluation fixed. On open-ended tasks, this can become a bottleneck: once the solver improves on the criteria a rubric measures, omitted dimensions remain invisible to the optimization signal. Simply evolving the rubric is also unreliable when updates are selected by the current solver's score, because apparent progress can come from making the rubric easier to satisfy. We introduce DecoEvo (Decoupled Co-Evolution), which co-evolves a solver skill and a rubric-generator skill under decoupled objectives without using gold rubrics during optimization. The solver skill is updated using criterion-level feedback, while the rubric-generator skill is revised through complementary audits of requirement coverage and response discrimination that are independent of aggregate solver score. This separation focuses generator updates on newly exposed solver weaknesses, reducing repeated emphasis on criteria the solver already satisfies. Under each benchmark's official evaluation, DecoEvo outperforms all compared methods across five benchmarks and three LLM backbones, yielding 2.8--5.0\% relative gains over SkillOpt in the five-benchmark average.
Read the original paperMore in Optimization
Browse all 36 papers →An $Ω(κ_y^8ε^{-6})$ Lower Bound for Stochastic NC-SC Bilevel Optimization with First-order Oracles
Zhihao Gu, Qilong Wu, Junchi Yang
This work proves that stochastic bilevel optimization fundamentally requires up to epsilon^{-6} oracle queries, showing existing methods are asymptotically optimal.
Hyper Algorithm Design Agent: Evolving Learnable Optimizer from Zero
Zipei Yu, Yue-Jiao Gong, Zeyuan Ma, Yuncheng Jiang, Zhiguang Cao
A pair of self-improving coding agents evolves new learnable optimization algorithms from a simple template, reducing the need for handcrafted optimizer design.
Tight Regret Bound for Online Inverse Linear Optimization via Multiscale Matrix Weights
Shinsaku Sakaue
A new multiscale matrix-weights algorithm learns hidden linear preferences online with provably optimal dimension-dependent regret.