NTH

From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

AuthorsQinsi Wang, Jing Shi, Huazheng Wang, Kun Wan, Yiran Wu, Bo Liu, Qingyun Wu, Hai Helen Li, Yiran Chen, Handong Zhao, Wentian Zhao

August 3, 2026 2 min read
Watch on YouTube
The one-line take

SpyRL turns open-ended LLM tasks into self-play games where verifiable voting outcomes provide scalable reinforcement-learning rewards.

Key results

75.4%
Qwen3-8B summarization win rate

GPT-4o pairwise win rate for SpyRL against the untrained Qwen3-8B backbone

77.3%
Qwen3-8B creative-writing win rate

GPT-4o pairwise win rate for SpyRL against the untrained Qwen3-8B backbone

36.7
GovReport ROUGE-L

Qwen3-4B with SpyRL, compared with 33.2 for Absolute Zero

79.5
Math500 accuracy

Qwen3-4B with SpyRL, compared with 68.2 for the base model

50.4
Reasoning average with RAE

Seven-benchmark Qwen3-4B average with Role-Advantage Estimation

What the paper found

This paper introduces Reinforcement Learning with Self-Verifiable Rewards, or RLSVR, to extend RLVR beyond domains such as mathematics and coding, which powered systems including OpenAI o1 and DeepSeek-R1. Its implementation, SpyRL, transforms summarization, creative writing, and mathematical reasoning into an information-asymmetric social-deduction game: several civilian agents receive the full input, one spy receives degraded information, and all generate outputs before voting to identify the spy. Because the environment predetermined the spy’s identity, detection rewards are exact and rule-based, while vote counts provide a quality-linked surrogate reward for performers without human labels, learned reward models, or external judges. Using alternating GRPO optimization and Role-Advantage Estimation, SpyRL trained Qwen3 models and achieved 75.4% summarization and 77.3% creative-writing win rates for Qwen3-8B. On GovReport, Qwen3-4B’s ROUGE-L rose from 33.2 to 36.7, while mathematical reasoning on Math500 increased from 68.2 to 79.5. Role-aware calibration was critical: the seven-benchmark reasoning average reached 50.4 with RAE versus 37.5 without it. The method’s vote signal correlated with GPT-4o quality rankings, and gains remained robust under Gemini-3.5-Flash evaluation, suggesting task transformation can create scalable, verifier-free self-improvement for open-ended LLM capabilities.

Original abstract

Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enabling large-scale optimization. However, its applicability remains largely limited to domains such as mathematics and coding, where correctness can be deterministically verifiable. Open-ended tasks instead often rely on human preferences, reward models, or LLM-based judges, introducing evaluation bias, judge capability bottlenecks, and additional inference costs. Drawing on the principle of self-supervised learning, which constructs pretext tasks to derive supervision from the data itself, we propose Reinforcement Learning with Self-Verifiable Rewards (RLSVR), a task-transformation-based training paradigm for extending RLVR to open-ended tasks. RLSVR transforms open-ended tasks into verifiable proxy environments whose internal rules and interaction outcomes automatically generate reward signals. We instantiate RLSVR with SpyRL, a Self-PlaY Reinforcement Learning method inspired by social deduction game Who Is the Spy?. Agents receive asymmetric information, complete the same target task, and vote to identify a designated spy. Because the spy identity is predetermined, voting outcomes provide fully verifiable rewards, while successful identification remains closely related to output quality. Experiments on text summarization, creative writing, and mathematical reasoning show that SpyRL outperforms existing self-improvement methods on non-verifiable tasks and yields consistent gains on verifiable reasoning tasks. These results demonstrate that task transformation can extend scalable RLVR-based self-improvement beyond inherently verifiable domains. Models and code have been released at https://github.com/wangqinsi1/RLSVR/tree/SpyRL.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →