From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement
AuthorsQinsi Wang, Jing Shi, Huazheng Wang, Kun Wan, Yiran Wu, Bo Liu, Qingyun Wu, Hai Helen Li, Yiran Chen, Handong Zhao, Wentian Zhao
SpyRL turns open-ended LLM tasks into self-play games where verifiable voting outcomes provide scalable reinforcement-learning rewards.
Key results
GPT-4o pairwise win rate for SpyRL against the untrained Qwen3-8B backbone
GPT-4o pairwise win rate for SpyRL against the untrained Qwen3-8B backbone
Qwen3-4B with SpyRL, compared with 33.2 for Absolute Zero
Qwen3-4B with SpyRL, compared with 68.2 for the base model
Seven-benchmark Qwen3-4B average with Role-Advantage Estimation
What the paper found
This paper introduces Reinforcement Learning with Self-Verifiable Rewards, or RLSVR, to extend RLVR beyond domains such as mathematics and coding, which powered systems including OpenAI o1 and DeepSeek-R1. Its implementation, SpyRL, transforms summarization, creative writing, and mathematical reasoning into an information-asymmetric social-deduction game: several civilian agents receive the full input, one spy receives degraded information, and all generate outputs before voting to identify the spy. Because the environment predetermined the spy’s identity, detection rewards are exact and rule-based, while vote counts provide a quality-linked surrogate reward for performers without human labels, learned reward models, or external judges. Using alternating GRPO optimization and Role-Advantage Estimation, SpyRL trained Qwen3 models and achieved 75.4% summarization and 77.3% creative-writing win rates for Qwen3-8B. On GovReport, Qwen3-4B’s ROUGE-L rose from 33.2 to 36.7, while mathematical reasoning on Math500 increased from 68.2 to 79.5. Role-aware calibration was critical: the seven-benchmark reasoning average reached 50.4 with RAE versus 37.5 without it. The method’s vote signal correlated with GPT-4o quality rankings, and gains remained robust under Gemini-3.5-Flash evaluation, suggesting task transformation can create scalable, verifier-free self-improvement for open-ended LLM capabilities.
Original abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enabling large-scale optimization. However, its applicability remains largely limited to domains such as mathematics and coding, where correctness can be deterministically verifiable. Open-ended tasks instead often rely on human preferences, reward models, or LLM-based judges, introducing evaluation bias, judge capability bottlenecks, and additional inference costs. Drawing on the principle of self-supervised learning, which constructs pretext tasks to derive supervision from the data itself, we propose Reinforcement Learning with Self-Verifiable Rewards (RLSVR), a task-transformation-based training paradigm for extending RLVR to open-ended tasks. RLSVR transforms open-ended tasks into verifiable proxy environments whose internal rules and interaction outcomes automatically generate reward signals. We instantiate RLSVR with SpyRL, a Self-PlaY Reinforcement Learning method inspired by social deduction game Who Is the Spy?. Agents receive asymmetric information, complete the same target task, and vote to identify a designated spy. Because the spy identity is predetermined, voting outcomes provide fully verifiable rewards, while successful identification remains closely related to output quality. Experiments on text summarization, creative writing, and mathematical reasoning show that SpyRL outperforms existing self-improvement methods on non-verifiable tasks and yields consistent gains on verifiable reasoning tasks. These results demonstrate that task transformation can extend scalable RLVR-based self-improvement beyond inherently verifiable domains. Models and code have been released at https://github.com/wangqinsi1/RLSVR/tree/SpyRL.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.