NTH

Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL

AuthorsYunhao Yang, Yuexin Bian, Yunjie Tian, Di Fu, Tianjin Huang, Yuanyuan Shi, Ziang Xiao, Nuno Vasconcelos, Yijiang Li

August 27, 2026 2 min read
Watch on YouTube
The one-line take

Co-RL trains diverse language and vision-language agents to improve reasoning by rewarding each other, achieving strong gains without ground-truth labels.

Key results

8.6%
Text benchmark gain

Upper end of Co-RL's reported average improvement across seven text-only reasoning benchmarks.

7.2%
Multimodal benchmark gain

Upper end of Co-RL's reported average improvement across four multimodal benchmarks.

4.0%
CoMAS improvement

Average improvement over CoMAS in the controlled multi-agent evaluation.

53.6
Qwen2.5-7B Co-RL average

Average score for Qwen2.5-7B with the different-family-plus Co-RL variant.

47.7
Llama-3.1-8B-Instruct Co-RL average

Average score for Llama-3.1-8B-Instruct, exceeding its 47.1 ground-truth-reward reference.

What the paper found

Co-RL introduces a label-free reinforcement-learning framework in which independently parameterized language or vision-language models supervise one another without sharing weights, gradients, ground-truth labels, external judges, or learned reward models. Each agent samples multiple answers, builds a majority-vote pseudo-label from a peer, and receives binary rewards for agreeing with that peer’s vote; optimization uses GRPO or REINFORCE++. The central finding is that cohort diversity—different model families, parameter scales, and rephrased prompts—reduces correlated errors and prevents the response homogenization and training collapse seen in self-rewarding methods such as TTRL. Experiments pair models including Qwen, Llama, Gemma, and InternVL, with DeepSeek-V3 generating semantically equivalent prompt rewrites. Across seven text reasoning benchmarks, Co-RL produces average gains of 3.0–8.6%, while multimodal experiments across MathVision, MathVerse, MathVista, and We-Math deliver 2.3–7.2% gains for models from 2B to 12B parameters. In a controlled CoMAS comparison, it improves average performance by 4.0% using half as many agents and no LLM judge. At larger scale, Co-RL raises Qwen2.5-7B from 49.0 to 53.6 average score and Llama-3.1-8B-Instruct from 44.7 to 47.7, with the latter exceeding ground-truth-reward training despite using no labels. Theory explains the advantage: two agents converge correctly when their initial answer probabilities sum above 1, enlarging the basin of correct convergence beyond self-rewarding.

Original abstract

Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-language models, yet its strongest successes still depend heavily on ground-truth supervision (e.g., verifiable reward). Such annotations are costly to obtain and become increasingly scarce as reasoning capabilities advance beyond what humans can reliably evaluate. Self-rewarding RL reduces this dependence by enabling models to derive reward signals from their own completions. However, training solely on self-generated feedback can reinforce existing biases and suboptimal behaviors, reduce response diversity, and ultimately lead to homogenized responses and training collapse. In this work, we show that unsupervised reasoning can emerge through cooperative multi-agent training. We introduce Co-RL, a framework in which multiple decoupled models, sharing no parameters, are simultaneously optimized through RL using rewards derived from their peers. We further show that increasing cohort diversity, through heterogeneous model families, sizes, and rephrased training samples, reduces the correlated errors that drive self-reinforcing feedback loops. This diversity consistently improves reasoning performance, maintains behavioral diversity, and mitigates training collapse. Across text-only and multimodal domains, Co-RL consistently outperforms the base models and prior label-free approaches, while matching or surpassing supervised methods, without access to any ground-truth labels. Concretely, Co-RL yields average gains of 3.0-8.6% across seven text-only benchmarks for LLMs and 2.3-7.2% across four multimodal benchmarks for VLMs. Code is available at https://github.com/DrStranded/Co-RL.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →