NTH

HuatuoGPT-3: RL-Only Domain Adaptation from Base Models

AuthorsJunying Chen, Xinyuan Xie, Ziniu Li, Wenyuan Gu, Jianquan Li, Xiang Wan, Guangjun Yu, Ruoyu Sun, Haizhou Li, Benyou Wang

AffiliationsXiang Wan, Guangjun Yu, Ruoyu Sun, Haizhou Li, Benyou Wang22footnotemark: 2 · The Chinese University of Hong Kong, Shenzhen · Shenzhen Research Institute of Big Data · Shenzhen Loop Area Institute · National Health Data Institute, Shenzhen

October 10, 2026 2 min read
Watch on YouTube
The one-line take

OnePO uses temporary teacher guidance and reinforcement learning alone to turn general LLMs into stronger medical specialists without the usual supervised fine-tuning stage.

Key results

20K
Medical training examples

Training set used for the controlled medical adaptation experiments.

7.4
HealthBench gain over pure RL

Points higher for OnePO using DeepSeek-V3.2 guidance.

0.1
Probability floor

OnePO setting used to strengthen learning on low-probability teacher tokens.

27B
HuatuoGPT-3 model size

Size of the scaled model variant.

70.1
HuatuoGPT-3-27B HealthBench Total

Reported HealthBench Total score.

71.4
HuatuoGPT-3-27B HealthBench Professional

Reported HealthBench Professional score.

What the paper found

HuatuoGPT-3 introduces One-stage Policy Optimization, or OnePO, an RL-only approach that uses teacher answers as temporary guidance instead of first supervised-fine-tuning the base model. It targets two problems in mixed-policy RL: low-probability teacher tokens can receive too little learning signal, and outdated teacher answers can hold back later exploration. OnePO’s Adaptive Objective Evolution applies a probability floor of 0.1 and rescales gradients to help the model learn useful teacher tokens; Teacher Retirement removes a teacher answer whenever its reward no longer exceeds the best on-policy answer. On Qwen3-8B-Base, training on 20K medical examples with DeepSeek-V3.2 reasoning outputs as guidance produces a HealthBench Total score of 67.2, 7.4 points above pure RL and 2.7 above SFT+RL. Scaling OnePO yields the open-source HuatuoGPT-3 series; its 27B model scores 70.1 on HealthBench Total and 71.4 on HealthBench Professional, exceeding the reported GPT-6 Astra score on the professional benchmark. The study also uses OpenAI’s GPT-5 Chat to generate training prompts and rubrics, combining rubric-based rewards for open-ended medical responses with verifiable rewards for multiple-choice questions.

Original abstract

Domain adaptation aims to turn a general-purpose large language model (LLM) into an expert for a target domain. While the dominant SFT+RL pipeline offers a convenient cold start, it may reduce exploration diversity and introduces additional complexity through multi-stage optimization. These limitations motivate RL-only adaptation. However, pure on-policy RL suffers from a cold-start problem, while mixed-policy RL still falls short: informative tokens in teacher outputs are learned too slowly in early training, and stale teacher outputs can hinder later improvement. We identify these two failure modes as Gradient Starvation and Teacher-Distribution Anchoring. To address them, we propose One-stage Policy Optimization (OnePO), which treats teacher outputs as transient guidance for policy improvement. OnePO combines Adaptive Objective Evolution to strengthen learning on informative low-probability teacher tokens and Teacher Retirement to discard teacher outputs once the current policy can surpass them. On medical adaptation, OnePO achieves 67.2 on HealthBench (Total) with only 20K training samples, outperforming SFT+RL and pure RL by 2.7 and 7.4 points, respectively. We further scale OnePO to produce HuatuoGPT-3, an open-source medical LLM series whose 27B variant reaches 70.1 on HealthBench (Total) and 71.4 on HealthBench Professional, surpassing frontier models such as GPT-6 Astra. Models and code are available at https://github.com/FreedomIntelligence/HuatuoGPT-3.

Read the original paper

More in Large Language Models

Browse all 84 papers →
01Llm

Learning to Learn a Language

Lennart Carstens-Behrens, Holger Fröhlich

A transformer trained on synthetic worlds learns to infer the hidden rules of real language and other sequential data without ever seeing language during training.

Read analysis
02Llm

On-Policy Distillation with Negative-Policy Rollouts

Jaehui Hwang, Dongyoon Han, Sangdoo Yun, Byeongho Heo

NP-OPD improves language-model distillation by exposing students to rollouts from a weaker policy, helping them learn from the teacher while moving away from inferior behavior.

Read analysis