📃 Paper | 🤗 HuatuoGPT-3-9B | 🤗 HuatuoGPT-3-27B | ⚖️ Grader | 📚 Data
Welcome to HuatuoGPT-3, the medical LLM series developed with OnePO!
One-stage Policy Optimization (OnePO) adapts language models to specialized domains in a single reinforcement-learning stage, without a preceding domain-specific supervised fine-tuning stage. It uses teacher responses as temporary guidance through two mechanisms:
- Adaptive Objective Evolution helps the model learn from low-probability teacher tokens.
- Teacher Retirement removes teacher guidance once the model's own responses match or exceed the teacher's reward.
This repository provides our medical models, OnePO training code, RL dataset, and rubric grader.
- Model Access
| Model | Backbone | HealthBench Professional ↑ | Access |
|---|---|---|---|
| HuatuoGPT-3-9B | Qwen3.5-9B | 68.3 | Hugging Face |
| HuatuoGPT-3-27B | Qwen3.8-27B | 71.4 | Hugging Face |
| HuatuoGPT-3-Grader-8B | Qwen3-8B | — | Hugging Face |
- Deploy
HuatuoGPT-3 can be used like Qwen3.5 and deployed with vLLM or SGLang.
OnePO-Medical-20K provides 20,338 medical RL tasks for OnePO, combining verifiable and rubric-based rewards.
| Task | Samples | Verification | Characteristics |
|---|---|---|---|
| Multiple-choice | 10,191 | Exact-match answer checking | Medical reasoning with a correct option label |
| Open-ended | 10,147 | Rubric-based grading | Free-form responses assessed against multiple clinical criteria |
Set TRAIN_FILE to onepo_medical_20K.json or your own training JSON.
- Installation
Prepare the GPU environment following the installation guide, then install the training code:
git clone https://github.com/FreedomIntelligence/HuatuoGPT-3.git
cd HuatuoGPT-3
python -m pip install -r requirements.txt- Run OnePO
We use HuatuoGPT-3-Grader-8B for efficient training-time rewards: it checks multiple rubric items in one generation, reducing grading overhead and API costs.
MODEL_PATH=/path/to/student_model \
TRAIN_FILE=/path/to/OnePO-Medical-20K/onepo_medical_20K.json \
VAL_FILE=/path/to/validation.json \
REWARD_MODEL_PATH=/path/to/HuatuoGPT-3-Grader-8B \
bash OnePO.shSee the training guide for GPU settings, hyperparameters, and using an existing grader service.
Track progress during training with your own validation tasks:
VAL_FILE: your validation JSON. Customize its questions and rubrics using the task schema.TEST_FREQ=30: evaluate every 30 optimizer steps.TEST_FREQ=0 TEST_STEPS='[50,100,200]': evaluate at selected steps.
Rubric-based feedback comes from the configured grader. For official HealthBench results, use the official evaluation code and the benchmark's grading protocol.
Explore our HuatuoGPT series:
- HuatuoGPT: Taming Language Models to Be a Doctor
- HuatuoGPT-II: One-stage Training for Medical Adaptation of LLMs
- HuatuoGPT-Vision: Injecting Medical Visual Knowledge into Multimodal LLMs at Scale
- HuatuoGPT-o1: Towards Medical Complex Reasoning with LLMs
- HuatuoGPT-3: OnePO: Direct One-stage Policy Optimization for SFT-free Domain Adaptation
@inproceedings{chen2026onepo,
title={OnePO: Direct One-stage Policy Optimization for SFT-free Domain Adaptation},
author={Chen, Junying and Xie, Xinyuan and Li, Ziniu and Wang, Benyou},
booktitle={Proceedings of the 43rd International Conference on Machine Learning},
year={2026}
}