Skip to content

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

OnePO: Direct One-stage Policy Optimization for SFT-free Domain Adaptation

HuatuoGPT-3

📃 Paper | 🤗 HuatuoGPT-3-9B | 🤗 HuatuoGPT-3-27B | ⚖️ Grader | 📚 Data

⚡ Introduction

Welcome to HuatuoGPT-3, the medical LLM series developed with OnePO!

One-stage Policy Optimization (OnePO) adapts language models to specialized domains in a single reinforcement-learning stage, without a preceding domain-specific supervised fine-tuning stage. It uses teacher responses as temporary guidance through two mechanisms:

  • Adaptive Objective Evolution helps the model learn from low-probability teacher tokens.
  • Teacher Retirement removes teacher guidance once the model's own responses match or exceed the teacher's reward.

OnePO framework

This repository provides our medical models, OnePO training code, RL dataset, and rubric grader.

👨‍⚕️ Model

  • Model Access
Model Backbone HealthBench Professional ↑ Access
HuatuoGPT-3-9B Qwen3.5-9B 68.3 Hugging Face
HuatuoGPT-3-27B Qwen3.8-27B 71.4 Hugging Face
HuatuoGPT-3-Grader-8B Qwen3-8B — Hugging Face
  • Deploy

HuatuoGPT-3 can be used like Qwen3.5 and deployed with vLLM or SGLang.

📚 Data

OnePO-Medical-20K provides 20,338 medical RL tasks for OnePO, combining verifiable and rubric-based rewards.

Task Samples Verification Characteristics
Multiple-choice 10,191 Exact-match answer checking Medical reasoning with a correct option label
Open-ended 10,147 Rubric-based grading Free-form responses assessed against multiple clinical criteria

Set TRAIN_FILE to onepo_medical_20K.json or your own training JSON.

🚀 Training

  • Installation

Prepare the GPU environment following the installation guide, then install the training code:

git clone https://github.com/FreedomIntelligence/HuatuoGPT-3.git
cd HuatuoGPT-3
python -m pip install -r requirements.txt
  • Run OnePO

We use HuatuoGPT-3-Grader-8B for efficient training-time rewards: it checks multiple rubric items in one generation, reducing grading overhead and API costs.

MODEL_PATH=/path/to/student_model \
TRAIN_FILE=/path/to/OnePO-Medical-20K/onepo_medical_20K.json \
VAL_FILE=/path/to/validation.json \
REWARD_MODEL_PATH=/path/to/HuatuoGPT-3-Grader-8B \
bash OnePO.sh

See the training guide for GPU settings, hyperparameters, and using an existing grader service.

🧐 Evaluation

Track progress during training with your own validation tasks:

  • VAL_FILE: your validation JSON. Customize its questions and rubrics using the task schema.
  • TEST_FREQ=30: evaluate every 30 optimizer steps.
  • TEST_FREQ=0 TEST_STEPS='[50,100,200]': evaluate at selected steps.

Rubric-based feedback comes from the configured grader. For official HealthBench results, use the official evaluation code and the benchmark's grading protocol.

🩺 HuatuoGPT Series

Explore our HuatuoGPT series:

  • HuatuoGPT: Taming Language Models to Be a Doctor
  • HuatuoGPT-II: One-stage Training for Medical Adaptation of LLMs
  • HuatuoGPT-Vision: Injecting Medical Visual Knowledge into Multimodal LLMs at Scale
  • HuatuoGPT-o1: Towards Medical Complex Reasoning with LLMs
  • HuatuoGPT-3: OnePO: Direct One-stage Policy Optimization for SFT-free Domain Adaptation

📖 Citation

@inproceedings{chen2026onepo,
  title={OnePO: Direct One-stage Policy Optimization for SFT-free Domain Adaptation},
  author={Chen, Junying and Xie, Xinyuan and Li, Ziniu and Wang, Benyou},
  booktitle={Proceedings of the 43rd International Conference on Machine Learning},
  year={2026}
}

About

HuatuoGPT-3: RL-Only Domain Adaptation from Base Models via Off-Policy Seeding

Resources

Stars

11 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages