2026-10-11 16:38 UTC

FreedomIntelligence claims its released HuatuoGPT-3 models and OnePO training stack adapt language models to medicine in one reinforcement-learning stage without domain-specific supervised fine-tuning, potentially simplifying reproducible specialization of open models.

state: watchingheat: lowuncertainty: highnovelscott: mediumopen-models reinforcement-learning domain-adaptationFreedomIntelligenceJunying ChenXinyuan XieZiniu LiBenyou Wang

What is this?

FreedomIntelligence, a group affiliated with CUHK-Shenzhen and the Shenzhen Research Institute of Big Data (Junying Chen, Xinyuan Xie, Ziniu Li, Benyou Wang), has released HuatuoGPT-3, a family of open-weight medical LLMs built on Qwen3-family backbones (8B/9B/27B/32B, plus a 7B variant on Huawei's openPangu-Embedded-7B per the 32B card) deployable via vLLM/SGLang. The distinctive claim is OnePO β€” referred to as 'SeedRL' on the newer cards without explanation of the naming relationship β€” an RL-only domain-adaptation paradigm that turns a base model into a medical expert in a single reinforcement-learning stage with no preceding domain-specific SFT; weights, training code, the OnePO-Medical-20K RL dataset, and a rubric grader are released (Apache-2.0), and the OnePO paper is cited as ICML 2026 in the publisher's own citation block. The snippets show this is a continuation of a multi-year house program β€” HuatuoGPT-II's 'one-stage training' (2023), then the explicitly two-stage SFT+PPO HuatuoGPT-o1 (2024) β€” making the RL-only claim a pivot from the group's own earlier pipelines. Nothing in the supplied material constitutes independent reproduction or third-party evaluation of the one-stage claim; every substantive line still originates from the publisher.

Why it matters to Scott

OnePO's SFT-free single-stage RL adaptation is new territory β€” Scott's canon holds no position on RL-only specialization β€” but it is a directly testable alternative to the curated-SFT-data route of his reddit fine-tuning-data factory (dev:project.reddit, dev:concept.synthetic-finetuning-dataset): if it trains as documented, it would undercut the premise that domain adaptation needs his kind of SFT corpus, so it warrants an evaluation run before it earns a claim either way. Until an independent reproduction exists, his own evidence-class ladder (ip:concept.evidence-class-ladder) caps this at evaluation lead, not a validated replacement.
dev:project.redditdev:concept.synthetic-finetuning-datasetip:concept.evidence-class-ladderradar:concept.post-trainingradar:concept.reinforcement-learningradar:concept.fine-tuningradar:concept.reproducibilityradar:financial-rlvr-10k-validationradar:qwen3-chat-template-thinking-loss
queries asked of Scott's wikis
  • synthetic fine-tuning data pipeline for domain adaptation
  • SFT vs single-stage RL post-training recipes
  • rubric grader LLM-as-judge reward models
  • open-weight domain specialization local inference economics
  • reproducibility of released training stacks evaluation criteria
  • base model selection for domain fine-tuning Qwen3

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 575h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-17 17:39 (minted)⭐ origin echo-reconstructedThe linked model card identifies this repository as the released HuatuoGPT-3 training code and describes OnePO as single-stage medical RL ad
FreedomIntelligence on github (echo) Β· attributed from reddit.post.1wiyu3e Β· published time unknown
β€”
09-17 16:30first on r/LocalLLaMA Β· published Β· lag ?FreedomIntelligence/HuatuoGPT-3-9B Β· Hugging Face
jacek2023
β€”
09-17 16:30amplified on r/LocalLLaMAreddit.post.1wiyu3e
jacek2023
peak 22 Β· 8 comments Β· 41% of case engagement
09-24 19:20amplified on r/LocalLLaMA πŸ‘‘reddit.post.1wpaxud
jacek2023
peak 40 Β· 3 comments Β· 59% of case engagement
09-17 17:20our radar first saw it Β· lag ?discovery anchor: reddit.post.1wiyu3eβ€”
pace: p63 vs 1032 stories at the 336h mark (now 575h old) β€” ahead of tencent-evie-visual-retrieval (1.0x), behind pentagon-anthropic-ban-reaffirmed (1.0x)

Evidence (3) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditFreedomIntelligence/HuatuoGPT-3-9B · Hugging Face
LocalLLaMA
Retrieved article excerpt

Open article Β· Retrieved 2026-09-17T17:26:09.459499+00:00

# [FreedomIntelligence](https://huggingface.co/FreedomIntelligence) / [HuatuoGPT-3-9B](https://huggingface.co/FreedomIntelligence/HuatuoGPT-3-9B) Like 1 Follow FreedomAI 1.08k

[Image-Text-to-Text](https://huggingface.co/models?pipeline_tag=image-text-to-text)[Transformers](https://huggingface.co/models?library=transformers)[Safetensors](https://huggingface.co/models?library=safetensors)[qwen3\_5](https://huggingface.co/models?other=qwen3_5)[medical](https://huggingface.co/models?other=medical)[reasoning](https://huggingface.co/models?other=reasoning)[conversational](https://huggingface.co/models?other=conversational)[onepo](https://huggingface.co/models?other=onepo)

License: apache-2.0

[Model card](https://huggingface.co/FreedomIntelligence/HuatuoGPT-3-9B)  [Files Files and versions  

xet](https://huggingface.co/FreedomIntelligence/HuatuoGPT-3-9B/tree/main)  [Community](https://huggingface.co/FreedomIntelligence/HuatuoGPT-3-9B/discussions)

 

Deploy

  Copy to bucket new   

Use this model

 

# 🩺 HuatuoGPT-3-9B

[🏠 GitHub](https://github.com/FreedomIntelligence/HuatuoGPT-3) |
[πŸ“„ Paper](https://openreview.net/pdf?id=M8eyUQldfx)

## Introduction

**HuatuoGPT-3-9B** is a medical LLM built on [Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) with **One-stage Policy Optimization (OnePO)**. OnePO adapts language models to medicine in a single reinforcement-learning stage, without preceding domain-specific supervised fine-tuning. Teacher responses provide temporary guidance and are retired as the model improves.

We release the [training code](https://github.com/FreedomIntelligence/HuatuoGPT-3), [medical RL dataset](https://huggingface.co/datasets/FreedomIntelligence/OnePO-Medical-20K), and [8B rubric grader](https://huggingface.co/FreedomIntelligence/HuatuoGPT-3-Grader-8B).

> **HuatuoGPT-3 requires thinking mode.** Keep `enable_thinking=True` during inference. The model generates reasoning before providing its final answer after `</think>`.

## Model Info

| Model | Backbone | Purpose | Access |
| --- | --- | --- | --- |
| HuatuoGPT-3-8B | Qwen3-8B-Base | Medical reasoning | [HF Link](https://huggingface.co/FreedomIntelligence/HuatuoGPT-3-8B) |
| **HuatuoGPT-3-9B** | **Qwen3.5-9B** | **Medical reasoning** | [HF Link](https://huggingface.co/FreedomIntelligence/HuatuoGPT-3-9B) |
| HuatuoGPT-3-32B | Qwen3-32B | Medical reasoning | [HF Link](https://huggingface.co/FreedomIntelligence/HuatuoGPT-3-32B) |
| HuatuoGPT-3-Grader-8B | Qwen3-8B | Rubric scoring | [HF Link](https://huggingface.co/FreedomIntelligence/HuatuoGPT-3-Grader-8B) |

## Usage

**HuatuoGPT-3-9B** can be used like [Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) and deployed with [vLLM](https://github.com/vllm-project/vllm) or [SGLang](https://github.com/sgl-project/sglang).

For direct text inference, use a Transformers version with Qwen3.5 support (`transformers>=5.4.0`) and `accelerate`:

```
from transformers import AutoModelForImageTextToText, AutoProcessor

model_id = "FreedomIntelligence/HuatuoGPT-3-9B"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
    model_id,
    dtype="auto",
    device_map="auto",
).eval()

messages = [{
    "role": "user",
    "content": [{"type": "text", "text": "What are the common causes of chest pain?"}],
}]
inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    add_generation_prompt=True,
    enable_thinking=True,
    return_dict=True,
    return_tensors="pt",
).to(model.device)

outputs = model.generate(**inputs, max_new_tokens=4096)
response = outputs[0, inputs["input_ids"].shape[-1]:]
print(processor.decode(response, skip_special_tokens=True))
```

## πŸ“– Citation

```
@inproceedings{chen2026onepo,
  title={OnePO: Direct One-stage Policy Optimization for SFT-free Domain Adaptation},
  author={Chen, Junying and Xie, Xinyuan and Li, Ziniu and Wang, Benyou},
  booktitle={Proceedings of the 43rd International Conference on Machine Learning},
  year={2026}
}
```

Downloads last month
:   -

 

Safetensors

Model size

9B params

Tensor type

BF16

Β·

Chat template

Files info

  

## Model tree for FreedomIntelligence/HuatuoGPT-3-9B

Base model

[Qwen/Qwen3.5-9B-Base](https://huggingface.co/Qwen/Qwen3.5-9B-Base)

Finetuned

  [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B)

Finetuned

 ([818](https://huggingface.co/models?other=base_model:finetune:Qwen/Qwen3.5-9B)) 

this model
jacek2023228
🟧 echo.github ⭐The linked model card identifies this repository as the released HuatuoGPT-3 training code and describes OnePO as single-stage medical RL adFreedomIntelligenceβ€”β€”
🟠 redditFreedomIntelligence/HuatuoGPT-3-27B · Hugging Face
LocalLLaMA
jacek2023383

Interpretation history

Decision trace