TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK
Multiple Choice Question Answering (MCQA)	MedMCQA	BioMedGPT-10B	Test Set (Acc-%)	0.514	# 6
Question Answering	MedQA	BioMedGPT-10B	Accuracy	50.4	# 13
Multiple Choice Question Answering (MCQA)	MMLU (Professional medicine)	BioMedGPT-LM-7B	Accuracy	51.1	# 4
Question Answering	PubChemQA	BioMedGPT-10B	BLEU-2	0.234	# 1
Question Answering	PubChemQA	BioMedGPT-10B	BLEU-4	0.141	# 1
Question Answering	PubChemQA	BioMedGPT-10B	ROUGE-1	0.386	# 1
Question Answering	PubChemQA	BioMedGPT-10B	ROUGE-2	0.206	# 1
Question Answering	PubChemQA	BioMedGPT-10B	ROUGE-L	0.332	# 1
Question Answering	PubChemQA	BioMedGPT-10B	MEATOR	0.308	# 1
Question Answering	PubMedQA	BioMedGPT-10B	Accuracy	76.1	# 11
Question Answering	UniProtQA	BioMedGPT-10B	BLEU-2	0.571	# 1
Question Answering	UniProtQA	BioMedGPT-10B	BLEU-4	0.535	# 1
Question Answering	UniProtQA	BioMedGPT-10B	ROUGE-1	0.743	# 1
Question Answering	UniProtQA	BioMedGPT-10B	ROUGE-2	0.759	# 1
Question Answering	UniProtQA	BioMedGPT-10B	ROUGE-L	0.622	# 1
Question Answering	UniProtQA	BioMedGPT-10B	MEATOR	0.754	# 1

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/biomedgpt-open-multimodal-generative-pre/question-answering-on-pubchemqa)](https://paperswithcode.com/sota/question-answering-on-pubchemqa?p=biomedgpt-open-multimodal-generative-pre)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/biomedgpt-open-multimodal-generative-pre/question-answering-on-uniprotqa)](https://paperswithcode.com/sota/question-answering-on-uniprotqa?p=biomedgpt-open-multimodal-generative-pre)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/biomedgpt-open-multimodal-generative-pre/multiple-choice-question-answering-mcqa-on-25)](https://paperswithcode.com/sota/multiple-choice-question-answering-mcqa-on-25?p=biomedgpt-open-multimodal-generative-pre)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/biomedgpt-open-multimodal-generative-pre/multiple-choice-question-answering-mcqa-on-21)](https://paperswithcode.com/sota/multiple-choice-question-answering-mcqa-on-21?p=biomedgpt-open-multimodal-generative-pre)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/biomedgpt-open-multimodal-generative-pre/question-answering-on-pubmedqa)](https://paperswithcode.com/sota/question-answering-on-pubmedqa?p=biomedgpt-open-multimodal-generative-pre)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/biomedgpt-open-multimodal-generative-pre/question-answering-on-medqa-usmle)](https://paperswithcode.com/sota/question-answering-on-medqa-usmle?p=biomedgpt-open-multimodal-generative-pre)`

BioMedGPT: Open Multimodal Generative Pre-trained Transformer for BioMedicine

18 Aug 2023 · Yizhen Luo, Jiahuan Zhang, Siqi Fan, Kai Yang, Yushuai Wu, Mu Qiao, Zaiqing Nie ·

Foundation models (FMs) have exhibited remarkable performance across a wide range of downstream tasks in many domains. Nevertheless, general-purpose FMs often face challenges when confronted with domain-specific problems, due to their limited access to the proprietary training data in a particular domain. In biomedicine, there are various biological modalities, such as molecules, proteins, and cells, which are encoded by the language of life and exhibit significant modality gaps with human natural language. In this paper, we introduce BioMedGPT, an open multimodal generative pre-trained transformer (GPT) for biomedicine, to bridge the gap between the language of life and human natural language. BioMedGPT allows users to easily ``communicate'' with diverse biological modalities through free text, which is the first of its kind. BioMedGPT aligns different biological modalities with natural language via a large generative language model, namely, BioMedGPT-LM. We publish BioMedGPT-10B, which unifies the feature spaces of molecules, proteins, and natural language via encoding and alignment. Through fine-tuning, BioMedGPT-10B outperforms or is on par with human and significantly larger general-purpose foundation models on the biomedical QA task. It also demonstrates promising performance in the molecule QA and protein QA tasks, which could greatly accelerate the discovery of new drugs and therapeutic targets. In addition, BioMedGPT-LM-7B is the first large generative language model based on Llama2 in the biomedical domain, therefore is commercial friendly. Both BioMedGPT-10B and BioMedGPT-LM-7B are open-sourced to the research community. In addition, we publish the datasets that are meticulously curated for the alignment of multi-modalities, i.e., PubChemQA and UniProtQA. All the models, codes, and datasets are available at \url{https://github.com/PharMolix/OpenBioMed}.

PDF Abstract

Code

Add Remove Mark official

pharmolix/openbiomed official

590

Tasks

Add Remove

Language Modelling

Multiple Choice Question Answering (MCQA)

Question Answering

Datasets

Introduced in the Paper:

UniProtQA PubChemQA

Used in the Paper:

MMLU

PubMedQA

MedQA

MedMCQA

Results from the Paper

Edit

Ranked #1 on Question Answering on PubChemQA

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Benchmark
Multiple Choice Question Answering (MCQA)	MedMCQA	BioMedGPT-10B	Test Set (Acc-%)	0.514	# 6	Compare
Question Answering	MedQA	BioMedGPT-10B	Accuracy	50.4	# 13	Compare
Multiple Choice Question Answering (MCQA)	MMLU (Professional medicine)	BioMedGPT-LM-7B	Accuracy	51.1	# 4	Compare
Question Answering	PubChemQA	BioMedGPT-10B	BLEU-2	0.234	# 1	Compare
			BLEU-4	0.141	# 1	Compare
			ROUGE-1	0.386	# 1	Compare
			ROUGE-2	0.206	# 1	Compare
			ROUGE-L	0.332	# 1	Compare
			MEATOR	0.308	# 1	Compare
Question Answering	PubMedQA	BioMedGPT-10B	Accuracy	76.1	# 11	Compare
Question Answering	UniProtQA	BioMedGPT-10B	BLEU-2	0.571	# 1	Compare
			BLEU-4	0.535	# 1	Compare
			ROUGE-1	0.743	# 1	Compare
			ROUGE-2	0.759	# 1	Compare
			ROUGE-L	0.622	# 1	Compare
			MEATOR	0.754	# 1	Compare

Methods

Add Remove

No methods listed for this paper. Add relevant methods here

Edit Social Preview

BioMedGPT: Open Multimodal Generative Pre-trained Transformer for BioMedicine

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Edit

Methods

Add Remove