TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK
Question Answering	COPA	KELM (finetuning BERT-large based single model)	Accuracy	78.0	# 39
Question Answering	MultiRC	KELM (finetuning BERT-large based single model)	F1	70.8	# 15
Question Answering	MultiRC	KELM (finetuning BERT-large based single model)	EM	27.2	# 10
Common Sense Reasoning	ReCoRD	KELM (finetuning RoBERTa-large based single model)	F1	89.6	# 14
Common Sense Reasoning	ReCoRD	KELM (finetuning RoBERTa-large based single model)	EM	89.1	# 11
Common Sense Reasoning	ReCoRD	KELM (finetuning BERT-large based single model)	F1	76.7	# 24
Common Sense Reasoning	ReCoRD	KELM (finetuning BERT-large based single model)	EM	76.2	# 21

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/kelm-knowledge-enhanced-pre-trained-language/common-sense-reasoning-on-record)](https://paperswithcode.com/sota/common-sense-reasoning-on-record?p=kelm-knowledge-enhanced-pre-trained-language)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/kelm-knowledge-enhanced-pre-trained-language/question-answering-on-multirc)](https://paperswithcode.com/sota/question-answering-on-multirc?p=kelm-knowledge-enhanced-pre-trained-language)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/kelm-knowledge-enhanced-pre-trained-language/question-answering-on-copa)](https://paperswithcode.com/sota/question-answering-on-copa?p=kelm-knowledge-enhanced-pre-trained-language)`

KELM: Knowledge Enhanced Pre-Trained Language Representations with Message Passing on Hierarchical Relational Graphs

9 Sep 2021 · Yinquan Lu, Haonan Lu, Guirong Fu, Qun Liu ·

Incorporating factual knowledge into pre-trained language models (PLM) such as BERT is an emerging trend in recent NLP studies. However, most of the existing methods combine the external knowledge integration module with a modified pre-training loss and re-implement the pre-training process on the large-scale corpus. Re-pretraining these models is usually resource-consuming, and difficult to adapt to another domain with a different knowledge graph (KG). Besides, those works either cannot embed knowledge context dynamically according to textual context or struggle with the knowledge ambiguity issue. In this paper, we propose a novel knowledge-aware language model framework based on fine-tuning process, which equips PLM with a unified knowledge-enhanced text graph that contains both text and multi-relational sub-graphs extracted from KG. We design a hierarchical relational-graph-based message passing mechanism, which can allow the representations of injected KG and text to mutually update each other and can dynamically select ambiguous mentioned entities that share the same text. Our empirical results show that our model can efficiently incorporate world knowledge from KGs into existing language models such as BERT, and achieve significant improvement on the machine reading comprehension (MRC) task compared with other knowledge-enhanced models.

PDF Abstract

Code

Add Remove Mark official

nlp-anonymous-happy/anonymous-kg-gu… official

Tasks

Add Remove

Common Sense Reasoning

Language Modelling

Machine Reading Comprehension

Question Answering

Reading Comprehension

World Knowledge

Datasets

SuperGLUE

COPA

NELL

MultiRC

ReCoRD

Results from the Paper

Edit

Ranked #11 on Common Sense Reasoning on ReCoRD

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Benchmark
Question Answering	COPA	KELM (finetuning BERT-large based single model)	Accuracy	78.0	# 39	Compare
Question Answering	MultiRC	KELM (finetuning BERT-large based single model)	F1	70.8	# 15	Compare
Question Answering	MultiRC	KELM (finetuning BERT-large based single model)	EM	27.2	# 10	Compare
Common Sense Reasoning	ReCoRD	KELM (finetuning RoBERTa-large based single model)	F1	89.6	# 14	Compare
Common Sense Reasoning	ReCoRD	KELM (finetuning RoBERTa-large based single model)	EM	89.1	# 11	Compare
Common Sense Reasoning	ReCoRD	KELM (finetuning BERT-large based single model)	F1	76.7	# 24	Compare
Common Sense Reasoning	ReCoRD	KELM (finetuning BERT-large based single model)	EM	76.2	# 21	Compare

Methods

Add Remove

Adam • Attention Dropout • BERT • Dense Connections • Dropout • GELU • Layer Normalization • Linear Layer • Linear Warmup With Linear Decay • Multi-Head Attention • Residual Connection • Scaled Dot-Product Attention • Softmax • Weight Decay • WordPiece

Edit Social Preview

KELM: Knowledge Enhanced Pre-Trained Language Representations with Message Passing on Hierarchical Relational Graphs

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Edit

Methods

Add Remove