TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK
Utterance-level pronounciation scoring	speechocean762	3MH	Pearson correlation coefficient (PCC)	0.811	# 1
Word-level pronunciation scoring	speechocean762	3MH	Pearson correlation coefficient (PCC)	0.694	# 1
Phone-level pronunciation scoring	speechocean762	3MH	Pearson correlation coefficient (PCC)	0.693	# 1

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/a-hierarchical-context-aware-modeling/utterance-level-pronounciation-scoring-on)](https://paperswithcode.com/sota/utterance-level-pronounciation-scoring-on?p=a-hierarchical-context-aware-modeling)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/a-hierarchical-context-aware-modeling/word-level-pronunciation-scoring-on)](https://paperswithcode.com/sota/word-level-pronunciation-scoring-on?p=a-hierarchical-context-aware-modeling)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/a-hierarchical-context-aware-modeling/phone-level-pronunciation-scoring-on)](https://paperswithcode.com/sota/phone-level-pronunciation-scoring-on?p=a-hierarchical-context-aware-modeling)`

A Hierarchical Context-aware Modeling Approach for Multi-aspect and Multi-granular Pronunciation Assessment

29 May 2023 · Fu-An Chao, Tien-Hong Lo, Tzu-I Wu, Yao-Ting Sung, Berlin Chen ·

Automatic Pronunciation Assessment (APA) plays a vital role in Computer-assisted Pronunciation Training (CAPT) when evaluating a second language (L2) learner's speaking proficiency. However, an apparent downside of most de facto methods is that they parallelize the modeling process throughout different speech granularities without accounting for the hierarchical and local contextual relationships among them. In light of this, a novel hierarchical approach is proposed in this paper for multi-aspect and multi-granular APA. Specifically, we first introduce the notion of sup-phonemes to explore more subtle semantic traits of L2 speakers. Second, a depth-wise separable convolution layer is exploited to better encapsulate the local context cues at the sub-word level. Finally, we use a score-restraint attention pooling mechanism to predict the sentence-level scores and optimize the component models with a multitask learning (MTL) framework. Extensive experiments carried out on a publicly-available benchmark dataset, viz. speechocean762, demonstrate the efficacy of our approach in relation to some cutting-edge baselines.

PDF Abstract

Code

Add Remove Mark official

No code implementations yet. Submit your code now

Tasks

Add Remove

Automatic Speech Recognition

Multi-Task Learning

Phone-level pronunciation scoring

Sentence

Utterance-level pronounciation scoring

Word-level pronunciation scoring

Datasets

LibriSpeech speechocean762

Results from the Paper

Add Remove

Ranked #1 on Utterance-level pronounciation scoring on speechocean762

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Benchmark
Utterance-level pronounciation scoring	speechocean762	3MH	Pearson correlation coefficient (PCC)	0.811	# 1	Compare
Word-level pronunciation scoring	speechocean762	3MH	Pearson correlation coefficient (PCC)	0.694	# 1	Compare
Phone-level pronunciation scoring	speechocean762	3MH	Pearson correlation coefficient (PCC)	0.693	# 1	Compare

Methods

Add Remove

APA • Convolution

Edit Social Preview

A Hierarchical Context-aware Modeling Approach for Multi-aspect and Multi-granular Pronunciation Assessment

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit Add Remove

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Add Remove

Methods

Add Remove