TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK
Speech Enhancement	VoiceBank + DEMAND	MANNER-S + MV-AT (8.1GF)	PESQ	3.12	# 13
Speech Enhancement	VoiceBank + DEMAND	MANNER-S + MV-AT (8.1GF)	CSIG	4.45	# 6
Speech Enhancement	VoiceBank + DEMAND	MANNER-S + MV-AT (8.1GF)	CBAK	3.61	# 4
Speech Enhancement	VoiceBank + DEMAND	MANNER-S + MV-AT (8.1GF)	COVL	3.82	# 8
Speech Enhancement	VoiceBank + DEMAND	MANNER-S + MV-AT (8.1GF)	STOI	95	# 5
Speech Enhancement	VoiceBank + DEMAND	MANNER-S + MV-AT (8.1GF)	Para. (M)	1.38	# 3

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/multi-view-attention-transfer-for-efficient/speech-enhancement-on-demand)](https://paperswithcode.com/sota/speech-enhancement-on-demand?p=multi-view-attention-transfer-for-efficient)`

Multi-View Attention Transfer for Efficient Speech Enhancement

22 Aug 2022 · WooSeok Shin, Hyun Joon Park, Jin Sob Kim, Byung Hoon Lee, Sung Won Han ·

Recent deep learning models have achieved high performance in speech enhancement; however, it is still challenging to obtain a fast and low-complexity model without significant performance degradation. Previous knowledge distillation studies on speech enhancement could not solve this problem because their output distillation methods do not fit the speech enhancement task in some aspects. In this study, we propose multi-view attention transfer (MV-AT), a feature-based distillation, to obtain efficient speech enhancement models in the time domain. Based on the multi-view features extraction model, MV-AT transfers multi-view knowledge of the teacher network to the student network without additional parameters. The experimental results show that the proposed method consistently improved the performance of student models of various sizes on the Valentini and deep noise suppression (DNS) datasets. MANNER-S-8.1GF with our proposed method, a lightweight model for efficient deployment, achieved 15.4x and 4.71x fewer parameters and floating-point operations (FLOPs), respectively, compared to the baseline model with similar performance.

PDF Abstract

Code

Add Remove Mark official

No code implementations yet. Submit your code now

Tasks

Add Remove

Knowledge Distillation

Speech Enhancement

Datasets

VoiceBank + DEMAND

Results from the Paper

Edit

Ranked #14 on Speech Enhancement on VoiceBank + DEMAND

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Benchmark
Speech Enhancement	VoiceBank + DEMAND	MANNER-S + MV-AT (8.1GF)	PESQ	3.12	# 13	Compare
			CSIG	4.45	# 6	Compare
			CBAK	3.61	# 4	Compare
			COVL	3.82	# 8	Compare
			STOI	95	# 5	Compare
			Para. (M)	1.38	# 3	Compare

Methods

Add Remove

Knowledge Distillation

Edit Social Preview

Multi-View Attention Transfer for Efficient Speech Enhancement

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Edit

Methods

Add Remove