TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK	REMOVE
Sign Language Recognition	AUTSL	STF+LSTM	Rank-1 Recognition Rate	0.9856	# 1
Audio-Visual Speech Recognition	LRW	2DCNN + BiLSTM + ResNet + MLF	Top-1 Accuracy	98.76	# 1

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/audio-visual-speech-and-gesture-recognition/sign-language-recognition-on-autsl)](https://paperswithcode.com/sota/sign-language-recognition-on-autsl?p=audio-visual-speech-and-gesture-recognition)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/audio-visual-speech-and-gesture-recognition/audio-visual-speech-recognition-on-lrw)](https://paperswithcode.com/sota/audio-visual-speech-recognition-on-lrw?p=audio-visual-speech-and-gesture-recognition)`

Audio-Visual Speech and Gesture Recognition by Sensors of Mobile Devices

Sensors 2023 · Dmitry Ryumin, Denis Ivanko, Elena Ryumina ·

Audio-visual speech recognition (AVSR) is one of the most promising solutions for reliable speech recognition, particularly when audio is corrupted by noise. Additional visual information can be used for both automatic lip-reading and gesture recognition. Hand gestures are a form of non-verbal communication and can be used as a very important part of modern human–computer interaction systems. Currently, audio and video modalities are easily accessible by sensors of mobile devices. However, there is no out-of-the-box solution for automatic audio-visual speech and gesture recognition. This study introduces two deep neural network-based model architectures: one for AVSR and one for gesture recognition. The main novelty regarding audio-visual speech recognition lies in fine-tuning strategies for both visual and acoustic features and in the proposed end-to-end model, which considers three modality fusion approaches: prediction-level, feature-level, and model-level. The main novelty in gesture recognition lies in a unique set of spatio-temporal features, including those that consider lip articulation information. As there are no available datasets for the combined task, we evaluated our methods on two different large-scale corpora—LRW and AUTSL—and outperformed existing methods on both audio-visual speech recognition and gesture recognition tasks. We achieved AVSR accuracy for the LRW dataset equal to 98.76% and gesture recognition rate for the AUTSL dataset equal to 98.56%. The results obtained demonstrate not only the high performance of the proposed methodology, but also the fundamental possibility of recognizing audio-visual speech and gestures by sensors of mobile devices.

PDF Abstract

Code

Add Remove Mark official

No code implementations yet. Submit your code now

Tasks

Add Remove

Audio-Visual Speech Recognition

Gesture Recognition

Lip Reading

Sign Language Recognition

speech-recognition

Speech Recognition

Visual Speech Recognition

Datasets

LRW WLASL

AUTSL

Results from the Paper

Add Remove

Ranked #1 on Sign Language Recognition on AUTSL

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Benchmark
Sign Language Recognition	AUTSL	STF+LSTM	Rank-1 Recognition Rate	0.9856	# 1	Compare
Audio-Visual Speech Recognition	LRW	2DCNN + BiLSTM + ResNet + MLF	Top-1 Accuracy	98.76	# 1	Compare

Methods

Add Remove

No methods listed for this paper. Add relevant methods here

Edit Social Preview

Audio-Visual Speech and Gesture Recognition by Sensors of Mobile Devices

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit Add Remove

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Add Remove

Methods

Add Remove