LRS3-TED

Introduced by Afouras et al. in LRS3-TED: a large-scale dataset for visual speech recognition

LRS3-TED is a multi-modal dataset for visual and audio-visual speech recognition. It includes face tracks from over 400 hours of TED and TEDx videos, along with the corresponding subtitles and word alignment boundaries. The new dataset is substantially larger in scale compared to other public datasets that are available for general research.

Source: LRS3-TED: a large-scale dataset for visual speech recognition

Homepage

Benchmarks

Add a new result Link an existing benchmark

Task	Dataset Variant	Best Model
Lipreading	LRS3-TED	CTC/Attention
Audio-Visual Speech Recognition	LRS3-TED	CTC/Attention
Visual Speech Recognition	LRS3-TED	CTC/Attention
Speech Recognition	LRS3-TED	AV-HuBERT Large
Visual Keyword Spotting	LRS3-TED	Transpotter
Automatic Speech Recognition (ASR)	LRS3-TED	CTC/Attention