TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK
3D Face Animation	BEAT2	CodeTalker	MSE	8.026	# 4
3D Face Animation	Biwi 3D Audiovisual Corpus of Affective Communication - B3D(AC)^2	CodeTalker	Lip Vertex Error	4.7914	# 4
3D Face Animation	Biwi 3D Audiovisual Corpus of Affective Communication - B3D(AC)^2	CodeTalker	FDD	4.1170	# 3

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/codetalker-speech-driven-3d-facial-animation/3d-face-animation-on-beat2)](https://paperswithcode.com/sota/3d-face-animation-on-beat2?p=codetalker-speech-driven-3d-facial-animation)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/codetalker-speech-driven-3d-facial-animation/3d-face-animation-on-biwi-3d-audiovisual)](https://paperswithcode.com/sota/3d-face-animation-on-biwi-3d-audiovisual?p=codetalker-speech-driven-3d-facial-animation)`

CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion Prior

CVPR 2023 · Jinbo Xing, Menghan Xia, Yuechen Zhang, Xiaodong Cun, Jue Wang, Tien-Tsin Wong ·

Speech-driven 3D facial animation has been widely studied, yet there is still a gap to achieving realism and vividness due to the highly ill-posed nature and scarcity of audio-visual data. Existing works typically formulate the cross-modal mapping into a regression task, which suffers from the regression-to-mean problem leading to over-smoothed facial motions. In this paper, we propose to cast speech-driven facial animation as a code query task in a finite proxy space of the learned codebook, which effectively promotes the vividness of the generated motions by reducing the cross-modal mapping uncertainty. The codebook is learned by self-reconstruction over real facial motions and thus embedded with realistic facial motion priors. Over the discrete motion space, a temporal autoregressive model is employed to sequentially synthesize facial motions from the input speech signal, which guarantees lip-sync as well as plausible facial expressions. We demonstrate that our approach outperforms current state-of-the-art methods both qualitatively and quantitatively. Also, a user study further justifies our superiority in perceptual quality.

PDF Abstract CVPR 2023 PDF CVPR 2023 Abstract

Code

Add Remove Mark official

Doubiiu/CodeTalker official

↳ Quickstart in

Colab

457

Tasks

Add Remove

3D Face Animation

regression

Datasets

VOCASET

BEAT2

Biwi 3D Audiovisual Corpus of Affective Communication - B3D(AC)^2

Results from the Paper

Edit

Ranked #4 on 3D Face Animation on BEAT2

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Benchmark
3D Face Animation	BEAT2	CodeTalker	MSE	8.026	# 4	Compare
3D Face Animation	Biwi 3D Audiovisual Corpus of Affective Communication - B3D(AC)^2	CodeTalker	Lip Vertex Error	4.7914	# 4	Compare
3D Face Animation		CodeTalker	FDD	4.1170	# 3	Compare

Methods

Add Remove

Absolute Position Encodings • Adam • BPE • Dense Connections • Dropout • Label Smoothing • Layer Normalization • Linear Layer • Multi-Head Attention • Position-Wise Feed-Forward Layer • Residual Connection • Scaled Dot-Product Attention • Softmax • Transformer • VQ-VAE

Edit Social Preview

CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion Prior

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Edit

Methods

Add Remove