TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK
Gesture Generation	BEAT	Trimodal	FID	177.2	# 2
Gesture Generation	BEAT2	Trimodal	FGD	1.241	# 5
Gesture Generation	TED Gesture Dataset	Trimodal	FGD	3.729	# 5

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/speech-gesture-generation-from-the-trimodal/gesture-generation-on-beat)](https://paperswithcode.com/sota/gesture-generation-on-beat?p=speech-gesture-generation-from-the-trimodal)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/speech-gesture-generation-from-the-trimodal/gesture-generation-on-beat2)](https://paperswithcode.com/sota/gesture-generation-on-beat2?p=speech-gesture-generation-from-the-trimodal)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/speech-gesture-generation-from-the-trimodal/gesture-generation-on-ted-gesture-dataset)](https://paperswithcode.com/sota/gesture-generation-on-ted-gesture-dataset?p=speech-gesture-generation-from-the-trimodal)`

Speech Gesture Generation from the Trimodal Context of Text, Audio, and Speaker Identity

4 Sep 2020 · Youngwoo Yoon, Bok Cha, Joo-Haeng Lee, Minsu Jang, Jaeyeon Lee, Jaehong Kim, Geehyuk Lee ·

For human-like agents, including virtual avatars and social robots, making proper gestures while speaking is crucial in human--agent interaction. Co-speech gestures enhance interaction experiences and make the agents look alive. However, it is difficult to generate human-like gestures due to the lack of understanding of how people gesture. Data-driven approaches attempt to learn gesticulation skills from human demonstrations, but the ambiguous and individual nature of gestures hinders learning. In this paper, we present an automatic gesture generation model that uses the multimodal context of speech text, audio, and speaker identity to reliably generate gestures. By incorporating a multimodal context and an adversarial training scheme, the proposed model outputs gestures that are human-like and that match with speech content and rhythm. We also introduce a new quantitative evaluation metric for gesture generation models. Experiments with the introduced metric and subjective human evaluation showed that the proposed gesture generation model is better than existing end-to-end generation models. We further confirm that our model is able to work with synthesized audio in a scenario where contexts are constrained, and show that different gesture styles can be generated for the same speech by specifying different speaker identities in the style embedding space that is learned from videos of various speakers. All the code and data is available at https://github.com/ai4r/Gesture-Generation-from-Trimodal-Context.

PDF Abstract

Code

Add Remove Mark official

ai4r/Gesture-Generation-from-Trimod… official

234

PantoMatrix/BEAT

↳ Quickstart in

Colab

Tasks

Add Remove

Gesture Generation

Datasets

BEAT TED Gesture Dataset

BEAT2

Results from the Paper

Edit

Ranked #2 on Gesture Generation on BEAT

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Benchmark
Gesture Generation	BEAT	Trimodal	FID	177.2	# 2	Compare
Gesture Generation	BEAT2	Trimodal	FGD	1.241	# 5	Compare
Gesture Generation	TED Gesture Dataset	Trimodal	FGD	3.729	# 5	Compare

Methods

Add Remove

No methods listed for this paper. Add relevant methods here

Edit Social Preview

Speech Gesture Generation from the Trimodal Context of Text, Audio, and Speaker Identity

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Edit

Methods

Add Remove