TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK
Motion Captioning	HumanML3D	TM2T	BLEU-4	22.3	# 2
Motion Synthesis	HumanML3D	Language2Pose	FID	11.02	# 25
Motion Synthesis	HumanML3D	Language2Pose	Diversity	7.676	# 23
Motion Synthesis	HumanML3D	Language2Pose	R Precision Top3	0.486	# 24
Motion Synthesis	HumanML3D	Text2Gesture	FID	5.012	# 23
Motion Synthesis	HumanML3D	Text2Gesture	Diversity	6.409	# 24
Motion Synthesis	HumanML3D	Text2Gesture	R Precision Top3	0.345	# 25
Motion Synthesis	HumanML3D	TM2T	FID	1.501	# 22
Motion Synthesis	HumanML3D	TM2T	Diversity	8.589	# 21
Motion Synthesis	HumanML3D	TM2T	Multimodality	2.424	# 7
Motion Synthesis	HumanML3D	TM2T	R Precision Top3	0.729	# 20
Motion Synthesis	KIT Motion-Language	Language2Pose	FID	6.545	# 21
Motion Synthesis	KIT Motion-Language	Language2Pose	R Precision Top3	0.483	# 20
Motion Synthesis	KIT Motion-Language	Language2Pose	Diversity	9.073	# 21
Motion Synthesis	KIT Motion-Language	TM2T	FID	3.599	# 19
Motion Synthesis	KIT Motion-Language	TM2T	R Precision Top3	0.587	# 19
Motion Synthesis	KIT Motion-Language	TM2T	Diversity	9.473	# 19
Motion Synthesis	KIT Motion-Language	TM2T	Multimodality	3.292	# 2
Motion Synthesis	KIT Motion-Language	Text2Gesture	FID	12.12	# 22
Motion Synthesis	KIT Motion-Language	Text2Gesture	R Precision Top3	0.338	# 22
Motion Synthesis	KIT Motion-Language	Text2Gesture	Diversity	9.334	# 20
Motion Captioning	KIT Motion-Language	TM2T	BLEU-4	18.4	# 2

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/tm2t-stochastic-and-tokenized-modeling-for/motion-captioning-on-humanml3d)](https://paperswithcode.com/sota/motion-captioning-on-humanml3d?p=tm2t-stochastic-and-tokenized-modeling-for)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/tm2t-stochastic-and-tokenized-modeling-for/motion-captioning-on-kit-motion-language)](https://paperswithcode.com/sota/motion-captioning-on-kit-motion-language?p=tm2t-stochastic-and-tokenized-modeling-for)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/tm2t-stochastic-and-tokenized-modeling-for/motion-synthesis-on-kit-motion-language)](https://paperswithcode.com/sota/motion-synthesis-on-kit-motion-language?p=tm2t-stochastic-and-tokenized-modeling-for)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/tm2t-stochastic-and-tokenized-modeling-for/motion-synthesis-on-humanml3d)](https://paperswithcode.com/sota/motion-synthesis-on-humanml3d?p=tm2t-stochastic-and-tokenized-modeling-for)`

TM2T: Stochastic and Tokenized Modeling for the Reciprocal Generation of 3D Human Motions and Texts

4 Jul 2022 · Chuan Guo, Xinxin Zuo, Sen Wang, Li Cheng ·

Inspired by the strong ties between vision and language, the two intimate human sensing and communication modalities, our paper aims to explore the generation of 3D human full-body motions from texts, as well as its reciprocal task, shorthanded for text2motion and motion2text, respectively. To tackle the existing challenges, especially to enable the generation of multiple distinct motions from the same text, and to avoid the undesirable production of trivial motionless pose sequences, we propose the use of motion token, a discrete and compact motion representation. This provides one level playing ground when considering both motions and text signals, as the motion and text tokens, respectively. Moreover, our motion2text module is integrated into the inverse alignment process of our text2motion training pipeline, where a significant deviation of synthesized text from the input text would be penalized by a large training loss; empirically this is shown to effectively improve performance. Finally, the mappings in-between the two modalities of motions and texts are facilitated by adapting the neural model for machine translation (NMT) to our context. This autoregressive modeling of the distribution over discrete motion tokens further enables non-deterministic production of pose sequences, of variable lengths, from an input text. Our approach is flexible, could be used for both text2motion and motion2text tasks. Empirical evaluations on two benchmark datasets demonstrate the superior performance of our approach on both tasks over a variety of state-of-the-art methods. Project page: https://ericguo5513.github.io/TM2T/

PDF Abstract

Code

Add Remove Mark official

EricGuo5513/TM2T official

Tasks

Add Remove

Machine Translation

Motion Captioning

Motion Synthesis

NMT

Datasets

HumanML3D KIT Motion-Language

Results from the Paper

Edit

Ranked #2 on Motion Captioning on KIT Motion-Language

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Benchmark
Motion Captioning	HumanML3D	TM2T	BLEU-4	22.3	# 2	Compare
Motion Synthesis	HumanML3D	Language2Pose	FID	11.02	# 25	Compare
			Diversity	7.676	# 23	Compare
			R Precision Top3	0.486	# 24	Compare
Motion Synthesis	HumanML3D	Text2Gesture	FID	5.012	# 23	Compare
			Diversity	6.409	# 24	Compare
			R Precision Top3	0.345	# 25	Compare
Motion Synthesis	HumanML3D	TM2T	FID	1.501	# 22	Compare
			Diversity	8.589	# 21	Compare
			Multimodality	2.424	# 7	Compare
			R Precision Top3	0.729	# 20	Compare
Motion Synthesis	KIT Motion-Language	Language2Pose	FID	6.545	# 21	Compare
			R Precision Top3	0.483	# 20	Compare
			Diversity	9.073	# 21	Compare
Motion Synthesis	KIT Motion-Language	TM2T	FID	3.599	# 19	Compare
			R Precision Top3	0.587	# 19	Compare
			Diversity	9.473	# 19	Compare
			Multimodality	3.292	# 2	Compare
Motion Synthesis	KIT Motion-Language	Text2Gesture	FID	12.12	# 22	Compare
			R Precision Top3	0.338	# 22	Compare
			Diversity	9.334	# 20	Compare
Motion Captioning	KIT Motion-Language	TM2T	BLEU-4	18.4	# 2	Compare

Methods

Add Remove

No methods listed for this paper. Add relevant methods here

Edit Social Preview

TM2T: Stochastic and Tokenized Modeling for the Reciprocal Generation of 3D Human Motions and Texts

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Edit

Methods

Add Remove