TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK
Multimodal Machine Translation	Multi30K	IKD-MMT	BLEU (EN-DE)	41.28	# 2
Multimodal Machine Translation	Multi30K	IKD-MMT	Meteor (EN-DE)	58.93	# 1
Multimodal Machine Translation	Multi30K	IKD-MMT	Meteor (EN-FR)	77.20	# 1

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/distill-the-image-to-nowhere-inversion/multimodal-machine-translation-on-multi30k)](https://paperswithcode.com/sota/multimodal-machine-translation-on-multi30k?p=distill-the-image-to-nowhere-inversion)`

Distill the Image to Nowhere: Inversion Knowledge Distillation for Multimodal Machine Translation

10 Oct 2022 · Ru Peng, Yawen Zeng, Junbo Zhao ·

Past works on multimodal machine translation (MMT) elevate bilingual setup by incorporating additional aligned vision information. However, an image-must requirement of the multimodal dataset largely hinders MMT's development -- namely that it demands an aligned form of [image, source text, target text]. This limitation is generally troublesome during the inference phase especially when the aligned image is not provided as in the normal NMT setup. Thus, in this work, we introduce IKD-MMT, a novel MMT framework to support the image-free inference phase via an inversion knowledge distillation scheme. In particular, a multimodal feature generator is executed with a knowledge distillation module, which directly generates the multimodal feature from (only) source texts as the input. While there have been a few prior works entertaining the possibility to support image-free inference for machine translation, their performances have yet to rival the image-must translation. In our experiments, we identify our method as the first image-free approach to comprehensively rival or even surpass (almost) all image-must frameworks, and achieved the state-of-the-art result on the often-used Multi30k benchmark. Our code and data are available at: https://github.com/pengr/IKD-mmt/tree/master..

PDF Abstract