TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK	REMOVE
Image-text matching	CommercialAdsDataset	Unicoder-VL	ADD(S) AUC	83.16	# 7
Image-to-Text Retrieval	MS COCO	Unicoder-VL	Recall@10	97.2	# 5

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/unicoder-vl-a-universal-encoder-for-vision/image-to-text-retrieval-on-coco)](https://paperswithcode.com/sota/image-to-text-retrieval-on-coco?p=unicoder-vl-a-universal-encoder-for-vision)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/unicoder-vl-a-universal-encoder-for-vision/image-text-matching-on-commercialadsdataset)](https://paperswithcode.com/sota/image-text-matching-on-commercialadsdataset?p=unicoder-vl-a-universal-encoder-for-vision)`

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training

16 Aug 2019 · Gen Li, Nan Duan, Yuejian Fang, Ming Gong, Daxin Jiang, Ming Zhou ·

We propose Unicoder-VL, a universal encoder that aims to learn joint representations of vision and language in a pre-training manner. Borrow ideas from cross-lingual pre-trained models, such as XLM and Unicoder, both visual and linguistic contents are fed into a multi-layer Transformer for the cross-modal pre-training, where three pre-trained tasks are employed, including Masked Language Modeling (MLM), Masked Object Classification (MOC) and Visual-linguistic Matching (VLM). The first two tasks learn context-aware representations for input tokens based on linguistic and visual contents jointly. The last task tries to predict whether an image and a text describe each other. After pretraining on large-scale image-caption pairs, we transfer Unicoder-VL to caption-based image-text retrieval and visual commonsense reasoning, with just one additional output layer. We achieve state-of-the-art or comparable results on both two tasks and show the powerful ability of the cross-modal pre-training.

PDF Abstract