TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK
Video Retrieval	ActivityNet	Ours	text-to-video R@1	25.4	# 28
Video Retrieval	ActivityNet	Ours	text-to-video R@5	59.1	# 23
Video Retrieval	ActivityNet	Ours	video-to-text R@1	26.1	# 13
Video Retrieval	ActivityNet	Ours	video-to-text R@5	60	# 11
Video Retrieval	LSMDC	Ours	text-to-video R@1	14.9	# 28
Video Retrieval	LSMDC	Ours	text-to-video R@5	33.2	# 23
Video Retrieval	LSMDC	Ours	video-to-text R@1	15.3	# 13
Video Retrieval	LSMDC	Ours	video-to-text R@5	34.1	# 10
Video Retrieval	MSR-VTT	Ours	text-to-video R@1	26	# 28
Video Retrieval	MSR-VTT	Ours	text-to-video R@5	56.7	# 21
Video Retrieval	MSR-VTT	Ours	text-to-video Median Rank	3	# 1
Video Retrieval	MSR-VTT	Ours	video-to-text R@1	26.7	# 10
Video Retrieval	MSR-VTT	Ours	video-to-text R@5	56.5	# 8
Video Retrieval	MSR-VTT	Ours	video-to-text Median Rank	3	# 4
Video-Guided Machine Translation	VATEX Chinese-to-English	Ours	BLEU-4	28.11	# 1
Video-Guided Machine Translation	VATEX English-to-Chinese	Ours	BLEU-4	32.34	# 1

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/video-and-text-matching-with-conditioned/video-guided-machine-translation-on-vatex-1)](https://paperswithcode.com/sota/video-guided-machine-translation-on-vatex-1?p=video-and-text-matching-with-conditioned)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/video-and-text-matching-with-conditioned/video-guided-machine-translation-on-vatex)](https://paperswithcode.com/sota/video-guided-machine-translation-on-vatex?p=video-and-text-matching-with-conditioned)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/video-and-text-matching-with-conditioned/video-retrieval-on-activitynet)](https://paperswithcode.com/sota/video-retrieval-on-activitynet?p=video-and-text-matching-with-conditioned)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/video-and-text-matching-with-conditioned/video-retrieval-on-lsmdc)](https://paperswithcode.com/sota/video-retrieval-on-lsmdc?p=video-and-text-matching-with-conditioned)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/video-and-text-matching-with-conditioned/video-retrieval-on-msr-vtt)](https://paperswithcode.com/sota/video-retrieval-on-msr-vtt?p=video-and-text-matching-with-conditioned)`

Video and Text Matching with Conditioned Embeddings

21 Oct 2021 · Ameen Ali, Idan Schwartz, Tamir Hazan, Lior Wolf ·

We present a method for matching a text sentence from a given corpus to a given video clip and vice versa. Traditionally video and text matching is done by learning a shared embedding space and the encoding of one modality is independent of the other. In this work, we encode the dataset data in a way that takes into account the query's relevant information. The power of the method is demonstrated to arise from pooling the interaction data between words and frames. Since the encoding of the video clip depends on the sentence compared to it, the representation needs to be recomputed for each potential match. To this end, we propose an efficient shallow neural network. Its training employs a hierarchical triplet loss that is extendable to paragraph/video matching. The method is simple, provides explainability, and achieves state-of-the-art results for both sentence-clip and video-text by a sizable margin across five different datasets: ActivityNet, DiDeMo, YouCook2, MSR-VTT, and LSMDC. We also show that our conditioned representation can be transferred to video-guided machine translation, where we improved the current results on VATEX. Source code is available at https://github.com/AmeenAli/VideoMatch.

PDF Abstract

Code

Add Remove Mark official

ameenali/videomatch official

Tasks

Add Remove

Machine Translation

Sentence

Text Matching

Translation

Video-Guided Machine Translation

Video Retrieval

Datasets

ActivityNet

MSR-VTT

HowTo100M

DiDeMo

YouCook2

LSMDC

Results from the Paper

Edit

Ranked #1 on Video-Guided Machine Translation on VATEX English-to-Chinese

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Benchmark
Video Retrieval	ActivityNet	Ours	text-to-video R@1	25.4	# 28	Compare
			text-to-video R@5	59.1	# 23	Compare
			video-to-text R@1	26.1	# 13	Compare
			video-to-text R@5	60	# 11	Compare
Video Retrieval	LSMDC	Ours	text-to-video R@1	14.9	# 28	Compare
			text-to-video R@5	33.2	# 23	Compare
			video-to-text R@1	15.3	# 13	Compare
			video-to-text R@5	34.1	# 10	Compare
Video Retrieval	MSR-VTT	Ours	text-to-video R@1	26	# 28	Compare
			text-to-video R@5	56.7	# 21	Compare
			text-to-video Median Rank	3	# 1	Compare
			video-to-text R@1	26.7	# 10	Compare
			video-to-text R@5	56.5	# 8	Compare
			video-to-text Median Rank	3	# 4	Compare
Video-Guided Machine Translation	VATEX Chinese-to-English	Ours	BLEU-4	28.11	# 1	Compare
Video-Guided Machine Translation	VATEX English-to-Chinese	Ours	BLEU-4	32.34	# 1	Compare

Methods

Add Remove

Triplet Loss

Edit Social Preview

Video and Text Matching with Conditioned Embeddings

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Edit

Methods

Add Remove