TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK
Video Question Answering	ActivityNet-QA	Just Ask (0-shot)	Accuracy	12.2	# 34
Video Question Answering	ActivityNet-QA	Just Ask (fine-tune)	Accuracy	38.9	# 26
Video Question Answering	How2QA	Just Ask	Accuracy	84.4	# 3
Video Question Answering	How2QA	Just Ask (0-shot)	Accuracy	51.1	# 8
Video Question Answering	iVQA	Just Ask (fine-tune)	Accuracy	35.4	# 5
Video Question Answering	iVQA	Just Ask (0-shot)	Accuracy	12.2	# 7
Visual Question Answering	MSRVTT-QA	Just Ask	Accuracy	0.415	# 2
Visual Question Answering	MSVD-QA	Just Ask	Accuracy	0.463	# 2
Video Question Answering	VideoQA	Just Ask (fine-tune)	Accuracy	15.6	# 1

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/just-ask-learning-to-answer-questions-from/video-question-answering-on-videoqa)](https://paperswithcode.com/sota/video-question-answering-on-videoqa?p=just-ask-learning-to-answer-questions-from)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/just-ask-learning-to-answer-questions-from/visual-question-answering-on-msrvtt-qa-2)](https://paperswithcode.com/sota/visual-question-answering-on-msrvtt-qa-2?p=just-ask-learning-to-answer-questions-from)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/just-ask-learning-to-answer-questions-from/visual-question-answering-on-msvd-qa-2)](https://paperswithcode.com/sota/visual-question-answering-on-msvd-qa-2?p=just-ask-learning-to-answer-questions-from)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/just-ask-learning-to-answer-questions-from/video-question-answering-on-how2qa)](https://paperswithcode.com/sota/video-question-answering-on-how2qa?p=just-ask-learning-to-answer-questions-from)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/just-ask-learning-to-answer-questions-from/video-question-answering-on-ivqa)](https://paperswithcode.com/sota/video-question-answering-on-ivqa?p=just-ask-learning-to-answer-questions-from)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/just-ask-learning-to-answer-questions-from/video-question-answering-on-activitynet-qa)](https://paperswithcode.com/sota/video-question-answering-on-activitynet-qa?p=just-ask-learning-to-answer-questions-from)`

Just Ask: Learning to Answer Questions from Millions of Narrated Videos

ICCV 2021 · Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, Cordelia Schmid ·

Recent methods for visual question answering rely on large-scale annotated datasets. Manual annotation of questions and answers for videos, however, is tedious, expensive and prevents scalability. In this work, we propose to avoid manual annotation and generate a large-scale training dataset for video question answering making use of automatic cross-modal supervision. We leverage a question generation transformer trained on text data and use it to generate question-answer pairs from transcribed video narrations. Given narrated videos, we then automatically generate the HowToVQA69M dataset with 69M video-question-answer triplets. To handle the open vocabulary of diverse answers in this dataset, we propose a training procedure based on a contrastive loss between a video-question multi-modal transformer and an answer transformer. We introduce the zero-shot VideoQA task and show excellent results, in particular for rare answers. Furthermore, we demonstrate our method to significantly outperform the state of the art on MSRVTT-QA, MSVD-QA, ActivityNet-QA and How2QA. Finally, for a detailed evaluation we introduce iVQA, a new VideoQA dataset with reduced language biases and high-quality redundant manual annotations. Our code, datasets and trained models are available at https://antoyang.github.io/just-ask.html.

PDF Abstract ICCV 2021 PDF ICCV 2021 Abstract

Code

Add Remove Mark official

antoyang/just-ask official

113

Tasks

Add Remove

Question Answering

Question Generation

Question-Generation

Video Question Answering

Visual Question Answering

Visual Question Answering (VQA)

Zero-Shot Learning

Datasets

Introduced in the Paper:

iVQA

HowToVQA69M

Used in the Paper:

HowTo100M

ActivityNet-QA MSRVTT-QA MSVD-QA

How2QA

Results from the Paper

Edit

Ranked #1 on Video Question Answering on VideoQA

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Benchmark
Video Question Answering	ActivityNet-QA	Just Ask (0-shot)	Accuracy	12.2	# 34	Compare
Video Question Answering	ActivityNet-QA	Just Ask (fine-tune)	Accuracy	38.9	# 26	Compare
Video Question Answering	How2QA	Just Ask	Accuracy	84.4	# 3	Compare
Video Question Answering	How2QA	Just Ask (0-shot)	Accuracy	51.1	# 8	Compare
Video Question Answering	iVQA	Just Ask (fine-tune)	Accuracy	35.4	# 5	Compare
Video Question Answering	iVQA	Just Ask (0-shot)	Accuracy	12.2	# 7	Compare
Visual Question Answering	MSRVTT-QA	Just Ask	Accuracy	0.415	# 2	Compare
Visual Question Answering	MSVD-QA	Just Ask	Accuracy	0.463	# 2	Compare
Video Question Answering	VideoQA	Just Ask (fine-tune)	Accuracy	15.6	# 1	Compare

Methods

Add Remove

No methods listed for this paper. Add relevant methods here

Edit Social Preview

Just Ask: Learning to Answer Questions from Millions of Narrated Videos

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Edit

Methods

Add Remove