Just Ask: Learning to Answer Questions from Millions of Narrated Videos

Recent methods for visual question answering rely on large-scale annotated datasets. Manual annotation of questions and answers for videos, however, is tedious, expensive and prevents scalability. In this work, we propose to avoid manual annotation and generate a large-scale training dataset for video question answering making use of automatic cross-modal supervision. We leverage a question generation transformer trained on text data and use it to generate question-answer pairs from transcribed video narrations. Given narrated videos, we then automatically generate the HowToVQA69M dataset with 69M video-question-answer triplets. To handle the open vocabulary of diverse answers in this dataset, we propose a training procedure based on a contrastive loss between a video-question multi-modal transformer and an answer transformer. We introduce the zero-shot VideoQA task and show excellent results, in particular for rare answers. Furthermore, we demonstrate our method to significantly outperform the state of the art on MSRVTT-QA, MSVD-QA, ActivityNet-QA and How2QA. Finally, for a detailed evaluation we introduce iVQA, a new VideoQA dataset with reduced language biases and high-quality redundant manual annotations. Our code, datasets and trained models are available at https://antoyang.github.io/just-ask.html.

PDF Abstract ICCV 2021 PDF ICCV 2021 Abstract

Datasets


Results from the Paper


Task Dataset Model Metric Name Metric Value Global Rank Uses Extra
Training Data
Result Benchmark
Video Question Answering ActivityNet-QA Just Ask (0-shot) Accuracy 12.2 # 34
Video Question Answering ActivityNet-QA Just Ask (fine-tune) Accuracy 38.9 # 26
Video Question Answering How2QA Just Ask Accuracy 84.4 # 3
Video Question Answering How2QA Just Ask (0-shot) Accuracy 51.1 # 8
Video Question Answering iVQA Just Ask (fine-tune) Accuracy 35.4 # 5
Video Question Answering iVQA Just Ask (0-shot) Accuracy 12.2 # 7
Visual Question Answering MSRVTT-QA Just Ask Accuracy 0.415 # 2
Visual Question Answering MSVD-QA Just Ask Accuracy 0.463 # 2
Video Question Answering VideoQA Just Ask (fine-tune) Accuracy 15.6 # 1

Methods


No methods listed for this paper. Add relevant methods here