The "something something" video database for learning and evaluating visual common sense

ICCV 2017 Raghav GoyalSamira Ebrahimi KahouVincent MichalskiJoanna MaterzyńskaSusanne WestphalHeuna KimValentin HaenelIngo FruendPeter YianilosMoritz Mueller-FreitagFlorian HoppeChristian ThurauIngo BaxRoland Memisevic

Neural networks trained on datasets such as ImageNet have led to major advances in visual object classification. One obstacle that prevents networks from reasoning more deeply about complex scenes and situations, and from integrating visual knowledge with natural language, like humans do, is their lack of common sense knowledge about the physical world... (read more)

PDF Abstract

Evaluation results from the paper

  Submit results from this paper to get state-of-the-art GitHub badges and help community compare results to other papers.