TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK	REMOVE
Action Recognition	Something-Something V2	model3D_1 with left-right augmentation and fps jitter	Top-1 Accuracy	51.33	# 115
Action Recognition	Something-Something V2	model3D_1 with left-right augmentation and fps jitter	Top-5 Accuracy	80.46	# 84

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/the-something-something-video-database-for/action-recognition-in-videos-on-something)](https://paperswithcode.com/sota/action-recognition-in-videos-on-something?p=the-something-something-video-database-for)`

The "something something" video database for learning and evaluating visual common sense

ICCV 2017 · Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzyńska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, Florian Hoppe, Christian Thurau, Ingo Bax, Roland Memisevic ·

Neural networks trained on datasets such as ImageNet have led to major advances in visual object classification. One obstacle that prevents networks from reasoning more deeply about complex scenes and situations, and from integrating visual knowledge with natural language, like humans do, is their lack of common sense knowledge about the physical world. Videos, unlike still images, contain a wealth of detailed information about the physical world. However, most labelled video datasets represent high-level concepts rather than detailed physical aspects about actions and scenes. In this work, we describe our ongoing collection of the "something-something" database of video prediction tasks whose solutions require a common sense understanding of the depicted situation. The database currently contains more than 100,000 videos across 174 classes, which are defined as caption-templates. We also describe the challenges in crowd-sourcing this data at scale.

PDF Abstract ICCV 2017 PDF ICCV 2017 Abstract

Code

Add Remove Mark official

jayleicn/singularity

124

bit-ml/dyreg-gnn

caspillaga/Conv3DSelfAttention

latte488/smth-smth-v2

akshyta/Human-Activity-Recognition

Tasks

Add Remove

Action Recognition

Common Sense Reasoning

General Classification

Video Prediction

Datasets

Introduced in the Paper:

Something-Something V2

Something-Something V1

Used in the Paper:

ImageNet

Results from the Paper

Edit

Ranked #115 on Action Recognition on Something-Something V2

Get a GitHub badge

Results from Other Papers

Task	Dataset	Model	Metric Name	Metric Value	Rank	Source Paper	Compare
Action Recognition	Something-Something V2	model3D_1 with left-right augmentation and fps jitter	Top-1 Accuracy	51.33	# 115		See all
Action Recognition	Something-Something V2	model3D_1 with left-right augmentation and fps jitter	Top-5 Accuracy	80.46	# 84		See all

Methods

Add Remove

No methods listed for this paper. Add relevant methods here

Edit Social Preview

The "something something" video database for learning and evaluating visual common sense

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit