TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK
Situation Recognition	imSitu	SituFormer	Top-1 Verb	44.2	# 3
Situation Recognition	imSitu	SituFormer	Top-1 Verb & Value	35.24	# 3
Situation Recognition	imSitu	SituFormer	Top-5 Verbs	71.21	# 3
Situation Recognition	imSitu	SituFormer	Top-5 Verbs & Value	55.75	# 3
Grounded Situation Recognition	SWiG	SituFormer	Top-1 Verb	44.2	# 3
Grounded Situation Recognition	SWiG	SituFormer	Top-1 Verb & Value	35.24	# 4
Grounded Situation Recognition	SWiG	SituFormer	Top-1 Verb & Grounded-Value	29.22	# 2
Grounded Situation Recognition	SWiG	SituFormer	Top-5 Verbs	71.21	# 3
Grounded Situation Recognition	SWiG	SituFormer	Top-5 Verbs & Value	55.75	# 3
Grounded Situation Recognition	SWiG	SituFormer	Top-5 Verbs & Grounded-Value	46	# 3

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/rethinking-the-two-stage-framework-for/situation-recognition-on-imsitu)](https://paperswithcode.com/sota/situation-recognition-on-imsitu?p=rethinking-the-two-stage-framework-for)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/rethinking-the-two-stage-framework-for/grounded-situation-recognition-on-swig)](https://paperswithcode.com/sota/grounded-situation-recognition-on-swig?p=rethinking-the-two-stage-framework-for)`

Rethinking the Two-Stage Framework for Grounded Situation Recognition

10 Dec 2021 · Meng Wei, Long Chen, Wei Ji, Xiaoyu Yue, Tat-Seng Chua ·

Grounded Situation Recognition (GSR), i.e., recognizing the salient activity (or verb) category in an image (e.g., buying) and detecting all corresponding semantic roles (e.g., agent and goods), is an essential step towards "human-like" event understanding. Since each verb is associated with a specific set of semantic roles, all existing GSR methods resort to a two-stage framework: predicting the verb in the first stage and detecting the semantic roles in the second stage. However, there are obvious drawbacks in both stages: 1) The widely-used cross-entropy (XE) loss for object recognition is insufficient in verb classification due to the large intra-class variation and high inter-class similarity among daily activities. 2) All semantic roles are detected in an autoregressive manner, which fails to model the complex semantic relations between different roles. To this end, we propose a novel SituFormer for GSR which consists of a Coarse-to-Fine Verb Model (CFVM) and a Transformer-based Noun Model (TNM). CFVM is a two-step verb prediction model: a coarse-grained model trained with XE loss first proposes a set of verb candidates, and then a fine-grained model trained with triplet loss re-ranks these candidates with enhanced verb features (not only separable but also discriminative). TNM is a transformer-based semantic role detection model, which detects all roles parallelly. Owing to the global relation modeling ability and flexibility of the transformer decoder, TNM can fully explore the statistical dependency of the roles. Extensive validations on the challenging SWiG benchmark show that SituFormer achieves a new state-of-the-art performance with significant gains under various metrics. Code is available at https://github.com/kellyiss/SituFormer.

PDF Abstract

Code

Add Remove Mark official

kellyiss/situformer official

Tasks

Add Remove

Grounded Situation Recognition

Object Recognition

Situation Recognition

Vocal Bursts Valence Prediction

Datasets

FrameNet

Results from the Paper

Edit

Ranked #3 on Situation Recognition on imSitu

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Benchmark
Situation Recognition	imSitu	SituFormer	Top-1 Verb	44.2	# 3	Compare
			Top-1 Verb & Value	35.24	# 3	Compare
			Top-5 Verbs	71.21	# 3	Compare
			Top-5 Verbs & Value	55.75	# 3	Compare
Grounded Situation Recognition	SWiG	SituFormer	Top-1 Verb	44.2	# 3	Compare
			Top-1 Verb & Value	35.24	# 4	Compare
			Top-1 Verb & Grounded-Value	29.22	# 2	Compare
			Top-5 Verbs	71.21	# 3	Compare
			Top-5 Verbs & Value	55.75	# 3	Compare
			Top-5 Verbs & Grounded-Value	46	# 3	Compare

Methods

Add Remove

Triplet Loss

Edit Social Preview

Rethinking the Two-Stage Framework for Grounded Situation Recognition

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Edit

Methods

Add Remove