TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK	REMOVE
Human-Object Interaction Detection	HICO-DET	RLIP-ParSe (ResNet-50)	mAP	32.84	# 16
Human-Object Interaction Detection	HICO-DET	ParSe (ResNet-101)	mAP	32.76	# 17

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/rlip-relational-language-image-pre-training/human-object-interaction-detection-on-hico)](https://paperswithcode.com/sota/human-object-interaction-detection-on-hico?p=rlip-relational-language-image-pre-training)`

RLIP: Relational Language-Image Pre-training for Human-Object Interaction Detection

5 Sep 2022 · Hangjie Yuan, Jianwen Jiang, Samuel Albanie, Tao Feng, Ziyuan Huang, Dong Ni, Mingqian Tang ·

The task of Human-Object Interaction (HOI) detection targets fine-grained visual parsing of humans interacting with their environment, enabling a broad range of applications. Prior work has demonstrated the benefits of effective architecture design and integration of relevant cues for more accurate HOI detection. However, the design of an appropriate pre-training strategy for this task remains underexplored by existing approaches. To address this gap, we propose Relational Language-Image Pre-training (RLIP), a strategy for contrastive pre-training that leverages both entity and relation descriptions. To make effective use of such pre-training, we make three technical contributions: (1) a new Parallel entity detection and Sequential relation inference (ParSe) architecture that enables the use of both entity and relation descriptions during holistically optimized pre-training; (2) a synthetic data generation framework, Label Sequence Extension, that expands the scale of language data available within each minibatch; (3) mechanisms to account for ambiguity, Relation Quality Labels and Relation Pseudo-Labels, to mitigate the influence of ambiguous/noisy samples in the pre-training data. Through extensive experiments, we demonstrate the benefits of these contributions, collectively termed RLIP-ParSe, for improved zero-shot, few-shot and fine-tuning HOI detection performance as well as increased robustness to learning from noisy annotations. Code will be available at https://github.com/JacobYuan7/RLIP.

PDF Abstract

Code

Add Remove Mark official

jacobyuan7/rlip official

jacobyuan7/rlipv2

jacobyuan7/ocn-hoi-benchmark

Tasks

Add Remove

Human-Object Interaction Detection

Relation

Synthetic Data Generation

Datasets

HICO-DET

V-COCO

HICO

Results from the Paper

Edit

Ranked #16 on Human-Object Interaction Detection on HICO-DET

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Result	Benchmark
Human-Object Interaction Detection	HICO-DET	RLIP-ParSe (ResNet-50)	mAP	32.84	# 16		Compare
Human-Object Interaction Detection	HICO-DET	ParSe (ResNet-101)	mAP	32.76	# 17		Compare

Methods

Add Remove

Visual Parsing

Edit Social Preview

RLIP: Relational Language-Image Pre-training for Human-Object Interaction Detection

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Edit

Methods

Add Remove