TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK
Video Instance Segmentation	OVIS validation	ROVIS (Swin-L)	mask AP	42.6	# 13
Video Instance Segmentation	OVIS validation	ROVIS (Swin-L)	AP50	64.7	# 16
Video Instance Segmentation	OVIS validation	ROVIS (Swin-L)	AP75	42.6	# 15
Video Instance Segmentation	OVIS validation	ROVIS (Swin-L)	AR1	18.4	# 9
Video Instance Segmentation	OVIS validation	ROVIS (Swin-L)	AR10	49.1	# 9

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/robust-online-video-instance-segmentation/video-instance-segmentation-on-ovis-1)](https://paperswithcode.com/sota/video-instance-segmentation-on-ovis-1?p=robust-online-video-instance-segmentation)`

Robust Online Video Instance Segmentation with Track Queries

16 Nov 2022 · Zitong Zhan, Daniel McKee, Svetlana Lazebnik ·

Recently, transformer-based methods have achieved impressive results on Video Instance Segmentation (VIS). However, most of these top-performing methods run in an offline manner by processing the entire video clip at once to predict instance mask volumes. This makes them incapable of handling the long videos that appear in challenging new video instance segmentation datasets like UVO and OVIS. We propose a fully online transformer-based video instance segmentation model that performs comparably to top offline methods on the YouTube-VIS 2019 benchmark and considerably outperforms them on UVO and OVIS. This method, called Robust Online Video Segmentation (ROVIS), augments the Mask2Former image instance segmentation model with track queries, a lightweight mechanism for carrying track information from frame to frame, originally introduced by the TrackFormer method for multi-object tracking. We show that, when combined with a strong enough image segmentation architecture, track queries can exhibit impressive accuracy while not being constrained to short videos.

PDF Abstract

Code

Add Remove Mark official

zitongzhan/mmtracking official

Tasks

Add Remove

Image Segmentation

Instance Segmentation

Multi-Object Tracking

Object Tracking

Segmentation

Semantic Segmentation

Video Instance Segmentation

Video Segmentation

Video Semantic Segmentation

Datasets

YouTube-VIS 2019

OVIS

UVO

Results from the Paper

Add Remove

Ranked #13 on Video Instance Segmentation on OVIS validation

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Benchmark
Video Instance Segmentation	OVIS validation	ROVIS (Swin-L)	mask AP	42.6	# 13	Compare
			AP50	64.7	# 16	Compare
			AP75	42.6	# 15	Compare
			AR1	18.4	# 9	Compare
			AR10	49.1	# 9	Compare

Methods

Add Remove

CLIP

Edit Social Preview

Robust Online Video Instance Segmentation with Track Queries

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit Add Remove

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Add Remove

Methods

Add Remove