Mask2Former for Video Instance Segmentation

Bowen Cheng, Anwesa Choudhuri, Ishan Misra, Alexander Kirillov, Rohit Girdhar, Alexander G. Schwing

Introduction

Video instance segmentation yang2019video differs from image segmentation in its goal to simultaneously segment and track objects in videos. Initial methods yang2019video; voigtlaender2019mots add “track embeddings” to per-pixel image segmentation architectures to track objects across time. Recent transformer-based methods vistr; ifc; wu2021seqformer use “object queries” detr to connect objects across frames. However, all these methods are specifically designed to only process video data, disconnecting image and video segmentation research.

Universal image segmentation methods cheng2021maskformer; cheng2021mask2former show that mask classification is able to address common image segmentation tasks (i.e. panoptic, instance and semantic) using the same architecture, and achieve state-of-the-art results. A natural question emerges: does this universality extend to videos? In this report, we find Mask2Former cheng2021mask2former achieves state-of-the-art results on video instance segmentation as well without modifying the architecture, the loss or even the training pipeline. To achieve this, we let Mask2Former directly attend to the 3D spatio-temporal features and predict a 3D volume to track each object instance across time (Fig. 1). This simple change results in state-of-the-art performance of 60.4 AP on the challenging YouTubeVIS-2019 data and 52.6 AP on YouTubeVIS-2021.

Background

Video instance segmentation (VIS) yang2019video requires tracking instances across time. In general, there are two ways to extend image segmentation models for the VIS task:

Per-frame methods (a.k.a. online methods) treat a video clip as a sequence of frames. They run image segmentation models on each frame independently and associate predicted instance masks across frames with a post-processing step. MaskTrack R-CNN yang2019video, TrackR-CNN voigtlaender2019mots and VPSNet kim2020video add per-instance track embedding prediction to Mask R-CNN he2017mask. VIP-DeepLab qiao2021vip extends Panoptic-DeepLab cheng2020panoptic by predicting instance center offsets across frames. However, it is not always straightforward to modify an image segmentation model, as architectural changes and additional losses are needed.

Per-clip methods (a.k.a. offline methods) treat a video clip as a 3D spatio-temporal volume and directly predict the 3D mask for each instance. STEm-Seg athar2020stem predicts and clusters spatio-temporal instance embeddings to generate 3D masks. More recently, inspired by the success of DETR detr, VisTR vistr, IFC ifc and SeqFormer wu2021seqformer design Transformer-based architectures to process the 3D volume via cross-attention. However, all these models are designed specifically for video instance segmentation, disconnecting image and video segmentation research. In this report, we show Mask2Former cheng2021mask2former, a state-of-the-art universal image segmentation model, can also achieve state-of-the-art video instance segmentation without modifying the architecture, the loss or even the training pipeline.

Mask2Former for Videos

We treat a video sequence as a 3D spatio-temporal volume of dimension T×H×WT\times H\times W, where TT is the number of frames and H,WH,W are height and width respectively. We make three changes to adapt Mask2Former to video segmentation: 1) we apply the masked attention to the spatio-temporal volume; 2) we add an extra positional encoding for the temporal dimension; and 3) we directly predict a 3D volume of an instance across time.

We apply masked attention to the 3D spatio-temporal features, i.e., we use

Moreover, the 3D attention mask \mathbfcalMl−1\mathbfcal{M}_{l-1} at feature location (t,x,y)(t,x,y) is

Here, Ml−1∈{0,1}N×THlWl\mathbf{M}_{l-1}\in\{0,1\}^{N\times TH_{l}W_{l}} is the binarized output (thresholded at 0.50.5) of the resized 3D mask prediction of the previous (l−1l-1)-th Transformer decoder layer.

2 Temporal positional encoding

To obtain compatibility with image segmentation models we decouple the temporal positional encoding from the spatial encoding, i.e., we use the positional encoding

3 Joint spatio-temporal mask prediction

Similarly to the masked attention, we obtain the 3D mask of the nthn^{\text{th}} query via a simple dot product, i.e.,

Note, computation of the classification does not change.

Experiments

We evaluation Mask2Former on YouTubeVIS-2019 and YouTubeVIS-2021 yang2019video. We do not modify the architecture, the loss or even the training pipeline.

We use Detectron2 wu2019detectron2 and follow the IFC ifc settings for video instance segmentation. More specifically, we use the AdamW loshchilov2018decoupled optimizer and the step learning rate schedule. We use an initial learning rate of 0.00010.0001 and a weight decay of 0.050.05 for all backbones. A learning rate multiplier of 0.10.1 is applied to the backbone and we decay the learning rate at 2/32/3 fractions of the total number of training steps by a factor of 1010. We train our models for 66k iterations with a batch size of 1616 for YouTubeVIS-2019 and 88k iterations for YouTubeVIS-2021. During training, each video clip is composed of T=2T=2 frames, which makes the training much more efficient (2 hours per training with 8 V100 GPUs), and the shorter spatial side is resized to either 360 or 480. All models are initialized with COCO lin2014coco instance segmentation models from cheng2021mask2former. Unlike wu2021seqformer, we only use YouTubeVIS training data and do not use COCO images for data augmentation.

2 Inference

Our model is able to handle video sequences of various length. During inference, we provide the whole video sequence as input to the model and obtain 3D mask predictions without any post-processing. We keep the top 10 predictions for each video sequence. If not stated otherwise, we resize the shorter side to 360 pixels during inference for ResNet he2016deep backbones and to 480 pixels for Swin Transformer liu2021swin backbones.

3 Results

We compare Mask2Former with state-of-the-art models on the YouTubeVIS-2019 dataset in Table 1 and the YouTubeVIS-2021 dataset in Table 2. Using the exact same training parameters, Mask2Former outperforms IFC ifc by more than 6 AP. Mask2Former also outperforms concurrent SeqFormer wu2021seqformer without using extra COCO images for data augmentation.

B. Cheng, A. Choudhuri and A. Schwing were supported in part by NSF grants #1718221, 2008387, 2045586, 2106825, MRI #1725729, NIFA award 2020-67021-32799 and Cisco Systems Inc. (Gift Award CG 1377144 - thanks for access to Arcetri).

References