Temporally Efficient Vision Transformer for Video Instance Segmentation
Shusheng Yang, Xinggang Wang, Yu Li, Yuxin Fang, Jiemin Fang, Wenyu Liu, Xun Zhao, Ying Shan
Introduction
Video Instance Segmentation (VIS) vis is a representative and challenging video understanding task that requires detecting, segmenting and tracking video instances across frames simultaneously. Similar to other instance-level video recognition tasks, making full use of temporal context information is critical for building high-performance VIS systems. Vision transformer (ViT) vit, which is based on self-attention transformer, has shown strong long-range context modeling ability and obtained great successes on image classification vit; deit; halonet; swintransformer; pvt; msgtransformer; shuffletransformer, object detection detr; defdetr; conddetr; sparsercnn, semantic segmentation segformer; segmenter; maskformer, instance segmentation queryinst; solq; knet, and video recognition timesformer; videoswin; tokshift; vivit; vidtr; mvit; vtn.
Recently, how to design ViTs for instance-level video understanding, especially VIS, becomes an emerging problem. Different from the detection transformers detr; defdetr; conddetr; sparsercnn; yolos, semantic segmentation transformers segformer; segmenter; maskformer, and instance segmentation transformers queryinst; solq; knet, which focus on D contextual information modeling, VIS transformers additionally require to perform temporal context modeling. To this end, VisTR vistr firstly proposes a transformer encoder to fuse patch features from a sequence of frames using a CNN backbone and leverages a query-based decoder to predict video instances, IFC ifc introduces memory tokens to store frame-level features and performs cross-frame feature interaction by computing self-attention among memory tokens, and then decodes instance-level results using a conditional mask head.
Experiments are conducted on three large-scale VIS datasets, i.e., YouTube-VIS- vis, YouTube-VIS- vis2021, and OVIS ovis. New state-of-the-art (SoTA) performance has been obtained, e.g., TeViT obtains AP with FPS on YouTube-VIS-. Our main contributions are summarized as follows.
TeViT is the first video instance segmentation transformer that can efficiently capture temporal contextual information at both frame level and instance level.
Benefiting from the flexibility of self-attention, the proposed temporal modeling modules, i.e., messenger token shift and spatiotemporal query interaction, both are friendly to the image-level pre-trained models, cost marginal extra computation overhead and parameters.
TeViT is a nearly convolution-free framework and obtains SoTA VIS results. In TeViT, the concepts of “early temporal feature fusion” and “a video instance as a query” shield lights on how to build effective video transformers for instance-level recognition tasks.
Related Work
Video instance segmentation. How to achieve efficiently temporal modeling is always the focus of video tasks, such as video object segmentation (VOS) uvos; uvos; stm, multi-object tracking and segmentation (MOTS) mots and VIS vis. Though VOS and MOTS are very related with VIS, MOTS mainly focuses on the urban scene understanding and VOS aims at tracking specific object by a given mask. Representative VIS works are reviewed as follows. MaskTrack R-CNN vis extends Faster R-CNN fasterrcnn and Mask R-CNN maskrcnn to VIS with a tracking branch and external memory that saves instance features across multiple frames. MaskProp maskprop builds on the Hybrid Task Cascade Network htc and propagates instance region features to adjacent frames to perform temporal modeling. STEm-Seg stemseg treats video clip as D spatiotemporal volume and captures temporal information by D convolutional backbone network. CompFeat compfeat refines temporal features at both frame-level and instance-level. CrossVIS crossvis introduces a crossover learning scheme upon fcos; condinst to make use of contextual information across video frames. SeqMask R-CNN seqmaskrcnn establishes temporal relation across frames by adding an extra sequence propagation head upon Mask R-CNN. Both VisRGNN visrgnn and VisSTG visstg model temporal information in VIS by a graph neural network. VisTR vistr proposes the first fully end-to-end VIS method upon DETR detr, temporal contexts are fused by the multi-head attention mechanism in transformer encoder layers. IFC ifc presents inter-frame communication to exchange frame-level information. In this paper, we present a temporally efficient framework to model temporal contexts at both frame-level and instance-level.
Vision Transformer. Transformer transformer is firstly proposed to model long-range sequence data in natural language process (NLP). ViT vit firstly adopts transformer to image domain. After that various high-performance vision transformers swintransformer; halonet; deit; pvt; pvtv2; msgtransformer; shuffletransformer have been proposed as backbones for image understanding. Beyond serving as backbone networks, transformer has motivated lots of novel object detection detr; defdetr; sparsercnn; conddetr; yolos, instance segmentation queryinst; solq; knet, and semantic segmentation segformer; segmenter; maskformer frameworks. Recently, VisTR vistr, IFC ifc, QueryTrack querytrack, and TCIS tcis bring transformer to video instance segmentation and achieve excellent performance. In this paper, we investigate how to efficiently model temporal context across video frames and propose TeViT. TeViT is a nearly convolution-free transformer while VisTR and IFC both use ResNet resnet backbone.
Method
The overall architecture of our VIS method TeViT is shown in Fig. 1, which contains a transformer-based backbone network and a query-driven head network. Given a sequence of video frames, the transformer backbone performs feature extraction and generates multi-scale pyramid features. The query-driven head network takes randomly initialized instance queries with backbone feature maps to predict video instances. Our whole network is end-to-end for both training and inference.
2 Messenger Shift Transformer Backbone
In previous VIS methods, the backbone networks only perform feature extraction in per-frame fashion vistr; ifc and neglect the rich contextual information inherent in video frames. In contrast, inspired by MSG-Transformer msgtransformer, we propose messenger shift transformer (MsgShifT) which performs highly efficient temporal context modeling in a bottom-up manner, as shown in Fig. 1 (left). Without loss of generality, we build MsgShifT based on the pyramid vision transformer (PVT) pvt; pvtv2.
where indicates the copycat of messenger tokens . The concatenated joint tokens are taken as inputs for our MsgShifT.
Next, a messenger shift manipulation performs temporal information exchange across video frames.
In short, the messenger shift mechanism takes temporal messenger tokens as inputs and builds temporal context modeling by shifting messenger tokens along the temporal axis. Fig. 2 gives a detailed illustration. First, messenger tokens are divided into groups and shifted along the temporal axis with different time steps ( or ) and direction (forward or backward). With various time steps and directions, messenger tokens are able to achieve temporal context exchange with both past and future frames. Moreover, for every two messenger shift operations, we apply an inverse operation to the second one, which implies the messenger tokens will be shifted back to their original corresponding frames after two contiguous messenger shift manipulations. This design aims to maintain a stable temporal receptive field as the network goes deeper.
After the above process, the messenger tokens and patch tokens go through one of four stages, and the output tokens are reshaped to feature maps which is smaller than the original image. In the same way, using the output messenger tokens and patch tokens of prior stage as inputs, we obtain the following pyramid feature maps , and , whose strides are , and pixels with respect to the input image. The pyramid feature maps will be used to predict video instances in the head network.
MsgShifT performs early temporal fusion in the backbone network, while the previous transformer-based VIS approaches vistr; querytrack; tcis; ifc only perform temporal feature fusion using transformer encoders after image-level feature extraction. It is almost parameter-free, friendly to image-level pre-training models, and brings negligible computation costs. The messenger tokens are randomly initialized and the shift manipulation has no parameter, so this module is insensitive to the pre-training process, which will be further discussed in the experiments in Tab. 8.
3 Spatiotemporal Query Interaction Head
MsgShifT achieves frame-level spatiotemporal context modeling. Meanwhile, in the VIS head network, our method still emphasizes temporally efficient spatiotemporal context modeling, but at the instance level. To this end, we propose a spatiotemporal query interaction (STQI) head network based on the recent SoTA query-based image-level instance segmentation method, i.e., QueryInst queryinst.
“:” denotes ranging from to . Enhanced instance queries are fed into a dynamic convolution module and perform interactions with instance region features. Its output serves as the input queries of the next head. Finally, task specific heads (i.e., classification head, box head and mask head) predict a sequence of video instances:
where , and denotes predicted confidence scores, bounding boxes and instance foreground masks, respectively.
4 Matching and Loss Function
The loss function is motivated by detr. We first compute the one-to-one assignment between predicted instances and ground-truth annotations. The ground-truth annotations are denoted as follows:
in which indicates the number of ground-truth video instances, , and indicates the category, bounding box and mask respectively. We then perform sequence-level bipartite matching between predictions and annotations by Hungarian algorithm hungarian. The cost matrix with size of between each predicted video instance and each annotation is defined as follows.
5 Online and Offline Inference
Our method is flexible for both offline and online inference. Under the offline scenario, our TeViT takes the whole video clips as inputs and then outputs all possible video instances with a single run. No post-tracking process is needed. When it comes to the near online stemseg scenario, an entire video is split into several overlapping segments. TeViT takes clips in time order and generates predictions. A rule-based post-tracking procedure is applied to linking instances across different video clips. For instances from two overlapping video clips, we first compute the similarity score between each instance, and then a Hungarian matcher gives the assignment according to the similarity matrix. The similarity score is defined as a combination of box IoU and mask IoU.
Experiments
We evaluate TeViT on three challenging video instance segmentation benchmarks, i.e., YouTube-VIS- vis, YouTube-VIS- vis2021, and OVIS ovis. YouTube-VIS- is the first dataset that focuses on the VIS problem. It contains common object categories, unique video instances and about high-quality instance-level annotations. YouTube-VIS- dataset is the new version of YouTube-VIS- with more video frames and more annotations. OVIS dataset aims to explore the VIS problem under high-occlusion scenarios. It consists of high-quality instance masks and instances per video from semantic categories. Following previous works, we report the performance on the validation set for all three datasets. We follow the standard VIS evaluation metrics defined in vis.
2 Implementation Details
TeViT is built upon the toolbox mmdetection. Unless otherwise noted, hyper-parameters follow the settings of QueryInst queryinst. We use video instance queries as ifc; queryinst. Due to the temporal efficient designs in TeViT, we do not need to create pseudo video data, e.g. stemseg; ifc, to train the temporal modeling parameters, instead, we first train a transformer-based QueryInst for image-level instance segmentation on the COCO dataset mscoco and then initialize TeViT with the COCO pre-trained QueryInst weights. Besides, we provide a MindSpore mindspore implementation of TeViT.
When training on the VIS datasets, we use the AdamW adam optimizer with an initial learning rate of , and a weight decay of . Especially, the backbone learning rate is slightly lower with a multiplier set to . We also apply gradient clipping with a maximal gradient norm of . TeViT is trained with a batch size of and a clip length of . The total training process contains epochs, and the learning rate is decreased by at the -th and -th epoch respectively. For example, our TeViT can be trained in about hours with Tesla V GPUs on YouTube-VIS-, which is much faster than previous transformer-based method (i.e., VisTR vistr). The number of instance queries is set to for all experiments. Following vis, all input frames are resized to in single-scale experiments. Settings of multi-scale training simply follow sipmask. For inference, all frames are resized to regardless of the training setups. During inference, we use for most results and report the near online results in ablation study. For main results, we evaluate our framework on YouTube-VIS-, YouTube-VIS-, and OVIS datasets, with PVT-B pvtv2 based MsgShifT as backbone.
3 Main Results
Main results on YouTube-VIS-2019 dataset. We compare our TeViT to state-of-the-art methods on YouTube-VIS- dataset in Tab. 1. The longest video in YouTube-VIS- dataset only contains frames, so that our TeViT executes fully offline inference on this dataset. Without bells and whistles, our TeViT achieves AP when using a single-scale training strategy and outperforms the previous state-of-the-art methods by a large margin. Multi-scale training strategy further boosts the performance to AP. Meanwhile, our method also achieves competitive inference speed. With about AP higher, our method is still faster than VisTR.
Main results on YouTube-VIS-2021 dataset. Tab. 3 shows the final results of several VIS methods and ours on YouTube-VIS- dataset. Due to the video length in YouTube-VIS- is longer than our inference clip length (), TeViT performs near online tracking described in Sec. 3.5 on this dataset. TeViT obtains AP, outperforming the previous state-of-the-art method by AP.
Main results on OVIS dataset. The results on the OVIS dataset are shown in Tab. 3. Our method also performs near online inference on OVIS dataset. TeViT achieves a relatively higher performance of AP on the split, surpassing previous state-of-the-art methods. Compared to CMaskTrack R-CNN ovis which presents an elaborate-designed feature calibration plug-in to alleviate occlusion, our TeViT still gains AP improvement, which shows that our temporal context modeling designs are helpful to segment occluded instances.
4 Ablation Study
Effect of frame-level & instance-level temporal context modeling. We investigate the effects of messenger shift mechanism and spatiotemporal query interaction individually and simultaneously in Tab. 5. Using messenger shift mechanism and spatiotemporal query interaction individually brings and AP improvements respectively. The results show that both frame-level and instance-level temporal context modeling can obviously improve VIS performance. In addition, the instance-level one brings more significant performance gain. The two designs together brings () AP improvements over a high-performance baseline. Besides the remarkable performance improvements, our designs only bring computation overhead on our baseline ( GFLOPs vs. GFLOPs), which demonstrates our design is very efficient.
Number of messenger tokens. In Tab. 8, we test our method with number of messenger tokens increases from to . Compared to less messenger tokens (), more messenger tokens () achieves better results. Unless specified, our experiments are conducted with messenger tokens.
Training and inference clip length. We also investigate the effects of clip length in both the training and testing phase. From Tab. 10, we find that: (1) Our method shows great tolerance to short length of training clip. Only trained with or frames, our method can effectively learn temporal context and obtains comparable results to previous methods. (2) The performance improvements by increasing the length of the training clip gradually gets saturated. Increasing training clip length from to and to brings AP and AP gains respectively while increasing training clip length from to only obtains slight AP profit. Besides, a longer training clip requires more training computations and memory budgets. To this end, we set the training clip length of our method to as a compromise between performance and training costs.
Tab. 10 gives the results of TeViT under different inference settings. “T” indicates the input clip length during the inference phase, and “S” indicates strides. It shows that our TeViT obtains promising performance under various inference setups. Even with and , TeViT still achieves AP, which indicates TeViT can serve as a strong baseline for both offline and online video understanding scenarios.
Revisiting messenger tokens. Inspired by MSG-Transformer msgtransformer, we re-initialize messenger tokens in inference phase and obvious the influence on performance in Tab. 12. As the results show, when we re-initialize messenger tokens to zero, the performance merely drops AP (compare Row to Row ). Randomly initialize the messenger token in inference phase leads to a similar performance decrease (Row ). We think this phenomenon implies that the messenger tokens contain only a few or not specific information in themselves. On the contrary, they play the role of summarizing frame-level contexts, and exchanging them across adjacent frames.
Conclusion
In this paper, we provide lightweight and effective solutions to fully exploit temporal context for VIS. Based on existing ViTs and query-based image-level instance segmentation methods, we proposes the TeViT VIS method that contains the messenger shift and spatiotemporal query interaction mechanisms. TeViT performs both frame-level and instance-level temporal feature interactions while only bringing a few parameters and marginal extra computational costs. Experiments on YouTube-VIS-, YouTube-VIS-, and OVIS show that TeViT can obtain remarkably better results than previous SoTA methods, e.g., IFC, VisTR, MaskProp, and STEm-Seg. We believe the proposed temporal context modeling mechanisms have great potential to be extended to other video understanding tasks.
Limitations. Although the extensive experiments have demonstrated the capacity and efficiency of our TeViT on temporal context modeling, it still suffers effects from occlusion, motion deformation and long time-span videos (i.e., results of TeViT in Tab. 3 and Tab. 3 are far from satisfying). We leave these promising directions as future work.
Broader impact. Although our research does not make direct negative impacts in society, it may be misused by illegal video applications, which could be a potential invasion to human privacy.
Acknowledgement. This work was in part supported by NSFC (No. 61876212 and No. 61733007) and CAAI-Huawei MindSpore Open Fund.