In Defense of Online Models for Video Instance Segmentation
Junfeng Wu, Qihao Liu, Yi Jiang, Song Bai, Alan Yuille, Xiang Bai
Introduction
Video instance segmentation aims at detecting, segmenting, and tracking object instances simultaneously in a given video. It attracted considerable attention after first defined in 2019 due to the huge challenge and the wide applications in video understanding, video editing, autonomous driving, augmented reality, etc. Current VIS methods can be categorized as online or offline methods. Online methods take as input a video frame by frame, detecting and segmenting objects per frame while tracking instances and optimizing results across frames. Offline methods , in contrast, take the whole video as input and generate the instance sequence of the entire video with a single step.
Despite high performance, the requirement of full videos limits the application scenarios of offline models, especially for scenarios involving long video sequences (videos exceed 50 frames on GPU of 32G RAM ) and ongoing videos (\eg, videos in autonomous driving and augmented reality). However, online models are usually inferior to the contemporaneous offline models by over 10 AP, which is a huge drawback. Previously, little work tries to explain the performance gap between these two paradigms, or gives insight into the high performance of the offline paradigm. A common attempt of the latter is made from the inherent advantages of offline models being able to skip the error-accumulating tracking steps and utilize the richer information provided by multiple frames to improve segmentation . However, does that really explain the high performance of current offline methods? What’s the main problem causing the poor performance of online models? Can online models achieve performance comparable to, or even better than, SOTA offline ones?
To deeply understand the performance of both online and offline models, we analyze in detail three SOTA methods (offline: IFC and SeqFormer , online: CrossVIS ) on two datasets (YouTube-VIS and OVIS ) that have different difficulty levels. ‘Simple video’ refers to the video in YouTube-VIS. These videos are much shorter and only contain very slight occlusions, simple motions, and smooth changes in illumination and object shapes. ‘Complex video’ refers to the video in OVIS. Please see Sec. 4.1 for details. The results (Fig. 1) of our oracle experiments give us a deep understanding of current SOTA methods:
From the perspective of instance segmentation, per-clip segmentation doesn’t outperform per-frame segmentation a lot in mask quality, and mask quality is also not the reason for the poor performance of online methods: CrossVIS even outperforms its contemporaneous work (\ieIFC) in frame oracle experiments on both datasets (results in Table 1). What’s more, per-clip segmentation of current SOTA methods is not effective and robust. Multiple frames do provide more information and improve the mask quality by 3.7 AP for IFC on YouTube-VIS (Fig. 1). But it only works for some cases: per-clip segmentation doesn’t improve the performance of SeqFormer a lot. In addition, when testing on more challenging datasets like OVIS, segmentation on multiple frames even degrades the performance by 1.8 and 2.2 AP on IFC and SeqFormer respectively when clip size becomes longer. Although in theory, per-clip segmentation has its inherent advantage to using multiple frames, it still requires further exploration, especially in how to utilize information in multiple frames and how to handle complex motion patterns, occlusion, and object deformation. Currently, we don’t see an obvious gap between the mask qualities of per-clip and per-frame segmentation.
From the perspective of association, a huge advantage of the offline methods is their ability to avoid the use of hand-designed association modules. It works well on simple cases of the YouTube-VIS dataset. We demonstrate that it is the main reason causing the performance gap between the current online and offline paradigms. However, this black-box association process done within offline models also gets worse rapidly when the video becomes complex (degrades the performance of IFC by 12.3 AP and SeqFormer by 20.9 AP on OVIS). In addition, when handling longer videos, \egvideos in the real world or from OVIS dataset, offline methods require splitting the input video into clips to avoid exceeding computational limits, and thus hand-designed clip matching is still inevitable, which will further decrease the overall performance. To sum up, matching/association is the main reasoning for the performance gap, and it is still inevitable and of great importance for offline models.
To improve the matching performance and thus bridge the performance gap, we propose a framework In Defense of OnLine models for video instance segmentation, termed IDOL. The key idea is to ensure, in the embedding space, the similarity of the same instance across frames and the difference of different instances in all frames, even for instances that belong to the same category and have similar appearances. It provides more discriminative instance features with better temporal consistency, which guarantees more accurate association results. However, previous method selects positive and negative samples by a hand-craft setting, introducing false positives in occlusions and crowded scenes, thus impairing contrastive learning. To address it, we formulate the problem of sample selection as an Optimal Transport problem in Optimization Theory, which reduces false positives and further improves the quality of the embedding. During inference, by using one-to-many temporally weighted softmax, we utilize the learned prior of the embedding to re-identify missing instances caused by occlusions and to enforce the consistency and integrality of associations.
Our thorough analysis gives us a deep understanding of current online and offline VIS methods. Based on our observation, we bridge the performance gap from the perspective of feature embeddings and propose IDOL. We conduct extensive experiments on YouTube-VIS 2019, YouTube-VIS 2021, and OVIS datasets. Despite its simplicity, our method sets a new state-of-the-art achieving 64.3 AP, 56.1 AP, and 42.6 AP on the validation set of these three benchmarks, respectively. More importantly, compared with previous online methods, we achieve a consistent improvement ranging from 13.2 AP to 14.7 AP on these datasets. We even surpass the previous SOTA offline method by up to 2.1 AP. We believe the simplicity and effectiveness of our method shall benefit further research. In addition, our thorough analysis provides insights for current methods and suggests useful directions for future work in both online and offline VIS.
Related Work
Online Video Instance Segmentation. Most online VIS methods are built upon image-level instance segmentation with an additional tracking head to associate instances across the video. The baseline method MaskTrack R-CNN is built upon Mask R-CNN and proposes to leverage multi cues such as appearance similarity, semantic consistency, spatial correlation, and detection confidence to determine the instance labels. Most online methods follow this pipeline. CrossVIS proposes a new learning scheme that uses the instance feature in the current frame to segment the same instance in another frame. Multi-Object Tracking and Segmentation (MOTS) aims to simultaneously segment and track all object instances of a given video sequence in real-time, which is similar to online VIS. MOTS methods are usually built upon multiple object trackers . Track R-CNN firstly extends the popular task of multi-object tracking to MOTS based on Mask R-CNN . PointTrack proposes a new online tracking-by-points paradigm with learning discriminative instance embeddings. Trackformer adopts the transformer architecture for MOTS and introduces track query embeddings that follow objects through a video sequence. Online models have a wider range of application scenarios, however, they are usually inferior to offline art by over 10 AP. We find that the current SOTA online models fail to achieve accurate associations, causing the performance gap. We aim to tackle this problem in this work.
Offline Video Instance Segmentation. Offline methods for VIS take the whole video as input and predict the instance sequence of the entire video (or video clip) with a single step. MaskProp and Propose-Reduce perform mask propagation in a video clip to improve mask and association. However, the propagation process is time-consuming, which limits its application. Recently, VisTR adopts the transformer to VIS and models the instance queries for the whole video. However, it learns an embedding for each instance of each frame, which makes it hard to apply to longer videos and more complex scenes. IFC proposes inter-frame communication transformers and significantly reduces computation and memory usage. SeqFormer dynamically allocates spatial attention on each frame and learns a video-level instance embedding, which greatly improves the performance. We deeply analyze the current SOTA offline models, IFC and SeqFormer, find that their improvement mainly comes from the black-box association between frames, but this advantage is gradually lost in complex scenarios. In contrast, our online method can be applied to both ongoing and long videos and complex scenarios, with more stable association quality and higher performance.
Contrastive learning has made significant progress in representation learning . MOCO and SimCLR use contrastive learning for image-level self-supervised training and learn strong feature representations for downstream tasks. Some methods extend the contrastive learning into multiple positive samples format to obtain better feature representations. We absorb ideas from contrastive learning and propose to learn contrastive embeddings between frames for each instance.
Method
Given a video clip that consists of multiple image frames, online VIS models utilize additional association head upon on instance segmentation models . We have already discussed that achieving more stable and discriminative instance embeddings between frames is the key to improve the performance of online models. To achieve this, we propose a contrastive learning framework to extract more discriminative features for instance association. We first introduce the instance segmentation pipeline in Sec. 3.1. Then the details of our contrastive learning framework and the cross-frame instance association strategy are introduced in Sec. 3.2 and Sec. 3.3 respectively.
For fair comparisons with the state-of-the-art offline method , we take DeformableDETR with dynamic mask head as our instance segmentation pipeline in this paper. Our method can be coupled with other instance segmentation methods with minor modifications.
Given an input frame of a video, a CNN backbone extracts multi-scale feature maps. The Deformable DETR module takes the feature maps with additional fixed positional encodings and learnable object queries as input. The object queries are first transformed into output embeddings by the transformer decoder. After that, they are decoded into box coordinates and class labels by 3-layer feed-forward network (FFN) separately. For per-frame mask generation, we employ an FPN-like mask branch to make the use of multi-scale feature maps from transformer encoder and generate feature map that are 1/8 resolution of the input frame. Another FFN encode outputs embeddings into parameters of mask head, which performs three-layer convolution on the given feature map :
Then we calculate pair-wise matching cost which takes into account both the class prediction and the similarity of predicted and ground truth boxes. For each ground truth, we assign multi predictions to it by selecting the top k predictions with the least cost by an optimal transport method . Finally, the whole model is optimized with a multi-task loss function
where loss weights and are set to 2.0 and 1.0 by default. For , we use a combination of loss and the generalized IoU loss . The is defined as a combination of the Dice loss and Focal loss . is the contrastive loss described in the next section.
2 Contrastive Learning between Frames.
More discriminative feature embeddings can help distinguish instances on different frames, thereby improving the quality of cross-frame association. To this end, we introduce contrastive learning between frames to make the embedding of the same object instance closer in the embedding space, and the embedding of different object instances farther away. Object queries are used to query the features of instances from each frame in our instance segmentation pipeline. Therefore, the output embeddings can be regarded as features of different instances. We employ an extra light-weighted FFN as a contrastive head to decode the contrastive embeddings from the instance features.
Given a key frame for instance segmentation training, we select a reference frame from the temporal neighborhood. The instances appearing on the key frame may have different positions and appearances on the reference frame, but their contrastive embeddings should be as close as possible in embedding space. For each instance in the key frame, we send the output embedding with the lowest cost to the contrastive head and get the contrastive embedding v. Different from previous method , which selects positive and negative samples by a hand-craft setting, if the same instance appears on the reference frame, we take top predictions with the least cost as positive and top predictions with the highest cost as negatives. and are calculated dynamically by the optimal transport method . Please refer to the supplementary for more details. The contrastive loss function for a positive pair of examples is defined as follows:
where and are positive and negative feature embeddings from the reference frame, respectively. We extend Eq. 3 to multiple positive scenarios:
3 Instance Association.
Previous online methods take semantic consistency, spatial correlation, and detection confidence as cures. They are then leveraged to determine the instance labels. Other clip-based nearly online methods match instances using the predicted masks of overlapping frames by masking soft IoU metric between clips. However, the online models perform instance segmentation on each frame independently, and therefore, the prediction quality on each frame is unstable. What makes it worse is the complex motion patterns, severe occlusions, false positives, duplicate predictions, error accumulation in long videos, and the frequently disappear and reappear objects, which makes instance association very challenging. Therefore, a strong instance association method should be robust to these cases. To this end, we propose a temporally weighted softmax score for instance matching and a memory bank-based association strategy to address these problems and improve the association quality of the online model.
Then we compute bi-directional similarity between predicted instance and memory instance by:
where is the existing time of instance in the memory, it serves as the confidence scores of each instance in the memory. By introducing the temporal contrastive embeddings and the confidence scores determined by the duration of existence, the learned prior information is able to reidentify missing instances caused by occlusions, enforcing the consistency and integrality of associations.
Association Strategy. To take full advantage of the learned contrastive embedding, we propose a new association strategy during inference. Given a test video, we initialize an empty memory bank for it and perform instance segmentation on each frame sequentially in an online scheme. For the prediction of each frame, we first perform inter-class duplicate removal by NMS with a threshold of 0.5. Then we compute matching scores between predictions and memory bank by Eq. 6, and search for the best assignment for instance by:
If , we assign the instance on current frame to the memory instance . For the prediction without an assignment but has a high class score, we start a new instance ID in the memory bank.
Experiments
We report our results on YouTube-VIS 2019 , YouTube-VIS 2021 , and OVIS datasets. YouTube-VIS 2019 is the first and largest dataset for video instance segmentation. It contains 2,238 training, 302 validation, and 343 test high-resolution YouTube video clips, with an average video duration of 4.61s. YouTube-VIS 2021 is an extended version of YouTube-VIS 2019. Both datasets have 40 object categories, but the category label set is slightly different. OVIS dataset is a relatively new and challenging dataset. It consists of 607 training videos, 140 validation videos, and 154 test videos. Compared with YouTube-VIS, its videos are much longer and last 12.77s on average, and more importantly, it contains much more videos that record objects with severe occlusion, complex motion patterns, and rapid deformation. All these features make OVIS an ideal dataset to evaluate and analyze different methods. We report standard metrics such as . IoU threshold is used during evaluation.
2 Implementation Details
Model settings. We use ResNet-50 as our backbone unless otherwise specified. For a fair comparison with SOTA offline method, we use the same setting for Deformable DETR and the dynamic mask head following SeqFormer . For the transformer, we use 6 encoders, 6 decoder layers of width 256 with bounding box refinement mechanism, and the number of object queries is set to 300.
Inference. During inference, the input frames are downscaled to 360p for YouTube-VIS 2019 and YouTube-VIS 2021 following previous work, and 720p for OVIS as its videos has a higher resolution. For the hyper-parameters of temporally weighted softmax, we set and by default.
3 Analysis of Current SOTA VIS Models
Since no annotation for the validation set is available, we split the original training set into custom training split and validation split. All models are trained on the training split and evaluated on the validation split. YouTube-VIS 2019 and OVIS are used. We analyze the results as follows:
Performance gaps between online and offline models: First, we compare the mask quality of two recent online and offline methods in Table. 1. They are both published in 2021 thus can be considered as work in the same period. When ground-truth instance ID is provided (the column of ‘frame oracle’), CrossVIS outperforms IFC on both YouTube-VIS and OVIS. However, when the instance ID is not provided (the column of ‘predicted’), the methods are required to match the results, and the performance of CrossVIS drops dramatically by 9.4 AP while the performance of IFC drops by 3.3 AP on YouTube-VIS, leading to the poor performance of CrossVIS. Offline methods enable the model to match predicted masks by itself and avoid using hand-designed association modules. It works well on simple datasets and benefits current offline models, but it still fails on challenging datasets like OVIS.
Analysis of current offline models: In Table. 2, we give detailed analyses for SOTA offline models, hoping to provide insights for future research. Compared with online methods, offline models in theory have two inherent advantages:
First, as we mentioned above, it avoids hand-designed association. However, this step is very sensitive to the occlusion and the complexity of videos, especially when the clip becomes longer. For example, when clip length is equal to 5, the black-box association degrades the performance of IFC and SeqFormer on OVIS by 3.4 AP (25.9 AP \vs22.5 AP) and 6.8 AP (31.8 AP \vs25.0 AP), respectively. When clip length is set to 30, the performance drops by 12.3 AP (22.8 AP \vs10.5 AP) and 20.9 AP (29.8 AP \vs8.9 AP), respectively. What’s more, even a clip length of 30 still doesn’t meet the requirement of real-world application. Clip matching is still inevitable, and it further decreases the overall performance.
Another inherent advantage of offline models is the ability to use multiple frames for instance segmentation, which provides more information to handle occlusion and optimize the results. However, current models still fail to fully utilize this feature. Currently, it only works for simple videos: compared with per-frame segmentation (clip length=1), it improves the mask quality by at most 3.7 AP for IFC (when clip length=30) and 0.3 AP for SeqFormer (when clip length=3). When testing on OVIS, multiple frames segmentation only improves the mask quality by at most 1.3 AP for IFC (clip length=5) and 0.3 AP for SeqFormer (clip length=3), and even degrades the performance by 1.8 AP for IFC and 2.2 AP for SeqFormer when clip size becomes longer (clip length=30). What’s more, the improvement is even less obvious in practice when no ground-truth instance ID is provided, even for simple videos, due to the association problem. It only improves the performance of IFC on YouTube-VIS by 2.0 AP, but degrades the performance in all the other experiments.
To prove the effectiveness of our method, we further analyze IDOL with oracle experiments in Fig.3. Since IDOL is an online model, the results of frame oracles with different clip lengths are the same as per-frame segmentation oracle results. For clip oracles, IDOL is required to do association within the clips by itself. Compared with SeqFormer, the gaps between frame oracles and clip oracles of IDOL are much smaller on OVIS, proving that IDOL performs a more robust association between frames than the offline model on challenging datasets.
4 Main Results
We compare IDOL against current online and offline SOTA methods on YouTube-VIS 2019, YouTube-VIS 2021, and OVIS validation sets. The results are reported in Tables 3, 4, and 5, respectively. We compare the methods with different backbones for a fair comparison. Notably, our method significantly surpasses all previous online methods by at least AP. In addition, we also outperform all previous offline methods under all evaluation metrics when training on the same data. More importantly, our method achieves an overall first place in the YouTube-VIS Challenge 2022, which proves our SOTA performance. In addition, our method only decreases the inference speed of the adopted instance segmentation pipeline by 1.1 FPS on an RTX-2080Ti, which proves our efficiency. In general, our method is simple and very effective compared with all baseline methods. Qualitative results on sample videos of the challenging OVIS dataset are shown in Fig. 4. More qualitative results can be found in the supplementary. We analyze the performance in detail as follows:
YouTube-VIS 2021: It is an extended version of YouTube-VIS 2019. It contains more videos with a larger number of instances and frames. As shown in Table 4, we achieve 43.9 AP with a ResNet-50 backbone, surpassing the previous best online method and offline method by 9.7 AP and 3.4 AP, respectively.
OVIS: As mentioned before, OVIS contains long video sequences with heavy occlusion and complex motion, thus it is extremely difficult for all algorithms and exceeds the capability of offline methods due to the limit of computational resources. STEm-Seg is the only offline method that can be directly evaluated on OVIS since its design enables it to run in a nearly online manner. To compare with the SOTA offline methods (\egIFC and SeqFormer ), we split the video into short clips and apply the clipping matching method provided in IFC on these two methods (SeqFormer doesn’t provide its matching method). The results are provided in Table 5. Note that the previous best method only gains 15.4 AP on the validation set. IDOL with the same ResNet-50 backbone achieves performance and gains 30.2 AP, surpassing the previous method by 14.8 AP. What’s more, when using a stronger backbone (\ie, Swin-L) to extract better features, IDOL achieves the state-of-the-art performance of 42.6 AP, which is a huge improvement over previous best results.
5 Ablation Study
In this section, we conduct extensive ablation experiments to study the importance of the core factors of our method on YouTube-VIS 2019 and OVIS. Previous SOTA online method uses an extra M-class classification head where equals to the number of all instances in the training set, termed as “ID Head” in Table 6. and 7. In the “Contrastive” setting, we use box IoU between predictions and ground truth for positive and negative embeddings selection following . We further evaluate our optimal transport method to dynamically select positive and negative embeddings, termed as “OT”. For ablation study on inference strategy, “multi-cues” setting combines semantic consistency, spatial correlation, detection confidence and appearance similarity together to perform association following . Our association strategy is termed “embedding”.
Contrastive Training. To evaluate the importance of our contrastive embeddings, we apply the same association method on the embeddings predicted by ID Head and contrastive head. As shown in Table 6, contrastive training only improves 1.5 AP with the multi-cues association but improves 8.9 AP when it comes to the embedding-based association. Our explanation is that contrastive training provides more discriminative embeddings for instance association, but other cues in multi-cues weaken the role of embeddings. In addition, on the more challenging OVIS dataset, contrastive embeddings increase AP from 11.0 to 18.4, which brings an improvement of 67.3%. This indicates that embedding-based association is more robust in longer videos and complex scenarios. Furthermore, optimal transport matching (OT) improves the results by 2.3 AP on YTVIS and 2.3 AP on OVIS, which indicates that the choice of positive and negative embeddings plays an important role in learning discriminative embeddings. OT provides a better selection of positive and negative embeddings during training, improving the quality of embeddings. We show the visualization of positive and negative embeddings selected by these two strategies in supplementary.
Association Strategy and Temporally Weighted Softmax. As shown in Table 6, compared with “multi-cues”, our embedding association strategy takes advantage of the discriminative embedding learned by contrastive learning, and improves the AP from 31.8 to 42.5 on YouTube-VIS. When it comes to OVIS in Table 7, our association strategy also improves the AP from 18.4 to 26.7. In addition, when temporally weighted softmax is added, it can be further improved by 1.6 AP on YouTube-VIS and 1.9 AP on OVIS. Utilizing information and priors from multiple previous frames improve robustness of association. Considering the problem of false positives and disappearing-and-reappearing that the online model needs to deal with, we believe this strategy helps maintain temporal consistency. We provide more visualization results, detailed analysis, and additional ablation experiments on OVIS in supplementary.
Conclusions
Online video instance segmentation methods have their inherent advantage in handling long/ongoing videos, but they are inferior to the offline models in performance. In this work, we aim to bridge the performance gap. We first deeply analyze the current online and offline models and find that the gap mainly comes from the error-prone association between frames. Based on this observation, we propose IDOL, which enables models to learn more discriminative and robust instance features for VIS tasks. It significantly outperforms all online and offline methods and achieves new SOTA on three benchmarks. We believe our insights on VIS methods will inspire future work in both online and offline methods.
Acknowledgements This work was supported by NSF 1763705. We thank the reviewers for their efforts and valuable feedback to improve our work.
References
Appendix 0.A Appendix
In this section, we show several qualitative results on the validation sets of YouTube-VIS and OVIS to demonstrate the following advantages of IDOL:
For instances that belong to the same category and have very similar appearances, our contrastive learning enables IDOL to segment and track these instances more accurately. (\egFig. 5)
Our method learns embedding with better temporal consistency, benefiting the tracking in videos with high-speed, large, and/or complex motions. (\egFig. 6)
With the help of more stable and discriminative embeddings, as well as our one-to-many temporally weighted softmax during inference, IDOL is more robust when handling crowded scenes with heavy occlusions and frequent position exchanges. (\egFig. 7)
A.2 Optimal Transport
Given a ground truth bounding box of an instance, the IoU-based method selects positive and negative samples by a hand-craft IoU threshold setting. A predicted box is defined as positive to an instance if they have an IoU higher than 0.7, or negative if they have an IoU lower than 0.3, which introduces false positives in occlusions and crowded scenes. As shown in Fig. 8 (a), in the case of occlusion between two pandas, IoU-based method would take the boxes belonging to the panda in the back as the positive samples of the front one, which causes false positives. To address it, we formulate the problem of sample selection as an Optimal Transport problem in Optimization Theory, which reduces false positives and further improves the quality of the embedding. For each ground truth, we sum the top 10 IoU values to get and the top 100 IoU values to get . Then we take top predictions with the lowest cost as positive and top predictions with the highest cost as negatives. As shown in Fig. 8 (c), the optimal transport provides a better selection of positive embeddings during training, and thus improves the quality of the embedding.
A.3 Temporally Weighted Softmax
In Fig. 9, we show qualitative results of the temporally weighted softmax in our association strategy. As shown in Fig. 9 (a) and (b), the bear with ‘id:1’ in (a) is occluded by another bear in some frames, and without temporally weighted softmax, it is assigned a new id when it reappears. As shown in Fig. 9 (c), the people with ‘id:3’ and elephant with ‘id:0’ disappear in the corner of the video, but they swap ids when they reappear after several frames, and this leads to classification errors. However, in Fig. 9 (d), temporally weighted softmax helps maintain temporal consistency of id for the sampe people and elephant.