DVIS++: Improved Decoupled Framework for Universal Video Segmentation
Tao Zhang, Xingye Tian, Yikang Zhou, Shunping Ji, Xuebo Wang, Xin Tao, Yuan Zhang, Pengfei Wan, Zhongyuan Wang, Yu Wu
Introduction
Video segmentation is a fundamental task in computer vision, playing a significant role in video understanding, video editing, and autonomous driving , among other applications . Most studies, such as , , and , focus on designing specialized architectures for specific subdomains of video segmentation. A few researches, and , introduce unified architectures but significantly underperform the specialized architectures. Therefore, in this paper, we concentrate on universal video segmentation and design an efficient unified architecture that can achieve state-of-the-art performance across various segmentation tasks, including video instance segmentation (VIS) , video semantic segmentation (VSS), and video panoptic segmentation (VPS) . Universal video segmentation requires the simultaneous tracking, segmentation, and identification of all instance-level “thing” objects (e.g., person, dog) and semantic-level “stuff” objects (e.g., sky, road) in the video. In this paper, we achieve universal video segmentation by adopting a unified perspective to describe both “thing” objects and “stuff” objects.
In the VIS community, previous methods , , , , , , , , , can be divided into two technical pipelines: offline pipeline (i.e., processing the entire video at once) and online pipeline (i.e., processing the video frame by frame). The offline pipeline focuses on studying how to effectively utilize the spatio-temporal information in videos, while the online pipeline focuses on studying how to associate objects frame-by-frame more stably and accurately.
Offline methods such as VisTR , IFC , SeqFormer , Mask2Former-VIS , and VITA have designed various mechanisms to extract stronger spatio-temporal features of objects. These methods have achieved satisfactory results on simple, short videos , but they face significant challenges when dealing with complex, long videos . This is because the complexity of object motion trajectories and probability of occlusion dramatically increase when dealing with complex, long videos, making it extremely challenging to directly model the spatio-temporal representation of objects from video features using prior queries.
On the other hand, online methods like Mask Track R-CNN , MinVIS , IDOL , GenVIS , and CTVIS focus on obtaining more discriminative representations of objects and designing more powerful object association algorithms. These online methods have achieved good results on both simple, short videos and complex, long videos. However, it is evident that these methods have not attempted to model the long-term spatio-temporal representation of objects. Therefore, there is still room for improvement in these online methods.
In this paper, we rethink the problems in modern video segmentation and propose a solution to overcome these challenges by decoupling the video segmentation task. Directly modeling the spatio-temporal representation of objects from the image features of all frames presents unimaginable challenges and often leads to failure when dealing with complex, long videos. However, if given temporally pre-aligned object representations, modeling the spatio-temporal representation of objects becomes much easier. Therefore, we propose dividing the video segmentation task into three sub-tasks: segmentation, tracking, and refinement. The difficulty of these sub-tasks is independent of the video’s length and complexity. Segmentation aims to extract all objects of interest, including “thing” and “stuff” elements, and obtain their representations from a single frame. Tracking seeks to establish the association of the same object across adjacent frames, significantly reducing complexity compared to the direct association across all frames as applied in previous offline methods. Refinement optimizes both segmentation and association results by utilizing the pre-aligned temporal information of the object.
The segmentation subtask has been well addressed by works on image segmentation , , and . This paper focuses on designing an effective tracker and refiner for the more challenging subtasks of tracking and refinement.
We propose a learnable tracker, termed the referring tracker, which models the tracking task as a reconstruction or denoising process. Specifically, by utilizing the representation of a specific object from the previous frame as a reference, the referring tracker generates the representation of the same object based on the object representations outputted by the segmenter in the current frame. The pre-aligned object representations between frames can be obtained through the referring tracker by following the above pipeline.
Additionally, we construct a temporal refiner using simple, naive self-attention and 1D convolution to effectively model the spatio-temporal representations of objects based on the pre-aligned representations.
Finally, we introduce the Decoupled VIdeo Segmentation (DVIS) framework by cascading the segmenter, referring tracker, and temporal refiner. This framework enables convenient and efficient modeling of spatio-temporal representations of objects and ultimately outperforms all contemporary methods , , , , and .
Furthermore, we have proposed DVIS++, which improves our previous conference work DVIS on tracking capabilities. The tracking sub-task is a prerequisite for the refinement sub-task, as good tracking results are fundamental for effectively modeling the spatio-temporal representations of objects. To further enhance the tracking ability of DVIS, we have introduced a denoising training strategy to simulate challenging cases and incorporated contrastive learning to obtain more discriminative object representations.
Specifically, the denoising training strategy, including three noise simulation approaches, mimics the most challenging identification swapping problems in video segmentation. The introduction of the denoising training strategy significantly enhances the tracking capability of the learnable referring tracker. Moreover, effective contrastive learning strategies are newly designed for the referring tracker and the temporal refiner.
We further validate the effectiveness and universality of DVIS++ under various settings. Initially, we investigate whether DVIS++ remains effective when not allowed to finetune the pre-trained backbone. To this end, we introduce a vision foundation model pre-trained with DINOv2 . Secondly, we test our method in an open-vocabulary setting. By integrating CLIP with DVIS++, we create OV-DVIS++, the first open-vocabulary universal video segmentation framework.
We conducted extensive experiments on six datasets, including OVIS , YouTube VIS 2019/2021/2022 , VIPSeg , and VSPW , to validate the effectiveness of the proposed DVIS and DVIS++. As shown in Figure 1, DVIS++ achieves comprehensive improvements compared to DVIS and outperforms previous state-of-the-art (SOTA) methods on all six benchmarks. It is worth noting that we have secured the championship in the PVUW Challenge at CVPR 2023 and the LSVOS Challenge at ICCV 2023 , utilizing only a subset of the strategies presented in this paper. OV-DVIS++ enables open-vocabulary universal video segmentation, achieving a significant improvement in zero-shot performance on VIS datasets compared to previous SOTA methods, as illustrated on the right side of Figure 1. We believe that DVIS, DVIS++ and OV-DVIS++ can serve as solid baselines in the field of video segmentation.
To summarize, our contributions are as follows:
Through rethinking problems in modern video segmentation, we propose a decoupling strategy to better model objects’ spatio-temporal representations. In line with this strategy, we design DVIS for universal video segmentation, which includes a segmenter, a novel referring tracker, and a novel temporal refiner.
To enhance the tracking capability, the foundation of modeling spatio-temporal representations, we incorporate a denoising training strategy and contrastive learning into DVIS, resulting in an improved version, DVIS++. Compared to DVIS, DVIS++ demonstrates enhanced robustness in tracking and segmentation capabilities.
We validate the effectiveness and universality of DVIS++ under various settings, including freezing pre-trained backbone and open-vocabulary settings. We introduce the first open-vocabulary universal video segmentation framework, OV-DVIS++, which achieves SOTA performance in the open-vocabulary setting. Furthermore, when utilizing a frozen pre-trained backbone, DVIS++ also works well and shows remarkable performance.
Related Work
Specialized Video Segmentation. Video segmentation is a fundamental task in the field of computer vision. In the past, traditional methods were employed by , , and to address video object segmentation and video matting. However, in recent years, deep learning has become the mainstream approach in video segmentation and has achieved remarkable success. There are two distinct communities in video segmentation: video semantic segmentation (VSS) and video instance segmentation (VIS). These communities utilize different technique pipelines and focus on different challenges. VSS methods naturally associate the segmentation results of different frames based on semantic categories. Therefore, VSS methods do not need to address the problem of object association. Instead, the VSS methods , , , and focus on improving the temporal consistency of the segmentation results. As a result, these VSS methods cannot be directly applied to universal video segmentation.
VIS faces the most challenging problem of consistently and accurately associating instance-level objects with the same identity in videos, given significant deformations, large-scale movements, and complex occlusions. The first VIS method, Mask Track R-CNN , achieves VIS by incorporating a tracking head onto Mask R-CNN and utilizing multiple cues to associate objects. Subsequently, SipMask , and CrossVIS improve performance by introducing a more powerful segmenter and a crossover learning scheme, respectively. Additionally, MinVIS and IDOL discover that query-based instance representations can be directly used for object association in adjacent frames, resulting in surprising performance. Furthermore, IDOL and CTVIS introduce contrastive learning into VIS to obtain more discriminative instance representations. GenVIS and GRAtt-VIS design learnable trackers to handle object association, exhibiting better performance than naive heuristic algorithms.
In contrast to the aforementioned online VIS methods, offline VIS methods VisTR , IFC , Mask2Former-VIS , SeqFormer , and VITA focus on enhancing the modeling of spatio-temporal representations of objects and have shown impressive performance on short and simple videos, but their effectiveness diminishes when confronted with complex and lengthy videos. The main reason behind this limitation is that these methods strive to link the same objects in all frames and model their spatio-temporal representations simultaneously, which becomes exceedingly difficult in the case of complex, long videos. As a result, it is not surprising that these VIS methods face challenges when dealing with complex, long videos . GenVIS and NOVIS adopt a different approach, decomposing long videos into multiple fixed-length video clips. The short clip typically contains fewer than eight frames. This strategy significantly reduces the temporal association challenge and thus enables the direct modeling of the spatio-temporal representation of the object within each clip. However, it is important to note that GenVIS and NOVIS do not offer significant advantages when limited spatio-temporal information is employed.
Our proposed methods, DVIS and DVIS++, address all the challenges above through the decoupling design. Firstly, we associate objects frame by frame and subsequently obtain spatio-temporal representations based on pre-aligned object representations. Consequently, our method significantly outperforms all current methods on complex, long videos.
Universal Video Segmentation. Thanks to the success of universal image segmentation methods such as , , , , and , universal video segmentation methods like Video K-Net , TubeFormer , Tube-Link , TarVIS , and DVIS have achieved comparable or even higher performance compared to specialized methods. Among these methods, Video K-Net achieves video segmentation by tracking and associating object kernels, while TubeFormer, Tube-Link, and TarVIS perform video segmentation by directly segmenting objects within clips and associating segmentation results between clips through overlapping frames or similarity clues. However, all the aforementioned methods utilize heuristic algorithms for object association and only leverage temporal information within short clips. In contrast, our proposed DVIS and DVIS++ employ a trainable referring tracker for more elegant and efficient object association. Additionally, they effectively model the spatio-temporal representation of objects throughout the entire video using a temporal refiner.
2 Image Segmentation
Universal Image Segmentation. Thanks to the success of the transformer in image detection , and the emergence of more generic representation forms , compared to per-pixel classification maps, many universal image segmentation methods , , , and have been able to unify semantic segmentation and instance segmentation by adopting the same representation approach. , , and achieve more capabilities, such as open-vocabulary segmentation and multi-granularity segmentation, while implementing universal image segmentation. In this paper, DVIS and DVIS++ adopt the classical universal segmentation method Mask2Former as the segmenter.
Open-Vocabulary Image Segmentation. Open-vocabulary segmentation aims to segment objects of any category, including those not present in the training set. Recent works , , , , , , have utilized large-scale visual language models , , to achieve open-vocabulary semantic segmentation. MaskCLIP combines a class-agnostic mask proposal network with a frozen CLIP encoder to achieve open-vocabulary panoptic segmentation. DenseCLIP achieves impressive zero-shot segmentation performance by fine-tuning the CLIP encoder. ODISE utilizes a pre-trained text-image diffusion model as a backbone to achieve open-vocabulary segmentation. OpenSeed achieves more powerful open-vocabulary segmentation performance by combining detection and segmentation datasets. Recently, FC-CLIP has achieved surprising open-vocabulary segmentation performance by training a mask generation decoder on the frozen CLIP encoder. In this paper, OV-DVIS++ is inspired by FC-CLIP and adopts the same approach to achieve the leading performance of open-vocabulary video segmentation.
METHODOLOGY
The overview of DVIS++. As shown in Figure 2, DVIS++ consists of four components: a segmenter, a noiser, a referring tracker, and a temporal refiner. Among these components, the noiser is non-trainable, whereas the other three are learnable modules. The segmenter extracts object representations from individual frames. During the training process, the noiser adds significant noise to the object representations produced by the segmenter, thus creating more challenging cases. The noiser effectively increases the difficulty of the object association subtask. During inference, the noiser is replaced by the Hungarian algorithm, which pre-matches the object representations of adjacent frames generated by the segmenter. The referring tracker learns to eliminate the noise introduced by the noiser, thereby achieving precise association of object representations in adjacent frames. Once the referring tracker has obtained temporally aligned object representations for all frames, the temporal refiner effectively utilizes temporal features to model spatio-temporal representations of objects. As a result, more accurate segmentation and tracking results are achieved.
The segmenter will be briefly introduced in Section 3.1, while the referring tracker and temporal refiner will be introduced in Sections 3.2 and 3.3, respectively. The noise training strategy and the details of the noiser will be presented in Section 3.4. In Section 3.5, we will explain of how the contrastive learning strategy is incorporated into the segmenter, referring tracker, and temporal refiner. In Section 3.6, we will describe how the visual foundation models are integrated into DVIS++ to allow DVIS++ to be evaluated in various settings, including open-vocabulary and freezing the pre-trained backbone. Finally, the objective functions will be introduced in Section 3.7.
will provide rich semantic information about objects for the referring tracker and temporal refiner. It should be noted that here, the object is a generic term for “thing” and “stuff,” thus covering semantic, instance, and panoptic segmentation.
2 Referring Tracker
The referring tracker models the inter-frame association as a task of referring denoising. Without the process of inter-frame object association, the i-th object representations in across all frames may not correspond to the same object. In other words, there is noise in the frame-by-frame association. Taking as the initial values, the referring tracker aims to eliminate this noise from the initial values and output the correct object representations, thereby achieving accurate temporal association results.
The overall architecture of the referring tracker is depicted as the blue part in Figure 2. It consists of L transformer denoising (TD) blocks connected in series. The referring tracker takes three inputs: , , and . Here, represents the reference, containing the object information propagated from the last frame. According to , the referring tracker aims to output aligned representations of objects in the T-th frame. is the object representation of the T-th frame outputted by the segmenter. represents the addition of noise to (the process will be detailed below), and the referring tracker aims to accurately filter out the noise added by noiser and the inherent noise in . The referring tracker outputs both the object representation and a new reference for the next frame. Additionally, also serves as the initial value for the temporal refiner.
The reference of the first frame is obtained by transforming using MLP:
The core of the referring tracker lies in the transformer denoising (TD) block. The architecture of the TD block is shown in Figure 3, which consists of a referring cross-attention (RCA), a self-attention module, and a feedforward neural network (FFN). The RCA plays a crucial role by effectively utilizing the similarity between corresponding objects represented in the previous and the current frame, as illustrated in the green part on the left side of Figure 3. The RCA takes three inputs: , , and . represents the initial values with noise, provides reference information, and provide object features. When is set to be equal to , the degenerates into a standard cross-attention.
3 Temporal Refiner
Finally, is used to predict the mask corresponding to the object in each frame, and is used to predict the object’s category throughout the entire video.
4 Denosing Training Strategy
In DVIS, the object representation outputted by the segmenter undergoes a coarse matching process using the Hungarian algorithm before being inputted into the referring tracker. This coarse matching process provides good initial values for the referring tracker, but it may hinder the effectiveness of the referring tracker’s training. The coarse matching results from the Hungarian algorithm are usually accurate because most available training video clips are simple. Therefore, the referring tracker learns very little using them as inputs, often leading to an identity mapping shortcut. That is to say, even without changing the inputs, the tracker can achieve a very low loss. To address this issue, we propose a denoising training strategy that enables the referring tracker to primarily learn how to handle challenging cases. Specifically, we have designed various noise simulation strategies to introduce strong noise into the input of the referring tracker. This eliminates the possibility of identity mapping shortcuts, thus compelling the referring tracker to acquire more stable and powerful tracking capabilities. As illustrated in Figure 2, during the training phase, the noiser introduces significant noise to the object representation. However, during the inference phase, noiser is not applied, and instead, the Hungarian algorithm is utilized for coarse matching of the object representation to provide a reliable initial value for the referring tracker.
The random weighted averaging strategy involves randomly combining each object representation with another randomly selected object representation:
The random cropping & concatenation strategy is applied to perform both random cropping and concatenating on the object representations:
The random shuffling strategy involves the random permutation of object representations:
5 Contrastive Learning
Contrastive learning aids the network in learning more discriminative object representations by minimizing the distance between anchor embedding and positive embeddings while concurrently enlarging the distance between anchor embedding and negative embeddings . The crucial aspect lies in the construction of contrastive items for the segmenter, referring tracker, and temporal refiner. Once the contrastive items are obtained, the contrastive loss can be computed using the following formula:
Contrastive Items of Segmenter. Figure 6 illustrates the construction of the contrastive item set . When selecting an object representation from the T-th frame as the anchor embedding, other object representations from the same frame are chosen as negative embeddings. Additionally, the representation of the same object from the previous frame is selected as the positive embedding. Furthermore, drawing inspiration from CTVIS , we calculate the momentum average of the object representations from previous frames using similarity-guided fusion and also utilize as a positive embedding. The process of calculating the momentum average is as follows:
where denotes the cosine similarity.
Contrastive Items of Referring Tracker. To enhance the between-frame consistency of the references, we construct the contrastive item set , as shown in Figure 6. In this process, when choosing a reference from the T-th frame as the anchor embedding, other references from the same frame are selected as negative embeddings. Additionally, positive embeddings are chosen as the reference from the previous frame and the reference from the subsequent frame.
Contrastive Items of Temporal Refiner. The temporal refiner effectively models the spatio-temporal representation of objects. The introduction of contrastive learning strengthens the discriminative nature of these representations. As shown in Figure 6, the object representation is chosen as the anchor embedding from the T-th frame. Negative embeddings are selected as other object representations from the same frame, while positive embeddings are chosen as the representations of the same object from other frames. To address the challenge of identification swapping (i.e., the mismatch of corresponding objects), we maintain a fixed-length memory bank that stores object representations from previous training batches. We then select object representations from the memory bank with the same category as the anchor embedding as additional negative embeddings.
6 Vision Foundation Models
In this section, we integrate some vision foundation models, such as DINOv2 and CLIP , into DVIS++ to enable evaluation under various settings, such as open-vocabulary and with a frozen pre-trained backbone.
DINOv2. Video segmentation is a dense prediction task that requires the generation of multi-scale features. However, the pre-trained VIT backbone of DINOv2 cannot output the required multi-scale features. To address this issue, here, the VIT-Adapter is incorporated to generate the necessary multi-scale features. It is important to note that video segmentation differs from image segmentation as it involves capturing inter-frame relationships from multiple frames during training. This results in higher GPU memory requirements. Thus, an efficient version of the VIT-Adapter is applied to reduce memory consumption. Figure 7 demonstrates this efficient version, where all injectors are removed, and the DINOv2 VIT backbone is frozen during training. These adaptations enable the DINOv2 pre-trained VIT backbone to be integrated with DVIS++.
CLIP. FC-CLIP demonstrates that training a mask generator from scratch based on a frozen CLIP backbone can achieve good open-vocabulary image segmentation capabilities. Inspired by FC-CLIP , we construct the segmenter to achieve open vocabulary video segmentation. As shown in Figure 8, the segmenter is built on the frozen CLIP backbone, where object category prediction is obtained by calculating the similarity between object representations and text embeddings. In addition, both the referring tracker and the temporal refiner are slightly adjusted to better adapt to open-vocabulary segmentation, as shown in Figure 9. Firstly, we observe that in some frames, objects may not display any discriminative features, such as a rabbit only showing its back with white fur. This may pose a challenge to object semantics recognition. To address this issue, we combine the reference and the query to predict the category in the referring tracker. Secondly, we no longer retrain the mask head and class head for the referring tracker and temporal refiner. Instead, we use the frozen, pre-trained mask head and class head from the segmenter, meaning that the referring tracker and temporal refiner only implement object representation mapping. Considering the improvements above, we have developed OV-DVIS++, which facilitates open-vocabulary universal video segmentation.
7 Objective Functions
Due to the requirement of modeling the relationship between multiple frames (usually more than 5 frames) for video segmentation, end-to-end training would require prohibitively high GPU memory requirements. Therefore, we adopt a separate training approach for the segmenter, referring tracker, and temporal refiner to alleviate the GPU memory demands. Specifically, we initially train the segmenter for the image segmentation task, followed by training the referring tracker with the frozen segmenter. Finally, we train the temporal refiner while keeping the segmenter and referring tracker frozen.
Objective Functions of Segmenter. The Segmenter is trained using the same objective functions as Mask2Former . Firstly, the predicted results are matched one-to-one with the ground truth using the Hungarian algorithm :
where is the matching cost used in . The predicted results that do not match with any ground truth will be assigned to the background.
The objective function is comprised of a contrastive item , a mask item , and a classification item . The mask item encompasses both the dice loss and the cross-entropy loss :
where , , , and are set as 2.0, 2.0, 5.0, and 5.0, respectively.
Objective Functions of Referring Tracker. The referring tracker tracks objects frame by frame, and as such, the network is supervised using an objective function that aligns with this paradigm. Specifically, the prediction results are only matched with the ground truth on the frame where the object first appears. To expedite convergence during the early training phase, the prediction results of the frozen segmenter are used for matching instead of the referring tracker’s prediction results.
where represents the frame in which the -th object first appears. The objective function is exactly the same as :
Implementation Details
Settings. DVIS++ employs Mask2Former as the segmenter. In DVIS++, the referring tracker utilizes six transformer denoising blocks, and the temporal refiner also employs six temporal decoder blocks. The channel dimensions of the object representations in both the tracker and refiner are consistent with those in the segmenter. The noiser applies a random weighted averaging strategy to introduce noise into the object representations output by the segmenter with probabilities of 0.5 and 0.8 when training for 40K and 160K iterations, respectively. These settings are uniform across all datasets.
OV-DVIS++ employs FC-CLIP as the segmenter. Besides the classification branch, the configurations of OV-DVIS++’s referring tracker and temporal refiner are identical to those of DVIS++. The classification process in OV-DVIS++ mirrors that of FC-CLIP, with the classification result derived from the concatenation of the query and reference. Additionally, FC-CLIP does not account for multi-dataset joint training. To facilitate this, we introduce a unique void embedding for each dataset, enabling joint training across multiple datasets.
Training & Inference. We employ the AdamW optimizer with an initial learning rate of 1e-4 and a weight decay of 5e-2 for training. The segmenter is initialized with weights pre-trained on the COCO dataset. In DVIS++, the segmenter undergoes fine-tuning on the video dataset before being frozen. Conversely, in OV-DVIS++, the segmenter is frozen directly without any fine-tuning. Notably, OV-DVIS++ is trained exclusively on the COCO dataset and performs zero-shot inference on video datasets without utilizing any video dataset-specific training.
For VIS datasets, such as YouTube-VIS 2019, 2021, and 2022, as well as OVIS, we utilize COCO pseudo videos for joint training with DVIS++. For VSS and VPS datasets, no additional datasets are employed. We implement random resize and random crop augmentations. During training, the videos in the VIS datasets are resized to a range of 320p to 640p, and to 480p for inference. For the VPS datasets, the videos are resized to a range of 480p to 800p during training and to 720p for inference. Unless otherwise specified, the referring tracker and temporal refiner are trained for 20K iterations, with the learning rate reduced to one-tenth at 14K iterations. If COCO pseudo videos are utilized for joint training, the number of iterations is increased to 40K, with the learning rate decay occurring at 28K iterations. The input for training the referring tracker comprises 5 consecutive frames sampled from the video while training the temporal refiner involves 21 consecutive frames.
Experiments
To validate the effectiveness of the proposed methods, we conducted extensive experiments on six benchmarks, as described below.
Youtube-VIS 2019. YouTube-VIS 2019 is a large-scale dataset for VIS that comprises 2,238/302/343 videos for training/validation/testing, respectively. The videos in YouTube-VIS 2019 are relatively short, and the object motion is relatively simple. This dataset encompasses 40 categories, out of which 22 categories overlap with the COCO dataset , while the remaining 18 categories are not included in the COCO dataset. The performance of VIS methods is evaluated using the AP (Average Precision) metric. The calculation process is similar to that in image segmentation, but with one key difference: it utilizes the IOU (Intersection over Union) of video segmentation results instead of image segmentation results. Further details can be found in .
Youtube-VIS 2021. YouTube-VIS 2021 was created by expanding the videos and refining the annotations from YouTube-VIS 2019. It consists of 2,985/421/453 videos for training/validation/testing, respectively. Additionally, it includes 40 object categories, which differ from the 2019 version. Out of these categories, 24 overlap with the COCO dataset, while the remaining 16 categories are not included in the COCO dataset.
Youtube-VIS 2022. The YouTube-VIS 2022 dataset utilizes the same training set as YouTube-VIS 2021 but incorporates extra-long videos in the validation and test sets. Additionally, YouTube-VIS 2022 employs different evaluation strategies, separately measuring AP on the long videos (APl) and the short videos (APs). The final performance is evaluated by taking the average of APl and APs.
OVIS. OVIS is a new and highly challenging VIS dataset comprising 25 object categories. Among these categories, 16 categories overlap with the COCO dataset, while the remaining 9 categories are unique to OVIS. The dataset includes 607/140/154 videos for training/validation/testing, which are longer and contain more instance annotations compared to the YouTube-VIS series. OVIS also features a significant number of videos with objects exhibiting severe occlusion, complex motion trajectories, and rapid deformation, thus making it more representative of real-world scenarios. Therefore, OVIS serves as an ideal benchmark for evaluating the performance of various VIS methods. Additionally, OVIS computes AP for objects with light, medium, and heavy occlusion, denoted as APl, APm, and APh, respectively.
VIPSeg. VIPSeg is a large-scale dataset for panoptic segmentation in the wild, which showcases a diverse array of real-world scenarios and encompasses 124 categories. These categories consist of 58 “thing” classes and 66 “stuff” classes. The dataset comprises 3,536 videos and 84,750 frames, with 2,806 videos allocated for training, 343 videos for validation, and 387 videos for testing. VPS methods are evaluated using the VPQ (Video Panoptic Quality) metric .
VSPW. VSPW is a large-scale video semantic segmentation (VSS) dataset that shares the same videos and categories as VIPSeg. The performance of VSS methods is evaluated using mIoU and VC (Video Consistency) .
2 Comparison with the State-of-the-art Methods
YouTube-VIS 2019 & 2021. YouTube-VIS 2019 and YouTube-VIS 2021 comprise videos with shorter durations and simpler scenes. Firstly, we compare our approach DVIS++ with SOTA VIS methods on these two datasets, and the results are presented in Table I. When using ResNet-50 as the backbone, DVIS++ achieves an AP of 55.5 and 50.0 in the online mode (without temporal refiner) on YouTube-VIS 2019 and YouTube-VIS 2021, respectively. This surpasses all contemporary online VIS methods. Additionally, by incorporating temporal information, DVIS++ achieves an AP of 56.7 and 52.0 in the offline mode (with temporal refiner) on YouTube-VIS 2019 and YouTube-VIS 2021, respectively. These results significantly outperform the previous SOTA methods (NOVIS and RefineVIS ) on these two datasets, with improvements of 3.9 AP (56.7 vs. 52.8) and 1.8 AP (52.0 vs. 50.2), respectively.
When using a frozen pre-trained VIT-L , DVIS++ achieves an AP of 67.7 and 62.3 on YouTube-VIS 2019 and YouTube-VIS 2021, respectively, in online mode, surpassing all previous SOTA methods, both online and offline. When inferring in offline mode, DVIS++ demonstrates even stronger performance by effectively utilizing spatio-temporal features, achieving an AP of 68.3 and 63.9 on YouTube-VIS 2019 and YouTube-VIS 2021, respectively. DVIS++ outperforms the previous SOTA methods UNINEXT and TCOVIS with improvements of 1.4 AP (68.3 compared to 66.9) and 2.7 AP (63.9 compared to 61.2), respectively, despite UNINEXT using a larger VIT-H backbone.
In addition, recent offline methods, such as GenVIS , MDQE , NOVIS , and RefineVIS , have only achieved comparable accuracy to online methods in semi-offline mode (processing the video clip by clip). They exhibit significant performance degradation in pure offline mode (processing the entire video at once). For instance, demonstrates GenVIS experiences approximately a 3 AP performance degradation when inferring in pure offline mode on YouTube-VIS 2021 compared to semi-offline mode. In contrast, our proposed DVIS and DVIS++ demonstrate a significant performance improvement when operating in pure offline mode, surpassing both semi-offline mode and online mode. This indicates that DVIS and DVIS++ effectively leverage spatio-temporal information.
OVIS. OVIS is a highly challenging VIS dataset, featuring lengthy videos (e.g., 500 frames) and a multitude of occlusion scenes, which brings it closer to real-world scenarios. As a result, OVIS serves as a more suitable benchmark for evaluating the performance of VIS methods in practical applications. Table I also presents a performance comparison of DVIS++ with other SOTA methods. When employing ResNet-50 as the backbone, DVIS++ achieves an AP of 37.2 in online mode and 41.2 in offline mode, surpassing all existing VIS methods. Notably, DVIS++ with ResNet-50 even outperforms the cutting-edge method MinVIS with the Swin-L backbone (41.2 vs. 39.4). It is worth highlighting that DVIS++ exhibits even more pronounced advantages in complex scenes by effectively modeling the spatio-temporal representation of objects. Specifically, it surpasses the previous SOTA method RefineVIS by 1.8 AP on the simple YouTube-VIS 2021 dataset and 7.5 AP on the complex OVIS dataset. Moreover, when employing a frozen pre-trained VIT-L backbone, DVIS++ achieves an AP of 49.6 in online mode and 53.4 in offline mode, outperforming the previous SOTA method GenVIS by 4.4 AP (49.6 vs. 45.2) and 8.0 AP (53.4 vs. 45.4), respectively.
In addition, only our proposed DVIS and DVIS++ demonstrate significant performance improvement between offline and online modes (48.6 vs. 45.9 and 53.4 vs. 49.6) in complex scenarios. Meanwhile, the contemporary methods GenVIS and RefineVIS exhibit similar or even worse performance in offline mode compared to that in the online mode (45.4 vs. 45.2 and 46.0 vs. 46.1). This proves the importance of our decoupling strategy in effectively utilizing spatio-temporal features.
YouTube-VIS 2022. YouTube-VIS 2022 has added a significant number of long videos on top of YouTube-VIS 2021. It has separately evaluated the performance of VIS methods using the APL metric on these long videos. Table II shows the performance comparison between DVIS++ and other SOTA methods. When using ResNet-50 as the backbone, DVIS++ achieves APL scores of 37.2 and 40.9, surpassing MinVIS with 13.9 APL (37.2 vs. 23.3) and GenVIS with 3.7 APL (40.9 vs. 37.2), respectively. When using a frozen pre-trained VIT-L as the backbone, DVIS++ achieves an APL score of 50.9 in offline mode, significantly outperforming VITA and GenVIS with 9.8 APL (50.9 vs. 41.1) and 6.6 APL (50.9 vs. 44.3), respectively. In the case of long videos, DVIS++ demonstrates a significant performance improvement in offline mode compared to online mode(40.9 vs. 37.2 with ResNet-50 and 50.9 vs. 37.5 with VIT-L), indicating the crucial importance of effectively utilizing temporal information for processing long videos.
2.2 Performance on Video Semantic Segmentation
VSPW. VSPW is a challenging large-scale video semantic segmentation dataset. We compare DVIS++ with other VSS methods on the validation set of VSPW, as shown in Table III. When using ResNet-50 as the backbone, DVIS++ achieves 92.3 mVC8, 91.1 mVC16, and 46.9 mIOU in online mode, surpassing all other methods in terms of video segmentation quality and consistency. DVIS++ outperformes Tube-Link by 3.1 VC8, 5.7 VC16, and 3.5 mIOU, as well as Video-kMax by 7.3 VC8, 9.7 VC16, and 2.6 mIOU. When running in offline mode, DVIS++ achieves 93.4 VC8, 92.4 VC16, and 48.6 mIOU, resulting in improvements of 1.1 VC8, 1.3 VC16, and 1.7 mIOU compared to online mode. When using a frozen pre-trained VIT-L as the backbone, DVIS++ achieves 95.0 VC8, 94.2 VC16, and 62.8 mIOU in online mode, and 95.7 VC8, 95.1 VC16, and 63.8 mIOU in offline mode, surpassing all competitors without any specific design for VSS.
2.3 Performance on Video Panoptic Segmentation
VIPSeg. VIPSeg is a large-scale video panoptic segmentation dataset with a wide range of real scenes and categories. The comparative results of DVIS++ and other methods are shown in Table IV. When using ResNet-50 as the backbone, DVIS++ achieves a VPQ of 41.9 in online mode, surpassing DVIS 2.5 VPQ and Tube-Link 2.7 VPQ. In offline mode, DVIS++ achieves a VPQ of 44.2 and an STQ of 43.6, surpassing all contemporaneous methods. When a frozen pre-trained VIT-L is used as the backbone, DVIS++ achieves a VPQ of 56.0 in online mode and 58.0 in offline mode, surpassing TarVIS 10.0 VPQ (58.0 vs. 48.0).
The experimental results demonstrate that DVIS and DVIS++ are powerful universal video segmentation baselines.
2.4 Performance on Open-Vocabulary Video Segmentation
VIS. We compare the zero-shot performance of OV-DVIS++ in open-vocabulary instance segmentation with other SOTA methods on the VIS datasets, and the results are presented in Table V. When utilizing ResNet-50 as the backbone and training solely on the COCO dataset, OV-DVIS++ achieves AP scores of 34.5, 30.9, and 14.8 on YouTube-VIS 2019, 2021, and OVIS datasets, respectively. Remarkably, OV-DVIS++ outperforms the previous SOTA method MindVLT by 11.4, 10.0, and 3.4 AP on YouTube-VIS 2019, 2021, and OVIS datasets, despite MindVLT being trained on the LVIS dataset , which encompasses more diverse categories than the COCO dataset. Furthermore, when employing ConvNext-L as the backbone, OV-DVIS++ achieves AP scores of 48.8, 44.5, and 24.0 on YouTube-VIS 2019, 2021, and OVIS datasets, respectively, surpassing all other open-vocabulary video instance segmentation methods.
When running in online mode, OV-DVIS++ outperforms FC-CLIP (a combination of FC-CLIP and MinVIS that we designed) in zero-shot performance by 5.6, 5.4, and 3.0 AP on the YouTube 2019, 2021, and OVIS datasets, respectively. Therefore, the referring tracker demonstrates its stronger tracking capabilities than heuristic algorithms, even with limited training data (training only on the COCO dataset). However, the pseudo-videos generated through affine transformations of images are too simplistic, lacking complex scenarios such as occlusions. Thus, when trained solely on the COCO dataset, OV-DVIS++ does not exhibit performance advantages in offline mode compared to online mode and even experiences performance degradation on the complex OVIS dataset (21.6 vs. 24.0).
When jointly trained on the image and video datasets, OV-DVIS++ achieves 60.1, 56.0, and 38.9 AP in online mode on the YouTube 2019, 2021, and OVIS datasets, respectively. In offline mode, OV-DVIS++ shows a performance improvement of 1.0, 0.7, and 1.7 AP on the YouTube 2019, 2021, and OVIS datasets compared to online mode.
VSS & VPS. As there are currently no methods for open-vocabulary video semantic and panoptic segmentation, we solely compare OV-DVIS++ with FC-CLIP. The results are presented in Table VI. When utilizing ResNet-50 as the backbone, OV-DVIS++ outperforms FC-CLIP in zero-shot performance by 3.3 mIOU and 2.1 VPQ. When ConvNext-L is employed as the backbone, OV-DVIS++ surpasses FC-CLIP by 5.4 mIOU and 1.0 VPQ. Besides the performance advantages, OV-DVIS++ exhibits significantly higher temporal consistency in segmentation results compared to FC-CLIP (91.3 mVC16 vs. 82.7 mVC16 and 93.0 mVC16 vs. 88.4 mVC16).
3 Ablation Study
We conduct ablation experiments on the validation set of OVIS to verify the effectiveness of the proposed components. The baseline, MinVIS ( in Table VII), consists of mask2former and a simple heuristic association algorithm. Table VII illustrates how DVIS++ is constructed based on this baseline and presents the impact of each component on performance.
Referring Tracker. The referring tracker is designed to replace heuristic association algorithms by modeling the tracking task as a reference denoising task. As shown in Table VII ( vs. ), this change results in a significant improvement of 7.0 AP, 4.2 APl, 7.0 APm, and 4.8 APh. This fully demonstrates that the referring tracker can learn association capabilities that far exceed those from the heuristic algorithm, especially for occluded objects.
In the referring tracker, the referring cross-attention we designed is the most crucial core component responsible for inter-frame information propagation. We replace referring cross-attention with standard cross-attention and observe a drastic drop in performance, as shown in Table VIII.
In addition, the initial value of the referring tracker’s input is also important. We attempt various initial value selections, and the results are shown in Table VIII. Firstly, the object representation outputted by the segmenter is directly used as the initial value, resulting in an AP of 31.9. When the matched obtained from the heuristic matching algorithm is used as the initial value, the model achieves an AP of 32.8. This slight improvement comes from the reduction of noise contained in the initial value.
We also attempt to use zero vector and learnable embedding as the initial value. In this case, the input for the referring tracker is the same for all objects, and the denoising task becomes a more challenging reconstruction task. When zero vector is used as the initial value, the referring tracker achieves a performance of 33.0 AP, surpassing the performance achieved by using matched as the initial value. When learnable embedding is used as the initial value, the referring tracker achieves a performance of 33.1 AP. It performs better on heavily occluded objects than zero initialization but worse on lightly occluded objects.
Through the attempts above, we discover that modeling more challenging tasks can enhance the tracking performance learned by the referring tracker. This revelation has motivated us to improve the referring tracker’s performance by incorporating simulated noise into the initial value, which will be elaborated upon in the following discussion of the denoising training strategy. Considering the synergistic effect of the denoising training strategy, despite using learnable embedding as the initial value yields the optimal outcomes, we ultimately opt for matched .
Temporal Refiner. The temporal refiner is designed to model the spatio-temporal representations of objects based on pre-aligned object representations from the tracker output. As shown in Table VII ( vs. ), the temporal refiner brings a significant improvement of 4.0 AP, 4.1 APm, and 4.4 APh. This indicates that modeling spatio-temporal representations of objects is crucial for challenging scenarios in the representative OVIS dataset .
We also conduct ablation experiments on the key components of the temporal refiner, and the results are presented in Table IX. The temporal refiner utilizes self-attention to model long-term temporal relationships. Removing long-term attention leads to a performance degradation of 3.2 AP. To model short-term temporal relationships, the temporal refiner employs 1D convolution. Although long-term attention theoretically covers this aspect, removing 1D convolution result in a performance degradation of 0.2 AP, indicating that 1D convolution is more effective in capturing short-term relationships. Furthermore, the temporal refiner leverages cross-attention to access image information provided by the segmenter, enabling the correction of potential errors such as segmentation and tracking errors. Removing cross-attention leads to a performance degradation of 0.8 AP, highlighting the reliance of the temporal refiner on the original information provided by the segmenter for correcting errors. The network heavily depends on self-attention to suppress confusion between different objects, and removing it results in a performance drop of 1.7 AP.
Denosing Training Strategy. The denoising training strategy is proposed to enhance the tracking capabilities of the referring tracker. As demonstrated in Table VII, the implementation of this strategy leads to notable performance improvements across various metrics: 3.7 AP, 2.7 APl, 4.3 APm, and 4.1 APh ( vs. ). Particularly, the denoising training strategy significantly enhances the tracker’s performance in challenging scenarios, including objects with moderate to heavy occlusion. As a result, the strategy yields more substantial performance improvements for heavily occluded objects compared to slightly occluded objects (+4.1 APh vs. +2.7 APl).
We conduct ablation experiments on the denoising training strategy to explore the effects of noise simulation strategies, noise injection probability, and iteration number. The results are shown in Table X. Firstly, different noise simulation strategies, including random weighted averaging (WA), random cropping and concatenation (CC), and random shuffling (RS), are used, and their performance is shown in to . Regardless of the noise simulation strategy used, we observe performance improvement. Among them, the WA strategy achieves the best performance, with a performance increase of 2.2 AP compared to no noise added. The RS strategy can be considered an extreme case of WA and CC, but its performance improvement is lower than that of the WA strategy. This is because although the RS strategy introduces strong noise, it significantly reduces the noise sampling space. The improvement of the denoising strategy is mainly reflected in occluded objects (with an increase of 3.3 APm and 2.9 APh), while it has no significant effect on slightly occluded objects (with a decrease of 0.4 APl).
In addition, we are surprised to find that the denoising training strategy significantly improves performance as the number of training iterations increases. When the training iterations are extended from 40K to 160K, the WA strategy achieves 36.7 AP, resulting in a performance gain of 1.4 AP.
The probability of adding noise, , also influences the effectiveness of the denoising training strategy. As demonstrated in to , the optimal performance is attained when is set to 0.5. However, when longer training iterations are employed, a higher probability of adding noise leads to improved network training. Ultimately, by adopting the WA strategy, setting to 0.8 and the iteration number to 160K, the model performance obtains a notable enhancement of 4.1 AP compared to not utilizing the denoising training strategy.
Contrastive Losses. We investigate the impact of contrastive learning on the segmenter, referring tracker, and temporal refiner. As shown in Table VII, utilizing contrastive loss during the training process of the segmenter results in more distinct object representations, leading to performance improvements of 0.8 AP, 8.0 APl, 1.3 APm, and 0.4 APh ( vs. ). Notably, contrastive learning has proved highly effective for objects with minor deformations, such as lightly occluded objects (8.0 APl improvement). However, it does not yield for heavily occluded objects with significant deformations (only 0.4 APh improvement).
When contrastive loss is implemented in the training process of the referring tracker to enhance the consistency of adjacent frame references, it results in improvements of 0.7 AP, 1.1 APl, 1.2 APm ( vs. ). The utilization of contrastive loss in the training of both the segmenter and referring tracker significantly improve the segmentation results for lightly occluded objects. However, contrastive loss does not enable satisfactory results for heavily occluded objects with large deformations.
Unfortunately, the use of contrastive loss during the training process of the temporal refiner results in a decrease in performance ( vs. ), with 0.6 AP, 0.7 APm, and 2.2 APh, despite a performance increase of 3.3 APl. The implementation of contrastive loss has the unintended consequence of suppressing the distinctions in those heavily occluded object representations across temporal frames. However, it has a beneficial impact on lightly occluded objects.
Qualitative Analysis. The video segmentation results for DVIS++ are presented in Figure 10. In the VIS prediction results, the three horses undergo significant deformation and severe occlusion. Despite these challenges, DVIS++ still manages to achieve flawless results. The perfect prediction results for VPS demonstrate DVIS++’s excellent capability in handling both ’thing’ and ’stuff’ objects. Furthermore, the prediction results for VSS underscore DVIS++’s exceptional segmentation quality and its remarkable temporal consistency.
The open-vocabulary segmentation results of OV-DVIS++ are shown in Figure 11. It can be observed that the model is capable of effectively tracking and segmenting new categories, such as “carrot”, “hay”, “lantern”, etc., even if they are not present in the training data.
However, some failure cases still exist in the prediction results of DVIS++, as shown in Figure 12. Firstly, DVIS++ cannot track fast-moving objects well. In the first video, when the bird with ID 2 suddenly takes off and moves quickly, DVIS++ mistakenly identifies it as a new object in subsequent frames. We think this problem can be alleviated by properly introducing the trajectory model information. Additionally, DVIS++ relies on a segmenter to perceive individual images, and when the segmenter does not work well, DVIS++ will fail. For example, in the second video, the segmenter fails to distinguish the closely clustered zebras, resulting in incorrect output from DVIS++. We will make efforts to address these issues in future work.
Conclusion
We introduce DVIS++, a novel universal video segmentation framework that effectively models the spatio-temporal representation of video objects through a decoupled design. It performs SOTA on six mainstream VIS, VSS, and VPS benchmarks. By leveraging CLIP, we also implement OV-DVIS++, an open-vocabulary universal video segmentation framework that achieves SOTA performance in zero-shot inference. Specifically, we decompose the video segmentation task into segmentation, tracking, and refinement. We propose novel referring tracker and temporal refiner modules to handle the tracking and refinement subtasks, respectively. To enhance the tracking capability of the referring tracker, we design a denoising training strategy. Furthermore, we investigate the impact of contrastive learning on the segmenter, referring tracker, and temporal refiner, highlighting its significance in video segmentation networks. Moreover, combining the visual foundation models, DVIS++ is evaluated under various settings. In the open-vocabulary setting, OV-DVIS++ achieves SOTA performance. Additionally, when evaluated with a frozen pre-trained backbone, DVIS++ works well and achieves higher performance.
We believe that DVIS, DVIS++, and OV-DVIS++ will serve as strong baselines for video universal segmentation, fostering future research in the fields of VIS, VSS, VPS, and related areas.