DVIS-DAQ: Improving Video Segmentation via Dynamic Anchor Queries

Yikang Zhou, Tao Zhang, Shunping Ji, Shuicheng Yan, Xiangtai Li

Introduction

Video segmentation aims at simultaneously identifying, segmenting, and tracking all objects of interest within a video , which has many applications, including autonomous driving, video editing, content analysis, and video understanding . Modern video segmentation methods utilize object queries to perform cross-frame association and achieve remarkable performance. Despite large-scale motion and transient occlusion, these methods yield satisfactory results in continuously appearing objects.

However, these query-based methods exhibit a notable limitation: they tend to underperform when dealing with newly emerging or disappearing objects, even with recent state-of-the-art (SOTA) methods . As shown in the left part of Fig. 1, these methods only achieve less than 45% recall ratio on emerging and disappearing objects. As illustrated in the middle part of Fig. 1, they often fail to detect newly emerging objects. And they may incorrectly track a different, similar object when the original object disappears. Given that many objects appear and disappear over time in real-world long videos, it is imperative to improve these methods to handle objects’ emergence and disappearance more effectively, thereby meeting the requirements of real-world application requirements.

A natural question arises: "What causes the huge performance gap between continuously appearing objects and emerging/disappearing objects?" We argue the answer lies in the unreasonable anchor queries used in current methods. As shown in Fig. 2, SOTA query-based methods use the background query as the anchor query and model the object’s emergence and disappearance as a feature transition between the background and foreground query. However, a significant transition gap between background and foreground queries poses considerable challenges during training. To illustrate, consider emergence as an example. It is difficult for the model to learn how to transform a background anchor query into a foreground query without any semantic similarity. Consequently, the model may fail to transform and retain the background anchor query, resulting in missing newly emerging objects. Similarly, for disappearing objects, the model may prefer to incorrectly transform their query into another object’s query with more semantic similarity.

In this paper, we address this challenge by introducing Dynamic Anchor Queries (DAQ) mechanism to detect the emergence and disappearance of objects. As illustrated on the right side of Fig. 2, we regard all objects segmented in the current frame as potential candidates for newly emerging objects. We dynamically generate emergence anchor queries based on the features of these candidates to reduce the transition gap significantly. In particular, we designate all tracked objects as candidates for disappearing objects and generate disappearance anchor queries dynamically using their features. These emergence and disappearance anchor queries and the representations of tracked objects are collectively inputted into the tracker to identify the objects of interest in the current frame. Objects identified by the emergence anchor queries are classified as newly appeared, while those pinpointed by the disappearance anchor queries are recognized as having disappeared, prompting their removal. The remaining queries are responsible for tracking objects present across both past and current frames.

The proposed DAQ mechanism enhances the network’s ability to handle the emergence and disappearance of objects. However, emergence and disappearance events are rare in existing video segmentation datasets. To further leverage the capabilities of DAQ, we propose a training time simulation strategy for the emergence and disappearance of objects, termed Emergence and Disappearance Simulation (EDS). This strategy simulates emergence and disappearance events by removing high-level object queries. We mimic object emergence by randomly eliminating some object queries propagated from the previous frame. Likewise, we simulate object disappearance by removing some of the current frame’s object queries. Our simple yet effective simulation strategy generates numerous instances of object emergence and disappearance during training, thereby enabling comprehensive training of the DAQ mechanism and maximizing its potential.

To verify the effectiveness of our proposed DAQ mechanism and EDS strategy, we integrate them into the current SOTA method, DVIS , to construct the DVIS-DAQ. As shown in the right part of Fig. 1, our proposed DVIS-DAQ achieves new SOTA performance on five mainstream video segmentation benchmarks and significantly outperforms previous SOTA methods.

We have found that current query-based methods fall short in managing objects’ emergence and disappearance due to the considerable transition gap between foreground and background queries. We propose a Dynamic Anchor Queries (DAQ) mechanism to address this challenge. The DAQ mechanism dynamically generates anchor queries for the emergence and disappearance of objects, thereby shortening the gap and effectively overcoming the challenge.

To further unleash the potential of the DAQ mechanism, we introduce the Emergence and Disappearance Simulation (EDS) strategy. This straightforward and efficient approach amplifies the number of emergence and disappearance cases, thereby ensuring comprehensive training of the DAQ.

We enhance the SOTA video segmentation method DVIS by integrating our proposed DAQ mechanism and EDS strategy, resulting in DVIS-DAQ. With extensive experiments on five video segmentation benchmarks, DVIS-DAQ achieves new SOTA performance on all benchmarks. We further present detailed ablation studies to verify the effectiveness of each component.

Related Work

Video Instance Segmentation (VIS). Earlier works extend image instance segmentation methods by adding tracking heads and learning the feature association. With the raise of vision transformer , current SOTA video instance methods adopt query-based designs. In particular, Video K-Net and IDOL directly learn the query association embeddings via contrastive learning. After that, association-based methods , accomplish the tracking of video targets between adjacent frames or adjacent video clips through object matching. Meanwhile, several models directly predict 3D volumes of video instances across times in a semi-online manner. On the other hand, propagation-based methods delegate the tracking problem to the transformer decoder. They take the object queries outputted from the previous frame or video clip as input and iteratively optimize them to segment the objects in the current frame. Note that these methods can also be applied for video panoptic segmentation (VPS) and video semantic segmentation (VSS) tasks by adding queries for segmenting the stuff. However, both poorly handle challenges such as new-emerging objects and target disappearance. Our approach improves the model’s ability to address these challenges by employing dynamic anchor queries to manage the emergence and disappearance of objects.

Universal Video Segmentation. Current efforts in the field of video segmentation directly draw upon the design of universal image segmentation models to develop universal video segmentation models , achieving comparable or even superior performance to specialized video segmentation methods. Video K-Net and TubeFormer are the first to unify all three video segmentation tasks. TarVIS extends the concept of mask classification from image segmentation to video segmentation, accomplishing universal video segmentation via one shared model. DVIS receives frame object query outputs from Mask2Former for universal video segmentation. Recently, several works unified both image and video segmentation in one model. However, these methods are still unaware of object emergence and disappearance problems, which are significant for long videos.

Video Object Tracking. Object tracking is a crucial task in VIS and VPS , and many works adopt the tracking-by-detection paradigm . These methods divide the task into two sub-tasks, where an object detector first detects objects and then associates them using a tracking algorithm. Recent works also perform clip-wise segmentation and tracking. However, the former only focuses on single-object mask tracking, while the latter considers global tracking via clip-level matching. Our work belongs to mask-based tracking, which is orthogonal to these existing tracking methods.

Method

Query-based Video Segmentation Formulations. Modern mainstream video segmentation methods rely on queries to represent objects and achieve cross-frame association of objects. Here, we summarize current query-based video segmentation methods with the following formulation:

First, a segmenter S\mathcal{S} extracts query representations of objects QSegTQ^{T}_{Seg} from a single-frame image. Then, an association component A\mathcal{A} establishes connections between the current frame’s object representations QSegTQ^{T}_{Seg} and QT−1Q^{T-1} to obtain associated object representations QTQ^{T}. Finally, the object representations QTQ^{T} are decoded to obtain predictions of object categories CT\mathcal{C}^{T} and segmentation masks MT\mathcal{M}^{T} for the current frame. In this pipeline, the association component A\mathcal{A} is the most critical and challenging part. Most video segmentation methods focus on designing more efficient association modules.

The current mainstream association modules A\mathcal{A} can be divided into association-based and propagation-based pipelines. They both model object emergence as a transition from background queries to foreground queries and model disappearance as a transition from foreground queries to background queries. Specifically, we can split QQ (Q={QBg,QFg}Q=\{Q_{Bg},Q_{Fg}\}) into QFgQ_{Fg} and QBgQ_{Bg}, and the tracking of continuously appearing objects, newly emerging objects, and disappearing objects can be described as follows:

Limitations of the Query-based Methods. We find that current query-based methods perform poorly on objects that emerge and disappear in the middle of a video, as shown in Fig. 1. For emerging objects, SOTA methods often miss detecting them. For disappearing objects, SOTA methods frequently fail to recognize that the object has disappeared and may incorrectly segment another object. We quantitatively analyzed the recall ratios of emerging and disappearing objects in preliminary experiments using several SOTA methods. We found that the recall ratios of emerging and disappearing objects are very low (below 45%) for these methods, as shown in Fig. 1.

Thorough Analysis and Our Motivation. Query-based methods have demonstrated satisfactory performance on consistently appearing objects (exceeding 60 AP on the YouTube-VIS ), but they exhibit poor performance on newly emerging and disappearing objects. As depicted in Fig. 2 and Eq. 2, in current query-based methods, there is a significant disparity in the modeling formulation between newly emerging or disappearing objects and consistently appearing objects. The former poses more challenges because of the significant feature gap in the transition between foreground and background queries. In this work, we dynamically generate anchor queries for newly emerging and disappearing objects, resulting in a lower feature transition gap than the background query. We term the background query used in previous methods as Static Anchor Query (SAQ), the query of continuously appearing objects as Continuously Tracked Query (CTQ), and our proposed dynamically generated anchor query as Dynamic Anchor Query (DAQ). We use our proposed DAQ to replace the SAQ to address the challenge above. The details of the DAQ will be introduced in Sec. 3.2.

Furthermore, we observed that only less than 20% of samples contain object emergence and disappearance in the existing datasets , including the most challenging OVIS dataset. Therefore, to fully unleash the potential of our proposed DAQ, we introduce a simple yet effective query-level object emergence and disappearance simulation strategy, which will be detailed in Sec. 3.3.

2 Dynamic Anchor Queries

The key idea of DAQ is to dynamically generate anchor queries using the features of candidate objects that may emerge or disappear, thereby reducing the gap between anchor queries and the target queries (the queries of actual newly emerging or disappearing objects). DAQ allows the network to easily transform anchor queries to target queries and effectively handle object emergence and disappearance.

In this subsection, we introduce how to design and utilize Dynamic Anchor Queries (DAQ) to address the challenges mentioned above. First, we explain how the DAQ for newly emerging and disappearing objects are generated. Then, we explore how to make minor adjustments to the current query-based tracker to accommodate our proposed DAQ.

Dynamic Anchor Queries for Emergence. When processing the Tth frame of a video, we consider all objects appearing in the current frame as candidates for newly emerging objects. The dynamic anchor queries for emergence DAQEmg are naturally generated based on the features of these candidates. As shown in Fig. 3 (a), DAQEmg consists of two parts: query feature DAQEmgfeat{}_{Emg}^{feat} and query positional embedding DAQEmgpos{}_{Emg}^{pos}. The query feature DAQEmgfeat{}_{Emg}^{feat} is a learnable embedding shared by all dynamic anchor queries. The query positional embedding DAQEmgpos{}_{Emg}^{pos} is obtained using the appearance features FmaskTF_{mask}^{T} of the candidates through simple Mask-Pooling:

where FTF^{T} represents the image feature of the T−thT-th frame, and MT\mathcal{M}^{T} represents the segmentation masks of the candidates outputted by the segmenter.

Dynamic Anchor Queries for Disappearance. When processing the Tth frame, we consider all currently tracked objects as candidates for disappearing objects. Similar to the dynamic anchor queries for emergence, the DAQDis for disappearing objects are naturally generated based on the features of these candidates. As shown in Fig. 3 (b), the DAQDis also consists of query feature DAQDisfeat{}_{Dis}^{feat} and query positional embedding DAQDispos{}_{Dis}^{pos}. We use the candidates’ momentum-weighted appearance features F‾maskT\overline{F}_{mask}^{T} as the query positional embedding DAQDispos{}_{Dis}^{pos}. The momentum-weighted function follows CTVIS :

Since the disappearance of an object needs to be modeled as a transition from an anchor query to a background query, to reduce the transition gap, we use initial queries of segmenter as the query feature DAQDisfeat{}_{Dis}^{feat}. Specifically, for each query feature, we match the closest one from the initial queries of the segmenter based on cosine similarity.

Continuously Tracked Queries. We also assign query positional embedding CTQpos for continuously tracked queries (CTQ) to maintain the format consistency with dynamic anchor queries. We use the momentum-weighted appearance feature of the tracked objects as CTQpos, the same process as Eq. 3, 4, and 5.

Incorporating Dynamic Anchor Queries into a Tracker. As depicted in Fig. 4, we can seamlessly integrate DAQ into mainstream propagation-based trackers , requiring no alterations to the tracker’s network architecture. We instantiate two trackers: Tracker 1, which takes CTQ and DAQEmg as inputs and is responsible for tracking continuously appearing and newly emerging objects, and Tracker 2, which receives DAQDis along with several learnable background embeddings as inputs and identifies disappearing objects in the current frame. It’s important to note that the SoftMax function in Tracker 2 operates along the Q-dimension to prevent multiple dynamic anchor queries from segmenting the same object. Through this pipeline, we can effectively and reasonably manage and model the emergence and disappearance of objects. Despite introducing an additional tracker, it’s worth highlighting that the computational overhead is minimal, constituting less than 2% of the overall computational cost , as processing occurs solely at the query level rather than on dense image features.

3 Emergence and Disappearance Simulation

As shown in Fig. 5, when using DAQ, it is possible to simulate object emergence and disappearance separately by removing parts of CTQ and QSegQ_{Seg}, respectively. We term this simple yet efficient simulation strategy as Emergence and Disappearance Simulation (EDS), which includes Emergence Simulation (ES) and Disappearance Simulation (DS). EDS can simulate numerous object emergence and disappearance cases during training without introducing additional costs, allowing the network to be fully trained and unleashing the potential of DAQ.

New Emergence Simulation. As shown in Fig. 5 (a), taking the green object as an example, we can simulate it as an emerging object by removing its corresponding query from CTQ. At this point, this green object is not present in CTQ but exists in the current frame, consistent with the state of a real newly emerged object colored in gray. Then, the Tracker needs to transition the corresponding anchor query from DAQEmg to a foreground query and predict this green object’s segmentation mask and class.

Disappearance Simulation. As shown in Fig. 5 (b), taking the yellow object as an example, by removing its corresponding query from QSegQ_{Seg}, we can simulate this object as a disappearing object because the Tracker only detects features from QSegQ_{Seg} through cross-attention to obtain image information. At this point, this blue object is not present in the current frame but exists in the set of tracked objects, consistent with the state of a real disappearing blue object. The Tracker needs to transition the corresponding anchor query to a background query and determine the disappearance of this object.

4 Overall Architecture

Architecture. Sec. 3.2 details how DAQ is integrated into the current query-based tracker without any requirement for architecture modification. We apply the SOTA video segmentation method DVIS as our baseline and replace the Referring Cross-Attention in its tracker with standard Cross-Attention to achieve a more straightforward and generalizable structure shown in Fig. 6. We then integrated DAQ into this simplified tracker, resulting in our architecture, DVIS-DAQ.

Objective Function. Our objective function differs from DVIS . We maintain the historical matching relationship between the predictions of CTQ and ground truth for continuously tracked objects and remove the matched items from the ground truth. For newly emerging and disappearing objects, since DAQ is dynamically generated based on the features of candidates (refer to Sec. 3.2), we assign the remaining ground truth to DAQ based on the matching relationship between candidates and the remaining ground truth. Following this process, we obtain the matching relationship {(i,σ(i))∣i∈[1,M]}\{(i,\sigma(i))|i\in[1,M]\} between ground truth and predictions, M is the number of ground truth. The not-matched predictions are assigned with the background class and an empty segmentation mask. Then, the loss between matched prediction-label pairs (P,G)(\mathcal{P},\mathcal{G}) can be calculated following the Mask2Former :

Here, Lcls\mathcal{L}_{cls} represents the classification loss term, implemented using Cross Entropy Loss. Ldice\mathcal{L}_{dice} and Lbce\mathcal{L}_{bce} denote the segmentation loss terms, implemented respectively using Dice loss and Binary Cross Entropy loss.

Experiment

We conduct experiments on five mainstream video segmentation benchmarks, including OVIS , Youtube-VIS 2019 & 2021 & 2022 , and VIPSeg . We employ AP and AR as evaluation metrics for the VIS datasets following . We utilize Video Panoptic Quality (VPQ) and Segmentation and Tracking Quality (STQ) for the VPS datasets as evaluation metrics following . More detailed settings can be found in the supplementary file.

2 Main Experiments

Performance on OVIS dataset. The OVIS dataset presents more challenging cases, such as occlusion, fast motion, and complex motion trajectories. We present the quantitative evaluation results in Tab. 1. When using ResNet-50 as the backbone, our method achieves an AP of 38.7 and 43.5 with the offline refiner, surpassing all existing methods. Compared to the current SOTA DVIS++ , our method shows an improvement of 1.5 AP (38.7 vs. 37.2) and 2.3 AP (43.5 vs. 41.2) with the offline refiner. When employing a larger backbone, our method achieves even more significant improvements. With Swin-L as the backbone, our method outperforms DVIS by 3.6 AP (49.5 vs. 45.9). When using VIT-L as the backbone, our method surpasses DVIS++ by 4.1 AP (53.7 vs. 49.6). Notably, our DVIS-DAQ even outperforms DVIS++ with the offline refiner (53.7 vs. 53.4), while DVIS-DAQ with the offline refiner has reached 57.1 AP. These quantitative results demonstrate that our DAQ design significantly enhances the capability of query-based video segmentation methods in processing complex scenes with many object emergence and disappearances.

Performance on Youtube-VIS 2019 & 2021 datasets. The videos in these two datasets are relatively short and feature simple scenes. As shown in Tab. 1, when using ResNet-50 as the backbone, our method achieves a 4.0 AP improvement compared to the baseline DVIS on the YTVIS 2019 and 2021 datasets. When using VIT-L as the backbone, we achieve comparable performance with DVIS++. When using the offline refiner, we outperform DVIS++ by 0.9 AP and 0.6 AP on the YTVIS 2019 and 2021 datasets.

Performance on Youtube-VIS 2022 dataset. Youtube-VIS 2022 has been expanded with 71 long videos based on the 2021 version. Therefore, we only report the APL for the long videos in Tab. 2. When using ResNet-50 as the backbone, DVIS-DAQ outperforms the baseline DVIS with 3.0 APL. Our method outperforms DVIS++ with 4.5 APL (42.0 vs. 37.5) when using VIT-L as the backbone.

Performance on VIPSeg dataset. VIPSeg is a large-scale video panoptic segmentation dataset containing various real-world scenes and more categories. Tab. 2 presents the performance comparison with SOTA methods. Our method achieves new SOTA performance with both ResNet-50 and VIT-L backbone settings. Our method achieves a VPQ of 42.1 when using ResNet-50 as the backbone, surpassing the baseline DVIS by 2.7 VPQ. With VIT-L as the backbone, our method achieves a VPQ of 57.4, demonstrating a 2.4 VPQ improvement over the previous SOTA DVIS++ in "thing" objects.

3 Ablation Studies and Analysis

We perform ablation experiments on the OVIS dataset using the ResNet-50 backbone and 40K training iterations. We use DVIS as the baseline. When aligning the training and testing settings with DVIS++ , the baseline achieves an AP of 35.4. More detailed settings are provided in the supplementary materials.

Effectiveness of DAQ and EDS. Tab. 3(a) presents the ablation studies about our proposed components. Using emergence dynamic anchor queries alone leads to a performance degradation of 2.1 AP, attributed to insufficient examples of emergence objects for training. When the emergence simulation strategy is used alone, performance drops by 3.1 AP, indicating that the unreasonable anchor design in the previous method makes it difficult for the network to learn to handle object emergence, even with sufficient training cases. When emergence dynamic anchor queries and emergence simulation are used together, the reasonable mechanism and sufficient training cases result in a 1.0 AP improvement compared to the baseline. Disappearing dynamic anchor queries contribute to a performance improvement of 0.2 AP. Introducing the Disappearing Simulation strategy to provide sufficient cases of disappearing objects further increases model performance by 0.5 AP.

Emergence dynamic anchor queries. DAQ comprises query feature embedding and positional embedding. Firstly, we explore different strategies for selecting emergency candidates, and the results are presented in Tab. 3(b). When choosing the top 100 object predictions of the segmenter as candidates, with only one shared learnable embedding, we obtain 36.4 AP. Setting the number of candidates to the top 50 and the number of learnable embeddings to 2 brings a performance gain of 0.1 AP. To explore potential new objects fully, we choose the top 100 scheme.

In Tab. 3(e), we investigate how to generate corresponding dynamic anchor queries based on the features of emerging candidates. We find that utilizing candidates’ appearance features yields the best results, surpassing the use of positional features and object queries with a hybrid of positional and appearance information by 1.6 AP and 0.4 AP, respectively. Furthermore, leveraging the features of candidates in the form of query positional embedding results in the best performance, outperforming addition and concatenation operations by 0.4 AP and 0.2 AP, respectively.

Disappearance dynamic anchor queries. In Tab. 3(f), we explore how to initialize the query feature embedding and positional embedding of disappearance DAQ. We find that initializing query feature of disappearance DAQ from the initial queries in the segmenter (as shown in Fig. 3) yields better AP than using a new shared learnable embedding for all disappearance DAQ (37.1 vs. 36.5). For generation of positional embedding of disappearance DAQ, we find utilizing the momentum-weighted appearance feature of tracked objects yield better AP than utilizing the query feature of tracked objects (37.1 vs. 36.8).

Emergence and disappearance simulation. We employ a threshold to filter out low-scoring tracked objects in emergence simulation. Tab. LABEL:tab:new_sim_thresh illustrates the results when using different thresholds. A minimal threshold (0.01) results in insufficient new emergence cases, whereas a larger threshold (0.5) disrupts the model’s ability to learn continuous tracking. Both result in poor performance (34.2 AP and 33.2 AP). A similar phenomenon is observed in the disappearance simulation, as shown in Tab. 3(d). The simulation of an object’s disappearance in a solitary frame within a video clip fails to generate sufficient disappearance cases. Conversely, simulating the disappearance across all frames compromises the model’s capability to learn continuous tracking. Consequently, we have determined that a threshold of 0.1 strikes an optimal balance for simulating emergence, and we opt for a stochastic selection of frames in a video clip to simulate disappearance. This strategy ensures ample emergence and disappearance cases while preserving the model’s tracking capability for continually appearing objects.

Qualitative results comparison. In Fig. 7, we present the visual comparison with our baseline. In this challenging scene, there are new-emerging, disappearing, and reappearing video objects. Our method accurately addresses all these scenarios, whereas DVIS suffers from issues including missed new-emerging objects and ID switches. Although DVIS detected object 1 in frame 2, it incorrectly switched the ID of another object to that of object 1. In the case of the reappearance of objects 1 and 2 in Frame 6, our method successfully maintained continuous tracking. In Frame 7, DVIS incorrectly switched the IDs of two nearby disappearing objects to objects 5 and 6, whereas our method correctly handled these disappearing cases.

Conclusion

We observed that current SOTA query-based video segmentation methods perform poorly on newly emerging and disappearing objects. We have analyzed and identified the significant transition gap hindering the network’s learning. We address this challenge by proposing dynamic anchor queries for shortening the transition gap. We also introduce a query-level simulation strategy for simulating object disappearance and emergence without additional costs. By integrating our proposed dynamic anchor queries and simulation strategy into DVIS, we obtained DVIS-DAQ, which achieved SOTA performance on mainstream video segmentation benchmarks. Our research will inspire the video community to bridge the gap between current advancements in the standard benchmarks and real-world application requirements.

Limitations. We have significantly enhanced the capability of current query-based methods to handle newly emerging and disappearing objects. However, we introduced an additional tracker specifically to address object disappearance. While the additional computational cost is minimal, it undoubtedly introduces more hyperparameters. In the future, we will explore using a single model to effectively handle continuously appearing, newly emerging, and disappearing objects. Additionally, our model reaches its performance peak after 160k training iterations. Although this issue is prevalent in other methods, such as GenVIS and DVIS++, we believe there is ample room for optimization to enhance the model’s training efficiency.

References

Appendix 0.A Appendix

Overview. In this section, we first introduce more implementation details and training settings of our model. To further substantiate the generalizability and effectiveness of our DAQ design, we integrate the DAQ into an alternative video segmentation framework, GenVIS , and conduct experiments on the OVIS dataset. Additionally, we provide more ablation studies to verify the effectiveness of the detailed components of the DAQ design. Lastly, we present a collection of qualitative results and supplementary video files to demonstrate the full performance of our method in challenging scenes.

Spatio-temporal padding. The tracker incorporating with DAQ outputs NtrcN_{trc} video objects for a video with TT frames, denoted as {{Qi}tiTi}i=1Ntrc\{\{Q^{i}\}^{T_{i}}_{t_{i}}\}^{N_{trc}}_{i=1}, where the temporal dimension lengths (Ti−ti+1T_{i}-t_{i}+1) of these NtrcN_{trc} video objects may vary. Therefore, before feeding into the temporal refiner, we pad each video object sequence with the momentum-weighted appearance feature at the time steps t∈{1,…,ti−1,Ti+1,…,T}t\in\{1,\ldots,t_{i}-1,T_{i}+1,\ldots,T\}, where the video object does not appear. The padded results are denoted as {{Qi}1T}i=1Ntrc\{\{Q^{i}\}^{T}_{1}\}^{N_{trc}}_{i=1}. In addition to temporal padding, we also perform padding in the spatial dimension. We first use the Hungarian algorithm to naively extract NN video object feature sequences {{Qnaivei}1T}i=1N\{\{Q_{naive}^{i}\}^{T}_{1}\}^{N}_{i=1} from the object queries QSeg∈RN×CQ_{Seg}\in R^{N\times C} of all frames, as MinVIS does. Then, we select the top N−NtrcN-N_{trc} video objects based on classification scores to pad {{Qi}1T}i=1Ntrc\{\{Q^{i}\}^{T}_{1}\}^{N_{trc}}_{i=1} in the spatial dimension to obtain {{Qi}1T}i=1N\{\{Q^{i}\}^{T}_{1}\}^{N}_{i=1}.

Training details. We employ the AdamW optimizer with an initial learning rate of 1e-4 and a weight decay of 5e-2 for training. For the VIS task, we use the COCO joint training setting to train the model. For the VPS task, no additional datasets are employed. When training in offline mode, we freeze all modules except the temporal refiner and initialize these modules with parameters trained in online mode. We set the lengths of training video clips as 5 and 15 for online and offline modes, respectively. For all datasets, we train for 160K iterations and apply learning rate decay at 112K iterations.

A.2 Additional Experiments

To further substantiate the efficacy of the Dynamic Attention Query (DAQ), we integrated it into an alternative video segmentation framework, GenVIS , culminating in the development of an augmented architecture dubbed GenVIS-DAQ. We employed ResNet-50 as the backbone and conducted experiments on the more challenging OVIS dataset. For a fair comparison, we used GenVIS without the instance prototype memory (IPM) as our baseline and adhered to the original training strategies utilized by GenVIS. The results are reported in Tab. A1. By incorporating DAQ, we achieved an improvement of 1.7 AP for GenVIS. This demonstrates that using DAQ to manage the emergence and disappearance of objects in videos can enhance the capability of query-based video segmentation methods to handle complex scenes.

A.3 More Ablation Studies

Softmax in the disappearance tracker. As depicted in Tab 2(a), compared to the standard cross-attention, which applies SoftMax on the key dimension, we observe that applying SoftMax on the query dimension yields better performance. This is because employing SoftMax on the query dimension prevents different disappearance DAQ from merging the features of the same candidate into the same target query.

Spatio-temporal padding. As shown in Tab. 2(b), prior to the integration of online results into the offline module, it is imperative to apply both temporal and spatial padding. This preparatory step guarantees that the input is appropriately aligned in both time and space to fulfill the requirement of the offline module. Opting to pad with momentum-weighted appearance features rather than zero or learnable padding affords the advantage of tailoring feature information to each individual video object. This tailored padding strategy has been instrumental in achieving a notable improvement in AP (42.7 vs. 38.4). Additionally, performing spatial padding subsequent to temporal padding can further enhance the AP (43.5 vs. 42.7).

A.4 More Qualitative Results

Fig. A1 presents qualitative results in complex scenes. To fully demonstrate the performance of our method in challenging scenes such as fast motion and occlusions, we also provide supplementary videos for reference.