Anchor Diffusion for Unsupervised Video Object Segmentation

Zhao Yang, Qiang Wang, Luca Bertinetto, Weiming Hu, Song Bai, Philip H. S. Torr

Introduction

Video object segmentation (VOS) is a fundamental task in many important areas such as autonomous driving , robotic manipulation , video surveillance and video editing . Contemporary literature typically considers this problem in either the semi-supervised or the unsupervised setting. In both cases the objective is to predict, in every frame, pixel-level masks delineating certain objects of interest.

Under the semi-supervised setting, at test time methods can rely on a mask that specifies the object to segment. In contrast, the unsupervised setting does not provide any initialisation. Without online supervision, the task might be considered ambiguous, as different objects could be considered of interest for different reasons, according to the application. Among researchers, the current consensus is to segment foreground objects where a human gaze is more likely to focus . In more practical terms, an object is generally considered as foreground if it is sufficiently large, in motion and centred in the scene. In certain datasets (e.g., FBMS and ViSal ), in the same video, multiple foreground objects are considered, while in DAVIS-2016 only a single object is considered.

With the aim of tracking temporal changes of target objects, current state-of-the-art unsupervised approaches generally model motion cues in a video sequence via optical flow or recurrent neural networks (RNNs) . Typically, these methods sequentially propagate features from the previous steps to the current one, thus making the current prediction depending on the entire history of the video.

Though having the potential of exploiting informative temporal cues, these approaches suffer from several limitations. RNNs often rely on training techniques such as truncated backpropagation through time to reduce the cost of parameter updates, which limits their long-term modelling capability . Moreover, while LSTM’s gating mechanism alleviates the issue of vanishing gradients , the phenomenon of exploding gradients often requires clipping or rescaling the norm of the gradients during training . Optical flow vectors only predict one-step motion cues at each frame in a video, which can accumulate errors over time. What is more, models relying on optical flow are typically trained on synthetic videos due to the high cost of per-frame and per-pixel labelling. Therefore, when applying these systems to real videos, the domain gap can cause the flow fields to contain several inaccuracies, especially when the foreground is nearly static .

In the video object segmentation community, the deterioration of performance over time in unsupervised VOS methods based on optical flow or RNNs is well known and has been widely discussed . For instance, Li et al. demonstrate that, as a regular optical flow-based model progresses through frames, foreground embeddings become increasingly closer in feature space to the first frame’s background as opposed to the foreground. Furthermore, Voigtlaender et al. observe that a simple static segmentation model can achieve competitive results in the unsupervised VOS setting, which further corroborates the case for steering away from the sequential modelling strategies used by established methods.

Motivated by the above observations, in this work we opt for a much simpler solution, which is based on learning the similarity between pixels belonging to frames that can be arbitrarily far apart in time. To ensure representation consistency and reduce long-term drift, we propagate the features of the first frame (the “anchor”) to the current frame via an aggregation technique inspired by the non-local operation introduced by Wang et al. . This approach allows us to forgo of sequential modelling, while at the same time enabling us to deal with long-term dependencies and achieve high robustness over time, as shown in our experiments.

Despite its simplicity and online operability, our method outperforms the current state of the art on the DAVIS-2016 leaderboard by a margin of (absolute) 2.2%2.2\% in terms of intersection-over-union, without resorting to auxiliary training data or post-processing. Moreover, it also achieves state-of-the-art results on FBMS and the ViSal video saliency benchmark. Code and pre-trained models are available at https://github.com/yz93/anchor-diff-VOS.

Related work

The problem of video object segmentation (VOS) is tackled by the computer vision community in the unsupervised or semi-supervised settings, which are defined by the level of supervision provided at test time.

Semi-supervised VOS methods are provided with a pixel-wise mask identifying the target object in the first frame of a video. When aiming at very high segmentation accuracy, methods generally perform online fine-tuning on the basis of this supervision , sometimes exploiting data-augmentation techniques or self-supervision . As online fine-tuning can take up to several minutes per video, many recently proposed methods renounce to it and instead aim at a faster online speed (e.g., ). These faster semi-supervised approaches come in many flavours. For instance, Chen et al. learn a metric space for pixel embeddings, which is then used to establish associations between pixels across frames, while Cheng et al. suggest to individually track object parts from the first frame with a visual object tracker and then aggregate them according to their similarity with the initialisation mask.

Unsupervised VOS methods, instead, cannot rely on any supervision at test time and are often based on optical flow and RNNs. The purely optical flow-based MP-Net discards appearance modelling and casts segmentation as foreground motion prediction, an approach which poorly deals with static foreground objects. To address this problem, several methods (e.g., LVO , SegFlow , MotAdapt and MBN ) suggest to integrate appearance-based and optical flow-based features together, leading to variations of the “two-stream model” presenting two dedicated parallel branches. The drawbacks of these methods are threefold. First, flow estimation networks are typically trained on synthetic datasets and can thus result in poor performance when deployed in the real world. Second, while modelling long-term temporal dependencies is critical for adapting to significant online changes, the vector fields can only model short-term one-step dependencies. Targeting this issue, Tokmakow et al. proposed to extend the horizon spanned by optical flow-based features by employing a convolutional gated recurrent unit . Third, vector fields cannot distinguish foreground and background objects when they move in a synchronised fashion (e.g., the cars in a traffic jam). Li et al. attempt to address this issue by employing a bilateral network for detecting the motion of background objects. Our investigations with a much simpler appearance-based approach show that optical flow may not be an essential component of unsupervised VOS systems.

RNN-based models are often challenged by the problems of exploding and vanishing gradients , which limit their long-term modelling capability. Among the methods that make use of recurrent connections, Song et al. propose a novel convolutional long short-term memory architecture, in which two atrous convolution layers are stacked along the forward axis and propagate features in opposite directions.

Recently, it has been shown that both recurrent and optical flow-based methods significantly suffer from a deterioration in the quality of their predictions over time. This has motivated the several approaches (including ours) that tackle video object segmentation by simply learning similarities between pixel embeddings (e.g., ). These methods first select a set of seed pixels that are most likely to belong to the foreground object and then classify all other pixels based on their similarities to these seeds, for instance by thresholding or by propagating labels between neighbours. Fathi et al. adopt this approach for semantic instance segmentation, in which the pairwise pixel similarity function measures the likelihood of two pixels belonging to the same instance. IET extends this concept to video sequences. Similarly, it selects a set of foreground and background seeds for each frame and organises them into tracks. It then segments each frame individually based on pixel similarities with the foreground and background seeds. Note that IET utilises pre-trained instance embeddings. MBN extends IET with a bilateral filtering network that filters false-positive foreground predictions using optical flow features and an energy minimisation procedure on a graph of seeds sampled from a few consecutive frames. When segmenting frame tt, MBN classifies each pixel by assigning it the label of the seed (sampled from frames t−1t{-}1, tt, and t+1t{+}1) with which it has the smallest embedding distance.

The main drawback of these methods is in the complexity involved in the procedures of seed selection, ranking and classification, critical for achieving good performance. Moreover, these algorithms also depend on multiple scores such as motion saliency and objectness that need to be carefully calibrated and combined into one final metric.

Albeit our proposal is related to this last class of approaches, it is considerably simpler. Instead of separately learning individual components from image datasets and classifying pixels based on similarities with seeds, our method performs similarity learning, feature propagation and binary segmentation in a single network.

Method

We are interested in the task of binary segmentation of a sequence of video frames, where the final performance is measured by the average segmentation quality of individual frames. Therefore, our method should perform well under two aspects. First, similarly to what is expected from a static segmentation model, it should be able to provide accurate segmentation masks of foreground objects in individual frames. Second, it should be able to well adapt to the appearance changes of the foreground objects throughout the whole video.

In the proposed anchor diffusion network (AD-Net) (Figure 2), we address both requirements in a single model trained end-to-end by leveraging the recently proposed non-local operations of Wang et al. . Closely related to the concept of self-attention , a non-local operation is a neural network building block that captures the dependencies within a set of input feature vectors.

To achieve our first goal, a non-local operation is applied to the encoding of the target frame, in a similar way it is applied for semantic image segmentation , forming the intra-frame branch of our overall model. To achieve our second goal, we propagate information between two frames: a fixed anchor frame and the current frame, forming the anchor-diffusion branch of our overall model. We name the branch this way to give relevance to its functionality of “diffusing” information from the anchor to the large number of target frames at test time, which encourages foreground embeddings of each target frame to be consistent over time.

In the following, we describe our pipeline in more detail.

The entire network is trained end-to-end with a binary cross-entropy loss. Though any frame could be selected as the anchor frame, in practice we always choose the first frame for computational convenience and because, in benchmarks, the first frame is guaranteed to contain the foreground objects. During training, the first frame and a random frame are sampled from the video.

As qualitatively illustrated in Figure 1 and Appendix D, this procedure significantly strengthens the foreground while weakening the background. It is worth noting that one can also simply use the concatenation of X0X_{0} and XtX_{t} to achieve this goal. However, we find in our experiments that the correspondence learning in Equation (1) can better localise the foreground objects.

Similarly to , the transition matrix is defined as

where X0XtTX_{0}X_{t}^{T} is a pairwise dot product similarity between each pair of pixel embeddings in X0X_{0} and XtX_{t}. Following , we scale the dot product with a factor z=cz=\sqrt{c}, where cc is the number of channels of X0X_{0} and XtX_{t}. The rationale being that, for embeddings with high dimensionality, dot products can be very large and thus push the output of the softmax to regions where gradients are small . The softmax function normalises each row of 1zX0XtT\frac{1}{z}X_{0}X_{t}^{T} to sum to one, thereby preserving scale invariance of the pixel embeddings. Without normalisation, multiplying 1zX0XtT\frac{1}{z}X_{0}X_{t}^{T} with XtX_{t} can entirely change the scale of the pixel embeddings.

In the case of the intra-frame branch, each output pixel embedding can be considered as a global aggregation of all input pixel embeddings weighted by pairwise appearance similarity. It has been shown that such use of non-local operations can harness long-range spatial information, which is beneficial for semantic segmentation . Empirically, as detailed in the ablation studies of Table 1, we found that incorporating this branch in addition to the anchor-diffusion branch further improves the performance of the model.

The intra-frame branch improves segmentation accuracy but does not address the temporal changes in a video sequence. Conversely, the anchor-diffusion branch models pairwise dependencies between frames, with the result of enhancing the consistency of pixel embeddings and reducing drift.

Qualitative analysis. As shown in Figure 1, each of the coloured pixels in the anchor frame finds desirable correspondences in the current frame. The foreground car pixel embedding (red) has high similarity with pixel embeddings of the foreground car in the current frame despite the appearance change and sets off a neat contrast with the background that precisely outlines the target object. Conversely, both heat maps of the distractor car pixel embedding (green) and the road pixel embedding (purple) have higher similarity values in the background region of the current frame, which is what expected for pixel embeddings of the background class. Moreover, as the distractor car is not present in the current frame, its pixel embedding only find weak and widespread correspondences in the general background region, with a weak separation between the foreground and the background. In contrast, the pixel embedding corresponding to the asphalt, which represents the common material appearing in both frames, shows a higher similarity with the road region and sets a larger separation between the foreground and the background. More qualitative results illustrating the similarity between pixel embeddings are showed in Appendix D.

Overall, these results show that the transition matrix PP learns a similarity metric that can well identify common objects/materials across two frames. Therefore, when used in Equation (1), PP can strengthen the signal from pixels which have strong correspondences in the anchor frame and weaken the signal from pixels which do not. As the foreground target object is almost always present in both frames while the background changes relatively quickly, our diffusion process generally strengthens the foreground and suppresses the background.

In Figure 3, instead, we report how foreground embeddings change over time by computing the average cosine distance between the foreground embeddings of a later frame and those of the first frame. Notice how the embeddings of the baseline quickly grow apart, while the ones learned with our proposed method are significantly stabler. This suggests that AD-Net is capable of preserving the foreground information from the first frame in a video over long time-frames.

Experiments

In the following, after discussing important implementation details regarding our architecture and training procedure, in Section 4.1 we illustrate the three benchmarks we adopted, in Section 4.2 we describe several ablation studies and in Section 4.3, we provide an extensive comparison with the state of the art.

Implementation details. We employ the fully-convolutional DeepLabv3 as the feature encoder, and initialise its ResNet101 backbone with weights pre-trained on ImageNet. The other layers in DeepLabv3 are randomly initialised. The configuration of the dilation rates follows the original model and presents a total stride of 88. We modify the number of output channels in the last layer to 128128, which corresponds to cc in Section 3.

In the anchor-diffusion step, the spatial dimensions of each image encoding are flattened and transposed where appropriate in order to perform batched matrix multiplication. The outputs of the three branches are concatenated along the channel axis and reduced to dimension 128128 via a 1×11{\times}1 convolution with LeakyReLU non-linearity and dropout rate 0.10.1. The final classification layer is implemented as a 1×11{\times}1 convolution with a single output channel followed by a sigmoid layer.

Training. Each training example consists of a pair of images. Given a randomly sampled video, we use the first frame as the anchor image and a randomly sampled frame as the second image. We also experimented with randomly sampling both frames and observed slightly worse performance. Each input frame is cropped to a randomly-sized region enclosing the ground-truth foreground. Random rotations are performed at 4545-degree increments, with a probability of 51%51\% of not rotating and equal probabilities of rotating to any of the remaining angles.

The model is trained with binary cross-entropy loss. Network parameters are optimised via stochastic gradient descent with a weight decay of 0.00050.0005. The initial learning rate is set to 0.0050.005 and follows a “poly” adjustment policy , where the initial learning rate is multiplied by (1−iter40,000)0.9(1-\frac{iter}{40,000})^{0.9} at each iteration. The model is trained for 3000030000 iterations with batch size 88. Raw predictions are upsampled via bilinear interpolation to the size of the ground-truth masks.

Inference. At test time, the features of the anchor frame are computed once and reused throughout the video. Multi-scale and mirrored inputs are employed to enhance the final performance. Each input image is scaled by factors of 0.750.75, 1.001.00 and 1.501.50 and horizontally flipped. The final heatmap is the mean of all output heatmaps. Thresholding at 0.50.5 produces the final binary labels.

Instance pruning. Since semantic segmentation approaches like the one we use lack the notion of instance and some videos from the DAVIS-2016 dataset present multiple objects that can be deemed as foreground, we experiment with a simple set of post-processing steps to prune “non-foreground” objects. As instance trajectories measure the spatial changes of an instance, they can be used to detect background instances which have distinct trajectory patterns than the foreground instance. First, we establish online temporal correspondences by using a pre-trained object detection model to predict the locations of all objects and track the trajectory of each detection across the entire video using an intersection-over-union criterion between consecutive bounding boxes. Once object tracks have been established, we use the cumulative area of instance masks across frames as a proxy to identify foreground objects, thus pruning small objects or objects that are only present in a fraction of the video. This process produces a filtering mask, which is multiplied element-wise with AD-Net predictions to obtain the final predictions. More details and hyper-parameters related to this process (which we refer to as instance pruning) are provided in Appendix E.

Datasets. DAVIS is a benchmark and yearly challenge for video object segmentation (VOS). Unsupervised methods are trained and evaluated with the DAVIS-2016 dataset, which annotates a single foreground entity. There are 30 videos for training and 20 videos for validation. We train our method on the training set and evaluate on the validation set.

The FBMS dataset is another challenging benchmark for unsupervised video object segmentation containing 2929 training videos and 3030 test videos. Following , we evaluate on the test set.

Finally, ViSal is a video salient object detection dataset containing 1717 video sequences. Despite our method has not been designed for the task of saliency, we can easily report results on this benchmark too.

Evaluation metrics. For DAVIS, we adopt the official evaluation metrics of mean region similarity J\mathcal{J}, which is the intersection-over-union of the prediction and ground truth, and mean contour accuracy F\mathcal{F}, which is the F-measure defined on contour points from the prediction and the ground truth. To provide more insights, we plot precision-recall (PR) curves on all three benchmark datasets. On the FBMS dataset, the main evaluation metric is the F-measure. On the ViSal dataset, we report the mean absolute error (MAE) and the F-measure. For definitions of MAE and the F-measure, we refer readers to .

2 Ablation studies

We conduct several ablations to evaluate the effectiveness of the anchor-diffusion procedure. First, we evaluate DeepLabv3 as-is, simply fine-tuning it on the DAVIS training set. This semantic segmentation baseline (designed for static images) performs on par with some state-of-the-art unsupervised VOS methods (see Table 2). This is in line with what described by Voigtlaender et al. , but it is rather curious that it still applies after two years of progress. Clearly, the competitive performance can be partially attributed to the high performance of DeepLabv3 for the similar task of semantic segmentation of static images. However, this result also shows that existing unsupervised VOS techniques are not able to successfully model and leverage temporal dependencies and that different approaches should be sought.

Starting from this baseline, we evaluate four variants that differ in the embeddings they consider at the terminal concatenation layer (see Figure 2). Each corresponds to a row below Baseline in Table 1. The first variant (“intra-frame”) computes non-local features within the same frame XtX_{t} and without the anchor-diffusion branch. The second (“anchor”) simply concatenates X0X_{0} to XtX_{t}. The third performs anchor diffusion on X0X_{0} and XtX_{t}, and concatenates the results with XtX_{t}, without features from the intra-frame branch. The fourth (our final model, AD-Net) concatenates both the output of the intra-frame branch and that of the anchor-diffusion branch with XtX_{t}.

The “intra-frame” variant improves over the baseline, which shows the potential of utilising context information within the current frame. The “anchor” variant demonstrates the general usefulness of an anchor frame, despite the apparent limitation that the fixed representation of the anchor frame does not adapt to changes in the current frame. The solid performance gains validate our motivation to further develop the anchor-diffusion mechanism. The “anchor-diffusion” variant illustrates the efficacy of the proposed feature diffusion mechanism across the anchor and current frames. It brings a performance boost of 2.022.02 (absolute) points over the baseline, larger than the contribution brought by the “intra-frame” and “anchor” variants.

3 Comparison with the state of the art

In Table 2, we evaluate AD-Net against state-of-the-art unsupervised VOS methods on the DAVIS public leader-board and also provide the performance of several popular semi-supervised methods as a term of reference. AD-Net attains the highest performance among all unsupervised methods on the DAVIS validation set, while also performing very competitively on the FBMS test set. In particular, on DAVIS we outperform the second-best method (MotAdapt ) by an absolute margin of 2.2%2.2\% in J\mathcal{J} and 0.8%0.8\% in F\mathcal{F} before applying the post-processing step of instance pruning. After applying instance pruning as described earlier, AD-Net achieves the final performance of 81.781.7 in J\mathcal{J} and 80.580.5 in F\mathcal{F}, leading the second-best method by 4.54.5 and 3.13.1 absolute points respectively. Also, despite being unsupervised at inference time, AD-Net outperforms many semi-supervised methods which instead require to be initialised with a mask in the first frame.

After our proposed AD-Net, the second and third-best ranking methods are MotAdapt and PDB , which are particularly representative of two classes of methods.

PDB is representative of top-performing RNN-based methods. Although, in theory, RNNs could model long-range time dependencies, in practice they are constrained to model relatively short sequences. First, as the computational graph of an (unrolled) RNN grows in depth with the length of a video sequence, backpropagation is typically limited to a few time steps (e.g., 55 in RGMP ). Such backpropagation cannot guarantee long-term dependency modelling . Second, despite the gating and memory mechanisms adopted by LSTMs and GRUs, long propagation paths of gradients still cause exploding or vanishing gradients .

Conversely, MotAdapt is representative of top-performing methods that employ optical flow. It consists of a two-stream architecture, which dedicates two network branches (trained jointly but with different parameters) to process RGB images and pre-computed optical flow fields. The two-branch network is further fine-tuned at inference time, with pseudo-labels generated by a teacher network. Although optical flow is an intuitive way to model inter-frame dependencies and aid segmentation, results in Tables 1 and 2 demonstrate that simply developing a better appearance-based model can overshadow the benefits of a dedicated optical flow branch. Moreover, the strategy of fine-tuning at inference time adopted by MotAdapt and many semi-supervised methods is a time-consuming process, taking many seconds up to minutes per video. In contrast, AD-Net leverages a simpler architecture, which makes it fast at inference time. Without instance pruning, it runs online and at 44 frames per second on an NVIDIA TITAN X GPU, with frames at the original DAVIS resolution of 854×480854{\times}480. Speed can be easily traded off at a small cost in performance, by using frames with lower resolution and/or a lighter architecture.

The precision-recall analysis of AD-Net is presented in Figure 4, where we demonstrate that our approach generally outperforms also existing salient object detection methods. AD-Net achieves superior performance in all regions of the PR curve on the DAVIS validation set, maintaining significantly higher precision at all recall thresholds. On the challenging FBMS test set, AD-Net maintains a clear advantage below the 90%90\% recall threshold. On the ViSal dataset, it is noteworthy that nearly perfect precision is maintained up until the 60%60\% recall rate, which is higher than the other methods.

Evaluation as video saliency. The definition of salient objects in a video for benchmarks like ViSal is very related to the one of “foreground objects” for benchmarks like DAVIS or FBMS (see Section 1). Annotations in salient object detection datasets can vary from coarse annotations such as bounding boxes to fine-grained pixel-level real-valued scores, and sometimes even take the form of human eye fixations. ViSal provides pixel-level annotations as binary labels, annotating large, moving objects as the foreground and everything else as the background. Despite the many types of annotations, evaluation metrics are fairly standard and use pixel-level annotations either in a binarised form (PR curve and F-measure) or as normalised saliency scores between and 11 (MAE), which are directly applicable to the scores produced by AD-Net.

As shown in Table 3, the proposed AD-Net improves the state of the art for both DAVIS and FBMS also for standard saliency scores, showing consistency with Table 2. The largest improvements lie in FBMS, where both MAE and F-measure significantly outperform previous records. On DAVIS, F-measure is the highest among all methods with a significant leading margin. On the ViSal dataset, AD-Net achieves best MAE (lower is better) among all video saliency models and obtains F-measure close to the overall best method. Remarkably, despite not having trained for the task of saliency prediction, we outperform previous saliency methods under saliency metrics on DAVIS and FBMS, and achieve very competitive results on ViSal.

Conclusion

In this paper, we proposed Anchor Diffusion Network (AD-Net), a method for unsupervised video object segmentation based on non-local operations. Instead of modelling temporal dependencies with recurrent connections or adopting pre-computed optical flow like contemporary work, we argue for a significantly simpler and more effective approach, which consists in establishing correspondences of pixel embeddings between a reference frame and the current one. With this strategy, we can easily model long-term temporal dependencies at a low computational cost. We show how, during inference, this procedure is able to suppress the background while preserving the foreground even when abrupt changes in appearance occur. Quantitative evaluations across three standard benchmarks demonstrate the advantage of our proposed method on the task of unsupervised video object segmentation with respect to the state of the art. Moreover, our method is also surprisingly competitive against the state of the art in semi-supervised video object segmentation and video saliency.

Acknowledgements. This work was supported by the ERC grant ERC-2012-AdG 321162-HELIOS, EPSRC grant Seebibyte EP/M013774/1, EPSRC/MURI grant EP/N019474/1, and Tencent. We would also like to acknowledge the Royal Academy of Engineering and Five AI.

References

Appendix A Global Comparison

Table 4 includes all metrics reported in the official DAVIS 2016 benchmark . Our method outperforms competing methods in the main evaluation metrics of mean region similarity J\mathcal{J} and mean contour accuracy F\mathcal{F}. The small decay measure for both J\mathcal{J} and F\mathcal{F} shows AD-Net’s long-term benefits on performance.

Appendix B Per-sequence Comparison

Figures 6 and 7 compare the per-sequence J\mathcal{J} and F\mathcal{F} of AD-Net against the top seven competing methods on the leaderboard. Our method performs well on videos presenting a variety of challenges, such as appearance change (Car-Shadow, Parkour), cluttered background (Car-Roundabout, Scooter-Black), occlusion (Libby, Bmx-Trees), fast motion (Bmx-Trees, Dog, Parkour), etc.

Appendix C Qualitative Analysis on FBMS and ViSal

In Figures 8 and 9, we visualise segmentation results on videos from the test sets of FBMS and ViSal respectively. The model is trained only with the DAVIS 2016 training set. We do not fine-tune it on the training set of FBMS or ViSal.

Appendix D Foreground Correspondence Analysis

In Figure 10, we visualise more examples of foreground pixel correspondences to pixels in the anchor frame. Most pixels are randomly selected from the foreground area on the last frame of the video (except when foreground becomes too small in the last frame, in which case another frame is randomly chosen).

Appendix E Instance Pruning

Algorithm 1 details the instance pruning procedure. First, SmallStaticSmallStatic returns a set of bounding boxes and the corresponding instance masks that represent small and nearly static instances. Then, GetPruningMaskGetPruningMask takes these instances and the original masks as inputs, and generates a pruning mask per frame, which incorporates all small and static instances that are much smaller than the largest instance in the current frame. Finally, each input mask is multiplied element-wise with the corresponding pruning mask to output the final predictions.