Decoupling Features in Hierarchical Propagation for Video Object Segmentation
Zongxin Yang, Yi Yang
Introduction
Video Object Segmentation (VOS), which aims at recognizing and segmenting one or multiple objects of interest in a given video, has attracted much attention as a fundamental task of video understanding. This paper focuses on semi-supervised VOS, which requires algorithms to track and segment objects throughout a video sequence given objects’ annotated masks at one or several frames.
Early VOS methods are mainly based on finetuning segmentation networks on the annotated frames or constructing pixel-wise matching maps . Based on the advance of attention mechanisms , many attention-based VOS algorithms have been proposed in recent years and achieved significant improvement. STM and the following works leverage a memory network to store and read the target features of predicted past frames and apply a non-local attention mechanism to match the target in the current frame. Furthermore, AOT introduces hierarchical propagation into VOS based on transformers and can associate multiple objects collaboratively by utilizing the IDentification (ID) mechanism . The hierarchical propagation can gradually propagate ID information from past frames to the current frame and has shown promising VOS performance with remarkable scalability.
Fig. 1(a) shows that AOT’s hierarchical propagation can transfer the current frame feature from an object-agnostic visual embedding to an object-specific ID embedding by hierarchically propagating the reference information into the current frame. The hierarchical structure enables AOT to be structurally scalable between state-of-the-art performance and real-time efficiency. Intuitively, the increase of ID information will inevitably lead to the loss of initial visual information since the dimension of features is limited. However, matching objects’ visual features, the only clues provided by the current frame, is crucial for attention-based VOS solutions. To avoid the loss of visual information in deeper propagation layers and facilitate the learning of visual embeddings, a desirable manner (Fig. 1(b)) is to decouple object-agnostic and object-specific embeddings in the propagation.
Based on the above motivation, this paper proposes a novel hierarchical propagation approach for VOS, i.e., Decoupling Features in Hierarchical Propagation (DeAOT). Unlike AOT, which shares the embedding space for visual (object-agnostic) and ID (object-specific) embeddings, DeAOT decouples them into different branches using individual propagation processes while sharing their attention maps. To compensate for the additional computation from the dual-branch propagation, we propose a more efficient module for constructing hierarchical propagation, i.e., Gated Propagation Module (GPM). By carefully designing GPM for VOS, we are able to use single-head attention to match objects and propagate information instead of the stronger multi-head attention , which we found to be an efficiency bottleneck of AOT .
To evaluate the proposed DeAOT approach, a series of experiments are conducted on three VOS benchmarks (YouTube-VOS , DAVIS 2017 , and DAVIS 2016 ) and one Visual Object Tracking (VOT) benchmark (VOT 2020 ). On the large-scale VOS benchmark, YouTube-VOS, the DeAOT variant networks remarkably outperform AOT counterparts in both accuracy and run-time speed as shown in Fig. 1(c). Particularly, our R50-DeAOT-L can achieve 86.0% at a nearly real-time speed, 22.4fps, and our DeAOT-T can achieve 82.0% at 53.4fps, which is superior compared to AOT-T (80.2%, 41.0fps). Without any test-time augmentations, our SwinB-DeAOT-L achieves top-ranked performance on four VOS/VOT benchmarks, i.e., YouTube-VOS 2018/2019 (86.2%/86.1%), DAVIS 2017 Val/Test (86.2%/82.8%), DAVIS 2016 (92.9%), and VOT 2020 (0.622 EAO).
Overall, our contributions are summarized below:
We propose a highly-effective VOS framework, DeAOT, by decoupling object-agnostic and object-specific features in hierarchical propagation. DeAOT achieves top-ranked performance and efficiency on four VOS/VOT benchmarks .
We design an efficient module, GPM, for constructing hierarchical matching and propagation. By using GPM, DeAOT variants are consistently faster than AOT counterparts, although DeAOT’s propagation processes are twice as AOT’s.
Related Work
Semi-supervised Video Object Segmentation. Given a video with one or several annotated frames (the first frame in general), semi-supervised VOS requires algorithms to propagate the mask annotations to the entire video. Traditional methods often solve an optimization problem with an energy defined over a graph structure . Based on deep neural networks (DNN), deep learning based VOS methods have achieved significant progress and dominated the field in recent years.
Finetuning-based Methods. Early DNN-based methods rely on fine-tuning pre-trained segmentation networks at test time to make the networks focus on the given object. Among them, OSVOS and MoNet propose to fine-tune pre-trained networks on the first-frame annotation. OnAVOS extends the first-frame fine-tuning by introducing an online adaptation mechanism. Following these approaches, MaskTrack and PReM further utilize optical flow to help propagate the segmentation mask from one frame to the next.
Template-based Methods. To avoid using the test-time fine-tuning, many researchers regard the annotated frames as templates and investigate how to match with them. For example, OSMN employs a network to extract object embedding and another one to predict segmentation based on the embedding. PML learns pixel-wise embedding with the nearest neighbor classifier, and VideoMatch uses a matching layer to map the pixels of the current frame to the annotated frame in a learned embedding space. Following these methods, FEELVOS and CFBI(+) extend the pixel-level matching mechanism by additionally doing local matching with the previous frame, and RPCM proposes a correction module to improve the reliability of pixel-level matching. Instead of using matching mechanisms, LWL proposes to use an online few-shot learner to learn to decode object segmentation.
Attention-based Methods. Based on the advance of attention mechanisms , STM and the following works (e.g., KMN and STCN ) leverage a memory network to embed past-frame predictions into memory and apply a non-local attention mechanism on the memory to propagate mask information to the current frame. Differently, SST proposes to calculate pixel-level matching maps based on the attention maps of transformer blocks . Recently, AOT introduces hierarchical propagation into VOS and can associate multiple objects collaboratively with the proposed ID mechanism.
Visual Transformers. Transformers was initially proposed to build hierarchical attention-based networks for natural language processing (NLP). Compared to RNNs, transformer networks model global correlation or attention in parallel, leading to better memory efficiency, and thus have been widely used in NLP tasks . Similar to Non-local Neural Networks , transformer blocks compute correlation with all the input elements and aggregate their information by using attention mechanisms . Recently, transformer blocks were introduced to computer vision and have shown promising performance in many tasks, such as image classification , object detection /segmentation , image generation , and video understanding .
Based on transformers, AOT proposes a Long Short-Term Transformer (LSTT) structure for constructing hierarchical propagation. By hierarchically propagating object information, AOT variants have shown promising performance with remarkable scalability. Unlike AOT, which shares the embedding space for object-agnostic and object-specific embeddings, we propose to decouple them into different branches using individual propagation processes. Such a dual-branch paradigm avoids the loss of object-agnostic information and achieves significant improvement. Besides, a more efficient structure, GPM, is proposed for hierarchical propagation.
Rethinking Hierarchical Propagation for VOS
Attention-based VOS methods are dominating the field of VOS. In these methods, STM and following algorithms uses a single attention layer to propagate mask information from memorized frames to the current frame. The use of only a single attention layer restricts the scalability of algorithms. Hence, AOT introduces hierarchical propagation into VOS by proposing the Long Short-term Transformer (LSTT) structure, which can propagate the mask information in a hierarchical coarse-to-fine manner. By adjusting the layer number of LSTT, AOT variants can be ranged from state-of-the-art performance to real-time run-time speed.
where the matching (or attention) map is calculated by the correlation function, .
Obviously, before all the propagation layers, the current frame feature, , is an object-agnostic feature extracted from an image encoder (e.g., ResNet-50 ). Nevertheless, the mask information will be gradually and hierarchically propagated into the current frame, and the output feature, , will become object-specific and can be decoded into the ID/mask prediction by a decoder network (e.g., FPN ). In other words, step by step, the hierarchical propagation transfers the current frame feature, , from an object-agnostic visual embedding to an object-specific ID embedding, as demonstrated in Fig. 1(a).
Intuitively, the absorption of object-specific ID information will inevitably lead to the oblivion of object-agnostic visual information within since the channel dimension of is limited. Such a phenomenon can also be observed by increasing the ID information directly. As shown in Fig. 2, the performance of AOT heavily drops as we increase the information amount of by containing more IDs inside. On the other hand, the significant progress of VOS in recent years is mainly based on matching object-agnostic visual embeddings (e.g., pixel-level matching methods and single-layer attention-based methods mentioned above). Hence, we argue that the loss of visual information in deeper propagation layers limits the performance of hierarchical propagation.
How to design a hierarchical propagation structure which can keep or even refine the initial object-agnostic visual information? Fig. 1(b) shows a simple, straightforward, and desirable approach, i.e., propagating object-agnostic and object-specific information in two different branches (Visual Branch and ID Branch). The object-agnostic branch is responsible for gathering visual information, refining visual features, and matching objects. By contrast, the object-specific branch is responsible for absorbing ID information propagated from memorized frames. These two branches share the attention maps used to match objects and propagate features. Compared to the single-branch LSTT, our dual-branch approach can keep and further refine visual features in the hierarchical propagation and thus can further facilitate the learning of visual embeddings.
Decoupling Features in Hierarchical Propagation
This section will introduce a new framework, Decoupling Features in Hierarchical Propagation (DeAOT), for solving semi-supervised video object segmentation. We show an overview of DeAOT in Fig. 3(a). Given a video with a reference frame annotation, DeAOT propagates the annotation to the entire video frame-by-frame. The multi-object annotation is encoded by the IDentification (ID) mechanism . Different from AOT, DeAOT decouples the hierarchical propagation of visual embedding and ID embedding, i.e., DeAOT propagates these two embeddings in two branches. Furthermore, DeAOT constructs the hierarchical propagation by using the proposed Gated Propagation Module (GPM), which is more efficient and effective than the LSTT block used in AOT.
Different from the previous attention-based VOS methods , DeAOT propagates objects’ visual features and mask features in two parallel branches. In detail, the visual branch is responsible for matching objects, gathering past visual information, and refining object features. To re-identify the objects, the ID branch reuses the matching maps (attention maps) calculated by the visual branch to propagate the ID embedding (encoded by the ID mechanism ) from past frames to the current frame. Both the branches share the same hierarchical structure with propagation layers.
Visual Branch is responsible for matching objects by calculating attention maps on patch-wise visual embeddings. The visual embeddings in the memorized frames will be propagated to the current frame regarding the attention maps. Since the propagation is not directly related to the object-specific ID embedding, the visual branch can learn to refine visual embeddings to be more contrastive but avoid being biased toward the given object-specific information. Let denote visual embeddings, we modify Eq. 2 into a layer of object-agnostic visual propagation,
which doesn’t leverage the object-specific ID embedding, . Thus, the visual branch can learn to keep and refine the visual embedding in the hierarchical propagation.
ID Branch is designed for propagating the object-specific information from past frames to the current frame. The prediction of object-specific segmentation is essential for VOS and can not be processed by the above object-agnostic visual propagation branch. Let denote the object-specific embeddings in our identification branch, the formulation of our object-specific ID propagation is,
2 Gated Propagation Module
Instead of using the LSTT block , which employs multi-head attention in propagation, we stack the hierarchical propagation based on the proposed Gated Propagation Module (GPM), which is designed based on more efficient single-head attention.
LSTT Block includes four parts, i.e., a long-term attention responsible for propagating information from the memorized frames (in ), a short-term attention responsible for propagating information from a spatial neighborhood in the previous () frame, a self-attention module for associating objects in the current () frame, and a feed-forward module. The three kinds of attention modules are built on the multi-head extension of Eq. 1 or Eq. 2. According to the experiments in Table 3(b), reducing the head number from multiple heads (8 heads in default) to a single head will decrease the performance of AOT but can significantly improve the run-time speed, which means the multi-head attention is an efficiency bottleneck of LSTT. Concretely, the computational complexity of long-term attention is , which is proportional to the head number since each head contains a correlation function, .
Gated Propagation Module consists of three kinds of gated propagation, self-propagation, long-term propagation, and short-term propagation. Compared with LSTT, GPM removes the feed-forward module for further saving computation and parameters. All the propagation processes employ the gated propagation function defined in Eq. 5. In DeAOT, both the propagation branches (i.e., visual branch and identification branch) are stacked by GPM as shown in Fig. 3(b).
Based on the formulation of visual propagation (Eq. 3) and ID propagation (Eq. 4), the Long-term Propagation can be formulated as
for the visual branch and ID branch, respectively. The ID propagation reuses the attention maps of the visual propagation as discussed in Eq. 4. Based on the long-term propagation, we can formulate the Short-term Propagation at spatial location to be
Finally, the Self-Propagation can also be formulated similar to the long-term propagation, i.e.,
where is a concatenation process on the channel dimension. In the self-propagations, both the visual embedding and ID embedding are used in the calculation of attention maps (i.e., ). Here, the object-specific performs like a positional embedding additional to the visual embedding . We found that such a process can help associate the objects in the current frame more effectively. Apart from this, the current frame segmentation is unavailable before being decoded and is not used in the ID self-propagation . For simplicity, we reuse the parameter symbols in Eq. 6 and 7, but the trainable parameters are not shared with long-term propagation.
Implementation Details
Network Details: Consistent with AOT , three kinds of encoders are used in our experiments, i.e., MobileNet-V2 (in default), ResNet-50 (R50) , and Swin-B . The decoder is the same FPN network. Besides, the spatial neighborhood size is set to 15, and the maximum object number within the ID embedding is 10. In our GPM module, the channel dimension of visual and ID embeddings is 256, the matching features’ dimension is 128, and the propagation features’ dimension is 512. Moreover, the kernel size of is 5, and the gating function is SiLU/Swish .
To make fair comparisons with AOT’s variants , we build corresponding DeAOT variants with different GPM number or long-term memory size . The hyper-parameters of these variants are: DeAOT-T: , ; DeAOT-S: , ; DeAOT-B: , ; DeAOT-L: , . DeAOT-T/S/B considers only the reference frame as the long-term memory, leading to consistent run-time speeds. DeAOT-L updates the long-term memory per (set to 2/5 for training/testing) frames as AOT-L .
Training Details: Following , we first pre-train DeAOT on synthetic video sequence generated from static image datasets by randomly applying multiple image augmentations . Then, we do main training on the VOS benchmarks by randomly applying video augmentations . Besides, we keep our optimization strategies and related hyper-parameters the same as AOT. More details are supplied in Supplementary.
Experimental Results
We conduct experiments on three popular VOS benchmarks (YouTube-VOS , DAVIS 2017 , and DAVIS 2016 ) and one challenging Visual Object Tracking (VOT) benchmark (VOT 2020 ), which gives segmentation annotations and can be used to evaluate VOS algorithms.
To validate DeAOT’s generalization ability, all the benchmarks share the same model parameters. When evaluating YouTube-VOS, we use the default 6fps videos, which are restricted to be smaller than resolution. On DAVIS, the default 480p 24fps videos are used. For evaluating VOT 2020, more details can be found in the supplementary material.
The evaluation metrics for VOS benchmarks include the score (calculated as the average IoU score between the prediction and the ground truth mask), the score (calculated as an average boundary similarity measure between the boundary of the prediction and the ground truth), and their mean value (denoted as &). As to VOT 2020, we use the official EAO criteria . We evaluate all the results on official evaluation servers or with official tools.
YouTube-VOS is a large-scale multi-object VOS benchmark, which contains 3471 videos in the training split with 65 categories and 474/507 videos in the Validation 2018/2019 split with additional 26 unseen categories. Table 1 shows that DeAOT variants remarkably outperforms AOT counterparts in both accuracy and run-time speed on YouTube-VOS 2018/2019. For example, our R50-DeAOT-L achieves 86.0%/85.9% (&) at 22.4fps, which is superior compared to R50-AOT-L (84.1%/84.1% at 14.9fps). Particularly, our SwinB-DeAOT-L achieves new state-of-the-art performance (86.2%/86.1%), surpassing previous methods by more than 1.7%/1.6%. In addition, our smallest variant, DeAOT-T, precedes SST (82.0%/82.0% vs 81.7%/81.8%) and runs about 15 faster than CFBI (53.4fps vs 3.4fps).
DAVIS 2017 is a multi-object extension of DAVIS 2016. The training/validation split consists of 60/30 videos with 138/59 objects, and the test split contains 30 more challenging videos with 89 objects. As shown in Table 1, DeAOT variants can generalize to DAVIS 2017 well. R50-DeAOT-L achieves 85.2%/80.7% on the validation/test split at a real-time speed (27fps), surpassing R50-AOT-L in accuracy and efficiency. Also, SwinB-DeAOT-L achieves the top-ranked performance on DAVIS 2017 (86.2%/82.8%).
DAVIS 2016 is a single-object benchmark containing 20 videos in the validation split, and we show related experiments in Table 2. Although AOT-like methods focus on multi-object scenarios, our DeAOT-L is faster and more robust than STCN , whose architecture was designed for single-object VOS. Besides, SwinB-DeAOT-L achieves 92.9% and outperforms all the VOS methods as well.
VOT 2020 consists of 60 single-object videos with challenging scenarios including fast motion, occlusion, etc. The average frame number of VOT 2020 is 327, which is much longer than the maximum video length of the above VOS benchmarks. DeAOT shows superior performance on VOT 2020 in Table 2. The DeAOT variants larger than DeAOT-T outperform MixFormer-L (the state-of-the-art tracker), RPT (VOT 2020 short-term challenge winner), and AlphaRef (VOT 2020 real-time challenge winner) in both EAO and real-time EAO scores. Specifically, SwinB-DeAOT-L achieves 0.622 EAO, outstandingly exceeding MixFormer-L by 0.067, and R50-DeAOT-L achieves 0.571 EAO under a real-time requirement, impressively overtaking AlphaRef by 0.085.
Qualitative results: Fig. 4 give qualitative comparisons to AOT. By introducing the dual-branch propagation, R50-DeAOT-L performs better than R50-AOT-L on tiny or scale-changing objects (ski poles or ski board). Nevertheless, R50-DeAOT-L still may fails to track multiple highly similar objects (dancer and cow) when serious occlusion happens.
2 Ablation Study
This section analyzes the necessity of dual-branch propagation and GPM of DeAOT in Table 3.
Propagation module: Table 3(a) shows that the performance of DeAOT drops from 82.5% to 81.5% by coupling the propagation of visual and ID embeddings (w/o De) like AOT. Furthermore, doubling the channel dimensions only partially relieves the performance loss. Moreover, the performance will be seriously degraded to 80.3% by replacing our GPM with the LSTT module of AOT. In conclusion, the dual-branch propagation approach and the GPM module are crucial in improving VOS performance.
Head number: According to the results in Table 3(b), the head number () of attention-based modules is negatively correlated with the efficiency of AOT/De-AOT. The single-head AOT (44.6fps) runs much faster than the default AOT (=8, 27.1fps) but loses 0.7% accuracy. By contrast, DeAOT is robust to the head number by using our proposed GPM module.
Attention map: Our DeAOT shares the attention maps between two propagation branches. Table 3(c) shows the study of different kinds of attention maps. Concretely, visual embeddings are essential in building attention maps in the long-term/short-term propagation, whose attention maps are used to match objects. Introducing ID embeddings does not help learn better visual embeddings and will decrease the performance (82.5% vs 82.1%). In the self-propagation, however, utilizing the ID embedding as a positional embedding will facilitate the association of objects (82.2% vs 82.5%) in the current frame.
Kernel size of : Large receptive fields have been proved to be critical in segmentation-related tasks . The depth-wise convolution, , is an important part of GPM for enlarging the receptive fields. Without , the performance of DeAOT drops from 82.5% to 81.1%, as shown in Table 3(d). We empirically found the best kernel size of is 5 among .
Conclusion
This paper proposes a highly effective and efficient framework, Decoupling Features in Hierarchical Propagation (DeAOT), for video object segmentation. Based on the rethinking of AOT-like hierarchical propagation, we propose to decouple the propagation of visual and ID embeddings into two network branches and thus avoid the loss of visual information in deep propagation layers. Besides, we propose the Gated Propagation Module (GPM), an efficient module for constructing hierarchical VOS propagation. Applying GPM to the dual-branch propagation, our DeAOT variant networks achieve new state-of-the-art performance on four VOS/VOT benchmarks with superior run-time speed compared to previous solutions.
Acknowledgements. This work is partly supported by the Fundamental Research Funds for the Central Universities (No. 226-2022-00051).