SwiftNet: Real-time Video Object Segmentation
Haochen Wang, Xiaolong Jiang, Haibing Ren, Yao Hu, Song Bai
Introduction
Given the first frame annotation, semi-supervised video object segmentation (one-shot VOS) localizes the annotated object(s) on pixel-level throughout the video. One-shot VOS generally adopts a matching-based strategy, where target objects are first modeled from historical reference frames, then precisely matched against the incoming query frame for localization. Being a video-based task, VOS finds vast applications in surveillance, video editing, and mobile visions, most of which ask for real-time processing .
Nonetheless, although pursued in fruitful endeavors , real-time VOS remains unsolved, as object variation over-time poses heavy demands for sophisticated object modeling and matching computations. As a compromise, most existing methods solely focus on improving segmentation accuracy while at the expense of speed. Amongst, memory-based methods reveal exceptional accuracy with comprehensively modeling object variations using all historical frames and expressive non-local reference-query matching. Unfortunately, deploying more reference frames and complicated matching scheme inevitably slow down segmentation. Accordingly, recent attempts seek to accelerate VOS with reduced reference frames and light-weight matching scheme . For the first aspect, solutions proposed in follow a mask-propagation strategy, where only the first and last historical frames are considered reference for current segmentation. For the second aspect, efficient pixel-wise matching , region-wise distance measuring , and correlation filtering are deployed to reduce computations. However, as shown in Fig.1, although these accelerated methods enjoy faster segmentation speed, they still barely meet real-time requirement, and more critically, they are far from state-of-the-art segmentation accuracy.
We argue that, the accurate solutions are less efficient due to the spatiotemporal redundancy inherently resides in matching-based VOS, and the fast solutions suffer degraded accuracy for reducing the redundancy indiscriminately. Considering its pixel-wise modeling, matching, and estimating nature, matching-based VOS manifests positive correlation between processing time and multiplied number of pixels and in reference and query frames as described in Eqa.1. The spatiotemporal redundancy denotes that is populated with pixels not beneficial for accurate segmentation. Temporally, existing methods carelessly involve all historical frames (mostly by periodic sampling) for reference modeling, resulting in the fact that static frames showing no object evolution are repeatedly modeled, while dynamic frames containing incremental object information are less attended. Spatially, full-frame modeling and matching are adopted as defaults , wherein most static pixels are redundant for segmentation. Illustration in Fig. 2 vividly advocates the above point. From this standpoint, explicitly compressing pixel-wise spatiotemporal redundancy is the best way to yield accurate and fast one-shot VOS.
Accordingly, we propose SwiftNet for real-time one-shot video object segmentation. Overall, as depicted in Fig.3, SwiftNet instantiates matching-based segmentation with an encoder-decoder architecture, where spatiotemporal redundancy is compressed within the proposed Pixel-Adaptive Memory (PAM) component. Temporally, instead of involving all historical frames indiscriminately as reference, PAM introduces a variation-aware trigger module, which computes inter-frame difference to adaptively activate memory update on temporally-varied frames while overlook the static ones. Spatially, we abolish full-frame operations and design pixel-wise update and match modules in PAM. For pixel-wise memory update, we explicitly evaluate inter-frame pixel similarity to identify a subset of pixels beneficial for memory, and incrementally add their feature representation into the memory while bypassing the redundant ones. For pixel-wise memory match, we compress the time-consuming non-local computation to accommodate the pixel-wise memory as reference, thus achieving efficient matching without degradation of accuracy. To further accelerate segmentation, PAM is equipped with a novel light-aggregation encoder (LAE), which eschews redundant feature extraction and enables multi-scale mask-frame aggregation leveraging reversed sub-pixel operations.
In summary, we highlight three main contributions:
We propose SwiftNet to set the new record in overall segmentation accuracy and speed, thus providing a strong baseline for real-time VOS with publicized source code.
We pinpoint spatiotemporal redundancy as the Achilles heel of real-time VOS, and resolve it with Pixel-Adaptive Memory (PAM) composing variation-aware trigger and pixel-wise update & matching. Light-Aggregation Encoder (LAE) is also introduced for efficient and thorough reference encoding.
We conduct extensive experiments deploying various backbones on DAVIS 2016 & 2017 and YouTube-VOS datasets, reaching the best overall segmentation accuracy and speed performance at 77.8% & and 70 FPS on DAVIS2017 validation set.
Related Work
One-shot VOS establishes a spatiotemporal matching problem, such that objects annotated in the first frame are localized in upcoming query frames by searching pixels best-matched to object template modeled in the reference frames. From this perspective, we categorize one-shot VOS methods w.r.t different reference modeling and reference-query matching strategies. Reference modeling builds object template by exploiting object evolution in historical frames, and methods either follows the last-frame or all-frame approaches. For the former one, utilize only the first and/or last frame as reference, demonstrate favorable segmentation speed but suffer uncompetitive accuracy due to inadequate modeling over object variation. For the latter one, methods proposed in leverage all previous frames and reveal improved accuracy, but they suffer slower speed for heavy computation overhead even with periodic sampling.
Considering reference-query matching, we classify methods as two-stage and one-stage basing on whether matching with proposed region-of-interest. Similar to object detection, two-stage method is more accurate while one-stage leads in speed. Besides, the key of matching is similarity measuring, where convolutional networks , cross correlation , and non-local computation are widely adopted. Amongst, non-local reveals best accuracy for capturing all-pairs pixel-wise dependency but are computationally heavy. In addition to matching-based VOS, propagation-based methods leverage temporal motion consistency to reinforce segmentation, which is highly effective when appearance matching fails due to severe variations. Additionally, time-consuming online fine-tuning are exploited in to improve segmentation accuracy, which however is impractical for real-time application.
2 Fast VOS
For efficiency, most fast VOS solutions deploy the single-frame reference strategy . Besides, methods proposed in employ segmentation-by-tracking where pixel-wise estimation is gated within tracked bounding-boxes to avoid full-frame estimation. To expedite time-consuming pixel-wise matching, RGMP computes similarity responses with convolutions; AGAME discriminates object from background with a probabilistic generative appearance model; RANet adopts cross-correlation on ranked pixel-wise features to match query with reference. In addition, OSNM proposes to spur VOS with network modulation.
3 Memory-based VOS
Memory-based VOS exploits all historical frames in an external memory for object modeling, an alternative approach for modeling all-frame evolution is via the implementation of recurrent neural networks . First proposed in , STM is the seminal memory-based method which boosts segmentation accuracy by a large margin. As follows, modify STM by introducing Siamese-based semantic similarity and motion-guided attention. To induce heavy computations, GCNet designs a global context module using attentions to reduce temporal complexity executed in the memory.
SwiftNet
In this section, we present SwiftNet by first briefly formulating the problem of matching-based one-shot VOS. As follows, PAM is discussed in details, including variation-aware trigger as well as pixel-wise memory update and match modules. LAE is explained afterwards.
Given a video sequence , its first frame is annotated with mask . The goal of one-shot VOS is to delineate objects from the background by generating mask for each frame . Particularly, matching-based VOS computes mask via object modeling and matching.
For object modeling at frame , historical information embedded in reference frames and is exploited to establish object model for up till frame :
here is an indicator function denoting whether frame involves in modeling, indicates reference encoder for feature extraction, and generalizes the object modeling process.
For reference-query matching, the task is to search within on a pixel-level and generate object affinity map :
Here denotes pixel-wise matching and refers to the query encoder. The final segmentation mask is produced by a decoder integrating encoded features and .
At test time with SwiftNet, upon the arrival of query frame , it is first processed by the query encoder and then passed into the pixel-wise memory match module. The matching output and encoded query features are aggregated in the decoder to generate mask . Subsequently, , , and are jointly fed into the variation-aware trigger module, and if triggered, they are then handled by LAE for later pixel-wise memory update. This overall workflow is illustrated in Fig.3.
2 Pixel-Adaptive Memory
As the core component of SwiftNet, PAM models object evolution and performs object matching with explicitly compressed spatiotemporal redundancy. PAM mainly composes the variation-aware trigger as well as the pixel-wise memory update and match modules.
Instead of utilizing merely the first and last frames for object modeling, incorporating all historical frames as reference help establish temporally-coherent object evolution . Nonetheless, this approach is rather impractical considering its prohibitive temporal redundancy and computation overhead. As a straightforward solution, previous methods sample historical frames at a predefined pace , which indiscriminately reduces temporal redundancy and leads to accuracy degradation.
To explicitly compress temporal redundancy, variation-aware trigger module evaluates inter-frame variation frame-by-frame, and activates memory update once the accumulated variation surpass threshold . Specifically, given , and and , we separately compute image difference and mask difference as:
at each pixel we update the overall running variation degree as:
Once exceeds , PAM triggers a new round of memory update as described in 3.2.2. Empirically, , and equal 200, 1, 0 respectively yields best performance.
2.2 Pixel-wise Memory Update
In terms of matching-based VOS, memory infers a temporally-maintained template which characterizes object evolution over time. In the existing literature, memory update and matching typically adopt full-frame operations, where reference frames are concatenated into memory and matched with query frame intactly . This strategy induces heavy storage and computation overhead, as redundant pixels showing no benefit for object modeling are incorporated without discrimination.
for each row in , we find the largest score as the feature similarity between pixel in query frame and in pixel in the memory. In formulation, we compute pixel similarity vector as:
we sort in increasing order of similarity (original index is kept), then the select top percents pixels for memory update. These set of pixels exhibit most severe feature variations. Here is a hyper-parameter controlling the balance between method efficiency and update comprehensiveness, and is experimentally set to 10% for the best performance. To execute the memory update, we find feature vectors of the selected set of pixels from and according to indexes as in , then directly add them into memory which is instantiated as an array of feature vectors.
2.3 Pixel-wise Memory Match
As illustrated in Fig. 3, segmentation mask is decoded utilizing query value and output of reference-query matching, which provides strong spatial prior w.r.t. the foreground object. In essence, this matching process computes similarity between pixels from the reference and query frames, and can be instantiated with cross-correlation , neural networks , distance measuring , and non-local computation , etc. Comparatively, non-local leads to excellent accuracy performance but suffers heavy computation expenses in the context of full-frame operations. In PAM, we implement pixel-wise matching to achieve efficient and accurate segmentation.
3 Light-Aggregation Encoder
In existing memory-based VOS solutions , the image encoding process is time-consuming as both query and reference encoders adopt heavy backbone networks. In SwiftNet we expedite encoding by sparing the backbone resides in . Particularly, we instantiate with ResNet-based networks, and after is encoded by , we buffer the generated feature maps. If frame is triggered for update, these buffered features are directly utilized in for reference encoding. Efficiency comparison w.r.t. encoding strategies are listed in Table. 1.
To facilitate the described encoding process, we design the novel light-aggregation encoder as shown in Fig. 5. The upper blue entities represent buffered feature maps encoded by , the bottom green ones show feature transformation hierarchy of the input mask. Features aligned vertically in the same column are with the same size and concatenated together to facilitate multi-scale aggregation. In particular, to instantiate feature transformation of the input mask, we implement reversed sub-pixel for down-samplings and convolutions for channel manipulation. Reversed sub-pixel technique is motivated by the popular up-sampling method in super-resolution , which shrinks spatial dimension of features without information loss.
Experiments
In this section we first discuss implementation details of the experiments, then elaborate the ablation study specifying contributions of different components proposed in SwiftNet. Comparisons with other state-of-the-art methods on DAVIS 2016 & 2017 and YouTube-VOS datasets are provided as follows, where SwiftNet demonstrates the best overall segmentation accuracy and inference speed. All experiments are implemented in PyTorch on 1 NVIDIA P100 GPU. Particularly, SwiftNet adopting both ResNet-18 and ResNet-50 backbones are experimented to show the favorable compatibility and efficacy of our method. We employs a decoder similarly constructed as in STM , which adopts three refinement modules to gradually restore the spatial scale of the segmentation mask.
DAVIS 2016 & 2017. DAVIS 2016 dataset contains in total 50 single-object videos with 3455 annotated frames. Considering its confined size and generalizability, it is soon supplemented into DAVIS 2017 dataset comprising 150 sequences with 10459 annotated frames, a subset of which exhibit multiple objects. Following the DAVIS standard, we utilize mean Jaccard index and mean boundary score, along with mean & to evaluate segmentation accuracy. We adopt the Frames-Per-Second (FPS) metric to measure segmentation speed.
YouTube-VOS. Being the largest dataset at the present, YouTube-VOS encompasses totally 4453 videos annotated with multiple objects. In particular, its validation set possesses 474 sequences covering 91 object classes, 26 of which are not visible in the training set, and thus facilitating evaluations w.r.t. seen and unseen object classes to reflect method generalizability. On YouTube-VOS we report & for accuracy assessment, the overall score is generated by averaging & on seen and unseen classes.
2 Training and Inference
SwiftNet is first pre-trained on simulated data generated upon MS-COCO dataset , then finetuned on DAVIS 2017 and YouTube-VOS Dataset respectively. In both training stages, input image size is set to 384 384, and we adopt Adam optimizer with learning rate starting at 1e-5. The learning rate is adjusted with polynomial scheduling using the power of 0.9. All batch normalization layers in the backbone are fixed at its ImageNet pre-trained value during training. We use batch size of 4, which is realized on 1 GPU via manual accumulation.
MS-COCO Pre-train. Considering the scarcity of video data and to ensure the generalizability of SwiftNet, we perform pre-training on simulated video clips generated upon MS-COCO dataset . Specifically, we randomly crop foreground objects from a static image, which are then pasted onto a randomly sampled background image to form a simulated image. Affine transformations such as rotation, resizing, sheering, and translation are applied to foreground and background separately to generate deformation and occlusion, and we maintain an implicit motion model to generate clips with length of 5. SwiftNet is trained with simulated clips for 150000 iterations and the & reaches 65.6 on DAVIS 2017 validation set, which demonstrates the efficacy of our simulated pre-training.
DAVIS 2017 YouTube-VOS Finetune. After pre-training, we finetune SwiftNet on DAVIS 2017 and YouTube-VOS training set for 200000 iterations. At each iteration, we randomly sampled 5 images consecutively (with random skipping step smaller than 5 frames) and estimate corresponding segmentation masks one after another. Pixel-wise memory update and match are executed on every frame within the 5-frame clip.
2.2 Inference
Given a test video accompanied by its first frame annotation mask, at inference time we frame-by-frame segment the video using SwiftNet. Particularly, memory at the first frame, , is initialized with feature maps output by the encoder given first frame image and mask, then it is updated online throughout the inference. At frame , we utilize memory and frame image to compute segmentation mask with SwiftNet. If frame is triggerd, is feed into the LAE and to update the memory for further computations.
3 Ablation Study
Ablation study is conducted on DAVIS 2017 validation set to show contributions of different SwiftNet modules.
To demonstrate the efficacy of the proposed LAE, we additionally develop two baseline reference encoders for comparison. The first baseline instantiates low-level aggregation as adopted in STM , where mask produced by the last frame is directly concatenated with raw image. This encoder maintains high-resolution mask but requires two separate encoders for reference and query frames, hence heavier model size. The second baseline implements high-level aggregation following CFBI , where segmentation mask is first down-sampled to the minimal feature resolution and then fused with query feature for foreground discovery. This baseline enables encoder reuse between reference query frames, but spatial details of the mask are lost during pooling-based down-samplings. As shown in Table 1, low-level baseline reveals better accuracy while high-level baseline runs faster, conforming to the fact that the low-level one involves more sophisticated feature aggregations between image and mask. Notably, LAE surpasses the low-level baseline in both & and FPS (by 19), and outperforms the high-level baseline by 4.2% in & while keeping comparable FPS. This results strongly suggest that LAE promotes thorough mask-frame aggregation and elevates segmentation speed.
3.2 Pixel-Adaptive Memory
In this section we showcase the efficacy of PAM in elevating accuracy and speed. Table 2 row-wise illustrates the contribution of pixel-wise memory update and match in eliminating spatial redundancy, where it significantly boosts processing speed by 30 and 18 PFS in both temporal strategies, and only experience around 0.4% drop in &. Column-wise reveals the contribution of variation-aware trigger in compressing temporal redundancy, where it raises segmentation speed by 17 and 5 FPS in both spatial strategies, and at most 0.1% & is reported. Notably here we experiment with periodic sampling at a pace of 5 frame, which is tested to be the optimal parameter as in . To provide a up-closer view of PAM, in Fig. 7 we illustrate the variation of & and FPS w.r.t. different spatial update ratio and temporal trigger strategy. As shown in green, & increases in accordance with enlarged , i.e. segmentation accuracy will grow if more percentage of pixels are updated. It is worth noting that, yields the best accuracy while larger value shows no significant improvement. Besides, temporal trigger brings minute effect in accuracy. The blue color draws variations w.r.t. FPS, where larger steadily decreases FPS, and variation-aware trigger constantly increases FPS in under different . Notably, the gap between blue curves are enlarged with larger , echos that more spatiotemporal redundancy are compressed by the trigger during heavy spatial update.
4 State-of-the-art Comparison
Comparison results on Davis 2017 validation set are listed in Table 3. As shown, both SwiftNet versions demonstrate better &, , and scores than all other real-time methods by a large margin. In particular, SwiftNet with ResNet-18 runs the fastest at 70 FPS, outperforming the second fastest SAT-fast in & by 8.3%. This considerable lead is because that SAT updates global feature with cropped regions containing heavy background noise, while SwiftNet updates memory with useful and discriminative pixels and filters out redundant and noise regions. SwiftNet with ResNet-50 not only meets real-time requirement, but also reaches 81.1 in & score, which ranks the second best in both real-time and slow methods. STM reports the best & at 81.8, which is 0.7% better than ours, while we run almost 4 times faster than STM. This significant improvement of SwiftNet is achieved by explicitly compressing spatiotemporal redundancy resides in STM, which adopts heavy periodical sampling and full-frame matching. In addition, GCNet also strives to accelerate memory-based VOS by designing light-weight memory reading and writing strategies. As shown, it runs at comparable speed with our ResNet-50 version, while we exceeds GCNet in term of & by 9.7%. Fig 6 shows qualitative results on DAVIS17 validation set produced by SwiftNet with ResNet-18. The first row demonstrates that SwiftNet is robust against deformation, the second to the fourth row reveal that SwiftNet is highly capable of handling fast motion, similar distractor, and tremendous occlusion, respectively.
4.2 DAVIS 2016
Results on DAVIS 2016 dataset is shown in Table. 4. Since DAVIS 2016 only contains single-object sequences, most methods experience considerable performance gains when transferred from DAVIS 2017, and the accuracy gap between ResNet-18 and ResNet-50 SwiftNet is reduced because the demands for highly semantical features are alleviated. It is worth noting that, SwiftNet with both ResNet-18 and ResNet-50 outperform all other methods in segmentation accuracy, where the ResNet-50 version leads the second best STM by 1.1% and 18.7 in terms of & and FPS.
4.3 Youtube-VOS
As testing on the large YouTube-VOS validation set is time-consuming, here we show comparison results with most representative methods. Besides, considering the varied formulation of test sequences, we omit FPS readings on YouTube-VOS by default. As shown in Table 5, SwiftNet with ResNet-50 considerably outperform all other real-time methods in accuracy, leading the second best GCNet by 4.6% in term of overall score , while running at a comparable speed with GCNet. SwiftNet with ResNet-18 performs comparably with GCNet, but runs significantly faster. Moreover, SwiftNet performs stably across seen and unseen classes, demonstrating its favorable generalizability.
Conclusion
We propose a real-time semi-supervised video object segmentation (VOS) solution, named SwiftNet, which delivers the best overall accuracy and speed performance. SwiftNet achieves real-time segmentation by explicitly compressing spatiotemporal redundancy via Pixel-Adaptive Memory (PAM). In PAM, temporal redundancy is reduced using variation-aware trigger, which adaptively selects incremental frames for memory update and ignores static ones. Spatial redundancy is eliminated with pixel-wise memory update and match modules, which abandon full-frame operations and incrementally process with temporally-varied pixels. Light-aggregation encoder is also introduced to promote thorough and expedite reference encoding. Overall, SwiftNet is effective and compatible, we hope it could set a strong baseline for real-time VOS.