Efficient Video Object Segmentation via Network Modulation

Linjie Yang, Yanran Wang, Xuehan Xiong, Jianchao Yang, Aggelos K. Katsaggelos

Introduction

Semantic segmentation plays an important role in understanding visual content of an image as it assigns pre-defined object or scene labels to each pixel and thus translates the image into a segmentation map. When dealing with video content, a human can easily segment an object in the whole video without knowing its semantic meaning, which inspired a research topic named semi-supervised video segmentation. In a typical scenario of semi-supervised video segmentation, one is given the first frame of a video along with an annotated object mask, and the task is to accurately locate the object in all following frames Perazzi2016davis; Li2013video. The ability of performing accurate pixel-level video segmentation with minimum supervision (e.g., one annotated frame) can foster a large amount of applications, such as accurate object tracking for video understanding, interactive video editing, augmented reality, and video-based advertisement. When the supervision is limited to only one annotated frame, researchers refer to this scenario as one-shot learning. In the recent years, we have witnessed a rising amount of interests in developing one-shot learning techniques for video segmentation Caelles2017osvos; Perazzi2017masktrack; Tsai2016objflow; Marki2016bilateral; Shin2017pixel; Cheng2017segflow. Most of these work share a similar two-stage paradigm: first, train a general-purpose Fully Convolutional Network (FCN) shelhamer2017fully to segment the foreground object; Second, fine-tune this network based on the first frame of the video for several hundred forward-backward iterations to adapt the model to the specific video sequence. Despite the high accuracies achieved by these approaches, the fine-tuning process is arguably time consuming, which makes it prohibited for real-time applications. Some of these approaches Cheng2017segflow Perazzi2017masktrack also utilize optical flow information, which is computationally heavy for state-of-the-art algorithmsRevaud2015epicflow Ilg2017flownet.

In order to alleviate the computational cost of semi-supervised segmentation, we propose a novel approach to adapt the generic segmentation network to the appearance of a specific object instance in one single feed-forward pass. We propose to employ another meta neural network called modulator to learn to adjust the intermediate layers of the generic segmentation network given an arbitrary target object instance. Fig. 1 shows an illustration of our approach. By extracting information from the image of the annotated object and the spatial prior of the object, the modulator produces a list of parameters, which are injected into the segmentation model for layer-wise feature manipulation. Without one-shot fine-tuning, our model is able to change the behavior of the segmentation network with minimum extracted information from the target object. We name this process network modulation.

Our proposed model is efficient, requiring only one forward pass from the modulator to produce all parameters needed for the segmentation model to adapt to the specific object instance. Network modulation guided by the spatial prior facilitates the model to track the object even with the presence of multiple similar instances. The whole pipeline is differentiable and can be learned end-to-end using the standard stochastic gradient descent. The experiments show that our approach outperforms previous approaches without one-shot fine-tuning by a large margin, and achieves comparable performance with these approaches after one-shot fine-tuning with a 70×\times speed up.

Related Work

Semi-supervised video object segmentation aims at tracking an object mask given from the first annotated frame throughout the rest of video. Many approaches have been proposed in the literature, including those propagating superpixels Jain2014youtube Tsai2016objflow, patches jumpcut, object proposals objproposals, or in bilateral space Marki2016bilateral, and graphical model based optimization is usually performed to consider multiple frames simultaneously. With the success of FCN on static image segmentation Hariharan2015hypercolumns, deep learning based methods Perazzi2017masktrack; Caelles2017osvos; Shin2017pixel; Tokmakov2017memory; Jampani2017vpn; Cheng2017segflow have been recently proposed for video segmentation and promising results have been achieved. To model the temporal motion information, some works heavily rely on optical flow Tokmakov2017memory Cheng2017segflow, and use CNNs to learn mask refinement of an object from current frame to the next one Perazzi2017masktrack, or combine the training of CNN with bilateral filtering between adjacent frames Jampani2017vpn. Chen et al. Cheng2017segflow use a CNN to jointly estimate the optical flow and provide the learned motion representation to generate motion consistent segmentation across time. Different from these approaches, Caelles et al. Caelles2017osvos combine offline and online training process on static images without using temporal information. While it saves the computation of optical flow and/or conditional random fields (CRF) Krahenbuhl2011crf involved in some previous methods, online fine-tuning still requires many iterations of optimization, which poses a challenge for real-world applications that need rapid inference.

Current success of deep learning relies on the ability of learning from large-scale labeled datasets through gradient descent optimization. However, if we want our model to learn many tasks adapted to many environments, it is not affordable to learn each task for each setting from scratch. Instead, we want our deep learning system to be able to learn new tasks very fast and from very limited quantities of data. In the extreme of “one-shot learning”, the algorithm needs to learn the new task with a single observation. One potential strategy for learning a versatile model is the notion of meta-learning, or learning to learn, which can date back to the late 80s. Recently, meta-learning has become a hot research topic with publications on neural network optimization Chen2017LearningToLearn, finding good network architectures, fast reinforcement learning, and few-shot image recognition Vinyals2016Matching; Ravi2016optimization; Hariharan2017ICCV; Finn2017ICML; Santoro2016ICML. Ravi and Larochelle Ravi2016optimization proposed a LSTM meta-learner to learn the update rules for few shot learning. The meta optimization over a large number of tasks in Finn2017ICML targets at learning a model that can quickly adapt to the new task with limited number of updates. Hariharan and Girschick Hariharan2017ICCV trained a learner that generated new samples and used new samples for training new tasks. Our approach shares the similarity with meta-learning that it learns to update the segmentation model rapidly with another meta learner, i.e. the modulator.

Several previous work try to incorporate modules to manipulate the behavior of a deep neural network, either to manipulate spatial arrangement of data Jaderberg2015spatial or filter connections Dai2017deformable. Our method is also heavily motivated by conditional batch normalization Dumoulin2017ICLR; Ghiasi2017BMVC; Huang2017ArbitraryST; Perez2017LearningVR, where the behavior of the deep model is manipulated by batch normalization parameters conditioned on a guidance input, e.g. a style image for image stylization or a language sentence for visual question answering.

Video Object Segmentation with Network Modulation

In our proposed framework, we utilize modulators to instantly adapt the segmentation network to a specific object, rather than performing hundreds of iterations of gradient descent. We can achieve similar accuracy by adjusting a limited number of parameters in the segmentation network, compared with the updating the whole network in one-shot learning approaches Perazzi2017masktrack; Caelles2017osvos. There are two important cues for video object segmentation: visual appearance and continuous motion in space. To use information from both visual and spatial domains, we incorporate two network modulators, namely visual modulator and spatial modulator, to learn to adjust intermediate layers in the main segmentation network, based on the annotated first frame and spatial location of the object, respectively.

Our approach is inspired by recent works using Conditional Batch Normalization (CBN) Vries2017ModulatingEV; Huang2017ArbitraryST; Perez2017LearningVR, where the scale and bias parameters of each batch-normalization layer are produced by a second controller network. These parameters are used to control the behavior of the main network for tasks such as image stylization and question answering. Mathematically, each CBN layer can be formulated as follows:

where xc\bm{x}_{c} and yc\bm{y}_{c} are the input and output feature maps in the cthc_{th} channel, and γc\gamma_{c} and βc\beta_{c} are the scale and bias parameters produced by the controller network, respectively. The mean and variance parameters are omitted for clarity.

2 Visual and spatial modulation

The CBN layer is a special case of the more general scale-and-shift operation on feature maps. Following each convolution layer, we define a new modulation layer with parameters generated by both visual and spatial modulators that are jointly trained. We design the two modulators such that the visual modulator produces channel-wise scale parameters to adjust the weights of different channels in the feature maps, while the spatial modulator generates element-wise bias parameters to inject spatial prior to the modulated features. Specifically, our modulation layer can be formulated as follows:

where γc\gamma_{c} and βc\bm{\beta}_{c} are modulation parameters from the visual and spatial modulators, respectively. γc\gamma_{c} is a scalar for channel-wise weighting, while βc\bm{\beta}_{c} is a two-dimensional matrix to apply point-wise bias values.

Fig. 2 shows an illustration of the proposed approach, which consists of three networks: a fully-convolutional main segmentation network, a visual modulator network, and a spatial modulator network. The visual modulator network is a CNN that takes the annotated visual object image as input and produces a vector of scale parameters for all modulation layers, while the spatial modulator network is a very efficient network that produces bias parameters based on the spatial prior input. We will discuss the two modulators in more detail in the following sections.

3 Visual modulator

The visual modulator is used to adapt the segmentation network to focus on a specific object instance, which is the annotated object in the first frame. The annotated object is referred to as visual guide hereafter for convenience. The visual modulator extracts semantic information such as category, color, shape, and texture, from the visual guide and generates corresponding channel-wise weights so as to re-target the segmentation network to segment the object. We use VGG16 simonyan2014very neural network as the model for the visual modulator. We modify its last layer trained for ImageNet classification to match the number of parameters in the modulation layers for the segmentation network.

The visual modulator implicitly learns an embedding of different types of objects. It should produce similar parameters to adjust the segmentation network for similar objects while different parameters for different objects. This is indeed true as we show in Sec. 4.2 that the embedding of the modulator outputs correlates with object appearance very well. One big advantage of using such a visual modulator is that we can potentially transfer the knowledge learned with a large number of object classes, e.g., ImageNet, in order to learn a good embedding.

4 Spatial modulator

Our spatial modulator takes a prior location of the object in the image as input. Since objects move continuously in a video, we set the prior to be the predicted location of the object mask in the previous frame. Specifically, we encode the location information as a heatmap with a two-dimensional Gaussian distribution on the image plane. The center and standard deviations of the Gaussian distribution are computed from the predicted mask of the previous frame. This heatmap is referred as spatial guide hereafter for convenience. The spatial modulator downsamples the spatial guide into different scales, to match the resolution of different feature maps in the segmentation network, and then applies a scale-and-shift operation on each downsampled heatmap to generate the bias parameters of the corresponding modulation layer. Mathematically,

Our method shares some similarities with the previous work MaskTrack Perazzi2017masktrack in utilizing information from the previous mask. Comparing with their approach that uses the exact foreground mask of the previous frame, we only use a very coarse location prior. It may seem that our method throws away more information from the previous frame. However, we argue that the rough position and size in the previous frame possess enough information to infer the object mask with the RGB image, and it prevents the model from relying too much on the mask and as a result the error propagation, which can be catastrophic when the object has large movements in the video. As a drawback of such over-utilization of the mask, MaskTrack has to apply plenty of well-engineered data augmentation to prevent over-fitting, while we only apply simple shift and scaling as augmentation.

5 Implementation details

Our FCN structure follows the one used by Caelles2017osvos, which is a VGG16 simonyan2014very model with a hyper-column structure Hariharan2015hypercolumns. Intuitively, we should add modulation layers after each convolution layer in the FCN. However, we found that adding modulation layers in-between the early convolution layers actually makes the model perform worse. One possible reason is that early layers extract low-level features that are very sensitive to the scale-and-shift operations introduced by the modulator. In our implementation, we add modulation operations to all convolution layers in VGG16 except the first four layers, which results in nine modulation layers.

Similar to MaskTrack Perazzi2017masktrack, we also utilize static images for training our model. Ideally, the visual modulator should learn a mapping from any object to modulation weights of different layers in a FCN, which requires the model to see all possible different objects. However, most video semantic segmentation datasets only contain a very limited number of categories. We tackle this challenge by using the largest public semantic segmentation dataset MS-COCO Lin2014mscoco, which has 80 object categories. We select objects that are larger than 3%3\% of the image size for training, resulting in a total number of 217,516217,516 objects. For preprocessing the input for the visual modulator, we first crop the object using the annotated mask, then set the background pixels to mean image values, and then resize the cropped image to a constant resolution of 224×224224\times 224. The object is also augmented with up to 10%10\% random scaling and 10∘10^{\circ} random rotation. For preprocessing the spatial guide as input to the spatial modulator, we first compute the mean and standard deviation of the mask, and then augment the mask with up to 20%20\% random shift and 40%40\% random scaling. For the whole image fed into the FCN, we use a random size from 320320, 400400, and 480480 with a square shape.

The visual modulator and segmentation network are both initialized with VGG16 model pretrained on the ImageNet Deng2009imagenet classification task. The modulation parameters {γc}\{\gamma_{c}\} are initialized to ones by setting the weights and biases of the last fully-connected layer of the visual modulator to zeros and ones, respectively. The weights of spatial modulator are initialized randomly. We used the same balanced cross-entropy loss as in Caelles2017osvos. A mini-batch size of 88 is used. We use Adam optimizer with default momentum 0.90.9 and 0.9990.999 for β1\beta_{1} and β2\beta_{2}, respectively. The model is first trained for 1010 epochs with learning rate 10−510^{-5} and then trained for another 55 epochs with learning rate 10−610^{-6}.

Further, in order to model appearance variations of moving objects in videos, the model can be finetuned on video segmentation dataset such as DAVIS 2017 Pont-Tuset2017davis. To be more robust to appearance variations, we randomly pick a foreground object from the whole video sequence as the visual guide for each frame. The spatial guide is obtained from the ground truth mask of the object in the previous frame. The same data augmentations are applied as training on MS-COCO. The model is finetuned for 2020 epochs with learning rate 10−610^{-6}.

Experiments

In this section, we will introduce three parts of experiment: the comparison of our approach with previous methods, the visualization of the modulation parameters, and ablation study. Our model is trained on MS-COCO Lin2014mscoco 2017 dataset, and is tested on several popular video segmentation datasets, including DAVIS Perazzi2016davis Pont-Tuset2017davis and YoutubeObjects Jain2014youtube.

In this section, we compare with traditional approaches including OFL Tsai2016objflow, BVSMarki2016bilateral, and deep learning-based approaches including PLM Shin2017pixel, MaskTrack Perazzi2017masktrack, OSVOS Caelles2017osvos, VPN Jampani2017vpn, SFL Cheng2017segflow, and ConvGRU Tokmakov2017memory.

First, we compare our approach with previous approaches on DAVIS 2016 and YoutubeObjects. Some approaches (MaskTrackPerazzi2017masktrack, SFL Cheng2017segflow and OSVOSCaelles2017osvos) reported results both with and without model fine-tuning on the target sequences. We include both of them and denote the variants without fine-tuning as MaskTrack-B, SFL-B, and OSVOS-B, respectively. Our model has two variants,with the first only trained on static images (Stage 1) and the second finetuned on video data (Stage 1&2). Since there are several popular add-ons for this line of research, such as optical flow and CRF Krahenbuhl2011crf, which both have a lot of variants and make a fair comparison hard, we only include the performances without optical flow and CRF if possible, and mark those with add-ons in Table 1.

In Table 1, by comparing our method with OFL Tsai2016objflow, an expensive graphical model based approach, we achieve better accuracy on both DAVIS 2016 and YoutubeObjects. Comparing with deep learning approaches without model fine-tuning, and therefore, similar speed as ours, our method achieves the best accuracy on both DAVIS 2016 and YoutubeObjects. Comparing with the four approaches using model fine-tuning on target videos (PLM, MaskTrack, SFL, and OSVOS), our approach achieves better performance than PLM and MaskTrack, and is on-par with SFL. OSVOS achieves higher accuracy but it also utilizes a boundary snapping approach which contributes 2.4%2.4\% in mean IU. Our method is 70×70\times faster than MaskTrack and OSVOS, 50×50\times faster than SFL. We measure the running time of MaskTrack-B, OSVOS-B, and our method on a NVIDIA Quadro M6000 GPU using Tensorflow tensorflow2015. Speed of other methods are derived from the corresponding papers Speed of ConvGRU is estimated with the expensive optical flow they use, speed of PLM is derived through communication with the authors..

In our method, the adaptation of the segmentation model by the modulators is done with one forward pass for visual modulator, so it is much more efficient than the approaches with model fine-tuning on target videos. The visual modulator only needs to be computed once for the whole video, while the spatial modulator needs to be computed for every frame but the overhead is negligible, i.e., the average speed of our model on a video sequence is about the same as FCN itself. Our method is the second fastest of all compared methods, with only MaskTrack-B and OSVOS-B achieving similar speed but with much worse accuracies.

1.2 DAVIS 2017

To further investigate the capability of our model, we conduct more experiments on DAVIS 2017 Pont-Tuset2017davis, which is the largest video segmentation dataset to date. DAVIS 2017 is more challenging than DAVIS 2016 and YoutubeObjects in that it has multiple objects for each video sequences and some of the objects are very similar. We compare our method with two most related approaches, MaskTrack Perazzi2017masktrack and OSVOS Caelles2017osvos. For fair comparison, we only use their single network and adds-on free versions. We directly use open source code of OSVOS and adapt MaskTrack model to Tensorflow tensorflow2015. For each video sequence, OSVOS and MaskTrack are finetuned with 10001000 iterations. To show that network modulation is capable of adapting different model structures to specific object instances, we also experiment with modified OSVOS and MaskTrack models by adding a visual modulator to each of them, which are named OSVOS-M and MaskTrack-M respectively. For these two models, we only update the weights of the visual modulators and keep the weights of the segmentation model fixed in training.

Table 2 shows the results of different approaches on DAVIS 2017. We utilize the official evaluation metrics of DAVIS dataset: mean, recall, and decay of region similarity J\mathcal{J} and contour accuracy F\mathcal{F}, respectively. Note J\mathcal{J} mean is equivalent to mean IU we used above. Again, our model outperforms OSVOS-B and MaskTrack-B with a large margin, while obtaining comparable performance with the two methods with model fine-tuning. OSVOS-M and MaskTrack-M are both better than their baseline implementations with a 18%18\% and 9.3%9.3\% gain in J\mathcal{J} mean, respectively. Since the weights of the segmentation model are fixed, the accuracy gain comes solely from the modulator, which proves that the visual modulator is capable of improving different model structures by manipulating the scales of the intermediate feature maps. Noticeably, our method obtains much lower decay rate for both region similarity and contour accuracy compared to OSVOS and MaskTrack. The accuracy changes of the different methods over time are illustrated in Fig. 4. In the beginning of the video, our method lags behind OSVOS and MaskTrack. However, when it proceeds to around 40%40\% of the video, our method is on par with OSVOS and outperforms MaskTrack towards the end of the video. With one-shot fine-tuning, OSVOS and MaskTrack fit to the first frame very well. They are able to obtain high accuracy in the beginning of the video since these frames are all similar to the first one. But as time goes on and the object turns into different poses and appearances, it gets harder for the fine-tuned model to generalize to new object appearances. Our model is more robust to the appearance changes since it learns a feature embedding (see Section 4.2) for the annotated object which is more tolerant to pose and appearance changes compared to one-shot fine-tuning.

Some qualitative results of our methods compared with the two previous approaches are shown in Fig. 3. Compared with MaskTrack, our method generally obtains more accurate boundaries, partially due to that the coarse spatial prior forces the model to explore more cues on the image rather than the mask in the previous frame. Compared with OSVOS, our method shows better results when there are multiple similar objects in the image, thanks to the tracking capability provided by the spatial modulator. On the other hand, our method is also shown to work well on unseen object categories in training data. In Fig. 3, the camel and the pigs are unseen object categories in MS-COCO dataset.

2 Visualization of the modulation parameters

Our model implicitly learns an embedding with the modulation parameters from the visual modulator for the annotated objects. Intuitively, similar objects should have similar modulation parameters, while different objects should have dramatically different modulation parameters. To visualize this embedding, we extract modulation parameters from 100100 object instances in 1010 object classes in MS-COCO, and visualize the parameters in a two-dimensional embedding space using multi-dimensional scaling in Fig. 5. We can see that objects in the same category are mostly clustered together, and similar categories are closer to each other than dissimilar categories. For example, cats and dogs, cars and buses are mixed up due to their similar appearance, while bicycles and dogs, buses and horses are far from each other due to the big visual difference. Mammal classes (cats, dogs, cows, horses, human) are generally clustered together, and man-made objects (cars, buses, bicycles, motorcycles, trucks) are clustered together.

We also investigate the magnitude of the modulation parameters in different layers. The modulation parameters {γc}\{\gamma_{c}\} changes according to the visual guide. Therefore, we compute the standard deviations of modulation parameters {γc}\{\gamma_{c}\} in each modulation layer for images in MS-COCO validation set and illustrate them in Fig. 6. An interesting observation is that towards deeper level of the network, the variations of modulation parameters get larger. This shows that the manipulation of feature maps is more dramatic in the last few layers than in early layers of the network. The last few layers of a deep neural network usually learn high-level semantic meanings Zeiler2014visualizing, which could be used to adjust the segmentation model to a specific object more effectively.

3 Ablation Study

We study the impact of different ingredients in our method. We conduct experiments on DAVIS 2017 and measure the performance using mean IU. For variants of model structures, we experiment with only using spatial or visual modulator. For data augmentation methods, we experiment with no random crop augmentation for the FCN input, and no affine transformation for the visual guide and the spatial guide. We experiment with CRF as a post-processing step. To investigate the effect of one-shot fine-tuning on our model, we also experiment with standard one-shot fine-tuning using a small number of iterations. Results are shown in Table 3.

By adding a CRF post-processing, our method achieves mIU (mean IU) of 54.454.4. By one-shot fine-tuning with only 100100 iterations for each sequence, our method achieves mIU of 60.860.8, which is 5.75.7 better than OSVOS with 10001000 iterations. With fine-tuning, our method is still relatively efficient with average running time around 11 s/frame. Without visual modulator, our model deteriorates to 33.033.0, while without spatial modulator, our model obtains mIU of 40.140.1, which shows that the visual guide is more important than the spatial guide. For data augmentation, without random crop, the accuracy drops by 1.91.9. Without affine data augmentation on the visual guide, the accuracy further decreases by 1.11.1. Without augmentation on the spatial guide, our model only obtains mIU of 35.635.6, which is a dramatic drop from 49.549.5. The results indicates that the spatial guide augmentation is the most significant on the performance. Without perturbation, the model might rely on the location of the spatial prior too much that it cannot deal with moving objects in real video sequences.

Conclusions

In this work, we propose a novel framework to process one-shot video segmentation efficiently. To alleviate the slow speed of one-shot fine-tuning developed by previous FCN-based methods, we propose to use a network modulation approach mimicking the fine-tuning process with one forward pass of the modulator network. We show in experiments that by injecting a limited number of parameters computed by the modulators, the segmentation model can be re-purposed to segment an arbitrary object. The proposed network modulation method is a general learning method for few-shot learning problems, which could be applied to other tasks such as visual tracking and image stylization. Our approach falls into the general category of meta-learning, and it would also be interesting to investigate other meta-learning approaches for video segmentation. Another piece of future work would be to learn a recurrent representation of the modulation parameters to manipulate the FCN based on temporal information.

References