Arbitrary Video Style Transfer via Multi-Channel Correlation
Yingying Deng, Fan Tang, Weiming Dong, Haibin Huang, Chongyang Ma, Changsheng Xu
Introduction
Style transfer is a significant topic in the industrial community and research area of artificial intelligence. Given a content image and an art painting, a desired style transfer method can render the content image into the artistic style referenced by the art painting. Traditional style transfer methods based on stroke rendering, image analogy, or image filtering (Efros and Freeman 2001; Bruckner and Gröller 2007; Strothotte and Schlechtweg 2002) only use low-level features for texture transfer (Gatys, Ecker, and Bethge 2016; Doyle et al. 2019). Recently, deep convolutional neural networks (CNNs) have been widely studied for artistic image generation and translation (Gatys, Ecker, and Bethge 2016; Johnson, Alahi, and Fei-Fei 2016; Zhu et al. 2017; Huang and Serge 2017; Huang et al. 2018; Jing et al. 2020).
Although existing methods can generate satisfactory results for still images, they lead to flickering effects between adjacent frames when applied to videos directly (Ruder, Dosovitskiy, and Brox 2016; Chen et al. 2017). Video style transfer is a challenging task that needs to not only generate good stylization effects per frame but also consider the continuity between adjacent frames of the video. Ruder, Dosovitskiy, and Brox (2016) added temporal consistency loss on the basis of the approach proposed in Gatys, Ecker, and Bethge (2016) to maintain a smooth transition between video frames. Chen et al. (2017) proposed an end-to-end framework for online video style transfer through a real-time optical flow and mask estimation network. Gao et al. (2020) proposed a multistyle video transfer model, which estimates the light flow and introduces temporal constraint. However, these methods highly depend on the accuracy of optical flow calculation, and the introduction of temporal constraint loss reduces the stylization quality of individual frames. Moreover, adding optical flow to constrain the coherence of stylized videos makes the network difficult to train when applied to an arbitrary style transfer model.
In the current work, we revisit the basic operations in state-of-the-art image stylization approaches and propose a frame-based Multi-Channel Correlation network (MCCNet) for temporally coherent video style transfer that does not involve the calculation of optical flow. Our network adaptively rearranges style representations on the basis of content representations by considering the multi-channel correlation of these representations. Through this approach, MCCNet is able to make style patterns suitable for content structures. By further merging the rearranged style and content representations, we generate features that can be decoded into stylized results with clear content structures and vivid style patterns. MCCNet aligns the generated features to content features, and thus, slight changes in adjacent frames will not cause flickering in the stylized video. Furthermore, the illumination variation among consecutive frames influences the stability of video style transfer. Thus, we add random Gaussian noise to the content image to simulate illumination varieties and propose an illumination loss to make the model stable and avoid flickering. As shown in Figure 1, our method is suitable for stable video style transfer, and it can generate single-image style transfer results with well-preserved content structures and vivid style patterns. In summary, our main contributions are as follows:
We propose MCCNet for framed-based video style transfer by aligning cross-domain features with input videos to render coherent results.
We calculate the multi-channel correlation across content and style features to generate stylized images with clear content structures and vivid style patterns.
We propose an illumination loss to make the style transfer process increasingly stable so that our model can be flexibly applied to videos with complex light conditions.
Related Work
Image style transfer has been widely studied in recent years. Essentially, it enables the generation of artistic paintings without the expertise of a professional painter. Gatys, Ecker, and Bethge (2016) found that the inner products of the feature maps in CNNs can be used to represent style and proposed a neural style transfer (NST) method through continuous optimization iterations. However, the optimization process is time consuming and cannot be widely used. Johnson, Alahi, and Fei-Fei (2016) put forward a real-time style transfer method to dispose of a specific style in one model. Dumoulin, Shlens, and Kudlur (2016) proposed conditional instance normalization (CIN), which allows the learning of multiple styles in one model by reducing a style image into a point in the embedded space. A number of methods achieve arbitrary style transfer by aligning the second-order statistics of style and content images (Huang and Serge 2017; Li et al. 2017; Wang et al. 2020b). Huang and Serge (2017) proposed an arbitrary style transfer method by adopting adaptive instance normalization (AdaIN), which normalizes content features using the mean and variance of style features. Li et al. (2017) used whitening and coloring transformation (WCT) to render content images with style patterns. Wang et al. (2020b) adopted deep feature perturbation (DFP) in a WCT-based model to achieve diversified arbitrary style transfer. However, these holistic transformations lead to unsatisfactory results. Park and Lee (2019) proposed a style-attention network (SANet) to obtain abundant style patterns in generated results but failed to maintain distinct content structures. Yao et al. (2019) proposed an attention-aware multi-stroke style transfer (AAMS) model by adopting self-attention to a style swap-based image transfer method, which highly relies on the accuracy of the attention map used. Deng et al. (2020) proposed a multi-adaptation style transfer (MAST) method to disentangle content and style features and combine them adaptively. However, some results are rendered with uncontrollable style patterns.
In the current work, we propose an arbitrary style transfer approach, which can be applied to video transfer with better stylized results than other state-of-the-art methods.
Video style transfer.
Most video style transfer methods rely on existing image style transfer methods (Ruder, Dosovitskiy, and Brox 2016; Chen et al. 2017; Gao et al. 2020). Ruder, Dosovitskiy, and Brox (2016) built on NST and added a temporal constraint to avoid flickering. However, the optimization-based method is inefficient for video style transfer. Chen et al. (2017) proposed a feed-forward network for fast video style transfer by incorporating temporal information. Chen et al. (2020) distilled knowledge from the video style transfer network with optical flow to a student network to avoid optical flow estimation in the test stage. Gao et al. (2020) adopted CIN for multi-style video transfer, and the approach incorporated one FlowNet and two ConvLSTM modules to estimate light flow and introduce a temporal constraint. The temporal constraint aforementioned is achieved by calculating optical flow, and the accuracy of optical flow estimation affects the coherence of the stylized video. Moreover, the diversity of styles is limited because of the used basic image style transfer methods used. Wang et al. (2020a) proposed a novel interpretation of temporal consistency without optical flow estimation for efficient zero-shot style transfer. Li et al. (2019) learned a transformation matrix for arbitrary style transfer. They found that the normalized affinity for generated features are the same as that for content features, and is thus suitable for frame-based video style transfer. However, the style patterns in their stylized results are not clearly noticeable.
To avoid using optical flow estimation while maintaining video continuity, we aim to design an alignment transform operation to achieve stable video style transfer with vivid style patterns.
Methodology
As shown in Figure 2, the proposed MCCNet adopts an encoder-decoder architecture. Given a content image and a style image , we can obtain corresponding feature maps and through the encoder. Through MCC calculation, we generate that can be decoded into stylized image .
We first formulate and analyze the proposed multi-channel correlation in Sections 3.1 and 3.2 and then introduce the configuration of MCCNet involving feature alignment and fusion in Section 3.3.
Cross-domain feature correlation has been studied for image stylization (Park and Lee 2019; Deng et al. 2020). It fuses the multiple content and style features by using several adaption/attention modules without considering the inter-channel relationship of the content features. In SANet (Park and Lee 2019), the generated features can be formulated as
Slight changes in input content features can lead to a large-scale variation in output features. MAST (Deng et al. 2020) has the same issue. Thus the coherence of input content features cannot be migrated to generated features, and the stylized videos will present flickering artifacts.
where . Finally, the channel-wise generated features are
However, the multi-channel correlation in style features is also important to represent style patterns (e.g., texture and color). We calculate the correlation between each content channel and every style channel. The -th channel for generated features can be rewritten as
where is the number of channels and represents the weights of the -th style channel. Finally, the generated features can be obtained by
where is the style information learned by training.
From Equation (6), we can conclude that MCCNet can help to generate stylized features that are strictly aligned to content features. Therefore, the coherence of input videos is naturally maintained to output videos and slight changes (e.g., objects motions) in adjacent frames cannot lead to violent changes in the stylized video. Moreover, through the correlation calculation between content and style representations, the style patterns are adjusted according to content distribution. The adjustment helps generate stylized results with stable content structures and appropriate style patterns.
2 Coherence Analysis
A stable style transfer requires that the output stylized video to be as coherent as the input video. As described in Wang et al. (2020a), a coherent video should satisfy the following constraint:
where are the -th and -th frames, respectively; and is the warping matrix from to . is a minimum so that humans are not aware of the minor flickering artifacts. When the input video is coherent, we obtain the content features of each frame. And the content features of the -th and -th frame satisfy
For corresponding output features and :
where is a minimum. Therefore, our network can migrate the coherence of the input video frame features to the stylized frame features without additional temporal constraints. We further demonstrate that the coherence of stylized frame features can be well-transited to the generated video despite the convolutional operation of the decoder in Section 4.4. Such observations prove that the proposed MCCNet is suitable for video style transfer tasks.
3 Network Structure and Training
As shown in Figure 4, our network is trained by minimizing the loss function defined as
The total loss function includes perceptual loss and , identity loss and illumination loss in the training procedure. The weights , , , and are set to , , , and to eliminate the impact of magnitude differences.
We use a pretrained VGG19 to extract content and style feature maps and compute the content and style perceptual loss similar to AdaIN (Huang and Serge 2017). In our model, we use layer to calculate the content perceptual loss and layers , , , and to calculate the style perceptual loss. The content perceptual loss is used to minimize the content differences between generated images and content images, where
The style perceptual loss is used to minimize the style differences between generated images and style images:
where denotes the features extracted from the -th layer in a pretrained VGG19, denotes the mean of features, and denotes the variance of features.
Identity loss.
We adopt the identity loss to constrain the mapping relation between style features and content features, and help our model to maintain the content structure without losing the richness of the style patterns. The identity loss is defined as:
where denotes the generated results using a common natural image as content image and style image and denotes the generated results using a common painting as content image and style image.
Illumination loss.
For a video sequence, the illumination may change slightly that is difficult to be discovered by humans. The illumination variation in video frames could influence the final transfer results and result in flicking. Therefore, we add random Gaussian noise to the content images to simulate light. The illumination loss is formulated as
where is our generation function, . With illumination loss, our method can be robust to complex light conditions in input videos.
Experiments
Typical video stylization methods use temporal constraint and optical flow to avoid flickering in generated videos. Our method focuses on promoting the stability of the transform operation in the arbitrary style transfer model on the basis of single frame. Thus, the following frame-based SOTA stylization methods are selected for comparison: MAST (Deng et al. 2020), CompoundVST (Wang et al. 2020a), DFP (Wang et al. 2020b), Linear (Li et al. 2019), SANet (Park and Lee 2019), AAMS (Yao et al. 2019), WCT (Li et al. 2017), AdaIN (Huang and Serge 2017), and NST (Gatys, Ecker, and Bethge 2016).
In this section, we start from the training details of the proposed approach and then move on to the evaluation of image (frame) stylization and the analysis of rendered videos.
We use MS-COCO (Lin et al. 2014) and WikiArt (Phillips and Mackintosh 2011) as the content and style image datasets for network training. At the training stage, the images are randomly cropped to pixels. At the inference stage, images in arbitrary size are acceptable. The encoder is a pretrained VGG19 network. The decoder is a mirror version of the encoder, except for the parameters that need to be trained. The training batch size is , and the whole model is trained through steps.
We measure our inference time for the generation of an output image and compare the result with those of SOTA methods using 16G TitanX GPU. The optimization-based method NST (Gatys, Ecker, and Bethge 2016) is trained for epochs. Table 1 shows the inference times of different methods using three scales of image size. The inference speed of our network is much faster than that in (Gatys, Ecker, and Bethge 2016; Li et al. 2017; Yao et al. 2019; Wang et al. 2020b, a; Deng et al. 2020). Our method can achieve a real-time transfer speed that is comparable to that of (Huang and Serge 2017; Li et al. 2019; Park and Lee 2019) for efficient video style transfer.
2 Image Style Transfer Results
CompoundVST is not selected for image stylization comparison because it is only used for video style transfer. The comparisons of image stylizations are shown in Figure 5. On the basis of the optimized training mechanism, NST may introduce failure results (the second row), and it cannot easily achieve a trade-off between the content structure and style patterns in the rendered images.
In addition to crack effects caused by over-simplified calculation in AdaIN, the absence of correlation between different channels causes a poor transfer of style color distribution badly transferred (the 4th row). As for WCT, the local content structures of generated results are damaged due to the global parameter transfer method. Although the attention map in AAMS helps to make the main structure exact, the patch splicing trace affects the overall generated effect. Some style image patches are transferred into the content image directly (the first row) by SANet and damage the content structures (the fourth row). By learning a transformation matrix for style transfer, the process of Linear is too simple to acquire adequate style textural patterns for rendering. DFP focuses on generating diversified style transfer results, but it may lead to failures similar to WCT. MAST may introduce unexpected style patterns in rendered results (the fourth and the fifth rows).
Our network calculates the multi-channel correlation of content and style features and rearranges style features on the basis of the correlation. The rearranged style features fit the original content features, and the fused generated features consist of clear content structures and controllable style pattern information. Therefore, our method can achieve satisfactory stylized results.
Quantitative analysis.
Two classification models are trained to assess the performance of different style transfer networks by considering their ability to maintain content structure and style migration. We generate several stylized images by using different style transfer methods aforementioned. Then we input the stylized images generated by different methods. Then, we input the stylized images generated by different methods into the style and content classification models. High-accuracy style classification indicates that the style transfer network can easily learn to obtain effective style information while high-accuracy content classification indicates that the style transfer network can easily learn to maintain the original content information. From Figure 6, we can conclude that AAMS, AdaIN, and MCCNet can establish a balance between content and style. However, our network is superior to the two methods in terms of visual effects.
User study
We conducted user studies to compare the stylization effect of our method with those of the aforementioned SOTA methods. We selected content images and style images to generate stylized images using different methods. First, we showed participants a pair of content and style images. Second, we showed them two stylized results generated by our method and a random contrast method. Finally, we asked the participants which stylized result has the best rendered effects by considering the integrity of the content structures and the visibility of the style patterns. We collected votes from participants and present the voting results in Figure 7(a). Overall, our method can achieve the best image stylization effect.
3 Video Style Transfer Results
Considering the size limitation for input images of SANet and generation diversity of DFP, we do not include SANet and DFP for video stylization comparisons. We synthesize stylized video clips by using the other methods and measure the coherence of the rendered videos by calculating the differences in the adjacent frames. As shown in Figure 8, the heat maps in the second row visualize the differences between two adjacent frames of the input or stylized videos. Our method can highly promote the stability of image style transfer. The differences of our results are closest to those of the input frames without damaging the stylization effect. MAST, Linear, AAMS, WCT, AdaIN, and NST fail to retain the coherence of the input videos. Linear can also generate a relatively coherent video, but the result continuity is influenced by nonlinear operation in deep CNNs.
Given two adjacent frames and in a T-frame rendered clip, we define and calculate the mean () and variance () of . As shown in Table 2, we can conclude that our method can yield the best video results with high consistency.
The global feature sharing used in CompoundVST increases the complexity of the model and limits the length of video clips that can be processed. Therefore, CompoundVST is not selected for comparison in this section. Then we conducted user studies to compare the video stylization effects of our method with those of the aforementioned SOTAs. First, we showed the participants an input video clip and a style image. Second, we showed them two stylized video clips generated by our method and a random contrast method. Finally, we asked the participants which stylized video clip is the most satisfying by considering the stylized effect and the stability of the videos. We collected votes from participants and present the the voting results in Figure 7(b). Overall, our method can achieve the best video stylization effect.
4 Ablation Study
Our network is based on multi-channel correlation calculation shown in Figure 3. To analyze the influence of multi-channel correlation on stylization effects, we change the model to calculate channel-wise correlation without considering the relationship between style channels. Figure 9 shows the results. Through channel-wise calculation, the stylized results show few style patterns (e.g., hair of the woman) and may maintain the original color distribution (the blue color in the bird’s tail). Meanwhile, the style patterns are effectively transferred by considering the multi-channel information in the style features.
Illumination loss.
The illumination loss is proposed to eliminate the impact of video illumination variation. We remove the illumination loss in the training stage and compare the results with ours in Figure 10. Without illumination loss, the differences between the two video frames increase, with the mean value being . With illumination loss, our method can be flexibly applied to videos with complex light conditions.
Network depth.
We use a shallow auto-encoder up to - instead of - to discuss the effects of the convolution operation of the decoder on our model. As shown in Figure 11, for image style transfer, the shallower network can not generate results with vivid style patterns (e.g., the circle element in the first row of our results). As shown in Figure 12, the depth of the network exerts little impact on the coherence of the stylized video. This phenomenon suggests that the coherence of stylized frames features can be well-transited to the generated video despite the convolution operation of the decoder.
Conclusion
In this work, we propose MCCNet for stable arbitrary video style transfer. The proposed network can migrate the coherence of input videos to stylized videos and thereby guarantee the stability of rendered videos. Meanwhile, MCCNet can generate stylized results with vivid style patterns and detailed content structures by analyzing the multi-channel correlation between content and style features. Moreover, the illumination loss improves the stability of generated video under complex light conditions.
Acknowledgements
This work was supported by National Key R&D Program of China under no. 2020AAA0106200, and by National Natural Science Foundation of China under nos. 61832016, U20B2070 and 61672520.