CAP-VSTNet: Content Affinity Preserved Versatile Style Transfer

Linfeng Wen, Chengying Gao, Changqing Zou

Introduction

Photorealistic style transfer aims to reproduce content image with the style from a reference image in a photorealistic way. To be photorealism, the stylized image should preserve clear content detail and consistent stylization of the same semantic regions. Content affinity preservation, including feature and pixel affinity preservation , is the key to achieve both clear content detail and consistent stylization in the transfer.

The framework of a deep learning based photorealistic style transfer generally uses such an architecture: an encoder module extracting content and style features, followed by a transformation module to adjust features statistics, and finally a decoder module to invert stylized feature back to stylized image. Existing photorealistic methods typically employ pre-trained VGG as encoder. Since the encoder is specifically designed to capture object-level information for the classification task, it inevitably results in content affinity loss. To reduce the artifacts, existing methods either use skip connection modules or build a shallower network . However, these strategies, limited by the image recovery bias, cannot achieve a perfect content affinity preservation on unseen images.

In this work, rather than use the traditional encoder-transformation-decoder architecture, we resort to a reversible framework based solution called CAP-VSTNet, which consists of a specifically designed reversible residual network followed by an unbiased linear transform module based on Cholesky decomposition that performs style transfer in the feature space. The reversible network takes the advantages of the bijective transformation and can avoid content affinity information loss during forward and backward inference. However, directly using the reversible network cannot work well on our problem. This is because redundant information will accumulate greatly when the network channel increases. It will further lead to content affinity loss and noticeable artifacts as the transform module is sensitive to the redundant information. Inspired by knowledge distillation methods , we improve the reversible network and employ a channel refinement module to avoid the redundant information accumulation. We achieve this by spreading the channel information into a patch of the spatial dimension. In addition, we also introduce cycle consistency loss in CAP-VSTNet to make the reversible network robust to small perturbations caused by numerical error.

Although the unbiased linear transform based on Cholesky decomposition can preserve feature affinity, it cannot guarantee pixel affinity. Inspired by , we introduce Matting Laplacian loss to train the network and preserve pixel affinity. Matting Laplacian may result in blurry images when it is used with another network like one with an encoder-decoder architecture. But it does not have this issue in CAP-VSTNet, since the bijective transformation of reversible network theoretically requires all information to be preserved.

CAP-VSTNet can be flexibly applied to versatile style transfer, including photorealistic and artistic image/video style transfer. We conduct extensive experiments to evaluate its performance. The results show it can produce better qualitative and quantitative results in comparison with the state-of-the-art image style transfer methods. We show that with minor loss function modifications, CAP-VSTNet can perform stable video style transfer and outperforms existing methods.

Related Work

Gatys et al. expose the powerful representation ability of deep neural networks and propose neural style transfer by matching the correlations of deep features. Feed-forward frameworks are proposed to address the issue of computational cost. To achieve universal style transfer, transformation modules are proposed to adjust statistics of deep features, such as the mean and variance and the inter-channel correlation .

Photorealistic style transfer requires that stylized image should be undistorted and consistently stylized. DPST optimizes stylized image with regularization term computed on Matting Laplacian to suppress distortion. PhotoWCT proposes post-processing algorithm by using Matting Laplacian as affinity matrix to reduce artifacts. However, both of these methods may blur the stylized images instead of preserving the pixel affinity. The following works mainly focus on preserving clear details and speeding up processing by designing skip connection module or shallower network. Content affinity preservation including feature and pixel affinity preservation remains an unsolved challenge.

Recently, versatile style transfer has received a lot of attention. Many approaches focus on exploring a general framework capable of performing artistic, photorealistic and video style transfer. Li et al. propose a linear style transfer network and a spatial propagation network for artistic and photorealistic style transfer, respectively. DSTN introduces a unified architecture with domain-aware indicator to adaptively balance between artistic and photorealistic stylization. Chiu et al. propose an optimization-based method to achieve fast artistic or photorealistic style transfer by simply adjusting the number of iterations. Chen et al. extend contrastive learning to artistic image and video style transfer by considering internal-external statistics. Wu et al. apply contrastive learning by incorporating neighbor-regulating scheme to preserve the coherence of the content source for artistic and photorealistic video style transfer. While achieving versatile style transfer, VGG-based networks suffer from inconsistent stylization due to content affinity loss. We show that preserving content affinity can improve image consistent stylization and video temporal consistency.

2 Reversible Network

Dinh et al. first propose an estimator that learns a bijective transform between data and latent space, which can be seen as a perfect auto-encoder pair as it naturally satisfies reconstruction term of auto-encoder . Follow-up work by Dinh et al. introduces new transformation that breaks the unit determinant of Jacobian to address volume-preserving mapping. Glow proposes a simple type of generative flow building on the works by Dinh . Since each layer’s activation of reversible network can be exactly reconstructed from the next layer’s, RevNet and Reformer present reversible residual layers to address memory consumption during deep network training. i-RevNet builds an invertible type of RevNet with invertible down-sampling module. i-ResNet inverts residual mapping by using Banach fixed point theorem to address the restriction of reversible network architecture.

Recently, An et al. apply flow-based model to address the content leak problem for artistic style transfer. However, content affinity may not be preserved due to transformation module and redundant information, which leads to noticeable artifacts. The proposed method addresses this issue via a new reversible residual network enhanced by a channel refinement module and a Matting Laplacian loss based training.

Method

The architecture of CAP-VSTNet is shown in Figure 2. Given a content image and a style image, our framework first maps the input content/style images to latent space through the forward inference of network after an injective padding module which increases the input dimension by zero-padding along the channel dimension. The forward inference is performed through cascaded reversible residual blocks and spatial squeeze modules. After that a channel refinement module is then used to remove the channel redundant information in content/style image features for a more effective style transformation. Then a linear transform module cWCT is used to transfer the content representation to match the statistics of the style representation. Lastly the stylized representation is inversely mapped back to the stylized image through the backward inference.

In our network design, each reversible residual block performs a function of a pair of inputs x1,x2→y1,y2x_{1},x_{2}\rightarrow y_{1},y_{2}, which can be expressed in the form:

Following Gome et al. , we use the channel-wise partitioning scheme that divides the input into two equal-sized parts along the channel dimension. Since the reversible residual block processes only half of the channel dimension at one time, it is necessary to perturb the channel dimension of the feature maps. We find that channel shuffling is effective and efficient: y=(y2,y1)y=(y_{2},y_{1}). Each block can be reversed by subtracting the residuals:

Figure 2 (a) and (b) show the illustration of the forward and backward inference of reversible residual block, respectively. The residual function FF is implemented by consecutive conv layers with kernel size 3. And each conv layer is followed by a relu layer, except for the last. We attain large receptive field by stacking multiple layers and blocks, in order to capture dense pairwise relations. We abandon the normalization layer as it poses a challenge to learn style representation. To capture large scale style information, the squeeze module is used to reduce the spatial information by a factor of 2 and increase the channel dimension by a factor of 4. We combine reversible residual blocks and squeeze modules to implement a multi-scale architecture.

2 Channel Refinement

The cascaded reversible residual block and squeeze module design in CAP-VSTNet leads to redundant information accumulation during forward inference as the squeeze module exponentially increases the channels. The redundant information will negatively affect the stylization. In , channel compression is used to address the redundant information problem and facilitate better stylization. In our network design, we instead use a channel refinement module (CR) which is more suitable for the connected cascaded reversible residual blocks.

As illustrated in Figure 3, the CR module first uses an injective padding module increasing latent dimension to ensure that the input content/style image feature channel can be divisible by the target channel. Then, it uses patch-wise reversible residual blocks to integrate large-field information, after that it spreads the channel information into a patch of the spatial dimension. There are several alternative design choices for this CR module. MLP based pointwise layers can also be used for information distillation. Our preliminary experiments have found aliasing artifacts may appear when pointwise layers (e.g., fully connected layers or inverted residuals ) are employed (see the results produced by CR-MLP and CR-IR in Figure 4), but the adopted CR-RRB design does not have this issue.

3 Transformation Module

Existing photorealistic methods typically employ WCT as transformation module, which contains whitening and coloring steps. Both of the above steps require the calculation of singular value decomposition (SVD). However, the gradient depends on the singular values σ\sigma by calculating 1minσ(i≠j)σi2−σj2\frac{1}{min_{\sigma(i\neq j)}\sigma_{i}^{2}-\sigma_{j}^{2}} . If the covariance matrix of content (style) feature map Σc=fcfcT(Σs=fsfsT)\Sigma_{c}=f_{c}f_{c}^{T}(\Sigma_{s}=f_{s}f_{s}^{T}) has the same singular values, or the distance between any two singular values is close to 0, the gradient becomes infinite. It will further cause the WCT module to fail and the model training to crash.

We use an unbiased linear transform based on Cholesky decomposition to address this problem. The Cholesky decomposition is derivable with gradient depending on 1σ\frac{1}{\sigma}. It does not require that the two singular values are not equal as SVD, thus is more stable. To avoid overflow, we can regularize it with an identity matrix: Σ^=Σ+ϵI\hat{\Sigma}=\Sigma+\epsilon I. Another advantage of Cholesky decomposition is that its computational cost is much lower than that of SVD. Therefore, the adopted Cholesky decomposition based WCT (cWCT for short) is more stable and faster. We show the comparison of various linear transformation modules in Table 1.

4 Training Loss

We train our network in an end-to-end manner with the integration of three types of losses:

where LmL_{m}, LsL_{s}, and LcycL_{cyc} denote Matting Laplacian loss, style loss, and cycle consistency loss, respectively. λm\lambda_{m} and λcyc\lambda_{cyc} are the weights corresponding to the losses.

The Matting Laplacian loss in our design can be formulated as:

where NN denotes the number of image pixels, Vc[Ics]V_{c}[I_{cs}] denotes the vectorization of the stylized image IcsI_{cs} in channel c, and MM denotes the Matting Laplacian matrix of the content image IcI_{c}.

Directly introducing Matting Laplacian loss in a network training could result in blurry images because Matting Laplacian loss will force the network to smooth the image rather than preserve pixel affinity. Fortunately, introducing Matting Laplacian loss in our reversible network does not have the issue. It is because the bijective transformation in our reversible network requires all information to be preserved during forward and backward inference. The reversible network does not trick the loss by smoothing the image as it results in information loss. When performing linear transform, it depends on covariance matrix Σs\Sigma_{s}. As the transformation of reversible network is deterministic, only a few style images with smooth texture may smooth the content structure. In this situation, it is reasonable to output a stylized image with the same smooth texture as we aim to transfer vivid style.

where IsI_{s} denotes style image, ϕi\phi_{i} denotes the ithi_{th} layer of the VGG-19 network (from ReLu1_1ReLu1\_1 to ReLu4_1ReLu4\_1), and μ\mu and σ\sigma denote the mean and variance of the feature maps, respectively.

The cycle consistency loss is calculated with L1L1 distance:

5 Video Style Transfer

Single-frame methods show that applying image algorithms that operate on each video frame individually is possible. Since our framework preserves the affinity of input videos, which is naturally consistent and stable, the content of stylized video is also visually stable. To constrain the style of stylized video , we have two strategies: adjust the style loss (Eq.5) with lower layers of VGG-19 network (from ReLu1_1ReLu1\_1 to ReLu3_1ReLu3\_1) or add the regularization to Eq.3 and fine-tune the model. Both strategies can achieve good temporal consistency. We choose the latter one as it can produce slightly better stylization effect.

Analysis

To show the advantages of preserving feature and pixel affinity, we compare the stylization results with three types of methods. As shown in Figure 6, LinearWCT applies linear transform to preserve feature affinity. However, the image details is unclear and the stylization is inconsistent as feature and pixel affinity could be damaged by VGG-base network. WCT2 aims to preserve spatial information rather than content affinity. While preserving clear details, it particularly relies on the precise masks, which otherwise produce noticeable seams. ArtFlow uses the flow-based model to address content leak problem. However, it typically generates noticeable artifacts as linear transform and redundant information damage content affinity. Compared with other methods, ours model not only preserves clear details, but also achieves seamless style transfer.

2 Ablation Study

We conduct an ablation study to quantitatively evaluate how much each component (i.e., channel refinement components and training losses) affects the visual effects. Table 2 shows the ablation study results. When all the design components are used, the network can obtain the best results. Replacing residual block (RRB) with inverted residuals degrades performance as the pointwise layer has smaller receptive field and damages content affinity. Removing injective padding (IP), the model fails to capture high-level content and style information from pixel image. Adding the channel refinement module (CR-RRB) helps remove redundant information for better content preservation and stylization effect. Implementing the channel refinement module with CR-MLP results in aliasing artifact, which degrades content affinity. Using VGG content loss (w/o Lm&LcycL_{m}\&L_{cyc}) cannot guarantee pixel affinity due to the linear transform. With cycle consistency loss (LcycL_{cyc}), the network achieves robustness to small perturbations.

Experiments

We implement a three-scale architecture with 30 blocks and 2 squeeze modules. For photorealistic style transfer, we sample content and style images from MS COCO dataset and randomly crop them to 256×\times256. We set the weight factors of loss function as: λm=1200\lambda_{m}=1200 and λcyc=10\lambda_{cyc}=10. We train the network for 160,000 iterations using Adam optimizer with batch size of 2. The initial learning rate is set to 1e-4 and decays at 5e-5. For artistic style transfer, we set λm=λcyc=1\lambda_{m}=\lambda_{cyc}=1 to allow more variation of image pixel and sample style images from WikiArt dataset . All the experiments are conducted on a single NVIDIA RTX 3090 GPU.

2 Photorealistic Image Style Transfer

Figure 7 shows the comparison of the stylization results with advanced photorealistic style transfer methods, including PhotoWCT , WCT2 , PhotoNet , DSTN and PCA-KD . We can see that PhotoWCT usually generates blurry images with loss of details. Although WCT2 faithfully preserves image spatial information, it produces noticeable seams. PhotoNet generates poor stylization effect due to discarding masks. DSTN stylizes images with noticeable artifacts and distorts image structure. PCA-KD is not able to produce consistent stylization. Compared with the existing methods, our method faithfully preserves image details and achieves better stylization effect. Besides, image stylization is consistent without artifacts, which greatly enhances photorealism.

Quantitative evaluation.

Following previous works , we use structural similarity (SSIM) to evaluate photorealism and Gram loss to evaluate stylization effect. We use all pairs of content and style images with semantic segmentation masks provided by DPST for quantitative evaluation. Table 3 shows the comparison of quantitative results. Our method not only preserves structure better, but also achieves stronger stylization effect. Since the reversible residual network naturally satisfies the reconstruction condition, we reduce network parameters and make it more lightweight than most of standard VGG-based networks. PCA-KD applies knowledge distillation method to crate the lightweight model for ultra-resolution style transfer. We note that our model is also applicable for ultra-resolution (i.e., 4K resolution) and achieves better performance as well.

3 Video Style Transfer

We compare our method with state-of-the-art methods . To visualize video stability, we show the heatmap of temporal error between the consecutive frames in Figure 8. To quantitatively evaluate, we collect 20 pairs of video clips of multiple scenes and semantically related style images from the Internet. Following , we adopt the temporal loss to measure temporal consistency. We use RAFT to estimate the optical flow for short-term consistency (two adjacent frames) and long-term consistency (9 frames in between) evaluation. Table 4 shows that our framework performs well against the other methods.

Artistic video style transfer.

Figure 9 shows the comparison with four advanced methods . To quantitatively evaluate, we use all the sequences of MPI Sintel dataset and collect 20 artworks of various types to stylize each video. For short-term consistency, MPI Sintel provides ground truth optical flows. For long-term consistency, we use PWC-Net to estimate the optical flow following . Table 5 shows that our framework achieves the best temporal consistency, thanks to the content affinity preservation. Our model also produces vivid stylization effect comparable to CCPL .

4 Limitation

Preserving content affinity helps to achieve consistent stylization. However, both our artistic and photorealistic models fail to capture complex texture and may generate artifacts (Figure 10). Generating realistic textures remains a challenge for style transfer and image generation tasks. Existing stylization methods typically build on small models (e.g., VGG). Since realistic texture requires much high-frequency details, an interesting direction is to investigate whether large models can solve this problem.

Conclusion

In this paper, we propose a new framework named CAP-VSTNet for versatile style transfer, which consists of a new effective reversible residual network and an unbiased linear transform. It can preserve two major content affinity: pixel and feature affinity with the introduction of Matting Laplacian training loss. We show that CAP-VSTNet achieves consistent and vivid stylization with clear details. CAP-VSTNet is also flexible for photorealistic and artistic video style transfer. Extensive experiments demonstrate the effectiveness and superiority of CAP-VSTNet in comparisons with state-of-the-art approaches.

Acknowledgement

This work was supported by the Natural Science Foundation of Guangdong Province, China (Grant No. 2022A1515011425).

References

Appendix A Pseudocode of CAP-VSTNet

Appendix B Unbiased Transformation Module

We show that the whitening and coloring transforms in cWCT is an unbiased style transfer module. Suppose we have a style transfer module fcs=C(fc)S(fs)f_{cs}=C(f_{c})S(f_{s}), where C,SC,S denote the content and style factor, and fc,fsf_{c},f_{s} denote the content and style feature. ArtFlow defines that fcsf_{cs} is an unbiased style transfer moudle if C(fcs)=C(fc)C(f_{cs})=C(f_{c}) and S(fcs)=S(fs)S(f_{cs})=S(f_{s}).

Without loss of generality, we assume both fcf_{c} and fsf_{s} are centered. We have,

where fcfcT=LcLcTf_{c}f_{c}^{T}=L_{c}L_{c}^{T}, fsfsT=LsLsTf_{s}f_{s}^{T}=L_{s}L_{s}^{T} and LL is a triangular matrix. Therefore,

Appendix C Style Interpolation

We investigate the linear interpolation of extracted style representations by CAP-VSTNet. Considering the feature covariance matrices ΣA\Sigma_{A} and ΣB\Sigma_{B} of style images IAI_{A} and IBI_{B}, the interpolated Σs\Sigma_{s} should be:

where α\alpha is the style ratio between the two. Figure 11 and 12 present the smooth transformation from one style image to another.

Appendix D Comparison with ArtFlow

Figure 13 and 14 show the comparisons with ArtFlow.

Appendix E Photorealistic Image Style Transfer

Figure 15 and 16 show the comparisons with advanced photorealistic image style transfer methods.

Appendix F Video Style Transfer

Figure 17, 18, 19 and 20 show the comparisons with advanced photorealistic video and artistic video style transfer methods.

Appendix G Ultra-Resolution Photorealistic Style Transfer

Figure 21 and 22 show the ultra-resolution (4K) photorealistic stylization results of CAP-VSTNet.