Domain Enhanced Arbitrary Image Style Transfer via Contrastive Learning
Yuxin Zhang, Fan Tang, Weiming Dong, Haibin Huang, Chongyang Ma, Tong-Yee Lee, Changsheng Xu
Introduction
If a picture is worth a thousand words, then an artwork tells the whole story. Art styles, which describe the way the artwork looks, are the manner in which the artist portrays his or her subject matter and how the artist expresses his or her vision. Style is determined by the characteristics that describe the artwork, such as the way the artist employs form, color, and composition. Artistic style transfer, as an efficient way to create a new painting by combining the content of a natural images and the style of an existing painting image, is a major research topic in computer graphics and computer vision (Liao et al. 2017; Jing et al. 2020b), with style representation as the most important issue.
Since Gatys et al. (Gatys et al. 2016) proposed to use Gram matrix as artistic style representation, high-quality visual results are generated by advanced neural style transfer networks. Despite the remarkable progress made in the field of arbitrary image style transfer, the second-order feature statistics (Gram matrix or mean/variance) style representation has restricted the further development and application. As shown in Figure 1, the appearances of different artwork styles vary considerably in terms of not only the colors and local textures but also the layouts and compositions. Figures 2(d) and 2(e) show the results of two recently proposed state-of-the-art style transfer approaches. We obverse that aligning the distributions of neural activation between images using second-order statistics results in difficulty to capture the color distribution or the special layouts, or imitate specific detailed brush effects of different styles.
In this paper, we revisit the core problem for neural style transfer, that is, the proper artistic style representation. The widely used second-order statistics as a global style descriptor can distinguish styles to some extent, but they are not the optimal way to represent styles. By second-order statistics, arbitrary stylization formulates styles through artificially designed image features and loss functions in a heuristic manner. In other words, the network learns to fit the second-order statistics of the style image and generated image, instead of the style itself. Exploring the relationship and distribution of styles directly from artistic images instead of using pre-defined style representations is worthwhile.
Toward this end, we propose to improve arbitrary style transfer with a novel style representation by contrastive learning-based optimization. Our key insight is that a person without artistic knowledge has difficulty defining the style if only one artistic image is given, but identifying the difference between different styles is relatively easy. Specifically, we present a novel Contrastive Arbitrary Style Transfer (CAST) framework for image style representation and style transfer. CAST consists of a backbone based on an encoder-transformation-decoder structure, a multi-layer style projector (MSP) module, and a domain enhancement (DE) module. We introduce contrastive learning to consider the positive and negative relationships between styles, and we use DE to learn the distribution of overall art image domains. To capture the style features at various scales, our MSP module projects the features of each layer of the style image to the corresponding style encoding space.
Our contribution can be summarized as follows:
We propose an MSP module for style encoding and a novel CAST model for encoder-transformation-decoder-based arbitrary style transfer without using the second-order statistics as style representations.
We introduce contrastive learning and domain enhancement by considering the relationships between positive and negative examples as well as the global distribution of styles, which solves the problem that existing style transfer models cannot fully utilize a large amount of style information.
Experiments show that our method achieves state-of-the-art style transfer results in terms of visual quality. A challenging subjective survey was conducted, as inspired by the Turing test, to show that output of CAST could mislead participants from telling the fake painting images from real ones.
Related Work
Traditional style transfer methods such as stroke-based rendering (Fišer et al. 2016) and image filtering (Wang et al. 2004) typically use low-level hand-crafted features. Gatys et al. (Gatys et al. 2016) and the follow-up variants (Gatys et al. 2017; Kolkin et al. 2019) demonstrate that the statistical distribution of features extracted from pre-trained deep convolutional neural networks can capture style patterns effectively. Although the results are remarkable, these methods formulate the task as a complex optimization problem, which leads to high computational cost. Some recent approaches rely on a learnable neural network to match the statistical information in feature space for efficiency. Per-style-per-model methods (Johnson et al. 2016; Gao et al. 2020; Puy and Pérez 2019) train a specific network for each individual style. Multiple-style-per-model methods (Chen et al. 2017; Zhang and Dana 2018; Dumoulin et al. 2017; Ulyanov et al. 2016) represent multiple styles using one single model.
Arbitrary style transfer methods (Li et al. 2017; Deng et al. 2020; Svoboda et al. 2020; Wu et al. 2021a; Deng et al. 2022) build more flexible feed-forward architectures to handle an arbitrary style using a unified model. AdaIN (Huang and Belongie 2017) and DIN (Jing et al. 2020a) directly align the overall statistics of content features with the statistics of style features and adopt conditional instance normalization. However, dynamic generation of affine parameters in the instance normalization layer may cause distortion artifacts. Instead, several methods follow the encoder-decoder manner, where feature transformation and/or fusion is introduced into an auto-encoder-based framework. For example, Li et al. (Li et al. 2019) learn a cross-domain feature linear transformation matrix (LST) to enable universal style transfer and generate the desired stylization results by decoding from the transformed features. Park et al. (Park and Lee 2019) introduce SANet to flexibly match the semantically nearest style features onto the content features. Deng et al. (Deng et al. 2021) propose MCCNet to fuse exemplar style features and input content features by multi-channel correlation for efficient style transfer. An et al. (An et al. 2021) propose reversible neural flows and an unbiased feature transfer module (ArtFlow) to prevent content leak during universal style transfer. Liu et al. (Liu et al. 2021b) present an adaptive attention normalization module (AdaAttN) to consider both shallow and deep features for attention score calculation. GAN-based methods (Zhu et al. 2017; Svoboda et al. 2020; Kotovenko et al. 2019b; Kotovenko et al. 2019a; Sanakoyeu et al. 2018a) have been successfully used in collection style transfer, which considers style images in a collection as a domain (Chen et al. 2021b; Xu et al. 2021; Lin et al. 2021).
Contrastive learning.
Contrastive learning has been used in many applications, such as image dehazing (Wu et al. 2021b), context prediction (Santa Cruz et al. 2019), geometric prediction (Liu et al. 2019) and image translation. Contrastive learning is introduced in image translation to preserve the content of the input (Han et al. 2021) and reduce mode collapse (Liu et al. 2021a; Jeong and Shin 2021; Kang and Park 2020). CUT (Park et al. 2020) proposes patch-wise contrastive learning by cropping input and output images into patches and maximizing the mutual information between patches. Following CUT, TUNIT (Baek et al. 2021) adopts contrastive learning on images with similar semantic structures. However, the semantic similarity assumption does not hold for arbitrary style transfer tasks, which leads the learned style representations to a significant performance drop. IEST (Chen et al. 2021a) applies contrastive learning to image style transfer based on feature statistics (mean and standard deviation) as style priors. The contrastive loss is calculated only within the generated results. Contrastive learning in IEST is an auxiliary method to associate stylized images sharing the same style, and the ability comes from the feature statistics from pre-trained VGG. Differently, we introduce contrastive learning for style representation by proposing a novel framework that uses visual features comprehensively to represent style for the task of arbitrary image style transfer.
Method
As shown in Figure 3, our framework consists of three key components: (1) a multi-layer style projector which is trained to project features of artistic image into style code; (2) a contrastive style learning module which is applied to guide both the training of the multi-layer style projector and the style image generation; and (3) a domain enhancement scheme to further help learn the distribution of artistic image domain. All these components are used for learning style representations to measure the difference between the input artistic images and generated results and thus, they could be applied to different kinds of arbitrary style transfer networks.
Our goal is to develop an arbitrary style transfer framework that can capture and transfer the local stroke characteristics and overall appearance of an artistic image to a natural image. A key component is to find a suitable style representation which can be used to distinguish different styles and further guide the generation of style images. To this end, we design an MSP module, which includes a style feature extractor and a multi-layer projector. Instead of using features from a specific layer or a fusion of multiple layers, our MSP projects features of different layers into separate latent style spaces to encode local and global style cues.
Specifically, we adapt VGG-19 (Simonyan and Zisserman 2014) and finetune the VGG-19 model pre-trained on ImageNet with a collection of 18,000 artistic images in 30 categories. We then select layers of feature maps in VGG-19 as input to our multi-layer projector (we use layers of ReLU1_2, ReLU2_2, ReLU3_3, and ReLU4_3 in all experiments). We use max pooling and average pooling to capture the mean and peak value of features. The multi-layer projector consists of pooling, convolution, and several multilayer perceptron layers, and it projects the style features into a set of -dimensional latent style code, as shown in Figure 4.
2. Contrastive Style Learning
As demonstrated above, the style code of an image can be used as the target for MSP training and the guidance for the style transfer network. However, we lack the ground-truth style code for supervised training. Therefore, we adopt contrastive learning and design a new contrastive style loss as an implicit measurement for network training.
where denotes the dot product of two vectors, and is a temperature scaling factor and is set to be in all of our experiments. Meanwhile, we maintain a large dictionary of 4096 negative examples using a memory bank architecture following MOCO (He et al. 2020). It is worth noting that we calculate the contrastive loss between images, as opposed to CUT (Park et al. 2020) which adopts contrastive learning by cropping images into patches and maximizing the mutual information between patches.
The contrastive representation also provides proper guidance for the generator to transfer styles between images. We adopt the same form of contrastive loss as used for learning MSP in Eq. (1), but compute the loss using the contrastive representations of the output image and the reference style image , then will have a style similar to :
3. Domain Enhancement
We introduce DE with adversarial loss to enable the network to learn the style distribution Recent style transfer models employ GAN (Goodfellow et al. 2014) to align the distribution of generated images with specific artistic images (Chen et al. 2021b; Lin et al. 2021). The adversarial loss can enhance the holistic style of the stylization results, while it strongly relies on the distribution of datasets. Even with the specific artistic style loss, the generation process is often not robust enough to be artifact-free.
Differently from these previous methods, we divide the images in the training set into realistic domain and artistic domain, and we use two discriminators and to enhance them respectively (see Figure 3). During the training process, we first randomly select an image from the realistic domain as the content image and another image from the artistic domain as the style image . and are used as the real samples of and , respectively. The generated image is used as the fake sample of . We exchange the content and style images to generate an image as the fake sample of . The adversarial loss is:
To maintain the content information of the content image in the process of style transfer between the two domains, we also add a cycle consistency loss:
4. Network Training
Our full objective function for training of the generator and discriminators and is formulated as:
where , , and are weights to balance different loss terms. We set , , and in all of our experiments.
We collect 100,000 artistic images in different styles from WikiArt (Phillips and Mackintosh 2011) and randomly sample 20,000 images as our artistic dataset. We averagely sample 20,000 images from Places365 (Zhou et al. 2018) as realistic image dataset. We train and evaluate our framework on those artistic and realistic images. In the training phase, all images are loaded with resolution. The number of feature map layers is set to be 4. The dimension of style latent code is set to 512, 1024, 2048, and 2048 for the four different layers, respectively. We use Adam (Kingma and Ba 2014) as optimizer with , , and a batch size of . The initial learning rate is set to and linear decayed linear for total iterations. The training process takes about hours on one NVIDIA GeForce RTX3090. We choose the same backbone as AdaIN (Huang and Belongie 2017) in our experiments for simplicity. The results of using other backbones are shown in the supplementary materials.
Experiments
We compare CAST with several state-of-the-art style transfer methods, including NST (Gatys et al. 2016), AdaIN (Huang and Belongie 2017), LST (Li et al. 2019), SANet (Park and Lee 2019), ArtFlow (An et al. 2021), MCCNet (Deng et al. 2021), AdaAttN (Liu et al. 2021b), and IEST (Chen et al. 2021a). All the baselines are trained using publicly available implementations with default configurations. The comparison of inference speed is shown in Table 1.
We first present qualitative results of our method against the selected state-of-the-art methods in Figure 5. The comparison shows the superiority of CAST in terms of visual quality. NST is likely to encounter the issue of unpleasant local minimum (e.g., the 1st, 5th and 8th rows). AdaIN often fails to generate sharp details and introduces undesired patterns that do not exist in style images (e.g., the 1st, 3rd, 6th and 8th rows). LST tends to transfer low-level style patterns like colors but the local details of strokes are often ignored (e.g., the 2nd-5th rows). SANet often generates repetitive patterns in the stylized images (e.g., the 2nd, 5th, 6th and 8th rows). ArtFlow sometimes generates unexpected colors or patterns in relatively smooth regions in some cases (e.g., the 1st-5th and 8th rows). MCCNet can effectively preserve the input content but may fail to capture the stroke details and often generates haloing artifacts around object contours (e.g., the 2nd, 4th, 6th-8th rows). AdaAttN cannot well capture some stroke patterns (e.g., the 1st, 3rd, 4th and 6th rows) and fails to transfer important colors of the style references to the results (the 2nd and 8th rows). Although the generated visual effects of IEST are of high quality, the usage of second-order statistics as style representation causes color distortion (e.g., the 1st row in Figure 2(e) and the 4th row in Figure 5) and cannot capture the detailed stylized patterns (e.g., the regions of sky in the 5th and 7th rows in Figure 5). In particular, these state-of-the-art methods cannot capture the leaving blank characteristic of Chinese painting style in the 1st row of Figure 5 and fail to generate results with a clean background.
2. Quantitative Evaluation
We use the content loss (Li et al. 2017), LPIPS (Chen et al. 2021a), and deception rate (Sanakoyeu et al. 2018b) and conduct two user studies to evaluate our method quantitatively. The two user studies are online surveys that cover art/computer science students/professors and civil servants.
For content loss and LPIPS, we use a pre-trained VGG-19 and compute the average perceptual distances between the content image and the stylized image. The statistics are shown in Table 1. For deception rate, we train a VGG-19 network to classify 10 styles on WikiArt. Then, the deception rate is calculated as the percentage of stylized images that are predicted by the pre-trained network as the correct target styles. We report the deception rate for the proposed CAST and the baseline models in the 2nd column of Table 1. As observed, CAST achieves the highest accuracy and surpasses other methods by a large margin. As a reference, the mean accuracy of the network on real images of the artists from WikiArt is .
We compare CAST with eight state-of-the-art style transfer methods to evaluate which method generates results that are most favored by humans. For each participant, 50 content-style pairs are randomly selected and the stylized results of CAST and one of the other methods are displayed in a random order. Then, we ask the participant to choose the image that learns the most characteristics from the style image. Participants were told that the consistency of content and style was the primary metrics. The style is subjective and the effectiveness of training also depends on their understanding ability. Finally, we collect 3,400 votes from 68 participants. We report the percentage of votes for each method in the 3rd column of Table 1. CAST obtains significantly higher preferences in categories of Sketch, Chinese painting, and Impressionism.
User Study II
We design a novel user study to evaluate the stylized images quantitatively, which is called the Stylized Authenticity Detection (SAD). For each question, we show participants ten artworks of similar styles, including two to four stylized fake painting and ask them to select the synthetic ones. Within each single question, the stylized paintings are generated by the same method. Each participant finished 25 questions. Finally, we collect 2125 groups of results from 85 participants and use the average precision and recall as the measurement for how likely the results will be recognized as synthetics. Table 1 shows the statistics. The paintings generated by CAST have the lowest chance to be decided by people as fake paintings. We also notice that the precision and recall of CAST is less than , which means that users could not tell the real ones from the fakes and tend to select more real paintings as synthetics when doing the testing.
3. Ablation Study
We replace the contrastive style loss with Gram matrix-based perceptual loss, i.e., the model includes perceptual loss, adversarial loss, and cycle consistency loss. As shown in Figures. 6(e) and 6(i), the model using Gram matrix instead of our contrastive style loss cannot capture the stroke characteristics of the style image compared with the full CAST model. The sharp pencil lines of the style image in the 1st row become large black blocks. The textural oil painting strokes of the style image in the 2nd row become smooth blocks, and the vivid colors becomes murky, while unexpected yellow color appears. With the contrastive style loss, our full model can faithfully transfer the brushstrokes, textures, and colors from the input style image.
Domain enhancement.
Our full CAST uses DE for realistic and artistic images separately. We train a simplified CAST model using one discriminator that mixes realistic and artistic images together (mix-DE). As shown in Figure 6(f), the results generated by mix-DE model are acceptable, but the stroke details in the generated images are weaker than the ones by the full CAST model. This fact is due to the existence of a significant gap between the artistic and realistic image domains. We further abandon all images from realistic domain for ablation. As shown in Figure 6(g), the results generated by one-DE model lack details.
Cycle consistency loss.
To better evaluate the improvement of the contrastive style loss on the style transfer task, we exclude the latent promotion of cycle consistency loss from network training. The reason is that the reconstruction process of artistic image may imply style information. We train CAST with an asymmetric cycle consistent loss, which only reconstructs the realistic images. The decoder of the style transfer network is unaffected by the reconstruction of the artistic image. As shown in Figure 6(h), removing realistic image reconstruction will lead to slightly degraded stylization results.
Conclusion and Future Work
In this work, we present a novel framework, namely CAST, for the task of arbitrary image style transfer. Instead of relying on second-order metrics such as Gram matrix or mean/variance of deep features, we use image features directly by introducing an MSP module for style encoding. We develop a contrastive loss function to leverage the available multi-style information in the existing collection of artwork and help train the MSP module and our generative style transfer network. We further propose a DE scheme to effectively model the distribution of realistic and artistic image domains. Extensive experimental results demonstrate that our proposed CAST method achieves superior arbitrary style transfer results compared with state-of-the-art approaches. In the future, we plan to improve the contrastive style learning process by considering artist and category information.