StylizedNeRF: Consistent 3D Scene Stylization as Stylized NeRF via 2D-3D Mutual Learning
Yi-Hua Huang, Yue He, Yu-Jie Yuan, Yu-Kun Lai, Lin Gao
Introduction
Controlling the appearance of complex 3D real scenes has attracted increasing attention in recent years. Numerous works have made great effort to this task, such as texture synthesis and semantic view synthesis . In this paper, we focus on the problem of stylizing complex 3D real scenes, which is useful for applications such as virtual reality and augmented reality. Thanks to the recently advanced 3D representation methods, complex 3D scenes can be represented as point clouds with appearance features or implicit fields by deep neural networks such as neural radiance fields (NeRF) . Compared with point clouds, NeRF can be more reliably obtained from multi-view images, and is continuous in 3D spaces, making learning easier.
In this paper, we aim to stylize a 3D scene following a given set of style examples. This allows generating stylized images of the scene from arbitrary novel views, while making sure rendered images from different views are consistent. To ensure consistency, we formulate the problem as stylizing a NeRF with a given set of style images. Some examples of our NeRF stylization method are presented in Fig. 1.
However, there are two challenges to leverage NeRF as the representation of a complex 3D scene in the task of stylization. Firstly, NeRF needs to query hundreds of sample points along the ray to render a single pixel. The memory limitation makes it intractable to render the whole image or even a big enough patch at one time which is important for calculating content and style losses . Therefore, straightforwardly training a stylized NeRF with perceptual style and content losses on small training patches (3232 for a single RTX 2080Ti GPU) leads to poor stylization results, as shown in Fig. 2. Secondly, directly adopting state-of-the-art image stylization methods to stylize rendered images from NeRF will generate inconsistent results across different views . This is because these stylization methods lack 3D information. Taking a representative Adaptive Instance Normalization (AdaIN) method for example, its results can be seen in the third and fourth columns in Fig. 2. On the other hand, training a NeRF with inconsistent 2D stylized images will cause blurriness in results, which will be further illustrated in Sec. 5.
In order to tackle the problems mentioned above, we propose a novel mutual learning framework between NeRF and a 2D image stylization method. An ordinary NeRF network is first trained to model the opacity field of the scene. The opacity field of NeRF has the inherence of geometric consistency and can estimate the 3D coordinates of the rendered pixels, which is distilled to the 2D stylization method through a consistency loss at a pre-training stage. To represent the stylized scene, we replace the module that predicts color in NeRF with a style module (referred to as stylized NeRF). We then co-train the novel stylized NeRF network (with density prediction fixed) with the pre-trained 2D stylization network for fine-tuning collaboratively. A mimic loss is introduced to align the outputs of the stylized NeRF and the 2D method, aiming to share the stylization knowledge of the 2D method and inherent geometry consistency of NeRF to update networks. However, the 2D stylization method cannot guarantee strict consistency, which leads to ambiguities when transferring a given style among multi-view frames of a certain 3D scene, resulting in blurry results of the stylized NeRF.
Inspired by NeRF-W , our style module takes learnable latent codes as conditioned inputs to handle the ambiguities of the 2D stylized results. Unlike NeRF-W, we build a novel probability model of the latent codes conditional to styles, which enables our model to handle the inconsistency of 2D results and meanwhile to stylize the scene conditionally. We first extract style features of style images with VGG , which are defined as the mean and variance of feature maps along the spatial dimensions . We then encode the style features into latent distributions using a pre-trained variational autoencoder (VAE) . The encoded distributions are conditioned to the encoding style features. Since the inconsistent stylized results generated from the 2D network can be considered to be different samples obeying the distributions conditional to styles, we parameterize 2D stylization results as latent codes obeying the distributions encoded by the corresponding styles. A minus log-likelihood of latent codes is then applied to constrain the conditional probability modeling of latent codes and further ensure the robustness of conditional stylization.
Our main technical contributions are as follows:
We propose a novel stylized NeRF approach for stylizing 3D scenes with given style images, outperforming existing methods in terms of visual quality and 3D consistency.
We propose a mutual learning strategy for the stylized NeRF and 2D stylization method, leveraging the stylization capability of 2D method and geometry consistency of NeRF.
A conditional probability modeling for learnable latent codes is proposed to handle the ambiguities of 2D stylized results while enabling conditional stylization.
Related Work
Novel View Synthesis. Various methods have been proposed to synthesize novel views of a scene with a given set of photographs. Traditional light field techniques interpolate the dense input images as 2D slices of a 4D function to render novel views. Multi Plane Image (MPI) uses RGBD layers of different depths to represent the light field of the scene. Novel views can be rendered by warping the layers and compositing them to an image. Some works utilize explicit 3D proxies, such as meshes , point clouds or voxels to reconstruct the scene. Based on the 3D proxy, existing work combines the geometry with methods representing appearance like colors , texture mapping , light fields or deep networks for neural rendering . Recently proposed image based works estimate the 3D proxy in the form of a point cloud and neurally render the novel views.
A recent trend of continual neural representations is to replace discrete 3D proxy representations with MLPs mapping 3D coordinates to the property of corresponding locations. Implicit functions model the implicit surfaces of the scenes. NeRF further models the radiance fields of the scene as particles emitting and blocking lights. The following works extend NeRF to octree structure , unbounded scenes , reflectance decomposition and uncontrolled real-world images . Please refer to for the recent developments in neural rendering.
Style Transfer Methods. Style transfer is a long-standing research topic in computer vision. Gatys et al. is a pioneering work in this field focusing on an optimization-based scheme. For faster stylization, follow-up works turn to leverage feed-forward neural networks, such as Avatar , and AdaIN . Li et al. proposed a method that embeds whitening and coloring transformation (WCT) to generate high resolution stylized images. Methods based on it sprung up like PhotoWCT , WCT2. Instead of optimizing on pixels, optimizes directly on parameterized brushstrokes. The effect is stunning when the style image is full of delicate strokes.
Video stylization is another topic in this field since it demands consistency between adjacent frames to ensure the stylized video is free of flickering. Most methods are based on optic flow or simply add a temporal constraint to existing methods . The work aligns cross-domain features with input videos and achieves coherent results and dynamically adjusts inter-channel distributions based on relaxation and regularization.
These 2D-based methods lack a spatial consistency constraint and 3D scene perception, thus do not have the ability to maintain long-term consistency in our task. Huang et al. extend stylization to 3D scenes where a 3D scene is represented as a point cloud. It ensures consistency even in 360∘unbounded scenes. firstly introduces NeRF to 3D scene stylization. The NeRF of is trained with style and content losses by sub-sampling patches to cope with the high complexity. As a result, the method tends to lose fine details. Some other works focus on transferring the pose style of 3D models or 2D human skeletons.
Preliminaries
To facilitate the fitting capability of the model, NeRF uses positional encoding to map inputs of the network and to their Fourier features containing signals of muti-scale frequencies:
where is a hyper parameter controlling the spectral bandwidth.
Method
We now illustrate our framework for stylizing a 3D scene with given style images. Given a collection of images from a scene with corresponding camera parameters, our goal is to generate stylized images following the given style from specified novel views while keeping geometry consistency. To achieve this, we propose a mutual learning scheme to optimize the newly introduced stylized NeRF and the 2D stylization network mutually through consistency and mimic losses. Even though the mutually learned stylized NeRF is inherent consistent, 2D stylization network cannot guarantee strict consistency in results, which will still cause blurriness in the results of the stylized NeRF. Therefore, we propose to view the inconsistent 2D stylization results as different samples obeying the distributions conditioned by the style and introduce latent codes that obey such conditioned distributions to handle the inconsistency. We model learnable latent codes with conditional probability through a minus log likelihood loss. In the next, we will first introduce the 2D stylization network we adopt in Sec. 4.1, and then discuss our stylized NeRF in Sec. 4.2. Finally, we describe how we build the mutual learning framework based on the two networks in Sec. 4.3.
We adopt AdaIN as our 2D stylization method, which consists of a VGG encoder, an adaptive instance normalization layer, and a CNN-based decoder. It should be noted that AdaIN is a representative method but may be replaced with other advanced image stylization methods. The feature maps are firstly extracted by the encoder from given input style and content images and then the adaptive instance normalization layer aligns the mean and variance of the content feature maps to the style feature maps. Finally the decoder decodes the aligned feature maps and generates output results with the target style. During the training process, only the decoder of AdaIN is learnable. We pre-train the decoder with distilled 3D consistency knowledge from NeRF through a consistency loss in addition to style and content losses. is calculated by warping stylized images from different views to a fixed one according to the geometry prior from NeRF:
where denotes the stylized result of view and style .
denotes the warping operation from view to view according to the depth estimated by NeRF and denotes the mask of warping and occlusion.
2 Stylized NeRF
An ordinary NeRF is trained to model the opacity field and original radiance color field , which is fixed in the following mutual learning process. To enable the stylization capability of NeRF, an MLP network is added as a style module to NeRF in place of the original color module, modeling the stylized radiance color of the scene. When querying the stylized radiance color of the scene in the training stage, the module takes the input of learnable latent codes in addition to position coordinates. Unlike the latents in NeRF-W that models random appearance and transients of the scene, the latent codes here both learn the style and ambiguities of the 2D stylization results, avoiding blurriness in results of the stylized NeRF and enabling it to conditionally stylize the scene. The stylized results of 2D methods on different views with the specified style can be regarded as samples of a conditional distribution. Those samples are different due to their inconsistency. We parameterize the conditional distributions of 2D stylized results through a pre-trained VAE . The VAE encodes style features extracted by VGG as Gaussian distributions . The conditional distributions of 2D stylized results are parameterized as the embedded Gaussian distributions, which is conditioned with style features. For the 2D stylized result of the -th view and -th style, a latent code initialized by sampling on is assigned to it. The latent codes are optimized during the mutual learning process. To constrain the latent codes to obey the distributions , a minus log likelihood loss is used:
where and are the indices of the training view and style image, respectively. and denote the mean and variance of the distribution embedded by the -th style image. Hence we parameterize the conditional distributions of 2D stylized results by constraining the learnable latent codes to obey the distributions conditioned on styles. At inference time, the mean of embedded distribution is used as input to stylize the scene. The loss constrains the latent codes to obtain better clustering and generalization, thus leads to better results, as we will later demonstrate in Fig. 9.
The style module takes in the embedded and 3D position coordinates to obtain the stylized color . The style module along with the opacity prediction module of the pre-trained NeRF forms our stylized NeRF. The rendering procedure follows Eq. 1, which uses the original opacity field.
3 Mutual Learning
The mutual learning starts with distilling spatial consistency prior knowledge from NeRF to the 2D stylization network through as described in 4.1. It is followed by the collaborative training of the learnable style module, a pre-trained decoder of AdaIN for fine-tuning and latent codes. To augment the training dataset, a series of views are rendered by the ordinary NeRF as training data. We denote the style images as . A training view and a given style together form a training instance, to which a latent code described in Sec. 4.2 is assigned. Using the original opacity, the image of the stylized NeRF is rendered through sampling points along the ray and approximating Eq. 1 by numerical quadrature, as discussed in :
where is the predicted stylized color of the pixel and is the Euler distance between the -th and -th sample points. The mimic loss is defined as the L2 distance between the stylization result from NeRF and from the 2D stylization method:
The mimic loss is introduced to best exchange the knowledge of different strengths between the NeRF and 2D stylization method. The perceptual content loss and style loss are determined by the results of the decoder , which allows larger patches within limited GPU memory. The objective function of the mutual learning process for the NeRF style module and latent codes is:
The objective function for fine-tuning the 2D stylization decoder can be written as
where , , are hyper parameters controlling the impact of terms.
Experiments
We conduct experiments to qualitatively and quantitatively evaluate our method, including comparisons between our method and the state-of-the-art methods of stylization for video and 3D scenes respectively. In quantitative evaluation, a user study is also conducted to collect user preferences, presented in the form of boxplot. We also perform an ablation study on the impact of the ingredients in our method and the effect of training procedure. The hyper parameters of , and are set to 1e-5, 1 and 10 respectively. The style module, latent codes and CNN-based decoder are collaboratively trained for 50k iterations on a single RTX 2080 Ti GPU. The decoder of the 2D stylization method is pre-trained with the consistency loss (Eq. 3) before the mutual learning process for 1k iterations and fixed in the first 20k iterations of the following collaborative training. We test our method on two types of datasets: forward-facing and 360∘unbounded Tanks & Templates (T&T) datasets .
LSNV. In Fig. 4, we qualitatively compare the stylized results of novel views generated by LSNV and our method. The geometry representation of LSNV comes from the voxelized point clouds of COLMAP Structure from Motion (SfM) reconstruction . The discrete representation results in the absence of fine geometry and the loss of precision, which further damages the stylization results. As is shown in framed yellow boxes, the fine-level shapes like the irregular walls, slender poles, thin chain, etc. are broken and the cracks on the truck are lost and filled in the geometry proxy of LSNV. In contrast, our method gives competitive results thanks to its better preservation of geometry.
Video stylization. In Fig. 5, we compare our results in multiple views with two state-of-the-art video stylization methods MCCNet and ReReVST . Due to the lack of space awareness, video stylization methods cannot guarantee the long-term consistency and even violate the geometry of the scene. Results of the fern in the st nd rows give the examples of long-term inconsistency of 2D methods, where stylized color of the other two methods in the framed areas changes obviously across long-term views. In the results shown in the rd th rows, the geometry of the slide is broken and fused with the pillar behind. Compared with video stylization methods, we conclude that our results are more visually consistent.
NeRF-based Stylization. In Fig. 6, we compare our results with , a pioneering work introducing NeRF into stylization. calculates the style and content losses on small sub-sampled patches which approximate large patches. However, such approximation degrades the preservation of content details, as shown in our comparisons. Our method produces results with better detail preservation and less artifacts, due to fundamental technical improvements.
2 Quantitative Results
Consistency Measurement. Following the measurement in , we measure the short and long term consistency using the warped LPIPS metric . A view is warped with the depth expectation estimated by NeRF. The score is formulated as:
where is the warping function and is the warping mask. When calculating the average distance across spatial dimensions in , only pixels within the mask are taken. We compute the evaluation values on 4 scenes in the T&T dataset, using 20 pairs of views for each scene. For each pair, we stylize the images with 10 style images respectively, thus achieve 200 data pairs to evaluate in total. The test views are upsampled three times of the training views to ensure the density of frames for video-based methods. We use view pairs of gap 5() and 35 () for short and long-range consistency calculation. The comparisons of short and long-range consistency are shown in Tab. 1 and Tab. 2, respectively. Our method outperforms other methods by a significant margin.
User study. A user study is conducted to compare the stylization and consistency quality of our method with other state-of-the-art methods. We stylize ten series of views of the 3D scenes in the T&T dataset, using different methods , , and invite 50 participants (including 28 males, 22 females, aged from 18 to 45). First we showed the participants a style image and two stylized videos generated by our method and a random compared method. Then we asked the participants their votes for the video in two evaluating indicators, quality of the stylized results and whether to keep the consistency. We collected 1000 votes for each evaluating indicator and present the result in Fig. 7 in the form of boxplot. Our scores stand out from other methods in both stylization quality and consistency.
3 Ablation Study
The impact of the learnable codes design and mutual training scheme with learnable 2D method. Compared to the choice of shared and fixed style latent codes (w/o LC), applying learnable latent codes (w/ LC) helps to handle the inconsistency of distilled knowledge from the 2D method. On the other hand, to train the decoder (MD) in the mutual learning process is another operation we take to raise the consistency level of 2D method’s outputs and thus to make it easier to train the latent codes. It can be obviously seen that without one or both of these designs mentioned above, the artifacts and blurriness appear in the results as shown in Fig. 8. It demonstrates the necessity and robustness of our design, whose result is of clearer object outlines and more reasonable stylization.
The impact of . At the inference time, the mean code of the encoded distribution is used as input to the style module of NeRF. Fig. 9 compares the inference results with and without . We use green boxes to frame obvious artifacts in the results without in the third column, while results of our complete network in the second column handles this well. This distribution loss constrains the learnable latent codes to obtain better clustering around the mean of the pre-trained distribution, and helps avoid artifacts during inference process.
The choice of outputs. Our method produces two results of 2D method and stylized NeRF at every iteration in a mutual-learning way. We compare the two obtained results as shown in the Fig. 10. Although the mutual learning process makes the 2D method (nd and rd columns) iterates towards a more consistent trend, its consistency is not strict enough and encounters flickering issues between long-range views. On the contrary, the rendered results of stylized NeRF (th column) keep the excellent consistency thanks to its physical volume rendering scheme. Hence we choose the outputs of stylized NeRF as our final results.
Conclusion
We present StylizedNeRF, a novel method for stylizing 3D scenes. A novel mutual-learning framework is proposed to best leverage the stylization ability of 2D method and the spatial consistency of NeRF. A consistency loss distilling spatial consistency prior from NeRF to 2D networks and a mimic loss aligning outputs of 2D networks and stylized NeRF are introduced. To further suppress the inconsistency of 2D method and enable the conditional stylization, we parameterize the inconsistent 2D stylized results as latent codes obeying the distributions conditioned on styles. Our StylizedNeRF outperforms state-of-the-art methods both in terms of visual quality and consistency. In future work, we will implement our proposed approach in Jittor , which is a fully just-in-time (JIT) compiled deep learning framework.
Acknowledgement
This work was supported by the Beijing Municipal Natural Science Foundation for Distinguished Young Scholars (No. JQ21013), the National Natural Science Foundation of China (No. 62061136007 and No. 61872440), Royal Society Newton Advanced Fellowship (No. NAF\R2\192151) and the Youth Innovation Promotion Association CAS.