Stylizing 3D Scene via Implicit Representation and HyperNetwork
Pei-Ze Chiang, Meng-Shiun Tsai, Hung-Yu Tseng, Wei-sheng Lai, Wei-Chen Chiu
Introduction
This paper focuses on the problem of stylizing complex 3D scenes. As shown in Figure 1, given a set of example images of a 3D scene and a reference image with the desired artistic style, we aim to render consistent stylized images at arbitrary novel views. The proposed framework enables various virtual reality (VR) and augmented reality (AR) applications. For instance, with the growing popularity of virtual tours, our method enables seamless switching between real-world scenes and virtual artistic styles, such as walking through the River Seine under Van Gogh’s starry night. 3D style transfer allows us to change the style of a scene and ensures style consistency across view angles.
Numerous efforts have been made for controlling the appearance of the rendered 3D target. For instance, Xiang et al. and Kanazawa et al. formulate it as a texture synthesis problem. Specifically, they render the 3D objects with the desired texture by aligning the coordinates of the 2D UV (texture) map to those of the target object. However, these methods are designed specifically for a single object and are not capable of stylizing complex 3D scenes. On the other hand, PSNet stylizes the point cloud of a 3D scene. Nevertheless, given a set of images of a 3D scene, it requires either the ground-truth 3D geometry or the estimated proxy geometry to build the point cloud. Moreover, the PSNet scheme suffers from the limited resolution issue due to the discrete characteristic of the point cloud representation. In contrast to point clouds that explicitly model 3D scenes, the recent neural radiance fields (NeRF) methods introduce an implicit representation that models a 3D scene using neural networks. Motivated by the high-quality novel view synthesis results, we aim to leverage NeRF for transferring arbitrary styles to complex 3D scenes.
Leveraging NeRF to stylize complex 3D scenes is challenging for two reasons. First, NeRF models lack the controllability to manipulate the appearance of the 3D scene. Since the implicit continuous volumetric representation is built on the deep network with millions of parameters, it is unclear which parameters control the style information of the 3D scene. To overcome this issue, one possible solution is combining existing image/video stylization approaches with novel view rendering techniques by first rendering novel view images and then performing image stylization. However, as shown in Figure 2, current image/video stylization methods do not consider the consistency across different viewpoints for the same scene. We empirically show that the inconsistency issue leads to various problematic results in Section 4. The second challenge is the memory limitation to apply the content and style losses for learning stylization on a NeRF model. Note that these losses are computed across holistic images or patches in order to extract meaningful semantic features. However, the NeRF model requires dense sampling along a camera ray to render a single pixel, which takes significantly more memory to render a patch for computing the losses and back-propagate the gradients (e.g., taking MB to render a patch of size ).
In this work, we propose a 3D scene style transfer approach based on NeRF to address the above mentioned challenges. Our method is able to 1) transfer arbitrary styles to complex 3D scenes, and 2) be optimized with the commonly-used content and stylization losses. The proposed method consists of a NeRF model and a hypernetwork. Our NeRF model has two branches: a geometry branch and an appearance branch. We first optimize the NeRF model to reconstruct the input 3D scene, i.e., learn the implicit scene representation. Then, we fix the parameters of the geometry branch and optimize the hypernetwork to predict the parameters of the appearance branch in order to render the 3D scene with the style of the reference image. Moreover, to alleviate the GPU memory issue, we design a patch sub-sampling algorithm to train the hypernetwork using the content and stylization loss functions. After optimizing for a specific scene, our model is able to 1) render novel views with arbitrary and unseen styles, and 2) generate consistent stylization results across various viewpoints. We evaluate the proposed method with quantitative metrics (e.g., measuring the consistency of stylization across different viewpoints) and subjective user studies, which demonstrate that our method performs favorably against the baseline approaches on rendering more consistent stylization results. The main contributions of this work include:
We propose a 3D scene style transfer approach to render novel views of a complex 3D scene with desired style.
We develop a hypernetwork to control the appearance-related weights of the NeRF model based on the given style image. Our model is universal and able to support arbitrary style images after optimization.
We demonstrate that our method can synthesize stylized images that are consistent across different view angles.
Related Work
Novel view synthesis aims to synthesize a target image at an arbitrary camera pose from a set of source images. Conventional approaches often model a scene with explicit 3D representations, e.g., 3D meshes or 3D voxels based on the multi-view geometry. This line of work relies on a large number of source images to ensure the quality of 3D models. Recently, structure from motion and multi-view stereo techniques are also widely used to build up a 3D model. With the rapid advance of deep learning techniques, several recent approaches learn to estimate the 3D representation of a scene, such as mesh , point cloud , and 3D voxel . However, these methods require supervision from the ground-truth 3D representations and are only able to reconstruct a single object. Another group of works builds the 3D representation without ground-truth supervisions. Image based rendering approaches integrate the image features with the 3D proxy geometry (which is often reconstructed by multi-view stereo approaches), and then warp the input images to synthesize the target view. Different from the explicit representations used in the above schemes, the neural radiance field approaches encode the 3D scene information into a multi-layer perceptron (MLP), which is an implicit 3D representation. This method takes the 3D coordinate and camera view direction as input to directly predict the RGB values and density, which does not rely on any pre-processing to obtain the proxy geometry. In this work, we also learn the implicit 3D representation using the neural radiance field model, but focus on transferring artistic styles to the rendered novel views.
2 Image and Video Style Transfer
Given a content image and a reference style image, style transfer methods aim to synthesize an output image which shows the style of the reference image while preserving the structure of the content image. As a seminal work, Gatys et al. iteratively optimize the output image via a pre-trained model to render the desired style. Afterwards, several methods develop feed-forward networks to significantly reduce the computational cost, but can only transfer a single or a set of pre-determined styles. To achieve arbitrary style transfer (i.e. universal style transfer), recent methods use the adaptive instance normalization (AdaIN) , whitening and coloring transform (WCT) , or linear transformation (LST) . Recently, TPFR approach disentangles an image into style and content codes, and designs a two-stage peer-regularized layer to transfer the target style into the style code of the content image.
On the other hand, applying image style transfer approaches to a video frame-by-frame often results in temporal flickering and instability, as a small perturbation in the input frame may lead to significant changes in the stylized frame. Therefore, video style transfer methods focus on addressing the temporal consistency across the video footage. Existing methods introduce the optical flow to calculate temporal losses or align intermediate feature representations in order to stabilize the model prediction across nearby video frames. Recent efforts further achieve consistent and real-time video style transfer through temporal regularization , multi-channel correlation , and bilateral learning . Although these methods have demonstrated impressive performance, they are designed specifically for stylizing images or video sequences. Since the consistency across various viewpoints of the same scene is not considered, directly using existing image/video stylization schemes for our problem has various issues (see Figure 2 and Section 4).
3 Texture Transfer
Texture transfer aims to change the texture or style of a 3D object while keeping the appearance consistent across different view angles. With the recent advance of deep learning techniques, different methods are proposed to learn the correspondences between 3D shape (e.g. meshes) and the texture space for realizing the texture transfer. For instance, Xiang et al. learn a mapping from the implicit 3D representation to the 2D texture map, such that one can change the appearance of a 3D model by swapping the 2D texture map. However, such a texture mapping approach may generate unnatural results if the texture image is mapped across object edges or boundaries. In addition to texture mapping, recently there are several works proposed to tackle the style transfer task on the 3D representations. For instance, Kato et al. propose to build a neural renderer where the rendering is integrated into neural networks, and demonstrate its application of style transfer on 3D models. PSNet performs the style transfer on the point cloud data via manipulating the point cloud features in latent space to change the style. However, these methods rely on explicit 3D representations, such as mesh and point cloud, and are limited to the object level instead of the entire scene. Moreover, they do not support synthesizing stylized images with high-quality, which further limits their applications in the real world (e.g. AR). In this work, we resort to the implicit 3D scene representation and focus on transferring style for real-world 3D scenes with a complex background to generate the stylized novel views.
Proposed Method
Given a set of images/photos of a static 3D scene taken from different camera poses , our goal is to render arbitrary novel views of the 3D scene with the style extracted from a reference image . The rendered images should have consistent appearance and stylization effect across different views. To this end, we propose a 3D style transfer method to enable the universal stylization of a complex 3D scene. We model a 3D scene with implicit representation by the neural radiance fields (Section 3.1), and learn to transfer arbitrary style using a hypernetwork (Section 3.2). To alleviate the training difficulties, we propose a two-stage training pipeline, where the geometric training stage learns the implicit representation of a 3D scene by disentangling the geometry and appearance into two branches, and the stylization training stage learns to predict the parameters of the appearance branch from the reference image (Section 3.3).
The model of neural radiance fields (NeRF) proposed in adopts a sparse set of input views of a 3D scene for learning to optimize the underlying continuous volumetric scene function. The basic idea behind NeRF can be illustrated in Figure 3(a). Given a camera observing the 3D scene at the viewing direction , we first march along the rays back-projected from the camera center through all the pixels on the image plane for obtaining the samples of 3D points . The scene function as an implicit scene representation then takes both and as input to output the volume density at and the corresponding RGB color emitted towards the viewing direction . To be detailed, the scene function consists of three multilayer perceptrons (MLPs): , , and . In practice, takes as input where the produced is either passed through to obtain the volume density , or further processed by together with to obtain the view-dependent color via . With accumulating the colors and densities via the volume rendering techniques, the high-quality 2D images as the observation of the 3D scene from various views can be generated. Please note that in practical implementation both and are firstly transformed into positional embeddings before being utilized by the MLPs (i.e. and ).
While the original NeRF model seems to demonstrate compelling capability in the view synthesis, when it is adopted to tackle the captures of unbounded and complex scenes, simultaneously modelling the nearby and far objects (related to foreground and background respectively) by the same volumetric scene function would cause problems for volume rendering as being required to handle the large dynamic depth range between objects, as pointed out by . To deal with such issue, NeRF++ separates the foreground and background objects into the inner volume and outer volume by a unit sphere, with having two NeRFs adopted to model them respectively, where an inverted sphere parameterization is particularly applied on the coordinate system of outer volume for bounding the unlimited distance between the background objects and the origin. In our work, we hence adopt the model of NeRF++ to represent the unbounded and complex 3D scenes.
2 Stylizing Implicit Representations
Without loss of generality, the operation of style transfer aims to keep the geometry/content of the target scene while modifying its appearance/texture according to the reference style image. From the design of NeRF++ models, we can see that the geometry and the appearance information of the scene are respectively represented by the volume densities and the view-dependent color values at each 3D location. Therefore, to stylize the 3D scene implicitly encoded by NeRF++, we propose a novel approach to modify the parameters of the MLP (i.e. , the appearance branch) which is responsible for predicting the color values, while keeping the other MLPs (i.e. and , the geometry branch) fixed to retain the geometry of the target scene. Basically, our stylization method on the NeRF-based scene representation is realized by two components: the style variational autoencoder (i.e., style-VAE) to extract the style latent vector from the style image , and the hypernetwork to estimate the weights for modifying according to (as illustrated in Figure 3). We detail these two components in the supplementary material.
3 Model Training
The training procedure for our 3D scene style transfer approach contains the geometric training stage and the stylization training stage. We describe these two stages and the corresponding objective functions in the following.
In this stage we aim to learn a NeRF-based representation of the target 3D scene from the given images taken from different camera poses . The NeRF++ model is adopted and its training process is briefly summarized as follows. At each optimization iteration, we randomly sample pixels from input images to form a batch of camera rays and 3D points (marching along the rays) based on the corresponding camera poses and camera intrinsic. With taking these 3D points and their corresponding view directions (i.e. the directions of the corresponding camera rays) as input to the scene function of NeRF++ to obtain the output set of volume densities and colors, the volume rendering technique is used to render the color value of each ray . The objective for training the MLPs of the scene function (i.e. , , and ) is the mean square error (MSE) between and the groundtruth color of the corresponding image pixel of ray .
Stylization Training Stage (Second Stage).
As in the previous geometric training stage, we have encoded the complete geometry and the original appearance of the target 3D scene into the NeRF++ model, now in the stylization training stage we focus on learning the hypernetwork to predict from the style latent vector the weights for updating MLP , where is extracted from the style reference image by the pre-trained encoders of style-VAE, i.e. . Note that as the geometry of the 3D scene should be retained during the stylization, both and are kept fixed during this training stage.
where and denote the mean and standard deviation respectively, and denotes the feature representation obtained from the -th layer of an ImageNet-pretrained VGG-19 network, basically relu1_1, relu2_1, relu3_1, and relu4_1 layers are used.
The overall objective function for the stylization training stage, in which the gradients are back-propagated to learn the hypernetwork , is then defined as:
where the hyperparameter controls the balance between the content loss and style loss, and we set for all our experiments.
Implementation Details.
In the geometric training stage, our NeRF model is trained for iterations and we set ; while in the stylization training stage, hypernetwork is trained for iterations and we set and to and . We adopt the Adam optimizer for both stages with learning rates set to and , respectively. Following , each style image used in our experiments is resized and randomly cropped to be the size of .
Experimental Results and Analysis
In this section, we present qualitative and quantitative results to validate the effectiveness of our proposed framework. Please refer to our project page1 for source code, pre-trained model and more qualitative results.
We conduct the experiments using five real-world 3D scenes collected in the Tanks and Temples dataset, i.e., Family, Francis, Horse, Playground and Truck. During the geometric training stage, we follow to use the COLMAP SfM method to estimate the camera poses and intrinsics of the input images for each 3D scene. On the other hand, we use images in the WikiArt dataset as the reference style images. Specifically, we randomly select images as the testing data and keep the others (i.e., totally ) for the stylization training stage.
Compared Methods.
To the best of our knowledge, there is no existing method that focuses on stylizing complex 3D scenes. Therefore, we combine different image/video stylization methods with the novel view synthesis (NVS) algorithms to build three types of the baseline approaches:
Image stylization NVS: we stylize the input images of the target scene, then perform novel view synthesis.
NVS image stylization: we perform image stylization on the novel view synthesis results.
NVS video stylization: we treat a series of novel view synthesis results (generally along a smooth camera path) as the video, then perform video stylization.
Specifically, we use NeRF model described in Figure 3 (a) as the novel view synthesis approach. The AdaIN , WCT , LST , and TPFR schemes are used for image stylization. Finally, we use two video stylization frameworks, i.e., ReReVST and MCCNet .
1 Qualitative Results
We present the qualitative comparisons in Figure 5 and 8. The baseline “image stylization NVS” produces blurry results, as shown in Figure 5. Since the input images are processed independently by the AdaIN approach, the stylized images are not consistent across different views of the same scene. Therefore, the optimization of the NeRF model with these inconsistent images leads to blurry results. Moreover, as the NeRF model is optimized for a specific style, this baseline method is not capable of transferring arbitrary style to the 3D scene. On the other hand, the baseline “NVS image stylization” also produces inconsistent results across different viewpoints, as highlighted in the red boxes in Figure 8. In particular, the baseline based on the WCT approach fails to preserve the content of the original 3D scene, while the other one based on the TPFR scheme does not transfer the desired style provided by the reference image. In contrast, the results synthesized by our method not only match the desired style, but also are consistent across various novel views.
We demonstrate the qualitative results by the baseline “NVS video stylization” in Figure 8. Although these video-stylization-based methods are trained to consider the short-term consistency, they fail to produce consistent results between two far-away viewpoints due to the error accumulation, e.g., the head and the back of the statue. In contrast, since our framework is trained to stylize the holistic 3D scene, it generates results that are consistent between both short-range or long-range viewpoints. More stylized results of our proposed method are provided in Figure 9.
2 Quantitative Results
To evaluate the quality of stylizing complex 3D scenes, we conduct a study to understand the user preference between the results rendered by the proposed and baseline methods. Here we focus on the comparison against “NVS image stylization” and “NVS video stylization” baselines as the results of “image stylization NVS” ones are generally blurry as shown in Figure 5.
There are users participated in this study. For each user, there are tests conducted for each comparison (i.e. proposed method versus one baseline in terms of stylization quality or temporal consistency). As shown in Figure 6, our proposed method performs favorably against the baseline schemes in terms of both the stylization quality and consistency. We also observe that the “NVS video stylization” baseline produces videos with less flickering compared to the “NVS image stylization” ones since they consider the temporal consistency. However, these approaches fail to preserve the consistency between two far-away viewpoints, as demonstrated in the following experiments.
Consistency.
In addition to the user preference study that evaluates the quality of the stylization results, we use the metric from Lai et al. to measure the consistency between different stylized novel view images. More details of the consistency metric are provided in the supplementary materials. In the following experiments, we evaluate the consistency from two different perspectives: 1) the short-range consistency between nearby novel views, and 2) the long-range consistency between far-away novel views.
Table 1 shows the short-range consistency scores. In this experiment, we use every two adjacent novel views, i.e., the and frames in the testing videos, to compute the consistency score. We observe that the results generated by the image stylization baseline methods are not consistent as the novel view images are processed independently. Moreover, while the TPFR approach achieves the lowest scores among all the baseline schemes, it fails to capture the desired style of the reference image in some cases, as shown in Figure 8.
We present the long-range consistency score in Table 2. Specifically, we use every two far-away views, i.e., the and frames in the testing videos, to compute the consistency score. Since the distance between two views is larger, the consistency scores of all methods in this experiment are higher than those in the short-range study. Although the video stylization baselines generally better preserve the short-range consistency than the image stylization ones, they fail to maintain the consistency between two far-away views due to the error accumulation. In contrast, the proposed method is capable of synthesizing images that are both short-range and long-range consistent.
3 Limitations
The quality of the stylization results is limited by the backbone NeRF model. As the red boxes shown in Figure 7, the proposed method produces blurry stylization results since the backbone model fails to capture the details of the trees. In contrast, the details of both the original and stylized wheels demonstrated in the green boxes are clear.
Conclusions
In this paper, we propose a NeRF model for transferring arbitrary styles to complex 3D scenes. We design a hypernetwork to predict the appearance-related parameters in the NeRF model to stylize the 3D scene according to the input reference (style) image. In addition, we develop a two-stage training strategy along with the patch sub-sampling algorithm to learn the hypernetwork. Qualitative and quantitative results validate that the proposed method renders high-quality novel view images with the desired style.
Acknowledgement. This project is supported by MediaTek Inc. and MOST 110-2636-E-009-001. We are grateful to the National Center for High-performance Computing for computer time and facilities.