ARF: Artistic Radiance Fields

Kai Zhang, Nick Kolkin, Sai Bi, Fujun Luan, Zexiang Xu, Eli Shechtman, Noah Snavely

Introduction

Creating artistic images often requires a significant amount of time and special expertise. Extending an artwork to dimensions beyond the 2D image plane, such as time (in the case of animation), or 3D space (in the case of sculptures or virtual environments), introduces new constraints and challenges. Hence, the styles employed by artists when moving their work beyond a static 2D canvas are constrained by the effort required to create a consistent visual experience.

We propose Artistic Radiance Fields (ARF), a novel approach that can transfer the artistic features from a single 2D image to a full, real-world 3D scene, leading to artistic novel view renderings that are faithful to the style image. Our method converts a photorealistic radiance field reconstructed from multiple images of complex, real-world scenes into a new, stylized radiance field that supports high-quality view-consistent stylized renderings from novel viewpoints, as shown in Fig. 1. The quality of these renderings is in contrast to previous 3D stylization works that often suffer from geometrically inaccurate reconstructions of point cloud or triangle meshes and the lack of style details.

We formulate the stylization of radiance fields as an optimization problem; we render images of the radiance fields from different viewpoints in a differentiable manner, and minimize a content loss between the rendered stylized images and the original captured images, and also a style loss between the rendered images and the style image. While previous methods apply the commonly-used Gram matrix-based style loss for 3D stylization, we observe that such a loss leads to averaged-out style details that degrades the quality of the stylized renderings.

This limitation motivates us to apply a novel style loss based on Nearest Neighbor Feature Matching (NNFM) that is better suited to the creation of high-quality 3D artistic radiance fields. In particular, for each feature vector in the VGG feature map of a rendered image, we find its nearest neighbor (NN) feature vector in the style image’s VGG feature map and minimize the distance between the two feature vectors. Unlike a Gram matrix describing global feature statistics across the entire image, NN feature matching focuses on local image descriptions, better capturing distinctive local details. Coupled with our style loss, we also enforce a VGG feature-based content loss – that balances stylization and content preservation – and a simple color transfer technique – that improves the color match between our final renderings and the input style.

Volumetric radiance field rendering consumes a lot of memory and often can only regress sparsely sampled pixels during training, and not the full images necessary for computing the VGG features used in many style losses. We contribute a practical innovation that allows us to perform optimization on high-resolution images. In particular, we devise a method we call deferred back-propagation that enables memory-efficient auto-differentiation of scene parameters with image losses computed on full-resolution images (e.g., VGG-based style losses) by accumulating cached gradients in a patch-wise fashion.

We demonstrate that ARF can robustly transfer detailed artistic features from diverse and challenging 2D style exemplars to a variety of complex 3D scenes, resulting in significantly better visual quality compared to previous methods, which tend to yield over-smoothed and blurry stylized novel views (see Figures 4, 5, and 6). In our user studies, our method is also consistently preferred over baselines.

A novel radiance field-based approach for 3D scene stylization that can faithfully transfer detailed style features from a 2D image to a 3D scene and produces consistent stylized novel views of high visual quality.

We find that Nearest Neighbor Feature Matching (NNFM) loss better preserves details in the style images than the Gram-matrix-based loss commonly used in prior 3D stylization works.

A deferred back-propagation method for differentiable volumetric rendering, allowing for computation of losses on full-resolution images while significantly reducing the GPU memory footprint.

Related Work

In this section, we review related work to provide context for our own work.

Image style transfer. Since Gatys et al. introduced neural style transfer, significant progress has been made towards artistic stylization , image harmonization , color matching , texture synthesis and beyond . These style transfer approaches leverage features extracted by a pre-trained convolutional neural network (e.g., VGG-19 ) and optimize for a set of loss functions (typically a content loss capturing an input photo’s features and a style loss matching a target image’s feature statistics, e.g., encoded in a Gram matrix) to achieve good performance for painterly style transfer. Depending on whether the style transfer is achieved via iterative optimization on a single input or with a forward pass from a pre-trained generative model, existing methods can be categorized as optimization-based and feed-forward-based:

Optimization-based style transfer. Gatys et al. perform style transfer via iterative optimization to minimize content and style losses. Many follow-up works have investigated alternative style loss formulations to further improve the quality of semantic consistency and high-frequency style details like brushstrokes. Unlike neural style transfer methods that encode statistics of style features with a single Gram matrix, Chen and Schmidt , CNNMRF , Deep Image Analogy and NNST propose to search for nearest neighbors and minimize distances between features extracted from corresponding content and style patches in a coarse-to-fine fashion. These methods achieve impressive 2D stylization quality when provided with source and target images that share similar semantics. Our approach draws inspiration from this line of work and is the first to introduce nearest neighbor feature matching (NNFM) for 3D stylization. Our NNFM loss is most similar to that proposed in for 2D style transfer. However, when stylizing 3D radiance fields, we find that we can achieve the same level of stylistic detail more efficiently by only applying stylization at the final scale (as opposed to coarse-to-fine) and by skipping the style image augmentations (rotation and/or scaling) used in .

Feed-forward style transfer. Rather than performing iterative optimization, feed-forward approaches train neural networks that can capture the style information of the style image and transfer it to the input image using a single forward pass. While fast, these methods often struggle to faithfully reproduce stylistic feature like colors and brushstrokes, and yield lower visual quality compared to optimization-based techniques. For the sake of creating high-quality artistic radiance fields, we do not pursue this direction as a component of ARF.

Video style transfer. Stylizing each video frame separately with a 2D style transfer method often leads to flickering artifacts in the resulting stylized videos. Video style transfer techniques address this problem by enforcing an additional temporal coherency loss across frames . Alternative approaches rely on aligning and fusing style features according to their similarity to content features to maintain temporal consistency. Despite sharing the similar challenge of consistency across views, stylizing a 3D scene is a distinct problem from video stylization, because it requires synthesizing novel views while maintaining style consistency, which in turn is best achieved through stylization in 3D rather than 2D image space.

3D style transfer. 3D style transfer aims to transform the appearance of a 3D scene so that its renderings from different viewpoints match the style of a desired image. Previous approaches represent real world scenes using point clouds or triangle meshes . For example, Huang et al. and Mu et al. use featurized 3D point clouds modulated with the style image, followed by a 2D CNN renderer to produce stylized renderings. Yin et al. create novel geometric and texture variations of 3D meshes by transferring the shape and texture style from one textured mesh to another. The performance of such methods is limited by the quality of the geometric reconstructions, which oftentimes contain noticeable artifacts for complex real-world scenes. In contrast, we perform style transfer on radiance fields which have been shown to more faithfully reproduce the appearance of real world scenes. A work closely relevant to ours is that of Chiang et al. , who apply neural radiance fields for scene representation and rely on pre-trained style hypernetworks for appearance stylization. However, their method produces over-smoothed and blurry stylization results, and cannot capture the detailed structures of the style image such as brushstrokes, due to the limitation of pre-trained feed-forward models. We show that our approach can more faithfully capture distinctive details in the style exemplar while preserving recognizable scene content.

Background of Radiance Fields

NeRF proposes neural radiance fields to model and reconstruct real scenes, achieving photo-realistic novel view synthesis results. In general, the radiance field representation can be seen as a 5D function that maps any 3D location x\mathbf{x} and viewing direction d\mathbf{d} to volume density σ\sigma and RGB color c\mathbf{c}:

This representation can be rendered from any viewpoint via differentiable volume rendering, and hence can be optimized to fit a collection of input photos captured from multiple views, and then later used for synthesizing photo-realistic novel views. We move beyond photo-realism and add an artistic feel to the radiance field by stylizing it using an exemplar style image, such as a painting or sketch.

Stylizing Radiance Fields

In this section, we describe our radiance fields stylization technique in detail. Given a photo-realistic radiance field reconstructed from photos of a real scene, our approach transforms it into an artistic style by stylizing the 3D scene appearance with a 2D style image. We achieve this by fine-tuning the radiance field using a novel nearest neighbor feature matching style loss (Sec. 4.1) that can transfer detailed local style structures. We also introduce a deferred back-propagation technique that enables radiance field optimization with full-resolution images (Sec. 3) in the face of limited GPU memory. We apply a view-consistent color transfer technique to further enhance our final visual quality (Sec. 4.3).

Artwork often features unique visual details; for instance, the Van Gogh’s The Starry Night is characterized by long and curvy brushstrokes. In general, neural features produced by pre-trained neural networks (like VGG) can effectively capture such details and have been successfully used for 2D style transfer . However, it is challenging to transfer such rich visual details to 3D scenes using prior VGG-based style losses, since the style information measured by such losses are generally based on global statistics that do not necessarily capture the local details well in a view-consistent way.

To address this, we propose to use the Nearest Neighbor Feature Matching (NNFM) loss to transfer complex high-frequency visual details from a 2D style image to a 3D scene (parameterized by a radiance field), consistently across multiple viewpoints. In particular, let Istyle\bm{I}_{\textrm{style}} denote the style image, and Irender\bm{I}_{\textrm{render}} denote an image rendered from the radiance field at a selected viewpoint. We extract the VGG feature map Fstyle\bm{F}_{\textrm{style}} and Frender{\bm{F}_{\textrm{render}}} for Istyle\bm{I}_{\textrm{style}} and Irender{\bm{I}}_{\textrm{render}}, respectively. Let Frender(i,j){\bm{F}}_{\textrm{render}}(i,j) denote the feature vector at pixel location (i,j)(i,j) of the feature map Frender{\bm{F}}_{\textrm{render}}. Our NNFM loss can be written as:

where NN is the number of pixels in Frender{\bm{F}_{\textrm{render}}}, and D(v1,v2)D(\bm{v}_{1},\bm{v}_{2}) computes the cosine distance between two vectors v1,v2\bm{v}_{1},\bm{v}_{2}:

In short, for each feature in Frender{\bm{F}_{\textrm{render}}}, we minimize its cosine distance (Eq. (3)) to its nearest neighbor in the style image’s VGG feature space (Fstyle\bm{F}_{\textrm{style}}).

Note that our loss does not rely on global statistics. This grants more flexibility to the optimization process, which can focus on adjusting the local scene appearance to perceptually match the style image in a given image rendered from a given training viewpoint.

where λ\lambda is a weight controlling stylization strength: a larger λ\lambda preserves more content, while a smaller λ\lambda leads to stronger stylization. Note that Frender,Fstyle,Fcontent{\bm{F}_{\textrm{render}}},\bm{F}_{\textrm{style}},\bm{F}_{\textrm{content}} are extracted by exactly the same feature extractors. (See Sec. 4.4 for details)

2 Deferred back-propagation

We propose a simple technique termed deferred back-propagation that can directly optimize on full-resolution images, allowing for more sophisticated and powerful image losses to be used in practice with radiance fields representation. As shown by Fig. 3, we first render a full-resolution image with auto-differentiation disabled; then we compute the image loss and its gradient with respect to the rendered image’s pixel colors, which produces a cached gradient image; finally, in a patch-wise manner, we re-render the pixel colors with enabled auto-differentiation, and back-propagate the cached gradients to the scene parameters for accumulation. In this way, gradient back-propagation is deferred from the full-resolution image rendering stage to the patch-wise re-rendering stage, reducing the GPU memory cost from that of rendering a full-resolution image to that of rendering a small patch. In our work we apply this technique for the stylization task by optimizing our style loss, and also by optimizing the standard Gram loss for comparison (see Fig. 7).

3 View-consistent color transfer

While our style and content losses can perceptually transfer styles and preserve the original content, we find they can lead to color mismatches between rendered images and the style image. We devise a simple technique to address this issue which leads to much better stylization quality (see Fig. 7). We first recolor the training views via color transfer from the style image. These recolored images are used to pre-optimize our artistic radiance field as initialization for our stylization optimization based on Eq. 4. These color transferred images are also used for our content preservation loss. Additionally, after the 3D stylization process, we again perform a color transfer to images rendered to the training viewpoints, and apply the same color transformation directly to the color values produced from rendering the radiance fields.

4 Implementation details

To represent a radiance field, our work primarily uses the recently-proposed Plenoxels for its fast reconstruction and rendering speed. However, our framework is agnostic to the radiance field representation. To demonstrate this, we apply our proposed techniques to stylize NeRF and TensoRF , and in each case achieve high visual quality with faithful style transfer, as shown in Fig. 8.

During stylization, we fix the density component of the initial photorealistic radiance field, and only optimize the appearance component when converting to an artistic radiance field. We also discard the view-dependent appearance modelling,In Plenoxels, radiance fields are represented as a mixture of spherical harmonics at each point. To discard view-dependence, we simply move all spherical harmonics components except the first one. For TensoRF, we zero out the view directions when inputting them to the MLP. To extract feature maps Frender,Fstyle,Fcontent{\bm{F}_{\textrm{render}}},\bm{F}_{\textrm{style}},\bm{F}_{\textrm{content}} in Eq. 4, we use a pretrained VGG-16 network that consists of 5 layer blocks: conv1, conv2, conv3, conv4, conv5.In VGG-16, each layer block begins with a max-pooling layer that downsamples the feature map by 2. Inside each layer block, feature maps are of the same spatial resolution and hence can be concatenated to form a single feature map for this block. We use the conv3 block as the feature extractor, because we empirically find that it captures style details better than the other blocks, as shown in Fig. 9. We set the content-preserving weight λ=0.001\lambda=0.001 in Eqn. 4 for all forward-facing captures, and λ=0.005\lambda=0.005 for all 360∘ captures. At each stylization iteration, we render an image for computing losses from a viewpoint randomly selected out of all the training viewpoints used for reconstructing photo-realistic radiance fields. We perform the stylization optimization for 10 epochs with learning rate exponentially decayed from 1e-1 to 1e-2.

Experiments

We evaluate our method by performing both quantitative and qualitative comparisons to baseline methods. We show stylization results for various real world scenes guided by different style images. The experimental results show that our method significantly outperforms baseline methods in terms of generating stylized renderings that are more faithful to the input style image, while maintaining the recognizable semantic and geometric features of the original scene. We invite readers to watch our supplemental videos for better assessment of 3D stylization quality.

Datasets. We conduct extensive experiments on multiple real-world scenes including four forward-facing captures: Flower, Orchids, Horns, Trex, from , and seven 360∘ captures: Family, Horse, Playground, Truck, M60, Train from the Tanks and Temples dataset , as well as the Real Lego dataset from . All scenes contain complex structures and intricate details that are difficult to reconstruct with previous triangle mesh or point cloud-based methods. We also experiment with a diverse set of style images including a neon tiger, Van Gogh’s The Starry Night, sketches, etc., to test our method’s ability to handle a diverse range of style exemplars.

Baselines. We compare our method to state-of-the-art methods for 3D style transfer quality. Specifically, Huang et al. adopt point clouds featurized by VGG features averaged across views as a scene representation, and transform the pointwise features by modulating them with the encoding vector of a style image for stylization. Chiang et al. use implict MLPs as in NeRF++ to reconstruct a radiance field for a scene, then update the weights of the radiance prediction branch using a hypernetwork that takes a style image as input. For both methods, we use their released code and pre-trained models. We chose not to compare to off-the-shelf video stylization methods, because prior work has demonstrated that they are less competitive compared to 3D style transfer approaches .

Qualitative comparisons. We show visual comparisons between methods in Fig. 4 (forward-facing captures) and Fig. 5 (360∘ captures). Visually, we see that our results exhibit a better style match to the exemplar image compared to the baselines. For instance, in the Flower scene in Fig. 4, our method faithfully captures both the color tone and the brushstrokes of The Starry Night, while the baseline method of Huang et al. generates over-smoothed results without less detailed structures. Moreover, Huang et al. also fails to recover complex geometric structures such as leaves of the plants due to inaccuracies in the reconstructed point cloud. In comparison, our method effectively reconstructs and preserves the geometric and semantic content of the original scene, thanks to the more robust radiance fields representation.

Chiang et al. only transfers the overall color tone of the style image to the scene and fails to recover the rich details that our method does. For example, in the Family statue scene in Fig. 5, our method captures the subtle textural details of the watercolor feather style image, and reproduces them in the stylized renderings. In contrast, the method of Chiang et al. generates blurry results with no such intricate structures, because their hypernetwork is trained on a fixed dataset of style images and often fails to capture the details of an unseen style input. Our method benefits from both the optimization-based framework as well as our NNFM style loss, which greatly boost the 3D stylization performance.

We show additional results from our method in Fig. 6. Our method is robust to different scenes with varying levels of complexity and also generates consistently superior results under a variety of styles. We refer the readers to the supplementary videos for more visual comparisons and results.

User study. We also perform a user study to compare our methods to baseline methods. A user is presented with a sequence of stylization results, where for each result the user is shown a style image, a video of the original scene, and two corresponding stylized videos produced with our method and a baseline method. The user is then asked to select the result that better matches the style of the given style image. In total, we collect ratings covering 25 randomly selected (scene, style) pairs. We divided the questions into 5 batches, each with 5 questions, and asked a large group of users to rate a randomly selected batch. We collected an average of ∼\sim12 ratings for each individual pair. We found that users prefer our method over the baseline Huang et al. 86.8%86.8\% of the time, and over the baseline Chiang et al. 94.1%94.1\% of the time. These results show a clear preference for our method.

Ablations. We perform ablation studies to justify our design choices. We first compare our NNFM loss to the prior Gram-based and CNNMRF losses. As we can see from Fig. 7, our NNFM loss generates significantly better results and more faithfully preserves the style details of the example images compared to the other two losses. In Fig. 7, we also validate the necessity of the color transfer stage. Without color transfer, the generated results tend to have different color tones from the style images, leading to a degraded style match. Our color transfer method effectively addresses this issue. Finally, we perform an ablation of using the feature map at different layers of the VGG-16 network for computing NNFM loss in Fig. 9. We find that our choice of the conv3 layer block preserves stylistic details better than other layers.

Our method has a few limitations. First, geometric artifacts, e.g., floaters, in the radiance fields can cause artifacts in both the photorealistic and our stylized renderings. Such floaters can be removed by adding additional regularizers on volume density to the loss during optimization . Second, although our artistic radiance fields can be rendered in real time once optimized, a relatively time-consuming optimization procedure is still required for every style image (∼\sim3 mins for forward-facing captures, and ∼\sim20 mins for 360 captures on a single NVIDIA RTX 3090 GPU). Third, our reconstructed artistic radiance fields do not support manual editing. Enabling artists to interactively edit them is highly desirable for the sake of facilitating creativity.

Conclusion

We have presented a method to reconstruct artistic radiance fields from photorealistic radiance fields given user-specified style exemplars. The reconstructed artistic radiance fields can then be used to render high-quality stylized novel views that faithfully mimic the input style image in terms of color tone and style details like brushstrokes, enabling an immersive experience of an artistic 3D scene. Key to our method’s success is the proposed coupling of the nearest neighbor featuring matching loss and view-consistent color transfer, rather than the commonly-used Gram loss. We demonstrate that our method achieves superior 3D stylization quality over baselines through evaluations across various 3D scenes and 2D styles.

We would like to thank Adobe artist Daichi Ito for helpful discussions about 3D artistic styles.

References

Appendix