Unified Implicit Neural Stylization

Zhiwen Fan, Yifan Jiang, Peihao Wang, Xinyu Gong, Dejia Xu, Zhangyang Wang

Introduction

Implicit Neural Representation (INR) has gained remarkable popularity in representing concise signal representation in computer vision and computer graphics . As an alternative to discrete grid-based signal representation, implicit representation is able to parameterize modern signals as samples of a continuous manifold, using multi-layer perceptions (MLP) to map between coordinates and signal values. Several seminal works have verified the effectiveness of INR in representing image, video, and audio. Followups further apply INR to more challenging tasks including novel-view synthesis , 3D-aware generative model , and inverse problem .

While implicit neural representation reveals multiple advantages compared to conventional discrete signals, a general question of curiosity might be: which and how modern visual signal processing approaches/tasks designed for discrete signals can also be applied to continuous representations? Research pursuing this answer has been conducted on implicit neural representation since its origin. Chen et al. apply a local implicit function to image super-resolution and they observe that INR can surpass bilinear and nearest upsampling. Sun et al. demonstrate the effectiveness of INR in the context of sparse-view X-ray CT. Dupont et al. propose to store the weights of a neural implicit function instead of pixel values, which surprisingly outperforms JPEG compression format. further demonstrates superior video compression using similar ideas.

We investigate a novel setting: to yield visually pleasing stylized examples under various 2D and 3D scenarios, using a generalized approach leveraging implicit neural representations. Note that, training a stylized implicit neural representation still faces many hurdles. On one hand, the aforementioned works mostly have the access to dense measurements or at least sparse clean data, which enables training an implicit neural network under the supervision of target signal. In contrast to those tasks/approaches, current image stylization mechanisms are mostly conducted in an unsupervised manner, due to the absence of stylized ground truth data. Consequently, it is still unknown whether coordinate-based MLP can be optimized without accessing corresponding ground truth signals. On the other hand, existing style images are mostly based on 2D scenes, which raises obstacles when being considered as the appearance of 3D implicit representation. Prior art attempted on marrying stylization with one specific type of Implicit Neural Representation, the neural radiance field (NeRF) . Nevertheless, it still captures the statistics of style information by a series of pre-trained convolution-based hypernetwork to generate model weights, rather than a direct implicitly encoding stylization. As indicated by recent literature , training a robust hypernetwork requires a large amount of training samples, while novel-view synthesis tasks commonly hold no more than hundreds of views, potentially jeopardizing the synthesized visual quality.

To conquer the aforementioned fragility, we propose a Unified Implicit Neural Stylization framework, coined as INS. Different from the vanilla implicit function which is built upon a single MLP network, the proposed framework divides an ordinary implicit neural representation to multiple individual components. Concretely speaking, we introduce a Style Implicit Module to the ordinary implicit representation, and coin the later one as Content Implicit Module in our framework. During the training process, the stylized information and content scene are encoded as one continuous representation, and then fused by another Amalgamation Module. To further regularize the geometry of given scenes, we utilize an additional self-distilled geometry consistency loss on top of the rendered density, for the stylization of NeRF. Eventually, INS is able to render view-consistent stylized scenes from novel views, with visually impressive texture details: a few examples are shown in Figure. 1.

We propose INS, a unified implicit neural stylization framework, consists of a style implicit module, a content implicit module, and an amalgamation module, which enables us to synthesize promising stylized scenes under multiple 2D and 3D implicit representations.

We conduct comprehensive experiments on several popular implicit representation frameworks in this novel stylization setting, including 2D coordinate-based framework (SIREN ), Neural Radiance Field (NeRF ), and Signed Distance Functions (SDF ). The rendering results are found to be more consistent, in both shape and style details, from different views.

We further demonstrate that INS is able to learn representations that are continuous not only with regard to spatial placements (including views), but also in the style space. This leads to effortlessly interpolating between different styles and generating images rendered by the new mixed styles.

Related Works

Recent research has exhibited the potential of Implicit Neural Representation (INR) to replace traditional discrete signals with continuous functions parameterized by multilayer perceptrons (MLP), in computer vision and graphics . The coordinate-based neural representations have become a popular representation for various tasks such as representing image/video , 3D reconstruction , and 3D-aware generative modelling . Analogously, as this representation is differentiable, prior works apply coordinate-based MLPs to many inverse problems in computational photography and scientific computing .

2 Implicit 3D Scene Representation

Traditional 3D reconstruction methods utilizes discrete representations such as point cloud , meshes , multi-plane images , depth maps and voxel grids . Recently, INR has also prevailed among 3D scene representation tasks, which simply adopt an MLP that maps from any continuous input 3D coordinate to the geometry of the scene, including signed distance function (SDF) , 3D occupancy network , and so on. In addition to representing shape, INR has also been extended to encode object appearance. Among them, Neural Radiance Field (NeRF ) is one of the most effective coordinate-based neural representations for photo-realistic view synthesis that represents a scene as a field of particles. Draw inspiration from the preliminary success made by NeRF, a lot of following works further improve and extend it to wider application . Different from the grid-based approaches, training a stylized implicit representation can not access ground truth signals, which further amplifies the difficulty of optimizing the implicit neural representation.

3 Stylization

Traditionally, image stylization is formulated as a painterly rendering process through stroke prediction . The first neural style transfer method, proposed by Gatys et al. , builds an iterative framework to optimize the input image in order to minimize the content and style loss defined by a pre-trained VGG network. Due to the frustratingly large cost of training time, a number of follow-ups further explore how to design a feed-forward deep neural networks , which obtain real-time performance without sacrificing too much style information. Recently, several works extend it to video stylization and 3D environment .

Ha-NeRF is proposed for recovering a realistic NeRF at a different time of day from a group of tourism images, with a CNN to encode the appearance latent code. The most related NeRF-based stylization works: Style3D and StylizedNeRF still require a CNN-based hypernetwork or decoder to generate the stylized parameters for neural radiance field. In comparison, our proposed INS framework can work for more general implicit representations beyond neural radiance field, and can also be extended to encoding multiple styles. Experiments demonstrate that INS generate more faithful stylization on NeRF compared with Style3D .

Preliminary

This section introduces the relevant background on several implicit representations and volumetric radiance representations, including image fitting , neural radiance field and signed distance function .

Neural Radiance Field:

In contrast to point-wisely regression of implicit fields, NeRF proposes to reconstruct a radiance field by inversing a differentiable rendering equation from captured images. Specifically, NeRF learns an MLP f:(x,θ)↦(c,σ)f:(\boldsymbol{x},\boldsymbol{\theta})\mapsto(\boldsymbol{c},\sigma) with parameters Θ\boldsymbol{\Theta}, where x\boldsymbol{x} is the spatial coordinate in 3D space and θ\boldsymbol{\theta} represents the view directions ∈[−π,π]2\in[-\pi,\pi]^{2}. The output c∈3\boldsymbol{c}\in{}^{3} indicates the predicted color of the sampled point, σ∈+\sigma\in{}_{+} signifies its density value. The pixel color intensity can be obtained using volume rendering by ray tracing, integrating the predicted color and density along the ray. To render a pixel on the image plane, NeRF casts a ray r=(o,d,θ)\boldsymbol{r}=(\boldsymbol{o},\boldsymbol{d},\boldsymbol{\theta}) through the pixel and accumulate the color and density of KK point samples along the view direction in the 3D space. The pixel color intensity can be estimated:

where (ck,σk)=f(xk,θ)(\boldsymbol{c}_{k},\sigma_{k})=f(\boldsymbol{x}_{k},\boldsymbol{\theta}), xk=o+tkd\boldsymbol{x}_{k}=\boldsymbol{o}+t_{k}\boldsymbol{d}, tkt_{k} are the marching distance of sampled points, and Tk=exp⁡(−∑l=1k−1σlΔtl)T_{k}=\exp(-\sum_{l=1}^{k-1}\sigma_{l}\Delta t_{l}) is known as the transmittance to model occlusion. Δtk=tk+1−tk\Delta t_{k}=t_{k+1}-t_{k} indicates the distance of sampled point in 3D space. With this approximated rendering pipeline, the model weights are optimized by minimizing the L2L_{2} distance between rendered ray colors C(r)\boldsymbol{C}(\boldsymbol{r}) and captured pixel colors C^(r)\widehat{\boldsymbol{C}}(\boldsymbol{r}) as follows:

Implicit Surface Representation:

Signed Distance Function (SDF) f:3→ℜf:{}^{3}\rightarrow\real is an implicit representation of 3D geometries. SDF specifies each spatial point with the signed distance to the implicit iso-surface, where the sign indicates whether the point is inside or outside the object. Recent works of propose to employ MLPs to represent this continuous field via direct supervision using point clouds. To optimize a textured SDF from multi-view images like NeRF , Yariv et al. proposes a neural rendering pipeline, named IDR, which enables rendering images from an SDF. With this framework, one can indirectly supervise SDF using its multi-view projections. Suppose given a camera pose, we can cast rays r=(o,v)\boldsymbol{r}=(\boldsymbol{o},\boldsymbol{v}) through each pixel to trace an intersected point with the surface:

where t0t_{0}, v0\boldsymbol{v}_{0} and f0f_{0} are initial states when performing ray tracing (see ). After obtaining the intersection x^\boldsymbol{\hat{x}} of ray and surface, IDR also lets the SDF network ff output an appearance embedding γ^\boldsymbol{\hat{\gamma}}, and computes the normal n^=∇f(x^)\boldsymbol{\hat{n}}=\nabla f(\boldsymbol{\hat{x}}). Then the ray color can be rendered by another rendering MLP conditioned on both point coordinate x^\boldsymbol{\hat{x}} and normal n^\boldsymbol{\hat{n}}:

Similar to NeRF , ff and rr are simultaneously optimized by photometric loss between captured image pixels and rendered rays (see Equation 2).

Method

We next illustrate the main pipeline of Unified Implicit Neural Stylization (INS). INS consists of a Style Implicit Module (SIM) to transform the input style embedding into implicit style representations, a Content Implicit Module (CIM) to map the input coordinates into implicit scene representations, and an Amalgamate Module (AM) which amalgamates the two representation to predict RGB intensity. To preserve the geometry fidelity while generating the stylized texture of rendered views, a self-distilled geometry consistency regularization is applied upon the INS framework.

Generating the stylized images Y\boldsymbol{Y} can be formulated as an energy minimization problem . It consists of a content loss and a style loss, defined under a pre-trained VGG network . We build upon the prior work and thus propose our implicit stylization framework for SIREN, SDF and NeRF.

where C\boldsymbol{C} denotes the content ground truth image and Y\boldsymbol{Y} denotes synthesized output, Fi,jF_{i,j} denotes the feature map extracted from a VGG-16 model pre-trained on ImageNet, ii represents its ii-th max pooling, and jj represents its jj-th convolutional layer after ii-th max pooling layer. Ci,jC_{i,j}, Wi,jW_{i,j} and Hi,jH_{i,j} are the dimensions of the extracted feature maps. We adapt the content loss to the intermediate layer of INS pipeline to preserve the content of the predicted color image patch, we choose ii = 2, jj = 2 by default.

Style Representation

where J\mathcal{J} are the indices of selected feature maps. In practice, we choose J={(1,2),(2,2),(3,3),(4,3)}\mathcal{J}=\{(1,2),(2,2),(3,3),(4,3)\} in our experiments.

Conditional INS Stylization

Conditional encoding has been widely applied in convolutional networks . Similarly, we propose the conditional implicit representation by input with style conditioned embeddings using and extracting style-dependent features to render stylized color and density, which is shown in Figure 2. In training, we prepare nn style images with an nn dimensional one-hot style-condition vector. A mini-batch is constructed with the combinations of one content training patch and all candidate style images. The one-hot vector is fed into SIM to extract the ww-dimensional style features, they are then concatenated with the implicit representations output from CIM. The following layers of AM take the two features, aggregate them to render the pixel intensity and scene geometry along the rays. A pre-trained VGG is appended on the top of the INS pipeline to apply implicit style and content constraints during training. During the inference stage, we discard the VGG network, the INS framework becomes a pure MLPs-based network.

2 Geometry Consistency for Neural Radiance Field

Neural Radiance Field casts a number of rays (typically not adjacent) from camera origin, intersecting the pixel, into the volume and accumulating the color based on density along the ray.

While our model input with rays intersected with an image patch of size P∈K×K\mathcal{P}\in K\times K, predicting the stylized patch P′∈K×K\mathcal{P^{\prime}}\in K\times K with its texture closed to the given style images. 2D style transfer methods typically crop the patch larger than 256×256256\times 256. However, it is too expensive for the neural radiance field as it queries the MLPs more than 256×256×N256\times 256\times N times of the MLP for each step , where NN indicates sampled points number along each ray.

Similar to , we adopt a Sampling-Stride Ray Sampling strategy to enlarge the receptive field of the sampled patch to capture a more global context. The illustration of the ray sampling can be found in our supplementary materials, where a sampling stride larger than 1 result in a large receptive field while keeping computational cost fixed.

3 Optimization

Experiments

As one representative example, we apply INS on fitting an image via SIREN MLPs. We reuse the original SIREN framework as CIM and follow its training recipes to fit images of 512×\times512 pixels. Besides that, we also incorporate the SIM and AM on SIREN. A pre-trained VGG-16 network is appended on the output to provide style and content supervisions during training. As seen in Figure 4, the proposed framework successfully in representing the images with the given style statistics in an implicit way.

2 Novel View Synthesis with NeRF

We train our INS framework on NeRF-Synthetic dataset and Local Light Field Fusion(LLFF) dataset. NeRF-Synthetic consists of complex scenes with 360-degree views, where each scene has a central object with 100 inward-facing cameras distributed randomly on the upper hemisphere. Both rendered images and ground truth meshes are provided in NeRF-Synthetic dataset. LLFF dataset consists of forward-facing scenes, with fewer images. We implement INS on the same architecture and training strategy with the original NeRF . λ1\lambda_{1}, λ2\lambda_{2} and λ3\lambda_{3} are set as zero in the first 150k iterations and then set as 1e6, 1 and 1e8 in the following 50k iterations. The self-distilled density supervision depicted in Figure 3 is generated from the CIM with 150k iterations pre-training. Adam optimizer is adopted with learning rates of 0.0005. Hyper-parameters are carefully tuned via grid searches and the best configuration is applied to all experiments. All experiments are trained on one NVIDIA RTX A6000 GPU. We retrain Style3D in NeRF-Synthetic and LLFF datasets using their provided code and setting. We train all methods using the same number of style images for fair comparisons.

Results

In Figure 5, we can see INS generates faithful and view-consistent results for new viewpoints, with rich textures across scenes and styles. We further compare INS with three state-of-the art methods, including 3D neural stylization , image-based stylization methods . As is shown in Figure 6, We can see that stylizations from image-based methods produce noisy and view-inconsistent stylization as they transfer styles based on a single image. Style3D generates blur results as it still relies on convolution networks (a.k.a. hypernetwork) to generate the MLP weights for the subsequent volume rendering. Our proposed implicit neural stylization method is trained to preserve correct scene geometry as well as capture global context, generating better view-consistent stylizations.

3 Stylization on Signed Distance Function

DeepSDF only learns the 3D geometry from given inputs. Later work IDR extend it to reconstruct both 3D surface and appearance. We follow IDR to implement the implicit neural stylization framework. To encode style statistic onto IDR, we project the learned textured SDF into multi-view images and implement our style loss on the rendered results. In the experiments, we picked 2 scenes from the DTU dataset , where each scene consists of 50 to 100 images and object masks captured from different angles. Similar to NeRF, we pre-train the IDR model for chosen scenes by minimizing the loss between the ground truth image and the rendered result. Then both the SDF network and rendering network are jointly optimized the proposed framework from projected views. Note that due to IDR’s architectural design, we are no longer able to impose self-distilled geometry consistency loss. Instead, we employ content/style loss in the masked region, which is similar to . Besides, we observe that the SDF representation is more sensitive to parameter variations. To maintain intact geometries, we adjust the learning rate for the SDF network to 10−1110^{-11} times smaller than the rendering network.

Results

As are shown in Figure 8, the visualizations of two-view SDF representation demonstrate that both the learned appearance and geometry have deformed to fit the given style statistics.

4 Conditional Style Interpolation

Training with style-conditioned one-hot embedding, we can interpolate between style images to mix multiple styles with arbitrary weights. Specifically, we train INS on NeRF with two style images along with a two-dimensional one-hot vector as conditional code. After training, we mix the two style statistics by using a weighted two-dimensional vector. As is shown in Figure 9, the synthesized results can smoothly transfer from the first style to the second style when we linearly mix the two style embeddings at inference time.

5 Ablation Study

The geometry modification in the density branch enables a more flexible stylization, by stylizing shape tweaks on the object surface. As shown in Figure. 11, only updating the color branch easily collapses to the low-level color transformation instead of the painterly texture transformation, violating our original goal. This phenomenon is also addressed in the previous 3D mesh stylization , where they explicitly model the 3D shape by deforming the mesh vertexes.

Effect of Self-distilled Geometry Consistency

To evaluate the effectiveness of the proposed geometry consistency regularizer, we visualize the front and back viewpoints of the synthesized color images and depth maps. As is shown in Figure 10, the proposed self-distilled geometry consistency learns a good trade-off between stylization and clean geometry.

Should INS Learned with Larger Receptive Field?

To investigate the effect by using the Sampling Stride (SS) Ray Sampling strategy, we conduct a comparison of ray sampling with and without sampling stride for NeRF stylization. For a fair comparison, we set the ray number as 64×\times64 in both settings. INS with sampling stride covers the content resolution of (64×\timess) ×\times (64×\timess) where ss indicates sampling strides depicted in Section 4.2 and here we set ss=4. Figure 7 shows that INS with SS achieves significantly better visual results, as it results in a higher receptive field in perceiving content statistics.

Conclusions

In this work, we present a Unified Implicit Neural Stylization framework (INS) to stylize complex 2D/3D scenes using implicit function. We conduct a pilot study on different types of implicit representations, including 2D coordinate-based mapping function, Neural Radiance Field, and Signed Distance Function. Comprehensive experiments demonstrate that the proposed method yields photo-realistic images/videos with visually consistent stylized textures. One limitation of our work lies in the training efficiency issue, similar to most implicit representation, rendering a style scene requires several hours of training, precluding on-device training. Addressing this issue could become a future direction.

References