StyleMesh: Style Transfer for Indoor 3D Scene Reconstructions

Lukas Höllein, Justin Johnson, Matthias Nießner

Introduction

Creating 3D content from RGB-D scans is a popular topic in computer vision . We tackle a novel use case in this area: stylization of a reconstructed mesh with an explicit RGB texture. Neural Style Transfer (NST) shows great results for stylization of images or videos, but stylization of 3D content like meshes has been underexplored. We synthesize a texture for the mesh which is a combination of observed RGB colors and a painting’s artistic style. After stylization, one could explore the space in VR and see it painted in the style of Van Gogh.

Our use case is similar to prior texture mapping methods which construct a texture from a set of posed RGB images, but we produce a stylized texture rather than directly matching input images. This is difficult since style transfer losses are typically defined on 2D image features , so NST does not immediately generalize to 3D meshes. Recently, style transfer has been combined with novel view synthesis to stylize arbitrary scenes with a neural renderer from a sparse set of input images . These model-based methods require a forward pass during inference and cannot directly be applied to meshes. Kato et al. and Mordvintsev et al. use differentiable rendering to bridge the gap between image style transfer and texture mapping: backpropagating image losses to a texture representation enables consistent mesh stylization.

However, applying these methods to room-scale geometry is challenging as the resulting stylization patterns are noisy and can contain view-dependent stretch and size artifacts. For example, optimizing a surface from a small grazing angle creates patterns in the image plane for that pose. Viewing the same surface from an orthogonal angle then shows stretched-out patterns due to the perspective distortion. Similarly, seeing an object from close and far-away viewpoints mixes small and large patterns on the same surface. Perceiving the depth thus becomes harder, due to inconsistent stylization sizes. These issues arise because 2D style transfer losses do not incorporate 3D data like surface normals and depth. Instead, textures are separately stylized in each pose’s image plane.

To this end, we formulate an energy minimization problem over the texture that combines texture mapping with style transfer (similar to ) and minimizes style transfer losses for each pose in a 3D-aware manner that avoids view-dependent artifacts. First, we utilize depth to render image patches at increasingly larger screen-space resolutions. By splitting the style loss calculation over these patches, we create larger stylization patterns in the foreground than the background. As a result, patterns have the same size in world-space and are optimized in a view-independent way. Second, we use the angle between the surface normal and view direction to determine the degree of stylization for each pixel. By calculating Gram matrices from different style image resolutions (similar to ) areas seen from small grazing angles are stylized with coarse details, which are later refined if they are seen from better angles. Third, we avoid discretization artifacts by scaling gradients with per-pixel angle and depth weights during backpropagation.

Compared to state-of-the-art 3D style transfer methods, our experiments show an improvement in terms of 3D consistent stylization both qualitatively and quantitatively. Additionally, our explicit texture representation allows for direct usage with traditional rendering pipelines.

Style transfer for room-scale indoor scene meshes with a new texture optimization, which results in 3D consistent textures and mitigates view-dependent artifacts.

A depth-aware optimization at different screen-space resolutions, that creates equally-sized stylization patterns in the world-space of the mesh.

An angle-aware optimization at different stylization details, that creates unstretched stylization patterns in the world-space of the mesh.

Related Work

Our approach is a NST method operating on the texture parametrization of a mesh. It is related to recent work on style transfer for videos and 3D objects, as well as texture generation from RGB-D images.

Texture Mapping. Many methods texture a reconstructed mesh from multiple RGB images, i.e., they map a texture onto the geometry that combines the color information of all images . These methods must handle inaccuracies in pose, geometry, color and distortions to find the best texture for the scene. In contrast, we aim to create a texture that is also styled to a specific image and avoid view-dependent stylization artifacts by introducing depth- and angle-awareness into the optimization.

Image Style Transfer. NST, first introduced in Gatys et al. , can be optimization-based or model-based . It is inherently defined in the image domain by matching CNN features either globally or in a local, patch-based manner . Thus, it cannot directly utilize 3D data like depth or surface normals of a mesh. This can lead to view-dependent stylization artifacts when optimizing a texture through multiple poses. We induce 3D-awareness into optimization-based NST by splitting the loss calculation across different image segments.

Video Style Transfer. Video style transfer (VST) methods consistently stylize RGB video frames with a given style. These methods are optimization-based or model-based and employ temporal consistency or optical flow constraints. Other methods combine features in a temporally consistent way, without using optical flow or depth constraints directly . VST methods can be combined with texture mapping to achieve consistent stylization of indoor scenes. However, since VST optimizations are unaware of the underlying 3D structures, the resulting textures are often blurry or low-detail.

3D Style Transfer. Lifting style transfer into 3D has been explored for texturing individual objects or faces . However, they focus on isolated objects (not room-scale scenes) and do not utilize 3D data. In contrast, our method stylizes complete indoor scenes in a 3D-aware way. Another line of work applies exemplar-based NST to 3D models , guiding the stylization process explicitly from (hand-crafted) examples. In contrast, we follow original NST by stylizing 3D scene models from artistic paintings and camera images. Cao et al. stylize indoor scenes using a point cloud that cannot be directly used to texture a mesh. Other methods combine novel view synthesis and NST for consistent stylization from only a few input images . In contrast, we do not require a network during inference to produce stylization results; our results can be rendered by a standard graphics pipeline.

Method

Our goal is to stylize the mesh of an indoor scene: we want to create a texture that is a mixture of original RGB colors and a style image. To avoid view-dependent artifacts, we formulate a depth- and angle-aware optimization problem over all images. We require a set of NN images {Ik}k=1N\{I_{k}\}_{k=1}^{N} captured at different poses. We also need a mesh reconstruction of the scene for which we create a texture parametrization, i.e., we need a uvuv coordinate per vertex. For each pose, we sample the texture with the corresponding uvuv map at multiple resolutions, yielding a render pyramid. Depending on the depth of each pixel, we split the image into multiple render parts, each belonging to one pyramid resolution. Each part is used in content and style losses, where we only stylize pixels with fine details, that are seen from good angles. Finally, we smooth the per-pixel gradients before backpropagating to the texture. The complete method is visualized in Fig. 2.

We optimize a stylized RGB texture T∗\mathcal{T}^{*} from all RGB images {Ik}k=1N\{I_{k}\}_{k=1}^{N} and a separate style image IsI_{s}. Similar to , we formulate a minimization problem with content and style losses Lc,Ls\mathcal{L}_{c},\mathcal{L}_{s} and add a regularization term Lr\mathcal{L}_{r}:

where P^i\hat{P}_{i} is the render pyramid for the current pose, sampled from the texture with the corresponding uvuv maps and λc,λs,λr\lambda_{c},\lambda_{s},\lambda_{r} are loss weights. The sampling operation is identical to traditional graphics and differentiable, i.e., we bilinearly interpolate each pixel from four neighboring texels. Similar to Thies et al. , we define our texture using a Laplacian Pyramid to regularize the texels in each layer with Lr\mathcal{L}_{r}. This helps avoid magnification and minification artifacts and reduces visible noise in the texture. For each pose, we optimize the subset of observed texels. Thus, we require a pose set covering most of the scene to optimize the texture completely. In contrast, stylizing in texture space directly is problematic for a room-scale texture parametrization, which may contain many seams.

2 Depth Level Render Parts

Style transfer operates on the CNN features of an image . This leads to a limited sense of depth when optimizing over multiple poses. Stylization patterns can appear equally large in the foreground and background, e.g., when parts of a surface are seen far-away and close-up (see Fig. 3). Observing the same surface from multiple poses thus mixes small and large patterns next to each other. As a result, renderings using the optimized texture do not convincingly capture depth. Liu et al. make style transfer depth-aware in the image plane with a depth-loss network. In contrast, we incorporate depth-awareness by optimizing at multiple screen-space resolutions. Larger patterns appear in the foreground than the background of an image, ultimately leading to equally large style in world space.

We make use of the relation that area in screen-space is inversely proportional to depth, i.e., when depth increases by a factor of pp, a given projected area decreases by p2p^{2}. On the other hand, style transfer is agnostic to the image resolution that it is applied on, i.e., when resolution increases by p2p^{2}, stylization patterns appear proportionally smaller (because the receptive field becomes smaller relative to the resolution) . We combine both relations to optimize stylization patterns having the same size in world space: when depth increases by pp, we increase image resolution by p2p^{2}.

We apply the relation to divide the image into parts, sampled from the render pyramid at increasingly larger resolutions. The content and style losses are then calculated independently for each part. To discretize into parts, we define a minimum depth value θd\theta_{d}, making the relation absolute. We calculate the optimal image height per-pixel as

where dxyd_{xy} is the depth at pixel (x,y)(x,y) and θmin\theta_{min} is the minimum resolution. We express resolution RxyR_{xy} as height in pixels and scale the width accordingly. We then map RxyR_{xy} to the nearest neighbor in the render pyramid, yielding its index as depth level per-pixel. Finally, we apply a 3×33{\times}3 erosion kernel to smooth the depth level map over all pixels.

3 Angle Filter

Style transfer in screen-space can create stretched-out stylization patterns (see Fig. 3). Patterns might look circular from one view, but are stretched-out ellipses in world space (e.g., when optimizing from a small grazing angle). To prevent this, we combine coarse and fine style losses and optimize fine details only for areas seen from good angles. Similar to previous work , we utilize the fact that the receptive field of high-resolution images is still small . As a result, stylization patterns appear coarser and less detailed, when optimized from a larger style image. We find that coarse patterns are less prone to stretch artifacts.

For each pixel, we calculate its normal-to-view angle αxy=∡(n⃗xy,v⃗)\alpha_{xy}=\measuredangle(\vec{n}_{xy},\vec{v}) where n⃗xy\vec{n}_{xy} is the interpolated surface normal at pixel (x,y)(x,y) and v⃗\vec{v} is the viewing direction. Only the pixels where αxy≤θa\alpha_{xy}\leq\theta_{a} are used for the style loss with a low-resolution style image that produces fine stylizations. We always use all pixels and a high-resolution style image to optimize coarse stylization patterns. This creates a combination of coarse and fine patterns without stretch artifacts.

4 Multi-Resolution Part-based Losses

Multiple content and style losses combine depth levels (Sec. 3.2) and angle filtering (Sec. 3.3) to optimize the texture without view-dependent artifacts. We encode the render pyramid P^\hat{P} with a pretrained VGG network into the feature pyramid F^\hat{F}. Using the depth level map, we only keep corresponding features in each layer of F^\hat{F}. We compute a coarse Gram matrix G^c\hat{G}_{c} and an angle-filtered fine one G^f\hat{G}_{f} from the features in every layer. Similarly, GcG_{c} and GfG_{f} correspond to the high- and low-resolution style images. We define the style loss as

which sums over all depth levels independently (part-based) and combines coarse and fine stylization (multi-resolution). We calculate the normalized weighting factor w^l\hat{w}_{l} as

where vlv_{l} is the visible and tlt_{l} the total number of pixels in depth level ll. Similarly, the content loss is defined as

where FlF^{l} are the features of the content image II, split in a similar way. For brevity, we omit different VGG layers and image indices from the notation. As proposed in Gatys et al. , we use the layers relu_{1-5}_1 for the style loss and relu_4_2 for the content loss. We calculate the losses independently for every VGG layer and sum them accordingly.

5 Per-Pixel Gradient Scaling

Depth levels (Sec. 3.2) and angle filtering (Sec. 3.3) impose hard thresholds on the image of each pose. To avoid discretization artifacts at decision boundaries, we scale the per-pixel gradients before backpropagating them to the texture. First, we calculate a weighting factor wxya=cos(αxy)w_{xy}^{a}=cos(\alpha_{xy}) from the normal-to-view angle αxy\alpha_{xy}. This controls the influence of a pose on each pixel by preferring orthogonal over small grazing viewing angles. Scaling features similar to Gatys et al. instead results in oversaturation artifacts.

Second, we adapt the idea of trilinear Mipmap interpolation . Each pixel contributes to the render parts of its nearest two pyramid layers, resulting in two per-pixel gradients. We calculate the distance to the nearest layer as

where RxyR_{xy} is the optimal resolution for pixel (x,y)(x,y) and Lxy1L_{xy}^{1}, Lxy2L_{xy}^{2} are the resolutions of the nearest and second nearest pyramid layers. Finally, we linearly interpolate between the per-pixel gradients as

where L1\mathcal{L}_{1} is the loss term for the nearest pyramid layer of pixel (x,y)(x,y) and L2\mathcal{L}_{2} for the second nearest, respectively.

6 Data Preprocessing

We use the ScanNet and Matterport3D datasets, which provide RGB-D images and reconstructed meshes (we use per-region meshes for Matterport3D ). We use the RGB images for optimization, but filter them with a Laplacian kernel to remove blurry images. We reduce each mesh’s complexity by merging vertices until ≤500\leq 500K faces remain. Then, we generate a texture parametrization with Blender’s smart uvuv project with an angle limit of 70∘70^{\circ}. We precompute the uvuv maps for each estimated pose.

Results

Implementation Details. We optimize textures at a resolution of 4096×40964096{\times}4096 as a Laplacian Pyramid with 4 layers and regularization strength λr=5000\lambda_{r}{=}5000. We use λc=70\lambda_{c}{=}70 and λs=0.0001\lambda_{s}{=}0.0001 for content and style loss weights. We optimize for 7 epochs and repeat each frame 10 times. We set θmin=32\theta_{min}{=}32 and use θl=4\theta_{l}{=}4 render pyramid layers at heights of {256,432,608,784}\{256,432,608,784\} pixels. We set θa=30∘\theta_{a}{=}30^{\circ}, θd=0.25\theta_{d}{=}0.25 meters for ScanNet and θa=40∘\theta_{a}{=}40^{\circ}, θd=0.2\theta_{d}{=}0.2 meters for Matterport3D . We incrementally halve the original style image resolution until either the width or height reaches a size of 256 pixels. We use the resulting image for the stylization of fine details and a two steps larger image for coarse details. We use Adam with batch size 1 and initial learning rate 11 which decays multiplicatively by 0.10.1 every 3 epochs. We tried L-BFGS which gave similar results. After optimization, we export the Laplacian Pyramid to a single texture image and use a standard rasterizer with Mipmaps and shading for rendering.

Evaluation Metrics. We conduct a user study to show the advantages of depth- and angle-awareness (Fig. 10). Additionally, we quantify them by stylizing with a “circle” image (Fig. 8). We calculate the correlation between circle size and depth in screen-space (Corr. 2D) and world-space (Corr. 3D), as well as circle stretch as the ratio of horizontal and vertical radius in world-space (Tab. 2). For quantifying 3D consistency, we calculate the L1L_{1} distance between source and reprojected target frames (Tab. 1). Please refer to the supplemental material for more details about metrics.

Our method competes with 3D style transfer methods that stylize a scene through an explicit or implicit representation. Specifically, we compare our method with DIP of Mordvintsev et al. and NMR of Kato et al. : like us they also optimize a texture, but they do not utilize angle or depth data. Additionally, we compare with LSNV of Huang et al. , which uses a neural renderer to stylize point clouds. We show results on the Matterport3D dataset in Fig. 5 and on the ScanNet dataset in Fig. 6. A visualization of textured meshes is given in Fig. 4. Please see the supplemental material for more examples.

Our results show that we are able to stylize scenes without view-dependent size or stretch artifacts. In contrast to the other methods, our approach creates sharp and detailed effects for the complete scene. Optimizing the complete texture is especially difficult for DIP and NMR , which both contain noisy texels. LSNV stylizes complete images, but their results are less detailed. To quantitatively evaluate our method and the related approaches, we compute the mean L1L_{1} distance between source frame and a reprojected target frame. The results are listed in Tab. 1.

2 Ablation Studies

Qualitative Comparison. Our method uses per-pixel angle and depth as input to optimize the texture in a 3D-aware manner. This helps avoid view-dependent stretch and size artifacts being optimized into the texture from different poses. We compare only using angle input (no render pyramid) and not using angle/depth (only 2D texture optimization with Laplacian Pyramid representation). In Fig. 7 we can see that using angle makes it easier to distinguish between surfaces like the wall and sofa in row 2. Adding depth creates smaller and detailed patterns in the background (e.g., the strokes in the background of row 1). Please see the supplemental material for more examples.

We optimize all ablation modes such that stylization patterns are equally strong, i.e., style should be similar for a fair comparison. A too low degree of stylization would reduce view-dependent artifacts because original RGB colors get more dominant. Similarly, a too high degree discards content features too much, which increases artifacts.

Quantitative Comparison. We measure the effects of angle- and depth-awareness as follows. We stylize a scene with a “circle” image using only the style loss (see Fig. 8). We then detect ellipses in the resulting images and measure their horizontal and vertical axis lengths. Naturally, NST creates ellipses of different shapes, but their overall distribution reveals the degree of 3D awareness for the complete scene. Inverse correlation between per-pixel depth and ellipse size in screen-space (Corr. 2D) indicates that stylized features are smaller in the background. A weak correlation in world-space (Corr. 3D) indicates that absolute size is independent of the observed poses. Both metrics together classify the depth-awareness. View-dependent stretch is larger if ellipse’s horizontal and vertical axes are of different lengths. The stylization is angle-aware if the stretch is reduced. We do not measure coarse and fine stylization this way, because the “circle” image contains too few high-resolution features. Please see the supplemental material for more details about metric computation. As can be seen in Tab. 2, using angle and depth improves our method.

Depth Scaling. A key piece of our method is the render pyramid of different image resolutions. By tuning the value of θd\theta_{d}, we change the threshold of when to sample from the next higher resolution. This increases (higher θd\theta_{d}) or decreases (lower θd\theta_{d}) the absolute stylization size, while still retaining relative change in size (see Fig. 9). This allows to fine-tune the complete scene until a desired look is obtained.

User Study. We conduct a user study on the effectiveness of our proposed depth- and angle-awareness. Users compared our method against each baseline separately by preferring one of two images. They judged in which image stylization patterns (a) have less visible stretch and (b) are smaller in the background. In total, 20 users each answered 70 questions, comparing against NMR , DIP and ours without angle- and depth-awareness (Only 2D). As can be seen in Fig. 10, our method is preferred in both categories.

3 Comparison to Video Style Transfer

As an alternative way to optimizing a stylized texture, one could combine video style transfer (VST) methods and RGB texture mapping to produce a stylized scene in two steps (see Fig. 11). We can obtain an RGB texture from all images of the scene and render arbitrary trajectories, that we stylize with a VST method (Tex→\rightarrowVST). However, we never obtain a stylized texture this way and thus need the VST method during inference for each novel pose. Stylization details are also much lower, due to missing details in the RGB texture and reconstructed geometry. By optimizing directly from camera images, we obtain sharper details.

Alternatively, we can stylize a trajectory of camera images with a VST method and optimize an RGB texture from these images (VST→\rightarrowTex). However, we might only have access to a sparse set of images in some scenarios. Due to inconsistencies between stylized frames (e.g., caused by illumination changes), the optimized texture is blurrier, as well. Our method is 3D-consistent by combining stylization and texture optimization over all available images directly.

4 Runtime Comparison

We propose an optimization-based NST method, that converges in roughly 3 hours on a single RTX 3090 GPU. After optimization, we can use the texture in traditional graphics pipelines and achieve real-time rendering, similar to . In contrast, model-based NST might take days to train and needs a forward-pass at inference. However, these methods can generalize across scenes, whereas we need to optimize a separate texture per-scene.

5 Limitations

By design, our method is a per-scene/per-style NST algorithm, i.e., we optimize each explicit texture image separately. Recent work in implicit texture representations could enable training generative models for our task. We do not disentangle lighting and albedo, i.e., view-dependent effects in camera images can be visible in the stylized texture. One could leverage neural rendering techniques to train a relightable stylization model . Incomplete mesh reconstructions lead to holes in rendered poses, which can be reduced by employing mesh completion techniques first . Similarly, an insufficient number of poses may lead to unobserved surfaces during optimization, i.e., we do not hallucinate texture. Inpainting techniques could be utilized to complete those texels.

Conclusion

We have shown a method to stylize the mesh of room-scale indoor scene reconstructions. We lift style transfer to the 3D domain by optimizing a texture only through 2D images. Our method makes use of depth and surface normals of the mesh to achieve uniform world space stylization without view-dependent artifacts. For that, we split the loss calculation into image parts and stylize coarse and fine details separately. The explicit texture representation allows for real-time rendering of the scene after optimization.

Acknowledgements

This project is funded by a TUM-IAS Rudolf Mößbauer Fellowship, the ERC Starting Grant Scan2CAD (804724), and the German Research Foundation (DFG) Grant Making Machine Learning on Static and Dynamic 3D Data Practical. We also thank Angela Dai for the video voice-over.

References

Appendix A Supplemental

We calculate reprojection error as the L1L_{1} distance between source frame and reprojected target frame. This allows to quantify the 3D consistency. While a texture representation is 3D consistent by design, unoptimized texels might still create visible noise artifacts when rendering a trajectory. We calculate reprojection error on the ScanNet dataset by using their captured camera trajectories. For each source frame, we select a target frame that is (a) two frames after the source frame (short-range consistency) or (b) 20 frames after the source frame (long-range consistency). Using the estimated poses and camera intrinsics, we warp the pixels of the target frame to the source view. We calculate L1L_{1} distance in the normalized image range $$ between reprojected target frame and source frame for all pixels that are visible in both views. Please see Fig. 12 for a visualization of the procedure.

Because the estimated poses are not perfectly accurate, the reprojection error may never sink below a certain threshold that captures this inaccuracy. Still, it allows to quantify the 3D consistency by measuring the additional inconsistencies caused by unoptimized textures, i.e., the error is higher for unoptimized textures that are less consistent.

A.2 User Study Setup

We conduct a user study on the effectiveness of our proposed depth- and angle-awareness. Users compared our method against each baseline separately by preferring one of two images. They judged in which image stylization patterns (a) have less visible stretch and (b) are smaller in the background. In total, 20 users each answered 70 questions, comparing against NMR , DIP and ours without angle- and depth-awareness (Only 2D). We show two sample questions, one for each type, in Fig. 13. As can be seen, users have the possibility to decide for one of two images or to answer that none of the two is better/worse. The order of questions and of the “A” and “B” images is random and different for each user. Users have the possibility to zoom-in on the images for better judgement. Additionally, we add rectangles on image regions that might be especially interesting for evaluation of the questions. Note that users still had to consider the whole image in their answer; the rectangles merely act as additional input.

A.3 Variation of Depth Levels

The number of depth levels θl\theta_{l} controls the depth variation that can be achieved within rendered poses of a scene. Setting θl=4\theta_{l}{=}4 is sufficient for our datasets, as larger scene extent is rarely captured by many pixels. We could precompute uv maps at larger resolutions to enable depth scaling at even larger depth values. Small scenes may not require the last layers, in which case they are simply not utilized during optimization. Adding more layers in-between maps less pixels to one layer, which can yield insufficient Gram matrices and is computationally more expensive. Decreasing θl\theta_{l} reduces size variation at different depths (see Fig. 14).

A.4 Circle-Stretch and -Size Metric

We describe in more detail the metrics and principles used in the main paper to quantify the effects of our depth and angle awareness. In order to measure the effects, we stylize a scene with a hand-crafted “circle” image (see main paper) and only use the (multi-resolution, part-based) style loss. After optimization, the red circles are stylized all over the scene and are well-suited to describe the two drawbacks of missing angle- and depth-awareness. For example, circles become ellipsoidal if a small grazing angle is used for stylization and circles change their radius inconsistently without depth awareness. We can now measure the degree we alleviate these issues by measuring the size and stretch of the circles/ellipses. Naturally, NST creates ellipses of different shapes, but their overall distribution reveals the degree of 3D awareness for the complete scene.

First, we automatically segment red ellipses from each image of the stylized scene (see Fig. 15). We first apply an HSV filter and only keep pixels in the ranges 0.6≤S,V≤1.00.6\leq S,V\leq 1.0, 0.0≤H≤0.080.0\leq H\leq 0.08 and 0.88≤H≤1.00.88\leq H\leq 1.0. Then we turn the filtered image into a binary mask by thresholding colors above 0.150.15 intensity and denoise it with OpenCV’s “fastNLMeansDenoising” function . Afterwards, we use OpenCV’s contour detection to get an edge map. We filter out all contours with maxd>2max_{d}>2, where maxdmax_{d} is the maximum deviation from a convex hull, as measured by “convexityDefects” . We now fit ellipses to the remaining contours with “fitEllipse”. We extract the pixel-radius as

for every fitted ellipse and calculate its pixel-stretch as

where hph_{p} is the pixel-length of the horizontal ellipse radius and vpv_{p} the vertical, respectively. We remove the remaining wrongfully detected ellipses with rp<10r_{p}<10, rp>1000r_{p}>1000 and sp>10s_{p}>10 to get a result like in Fig. 15. We use these ellipse characteristics to calculate metrics for depth- and angle-awareness.

A.4.2 Calculation of Depth Metrics

We calculate the correlation between per-pixel depth dxyd_{xy} and ellipse radius rpr_{p} to quantify the effect of our depth-awareness in the 2D image plane (Corr. 2D). For each detected ellipse we use the depth value of the pixel corresponding to the ellipse center. A high negative correlation (e.g., −0.5-0.5) signals, that ellipse size decreases with increasing depth, whereas a low correlation (e.g., −0.05-0.05) signals, that ellipse size is independent of changes in depth. A method that is able to stylize a scene depth-aware would create ellipses with smaller size in the background and thus have a high negative correlation in the 2D image plane.

To quantify the correlation in 3D, we backproject hph_{p} and vpv_{p} to world-space using the estimated pose and camera intrinsics and calculate the world-space radius as

where hwh_{w} and vwv_{w} are the backprojected axis lengths. We then calculate the correlation between rwr_{w} and the per-pixel depth dxyd_{xy} (Corr. 3D). A high negative correlation (e.g., −0.5-0.5) signals, that ellipse size in world-space still decreases with increasing depth, whereas a low correlation (e.g., −0.05-0.05) signals, that ellipse size in world-space is independent of changes in depth. A method that is able to stylize a scene depth-aware would create ellipses with uniformly distributed size in world-space (because the ellipse size should only change when rendering a scene from different poses, due to perspective projection).

Note that the stylized ellipses naturally vary in their sizes (e.g., ellipses can be smaller and larger independent of depth). Therefore, the correlations will be precise up to a certain threshold. However, the distribution of all segmented ellipses across the whole scene still allows to quantify the depth-awareness.

A.4.3 Calculation of Angle Metric

We backproject the pixel-stretch sps_{p} back to world-space as

. Then we calculate the arithmetic mean over all sws_{w} values for all detected ellipses. A higher mean value means that overall we have more stretch, whereas a lower value signals a more uniform stylization result. A method that is able to stylize a scene angle-aware would create ellipses with small stretch.

A.5 Additional Qualitative Results

We show additional qualitative results for our method.

Additional comparisons on the ScanNet dataset can be found in Fig. 16 and Fig. 17.

Additional comparisons for our ablation study can be found in Fig. 18.

A.6 Style Image Assets

Throughout the main paper and the supplemental material, we use style images created by artists. In Fig. 19 we list all images and give credit to their respective creators.