VoxGRAF: Fast 3D-Aware Image Synthesis with Sparse Voxel Grids

Katja Schwarz, Axel Sauer, Michael Niemeyer, Yiyi Liao, Andreas Geiger

Introduction

Generating photorealistic renderings of scenes at high resolution is a long-standing goal in computer vision and graphics. The primary paradigm is to carefully design 3D models, which are then rendered using realistic camera and illumination models. In recent years, the computer vision community has made significant headway towards reducing these design efforts by approaching content generation from a data-centric perspective. Generative Adversarial Networks (GANs) have emerged as a powerful class of generative models for photorealistic high-resolution image synthesis . One benefit of these 2D models is that they can be trained with large collections of images which are readily available. However, scaling GANs to 3D is non-trivial because 3D supervision is difficult to obtain. Recently, 3D-aware GANs have emerged to address the gap between handcrafted 3D models and image synthesis with 2D GANs which lack 3D constraints . 3D-aware GANs combine 3D generators, differentiable rendering and adversarial training to synthesize novel images with explicit control over the camera pose and, potentially, other scene properties like object shape and appearance.

Early 3D-aware GANs explored voxel-based 3D representations . To compensate for the cubic memory growth of voxel grids, HoloGAN generates features on a small 3D grid and uses a neural network to map 3D features to a 2D image. While such a neural renderer allows scaling to higher image resolutions, it may also entangle viewpoint and generated content . Consequently, early voxel-based approaches were either limited in image resolution or lacking 3D consistency. About the same time, Neural Radiance Fields (NeRF) emerged in the context of view synthesis as a powerful alternative 3D representation. In their seminal work, Mildenhall et al. represent a scene as a function of color and density, parameterized by a coordinate-based MLP. The predicted color and density values are then projected to an image with differentiable volume rendering. GRAF adapts NeRF’s coordinate-based representation to Generative Radiance Fields (GRAF) and proposes a 3D-aware GAN using a coordinate-based MLP and volume rendering. This propelled 3D-aware image synthesis to higher image resolutions while better preserving 3D consistency due to the physically-based and parameter-free rendering. These benefits led to the establishment of coordinate-based MLPs as new de facto standard for 3D-aware image synthesis .

While recent 3D-aware GANs have started to attain image fidelity and resolution similar to 2D GANs, training and inference is computationally expensive as the MLP must be queried at multiple points along each ray for volume rendering. However, querying 3D space densely is prohibitively costly. For example, rendering an image at resolution 2562256^{2} using 4848 sample points along each ray requires to query the neural network 2562⋅48≈3M256^{2}\cdot 48\approx 3M times. As a result, most recent works combine neural and volume rendering to ease computational cost at high resolutions . Consequently, viewpoint and 3D content are often entangled, such that changing the camera pose might result in unwanted changes of the geometry or the appearance of the 3D scene. Further, for many downstream applications, e.g. integrating assets into physics engines, it is desirable to generate 3D content at high resolution directly. These shortcomings identify the need for a 3D representation that can be efficiently rendered at high resolution with a model that is 3D-consistent by design.

Recently, there has been significant progress towards accelerated training of novel view synthesis models for single scenes by removing the MLP from the representation. In particular, DVGO and Plenoxels directly optimize a sparse voxel grid for novel-view synthesis, demonstrating that the visual fidelity attained by NeRF is not primarily attributed to its MLP-based representation but rather to volumetric rendering and gradient-based optimization. In addition to impressive image quality, obtain large rendering speedups due to their fast density and color queries. Taking inspiration from these works, we revisit voxel-based representations for 3D-aware GANs in the context of volumetric rendering. To circumvent the cubic memory growth that limits early voxel-based approaches , we explore sparse voxel grids for the generative settings, see Fig. 1. We observe that sparsity is key to enable scaling the 3D representation to higher resolution and to combine it with volume rendering. Specifically, we propose a 3D-aware GAN with a sparse voxel grid generator at its core. As a result, our approach inherits fast rendering and trilinear interpolation while being 3D-consistent by design, separating it from other recent 3D-aware GANs which require a forward pass for every point along each camera ray of every view. Another difference to existing 3D-aware GANs is that sparsity is a built-in feature of our representation, mitigating the need for exploiting sophisticated strategies to sample points along camera rays as in . Our final model achieves image fidelity similar to recent 3D-aware GANs leveraging neural rendering while generating high-resolution geometry and improving 3D consistency. During inference, our model only requires a single forward pass which takes up most of the inference time. Once the 3D scene is generated, images can be rendered within milliseconds while existing approaches require another forward pass which is two orders of magnitude slower. We refer to our model as VoxGRAF.

Related Work

2D GANs. Rapid progress on Generative Adversarial Networks now enables photorealistic synthesis up to megapixel resolution . While the disentangled style-space of StyleGANs allows for control over the viewpoint of the generated images to some extent , gaining precise 3D-consistent control is still non-trivial due to its lack of physical interpretation and operation in 2D. In contrast, in this work we aim for explicit control over the camera pose by incorporating a 3D representation into the generator.

3D-Aware GANs. The first 3D-aware GANs, i.e. GANs that incorporate a 3D representation into the generator model, were voxel-based approaches. Dense grid-based approaches are limited to a lower grid resolution due to their cubic memory growth. Other works combine lower-resolution grids with neural rendering which scale to higher resolutions, but the generated images lack 3D consistency . Recently, GRAF and π\pi-GAN combine volume rendering and coordinate-based representations allowing to scale 3D-aware GANs with physically inspired rendering to high resolutions. However, dense ray marching remains computationally expensive and limits image fidelity. GIRAFFE therefore proposes a hybrid rendering approach. They render a low-resolution feature map with ray marching and use a neural renderer to decode it into a high-resolution image. Due to efficiency, this approach has been widely adopted in subsequent works . While some approaches try to counteract introduced inconsistencies, e.g. via dual discrimination or a reconstruction loss , we instead propose a model that is 3D-consistent by design and can generate the 3D object at high-resolution.

As an alternative to hybrid rendering, GOF , ShadeGAN and GRAM focus on reducing the number of query points for volume rendering. While aforementioned methods aim for reducing the number of sample points as querying a large MLP is computationally expensive, we instead use a sparse voxel grid as 3D representation. This allows us to speed up rendering without reducing the sample size as feature querying via trilinear interpolation is fast and can be efficiently implemented via custom CUDA kernels.

Sparse 3D Representations. As NeRF requires an optimization time in the order of multiple days per scene, a series of follow-up works propose techniques to speed up this process. The recent works Plenoxels and DVGO demonstrate that sparse voxel grid representations can achieve even faster convergence and higher rendering speed. In addition to efficient rendering, sparse voxel grids enable fast trilinear interpolation when queried beyond their grid resolution. Building on this representation, our approach inherits these benefits. We remark that very recently Instant-NGP achieves even faster rendering by combining small MLPs with a multi-resolution hash table. Exploring this representation might be an interesting avenue for future extensions of our work. Note that all of these works focus on novel-view-synthesis for single scenes and require multi-view image supervision. Instead, we propose a generative model that trains with raw image collections and that can generate multiple novel instances at inference.

Method

We first provide the necessary background by summarizing the currently dominating paradigm for designing 3D-aware GANs which combines an MLP scene representation and volume rendering as introduced in GRAF . Next, we introduce our sparse voxel-based scene representation which boosts rendering speed while retaining 3D-consistency by design. We refer to our model as VoxGRAF.

In this paper, we challenge this paradigm and investigate a sparse 3D CNN instead of a coordinate-based MLP as generator, as described in the next section.

The radiance field is rendered by approximating the intractable volumetric projection integral via numerical integration. First, the generator is queried at NN sampling points along each camera ray rr yielding colors and densities {(cri,σri)}i=1N\{(\mathbf{c}_{r}^{i},\sigma_{r}^{i})\}_{i=1}^{N}. For each camera ray rr, these points are projected to an RGB color value cr\mathbf{c}_{r} and optionally an alpha mask ar\mathbf{a}_{r} using alpha composition

where TriT_{r}^{i} and αri\alpha_{r}^{i} denote the transmittance and alpha value of sample point ii along ray rr and δri=∥xri+1−xri∥2\delta_{r}^{i}=\left\|\mathbf{x}_{r}^{i+1}-\mathbf{x}_{r}^{i}\right\|_{2} is the distance between neighboring sample points. As volume rendering has proven a powerful tool for high-fidelity reconstruction, VoxGRAF retains this rendering mechanism but reduces its computational cost by leveraging sparse scene representations.

2 VoxGRAF: Generating Radiance Fields on Sparse Voxel Grids

Our goal is to design a 3D-aware GAN based on a sparse scene representation that allows for efficient rendering. Fig. 2 shows an overview over our approach. In contrast to recent works , we do not use a coordinate-based MLP to parameterize the radiance field. Instead, inspired by recent work on novel view synthesis , we generate values on a sparse voxel grid using a 3D convolutional neural network. Therefore, our generator requires only a single forward pass to generate a 3D scene. To disentangle 3D content from the background we combine a 3D foreground generator Gθffg\mathbf{G}_{\theta_{f}}^{fg} with a 2D background generator Gθbbg\mathbf{G}_{\theta_{b}}^{bg}. Gθffg\mathbf{G}_{\theta_{f}}^{fg} takes a camera matrix K\mathbf{K}, camera pose ξ\mathbf{\xi} and a latent code z\mathbf{z} as input and predicts colors c\mathbf{c} and density values σ\mathbf{\sigma} on a sparse voxel grid. For unstructured images, camera matrix K\mathbf{K} and pose ξ\mathbf{\xi} can be determined e.g. with an off-the-shelf pose detector as done in . Volume rendering yields a foreground image and an alpha mask. The background generator maps the latent code z\mathbf{z} to a background image which is then combined with the foreground image using alpha composition. We train our model in an adversarial setting using a 2D discriminator on the full image.

where θf\theta_{f} and RGR_{G} denote the learnable parameters and the resolution of the voxel grid, respectively. Note that in contrast to most existing coordinate-based 3D-aware GANs, see Eq. (1), we do not condition the 3D generator on the view direction (per-ray). Instead, we follow and condition it on the pose ξ\boldsymbol{\xi} (per-image) to model directional dependencies. For rendering, we compute (cri,σri)(\mathbf{c}_{r}^{i},\sigma_{r}^{i}) via trilinear interpolation of densities and colors stored at the nearest eight vertices . We use the same rendering formulation outlined in 3.1 and leverage custom CUDA kernelsWe build on the kernels from https://github.com/sxyu/svox2.git for efficiency.

To generate voxel grids instead of images, we replace all 2D operations of the StyleGAN2 generator with their 3D equivalent, e.g., 3D modulated convolutions and 3D upsampling. At resolutions beyond 32332^{3}, we investigate sparse convolutions instead of dense ones as they can be more computationally efficient. We compare the computational efficiency of sparse convolutionsWe use the Minkowski Engine library https://github.com/NVIDIA/MinkowskiEngine.git to dense convolutions for which we zero out the values of pruned voxels. For our architecture, using sparse convolutions reduces the memory consumption but due to the computational overhead for managing coordinates increases the runtime, see Table 1. To sparsify the representation, we combine progressive growing with pruning as illustrated in Fig. 3(a). Specifically, we start with training a dense model at resolution 32332^{3}. After sufficient training, we add the next layer of convolutions and prune its inputs based on the rendered view. Intuitively, the next layer should only operate on voxels visible in the rendered view to yield a sparse representation. Consequently, we prune voxels that are either occluded or have a low density. Following Eq. (2), the rendered alpha value is ar=∑i=1N Tri αri\mathbf{a}_{r}=\sum_{i=1}^{N}\,T_{r}^{i}\,\alpha_{r}^{i}. Accordingly, occluded voxels (low transmittance TiT_{i}) and empty voxels (low density σi\sigma_{i}) do not contribute to the final image. The pruning operator ρ\rho discards all voxels along a camera ray rr for which transmittance TiT^{i} or density σi\sigma_{i} is smaller than threshold τT\tau_{T} or τσ\tau_{\sigma}, respectively. Let Vr\mathcal{V}_{r} denote the set of voxels intersecting with ray r\mathbf{r}, then

where Vrp\mathcal{V}_{r}^{p} is the set of the retained voxels. After pruning, we upsample the kept voxels via sparse transposed convolutions. After the newly-added layer is sufficiently trained, we repeat this process for the next added stage. Fig. 3(b) shows images and voxel grids at different resolutions. We implement this on-the-fly pruning operation using efficient custom CUDA kernels.

In agreement with , we observe that it is crucial to account for view dependent effects or pose-correlated attributes in the training data, like eyes always looking into the camera. We take two measures to increase the flexibility of our model: Following , Gθffg\mathbf{G}_{\theta_{f}}^{fg} is conditioned on the pose ξ\xi which corresponds to the rendering pose ξ′\xi^{\prime} in 50%50\% of the cases and is randomly chosen otherwise. Formally, ξ∼pξ∣p(ξ=ξ′)=0.5\xi\sim p_{\xi}|_{p(\xi=\xi^{\prime})=0.5}. At inference, the pose-conditioning is fixed to retain 3D consistency. However, we find that pose conditioning alone is not always sufficient to account for strong correlations in the data. Depending on the dataset, we optionally refine the rendered image with a shallow 2D CNN with 22 hidden layers of dimension 1616 and kernel size 33. While this refinement is powerful enough to model dataset biases, it is considerably less flexible than the neural rendering used in : Our 2D CNN operates on the rendered image instead of rendered features at smaller resolution. Operating on 3 channels at full resolution allows to keep its capacity at a minimum and, due to not using any upsampling operation, results in a local receptive field.

Background Generator. We consider datasets with a single object per image. Modeling only the object in 3D saves computation and is advantageous for potential downstream tasks, e.g., integrating generated assets into new environments Hence, we model the object in 3D and generate the background of the image with a 2D GAN. Specifically, we use the StyleGAN2 generator with reduced channel size as modeling the background requires less capacity than generating the full image:

We choose the same latent code z\mathbf{z} for foreground and background to allow for modeling correlations like lighting between the two. Note that, unlike the foreground generator, the background generator is not conditioned on the camera pose. As the background remains fixed when changing the camera pose, the generator is encouraged to model pose-dependent content with the foreground generator leading to disentanglement (see supp. mat. for qualitative disentanglement results).

The final image is obtained using alpha composition

Regularization. For fast rendering, it is crucial that most generated voxels are either fully opaque or empty such that early stopping and empty space skipping are effective. By regularizing the variance of the expected depth z^r\hat{z}_{r} along each ray rr, the foreground generator is encouraged to generate a single, sharp surface:

where τ\tau is a hyperparameter that defines the thickness of the surface. We find that thresholding the loss is important to avoid an empty foreground. In addition, we find that adding the grid TV regularization LTV\mathcal{L}_{TV} from and fore- and background coverage regularization Lcvgfg\mathcal{L}_{cvg}^{fg} and Lcvgbg\mathcal{L}_{cvg}^{bg} from further stabilizes training (see sup. mat. for details). The full regularization term of our generator is

Discriminator. We use the StyleGAN2 discriminator and condition it on the camera pose as proposed in . Similarly, we find that conditioning guides the generator to learn correct 3D priors and a canonical representation. Since rendering our sparse representation is fast, our discriminator is able to operate on the full image and does not need to consider image patches as done in GRAF .

3 Training

Given images I\mathbf{I} from the data distribution pDp_{\mathcal{D}} with known camera extrinsics ξI\boldsymbol{\xi}_{\mathbf{I}} and intrinsics KI\mathbf{K}_{\mathbf{I}} and latent codes z∈N(0,1)\mathbf{z}\in\mathcal{N}(\mathbf{0},\mathbf{1}), we train our model using a GAN objective with R1-regularization

where f(t)=−log⁡(1+exp⁡(−t))f(t)=-\log(1+\exp(-t)) and λ\lambda controls the strength of the R1-regularizer. GθG_{\theta} and DϕD_{\phi} are trained with alternating gradient descent combining the GAN objective with the regularization terms:

In practice, we optimize the generator with a non-saturating variant of Eq. (12) . We train our approach with Adam using a batch size of 64 at grid resolution RG=32,64R_{G}=32,64 and 32 at RG=128R_{G}=128. We use a learning rate of 0.00250.0025 for the generator and 0.0020.002 for the discriminator. For faster training, we first grow RIR_{I} from 3232 to 128128 while keeping RGR_{G} at 3232. Then, we alternately increase the grid resolution and the image resolution until the dataset resolution and RG=128R_{G}=128 are reached. For synthetic datasets, i.e. Carla , we do not add any refinement layers. Depending on the dataset, we train our models for 3 to 7 days on 8 Tesla V100 GPUs. Details on the network architectures can be found in the supplemental material.

Results

Datasets. We validate our approach on standard benchmark datasets for 3D-aware image synthesis. The synthetic Carla dataset contains 1010k images and camera poses of 1818 car models with randomly sampled colors. FFHQ comprises 7070k aligned face images. AFHQv2 Cats consists of 48344834 cat faces. Following , we estimate camera poses for both datasets with off-the-shelf pose estimators and augment all datasets with horizontal flips. Due to the limited number of images in AFHQv2 and Carla, we use adaptive discriminator augmentation for these datasets.

Evaluation Metrics. We measure image fidelity by calculating the Fréchet Inception Distance (FID) between 2020k generated images and the full dataset. For all runtime comparisons, we report times on a single Tesla V100 GPU with a batch size of 1.

We investigate the sparsity of the generated voxel grids and validate the importance of the depth variance loss, see Eq. (8). Sparsity is evaluated by fusing the pruned voxel grids from 16 equally spaced camera views and reporting the number of empty voxels divided by the total number of voxels RG3R_{G}^{3}. Table 1 shows results averaged over 100100 instances. The depth variance loss increases sparsity from 74%74\% to 95%95\%, lowering the memory consumption. With the sparser representation, the times needed for scene generation tGfg+bgt_{G^{fg+bg}} and rendering tπt_{\pi} are reduced significantly. We further compare an implementation with dense instead of sparse convolutions where we zero out the values of pruned voxels. While this increases memory consumption, it reduces the runtime due to the computational overhead of managing coordinates for sparse convolutions. We prioritize faster training over memory and hence train our models with dense convolutions where we set values of pruned voxels to zero.

2 Baseline Comparison

Baselines. We group the baselines into two categories: (i) methods that render low-resolution features and use 2D-upsampling, i.e. a neural renderer, to obtain the final image, e.g. StyleNeRF , and (ii) methods that render the 3D representation directly at the final image resolution, e.g. π\pi-GAN . VoxGRAF falls into the second group. We use the official code release for all methods. Further, we reference numbers reported by concurrent works GRAM and EG3D .

Fig. 4 shows samples from multiple views for StyleNeRF, GRAM and VoxGRAF on FFHQ at resolution 2562256^{2}. StyleNeRF can generate additional faces or strands of hair across views, as indicated by the red frames in Fig. 4. These inconsistencies under viewpoint changes are introduced by StyleNeRF’s powerful neural renderer. GRAM achieves more consistent results, but its plane representation creates stripe artifacts for large viewpoint ranges. Due to the shallow 2D CNN, VoxGRAF can model the dataset bias of eyes looking into the camera but otherwise achieves high 3D-consistency even under large viewing angles. Additional samples of our method and the corresponding sparse voxel grids for all datasets are provided in Fig. 1 and Fig. 5.

Table 3 reports FID on all datasets. As expected, methods that use neural rendering with upsampling, i.e., StyleNeRF and EG3D, perform best in terms of image fidelity. This is expected as a neural renderer can add flexibility. But, as shown in Fig. 4, for StyleNeRF it reduces 3D-consistency. Among methods without a neural renderer, VoxGRAF significantly improves over π\pi-GAN and GOF and surpasses the current state-of-the-art approach GRAM.

Runtime Comparison. Lastly, we compare the rendering times at inference for methods with and methods without a neural renderer. In contrast to all baselines, VoxGRAF requires only a single forward pass to generate the scene, which can then be rendered from different viewpoints efficiently. Note that at inference, voxels are pruned solely based on their density to amortize the rendering costs per scene. We find that this does not visibly affect the rendered images. We report the times for generating the scene and rendering one view separately in Table 3. One of the earliest works, GIRAFFE, is the fastest among all approaches as it renders a low-resolution feature volume and uses comparably small neural networks. StyleNeRF significantly increases the neural renderer’s size to improve image fidelity, which comes at the cost of speed compared to GIRAFFE. Yet, StyleNeRF is the fastest approach for generating a single image among the best-performing methods. However, a potential application of 3D-aware GANs is generating novel views of a single instance in real-time. In this setting, at resolution 2562256^{2}, VoxGRAF generates novel views at 167 FPS, whereas StyleNeRF runs at 20 FPS as its rendering costs are not amortized per scene.

Limitations and Discussion

In this work, we investigate sparse voxel grids as representation for 3D-aware image synthesis. We find that the key to generating sparse voxel grids is to combine progressive growing, pruning, and regularization to encourage a sharp surface that can be rendered efficiently. Our approach outperforms all methods that do not employ a neural renderer. Instead of discarding neural rendering entirely, we find it advantageous to utilize a shallow CNN for refinement. This CNN can model dataset bias but is significantly weaker than standard neural rendering approaches that upsample low resolution feature maps. Our approach can reduce the gap to models that build heavily on neural rendering, yet a trade-off between 3D-consistency and image fidelity remains. Whether a certain amount of neural rendering is inherently needed to reach best performance is an important direction for future research. Lastly, the speed of our method depends on the sparsity of the modeled scene. Therefore, rendering times will likely increase on more complex datasets than those commonly used in literature.

Acknowledgments and Disclosure of Funding

We acknowledge the financial support by the BMWi in the project KI Delta Learning (project number 19A19013O), the support from the BMBF through the Tuebingen AI Center (FKZ:01IS18039A), and the support of the DFG under Germany’s Excellence Strategy (EXC number 2064/1 - Project number 390727645). Andreas Geiger and Michael Niemeyer were supported by the ERC Starting Grant LEGO-3D (850533). We thank the International Max Planck Research School for Intelligent Systems (IMPRS-IS) for supporting Katja Schwarz and Michael Niemeyer. This work was supported by an NVIDIA research gift. We thank Christian Reiser for the helpful discussions and suggestions. Lastly, we would like to thank Nicolas Guenther for his general support.

References

Checklist

Do the main claims made in the abstract and introduction accurately reflect the paper’s contributions and scope? [Yes]

Did you describe the limitations of your work? [Yes]

Did you discuss any potential negative societal impacts of your work? [Yes] See supplemental material.

Have you read the ethics review guidelines and ensured that your paper conforms to them? [Yes]

If you are including theoretical results…

Did you state the full set of assumptions of all theoretical results? [N/A]

Did you include complete proofs of all theoretical results? [N/A]

Did you include the code, data, and instructions needed to reproduce the main experimental results (either in the supplemental material or as a URL)? [No] We will provide code upon acceptance.

Did you specify all the training details (e.g., data splits, hyperparameters, how they were chosen)? [Yes]

Did you report error bars (e.g., with respect to the random seed after running experiments multiple times)? [No] Multiple runs are computationally infeasible for our models.

Did you include the total amount of compute and the type of resources used (e.g., type of GPUs, internal cluster, or cloud provider)? [Yes] We include total training time and the type of GPU we used.

If you are using existing assets (e.g., code, data, models) or curating/releasing new assets…

If your work uses existing assets, did you cite the creators? [Yes]

Did you mention the license of the assets? [N/A]

Did you include any new assets either in the supplemental material or as a URL? [No]

Did you discuss whether and how consent was obtained from people whose data you’re using/curating? [N/A]

Did you discuss whether the data you are using/curating contains personally identifiable information or offensive content? [N/A]

If you used crowdsourcing or conducted research with human subjects…

Did you include the full text of instructions given to participants and screenshots, if applicable? [N/A]

Did you describe any potential participant risks, with links to Institutional Review Board (IRB) approvals, if applicable? [N/A]

Did you include the estimated hourly wage paid to participants and the total amount spent on participant compensation? [N/A]

Appendix A Implementation

Foreground Generator The foreground generator builds on the StyleGAN2 generator replacing 2D operations with their 3D equivalent as described in the main paper. For faster training, we consider the layers of StyleGAN2 instead of their alias-free version proposed in StyleGAN3 . The mapping network has 2 layers with 64 channels. Since 3D convolutions have more parameters than 2D convolutions with the same channel size, we reduce the channel basesee https://github.com/NVlabs/stylegan3.git from 3276832768 for StyleGAN2 to 40004000. To facilitate progressive growing we choose an architecture with skip connections which adds an upsampled version of the output grid of the previous layer to the input of the next layer. The skip architecture is equivalent to the 2D variant proposed in .

Following , we condition the generator on a camera pose. Specifically, we condition the generator on a rotation matrix and a translation vector, yielding a 1212-dimensional vector as input to the conditioning.

The foreground generator predicts color and density values on a sparse voxel grid. Following Plenoxels , the generator outputs the coefficients of spherical harmonics. We choose spherical harmonics of degree , i.e., a single coefficient for each color channel. For a sharp surface and efficient rendering, the foreground generator needs to predict high values for the density. We facilitate generating high values by multiplying the density output of the network with a factor of 3030. A similar idea was proposed in where the learning rate for the density is set to a higher value than the learning rate for the color.

Sparse Convolutions We investigate sparse convolutions for the foreground generator but find that the computational overhead of managing coordinates increases runtime for our architecture. We therefore use dense convolutions and zero out values in the feature maps for pruned voxels. We also compare the difference for both implementations on performance. In general, we observe similar training behavior for both implementations. However, the faster dense implementation allows us to train the model for 60M iterations compared to 30M for the model with sparse convolutions. This improves FID from 14.4 for the sparse implementation to 9.0 for the dense implementation on FFHQ 256.

Background Generator We use the StyleGAN2 generator with a 2-layer mapping network with 6464 channels, and a synthesis network with channel base 20482048 and a maximum of 6464 channels per layer.

2D Refinement Layers Depending on the dataset, we optionally refine the rendered image with a shallow 2D CNN with 22 hidden layers of dimension 1616 and kernel size 33. To avoid texture sticking under viewpoint changes we use alias-free layers with critical sampling.

Regularization Without regularization, volume rendering is prone to result in semi-opaque voxels and floating artifacts and struggles to accurately represent sharp surfaces . Therefore, we regularize both the variance of the depth, as described in the main paper, and the total variation of the predicted density. For the depth variance loss LDV\mathcal{L}_{DV}, we set τ=(1.5δ0)2\tau=\left(1.5\delta_{0}\right)^{2} where δ0\delta_{0} is the size of one voxel and λDV=0.01\lambda_{DV}=0.01.

Following Plenoxels , we regularize the total variation of the predicted density values in the set of all voxels V\mathcal{V} for a compact, smooth geometry

with Δx2(σ)\Delta^{2}_{x}(\sigma) shorthand for (σi,j,k−σi+1,j,k)2(\sigma_{i,j,k}-\sigma_{i+1,j,k})^{2} and analogously for Δy2(σ)\Delta^{2}_{y}(\sigma) and Δz2(σ)\Delta^{2}_{z}(\sigma). For efficiency, we evaluate the loss stochastically on random contiguous segments of voxels as proposed in and set λTV=10−5\lambda_{TV}=10^{-5} in all experiments.

To avoid that the full image is generated by either background or foreground generator, we use a hinge loss on the mean mask value as proposed in GIRAFFE-HD

where ηfg\eta^{fg} and ηbg\eta^{bg} denote the minimum fraction that should be covered by foreground and background, respectively. To ensure that both models are used, we set ηfg=0.4\eta^{fg}=0.4 for AFHQ and FFHQ and ηfg=0.1\eta^{fg}=0.1 for Carla because Carla’s objects cover a much smaller fraction of the image. We use ηbg=0.1\eta^{bg}=0.1 and λcvgfg=λcvgbg=0.1\lambda_{cvg}^{fg}=\lambda_{cvg}^{bg}=0.1 for all datasets.

Discriminator We use the StyleGAN2 discriminator with conditional input as in . To facilitate progressive growing we choose a skip architecture which adds a downsampled version of the input image to the input of each layer as introduced in .

Rendering For efficient rendering, we leverage custom CUDA kernels building on the official code release of . We select equidistant sampling points for volume rendering in steps of 0.50.5 voxels but skip voxels with σi<10−10\sigma_{i}<10^{-10} and stop rendering early if Ti<10−7T_{i}<10^{-7} as in .

Implementation and Training Our code base builds on the official PyTorch implementation of StyleGAN2 available at https://github.com/NVlabs/stylegan3. Similar to StyleGAN2, we train with equalized learning rates for the trainable parameters and a minibatch standard deviation layer at the end of the discriminator and apply an exponential moving average of the generator weights. For faster training, we use mixed-precision for both the generator and the discriminator as proposed in . Unlike , we do not train with path regularization or style-mixing. To reduce computational cost and overall memory usage R1-regularization is applied only once every 4 minibatches. We use a regularization strength of γ=1\gamma=1 for all datasets. Due to the small size of AFHQ, we follow and finetune a generator that is pretrained on FFHQ with RI=128R_{I}=128 and RG=32R_{G}=32, i.e., before the representation is pruned.

Appendix B Baselines

Qualitative Results We provide qualitative results for StyleNeRF and GRAM in both the main paper and the supplemental material. For StyleNeRF, we obtain samples from the pretrained FFHQ model available at https://github.com/facebookresearch/StyleNeRF.git. For GRAM, its authors kindly provided unpublished code and pretrained models which we use for evaluation. In the qualitative comparisons, i.e., Fig. 4 of the main paper and Fig. 8, we use a truncation of ψ=0.7\psi=0.7 for all methods. For GRAM and our approach, we show samples from −40∘-40^{\circ} to +40∘+40^{\circ} which roughly corresponds to 22 standard deviations of the pose distribution. We find that StyleNeRF does not necessarily adhere to the input pose. Hence, we manually define the range to be −60∘-60^{\circ} to +60∘+60^{\circ} such that the rendered images roughly align with the other methods.

Quantitative Results Table 2 of the main paper shows a quantitative comparison for all baselines and our method in terms of FID. For EG3D and GIRAFFE , we report the numbers from . For StyleNeRF , we take the numbers from . From we further reference the results on FFHQ and AFHQ for GRAF and π\pi-GAN as these datasets were not considered in the original publications. On Carla, we report the results from and , respectively. For GOF , we reference the numbers from on Carla and train new models to obtain results on FFHQ and AFHQ using the official code release available at https://github.com/SheldonTsui/GOF_NeurIPS2021.git. For GRAM , we report the results from for FFHQ. As AFHQ is not considered in , we train GRAM on AFHQ using their unpublished code. Similar to our approach, we finetune a generator that was pre-trained on FFHQ. We remark that across all reported values for FID the number of generated image varies where most methods report values considering either 20k or 50k generated images. An overview is provided in Table 4 which is discussed in more detail at the end of Section C.

Appendix C Results

Background and Foreground Disentanglement.

Fig. 6 illustrates foreground masks, foreground and background image, and the image after alpha composition. As the background remains fixed under viewpoint changes, the generator is encouraged to model pose-dependent content with the foreground generator. The regularization in Eq. (14) and Eq. (15) encourages the generator to use both the background and the foreground generator to synthesize the full image.

Corresponding to Fig. 4 of the main paper, we provide additional qualitative comparisons on multi-view consistency in Fig. 8. For StyleNeRF , the red boxes highlight inconsistencies, like changing eye shape (first row on the left), moving strands of hair (second row, third row on the left) and distortion of the face shape (first and third row on the right). For GRAM , red boxes indicate layered artifacts stemming from its manifold representation. In contrast, our method leads to more multi-view consistent results. We further include a quantitative evaluation on consistency in Table 6. We implemented our own version of the depth and pose metric following the description in as their evaluation code is not publicly available. We report the results for our approach, GRAM, StyleNeRF and EG3D for reference. We also report the standard deviation across the 10241024 samples used for evaluation. While the results in Table 6 agree with our qualitative analysis of view consistency in Fig. 4 and the supplementary video, we find that both metrics are very sensitive to the latent code and the sampled poses, as indicated by the large standard deviations. As no established evaluation pipeline exists, our results should not be directly compared to the numbers in as the implementation and pose sampling might differ.

Pose Conditioning. Fig. 9 illustrates the impact of the pose conditioning on the generated images. Pose conditioning is not only used to model slight changes, e.g. of the eyes or the smile, but can alter the general appearance of the generated instance. However, by fixing the pose conditioning during inference, view-consistent images can be generated.

Regularization. We ablate the effect of regularization in terms of end-to-end performance measured in FID. Table 6 shows that the regularizers do not significantly change FID. Nonetheless, LDV\mathcal{L}_{DV} speeds up training (see Table 1 of the main paper) and we find the remaining losses to be helpful for stabilizing training.

Failure Cases. Fig. 7 illustrates failure cases of out method. For some samples, we observe that the background and the foreground are not disentangled properly. The left column in Fig. 7 shows an example where the foreground generates parts of the background, as indicated by the red boxes. In turn, the background models the body of the person which should be part of the foreground. Learning to disentangle background and foreground without direct supervision, e.g., instance masks, is often ambiguous which makes it a challenging task. The middle column in Fig. 7 displays a common failure case of our model on AFHQ: The whiskers of the cat are connected to its body. This is likely a consequence from the depth variance and total variation regularization we apply to obtain a single sharp surface and compact geometry. The right column in Fig. 7 shows an occasional failure on FFHQ. For some samples, we observe that the hair is directed inward to the head instead of outward. In the rendered image this effect is not visible which suggests that it likely results from ambiguity in the training data.

Uncurated Samples. We provide additional samples of our method for FFHQ in Fig. 10, AFHQ in Fig. 11, and Carla in Fig. 12.

Quantitative Comparison. As FID is biased towards the number of images , we explicitly annotate the number of generated images for the results reported in Table 2 of the main paper in Table 4. Most works evaluate FID either using 20k or 50k generated images. We evaluate our method for both cases. For our method and the considered datasets, the difference between using 20k or 50k generated images is reasonably small.

Appendix D Societal Impact

This work considers the task of generating photorealistic renderings of scenes with data-driven approaches which has potential downstream applications in virtual reality, augmented reality, gaming and simulation. While many use-cases are possible, we believe that in the long run this line of research could support designers in creating renderings of 3D models more efficiently. However, generating photorealistic 3D-scenarios also bears the risk of manipulation, e.g., by creating edited imagery of real people. Further, like all data-driven approaches, our method is susceptible to biases in the training data. Such biases can, e.g., result in a lack of diversity for the generated faces and have to be addressed before using this work for any downstream applications.