EpiGRAF: Rethinking training of 3D GANs

Ivan Skorokhodov, Sergey Tulyakov, Yiqun Wang, Peter Wonka

Introduction

Generative models for image synthesis achieved remarkable success in recent years and now enjoy a lot of practical applications . While initially they mainly focused on 2D images , recent research explored generative frameworks with partial 3D control over the underlying object in terms of texture/structure decomposition, novel view synthesis or lighting manipulation (e.g., ). These techniques are typically built on top of the recently emerged neural radiance fields (NeRF) to explicitly represent the object (or its latent features) in 3D space.

People address these scaling issues of NeRF-based GANs in different ways. The dominating approach is to train a separate 2D decoder to produce a high-resolution image from a low-resolution image or feature grid rendered from a NeRF backbone . During the past six months, there appeared more than a dozen of methods that follow this paradigm (e.g., ). While using the upsampler allows scaling the model to high resolution, it comes with two severe limitations: 1) it breaks the multi-view consistency of a generated object, i.e., its texture and shape change when the camera moves; and 2) the geometry gets only represented in a low resolution (≈643{\approx}64^{3}). In our work, we show that by dropping the upsampler and using a simple patch-wise optimization scheme, one can build a 3D generator with better image quality, faster training speed, and without the above limitations.

Patch-wise training of NeRF-based GANs was initially proposed by GRAF and got largely neglected by the community since then. The idea is simple: instead of training the generative model on full-size images, one does this on small random crops. Since the model is coordinate-based , it does not face any issues to synthesize only a subset of pixels. This serves as an excellent way to save computation for both the generator and the discriminator since it makes them both operate on patches of small spatial resolution. To make the generator learn both the texture and the structure, crops are sampled to be of variable scales (but having the same number of pixels). In some sense, this can be seen as optimizing the model on low-resolution images + high-resolution patches.

In our work, we improve patch-wise training in two crucial ways. First, we redesign the discriminator by making it better suited to operating on image patches of variable scales and locations. Convolutional filters of a neural network learn to capture different patterns in their inputs depending on their semantic receptive fields . That’s why it is detrimental to reuse the same discriminator to judge both high-resolution local and low-resolution global patches, inducing additional burden on it to mix filters’ responses of different scales. To mitigate this, we propose to modulate the discriminator’s filters with a hypernetwork , which predicts which filters to suppress or reinforce from a given patch scale and location.

Second, we change the random scale sampling strategy from an annealed uniform to an annealed beta distribution. Typically, patch scales are sampled from a uniform distribution s∼U[s(t),1]s\sim\mathcal{U}[s(t),1] , where the minimum scale s(t)s(t) is gradually decreased (i.e. annealed) till some iteration TT from s(0)=0.9s(0)=0.9 to a smaller value s(T)s(T) (in the interval [0.125−0.5][0.125-0.5]) during training. This sampling strategy prevents learning high-frequency details early on in training and puts too little attention on the structure after s(t)s(t) reaches its final value s(T)s(T). This makes the overall convergence of the generator slower and less stable that’s why we propose to sample patch scales using the beta distribution Beta(1,β(t))\text{Beta}(1,\beta(t)) instead, where β(t)\beta(t) is gradually annealed from β(0)≈0\beta(0)\approx 0 to some maximum value β(T)\beta(T). In this way, the model starts learning high-frequency details immediately with the start of training and focuses more on the structure after the growth finishes. This simple change stabilizes the training and allows it to converge faster than the typically used uniform distribution .

We use those two ideas to develop a novel state-of-the-art 3D GAN: Efficient patch-informed Generative Radiance Fields (EpiGRAF). We employ it for high-resolution 3D-aware image synthesis on four datasets: FFHQ , Cats , Megascans Plants, and Megascans Food. The last two benchmarks are introduced in our work and contain 360∘360^{\circ} renderings of photo-realistic scans of different plants and food objects (described in §4). They are much more complex in terms of geometry and are well-suited for assessing the structural limitations of modern 3D-aware generators.

Our model uses a pure NeRF-based backbone, that’s why it represents geometry in high resolution and does not suffer from multi-view synthesis artifacts, as opposed to upsampler-based generators. Moreover, it has higher or comparable image quality (as measured by FID ) and 2.5×2.5\times lower training cost. Also, in contrast to upsampler-based 3D GANs, our generator can naturally incorporate the techniques from the traditional NeRF literature. To demonstrate this, we incorporate background separation into our framework by simply copy-pasting the corresponding code from NeRF++ .

Related work

Neural Radiance Fields. Neural Radiance Fields (NeRF) is an emerging area , which combines neural networks with volumetric rendering techniques to perform novel-view synthesis , image-to-scene generation , surface reconstruction and other tasks . In our work, we employ them in the context of 3D-aware generation from a dataset of RGB images .

3D generative models. A popular way to learn a 3D generative model is to train it on 3D data or in an autoencoder’s latent space (e.g., ). This requires explicit 3D supervision and there appeared methods which train from RGB datasets with segmentation masks, keypoints or multiple object views . Recently, there appeared works which train from single-view RGB only, including mesh-generation methods and methods that extract 3D structure from pretrained 2D GANs . And recent neural rendering advancements allowed to train NeRF-based generators from purely RGB data from scratch, which became the dominating direction since then and which are typically formulated in the GAN-based framework .

NeRF-based GANs. HoloGAN generates a 3D feature voxel grid which is projected on a plane and then upsampled. GRAF trains a noise-conditioned NeRF in an adversarial manner. π\pi-GAN builds upon it and uses progressive growing and hypernetwork-based conditioning in the generator. GRAM builds on top of π\pi-GAN and samples ray points on a set of learnable iso-surfaces. GNeRF adapts GRAF for learning a scene representation from RGB images without known camera parameters. GIRAFFE uses a composite scene representation for better controllability. CAMPARI learns a camera distribution and a background separation network with inverse sphere parametrization . To mitigate the scaling issue of volumetric rendering, many recent works train a 2D decoder under different multi-view consistency regularizations to upsample a low-resolution volumetrically rendered feature grid . However, none of such regularizations can currently provide the multi-view consistency of pure-NeRF-based generators.

Patch-wise generative models. Patch-wise training had been routinely utilized to learn the textural component of image distribution when the global structure is provided from segmentation masks, sketches, latents or other sources (e.g., ). Recently, there appeared works which sample patches at variable scales, in which way a patch can carry global information about the whole image. Recent works use it to train a generative NeRF , fit a neural representation in an adversarial manner or to train a 2D GAN on a dataset of variable resolution .

Model

We build upon StyleGAN2 , replacing its generator with the tri-plane-based NeRF model and using its discriminator as the backbone. We train the model on r×rr\times r patches (we use r=64r=64 everywhere) of random scales instead of the full images of resolution R×RR\times R. Scales s∈[rR,1]s\in[\frac{r}{R},1] are randomly sampled from a time-varying distribution s∼pt(s)s\sim p_{t}(s).

2 2D scale/location-aware discriminator

Our discriminator D\mathsf{D} is built on top of StyleGAN2 . Since we train the model in a patch-wise fashion, the original backbone is not well suited for this: convolutional filters are forced to adapt to signals of very different scales and extracted from different locations. A natural way to resolve this problem is to use separate discriminators depending on the scale, but that strategy has three limitations: 1) each particular discriminator receives less overall training signal (since the batch size is limited); 2) from an engineering perspective, it is more expensive to evaluate a convolutional kernel with different parameters on different inputs; 3) one can use only a small fixed amount of possible patch scales. This is why we develop a novel hypernetwork-modulated discriminator architecture to operate on patches with continuously varying scales.

This suppresses and reinforces different convolutional filters of the layer depending on the patch scale and location. And to incorporate even stronger conditioning, we also use the projection strategy in the final discriminator block. We depict our discriminator architecture in Fig 3. As we show in Tab 2, it allows us to obtain ≈15{\approx}15% lower FID compared to the standard discriminator.

3 Patch-wise optimization with Beta-distributed scales

Training NeRF-based GANs is computationally expensive because rendering each pixel via volumetric rendering requires many evaluations (e.g., in our case, 96) of the underlying MLP. For scene reconstruction tasks, it does not create issues since the typically used L2\mathcal{L}_{2} loss can be robustly computed on a sparse subset of the pixels. But for NeRF-based GANs, it becomes prohibitively expensive for high resolutions since convolutional discriminators operate on dense full-size images. The currently dominating approach to mitigate this is to train a separate 2D decoder to upsample a low-resolution image representation rendered from a NeRF-based MLP. But this breaks multi-view consistency (i.e., object’s shape and texture change when the camera is moving) and learns the 3D geometry in a low resolution (from ≈162{\approx}16^{2} to ≈1282{\approx}128^{2} ). This is why we build upon the multi-scale patch-wise training scheme and demonstrate that it can give state-of-the-art image quality and training speed without the above limitations.

Patch-wise optimization works the following way. On each iteration, instead of passing the full-size R×RR\times R image to D\mathsf{D}, we instead input only a small patch with resolution r×rr\times r of random scale s∈[r/R,1]s\in[r/R,1] and extracted with a random offset (δx,δy)∈[0,1−s]2(\delta_{x},\delta_{y})\in[0,1-s]^{2}. We illustrate this procedure in Fig 3. Patch parameters are sampled from distribution:

where tt is the current training iteration. In this way, patch scales depend on the current training iteration tt, and offsets are sampled independently after we know ss. As we show next, the choice of distribution pt(s)p_{t}(s) has a crucial influence on the learning speed and stability.

Typically, patch scales are sampled from the annealed uniform distribution ss:

To mitigate this, we propose a small change in the pipeline by simply replacing the uniform scale sampling distribution with:

where β(t)\beta(t) is gradually annealed from β(0)\beta(0) to some final value β(T)\beta(T). Using beta distribution instead of the uniform one gives a very convenient knob to shift the training focus between large patch scales s→1s\to 1 (carrying the global information about the whole image) and small patch scales r→r/Rr\to r/R (representing high-resolution local crops).

A natural way to do the annealing is to anneal from 0 to 10\text{ to }1: at the start, the model focuses entirely on the structure, while at the end, it transforms into the uniform distribution (See Fig 4). We follow this strategy, but from the design perspective, set β(T)\beta(T) to a value that is slightly smaller than 1 (we use β(T)=0.8\beta(T)=0.8 everywhere) to keep more focus on the structure at the end of the annealing as well. In our initial experiments, β(T)∈[0.7,1]\beta(T)\in[0.7,1] performs similarly. The scales distributions comparison between beta and uniform sampling is provided in Fig 4 and the convergence comparison in Fig 7.

4 Training details

We inherit the training procedure from StyleGAN2-ADA with minimal changes. The optimization is performed by Adam with a learning rate of 0.002 and betas of 0 and 0.99 for both G\mathsf{G} and D\mathsf{D}. We use β(T)=0.8\beta(T)=0.8 for T=10000T=10000, z∼N(0,I)\bm{z}\sim\mathcal{N}(\bm{0},\bm{I}) and set Rp=512R_{p}=512. D\mathsf{D} is trained with R1 regularization with γ=0.05\gamma=0.05. We train with the overall batch size of 64 for ≈15{\approx}15M images seen by D\mathsf{D} for 2562256^{2} resolution and ≈20{\approx}20M for 5122512^{2}. Similar to previous works , we use pose supervision for D\mathsf{D} for the FFHQ and Cats dataset to avoid geometry ambiguity. For this, we take the rotation and elevation angles, encode them with positional embeddings and feed them into a 2-layer MLP. After that, we multiply the obtained vector with the last hidden representation in the discriminator, following the Projection GAN strategy from StyleGAN2-ADA . We train G\mathsf{G} in full precision and use mixed precision for D\mathsf{D}. Since FFHQ has too noticeable 3D biases, we use generator pose conditioning for it . Further details can be found in the source code.

Experiments

Benchmarks. In our study, we consider four benchmarks: 1) FFHQ in 2562256^{2} and 5122512^{2} resolutions, consisting of 70,000 (mostly front-view) human face images; 2) Cats 2562256^{2} , consisting of 9,998 (mostly front-view) cat face images; 3) Megascans Food (M-Food) 2562256^{2} consisting of 231 models of different food items with 128 views per model (25472 images in total); and 4) Megascans Plants (M-Plants) 2562256^{2} consisting of 1166 different plant models with 128 views per model (141824 images in total). The last two datasets are introduced in our work to fix two issues with the modern 3D generation benchmarks. First, existing benchmarks have low variability of global object geometry, focusing entirely on a single class of objects, like human/cat faces or cars, that do not vary much from instance to instance. Second, they all have limited camera pose distribution: for example, FFHQ and Cats are completely dominated by the frontal and near-frontal views (see Appx E). That’s why we obtain and render 1307 Megascans models from Quixel, which are photo-realistic (barely distinguishable from real) scans of real-life objects with complex geometry. Those benchmarks and the rendering code will be made publicly available.

Metrics. We use FID to measure image quality and estimate the training cost for each method in terms of NVidia V100 GPU days needed to complete the training process.

Baselines. For upsampler-based baselines, we compare to the following generators: StyleNeRF , StyleSDF , EG3D , VolumeGAN , MVCGAN and GIRAFFE-HD . Apart from that, we also compare to pi-GAN and GRAM , which are non-upsampler-based GANs. To compare on Megascans, we train StyleNeRF, MVCGAN, pi-GAN, and GRAM from scratch using their official code repositories (obtained online or requested from the authors), using their FFHQ or CARLA hyperparameters, except for the camera distribution and rendering settings. We also train StyleNeRF, MVCGAN, and π\pi-GAN on Cats 2562256^{2}. GRAM restricts the sampling space to a set of learnable iso-surfaces, which makes it not well-suited for datasets with varying geometry.

2 Results

EpiGRAF achieves state-of-the-art image quality. For Cats 2562256^{2}, M-Plants 2562256^{2} and M-Food 2562256^{2}, EpiGRAF outperforms all the baselines in terms of FID except for StyleNeRF, performing very similar to it on all the datasets even though it does not have a 2D upsampler. For FFHQ, our model attains very similar FID scores as the other methods, ranking 4/9 (including older π\pi-GAN ), noticeably losing only to EG3D , which trains and evaluates on a different version of FFHQ and uses pose conditioning in the generator (which potentially improves FID at the cost of multi-view consistency). We provide a visual comparison for different methods in Fig 5.

EpiGRAF is much faster to train. As reported in Tab 1, existing methods typically train for ≈1{\approx}1 week on 8 V100s, EpiGRAF finishes training in just 2 days for 2562256^{2} and 33 days for 5122512^{2} resolutions, which is 2−3×2-3\times faster. Note that this high training efficiency is achieved without using an upsampler, which initially enabled the high-resolution synthesis of 3D-aware GANs. As to the non-upsampler methods, we couldn’t train GRAM or π\pi-GAN on 5122512^{2} resolution due to the memory limitations of the setup with 8 NVidia V100 32GB GPUs (i.e., 256GB of GPU memory in total).

EpiGRAF learns high-fidelity geometry. Using a pure NeRF-based backbone has two crucial benefits: it provides multi-view consistency and allows learning the geometry in the full dataset resolution. In Fig 6, we visualize the learned shapes on M-Food and M-Plants for 1) π\pi-GAN: a pure NeRF-based generator without the geometry constraints; 2) MVC-GAN : an upsampler-based generator with strong multi-view consistency regularization; 3) our model. We provide the details and analysis in the caption of Fig 6. We also provide the geometry comparison with EG3D on FFHQ 5122512^{2} in Fig 2.

EpiGRAF easily capitalizes on techniques from the NeRF literature. Since our generator is purely NeRF based and renders images without a 2D upsampler, it is well coupled with the existing techniques from the NeRF scene reconstruction field. To demonstrate this, we adopted background separation from NeRF++ using the inverse sphere parametrization by simply copy-pasting the corresponding code from their repo. We depict the results in Fig 1 and provide the details in Appx B.

3 Ablations

We report the ablations for different discriminator architectures and patch sizes on FFHQ 5122512^{2} and M-Plants 2562256^{2} in Tab 2. Using a traditional discriminator architecture results in ≈15%{\approx}15\% worse performance. Using several ones (via the group-wise convolution trick ) results in a noticeably slower training time and dramatically degrades the image quality. We hypothesize that the reason for it was the reduced overall training signal each discriminator receives, which we tried to alleviate by increasing their learning rate, but that did not improve the results. A too-small patch size hampers the learning process and produces a ≈80{\approx}80% worse FID. A too-large one provides decent image quality but greatly reduces the training speed. Using a single scale/position-aware discriminator achieves the best performance, outperforming the standard one by ≈15%{\approx}15\% on average.

To assess the convergence of our proposed patch sampling scheme, we compared against uniform sampling on Cats 2562256^{2} for T∈{1000,5000,10000}T\in\{1000,5000,10000\}, representing different annealing speeds. We show the results for it in Fig 7: our proposed beta scale sampling strategy with T=10T=10k schedule robustly converges to lower values than the uniform one with T=5T=5k or T=10T=10k and does not fluctuate much compared to the T=1kT=1k uniform one (where the model reached its final annealing stage in just 1k kilo-images seen by D\mathsf{D}).

To analyze how hyper-modulation manipulates the convolutional filters of the discriminator, we visualize the modulation weights σ\bm{\sigma}, predicted by H\mathsf{H}, in Fig 8 (see the caption for the details). These visualizations show that some of the filters are always switched on, regardless of the patch scale; while others are always switched off, providing potential room for pruning . And ≈40{\approx}40% of the filters are getting switched on and off depending on the patch scale, which shows that H\mathsf{H} indeed learns to perform meaningful modulation.

Limitations

Performance drop for 2D generation. Before switching to training 3D-aware generators, we spent a considerable amount of time, exploring our ideas on top of StyleGAN2 for traditional 2D generation since it is faster, less error-prone and more robust to a hyperparameters choice. What we observed is that despite our best efforts (see C) and even with longer training, we couldn’t obtain the same image quality as the full-resolution StyleGAN2 generator.

A range of possible patch sizes is restricted. Tab 2 shows the performance drop when using the 32232^{2} patch size instead of the default 64264^{2} one without any dramatic improvement in speed. Trying to decrease it further would produce even worse performance (imagine training in the extreme case of 222^{2} patches). Increasing the patch size is also not desirable since it decreases the training speed a lot: going from 64264^{2} to 1282128^{2} resulted in 30% cost increase without clear performance benefits. In this way, we are very constrained in what patch size one can use.

Discriminator does not see the global context. When the discriminator classifies patches of small scale, it is forced to do so without relying on the global image information, which could be useful for this. Our attempts to incorporate it (see Appx C) did not improve the performance (though we believe we under-explored this).

Low-resolution artifacts. While our generator achieves good FID on FFHQ 5122512^{2}, we noticed that it has some blurriness when one zooms-in into the samples. It is not well captured by FID since it always resizes images to the 299×299299\times 299 resolution. We attribute this problem to our patch-wise training scheme, which puts too much focus on the structure and believe that it could be resolved.

Conclusion

In this work, we showed that it is possible to build a state-of-the-art 3D GAN framework without a 2D upsampler, but using a pure NeRF-based generator trained in a multi-scale patch-wise fashion. For this, we improved the traditional patch-wise training scheme in two important ways. First, we proposed to use a scale/location-aware discriminator with convolutional filters modulated by a hypernetwork depending on the patch parameters. Second, we developed a schedule for patch scale sampling based on the beta distribution, that leads to faster and more robust convergence. We believe that the future of 3D GANs is a combination of efficient volumetric representations, regularized 2D upsamplers, and patch-wise training. We propose this avenue of research for future work.

Our method also has several limitations. Before switching to training 3D-aware generators, we spent a considerable amount of time exploring our ideas on top of StyleGAN2 for traditional 2D generation, which always resulted in higher FID scores. Further, the discriminator loses information about global context. We tried multiple ideas to incorporate global context, but it did not lead to an improvement. Next, our current patch-wise training scheme might cause some low-res artifacts. Finally, 3D GANs generating faces and humans may have negative societal impact as discussed in Appx H.

Acknowledgements

We would like to acknowledge support from the SDAIA-KAUST Center of Excellence in Data Science and Artificial Intelligence.

References

Checklist

Do the main claims made in the abstract and introduction accurately reflect the paper’s contributions and scope? [Yes]

Did you describe the limitations of your work? [Yes] See §5 and Appx A.

Did you discuss any potential negative societal impacts of your work? [Yes] We do this in Appendix G.

Have you read the ethics review guidelines and ensured that your paper conforms to them? [Yes] We discuss the potential ethical concerns of using our model in Appendix G.

If you are including theoretical results…

Did you state the full set of assumptions of all theoretical results? [N/A]

Did you include complete proofs of all theoretical results? [N/A]

Did you include the code, data, and instructions needed to reproduce the main experimental results (either in the supplemental material or as a URL)? [Yes] We provide the code/data and additional visualizations on https://universome.github.io/epigraf(as specified in the introduction).

Did you specify all the training details (e.g., data splits, hyperparameters, how they were chosen)? [Yes] We provide the most important training details in §3.4. The rest of the details are provided in Appx B and the provided source code.

Did you report error bars (e.g., with respect to the random seed after running experiments multiple times)? [No] . That’s too computationally expensive and single-run results are typically reliable in the GAN field.

Did you include the total amount of compute and the type of resources used (e.g., type of GPUs, internal cluster, or cloud provider)? [Yes] We report this numbers in Appx B.

If you are using existing assets (e.g., code, data, models) or curating/releasing new assets…

If your work uses existing assets, did you cite the creators? [Yes] We cite all the sources of the datasets which were used or mentioned in our submission.

Did you mention the license of the assets? [Yes] In this work, we release two new datasets: Megascans Plants and Megascans Food. We discuss their licensing in Appx E.

Did you include any new assets either in the supplemental material or as a URL? [Yes] We provide our datasets on the project website: https://universome.github.io/epigraf.

Did you discuss whether and how consent was obtained from people whose data you’re using/curating? [Yes] We specify the information on dataset collection in Appx E.

Did you discuss whether the data you are using/curating contains personally identifiable information or offensive content? [Yes] As discussed in Appx E, the released data does not contain personally identifiable information or offensive content.

If you used crowdsourcing or conducted research with human subjects…

Did you include the full text of instructions given to participants and screenshots, if applicable? [N/A]

Did you describe any potential participant risks, with links to Institutional Review Board (IRB) approvals, if applicable? [N/A]

Did you include the estimated hourly wage paid to participants and the total amount spent on participant compensation? [N/A]

Appendix A Additional limitations due to aliasing

Current patch-wise training strategies (both ours and in prior works ) do not take aliasing into account when extracting patches. This leads to additional problems, which we illustrate in Figure 9. Basically, sampling patches from images (when performed naively) is prone to aliasing, which can potentially result in learning an incorrect distribution.

If we rely on the GAN convergence theorem , stating that we recover the training distribution as the solution of the minimax problem, then G\mathsf{G} will learn to approximate all the possible marginal distributions p(xp)p(\bm{x}_{p}) instead of the full joint distribution p(x)p(\bm{x}), that we seek.

Sampling patches in an alias-free manner is tricky for our generator, since we cannot obtain intermediate high-resolution representations (which are needed to address aliasing) from tri-planes due to the computational overhead. We leave the development of patch sampling schemes for NeRF-based generators for future work.

Appendix B Training details

We inherit most of the hyperparameters from the StyleGAN2-ADA repo repo which we build on tophttps://github.com/NVlabs/stylegan2-ada-pytorch. In this way, we use the dimensionalities of 512512 for both z\bm{z} and w\bm{w}. The mapping network has 2 layers of dimensionality 512 with LeakyReLU non-linearities with the negative slope of −0.2-0.2. Synthesis network S\mathsf{S} produced three 5122512^{2} planes of 3232 channels each. We use the SoftPlus non-linearity instead of typically used ReLU as a way to clamp the volumetric density. Similar to π\pi-GAN, we also randomize

For FFHQ and Cats, we also use camera conditioning in D\mathsf{D}. For this, we encode yaw and pitch angles (roll is always set to 0) with Fourier positional encoding , apply dropout with 0.5 probability (otherwise, D\mathsf{D} can start judging generations from 3D biases in the dataset, hurting the image quality), pass through a 2-layer MLP with LeakyReLU activations to obtain a 512-dimensional vector, which is finally as a projection conditioning . Cameras positions were extracted in the same way as in GRAM .

We optimize both G\mathsf{G} and D\mathsf{D} with the batch size of 64 until D\mathsf{D} sees 25,000,000 real images, which is the default setting from StyleGAN2-ADA. We the default setup of adaptive augmentations, except for random horizontal flipping, since it would require the corresponding change in the yaw angle at augmentation time, which was not convenient to incorporate from the engineering perspective. Instead, random horizontal flipping is used non-adaptively as a dataset mirroring where flipping the yaw angles is more accessible. We train G\mathsf{G} in full precision, while D\mathsf{D} uses mixed precision.

For the background separation experiment, we adapt the neural representation MLP from INR-GAN , but passing 4 coordinates (for the inverse sphere parametrization ) instead of 2 as an input. It consists of 2 blocks with 2 linear layers each. We use 16 steps per ray for the background without hierarchical sampling.

Further details could be found in the accompanying source code.

B.2 Utilized computational resources

While developing our model, we had been launching experiments on 4×4\times NVidia A100 81GB or Nvidia V100 32GB GPUs with the AMD EPYC 7713P 64-Core processor. We found that in practice, running the model on A100s gives a 2×2\times speed-up compared to V100s due to the possibility of increasing the batch size from 32 to 64. In this way, training EpiGRAF on 4×4\times A100s gives the same training speed as training it 8×8\times V100s.

For the baselines, we were running them on 4-8×\times V100s GPUs as was specified by the original papers unless the model could fit into 4 V100s without decreasing the batch size (it was only possible for StyleNeRF ).

For rendering Megascans, we used 4×4\times NVIDIA TITAN RTX with 24GB memory each. But resource utilization for rendering is negligible compared to training the generators.

In total, the project consumed ≈4{\approx}4 A100s GPU-years, ≈4{\approx}4 V100s GPU-years, and ≈20{\approx}20 TITAN RTX GPU-days. Note, that out of this time, training the baselines consumed ≈1.5{\approx}1.5 V100s GPU-years.

B.3 Annealing schedule details

As being said in §3.3, the existing multi-scale patch-wise generators use uniform distribution U[smin⁡(t),1]U[s_{\min(t)},1] to sample patch scales, where smin⁡(t)s_{\min(t)} is gradually annealed during training from 0.9 (or 0.8 ) to r/Rr/R with different speeds. We visualize the annealing schedule for both GRAF and GNeRF on Fig 10, which demonstrates that their schedules are very close to lerp-based one, described in §3.3.

Appendix C Failed experiments

Modern GANs are a lot of engineering and it often takes a lot of futile experiments to get to a point where the obtained performance is acceptable. We want to enumerate some experiments which did not work out (despite looking like they should work) — either because the idea was fundamentally flawed on its own or because we’ve under-explored it (or both).

Conditioning D\mathsf{D} on global context worsened the performance. In Appx 5, we argued that when D\mathsf{D} processes a small-scale patch, it does not have access to the global image information, which might be a source of decreased image quality. We tried several strategies to compensate for this. Our first attempt was to generate a low-resolution image, bilinearly upsample it to the target size, and then “grid paste” a high-resolution patch into it. The second attempt was to simply always concatenate a low-resolution version of an image as 3 additional channels. However, in both cases, generator learned to produce low-resolution version of images well, but the texture was poor. We hypothesize that it was due to D\mathsf{D} starting to produce its prediction almost entirely based on the low-resolution image, ignoring the high-resolution patches since they are harder to discriminate.

Patch importance sampling did not work. Almost all the datasets used for 3D-aware image synthesis have regions of difficult content and regions with simpler content — it is especially noticeable for CARLA and our Megascans datasets, which contain a lot of white background. That’s why, patch-wise sampling could be improved if we sample patches from the more difficult regions more frequently. We tried this strategy in the GNeRF problem setup on the NeRF-Synthetic dataset of fitting a scene without known camera parameters. We sampled patches from regions with high average gradient norm more frequently. For some scenes, it helped, for other ones, it worsened the performance.

View direction conditioning breaks multi-view consistency. Similar to the prior works , our attempt to condition the radiance (but not density) MLP on ray direction (similar to NeRF ) led to poor multi-view consistency with radiance changing with camera moving. We tested this on FFHQ , which has only a single view per object instance and suspect that it wouldn’t be happening on Megascans, where view coverage is very rich.

Tri-planes produced from convolutional layers are much harder to optimize for reconstruction. While debugging our tri-plane representation, we found that tri-planes produced with convolutional layers are extremely difficult to optimize for reconstruction. I.e., if one fits a 3D scene while optimizing tri-planes directly, then everything goes smoothly, but when those tri-planes are being produced by the synthesis network of StyleGAN2 , then PNSR scores (and the loss values) are plateauing very soon.

Appendix D Additional patch size ablation

Table 2 shows that the generator achieves the best performance for the 64264^{2} patch resolution. The comparison between patch sizes was performed while keeping all other hyperparameters fixed. This creates an issue since StyleGAN-based generators should use different values for the R1 regularization weight γ\gamma depending on the training resolutions.See https://github.com/NVlabs/stylegan2-ada-pytorch/blob/main/train.py#L173. This is why in Table 4, we provide the results for a 3×33\times 3 grid search over patch sizes and R1 regularization gamma.

And this better aligns with intuition: increasing the patch size should improve the performance (at the loss of the training speed) since the model uses more information during training. The main reason why we fixed the patch resolution to 64264^{2} is because we considered the computational overhead not to be worth the quality improvements it brings: while it is not expensive to run several individual experiments with the 1282128^{2} patch resolution, it is expensive to develop the whole project around the 1282128^{2} patch resolution generator.

Appendix E Datasets details

Modern 3D-aware image synthesis benchmarks have two issues: 1) they contain objects of very similar global geometry (like, human or cat faces, cars and chairs), and 2) they have poor camera coverage. Moreover, some of them (e.g., FFHQ), contain 3D-biases, when an object features (e.g., smiling probability, gaze direction, posture or haircut) correlate with the camera position . As a result, this does not allow to evaluate a model’s ability to represent the underlying geometry and makes it harder to understand whether performance comes from methodological changes or better data preprocessing.

To mitigate these issues, we introduce two new datasets: Megascans Plants (M-Plants) and Megascans Food (M-Food). To build them, we obtain ≈1,500{\approx}1,500 models from Quixel Megascanshttps://quixel.com/megascans from Plants, Mushrooms and Food categories. Megascans are very high-quality scans of real objects which are almost indistinguishable from real. For Mushrooms and Plants, we merge them into the same Food category since they have too few models on their own.

We render all the models in Blender with cameras, distributed uniformly at random over the sphere of radius 3.5 and field-of-view of π/4\pi/4. While rendering, we scale each model into 3^{3} cube and discard those models, which has the dimension produce of less than 2. We render 128 views per object from a fixed distance to the object center from uniformly sampled points on the entire sphere (even from below). For M-Plants, we additionally remove those models which have less than 0.03 pixel intensity on average (computed as the mean alpha value over the pixels and views). This is needed to remove small grass or leaves which will be occupying a too small amount of pixels. As a result, this procedure produces 1,108 models for the Plants category and 199 models for the Food category.

We include the rendering script as a part of the released source code. We cannot release the source models or textures due to the copyright restrictions. We release all the images under the CC BY-NC-SA 4.0 licensehttps://creativecommons.org/licenses/by-nc-sa/4.0. Apart from the images, we also release the class categories for both M-Plants and M-Food.

The released datasets do not contain any personally identifiable information or offensive content since it does not have any human subjects, animals or other creatures with scientifically proven cognitive abilities. One concern that might arise is the inclusion of Amanita muscariahttps://en.wikipedia.org/wiki/Amanita_muscaria into the Megascans Food dataset, which is poisonous (when consumed by ingestion without any specific preparation). This is why we urge the reader not to treat the included objects as edible items, even though they are a part of the “food” category. We provide random samples from both of them in Fig 11 and Fig 12. Note that they are almost indistinguishable from real objects.

E.2 Datasets statistics

We provide the datasets statistics in Tab 5. For CARLA , we provide them for comparison and do not use this dataset as a benchmark since it is small, has simple geometry and texture.

Appendix F Additional samples

We provide random non-cherry-picked samples from our model in Fig 14, but we recommend visiting the website for video illustrations: https://universome.github.io/epigraf.

To demonstrate GRAM’s mode collapse, we provide more its samples in Fig 15.

Appendix G Potential negative societal impacts

Our developed method is in the general family of media synthesis algorithms, that could be used for automatized creation and manipulation of different types of media content, like images, videos or 3D scenes. Of particular concern is creation of deepfakeshttps://en.wikipedia.org/wiki/Deepfake — photo-realistic replacing of one person’s identity with another one in images and videos. While our model does not yet rich good enough quality to have perceptually indistinguishable generations from real media, such concerns should be kept in mind when developing this technology further.

Appendix H Ethical concerns

We have reviewed the ethics guidelineshttps://nips.cc/public/EthicsGuidelines and confirm that our work complies with them. As discussed in Appx E.1, our released datasets are not human-derived and hence do not contain any personally identifiable information and are not biased against any groups of people.