DiffRF: Rendering-Guided 3D Radiance Field Diffusion

Norman Müller, Yawar Siddiqui, Lorenzo Porzi, Samuel Rota Bulò, Peter Kontschieder, Matthias Nießner

Introduction

In recent years, Neural Radiance Fields (NeRFs) have emerged as a powerful representation for fitting individual 3D scenes from posed 2D input images. The ability to photo-realistically synthesize novel views from arbitrary viewpoints while respecting the underlying 3D scene geometry has the potential to disrupt and transform applications like AR/VR, gaming, mapping, navigation, etc. A number of recent works have introduced extensions for making NeRFs more sophisticated, by e.g., showing how to incorporate scene semantics , training models from heterogeneous data sources , or scaling them up to represent large-scale scenes . These advances are testament to the versatility of ML-based scene representations; however, they still fit to specific, individual scenes rather than generalizing beyond their input training data.

In contrast, neural field representations that generalize to multiple object categories or learn priors for scenes across datasets appear much more limited to date, despite enabling applications like single-image 3D object generation and unconstrained scene exploration . These methods explore ways to disentangle object priors into shape and appearance-based components, or to decompose radiance fields into several small and locally-conditioned radiance fields to improve scene generation quality; however, their results still leave significant gaps w.r.t. photorealism and geometric accuracy.

Directions involving generative adversarial networks (GANs) that have been extended from the 2D domain to 3D-aware neural fields generation are demonstrating impressive synthesis results . Like regular 2D GANs, the training objective is based on discriminating 2D images, which are obtained by rendering synthesized 3D radiance fields.

At the same time, diffusion-based models have recently taken the computer vision research community by storm, performing on-par or even surpassing GANs on multiple 2D benchmarks, and are producing photo-realistic images that are almost indistinguishable from real photographs. For multi-modal or conditional settings such as text-to-image synthesis, we currently observe unprecedented output quality and diversity from diffusion-based approaches. While several works address purely geometric representations , lifting the denoising-diffusion formulation directly to 3D volumetric radiance fields remains challenging. The main reason lies in the nature of diffusion models, which require a one-to-one mapping between the noise vector and the corresponding ground truth data samples. In the context of radiance fields, such volumetric ground truth data is practically infeasible to obtain, since even running a costly per-sample NeRF optimization results in incomplete and imperfect radiance field reconstructions.

In this work, we present the first diffusion-based generative model that directly synthesizes 3D radiance fields, thus unlocking high-quality 3D asset generation for both shape and appearance. Our goal is to learn such a generative model trained across objects, where each sample is given by a set of posed RGB images.

To this end, we propose a 3D denoising model directly operating on an explicit voxel grid representation (Fig. 1, left) producing high-frequency noise estimates. To address the ambiguous and imperfect radiance field representation for each training sample, we propose to bias the noise prediction formulation from Denoising Diffusion Probabilistic Models (DDPMs) towards synthesizing higher image quality by an additional volumetric rendering loss on the estimates. This enables our method to learn radiance field priors less prone to fitting artifacts or noise accumulation during the sampling process. We show that our formulation leads to diverse and geometrically-accurate radiance field synthesis producing efficient, realistic, and view-consistent renderings. Our learned diffusion prior can be applied in an unconditional setting where 3D object synthesis is obtained in a multi-view consistent way, generating highly-accurate 3D shapes and allowing for free-view synthesis. We further introduce the new task of conditional masked completion – analog to shape completion – for radiance field completion at inference time. In this setting, we allow for realistic 3D completion of partially-masked objects without the need for task-specific model adaptation or training (see Fig. 1, right).

We summarize our contributions as follows:

To the best of our knowledge, we introduce the first diffusion model to operate directly on 3D radiance fields, enabling high-quality, truthful 3D geometry and image synthesis.

We introduce the novel application of 3D radiance field masked completion, which can be interpreted as a natural extension of image inpainting to the volumetric domain.

We show compelling results in unconditional and conditional settings, e.g., by improving over GAN-based approaches on image quality (from 16.54 to 15.95 in FID) and geometry synthesis (improving MMD from 5.62 to 4.42), on the challenging PhotoShape Chairs dataset .

Related work

Diffusion models. Since the seminal work by Sohl-Dickstein et al. on generative diffusion modeling, two classes of generative models have been proposed that perform inversion of a diffusion process: Denoising Score Matching (DSM) and Denoising Diffusion Probablistic Models (DDPMs) . Both approaches have shown to be flavours of a single framework, coined Score SDE, in the work of Song et al. . The term “Diffusion models” is now being used as an all-encompassing name for this constellation of methods. Several main directions being studied include devising different sampling schemes and noising models , exploring alternative formulations and training algorithms , and improving efficiency . An overview of current research results is given by Karras et al. .

Diffusion models have been used to obtain state-of-the-art results in many domains such as text-to-image and guided synthesis , 3D shape generation , molecule prediction , and video generation . Interestingly, diffusion models have shown to outperform Generative Adversarial Networks (GANs) in high-resolution image generation tasks , achieving unprecedented results in conditional image generation . Furthermore, compared to GANs, which are often prone to divergence and mode collapse , diffusion models have been observed to be much easier to train, although training time is still relatively long.

3D generation. Initially developed for 2D image synthesis, adversarial approaches have also found success in 3D, e.g., for generating meshes , 3D textures , voxelized representations , or Neural Radiance Fields (NeRFs) . In particular, methods in this last category have received much attention in recent years, as they can be trained purely from collections of 2D images, without any form of 3D supervision, and enable for the first time photo-realistic novel view synthesis of the generated 3D objects. Pi-GAN and GRAF propose similar approaches where a standard NeRF model is cast in a GAN setting by adding a form of stochastic conditioning, trained with an adversarial loss. These approaches are partly limited by the high training-time memory cost of NeRF-style volumetric rendering, forcing them to use low-resolution image patches. CIPS-3D and GIRAFFE solve this issue by letting the volume rendering component output a low-resolution 2D feature map, which is then upsampled by an efficient convolutional network to produce the final image. This approach drastically improves the quality and resolution of the rendered images, but also introduces 3D inconsistencies, as the convolutional stage can process different views of the same object in arbitrarily different ways. StyleNeRF partially addresses this problem by carefully designing the convolutional stage to minimize inconsistencies, while EG3D further improves on training efficiency by replacing the MLP-based NeRF with a light-weight tri-plane volumetric model generated by a convolutional network. In contrast to our method, these GAN-based approaches do not naturally support conditional synthesis or completion.

Compared to GANs, diffusion models are relatively under-explored as a tool for 3D synthesis, but a few works have emerged in the past two years. Some diffusion based-generators have been proposed for 3D point clouds , showing promising results for conditional synthesis, completion, and other related tasks . DreamFusion , 3DDesigner , and GAUDI , like our work, employ diffusion models in conjunction with radiance fields, with applications in both conditional (on text and images) and unconditional 3D generation. DreamFusion presents an algorithm to generate NeRFs, augmented with an illumination component, by optimizing a loss defined by a pre-trained 2D text-conditional diffusion model. GAUDI builds a 3D scene generator by first training a conditional NeRF to reconstruct a set of indoor videos given scene-specific latents, and then fitting a diffusion model to capture the learned latent space. Concurrent works like apply the denoising-diffusion approach to factorized radiance representations. In contrast, our diffusion model operates directly in the space of radiance fields which directly enables 3D-conditional tasks like shape completion (Sec. 4.2) by leveraging the learned volumetric prior.

Method

Our method consists of a generative model for 3D objects that builds on recent state-of-the-art diffusion probabilistic models . It is trained to revert a process that gradually corrupts 3D objects by injecting noise at different scales. In our case, 3D objects are represented as radiance fields , so the learned denoising process allows our method to generate object radiance fields from noise.

Since we target the generation of 3D objects as radiance fields, we begin with a brief overview of this representation before delving into the details of our method.

Following a ray casting logic, the radiance field can be rendered along a given ray rr yielding an RGB color crc_{r} with the following equation

where a ray rr is a linear curve parametrized by ss with unit velocity, rs∈Xr_{s}\in\mathcal{X} denotes the point along the ray at ss, and τs(r)\tau_{s}(r) is the transmittance probability at ss, which is given by

Using the ray rendering equation above, we can render the radiance field from any given camera, yielding an image of the novel view. It is sufficient to turn the camera into a proper set of rays expressed in world coordinates.

There exist various ways of implementing a radiance field, ranging from neural networks to explicit voxel grids . In this work, we opt for the latter, since it enables good rendering quality along with faster training and inference. The explicit grid can be queried at continuous positions via bilinear interpolation of the voxel vertices.

Under the explicit representation, radiance fields become 4D tensors, where the first three dimensions index a grid spanning X\mathcal{X}, whereas the last dimension indexes the density and color channels.

2 Generating Radiance Fields

Following recent advancements in the context of denosing-based generative methods , we formulate our generative model of radiance fields as a denoising diffusion probabilistic model.

Generation process. The generation (a.k.a. denoising) process is governed by a discrete-time Markov chain defined on the state space F\mathcal{F} of all possible pre-activated radiance fields expressed as flattened 4D tensors of fixed size.We consider pre-activated radiance fields, where both density and RGB color channels span a linear space and we assume a proper activation function will be applied at the time of rendering. This is required to have additive noise, while preserving a valid radiance field representation. The chain has a finite number of time steps {0,…,T}\{0,\ldots,T\}. The denoising process starts by sampling a state fTf_{T} from a standard, multivariate normal distribution p(fT)≔N(fT∣0,I)p(f_{T})\coloneqq\mathcal{N}(f_{T}|0,I), and generates states ft−1f_{t-1} from ftf_{t} by leveraging reversed transition probabilities pθ(ft−1∣ft)p_{\theta}(f_{t-1}|f_{t}) that are Gaussian with learned parameters θ\theta. Specifically we have

The generation process iterates up to the final state f0f_{0}, which represents the radiance field of a 3D object generated by our method. The mean of the Gaussian in (3) can be directly modeled with a neural network. However, as we will see later, it is more convenient to consider the following reparametrization of it

where ϵθ(ft,t)\epsilon_{\theta}(f_{t},t) is the noise that has been used to corrupt ft−1f_{t-1} predicted by e.g. a neural network, whereas ata_{t} and btb_{t} are pre-defined coefficients. Also the covariance Σt\Sigma_{t} takes a pre-defined value, although it could be data-dependent. Additional details about the value that the pre-defined variables take are given in Sec. 3.3.

Diffusion process. While the generation process works by iteratively denoising a completely random radiance field, the diffusion process works the other way around and iteratively corrupts samples from the distribution of 3D objects we want to model. We introduce it because it plays a fundamental role in the training scheme of the generation process. The diffusion process is governed by a discrete-time Markov chain with the same state space and time bounds mentioned in the generation process but with Gaussian transition probabilities that are pre-defined and given by

where αt≔1−βt\alpha_{t}\coloneqq 1-\beta_{t} and 0≤βt≤10\leq\beta_{t}\leq 1 are predefined coefficients implementing a schedule for the injected noise variance. The process starts by picking f0f_{0} from the distribution q(f0)q(f_{0}) of 3D object radiance fields we want to model, iteratively samples ftf_{t} given ft−1f_{t-1} yielding a scaled and noise-corrupted version of the latter, and stops with fTf_{T} being typically close to completely random depending on the implemented noise-variance schedule. By exploiting properties of the Gaussian distribution, we can conveniently express the distribution of ftf_{t} conditioned on f0f_{0} directly as a Gaussian distribution, yielding

where αˉt≔∏i=1tαi\bar{\alpha}_{t}\coloneqq\prod_{i=1}^{t}\alpha_{i}. This relation will be useful to quickly generate diffused data points at arbitrary time steps.

3 Training Objective

Our training objective comprises two complementary losses: i) A loss LRFL_{\mathtt{RF}} that penalizes the generation of radiance fields that do not fit the data distribution, and ii) an RGB loss LRGBL_{\mathtt{RGB}} geared towards improving the quality of renderings from generated radiance fields.

Radiance field generation loss. Following , we derive the training objective for our model starting from a variational upper-bound on the Negative Log-Likelihood (NLL). This upper-bound requires specifying a surrogate distribution that we refer to as qq because it indeed corresponds to the distribution qq governing the diffusion process, establishing the anticipated fundamental link with the generation process. We provide here some key steps of the derivation of the bound and refer to for more details about the intermediate ones. By Jensen inequality, the NLL of a data point f0∈Ff_{0}\in\mathcal{F} can be upper-bounded by leveraging qq as follows:

where ft1:t2f_{t_{1}:t_{2}} stands for (ft1,…,ft2)(f_{t_{1}},\ldots,f_{t_{2}}). The loss LRF(f0∣θ)L_{\mathtt{RF}}(f_{0}|\theta) bounding the NLL can be further decomposed into the following sum, up to a constant independent from θ\theta

Here, LRFt(f0∣θ)L_{\mathtt{RF}}^{t}(f_{0}|\theta) takes a simple and intuitive form if we set at≔1αta_{t}\coloneqq\frac{1}{\sqrt{\alpha_{t}}} and bt≔βt1−αˉtb_{t}\coloneqq\frac{\beta_{t}}{\sqrt{1-\bar{\alpha}_{t}}} in (4), and pick Σt≔βt22αt(1−αˉt)I\Sigma_{t}\coloneqq\frac{\beta_{t}^{2}}{2\alpha_{t}(1-\bar{\alpha}_{t})}I. Indeed, it yields

where ϕ(ϵ)≔N(ϵ∣0,I)\phi(\epsilon)\coloneqq\mathcal{N}(\epsilon|0,I) is the probability distribution of ϵ\epsilon, namely a normal multivariate, and the last equality follows from (6).

Radiance field rendering loss. We complement the previous loss with an additional RGB loss LRGB(f0∣θ)L_{\mathtt{RGB}}(f_{0}|\theta), aimed at improving the quality of renderings from generated radiance fields. Indeed, the Euclidean metric on the representation that is implicitly used in the previous loss to assess the quality of generated radiance fields does not necessarily ensure the absence of artifacts once we try to render the radiance field. We define LRGB(f0∣θ)L_{\mathtt{RGB}}(f_{0}|\theta) as a sum of time-specific terms LRGBt(f0∣θ)L^{t}_{\mathtt{RGB}}(f_{0}|\theta) similar to (8), yielding

where the expectation is taken with respect to a prior distribution ψ\psi for the viewpoint vv and ϵ∼ϕ(ϵ)\epsilon\sim\phi(\epsilon). Since the approximation is reasonable only with steps tt close to zero, we introduce a weight wtw_{t} that decays as step values increase (e.g. we use ωt≔αˉt2\omega_{t}\coloneqq\bar{\alpha}_{t}^{2}). We provide evidence in the experimental section that despite being an approximation, the proposed loss contributes to significantly improving the results.

Final loss. To summarize, the final training loss per data point f0f_{0} is given by the weighted combination of the radiance field generation and rendering losses introduced before, with a small variation that enables stochastic sampling of the step tt from a uniform distribution κ(t)\kappa(t):

Implementation details We implement ϵθ\epsilon_{\theta} as a 3D-UNet, which is based on the 2D-UNet architecture introduced in by replacing 2D convolutions and attention layers with corresponding 3D operators. For training, we uniformly sample timesteps t=1,…,T=1000t=1,\ldots,T=1000 for all experiments with variances of the diffusion process linearly increasing from β1=0.0015\beta_{1}=0.0015 to βT=0.05\beta_{T}=0.05, and choose to weight the rendering loss LRGBtL_{\mathtt{RGB}}^{t} with wt=αˉt2w_{t}=\bar{\alpha}_{t}^{2}. We refer to the supplementary material for additional details.

Experiments

In this section, we evaluate the performance of our method on both unconditional and conditional radiance field generation.

Datasets. We run experiments on the PhotoShape Chairs and on the Amazon Berkeley Objects (ABO) Tables dataset . For PhotoShape Chairs, we render the provided 15,576 chairs using Blender Cycles from 200 views on an Archimedean spiral. For ABO Tables, we use the provided 91 renderings with 2-3 different environment map settings per object, resulting in 1676 tables. Since both datasets do not provide radiance field representations of 3D objects, we generate them using a voxel-based approach at a resolution of 32332^{3} from the multi-view renderings.

Metrics. We evaluate image quality using the Fréchet Inception Distance (FID) and Kernel Inception Distance (KID) using . For comparison of the geometrical quality, we follow and compute the Coverage Score (COV) and Minimum Matching Distance (MMD) using Chamfer Distance (CD). While the Coverage Score measures the diversity of the generated samples, MMD assesses the quality of the generated samples. All metrics are evaluated at a resolution of 128×128128\times 128.

Comparison against state of the art. We quantitatively evaluate our approach on the task of unconditional 3D synthesis on PhotoShape in Tab. 1 and ABO Tables in Tab. 2. We compare against leading methods for 3D-aware image synthesis: π\pi-GAN and EG3D . Both our method and GAN-based approaches use the same set of rendered images for training. While we pre-process the rendered images to create a radiance field representation for each of the shape samples, the GAN-based methods are trained directly on the rendered images.

Compared to these approaches, our method yields overall better image quality while achieving significant improvements in geometrical quality and diversity. Fig. 3 and Fig. 4 show a qualitative comparison of our method with π\pi-GAN and EG3D. While EG3D achieves good image quality, it tends to produce inaccurate shapes and view-dependent image artifacts, like adding or removing armrests or changing the supporting structure. As the training objective (see Sec. 3.3) is to invert the diffusion process to denoise towards detailed volumetric representations, we observe that DiffRF reliably generates radiance fields with fine photo-metric and geometrical details.

Contribution of the rendering loss. Tab. 1 and Tab. 2 show ablation results where we evaluate the effects of 2D supervision on the radiance synthesis. Removing 2D supervision (“DiffRF w/o 2D”, line 3 in the tables), as expected, has a noticeable effect on FID, which increases by ≈2.3\approx 2.3 for PhotoShape and by ≈8.8\approx 8.8 for ABO Tables. This shows that biasing the noise prediction formulation from DDPM by a volumetric rendering loss leads to higher image quality. Qualitative comparisons can be found in the appendix. We notice a decrease in the Coverage Score that we explain by the fact that the rendering loss guides the denoising model towards learning radiance fields with less artifacts, thus reducing the amount of diverse but spurious shapes.

2 Conditional Generation

GANs need to be trained in order to be conditioned on a particular task, while diffusion models can be effectively conditioned at test-time . We leverage this property for the novel task of masked radiance field completion.

Masked Radiance Field Completion. Shape completion and image inpainting are well-studied tasks aiming to fill missing regions within a geometrical representation or in an image, respectively. We propose to combine both in the novel task of masked radiance field completion: Given a radiance field and a 3D mask, synthesize a completion of the masked region that harmonizes with the non-masked region. Inspired by RePaint , we perform conditional completion by gradually guiding the unconditional sampling process in the known region to the input finf^{in}

where mm is a binary mask applied to the input (light blue in Fig. 6) and ⊙\odot denotes element-wise multiplication on the voxel grid.

A quantitative analysis of the masking performance is shown in Tab. 3, where we compare our method against EG3D at different levels of masking. At each level, we randomly mask 200 samples and evaluate the completion performance in terms of FID (by rendering from 10 random views), as well as the photo-metric accuracy of unmasked regions (mPSNR). For EG3D, we use masked GAN inversion, where we perform global latent optimization (GLO) to minimize the photo-metric error on the re-projected, unmasked regions. We notice that, due to the single latent code representation, EG3D struggles to faithfully reconstruct the unmasked region of the input sample, and regularization is needed to not corrupt the overall representation. Fig. 5 further shows a qualitative comparison of the masking performance against EG3D.

Image-to-Volume Synthesis. DiffRF can be used to obtain 3D radiance fields from single view images by steering the sampling process using volumetric rendering. For this, we adopt the Classifer Guidance formulation from to guide the denoising process towards minimizing the rendering error against a posed RGB image with corresponding object mask (obtained, for example, with off-the-shelf segmentation networks). Fig. 7 shows qualitative results for this single-image reconstruction task on chairs from ScanNet . Furthermore, we show results of our model conditioned on CLIP-embeddings in the appendix.

3 Limitations

While our method shows promising results on the tasks of conditional and unconditional radiance field synthesis, several limitations remain. Compared to GAN-based approaches, our radiance fields-based approach requires a sufficient number of posed views in order to generate good training samples and suffers from lower sampling times. In this context, it would be interesting to explore leveraging faster sampling methods . Finally, our model is constrained in the maximum grid resolution by training-time memory limitations. These could be addressed by exploring adaptive or sparse grid structures as well as factorized neural fields representations .

Conclusions

We introduce DiffRF – a novel approach for 3D radiance field synthesis based on denoising diffusion probabilistic models. To the best of our knowledge, DiffRF is the first generative diffusion-based method to operate directly on volumetric radiance fields. Our model learns multi-view consistent priors from collections of posed images, enabling free-view image synthesis and accurate shape generation. We evaluated DiffRF on several object classes, comparing its performance against state-of-the-art GAN-based approaches, and demonstrating its effectiveness in both conditional and unconditional 3D generation tasks.

Acknowledgements

This work was done during Norman’s and Yawar’s internships at Meta Reality Labs Zurich as well as at TUM, funded by a Meta SRA. Matthias Nießner was also supported by the ERC Starting Grant Scan2CAD (804724).

References

Appendix

In this supplementary document, we discuss additional details about our method, the data used for training and evaluation, and show further qualitative results. We also refer to our for a comprehensive overview with further qualitative results.

Appendix A Additional qualitative results

We provide additional qualitative results on PhotoShape Chairs in Fig. 8 as well as on ABO Tables in Fig. 9. Furthermore, we provide a qualitative comparison of our method when removing the rendering loss in Fig. 10 and Fig. 11.

Appendix B Implementation detail

We base our architecture on 2D U-Net structure from : For this, we replace the 2D convolutions in the ResNet and attention blocks with corresponding 3D convolutions, preserving kernel sizes and strides. Furthermore, we use 3D average pooling layers instead of 2D in the downsampling steps. Our U-Net consists of 44 scaling blocks with two ResNet blocks per scale, where we linearly increase the initial feature channel dimension of 6464 to 256256. We use skip attention blocks at the scaling factors 22, 44, and 88 with 3232 channels per head.

Training details

We train all models with a batch size of 8 and use the Adam optimizer with an initial learning of 10−410^{-4}. We apply a linear beta scheduling from 0.00150.0015 to 0.050.05 at 10001000 timesteps. From 44 random training views at a resolution of 128×128128\times 128, we sample 81928192 random pixels for the rendering supervision (with 9292 z-steps for volumetric rendering) and weight the rendering loss with ωt=αˉt2\omega_{t}=\bar{\alpha}_{t}^{2}. We train for 3.03.0m iterations with a decaying LR scheduling for 10−410^{-4} to 10−610^{-6} at a voxel grid resolution of 3232 on 22 GPUs on every data set.

Sampling time

We perform DDPM sampling for 10001000 iterations leading to a run time of 48.648.6s per sample on an NVIDIA RTX 2080 TI. Once synthesized, our explicit representation enables rendering at 128×128128\times 128 resolution with over 380 FPS.

Appendix C Data

For PhotoShape Chairs, we render the provided 15,576 chairs using Blender Cycles from 200 views on an Archimedean spiral at a fixed radius of 2.52.5 units with pitch starting from −20∘-20^{\circ} to 60∘60^{\circ}. For ABO Tables, we use the provided 91 renderings with 2-3 different environment map settings per object, resulting in 1676 tables. For PhotoShape Chairs, we hold out 10%10\% of the samples for testing based on shape ids selected randomly, whereas for ABO Tables, we use the official data split. We fit explicit voxel grids at a resolution of 32332^{3} using volumetric rendering with spherical harmonics of degree 22 for an initial fit. We then fine-tune our representations for spherical harmonics of degree , which we found to lead to sharper geometry compared to directly optimizing density and color features. We furthermore bound the feature space to $$ which we found to stabilize the sampling process noticeably affecting the rendering quality.

Evaluation

For image quality evaluation, we calculate FID and IS by sampling 10k views by rendering 10001000 samples from 1010 random views at a resolution of 128×128128\times 128. We follow and evaluate the geometric quality by computing the Coverage Score (COV) and Minimum Matching Distance (MMD) using Chamfer Distance (CD)

on a reference set SrS_{r} (the test samples) and a generated set SgS_{g} twice as large as the reference set. We extract meshes using marching cubes and sample 20482048 points on the faces. To account for potentially different scaling of the samples produced by the 3D-aware GAN models, we normalize all point clouds by centering in the origin and an-isotropic scaling of the extent to $$.

For evaluation of the masked radiance field completion, we additionally compute a masked peak signal-to-noise ratio (mPSNR): Given a binary mask mm of the input radiance field finf^{in}, we compute for each corresponding input image the non-masked area by depth-based projection into the image plane using depth estimated from the input radiance field. We then compute the mPSNR by averaging the PSNR on the non-masked pixels for all evaluation views (we choose 10 views randomly).

Appendix D Conditional sampling

Since the generator of EG3D is trained via 2D discriminator guidance, we perform 3D masked completion via GAN inversion. For this, we start from a random initial latent code and repeat the following steps for 200 iterations on each masked sample: We render the current synthesized sample from 88 views and project the 3D input mask onto the synthesized views using the predicted depths. On the remaining non-masked regions, we compute the photometric error with the input images. We use the Adam optimizer with a learning rate of 10−210^{-2} with a small L2L_{2} regularization term on the code (weighted with 5×10−25\times 10^{-2}) in order to update the latent code.

Image-to-Volume Synthesis

Appendix E CLIP conditioning

Following related work , we additionally augment our model to condition on embeddings derived from text or single-image encodings obtained from CLIP ViT-B/32 using cross-attention layers. For training, we use random single training views encoded by the frozen CLIP model to condition the denoiser. Here, we adapt the cross-attention mechanism from for the 3D U-Net and do not train the image encoder in order to preserve the image-text-correspondence of CLIP. We show examples on single-image PhotoShape samples in Fig. 14 as well as on real-world image from the Pix3D dataset in Fig. 13.

As these codes have strong correspondences to text samples by design, we can guide the sampling process by text prompts, examples are shown in Fig. 12 without the need for training on text-radiance field pairs.