GET3D: A Generative Model of High Quality 3D Textured Shapes Learned from Images
Jun Gao, Tianchang Shen, Zian Wang, Wenzheng Chen, Kangxue Yin, Daiqing Li, Or Litany, Zan Gojcic, Sanja Fidler
Introduction
Diverse, high-quality 3D content is becoming increasingly important for several industries, including gaming, robotics, architecture, and social platforms. However, manual creation of 3D assets is very time-consuming and requires specific technical knowledge as well as artistic modeling skills. One of the main challenges is thus scale – while one can find 3D models on 3D marketplaces such as Turbosquid or Sketchfab , creating many 3D models to, say, populate a game or a movie with a crowd of characters that all look different still takes a significant amount of artist time.
To facilitate the content creation process and make it accessible to a variety of (novice) users, generative 3D networks that can produce high-quality and diverse 3D assets have recently become an active area of research . However, to be practically useful for current real-world applications, 3D generative models should ideally fulfill the following requirements: (a) They should have the capacity to generate shapes with detailed geometry and arbitrary topology, (b) The output should be a textured mesh, which is a primary representation used by standard graphics software packages such as Blender and Maya , and (c) We should be able to leverage 2D images for supervision, as they are more widely available than explicit 3D shapes.
Prior work on 3D generative modeling has focused on subsets of the above requirements, but no method to date fulfills all of them (Tab. 1). For example, methods that generate 3D point clouds typically do not produce textures and have to be converted to a mesh in post-processing. Methods generating voxels often lack geometric details and do not produce texture . Generative models based on neural fields focus on extracting geometry but disregard texture. Most of these also require explicit 3D supervision. Finally, methods that directly output textured 3D meshes typically require pre-defined shape templates and cannot generate shapes with complex topology and variable genus.
Recently, rapid progress in neural volume rendering and 2D Generative Adversarial Networks (GANs) has led to the rise of 3D-aware image synthesis . However, this line of work aims to synthesize multi-view consistent images using neural rendering in the synthesis process and does not guarantee that meaningful 3D shapes can be generated. While a mesh can potentially be obtained from the underlying neural field representation using the marching cube algorithm , extracting the corresponding texture is non-trivial.
In this work, we introduce a novel approach that aims to tackle all the requirements of a practically useful 3D generative model. Specifically, we propose GET3D, a Generative model for 3D shapes that directly outputs Explicit Textured 3D meshes with high geometric and texture detail and arbitrary mesh topology. In the heart of our approach is a generative process that utilizes a differentiable explicit surface extraction method and a differentiable rendering technique . The former enables us to directly optimize and output textured 3D meshes with arbitrary topology, while the latter allows us to train our model with 2D images, thus leveraging powerful and mature discriminators developed for 2D image synthesis. Since our model directly generates meshes and uses a highly efficient (differentiable) graphics renderer, we can easily scale up our model to train with image resolution as high as , allowing us to learn high-quality geometric and texture details.
We demonstrate state-of-the-art performance for unconditional 3D shape generation on multiple categories with complex geometry from ShapeNet , Turbosquid and Renderpeople , such as chairs, motorbikes, cars, human characters, and buildings. With explicit mesh as output representation, GET3D is also very flexible and can easily be adapted to other tasks, including: (a) learning to generate decomposed material and view-dependent lighting effects using advanced differentiable rendering , without supervision, (b) text-guided 3D shape generation using CLIP embedding.
Related Work
We review recent advances in 3D generative models for geometry and appearance, as well as 3D-aware generative image synthesis.
In recent years, 2D generative models have achieved photorealistic quality in high-resolution image synthesis . This progress has also inspired research in 3D content generation. Early approaches aimed to directly extend the 2D CNN generators to 3D voxel grids , but the high memory footprint and computational complexity of 3D convolutions hinder the generation process at high resolution. As an alternative, other works have explored point cloud , implicit , or octree representations. However, these works focus mainly on generating geometry and disregard appearance. Their output representations also need to be post-processed to make them compatible with standard graphics engines.
More similar to our work, Textured3DGAN and DIBR generate textured 3D meshes, but they formulate the generation as a deformation of a template mesh, which prevents them from generating complex topology or shapes with varying genus, which our method can do. PolyGen and SurfGen can produce meshes with arbitrary topology, but do not synthesize textures.
Inspired by the success of neural volume rendering and implicit representations , recent work started tackling the problem of 3D-aware image synthesis . However, neural volume rendering networks are typically slow to query, leading to long training times , and generate images of limited resolution. GIRAFFE and StyleNerf improve the training and rendering efficiency by performing neural rendering at a lower resolution and then upsampling the results with a 2D CNN. However, the performance gain comes at the cost of a reduced multi-view consistency. By utilizing a dual discriminator, EG3D can partially mitigate this problem. Nevertheless, extracting a textured surface from methods that are based on neural rendering is a non-trivial endeavor. In contrast, GET3D directly outputs textured 3D meshes that can be readily used in standard graphics engines.
Method
We now present our GET3D framework for synthesizing textured 3D shapes. Our generation process is split into two parts: a geometry branch, which differentiably outputs a surface mesh of arbitrary topology, and a texture branch that produces a texture field that can be queried at the surface points to produce colors. The latter can be extended to other surface properties such as for example materials (Sec. 4.3.1). During training, an efficient differentiable rasterizer is utilized to render the resulting textured mesh into 2D high-resolution images. The entire process is differentiable, allowing for adversarial training from images (with masks indicating an object of interest) by propagating the gradients from the 2D discriminator to both generator branches. Our model is illustrated in Fig. 2. In the following, we first introduce our 3D generator in Sec 3.1, before proceeding to the differentiable rendering and loss functions in Sec 3.2.
We aim to learn a 3D generator to map a sample from a Gaussian distribution to a mesh with texture .
We design our geometry generator to incorporate DMTet , a recently proposed differentiable surface representation. DMTet represents geometry as a signed distance field (SDF) defined on a deformable tetrahedral grid , from which the surface can be differentiably recovered through marching tetrahedra . Deforming the grid by moving its vertices results in a better utilization of its resolution. By adopting DMTet for surface extraction, we can produce explicit meshes with arbitrary topology and genus. We next provide a brief summary of DMTet and refer the reader to the original paper for further details.
After obtaining and for all the vertices, we use the differentiable marching tetrahedra algorithm to extract the explicit mesh. Marching tetrahedra determines the surface topology within each tetrahedron based on the signs of . In particular, a mesh face is extracted when , where denotes the indices of vertices in the edge of tetrahedron, and the vertices of that face are determined by a linear interpolation as . Note that the above equation is only evaluated when , thus it is differentiable, and the gradient from can be back-propagated into the SDF values and deformations . With this representation, the shapes with arbitrary topology can easily be generated by predicting different signs of .
1.2 Texture Generator
We represent our texture field using a tri-plane representation, which is efficient and expressive in reconstructing 3D objects and generating 3D-aware images . Specifically, we follow and use a conditional 2D convolutional neural network to map the latent code to three axis-aligned orthogonal feature planes of size , where denotes the spatial resolution and the number of channels.
2 Differentiable Rendering and Training
In order to supervise our model during training, we draw inspiration from Nvdiffrec that performs multi-view 3D object reconstruction by utilizing a differentiable renderer. Specifically, we render the extracted 3D mesh and the texture field into 2D images using a differentiable renderer , and supervise our network with a 2D discriminator, which tries to distinguish the image from a real object or rendered from the generated object.
We assume that the camera distribution that was used to acquire the images in the dataset is known. To render the generated shapes, we randomly sample a camera from , and utilize a highly-optimized differentiable rasterizer Nvdiffrast to render the 3D mesh into a 2D silhouette as well as an image where each pixel contains the coordinates of the corresponding 3D point on the mesh surface. These coordinates are further used to query the texture field to obtain the RGB values. Since we operate directly on the extracted mesh, we can render high-resolution images with high efficiency, allowing our model to be trained with image resolution as high as 10241024.
We train our model using an adversarial objective. We adopt the discriminator architecture from StyleGAN , and use the same non-saturating GAN objective with R1 regularization . We empirically find that using two separate discriminators, one for RGB images and another one for silhouettes, yields better results than a single discriminator operating on both. Let denote the discriminator, where can either be an RGB image or a silhouette. The adversarial objective is then be defined as follows:
where is defined as , is the distribution of real images, denotes rendering, and is a hyperparameter. Since is differentiable, the gradients can be backpropagated from 2D images to our 3D generators.
To remove internal floating faces that are not visible in any of the views, we further regularize the geometry generator with a cross-entropy loss defined between the SDF values of the neighboring vertices :
The overall loss function is then defined as:
where is a hyperparameter that controls the level of regularization.
Experiments
We conduct extensive experiments to evaluate our model. We first compare the quality of the 3D textured meshes generated by GET3D to the existing methods using the ShapeNet and Turbosquid datasets. Next, we ablate our design choices in Sec. 4.2. Finally, we demonstrate the flexibility of GET3D by adapting it to downstream applications in Sec. 4.3. Additional experimental results and implementation details are provided in Appendix.
For evaluation on ShapeNet , we use three categories with complex geometry – Car, Chair, and Motorbike, which contain 7497, 6778, and 337 shapes, respectively. We randomly split each category into training (70%), validation (10 %), and test (20 %), and further remove from the test set shapes that have duplicates in the training set. To render the training data, we randomly sample camera poses from the upper hemisphere of each shape. For the Car and Chair categories, we use 24 random views, while for Motorbike we use 100 views due to less number of shapes. As models in ShapeNet only have simple textures, we also evaluate GET3D on an Animal dataset (442 shapes) collected from TurboSquid , where textures are more detailed and we split it into training, validation and test as defined above. Finally, to demonstrate the versatility of GET3D, we also provide qualitative results on the House dataset collected from Turbosquid (563 shapes), and Human Body dataset from Renderpeople (500 shapes). We train a separate model on each category.
We compare GET3D to two groups of works: 1) 3D generative models that rely on 3D supervision: PointFlow and OccNet . Note that these methods only generate geometry without texture. 2) 3D-aware image generation methods: GRAF , PiGAN , and EG3D .
To evaluate the quality of our synthesis, we consider both the geometry and texture of the generated shapes. For geometry, we adopt metrics from and use both Chamfer Distance (CD) and Light Field Distance (LFD) to compute the Coverage score and Minimum Matching Distance. For OccNet , GRAF , PiGAN and EG3D , we use marching cubes to extract the underlying geometry. For PointFlow , we use Poisson surface reconstruction to convert a point cloud into a mesh when evaluating LFD. To evaluate texture quality, we adopt the FID metric commonly used to evaluate image synthesis. In particular, for each category, we render the test shapes into 2D images, and also render the generated 3D shapes from each model into 50k images using the same camera distribution. We then compute FID on the two image sets. As the baselines from 3D-aware image synthesis do not directly output textured meshes, we compute FID score in two ways: (i) we use their neural volume rendering to obtain 2D images, which we refer to as FID-Ori, and (ii) we extract the mesh from their neural field representation using marching cubes, render it, and then use the 3D location of each pixel to query the network to obtain the RGB values. We refer to this score, that is more aware of the actual 3D shape, as FID-3D. Further details on the evaluation metrics are available in the Appendix B.3.
We provide quantitative results in Table. 2 and qualitative examples in Fig. 3 and Fig. 4. Additional results are available in the supplementary video. Compared to OccNet that uses 3D supervision during training, GET3D achieves better performance in terms of both diversity (COV) and quality (MMD), and our generated shapes have more geometric details. PointFlow outperforms GET3D in terms of MMD on CD, while GET3D is better in MMD on LFG. We hypothesize that this is because PointFlow directly optimizes on point locations, which favours CD. GET3D also performs favourably when compared to 3D-aware image synthesis methods, we achieve significant improvements over PiGAN and GRAF in terms of all metrics on all datasets. Our generated shapes also contain more detailed geometry and texture. Compared with recent work EG3D . We achieve comparable performance on generating 2D images (FID-ori), while we significantly improve on 3D shape synthesis in terms of FID-3D, which demonstrates the effectiveness of our model on learning actual 3D geometry and texture.
Since we synthesize textured meshes, we can export our shapes into Blender We use xatlas to get texture coordinates for the extracted mesh, from where we can warp our 3D mesh into a 2D plane and obtain the corresponding 3D location on the mesh surface for any position on the 2D plane. We then discretize the 2D plane into an image, and for each pixel, we query the texture field using corresponding 3D location to obtain the RGB color to get the texture map.. We show rendering results in Fig. 1 and 5. GET3D is able to generate shapes with diverse and high quality geometry and topology, very thin structures (motorbikes), as well as complex textures on cars, animals, and houses.
Shape Interpolation GET3D also enables shape interpolation, which can be useful for editing purposes. We explore the latent space of GET3D in Fig. 6, where we interpolate the latent codes to generate each shape from left to right. GET3D is able to faithfully generate a smooth and meaningful transition from one shape to another. We further explore the local latent space by slightly perturbing the latent codes to a random direction. GET3D produces novel and diverse shapes when applying local editing in the latent space (Fig. 7).
2 Ablations
We ablate our model in two ways: 1) w/ and w/o volume subdivision, 2) training using different image resolutions. Further ablations are provided in the Appendix C.3.
As shown in Tbl. 2, volume subdivision significantly improves the performance on classes with thin structures (e.g., motorbikes), while not getting gains on other classes. We hypothesize that the initial tetrahedral resolution is already sufficient to capture the detailed geometry on Chairs and Cars, and hence the subdivision cannot provide further improvements.
We ablate the effect of the training image resolution in Tbl. 3. As expected, increased image resolution improves the performance in terms of FID and shape quality, as the network can see more details, which are often not available in the low-resolution images. This corroborates the importance of training with higher image resolution, which are often hard to make use of for implicit-based methods.
3 Applications
We provide qualitative results of generated surface materials in Fig. 8. Despite unsupervised, GET3D discovers interesting material decomposition, e.g., the windows are correctly predicted with a smaller roughness value to be more glossy than the car’s body, and the car’s body is discovered as more dielectric while the window is more metallic. Generated materials enable us to produce realistic relighting results, which can account for complex specular effects under different lighting conditions.
3.2 Text-Guided 3D Synthesis
Similar to image GANs, GET3D also supports text-guided 3D content synthesis by fine-tuning a pre-trained model under the guidance of CLIP . Note that our final synthesis result is a textured 3D mesh. To this end, we follow the dual-generator design from styleGAN-NADA , where a trainable copy and a frozen copy of the pre-trained generator are adopted. During optimization and both render images from 16 random camera views. Given a text query, we sample 500 pairs of noise vectors and . For each sample, we optimize the parameters of to minimize the directional CLIP loss (the source text labels are “car”, “animal” and “house” for the corresponding categories), and select the samples with minimal loss. To accelerate this process, we first run a small number of optimization steps for the 500 samples, then choose the top 50 samples with the lowest losses, and run the optimization for 300 steps. The results and comparison against a SOTA text-driven mesh stylization method, Text2Mesh , are provided in Fig. 9. Note that, requires a mesh of the shape as an input to the method. We provide our generated meshes from the frozen generator as input meshes to it. Since it needs mesh vertices to be dense to synthesize surface details with vertex displacements, we further subdivide the input meshes with mid-point subdivision to make sure each mesh has 50k-150k vertices on average.
Conclusion
We introduced GET3D, a novel 3D generative model that is able to synthesize high-quality 3D textured meshes with arbitrary topology. GET3D is trained using only 2D images as supervision. We experimentally demonstrated significant improvements on generating 3D shapes over previous state-of-the-art methods on multiple categories. We hope that this work brings us one step closer to democratizing 3D content creation using A.I..
While GET3D makes a significant step towards a practically useful 3D generative model of 3D textured shapes, it still has some limitations. In particular, we still rely on 2D silhouettes as well as the knowledge of camera distribution during training. As a consequence, GET3D was currently only evaluated on synthetic data. A promising extension could use the advances in instance segmentation and camera pose estimation to mitigate this issue and extend GET3D to real-world data. GET3D is also trained per-category; extending it to multiple categories in the future, could help us better represent the inter-category diversity.
We proposed a novel 3D generative model that generates 3D textured meshes, which can be readily imported into current graphics engines. Our model is able to generate shapes with arbitrary topology, high quality textures and rich geometric details, paving the path for democratizing A.I. tool for 3D content creation. As all machine learning models, GET3D is also prone to biases introduced in the training data. Therefore, an abundance of caution should be applied when dealing with sensitive applications, such as generating 3D human bodies, as GET3D is not tailored for these applications. We do not recommend using GET3D if privacy or erroneous recognition could lead to potential misuse or any other harmful applications. Instead, we do encourage practitioners to carefully inspect and de-bias the datasets before training our model to depict a fair and wide distribution of possible skin tones, races or gender identities.
Disclosure of Funding
This work was funded by NVIDIA. Jun Gao, Tianchang Shen, Zian Wang and Wenzheng Chen acknowledge additional revenue in the form of student scholarships from University of Toronto and the Vector Institute, which are not in direct support of this work.
References
Appendix
In this Appendix, we first provide detailed description of the GET3D network architecture (Sec. A.1- A.4) along with the training procedure and hyperparameters (Sec. A.6). We then describe the datasets (Sec. B.1), baselines (Sec. B.2), and evaluation metrics (Sec. B.3). Additional qualitative results, ablation studies, robustness analysis, and results on the real dataset are available in Sec. C. Details and additional results of the material generation for view-dependent lighting effects are provided in Sec. D. Sec E contains more information about the text-guided shape generation experiments as well as more additional qualitative results. The readers are also kindly referred to the accompanying video (demo.mp4) that includes 360-degree renderings of our results (more than 400 generated shapes for each category), detailed zoom-ins, interpolations, material generation, and shapes generated with text-guidance.
A Details of Our Model
In Sec. 3 we have provided a high level description of GET3D. Here, we provide the implementation details that were omitted due to the lack of space. Please consult the Figure B and Figure 2 in the main paper for more context. Source code is available at our project webpage
A.2 Geometry Generator
where and are the original and modulated weight, respectively. is the style corresponding to the th input channel, is the output channel dimension, and denote the spatial dimension of the 3D convolutional filter.
Note that for simplicity, we remove all the noise vector from StyleGAN and only have stochasticity in the input z. Furthermore, following practices from DEFTET and DMTET , we us two copies of the geometry generator. One generates the vertex offsets , while the other outputs the SDF values . The architecture of both is the same, except for the output dimension and activation function of the last layer.
In cases where modeling at a high-resolution is required (e.g. motorbike with thin structures in the wheels), we further use volume subdivision following DMTET . As illustrated in Fig. A, we first subdivide the tetrahedral grid and compute SDF values of the new vertices (midpoints) by averaging the SDF values on the edge. Then we identify tetrahedra that have vertices with different SDF signs. These are the tetrahedra that intersect with the underlying surface encode by SDF. To refine the surface at increased grid resolution after subdivision, we further predict the residual on SDF values and deformations to update and of the vertices in identified tetrahedra. Specifically, we use an additional 3D convolutional layer to upsample feature volume to of shape conditioned on . Then, following the steps described above, we use trilinear interpolation to obtain per-vertex feature, concatenate it with PE and decode the residuals and using conditional FC layers. The final SDF and vertex offset are computed as:
A.3 Texture Generator
The output of the last tTPF layer is then reshaped into three axis-aligned feature planes of size .
A.4 2D Discriminator
We use two discriminators to train GET3D: one for the RGB output and one for the 2D silhouettes. For both, we use exactly the same architecture as the discriminator in StyleGAN . Empirically, we have observed that conditioning the discriminator on the camera pose leads to canonicalization of the shape orientations. However, discarding this conditioning only slightly affects the performance, as shown in Section C.3. In fact, we primarily use this conditioning to enable the evaluation of geometry using evaluation metrics, which assume that the shapes are generated in the canonical frame.
A.5 Improved Generator
The motivation for sampling two noise vectors (, ) in the generator is to enable disentanglement of the geometry and texture, where geometry is to be treated as a first-class citizen. Indeed, the geometry should only be controlled by the geometry latent code, while the texture should be able to not only adapt to the changes in the texture latent code, but also to the changes in geometry, i.e. a change in the geometry latent should propagate to the texture. However, in the original design of the GET3D generator (c.f. Sec. 3 and Fig. 2) the information flow from the geometry to the texture generator is very limited—concatenation of the two latent codes (Fig. B). Such a weak connection makes it hard to learn the disentanglement of geometry and texture and the texture generator can even learn to ignore the texture latent code (Fig. D.).
This empirical observation motivated us to improve the design of the generator network, after the initial submission, by improving the information flow, which in turn better supports the disentanglement of the geometry and texture. To this end, our improved generator shares the same backbone network for both geometry and texture generation, as shown in Fig. C. In particular, we follow SemanticGAN and use StyleGAN2 backbone. Each ModBlock2D (modulated with the geometry latent code ), now has two tTPF branches, one for generating the geometry feature (tGEO), and the other for generating texture features (tTEX). The output of this backbone network are two feature triplanes, one for geometry and one for texture. To predict the SDF value and deformation for each vertex in the tetrahedral grid, we project the vertex onto each of the geometry triplanes, obtain its feature vector using Eq. 7, and finally use a ModFC to decode and . The prediction of the color in the texture branch remains unchanged.
Qualitative result of the geometry and texture disentanglement achieved with this improved generator is depicted in Fig. E and F. Shared backbone network allows us to achieve much better disentanglement of geometry and texture (Fig. D vs Fig. E), while also achieving better quantitative metrics on the task of unconditional generation (Tab. 2).
A.6 Training Procedure and Hyperparameters
We implement GET3D on top of the official PyTorch implementation of StyleGAN2 StyleGan3: https://github.com/NVlabs/stylegan3 (NVIDIA Source Code License). Our training configuration largely follows StyleGAN2 including: using a minibatch standard deviation in the discriminator, exponential moving average for the generator, non-saturating logistic loss, and R1 Regularization. We train GET3D along with the 2D discriminators from scratch, without progressive training or initialization from pretrained checkpoints. Most of our hyper-parameters are adopted form styleGAN2 . Specifically, we use Adam optimizer with learning rate 0.002 and . For R1 regularization, we set the regularization weight to 3200 for chair, 80 for car, 40 for animal, 80 for motorbike, 80 for renderpeople, and 200 for house. We follow StyleGAN2 and use lazy regularization, which applies R1 regularization to discriminators only every 16 training steps. Finally, we set the hyperparameter that controls the SDF regularization to 0.01 in all the experiments. We train our model using a batch size of 32 on 8 A100 GPUs for all the experiments. Training a single model takes about 2 days to converge.
B Experimental Details
We evaluate GET3D on ShapeNet , TurboSquid , and RenderPeople datasets. In the following, we provide their detailed description and the preprocessing steps that were used in our evaluation. Detailed statistic of the datasets is available in Table A.
ShapeNet The ShapeNet license is explained at https://shapenet.org/terms contains more than 51k shapes from 55 different categories and is the most commonly used dataset for benchmarking 3D generative models Herein, we used ShapeNet v1 Core subset obtained from https://shapenet.org/. Prior work typically uses the categories Airplane, Car, and Chair for evaluation. Herein, we replace the category Airplane with Motorcycle, which has more complex geometry and contains shapes with varying genus. Car, Chair, and Motorcycle contain 7497, 6778, and 337 shapes, respectively. We random split the shapes of each category into training (70%), validation (10%), and test (20%) and remove from the test set shapes that have duplicates in the training set.
TurboSquid https://www.turbosquid.com, we obtain consent via an agreement with TurboSquid, and following license at https://blog.turbosquid.com/turbosquid-3d-model-license/ is a large collection of various 3D shapes with high-quality geometry and texture, and is thus well suited to evaluate the capacity of GET3D to generate shapes with high-quality details. To this end, we use the category Animal that contains 442 textured shapes with high diversity ranging from cats, dogs, and lions, to bears and deer . We again randomly split the shapes into training (70%), validation (10%), and test (20%) set. Additionally, we provide qualitative results on the category House that contains 563 shapes. Since we perform only qualitative evaluation on House, we use all the shapes for training.
RenderPeople We follow the license of Renderpeople https://renderpeople.com/general-terms-and-conditions/ is a large dataset containing photorealistic 3D models of real-world humans. We use it to showcase the capacity of GET3D to generate high-quality and diverse characters that can be used to populate virtual environments, such as games or even movies. In particular, we use 500 models from the whole dataset for training and only perform qualitative analysis.
To generate the data, we first scale each shape such that the longest edge of its bounding-box equals , where for Car, Motorcycle, and Human, for House, and for Chair and Animal. For methods that use 2D supervision (Pi-GAN, GRAF, EG3D, and our model GET3D), we then render the RGB images and silhouettes from camera poses sampled from the upper hemisphere of each object. Specifically, we sample 24 camera poses for Car and Chair, and 100 poses for Motorcycle, Animal, House, and Human. The rotation and elevation angles of the camera poses are sampled uniformly from a specified range (see Table A). For all camera poses, we use a fixed radius of 1.2 and the fov angle of . We render the images in Blender using a fixed lighting, unless specified differently.
For the methods that rely on 3D supervision, we follow their preprocessing pipelines . Specifically, for Pointflow we randomly sample 15k points from the surface of each shape, while for OccNet we convert the shapes into watertight meshes by rendering depth frames from random camera poses and performing TSDF fusion.
B.2 Baselines
PointFlow is a 3D point cloud generative model based on continuous normalizing flows. It models the generative process by learning a distribution of distributions. Where the former, denotes the distribution of shapes, and the latter the distribution of points given a shape . PointFlow generates only the geometry, which is represented in the form of a point cloud. To generate the results of , we use the original source code provided by the authors PointFlow: https://github.com/stevenygd/PointFlow (MIT License) and train the models on our data. To compute the metrics based on LFD, we convert the output point clouds (10k points) to a mesh representation using Open3D implementation of Poisson surface reconstruction .
OccNet is an implicit method for 3D surface reconstruction, which can also be applied to unconditional generation of 3D shapes. OccNet is an autoencoder that learns a continuous mapping from 3D coordinates to occupancy values, from which an explicit mesh can be extracted using marching cubes . When applied to unconditional 3D shape generation, OccNet is trained as a variational autoencoder. To generate the results of , we use the original source code provided by the authors OccNet: https://github.com/autonomousvision/occupancy_networks (MIT License) and train the models on our data.
GRAF is a generative model that tackles the problem of 3D-aware image synthesis. GRAF’s underlying representation is a neural radiance field—conditioned on the shape and appearance latent codes—parameterized using a multi-layer perceptron with positional encoding. To synthesize novel views, GRAF utilizes a neural volume rendering approach similar to Nerf . In our evaluation, we use the source code provided by the authors GRAF: https://github.com/autonomousvision/graf (MIT License) and train GRAF models on our data.
Pi-GAN similar to GRAF, Pi-GAN also tackles the problem of 3D-aware image synthesis, but uses a Siren network—conditioned on a randomly sampled noise vector—to parameterize the neural radiance field. To generate the results of Pi-GAN , we use the original source code provided by the authors Pi-GAN: https://github.com/marcoamonteiro/pi-GAN (License not provided) and train the models on our data.
EG3D is a recent model for 3D-aware image synthesis. Similar to our method, EG3D builds upon the StyleGAN formulation and uses a tri-plane representation to parameterize the underlying neural radiance field. To improve the efficiency and to enable synthesis at higher resolution, EG3D utilizes neural rendering at a lower resolution and then upsamples the output using a 2D CNN. The source code of EG3D was provided to us by the authors. To generate the results, we train and evaluate EG3D on our data.
B.3 Evaluation Metrics
To evaluate the performance, we compare both the texture and geometry of the generated shapes to the reference ones .
To evaluate the geometry, we use all shapes of the test set as , and synthesize five times as many generated shapes, such that , where denotes the cardinality of a set. Following prior work , we use Chamfer Distance and Light Field Distance to measure the similarity of the shapes, which is in turn used to compute Coverage (COV) and Minimum Matching Distance (MMD) evaluation metrics.
While Chamfer distance has been widely used in the field of 3D generative models and reconstruction , LFD has received a lot attention in computer graphics . Inspired by human perception, LFD measures the similarity between the 3D shapes based on their appearance from different viewpoints. In particular, LFD renders the shapes and (represented as explicit meshes) from a set of selected viewpoints, encodes the rendered images using Zernike moments and Fourier descriptors, and computes the similarity over these encodings. Formal definition of LFD is available in . In our evaluation, we use the official implementation to compute LFD: https://github.com/Sunwinds/ShapeDescriptor/tree/master/LightField/3DRetrieval_v1.8/3DRetrieval_v1.8 (License not provided).
We combine these similarity measures with the evaluation metrics proposed in , which are commonly used to evaluate 3D generative models:
Coverage (COV) measures the fraction of shapes in the reference set that are matched to at least one of the shapes in the generated set. Formally, COV is defined as
where the distance metric D can be either or . Intuitively, COV measures the diversity of the generated shapes and is able to detect mode collapse. However, COV does not measure the quality of individual generated shapes. In fact, it is possible to achieve high COV even when the generated shapes are of very low quality.
Minimum Matching Distance (MMD) complements COV metric, by measuring the quality of the individual generated shapes. Formally, MMD is defined as
where D can again be either or . Intuitively, MMD measures the quality of the generated shapes by comparing their geometry to the closest reference shape.
B.3.2 Evaluating the Texture and Geometry
To evaluate the quality of the generated textures, we adopt the Fréchet Inception Distance (FID) metric, commonly used to evaluate the synthesis quality of 2D images. In particular, for each category, we render 50k views of the generated shapes (one view per shape) from the camera poses randomly sampled from the predefined camera distribution, and use all the images in the test set. We then encode these images using a pretrained Inception v3 model Inception network checkpoint path: http://download.tensorflow.org/models/image/imagenet/inception-2015-12-05.tgz, where we consider the output of the last pooling layer as our final encoding. The FID metric can then be computed as:
where Tr denotes the trace operation. and are the mean value and covariance matrix of the generated image encoding, while and are obtained from the encoding of the test images.
As briefly discussed in the main paper, we use two variants of FID, which differ in the way in which the 2D images are rendered. In particular, for FID-Ori, we directly use the neural volume rendering of the 3D-aware image synthesis methods to obtain the 2D images. This metric favours the baselines that were designed to directly generate valid 2D images through neural rendering. Additionally, we propose a new metric, FID-3D, which puts more emphasis on the overall quality of the generated 3D shape. Specifically, for the baselines which do not output a textured mesh, we extract the geometry from their underlying neural field using marching cubes . Then, we find the intersection point of each pixel ray with the generated mesh and use the 3D location of the intersected point to query the RGB value from the network. In this way, the rendered image is a more faithful representation of the underlying 3D shape and takes the quality of both geometry and texture into account. Note that FID-3D and FID-Ori are identical for methods that directly generate textured 3D meshes, as it is the case with GET3D.
C Additional Results on the Unconditioned Shape Generation
In this section we provide additional results on the task of unconditional 3D shape generation. First, we perform additional qualitative comparison of GET3D with the baselines in Section C.1. Second, we present further qualitative results of GET3D in Section C.2. Third, we provide additional ablation studies in Section C.3. We also analyse the robustness and effectiveness of GET3D. Specifically, in Sec. C.4 and C.5, we evaluate GET3D trained with noisy cameras and 2D silhouettes predicted by 2D segmentation networks. We further provide addition experiments on StyleGAN generated realistic dataset from GANverse3D in Sec. C.6. Finally, we provide additional comparison with EG3D on human character generation in Sec. C.7.
We provide additional visualization of the 3D shapes generated by GET3D and compare them to the baseline methods in Figure Q. GET3D is able to generate shapes with complex geometry, different topology, and varying genus. When compared to the baselines, the shapes generated by GET3D contain more details and are more diverse.
We provide additional results on the task of 2D image generation in Figure R. Even though GET3D is not designed for this task, it produces comparable results to the strong baseline EG3D , while significantly outperforming other baselines, such as PiGAN and GRAF . Note that GET3D directly outputs 3D textured meshes, which are compatible with standard graphics engines, while extracting such representation from the baselines is non-trivial.
C.2 Additional Qualitative Results of GET3D
We provide additional visualizations of the generated geometry and texture in Figures S-X. GET3D can generate high quality shapes with diverse textures across all the categories, from chairs, cars, and animals, to motorbikes, humans, and houses. Accompanying video (demo.mp4) contains further visualizations, including detailed turntable animations for 400+ shapes and interpolation results.
To demonstrate that GET3D is capable of generating novel shapes, we perform shape retrieval for our generated shapes. In particular, we retrieve the closest shape in the training set for each of shapes we showed in the Figure 1 by measuring the CD between the generated shape and all training shapes. Results are provided in Figure G. All generated shapes in Figure 1 significantly differ from their closest shape in the training set, exhibiting different geometry and texture, while still maintaining the quality and diversity.
We provide further qualitative results highlighting the benefits of volume subdivision in Figure Y. Specifically, we compare the shapes generated with and without volume subdivision on ShapeNet motorbike category. Volume subdivision enables GET3D to generate finer geometric details like handle and steel wire, which are otherwise hard to represent.
C.3 Additional Ablations Studies
We now provide additional ablation studies in an attempt to further justify our design choices. In particular, we first discuss the design choice of using two dedicated discriminators for RGB images and 2D silhouettes, before ablating the impact of adding the camera pose conditioning to the discriminator.
We empirically find that using a single discriminator on both RGB image and silhouettes introduces significant training instability, which leads to divergence when training GET3D. We provide a comparison of the training dynamics in Figure H and I, where we depict the loss curves for the generator and discriminator. We hypothesize that the instability might be caused by the fact that a single discriminator has access to both geometry (from 2D silhouettes) and texture (from RGB image) of the shape, when classifying whether the image is real or not. Since we randomly initialize our geometry generator, the discriminator can quickly overfit to one aspect—either geometry or texture—and thus produces bad gradients for the other branch. A two-stage approach in which two discriminators would be used in the first stage of the training, and a single discriminator in the later stage, when the model has already learned to produce meaningful shapes, is an interesting research direction, which we plan to explore in the future.
C.3.2 Ablation on Using Camera Condition for Discriminator
Since we are mainly operating on synthetic datasets in which the shapes are aligned to a canonical direction, we condition the discriminators on the camera pose of each image. In this way, GET3D learns to generate shapes in the canonical orientation, which simplifies the evaluation when using metrics that assume that the input shapes are canonicalized. We now ablate this design choice. Specifically, we train another model without the conditioning and evaluate its performance in terms of FID score. Quantitative results are given in Table. B. We observe that removing the camera pose conditioning, only slightly degrades the performance of GET3D (-1.38 FID). This confirms that our model can be successfully trained without such conditioning, and that the primary benefit of using it is the easier evaluation.
C.4 Robustness to Noisy Cameras
To demonstrate the robustness of GET3D to imperfect cameras poses, we add Gaussian noises to the camera poses during training. Specifically, for the rotation angle, we add a noise sampled from a Gaussian distribution with zero mean, and 10 degrees variance. For the elevation angle, we also add a noise sampled from a Gaussian distribution with zero mean, and 2 degrees variance. We use ShapeNet Car dataset in this experiment.
The quantitative results are provided in Table C and qualitative examples are depicted in Figure J. Adding camera noise harms the FID metric, whereas we observe only little degradation in visual quality. We hypothesize that the drop in the FID is a consequence of the camera pose distribution mismatch, which occurs as result of rendering the testing dataset, used to calculate the FID score, with a camera pose distribution without added noise. Nevertheless, based on the visual quality of the generated shapes, we conclude that GET3D is robust to a moderate level of noise in the camera poses.
C.5 Robustness to Imperfect 2D Silhouettes
To evaluate the robustness of GET3D when trained with imperfect 2D silhouettes, we replace ground truth 2D masks with the ones obtained from Detectron2 https://github.com/facebookresearch/detectron2 using pretrained PointRend checkpoint, mimicking how one could obtain the 2D segmentation masks in the real world. Since our training images are rendered with the black background, we use two approaches to obtain the 2D silhouettes: i) we directly feed the original training image into Detectron2 to obtain the predicted segmentation mask (we refer to this as Mask-Black), and ii) we add a background image, randomly sampled from PASCAL-VOC 2012 dataset (we refer to this as Mask-Random). In this setting, the pretrained Detectron2 model achieved 97.4 and 95.8 IoU for the Mask-Black and Mask-Random versions, respectively. We again use the Shapenet Car dataset in this experiment.
Quantitative results are summarized in Table C, with qualitative examples provided in Figures K and L. Although we observe drop in the FID scores, qualitatively the results are still similar to the original results in the main paper. Our model can generate high quality shapes even when trained with the imperfect masks. Note that, in this scenario, the training data for GET3D is different from the testing data that is used to compute the FID score, which could be one of the reasons for worse performance.
C.6 Experiments on "Real" Image
Since many real-world datasets lack camera poses, we follow GANverse3D and utilize pretrained 2D StyleGAN to generate a realistic car dataset. We train GET3D on this dataset to demonstrate the potential applications to real-world data.
Following GANverse3D , we manipulate the latent codes of 2D StyleGAN and generate multi-view car images. To obtain the 2D segmentation of each image, we use DatasetGAN to predict the 2D silhouette. We then use SfM to obtain the camera initialization for each generated image. We visualize some examples of this dataset in Fig N and refer the reader to the original GANverse3D paper for more details. Note that, in this dataset both cameras and 2D silhouettes are imperfect.
We provide qualitative examples in Fig. M. Even when faced with the imperfect inputs during training, GET3D is still capable of generating reasonable 3D textured meshes, with variation in geometry and texture.
C.7 Comparison with EG3D on Human Body
Following the suggestion of the reviewer, we also train EG3D model on the Human Body dataset rendered from Renderpeople and compare it to the results of GET3D.
Quantitative results are available in Table D and qualitative comparisons in Figure O. GET3D achieves comparable performance to EG3D in terms of generated 2D images (FID-ori), while significantly outperforming it on 3D shape synthesis (FID-3D). This once more demonstrates the effectiveness of our model in learning actual 3D geometry and texture.
D Material Generation for View-dependent Lighting Effects
In modern computer graphics engines such as Blender and Unreal Engine , surface properties are represented by material parameters crafted by graphics artists. To make the generated assets graphics-compatible, one direct extension of our method is to also generate surface material properties. In this section, we describe how GET3D is able to incorporate physics-based rendering models, predicting SVBRDF to represent view-dependent lighting effects such as specular surface reflections.
Similar to the original training pipeline, we randomly sample light from a set of real-world outdoor HDR panoramas (detailed in the following “Datasets” paragraph) and render the generated 3D assets into 2D images using cameras sampled from the camera distribution of training set. We train the model using the same method as in the main paper by adopting the discriminators to encourage the perceptual realism of the rendered images under arbitrary real-world lighting, along with a second discriminator on the 2D silhouettes to learn the geometry. Note that no supervision from material ground truth is used during training, and the material decomposition emerges in a fully unsupervised manner. When equipped with a physics-based rendering models, GET3D successfully predicts reasonable surface material parameters, generating delicate models which can be directly utilized in stand rendering engines like Blender and Unreal .
We collect a set of 724 outdoor HDR panoramas from HDRIHaven polyhaven.com/hdris (License: CC0), DoschDesign doschdesign.com (License: doschdesign.com/information.php?p=2) and HDRMaps hdrmaps.com (License: Royalty-Free), which cover a diverse range of real-world lighting distribution for outdoor scenes. We also apply random flipping and random rotation along azimuth as data augmentation. During training, we convert all the environment maps to SG lighting representations, where we adopt 32 SG lobes, optimizing their directions, sharpness and amplitudes such that the approximated lighting is close to the environment map. We optimize 7000 iterations with MSE loss and Adam optimizer. The converged SG lighting can preserve the most contents in the environment map.
As ShapeNet dataset does not contain consistent material definition, we additionally collect 1494 cars from Turbosquid with materials consistently defined with Disney BRDF. To render the dataset using Blender , we follow the camera configuration of ShapeNet Car dataset, and randomly select from the collected set of HDR panoramas as lighting. In the dataset, the groundtruth roughness for car windows is in the range of and the metallic is set to ; for car paint, the groundtruth roughness is in the range of and the metallic is set to . We disable complex materials such as the transparency and clear coat effects, such that the rendered results can be interpreted by the basic Disney BRDF properties including base color, metallic and roughness.
Since we aim to generate 3D assets that can be used in graphics workflow to produce realistic 2D renderings, we quantitatively evaluate the realism of the 2D rendered images under real-world lighting using FID score.
To the best of our knowledge, up to date no generative model can directly generate complex geometry (meshes) with material information. We therefore only compare different version of our model. In particular, we compare the results to the texture prediction version of GET3D, where we do not use material and directly predict RGB color for the surface points. We then ablate the effects of using real-world HDR panoramas for lighting, which are typically hard to obtain. To this end, we manually use two spherical Gaussians for ambient lighting and a random directions to simulate the lighting when rendering the generated shapes during training, and try to learn the materials under this simulated lighting.
The quantitative FID scores are provided in Table E. With material generation, the FID score improves by more than 2 points when compared to the texture prediction baseline (18.53 vs 20.78). This indicates that the material generation version of GET3D has better capacity and improved realism compared to the texture only baseline. When using the simulated lighting, instead of real-world HDR panorama, the FID score gets slightly worse but still produces reasonable performance. We further provide additional qualitative results in Fig. P visualizing rendering results of generated assets under different real-world lighting conditions. We import our generated assets in Blender and show animated visualization in the accompanied video (demo.mp4).
E Text-Guided 3D Synthesis
As briefly described in Sec. 4.3.2, our text-guided 3D synthesis method follows the dual-Generator design from StyleGAN-NADA , and uses the directional CLIP loss . In particular, at each optimization iteration, we randomly sample camera views and render paired images using two generators: the frozen one () and the trainable one (). The directional CLIP loss can then be computed as:
where is the translation of the CLIP embeddings () from the rendering with to the rendering with , under camera and is the CLIP embedding translation from the class text label to the provided query text. In our implementation, we used two pre-trained CLIP models with different Vision Transformers (‘ViT-32/B’ and ‘ViT-B/16’) for different level of details, and follow the text augmentation as in the StyleGAN-NADA codebase https://github.com/rinongal/StyleGAN-nada (MIT License).