EVA3D: Compositional 3D Human Generation from 2D Image Collections

Fangzhou Hong, Zhaoxi Chen, Yushi Lan, Liang Pan, Ziwei Liu

Introduction

Inverse graphics studies inverse-engineering of projection physics, which aims to recover the 3D world from 2D observations. It is not only a long-standing scientific quest, but also enables numerous applications in VR/AR and VFX. Recently, 3D-aware generative models (Chan et al., 2021; Or-El et al., 2022; Chan et al., 2022; Deng et al., 2022) demonstrate great potential in inverse graphics by learning to generate 3D rigid objects (e.g. human/animal faces, CAD models) from 2D image collections. However, human bodies, as articulated objects, have complex articulations and diverse appearances. Therefore, it is challenging to learn 3D human generative models that can synthesis animatable 3D humans with high-fidelity textures and vivid geometric details.

To generate high-quality 3D humans, we argue that two main factors should be properly addressed: 1) 3D human representation; 2) generative network training strategies. Due to the articulated nature of human bodies, a desirable human representation should be able to explicitly control the pose/shape of 3D humans. An articulated 3D human representation need to be designed, rather than the static volume modeling utilized in existing 3D-aware GANs. With an articulated representation, a 3D human is modeled in its canonical pose (canonical space), and can be rendered in different poses and shapes (observation space). Moreover, the efficiency of the representation matters in high-quality 3D human generation. Previous methods (Noguchi et al., 2022; Bergman et al., 2022) fail to achieve high resolution generation due to their inefficient human representations.

In addition, training strategies could also highly influence 3D human generative models. The issue mainly comes from the data characteristics. Compared with datasets used by Noguchi et al. (2022) (e.g. AIST (Tsuchida et al., 2019)), fashion datasets (e.g. DeepFashion (Liu et al., 2016)) are more aligned with real-world human image distributions, making a favorable dataset choice. AIST only has 40 dancers, which are mostly dressed in black. In contrast, DeepFashion contains much more different persons wearing various clothes, which benefits the diversity and quality of 3D human generation. However, fashion datasets mostly have very limited human poses (most are similar standing poses), and highly imbalanced viewing angles (most are front views). This imbalanced 2D data distribution could hinder unsupervised learning of 3D GANs, leading to difficulties in novel view/ pose synthesis. Therefore, a proper training strategy is in need to alleviate the issue.

In this work, we propose EVA3D, an unconditional high-quality 3D human generative model from sparse 2D human image collections only. To facilitate that, we propose a compositional human NeRF representation to improve the model efficiency. We divide the human body into 16 parts and assign each part an individual network, which models the corresponding local volume. Our human representation mainly provides three advantages. 1) It inherently describes the human body prior, which supports explicit control over human body shapes and poses. 2) It supports adaptively allocating computation resources. More complex body parts (e.g. heads) can be allocated with more parameters. 3) The compositional representation enables efficient rendering and achieves high-resolution generation. Rather than using one big volume (Bergman et al., 2022), our compositional representation tightly models each body part and prevents wasting parameters on empty volumes. Moreover, thanks to the part-based modeling, we can efficiently sample rays inside local volumes and avoid sampling empty spaces. With the compact representation together with the efficient rendering algorithm, we achieve high-resolution (512×256512\times 256) rendering and GAN training without using super-resolution modules, while existing methods can only train at a native resolution of 1282128^{2}.

Moreover, we carefully design training strategies to address the human pose and viewing angle imbalance issue. We analyze the head-facing angle distribution and propose a pose-guided sampling strategy to help effective 3D human geometry learning. Besides, we utilize the SMPL model to leverage its human prior during training. Specifically, we use SMPL skinning weights to guide the transformation between canonical and observation spaces, which shows good robustness to the pose distribution of the dataset and brings better generalizability to novel pose generation. We further use SMPL as the geometry template and predict offsets to help better geometry learning.

Quantitative and qualitative experiments are performed on two fashion datasets (Liu et al., 2016; Fu et al., 2022) to demonstrate the advantages of EVA3D. We also experiment on UBCFashion (Zablotskaia et al., 2019) and AIST (Tsuchida et al., 2019) for comparison with prior work. Extensive experiments on our method designs are provided for further analysis. In conclusion, our contributions are as follows: 1) We are the first to achieve high-resolution high-quality 3D human generation from 2D image collections; 2) We propose a compositional human NeRF representation tailored for efficient GAN training; 3) Practical training strategies are introduced to address the imbalance issue of real 2D human image collections. 4) We demonstrate applications of EVA3D, i.e. interpolation and GAN inversion, which pave way for further exploration in 3D human GAN.

Related Work

3D-Aware GAN. Generative Adversarial Network (GAN) (Goodfellow et al., 2020) has been a great success in 2D image generation (Karras et al., 2019; 2020). Many efforts have also been put on 3D-aware generation. Nguyen-Phuoc et al. (2019); Henzler et al. (2019) use voxels, and Pan et al. (2020) use meshes to assist the 3D-aware generation. With recent advances in NeRF (Mildenhall et al., 2020; Tewari et al., 2021), many have build 3D-aware GANs based on NeRF (Schwarz et al., 2020; Niemeyer & Geiger, 2021; Chan et al., 2021; Deng et al., 2022). To increase the generation resolution, Gu et al. (2021); Or-El et al. (2022); Chan et al. (2022) use 2D decoders for super resolution. Moreover, it is desirable to lift the raw resolution, by improving the rendering efficiency, for more detailed geometry and better 3D consistency (Skorokhodov et al., 2022; Xiang et al., 2022). We also propose an efficient 3D human representation to allow high resolution training.

Human Generation. Though great success has been achieved in generating human faces, it is still challenging to generate human images for the complexity in human poses and appearances (Sarkar et al., 2021b; Lewis et al., 2021; Sarkar et al., 2021a; Jiang et al., 2022c). Recently, Fu et al. (2022); Frühstück et al. (2022) scale-up the dataset and achieve impressive 2D human generation results. For 3D human generation, Chen et al. (2022) generate human geometry using 3D human dataset. Some also attempt to train 3D human GANs using only 2D human image collections. Grigorev et al. (2021); Zhang et al. (2021) use CNN-based neural renderers, which cannot guarantee 3D consistency. Noguchi et al. (2022) use human NeRF (Noguchi et al., 2021) for this task, which only trains at low resolution. Bergman et al. (2022); Zhang et al. (2022) propose to increase the resolution by super-resolution, which still fails to produce high-quality results. Hong et al. (2022b) generate 3D avatars from text inputs.

3D Human Representations. 3D human representations serve as fundamental tools for human related tasks. Loper et al. (2015); Pavlakos et al. (2019b); Hong et al. (2021) create parametric models for explicit modeling of 3D humans. To model human appearances, Habermann et al. (2021); Shysheya et al. (2019); Yoon et al. (2021); Liu et al. (2021) further introduce UV maps. Parametric modeling gives robust control over the human model, but less realism. Palafox et al. (2021) use implicit functions to generate realistic 3D human body shapes. Embracing the development of NeRF, the number of works about human NeRF has also exploded (Peng et al., 2021b; Zhao et al., 2021; Peng et al., 2021a; Xu et al., 2021; Noguchi et al., 2021; Weng et al., 2022; Chen et al., 2021; Su et al., 2021; Jiang et al., 2022a; b; Wang et al., 2022). Hong et al. (2022a) propose to learn modal-invariant human representations for versatile down-stream tasks. Cai et al. (2022) contribute a large-scale multi-modal 4D human dataset. Some propose to model human body in a compositional way (Mihajlovic et al., 2022; Palafox et al., 2022; Su et al., 2022), where several submodules are used to model different body parts, and are more efficient than single-network ones.

Methodology

NeRF (Mildenhall et al., 2020) is an implicit 3D representation, which is capable of photorealistic novel view synthesis. NeRF is defined as {c,σ}=FΦ(x,d)\{\bm{c},\sigma\}=F_{\Phi}(\bm{x},\bm{d}), where x\bm{x} is the query point, d\bm{d} is the viewing direction, c\bm{c} is the emitted radiance (RGB value), σ\sigma is the volume density. To get the RGB value C(r)C(\bm{r}) of some ray r(t)=o+td\bm{r}(t)=\bm{o}+t\bm{d}, namely volume rendering, we have the following formulation, C(r)=∫tntfT(t)σ(r(t))c(r(t),d)dtC(\bm{r})=\int_{t_{n}}^{t_{f}}T(t)\sigma(\bm{r}(t))\bm{c}(\bm{r}(t),\bm{d})dt , where T(t)=exp(−∫tntσ(r(s))ds)T(t)=\text{exp}(-\int_{t_{n}}^{t}\sigma(\bm{r}(s))ds) is the accumulated transmittance along the ray r\bm{r} from tnt_{n} to tt. tnt_{n} and tft_{f} denotes the near and far bounds. To get the estimation of C(r)C(\bm{r}), it is discretized as

For better geometry, Or-El et al. (2022) propose to replace the volume density σ(x)\sigma(\bm{x}) with SDF values d(x)d(\bm{x}) to explicitly define the surface. SDF can be converted to the volume density as σ(x)=α−1sigmoid(−d(x)/α)\sigma(\bm{x})=\alpha^{-1}\text{sigmoid}\left(-d(\bm{x})/\alpha\right), where α\alpha is a learnable parameter. In later experiments, we mainly use SDF as the implicit geometry representation, which is denoted as σ\sigma for convenience.

SMPL (Loper et al., 2015), defined as M(β,θ)M(\bm{\beta},\bm{\theta}), is a parametric human model, where β,θ\bm{\beta},\bm{\theta} controls body shapes and poses. In this work, we use the Linear Blend Skinning (LBS) algorithm of SMPL for the transformation from the canonical space to observation spaces. Formally, point x\bm{x} in the canonical space is transformed to an observation space defined by pose θ\bm{\theta} as x′=∑k=1KwkGk(θ,J)x\bm{x}^{\prime}=\sum_{k=1}^{K}w_{k}\bm{G_{k}}(\bm{\theta},\bm{J})\bm{x}, where KK is the joint number, wkw_{k} is the blend weight of x\bm{x} against joint kk, Gk(θ,J)\bm{G_{k}}(\bm{\theta},\bm{J}) is the transformation matrix of joint kk. The transformation from observation spaces to the canonical space, namely inverse LBS, takes a similar formulation with inverted transformation matrices.

2 Compositional Human NeRF Representation

m,nm,n are chosen empirically. Different from Palafox et al. (2022); Su et al. (2022), we only query subnetworks whose bounding boxes contain query points. It increases the efficiency of the query process and saves computational resources.

Taking advantages of the compositional representation, we also adopt an efficient volume rendering algorithm. Previous methods need to sample points, query, and integrate for every pixel of the canvas, which wastes large amounts of computational resources on backgrounds. In contrast, for the compositional representation, we have pre-defined bounding boxes to filter useful rays, which is also the key for our method being able to train on high resolution.

3 3D Human GAN Framework

4 Training

Delta SDF Prediction. Real-world 2D human image collections, especially fashion datasets, usually have imbalanced pose distribution. For example, as shown in Fig. 6, we plot the distribution of facing angles of DeepFashion. Such heavily imbalanced pose distribution makes it hard for the network to learn correct 3D information in an unsupervised way. Therefore, we propose to introduce strong human prior by utilizing the SMPL template geometry dT(x)\bm{d}_{T}(\bm{x}) as the foundation of our human representation. Instead of directly predicting the SDF value d(x)\bm{d}(\bm{x}), we predict an SDF offset Δd(x)\Delta\bm{d}(\bm{x}) from the template (Yifan et al., 2022). Then dT(x)+Δd(x)\bm{d}_{T}(\bm{x})+\Delta\bm{d}(\bm{x}) is used as the actual SDF value of point x\bm{x}.

Pose-guided Sampling. To facilitate effective 3D information learning from sparse 2D image collections, other than introducing a 3D human template, we propose to balance the input 2D images based on human poses. The intuition behind the pose-guided sampling is that different viewing angles should be sampled more evenly to allow effective learning of geometry. Empirically, among all human joints, we use the angle of the head to guide the sampling. Moreover, facial areas contain more information than other parts of the head. Front-view angles should be sampled more than other angles. Therefore, we choose to use a Gaussian distribution centered at the front-view angle μθ\mu_{\theta}, with a standard deviation of σθ\sigma_{\theta}. Specifically, MM bins are divided on the circle. For an image with the head angle falling in bin mm, its probability pmp_{m} of being sampled is defined as

We visualize the balanced distribution in Fig. 6. The network now has higher chances of seeing more side-views of human bodies, which helps better geometry generation.

Loss Functions. For the adversarial training, we use the non-saturating GAN loss with R1 regularization (Mescheder et al., 2018), which is defined as

where f(u)=−log(1+exp(−u))f(u)=-\text{log}(1+\text{exp}(-u)). Other than the adversarial loss, some regularization terms are introduced for the delta SDF prediction. Firstly, we want minimum offset from the template mesh to maintain plausible human shape, which gives the minimum offset loss Loff=Ex[∥Δd(x)∥22]\mathcal{L}_{\text{off}}=\bm{E}_{\bm{x}}[\|\Delta d(\bm{x})\|_{2}^{2}]. Secondly, to ensure that the predicted SDF values are physically valid (Gropp et al., 2020), we penalize the derivation of delta SDF predictions to zero Leik=Ex[∥∇(Δd(x))∥22]\mathcal{L}_{\text{eik}}=\bm{E}_{\bm{x}}[\|\nabla(\Delta d(\bm{x}))\|_{2}^{2}]. The overall loss is defined as L=Ladv+λoffLoff+λeikLeik\mathcal{L}=\mathcal{L}_{\text{adv}}+\lambda_{\text{off}}\mathcal{L}_{\text{off}}+\lambda_{\text{eik}}\mathcal{L}_{\text{eik}}, where λ∗\lambda_{*} are loss weights defined empirically.

Experiments

Datasets. We conduct experiments on four datasets: DeepFashion (Liu et al., 2016), SHHQ (Fu et al., 2022), UBCFashion (Zablotskaia et al., 2019) and AIST (Tsuchida et al., 2019). The first two are sparse 2D image collections, meaning that each image has different identities and poses are sparse, which makes them more challenging. The last two are human video datasets containing different poses/ views of the same identities, which is easier for the task but lacks diversity.

Comparison Methods. We mainly compare with three baselines. ENARF-GAN (Noguchi et al., 2022) makes the first attempt at human NeRF generation from 2D image collections. EG3D (Chan et al., 2022) and StyleSDF (Or-El et al., 2022) are state-of-the-art methods for 3D-aware generation, both requiring super-resolution modules to achieve high-resolution generation.

Evaluation Metrics. To evaluate the quality of rendered images, we adopt Frechet Inception Distance (FID) (Heusel et al., 2017) and Kernel Inception Distance (KID) (Bińkowski et al., 2018). Following ENARF-GAN, we use Percentage of Correct Keypoints (PCKh@0.5) (Andriluka et al., 2014) to evaluate the correctness of generated poses. Note that PCKh@0.5 can only be calculated on methods that can control generated poses, i.e. ENARF-GAN and EVA3D. To evaluate the correctness of geometry, we use an off-the-shelf tool (Ranftl et al., 2022) to estimate depth from the generated images and compare it with generated depths. 50K50K samples padded to square are used to compute FID and KID. PCKh@0.5 and Depth are evaluated on 5K5K samples.

2 Qualitative Evaluations

Generation Results and Controlling Ability of EVA3D. As shown in Fig. 4 a), EVA3D is capable of generating high-quality renderings in novel views and remain multi-view consistency. Due to the inherent human prior in our model design, EVA3D can control poses and shapes of the generated 3D human by changing β\bm{\beta} and θ\bm{\theta} of SMPL. We show novel pose and shape generation results in Fig. 4 b)& c). We refer readers to the supplementary PDF and video for more qualitative results.

Comparison with Baseline Methods. We show the renderings and corresponding meshes generated by baselines and our method in Fig. 5. EG3D trained on DeepFashion, as well as StyleSDF trained on SHHQ, generate reasonable RGB renderings and geometry. However, without explicit human modeling, complex human poses make it hard to align and model 3D humans in observation spaces, which leads to distorted generation. Moreover, because of the use of super resolution, their geometry is only trained under low resolution (64264^{2}) and therefore lacks details. EG3D trained on SHHQ and StyleSDF trained on DeepFashion fail to capture 3D information and collapse to the trivial solution of painting on billboards. Limited by the inefficient representation and computational resources, ENARF-GAN can only be trained at a resolution of 1282128^{2}, which leads to low-quality rendering results. Besides, lacking human prior makes ENARF-GAN hard to capture correct 3D information of human from sparse 2D image collections, which results in broken meshes. EVA3D, in contrast, generates high-quality human renderings on both datasets. We also succeeded in learning reasonable 3D human geometry from 2D image collections with sparse viewing angles and poses, thanks to the strong human prior and the pose-guided sampling strategy. Due to space limitations, we only show results of DeepFashion and SHHQ here. For visual comparisons on UBCFashion and AIST, please refer to the supplementary material.

3 Quantitative Evaluations

As shown in Tab. 1, our method leads all metrics in four datasets. EVA3D outperforms ENARF-GAN in all settings thanks to our high-resolution training ability. EG3D and StyleSDF, as the SOTA methods in the 3D generation, can achieve reasonable scores in some settings (e.g. StyleSDF achieves 18.52 FID on UBCFashion) for their super-resolution modules. But they also fail on some datasets (e.g. StyleSDF fails on AIST with 199.5 FID) for complexity in human poses. In the contrast, EVA3D achieves the best FID/KID scores under all settings. Moreover, unlike EG3D or StyleSDF, EVA3D can control the generated pose and achieve higher PCKh@0.5 score than ENARF-GAN. For the geometry part, we also achieve the lowest depth error, which shows the importance of natively high-resolution training.

4 Ablation Studies

Ablation on Method Designs. To validate the effectiveness of our designs on EVA3D, we subsequently add different designs on a baseline method, which uses one large network to model the canonical space. Experiments are conducted on DeepFashion. The results are reported in Tab. 3. Limited by the inefficient representation, the baseline (“Baseline”) can only be trained at 256×128256\times 128, which results in the worst FID score. Adding compositional design (“+Composite”) makes the network efficient enough to be trained at a higher resolution of 512×256512\times 256 and achieve higher generation quality. We further introduce human prior by predicting delta SDF (“+Delta SDF”), which gives the best FID score and lower depth error. Finally, using the pose-guided sampling (“+Pose-guide”), we further decrease the depth error, which means better geometry. However, the FID score slightly increases, which is discussed in the next paragraph. We refer readers to the supplementary material for qualitative evaluations of ablation studies.

Analysis on Pose-Guided Sampling. We further analyze the importance of the sampling strategy in 3D human GAN training. Three types of distributions pestp_{est} are experimented, including the original dataset distribution (“Original”), pose-guided Gaussian distribution (“σθ=∗\sigma_{\theta}=*”), and pose-guided uniform distribution (“Uniform”). The results are reported in Tab. 3. Firstly, uniform sampling is not a good strategy, as shown by its high FID score. This is because the information density is different between different parts of human. Faces require more training iterations. Secondly, the original distribution gives the best visual quality but the worst geometry. As shown in Fig. 6, the original distribution leads to the network mostly being trained on front-view images. It could result in the trivial solution of painting on billboards. Though using delta SDF prediction can alleviate the problem to some extent, the geometry is still not good enough. Thirdly, the pose-guided Gaussian sampling can avoid damaging visual quality too much and improve geometry learning. As the standard deviation σθ\sigma_{\theta} increases, FID increases while the depth error decreases. Therefore, it is a trade-off between visual quality and geometry quality. In our final experiments, we choose σθ=15∘\sigma_{\theta}=15^{\circ} which is a satisfying balance between the two factors.

5 Applications

Interpolation on Latent Space. As shown in Fig. 7 a), we linearly interpolate two latent codes to generate a smooth transition between them, showing that the latent space learned by EVA3D is semantically meaningful. More results are provided in the supplementary video.

Inversion. We use Pivotal Tuning Inversion (PTI) (Roich et al., 2021) to inverse the target image and show the results in Fig. 7 b). Reasonable novel view synthesis results can be achieved. The geometry, however, fails to capture geometry details corresponding to RGB renderings, which can be caused by the second stage generator fine-tuning of PTI. Nevertheless, we demonstrate the potential of EVA3D in more related downstream tasks.

Discussion

To conclude, we propose a high-quality unconditional 3D human generation model EVA3D that only requires 2D image collections for training. We design a compositional human NeRF representation for efficient GAN training. To train on the challenging 2D image collections with sparse viewing angles and human poses, e.g. DeepFashion, strong human prior and pose-guided sampling are introduced for better GAN learning. On four large-scale 2D human datasets, we achieve state-of-the-art generation results at a high resolution of 512×256512\times 256.

Limitations: 1) There still exists visible circular artifacts in the renderings, which might be caused by the SIREN activation. A better base representation, e.g. tri-plane of EG3D, and a 2D decoder might solve the issue. 2) The estimation of SMPL parameters from 2D image collections is not accurate, which leads to a distribution shift from the real pose distribution and possibly compromises generation results. Refining SMPL estimation during training would make a good future work. 3) Limited by our tight 3D human representation, it is hard to model loose clothes like dresses. Using separate modules to handle loose clothes might be a promising direction.

Although the results of EVA3D are yet to the point where they can fake human eyes, we still need to be aware of its potential ethical issues. The generated 3D humans might be misused to create contents that are misleading. EVA3D can also be used to invert real human images, which can be used to create fake videos of real humans and cause negative social impacts. Moreover, the generated 3D humans might be biased, which is caused by the inherent distribution of training datasets. We make our best effort to demonstrate the impartiality of EVA3D in Fig. 1.

Reproducibility Statement

Our method is thoroughly described in Sec. 3. Together with implementation details included in the supplementary material, the reproducibility is ensured. Moreover, our code will be released upon acceptance.

Acknowledgments

This work is supported by NTU NAP, MOE AcRF Tier 2 (T2EP20221-0033), and under the RIE2020 Industry Alignment Fund – Industry Collaboration Projects (IAF-ICP) Funding Initiative, as well as cash and in-kind contribution from the industry partner(s).

References

Appendix A Appendix

This is the supplementary material for EVA3D. We provide more information of the training datasets in Sec. A.1. More implementation details are introduced in Sec. A.2. More visual results and comparisons are provided in Sec. A.3. We also attached a demo video for better viewing experience.

DeepFashion (Liu et al., 2016) collects fashion images from the internet. We only use images that contain the full body and not wearing dresses, which results in 8,036 images for training. We use SMPLify-X (Pavlakos et al., 2019b) to estimate SMPL parameters and camera parameters. All images are resized to 512×256512\times 256 for training. The alignment of the human body is the same as that proposed by Jiang et al. (2022c).

SHHQ (Fu et al., 2022) collects a larger-scale fashion dataset from the internet. It is processed similarly as DeepFashion, which results in 120,865 images in resolutions of 512×256512\times 256. In our experiments, we find that models trained using SMPL estimated by SMPLify-X performs better than that of SPIN (Kolotouros et al., 2019). The head direction distribution of SHHQ, like DeepFashion, is also heavily imbalanced, as shown in Fig. 8 a). The accompanying blue line is the distribution of the proposed pose-guided sampling.

UBCFashion (Zablotskaia et al., 2019) is a fashion video dataset containing 500 sequences of models posing in front of the camera. Most models wear dresses in this dataset. We estimate SMPL sequences from videos by VIBE (Kocabas et al., 2020). We use all frames of the 500 videos and crop them to 512×256512\times 256 for training, which leads to 192,179 samples. Although most models spin in front of the camera, the head direction of UBCFashion is still heavily imbalanced, as shown in Fig. 8 b).

AIST (Tsuchida et al., 2019) is a multi-view human dancing video dataset that provides rich poses and accurate SMPL estimations. We directly use the dataset processing scripts provided by ENARF-GAN (Noguchi et al., 2022) and get 72,000 samples. Each sample is resized to 256×256256\times 256 for training.

A.2 Implementation Details

As introduced in the main paper, we split the whole body into 16 parts, which is shown in Fig. 9 in specific. For each part, a subnetwork is assigned, which is developed based on StyleSDF (Or-El et al., 2022). The architecture of each subnetwork is shown in Fig. 9 b). For each subnetwork, multiple MLP and FiLM SIREN (Chan et al., 2021) activation layers are stacked alternatively. At the end of each subnetwork, two branches are used to separately estimate SDF value and RGB value. We assign different numbers of network layers for different body part empirically. Specific numbers are listed on Fig. 9 a). For the discriminator, we use the same architecture as that of StyleSDF (Or-El et al., 2022).

A.2.2 Training Settings

Hyperparameters. We use Adam optimizer (Kingma & Ba, 2014) for the optimization of both generator and discriminator. The learning rate for generator is 2×10−52\times 10^{-5}. The learning rate for discriminator is 2×10−42\times 10^{-4}. The loss weights as set empirically as λoff=1.5\lambda_{\text{off}}=1.5 and λeik=0.5\lambda_{\text{eik}}=0.5. For one ray, 28 query points are sampled. For the pose-guided sample, we choose to use σθ=15∘\sigma_{\theta}=15^{\circ}.

R1 Scheduler. R1 regularization is used during training to penalize gradients of discriminator. Because it is highly challenging for the generator to learn plausible human appearance, the discriminator tends to overfit quickly if low R1 is set. But too high of R1 value would harm the final generation quality. Therefore, we set a R1 scheduler empirically, where R1 decrease from 300 to 18.5. R1 is cut in half every 50,000 iterations.

Augmentation. Inevitably, the SMPL estimations for 2D human images are not accurate for most samples. To compensate for the estimation error, we adopt small augmentations on real and fake samples before sent to the discriminator. The augmentation includes random panning, scaling and rotation in small ranges.

Runtime Analysis. The models are trained on 8 NVIDIA V100 GPUs for 5 days, with a batch size of 8. At test time, our model runs at ∼5\sim 5 FPS on one NVIDIA V100 GPU.

A.3 More Qualitative Results

Visual Comparison on UBCFashion & AIST. We further show renderings and corresponding meshes of three baseline methods and EVA3D trained on UBCFashion and AIST in Fig. 10. UBCFashion has dense views and simple human poses. Therefore, EG3D and StyleSDF succeeded in generating reasonable renderings. But the corresponding meshes lack details due to training at low native resolution (64×6464\times 64). EVA3D gives the best visual results among the baseline methods and also generates plausible meshes with reasonable details. Due to complex human poses, StyleSDF fails on AIST. EG3D manages to generate reasonable 3D human, but fails to capture correct human structure in some cases. ENARF-GAN, for its low-resolution training, loses most details and generates rough meshes. EVA3D not only gets the best RGB renderings, but also generates meshes that preserve details like brims.

Qualitative Evaluations on Ablation Studies. As shown in Fig. 11, we visualize renderings and geometry generated by baseline methods described in the ablation studies in the main paper. The “Baseline”, due to being trained at lower resolution (256×128256\times 128), generates blurry renderings. The geometry fails to capture correct human structure (see broken knees). The compositional 3D human representation (“+ Composite”) facilitates high resolution training (512×256512\times 256). But lack of human prior leads to low-quality geometry (see unreasonable “wrinkles” on the upper bodies). By introducing a 3D human template and predicting delta SDF (“+ Delta SDF”), the visual quality increases and the geometry is mostly reasonable. However, the facial area is still flat due to the highly imbalanced viewing angle distribution. By using the pose-guided sampling (“+ Pose-Guided Sample”), we alleviate the imbalance issue and generate both high fidelity renderings and plausible geometry. To further validate our choice of Gaussian distribution in the pose-guided sampling, we visualize the results of models trained using uniform distribution (“+ Uniform Sample”). The middle of generated heads have severe artifacts.

More Qualitative Results of EVA3D. More qualitative results of EVA3D on four datasets are shown in Fig. 12, 13, 14, 15. For each sample, we show its novel view renderings and novel pose rendering.