GRAM: Generative Radiance Manifolds for 3D-Aware Image Generation

Yu Deng, Jiaolong Yang, Jianfeng Xiang, Xin Tong

Introduction

Learning 3D-aware image generation with Generative Adversarial Networks (GAN) has attracted a surge of attention in recent years . Given an unstructured 2D image collection, GANs are trained to synthesize geometrically-consistent multiview imagery of novel instances. In particular, methods that use the volumetric rendering paradigm to composite an output image have demonstrated impressive results with more “strict” 3D consistency by virtue of an explicit, physics-based rendering process.

Notwithstanding the promising results shown by these methods, the image quality still lags far behind traditional 2D image synthesis, for which state-of-the-art GAN models can generate high-resolution and photorealistic images. One prominent hurdle is the high computation and memory requirements for training a volumetric representation. Methods that use neural radiance field (NeRF) generators can greatly reduce the complexity of voxel-based approaches , but the volume integrations approximated by sampling points along viewing rays are still costly for both training and inference.

This problem becomes even more pronounced in GAN training where a full image (rather than sparse pixels) needs to be rendered to train the discriminator. One workaround is to render patches during training , but using a patch discriminator may lead to inferior image generation quality. With an image discriminator, the state-of-the art method can only afford training on smaller image resolution and with significantly reduced number of sampling points per ray (typically a few dozens) compared to standard NeRF . However, we observed that radiance integration using Monte Carlo sampling becomes unstable with insufficient samples. The integrated colors among adjacent pixels suffer from intractable noise patterns that are detrimental to GAN training (e.g., see Fig. 11). An even worse issue is that optimizing a full radiance volume requires the sampling to cover both low-frequency regions and high-frequency details, leading to even less sample budget for the latter. Consequently, it is extremely difficult to generate fine details as they simply can be missed by the sampling.

This paper presents a novel method named Generative Radiance Manifolds (GRAM). Different from the previous methods, we constrain our point sampling and radiance field learning on 2D manifolds, embodied as a set of implicit surfaces. These implicit surfaces are shared for the trained object category, jointly learned with GAN training, and fixed at inference time. To generate an image, we accumulate the radiance along each ray using ray-surface intersections as point samples.

There are several advantages of our GRAM method. First, by confining sampling and radiance learning in a reduced space rather than anywhere in the volume, it greatly facilitates fine detail learning. The network can easily learn to generate thin structures and texture details on the surface manifolds which are guaranteed to have projections on the image and receive supervision during GAN training. Besides, our generated images are free from the noise pattern caused by inadequate Monte Carlo sampling, as the ray-surface intersections are deterministically calculated and smoothly varying across rays. Even with very few point samples (i.e., learning very few surfaces), our method can still learn to generate high-quality results. As a byproduct, at inference time we can render a generated instance in real time by pre-extracting the surfaces with their radiance.

Our implicit surfaces are defined as a set of isosurfaces in a scalar field predicted by a light-weight MLP network. Another MLP for radiance generation is employed, for which we use a structure similar to . We extract ray-surface intersections in a differentiable manner, and the whole framework is trained end-to-end using adversarial learning. Orthogonal to our novel radiance manifold design, we also explore network architecture and training method enhancements. In particular, we modify the network structure of inspired by and remove the progressive growing strategy used therein. Progressive growing not only introduces additional hyperparameters to tune but may also lead to degraded image quality shown in traditional 2D GAN . We also empirically find that our method generates better results by removing it.

Our method is evaluated on multiple datasets including FFHQ , Cats , and CARLA . We show that our 3D-aware generation method significantly outperforms the prior art. It can synthesize highly realistic images with geometrically-consistent fine details, which are unseen in previous results. We believe our method makes a significant step towards diminishing the quality gap between 3D-aware generation and traditional 2D image generation.

Related Work

For scene representation and synthesis, a large volume of works adopt neural networks as a new type of rendering tool due to their ability to synthesize high-quality images without requiring excessive human labor. Among them, earlier works employ convolutional networks for a variety of applications such as novel view synthesis , image-to-image translation , and controllable image manipulation .

More recently, plenty of works leverage implicit neural representations to model 3D scenes using Multi-Layer Perceptrons (MLP). The continuous representation of MLPs brings them the superiority at 3D-level control of image synthesis compared to conventional CNN-based methods. Among these approaches, NeRF shows promising results in capturing complex scene structures and synthesizing 3D-consistent images with fine details. Most of the NeRF-based methods focus on scene-specific learning tasks where a network is trained to fit a set of posed images of a certain scene. Only a few recent methods work on the image generation task using unconstrained 2D images for supervision. This paper proposes a new generative model for improving the image generation quality while maintaining the 3D consistency of generated contents.

D-Aware Image Generation.

Given uncontrolled 2D image collections, 3D-aware image generation methods aim to learn a generative model that can explicitly control the camera viewpoint of the generated content. To achieve this goal, the literature mainly follows two directions. The first line of works utilize 3D-aware features to represent a scene, and apply a neural renderer, typically a CNN, on top of them for realistic image synthesis. For example, HoloGAN and BlockGAN learn low-resolution voxel features for objects, project them onto 2D image plane, and apply a StyleGAN-like CNN to generate higher-resolution images. Liao et al. first generate 3D primitives using a 3D generator and then apply a 2D generator with an encoder-decoder structure on the projected features. Giraffe and GANcraft instead use 3D volumetric rendering to generate 2D feature maps for the subsequent image generation. Following a similar idea, some works concurrent to ours focus on designing better rendering networks to enable 3D-aware image generation at very high resolution. Nevertheless, an inevitable problem of these methods is the sacrifice of exact multi-view consistency due to the learned black-box rendering.

Another group of works seek to learn direct 3D representation of scenes and synthesize images under physical-based rendering process to achieve more strict 3D consistency. and adopt a mesh-based representation and generate images via rasterization. However, they cannot well handle complicated structures with non-Lambertian reflectance such as hair and fur. Recent methods use the NeRF representation to synthesize images with high 3D consistency. Still, the expensive computational cost of volumetric representation learning prevents them from generating images with adequate details. In this work, we propose a novel approach to learn a generative radiance field on 2D manifolds, and we achieve more realistic image generation with finer details significantly outperforming the previous methods.

Approach

Figure 2 shows the overall structure of GG, which consists of a manifold predictor M\mathcal{M} and a radiance generator Φ\Phi. The manifold predictor M\mathcal{M} defines a scalar field which derives a reduced domain for radiance generation, which is composed of multiple implicit isosurfaces (Sec. 3.1). Given a latent code z{\bm{z}}, the radiance generator Φ\Phi generates the occupancy and color for points on the manifolds (Sec. 3.2). Images are then generated by integrating the color of the manifold points along each viewing ray (Sec. 3.3). The whole method is trained end-to-end in an adversarial learning framework (Sec. 3.4). After training, GRAM can render high-quality and 3D-consistent images from different viewpoints.

Our manifold predictor M\mathcal{M} predicts a reduced space for point sampling and radiance field learning, which is shared across all generated instances. We implement it as a scalar field function which determines a set of isosurfaces. Specifically, M\mathcal{M} is a light-weight MLP which takes a point x\bm{x} as input and predicts a scalar value ss:

Given the predicted scalar field, we obtain NN isosurfaces {Si}\{\mathcal{S}_{i}\} with different levels {li}\{l_{i}\}:

These levels are predefined constant values. Note that although the scalar field is defined in the 3D volume of the scene to be rendered, the scalar values per se have no physical meaning and the levels {li}\{l_{i}\} can be trivially chosen.

We define the input domain of the radiance generator to be on these surfaces. Let {xi}\{{\bm{x}}_{i}\} be the NN intersections between a camera ray r={o+td,t∈[tn,tf]}\bm{r}=\{\bm{o}+t\bm{d},t\in[t_{n},t_{f}]\} and {Si}\{\mathcal{S}_{i}\}:

where oo and dd are ray origin and direction, and tnt_{n} and tft_{f} are the near plane and far plane parameters. We only pass {xi}\{{\bm{x}}_{i}\} to the radiance generator Φ\Phi for radiance generation and final rendering, as shown in Fig. 2. Since there is no prior knowledge for optimal isosurfaces, we learn them jointly in the generative adversarial training process.

Training the manifold predictor M\mathcal{M} with GAN necessitates a differentiable scheme for ray-surface intersection computation in order to backpropagate the adversarial loss. To this end, we follow Niemeyer et al. ’s strategy to calculate the intersections. As shown in Fig. 3, we evenly sample points along a ray between the near and far planes and feed them to M\mathcal{M} to obtain their values ss. Then we search for the first interval that a certain scalar level lil_{i} falls in, and calculate the intersection using linear interpolation between the two endpoints of the interval via:

We implement M\mathcal{M} as a light-weight MLP with 3 hidden layers, and thus dense points (64 points in our implementation) can be sampled to get accurate intersections using Eq. (5).

Random initialization of M\mathcal{M} may give rise to highly irregular isosurfaces which is unfavourable for the training process. In this work, we adopt the geometric initialization strategy proposed by Atzmon et al. with which the initial isosurfaces are close to spheres.

2 Radiance Generator

Since radiance is defined on surface manifolds instead of the whole volume in our method, we generate occupancy α\alpha instead of volume density σ\sigma in NeRF, following .

The network structure of Φ\Phi is adapted from the FiLM SIREN backbone of with some modifications, as presented in Fig. I. Inspired by StyleGAN2 , we use skip connections between output layers at different levels instead of only predicting occupancy and color at the final layer as done in previous methods . In this way, different levels of details are now predicted by different output layers and combined together to form the final results. This change not only removes the necessity of the progressive growing strategy used in previous methods, but also yields better results in our method as shown in the experiments.

3 Manifold Rendering

For a camera ray r\bm{r} which intersects the surface manifolds at points {xi}\{{\bm{x}}_{i}\} sorted from near to far following Eq. (4), the rendering equation can be written as :

Our rendering scheme is clearly different from the original volume rendering in NeRF which applies a hierarchical random sampling strategy (NeRF-H). NeRF-H’s sampling points may vary significantly across adjacent rays due to sampling randomness, resulting in noise patterns on the rendered image (see Fig. 11). By contrast, we only use intersections between camera rays and surface manifolds which are deterministically calculated and smoothly varying across rays, instead of selecting points in the whole volume space in a Monte Carlo fashion. This helps us eliminate the randomness in image generation and enable training a generator with fewer point samples per ray. Moreover, it greatly facilitates fine detail learning as high-frequency structures and textures can be easily generated on the surface manifolds (see Table 2 and Table 3).

4 Training Strategy

At training stage, we randomly sample latent code z{\bm{z}} and camera pose θ{\bm{\theta}} from prior distributions pzp_{z} and pθp_{\theta}. The generator GG synthesizes images with corresponding latent codes and poses as input. We also sample real images from the training data with prior distribution prealp_{real}. As in standard GAN , a discriminator DD receives the generated images as well as real images and judge if they are fake or real, for which we use the same CNN structure as in . We train all the networks, including the manifold predictor M\mathcal{M}, the radiance generator Φ\Phi and the discriminator DD, using non-saturating GAN loss with R1 regularization :

where f(u)=log⁡(1+exp⁡(u))f(u)=\log(1+\exp(u)) is the Softplus function.

In addition, we find that for certain objects, the training process with only adversarial loss is sometimes sensitive to random initialization. In a few occasions, the learned 3D geometry of convex objects could become concave (see Sec. B.3). To tackle this issue, we can optionally add a pose regularization term to enforce the generator to generate images under correct pose:

where DpD_{p} is an additional branch of the discriminator DD that predicts the camera pose of a given image, and θ^\hat{{\bm{\theta}}} is the pose label of a real image. We find that this loss can also slightly improve the image generation quality for objects without the concave geometry issue observed.

Experiments

We use three datasets for evaluation: FFHQ , Cats , and CARLA , which contain 70K high-resolution face images, 10K cat images with various resolutions, and 10K synthetic car images of 16 car models, respectively. For all experiments, we use the Adam optimizer , and the learning rates are set to 2e−52e{-5} for the generator and 2e−42e{-4} for the discriminator. The models are trained on 8 NVIDIA Tesla V100 GPUs with 32GB memory. More details can be found in Sec. A.

1 Generation Results

Some random image samples generated by our method are shown in Fig. 1, 5, and 9. For face and cat, the model is trained with 2562256^{2} resolution and 2424 manifold surfaces (i.e., 2424 point samples per ray). For the car images, we train on 1282128^{2} resolution and use 4848 manifold surfaces. As we can see, GRAM is able to generate high-quality images with fine details. Moreover, it allows an explicit control of camera viewpoint and achieves highly consistent results across different views. It even maintains strong visual 3D consistency for very thin structures such as bangs of hair, eyeglass, and whiskers of cat, which show correct parallax corresponding to realistic 3D geometry. Note that 3D consistency is best viewed with animations, which can be found on our project page.

Figure 6 shows the learned surface manifolds on the three datasets. Initially, the surfaces have near-spherical shapes and are positioned across the whole volume. After training, the surfaces for face and cat are tightened and exhibit small curvatures. The surfaces for car are also tightened but maintain a curving structure that covers the car geometry. The face and cat images from FFHQ and Cats only have small angle variations; most of them are nearly frontal. In this case, near-planar surfaces are enough to render a generated instance. In contrast, the camera viewpoints of the car images from CARLA are uniformly distributed on the upper hemisphere (i.e., 360∘360^{\circ} azimuth and 90∘90^{\circ} elevation angles). Such a wide viewpoint range necessities curved surfaces to ensure good rendering results from different views.

Figure 7 shows the radiance predicted on the manifolds with two examples. We evenly sample surfaces from front to back and render the color patterns on them with their contribution to the final image as opacity. As shown in the figures, the network is able to learn high-frequency details and thin structures (e.g. whiskers) on the manifolds.

Visualization of 3D geometry.

Although our method confines the input domain of the radiance field on 2D manifolds, we can still extract proxy 3D shapes of the generated objects using the volume-based marching cubes algorithm . Figure 8 shows the proxy 3D shapes of several generated instances. It can be observed that our method produces high-quality geometry with detailed structures well depicted, which is the key to achieve strong visual 3D consistency across different views for not only low-frequency regions but also fine details.

2 Comparison with Previous Methods

We compare GRAM with three state-of-the-art 3D-aware image generation approaches: GRAF , pi-GAN , and GIRAFFE . Experiments are conducted using the official implementation provided by the authors. For GRAF and GIRAFFE, we modify the camera pose distribution according to different datasets, and leave other configurations unchanged. For pi-GAN, we follow the authors’ settings that use 24, 48, and 96 sampling points for FFHQ, Cats, and CARLA respectively, for both training and testing. Note that for our method, we use 24 surfaces for FFHQ and Cats, and 48 surfaces for CARLA.

We further compare GRAM with a face-specific controllable image generation approach: DiscofaceGAN , which uses a 2D CNN as the generator and achieves pose control with the guidance of a prior 3D face model .

Figure 9 shows the visual comparison between GRAM and other methods. As we can see, GRAF and pi-GAN struggle to generate high-frequency details such as the texture of hair and fur. GIRAFFE produces images with finer details, but it suffers from 3D inconsistency (e.g., see hair region of the woman) due to the use of a CNN renderer . Our method achieves the best visual quality with realistic details and remarkable 3D consistency. See Fig. VII and our project page for more results.

Figure 10 shows the qualitative comparison between GRAM and DiscofaceGAN. While DiscofaceGAN can generate realistic face images and explicitly control their camera poses, it cannot well maintain the 3D consistency (e.g., see the bangs). By contrast, GRAM achieves strong 3D consistency under comparable generation quality without requiring extra 3D face priors.

Quantitative comparison.

We evaluate the image quality using the Fréchet Inception Distances (FID) and Kernel Inception Distances (KID) between 2020K randomly generated images and 2020K sampled real images. Table 1 shows that we significantly improve the two metrics compared to GRAF and pi-GAN, which also use NeRF generators. We even achieve lower FID and KID compare to GIRAFFE which applies a refinement CNN after the NeRF rendering to achieve better image quality. GIRAFFE is trained on a single GPU following its original implementation.

3 Ablation Study

We further conduct ablation study to validate the efficacy of our method designs. For efficiency, all experiments are conducted on FFHQ with 1282128^{2} resolution. Unless otherwise specified, we use 2424 points per ray for these experiments.

We compare our manifold sampling strategy with several baseline methods as shown in Table 2. NeRF-H is the original hierarchical sampling strategy used in NeRF and pi-GAN . Planes denotes using intersections between camera rays and multiple parallel planes placed across the volume. Spherical (init) denotes sphere-like surfaces obtained from the geometric initialization and fixed during training. Compare to the alternatives, our learnable manifolds yield the best image quality in terms of FID metrics. NeRF-H has a large performance gap with the others, indicating its deficiency under limited sample points. Our method outperforms Planes and Spherical (init), which demonstrates the advantage of using learnable surfaces that can better fit the trained object category.

Number of surface manifolds.

We further evaluate the generation quality of GRAM when training with different number of surfaces. For a reference, we also train models using the hierarchical sampling strategy NeRF-H with same number of sampling points for each ray. Table 3 shows that our method can generate high quality results using as few as 6 surfaces, and adding more gradually improves the quality. In contrast, training with NeRF-H largely fails with less than 12 points as indicated by the high FIDs, due to the difficulty to handle high-frequency details as well as the noise brought by inadequate sampling (Fig. 11). Even using 48 points, its generation quality is still worse than ours with 6 surfaces. In addition, it tends to learn unreasonable geometry with concave human foreheads, which rarely happens in our case (see Fig. VIII for visual results).

Influence of pose regularization.

Table 4 shows the effect of using pose labels of real images in Eq. (9) during training. For human face, our method produces slightly better results using the real pose regularization. In contrast, the hierarchical sampling strategy is unstable without real pose as guidance, leading to much worse results.

Training strategy and network structure.

As shown in Table 5, we first train our GRAM model with the network structure proposed in and the progressive growing strategy from 32232^{2} resolution following , which is the Base setting. Then we switch to the non-progressive growing strategy by training a model from scratch using 1282128^{2} resolution. Finally, we add skip connections in the network structure as depicted in Fig. I. The improvements on FID clearly demonstrate the advantages of our design.

4 Real-time multiview synthesis.

For objects generated by GRAM, we can achieve real-time free-view rendering thanks to our radiance manifold design. Spefically, we pre-extract the surface manifolds using marching cubes and store the radiance on them. With an efficient mesh rasterizer , we achieve 180FPS free-view rendering of 2562256^{2} images on a Nvidia Tesla V100 GPU.

Conclusions

We presented a novel approach for 3D-aware image generation. The core idea is to regulate point sampling and radiance learning on 2D manifolds for the radiance generator. Extensive experiments have shown its superiority over previous methods on both generation quality and 3D consistency. We believe our method takes a large step towards generating 3D-aware virtual contents for real applications.

The goal of this paper is to study generative modelling of the 3D objects from 2D images, and to provide a method for generating multi-view images of non-existing, virtual objects. It is not intended to manipulate existing images nor to create content that is used to mislead or deceive. This method does not have understanding and control of the generated content. Thus, adding targeted facial expressions or mouth movements is out of the scope of this work. However, the method, like all other related AI image generation techniques, could still potentially be misused for impersonating humans. Currently, the images generated by this method contain visual artifacts, unnatural texture patterns, and other unpredictable failures that can be spotted by humans and fake image detection algorithms. We also plan to investigate applying this technology for advancing 3D- and video-based forgery detection.

Limitations and future works.

Under constrained sampling budgets, our shared surfaces across the whole class can cause certain artifacts (see Sec. B.3) and limit our method to object categories sharing similar geometry. It may not well handle complex 3D scenes of multiple subjects with diverse structures. Learning instance-specific manifolds is a possible solution in the future. Besides, the generation quality and speed of GRAM still falls behind traditional 2D GANs. Better representations could be explored to further improve the fidelity and efficiency.

Acknowledgements.

We thank Harry Shum for the fruitful advice and discussion to improve the paper.

References

Appendix A More Implementation Details

We align the face images in FFHQ using 5 facial landmarks to centralize the faces and normalize their scales. Specifically, we first detected 5 facial landmarks of the images using an off-the-shelf landmark detector . Then we follow to resize and crop the images by solving a least square problem between the detected keypoints and corresponding 3D keypoints derived from a 3D face model . For pose distribution estimation, the face reconstruction method of is applied to extract the face poses for all the training images. Gaussian distributions are then fitted on the extracted poses, which are defined by the yaw and pitch angles (standard deviation 0.30.3 radians and 0.150.15 radians, respectively). During GAN training, we sample camera pose from the distributions and generate images accordingly. The extracted poses also serve as the pseudo labels for the pose regularization term defined in Eq. (9) of the main paper.

Cats [71].

For the cat images, we follow a similar procedure to align and resize the images using landmarks provided by the dataset . We also estimate the camera pose by solving the least square problem between the provided 2D landmarks and a set of manually-selected 3D landmarks on a 3D cat mesh. We found the pose distribution is very close to face images in FFHQ, and thus we simply use the same Gaussian to sample poses during training.

CARLA [16, 58].

We directly resize the car images rendered by to 1282128^{2} resolution without any alignment. Following , we uniformly sample camera pose from the upper hemisphere during training.

A.2 Network Structure

Figure I (a) shows the structure of the manifold predictor, which is an MLP with three hidden layers and an output layer. We set the channel dimension of the hidden layers to 128, 64, and 256 for FFHQ, Cats, and CARLA, respectively. These channel dimensions are empirically chosen without careful tuning.

Radiance generator ΦΦ\Phi.

Figure I (b) shows the detailed structure of the radiance generator, which consists of a mapping network and a synthesis network. The mapping network is an MLP with three hidden layers of dimension 256. The synthesis network consists of 88 FiLM SIREN blocks of dimension 256, and one FiLM SIREN block of dimension 259 which receives an extra view direction as input.

A.3 More Training Details

During training, we randomly sample latent code z{\bm{z}} from the normal distribution and camera pose θ{\bm{\theta}} from the known or estimated distributions of the training datasets. We jointly learn the manifold predictor M\mathcal{M}, the radiance generator Φ\Phi, and the discriminator DD using the losses described in the main paper. Geometric initialization is applied for the weights of M\mathcal{M} to obtain sphere-like initial isosurfaces. For FFHQ and Cats, we set the sphere center to (0,0,−1.5)(0,0,-1.5) for human face and cat centered in the 3^{3} cube. For CARLA, we set the center to (0,0,0)(0,0,0) to obtain hemispherical manifolds, as shown in Fig. 6 of the main paper. The {li}\{l_{i}\} are set to generate initial isosurfaces evenly positioned across the whole 3D volume. In addition, for FFHQ and Cats, we set the farmost surface to be a fixed plane to represent background. To calculate ray-surface intersections, we uniformly sample 64 points along each ray and calculate the intersections via Eq. (5) in the main paper. The weights of the radiance generator Φ\Phi and the discriminator DD are initialized following .

To enable training at 2562256^{2} resolution, we use PyTorch’s Automatic Mixed Precision (AMP) to reduce memory cost. We also use the mini-batch aggregation strategy similar to to ensure a relatively large batch size (16 for 2562256^{2} resolution and 32 for 1282128^{2} resolution) during training. We train GRAM for 120120K iterations, 8080K iterations, and 7070K iterations on FFHQ, Cats, and CARLA, respectively. Training took 3 to 7 days depending on the dataset and image resolution.

Appendix B More Results

Figure IV, V, and VI show more visual results of GRAM. Our method can generate realistic images with strong multiview consistency. Animation results can be found on the project page.

B.2 Comparisons

Figure VII shows more visual comparisons between GRAM and the previous 3D-aware image generation methods . Our method achieves the best result in terms of image quality and 3D consistency. Animations can be found on the project page.

More comparisons with NeRF-H sampling.

Figure VIII shows the visual comparisons between our manifold sampling strategy and the original NeRF-H sampling strategy. Our method achieves better visual quality with finer details. More importantly, NeRF-H fails to learn reasonable 3D structures of the generated instances with a number of sampling points fewer than 12. It still produces undesired artifacts (e.g., the concave forehead geometry which creates hollow-face illusion), even trained with 48 sampling points. In contrast, our method can learn reasonable 3D geometry with as few as 6 points (surfaces). We hardly observe the concave forehead issue for the generated instances in our cases.

B.3 Failure Cases

We empirically found that for cats, dropping pose regularization sometimes led to unstable training and yielded wrong pose and geometry (which is known as the “hollow-face illusion”; see Fig. II). Training on faces and cars were quite stable no matter pose regularizations were used or not.

Exaggerated parallax artifacts.

When varying camera poses, some contents (e.g. hair fringes) on certain generated subjects could float away from their expected positions, as shown in Fig. III. This is due to that the fixed and limited number of surface manifolds across the whole category cannot provide accurate depth for all structures on every single subject. The problem could be alleviated when using instance-specific surfaces, which we will explore in future works.

B.4 Camera Zoom

As shown in Fig. IX, GRAM can generate reasonable results with camera zoom-in and zoom-out effects. Animations can be found on the project page.

B.5 Latent Space Interpolation

We show the results of latent code interpolation in Fig. X. The continuous semantic changes between adjacent images demonstrate the reasonable latent space learned by GRAM.

B.6 Style Mixing

Figure XI shows the style mixing results between source subjects and target subjects. Similar to , styles in shallower layers (layer 1 to 5) of GRAM mainly control geometry, while styles in deeper layers (layer 6 to 9) control appearance. Note that our method is not trained with the style mixing strategy.