GANcraft: Unsupervised 3D Neural Rendering of Minecraft Worlds

Zekun Hao, Arun Mallya, Serge Belongie, Ming-Yu Liu

Introduction

Imagine a world where every Minecrafter is a 3D painter!

Advances in 2D image-to-image translation have enabled users to paint photorealistic images by drawing simple sketches similar to those created in Microsoft Paint. Despite these innovations, creating a realistic 3D scene remains a painstaking task, out of the reach of most people. It requires years of expertise, professional software, a library of digital assets, and a lot of development time. In contrast, building 3D worlds with blocks, say physical LEGOs or their digital counterpart, is so easy and intuitive that even a toddler can do it. Wouldn’t it be great if we could build a simple 3D world made of blocks representing various materials (like Fig. 1 (insets)), feed it to an algorithm, and receive a realistic looking 3D world featuring tall green trees, ice-capped mountains, and the blue sea (like Fig. 1)? With such a method, we could perform world-to-world translation to convert the worlds of our imagination to reality. Needless to say, such an ability would have many applications, from entertainment and education, to rapid prototyping for artists.

In this paper, we propose GANcraft, a method that produces realistic renderings of semantically-labeled 3D block worlds, such as those from Minecraft (www.minecraft.net). Minecraft, the best-selling video game of all time with over 200 million copies sold and over 120 million monthly users , is a sandbox video game in which a user can explore a procedurally-generated 3D world made up of blocks arranged on a regular grid, while modifying and building structures with blocks. Minecraft provides blocks representing various building materials—grass, dirt, water, sand, snow, etc. Each block is assigned a simple texture, and the game is known for its distinctive cartoonish look. While one might discount Minecraft as a simple game with simple mechanics, Minecraft is, in fact, a very popular 3D content creation tool. Minecrafters have faithfully recreated large cities and famous landmarks including the Eiffel Tower! The block world representations are intuitive to manipulate and this makes it well-suited as the medium for our world-to-world translation task. We focus on generating natural landscapes, which was also studied in several prior work in image-to-image translation .

At first glance, generating a 3D photorealistic world from a semantic block world seems to be a task of translating a sequence of projected 2D segmentation maps of the 3D block world, and is a direct application of image-to-image translation . This approach, however, immediately runs into several serious issues. First, obtaining paired ground truth training data of the 3D block world, segmentation labels, and corresponding real images is extremely costly if not impossible. Second, existing image-to-image translation models do not generate consistent views . Each image is translated independent of the others.

While the recent world-consistent vid2vid work overcomes the issue of view-consistency, it requires paired ground truth 3D training data. Even the most recent neural rendering approaches based on neural radiance fields such as NeRF , NSVF , and NeRF-W , require real images of a scene and associated camera parameters, and are best suited for view interpolation. As there is no paired 3D and ground truth real image data, as summarized in Table 1, none of the existing techniques can be used to solve this new task. This requires us to employ ad hoc adaptations to make our problem setting as similar to these methods’ requirements as possible, e.g. training them on real segmentations instead of Minecraft segmentations.

In the absence of ground truth data, we propose a framework to train our model using pseudo-ground truth photorealistic images for sampled camera views. Our framework uses ideas from image-to-image translation and improves upon work in 3D view synthesis to produce view-consistent photorealistic renderings of input Minecraft worlds as shown in Fig. 1. Although we demonstrate our results using Minecraft, our method works with other 3D block world representations, such as voxels. We chose Minecraft because it is a popular platform available to a wide audience.

The novel task of producing view-consistent photorealistic renderings of user-created 3D semantic block worlds, or world-to-world translation, a 3D extension of image-to-image translation.

A framework for training neural renderers in the absence of ground truth data. This is enabled by using pseudo-ground truth images generated by a pretrained image synthesis model (Section 3.1).

A new neural rendering network architecture trained with adversarial losses (Section 3.2), that extends recent work in 2D and 3D neural rendering to produce state-of-the-art results which can be conditioned on a style image (Section 4).

Related Work

2D image-to-image translation. The GAN framework has enabled multiple methods to successfully map an image from one domain to another with high fidelity, e.g., from input segmentation maps to photorealistic images. This task can be performed in the supervised setting , where example pairs of corresponding images are available, as well as the unsupervised setting , where only two sets of images are available. Methods operating in the supervised setting use stronger losses such as the L1 or perceptual loss , in conjunction with the adversarial loss. As paired data is unavailable in the unsupervised setting, works typically rely on a shared-latent space assumption or cycle-consistency losses . For a comprehensive overview of image-to-image translation methods, please refer to the survey of Liu et al. .

Our problem setting naturally falls into the unsupervised setting as we do not possess real-world images corresponding to the Minecraft 3D world. To facilitate learning a view-consistent mapping, we employ pseudo-ground truths during training, which are predicted by a pretrained supervised image-to-image translation method.

Pseudo-ground truths were first explored in prior work on self-training, or bootstrap learning See https://ruder.io/semi-supervised/ for an overview.. More recently, this technique has been adopted in several unsupervised domain adaptation works . They use a deep learning model trained on the ‘source’ domain to obtain predictions on the new ‘target’ domain, treat these predictions as ground truth labels, or pseudo labels, and finetune the deep learning model on such self-labeled data.

In our problem setting, we have segmentation maps obtained from the Minecraft world but do not possess the corresponding real image. We use SPADE , a conditional GAN model, trained for generating landscape images from input segmentation maps to generate pseudo ground truth images. This yields the pseudo pair: input Minecraft segmentation mask and the corresponding pseudo ground truth image. The pseudo pairs enable us to use stronger supervision such as L1, L2, and perceptual losses in our world-to-world translation framework, resulting in improved output image quality. This idea of using pretrained GAN models for generating training data has also been explored in the very recent works of Pan et al. and Zhang et al. , which use a pretrained StyleGAN as a multi-view data generator to train an inverse graphics model.

3D neural rendering. A number of works have explored combining the strengths of the traditional graphics pipeline, such as 3D-aware projection, with the synthesis capabilities of neural networks to produce view-consistent outputs. By introducing differentiable 3D projection and using trainable layers that operate in the 3D and 2D feature space, several recent methods are able to model the geometry and appearance of 3D scenes from 2D images. Some works have successfully combined neural rendering with adversarial training , thereby removing the constraint of training images having to be posed and from the same scene. However, the under-constrained nature of the problem limited the application of these methods to single objects, synthetic data, or small-scale simple scenes. As shown later in Section 4, we find that adversarial training alone is not enough to produce good results in our setting. This is because our input scenes are larger and more complex, the available training data is highly diverse, and there are considerable gaps in the scene composition and camera pose distribution between the block world and the real images.

Most recently, NeRF demonstrated state-of-the-art results in novel view synthesis by encoding the scene in the weights of a neural network that produces the volume density and view-dependent radiance at every spatial location. The remarkable synthesis ability of NeRF has inspired a large number of follow-up works which have tried to improve the output quality , make it faster to train and evaluate , extend it to deformable objects , account for lighting and compositionality , as well as add generative capabilities .

Most relevant to our work are NSVF , NeRF-W , and GIRAFFE . NSVF reduces the computational cost of NeRF by representing the scene as a set of voxel-bounded implicit fields organized in a sparse voxel octree, which is obtained by pruning an initially dense cuboid made of voxels. NeRF-W learns image-dependent appearance embeddings allowing it to learn from unstructured photo collections, and produce style-conditioned outputs. These works on novel view synthesis learn the geometry and appearance of scenes given ground truth posed images. In our setting, the problem is inverted — we are given coarse voxel geometry and segmentation labels as input, without any corresponding real images.

Similar to NSVF , we assign learnable features to each corner of the voxels to encode geometry and appearance. In contrast, we do not learn the 3D voxel structure of the scene from scratch, but instead implicitly refine the provided coarse input geometry (e.g. shape and opacity of trees represented by blocky voxels) during the course of training. Prior work by Riegler et al. also used a mesh obtained by multi-view stereo as a coarse input geometry. Similar to NeRF-W , we use a style-conditioned network. This allows us to learn consistent geometry while accounting for the view inconsistency of SPADE . Like neural point-based graphics and GIRAFFE , we use differentiable projection to obtain features for image pixels, and then use a CNN to convert the 2D feature grid to an image. Like GIRAFFE , we use an adversarial loss in training. We, however, learn on large, complex scenes and produce higher-resolution outputs (1024×\times2048 original image size in Fig. 1, v/s 64×\times64 or 256×\times256 pixels in GIRAFFE), in which case adversarial loss alone fails to produce good results.

Neural Rendering of Minecraft Worlds

Our goal is to convert a scene represented by semantically-labeled blocks (or voxels), such as the maps from Minecraft, to a photorealistic 3D scene that can be consistently rendered from arbitrary viewpoints (as shown in Fig. 1). In this paper, we focus on landscape scenes that are orders of magnitude larger than single objects or scenes typically used in the training and evaluation of previous neural rendering works. In all of our experiments, we use voxel grids of 512×\times512×\times256 blocks (512×\times512 blocks horizontally, 256 blocks tall vertically). Given that each Minecraft block is considered to have a size of 1 cubic meter , each scene covers an area equivalent to 262,144 m2m^{2} (65 acres, or the size of 32 soccer fields) in real life. At the same time, our model needs to learn details that are much finer than a single block, such as tree leaves, flowers, and grass, that too without supervision. As the input voxels and their labels already define the coarse geometry and semantic arrangement of the scene, it is necessary to respect and incorporate this prior information into the model. We first describe how we overcome the lack of paired training data by using pseudo-ground truths. Then, we present our novel sparse voxel-based neural renderer.

The most straightforward way of training a neural rendering model is to utilize ground truth images with known camera poses. A simple L2 reconstruction loss is sufficient to produce good results in this case . However, in our setting, ground truth real images are simply unavailable for user-generated block worlds from Minecraft.

An alternative route is to train our model in an unpaired, unsupervised fashion like CycleGAN , or MUNIT . This would use an adversarial loss and regularization terms to translate Minecraft segmentations to real images. However, as shown in the ablation studies in Section 4, this setting does not produce good results, for both prior methods, and neural renderers. This can be attributed to the large domain gap between blocky Minecraft and the real world, as well as the label distribution shift between worlds.

To bridge the domain gap between the voxel world and our world, we supplement the training data with pseudo-ground truth that is generated on-the-fly. For each training iteration, we randomly sample camera poses from the upper hemisphere and randomly choose a focal length. We then project the semantic labels of the voxels to the camera view to obtain a 2D semantic segmentation mask. The segmentation mask, as well as a randomly sampled style code, is fed to a pretrained image-to-image translation network, SPADE in our case, to obtain a photorealistic pseudo-ground truth image that has the same semantic layout as the camera view, as shown in the left part of Fig. 2. This enables us to apply reconstruction losses, such as L2, and the perceptual loss , between the pseudo-ground truth and the rendered output from the same camera view, in additional to the adversarial loss. This significantly improves the result.

The generalizability of the SPADE model trained on large-scale datasets, combined with its photorealistic generation capability helps reduce both, the domain gap and the label distribution mismatch. Sample pseudo-pairs are shown in the right part of Fig. 2. While this provides effective supervision, it is not perfect. This can be seen especially in the last two columns in the right part of Fig. 2. The blockiness of Minecraft can produce unrealistic images with sharp geometry. Certain camera poses and style code combinations can also produce images with artifacts. We thus have to be careful to balance reconstruction and adversarial losses to ensure successful training of the neural renderer.

2 Sparse voxel-based volumetric neural renderer

Voxel-bounded neural radiance fields. Let KK be the number of occupied blocks in a Minecraft world, which can also be represented by a sparse voxel grid with KK non-empty voxels given by V={V1,...,VK}\mathcal{V}=\{V_{1},...,V_{K}\}. Each voxel is assigned a semantic label {l1,...,lK}\{l_{1},...,l_{K}\}. We learn a neural radiance field per voxel. The Minecraft world is then represented by the union of voxel-bounded neural radiance fields given by

where FF is the radiance field of the whole scene and FiF_{i} is the radiance field bounded by ViV_{i}. Querying a location in the neural radiance field returns a feature vector (or color in prior work ) and a density value. At the location where a block does not exist, we have the null feature vector 0\mathbf{0} and zero density . To model diversified appearance of the same scene, e.g. day and night, the radiance fields are conditioned on style code z\mathbf{z}.

The voxel-bounded neural radiance field FiF_{i} is given by

where gi(p)g_{i}(\mathbf{p}) is the location code at p\mathbf{p} and li≡l(p)l_{i}\equiv l(\mathbf{p}) is a short-hand for the label of the voxel that p\mathbf{p} belongs to. The multi-layer perceptron (MLP) GθG_{\theta} is used to predict the feature c\mathbf{c}, and volume density σ\sigma at the location p\mathbf{p}. We note that GθG_{\theta} is shared amongst all voxels. Inspired by NeRF-W , c\mathbf{c} additionally depends on the style code, while the density σ\sigma does not. To obtain the location code, we first assign a learnable feature vector to each of the eight vertices of a voxel ViV_{i}. The location code at p\mathbf{p}, gi(p)g_{i}(\mathbf{p}), is then derived through trilinear interpolation. Here, we assume that each voxel has a shape of 1×\times1×\times1, and the coordinate axes are aligned to the voxel grid axes. Vertices and their feature vectors are shared for adjacent voxels. This allows for a smooth transition of features when crossing the voxel boundaries, preventing discontinuities in the output. We compute Fourier features from gi(p)g_{i}(\mathbf{p}), similar to NSVF , and also append the voxel class label. Our method can be interpreted as a generalization of NSVF to use a style and semantic label conditioning.

Volumetric rendering. Here, we describe how a scene represented by the above-mentioned neural radiance fields and sky dome can be converted to 2D feature maps via volumetric rendering. Under a perspective camera model, each pixel in the image corresponds to a camera ray r(t)=o+tv\mathbf{r}(t)=\mathbf{o}+t\mathbf{v}, originating from the center of projection o\mathbf{o} and advances in direction v\mathbf{v}. The ray travels through the radiance field while accumulating features and transmittance,

C(r,z)C(\mathbf{r},\mathbf{z}) denotes the accumulated feature of ray r\mathbf{r}, and T(t)T(t) denotes the accumulated transmittance when the ray travels a distance of tt. As the radiance field is bounded by a finite number of voxels, the ray will eventually exit the voxels and hit the sky dome. We thus consider the sky dome as the last data point on the ray, which is totally opaque. This is realized by the last term in Eq. 3.2. The above integral can be approximated using discrete samples and the quadrature rule, a technique popularized by NeRF . Please refer to NeRF or our supplementary for the full equations.

We use the stratified sampling technique from NSVF to randomly sample valid (voxel bounded) points along the ray. To improve efficiency, we truncate the ray so that it will stop after a certain accumulated distance through the valid region is reached. We regularize the truncated rays to encourage their accumulated opacities to saturate before reaching the maximum distance. We adopt a modified Bresenham method for sampling valid points, which has a very low complexity of O(N)O(N), where NN is the longest dimension of the voxel grid. Details are in the supplementary.

Hybrid neural rendering architecture. Prior works directly produce images by accumulating colors using the volumetric rendering scheme described above instead of accumulating features. Unlike them, we divide rendering into two parts: 1) We perform volumetric rendering with an MLP to produce a feature vector per pixel instead of an RGB image, and 2) We employ a CNN to convert the per-pixel feature map to a final RGB image of the same size. The overall framework is shown in Fig. 3. We perform activation modulation conditioned on the input style code for both the MLP and CNN. The individual networks are described in detail in the supplementary.

Apart from improving the output image quality as shown in Section 4, this two-stage design also helps reduce the computational and memory footprint of rendering. The MLP modeling the 3D radiance field is evaluated on a per-sample basis, while the image-space CNN is only evaluated after multiple samples along a ray are merged into a single pixel. The number of samples to the MLP scales linearly with the output height, width, and number of points sampled per ray (24 in our case), while the size of the feature map only depends on output height and width. However, unlike MLPs that operate pre-blending, the image-space CNN is not intrinsically view consistent. We thus use a shallow CNN with a receptive field of only 9×\times9 pixels to constrain its scope to local manipulations. A similar idea of combining volumetric rendering and image-space rendering has been used in GIRAFFE . Unlike us, they also rely on the CNN to upsample a low-resolution 16×\times16 feature map.

Losses and regularizers. We train our model with both reconstruction and adversarial losses. The reconstruction loss is applied between the predicted images and the corresponding pseudo-ground truth images. We use a combination of the perceptual , L1, and L2 losses. For the GAN loss, we treat the predicted images as ‘fake’, and both the real images, and the pseudo-ground truth images as ‘real’. We use a discriminator conditioned on the semantic segmentation maps, based on Liu et al. and Schönfeld et al. . We use the hinge loss as the GAN training objective. Following previous works on multimodal image synthesis , we also include a style encoder which predicts the posterior distribution of the style code given a pseudo-ground truth image. The reconstruction loss, in conjunction with the style encoder, makes it possible to control the appearance of the output image with a style image.

Experiments

The previous section described how we obtain pseudo-ground truths in the absence of paired Minecraft–real training data, and the architecture of our neural renderer. Here, we validate our framework by comparing with prior work on multiple diverse large Minecraft worlds.

Datasets. We collected a dataset of ∼\sim1M landscape images from the internet with a minimum side of at least 512 pixels. For each image, we obtained 182-class COCO-Stuff segmentation labels by using DeepLabV2 . This formed our training set of paired real segmentation maps and images. We set aside 5000 images as a test set. We generated 5 different Minecraft worlds of 512×\times512×\times256 blocks each. We sampled worlds with various compositions of water, sand, forests, and snow, to show that our method works correctly under significant label distribution shifts.

Baselines. We compare against the following, which are representative methods under different data availability regimes.

MUNIT . This is an image-to-image translation method trainable in the unpaired, or unsupervised setting. Unlike CycleGAN and UNIT , MUNIT can learn multimodal translations. We learn to translate Minecraft segmentation maps to real images.

SPADE . This is an image-to-image translation method that is trained in the paired ground truth, or supervised setting. We train this by translating real segmentation maps to corresponding images and test it on Minecraft segmentations.

wc-vid2vid . Unlike the above two methods, this can generate a sequence of images that are view-consistent. wc-vid2vid projects the pixels from previous frames to the next frame to generate a guidance map. This serves as a form of memory of the previously generated frames. This method also requires paired ground truth data, as well as the 3D point clouds for each output frame. We train this to translate real segmentation maps to real images, while using the block world voxel surfaces as the 3D geometry.

NSVF-W . We combine the strengths of two recent works on neural rendering, NSVF , and NeRF-W , to create a strong baseline. NSVF represents the world as voxel-bounded radiance fields, and can be modified to accept an input voxel world, just like our method. NeRF-W is able to learn from unstructured image collections with variations in color, lighting, and occlusions, making it well-suited for learning from our pseudo-ground truths. Combining the style-conditioned MLP generator from NeRF-W with the voxel-based input representation of NSVF, we obtain NSVF-W. This resembles the neural renderer used by us, with the omission of the image-space CNN. As these methods also require paired ground truth, we train NSVF-W using pseudo-ground truths generated by the pretrained SPADE model.

MUNIT, SPADE, and wc-vid2vid use perceptual and adversarial losses during training, while NSVF, NeRF-W, and thus NSVF-W use the L2 loss. Details are in the supplementary.

Implementation details. We train our model at an output resolution of 256×\times256 pixels. Each model is trained on 8 NVIDIA V100 GPUs with 32GB of memory each. This enables us to use a batch size of 8 with 24 points sampled per camera ray. Each model is trained for 250k iterations, which takes approximately 4 days. All baselines are also trained for an equivalent amount of time. Additional details are available in the supplementary.

Evaluation metrics. We use both quantitative and qualitative metrics to measure the quality of our outputs.

Fréchet Inception Distance (FID) and Kernel Inception Distance (KID). We use FID and KID to measure the distance between the distributions of the generated and real images, using Inception-v3 . We generate 1000 images for each of the 5 worlds from arbitrarily sampled camera view points using different style codes, for a total of 5000 images. We then generate outputs from each method for the same pair of view points and style code for a fair comparison. We use a held-out set of 5000 real landscape images to compute the metrics. For both metrics, a lower value indicates better image quality.

Human preference score. Using Amazon Mechanical Turk (AMT), we perform a subjective visual test to gauge the relative quality of generated videos with top-qualified turkers. We ask turkers to choose 1) the more temporally consistent video, and 2) the video with overall better realism. For each of the two questions, a turker is shown two videos synthesized by two different methods and asked to choose the superior one according to the criteria. We generate 64 videos per world, total of 320 per method, and each comparison is evaluated by 3 workers.

Main results. Fig. 4 shows output videos generated by different methods. Each row is a unique world, generated using the same style-conditioning image for all methods. We can observe that our outputs are more realistic and view-consistent when compared to baselines. MUNIT and SPADE demonstrate a lot of flickering as they generate one image at a time, without any memory of past outputs. Further, MUNIT also fails to learn the correct mapping of segmentation labels to textures as it does not use paired supervision. While wc-vid2vid is more view-consistent, it fails for large motions as it incrementally inpaints newly explored parts of the world. NSVF-W and GANcraft are both inherently view-consistent due to their use of volumetric rendering. However, due to the lack of a CNN renderer and the use of the L2 loss, NSVF-W produces dull and unrealistic outputs with artifacts. The use of an adversarial loss is key to ensuring vivid and realistic results, and this is further reinforced in the ablations presented below. Our method is also capable of generating higher resolution outputs as shown in Fig. 1, by sampling more rays.

We sample novel camera views from each world and compute the FID and KID against a set of held-out real images. As seen in Table 2, our method achieves FID and KID close to that of SPADE, which is a very strong image-to-image translation method, while beating other baselines. Note that wc-vid2vid uses SPADE to generate the output for first camera view in a sequence and is thus ignored in this comparison. Further, as summarized in Table 3, users consistently preferred our method and chose its predictions as the more view-consistent and realistic videos. More high-resolution results and comparisons as well as some failure cases are available in the supplementary.

Ablations. We train ablated versions of our full model on one Minecraft world due to computational constraints. We show example outputs from them in Fig. 5. Using no pseudo-ground truth at all and training with just the GAN loss produces unrealistic outputs, similar to MUNIT . Directly producing images from volumetric rendering, without using a CNN, results in a lack of fine detail. Compared to the full model, skipping the GAN loss on real images produces duller images, and skipping the GAN loss altogether produces duller, blurrier images resembling NSVF-W outputs. Qualitative analysis is available in the supplementary.

Discussion

We introduced the novel task of world-to-world translation and proposed GANcraft, a method to convert block worlds to realistic-looking worlds. We showed that pseudo-ground truths generated by a 2D image-to-image translation network provide effective means of supervision in the absence of real paired data. Our hybrid neural renderer trained with both real landscape images and pseudo-ground truths, and adversarial losses, outperformed strong baselines.

There still remain a few exciting avenues for improvements, including learning a smoother geometry in spite of coarse input geometry, and using non-voxel inputs such as meshes and point clouds. While our method is currently trained on a per-world basis, we hope future work can enable feed-forward generation on novel worlds.

Acknowledgements. We thank Rev Lebaredian for challenging us to work on this interesting problem. We thank Jan Kautz, Sanja Fidler, Ting-Chun Wang, Xun Huang, Xihui Liu, Guandao Yang, and Eric Haines for their feedback during the course of developing the method.

References

A Supplementary video

Our project website is available at https://nvlabs.github.io/GANcraft/. This includes an overview of the method as well as additional results.

We also provide a video, including more visual results and discussion of our work. Specifically, it contains:

Additional high-resolution video results rendered at 1024×\times2048 pixels and 30 frames per second

Additional comparisons with baseline methods

Please make sure to check it out at https://www.youtube.com/watch?v=1Hky092CGFQ.

B Method details

Here, we provide additional details of our approach.

The integral in Equation 2 of the main paper can be approximated with discrete samples via quadrature . Assume that we sample N+1N+1 points at t1,...,tN+1t_{1},...,t_{N+1} along a camera ray r(t)=o+tv\mathbf{r}(t)=\mathbf{o}+t\mathbf{v}. We define

B.2 Point sampling algorithm

In this section, we describe the method we use to efficiently sample points from the sparse voxel grid along a camera ray. Instead of relying on rejection sampling (as in Liu et al. ) to remove points that have not landed inside any voxel, we first traverse the voxel grid along the ray to obtain the entrance and exit points of each valid voxel that the ray has gone through, and then sample points only on the segments that are inside voxels.

For voxel grid traversal, We implement a 3D version of Bresenham’s line algorithm , which has a very low computational cost of O(N)O(N), where NN is the longest dimension of the voxel grid. Its working principle is as follows: Starting from the voxel position where the camera resides, for each step, we traverse to the next voxel which is adjacent to the current voxel by the face which the ray exits from.

B.3 Network Architecture

GANcraft contains 6 trainable neural networks. Here are their descriptions and their respective network architectures:

Per-sample MLP. This is the MLP for representing the implicit radiance field, in conjunction with the voxel features. The network architecture is illustrated in Fig. 9. We condition the output feature on the style code via weight modulation . The detailed implementation of weight modulation is shown in Fig. 10.

Neural sky dome. The sky is modeled with an MLP (Fig. 9) which takes ray direction (represented as a normalized 3D vector) input and produce the color feature for that ray. The network is also conditional on the style feature.

Image space renderer. This is a CNN for converting feature map to RGB image (Fig. 9). As discussed in the main paper, we use very small kernel sizes to reduce the receptive field in order to encourage view consistency. The network is conditional on the style feature.

Style network. Following StyleGAN2 , we use an MLP that is shared across all the style conditioning layers to convert the input style code to an intermediate style feature. Its architecture is shown in Fig. 9.

Style encoder. The style encoder is a CNN that predicts the style code given an image. In conjunction with pseudo ground truth and reconstruction loss, this allows GANcraft to produce images that follows the style of a given image. Our style encoder is taken from SPADE , which is a 6-layer CNN followed by a linear layer and VAE reparameterization. Please refer to the original paper for the details.

Label-conditional discriminator. The discriminator we use is based on feature pyramid semantics-embedding (FPSE) discriminator . Its construction is shown in Fig. 11. Compared to the patch discriminator used in SPADE , the FPSE discriminator is more robust to the distribution mismatch in the label map domain. A patch discriminator which takes the concatenated image and label map as input sometimes lead to training collapse almost immediately after the training starts.

B.4 Label Translation

There is significant difference between Minecraft voxel labels, which we use as the starting point of GANcraft, and COCO-Stuff labels, which is the format the pretrained DeepLabV2 model produces and the pretrained SPADE model accepts. In Minecraft Java edition, there are 680 labels in total, mostly describing raw materials (dirt, sand, log, water, etc.) useful for building objects. While in COCO-Stuff, there are 182 higher level labels of common objects such as mountain, tree, river, and sea. Due to the drastic difference in the level of abstraction, it is very difficult to find a one-to-one mapping between Minecraft label and COCO-Stuff label. For example, the water material in Minecraft can be mapped to either sea or river label in COCO-Stuff; the tree material consists of both log and leaf label in COCO-Stuff. We solve the labeling difference in two ways. For the label-conditional discriminator, we introduce a new set of 12 classes with high level of abstraction: ignore, sky, tree, dirt, flower, grass, gravel, water, rock, stone, sand, and snow. We then classify every Minecraft and COCO-Stuff label into one of the 12 classes, and use the translated semantic segmentation mask as the conditional input to the discriminator. For generating pseudo-ground truth, however, we will have to convert Minecraft labels to COCO-Stuff labels in order to be recognized by the pretrained SPADE generator. We achieve this by first translating the Minecraft labels to one of the 12 labels, and then map them to COCO-Stuff labels randomly, with equal chance across all the candidate labels. Note that we use the same mapping scheme within a segmentation map. We are able to obtain good result from such a simple measure, as the style encoder is able to explain away the randomness in the mapping.

B.5 Voxel Preprocessing

Minecraft voxel world has a sea level of 62, below which most of the voxels are not visible from above. It will be a waste of memory if we still assign voxel features to those invisible voxels. Thus we preprocess the voxel by removing the interior voxels, leaving a 4 voxel thick thin shell. This operation reduces the occupancy of a typical voxel world from 28% to 3%. The effect of preprocessing can be seen at the borders of the voxel worlds in Fig. 12. Note that the preprocessing step is not only useful for Minecraft world. It is applicable to any types of voxel grids.

C Experiment details

We use 5 different Minecraft worlds for our experiments. An overview of these worlds is shown in Fig. 12. As can be seen, the label distribution of each specific world is very different from that of a collection of real images, e.g. the first world is >>50% sand, and the second is >>90% water. Our method works for all these worlds despite the domain gap, indicating the robustness of our framework.

C.2 GANcraft settings

During training, we generate images at a resolution of 256×\times256. We sample 24 points along each ray, and truncate the rays to a maximum distance of 3 (distance traveled outside voxels doesn’t count). We use a learning rate of 1e-4 for the generator networks, and 4e-4 for the discriminator. For voxel features, we use a higher learning rate of 5e-3. We use a combination of GAN loss, L2 loss, L1 loss and perceptual loss, with their weights being 1.0, 10.0, 1.0 and 10.0, respectively. For regularization terms, we use a weight of 0.5 for the opacity regularization, and a weight of 0.05 for the KL divergence needed by the style encoder. We also clip the per-sample feature c\mathbf{c} to a range of $$ before blending to reduce the ambiguity between the opacity and the scale of feature. For random camera pose sampling, we sample two points that are slightly above ground, and use one of the as the camera location and the other one as the point that the camera looks at. We reject any camera pose that produces a depth map with a mean depth below 2 or that produces a segmentation mask with label entropy below 0.75. This guarantees that the segmentation mask along can provide enough scene geometry hint to the SPADE generator for generating a pseudo-ground truth that corresponds well to the actual scene geometry.

During evaluation, we increase the sample count to 32 points per ray. On an NVIDIA Titan V, this takes approximately 10 seconds to render a 1024×\times2048 frame.

C.3 Baseline settings

For fair comparison, the settings used in the NSVF-W baseline largely resembles GANcraft except for the following differences:

Only L2 loss and KL divergence is used during training.

The weight for KL divergence is reduced to 0.01 to avoid handicapping the style encoder too much in the absence of other reconstruction losses.

The image space CNN renderer is removed, and the per-sample MLP directly produces an RGB radiance (clipped by a sigmoid function) instead of a feature.

D Additional results

Here, we present quantitative results for the ablated versions of our full model. Sample outputs from these ablations were shown in Fig. 5 of the main paper. We trained all ablations on one world only, due to computational constraints (each model takes 4 days on 8 NVIDIA V100 GPUs).

The results of automated metric evaluation as shown in Table 4. We computed the FID and KID values with 2000 images generated from random camera poses and 5000 held-out real images. As expected, all ablated versions obtain higher FID and KID scores indicating worse quality. An exception is the model trained without any pseudo-ground truth images, i.e. trained with GAN loss between outputs and real images only. Surprisingly, it obtains a lower FID and KID than our full model. However, when we visually inspect the outputs, shown in Fig. 14, it is clear that the model fails to learn a meaningful mapping from Minecraft segmentations to real images. The model seems to have learned to produce unrealistic images that optimize the metrics due to training with the GAN loss. However, similar to MUNIT , the outputs are both unrealistic and incorrectly map Minecraft segmentation labels to real images.

We observed that our method can fail in two ways — either producing blocky outputs or producing unrealistic outputs. In the input block world, all objects and regions are made of blocks. Due to this coarse geometry, the method is sometimes unable to learn realistic geometries in the translated world. As a result, boundaries can often appear jagged, as shown in Fig. 14. Further, certain combinations of worlds and style-conditioning images can produce unrealistic outputs as shown in Fig. 15. For example, a forest world paired with a conditioning image of a red sunset can produce unrealistic, or overly dark outputs. As the style encoder is trained exclusively with pseudo-ground truth images that have the same label distribution as the rendered Minecraft images, it has never encountered such combinations.