Shape Inpainting using 3D Generative Adversarial Network and Recurrent Convolutional Networks

Weiyue Wang, Qiangui Huang, Suya You, Chao Yang, Ulrich Neumann

Introduction

Data collected by 3D sensors (e.g. LiDAR, Kinect) are often impacted by occlusion, sensor noise, and illumination, leading to incomplete and noisy 3D models. For example, a building scan occluded by a tree leads to a hole or gap in the 3D building model. However, a human can comprehend and describe the geometry of the complete building based on the corrupted 3D model. Our 3D inpainting method attempts to mimic this ability to reconstruct complete 3D models from incomplete data.

Convolutional Neural Network (CNN) based methods yield impressive results for 2D image generation and image inpainting. Generating and inpainting 3D models is a new and more challenging problem due to its higher dimensionality. The availability of large 3D CAD datasets and CNNs for voxel (spatial occupancy) models enabled progress in learning 3D representation, shape generation and completion. Despite their encouraging results, artifacts still persists in their generated shapes. Moreover, their methods are all based on 3D CNN, which impedes their ability to handle higher resolution data due to limited GPU memory.

In this paper, a new system for 3D object inpainting is introduced to overcome the aforementioned limitations. Given a 3D object with holes, we aim to (1) fill the missing or damaged portions and reconstruct a complete 3D structure, and (2) further predict high-resolution shapes with fine-grained details. We propose a hybrid network structure based on 3D CNN that leverages the generalization power of a Generative Adversarial model and the memory efficiency of Recurrent Neural Network (RNN) to handle 3D data sequentially. The framework is illustrated in Figure 1.

More specifically, a 3D Encoder-Decoder Generative Adversarial Network (3D-ED-GAN) is firstly proposed to generalize geometric structures and map corrupted scans to complete shapes in low resolution. Like a variational autoencoder (VAE) , 3D-ED-GAN utilizes an encoder to map voxelized 3D objects into a probabilistic latent space, and a Generative Adversarial Network (GAN) to help the decoder predict the complete volumetric objects from the latent feature representation. We train this network by minimizing both contextual loss and an adversarial loss. Using GAN, we can not only preserve contextual consistency of the input data, but also inherit information from data distribution.

Secondly, a Long-term Recurrent Convolutional Network (LRCN) is further introduced to obtain local geometric details and produce much higher resolution volumes. 3D CNN requires much more GPU memory than 2D CNN, which impedes volumetric network analysis of high-resolution 3D data. To overcome this limitation, we model the 3D objects as sequences of 2D slices. By utilizing the long-range learning capability from a series of conditional distributions of RNN, our LRCN is a Long Short-term Memory Network (LSTM) where each cell has a CNN encoder and a fully-convolutional decoder. The outputs of 3D-ED-GAN are sliced into 2D images, which are then fed into the LRCN, which gives us a sequence of high-resolution images.

Our hybrid network is an end-to-end trainable network which takes corrupted low resolution 3D structures and outputs complete and high-resolution volumes. We evaluate the proposed method qualitatively and quantitatively on both synthesized and real 3D scans in challenging scenarios. To further evaluate the ability of our model to capture shape features during 3D inpainting, we test our network for 3D object classification tasks and further explore the encoded latent vector to demonstrate that this embedded representation contains abundant semantic shape information.

The main contributions of this paper are:

a 3D Encoder-Decoder Generative Adversarial Convolutional Neural Network that inpaints holes in 3D models, which can further help 3D shape feature learning and help object recognition.

a Long-term Recurrent Convolutional Network that produces high resolution 3D volumes with fine-grained details by modeling volumetric data as sequences of 2D images to overcome GPU memory limitation.

an end-to-end network that combines the above two ideas and completes corrupted 3D models, while also producing high resolution volumes.

Related Work

Generative Adversarial Network (GAN) generates images by jointly training a generator and a discriminator. Following this pioneering work, a series of GAN models were developed for image generation tasks. Pathak et al. developed a context encoder in an unsupervised learning algorithm for image inpainting. Generative adversarial loss in their autoencoder-like network architecture achieves impressive performance for image inpainting.

With the introduction of 3D CAD model datasets , recent developments in 3D generative models use data-driven methods to synthesize new objects. CNN is used to learn embedded object representations. Bansal et al. introduced a skip-network model to retrieve 3D models for objects depicted in 2D images of CAD data. Choy et al. used a recurrent network with multi-view images for 3D model reconstruction. Girdhar proposed a TL-embedding network to learn an embedding space that can be generative in 3D and predicative from 2D rendered images. Wu et al. showed that the learned latent vector by 3D GAN can generate high-quality 3D objects and improve object recognition accuracy as a shape descriptor. They also added an image encoder to 3D GAN to generate 3D model from 2D images. Yan et al. formulated an encoder-decoder network with a loss by perspective transformation for predicting 3D models from a single-view 2D image.

2 3D Completion

Recent advances in deep learning have shown promising results in 3D completion. Wu et al. built a generative model with Convolutional Deep Belief Network by learning a probabilistic distribution from 3D volumes for shape completion from 2.5D depth maps. Sharma introduced a fully convolutional autoencoder that learns volumetric representation from noisy data by estimating voxel occupancy grids. This is the state of the art for 3D volumetric occupancy grid inpainting to the best of our knowledge. An important benefit of our 3D-ED-GAN over theirs is that we introduce GAN to inherit information from the data distribution. Dai et al. introduced a 3D-Encoder-Predictor Network to predict and fill missing data for 3D distance field and proposed a 3D synthesis procedure to obtain high-resolution objects. This is the state-of-the-art method for high-resolution object completion. However, instead of an end-to-end network, their shape synthesis procedure requires iterating every sample from the dataset. Since we are using occupancy grids to represent 3D shapes, we do not compare with them in our experiment. Song et al. synthesized a 3D scenes dataset and proposed a semantic scene completion network to produce complete 3D volumes and semantic labels for a scene from single-view depth map. Despite the encouraging results of the works mentioned above, these methods are mostly based on 3D CNN, which requires much more GPU memory than 2D convolution and impedes handling high-resolution data.

3 Recurrent Neural Networks

RNNs have been shown to excel at hard sequence problems ranging from natural language translation , to video analysis . By implicitly conditioning on all previous variables and preserving long-range contextual dependencies, RNNs are also suitable for dense prediction tasks such as semantic segmentation , and image completion . Donahue et al. applied 2D CNN and LSTM on 3D data (video) and developed a recurrent convolutional architecture for video recognition. Oord et al. presented a deep network that sequentially predicts the pixels in an image along two spatial dimensions. Choy et al. used a recurrent network and a CNN to reconstruct 3D models from a sequence of multi-view images. Followed by these pioneer works, we apply RNN on 3D object data and predict dense volume as sequences of 2D pixels.

Methods

The goal of this paper is to take a corrupted 3D object in low resolution as input and produce a complete high-resolution model as output. The 3D model is represented as volumetric occupancy grids. To fill the missing data requires an approach that can make conceivable predictions from data distributions as well as preserve structural context of the imperfect input.

We introduce an 3D Encoder-Decoder CNN by extending a 3D Generative Adversarial Network , namely 3D Encoder-Decoder Generative Adversarial Network (3D-ED-GAN), to accomplish the 3D inpainting task. Since 3D CNN is memory consuming and applying 3D-ED-GAN on a high-resolution volume is improbable, we only use 3D-ED-GAN to operate low-resolution voxels (say 32332^{3}). Then we treat 3D volume output of 3D-ED-GAN as a sequence of 2D images and reconstruct the object slice by slice. A Long-term Recurrent Convolutional Network (LRCN) based on LSTM is proposed to recover fine-grained details and produce high-resolution results. LRCN functions as an upsampling network while completing details by learning from the dataset.

We now describe our network structure of 3D-ED-GAN and LRCN respectively and the details of the training procedure.

The Generative Adversarial Network (GAN) consists of a generator GG that maps a noise distribution Z\mathbf{Z} to the data space X\mathbf{X}, and a discriminator DD that classifies whether the generated sample is real or fake. GG and DD are both deep networks that are learned jointly. DD distinguishes real samples from synthetic data. GG tries to generate ”real” samples to confuse DD. Concretely, the objective of GAN is to achieve the following optimization:

where pdatap_{data} is data distribution and pzp_{z} is noise distribution.

3D-ED-GAN extends the general GAN framework by modeling the generator GG as a fully-convolutional Encoder-Decoder network, where the encoder maps input data into a latent vector z\mathbf{z}. Then the decoder maps z\mathbf{z} to a cube. The 3D-ED-GAN consists of three components: an encoder, a decoder and a discriminator. Figure 2 depicts the algorithmic architecture of 3D-ED-GAN.

The encoder takes a corrupted 3D volume x′\mathbf{x}^{\prime} of size dl3{d_{l}}^{3} (say dl=32d_{l}=32) as input. It consists of three 3D convolutional layers with kernel size 5 and stride 2, connected via batch normalization (BN) and ReLU layers. The last convolutional layer is reshaped into a vector zz, which is the latent feature representation. There is no fully-connection (fc) layers. The noise vector in GAN is replaced with zz. Therefore, the 3D-ED-GAN network conditions zz using the 3D encoder. We show that this latent vector carries informative features for supervised tasks in Section 4.2.

The decoder has the same architecture as GG in GAN, which maps the latent vector zz to a 3D voxel of size dl3{d_{l}}^{3}. It has three volumetric full-convolution (also known as deconvolution) layers of kernel size 5 and strides 2 respectively, with BN and ReLU layers added in between. A tanh⁡\tanh activation layer is added after the last layer. The Encoder-Decoder network is a fully-convolutional neural network without linear or pooling layers.

The discriminator has the same architecture as the encoder with an fc layer and a sigmoid layer at the end.

The generator GG in 3D-ED-GAN is modeled by the Encoder-Decoder network. This can be viewed as a conditional GAN, in which the latent distribution is conditioned on given context data. Therefore, the loss function can been derived by reformulating the objective function in Equation 1

where Fed(⋅):X→XF_{ed}(\cdot):\mathbf{X}\rightarrow\mathbf{X} is the Encoder-Decoder network, and x′\mathbf{x}^{\prime} is the corrupted model of complete volume x\mathbf{x}.

Similar to , we add an object reconstruction Cross-Entropy loss, LreconL_{recon}, defined by

where N=dl3N={d_{l}}^{3}, xix_{i} represents for the iith voxel of the complete volume x\mathbf{x} and Fed(x′)iF_{ed}(\mathbf{x}^{\prime})_{i} is the iith voxel of the generated volume. In this way, the output of the Encoder-Decoder network Fed(x′)F_{ed}(\mathbf{x}^{\prime}) is the probability of a voxel being filled.

The overall loss function for 3D-ED-GAN is

where α1\alpha_{1} and α2\alpha_{2} are weight parameters.

The loss function can effectively infer the structures of missing regions to produce conceivable reconstructions from the data distribution. Inpainting requires maintaining coherence of given context and producing plausible information according to the data distribution. 3D-ED-GAN has the capability of capturing the correlation between a latent space and the data distribution, thus producing appropriate plausible hypothesis.

2 Long-term Recurrent Convolutional Network (LRCN) Model

3D CNN consumes much more GPU memory than 2D CNN. Extending 3D-ED-GAN by adding 3D convolution layers to produce high resolution output is improbable due to memory limitation. We take advantage of the capability of RNN to handle long-term sequential dependencies and treat the 3D object volume as slices of 2D images. The network is required to map a volume with dimension dl3{d_{l}}^{3} to a volume with dimension dh3{d_{h}}^{3} (we have dl=32,dh=128d_{l}=32,d_{h}=128). For a sequence-to-sequence problem with different input and output dimensions, we integrate an encoder-decoder pair to the LSTM cell inspired by the video processing work . Our LRCN model combines an LSTM, a 3D CNN, and 2D deep fully-convolutional network. It works by passing each 2D slice with its neighboring slices through a 3D CNN to produce a fixed-length vector representation as input to LSTM. The output vector of LSTM is passed through a 2D fully-convolutional decoder network and mapped to a high-resolution image. A sequence of high-resolution 2D images formulate the output 3D object volume. Figure 3 depicts our LRCN architecture.

In order to obtain the maximal amount of contextual data from each 3D object volume, we would like to maximize the number of nonempty slices for the volume. So given a 3D object volume of dimension dl3{d_{l}}^{3}, we firstly use principle component analysis (PCA) to align the 3D object and denote the aligned volume as I\mathbf{I} and its first principle component as direction l→\overrightarrow{l}In our experiment implementation, we use PCA to align the corrupted objects instead of the output of 3D-ED-GAN.. Then I\mathbf{I} is treated as a sequence of dl×dld_{l}\times d_{l} 2D images along l→\overrightarrow{l}, denoted as {I1,I2,...,Idl}\{I_{1},I_{2},...,I_{d_{l}}\}. Since the output of LRCN is a sequence with length dhd_{h}, the input sequence length should also be dhd_{h}. As illustrated in Figure 3, for each step, a slice with its 4 neighboring slices (so 5 slices total) is formed into a thin volume and fed into the network, say for step tt. And slices with negative indices, or indices beyond dld_{l}, are 0-padded. The input of the 3D CNN is then vt′={Itdh/dl−2,Itdh/dl−1,Itdh/dl,Itdh/dl+1,Itdh/dl+2}\mathbf{v}_{t}^{\prime}=\{I_{{\frac{t}{d_{h}/d_{l}}}-2},I_{{\frac{t}{d_{h}/d_{l}}}-1},I_{\frac{t}{d_{h}/d_{l}}},I_{{\frac{t}{d_{h}/d_{l}}}+1},I_{{\frac{t}{d_{h}/d_{l}}}+2}\}.

As illustrated in Figure 3, the 3D CNN encoder takes a dl×dl×cd_{l}\times d_{l}\times c volume as input, where cc represents number of slices (we have c=5c=5). At step tt, the 3D CNN transforms cc slices of 2D images vt′\mathbf{v}_{t}^{\prime} into a 200D200D vector vtv_{t}. The 3D CNN encoder has the same structure with the 3D encoder in 3D-ED-GAN with an fc layer at the end. After the 3D CNN, the recurrent model LSTM takes over. We use the LSTM cell as described in : Given input vtv_{t}, the LSTM updates at timestep tt are:

where σ\sigma is the logistic sigmoid function, i,f,o,ci,f,o,c are respectively inputgate,forgetgate,outputgate,cellgateinputgate,forgetgate,outputgate,cellgate, Wvi,vf,vc,vo,hi,hf,hc,ho,ci,co,cfW_{vi,vf,vc,vo,hi,hf,hc,ho,ci,co,cf} and bi,f,c,ob_{i,f,c,o} are parameters.

The output vector of LSTM oto_{t} is further going through a 2D fully-convolutional neural network to generate a dh×dhd_{h}\times d_{h} image. It has two fully-convolutional layers of kernel size 5 and stride 2, with BN and ReLU in between followed by a tanh⁡\tanh layer at the end.

We experimented with both l1l_{1} and l2l_{2} losses and found the l1l_{1} loss obtains higher-quality results. In this way, the l1l_{1} loss is adopted to train our LRCN, denoted as LLRCNL_{LRCN}.

The overall loss to jointly train the hybrid network (combination of 3D-ED-GAN and LRCN) is

where α3\alpha_{3} and α4\alpha_{4} are weight parameters.

Although the LRCN contains a 3D CNN encoder, the thin input slices makes the network sufficiently small compared to a regular volumetric CNN. By taking advantage of RNN’s ability to manipulate sequential data and long-range dependencies, our memory efficient network is able to produce high-resolution completion result.

3 Training the hybrid network

Training our 3D-ED-GAN and LRCN both jointly and from scratch is a challenging task. Therefore, we propose a three-phase training procedure.

In the first stage, 3D-ED-GAN is trained independently with corrupted 3D input and complete output references in low resolution. Since the discriminator learns much faster than the generator, we first train the Encoder-Decoder network independently without discriminator (with only reconstruction loss). The learning rate is fixed to 10−510^{-5}, and 2020 epochs are trained. Then we jointly train the discriminator and the Encoder-Decoder as in for 100100 epochs. We set the learning rate of the Encoder-Decoder to 10−410^{-4}, and DD to 10−610^{-6}. Then α1\alpha_{1} and α2\alpha_{2} are set to 0.001 and 0.999 respectively. For each batch, we only update the discriminator if its accuracy in the last batch is not higher than 80% as in . ADAM optimization is employed with β=0.5\beta=0.5 and a batch size of 44.

In the second stage, LRCN is trained independently with perfect 3D input in low resolution and high-resolution output references for 100100 epochs. We use a learning rate of 10−410^{-4} and a batch size of 44. In this stage, LRCN works as an upsampling network capable of predicting fine-grained details from trained data distributions.

In the final training phase, we jointly finetune the hybrid network on the pre-trained networks in the first and second stages with loss defined as in Equation 6. The learning rate of the discriminator is 10−710^{-7} and the learning rate of the remaining network is set to be 10−610^{-6} with batch size 11. Then α3\alpha_{3} and α4\alpha_{4} are both set to 0.5. We observe that most of the parameter updates happen in LRCN. The input of LRCN in this stage is imperfect and the output reference is still complete high-resolution model, which indicates that LRCN works as a denoising network while maintaining its power of upsampling and preserving details.

For convenience, we use the aforementioned PCA method to align all models before training instead of aligning the predictions of 3D-ED-GAN.

Experiments

Our network architecture is implemented using the deep learning library Tensorflow . We extensively test and evaluate our method using various datasets canonical to 3D inpainting and feature learning.

We split each category in the ShapeNet dataset to mutually-excluded 80 training points and 20 testing points. Our network is trained on the training points as stated in Section 3.3. We train separate networks for seven major categories (chairs, sofas, tables, boats, airplanes, lamps, dressers, and cars) without fine-tuning on any existing models. 3D meshes are voxelized into 32332^{3} grids for low-resolution input and 1283128^{3} grids for high-resolution output reference. The input 3D volumes are synthetically corrupted to simulate the imperfections of a real-world 3D scanner.

The following experiments are conducted with the trained model: We firstly evaluate the inpainting performance of 3D-ED-GAN on both real-world 3D range scans data and the ShapeNet 20-point testing set with various injected noise. Ablation experiments are conducted to assess the capability of producing high-resolution completion results from the combination of 3D-ED-GAN and LRCN. We also compare with the state-of-the-art method. Then, we evaluate the capability of 3D-ED-GAN as a feature learning framework. Please refer to the supplementary material for more results and comparisons.

Our hybrid network has 26.3M26.3M parameters and requires 7.437.43GB GPU memory. If we add two more full-convolution layers in the decoder and two more convolution layers in the discriminator of 3D-ED-GAN to produce high-resolution, the network has 116.4M116.4M parameters and won’t fit into the GPU memory. For comparison between low-resolution and high-resolution results, we simply upsample the prediction of 3D-ED-GAN and do numerical comparisons.

We test 3D-ED-GAN and LRCN on both real-world and synthetic data. The real-world scans are from the work of . They reconstructed 3D mesh from RGB-D data and we voxelized these 3D meshes into 32332^{3} grids for test. Our network is trained on ShapeNet dataset as in Section 3.3. Before testing, all shapes are aligned using PCA as stated in Section 3.2. Figure 4 shows shape completion examples on real-world scans for both low-resolution and high-resolution outputs. We use 3D-ED-GAN to represent the low-resolution output of 3D-ED-GAN and Hybrid to denote the high-resolution output of the combination of 3D-ED-GAN and LRCN. As we can see, our network is able to produce plausible completion results even with large missing area. The 3D-ED-GAN itself can result conceivable outputs while LRCN further improves fine-grained details.

1.2 Random Noise

We then assess our model with the splitted testing data of ShapeNet. This is applicable in cases where capturing the geometry of objects with 3D scanners results in holes and incomplete shapes. Since it is hard to obtain ground truth for real-world objects, we rely on the ShapeNet dataset where complete object geometry of diversified categories is available, and we test on data with simulated noises.

Because it is hard to predict the exact noise from 3D scanning, we test different noise characteristics and show the robustness of our trained model. We do the following ablation experiments:

3D-ED-GAN: 3D-ED-GAN is trained the first training stage of Section 3.3.

LRCN: After the LRCN is pre-trained as the second training stage in Section 3.3, we directly feed the partially scanned 3D volume into the LRCN as input and train LRCN independently for 100100 epochs with a learning rate of 10−510^{-5} and a batchsize 4. We test the shape completion ability of this single network.

Hybrid: Our overall network, i.e. the combination of 3D-ED-GAN and LRCN, is trained with the aforementioned procedure.

To have a better understanding of the effectiveness of our generative adversarial model, we also compare qualitatively and quantitatively with VConv-DAE . They adopted a full convolutional volumetric autoencoder network architecture to estimate voxel occupancy grids from noisy data. The major difference between 3D-ED-GAN and VConv-DAE is introduction of GAN. In our implementation of VConv-DAE, we simply remove the discriminator from 3D-ED-GAN and compare the two networks with the same parameters.

We first evaluate our model on test data with random noise. As stated above, we adopted simulated scanning noise in our training procedure. With random noise, volumes have to be recovered from limited given information, where the testing set and the training set have different patterns. Figure 5 shows the results of different methods for shape completion with 50%50\% noise injected.

We also vary the amount of noise injected to the data. For numerical comparison, the number nn of generated voxels (at 1283128^{3} resolution) which differ from ground truth (object volume before corruption) is counted for each sample. The reconstruction error is nn divided by total number of grids 1283128^{3}. For 3D-ED-GAN and VConv-DAE, their predictions are computed by upsampling the low resolution output. We use mean error for different object categories as our evaluation metric. The results are reported in Figure 6. It can be seen from Figure 5 that different methods produce similar results. Even though 50%50\% noise is injected, the corrupted input still maintains the global semantic structure of the original 3D shape. In this way, this experiment measures the denoising ability of these models. As illustrated in Figure 6, these models introduce noise when 0%0\% noise injected. LRCN performs better than the other three when the noise percentage is low. When the input gets more corrupted, 3D-ED-GAN tends to perform better than others.

1.3 Simulated 3D scanner

We then evaluate our network on completing shapes for simulated scanned objects. 3D scanners such as Kinect can only capture object geometry from a single view at one time. In this experiment, we simulate these 3D scanners by scanning objects in the ShapeNet dataset from a single view and evaluate the reconstruction performance of our method from these scanned incomplete data. This is a challenging task since the recovered region must contain semantically correct content. Completion results can be found in Figure 7. Quantitative comparison results are shown in Table 1.

As illustrated in Figure 7 and Table 1, our model performs better than 3D-ED-GAN, VConv-DAE and LRCN. For VConv-DAE, small or thin components of objects, such as the pole of a lamp tend to be filtered out even though these parts exist in the input volume. With the help of the generative adversarial model, our model is able to produce reasonable predictions for the large missing areas that are consistent with the data distribution. The superior performance of 3D-ED-GAN over VConv-DAE demonstrates our model benefits from the generative adversarial structure. Moreover, by comparing the results of 3D-ED-GAN and the hybrid network, we can see the capability of LRCN to recover local geometry. LRCN alone has difficulty capturing global context structure of 3D shapes. By combining 3D-ED-GAN and LRCN, our hybrid network is able to predict global structure as well as local fine-grained details.

Overall, our hybrid network performs best by leveraging 3D-ED-GAN’s ability to produce plausible predictions and LRCN’s power to recover local geometry.

2 Feature Learning

We now evaluate the transferability of unsupervised learned features obtained from inpainting to object classification. We use the popular benchmark ModelNet10 and ModelNet40, which are both subsets of the ModelNet dataset . Both ModelNet10 and ModelNet40 are split into mutually exclusive training and testing sets. We conduct three experiments.

Our-FT: We train 3D-ED-GAN as the first training stage stated in Section 3.3 on all samples of ShapeNet dataset as pre-training and treat the encoder component (with a softmax layer added on top of zz as a loss layer) as our classifier. We fine-tune this CNN classifier on ModelNet10 and ModelNet40.

RandomInit: We directly train the classifier mentioned in Our-FT with random initialization on ModelNet10 and ModelNet40.

Our-SVM: We generate zz (of dimension 1638416384) with the trained 3D-ED-GAN in Section 3.3 for samples on ModelNet10 and ModelNet40 and train a linear SVM classifier with zz as the feature vector.

We also compare our algorithm with the state-of-the-art methods . VRN , MVCNN , MVCNN-Multi are designed for object classification. 3DGAN , TL-network , and VConv-DAE-US learned a feature representation for 3D objects, and trained a linear SVM as classifier for this task. VConv-DAE and VRN adopted a VAE architecture with pre-training. We report the testing accuracy in Table 2.

Although our framework is not designed for object recognition, our results with 3D-ED-GAN pre-training is competitive with existing methods including models designed for recognition . By comparing RandomInit and Ours-FT, we can see unsupervised 3D-ED-GAN pre-training is able to guide the CNN classifier to capture the rough geometric structure of 3D objects. The superior performance of Our-SVM training over other vector representation methods demonstrate the effectiveness of our method as a feature learning architecture.

2.2 Shape Arithmetic

Previous works in embedding representation learning have shown the phenomena of the capability of shape transformation by performing arithmetic on the latent vectors. Our 3D-ED-GAN also learns a latent vector zz. To this end, we randomly chose two different instances and fed it into the encoder to produce two encoded vectors z′z^{\prime} and z′′z^{\prime\prime} and feed the interpolated vector z′′′=γz′+(1−γ)z′′z^{\prime\prime\prime}=\gamma z^{\prime}+(1-\gamma)z^{\prime\prime} (0<γ<10<\gamma<1) to the decoder to produce volumes. The results for the interpolation are shown in Figure 8. We observe smooth transitions in the generated object domain with gradually increasing γ\gamma.

Conclusion and Future Work

In this paper, we present a convolutional encoder-decoder generative adversarial network to inpaint corrupted 3D objects. A long-term recurrent convolutional network is further introduced, where the 3D volume is treated as a sequence of 2D images, to save GPU memory and complete high-resolution 3D volumetric data. Experimental results on both real-world and synthetic scans show the effectiveness of our method.

Since our model is easy to fit into GPU memory compared with other 3D CNN methods . A potential direction is to complete more complex 3D structures, such as indoor scenes , with much higher resolutions. Another interesting future avenue is to utilize our model on other 3D representations like 3D mesh, distance field etc.

References