ShapeFormer: Transformer-based Shape Completion via Sparse Representation
Xingguang Yan, Liqiang Lin, Niloy J. Mitra, Dani Lischinski, Daniel Cohen-Or, Hui Huang
Introduction
Shapes are typically acquired with cameras that probe and sample surfaces. The process relies on line of sight and, at best, can obtain partial information from the visible parts of objects. Hence, sampling complex real-world geometry is inevitably imperfect, resulting in varying sampling densities and missing parts. This problem of surface completion has been extensively investigated over multiple decades . The central challenge is to compensate for incomplete data by inspecting non-local hints in the observed data to infer missing parts using various forms of priors.
Recently, deep implicit function (DIF) has emerged as an effective representation for learning high-quality surface completion. To learn shape priors, earlier DIFs encode each shape using a single global latent vector. Combining a global code with region-specific local latent codes can faithfully preserve geometric details of the input in the completion. However, when presented with ambiguous partial input, for which multiple plausible completions are possible (see Fig. 1), the deterministic nature of local DIF usually fails to produce meaningful completions for unseen regions. A viable alternative is to combine generative models to handle the input uncertainty. However, for representations that contain enormous statistical redundancy, as in the case of current local methods, such combination excessively allocates model capacity towards perceptually irrelevant details .
We present ShapeFormer, a transformer-based autoregressive model that learns a distribution over possible shape completions. We use local codes to form a sequence of discrete, vector quantized features, greatly reducing the representation size while keeping the underlying structure. Applying transformer-based generative models toward such sequences of discrete variables have been shown to be effective for generative pretraining , generation and completion in image domain.
However, directly deploying transformers to 3D feature grids leads to a sequence length cubic in the feature resolution. Since transformers have an innate quadratic complexity on sequence length, only using overly coarse feature resolution, while feasible, can barely represent meaningful shapes. To mitigate the complexity, we first introduce Vector Quantized Deep Implicit Functions (VQDIF), a novel 3D representation that is both compact and structured, that can represent complex 3D shapes with acceptable accuracy, while being rather small in size. The core idea is to sparsely encode shapes as sequences of discrete 2-tuples, each representing both the position and content of a non-empty local feature. These sequences can be decoded to deep implicit functions from which high-quality surfaces can subsequently be extracted. Due to the sparse nature of 3D shapes, such encoding reduces the sequence length from cubic to quadratic in the feature resolution, thus enabling effective combination with generative models.
ShapeFormer completes shapes by generating complete sequences, conditioned on the sequence for partial observation. It is trained by sequentially predicting the conditional distribution of both location and content over the next element. Unlike image completion , where the model is trained with the BERT objective to only predict for unseen regions, in the 3D shape completion setting, the input features may also come from both noisy and incomplete observations, and keeping them intact necessarily yields noisy results. Hence, in order to generate whole complete sequences from scratch while being faithful to the partial observations, we adapt the auto-regressive objective and prepend the partial sequence to the complete one to achieve conditioning. This strategy has been proved effective for conditional synthesis for both text and images .
We demonstrate the ability of ShapeFormer to produce diverse high-quality completions for ambiguous partial observations of various shape types, including CAD models and human bodies, and of various incomplete sources such as real-world scans with missing parts. In summary, our contributions include: (i) a novel DIF representation based on sequences of discrete variables that compactly represents satisfactory approximations of 3D shapes; (ii) a transformer-based autoregressive model that uses our new representation to predict multiple high-quality completed shapes conditioned on the partial input; and (iii) state-of-the-art results for multi-modal shape completion in terms of completion quality and diversity. The FPD score on PartNet is improved by at most 1.7 compared with prior multi-modal method cGAN .
Related Work
3D reconstruction is a longstanding ill-posed problem in computer vision and graphics. Traditional methods can produce faithful reconstruction from complete input such as point cloud , or images . Recently, neural network-based methods have demonstrated an impressive performance toward reconstruction from partial input , where the unseen regions are completed with the help of data priors. They can be classified according to their output representation, such as voxels, meshes, point clouds, and deep implicit functions. Since voxels can be processed or generated easily through 3D convolutions thanks to their regularity, they are commonly used in earlier works . However, since their cubic complexity toward resolution, the predicted shapes are either too coarse or too heavy in size for later applications. While meshes are more data-efficient, due to the difficulty of handling mesh topology, mesh-based methods have to either use shape template , limiting to a single topology, or produce self-intersecting meshes . Point clouds, in contrast, do not have such a problem and are popularly used lately for generation and completion . However, point clouds need to be non-trivially post-processed using classical methods to recover surfaces due to their sparse nature. Recent works that represent shapes as deep implicit functions have been shown to be effective for high-quality 3D reconstruction . By leveraging local priors, follow-up works can further improve the fidelity of geometric details. However, most current methods are not effective toward ambiguous input due to their deterministic nature. Other methods handle such input by leveraging generative models. They learn the conditional distribution of complete shapes represented as either a single global code , which, due to their lack of spatial structure, leads to completions misaligned with the input, or raw point cloud , which, due to its statistical redundancy, is only effective for completing simple shapes with a limited number of points. In this paper, we show how building generative models upon our new compact, structured representation enables multi-modal high-quality reconstruction for complex shapes.
Autoregressive models and Transformers
Autoregressive models are generative models that aim to model distributions of high dimensional data by factoring the joint probability distribution to a series of conditional distributions via the chain rule . Using neural networks to parameterize the conditional distribution has been proved to be effective in general, and more specifically to image generation . Transformers , known for their ability to model long-range dependencies through self-attentions, have shown the power of autoregressive models in natural languages , image generation . Contrary to deterministic masked auto-encoders , Transformers can produce diverse image completions that are sharp in masked regions by adopting the BERT training objective. In the 3D domain, autoregressive models have been used to learn the distribution of point clouds and meshes . However, these models can only generate small point clouds or meshes restricted to 1024 vertices due to the lack of efficient representation. In contrast, by eliminating statistical redundancy, a compressed discrete representation enables generative models to focus on data dependencies at a more salient level and recently allows high-resolution image synthesis . Follow-up works utilize data sparsity to obtain even more compact representations . We explore this direction in the context of surface completion. Concurrently with our work, AutoSDF trains Transformers to complete and generation shapes with dense grid. And Point-BERT adopts generative pre-training for several downstream tasks.
Method
With such compact representation, the conditional distribution becomes , where and are the sequence encoding of the partial point cloud and the complete shape, respectively. Once such distribution is modeled, we can sample multiple complete sequences , from which different surface reconstructions can be obtained through decoding. This process is illustrated in Fig. 2.
We propose VQDIF, whose goal is to approximate 3D shapes with a shape dictionary, with each entry describing a particular type of local shape part inside a cell of volumetric grid with resolution . With such a dictionary, shapes can be encoded as short sequences of entry indices, describing the local shapes inside all non-empty grid cells, enabling transformers to efficiently model the global dependencies.
We design an auto-encoder architecture to achieve this. The encoder first maps the input point cloud to a 64 resolution feature grid with local-pooled PointNet and then downsample it to resolution . Unlike the previous strategy for image synthesis , the encoder parameters are carefully set to have the least receptive field, reducing the number of non-empty features to the number of sparse voxels of the voxelized input point cloud at resolution . Then these non-empty features are flattened to a sequence of length in row-major order. Since these features are sparse, we record their locations with their flattened index . Other orderings are also possible, but for generation they are not as effective as row-major order .
Following the idea of neural discrete representation learning , we compress the bit size of the feature sequence through vector quantization, that is, clamping it to its nearest entry in a dictionary of embeddings and we save the indices of these entries:
Thus, we get a compact sequence of discrete 2-tuples representing the 3D shape . Finally, the decoder projects this sequence back to a feature grid and, through a 3D-Unet , decodes it to a local deep implicit function , whose iso-surface is the reconstruction .
We train the VQDIF by simultaneously minimizing the reconstruction loss and updating the dictionary using exponential moving averages , where dictionary embeddings are gradually pulled toward the encoded features. We also adopt commitment loss to encourage encoded features to stay close to their nearest entry in the dictionary, with index , thus keeping the range of the embeddings bounded. We define the loss as,
where sg stands for stop gradient operator which prevents the embedding being affected by this loss.
The full training objective for VQDIF is the combination of reconstruction loss of with weighting factor :
Here, is the size of the target set and BCE is the binary cross-entropy loss which measures the discrepancy between the predicted and the ground truth occupancy at target point . During training, we select the target set and its occupancy values in a similar fashion to prior work .
2 Sequence generation for shape completion
We autoregressively model the distribution , by predicting the distribution of the next element conditioned on the previous elements. We also factor out the tuple distribution for each element: . The final factored sequence distribution is as follows:
Here, indicates model parameters and and are the distributions of the coordinate and the index value of the -th element of , conditioned on previously generated elements and the partial sequence . Note that is also conditioned on the current coordinate .
Different approaches have been applied to build a transformer model that can predict tuple sequences. Instead of flattening them , which in our case doubles the sequence length, we stack two decoder-only transformers to predict the and respectively in a similar way to prior works , as illustrated in Fig. 3. Unlike in the image completion case , where the partial sequence is strictly a part of the complete sequence so that only the missing regions need to be completed. For our case, however, due to the noise or incompleteness of local observations, we would like to predict complete sequences from scratch to fix such data deficiencies. And thanks to the autoregressive structure of the decoder-only transformer, we can achieve conditioning by simply prepending before to generate complete sequences that are in coordination with the partial one. We also append an additional end token to both sequences to help learning.
The training objective of ShapeFormer is to maximize the log-likelihood given both and : After the model is trained, ShapeFormer performs shape completion by sequentially sampling the next element of the complete sequence until an end token ([END]) is encountered. Given the partial sequence, we alternatively sample the new coordinate and value index using top-p sampling , where only a few top choices, for which the sum of probabilities exceeds a threshold , are kept. Also, we mask out the invalid choices for coordinate to guarantee monotonicity.
Results and Evaluation
In this section, we demonstrate our method outperforms prior arts for shape completion from ambiguous scans and part-level incompleteness (Sec. 4.1). Then we show our approach can effectively handle a variety of shape types, out-of-distribution shapes, and real-world scans from the Redwood dataset (Sec. 4.2). Lastly, we show our VQDIF representation has a significantly smaller size compared with prior DIFs while achieving similar accuracy (Sec. 4.3).
Throughout all these experiments, we use feature resolution for VQDIF and set its loss balancing factor . We also set the vocabulary of the dictionary to be . We use and blocks for Coordinate and Value Transformers, respectively. All of these blocks have heads self-attention, and the embedding dimension is . We find that a maximum sequence length of is enough for all of our experiments. We set the default probability factor for sampling. Further implementation details such as architecture and training statistics are provided in the supplementary.
We consider two datasets: 1) ShapeNet for testing on partial scan and 2) PartNet for testing on part-level incompleteness; we follow the same setting as in cGAN . For ShapeNet, following prior works , we use 13 classes of the ShapeNet with train/val/test split from 3D-R2N2 . The data are processed and sampled similarly to IMNet and we create partial input for training via random virtual scanning. For evaluation, we first measure the ambiguity score of a partial point cloud to its complete counterpart as the mean ratio of the distance of each point with its nearest neighbor in to its distance toward furthest neighbor in . We uniformly sample 70 viewpoints on a sphere for each shape. Then we create two setups for the dataset according to ambiguity. The high scan ambiguity setup selects scans with the top half ambiguity score and vice versa. More details about this score are provided in the supplementary material.
Metric
For the low ambiguity setting, we use Chamfer Distance (CD) and F-score%1 (F1) to measure how accurate the completion is; this is similar to the previous setup . And to evaluate completion quality for high ambiguity setting, we follow prior work to use pre-trained PointNet classifier as a feature extractor to compute the Fréchet Point Cloud Distance (FPD) between the set of completion results and ground truth shapes. Additionally, for the PartNet dataset, we follow cGAN and use Unidirectional Hausdorff Distance (UHD) to measure faithfulness toward input, Total Mutual Difference (TMD) to measure diversity, and Minimal Matching Distance (MMD).
Baselines
We compare our model with a global DIF method OccNet , two local DIF methods ConvONet and IF-Net , PoinTr , which adopts Transformers without autoregressive learning, and multi-modal completion method cGAN . We also compare our VQDIF-only model to illustrate the necessity of ShapeFormer. We train these methods for shape completion in our dataset setting with their official implementation.
Results on ShapeNet
As shown in Fig. 4, methods incorporating structured local features can better preserve the input details than those that only operate on global features (OccNet , cGAN ) And deterministic methods tend to produce averaged shape since they are unable to handle multi-modality. Notice that PoinTr also utilizes the power of Transformers, but they can not alleviate this problem by adopting Transformers without generative modeling. This phenomenon is more apparent for the chair example, which has higher ambiguity. Our VQDIF-only model also fails to produce good completion in this case. Based on VQDIF, our ShapeFormer resolves ambiguity by factoring the estimation into a distribution, with each sampled shape sharp and plausible. In contrast, the multi-modal method cGAN is unable to produce high-quality shapes due to their unstructured representation. Further, we generate one completion per input with top-p sampling for quantitative evaluation. As shown in Tab. 1, our method has a much better FPD for high ambiguity scans. Notice CD is not reliable when ambiguity is high since it often treats plausible completions as significant errors. For low ambiguity scans, our method is also competitive toward previous state-of-the-art completion methods in terms of accuracy.
Results on PartNet
We compare our model with cGAN and ShapeInversion on PartNet. The latter method achieves multiple completions through GAN inversion. The quantitative and qualitative comparisons are shown in Tab. 2 and Fig. 5, respectively. Thanks to our structured representation, we achieve much better faithfulness (UHD) and can generate more varied (TMD) high-quality shapes (MMD and FPD) than these GAN-based methods.
2 More results
We further investigate how our model pre-trained on ShapeNet can be applied to scans of real objects. We test our model on partial point clouds converted from RGBD scans of the Redwood 3D Scans dataset . Figure 6 shows the results for a sofa and a table, both of them have two scans from different views. Notice that our model sensitively captures the uncertainty of a scan, producing a distribution of completions that are faithful to the scan and plausible in unobserved regions. We also show results for a sports car in Fig. 2.
Results on out-of-distribution objects
We further evaluate ShapeFormer’s generalization by testing scans of unseen types of shapes on our trained model of Sec. 4.1. We pick the novel shapes from the "Famous" dataset collected by Erler et al. which includes many famous geometries for testing, such as the "Utah teapot," and apply virtual scan to get the partial point cloud. Fig. 7 demonstrates our ShapeFormer can grasp general concepts such as symmetry or hollow and filled. Even the model is only trained on the 13 ShapeNet categories, without ever seeing any cups or teapot, it can still successfully produce multiple reasonable completions from the partial scan. Moreover, in the second row, we see the completions of a one-side scan of a cup contain two distinct features: the cups might be solid or empty. These examples show the ShapeFormer’s potential for general-purpose shape completion, where once we have it trained, we can apply it for all types of shapes.
Results on human shapes
In addition to CAD models, we qualitatively evaluate our completion results on scans of human shapes (D-FAUST dataset ) using the same setting as Niemeyer et al. . Human shapes are very challenging due to their thin structures and the wide variety of poses. To simulate part level incompleteness, we randomly select a point from the complete cloud and only keep neighboring points within a ball of a fixed radius as partial input. Fig. 8 shows examples of our results. We can see that our completions keep the pose of the observed body parts and generate various possible poses for the unobserved body parts.
3 Surface reconstruction with VQDIF
Our final experiment evaluates the representation size and reconstruction accuracy of VQDIF. We compare VQDIF of different feature resolutions (, , ) with OccNet, ConvONet, IF-Net, which are retrained to auto-encode the complete shape with their released implementations. As shown in Fig. 9, achieves similar accuracy to the local implicit approach IF-Net while being significantly smaller in size thanks to the sparse and discrete VQDIF features. The minimum receptive field of our encoder keeps the feature as local as possible, which greatly reduces the feature amount. Then the multi-dimensional feature vectors are quantized and can be referred to using a single integer index, which further reduces the size. The accuracy loss is only salient for lower feature resolution, as seen in the w/o quant. comparison, where we train VQDIF without vector quantization. These together allow transformers to effectively model the distribution of shapes. We adopt for ShapeFormer since it only has an average length of 217 (see Tab. 3) and its accuracy is already comparable with ConvONet (see Fig. 10).
Conclusions
We have presented ShapeFormer, a transformer-based architecture that learns a conditional distribution of completions, from which multiple plausible completed shapes can be sampled. By explicitly modeling the underlying distribution, our method produces sharp output instead of regressing to the mean producing a blurry result. To facilitate generative learning for 3D shape, we propose a new 3D representation VQDIF that can significantly compress the shapes into short sequences of sparse, discrete local features, which in turn enables producing better results, both in terms of quality and diversity, than previous methods.
The major factor limiting our method to be applied in fields like robotics is the sampling speed, which is currently 20 seconds per generated complete shape. In the future, we would also like to explore utilizing a more efficient attention mechanism to allow Transformers to learn VQDIF with smaller size, producing even higher quality completions. Moreover, the current method is generic, leveraging advances in language models. More research is required to include geometric or physical reasoning in the process to better deal with ambiguities.
We thank the reviews for their comments. We thank Ziyu Wan, Xuelin Chen and Jiahui Lyu for discussions. This work was supported in parts by NSFC (62161146005, U21B2023, U2001206), GD Talent Program (2019JC05X328), GD Science and Technology Program (2020A0505100064), DEGP Key Project (2018KZDXM058, 2020SFKC059), Shenzhen Science and Technology Program (RCJC20200714114435012, JCYJ20210324120213036), Royal Society (NAF-R1-180099), ISF (3441/21, 2492/20) and Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ).
References
Appendix A Implementation Details
The ambiguity for a partial point cloud measures the variety of its potential complete shapes. However, the direct measurement for ambiguity is difficult, if not impossible. In contrast, the incompleteness of a partial cloud toward its complete shape is relatively easy to compute. Although it can not fully reflect ambiguity (e.g., a top scan of a table as incomplete as a bottom scan could have a much greater ambiguity), the ambiguity is still strongly correlated.
Hence, we seek to find a metric on the incompleteness of such a point cloud to indicate its ambiguity. Intuitively, we could use metrics like F-score to measure the ratio of the approximate partial surface area toward the complete area. But as indicated in the inset figure, such measures will fail to differentiate the coverage difference of the partial cloud (red dots) to the complete one (in blue). Instead, we propose to use a metric based on Chamfer-, which goes larger as the partial point cloud misses more global structure. Since the partial to complete distance is always negligible, we can only calculate the complete to partial distance. And to compare the ambiguity of scans on different shapes, we normalize the distance of a point according to its farthest distance in the complete shape. More specifically, we define the metric evaluating the ambiguity of scan given the complete point cloud as as:
Where is the number of points in the complete cloud.
We sample 70 views for each shape, 64 of which are evenly sampled from the view sphere (via Fibonacci sampling), and the rest are the six orthogonal views. Then we sort these views according to the score. In Fig. 11, we use a teapot as an example to show the score distribution of these 70 scans. For scans with low ambiguity scores, the underlying shape’s global structure is either captured or is clearly indicated by the captured shape salient features. For example, the scan covering the teapot’s mouth, handle, and body can be completed easily. However, it would be more difficult to infer the complete shape when the score is high since it may have different global structures, and a single explanation is not satisfactory. As shown in the main paper, our method can better handle such scans than existing shape completion methods.
A.2 Architectures
We show the detailed architecture of VQDIF and ShapeFormer in Figs. 12 and 13, respectively. and the parameters of their sub-modules are listed in Tab. 4
The encoder first processes the input cloud with a local pooled PointNet to obtain a feature grid. Similar to prior work , the local pooled PointNet aggregates features within a grid cell in contrast to the original PointNet, where all point features are pooled together to obtain a global feature. Specifically, we use a grid of resolution with a feature size of .
Next, to reduce the number of local features, the high-resolution feature grid is down-sampled to lower resolution , using several consecutive strided convolution blocks. As shown in Tab. 4, the parameters of these blocks are carefully set to have the least receptive field since a large receptive field lets each grid feature cover a larger region, reducing the sparsity of the representation. We can then extract the non-empty features by directly masking the encoded feature grid with the voxelized input point cloud (resolution ) thanks to the minimum receptive field. After flattening and quantizing the features (see the main paper), we get the 2-tuple sequence representation directly sent to the decoder. Note that we also save the "empty" feature to project the sequence back to the feature grid in the decoder.
The decoder consists of a 3D U-Net , an up-sampler, and an implicit decoder. It first projects the quantized sparse sequence back to a 3D feature grid, which serves as the input for the 3D U-Net. In contrast to the encoder, the decoder is designed to have a large receptive field. This is because, in order for the implicit decoder to infer whether a probe lies inside or outside of the shape, we need global knowledge. This is in alignment with prior works . More specifically, we use a 3-step U-Net to increase the receptive field, which integrates both local and global information. The up-sampler has the same number of scaling stages as the down-sampler, but it has a larger receptive field by design. Lastly, similarly to prior work , the implicit decoder consists of multiple ResNet blocks. It takes querying probe points and predicts their occupancy probability .
ShapeFormer
In Fig. 13, we show the detailed architecture of ShapeFormer. The input to the ShapeFormer consists of the concatenated sequence of and . Since these sequences both have variable lengths, we append an end-token ([END]) to each sequence to indicate when the sequence terminates. Next, as in prior works , all these indices are turned into learnable embeddings and are additively combined as the input embedding for ShapeFormer.
The main components of ShapeFormer are two causally-masked transformers, which consist of multiple decoder-only transformer blocks . The first transformer learns to predict the coordinate of the next tuple, conditioned on previous tuples, while the second one learns to predict the value of the next element conditioned on previous tuples and the (predicted) coordinate index of the next element. Thus, the output feature of the first transformer is additively mixed with the input embedding of the second transformer delivering the encoded sequence information.
Each transformer is followed by an output head, which converts the feature produced by the transformer into a categorical distribution of the next sequence element. Both output heads consist of two fully connected layers, followed by a softmax layer to produce categorical conditional distributions for each of the sequence elements: . Note that this essentially shifts the complete sequence to the right by one element. For training, we also empirically find randomly masking out the partial sequence will improve generalization.
A.3 Details on training and sampling
We use Adam optimizer for training both VQDIF and ShapeFormer, and we set the learning rate as for VQDIF and for ShapeFormer. We use step decay for VQDIF with step size equal to 10 and and do not apply learning rate scheduling for ShapeFormer. We train our network on a deep learning server with Intel Xeon CPU E5-2680 v4 CPU*56 and 256GB memory with 10 Nvidia Quadro P6000 graphics cards with a GPU memory size of 24GB. It takes 30 hours for our model to converge on our virtual scan dataset and 8 hours on the PartNet dataset. For D-Faust, the converging time is 16 hours. For sampling, we can obtain a single sample sequence in roughly 20 seconds, and we can also sample 24 sequences in parallel in 5 minutes.
Appendix B More comparisons
We show more visual comparisons between our method and prior state-of-the-art methods in Figs. 14, 15 and 16. Figs. 14 and 15 illustrates results on high-ambiguity scans, In these examples, we can see the averaging effect of the deterministic methods (See the scattering effect in ambiguous regions of the completions of PoinTr ). Our method produces significantly better results in terms of quality and diversity.
Also, we demonstrate our method can also achieve competitive accuracy for low-ambiguity scans in Fig. 16. Since there is limited ambiguity for such scans and the goal is to achieve accuracy toward ground truth, we put the ground truth in the first row and only sample 1 completion for each of our sampling strategies (Ours: top-.4 sampling, Ours*: top-.0, e.g., best sampling). Also, we only compare state-of-the-art deterministic methods: ConvONet , IF-Net , and PoinTr in these examples. As we can see, even the scans cover most areas of the ground truth shape; prior works can still produce unsatisfactory results for unseen regions. In contrast, our method can always produce more accurate, high-quality completions. Moreover, since Ours* always picks the coordinate and value indices with the highest probability, it often produces slightly more accurate shapes.
Appendix C More analysis
ShapeFormer inherits the typical limitations of transformer-based autoregressive models. Mainly, the representation length cannot be too long, and thus the method currently can only use VQDIF with , which may fail to complete and reconstruct shapes with intricate structures; an example is shown in Figure 17. Another related limitation is the sampling speed, which prevents interactive applications.
There are two research avenues to alleviate these problems: (i) Investigating more efficient attention mechanisms to reduce the transformer’s quadratic complexity in the sequence length to or even . (ii) Designing an adaptive quantization scheme for the point clouds, which enables Transformers to focus dependencies on a lower local level while using higher-level features for faraway regions. (iii) Adopt advanced sampling techniques for autoregressive models such as parallel sampling .
Moreover, since we generate sequences of complete shapes from scratch, our results may slightly alter the input geometry to overcome the potential sparsity and noise. Besides using higher resolution quantized features to obtain more accurate generation, another possible improvement to this issue is to include high-resolution features of the input in the decoding procedure as in a recent image inpainting technique .