Deformable Shape Completion with Graph Convolutional Autoencoders
Or Litany, Alex Bronstein, Michael Bronstein, Ameesh Makadia
Introduction
The problem of reconstructing 3D shapes from partial observations is central to a broad spectrum of applications, ranging from virtual and augmented reality to robotics and autonomous navigation. Of particular interest is the setting where objects may undergo articulations or more generally non-rigid deformations. While several methods based on (volumetric) convolutional neural networks have been proposed for completing man-made rigid objects (see ), they struggle at handling deformable shapes. However, this is not a limitation specific to volumetric approaches. The same difficulties with deformable shapes, irrespective of the completion task, are present for other 3D shape representations utilized in deep learning frameworks, such as view-based and point clouds .
The main reason for this is that for methods based on Euclidean convolutional operations (e.g. volumetric or view-based deep neural networks), an assumption of self-similarity under rigid transformations (in most cases, axis-aligned) is implied. For example a chair seat will always be parallel to the floor. Non-rigid deformations violate this assumption, effectively making each pose a novel object. Thus, tackling such data with a standard CNN requires many network parameters and a prohibitively large amount of training. Although model-based methods such as have shown good performance, they are restricted to a specific class of shape with manually constructed models.
To explicitly enable robustness towards non-rigid deformations, the approach advocated in this paper adopts recent advances for in CNNs on graphs which directly exploit the 3D mesh structure. This allows the learning of a powerful non-rigid shape representation from data without an explicit model.
Another shortcoming of deep learning shape completion techniques stems from their end-to-end design. A network trained to perform completion would be biased towards the type of missing data introduced at training, and may not generalize well to previously unseen types of missing information. To allow generalization to any style of partiality we choose to separate the task of completion from the training procedure altogether. As a result, we also avoid a significant amount of preprocessing and augmentation that is typically done on the training data.
Finally, when a complete mesh is desired as the output, producing a triangulation from point clouds or volumetric grids is itself a challenging problem and may introduce undesired artifacts (although recent advances such as address this by directly producing implicit surfaces). Conversely, by utilizing a mesh-convolutional network our method will produce complete and plausible surfaces by design.
The main contribution of this work is a method for deformable shape completion that decouples the task of partial shape completion from the task of learning a generative shape model, for which we introduce a novel graph convolutional autoencoder architecture. Compared to previous works the proposed method has several advantages. First, it can handle any style of partiality without needing to see any partial shapes during training. Second, the method is not limited to a specific class of shapes (e.g. humans) and can be applied to any kind of 3D data. Third, shape completion is an inherently ill-posed problem with potentially multiple valid solutions fitting the data (this is especially true for articulated and deformable shapes), thus making deterministic solutions inadequate. The proposed method reflects the inherent ambiguities of the problem by producing multiple plausible solutions.
Related work
The application addressed in this paper is a very active research area in computer vision and graphics, ranging from completion of small holes and larger missing regions in individual objects , to entire scenes . Completion guided by geometric priors has been explored, for example Poisson filling and self-similarity . However, such methods work only for small missing regions, and dealing with bigger occlusions requires stronger priors. A viable alternative is model-based approaches, where a parametric morphable model describing the variability of a certain class of objects can be fit to the observed data .
The setting of non-rigid shape completion differs from its rigid counterpart in that at inference, the input partial shape may admit a deformation unseen in the training data. This distinction becomes crucial as large missing regions force the priors to become more complex (see for example the human model designed in ).
Generative methods for non-rigid shapes.
The state-of-the-art in generative modeling has rapidly advanced with the introduction of Variational Autoencoders (VAE ), Generative Adversarial Networks , and related variations (e.g. VAEGAN ). These advances have been adopted by the 3D shape analysis community for dynamic surface generation through VAE and image-to-shape generation through VAEGAN . In , a VAE for non-rigid shapes is proposed. This work differs from ours in that the core operations of our network are graph-convolutional operations as opposed to fully-connected layers, and our network operates directly on raw 3D vertex positions rather than relying on hand-crafted features.
Geometric deep learning.
This paper is closely related to a broad area of active research in geometric deep learning (see for a summary). The success of deep learning (in particular, convolutional architectures ) in computer vision has brought a keen interest in the computer graphics community to replicate this progress for applications dealing with geometric 3D data. One of the key difficulties is that for such data it requires great care to define the basic operations constituting deep neural networks, such as convolution and pooling.
Several works avoid this problem by using a Euclidean representation of 3D shapes, such as rendering a collection of 2D views , volumetric representations , or point cloud . One of the main drawbacks of such extrinsic deep learning methods is their difficulty to deal with shape deformations as discussed earlier. Additionally, voxel representations are often memory intensive and suffer from poor resolution , although recent models have been proposed to address these issues: implicit surface representation , sparse octree networks , encoder-decoder CNN for patch-level geometry refinement , and a long-term recurrent CNN for upsampling coarse shapes . Regarding point cloud representations, the PointNet model applies identical operations to the coordinates of each point and aggregates this local information without allowing for interaction between different points which makes it difficult to capture local surface properties. PointNet++ addresses this by proposing a spatially hierarchical model. Additionally, for PointNet to be invariant to rigid transformations the input point clouds are aligned to a canonical space. This is achieved by a small network that predicts the appropriate affine transformation, but in general such an alignment would be difficult for articulated and deformable shapes.
An alternative strategy is to redefine the basic ingredients of deep neural networks in a geometrically meaningful or intrinsic manner. The first intrinsic CNN-type architectures for 3D shapes were based on local charting techniques generalizing the notion of “patches” to non-Euclidean and irregularly-sampled domains . The key advantage of this approach is that the generalized convolution operations are defined intrinsically on the manifold, and thus automatically invariant to its isometric deformations. As a result, intrinsic CNNs are capable of achieving correspondence results with significantly less parameters and a very small training set. Related independent efforts developed CNN-type architectures for general graphs .
Recently, suggested a dynamic filter in which the assignment of each filter to each member of the k-ring in a graph neighborhood is determined by its feature values. Importantly, this method demonstrated state-of-the-art performance working directly on the embedding features. Thus, in our work we build upon as a basic building block for convolution operations.
Partial shape correspondences.
Dense non-rigid shape correspondence is a fundamental challenge as it is an enabler for many high level tasks like pose or texture transfer across surfaces. We refer the interested reader to for a detailed review of the literature. The proposed method in this work builds upon correspondence between a partial input and a canonical shape of the same class, and related to this are several methods that explore partial shape correspondence and matching . The approaches demonstrating state-of-the-art performance on partial human shapes (e.g. ) treat correspondence as a vertex classification task. Recently has shown impressive results for correspondence across different human subjects in varied pose and clothing.
Inpainting.
The 3D shape completion task is closely related to the analogous structured prediction task of image inpainting . However, our proposed optimization scheme is more reminiscent of style transfer techniques. In our setting we optimize only for the best complete shape with no constraints on the internal feature representation.
Method
We propose a shape completion method that detaches the process of learning to generate 3D shapes from the task of partial shape completion. Our method requires a generative model for complete 3D shapes which we construct by training a graph-convolutional variational autoencoder (VAE ). Partial shapes can be completed by identifying the shape in the output space of the VAE’s generator which best aligns with the partial input. We propose an optimization in the latent space that iteratively deforms (non-rigidly) a randomly generated shape to align with a partial input. In what follows, we describe in more detail both ingredients of our process, the VAE generator and the partial shape completion scheme. A schematic rendition of the method is depicted in Figure 1.
The choice to measure shape reconstruction loss with pointwise distances is not the only option. For example, the VAE can be combined with a Generative Adversarial Network (VAE-GAN) as in , thus introducing an additional discriminator loss on the reconstructed shape. We do not consider a discriminator in the scope of this work to avoid additional model complexity but leave it as future work to investigate different loss functions that can be imposed on reconstructed shapes.
Partial shape completion.
Shape completion is an inherently ill-posed problem that can have multiple plausible solutions. In cases where there exists more than one solution consistent with the data, sampling a result from our proposed generative model allows us to explore this space. The results in Section 4.2 illustrate the variability in completed shapes when repeating the optimization procedure (2) with random initializations.
Experiments
The majority of our experiments are performed on human shapes. The VAE is trained on registered 4D scans from the DFAUST dataset comprising human subjects performing different activities. Scans are captured at a high frame rate and registered to a canonical topology. Due to the high frame rate, we subsample the data temporally by a factor of . We consistently subsample each mesh by the factor of down to vertices. Refer to the supplemental info for details on the data processing. The training set is created by holding out all scans for two human subjects and two activities leaving approximately training shapes. Details for additional experiments with face meshes is provided in Section 4.6.
Network parameters.
The structure of our graph-convolutional VAE is illustrated in Figure 1. We evaluated a number of model parameters on a subset of the training set to inform our final design choices. We use and latent dimensionality of for all our DFAUST experiments. A more important and delicate decision is the selection of the parameter controlling the emphasis on pushing the variational latent distribution towards the Gaussian prior. Our experiments show, as expected, that a higher weight for the Gaussian prior causes randomly sampled latent vectors to generate realistic shapes more likely, while a lower improves reconstruction accuracy over a wider variety of shapes. In the context of our problem, it is more important for the latent space to represent and for the decoder to be able to generate a wide variety of shapes accurately. Sampling from the latent space is less important since the final latent vectors are obtained by means of solving the optimization problem (2). Consequently, we selected at training (see supplemental material for empirical analysis motivating these choices).
Implementation details.
We train the model directly on the input meshes from the DFAUST dataset as described above; the sparse adjacency matrices (we use a vertex 2-ring as the neighborhood size) are passed as side information to define the graph convolutional layers. Data are augmented by adding normally distributed noise to the vertex positions as well as a global planar translations and scalings. We use the ADAM optimizer with the learning rate set to , momentum to , batch size to , Xavier initialization for all weights, and train for iterations. For shape completion optimization we use an SGD optimizer with a learning rate. All additional data for training and evaluation will be provided on the authors’ websites.
1 Representation quality
2 Completion variability
As explained in Section 3, given a partial input with more than one solution consistent with the data, we may explore this space of completions by sampling the initialization of problem 2 at random from the Gaussian prior. For evaluation we consider several test subjects with removed limbs. Figure 4 shows unique plausible completions of the same partial input achieved by random initializations.
3 Synthetic range scans completion
The following experiment considers the common practical scenario of range scan completion. We utilize a test-set of virtual scans produced from viewpoints around human subjects exhibiting different poses. The full shapes were taken from FAUST , and are completely disjoint from our train set, as they contain novel subjects and poses. Furthermore, the data is suitable for quantitative comparison as sufficient information is given in the partial shape to make the completion problem nearly deterministic. Keeping the ground truth correspondence from each view to the full shape, we report the mean completion error in table 1 as Ours (ground truth). More interesting are the results of end-to-end completion using partial correspondence obtained by MoNet (reported as Ours (MoNet)). For reference we report the performance of other shape completion methods: 3D-EPN which has shown state-of-the-art performance for shape completion using volumetric networks, Poisson reconstruction , and nearest neighbor (NN). Note, in order to comply with the architecture of 3D-EPN, we also provide viewpoint information, which is unknown for our method. For NN the completion is considered to be the closest shape from the entire training using the ground truth correspondences. Results in table 1 show mean Euclidean distance (in cm) and relative volumetric error (in %) for the missing region. More results are shown in Figures 5 and 6.
4 Dynamic Fusion
For a quantitative analysis, we perform fusion on three partial views from a static shape. We use the same FAUST shapes used for testing in Section 4.3. Table 2 shows mean reconstruction errors for all test shapes when fusing three different partial views. The results show how reconstruction accuracy changes according to the viewpoint, and consistently improves with latent space fusion. A qualitative evaluation of the fusion problem is shown for the dynamic setting in Figure 7. Each row shows three partial views of the same human subject from a different viewpoint and a different pose. The latent space fusion of the completed shapes is shown in column 4.
5 Real range scan completion
The MHAD dataset provides Kinect scans from viewpoints of subjects performing a variety of actions. We apply our completion method to the extracted point cloud (correspondences were initialized through coarse alignment to a training shape, see the supplemental for details). Figure 8 depicts examples of scan completion on the Kinect data as well as on real scans from the DFAUST dataset.
6 Face completion
A strength of our fully data-driven approach is that by avoiding explicit shape modeling it generalizes easily to different classes of shapes. This is illustrated by an evaluation on deformable faces. training face meshes, each with vertices, are generated from the model provided by . These face models exhibit less variability in pose relative to the human meshes, so we use a much smaller VAE network (only two convolutional layers and a latent dimensionality of ). Figure 9 shows completion for different styles of simulated partiality as well as simulated correspondence noise (see the supplemental for more details).
Conclusions and future work
This paper introduces a novel graph-convolutional method for shape completion. Its important properties include a model robust to non-rigid deformations, small sample complexity when training, and the ability to reconstruct any style of missing data. Evaluations indicate this is a promising first step towards shape completion from real-world scans, and the analysis reveals directions for future work. Firstly, exploring a representation that disentangles shape and pose would allow for more control in the completion and likely improve dynamic fusion results. Secondly, for initialization we require correspondences between the partial and canonical shape model. Although we show resilience to poor correspondences, improving this initialization for noisy real-world data would be beneficial. Finally, the proposed formulation assumes the desired shape topology (i.e. vertex connectivity) is known when decoding shapes. We leave to future work the task of completion with unknown topology.
Acknowledgement
The authors wish to thank Emanuele Rodolà, Federico Monti, Vikas Sindhwani, and Leonidas Guibas for useful discussions. Much appreciated is the DFAUST scan data provided by Federica Bogo and Senya Polikovsky.