GEOMetrics: Exploiting Geometric Structure for Graph-Encoded Objects
Edward J. Smith, Scott Fujimoto, Adriana Romero, David Meger
Introduction
Surfaces in our physical world exhibit highly non-uniform curvature; compare a plane’s wing to a microscopic screw that fastens its engine. Traditionally, deep 3D understanding systems have relied upon representations such as voxels and point clouds, which capture structure with either uniform volumetric or surface detail (Choy et al., 2016; Wu et al., 2016; Fan et al., 2017). By representing unimportant, or uninteresting, regions with high detail, these systems scale poorly to higher resolutions and object complexity (see Figure 1a and 1b). While efforts have been made to rectify this issue through intermediary representations (Tatarchenko et al., 2017; Häne et al., 2017; Smith et al., 2018; Zhang et al., 2018), these either continue to rely on sparse volumetric units, or maintain uniform detail through alternative representations.
A triangle mesh is a graph-based shape representation that encodes 3D structure through a set of vertices and corresponding planar faces. Recently, advances in deep learning on graphs have enabled mesh-based 3D shape methods (Kato et al., 2017; Wang et al., 2018; Kanazawa et al., 2018; Jack et al., 2018; Henderson & Ferrari, 2018; Groueix et al., 2018b). However, these approaches produce mesh predictions which uniformly space vertices and faces over their surface, providing no significant improvement over the previously highlighted representations (see Figure 1c) and hence, not exploiting the mesh representation advantages. In particular, by placing many vertices in regions of fine detail while using large triangles to summarize nearly planar regions, one could define adaptive meshes (see Figure 1d), enabling flexible scaling by effectively localizing complexity in object surfaces, and allowing for the 3D structure of complicated shapes to be encoded with smaller space requirements and higher precision.
Moreover, these deep learning mesh-based approaches rely on Graph Convolutional Networks (GCNs) (Kipf & Welling, 2016; Defferrard et al., 2016; Cucurull et al., 2018; Wang et al., 2018). Although effective in many node/graph classification and regression tasks, we argue that GCNs may be inadequate for understanding, reconstructing or generating 3D structure as they may induce over-smoothing while aggregating neighboring information at vertex level (Li et al., 2018). This aggregation bias could in turn lead to a harder learning problem when vital information held at each vertex cannot be derived from its neighbors, and as a direct consequence must not be lost.
Last, an important question when reconstructing 3D objects is how to define a loss between a prediction and its target. A common approach is to employ the Chamfer Distance over some parametrization of the two surfaces (Barrow et al., 1977; Insafutdinov & Dosovitskiy, 2018; Wang et al., 2018; Groueix et al., 2018a; Sun et al., 2018; Fan et al., 2017). However, this loss penalizes the point positions exclusively, and thus, its direct application to mesh vertex positions leads poor accuracy, as no information of the faces they define is provided and the placement of vertices over a surface is, to a large degree, arbitrary. In addition, this local loss function takes no consideration of the global structure of the predicted object, preventing class-specific attributes to emerge and creating global inconsistencies.
Therefore in this paper, we aim to address the above-mentioned limitations by introducing an adaptive mesh reconstruction system, called Geometrically Exploited Object Metrics (GEOMetrics), which properly capitalizes on the advantages and geometric structure of graph-encoded objects. GEOMetrics reformulates graph convolutional layers to prevent vertex smoothing. Moreover, it incorporates an adaptive face splitting heuristic allowing non-uniform detail to emerge. Finally, it introduces a training objective operating both on the local surfaces defined by vertices, via a differentiable sampling procedure, as well as the global structure defined by the graph, through a perceptual loss reminiscent of that of style transfer applications (Gatys et al., 2016; Johnson et al., 2016). To the best of our knowledge, our system is the first deep approach to describing shape as an adaptive mesh, through advances in geometrically-aware graph operations. We extensively evaluate our system on the task of 3D object reconstruction from single RGB images and show that the interplay of our introduced components encourages mesh reconstructions, which properly localize detail, while maintaining structural consistency. As a result, we are able to obtain mesh predictions which outperform previous methods and have far smaller space requirements.
The contributions of this paper can be summarized as:
We introduce the Zero-Neighbor GCN (0N-GCN), an extension of Kipf & Welling (2016), which allows the information at each vertex to be maintained, and as a result better suits the understanding and reconstruction of 3D meshes.
We present an adaptive face splitting procedure to encourage local complexity to emerge when reconstructing meshes, taking advantage of the mesh flexible scaling (see Figure 1d).
We propose a training objective, which operates locally and globally over the surface to produce mesh reconstructions, which are highly accurate and benefit from the graceful scaling of mesh representations.
We highlight through extensive evaluation the substantial benefits provided by the previous contributions and show, on the task of 3D object reconstruction from single RGB images, that by properly exploiting the meshs’ properties and geometry, our GEOMetrics system is able to notably outperform prior methods visually and quantitatively, while requiring far less vertices/faces.
Note that the above-mentioned contributions are not specific to the reconstruction system nor the chosen task and thus, can be easily adapted to arbitrary mesh problems. Code for our system is publicly available on a GitHub repository, to ensure reproducible experimental comparison.https://github.com/EdwardSmith1884/GEOMetrics
Related Work
3D Mesh Reconstruction. Mesh models have only recently been used in generation and reconstruction tasks due to the challenging nature of their complex definition (Wang et al., 2018). Recent mesh approaches rely on graph representations of meshes, and use GCNs (Kipf & Welling, 2016) to effectively process them. Our work most closely relates to Neural 3D Mesh Renderer (Kato et al., 2017) and Pixel2Mesh (Wang et al., 2018), which use deformations of a generic pre-defined input mesh, generally a sphere, to form 3D structures. Similarly, Atlas-Net (Groueix et al., 2018a) uses deformations over a set of primitive square faces to form 3D shapes. Conceptually similar, there exists numerous papers using class-specific input meshes which are deformed with respect to the given input image (Pontes et al., 2017; Kanazawa et al., 2018; Jack et al., 2018; Henderson & Ferrari, 2018; Groueix et al., 2018b; Kar et al., 2015). While effective, these approaches require prior knowledge on the target class or access to a model repository.
Graph Convolutional Networks. The great success of convolutional neural networks in numerous image-based tasks (He et al., 2016, 2017; Huang et al., 2017; Jégou et al., 2017; Casanova et al., 2018) has led to increasing efforts to extend deep networks to domains where graph-structured data is ubiquitous.
Early attempts to extend neural networks to deal with arbitrarily structured graphs relied on recursive neural networks (Frasconi et al., 1998; Gori et al., 2005; Scarselli et al., 2009). Recently, spectral approaches have emerged as an effective alternative which formulates the convolution as an operation on the spectrum of the graph (Henaff et al., 2015; Bruna et al., 2014; Bronstein et al., 2017; Levie et al., 2017). Methods operating directly on the graph domain have also been presented. Defferrard et al. (2016) proposed to approximate the filters using the Chebyshev polynomials applied on the Laplacian operator. This approximation was further simplified by Kipf & Welling (2016). Finally, several works have been introduced exploring well-established deep learning ideas and improving previously reported results (Duvenaud et al., 2015; Hamilton et al., 2017; Monti et al., 2017; Veličković et al., 2018).
3D Object Representation. Deep learning approaches for understanding 3D shapes have, for a long time, employed voxels as a default 3D object representation (Choy et al., 2016; Wu et al., 2016; Smith & Meger, 2017; Tulsiani et al., 2017; Wu et al., 2017, 2018). While straightforward to use, voxels induce a cubic computational cost, scaling poorly to higher resolutions and complex objects. Numerous computationally efficient approaches have arisen, such as octree methods (Riegler et al., 2017; Tatarchenko et al., 2017; Häne et al., 2017), which represent voxel objects with adaptive degrees of detail. Most similar to mesh models are point clouds methods, which represent 3D objects through a set of points in 3D space (Fan et al., 2017; Qi et al., 2017; Insafutdinov & Dosovitskiy, 2018; Novotny et al., 2017). Point clouds represent only the surface information of 3D objects, making them more efficient and scalable. However, as they do not define surface information beyond each point’s local neighborhood they must be uniformly sampled over a surface, and so to encode high levels of detail, the sampling density over the entire surface must increase.
Background
In this section, we review GCNs (Kipf & Welling, 2016), a key component for mesh generation and reconstruction systems, and outline how the Chamfer Distance has been previously employed as a loss for mesh reconstruction.
2 Chamfer Loss: Vertex-To-Point Loss
The Chamfer Distance between predicted and ground truth objects has become a standard metric for 3D reconstruction (Wang et al., 2018; Insafutdinov & Dosovitskiy, 2018; Groueix et al., 2018a; Sun et al., 2018; Fan et al., 2017). This loss is defined as:
and is computed between two sets of points, and , sampled from the predicted surface and the ground truth surface. As discussed in Section 1, this metric performs poorly when directly applied to two sets of mesh vertices as it does not take into account the faces which they define, and because of the difficult learning problem associated with matching highly arbitrary vertex placement on a surface. To avoid these issues, Wang et al. (2018) define as a large set of predicted vertex positions, and as a large, pre-computed set of points sampled from the ground truth surface (see Figure 3a). Defining the ground truth set as a dense uniform sampling over the target surface avoids issues with inconsistent vertex positions across similar objects. However, this leads to predictions with a high number of vertices packed tightly over the full surface. In this way, the mesh predictions resemble a point cloud, and thus fail to take advantage of the graceful scaling properties of their representation.
GEOMetrics Mesh Reconstruction
In this section, we describe our pipeline for reconstructing adaptive meshes from single images and outline our proposed 0N-GCN as well as our suggested adaptive face splitting.
Figure 2 depicts our mesh reconstruction module, which takes as input a mesh model, defined by a set of vertex positions and an adjacency matrix, together with an RGB image depicting an object view and outputs a new mesh prediction. The module is composed of three distinct phases: feature extraction, mesh deformation and face splitting, which are cascaded times to obtain incrementally refined mesh predictions. Note that the initial module takes as input a predefined mesh model (e.g. a sphere), whereas each subsequent module is fed the preceding module’s prediction. In this manner, the initial mesh is iteratively deformed and updated to match the input image.
Our feature extraction is based on the method proposed by Wang et al. (2018), where the input image is passed through a deep CNN and the features from 4 intermediary layers are outputted. The feature vector for each vertex of the mesh is then defined by projecting the vertices of the input mesh onto the CNN outputs and extracting their corresponding features. In addition, each vertex feature vector is provided its 3D coordinates and also, if available, the final feature vector it possessed in the preceding reconstruction module. Our mesh deformation consists of a graph convolutional model, which takes as input a mesh and deforms it by making a residual prediction for the position of each vertex. The residual prediction is then added to the original position to complete the deformation. The graph convolutional model of this mesh deformation phase is made up of a series of the proposed 0N-GCN layers (see Subsection 4.2 for details). Finally, our face splitting phase, described in Subsection 4.3, encourages local complexity to emerge in regions that require additional detail.
2 Zero-Neighbor Graph Convolutional Networks
As described in Subsection 3.1, a potential shortcoming of the standard GCN formulation is that a vertex has no capacity to maintain and directly draw conclusions upon its own information, as this information is smoothed with outside influence at each layer. This outside influence, while useful in global graph understanding contexts, may be detrimental when vital information held at each vertex cannot be derived from its neighbors. This situation is exemplified by meshes, where, if optimally defined, every vertex defines some new surface structure (see Figure 1d). To rectify this problem, we define a Zero-Neighbor update, in which a fraction of a vertex’s feature vector are not updated with the neighbors’ information. This is accomplished by, instead of applying higher powers of the adjacency matrix to reach further depths, taking the adjacency matrix to the power (equivalent to the identity matrix) to exchange with no further depths:
where denotes concatenation between vectors and is a feature index. This 0N-GCN provides a soft middle ground between full exchange of information and no vertex communication, where the network can choose how heavily a portion of the features of a given vertex will be influenced by the rest of the graph.
3 Adaptive Face Splitting
In the final step of each reconstruction module, the mesh’s set of vertices is redefined over its surface, by adding vertices in regions of high detail. To do so, we introduce a face splitting method, which adaptively increases the set of vertices and the connections between them by analyzing the local curvature of the surface at each mesh face. The curvature at each face is computed by taking the average of the angle between a face’s normal and its neighboring faces’ normals. For a given face , made up of vertices , its face normal is calculated as:
where and . The curvature at face is then computed as:
where is the set of neighboring faces of . All faces with curvature over a given threshold, , are then selected to be updated. A selected face is updated by adding a new vertex to its center and connecting it to its 3 original vertices, creating 3 new faces. As the new vertex positions are defined by the positions of already existing vertices, the gradients from each vertex are easily defined to flow back through all previous modules. In this way, each reconstruction module is able to identify areas of the current mesh which require increased detail and prescribe them a higher vertex density. This allows the mesh to fully take advantage of the scaling properties of its representation, by concentrating the generation process in areas of high detail.
GEOMetrics Losses
In this section, we describe the key contributions made to the mesh prediction task when considering 3D geometry. In particular, we introduce a training objective, considering the local topology and the global structure to produce mesh predictions that properly benefit from the graph representation.
We introduce a differentiable sampling procedure which enables us to penalize vertices by the surface they implicitly define, rather then their explicit position. This approach allows predicted meshes to match the target surface without emulating the target vertex positions, which are entirely arbitrary when randomly sampled from the ground truth, while also optimally positioning their vertices and faces.
To do so, we define a discrete probability distribution based on the relative area of each face and sample times from this distribution to determine the number of points to sample per face. Then, we sample the previously chosen number of points uniformly over each corresponding surface.More precisely, given a triangular face defined by vertices , following Osada et al. (2002), a point can be sampled uniformly from the surface of the triangle as:
where . This formulation allows us to differentiate through the random selection via the reparametrization trick (Kingma & Welling, 2013; Rezende et al., 2014), as the sampling procedure is defined by a deterministic transformation on the vertex coordinates and the independent stochastic terms.
We apply this sampling procedure to both the predicted mesh and the ground truth and define an alternative Chamfer loss, which operates over the previously sampled points, rather than the predicted vertices (see Figure 3b):
where and are the sampled points of the predicted mesh and the ground truth, respectively. An algorithmic description of the entire process is provided in the supplementary material. Note this loss differs from the vertex-to-point loss in that it properly penalizes the surface of the predicted mesh instead of the predicted vertex’s positions.
Building on these ideas, we define an improved loss term to more accurately compare the surfaces of two meshes:
where and are the predicted and ground truth meshes, and the faces, and the set of points sampled from the surfaces of and , and is a function computing the distance between a point and a triangular faceCalculated using an optimized adaptation of the Distance Between Point and Triangle in 3D algorithm (Eberly, 1999), details provided in the supplementary material.. This loss is shown in Figure 3c, and an algorithmic description can be found in the supplementary material. Note that this loss provides an exact measure of the distance between a point and a mesh surface. This is in contrast to the Chamfer loss in Eq. 2 and the point-to-point loss in Eq. 8, which are faster to compute, but can drastically lose accuracy if too few points are sampled on either surface. A quantitative analysis of the improvement from these losses is provided in the supplementary material through a toy problem.
2 Global Encoding of Graphs
In order to consider the global structure of an object during the reconstruction process, we introduce a global mesh loss. This loss relies on features extracted from a pre-trained mesh-to-voxel model, which is designed as an encoder-decoder network. The mesh-to-voxel encoder takes as input a mesh graph and produces a latent embedding, from which the 3D object is reconstructed, through the decoder, in a voxelized format. In this manner, objects with structural similarity in voxel space, will have similar latent representations, without requiring similar placement of vertices. The proposed global mesh loss is then defined as
where corresponds to the encoder function of the mesh-to-voxel network. This process is depicted in Figure 4.
The encoder network of the mesh-to-voxel model is built by stacking 0N-GCN layers, followed by a max pooling operation applied to the set of vertices as in Ciregan et al. (2012) to produce a single fixed length latent representation. The decoder is a 3D deconvolutional network (Choy et al., 2016), in following with the network defined in Smith & Meger (2017), to perform image to voxel mappings. The complete mesh-to-voxel network is pre-trained by minimizing the mean squared error on the voxelized representations prior to being used in the GEOMetrics system.
3 Optimization Details
Finally, we present the complete training objective for our mesh reconstruction system. This function combines our differential surface sampling losses, our global structure loss, along with two regularization techniques defined in Wang et al. (2018): an edge length minimizing regularizer and a Laplacian-maintaining regularizer , pushing the predicted mesh to be smooth and visually appealing.
The final loss function of our system is defined as:
where are hyper-parameters weighting the importance of each term. During the initial stages of training we approximate the term using the defined loss function for faster computation. Note that the loss is applied to the output of each mesh reconstruction module.
Experiments
In this section, we demonstrate our algorithm’s ability to reconstruct the surface information of 3D objects from single RGB images by taking advantage of the benefits of the mesh representation. We evaluate on this task across 13 classes of the ShapeNet (Chang et al., 2015) dataset. In addition, we present an ablation study to demonstrate how our algorithm’s individual components contribute to its overall performance.
The dataset consists of mesh models, voxel models, and RGB images computed from 13 large classes of CAD models found in the ShapeNet dataset (Chang et al., 2015). Mesh models were computed from the CADs by removing all texture information and downscaling their size so that each model possesses less then 2000 vertices, where possible. From these mesh models, voxelized counter parts were produced at resolution. From each CAD model, 24 RGB images were produced, from random viewpoints, with the camera projection matrix recorded for use in the feature selection method. The data in each class was then split into a training, validation and test set with a ratio of 70:10:20, respectively. This matches the dataset used for empirical evaluation by Wang et al. (2018) and Smith et al. (2018).
2 Implementation Details
Mesh-to-Voxel Mapping For each class, we train a mesh-to-voxel mapping from the mesh and voxel ground truths, for use in our latent loss. These mappings are trained with Adam optimizer (Kingma & Ba, 2014) ( = 0.9, = 0.999), a learning rate of , and a mini-batch size of 16. We train for iterations and practice early stopping, with the best model selected from evaluating on the validation set every 100 iterations.
GEOMetrics We train the full system on each class in our dataset with Adam optimizer (Kingma & Ba, 2014), at learning rate of for k iterations, and then again for k iterations at a learning rate of , with mini-batch size of 5. We practice early stopping by evaluating on the validation set every 50 iterations. The hyper-parameter settings used, as described in Eq. (10), are , , , and . As mentioned above, is employed as a faster approximation to the loss, specifically for the first 300k iteration. This is because is slow to compute and we found it sufficient to only use it to finetune pre-trained models. All hyper-parameters were initially tuned on the validation set of the chair class. The generic pre-defined mesh fed to the first reconstruction module is an ellipsoid. A face is split at the end of each module only if the curvature at that face is greater than 70°. Architecture details for all networks is provided in the supplementary.
3 Single Image Reconstruction
We evaluate our method’s performance quantitatively by comparing its ability to reconstruct mesh surfaces from single RGB image to an array of high performing 3D object reconstruction algorithms. To do this, we sample points from both the surface of the predicted object and the ground-truth object and compute the F1 score. In following with (Wang et al., 2018), precision and recall are calculated using the percentage of sampled points which exists within a threshold of any sampled point in the compared surface. State of the art results of mesh approaches, N3MR (Kato et al., 2017) and Pixel2Mesh (Wang et al., 2018), a point cloud method, PSG (Fan et al., 2017) and a voxel baseline, 3D-R2N2 (Choy et al., 2016), are reported from Wang et al. (2018). We also compare mesh-based approaches in terms of space requirements (number of vertices). The results of this comparison are summarized in Table 1. As shown, our GEOMetrics system boasts far higher performance than previous approaches, with an average increase in F1 score of points across all classes, and improved score in all classes but one, where we experience a negligible drop of points. In addition, in all cases, our system requires notably less vertices than the previous mesh-based state of the art Pixel2Mesh, e.g. cellphone objects require less vertices, whereas lamp objects require less vertices. With an average of vertices used across all classes, the vertex requirements drop as much as on average, highlighting the potential of the adaptive face splitting. Moreover, when compared to point cloud and voxel baselines, we also exhibits state of the art results.
Qualitative reconstruction results for each of the 13 classes are displayed in Figure 7See supplementary material for additional visualizations.. We boast highly accurate reconstructions of the input object, effectively capturing both global structure and local detail. In addition, we render an un-smoothed chair reconstruction in Figure 5 with its edges heavily outlined, demonstrating the obtained diverse vertex density across a single object and highlighting the way our system represents simple surfaces with a small number of faces, and shifts to higher density where required. Lastly, Figure 6 depicts a visual comparison between GEOMetrics and Pixel2Mesh reconstructions, where we can observe how GEOMetrics is able to provide reconstructions with higher detail (e.g. the sharpness of the chair legs).
4 Single Image Reconstruction Ablation Study
In this subsection, we study the influence of our system’s components and demonstrate their individual importance by comparing our full method’s results on the chair class to ablated versions of the system. We assess the impact of our 0N-GCN layers by replacing them with standard GCNs (Kipf & Welling, 2016) in both the mesh reconstruction as well as the mesh-to-voxel models. We validate the effectiveness of the proposed adaptive face splitting by substituting it by a procedure in which all faces are split at the end of each reconstruction module, keeping uniform detail and maintaining approximately the same the number of vertices as the full approach. We then check the importance of each one of the newly introduced losses by removing them at training time. Note that, when the face sampling losses and are removed, we replace them by the vertex-to-point loss () proposed by Wang et al. (2018). Finally, we compare our method to the Pixel2Mesh, when roughly equivalent in terms of number of vertices.
The results of this ablation study are reported in Table 2. As shown in the table, the biggest effect comes from the introduction of the adaptive face splitting, with a drop of points when replacing it with a uniform splitting heuristic. Moreover, assisting the model to give more importance to vertex features through the 0N-GCN also appears to be relevant. The losses proposed to train the whole system also play an important role, as ignoring them leads to a decrease in performance of points and points, respectively. Moreover, training Pixel2Mesh baseline to use as few vertices as GEOMetrics leads to notably worse performance. These results empirically justify the contributions of our GEOMetrics system.
Conclusion
In this paper, we presented GEOMetrics, a novel approach for adaptive mesh reconstruction, which focuses on exploiting the geometry of the mesh representation. The GEOMetrics system reformulates GCNs to explicitly preserve local vertex information and incorporates an adaptive face splitting procedure to enhance local complexity when necessary. Furthermore, the system is trained by introducing a training objective which operates both locally and globally at mesh level, and capitalizes on the geometric structure of graph-encoded objects. We demonstrated the potential of the approach through extensive evaluation on the challenging task of 3D object reconstruction from single images of the ShapeNet dataset. Finally, we reported visually appealing state of the art results, outperforming existing mesh-based methods by a large margin, while requiring (on average) as many as less vertices. Future research directions include addressing the restrictive constant topology prescribed by the initial mesh object through reconstruction and generation methods, which adapt the topology to match the target mesh.
References
Appendix A Point to Surface Loss
In this section, we describe the the Distance Between Point and Triangle in 3D algorithm (Eberly, 1999). For a given point and triangle , the algorithm computes the minimum distance between the point and any point contained within the triangle. Assuming the triangle is defined by corner point and directions and , then any point contained in the triangle can be defined by a pair of scalars such that , where . We can now define the squared distance between the point and any point in the triangle by the following quadratic function:
where for clarity we denote , , , , , and . Selecting which minimizes provides the minimum distance between the point and triangle . As is continuously differentiable, can be found at an interior point where or at the boundary of the set .
In the first case, note if and only if and satisfy the following:
Then if , we have the minimum distance.Otherwise, the distance minimizing must lie on the boundary of D, where either , , or . In each case , can be reduced to quadratic of one unknown variable, which can be minimized by setting the gradient to .
Appendix B Mesh-to-Voxel Mapping Ablation
In this section, we perform an ablation study over the use of 0N-GCN as building block for our Mesh-to-Voxel Mapping network to highlight its impact with respect to the standard GCN layers. To that end, we compare our model on 3 different object classes to an analogous network composed of standard GCN layers with the same number of parameters. Additionally, we assess the influence of pooling across a set of vertices by comparing it to other forms of aggregation such as the one introduced by the Neural Graph Fingerprint (NGF) model (Duvenaud et al., 2015). The results of this ablation study can be found in Table 3 in terms of mean squared error (MSE). As shown in the table, results demonstrate the benefits of the 0N-GCN layers, as well as the max-pooling vertex set aggregation, for this mesh understanding task.
Appendix C Differentiable Surface Loss Algorithms
This section provides the algorithmic details of both the point-to-point loss (Algorithm 1) as well as the point-to-surface loss (Algorithm 2).
Appendix D Loss Analysis
In this section, we present further analysis of GEOMetrics losses to emphasize the benefits of the introduced point-to-point and surface-to-point losses over the vertex-to-point loss. To that end, we design a toy problem, which consists in optimizing the placement of the vertices of an initial square surface to match the surface area of a target triangle in 2D. Figure 9 (top) depicts the above-mentioned initial and target surfaces. We optimize the placement of the vertices of the initial square by performing gradient descent on each of the losses independently, and calculate the intersection over union (IoU) of the predicted object and the target triangle. Moreover, in order to assess the impact of the number of points sampled, we repeat this experiment times, increasing the number of sampled points from to . Figure 8 shows the results of this experiment. Firstly, we observe that the vertex-to-point loss fails to match the target surface entirely, no matter the number of sampled points. Secondly, we observe that the point-to-point loss performance is notably affected by the number of sampled points. While it exhibits poor performance for lower number of sampled points (e.g. below ), it rapidly improves as the number of sampled points increases, and ultimately, converges to an average performance, which is only slightly lower than that of the point-to-surface loss. Finally, the point-to-surface loss begins with a far higher IoU and remains the stronger option for nearly all numbers of sampled points.
Figure 9 illustrates qualitative results for the three compared losses when optimizing with 50 points sampled. As can be seen, the point-to-surface deformation of the square better matches the target triangle shape, followed by point-to-point, which somewhat emulates the triangle, and vertex-to-point, which exhibits the poorest results.
Appendix E Network Architectures
In this section, we provide details on the architectures of the networks used in the paper. Table 4 describes the feature extractor network of the mesh reconstruction module. Similarly, Table 5 specifies the mesh deformation network of the reconstruction module. Finally, Table 6 and 7 detail the mesh-to-voxel encoder and decoder architectures, respectively.
Appendix F Single Image Reconstruction Visualizations
Figures 10 and 11 depict additional reconstruction results from each ShapeNet object class, with three objects shown per class.