Deep Marching Tetrahedra: a Hybrid Representation for High-Resolution 3D Shape Synthesis
Tianchang Shen, Jun Gao, Kangxue Yin, Ming-Yu Liu, Sanja Fidler
Introduction
Fields such as simulation, architecture, gaming, and film rely on high-quality 3D content with rich geometric details and complex topology. However, creating such content requires tremendous expert human effort. It takes a significant amount of development time to create each individual 3D asset. In contrast, creating rough 3D shapes with simple building blocks like voxels has been widely adopted. For example, Minecraft has been used by hundreds of millions of users for creating 3D content. Most of them are non-experts. Developing A.I. tools that enable regular people to upscale coarse, voxelized objects into high resolution, beautiful 3D shapes would bring us one step closer to democratizing high-quality 3D content creation. Similar tools can be envisioned for turning 3D scans of objects recorded by modern phones into high-quality forms. Our work aspires to create such capabilities.
A powerful 3D representation is a critical component of a learning-based 3D content creation framework. A good 3D representation for high-quality reconstruction and synthesis should capture local geometric details and represent objects with arbitrary topology while also being memory and computationally efficient for fast inference in interactive applications.
Recently, neural implicit representations , which use a neural network to implicitly represent a shape via a signed distance field (SDF) or an occupancy field (OF), have emerged as an effective 3D representation. Neural implicits have the benefit of representing complex geometry and topology, not limited to a predefined resolution. The success of these methods has been shown in shape compression , single-image shape generation , and point cloud reconstruction . However, most of the current implicit approaches are trained by regressing to SDF or OF values and cannot utilize an explicit supervision on the target surface, which imposes useful constraints for training. To mitigate this issue, several works proposed to utilize iso-surfacing techniques such as the Marching Cubes (MC) algorithm to extract a surface mesh from the implicit representation, which, however, is computationally expensive.
In this work, we introduce DMTet, a deep 3D conditional generative model for high-resolution 3D shape synthesis from user guides in the form of coarse voxels. In the heart of DMTet is a new differentiable shape representation that marries implicit and explicit 3D representations. In contrast to deep implicit approaches optimized for predicting sign distance (or occupancy) values, our model employs additional supervision on the surface, which empirically renders higher quality shapes with finer geometric details. Compared to methods that learn to directly generate explicit representations, such as meshes , by committing to a preset topology, our DMTet can produce shapes with arbitrary topology. Specifically, DMTet predicts the underlying surface parameterized by an implicit function encoded via a deformable tetrahedral grid. The underlying surface is converted into an explicit mesh with a Marching Tetrahedra (MT) algorithm, which we show is differentiable and more performant than the Marching Cubes. DMTet maintains efficiency by learning to adapt the grid resolution by deforming and selectively subdividing tetrahedra. This has the effect of spending computation only on the relevant regions in space. We achieve further gains in the overall quality of the output shape with learned surface subdivision. Our DMTet is end-to-end differentiable, allowing the network to jointly optimize the geometry and topology of the surface, as well as the hierarchy of subdivisions using a loss function defined explicitly on the surface mesh.
We demonstrate our DMTet on two challenging tasks: 3D shape synthesis from coarse voxel inputs and point cloud 3D reconstruction. We outperform existing state-of-the-art methods by a significant margin while being 10 times faster than alternative implicit representation-based methods at inference time. In summary, we make the following technical contributions:
We show that using Marching Tetrahedra (MT) as a differentiable iso-surfacing layer allows topological change for the underlying shape represented by a implicit field, in contrast to the analysis in prior works .
We incorporate MT in a DL framework and introduce DMTet, a hybrid representation that combines implicit and explicit surface representations. We demonstrate that the additional supervision (e.g. chamfer distance, adversarial loss) defined directly on the extracted surface from implicit field improves the shape synthesis quality.
We introduce a coarse-to-fine optimization strategy that scales DMTet to high resolution during training. We thus achieves better reconstruction quality than state-of-the-art methods on challenging 3D shape synthesis tasks, while requiring a lower computation cost.
Related Work
We review the related work on learning-based 3D synthesis methods based on their 3D representations.
Early work represented 3D shapes as voxels, which store the coarse occupancy (inside/outside) values on a regular grid, which makes powerful convolutional neural networks native and renders impressive results on 3D reconstruction and synthesis . For high-resolution shape synthesis, DECOR-GAN transfers geometric details from a high-resolution shape represented in voxel to a low-resolution shape by utilizing a discriminator defined on 3D patches of the voxel grid. However, the computational and memory costs grow cubically as the resolution increases, prohibiting the reconstruction of fine geometric details and smooth curves. One common way to address this limitation is building hierarchical structures such as octrees , which adapt the grid resolution locally based on the underlying shape. In this paper, we adopt a hierarchical deformable tetrahedral grid to utilize the resolution better. Unlike octree-based shape synthesis, our network learns grid deformation and subdivision jointly to better represent the surface without relying on explicit supervision from a pre-computed hierarchy.
represent a 3D shape as a zero level set of a continuous function parameterized by a neural network . This formulation can represent arbitrary typology and has infinite resolution. DIF-based shape synthesis approaches have demonstrated strong performance in many applications, including single view 3D reconstruction , shape manipulation, and synthesis . However, as these approaches are trained by minimizing the reconstruction loss of function values at a set of sampled 3D locations (a rough proxy of the surface), they tend to render artifacts when synthesizing fine details. Furthermore, if one desires a mesh to be extracted from a DIF, an expensive iso-surfacing step based on Marching Cubes or Marching Tetrahedra is required. Due to the computational burden, iso-surfacing is often done on a smaller resolution, hence prone to quantization errors. Lei et al. proposes an analytic meshing solution to reduce the error, but is only applicable to DIFs parametrized by MLPs with ReLU activation. Our representation scales to high resolution and does not require additional modification to the backward pass for training end-to-end. DMTet can represent arbitrary typology, and is trained via direct supervision on the generated surface. Recent works learn to regress unsigned distance to triangle soup or point cloud. However, their iso-surfacing formulation is not differentiable in contrast to DMTet.
directly predict triangular meshes and have achieved impressive results for reconstructing and synthesizing simpler shapes . Typically, they predefined the topology of the shape, e.g. equivalent to a sphere , or a union of primitives or a set of segmented parts . As a result, they can not model a distribution of shapes with complex topology variations. Recently, DefTet represents a mesh with a deformable tetrahedral grid where the grid vertex coordinates and the occupancy values are learned. However, similar to voxel-based methods, the computational costf increases cubically with the grid resolution. Furthermore, as the occupancy loss for supervising topology learning and the surface loss for supervising geometry learning do not support joint training, it tends to generate suboptimal results. In contrast, our method is able to synthesize high-resolution 3D shapes, not shown in previous work.
Deep Marching Tetrahedra
We now introduce our DMTet for synthesizing high-quality 3D objects. The schematic illustration is provided in Fig. 1. Our model relies on a new, hybrid 3D representation specifically designed for high-resolution reconstruction and synthesis, which we describe in Sec. 3.1. In Sec. 3.2, we describe the neural network architecture of DMTet that predicts the shape representation from inputs such as coarse voxels. We provide the training objectives in Sec. 3.3.
We represent a shape using a sign distance field (SDF) encoded with a deformable tetrahedral grid, adopted from DefTet . The grid fully tetrahedralizes a unit cube, where each cell in the volume is a tetahedron with 4 vertices and faces. The key aspect of this representation is that the grid vertices can deform to represent the geometry of the shape more efficiently. While the original DefTet encoded occupancy defined on each tetrahedron, we here encode signed distance values defined on the vertices of the grid and represent the underlying surface implicitly (Sec. 3.1.1). The use of signed distance values, instead of occupancy values, provides more flexibility in representing the underlying surface. For greater representation power while keeping memory and computation manageable, we further selectively subdivide the tetrahedra around the predicted surface (Sec. 3.1.2). We convert the signed distance-based implicit representation into a triangular mesh using a marching tetrahedra layer, which we discuss in Sec. 3.1.3. The final mesh is further converted into a parameterized surface with a differentiable surface subdivision module, described in Sec. 3.1.4.
We adopt and extend the deformable tetrahedral grid introduced in Gao et al. , which we denote with , where are the vertices in the tetrahedral grid . Following the notation in , each tetrahedron is represented with four vertices , with , where is the total number of tetrahedra and .
We represent the sign distance field by interpolating SDF values defined on the vertices of the grid. Specifically, we denote the SDF value in vertex as . SDF values for the points that lie inside the tetrahedron follow a barycentric interpolation of the SDF values of the four vertices that encapsulates the point.
1.2 Volume Subdivision
We represent shape in a coarse to fine manner for efficiency. We determine the surface tetrahedra by checking whether a tetrahedron has vertices with different SDF signs – indicating that it intersects the surface encoded by the SDF. We subdivide as well as their immediate neighbors and increase resolution by adding the mid point to each edge. We compute SDF values of the new vertices by averaging the SDF values on the edge (Fig. 2).
1.3 Marching Tetrahedra for converting between an Implicit and Explicit Representation
We use the Marching Tetrahedra algorithm to convert the encoded SDF into an explicit triangular mesh. Given the SDF values of the vertices of a tetrahedron, MT determines the surface typology inside the tetrahedron based on the signs of , which is illustrated in Fig. 3. The total number of configurations is , which falls into 3 unique cases after considering rotation symmetry. Once the surface typology inside the tetrahedron is identified, the vertex location of the iso-surface is computed at the zero crossings of the linear interpolation along the tetrahedron’s edges, as shown in Fig. 3.
Prior works argue that the singularity in this formulation, i.e. when , prevents the change of surface typology (sign change of ) during training. However, we find that, in practise, the equation is only evaluated when . Thus, during training, the singularity never happens and the gradient from a loss defined on the extracted iso-surface (Sec. 3.3), can be back-propagated to both vertex positions and SDF values via the chain rule. A more detailed analysis is in the Appendix.
1.4 Surface Subdivision
Having a surface mesh as output allows us to further increase the representation power and the visual quality of the shapes with a differentiable surface subdivision module. We follow the scheme of the Loop Subdivision method , but instead of using a fixed set of parameters for subdivision, we make these parameters learnable in DMTet. Specifically, learnable parameters include the positions of each mesh vertex , as well as which controls the generated surface via weighting the smoothness of neighbouring vertices. Note that different from Liu et al. , we only predict the per-vertex parameter at the beginning and carry it over to subsequent subdivision iterations to attain a lower computational cost. We provide more details in Appendix.
2 DMTet: 3D Deep Conditional Generative Model
Our DMTet is a neural network that utilizes our proposed 3D representation and aims to output a high resolution 3D mesh from input (a point cloud or a coarse voxelized shape). We describe the architecture (Fig. 4) of the generator for each module of our 3D representation in Sec. 3.2.1, with the architecture of the discriminator presented in Sec. 3.2.2. Further details are in Appendix.
We predict the SDF value for each vertex in the initial deformable tetrahedral grid using a fully-connected network . The fully-connected network additionally outputs a feature vector , which is used for the surface refinement in the volume subdivision stage.
After obtaining the initial SDF, we iteratively refine the surface and subdivide the tetrahedral grid. We first identify surface tetrahedra based on the current value. We then build a graph , where correspond to the vertices and edges in . We then predict the position offsets and SDF residual values for each vertex in using a Graph Convolutional Network (GCN):
where is the total number of vertices in and is the updated per-vertex feature. The vertex position and the SDF value for vertex are updated as and . This refinement step can potentially flip the sign of the SDF values to refine the local typology, and also move the vertices thus improving the local geometry.
After surface refinement, we perform the volume subdivision step followed by an additional surface refinement step. In particular, we re-identify and subdivide and their immediate neighbors. We drop the unsubdivided tetrahedra from the full tetrahedral grid in both steps, which saves memory and computation, as the size of the is proportional to the surface area of the object, and scales up quadratically rather than cubically as the grid resolution increases.
Note that the SDF values and positions of the vertices are inherited from the level before subdivision, thus, the loss computed at the final surface can back-propagate to all vertices from all levels. Therefore, our DMTet automatically learns to subdivide the tetrahedra and does not need an additional loss term in the intermediate steps to supervise the learning of the octree hierarchy as in the prior work .
After extracting the surface mesh using MT, we can further apply learnable surface subdivision. Specifically, we build a new graph on the extracted mesh, and use GCN to predict the updated position of each vertex , and for Loop Subvidision. This step removes the quantization errors and mitigates the approximation errors from the classic Loop Subdivision by adjusting , which are fixed in the classic method.
2.2 3D Discriminator
3 Loss Function
DMTet is end-to-end trainable. We supervise all modules to minimize the error defined on the final predicted mesh . Our loss function contains three different terms: a surface alignment loss to encourage the alignment with ground truth surface, an adversarial loss to improve realism of the generated shape, and regularizations to regularize the behavior of SDF and vertex deformations.
We sample a set of points from the surface of the ground truth mesh . Similarly, we also sample a set of points from to obtain , and minimize the L2 Chamfer Distance and the normal consistency loss between and :
where is the point that corresponds to when computing the Chamfer Distance, and denotes the normal direction at point .
We use the adversarial loss proposed in LSGAN :
The above loss functions operate on the extracted surface, thus, only the vertices that are close to the iso-surface in the tetrahedral grid receive gradients, while the other vertices do not. Moreover, the surface losses do not provide information about what is inside/outside, since flipping the SDF sign of all vertices in a tetrahedron would result in the same surface being extracted by MT. This may lead to disconnected components during training. To alleviate this issue, we add a SDF loss to regularize SDF values:
where denotes the SDF value of point to the mesh . In addition, we apply the regularization loss on the predicted vertex deformations to avoid artifacts: .
The final loss is a weighted sum of all five loss terms:
where are hyperparameters (provided in the Supplement).
Experiments
We first evaluate DMTet in the challenging application of generating high-quality animal shapes from coarse voxels. We further evaluate DMTet in reconstructing 3D shapes from noisy point clouds on ShapeNet by comparing to existing state-of-the-art methods.
We collected 1562 animal models from the TurboSquid websitehttps://www.turbosquid.com, we obtain consent via an agreement with TurboSquid, and following license at https://blog.turbosquid.com/turbosquid-3d-model-license/ . These models have a wide range of diversity, ranging from cats, dogs, bears, giraffes, to rhinoceros, goats, etc. We provide visualizations in Supplement. Among 1562 shapes, we randomly select 1120 shapes for training, and the remaining 442 shapes for testing. We follow the pipeline in Kaolin to convert shapes to watertight meshes. To prepare the input to the network, we first voxelize the mesh into the resolution of , and then sample 3000 points from the surface after applying marching cubes to the voxel grid. Note that this preprocessing is agnostic to the representation of the input coarse shape, allowing us to evaluate on different resolution voxels, or even meshes.
We compare our model with the official implementation of ConvOnet , which achieved SOTA performance on voxel upsampling. We also compare to DECOR-GAN , which obtained impressive results on transferring styles from a high-resolution voxel shape to a low-resolution voxel. Note that the original setting of DECOR-GAN is different from ours. For a fair comparison, we use all 1120 training shapes as the high-resolution style shapes during training, and retrieve the closet training shape to the test shape as the style shape during inference, which we refer as DECOR-Retv. We also compare against a randomly selected style shape as reference, denoted as DECOR-Rand.
We evaluate L2 and L1 Chamfer Distance, as well as normal consistency score to assess how well the methods reconstruct the corresponding high-resolution shape following . We also report Light Field Distance (LFD) which measures the visual similarity in 2D rendered views. In addition, we evaluate Cls score following . Specifically, we render the predicted 3D shapes and train a patch-based image classifier to distinguish whether images are from the renderings of real or generated shapes. The mean classification accuracy of the trained classifier is reported as Cls (lower is better). More details are in the Supplement.
We provide quantitative results in Table 1 with qualitative examples in Fig. 5. Our DMTet achieves significant improvements over all baselines in terms of all metrics. Compared to both ConvOnet and DECOR-GAN , our DMTet reconstructs shapes with better quality when training without adversarial loss (5th column in Fig. 5). Further geometric details, including nails, ears, eyes, mouths, etc, are captured when trained with the adversarial loss (6th column in Fig. 5), significantly improving the realism and visual quality of the generated shape. To demonstrate the generalization ability of our DMTet, we collect human-created low-resolution voxels from Turbosquid (shapes unseen in training). We provide qualitative results in Fig. 6. Despite the fact that these human-created shapes have noticeable differences with our coarse voxels used in training, e.g., different ratios of body parts compared with our training shapes (larger head, thinner legs, longer necks), our model faithfully generates high-quality 3D details conditioned on each coarse voxel – an exciting result.
We conduct user studies via Amazon Machanical Turk (AMT) to further evaluate the performance of all methods. In particular, we present two shapes that are predicted from two different models to the AMT workers and ask them to evaluate which one is a better looking shape and which one features more realistic details. Detailed experimental settings are provided in the Supplement. We compare DMTet against ConvONet , DECOR -Retv, as well as DMTet without adversarial loss (w.o. Adv.). Quantitative results are reported in Table 2. Human judges agree that the shapes generated from our model have better details, compared to all baselines, in a vast majority of the cases. Ablations on using adversarial loss demonstrate the effectiveness of generating higher quality geometry using a discriminator during training.
To evaluate the effectiveness of our volume subdivision and surface subdivision modules, we ablate by sequentially introducing them to the base model (we refer as DMTetB) which we train on 100-resolution uniform tetrahedral grid without both volume and surface subdivision modules and adversarial loss. We conduct user studies to evaluate the improvement after each step using the protocol described in the above paragraph. We first reduce the initial resolution to 70 and employ volume subdivision to support higher output resolution (we refer this model as DMTetV) and compare with DMTetB. Predictions by DMTetV wins 78% of cases over DMTetB for better looking, and 61% of cases for realistic details, showing that the volume subdivision module is effective in synthesizing shape details. We then add surface subdivision on top of the DMTetV and compare with it. The new model wins 62% of cases over DMTetV for better looking, and 62% of cases for realistic details as well, demonstrating the effect of surface subdivision module in enhancing the shape details.
2 Point Cloud 3D Reconstruction
We follow the setting from DefTet , and use all 13 categories in ShapeNet core dataThe ShapeNet license is explained at https://shapenet.org/terms, which we pre-process using Kaolin to watertight meshes. We sample 5000 points for each shape and add Gaussian noise with zero mean of standard deviation . For quantitative evaluation, we report the L1 Chamfer Distance in the main paper, and refer readers to the Supplement for results in other metrics (3D IoU, L2 Chamfer Distance and F1 score). We additionally report average inference time on the same Nvidia V100 GPU.
We compare DMTet against state-of-the-art 3D reconstruction approaches using different representations: voxels , deforming a mesh with a fixed template , deforming a mesh generated from a volumetric representation , DefTet , and implicit functions . For a fair comparison, we use the same point cloud encoder for all the methods, and adopt the decoders in the original papers to generate shapes in different representations. We also remove the adversarial loss in this application, since baselines also do not have it. We further compare with oracle performance of MC/MT where the ground truth SDF is utilized to extract iso-surface using MC/MT.
Quantitative results are summarized in Table 3, with a few qualitative examples shown in Fig. 7. Compared to DMC , which also predicts the SDF values and supervises with a surface loss, DMTet achieves much better reconstruction quality since training using the marching tetrahedra layer is more efficient than calculating an expectation over all possible configurations within one grid cell as done in DMC . Compared to a method that deforms a fixed template (sphere) , we reconstruct shapes with different topologies, achieving more faithful results compared to the ground truth shape. When compared with other explicit surface representations that also support different topology , our method achieves higher quality results for local geometry, benefiting from the fact that the typology is jointly optimized with the geometry, whereas it is separately supervised by an occupancy loss in . Compared to a neural implicit method , we generate higher quality shapes with less artifacts, while running significantly faster at inference. Finally, compared to a voxel-based method at the same resolution, our method recovers more geometric details, benefiting from the predicted vertex deformations as well as the surface loss.
2.1 Analysis
We investigate how each component in our representation affects the performance and reconstruction quality.
We first demonstrate the effect of learning on explicit surface via MT. We compare with the oracle performance of extracting the iso-surface with MT/MC from the ground truth signed distance fields on the Chair test set in ShapeNet, which contains diverse high-quality details. Specifically, for MC/MT, we first compute the discretized SDF at different grid resolutions, and compare the extracted surface to the ground truth surface.
As shown in Fig. 8, MT consistently outperforms MC when querying the same number of points. We found the staggered grids pattern in tetrahedral grid better captures thin structures at a limited resolution (Fig. 9). This makes MT a better choice for efficiency reasons. The usage of tetrahedral mesh in DMTet follows this motivation. Without deforming the grid, DMTet outperforms the oracle performance of MT by a large margin when querying the same number of points, although DMTet predicts the surface from noisy point cloud. This demonstrates that directly optimizing the reconstructed surface can mitigate the discretization errors imposed by MT to a large extent.
We further provide ablation studies on the entire ShapeNet test set, which is summarized in Tab. 3. We first compare the version where we only predict SDF values without learning to deform the vertices and volume/surface subdivision with the version that predicts both SDF and the deformation. Predicting deformation along with SDF is significantly more performant, since vertex movements allow for a better reconstruction of the underlying surface. This is especially true for categories with thin structures (e.g. lamp) where the grid vertices are desired to align with them. We further ablate the use of volume subdivision and surface subdivision. We show that each component provides an improvement. In particular, volume subdivision has a significant improvement for object categories with fine-grained structural details, such as airplane and lamp, which require higher grid resolutions to model the occupancy change. Surface subdivision generates shapes with a parametric surface, avoiding the quantization errors in the planar faces and produces more visually pleasing results.
Conclusion
In this paper, we introduced a deep 3D conditional generative model that can synthesize high-resolution 3D shapes using simple user guides such as coarse voxels. Our DMTet features a novel 3D representation that marries implicit and explicit representations by leveraging the advantages of both. We experimentally show that our approach synthesizes significantly higher quality shapes with better geometric details than existing methods, confirmed by quantitative metrics and an extensive user study. By showcasing the ability to upscale coarse voxels such as Minecraft shapes, we hope that we take one step closer to democratizing 3D content creation.
Broad Impact
Many fields such as AR/VR, robotics, architecture, gaming and film rely on high-quality 3D content. Creating such content, however, requires human experts, i.e., experienced artists, and a significant amount of development time. In contrast, platforms like Minecraft enable millions of users around the world to carve out coarse shapes with simple blocks. Our work aims at creating A.I. tools that would enable even novice users to upscale simple, low-resolution shapes into high resolution, beautiful 3D content. Our method currently focuses on 3D animal shapes. We are not currently aware of and do not foresee nefarious use cases of our method.
Disclosure of Funding
This work was funded by NVIDIA. Tianchang Shen and Jun Gao acknowledge additional revenue in the form of student scholarships from University of Toronto and the Vector Institute, which are not in direct support of this work.