Neural Mesh Flow: 3D Manifold Mesh Generation via Diffeomorphic Flows

Kunal Gupta, Manmohan Chandraker

Introduction

Polygon meshes allow an efficient virtual representation of 3D objects, enabling applications in graphics rendering, simulations, modeling and manufacturing. Consequently, mesh generation or reconstruction from images or point sets has received significant recent attention. While prior approaches have primarily focused on obtaining geometrically accurate reconstructions, we posit that physically-based applications require meshes to also satisfy manifold properties. Intuitively, a mesh is manifold if it can be physically realized, for example, by 3D printing. Typically, reconstructed meshes are post-processed with humans in the loop for manifoldness, in order to enable ray tracing, slicing or Boolean operations. In contrast, we propose a novel deep network that directly generates manifold meshes (Fig. 1), alleviating the need for manual post-processing.

A manifold is a topological space that locally resembles Euclidean space in the neighbourhood of each point. A manifold mesh is a discretization of the manifold using a disjoint set of simple 2D polygons, such as triangles, which allows designing simulations, rendering and other manifold calculations. While a mesh data structure can simply be defined as a set (V,E,F)(\mathcal{V},\mathcal{E},\mathcal{F}) of vertices V\mathcal{V} and corresponding edges E\mathcal{E} or face F\mathcal{F}, not every mesh (V,E,F)(\mathcal{V},\mathcal{E},\mathcal{F}) is manifold. Mathematically, we list various constraints on a singly connected mesh with the set (V,E,F)(\mathcal{V},\mathcal{E},\mathcal{F}) that enables manifoldnessIn the scope of this work, meshes do not exhibit defects like duplicate elements, isolated vertices, degenerate faces and inner surfaces that can also cause a mesh to be non-manifold..

Each edge e∈Ee\in\mathcal{E} is common to exactly 2 faces in F\mathcal{F} (Fig. 2a)

Each vertex v∈Vv\in\mathcal{V} is shared by exactly one group of connected faces (Fig. 2b)

Adjacent faces Fi,FjF_{i},F_{j} have normals oriented in same direction (Fig. 2c)

The above mentioned constraints on a mesh (V,E,F)(\mathcal{V},\mathcal{E},\mathcal{F}) guarantee it to be a manifold in the limit of infinitesimally small discretization. That is not the case when dealing with practical meshes with large and non-uniformly distributed triangles. To ensure physical realizability, we tighten the definition with a fourth practical constraint that no two triangles may intersect (Fig. 2d).

In this work, we pose the task of 3D shape generation as learning a diffeomorphic flow from a template genus-0 manifold mesh to a target mesh. Our key insight is that manifoldness is conserved under a diffeomorphic flow due to their uniqueness 9, 10 and orientation preserving property 11, 12. In contrast to methods that learn “deformations” of a template manifold using an MLP or graph-based network 3, 4, 5, our approach ensures manifoldness of the generated mesh. We use Neural ODEs 1 to model the diffeomorphic flow, however, must overcome their limited capability to represent a wide variety of shapes 9, 10, 13, which has restricted prior works to single-category representations 14, 15. We propose novel architectural features such as an instance normalization layer that enables generating 3D shapes across multiple categories and a series of diffeomorphic flows to gradually refine the generated mesh. We show quantitative comparisons to prior works and more importantly, compare resulting meshes on physically meaningful tasks such as rendering, simulation and 3D printing to highlight the importance of manifoldness.

Consider the task of deforming a template unit spherical mesh SS (Fig. 3a) into a target star mesh TT (Fig. 3b). We approximate the deformation with a multi-layer perceptron (MLP) fθf_{\theta} with a unit hidden layer of 256256 neurons with relurelu and output layer with tanhtanh activation. We train fθf_{\theta} by minimizing various losses over the points sampled from S,TS,T. A conventional approach involves minimizing the Chamfer Distance LcL_{c} between S,TS,T, leading to accurate point predictions but several edge-intersections (Fig. 3c). By introducing edge length regularization 4 LeL_{e}, we get fewer edge-intersections (Fig. 3d) but the solution is also geometrically sub-optimal. We can further reduce edge-intersections with Laplacian regularization 4 (Fig. 3e), but this takes a bigger toll on geometric accuracy. Thus, attempting to reduce self-intersections by explicit regularization not only makes the optimization hard, but can also lead to predictions with lower geometric accuracy. In contrast, our proposed use of NODE (with dynamics fθf_{\theta}) is designed by construction 9, 10 to prevent self-intersections without explicit regularization (Fig. 3f).

In summary, we make the following contributions:

A novel approach to 3D mesh generation, Neural Mesh Flow (NMF), with a series of NODEs that learn to deform a template mesh (ellipsoid) into a target mesh with greater manifoldness.

Extensive comparisons to state-of-the-art mesh generation methods for physically based rendering and simulation (see supplementary video), highlighting the advantage of NMF’s manifoldness.

New metrics to evaluate manifoldness of 3D meshes and demonstration of applications to single-view reconstruction, 3D deformation, global parameterization and correspondence.

Related Work

Existing learning based mesh generation methods, while yielding impressive geometric accuracy, do not satisfy one or more manifoldness conditions (Fig. 1b). While indirect approaches 6, 7, 8, 16, 17, 18 suffer from the non-manifoldness of the marching cube algorithm 19, direct methods 2, 3, 4, 5 are faced with the regularizer’s dilemma on the trade-off between geometric accuracy and higher manifoldness, illustrated in Fig. 3 and discussed in Sec. 1.

Indirect approaches predict the 3D geometry as either a distribution of voxels 20, 21, 22, 23, 24, 25, 26, 27, point clouds 7, 28 or an implicit function representing signed distance from the surface 8, 16, 18. Both voxel and point set prediction methods struggle to generate high resolution outputs which later makes the iso-surface extraction tools ineffective or noisy 3. Implicit methods feed a neural network with a latent code and a query point, encoding the spatial coordinates 8, 16, 18 or local features 29, to predict the TSDF value 16 or the binary occupancy of the point 8, 18. However, these approaches are computationally expensive since in order to get a surface from the implicit function representation, several thousands of points must be sampled. Moreover, for shapes such as chairs that have thin structures, implicit methods often fail to produce a single connected component.

All the above methods depend on the marching cube algorithm 19 for iso-surface extraction. While marching cubes can be applied directly to voxel grids, point clouds first regress the iso-surface using surface normals. Implicit function representations must regress TSDF values per voxel and then perform extensive query to generate iso-surface based on a threshold τ\tau. This is used to classify grid vertices vi∈Vv_{i}\in\mathcal{V} as ‘inside’ (TSDF(vi)≤τTSDF(v_{i})\leq\tau) and ‘outside’ (TSDF(vi)≥τTSDF(v_{i})\geq\tau). For each voxel, based on the arrangement of its grid vertices, marching cubes 19, 30, 31, 32 follows a lookup-table to find a triangle arrangement. Since this rasterization of iso-surface is a purely local operation, it often leads to ambiguities 30, 31, 32, resulting in meshes being non-manifold.

Direct Mesh Prediction

A mesh based representation stores the surface information cheaply as list of vertices and faces that respectively define the geometric and topological information. Early methods of mesh generation relied on predicting the parameters of category based mesh models 33, 34, 35. While these methods output manifold meshes, they work only for object category with available parameterized manifold meshes. Recently, meshes have been successfully generated for a wide class of categories using topological priors 3, 4. Deep networks are used to update the vertices of initial mesh to match that of the final mesh. AtlasNet 3 uses Chamfer distance applied on the vertices for training, while Pixel2Mesh 4 uses a coarse-to-fine deformation approach using vertex Chamfer loss. However, using a point set training scheme for meshes leads to severe topological issues and produced meshes are not manifold. Some recent works have proposed to use mesh regularizers like Laplacian 2, 4, 5, 36, edge lengths 2, 5, normal consistency 2 or pose it as a linear programming problem 37 to constrain the flexibilty of vertex predictions. They suffer from the regularizer’s dilemma discussed in Fig. 3, as better geometric accuracy comes at a cost of manifoldness.

In contrast to the above approaches, the proposed NMF achieves high resolution meshes with a high degree of manifoldness across a wide variety of shape categories. Similar to previous approaches 3, 4, 5, an initial ellipsoid is deformed by updating its vertices. However, instead of using explicit mesh regularizers, NMF uses NODE blocks to learn the diffeomorphic flow to implicitly discourage self-intersections, maintain the topology and thereby achieve better manifoldness of generated shape. The method is end-to-end trainable without requiring any post-processing.

Neural Mesh Flow

We now introduce Neural Mesh Flow (Fig 4), which learns to auto-encode 3D shapes. NMF broadly consists of four components. First, the target shape MT\mathcal{M}_{T} is encoded by uniformly sampling NN points from its surface and feeding them to a PointNet 38 encoder to get the global shape embedding zz of size kk. Second, NODE blocks diffeomorphically flow the vertices of template sphere towards target shape conditioned on shape embedding zz. Third, the instance normalization layer performs non-uniform scaling of NODE output to ease cross-category training. Finally, refinement flows provide gradual improvement in quality. We start with a discussion of NODE and its regularizing property followed by details on each component.

Diffeomorphic Conditional Flow.

The standard NODE 1 formulation cannot be used directly for the task of 3D mesh generation since they lack any means to feed in shape embedding and are therefore restricted to learning a few shape. A naive way would be to concatenate features to point coordinates like is done with traditional MLPs 4, 5 but this destroys the shape regularization properties due to several augmented dimensions 9, 10. Our key insight is that instead of a fixed NODE dynamics fΘf_{\Theta} we can use a family of dynamics fΘ∣zf_{\Theta|z} parameterized by zz while still retaining the uniqueness property as long as zz is held constant for the purpose of solving IVP with initial conditions {x0,xT}\{x_{0},x_{T}\}.

The objective of conditional flow (NODE Block) therefore is to learn a mapping FΘ∣zF_{\Theta|z} (1) given the shape embedding zz and initial values {(pIi,pOi)∣pIi∈MI,pOi∈MO}\{(p^{i}_{I},p^{i}_{O})|p^{i}_{I}\in\mathcal{M}_{I},p^{i}_{O}\in\mathcal{M}_{O}\} where MI,MO\mathcal{M}_{I},\mathcal{M}_{O} are respectively the input and output point clouds.

Instance Normalization.

Overall Architecture.

A single NODE block is often not sufficient to get desired quality of results. We therefore stack up two NODE blocks in a sequence followed by an instance normalization layer and call the collection a deformation block. While a single deformation block is capable of achieving reasonable results (as shown by Mp0\mathcal{M}_{p0} in Fig. 4) we get further refinement in quality by having two additional deformation blocks. Notice how the Mp1\mathcal{M}_{p1} has a better geometric accuracy than Mp0\mathcal{M}_{p0} and Mp2\mathcal{M}_{p2} is sharper compared to Mp1\mathcal{M}_{p1} with additional refinement. We report the geometric accuracy, manifoldness and inference time for different amounts of refinement in Table 1. The reported quantities are averaged over the 11 Shapnet categories (this excludes watercraft and lamp where NMF struggles with thin structures). For details on per category ablation, please see the supplementary material. To summarize, the entire NMF pipeline can be seen as three successive diffeomorphic flows {FΘ∣z0,FΘ∣z1,FΘ∣z2}\{F^{0}_{\Theta|z},F^{1}_{\Theta|z},F^{2}_{\Theta|z}\} of the initial spherical mesh to gradually approach the final shape.

Loss Function.

We compute chamfer distances Lp1,Lp2\mathcal{L}_{p1},\mathcal{L}_{p2} for meshes after deformation blocks FΘ∣z1F^{1}_{\Theta|z} and FΘ∣z2F^{2}_{\Theta|z}. For meshes generated from FΘ∣z0F^{0}_{\Theta|z} we found that computing chamfer distance Lv\mathcal{L}_{v} on the vertices gave better results since it encourages predicted vertices to be more uniformly distributed (like points sampled from target mesh). We thus arrive at the overall loss function to train NMF.

Here we take w0=0.1,w1=0.2,w3=0.7w_{0}=0.1,w_{1}=0.2,w_{3}=0.7 so as to enhance mesh prediction after each deformation block. The adjoint sensitivity 48 method is employed to perform the reverse-mode differentiation through the ODE solver and therefore learn the network parameters Θ\Theta using the standard gradient descent approaches.

Dynamics Equation.

Implementation Details

For the Neural Mesh Flow architecture, both the mesh vertices and NODE dynamics operate in n=3n=3 dimensions. We uniformly sample N=2520N=2520 from the target mesh and using PointNet 38 encoder, get a shape embedding zz of size k=1000k=1000. During training, the NODE is solved with a tolerance of 1e−51e^{-5} and interval of integration set to t=0.2t=0.2 for deforming an icosphere with 622 vertices. The integration time was empirically determined to be large enough for flow to work, but not too large to cause overfitting. At test time, we use an icosphere of 2520 vertices and tolerance of 1e−51e^{-5}. We train NMF for 125 epochs using Adam 50 optimizer with a learning rate of 10−510^{-5}, weight decay of 0.950.95 after every 250250 iterations and a batch size of 250, on 5 NVIDIA 2080Ti GPUs for 2 days. For single view reconstruction, we train an image to point cloud predictor network with pretrained ResNet encoder of latent code 1000 and a fully-connected decoder with size 1000,1000,3072 with relu non-linearities. The point predictor is trained for 125 epochs on the same split as NMF auto-encoder.

Experiments

In this section we show qualitative and quantitative results on the task of auto-encoding and single view reconstruction of 3D shapes with comparison against several state of the art baselines. In addition to these tasks, we also demonstrate several additional features and applications of our approach including latent space interpolation texture mapping, consistent correspondence and shape deformations in the supplementary material.

We evaluate our approach on the ShapeNet Core dataset 51, which consists of 3D models across 13 object categories which are preprocessed with 52 to obtain manifold meshes. We use the training, validation and testing splits provided by 6 to be comparable to other baselines. We use rendered views from 6.

Evaluation criteria

We detect non-manifold vertices (Fig. 2(b)) and edges (Fig. 2(a)) using 53 and report the metrics ‘NM-vertices’, ‘NM-edges’ respectively as the ratio(×105\times 10^{5}) of number of non-manifold vertices and edges to total number of vertices and edges in a mesh. To calculate non-manifold faces, we count number of times adjacent face normals have a negative inner product, then the metric ‘NM-Faces’ is reported as its ratio(%) to the number of edges in the mesh. To calculate the number of instances of self-intersection, we use 54 and report the ratio(%) of number of intersecting triangles to total number of triangles in a mesh. Only the mean over all ShapeNet categories are reported in this paper with category specific details can be found in the supplementary.

For qualitative evaluation, we render predicted meshes via a physically based renderer 55 with dielectric and metallic materials to highlight artifacts due to non-manifoldness. While we render NMF meshes directly, other methods render poorly due to non-manifoldness and are smoothed prior to rendering to obtain better visualizations. Please see supplementary for further visualizations.

Baselines

We compare with official implementations for Pixel2Mesh 2, 4, MeshRCNN 2 and AtlasNet 3. We use pretrained models for all these baselines motioned in this paper since they share the same dataset split by 6. We use the implementation of Pixel2Mesh provided by MeshRCNN, as it uses a deeper network that outperforms the original implementation. We also consider AtlasNet-O which is a baseline proposed in 3 that uses patches sampled from a spherical mesh, making it closer to our own choice of initial template mesh. We also create a baseline of our own called NMF-M, which is similar in architecture to NMF but trained with a larger icosphere of 2520 vertices, leading to slight differences in test time performance. To account for possible variation in manifoldness due to simple post processing techniques, we also report outputs of all mesh generation methods with 3 iterations of Laplacian smoothing. Further iterations of smoothing lead to loss of geometric accuracy without any substantial gain in manifoldness. We also compare with occupancy networks 8, a state-of-the-art indirect mesh generation method based on implicit surface representation. We compare with several variants of OccNet based on the resolution of Multi Iso-Surface Extraction algorithm 8. To this end, we create OccNet baselines OccNet-1, OccNet-2 and OccNet-3 with MISE upsampling of 1, 2 and 3 times respectively. For fair comparison to other baselines, we use OccNet’s refinement module to output its meshes with 5200 faces.

Auto-encoding 3D shapes

We now evaluate NMF’s ability to generate a shape given an input 3D point cloud and compare against AtlasNet 3 and AtlasNet-O3 in Table 2. We note that NMF outperforms AtlasNet in terms of manifoldness with 20 times less self-intersections. NMF generates meshes with a higher normal consistency, leading to more realistic results in simulations and physically-based rendering. All the three methods have manifold vertices. While both NMF and AtlasNet-O have no non-manifold edges, AtlasNet yields a constant value of 74007400 due to its constituent 25 non-manifold open templates. Visualizations in Fig. 6 show severe self-intersections and flipped normals for AtlasNet baselines which are absent for NMF. This leads to NMF giving more realistic physically based rendering results. Note the reflection of red box and green ball through NMF mesh, which are either distorted or absent for AtlasNet. The blue ball’s reflection on conductor’s surface is closer to ground truth for NMF due to higher manifoldness.

Single-view reconstruction

We evaluate NMF for single-view reconstruction and compare against state-of-the-art methods in Table 3. We note significantly lower self-intersections for NMF compared to the best baseline even after smoothing. Our method again results in fewer than 50% non-manifold faces compared to the best baseline. NMF-M also gets the highest normal consistency performance. Due to the cubify step as part of the MeshRCNN 2 pipeline which converts a voxel grid into a mesh, the method has several non-manifold vertices and edges compared to deformation based methods Pixel2Mesh 4, 2, AtlasNet-O 3 and NMF. AtlasNet suffers from the most number of non manifold edges, almost 100 times that of MeshRCNN. We note that MeshRCNN2 better performance in Chamfer Distance come at a cost of other metrics. We qualitatively show the effects of non-manifoldness in Figure 7 and supplementary material. We observe that for dielectric material (second row), NMF is able to transmit background colors closest to the ground truth, whereas other baselines only reflect the white sky due to the presence of flipped normals.

Soft body simulation (watch supplementary video for better understanding)

To further demonstrate the usefulness of manifoldness, we qualitatively evaluate predicted meshes with soft body simulation in Fig. 8. Here we simulate dropping meshes on the floor using Blender 56, with settings pull=0.9, push=0.9, bending=10pull=0.9,\ push=0.9,\ bending=10 to represent a rubber-like material. We note that AtlasNet 3 breaks into its constituent 25 independent meshlets upon hitting the floor. This behaviour is expected of methods that predict shapes as a set of n-connected components 57. Both Pixel2Mesh 4 and AtlasNet-O 3 yield unrealistic simulations due to the presence of severe self-intersection artifacts. We note that MeshRCNN 2 suffers from over-bounciness due to non-manifoldness and poor normal consistency. In contrast, NMF yields simulations with properties that are closest to the ground truth.

D printing

We now show in Fig. 9 a few renders of a 3D printed shape predicted by NMF using image from Figure 1. Since NMF predicts a manifold mesh, we can 3D print the predicted shapes without any post processing or repair efforts, obtaining satisfactory printed products.

Comparison with implicit representation method

We evalute NMF against state-of-the-art indirect mesh generation method OccNet 8 for the task of single view reconstruction in Table 4. We observe that NMF outperforms the best baseline OccNet-3 in terms of geometric accuracy. This is primarily because NMF predicts a singly connected mesh object as opposed to OccNet which leads to several disconnected meshes. Moreover, due to the limitations imposed by the marching cubes algorithm discussed in Section 2, OccNet-1,2,3 have several non-manifold vertices and edges where as by construction, NMF doesn’t suffer from such limitation. An example of non-manifold edge is shown in Fig. 10. For sake of completeness, we also show the mesh generated by MeshRCNN 2 that suffers from non-manifold vertices and edges. NMF is also competitive with OccNet in terms of self-intersections since both methods become practically intersection-free with Laplacian smoothing . While OccNet outperforms NMF in terms of non-manifold faces, we argue that this comes at a cost of higher inference time. For reference, the fastest version of OccNet has comparable non-manifold faces and self-intersections but suffers relatively in terms of other metrics.

Conclusions

In this paper, we have considered the problem of generating manifold 3D meshes using point clouds or images as input. We define manifoldness properties that meshes must satisfy to be physically realizable and usable in practical applications such as rendering and simulations. We demonstrate that while prior works achieve high geometric accuracy, such manifoldness has previously not been sought or achieved. Our key insight is that manifoldness is conserved under a diffeomorphic flow that deforms a template mesh to the target shape, which can be modeled by exploiting properties of Neural ODEs 1. We design a novel architecture, termed Neural Mesh Flow, composed of deformation blocks with instance normalization and refinement flows, to achieve manifold meshes without any post-processing. Our results in the paper and supplementary material demonstrate the significant benefits of NMF for real-world applications.

Broader Impact

The broader positive impact of our work would be to inspire methods in computer graphics and associated industries such as gaming and animation, to generate meshes that require significantly less human intervention for rendering and simulation. The proposed NMF method addresses an important need that has not been adequately studied in a vast literature on 3D mesh generation. While NMF is a first step in addressing that need, it tends to produce meshes that are over-smooth (also reflected in other methods sometimes obtaining greater geometric accuracy), which might have potential negative impact in applications such as manufacturing. Our code, models and data will be publicly released to encourage further research in the community.

Acknowledgement

We would like to thank Krishna Murthy Jatavallabhula and anonymous reviewers for valuable discussions and feedback. We would also like to thank Pengcheng Cao with UCSD CHEI for providing 3D printed models and Shreyam Natani for helping with Blender.

References