Deep Implicit Templates for 3D Shape Representation

Zerong Zheng, Tao Yu, Qionghai Dai, Yebin Liu

Introduction

Representing 3D objects effectively and efficiently in neural networks is fundamental for many tasks in computer vision, including 3D model reconstruction, matching, manipulation and understanding. In the pioneering studies, researchers have adopted various traditional geometry representations, including voxel grids , point clouds and meshes . In the past several years, deep implicit functions (DIFs) have been proposed as an alternative . Compared to traditional representations, DIFs show expressive and flexible capacity for representing complex shapes and fine geometric details, even in challenging tasks like human digitization .

Unfortunately, the implicit nature of DIFs is also its Achilles’ Heels: although DIFs are good at approximating individual shapes, they provide no information about the relationship between two different ones. One can easily establish vertice-to-vertice correspondences between two shapes when using mesh templates , but that is difficult in DIFs. The lack of semantic relationship in DIFs poses significant challenges for using DIFs in downstream applications such as shape understanding and editing.

To overcome this limitation, we propose Deep Implicit Templates, a new way to interpret and implement DIFs. The key idea is to decompose a conditional deep implicit function into two components: a template implicit function and a conditional spatial warping function. The template implicit function represents the “mean shape” for a category of objects, while the spatial warping function deforms the template implicit function to form specific object instances. On one hand, as both the template and the warping field are defined in an implicit manner, the advantages of deep implicit representations (compactness and efficiency) are preserved. On the other hand, with the template implicit function as an intermediate shape, the warping function automatically establishes dense correspondences across different object.

More recently, some techniques use a set of primitives to represent 3D shapes in order to capture structure-level semantics . The primitives can be either manually defined or learned from data . We emphasize that our method is essentially different from them in two ways. First, our method decomposes the implicit representations into a template implicit function and a continuous warping field. Compared to element-based methods, our decomposition not only provides a complete, global template for the training data, enabling many applications such as uv mapping and keypiont labeling, but also makes the latent shape space more interpretable as we can inspect how shapes deform. Second, our method directly builds up accurate correspondences in the whole 3D space, while element-based methods rely on interpolation or feature matching to compute dense correspondences. Overall, our method provides more flexibility and scalability to control the template and/or its deformation: one can, for example, replace the template in our framework with a custom designed one without losing the representation power, or apply additional semantic constraints on the spatial deformation for specific problems such as dynamic human modeling (Sec.6).

Training deep implicit templates is not straight-forward, because we have no access to either the ground-truth mapping between the templates and shape instances, or dense correspondence annotations across different shapes. Our ultimate goal is to make deep implicit templates an effective representation that can: 1) represent training shapes accurately, 2) establish plausible correspondences across shapes and 3) generalize to unseen data. However, without proper design and regularization, the network may not be able to learn such a representation in an unsupervised manner. We make several technical contributions to address these challenges. In terms of network architecture, we propose Spatial Warping LSTM, which decomposes the conditional spatial warping into multi-step point-wose transformation, guaranteeing the generalization capacity and the representation power of our warping function. In addition, we introduce a progressive reconstruction loss for our Spatial Warping LSTM, which further improves the reconstruction accuracy. Two-level regularization is also proposed to obtain plausible templates with accurate correspondences in an unsupervised manner. As shown in the experiments, our method can learn a plausible implicit template for a set of shapes, with conditional warping fields that accurately represent shapes while establishing dense correspondences among them without any supervision (See Fig.1). Overall, the proposed Deep Implicit Templates significantly expands the capability of DIFs without losing its advantages, making it a more effective implicit representation for 3D learning tasks. Code is available at https://github.com/ZhengZerong/DeepImplicitTemplates.

Related Work

Deep Implicit Functions (DIFs). Implicit functions represent shapes by constructing a continuous volumetric field and embedding shapes as its iso-surface . In recent years, implicit functions have been introduced into neural networks and show promising results. For example, DeepSDF proposed to learn an implicit function where the network output represents the signed distance of the point to its nearest surface. Other approaches defined the implicit functions as 3D occupancy probability and turned shape representation into a point classification problem . Some latest studies proposed to blend multiple local implicit functions in order to generalize to more complex scenes as well as to capture more geometric details . The training loss used in DeepSDF is also improved for more accurate reconstruction . DualSDF extended DeepSDF by introducing a coarse layer to support shape manipulation. Occupancy Flow , from another aspect, extended 3D occupancy functions into 4D domains, but this method is restricted to represent temporally continuous 4D sequences. Overall, implicit functions are promising for representing complex shapes and detailed surfaces, but it remains difficult to reason dense correspondences between different shapes represented by DIFs. In contrast, our method and other concurrent works overcomes this limitation and expands the capability of DIFs by introducing dense correspondences across shapes into DIFs.

Elementary Structures. Elementary representations, which aim to describe complex shapes uing a collection of simple shape elements, have been extensively studied for many years in computer vision and graphics . In this direction, previous methods usually require complicated non-convex optimization for primitive fitting. In order to improve the efficiency and effectiveness, various deep learning techniques were adopted, such as recurrent networks , differential model estimation and unsupervised training losses . Some recent approaches introduced more complex shape elements, such as multiple charts , part segmentations , axis-aligned 3D Gaussians , local deep implicit functions , superquadrics , convex decomposition , box-homeomorphic parts or learnable shape primitives . Although elementary representation is compact, it is challenging to produce consistent fitting across different shapes. In addition, local elements cannot be used in many tasks (e.g., shape completion) due to the lack of global knowledge.

Template Learning. Fitting global shape templates instead of local primitives has also attracted a lot of research efforts, as mesh-based templates are able to efficiently represent similar shapes such as articulated human bodies . They are very popular in many data-driven 3D reconstruction studies. For example, articulated human body templates, such as SMPL , are now widely used in a lot of human modeling studies . Ellipsoid meshes as a much simpler template can also be deformed to represent various objects . Given a predefined template, some recent techniques like 3D-Coded learned to perform shape matching in unsupervised manner. Although using mesh-based templates is convenient, mesh-based templates are unable to deal with topological changes and require dense vertices to recover surface details . Moreover, when representing shapes that are highly different from the template, mesh-based templates suffer from deformation artifacts (e.g., triangle intersection, extreme distortion, etc.) due to large number of degrees of freedom for mesh vertex coordinates. In contrast, the template in our method is defined in an implicit manner, thus possessing more powerful representation capacity than mesh-based templates. To the best of our knowledge, our work is the first one to learn implicit function-based templates for a collection of shapes.

Overview

Our Deep Implicit Template representation is designed on the basis of DeepSDF , which is a popular DIF-based 3D shape representation. In this section, we first review DeepSDF for clarity and then describe the overall framework of our approach.

where c∈X\bm{c}\in\mathcal{X} is the condition variable that encodes the shape of a specific object and can be custom designed in accordance of applications . With this SDF represented by F\mathcal{F}, the object surface can be extracted using Marching Cube . In DeepSDF , the condition variable c\bm{c} is a high-dimensional latent code and each shape instance has a unique code. All latent codes are firstly initialized with Gaussian noise and then optimized in parallel with network training.

2 Deep Implicit Templates

In DeepSDF, the shape variance is directly represented by the changes of SDFs themselves. Different from this formulation, we think that, given a category of shapes represented by SDFs, their shape variance can be reflected by the differences of these SDFs relative to a template SDF that captures their common structure. This key idea leads to our formulation of Deep Implicit Templates, which decomposes the conditional signed distance function F\mathcal{F} into F=T∘W\mathcal{F}=\mathcal{T}\circ\mathcal{W}, i.e.,

Intuitively, W\mathcal{W} is a conditional spatial warping function that warps the input points according to the latent code c\bm{c}, while the function T\mathcal{T} itself is an implicit function representing a common SDF which is irrelevant to c\bm{c}. Therefore, for a set of objects, the shape represented by T(⋅)\mathcal{T}(\cdot) can be regarded as their common template SDF; in the following context we call it an implicit template. To query the signed distance at p\bm{p} for a specific object defined by c\bm{c}, the spatial warping function W\mathcal{W} first transforms p\bm{p} to its canonical position in the implicit template, followed by T\mathcal{T} querying its signed distance. In other words, the implicit template is warped according to the latent codes to model different SDFs. In the strict sense of the term, a warping of an SDF is not an SDF in general. However, we adopt two constraints to make sure that the warped SDF can still approximate the target SDF: (1) we use truncated SDF that only varies in a small band near the surface, and (2) we normalize all the meshes into the same scale. Therefore, we loosely use the term ”SDF” in this paper and regard a warping of an SDF as another SDF.

Compared to the original formulation in Eqn.(1), the main advantage of the decomposition in Eqn.(2) is that it naturally induces correspondences between the implicit template and object instances, and accordingly, correspondences across different object instances. As shown in Sec.6, this feature offers more possibility for applying deep implicit functions in many applications.

However, implementing and training the decomposed network is not straight-forward. Specifically, without proper design and regularization, the network tends to overfit to a complicated transformer with an over-simplified implicit template, which further result in inaccurate correspondences. Our goal is to learn an optimal template that can represent the common structure for a set of objects, together with a spatial transformer that establishes accurate dense correspondences between the template and the object instances. Moreover, the learned Deep Implicit Templates should also preserve the representation power and the generalization capacity of DeepSDF, and thus support mesh interpolation and shape completion. In Sec.4 we will discuss how we achieve these goals.

Methodology

Similar to DeepSDF , we implement our implicit template, T\mathcal{T} in Eqn.(2), using a fully-connected network. For the spatial warping function W\mathcal{W}, we empirically found that an MLP implementation leads to unsatisfactory results (Sec.5.4). To deal with this challenge, we introduce a Spatial Warping LSTM, which decomposes the spatial transformation for a point p\bm{p} into multi-step point-wise transformations:

where ⊙\odot means element-wise product and p(i)=p\bm{p}^{(i)}=\bm{p}. We iterate this process in Eqn.(3-4) for SS steps (in all experiments we set S=8S=8), and the point coordinate at the final step yields the output of the warping function W\mathcal{W}:

This multi-step formulation takes inspiration from the iterative error feedback (IEF) loop , where progressive changes are made recurrently to the current estimate. The difference is that our LSTM implementation allows us to aggregate information from multiple previous steps while IEF makes independent estimations in each step. The network architecture is illustrated in Fig.2.

2 Network Training

The training loss for our network is composed of two components, a reconstruction loss and a regularization loss:

As we decompose the spatial warping field into multiple steps of point-wise transformation using our Spatial Warping LSTM, we expect that the network learns a progressive warping function, which starts from obtaining smooth shape approximations and then gradually strives for more local details. To this end, we take inspiration from the shape curriculum in and impose a progressive reconstruction loss upon the outputs of our network for different numbers of warping steps, aiming that the network recovers more geometric details when taking more transformation steps. Mathematically, the loss term for the outputs with ss warping steps is defined as:

where p(s)\bm{p}^{(s)} is defined as in Eqn.4, vk,iv_{k,i} the ground-truth SDF value of pi\bm{p}_{i} for the kk-th shape, NN the number of SDF samples for one shape and KK the number of shapes. Lϵ,λ(⋅,⋅)L_{\epsilon,\lambda}(\cdot,\cdot) is a curriculum training loss with ϵ\epsilon and λ\lambda controlling its smoothness level and hard example weights; please refer to for detail definition. In Tab.1 we present the parameter settings for different warping steps. Our progressive reconstruction loss is the sum of all levels of Lrec(s)\mathcal{L}_{rec}^{(s)}:

2.2 Regularization Loss

Ideally, the spatial warping function is supposed to establish plausible correspondences between the template and the object instances, while the implicit field template should capture the common structure for a set of objects. To achieve this goal, we introduce two regularization terms on point-wise and point-pair levels for the warping function.

Point-wise regularization. We assume that all meshes are normalized to a unit sphere and aligned in a canonical pose. Therefore, we introduce a point-wise regularization loss that is used to constrain the position shifting of points after warping. It is defined as:

where h(⋅)h(\cdot) is the Huber kernel with its hyper-parameter δh\delta_{h}.

Point pair regularization. Although spatial distortion is inevitable during template deformation, extreme distortions should be avoided. To this end, we introduce a novel regularization loss for the point pairs in each shape:

where Δp=W(p,c)−p\Delta\bm{p}=\mathcal{W}(\bm{p},\bm{c})-\bm{p} is the position shift of p\bm{p} and ϵ\epsilon is a parameter controlling the distortion tolerance. Intuitively, if ϵ=0\epsilon=0, Lpp\mathcal{L}_{pp} is reduced to a strict smoothness constraint enforcing translation consistency on neighboring points. We set ϵ=0.5\epsilon=0.5 in our experiments to construct a relaxed formulation of smoothness loss, which is found important to prevent shape structures from collapsing (Sec.5.4).

Our final regularization loss is defined as:

where the last term is the same magnitude constraint on the latent codes as in DeepSDF .

Experiments

We train Deep Implicit Templates on ShapeNet dataset , following for data pre-processing. In Sec.5.2, we first show that our model is capable of representing shapes in high quality with dense correspondences. For comparison in Sec.5.3, we select several strong baselines that use different types of representations: DeepSDF (deep implicit functions) , SIF (structured implicit functions) , AtlasNet (mesh parameterization) , PointFlow (point clouds) and DualSDF (two-level implicit functions) . We mainly evaluate them and our method in terms of reconstruction, correspondence and interpolation. Finally we evaluate our technical contributions in Sec.5.4. More results, experiments and details are presented in the supplemental materials.

2 Results

We demonstrate some results of our approach in Fig.1, Fig.3 and Fig.4. In Fig.1, we present the learned template shapes as well as the dense correspondences between the templates and object instances. The results show that our method can learn to abstract a template that captures the common structure for a collection of shapes, and also establish plausible dense correspondences that indicates the semantic relationship across different shapes. The results also show that our representations can deal with large deformations and describe objects with completely different structures. In Fig.3, we provide a more clear landscape on how the templates deforms to describe different objects. Fig.4 further presents the representation power of our method as well as the accurate correspondences established by our method.

3 Comparison

Reconstruction. We report the reconstruction results for known and unknown shapes (i.e., shapes belonging to the train and test sets) in Tab.2. We use two metrics for accuracy measurement, i.e., Chamfer Distance (CD) and Earth Mover Distance (EMD). As the numeric results show, our method is able to achieve comparable reconstruction performance when compared to state-of-the-art methods. Furthermore, we observe that our method is able to produce slightly better results for object classes that have a strong template prior (e.g., cars, airplanes and sofas), but degenerates for objects that vary enormously in shape structures (e.g., chairs).

Interpolation. Similar to DeepSDF, our learned shape embedding is continuous and supports shape interpolation in the latent space (Fig.5). Note that unlike DeepSDF that directly interpolates signed distance fields, our representation actually interpolates the spatial warping fields because the template is independent from the latent code. The results in Fig.5 (bottom row) show that the generative capability of DeepSDF is well preserved in our representation.

Correspondences. With the dense correspondence provided by our method, one can perform keypoint detection on point clouds by transferring the keypoint labels in training shapes to test ones. Therefore, we use this as a surrogate to evaluate the accuracy of correspondences, as there is no large-scale dense correspondence annotation available. We collect keypoint annotations from KeypointNet and use the percentage of correct keypoints (PCK) as the metric for our experiments. Tab.3 presents the PCK scores under different error distance threshold, showing that our method outperforms other representations. In Fig.6, we can see that, although our network has no access to the keypoint annotations or correspondence annotations during training, it still establishes accurate corresponding relationship across different objects in the same category.

4 Ablation Study

Spatial Warping LSTM. We evaluate our choice of spatial warping LSTM by replacing it with an MLP-based implementation. The numbers of the MLP neurons are (259,512,512,512,512,512,6)(259,512,512,512,512,512,6). The comparison on interpolation capacity of different network architectures is presented in Fig.7. We empirically find that the MLP implementation is prone to over-fitting and cannot generalize well for latent space interpolation. We think that it is because the large non-uniform deformations between the template and object instances are more easily to learn in a gradual manner, but much harder when an MLP try to learn the transformation in a single step.

Point-pair Regularization. To evaluate our point-pair regularization loss, we train a baseline network without it. As shown in Fig.8, without the point-pair regularization, the networks tends to learn over-simplified templates. Although this phenomenon has no impact on the reconstruct accuracy, it leads to inaccurate correspondences since the detailed structures like plane wings and sofa backrests collapse into tiny regions on the templates, as proven in Tab.3 (last two rows). Therefore, the proposed two-level regularization is essential in our representation of deep implicit templates.

Extensions and Applications

To prove the flexibility and the potential of our Deep Implicit Templates, we show how our method can be easily extended when more constraints are available.

User-defined templates. For example, we can replace the learned template in our method with a manually-specified one. In this case, we just need to construct an additional loss to train the template implicit function T\mathcal{T}. Specifically, the loss is defined as:

where (pi,si)(\bm{p}_{i},s_{i}) are the SDF training pairs extracted from the specified template model. In Fig.9, we demonstrate an example where we specify SMPL model as the template and train our network to model various clothed humans. With the correspondences we can transfer skinning weights from SMPL model to clothed humans, allowing animation for various clothed human models as in Fig.10.

Correspondence annotations. Furthermore, although our method is designed to learn dense correspondence without any supervision, our method can incorporate with sparse/dense correspondence annotations when provided. To this end, we just need to introduce an additional loss to enforce the known correspondences:

2 Applications

To show the effectiveness and practicability of our method, we investigate how Deep Implicit Templates can be used in applications. With the dense correspondences between the template and the object instances, we can transfer mesh attributes, such as texture (Fig.12(a)), or mesh stretching operation (Fig.12(b)) to multiple objects. Our method can also be used in the applications that have been demonstrated in DeepSDF such as shape completion.

Conclusion

We have presented Deep Implicit Templates, a new 3D shape representation that factors out implicit templates from deep implicit functions. To the best of our knowledge, this is the first method that explicitly interprets the latent space for a class of objects as a meaningful pair: an implicit template with its conditional deformations. Several technical contributions on network architectures and training losses are proposed to not only recover accurate dense correspondences without losing the properties of deep implicit functions, but also learn a plausible template capturing common geometry structures. The experiments further demonstrate some promising applications of Deep Implicit Templates. To conclude, we have demonstrated that a semantic interpretation of deep implicit functions leads to a more powerful representation, which we believe is an inspiring direction for future research.

Acknowledgement This paper is supported by the NSFC No.61827805.

References