Learning shape correspondence with anisotropic convolutional neural networks
Davide Boscaini, Jonathan Masci, Emanuele Rodolà, Michael M. Bronstein
Introduction
In geometry processing, computer graphics, and vision, finding intrinsic correspondence between 3D shapes affected by different transformations is one of the fundamental problems with a wide spectrum of applications ranging from texture mapping to animation [KZHCO10]. Of particular interest is the setting in which the shapes are allowed to deform non-rigidly. Recently, the emergence of 3D sensing technology has brought the need to deal with acquisition artifacts, such as missing parts, geometric, and topological noise, as well as matching 3D shapes in different representations, such as meshes and point clouds. The main topic of this paper is establishing dense intrinsic correspondence between non-rigid shapes in such challenging settings.
Traditional correspondence approaches try to find a point-wise matching between (a subset of) the points on two or more shapes. Minimum-distortion methods establish the matching by minimizing some structure distortion, which can include similarity of local features [OMMG10, ASC11, ZBVH09], geodesic [MS05, BBK06, CK15] or diffusion distances [CLL∗05], or a combination thereof [TKR08], or higher-order structures [Z∗10]. Typically, the computational complexity of such methods is high, and there have been several attempts to alleviate the computational complexity using hierarchical [SY11] or subsampling [T∗11] methods. Several approaches formulate the correspondence problem as quadratic assignment and employ different relaxations thereof [Ume88, LH05, RBA∗12, CK15, KKBL15].
Embedding methods try to exploit some assumption on the shapes (e.g. approximate isometry) in order to parametrize the correspondence problem with a small number of degrees of freedom. Elad and Kimmel [EK01] used multi-dimensional scaling to embed the geodesic metric of the matched shapes into a low-dimensional Euclidean space, where alignment of the resulting “canonical forms” is then performed by simple rigid matching (ICP) [CM91, BM92]. The works of [MHK∗08, SK13] used the eigenfunctions of the Laplace-Beltrami operator as embedding coordinates and performed matching in the eigenspace. Lipman et al.[LD11, KLCF10, KLF11] used conformal embeddings into disks and spheres to parametrize correspondences between homeomorphic surfaces as Möbius transformations.
As opposed to point-wise correspondence methods, soft correspondence approaches assign a point on one shape to more than one point on the other. Several methods formulated soft correspondence as a mass-transportation problem [M1́1, SNB∗12]. Ovsjanikov et al.[OBCS∗12] introduced the functional correspondence framework, modeling the correspondence as a linear operator between spaces of functions on two shapes, which has an efficient representation in the Laplacian eigenbases. This approach was extended in several follow-up works [P∗13, KBBV15, ADK16, RCB∗16] .
In the past year, we have witnessed the emergence of learning-based approaches for 3D shape correspondence. The dramatic success of deep learning (in particular, convolutional neural networks [Fuk80, LBD∗89]) in computer vision [KSH12] has lead to a recent keen interest in the geometry processing and graphics communities to apply such methodologies to geometric problems [MBBV15, SMKLM15, WSK∗15, BMM∗15, WHC∗16].
Extrinsic deep learning.
Many machine learning techniques successfully working on images were tried “as is” on 3D geometric data, represented for this purpose in some way “digestible” by standard frameworks. Su et al.[SMKLM15] used CNNs applied to range images obtained from multiple views of 3D objects for retrieval and classification tasks. Wei et al.[WHC∗16] used view-based representation to find correspondence between non-rigid shapes. Wu et al.[WSK∗15] used volumetric CNNs applied to rasterized volumetric representation of 3D shapes.
The main drawback of such approaches is their treatment of geometric data as Euclidean structures (see Figure 1). First, for complex 3D objects, Euclidean representations such as depth images or voxels may lose significant parts of the object or its fine details, or even break its topological structure (in particular, due to computational reasons, the volumetric CNNs [WSK∗15] used a cube, allowing only a very coarse representation of 3D geometry). Second, Euclidean representations are not intrinsic, and vary as the result of pose or deformation of the object. Achieving invariance to shape deformations, a common requirement in many applications, is extremely hard with the aforementioned methods and requires complex models and huge training sets due to the large number of degrees of freedom involved in describing non-rigid deformations.
Intrinsic deep learning
approaches try to apply learning techniques to geometric data by generalizing the main ingredients such as convolutions to non-Euclidean domains. One of the first attempts to learn spectral kernels for shape recognition was done in [ABBK11]. Litman and Bronstein [LB14] learned optimal spectral descriptors that generalize the popular “handcrafted” heat- [SOG09, GBAL09] and wave-kernel signatures [ASC11] and show performance superior to both. Their construction was recently extended in [BMR∗16] using anisotropic spectral kernels, referred to as Anisotropic Diffusion Descriptors (ADD), based on the anisotropic Laplace-Beltrami operator [ARAC14]. A key advantage of the resulting approach is the ability to disambiguate intrinsic symmetries [OSG08], to which most of the standard spectral descriptors are agnostic. Corman et al.[COC14] used descriptor learning in the functional maps framework. Rodolà et al.[RRBW∗14] proposed learning correspondences between non-rigid shapes using random forests applied to WKS descriptors.
The first intrinsic convolutional neural network architecture (Geodesic CNN) was presented in [MBBV15]. GCNN is based on a local intrinsic charting procedure from [KBLB12], and while producing impressive results on several shape correspondence and retrieval benchmarks, has a number of significant drawbacks. First, the charting procedure is limited to meshes, and second, there is no guarantee that the chart is always topologically meaningful.
Another intrinsic CNN construction (Localized Spectral CNN) using an alternative charting technique based on the windowed Fourier transform [SRV13] was proposed in [BMM∗15]. This method is a generalization of a previous work [BZSL14] on spectral deep learning on graphs. One of the key advantages of LSCNN is that the same framework can be applied to different shape representations, in particular, meshes and point clouds. A drawback of this approach is its memory and computation requirements, as each window needs to be explicitly produced.
2 Main contributions
In this paper, we present Anisotropic Convolutional Neural Networks (ACNN), a method for intrinsic deep learning on non-Euclidean domains. Though it is a generic framework that can be used to handle different tasks, we focus here on learning correspondence between shapes.
Our approach is related to two previous methods for deep learning on manifolds, GCNN [MBBV15] and ADD [BMR∗16]. Compared to [BMR∗16], where a learned spectral filter applied to the eigenvalues of anisotropic Laplace-Beltrami operator, we use anisotropic heat kernels as spatial weighting functions allowing to extract a local intrinsic representation of a function defined on the manifold. Unlike ADD, our ACNN is a convolutional neural network architecture. Compared to GCNN, our construction of the “patch operator” is much simpler, does not depend on the injectivity radius of the manifold, and is not limited to triangular meshes. Overall, ACNN combines all the best properties of the previous approaches without inheriting their drawbacks. We show that the proposed framework beats GCNN, ADD, and other state-of-the-art approaches on challenging correspondence benchmarks.
The rest of the paper is organized as follows. In Section 2 we overview the main notions related to spectral analysis on manifolds and define anisotropic Laplacians and heat kernels. In Section 3 we briefly discuss previous approaches to intrinsic deep learning on manifolds and their drawbacks. Section 4 describes the proposed ACNN construction for learning intrinsic dense correspondence between shapes. Section 5 discussed the discretization and numerical implementation. In Section 6 we evaluate the proposed approach on standard benchmarks and compare it to previous state-of-the-art methods. Finally, Section 7 concludes the paper.
Background
Curvature.
Given an embedding of the surface, the second fundamental form, represented as a matrix at each point, describes how the surface locally differs from a plane. The eigenvalues of the second fundamental form are called the principal curvatures; the corresponding eigenvectors called the principal curvature directions form an orthonormal basis on the tangent plane.
Differential operators on manifolds.
where denotes the standard Euclidean gradient acting in the tangent plane. The intrinsic gradient can be interpreted as the direction (tangent vector on ) in which changes the most at point . First-order Taylor expansion takes the form , where the second term is the directional derivative of along the tangent vector .
Given a smooth vector field , the intrinsic divergence is an operator acting on vector fields producing scalar fields, defined as the negative adjoint of the intrinsic gradient operator,
where the area element is induced by the Riemannian metric. Combining the two, we define the Laplace-Beltrami operator as
2 Spectral analysis on manifolds
is a positive-semidefinite operator, admitting real eigendecomposition
The Laplacian eigenfunctions are a generalization of the classical Fourier basis to non-Euclidean domains: a function can be represented as the Fourier series
where the eigenvalues play the role of frequencies (the first eigenvalue corresponds to a constant ‘DC’ eigenvector) and the Fourier coefficients can be interpreted as the Fourier transform of .
Heat diffusion on manifolds
is governed by the heat equation, which, in the simplest case of homogeneous and isotropic heat conductivity properties of the surface, takes the form
where denotes the temperature at point at time , and appropriate boundary conditions are applied if necessary. Given some initial heat distribution , the solution of heat equation (6) at time is obtained by applying the heat operator to ,
where is called the heat kernel, and the above equation can be interpreted as a non-shift-invariant version of convolution.In the Euclidean case, the heat kernel has the form and the solution is given as . In signal processing terms, the heat kernel in the Euclidean case is the impulse response of a linear shift-invariant system. In the spectral domain, the heat kernel is expressed as
Appealing again to the signal processing intuition, acts as a low-pass filter (larger corresponding to longer diffusion results in a filter with a narrower pass band).
Spectral descriptors.
The diagonal of the heat kernel , also known as autodiffusivity, for a range of values , was used in [SOG09, GBAL09] as a local intrinsic shape descriptor referred to as the Heat Kernel Signature (HKS). The Wave Kernel Signature (WKS) [ASC11] uses another set of band-pass filters instead of the low-pass filters . The Optimal Spectral Descriptors (OSD) [LB14] approach suggested to learn a set of optimal tasks-specific filters instead of the “handcrafted” low- or band-pass filters.
3 Anisotropic heat kernels
In a more general setting, the heat equation has the form
where is the thermal conductivity tensor ( matrix) applied to the intrinsic gradient in the tangent plane. This formulation allows modeling heat flow that is position- and direction-dependent (anisotropic). The special case of equation (6) assumes .
Andreux et al.[ARAC14] considered anisotropic diffusion driven by the surface curvature. Boscaini et al.[BMR∗16], assuming that at each point the tangent vectors are expressed w.r.t. the orthogonal basis of principal curvature directions, used a thermal conductivity tensor of the form
where the matrix performs rotation of w.r.t. to the maximum curvature direction , and is a parameter controlling the degree of anisotropy ( corresponds to the classical isotropic case).
Anisotropic Laplacian.
as the anisotropic Laplacian, and denote by its eigenfunctions and eigenvalues (computed, if applicable, with the appropriate boundary conditions). Analogously to equation (8), the anisotropic heat kernel is given by
This construction was used in Anisotropic Diffusion Descriptors (ADD) [BMR∗16] to generalize the OSD approach using anisotropic heat kernels (considering the diagonal and learning a set of optimal task-specific spectral filters replacing the low-pass filters ).
Intrinsic deep learning
This paper deals with the extension of convolutional neural networks (CNN) to non-Euclidean domains. CNN [LBD∗89] have recently become extremely popular in the computer vision community due to a series of successful applications in many classically difficult problems in that domain. A typical convolutional neural network architecture is hierarchical, composed of alternating convolutional-, pooling- (i.e. averaging), linear and non-linear layers. The parameters of different layers are learned by minimizing some task-specific cost function. The key feature of CNNs is the convolutional layer, implementing the idea of “weight sharing”, wherein a small set of templates (filters) is applied to different parts of the data.
In image analysis applications, the input into the CNN is a function representing pixel values given on a Euclidean domain (plane); due to shift-invariance the convolution can be thought of as passing a template across the plane and recording the correlation of the template with the function at that location. One of the major problems in applying the CNN paradigm to non-Euclidean domains is the lack of shift-invariance, making it impossible to think of convolution as correlation with a fixed template: the template now has to be location-dependent.
There have recently been several attempts to develop intrinsic CNNs on non-Euclidean domain, which we overview below. The advantage of intrinsic CNN models over descriptor learning frameworks such as OSD [LB14] or ADD [BMR∗16] is that they accept as input any information on the surface, which can represent photometric properties (texture), some geometric descriptor, motion field, etc. Conversely, learnable spectral descriptors try to learn the best spectral kernel that acts on the Laplacian eigenvalues, thus limited to geometric data of the manifold only.
was introduced by Masci et al.[MBBV15] as a generalization of CNN to triangular meshes based on geodesic local patches. The core of this method is the construction of local geodesic polar coordinates using a procedure previously employed for intrinsic shape context descriptors [KBLB12]. The patch operator in GCNN maps the values of the function around vertex into the local polar coordinates , leading to the definition of the geodesic convolution
which follows the idea of multiplication by template, but is defined up to arbitrary rotation due to the ambiguity in the selection of the origin of the angular coordinate. Taking the maximum over all possible rotations of the template is necessary to remove this ambiguity. Here, and in the following, is some feature vector that is defined on the surface (e.g. texture, geometric descriptors, etc.)
There are several drawbacks to this construction. First, the charting method relies on a fast marching-like procedure requiring a triangular mesh. The method is relatively insensitive to the triangulation [KBLB12], but may fail if the mesh is very irregular. Second, the radius of the geodesic patches must be sufficiently small compared to the injectivity radius of the shape, otherwise the resulting patch is not guaranteed to be a topological disk. In practice, this limits the size of the patches one can safely use, or requires an adaptive radius selection mechanism.
Spectral CNN (SCNN).
Bruna et al.[BZSL14] defined a generalization of convolution in the spectral domain, appealing to the Convolution Theorem, stating that in the Euclidean case, the convolution operator is diagonalized in the Fourier basis. This allows defining a non-shift-invariant convolution as the inverse Fourier transform of the product of two Fourier transforms,
where the Fourier transform is understood as inner products of the function with the Laplace-Beltrami orthogonal eigenfunctions. The filter in this formulation is represented in the frequency domain, by the set of Fourier coefficients . The SCNN is essentially a classical CNN where standard convolutions are replaced by definition (13) and the frequency representations of the filters are learned.
The key drawback of this approach is the lack of generalizability, due to the fact that the filter coefficients depend on the basis ; as a result, applying a filter on two different shapes may produce two different results. Therefore, SCNN can be used to learn on a non-Euclidean domain, but not across different domains. Secondly, the filters defined in the frequency domain lack a clear geometric interpretation. Third, the filters are not guaranteed to be localized in the spatial domain.
Localized Spectral CNN (LSCNN).
Boscaini et al.[BMM∗15] proposed an approach that combines the ideas of GCNN (spatial “patch operator”) and SCNN (frequency-domain filters). The key concept is based on the Windowed Fourier Transform (WFT) [SRV13], a generalization to manifolds of a classical tool of signal processing suggesting to apply frequency analysis in a small window. The WFT boils down to projecting the function on a set of modulated local windows (“atoms”) ,
where are the Fourier coefficients of the window. The WFT can be regarded as a “patch operator”, representing the local values of around a point in the frequency domain.
The main advantage of this approach is that being a spectral construction, it is easily applicable to any representations of the shape (mesh, point cloud, etc.), provided an appropriate discretization of the Laplace-Beltrami operator. Since the patch operator is constructed in the frequency domain using the WFT, there is also no issue related to the topology of the patch like in the GCNN. Yet, unlike the geodesic patches used in GCNN, the disadvantage of LSCNN is the lack of oriented structures which tend to be important in capturing the local context. From the computational standpoint, a notable disadvantage of the WFT-based construction is the need to explicitly produce each window, which result in high memory and computational requirements.
Anisotropic convolutional neural networks
The construction presented in this paper aims at benefiting from all the advantages of the aforementioned intrinsic CNN approaches, without inheriting their drawbacks.
The key idea of the Anisotropic CNN presented in this paper is the construction of a patch operator using anisotropic heat kernels. We interpret heat kernels as local weighting functions and construct
for some anisotropy level . This way, the values of around point are mapped to a local system of coordinates that behaves like a polar system (here denotes the scale of the heat kernel and is its orientation).
Note that unlike the arbitrarily oriented geodesic patches in GCNN, necessitating to take a maximum over all the template rotations (13), in our construction it is natural to use the principal curvature direction as the reference .
Such an approach has a few major advantages compared to previous intrinsic CNN models. First, being a spectral construction, our patch operator can be applied to any shape representation (like LSCNN and unlike GCNN). Second, being defined in the spatial domain, the patches and the resulting filters have a clear geometric interpretation (unlike LSCNN). Third, our construction accounts for local directional patterns (like GCNN and unlike LSCNN). Fourth, the heat kernels are always well defined independently of the injectivity radius of the manifold (unlike GCNN). We summarize the comparative advantages in Table 1.
1 ACNN architecture
Similarly to Euclidean CNNs, our ACNN consists of several layers that are applied subsequently, i.e. the output of the previous layer is used as the input into the subsequent one. The network is called deep if many layers are employed. ACNN is applied in a point-wise manner on a function defined on the manifolds, producing a point-wise output that is interpreted as soft correspondence, as described below. We distinguish between the following types of layers:
The output of the FC layer is optionally passed through a non-linear function such as the ReLU[NH10], .
Intrinsic convolution (ICQ𝑄Q)
layer replaces the convolutional layer used in classical Euclidean CNNs with the construction (16). The IC layer contains filters arranged in banks ( filters in banks); each bank corresponds to an output dimension. The filters are applied to the input as follows,
where are the learnable coefficients of the th filter in the th filter bank.
Softmax
layer is used as the output layer in a particular architecture employed for learning correspondence; it applies the softmax function
to the -dimensional input. The result is a vector that can be interpreted as a probability distribution.
Dropout(π𝜋\pi)
At test time, in order to do inference one would have to integrate over all possible binary masks. However, it has been shown that rescaling the input by the drop probability of the layer,
is a good approximation applicable for real applications.
Batch normalization
layer is another fixed layer recently introduced in [IS15] to reduce training times of very large CNN models. It normalizes each mini-batch during stochastic optimization to have zero mean and unit variance, and then performs a linear transformation of the form
where and are, respectively, the mean and the variance of the data estimated on the training set using exponential moving average; is a small positive constant to avoid numerical errors. After training, one can re-estimate the statistics on the test set or simply keep the training set estimates.
Overall, the ACNN architecture combining several layers of different type, acts as a non-linear parametric mapping of the form at each point of the shape, where denotes the set of all learnable parameters of the network. The choice of the parameters is done by an optimization process, minimizing a task-specific cost. Here, we focus on learning shape correspondence.
2 Learning correspondence
Finding the correspondence in a collection of shapes can be posed as a labelling problem, where one tries to label each vertex of a given query shape with the index of a corresponding point on some common reference shape [RRBW∗14]. Let and denote the number of vertices in and , respectively. For a point on a query shape, the output of ACNN is -dimensional and is interpreted as a probability distribution (‘soft correspondence’) on . The output of the network at all the points of the query shape can be arranged as an matrix with elements of the form , representing the probability of mapped to .
Let us denote by the ground-truth correspondence of on the reference shape. We assume to be provided with examples of points from shapes across the collection and their ground-truth correspondence, . The optimal parameters of the network are found by minimizing the multinomial regression loss
which represents the Kullback-Leibler divergence between the probability distribution produced by the network and the groundtruth distribution .
3 Correspondence refinement
The most straightforward way to convert the soft correspondence produced by ACNN into a point-wise correspondence is by assigning to
The value can be interpreted as the confidence of the prediction: the closer the distribution produced by the network is to a delta-function (in which case ), the better it is.
Partial correspondence.
A similar procedure is employed in the setting of partial correspondence, where instead of the computation of a functional map, we use the recently introduced partial functional map [RCB∗16].
Numerical implementation
where the shear matrix encodes the anisotropic scaling up to an orthogonal basis change. Note that in the isotropic case () we have , such that the -weighted inner product simplifies to the standard inner product .
The discretization of the anisotropic Laplacian takes the form of an sparse matrix . The mass matrix is a diagonal matrix of area elements , where denotes the area of triangle . The stiffness matrix is composed of weights
To obtain the general case , it is sufficient to rotate the basis vectors on each triangle around the respective normal by the angle , equal for all triangles (see Figure 2, red). Denoting by the corresponding rotation matrix, this is equivalent to modifying the -weighted inner product with the directed shear matrix . The resulting weights in equation (34) are thus obtained by using the inner products .
2 Heat kernels
Results
In this section, we evaluate the proposed ACNN method and compare it to state-of-the-art approaches on the FAUST [BRLB14] and SHREC’16 Partial Correspondence [CRB∗16] benchmarks.
Isotropic Laplacians were computed using the cotangent formula [Mac49, Duf59, PP93, MDSB03]; anisotropic Laplacians were computed according to (34). Heat kernels were computed in the frequency domain using all the eigenpairs. In all experiments, we used orientations and the anisotropy parameter .
Neural networks were implemented in Theano [B∗10]. The ADAM [KB15] stochastic optimization algorithm was used with initial learning rate of , , and . As the input to the networks, we used the local SHOT descriptor [STDS14] with dimensions and using default parameters. For all experiments, training was done by minimizing the loss (23). The code to reproduce all the experiments in this paper, and the full framework will be released upon publication.
Timing.
The following are typical timings for FAUST shapes with 6.9K vertices. Laplacian computation and eigendecomposition took sec and seconds per angle, respectively on a desktop workstation with 64Gb of RAM and i7-4820K CPU. Forward propagation of the trained model takes approximately sec to produce the dense soft correspondence for all the vertices.
2 Full mesh correspondence
In the first experiment, we used the FAUST humans dataset [BRLB14], containing meshes of scanned subjects, each in different poses. The shapes in the collection manifest strong non-isometric deformations. Vertex-wise groundtruth correspondence is known between all the shapes. The zeroth FAUST shape containing vertices was used as reference; for each point on the query shape, the output of the network represents the soft correspondence as a -dimensional vector which was then converted to point correspondence with the technique explained in Section 4. First shapes for training and the remaining for testing, following verbatim the settings of [MBBV15]. Batch normalization [IS15] was used to speed up the training; we did not experience any noticeable difference in raw performance of the produced soft correspondence compared to un-normalized setting.
Methods.
Batch normalization allows us to effectively train larger and deeper networks, for this experiment we adopted the following architecture inspired by GCNN [MBBV15]: FC+IC+IC+IC+FC+FC+Softmax. Additionally, we compared our method to Random Forests (RF) [RRBW∗14], Blended Intrinsic Maps (BIM) [KLF11], Localized Spectral CNN (LSCNN) [BMM∗15], and Anisotropic Diffusion Descriptors (ADD) [BMR∗16].
Results.
Figure 5 shows the performance of different methods. The performance was evaluated using the Princeton protocol [KLF11], plotting the percentage of matches that are at most -geodesically distant from the groundtruth correspondence on the reference shape. Two versions of the protocol consider intrinsically symmetric matches as correct (symmetric setting) or wrong (asymmetric, more challenging setting). Some methods based on intrinsic structures (e.g. LSCNN or RF applied on WKS descriptors) are invariant under intrinsic symmetries and thus cannot distinguish between symmetric points. The proposed ACNN method clearly outperforms all the compared approaches and also perfectly distinguishes symmetric points.
Figure Learning shape correspondence with anisotropic convolutional neural networks visualizes some correspondences obtained by ACNN using texture mapping. The correspondence show almost no noticeable artifacts. Figure 4 shows the pointwise geodesic error of different correspondence methods (distance of the correspondence at a point from the groundtruth). ACNN shows dramatically smaller distortions compared to other methods. Over of matches are exact (zero geodesic error), while only a few points have geodesic error larger than of the geodesic diameter of the shape.
3 Partial correspondence
In the second experiment, we used the recent very challenging SHREC’16 Partial Correspondence benchmark [CRB∗16], consisting of nearly-isometrically deformed shapes from eight classes, with different parts removed. Two types of partiality in the benchmark are cuts (removal of a few large parts) and holes (removal of many small parts). In each class, the vertex-wise groundtruth correspondence between the full shape and its partial versions is given. The dataset was split into training and testing disjoint sets. For cuts, training was done on 15 shapes per class; for holes, training was done on 10 shapes per class.
Methods.
Cuts.
Figure 6 compares the performance of different partial matching methods on the SHREC’16 Partial (cuts) dataset. ACNN outperforms other approaches with a significant margin. Figure 8 (top) shows examples of partial correspondence on the horse shape as well as the pointwise geodesic error (bottom). We observe that the proposed approach produces high-quality correspondences even in such a challenging setting.
Holes.
Figure 7 compares the performance of different partial matching methods on the SHREC’16 Partial (holes) dataset. In this setting as well, ACNN outperforms other approaches with a significant margin. Figure 9 (top) shows examples of partial correspondence on the dog shape as well as the pointwise geodesic error (bottom).
Conclusions
We presented Anisotropic CNN, a new framework generalizing convolutional neural networks to non-Euclidean domains, allowing to perform deep learning on geometric data. Our work follows the very recent trend in bringing machine learning methods to computer graphics and geometry processing applications, and is currently the most generic intrinsic CNN model. Our experiments show that ACNN outperforms previously proposed intrinsic CNN models, as well as additional state-of-the-art methods in the shape correspondence application in challenging settings. Being a generic model, ACNN can be used for many other applications. We believe it would be useful for the computer graphics and geometry processing community, and will release the code.
Acknowledgment
This research was supported by the ERC Starting Grant No. 307047 (COMET), a Google Faculty Research Award, and Nvidia equipment grant.