Continuous Surface Embeddings

Natalia Neverova, David Novotny, Vasil Khalidov, Marc Szafraniec, Patrick Labatut, Andrea Vedaldi

Introduction

Understanding the geometry of natural objects, such as humans and other animals, must start from the notion of correspondence. Correspondences tell us which parts of different objects are geometrically equivalent, and thus form the basis on which an understanding of geometry can be developed. In this paper, we are interested in particular in learning and computing correspondence starting from 2D images of the objects, a preliminary step for 3D reconstruction and other applications.

While the correspondence problem has been considered many times before, most solutions still involve a significant amount of manual work. Consider for example a state-of-the-art method such as DensePose . Given a new object category to model with DensePose, one must start by defining a canonical shape SS, a sort of ‘average’ 3D shape used as a reference to express correspondences. Then, a dataset of images of the object must be collected and annotated with millions of manual point correspondences between the images and the canonical 3D model. Finally, the model must be manually partitioned into a number of parts, or charts, and a deep neural network must be trained to segment the image and regress the uvuv coordinates for each chart, guided by the manual annotations, yielding a DensePose predictor. Given a new category, this process must be repeated from scratch.

There are some obvious scalability issues with this approach. The most significant one is that the entire process must be repeated for each new object category one wishes to model. This includes the laborious step of collecting annotations for the new class. However, categories such as animals share significant similarities between them; for instance, recently has shown that DensePose trained on humans transfers well on chimpanzees. Thus, a much better scalable solution can be obtained by sharing training data and models between classes. This brings us to the second shortcoming of DensePose: the nature of the model makes it difficult to realize this information sharing. In particular, the need for breaking the canonical 3D models into different charts makes relating different models cumbersome, particularly in a learning setup.

One important contribution of this paper is to introduce a better and more flexible representation of correspondences that can be used as a drop-in replacement in architectures such as DensePose. The idea is to introduce a learnable positional embedding. Namely, we associate each point XX of the canonical model SS to a compact embedding vector eXe_{X}, which provides a deformation-invariant representation of the point identity. We also note that the embedding can be interpreted as a smoothly-varying function defined over the 3D model SS, interpreted as a manifold. As such, this allows us to use the machinery of functional maps to work with the embeddings, with two important advantages: (1) being able to significantly reduce the dimensionality of the representation and (2) being able to efficiently relate representations between models of different object categories.

Empirically, we show that we can learn a deep neural network that predicts, for each pixel in a 2D image, the embedding vector of the corresponding object point, therefore establishing dense correspondences between the image pixels and the object geometry. For humans, we show that the resulting correspondences are as or more accurate than the reference state-of-the-art DensePose implementation, while achieving a significant simplification of the DensePose framework by removing the need of charting the model. As an additional bonus, this removes the ‘seams’ between the parts that affect DensePose. Then, we use the ability of the functional maps to relate different 3D shapes to help transferring information between different object categories. With this, and a very small amount of manual training data, we demonstrate for the first time that a single (universal) DensePose network can be extended to capturing multiple animal classes with a high degree of sharing in compute and statistics. The overview of our method with learning continuous surface embeddings (CSE) is shown in Fig. 1b (for comparison, the DensePose setup (IUV) is shown in Fig. 1a).

Related work

With deep learning, image-based human pose estimation has made substantial progress , also due to the availability of large datasets such as COCO , MPII , Leeds Sports Pose Dataset (LSP) , PennAction , or PoseTrack . Our work is most related with DensePose , which introduced a method to establish dense correspondences between image pixels and points on the surface of the average SMPL human mesh model .

Unsupervised pose recognition.

Most pose estimators require full supervision, which is expensive to collect, especially for a model such as DensePose. A handful of works have tackled this issue by seeking unsupervised and weakly-supervised objectives, using cues such as equivariance to synthetic image transformations. The most relevant to us is Slim DensePose , which showed that DensePose annotations can be significantly reduced without incurring a large performance penalty, but did not address the issue of scaling to multiple classes.

Animal pose recognition.

Compared to humans, animal pose estimation is significantly less explored. Some works specialise on certain animals (tigers , cheetahs or drosophila melanogaster flies ). Tulsiani et al. transfer pose between annotated animals and un-annotated ones that are visually similar. Several works have focused on animal landmark detection. Rashid et al. and Yang et al. studied animal facial keypoints. A well explored class are birds due to the CUB dataset . Some works proposed various types of detectors for birds, while others explored reconstructing sparse and dense 3D shapes. Beyond birds, Zuffi et al. have explored systematically the problem of reconstructing 3D deformable animal models from image data. They utilise the SMAL animal shape model, which is an analogue for animals of the more popular human SMPL model for humans . While are based on fitting 2D keypoints at test time, Biggs et al. directly regresses the 3D shape parameters instead. Recently, Kulkarni et al. leveraged canonical maps to perform the 3D shape fitting as well as establishing of dense correspondences across different instances of an animal species.

D shape analysis.

Our work is also related to the literature that studies the intrinsic geometry of 3D shapes. Early approaches analysed shapes by performing multi-dimensional scaling of the geodesic distances over the shape surface, as these are invariant to isometric deformations. Later, Coifman and Lafon popularized the diffusion geometry due to its increased robustness to small perturbations of the shape topology. The seminal work of Rustamov proposed to use the eigenfunctions of the Laplace-Beltrami operator (LBO) on a mesh to define a basis of functions that smoothly vary along the mesh surface. The LBO basis was later leveraged in other diffusion descriptors such as the heat kernel signature (HKS) , wave kernel signature (WKS) or Gromov-Hausdorff descriptors . A scale-invariant version of HKS was introduced in , while proposed an HKS-based equivalent of the image BoW descriptor . While HKS/WKS establish ‘hard’ correspondences between individual points on shapes, Ovsjanikov et al. introduced the functional maps (FM) that align shapes in a soft manner by finding a linear map between spaces of functions on meshes. Interestingly, has revealed an intriguing connection between FMs and their efficient representation using the LBO basis. The FM framework became popular and was later extended in . Relevantly to us, proposed ZoomOut, a method that estimates FM in a multi-scale fashion, which we improve for our species-to-species mesh correspondences.

Method

In this formulation, the embedding function eXe_{X} is learnable just like the network Φ\Phi. The simplest way of implementing this idea is to approximate the surface SS with a mesh with vertices X1,…,XK∈SX_{1},\dots,X_{K}\in S, obtaining a discrete variant of this model:

where ek=eXke_{k}=e_{X_{k}} is a shorthand notation of the embedding vector associated to the kk-th vertex of the mesh and EE is the K×DK\times D matrix with all the embedding vectors {ek}k=1K\{e_{k}\}_{k=1}^{K} as rows.

Given a training set of triplets (I,x,k)(I,x,k), where II is an image, xx a pixel, and kk the index of the corresponding mesh vertex, we can learn this model by minimizing the cross-entropy loss:

We found it beneficial to modify this loss to account for the geometry of the problem, minimizing the cross entropy between a ‘Gaussian-like’ distribution centered on the ground-truth point kk and the predicted posterior:

1 Injecting geometric knowledge via spectral analysis

There are three issues with the formulation we have given so far. First, the embedding matrix EE is large, as it contains DD parameters for each of the KK mesh vertices. Second, the representation depends on the discretization of the mesh, so for example it is not clear what to do if we resample the mesh to increase its resolution. Third, it is not obvious, given two surfaces SS and S′S^{\prime} for two different object categories (e.g. humans and chimps), how their embeddings ee and e′e^{\prime} can be related.

2 Relating different categories

When shapes SS and S′S^{\prime} are approximately isometric, we can resort to automatic methods to establish correspondences Π\Pi (or CC) between them. When SS and S′S^{\prime} are not, this is much harder. Instead, we start from a small number of manual correspondences (kj,kj′)(k_{j},k_{j}^{\prime}), j=1,…,Qj=1,\dots,Q between surfaces and interpolate those using functional maps. To do this, we use a variant of the ZoomOut method due to its simplicity. This amounts to alternating two steps: given a matrix CC of order M1×M1M_{1}\times M_{1}, we decode this as a set of point-to-point correspondences (kj,kj′)(k_{j},k_{j}^{\prime}); then, given (kj,kj′)(k_{j},k_{j}^{\prime}), we estimate a matrix CC of order M2>M1M_{2}>M_{1}, increasing the resolution of the match. This is done for a sequence Mt=12,16,…,256M_{t}=12,16,\dots,256 until the desired resolution is achieved (see Sec. A.2 for details).

3 Cross-species DensePose with functional maps

We are now ready to describe how functional maps can be used to facilitate transferring and sharing a single DensePose predictor Φ\Phi between categories with different canonical shapes S,S′S,S^{\prime} and corresponding per-vertex embeddings E,E′E,E^{\prime}.

To this end, assume that we have learned the pose regressor Φ\Phi on a source category (S,E)(S,E) (e.g. humans). We are free to apply Φ\Phi to an image I′I^{\prime} of a different category (S′,E′)(S^{\prime},E^{\prime}) (e.g. chimps), but, while we might know S′S^{\prime}, we do not know E′E^{\prime}. However, if we assume that the regressor Φ\Phi can be shared among categories, then it is natural to also share their positional embeddings as well. Namely, we assume that E′E^{\prime} is approximately the same as EE up to a remapping of the embedding vectors from shape SS to shape S′S^{\prime}. Based on the last section, we can thus write E^′=CE^\hat{E}^{\prime}=C\hat{E}, or, equivalently, E′=TEE^{\prime}=TE where T=U′CU⊤AT=U^{\prime}CU^{\top}A. With this, we can simply replace EE with E′E^{\prime} in eq. 2 to now regress the pose of the new category using the same regressor network Φ\Phi. For training, we optimise the same cross entropy loss L(E,Ψ)\mathcal{L}(E,\Psi) in eq. 3, just combining images and annotations from the two object categories and swapping EE and E′E^{\prime} depending on the class of the input image.

The procedure above can be easily generalised to any number of categories S1,…,SKS^{1},\dots,S^{\mathcal{K}} with reference to the same source category SS and functional maps C1,…,CKC^{1},\dots,C^{\mathcal{K}}. In our case, we select humans as source category as they contain the largest number of annotations.

Datasets

We rely on the DensePose-COCO dataset for evaluation of the proposed method on the human category and comparison with the DensePose (IUV) training. For the multi-class setting, we make use of a recent DensePose-Chimps test benchmark containing a small number of annotated correspondences for chimpanzees. We split the set of annotated instances of into 500 training and 430 test samples containing 1354 and 1151 annotated correspondences respectively.

Additionally, we collect correspondence annotations on a set of 9 animal categories of the LVIS dataset . Based on images from the COCO dataset , LVIS features significantly more accurate object masks. We refer to this data as DensePose-LVIS. The annotation statistics for the collected animal correspondences are given in Table 1. Note that compared to the original DensePose-COCO labelling effort that produced 5 million annotated points for the human category (96%96\% coverage of the SMPL mesh), our annotations are three orders of magnitude smaller and only 18%18\% of vertices of animal meshes, on average, have at least one ground truth annotation.

Experiments

Our networks are implemented in PyTorch within the Detectron2 framework. The training is performed on 8 GPUs for 130k iterations on DensePose-COCO (standard s1x schedule ) and 5k iterations on DensePose-Chimps and DensePose-LVIS. The code, trained models and the dataset will be made publicly available to ensure reproducibility.

Prior to benchmarking the CSE setup, we carefully optimized all core architectures for dense pose estimation and introduced the following changes (similarly to ): (1) single channel instance mask prediction as a replacement for the coarse segmentation of ; (2) optimized weights for (i,u,v)(i,u,v) components; (3) DeepLab head and Panoptic FPN (similarly to ). The experimental results are reported following the updated protocol based on GPSm scores . More details on the network architectures and training hyperparameters are given in the supplementary material.

Comparison of CSE vs IUV training.

The comparison of the state-of-the-art methods for dense pose estimation and our optimized architectures for both IUV and CSE training is provided in Table 2. The CSE-trained models perform better or on par with their IUV-trained counterparts, while producing a more compact representation (D=16D=16 vs D=75D=75) and requiring only simplified supervision (single vertex indices vs (i,u,v)(i,u,v) annotations).

Influence of hyperparameters.

In Table 3 we investigate the CSE network sensitivity to the size of the LBO basis, MM (left), and the output embedding, DD (right), given training losses L\mathcal{L} or Lσ\mathcal{L}_{\sigma}. The value M=256M=256 represents the tradeoff between mapping’s smoothness and its fidelity. It does not seem beneficial to increase the embedding size beyond D=16D=16, so we adapt this value for the rest of the experiments. The smoothed loss Lσ\mathcal{L}_{\sigma} yields better performance in a low dimensional setting.

Low data regime.

Prior to proceeding with the multi-class experiments on animal classes, we investigate changes in models performance as a function of the amount of ground truth annotations by training on subsets of the DensePose-COCO dataset. As shown in Table 3 on the right, Lσ\mathcal{L}_{\sigma}-based training scales down more gracefully and is significantly more robust than L\mathcal{L}.

Multi-surface training.

The results on DensePose-Chimps and DensePose-LVIS datasets are reported in Tables 4 and 5 (DP-RCNN* (R50), M=256M=256, D=16D=16). In both cases, training from scratch results in poor performance, especially in a single class setting. Initializing the predictor Φ\Phi with the human class trained weights together with the alignment of mesh-specific vertex embeddings (as described in 3.3) gives a significant boost. Interestingly, class agnostic training by mapping all class embeddings to the shared space turns out to be more effective than having a separate set of output planes for each category (the latter is denoted as multiclass in Table 5). Quantitative results produced by the best predictor are shown in Figure 4.

Conclusion.

In this work, we have made an important step towards designing universal networks for learning dense correspondences within and across different object categories (animals). We have demonstrated that training joint predictors in the image space with simultaneous alignment of canonical surfaces in 3D results in an efficient transfer of knowledge between different classes even when the amount of ground truth annotations is severely limited.

Broader impact

In our paper, we help improve the ability of machines to understand the pose of articulated objects such as humans in images. In particular, we make the process of learning new object categories much more efficient.

An application of our method is the observation of the human body. This may come with some concerns on possible negative uses of the technology. However, we should note that our approach cannot be considered biometrics, because from pose alone, even if dense, it is not possible to ascertain the identity of an individual (in particular, we do not perform 3D reconstruction, nor we reconstruct facial features). This mitigates the potential risk when our method is applied to humans.

We believe that our work has significant opportunities for a positive impact by opening up the possibility that machines could ultimately understand the pose of thousands of animal classes. In addition to numerous applications in VR, AR, marketing and the like, such a technology can benefit animal-human-machine interaction (e.g. in aid of the visually impaired), can be used to better safeguard animals on the Internet (e.g. by detecting animal abuse), and, perhaps most importantly, can allow conservationists and other researchers to observe animals in the wild at an unprecedented scale, automatically analysing their motion and activities, and thus collecting information on their number, state of health, and other statistics. Thus, while we acknowledge that this technology may find negative uses (as almost any technology does), we believe that the positives far outweigh them.

References

Appendix A Appendix A

In order to construct the discrete Laplace-Beltrami Operator (LBO), and with reference to the notation introduced in the main manuscript, we assume that the points XkX_{k} are the vertices of a simplicial mesh, i.e. the union Sˉ=∪f∈Ff\bar{S}=\cup_{f\in F}f of a finite set FF of triangular faces ff, approximating the 3D surface SS. With slight abuse of notation, we denote each face f=(X1,X2,X3)f=(X_{1},X_{2},X_{3}) as a triplet of vertices oriented in clockwise order with respect to the normal NfN_{f} of the face. If we assume that the function rr is continuous and linear within each face, then the samples r\mathbf{r} fully specify the function. Let b1=X3−X2b_{1}=X_{3}-X_{2}, b2=X1−X3b_{2}=X_{1}-X_{3} and b3=X2−X1b_{3}=X_{2}-X_{1}, be the edge vectors opposite to each vertex of the triangle ff and let AfA_{f} be its area. The gradient of rr, which is constant on each face ff, is given by:

We can verify the expression eq. 5 for the gradient as follows. The gradient dotted with an edge vector bib_{i} must give the function change along that edge. For example, for edge b1b_{1} we have:

We can write matrix GfG_{f} in eq. 5 much more compactly as:

Cotangent weight matrix W𝑊W.

Given the expression for GfG_{f}, we can find a compact expression fo WfW_{f} in the LBO:

Note that Bf⊤N^f⊤N^fBf=Bf⊤BfB_{f}^{\top}\hat{N}_{f}^{\top}\hat{N}_{f}B_{f}=B_{f}^{\top}B_{f} because bi⊥Nfb_{i}\perp N_{f} and thus:

WfW_{f} matches the usual cotangent discretization of the Laplace-Beltrami operator: BfB_{f} contains dot products of edges, 2Af2A_{f} the norm of their cross products, and the ratio of these two are cotangents.

Divergence operator D𝐷D.

A.2 Spectra interpolation of correspondences (ZoomOut)

Assume that we have complete correspondences for the mesh S′S^{\prime}, in the sense that kj′=j,k^{\prime}_{j}=j, j=1,…,K′j=1,\dots,K^{\prime}. We can encode those as a permutation matrix Π\Pi such that Πij=δki=j\Pi_{ij}=\delta_{k_{i}=j}, mapping functions r\mathbf{r} on SS to function r′=Πr\mathbf{r}^{\prime}=\Pi\mathbf{r} on S′S^{\prime} (this is analogous to backward warping). This can be rewritten in ‘Fourier’ space as r′=U′r^′=U′Cr^=Πr=ΠUr^,\mathbf{r}^{\prime}=U^{\prime}\hat{\mathbf{r}}^{\prime}=U^{\prime}C\hat{\mathbf{r}}=\Pi\mathbf{r}=\Pi U\hat{\mathbf{r}}, which gives us the constraint U′C=ΠUU^{\prime}C=\Pi U. We can use this equation to find CC given Π\Pi, or to find Π\Pi given CC. Finding Π\Pi is done in a greedy manner, searching, for each row of U′CU^{\prime}C, the best matching row in UU (in L2L^{2} distance). Finding CC is done by minimizing ∥U′C−ΠU∥A′2\|U^{\prime}C-\Pi U\|^{2}_{A^{\prime}}, which results in C=(U′)⊤A′ΠUC=(U^{\prime})^{\top}A^{\prime}\Pi U.

In practice, we found it beneficial to add three more standard constraints when resolving for CC. First, let Γ∈{0,1}K×K\Gamma\in\{0,1\}^{K\times K} be the symmetry matrix mapping each vertex of mesh SS to its symmetric counterpart (this is trivially determined for our canonical models), and let Γ′\Gamma^{\prime} be the same for S′S^{\prime}. Then a correct correspondence Π\Pi between meshes must preserve symmetry, in the sense that Γ′Π=ΠΓ\Gamma^{\prime}\Pi=\Pi\Gamma; this constraint can be rewritten in Fourier space as Γ^′C=CΓ^\hat{\Gamma}^{\prime}C=C\hat{\Gamma}, where Γ^=UΓU†\hat{\Gamma}=U\Gamma U^{\dagger}. For isometric meshes, the exact same reasoning applies to the LBO L=A−1WL=A^{-1}W because the LBO is an intrinsic property of the surface (i.e. invariant to isometry). Our meshes are not isometric, but, after resizing them to have the same total area, we can use the constraint L′Π≈ΠLL^{\prime}\Pi\approx\Pi L in a soft manner for regularization; it is easy to show that this reduces to Λ′C≈CΛ\Lambda^{\prime}C\approx C\Lambda where Λ\Lambda is the matrix of eigenvalues of the LBO. In practices, this encourages CC to be roughly diagonal. Finally, we use the method just described twice, to estimate jointly a mapping CC from mesh SS to S′S^{\prime}, and another C′C^{\prime} going in the other direction, and enforce CC′≈ICC^{\prime}\approx I (cycle consistency).

Appendix B Appendix B

We are following an annotation protocol similar to the one described in the original DensePose work . We start with instance mask annotations provided in the LVIS dataset and crop images around each instance. We only annotate instances with bounding boxes larger than 75 pixels. We do not collect annotations for body segmentation: instead, the points are sampled from the whole foreground region represented by the object mask. The annotators are then shown randomly sampled points displayed on the image and are asked to click on corresponding points in multiple views rendered from a 3D model representing the given species. Each worker is asked to annotate 3 points on a single object instance. The points on the rendered views are mapped directly to vertex indices of the corresponding model. Each mesh is normalised to have approximately 5k vertices.

B.2 Implementation details

Compared to the original DensePose models, we introduced the following changes:

single channel mask supervision as a replacement of the 15-way segmentation;

RoI pooling size for the DensePose task is set to 28×2828\times 28;

a decoder module based on Panoptic FPN, as implemented in ;

for the DeepLab models, the head architecture corresponds to ;

IUV training, weights of individual loss terms: wmask=5.0w_{mask}=5.0 (body mask), wi=1.0w_{i}=1.0 (point body indices), wuv=0.01w_{uv}=0.01 (uv coordinates);

CSE training, weight on the embedding loss term: we=0.6w_{e}=0.6.

For the DensePose-COCO dataset, all models are trained with the standard s1x schedule for 130k iterations. On the LVIS and DensePose-Chimps datasets, the models are trained for 5k iterations with the learning rate drop by the factor of 10 after 4000k and 4500k iterations.

For the evaluation purposes all 3D meshes are normalised in size to have the same geodesic distance between the pair of most distant points as the SMPL model (Pdist.,max=2.5P_{\text{dist.,max}}=2.5). For the animal classes, we do not employ part specific normalisation coefficients, as done in the updated DensePose evaluation protocol.

The code, the pretrained models and the annotations for the LVIS dataset will be publicly released.