NASA: Neural Articulated Shape Approximation

Boyang Deng, JP Lewis, Timothy Jeruzalski, Gerard Pons-Moll, Geoffrey Hinton, Mohammad Norouzi, Andrea Tagliasacchi

Introduction

There has been a surge of recent interest in representing 3D geometry using implicit functions parameterized by neural networks . Such representations are flexible, continuous, and differentiable. Neural implicit functions are useful for “inverse graphics” pipelines for scene understanding , as back propagation through differentiable representations of 3D geometry is often required. That said, neural models of articulated objects have received little attention. Articulated objects are particularly important to represent animals and humans, which are central in many applications such as computer games and animated movies, as well as augmented and virtual reality.

Although parametric models of human body such as SMPL have been integrated into neural network frameworks for self-supervision , these approaches depend heavily on polygonal mesh representations. Mesh representations require expert supervision to construct, and are not flexible for capturing topology variations. Furthermore, geometric representations often should fulfill several purposes simultaneously such as modeling the surface for rendering, or representing the volume to test intersections with the environment, which are not trivial when polygonal meshes are used . Although neural models have been used in the context of articulated deformation , they relegate query execution to classical acceleration data structures, thus sacrificing full differentiability.

Our method represents articulated objects with a neural model, which outputs a differentiable occupancy of the articulated body in a specific pose. Like previous geometric learning efforts , we represent geometry by indicator functions – also referred to as occupancy functions – that evaluate to 11 inside the object and otherwise. Unlike previous approaches, which focused on collections of static objects described by (unknown) shape parameters, we look at learning indicator functions as we vary pose parameters, which will be discovered by training on animation sequences. We show that existing methods cannot encode pose variation reliably, because it is hard to learn the occupancy of every point in space as a function of a latent pose vector.

Instead, we introduce NASA, a neural decoder that exploits the structure of the underlying deformation driving the articulated object. Exploiting the fact that 3D geometry in local body part coordinates does not significantly change with pose, we classify the occupancy of 3D points as seen from the coordinate frame of each part. Our main architecture combines a collection of per-part learnable indicator functions with a per-part pose encoder to model localized non-rigid deformations. This leads to a significant boost in generalization to unseen poses, while retaining the useful properties of existing methods: differentiability, ease of spatial queries such as intersection testing, and continuous surface outputs. To demonstrate the flexibility of NASA, we use it to track point clouds by finding the maximum likelihood estimate of the pose under NASA’s occupancy model. In contrast to mesh based trackers which are complex to implement, our tracker requires a few lines of code and is fully differentiable. Overall, our contributions include:

We propose a neural model of articulated objects to predict differentiable occupancy as a function of pose – the core idea is to model shapes by networks that encode a piecewise decomposition;

The results on learning 3D body deformation outperform previous geometric learning algorithms , and our surface reconstruction accuracy approaches that of mesh-based statistical body models ;

The differentiable occupancy supports constant-time queries (.06.06 ms/query on an NVIDIA GTX 1080), avoiding the need to convert to separate representations, or the dynamic update of spatial acceleration data structures;

We derive a technique that employs occupancy functions for tracking 3D geometry via an implicit occupancy template, without the need to ever compute distance functions.

Related work

Neural articulated shape approximation provides a single framework that addresses problems that have previously been approached separately. The related literature thus includes a number of works across several different research topics.

Efficient articulated deformation is traditionally accomplished with a skinning algorithm that deforms vertices of a mesh surface as the joints of an underlying abstract skeleton change. The classic linear blend skinning (LBS) algorithm expresses the deformed vertex as a weighted sum of that vertex rigidly transformed by several adjacent bones; see for details. LBS is widely used in computer games, and is a core ingredient of some popular vision models . Mesh sequences of general (not necessarily articulated) deforming objects have also been represented with skinning for the purposes of compression and manipulation, using a collection of non-hierarchical “bones” (i.e. transformations) discovered with clustering . LBS has well-known disadvantages: the deformation has a simple algorithmic form that cannot produce pose-dependent detail, it results in characteristic volume-loss effects such as the “collapsing elbow” and “candy wrapper” artifacts [31, Figs. 2,3], and for best results the weights must be manually painted by artists. It is possible to add pose-dependent detail with a shallow or deep net regression , but this process operates as a correction to classical LBS deformation.

Object intersection queries

Registration, template matching, 3D tracking, collision detection, and other tasks require efficient inside/outside tests. A disadvantage of polygonal meshes is that they do not efficiently support these queries, as meshes often contain thousands of individual triangles that must be tested for each query. This has led to the development of a variety of spatial data structures to accelerate point-object queries , including voxel grids, octrees, kdtrees, and others. In the case of deforming objects, the spatial data structure must be repeatedly rebuilt as the object deforms. A further problem is that typically meshes may be constructed (or deformed) without regard to being “watertight” and thus do not have a clearly defined interior .

Part-based representations

For object intersection queries on articulated objects, it can be more efficient to approximate the overall shape in terms of a moving collection of rigid parts, such as spheres or ellipsoids, that support efficient querying ; see supplementary material for further discussion. Unfortunately this has the drawback of introducing a second approximate representation that does not exactly match the originally desired deformation. A further core challenge, and subject of continuing research, is the automatic creation of such part-based representations . Unsupervised part discovery has been recently tackled by a number of deep learning approaches . In general these methods address analysis and correspondence across shape collections, but do not target accurate representations of articulated objects, and do not account for pose-dependent deformations.

Neural implicit object representation

Finally, several recent works represent objects with neural implicit functions . These works focus on the neural representation of static shapes in an aligned canonical frame and do not target the modeling of transformations. Our core contributions are to show that these architectures have difficulties in representing complex and detailed articulated objects (e.g. human bodies), and that a simple architectural change can address these shortcomings. Comparisons to these closely related works will be revisited in more depth in Section 6.

Neural Articulated Shape Approximation

This paper focuses on building an expressive model of p(O∣θ)p(\mathcal{O}|\boldsymbol{\theta}), that is, occupancy conditioned on pose. Figure 2 illustrates this problem for d=2d{=}2, and clarifies the notation. There is extensive existing research on pose priors p(θ)p(\boldsymbol{\theta}) for human bodies and other articulated objects . Our work is orthogonal to such prior models, and any parametric or non-parametric p(θ)p(\boldsymbol{\theta}) can be combined with our p(O∣θ)p(\mathcal{O}|\boldsymbol{\theta}) to obtain the joint distribution p(θ,O)p(\boldsymbol{\theta},\mathcal{O}). We delay the discussion of pose priors until Section 5.2, where we define a particularly simple prior that nevertheless supports sophisticated tracking of moving humans.

In what follows we describe different ways of building a pose conditioned occupancy function, denoted Oω(x∣θ)\mathcal{O}_{\omega}(\mathbf{x}|\boldsymbol{\theta}), which maps a 3D point x\mathbf{x} and a pose θ\boldsymbol{\theta} onto a real valued occupancy value. Our goal is to learn a parametric occupancy Oω(x∣θ)\mathcal{O}_{\omega}(\mathbf{x}|\boldsymbol{\theta}) that mimics a ground truth occupancy O(x∣θ)\mathcal{O}(\mathbf{x}|\boldsymbol{\theta}) as closely as possible, based on the following probabilistic interpretation:

where we assume a standard normal distribution around the predicted real valued occupancy Oω(x∣θ)\mathcal{O}_{\omega}(\mathbf{x}|\boldsymbol{\theta}) to score an occupancy O(x∣θ)\mathcal{O}(\mathbf{x}|\boldsymbol{\theta}).

Given pose parameters θ\boldsymbol{\theta}, we desire to query the corresponding indicator function O(x∣θ)\mathcal{O}(\mathbf{x}|\boldsymbol{\theta}) at a point x\mathbf{x}. This task is more complicated than might seem, as in the general setting this operation requires the computation of generalized winding numbers to resolve ambiguous configurations caused by self-intersections and non-necessarily watertight geometry . However, when given a database of poses Θ={θt}t=1T\boldsymbol{\Theta}{=}\{\boldsymbol{\theta}_{t}\}_{t=1}^{T} and corresponding ground truth indicator {O(x∣θt)}t=1T\{\mathcal{O}(\mathbf{x}|\boldsymbol{\theta}_{t})\}_{t=1}^{T}, we can formulate our problem as the minimization of the objective:

Pose conditioned occupancy 𝒪​(𝐱|𝜽)𝒪conditional𝐱𝜽\mathcal{O}(\mathbf{x}|\boldsymbol{\theta})

We investigate several neural architectures for the problem of articulated shape approximation; see Figure 3. We start by introducing an unstructured architecture (U) in Section 4.1. This baseline variant does not explicitly encode the knowledge of articulated deformation. However, typical articulated deformation models express deformed mesh vertices V{\mathbf{V}} reusing the information stored in rest vertices Vˉ\bar{\mathbf{V}}. Hence, we can assume that computing the function O(x∣θ)\mathcal{O}(\mathbf{x}|\boldsymbol{\theta}) in the deformed pose can be done by reasoning about the information stored at rest pose O(x∣θˉ)\mathcal{O}(\mathbf{x}|\bar{\boldsymbol{\theta}}). Taking inspiration from this observation, we investigate two different architecture variants, one that models geometry via a piecewise-rigid assumption (Section 4.2), and one that relaxes this assumption and employs a quasi-rigid decomposition, where the shape of each element can deform according to the pose (Section 4.3); see Figure 4.

Recently, a series of papers tackled the problem of modeling occupancy across shape datasets as Oω(x∣β)\mathcal{O}_{\omega}(\mathbf{x}|\boldsymbol{\beta}), where β\boldsymbol{\beta} is a latent code learned to encode the shape. These techniques employ deep and fully connected networks, which one can adapt to our setting by replacing the shape β\boldsymbol{\beta} with pose parameters θ\boldsymbol{\theta}, and using a neural network that takes as input [x,θ][\mathbf{x},\boldsymbol{\theta}]. Leaky ReLU activations are used for inner layers of the neural net and a sigmoid activation is used for the final output so that the occupancy prediction lies in the $$ range.

To provide pose information to the network, one can simply concatenate the set of affine bone transformations to the query point to obtain [x,{Bb}][\mathbf{x},\{\mathbf{B}_{b}\}] as the input. This results in an input tensor of size 3+16×B3{+}16{\times}B. Instead, we propose to represent pose as {Bb−1t0}\{\mathbf{B}_{b}^{-1}\mathbf{t}_{0}\}, where t0\mathbf{{t}}_{0} is the translation vector of the root bone in homogeneous coordinates, resulting in a smaller input of size 3+3×B3{+}3{\times}B; we ablate this choice against other alternatives in the supplementary material. Our unstructured baseline takes the form:

2 Piecewise rigid model – “R”

The simplest structured deformation model for articulated objects assumes objects can be represented via a piecewise rigid composition of elements; e.g. :

We observe that if these elements are related to corresponding rest-pose elements through the rigid transformations {Bb}\{\mathbf{B}_{b}\}, then it is possible to query the corresponding rest-pose indicator as:

where, similar to (4), we can represent each of components via a learnable indicator Oˉωb(.)=MLPωb(.)\bar{\mathcal{O}}^{b}_{\omega}(.){=}\text{MLP}^{b}_{\omega}(.). This formulation assumes that the local shape of each learned bone component stays constant across the range of poses when viewed from the corresponding coordinate frame, which is only a crude approximation of the deformation in realistic characters, and other deformable shapes.

3 Piecewise deformable model – “D”

We can generalize our models by combining the model of (4) to the one in (6), hence allowing the shape of each element to be adjusted according to pose:

Similar to (\refeq:piecewise)(\ref{eq:piecewise}) we use a collection of learnable indicator functions in rest pose {Oωb}\{\mathcal{O}_{\omega}^{b}\}, and to encode pose conditionals we take inspiration from (4). More specifically, we express our model as:

4 Technical details

The overall training loss for our model is:

where λ=5e−1\lambda{=}5e^{-1} was found through hyper-parameter tuning. We now detail the weights auxiliary loss, the architecture backbones, and the training procedure.

As most deformable models are equipped with skinning weights, we exploit this additional source of information to facilitate learning of the part-based models (i.e. “R” and “D”). We label each mesh vertex v\mathbf{v} with the index of the corresponding highest skinning weight value b∗(v)=arg max⁡bw(v)[b]b^{*}(v){=}\operatorname*{arg\,max}_{b}w(v)[b], and use the loss:

where Ib(v)=0.5\mathcal{I}_{b}(\mathbf{v}){=}0.5 when b=b∗b{=}b^{*}, and Ib(v)=0\mathcal{I}_{b}(\mathbf{v}){=}0 otherwise – recall that by convention the 0.50.5 level set is the surface represented by the occupancy function. Without such a loss, we could end up in the situation where a single (deformable) part could end up being used to describe the entire deformable model, and the trivial solution (zero) would be returned for all other parts.

Network architectures

To keep our experiments comparable across baselines, we use the same network architecture for all the models while varying the width of the layers. The network backbone is similar to DeepSDF , but simplified to 4 layers. Each layer has a residual connection, and uses the Leaky ReLU activation function with the leaky factor 0.1. All layers have the same number of neurons, which we set to 960960 for the unstructured model and 4040 for the structured ones. For the piecewise (6) and deformable (8) models the neurons are distributed across B=24B{=}24 different channels (note B×40=960B{\times}40=960). Similar to the use of grouped filters/convolutions , such a structure allows for significant performance boosts compared to unstructured models (4), as the different branches can be executed in parallel on separate compute devices.

Training

All models are trained with the Adam optimizer, with batch size 1212 and learning rate 1e−41e-4. For better gradient propagation, we use softmax whenever a max was employed in our expressions. For each optimization step, we use 10241024 points sampled uniformly within the bounding box and 10241024 points sampled near the ground truth surface. We also sample 20482048 vertices out of 68906890 mesh vertices at each step for Lweights\mathcal{L}_{weights}. The models are trained for 200200K iterations for approximately 66 hours on a single NVIDIA Tesla V100.

Dense articulated tracking

Following the probabilistic interpretation of Section 3, we introduce an application of NASA to dense articulated 3D tracking; see . Note that this section does not claim to beat the state-of-the-art in tracking of deformable objects , but rather it seeks to show how neural occupancy functions can be used effectively in the development of dense tracking techniques. Taking the negative log of the joint probability in (1), the tracking problems can be expressed as the minimization of a pair of energies :

If we could compute the Signed Distance Function (SDF) Φ\Phi of an occupancy O\mathcal{O} at a query point x\mathbf{x}, then the fitness of O\mathcal{O} to input data could be measured as:

The time complexity of computing SDF from an occupancy O\mathcal{O} that is discretized on a grid is linear in the number of voxels . However, the number voxels grows as O(nd)O(n^{d}), making naive SDF computation impractical for high resolutions (large nn) or high dimensions (in practice, d\text{\scriptsize\geq}3 is already problematic). Spatial acceleration data structures (kdtrees and octrees) are commonly employed, but these data structures still require an overall O(nlog⁡(n))O(n\log(n)) pre-processing (where nn is the number of polygons), and they need to be re-built at every frame (as θ\boldsymbol{\theta} changes), and do not support implicit representations.

Recently, Dou et al. proposed to smooth an occupancy function with a Gaussian blur kernel to approximate Φ\Phi in the near field of the surface. Following this idea, our fitting energy can be re-expressed as:

where N0,σ2\mathcal{N}_{0,\sigma^{2}} is a Gaussian kernel with a zero mean and a variance σ2\sigma^{2}, and ⊛\circledast is the convolution operator. This approximation is suitable for tracking, as large values of distance should be associated with outliers in a registration optimization , and therefore ignored. Further, this approximation can be explained via the algebraic relationship between heat kernels and distance functions . Note that we intentionally use O\mathcal{O} instead of Oω\mathcal{O}_{\omega}, as what follows is applicable to any implicit representation, not just our neural occupancy Oω\mathcal{O}_{\omega}.

Dou et al. used (13) being given a voxelized representation of O\mathcal{O}, and relying on GPU implementations to efficiently compute 3D convolutions ⊛\circledast. To circumvent these issues, we re-express the convolution via stochastic sampling:

Overall, Equation 17 allows us to design a tracking solution that directly operates on occupancy functions, without the need to compute signed distance functions , closest points , or 3D convolutions . It further provides a direct cost/accuracy control in terms of the number of samples used to approximate the expectation in (17). However, the gradients ∇θ\nabla_{\boldsymbol{\theta}} of (17) also need to be available – we achieve this by applying the re-parameterization trick to (17):

2 Pose prior energy

An issue of generative tracking is that once the model is too far from the target (e.g. fast motion) there will be no proper gradient to correct it. If we directly optimize for transformation without any constraints, there is a high chance that the model will degenerate into such a case. To address this, we impose a prior:

where E\mathcal{E} is the set of directed edges (b1,b2)(b_{1},b_{2}) on the pre-defined directed rig with b1b_{1} as the parent, and recall tb\mathbf{t}_{b} is the translation vector of matrix Bb\mathbf{B}_{b}. One can view this loss as aligning the vector pointing to tb2t_{b_{2}} at run-time with the vector at rest pose, i.e. (tˉb2−tˉb1)(\bar{\mathbf{t}}_{b_{2}}-\bar{\mathbf{t}}_{b_{1}}). We emphasize that more sophisticated priors exist, and could be applied, including employing a hierarchical skeleton , or modeling the density of joint angles . The simple prior used here is chosen to highlight the effectiveness of our neural occupancy model independent of such priors.

3 Iterative optimization

One would be tempted to use the gradients of (13) to track a point cloud via iterative optimization. However, it is known that when optimizing rotations centering the optimization about the current state is heavily advisable . Indexing time by (t)(t) and given the update rule θ(t)=θ(t−1)+ Δθ(t)\boldsymbol{\theta}^{(t)}{=}\boldsymbol{\theta}^{(t-1)}{+}\,\Delta\boldsymbol{\theta}^{(t)}, the iterative optimization of (13) can be expressed as:

where in what follows we omit the index (t)(t) for brevity of notation. As the pose θ\boldsymbol{\theta} is represented by matrices, we represent the transformation differential as:

where we parameterize the rotational portion of elements in the collection {ΔBb−1}\{\Delta\mathbf{B}_{b}^{-1}\} by two (initially orthogonal) vectors , and re-orthogonalize them before inversion within each optimization update Bb(i+1)=Bb(i)(ΔCb(i))−1\mathbf{B}_{b}^{(i+1)}{=}\mathbf{B}_{b}^{(i)}(\Delta\mathbf{C}_{b}^{(i)})^{-1}, where ΔCb=ΔBb−1\Delta\mathbf{C}_{b}{=}\Delta\mathbf{B}_{b}^{-1} are the quantities the solver optimizes for. In other words, we optimize for the inverse of the coordinate frames in order to avoid back-propagation through matrix inversion.

Results and discussion

We describe the training data (Section 6.1), quantitatively evaluate the performance of our neural 3D representation on several datasets (Section 6.2), as well as demonstrate its usability for tracking applications (Section 6.3). We conclude by contrasting our technique to recent methods for implicit-learning of geometry (Section 6.4). Ablation studies validating each of our technical choices can be found in the supplementary material.

Our training data consists of sampled indicator function values, transformation frames (“bones”) per pose, and skinning weights. The samples used for training (3) come from two sources (each comprising a total of 100,000100,000 samples): \raisebox{-.6pt}{1}⃝ we randomly sample points uniformly within a bounding box scaled to 110% of its original diagonal dimension; \raisebox{-.6pt}{2}⃝ we perform Poisson disk sampling on the surface, and randomly displace these points with isotropic normal noise with σ=.03\sigma{=}.03. The ground truth indicator function at these samples are computed by casting randomized rays and checking the parity (i.e. counting the number of intersections) – generalized winding numbers or sign-agnostic losses could also be used for this purpose. The test reconstruction performance is evaluated by comparing the predicted indicator values against the ground truth samples on the full set of 100,000100,000 samples. We evaluate using mean Intersection over Union (IoU), Chamfer-L1 and F-score (F%\%) with a threshold set to 0.00010.0001. The meshes are obtained from the “DFaust” and “Transitions” sub-datasets of AMASS , as detailed in Section 6.2.

2 Reconstruction

We employ the “DFaust” portion of the AMASS dataset to verify that our model can be used effectively across different subjects. This dataset contains 1010 subjects, 1010 sequences/subject, and ≈300{\approx}300 frames/sequence on average. We train 100100 different models by optimizing (9): for each subject we use 99 sequences for training, leaving one out for testing to compute our metrics. We average these metrics across the 100100 runs, and report these in Table 1. Note how learning a deformable model via decomposition provides striking advantages, as quantified by the fact that the rigid (R) baseline is consistently better than the unstructured (U) baseline under any metric – a +49%\mathbf{+49\%} in F-score. Similar improvements can be noticed by comparing the rigid (R) to the deformable (D) model, where the latter achieves an additional +5%\mathbf{+5\%} in F-score. Figure 5 (second row) gives a qualitative visualization of how the unstructured models struggles in generalizing to poses that are sufficiently different from the ones in the training set.

We employ the “Transitions” portion of the AMASS dataset to further study the performance of the model when more training data and a larger diversity of motions) is available for a single subject. This dataset contains 110110 sequences of one individual, with ≈1000+{\approx}1000+ frames/sequence. We randomly sample ≈250{\approx}250 frames from each sequence, randomly select 8080 sequences for training, and keep the remaining 3030 sequences for testing; see our supplementary material. Results are shown in Table 2. The conclusions are analogous to the ones we made from DFaust. Further, note that in this more difficult dataset containing a larger variety of more complex motions, the unstructured model struggles even more significantly (U→RU{\rightarrow}R: +70%\mathbf{+70\%} in F-score). As the model is exposed to more poses compared to DFaust, the reconstruction performance is also improved. Moving from DFaust to Transitions results in a +1%\mathbf{+1\%} in F-score for the deformable model.

3 Tracking

We validate our tracking technique on two sequences from the DFaust dataset; see Figure 6. Note these are test sequences, and were not used to train our model. The prior p(θ)p(\boldsymbol{\theta}) (Section 5.2) and the stochastic optimization ⊛\circledast (Section 5.1) can be applied to both unstructured (U) and structured (D) representations, with the latter leading to significantly better tracking performance. The quantitative results reported for the “easy” (Table 3) and “hard” (Table 4) tracking sequences are best understood by watching our supplementary video. It is essential to re-state that we are not trying to beat traditional baselines, but rather seek to illustrate how NASA, once trained, can be readily used as a 3D representation for classical vision tasks. For the purpose of this illustration, we use only noisy panoptic point clouds (i.e. complete rather than incomplete data), and do not use any discriminative per-frame re-initializer as would typically be employed in a contemporary tracking system.

4 Discussion

The recent success of neural implicit representations of geometry, introduced by , has heavily relied on the fact that the geometry in ShapeNet datasets is canonicalized: scaled to unit ranges and consistently oriented. Research has highlighted the importance of expressing information in a canonical frame , and one could interpret our method as a way to achieve this within the realm of articulated motion. To understand the shortcomings of unstructured models, one should remember that as an object moves, much of the local geometric details remain invariant to articulation (e.g. the geometry of a wristwatch does not change as you move your arm). However, unstructured pose conditioned models are forced to memorize these details in any pose they seek to reconstruct. Hence, as one evaluates unstructured models outside of their training manifold, their performance collapses – as quantified by the +49%+49\% performance change as we move from unstructured to rigid models; see Table 1. One could also argue that given sufficient capacity, a neural network should be able to learn the concept of coordinate frames and transformations. However, multiplicative relationships between inputs (e.g. dot products) are difficult to learn for neural networks [26, Sec. 2.3]. As changes of coordinate frames are nothing but collections of dot products, one could use this reasoning to justify the limited performance of unstructured models. We conclude by clearly contrasting our method, targeting the modeling of O(x∣θ)\mathcal{O}(\mathbf{x}|\boldsymbol{\theta}) to those that address shape completion O(x∣0pt)\mathcal{O}(\mathbf{x}|0pt) . In contrast to these, our solution, to the best of our knowledge, represents the first attempt to create a “neural implicit rig” – from a computer graphics perspective – for articulated deformation modeling.

Conclusions

We introduce a novel neural representation of a particularly important class of 3D objects: articulated bodies. We use a structured neural occupancy approach, enabling both direct occupancy queries and deformable surface representations that are competitive with classic hand-crafted mesh representations. The representation is fully differentiable, and enables tracking of realistic articulated bodies – traditionally a complex task – to be almost trivially implemented. Crucially, our work demonstrates the value of incorporating a task-appropriate inductive bias into the neural architecture. By acknowledging and encoding the quasi-rigid part structure of articulated bodies, we represent this class of objects with higher quality, and significantly better generalization.

References

Acknowledgments

We would like to acknowledge Hugues Hoppe, Paul Lalonde, Erwin Coumans, Angjoo Kanazawa, Alec Jacobson, David Levine, and Christopher Batty for the insightful discussions. Gerard Pons-Moll is funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) - 409792180. Andrea Tagliasacchi is funded by NSERC Discovery grant RGPIN-2016-05786, NSERC Collaborative Research and Development grant CRDPJ 537560-18, and NSERC Research Tool Instruments RTI-16-2018.

Supplementary material

In dense real-time tracking applications there are examples of articulated models that use hybrid representations. Similarly to our work, they also provide constant-time occupancy queries without the need for acceleration data structures. However, these shape models consist of simple rigid primitives, or use an artist-designed template. In contrast, we learn a deformable shape model from data, and allow users to query the pose-corrected occupancy anywhere in space. In more detail:

Schmidt et al. : relies on a discretization of the SDF on grids and/or primitives to answer distance queries, only models piecewise rigid parts (compare our rigid (R) variant), and requires the construction of the parts SDF before tracking – a process that relies on user interaction.

Tkach et al. : the quality of approximation produced by our method is vastly superior to that achieved by sphere-meshes. Further, similarly to , the template is specified manually.

Taylor et al. : similarly to , the tracking template is specified manually, and the technique requires a mixture of mesh-based closest point queries and implicit function queries.

Thiery et al. : the approximation quality argument is analogous to that of , while to construct such models one need a consistently meshed motion sequence – an extremely stringent requirement in practice.

2 Ablation studies

Please see the animated results of reconstruction and tracking in the supplementary video. Please see the details of the dataset in the supplementary data split files. Note that, except Figure 9 where we use AMASS/Transitions due to its diversity of poses, we adopt AMASS/DFaust for all the other studies. Also note that due to computational limitations, we evaluate on one motion sequence only in Figure 10. We select a sequence that has a median reconstruction performance as a representative example.

One can view O(x∣θ)\mathcal{O}(\mathbf{x}|\boldsymbol{\theta}) as a binary classifier that aims to separate the interior of the shape from its exterior. Accordingly, one can use a binary cross-entropy loss for optimization, but our experiments suggest that an L2 loss perform slightly better. Hence, we employ the L2 loss for all of our experiments; see Table 5. We also validate the importance of the skinning weights loss in Table 6 and observe a big improvement when Lweights\mathcal{L}_{\text{weights}} is included.

Linear subspace projection ΠΠ\Pi – Figure 8

Note that the rigid model (R) actually outperforms the deformable model (D) if one removes the learnt linear dimensionality reduction (D∖ΠD{\setminus}\Pi); see Table 7. This is a result only observed on the test set, while on the training set D∖ΠD{\setminus}\Pi performs comparably. In other words, Π\Pi helps our model to achieve better generalization by enforcing a sparse representation of pose. In Table 8, we report the results of an ablation study on the dimensionality of the projection, which was the basis for the selection of D=4D{=}4.

Analysis of pose representations – Figure 9

In Table 9, we ablate several representations for the pose θ\boldsymbol{\theta} used by the deformable model. We start by just using the collection of homogeneous transformations {Bb−1}\{\mathbf{B}_{b}^{-1}\}. Note that the query point encoded in various coordinate frames is also an effective pose representation {Bb−1x}\{\mathbf{B}_{b}^{-1}\mathbf{x}\}, which has a much lower dimensionality. Finally, we notice that rather than using the query point, one can pick a fixed point to represent pose. While any fixed point can be used, we select the origin of the model t0\mathbf{t}_{0} for simplicity, resulting in {Bb−1t0}\{\mathbf{B}_{b}^{-1}\mathbf{t}_{0}\}. The resulting representation is compact and effective. Table 10 shows a similar analysis for the unstructured model (metrics for D provided for reference). Note how the performance of U can be significantly improved by providing the network with the encoding of the query point x\mathbf{x} in various coordinate frames – that is, the network is no longer required to “learn” the concept of changes of coordinates.

Analysis of model size – Figure 10

Both rigid (R) and deformable (D) models significantly outperform the results of the unstructured (U) model, as we increase the neural network’s layer size and approach the network capacity employed by .

Tracking ablations – Figure 11

In the tracking application, we ablate with respect to the pose prior (p(θ)p(\boldsymbol{\theta})) and the use of random perturbations to approximate the distance function via convolution (⊛\circledast). First, note that the best results are achieved when both of these components are enabled, across all metrics. We compare the performance of our models on easy vs. hard sequences. Hard sequences more clearly illustrate the advantages of the algorithms proposed. We validate the usefulness of the pose prior in avoiding tracking failure (e.g. IoU:44.31%→86.15%IoU:44.31\%\rightarrow 86.15\%). The use of random perturbations allow the optimization to converge more precisely (Chamfer: .00258→.00006.00258\rightarrow.00006).

Metrics distribution on AMASS/DFaust

Rather than reporting aggregated statistics, we visualize the IoU errors of all of the 100100 DFaust experiments, and sort them by the performance achieved by the deformable model (D). Note how the deformable model achieves consistent performance across the dataset. There are only two sequences where the rigid model performs better than the deformable model.