On Self-Contact and Human Pose

Lea Müller, Ahmed A. A. Osman, Siyu Tang, Chun-Hao P. Huang, Michael J. Black

Introduction

Self-contact takes many forms. We touch our bodies both consciously and unconsciously . For the major limbs, contact can provide physical support, whereas we touch our faces in ways that convey our emotional state. We perform self-grooming, we have nervous gestures, and we communicate with each other through combined face and hand motions (e.g. “shh”). We may wring our hands when worried, cross our arms when defensive, or put our hands behind our head when confident. A Google search for “sitting person” or “thinking pose” for example, will return images, the majority of which, contain self-contact.

Although self-contact is ubiquitous in human behavior, it is rarely explicitly studied in computer vision. For our purposes, self-contact comprises “self touch” (where the hands touch the body) and contact between other body parts (e.g. crossed legs). We ignore body parts that are frequently in contact (e.g. at the crotch or armpits) and focus on contact that is communicative or functional. Our goal is to estimate 3D human pose and shape (HPS) accurately for any pose. When self-contact is present, the estimated pose should reflect the true 3D contact.

Unfortunately, existing methods that compute 3D bodies from images perform poorly on images with self-contact; see Fig. 1. Body parts that should be touching generally are not. Recovering human meshes from images typically involves either learning a regressor from pixels to 3D pose and shape , or fitting a 3D model to image features using an optimization method . The learning approaches rely on labeled training data. Unfortunately, current 2D datasets typically contain labeled keypoints or segmentation masks but do not provide any information about 3D contact. Similarly, existing 3D datasets typically avoid capturing scenarios with self-contact because it complicates mesh processing. What is missing is a dataset with in-the-wild images and reliable data about 3D self-contact.

To address this limitation, we introduce three new datasets that focus on self-contact at different levels of detail. Additionally, we introduce two new optimization-based methods that fit 3D bodies to images with contact information. We leverage these to estimate pseudo ground-truth 3D poses with self-contact. To make reasoning about contact between body parts, the hands, and the face possible, we represent pose and shape with the SMPL-X body model, which realistically captures the body surface details, including the hands and face. Our new datasets then let us train neural networks to regress 3D HPS from images of people with self-contact more accurately than state-of-the-art methods.

To begin, we first construct a 3D Contact Pose (3DCP) dataset of 3D meshes where body parts are in contact. We do so using two methods. First, we use high-quality 3D scans of subjects performing self-contact poses. We extend previous mesh registration methods to cope with self-contact and register the SMPL-X mesh to the scans. To gain more variety of poses, we search the AMASS dataset for poses with self-contact or “near” self-contact. We then optimize these poses to bring nearby parts into full contact while resolving interpenetration. This provides a dataset of valid, realistic, self-contact poses in SMPL-X format.

Second, we use these poses to collect a novel dataset of images with near ground-truth 3D pose. To do so, we show rendered 3DCP meshes to workers on Amazon Mechanical Turk (AMT). Their task is to Mimic The Pose (MTP) as accurately as possible, including the contacts, and submit a photograph. We then use the “true” pose as a strong prior and optimize the pose in the image by extending SMPLify-X to enforce contact. A key observation is that, if we know about self-contact (even approximately), this greatly reduces pose ambiguity by removing degrees of freedom. Thus, knowing contact makes the estimation of 3D human pose from 2D images more accurate. The resulting method, SMPLify-XMC (for SMPLify-X with Mimicked Contact), produces high-quality 3D reference poses and body shapes in correspondence with the images.

Third, to gain even more image variety, we take images from three public datasets and have them labeled with discrete body-part contacts. This results in the Discrete Self-Contact (DSC) dataset. To enable this, we define a partitioning of the body into regions that can be in contact. Given labeled discrete contacts, we extend SMPLify to optimize body shape using image features and the discrete contact labels. We call this method SMPLify-DC, for SMPLify with Discrete Self-Contact.

Given the MTP and DSC datasets, we finetune a recent HPS regression network, SPIN . When we have 3D reference poses, i.e. for MTP images, we use these as though they were ground truth and do not optimize them in SPIN. When discrete contact annotations are available, i.e. for DSC images, we use SMPLify-DC to optimize the fit in the SPIN training loop. Fine-tuning SPIN on MTP and DSC significantly improves accuracy of the regressed poses when there is contact (evaluated on 3DPW ). Surprisingly, the results on non-self-contact poses also improve, suggesting that (1) gathering accurate 3D poses for in-the-wild images is beneficial, and (2) that self-contact can provide valuable constraints that simplify pose estimation.

We call our regression method TUCH (Towards Understanding Contact in Humans). Figure 1 illustrates the effect of exploiting self-contact in 3D HPS estimation. By training with self-contact, TUCH significantly improves the physical plausibility.

In summary, the key contributions of this paper are: (1) We introduce TUCH, the first HPS regressor for self-contact poses, trained end-to-end. (2) We create a novel dataset of 3D human meshes with realistic contact (3DCP). (3) We define a “Mimic The Pose” MTP task and a new optimization method to create a novel dataset of in-the-wild images with accurate 3D reference data. (4) We create a large dataset of images with reference poses that use discrete contact labels. (5) We show in experiments that taking self-contact information into account improves pose estimation in two ways (data and losses), and in turn achieves state-of-the-art results on 3D pose estimation benchmarks. (6) The data and code are available for research purposes.

Related Work

3D pose estimation with contact. Despite rapid progress in 3D human pose estimation , and despite the role that self-contact plays in our daily lives, only a handful of previous works discuss self-contact. Information about contact can benefit 3D HPS estimation in many ways, usually by providing additional physical constraints to prevent undesirable solutions such as interpenetration between limbs.

Body contact. Lee and Chen approximate the human body as a set of line segments and avoid collisions between the limbs and torso. Similar ideas are adopted in where line segments are replaced with cylinders. Yin et al. build a pose prior to penalize deep interpenetration detected by the Open Dynamics Engine . While efficient, these stickman-like representations are far from realistic. Using a full 3D body mesh representation, Pavlakos et al. take advantage of physical limits and resolve interpenetration of body parts by adding an interpenetration loss. When estimating multiple people from an image, Zanfir et al. use a volume occupancy exclusion loss to prevent penetration. Still, other work has exploited textual and ordinal descriptions of body pose . This includes constraints like “Right hand above the hips”. These methods, however, do not consider self-contact.

Most similar to us is the work of Fieraru et al. , which utilizes discrete contact annotations between people. They introduce contact signatures between people based on coarse body parts. This is similar to how we collect the DSC dataset. Contemporaneous with our work, Fieraru et al. extend this to self-contact with a 2-stage approach. They train a network to predict “self-contact signatures”, which are used for optimization-based 3D pose estimation. In contrast, TUCH is trained end-to-end to regress body pose with contact information.

World contact. Multiple methods use the 3D scene to help estimate the human pose. Physical constraints can come from the ground plane , an object , or contextual scene information . Li et al. use a DNN to detect 2D contact points between objects and selected body joints. Narasimhaswamy et al. categorize hand contacts into self, person-person, and object contacts and aim to detect them from in-the-wild images. Their dataset does not provide reference 3D poses or shape.

All the above works make a similar observation: human pose estimation is not a stand-alone task; considering additional physical contact constraints improves the results. We go beyond prior work by addressing self-contact and showing how training with self-contact data improves pose estimation overall.

3D body datasets. While there are many datasets of 3D human scans, most of these have people standing in an “A” or “T” pose to explicitly minimize self-contact . Even when the body is scanned in varied poses, these poses are designed to avoid self-contact . For example, the FAUST dataset has a few examples of self-contact and the authors identify these as the major cause of error for scan processing methods . Recently, the AMASS dataset unifies 15 different optical marker-based motion capture (mocap) datasets within a common 3D body parameterization, offering around 170k meshes with SMPL-H topology. Since mocap markers are sparse and often do not cover the hands, such datasets typically do not explicitly capture self-contact. As illustrated in Table 1, none of these datasets explicitly addresses self-contact.

Pose mimicking. Our Mimic-The-Pose dataset uses the idea that people can replicate a pose that they are shown. Several previous works have explored this idea in different contexts. Taylor et al. crowd-source images of people in the same pose by imitation. While they do not know the true 3D pose, they are able to train a network to match images of people in similar poses. Marinoiu et al. motion capture subjects reenacting a 3D pose from a 2D image. They found that subjects replicated 3D poses with a mean joint error of around 100mm. This is on par with existing 3D pose regression methods, pointing to people’s ability to approximately recreate viewed poses. Fieraru et al. ask subjects to reproduce contact from an image in a lab setting. They manually annotate the contact, whereas our MTP task is done in people’s homes and SMPLify-XMC is used to automatically optimize the pose and contact.

Self-Contact

An intuitive definition of contact between two meshes, e.g. a human and an object, is based on intersecting triangles. Self-contact, however, must be formulated to exclude common, but not functional, triangle intersections, e.g. at the crotch or armpits. Intuitively, vertices are in self-contact if they are close in Euclidean distance (near zero) but distant in geodesic distance, i.e. far away on the body surface.

Given a mesh MM with vertices MVM_{V}, we define two vertices viv_{i} and vjv_{j} ∈MV\in M_{V} to be in self-contact, if (i) ∥vi−vj∥<teucl\left\|v_{i}-v_{j}\right\|<t_{\mathit{eucl}}, and (ii) geo(vi,vj)>tgeo\mathit{geo}(v_{i},v_{j})>t_{\mathit{geo}}, where teuclt_{\mathit{eucl}} and tgeot_{\mathit{geo}} are predefined thresholds and geo(vi,vj)\mathit{geo}(v_{i},v_{j}) denotes the geodesic distance between viv_{i} and vjv_{j}. We use shape-independent geodesic distances precomputed on the neutral, mean-shaped SMPL and SMPL-X models.

Following this definition, we denote the set of vertex pairs in self-contact as MC:={(vi,vj)∣vi,vj∈MV and vi,vj satisfy Definition \refdefinition-selfcontact}\mathnormal{M}_{C}:=\{(v_{i},v_{j})|v_{i},v_{j}\in M_{V}\text{ and }v_{i},v_{j}\text{ satisfy Definition~{}\ref{definition-selfcontact}}\}. MM is a self-contact mesh when ∣MC∣>0|\mathnormal{M}_{C}|>0. We further define an operator U(⋅)\mathcal{U}(\cdot) that returns a set of unique vertices in MCM_{C}, and an operator fg(⋅)f_{g}(\cdot) that takes viv_{i} as input and returns the Euclidean distance to the nearest vjv_{j} that is far enough in the geodesic sense. Formally, U(MC)={v0,v1,v2,…,vn}\mathcal{U}(M_{C})=\{v_{0},v_{1},v_{2},\dots,v_{n}\}, where ∀vi∈U(MC)  ,∃vj∈U(MC)\forall v_{i}\in\mathcal{U}(M_{C})\,\ ,\exists v_{j}\in\mathcal{U}(M_{C}), such that (vi,vj)∈MC(v_{i},v_{j})\in M_{C}. fg(vi):=min⁡vj∈MG(vi)∥vi−vj∥f_{g}(v_{i}):=\min_{v_{j}\in\mathnormal{M}_{G}(v_{i})}\left\|v_{i}-v_{j}\right\|, where MG(vi):={vj∣geo(vi,vj)>tgeo}\mathnormal{M}_{G}(v_{i}):=\{v_{j}|\mathit{geo}(v_{i},v_{j})>t_{\mathit{geo}}\}.

We further cluster self-contact meshes into distinct types. To that end, we define self-contact signatures S∈{0,1}K×K\mathbf{S}\in\{0,1\}^{K\times K}; see for a similar definition. We first segment the vertices of a mesh into KK regions RkR_{k}, where Rk∩Rl=∅R_{k}\cap R_{l}=\emptyset for k≠lk\neq l and ⋃k=1KRk=MV\bigcup_{k=1}^{K}R_{k}=M_{V}. We use fine signatures to cluster self-contact meshes from AMASS (see Sup. Mat.) and rough signatures (see Fig. 18) for human annotation.

Two regions RkR_{k} and RlR_{l} are in contact if ∃(vi,vj)∈MC\exists(v_{i},v_{j})\in M_{C}, such that vi∈Rkv_{i}\in R_{k} and vj∈Rlv_{j}\in R_{l} holds. If RkR_{k} and RlR_{l} are in contact, Skl=Slk=1\mathbf{S}_{kl}=\mathbf{S}_{lk}=1. MSM_{\mathbf{S}} denotes the contact signature for mesh MM.

Self-Contact Datasets

Our goal is to create datasets of in-the-wild images paired with 3D human meshes as pseudo ground truth. Unlike traditional pipelines that collect images first and then annotate them with pose and shape parameters , we take the opposite approach. We first curate meshes with self-contact and then pair them with images through a novel pose mimicking and fitting procedure. We use SMPL-X to create the 3DCP and MTP dataset to better fit contacts between hands and bodies. However, to fine-tune SPIN , we convert MTP data to SMPL topology, and use SMPLify-DC when optimizing with discrete contact.

We create 3D human meshes with self-contact in two ways: with 3D scans and with motion capture data.

3DCP Mocap. While mocap datasets are usually not explicitly designed to capture self-contact, it does occur during motion capture. We therefore search the AMASS dataset for poses that satisfy our self-contact definition. We find that some of the selected meshes from AMASS contain small amounts of self-penetration or near contact. Thus, we perform self-contact optimization to fix this while encouraging contact, as shown in Fig. 3; see Sup. Mat. for details.

2 Mimic-The-Pose (MTP) Data

To collect in-the-wild images with near ground-truth 3D human meshes, we propose a novel two-step process (see Fig. 4). First, using meshes from 3DCP as examples, workers on AMT are asked to mimic the pose as accurately as possible while someone takes their photo showing the full body (the mimicked pose). Mimicking poses may be challenging for people when only a single image of the pose is presented . Thus, we render each 3DCP mesh from three different views with the contact regions highlighted (the presented pose). We allot 3 hours time for ten poses. Participants also provide their height and weight. All participants gave informed consent for the capture and the use of their imagery. Please see Sup. Mat. for details.

In the second and third stage, we jointly optimize θ\theta, β\beta, and Π\Pi to minimize

The third stage actives LS\mathcal{L}_{S} for fine-grained self-contact optimization, which resolves interpenetration while encouraging contact. The objective is LS=λCLC+λPLP+λALA\mathcal{L_{S}}=\lambda_{C}\mathcal{L}_{C}+\lambda_{P}\mathcal{L}_{P}+\lambda_{A}\mathcal{L}_{A}. Vertices in contact are pulled together via a contact term LC\mathcal{L}_{C}; vertices inside the mesh are pushed to the surface via a pushing term LP\mathcal{L}_{P}, and LA\mathcal{L}_{A} aligns the surface normals of two vertices in contact.

To compute these terms, we must first find which vertices are inside, MI⊂MVM_{I}\subset M_{V}, or in contact, MC⊂MVM_{C}\subset M_{V}. MCM_{C} is computed following Definition 3.1 with tgeo=30t_{\mathit{geo}}=30cm and teucl=2t_{\mathit{eucl}}=2cm. The set of inside vertices MIM_{I} is detected by generalized winding numbers . SMPL-X is not a closed mesh and thus complicating the test for penetration. Consequently, we close it by adding a vertex at the back of the mouth. In addition, neighboring parts of SMPL and SMPL-X often intersect, e.g. torso and upper arms. We identify such common self-intersections and filter them out from MIM_{I}. See Sup. Mat. for details. To capture fine-grained contact, we map the union of inside and contact vertices onto the HD SMPL-X surface, i.e. MD=HD(MI∪MC)M_{D}=\mathit{HD}(M_{I}\cup M_{C}), which is further segmented into an inside MDIM_{D_{I}} and outside MDI∁M_{D_{I}^{\complement}} subsets by testing for intersections. The self-contact objectives are defined as

fgf_{g} denotes the function that finds the closest point pj∈MDp_{j}\in M_{D}. MDCM_{D_{C}} is the subset of vertices in contact in MDM_{D}. We use α1=α2=0.005\alpha_{1}=\alpha_{2}=0.005, β1=1.0\beta_{1}=1.0, and β2=0.04\beta_{2}=0.04 and visualize the contact and pushing functions in the Sup. Mat. Fig. 5 shows examples of our pseudo ground-truth meshes.

3 Discrete Self-Contact (DSC) Data

Images in the wild collected for human pose estimation normally come with 2D keypoint annotations, body segmentation, or bounding boxes. Such annotations lack 3D information. Discrete self-contact annotation, however, provides useful 3D information about pose. We use K=24K=24 regions and label their pairwise contact for three publicly available datasets, namely Leeds Sports Pose (LSP), Leeds Sports Pose Extended (LSPet), and DeepFashion (DF). An example annotation is visualized in Fig. 18. Of course, such labels are noisy because it can be difficult to accurately determine contact from an image. See Sup. Mat. for details.

4 Summary of the Collected Data

Our 3DCP human mesh dataset consists of 190 meshes containing self-contact from 6 subjects, 159 SMPL-X bodies fit to commercial scans from AGORA , and 1304 self-contact optimized meshes from mocap data. From these 1653 poses, we collect 3731 mimicked pose images from 148 unique subjects (52 female; 96 male) for MTP and fit pseudo ground-truth SMPL-X parameters. MTP is diverse in body shapes and ethnicities. Our DSC dataset provides annotations for 30K images.

TUCH

Finally, we train a regression network that has the same design as SPIN . At each training iteration, the current regressor estimates the pose, shape, and camera parameters of the SMPL model for an input image. Using ground-truth 2D keypoints, an optimizer refines the estimated pose and shape, which are used, in turn, to supervise the regressor. We follow this regression-optimization scheme for DSC data, where we have no 3D ground truth. To this end, we adapt the in-the-loop SMPLify routine to account for discrete self-contact labels, which we term SMPLify-DC. For MTP images, we use the pseudo ground truth from SMPLify-XMC as direct supervision with no optimization involved. We explain the losses of each routine below.

Regressor. Similar to SPIN, the regressor of TUCH predicts pose, shape, and camera, with the loss function:

EJE_{J} denotes the joint re-projection loss. LP\mathcal{L}_{P} and LC\mathcal{L}_{C} are self-contact loss terms used in LS\mathcal{L}_{S} in SMPLify-XMC, where LP\mathcal{L}_{P} penalizes mesh intersections and LC\mathcal{L}_{C} encourages contact. Further, EθE_{\theta} and EβE_{\beta} are L2-Losses that penalize deviation from the pseudo ground-truth pose and shape.

Optimizer. We develop SMPLify-DC to fit pose θopt\theta_{opt}, shape βopt\beta_{opt}, and camera Πopt\Pi_{opt} to DSC data, taking ground-truth keypoints and contact as constraints.

Typically, in human mesh optimization methods the camera is fit first, then the model parameters follow. However, we find that this can distort body shape when encouraging contact. Therefore, we optimize shape and camera translation first, using the same camera fitting loss as in . After that, body pose and global orientation are optimized under the objective

The discrete contact loss, LD\mathcal{L_{D}}, penalizes the minimum distance between regions in contact. Formally, given a contact signature S\mathbf{S} where Sij=Sji=1\mathbf{S}_{ij}=\mathbf{S}_{ji}=1 if two regions RiR_{i} and RjR_{j} are annotated to be in contact, we define

Given the optimized pose θopt\theta_{opt}, shape βopt\beta_{opt}, and camera Πopt\Pi_{opt}, we compute the re-projection error and the minimum distance between the regions in contact. When the re-projection error improves, and more regions with contact annotations are closer than before, we keep the optimized pose as the current best fit. When no ground truth is available, the current best fits are used to train the regressor.

We make three observations: (1) The optimizer is often able to fix incorrect poses estimated by the regressor because it considers the ground-truth keypoints and contact (see Fig. 7). (2) Discrete contact labels bring overall improvement by helping resolve depth ambiguity (see Fig. 8). (3) Since we have mixed data in each mini-batch, the direct supervision of MTP data improves the regressor, which benefits SMPLify-DC by providing better initial estimates.

Implementation details. We initialize our regression network with SPIN weights . For SMPLify-DC, we run 10 iterations per stage and do not use the HD operator to speed up the optimization process. For the 2D re-projection loss, we use ground-truth keypoints when available and, for MTP and DF images, OpenPose detections weighted by confidence. From DSC data we only use images where the full body is visible and ignore annotated region pairs that are connected in the DSC segmentation (see Sup. Mat.).

Evaluation

We evaluate TUCH on the following three datasets: 3DPW , MPI-INF-3DHP , and 3DCP Scan. This latter dataset consists of RGB images taken during the 3DCP Scan scanning process. While TUCH has never seen these images or subjects, the contact poses were mimicked in creation of MTP, which is used in training.

We use standard evaluation metrics for 3D pose, namely Mean Per-Joint Position Error (MPJPE) and the Procrustes-aligned version (PA-MPJPE), and Mean Vertex-to-Vertex Error (MV2VE) for shape and contact. Tables 2 and 3 summarize the results of TUCH on 3DPW and 3DCP Scan. Interestingly, TUCH is more accurate than SPIN on 3DPW. See Sup. Mat. for results of fine-tuning EFT.

We further evaluate our results w.r.t. contact. To this end, we divide the 3DPW test set into subsets, namely for tgeo=50t_{geo}=50cm: self-contact (teucl<1t_{eucl}<1cm), no self-contact (teucl>5t_{eucl}>5cm), and unclear (11cm <teucl<5<t_{eucl}<5cm). For 3DPW we obtain 8752 self-contact, 16752 no self-contact, and 9491 unclear poses. Table 4 shows a clear improvement on poses with contact and unclear poses compared to a smaller improvement on poses without contact.

To further understand the improvement of TUCH over SPIN, we break down the improved MPJPE in 3DPW self-contact into the pairwise body-part contact labels defined in the DSC dataset. Specifically, for each contact pair, we search all poses in 3DPW self-contact that have this particular self-contact. We find a clear improvement for a large number of contacts between two body parts, frequently between arms and torso, or e.g. left hand and right elbow, which is common in arms-crossed poses (see Fig. 9).

TUCH incorporates self-contact in various ways: annotations of training data, in-the-loop fitting, and in the regression loss. We evaluate the impact of each in Table 5. S+ is SPIN but it sees MTP+DSC images in fine-tuning and runs standard in-the-loop SMPLify with no contact information. S++ is S+ but uses pseudo ground truth computed with SMPLify-XMC on MTP images; thus self-contact is used to generate the data but nowhere else. S+ vs. SPIN suggests that, while poses in 3DCP Scan appear in MTP, just seeing similar poses for training and testing does not yield improvement. S+ vs. TUCH is a fair comparison as both see the same images during training. The improved results of TUCH confirm the benefit of using self-contact.

Conclusion

In this work, we address the problem of HPS estimation when self-contact is present. Self-contact is a natural, common occurrence in everyday life, but SOTA methods fail to estimate it. One reason for this is that no datasets pairing images in the wild and 3D reference poses exist. To address this problem we introduce a new way of collecting data: we ask humans to mimic presented 3D poses. Then we use our new SMPLify-XMC method to fit pseudo ground-truth 3D meshes to the mimicked images, using the presented pose and self-contact to constrain the optimization. We use the new MTP data along with discrete self-contact annotations to train TUCH; the first end-to-end HPS regressor that also handles poses with self-contact. TUCH uses MTP data as if it was ground truth, while the discrete, DSC, data is exploited during SPIN training via SMPLify-DC. Overall, incorporating contact improves accuracy on standard benchmarks like 3DPW, remarkably, not only for poses with self-contact, but also for poses without self-contact.

Acknowledgments: We thank Tsvetelina Alexiadis and Galina Henz for their help with data collection and Vassilis Choutas for the SMPL-X body measurements and his implementation of Generalized Winding Numbers. The authors thank the International Max Planck Research School for Intelligent Systems (IMPRS-IS) for supporting Lea Müller and Ahmed A. A. Osman.

Disclosure: MJB has received research gift funds from Adobe, Intel, Nvidia, Facebook, and Amazon. While MJB is a part-time employee of Amazon, his research was performed solely at, and funded solely by, Max Planck. MJB has financial interests in Amazon, Datagen Technologies, and Meshcapade GmbH.

References

Appendix A Self-Contact Datasets

Raw scans have varying topology. To bring a corpus of scans to a common topology is the process of “registration”. Most traditional registration methods ignore interpenetration and self-contact. Registering our self-contact scans without modeling self-contact would result in self-penetration, particularly where the extremities contact the body. We address this by modifying the registrations objective function to encourage self-contact without penetration.

Specifically, the fitting objective includes a data term ESE_{S} evaluating the goodness of fit of the the vertices xvx_{v} on the template VV to nn randomly sampled points, xsx_{s} on the surface of the scan SS

where ρ\rho is the Geman-McClure robust penalty function.

Additionally, we introduce a self-contact preserving energy term ECE_{C} to the objective function. The term ECE_{C} helps to minimize and preserve the point-to-plane distance between body parts that are in contact. ECE_{C} considers the set of contacting vertex pairs MCM_{C} defined by Definition 3.1 in the main paper. For each tuple (vi,vj)(v_{i},v_{j}) in MCM_{C}, we minimize the point-to-plane distance between triangles including viv_{i} and the triangular planes including vjv_{j}. The contact energy term ensures that body parts that are in contact remain in contact.

A.1.2 3DCP Mocap.

Sampling meshes from AMASS. First, each mocap. sequence is sampled at half of its original frame rate. For each sampled mesh, we compute the contact signatures MSM_{\mathbf{S}} with teucl=3t_{\mathit{eucl}}=3cm, tgeo=30t_{\mathit{geo}}=30cm and K=98K=98. The regions are visualized in Fig. 11. We select only one pose for each unique signature, while ignoring contact when it occurs in more than 1% of the data. We obtain a subset of 20,114 poses with unique self-contact signatures, as shown in Fig. 12.

Self-Contact Optimization. Here we provide details of the self-contact optimization for body meshes from the AMASS dataset. In this optimization, vertex pairs in MCM_{C} are further pulled together via a contact term LC\mathcal{L}_{C} and vertices inside the mesh are pushed to the surface via a pushing term LP\mathcal{L}_{P}, while LO\mathcal{L}_{O} ensures that vertices far away from contact regions stay in place. Note that LP\mathcal{L}_{P} and LC\mathcal{L}_{C} are slightly different from the loss terms in the main paper. LH\mathcal{L}_{H} is a prior for contact between hand and body and LA\mathcal{L}_{A} aligns the vertex normals when contact happens.

where θh\theta_{h} denote the hand pose vector of the SMPL-X model. Further,

where min⁡vj∈U(MC)geo(vi,vj)=1\min_{v_{j}\in\mathcal{U}(M_{C})}geo(v_{i},v_{j})=1 if MC=∅M_{C}=\emptyset and δ2=4\delta_{2}=4. Lastly, we use a term, LA\mathcal{L}_{A}, that encourages the vertex normals N(v)N(v) of vertices in contact to be aligned but in opposite directions:

Hand-on-Body Prior. Hands and fingers play an important role as they frequently make contact with the body. However, they have many degrees of freedom, which makes their optimization challenging. Therefore, we learn a hand-on-body prior from 1279 self contact registrations. For this, we use only poses where the minimum point-to-mesh distance between hand and body is <1mm<1mm. These are 718 and 701 poses for the right and left hand, respectively. Since left and right hand are symmetric in SMPL-X, we unite left and right hand poses. Across the 1429 poses, the mean distances per hand vertex to the body surface, dm(vi)d_{m}(v_{i}) ranges per vertex from 1.791.79 to 5.525.52 cm, as visualized in Fig. 13. To obtain the weights hvih_{v_{i}} in LH\mathcal{L}_{H}, we normalize dm(vi)d_{m}(v_{i}) to $,denotedas, denoted ass(d_{m}(v_{i})),andobtainthevertexweightby, and obtain the vertex weight byh_{v_{i}}=-s(d_{m}(v_{i}))+1$.

A.2 Mimic-The-Pose (MTP) Data

AMT task details. It can be challenging to mimic a pose precisely. To simplify the process for workers on AMT, we give detailed instructions, add thumbnails to compare the own image with the presented one and, most importantly, highlight the contact areas; see Fig. 14. To gain more variety, we also request that participants make small changes in the environment for each image, e.g. by rotating the camera, changing clothes, or turning lights on/off. We also ask participants to mimic the global orientation of the center image. For more variety in global orientation, we vary body roll from −90∘-90^{\circ} to 90∘90^{\circ} in 30∘30^{\circ} steps, resulting in seven different presented global orientations. For example, in the first and third row of Fig. 14, the center image shows the presented pose from a frontal view. In the second and fourth row, the center body has different orientations. We also ask participants for their height, weight, and gender (M, F, and Non-Binary).

SMPLify-XMC. In the first stage, we optimize body shape β\beta and camera Π\Pi (focal length, rotation and translation), and body global orientation θg\theta_{g}, using ground-truth height in meters, hgth_{gt}, and weight in kg, wgtw_{gt}. The objective function of the first stage is given as

LM=e100∣Mh−hgt∣+e∣Mw−wgt∣\mathcal{L}_{M}=e^{100|M_{h}-h_{gt}|}+e^{|M_{w}-w_{gt}|} is the measurements loss, where MhM_{h} and MwM_{w} are height and weight of mesh MM. We compute height and weight from mesh v-template in a zero pose (T-pose). For height, we compute the distance between the top of the head and the mean point between left and right heel. For weight, we compute the mesh volume and multiply it by 985 kg/m3kg/m^{3}, which approximates human body density. Lθg\mathcal{L}_{\theta_{g}} is a loss that allows rotation around the y-axis, but not around x and z.

In Fig. 15 we visualize the pushing and pulling terms used in the SMPLify-XMC objective. We use 6 PCA components for the hand pose space and initialize the fitting with a mean hand pose. In contrast to SMPLify-X we do not ignore hip joints and double the joint weights for knees and elbows. Before optimization, we resize images and keypoints to a maximum height or width of 500 pixel. Similar to SMPLify-X we use the PyTorch implementation of fast L-BFGS with strong Wolf line search as the optimizer . We do not use the VPoser pose prior for SMPLify-XMC because we have a strong prior from the presented pose.

We notice that the presented global orientation is not always mimicked well. For example, in row 4 of Fig. 14 the presented global orientation has a 60 degree rotation, whereas the mimicked image is taken from a frontal view. To better initialize the optimization, we select the best body orientation, θg\theta_{g}, among the seven presented ones based on their re-projection errors; then we compute the camera translation by again minimizing the re-projection error. We set the initial focal length, fxf_{x} and fyf_{y}, to 2170, which is the average of available EXIF data. These values, along with mean shape and presented pose are used to initialize the optimization.

In addition, SMPL and SMPL-X have not been trained to avoid self intersection. Therefore, we identify seven body segments that tend to intersect themselves, e.g. torso and upper arms (see Fig. 16). We test each segment for self intersection and thereby filter irrelevant intersections from MIM_{I}.

MTP Dataset Details. We sample meshes from 3DCP Scan, 3DCP Mocap., and AGORA to comprise the presented meshes in MTP datatset. In total, we present 1653 different meshes, from which 1498 (90%) are contact poses following Definition 3.1 in the main document. Of the 1653 meshes, 110 meshes are from 3DCP Scan, 1304 meshes are from 3DCP Mocap., and 159 are from AGORA. We collect at least one image for each mesh. From the 3731 collected images, 3421 (92%) images show a person mimicking a contact pose. Figure 17 shows how many image we collected per subset.

A.3 Discrete Self-Contact (DSC) Data.

Image selection. Discrete self-contact annotation may be ambiguous and we find some annotations that we do not consider to be functional self-contact. For example, in Fig. 18, some annotators label the left lower arm and left upper arm to be in contact, because of the slight skin touching at the elbow; we do not treat these as in self-contact. Therefore, we leverage the kinematic tree structure provided by SMPL-X and, in order to train TUCH, ignore the following annotations: left hand - left lower arm, left lower arm - left elbow, left lower arm - left upper arm, left elbow - left upper arm, left upper arm - torso, left foot - left lower leg, left lower leg - left knee, left lower leg - left upper leg, left knee - left upper leg, right hand - right lower arm, right lower arm - right elbow, right lower arm - right upper arm, right elbow - right upper arm, right upper arm - torso, right foot - right lower leg, right lower leg - right knee, right lower leg- right upper leg, right knee - right upper leg.

Appendix B TUCH

Here we provide details of the SMPLify-XMC and SMPLify-DC methods and how we apply them on MTP and DSC data respectively.

SMPLify-XMC is explained in Sec. 4.2 of the main paper. It is applied, before the training, to all MTP images to obtain gender-specific pseudo ground-truth SMPL-X fits. To use these fits for TUCH training, two pre-processing steps are necessary. First, they are converted to neutral SMPL fits. Second, we transform the converted SMPL fits to the camera coordinate frame estimated during SMPLify-XMC. This is necessary since SPIN assumes an identity camera rotation matrix. After that, the data is treated as ground truth during training, which means we apply the regressor loss directly on the converted SMPL pose and shape parameters without in-the-loop fitting.

On the contrary, SMPLify-DC is applied during TUCH training to images with discrete self-contact annotations. We run 10 iterations of SMPLify-DC for each image in a mini batch.

MTP and the DeepFashion subset of DSC do not have ground-truth 2D keypoints but we find OpenPose detections good enough in both cases. For the 2D re-projection loss, we use ground-truth keypoints (if available) and OpenPose detections weighted by the detection confidence. Each mini batch consists of 50% DSC and 50% MTP data.

Implementation details: We initialize our regression network with SPIN weights . We use the Adam optimizer and a learning rate of 1e−51e-5.

One disadvantage of training with fitting in the loop is that it is relatively slow. As an alternative, we also explore Exemplar Fine-Tuning (EFT) , which is a regression based method for fitting 3D meshes to a single image. The fitted SMPL meshes may then be used as pseudo annotations to train a regressor without in-the-loop optimization. With this approach, the authors train HMR-EFT, with which they achieve good results on 3DPW and MPI-INF-3DHP.

The idea of using discrete contact annotations is not limited to optimization based approaches. We show that they can also be applied in combination with EFT. Specifically, we extend the regressor loss of EFT with the contact terms from SMPLify-DC. We denote such an “EFT + contact loss” approach as EFT-C. Note that the original EFT loss uses a 2D orientation term to match the lower legs orientation, which we do not use here.

Each image in DSC is then paired with a pre-computed pseudo ground truth from EFT-C, and we denote the dataset as [DSC]EFT-C. Then, we finetune the HMR-EFT network on MTP, [DSC]EFT-C, as well as other training data from . This new model is called TUCH with EXemplar Finetuning, TUCHEX\text{TUCH}_{\mathit{EX}}. Unlike TUCH that still performs SMPLify-DC in the training loop, TUCHEX\text{TUCH}_{\mathit{EX}} is supervised only by pre-computed fits so it can be trained faster.

Implementation details. We initialize our network with state-of-the-art HMR-EFT weights. We train TUCHEX\text{TUCH}_{\mathit{EX}} on [COCO-All]EFT (CAE), H36M, MPI-INF-3DHP (MI), [DSC]EFT-C, and MTP. [COCO-All]EFT denotes the COCO dataset after EFT processing, as described in . In each batch we use a 10% CAE, 20% H36M, 10% MI, 20% 3DPW, 20% [DSC]EFT-C, and 20% MTP. The remaining details are the same as in the TUCH implementation. For the DSC dataset, we only consider images where the full body is visible. To identify these images, we test whether the OpenPose detection confidence of ankles, hips, shoulders, and knees is ≥\geq 0.2. We also ignore discrete contact annotations for connected body parts, as defined in A.3.

Appendix D Evaluation

3DCP Scan test images. During the scanning process when creating 3DCP Scan, we also take RGB photos of subjects being scanned, as shown in Figure 19. These images have high-fidelity ground-truth poses and shapes from the registration process described in Sec. A.1.1, making them a good test set for evaluation purposes. It is worth noting again that TUCH has never seen these images or subjects, but the contact poses were mimicked in creation of MTP, which is used in training TUCH.

TUCH. In Fig. 20 we visualize the improvement of TUCH over SPIN qualitatively. One can see that TUCH reconstructs bodies with better self-contact and less interpenetration (row 1 and row 2). Fig. 21, on the other hand, shows examples where SPIN is better than TUCH. Four of the images in Fig. 21 do not show the full body (rows 3, 4, 5, and 8). A possible reason why SPIN is better than TUCH in these cases is that MTP images always show the full body of a person, thus TUCH could be more sensitive to occlusion than SPIN.

We also evaluate the contribution of MTP data by finetuning SPIN only with it. The results are reported in Table 7, where TUCH (MTP+DSC) is the same as reported in Table 3 of the main paper. This experiment shows that MTP data alone is already sufficient to significantly improve state-of-the-art (SOTA) methods on 3DPW benchmarks. This suggests that the MTP approach is a useful new tool for gathering data to train neural networks.

For an additional comparison with SOTA EFT , we evaluate our TUCHEX\text{TUCH}_{\mathit{EX}} model on the same datasets (3DPW, MPI-INF-3DHP (MI), and 3DCP Scan) with error measures (MPJPE, PA-MPJPE, and MV2VE) like TUCH, see Tables 6, 8, and 9.

The MPJPE of TUCHEX\text{TUCH}_{\mathit{EX}} improves over HMR-EFT when evaluated on 3DPW. PA-MPJPE improves for contact poses and is overall on-par. Also the results on MPI-INF-3DHP improve. For the 3DCP Scan test set, PA-MPJPE improves. This shows that our data can not only be used with optimization based approaches, but also with exemplar fine-tuning, and that it allows us to improve the latest models in terms of estimating poses with contact.