Capturing and Inferring Dense Full-Body Human-Scene Contact

Chun-Hao P. Huang, Hongwei Yi, Markus Höschle, Matvey Safroshkin, Tsvetelina Alexiadis, Senya Polikovsky, Daniel Scharstein, Michael J. Black

Introduction

Understanding human actions and behaviors has long been studied in computer vision, with applications in robotics, healthcare, virtual try-on, AR/VR, and beyond. Remarkable progress has been made in both 2D human pose detection and 3D human pose and shape estimation (HPS) from a single image , thanks to realistic datasets annotated with 2D keypoints and 3D data . Despite this progress, something important is missing. Even the most basic human activities, such as walking, involve interaction with the surrounding environment. Fundamentally, human-scene interaction (HSI) involves the contact relationships between a 3D human and a 3D scene, i.e., human-scene contact (HSC). Existing HPS methods, however, largely ignore the scene and estimate human poses and/or shapes in isolation, often leading to physically implausible results.

Since reconstructing the full 3D scene from a single image is challenging, recent HPS methods tackle this problem by making several simplifying assumptions about the scene and/or body. Many methods consider only the contact between feet and ground , or assume the ground is a even plane , which is often violated, e.g., walking up stairs. To infer contact, many state-of-the-art (SOTA) methods use MoCap datasets to train a contact detector . Others exploit physics simulation or physics-inspired objectives but reduce the body representation to a small set of primitives. Surprisingly, none of these methods use image evidence when predicting human-scene contact. This is primarily due to the lack of datasets with images and 3D contact ground truth.

Many methods do estimate human object interaction (HOI) from images but constrain the reasoning to 2D image regions . That is, they estimate bounding boxes or heatmaps in the image corresponding to contact but do not relate these to the 3D body.

In this work, we address this problem with a framework that estimates 3D contact on the body directly from a single image. We make two main contributions. First, we create a new dataset that accurately captures human-scene contact by extending a markerless MoCap method to markerless HSC capture. Specifically, we capture multiview video sequences at 4K resolution in both indoor and outdoor environments. We also capture the precise 3D geometry of the scene using a laser scanner. Additionally, we capture high-resolution 3D scans of all subjects in minimal clothing and fit the SMPL-X body model to the scans. Our markerless HSC approach allows us to compute accurate per-vertex scene contact, as visualized in Fig. 1c.

Compared to the PROX dataset , which captures HSC with monocular RGB-D input, multiview data has two advantages: (1) it effectively resolves occlusions, leading to better reconstructed bodies and consequently more accurate scene contact; (2) it works for outdoor environments, as shown in Fig. 1.

The resulting dataset, called RICH (“Real scenes, Interaction, Contact and Humans”), provides: (1) high-resolution multiview images of single or multiple subjects interacting with a scanned 3D scene, (2) dense full-body scene-contact labels, (3) high-quality outdoor/indoor scene scans, (4) high-quality 3D human shapes and poses, and (5) dynamic backgrounds and moving cameras.

To estimate vertex-level HSC from a single color image, we develop BSTRO (Body-Scene contact TRansfOrmer), and train it with RICH. Our key insight in building BSTRO is that contact is not directly observable in images due to occlusion; thus, to infer contact, the network architecture must be able to explore the whole image for evidence. The transformer architecture enables BSTRO to learn non-local relationships and use scene information to “hallucinate” unobserved contact. We employ a multi-layer transformer , which has been successfully employed for natural-language processing and HPS estimation with occlusion .

In summary, our key contributions are: (1) We present RICH, a novel dataset that captures people interacting with complex scenes. It is the first dataset that provides both scans of outdoor scenes and images for monocular HSC estimation, unlike existing methods , which lack one or the other. (2) We propose BSTRO, a monocular HSC detector. It is body-centric so it does not require 3D scene reconstructions to infer contact. Unlike POSA , which is also body-centric, BSTRO directly estimates dense scene contact from the input image without reconstructing bodies. (3) We evaluate recent HSC methods and show that BSTRO gives SOTA results. (4) Since RICH has pseudo-ground-truth body fits, we also evaluate SOTA HPS methods and analyze their performance with respect to scene-contact, which is not supported by existing HPS datasets . We confirm that the performance of a SOTA HPS method degrades in the presence of scene contact.

Related Work

We review existing methods that consider contact between humans and scenes. Since many of them employ a 3D body reconstruction method as a backbone in the pipeline, we first briefly discuss recent HPS trends and then focus on how the prior art incorporates scene contact.

Monocular HPS methods reconstruct 3D human bodies from a single color image. Many methods output the parameters of statistical 3D body models . SMPLify fits the SMPL model to the output of a 2D keypoint detector and we build on it here.

In contrast, deep neural networks regress body-model parameters directly from pixels . To deal with the lack of in-the-wild 3D ground truth, some methods use 2D keypoints or linguistic attributes as weak supervision, while some directly fine-tune the network w.r.t. an input image at test time . Kolotouros et al. combine HMR and SMPLify in a training loop for better 3D supervision. On the other hand, non-parametric or model-free approaches directly estimate 3D vertex locations without body parameters . We refer readers to for a comprehensive review. None of the above methods estimate HSC.

Markerless MoCap exploits synchronized videos from multiple calibrated cameras and has a long history with commercial solutions, but these focus on estimating a 3D skeleton. To model HSC, we need to extract a full 3D body shape and, therefore, focus on such methods here. Early methods, either bottom-up or top-down , are fragile, need subject-specific templates and manual input, and do not generalize well to in-the-wild images.

Powered by CNNs, recent methods leverage multiview consistency to improve keypoint detection , to re-identify subjects across views or across view and time , but they estimate only joints, not body meshes. Dong et al. reconstruct SMPL bodies for multiple subjects and Zhang et al. additionally estimate hands and facial expressions. They demonstrate results for lab scenarios, while our HSC capture method in Sec. 3.1 works in less constrained outdoor scenes.

All methods above reconstruct human bodies in isolation without taking into account the interaction with scenes. Consequently, the results often contain physically implausible artifacts such as foot skating and ground penetration.

2 Human Scene Interaction (HSI)

2D Human-Object Interaction (HOI) methods localize 2D image regions with HOI and recognize the semantic interactions in them. Most methods represent humans and objects very roughly as bounding boxes ; only a few use body meshes for humans and spheres for objects .

3D Contact. Knowing which part of the body and scene are in contact provides compact yet rich information that enables many applications, such as HSI recognition or placing virtual humans into a scene . The upper part of Table 1 summarizes how body-scene contact gets incorporated in methods of different goals and tasks.

Early work uses scene contact as part of the HSI feature but represents a human body roughly as a stick figure. Recent HPS methods use contact to improve the estimated body poses. Ideally, when both the body and scene are “perfectly reconstructed,” applying a threshold to the 3D Euclidean distances between them is sufficient to infer accurate contact. Prior work takes this thresholding approach to annotate contact . At test time, PROX assumes scene scans to be known a priori; PHOSA estimates 3D objects, 3D people, and the contacts between them but only for a limited class of objects. Since reconstructing a 3D scene in high quality with correct layout and spatial arrangement is still an open challenge , monocular HSC detection methods resort to other heuristics. The most common one is a zero-velocity assumption; i.e., surfaces in contact should not slide relative to each other. This assumption is widely employed to reduce foot-skating artifacts . Some of these detect contact with a separate neural network at test time, taking the 2D/3D joints in a temporal window as input , while others integrate it in a body motion prior . These approaches use MoCap datasets such as AMASS and Mixamo to build training data, where contact is automatically labelled via thresholding the distance to the ground and/or the velocity.

POSA observes that scene contact is correlated with body poses and introduces a generative model to sample contact given a posed mesh. Some methods apply physics to encourage foot-ground contact and ensure physically plausible motions. However, they have to approximate the body as a set of boxes, cylinders or spheres. MOVER uses human scene contact to improve monocular estimation of 3D scene layout.

All these approaches first reconstruct bodies (2D or 3D), and then reason about contact, effectively ignoring valuable image information. To go further, we need a dataset consisting of natural images and 3D body-scene contact labels. As summarized in the lower part of Table 1, many existing contact-related datasets consider self contact or person-person contact , but not HSC. The most relevant datasets for HSC are and PROX . The former provides egocentric images for localization, which are not suitable for HSC detection from images. PROX can be used for our task but it consists of only indoor scenes and is of lower quality. The ground-truth bodies in PROX are computed by fitting to RGBD data, which is sensitive to occlusions. This not only limits the type of HSI in the dataset (mostly walking, sitting, lying) but also influences the quality of body fits.

Methods: RICH Dataset

Given videos captured by CC synchronized cameras, we first identify each subject across views and across time with . For each identified subject, we reconstruct a SMPL-X body by a multiview fitting method that is robust to noisy 2D keypoint detections, and we place it in a pre-scanned scene to compute body-scene contact (Sec. 3.1). With this approach, we build a monocular body-scene interaction dataset (RICH) comprising 577K images paired with SMPL-X parameters and scene contact labels (Sec. 5).

We first track subjects temporally in each video with AlphaPose , followed by MvPose to match the tracklets across views. Other methods that build such 4D associations could also be applied here.

At time tt, we now have at most CC bounding boxes of the same person and we aim to reconstruct the body. To this end, we adapt SMPLify-X to accommodate multiview data. SMPLify-X optimizes the pose θ\theta, shape β\beta and facial expression ψ\psi of SMPL-X to match the observed 2D keypoints by minimizing the following objective:

Multiview per-person reconstruction. For each person, we compute 2D keypoints in each camera cc. Instead of fitting them using SMPLify-X in each view, we combine all 2D landmarks in a multiview energy term: ∑cEJc\sum_{c}E_{J}^{c}. Unlike in , where one needs to estimate camera translation first, the perspective projection here is well defined by the pre-calibrated intrinsics and extrinsics. To pursue high-quality fits, body shape β\beta is estimated in advance by registering a SMPL-X template to minimally-clothed 3D scans following . β\beta is hence no longer a free variable in Eq. 1 and we set λβ=0\lambda_{\beta}=0. In addition to EJE_{J}, which measures joint errors, we also use EOE_{O} that measures errors in “bone orientations.” Figure 2(a) illustrates the intuition behind this term. Since posing human bodies requires traversing a kinematic chain, with the joint term EJE_{J}, the error of parent joints ϵ1\epsilon_{1} is accumulated in the error of child joints ϵ2\epsilon_{2}. When ∥ϵ2∥\|\epsilon_{2}\| gets too large, the influence is downweighted because our robust loss treats it as an outlier. Instead, EOE_{O} factors out the errors of ancestors and focuses on the error of the joint per se. Our final objective is Emv(θ,ψ)=∑cEJc+∑cEOc+EregE_{\text{mv}}(\theta,\psi)=\sum_{c}E_{J}^{c}+\sum_{c}E_{O}^{c}+E_{\text{reg}}.

Due to noisy 2D detections, keypoints in each view often disagree with each other. One may count on the robustifier to identify outliers and reduce their contribution. This depends, however, on the current estimated body in the optimization, so it assumes good initialization. Instead, we check the multiview consistency of landmarks as illustrated in Fig. 2(b). For each joint, we take the detections in two views (blue), triangulate a 3D point and project it to the third view (green). If the distance between the projected point (red) and the detection (green) in the third view is large, that means the three detections do not agree with each other and at least one of them is wrong. Instead of making hard decision separating outliers from inliers, we exhaustively compute all triplets of views, accumulate the reprojection error and downweight the contribution in ∑cEJc\sum_{c}E_{J}^{c} for views with high errors. We term this multiview consensus, as it behaves like a soft majority voting mechanism. As long as there are more correct detections than wrong ones, it can reduce the influence of noisy landmarks, independent of the current body estimate.

To further avoid local minima, we apply a state-of-the-art in-the-wild body regressor (PARE ) to initialize θ\theta. We run PARE on the bounding box from each view, fuse the results by averaging the poses, and covert the fused body from SMPL to SMPL-X. The SMPL-X body pose gives the initial value of θ\theta for minimizing EmvE_{\text{mv}}. We first solve EmvE_{\text{mv}} in a frame-wise manner and then refine a batch of TT frames jointly with two additional terms on body and hand motions: Ebatch(θ1,⋯ ,θT)=∑t=1TEmvt+λsmbEsmb+λsmhEsmhE_{\text{batch}}(\theta_{1},\cdots,\theta_{T})=\sum_{t=1}^{T}E_{\text{mv}}^{t}+\lambda_{\text{sm}}^{b}E_{\text{sm}}^{b}+\lambda_{\text{sm}}^{h}E_{\text{sm}}^{h}. EsmbE_{\text{sm}}^{b} is the smoothness term in and EsmhE_{\text{sm}}^{h} encourages neighboring frames to have similar hand-pose PCA vectors ZhZ_{h}.

We place the reconstructed bodies into pre-scanned 3D scenes to estimate the body-scene contact. The scene mesh and HDR textures were acquired using an industrial laser scanner, Leica RTC360. To put the bodies in the scene, we solve the rigid transformation between camera coordinates and scan coordinates with manually identified correspondences. To annotate human-scene contact automatically, our approach is similar to POSA . Specifically, for each vertex on the body mesh, we compute the point-to-surface distance to the scene scan. If the distance is lower than a threshold and the normal is compatible, we accept the hypothesis that it is in contact. Considering the thickness of shoe soles, the threshold is 5cm for the vertices at the bottom of feet and 2.5cm for the rest of body. This is different from POSA, which uses 5cm for the whole body to collect training data from PROX . Furthermore, the pseudo-ground-truth body poses in PROX are obtained by fitting the SMPL-X template to monocular RGBD data. As shown in the bottom row of Fig. 5, PROX accuracy suffers from occlusion, sometimes resulting in severe penetration with the scene. The errors in body fits are carried over to the ground-truth HSC data for POSA. In contrast, in RICH, bodies are recovered from multiview data, which reduces the issues caused by occlusion and depth ambiguity.

Methods: BSTRO

Here we introduce BSTRO for dense HSC estimation from a single image. This relies on RICH, described in Sec. 5 in detail. Existing HSC methods usually take a multi-stage approach. Given an input image, they first reconstruct the body mesh and use it as a proxy to infer contact. Formally, let ff denote the function recovering a body mesh MM from the input image II, M=f(I)M=f(I). ff can be an energy-minimization process such as or a neural network as in . To estimate contact, SOTA methods differ from each other in two ways: (1) the features extracted from MM, e.g., Euclidean distance to the 3D scene, velocity and body poses (cf. Table 1); (2) the prediction functions, e.g., simple thresholding, neural network, or physics engine. With a slight abuse of notation, we denote these feature extraction and contact estimation processes collectively as gg, which takes the body MM as input and predicts a contact vector c=g(M)\mathbf{c}=g\left(M\right). Each element in c\mathbf{c} is 11 if the corresponding part of the body (vertex, joint or body part) is in contact with the scene, and otherwise. For example, gg represents the decoder of a conditional VAE in POSA , taking the vertex locations of MM as input, while in , gg is a MLP operating on the motion of MM.

With this formulation, the body-scene contact c\mathbf{c}, whether defined on a dense mesh or on a set of sparse joints/parts, is a composite function of gg and ff: c=g∘f(I)\mathbf{c}=g\circ f(I), where gg is agnostic to the input image. In contrast, our goal is to detect dense body-scene contact directly from the input II: c=g(I)\mathbf{c}=g(I). To our knowledge, this was explored only for self-contact and person-person contact and only at a coarse region level, not the vertex level.

We use SMPL as the body representation for BSTRO, hence c∈{0,1}V\mathbf{c}\in\{0,1\}^{V}, where VV=6890 is the number of vertices on a SMPL mesh, as opposed to VV=10475 on a SMPL-X mesh. The reason for this choice is that a SMPL-X mesh has nearly 50% of the vertices on the head, which rarely participates in natural body-scene contact, so we would like to reduce the dimensionality of the output space. See Sup. Mat. for more discussion of this design choice.

We model gg as a neural network and train it end-to-end in a supervised way with the (I,c)(I,\mathbf{c}) pairs sampled from RICH. The network architecture is designed based on our key observation. That is, regions in contact are not directly observable due to occlusion. However, there is rich information in the image to tell which parts of the body are in contact with the scene. Estimating HSC from images is therefore inherently a “hallucination” task. Without really “seeing” the regions in contact, the network needs to explore the image freely and attend to regions it finds informative.

Training. We apply the binary cross entropy loss between the ground truth contact and the predicted contact probability pvp_{v}. One can think of this as a multi-label classification problem, where each category (vertex) has its own probability of being true (in contact) or not.

To gain robustness to occlusion, we employ Masked Vertex Modeling (MVM) . Specifically, at each iteration, we randomly mask out some queries in QQ and still ask the transformer to estimate contact for all vertices. In order to predict the output of a missing query, the model has to explore other relevant queries. This simulates occlusions where bodies are only partially visible and also encourages the network to hallucinate contact.

RICH Dataset

We capture 22 subjects performing various human-scene interactions in 5 static 3D scenes with 6-8 static cameras and, in some scenes, with an additional (untracked) moving camera (Fig. 4 rightmost scene). Subjects gave prior written informed consent for the capture, use, and distribution of their data for research purposes. The experimental methodology has been reviewed by the University of Tübingen Ethics Committee with no objections.

RICH has in total 142 single or multi-person multiview videos, with a total of 90K posed 3D body meshes, together with 90K dense full-body contact labels in both SMPL-X and SMPL mesh topology, and 577K high resolution (4K) images. Compared to PROX, RICH consists of mostly outdoor environments with areas of roughly 60m2. The images in RICH are real, not limited to a single subject, have dynamic backgrounds and varied viewpoints. All these features make it suitable for training and evaluating monocular HSC methods. Figure 4 shows several examples of RICH.

In addition, since RICH provides SMPL-X fits, i.e., pseudo-ground-truth human poses and shapes, it can also serve as a monocular or multiview HPS benchmark. It contains more subjects than 3DPW , more accurate body shapes than AGORA , and real human-scene interaction unlike Human3.6M . In our experiments we analyze the performance of SOTA HPS methods with respect to body-scene contact. Such analyses are not feasible with existing HPS datasets.

Experiments

We split 142 multiview videos in RICH into 62, 28, 52 for training, validation, and testing purposes, respectively. The test set consists of several subsets designed for varied evaluation protocols. Each subset is defined by whether or not each of three attributes has been observed in training: scene, human-scene interaction, and subject. The most challenging subset is when they are all unseen in RICH-train. The split ensures there is one completely withheld scene and 7 unseen subjects in the test set. See Sup. Mat. for more breakdowns in terms of 3D bodies and images.

2 Evaluation Metrics and Baselines

We apply standard detection metrics (precision, recall, and F1 score) to evaluate the estimated dense HSC. Since vertex density varies over the SMPL template, the same number of false positives, say, on the palm and on the thigh correspond to different areas on the body surface, but this is not reflected in the scores above. To better understand how well an HSC method estimates contact, we additionally consider a measure that translates the count-based scores to errors in metric space. Specifically, for each vertex predicted in contact, we compute its shortest geodesic distance to a ground-truth vertex in contact. If it is a true positive, this distance is zero; if not, this distances indicates the amount of prediction error along the body.

We evaluate three HSC baselines on the RICH-test. Zou et al. use the velocity of 4 2D keypoints on the feet to predict contact; HuMoR estimates contact for 8 joints while reconstructing human motions. These two methods estimate contact for sparse joints, not dense vertices, so we mark all vertices that correspond to a joint as contact when the method predicts the joint is in contact. POSA requires a 3D body mesh in the canonical space as input to sample dense body contact. We consider two choices of 3D bodies for POSA: (1) using the results from a SOTA body regressor PIXIE , or (2) using ground-truth bodies to evaluate the impact of errors in estimated body pose.

3 Main Results

The results on RICH-test are reported in Table 2. We see that HuMoR yields overall lowest detection scores and highest geodesic errors. This is partially due to the fact that it only considers contact with an even ground plane, while RICH-test contains more varied real scene interactions.

POSA, in general, has higher recall compared to other methods. This, however, comes with a cost of precision, meaning that there are many false positives. Comparing rows (c) and (d) we see that recall is significantly better when using ground-truth bodies. BSTRO yields significantly better precision but with lower recall than POSA. Still, it has the highest F1 score and lowest geodesic error, which shows that it strikes a good balance between precision and recall. Figure 6 shows some visual examples. RICH has accurately fitted SMPL-X bodies and body-scene contact. Given an input image, BSTRO estimates scene contact that is closer to the ground truth, whereas POSAPIXIE yields false positives frequently (red circles) and sometimes misses the contact on the hands. While the training dataset is limited, BSTRO also works on in-the-wild images, as shown in the right part of Fig. 6.

4 Generalization

To analyze how well BSTRO generalizes, we split RICH-test into several subsets. Each subset represents whether BSTRO has observed similar images of the three attributes: scene, human-scene interaction (HSI), and subject. This allows us to inspect the importance of each attribute, and to know which aspect future methods should focus on. Note that this is a unique feature of RICH, as existing HSC datasets from MoCap and HPS datasets do not support such an analysis.

In Table 3, ✓ means BSTRO has seen similar images of that attribute during training, while ✗ means it has not. For example, images in row (a) share the same subjects and similar HSI with training data but the scenes are new. Intuitively, this is an easy subset and indeed the scores are high in this scenario. Once HSI is withheld, the performance drops (row (e)). This drop is more pronounced than the drop caused by withholding a subject (row (c)). Comparing each of the rows (c,d,e) to row (f), we observe that seeing similar HSI at training helps the most. Seeing the same scenes or same subjects does not guarantee gains in performance. Finally, row (f) represents the most challenging subset, where scene, HSI, and subjects are all unseen during training. We see that BSTRO still yields results that are comparable to other subsets. Subset (c) contains many images with person-person occlusion, e.g., Fig. 6 bottom left, which partially explains why it is the most challenging.

5 HPS Evaluation on RICH-test

Besides evaluating human-scene contact, RICH can also serve as a benchmark for monocular HPS methods. Unlike existing HPS benchmarks with real images such as 3DPW or Human3.6M , the real scene contact in RICH enables a new way of analyzing the performance of an HPS method. In particular, we use PIXIE , a recent monocular HPS method, to regress SMPL-X bodies from RICH-test. We compare the estimated SMPL-X bodies with the pseudo-ground-truth SMPL-X fits from Sec. 3.1, and compare the error when body-scene contact is present or absent.

We consider Mean Per-Joint Position Error (MPJPE) and Vertex-to-Vertex Error (V2V) to measure the discrepancies in joints and body meshes respectively. For freely moving cameras, we apply Procrustes alignment (PA) before calculating the two errors, hence PA-MPJPE and PA-V2V. Procrustes alignment factors out differences in rotation, scale and translation, focusing on measuring the difference in “pure body poses.” PA hides many sources of errors so we use it only when ground-truth camera extrinsic parameters are not available. For calibrated cameras, on the other hand, we factor out only translation by aligning the estimated and ground-truth bodies to their pelvis locations, denoted with a prefix “TR.” We ignore foot-ground contact, which is ubiquitous, and compare the results when there is meaningful scene contact vs. no scene contact.

On average, images containing meaningful scene contact yield 214.0mm/172.81mm TR-MPJPE/TR-V2V, higher than 161.81mm/121.71mm for images with no contact other than foot-ground contact. This is partially due to the fact that scene contact usually comes with scene occlusion, and this shows a direction where monocular HPS methods can improve. The corresponding errors in moving cameras are 84.15mm/83.16mm PA-MPJPE/PA-V2V for images with meaningful contact and 63.67mm/64.37mm for those without. We again observe that the presence of scene contact makes HPS more challenging, yielding higher errors. This shows that scene contact impacts all aspects of the problem: from pure body poses to global orientation and translation.

Conclusion

While there is rapid progress on estimating 3D human pose and shape from images, much of this work ignores the scene and the interaction of the body with that scene. Capture and analysis of body-scene contact, however, is critical to understanding human action in detail. To address this, and to help the research community study this problem, we created RICH, a new dataset with challenging natural video sequences, high-resolution 3D scene scans, ground-truth body shapes, high-quality reference poses, and detailed 3D contact labels. We use the contact information to train a new method (BSTRO) that takes a single image of a person interacting with a scene and infers the 3D contacts on their body. We also use the dataset to evaluate human pose estimation and find that scenes with significant contact cause problems for the state of the art. The dataset and code are available for research purposes.

Limitations and future work. RICH considers only contact with static scenes so does not account for the body contact with dynamic scenes, e.g., with hand-held objects, or human-human interaction. One extension would estimate the rigid-body pose of an object given its 3D model and simultaneously reconstruct the hand/body that interacts with it. Another interesting direction would jointly estimate the body pose, shape, and scene contact in one single network.

Acknowledgments. We thank Taylor McConnell, Claudia Gallatz, Mustafa Ekinci, Camilo Mendoza, Galina Henz, Tobias Bauch, and Mason Landry for the data collection and cleaning. We thank all participants who contributed to the dataset, Paola Forte for PIXIE experiment and Benjamin Pellkofer for IT support. Daniel Scharstein was supported by NSF grant IIS-1718376.

Disclosure: https://files.is.tue.mpg.de/black/CoI_CVPR_2022.txt

References

Appendix A SMPL-X vs. SMPL HSC labels

We build RICH by fitting a SMPL-X template to multi-view data and compute the human-scene contact (HSC) as explained in the Sec. 3 and Sec. 5 of the main paper (Fig. 7). The contact labels are defined in SMPL-X format and we map them to SMPL format for training BSTRO. This is feasible since there is an 1-to-1 correspondence between SMPL-X and SMPL vertices below the neck, as shown in Fig. 8.

With this mapping, we convert the ground-truth HSC labels from SMPL-X to SMPL without losing information. As a result, we benefit from realistic hand articulation in SMPL-X and still keep the dimension of the output space small (SMPL). Such a mapping also makes RICH a suitable HSC benchmark for both body models. Since the two models share the set of vertices of interest, choosing either of them does not influence the detection scores or errors.

On the other hand, the human pose and shape (HPS) parameters of the two models differ. Converting HPS parameters between SMPL-X and SMPL requires extra processing and one always loses the hand articulation when converting from SMPL-X to SMPL. Therefore, RICH provides only SMPL-X as pseudo ground truth. To evaluate methods that regress SMPL parameters using RICH, users should convert SMPL to SMPL-X, which does not result in a loss of information.

Appendix B RICH Dataset

The 142 multi-view videos in RICH are recorded at a rate of 30 frames per second. We separate them into subsets of 62, 28, 52 for training, validation, and testing purposes, respectively. This amounts to 303K, 149K, 125K images of 4K resolutions (in total 577K), and 40K, 18K, 32K 3D SMPL-X bodies along with dense scene-contact labels (in total 577K) in each subset. By “body” here, we mean any SMPL-X mesh. Note that the number of unique “people” in the dataset is much smaller than the number of bodies because every posed mesh constitutes a separate body.

Compared to the recent HPS dataset AGORA , RICH has more 3D bodies (90K vs. 4K), more images (577K vs. 19K) and more accurate body shapes (registrations to minimally-closed scans vs. clothed scans ). It has more subjects in varied body shapes than 3DPW (22 vs. 18) and subjects are in natural clothing as opposed to those in Human3.6M . Last but not least, RICH provides high-quality scene scans and scene contact labels that none of the above datasets provides.

Following the illustration in Fig. 2(a) of the main paper, the bone-orientation term EOE_{O} factors out the residual of the parent joint ϵ1\epsilon_{1} from the residual of the child joint ϵ2\epsilon_{2}:

where b2′=j2′−j1′b^{\prime}_{2}=j^{\prime}_{2}-j^{\prime}_{1} and b2=j2−j1b_{2}=j_{2}-j_{1} denote the “bone vector” of target points (detected landmarks) and estimated SMPL-X joints respectively. It follows that

Since b2′b^{\prime}_{2} involves only the detected landmarks and ∥b2∥\|b_{2}\| is fixed given a constant body shape β\beta, the first two terms are constant when optimizing the multi-view objective EmvE_{\text{mv}}. ∥r2∥22\|r_{2}\|_{2}^{2} is therefore minimized when b2⊤b2′b_{2}^{\top}b^{\prime}_{2} is maximized, i.e., when b2b_{2} has the same orientation as b2′b^{\prime}_{2}.

Appendix D BSTRO Implementation Details

We sample RICH-train to build the image-HSC pairs (I,c)(I,\mathbf{c}) for training BSTRO. For each sequence, we consider only every other frame, and for each frame, we use the dynamic view and one randomly selected static view, or two static views if no moving camera is available. This sampling strategy ensures sufficient variations in viewpoints and background, while keeping the total number of the training pairs tractable.

We train with in total 24K (I,c)(I,\mathbf{c}) pairs from RICH-train and use Adam optimizer with an initial learning rate of 1e-4 for 100 epochs. The HR-Net backbone is initialized with the weights pre-trained on ImageNet , Human3.6M or 3DPW . The best checkpoint is selected by the best performance on RICH-validation with RICH-test completely withheld. We refer interested readers to for the architecture of the multi-layer transformer.