EgoBody: Human Body Shape and Motion of Interacting People from Head-Mounted Devices

Siwei Zhang, Qianli Ma, Yan Zhang, Zhiyin Qian, Taein Kwon, Marc Pollefeys, Federica Bogo, Siyu Tang

Introduction

Humans constantly interact and communicate with each other; understanding our social interaction partners’ motions, intentions and emotions is almost instinctive for us. However, the same does not hold for machines. A first step towards automated human interaction understanding is the estimation of the 3D body pose, shape and motion of the social interaction partner (“interactee”) from egocentric views, e.g. from head-mounted devices (HMD). Addressing this challenging problem is crucial for many applications, ranging from assistive robotics to Augmented and Virtual Reality (AR/VR), where sensors typically perceive the interactee from the egocentric view. Despite its importance, the problem has received little attention in the literature so far. While there are a large number of methods for full-body pose (and sometimes also shape) estimation from RGB(D) frames , they tend to perform poorly on data captured with an HMD (see Sec. 5). Indeed, this setup brings its own unique challenges, which most methods have not explicitly addressed so far. Any method aiming at understanding the pose and shape of the interactee must deal with severe body truncations, motion blur (exacerbated by the embodied movement of the HMD), people entering/exiting the field of view, to name a few.

A reason for such limited attention is the lack of data. On one hand, most human motion datasets are captured by third-person-view cameras without egocentric frames , which do not faithfully replicate AR/VR scenarios; most capture only one subject at a time, without interactions . On the other hand, existing egocentric datasets are limited in terms of annotation modalities, scale and interaction diversity. They either focus on coarse-level interaction/action labels , or provide only the camera wearer’s pose without considering , or with very limited data involving , the interactee. You2Me collects egocentric RGB frames of two-people interactions, annotated with 3D skeletons, without 3D scene context, nor the body shape. Recently, Ego4D collects a large amount of egocentric videos for various tasks including action and social interaction understanding, but without 3D ground truth for human pose, shape and motions.

To fill this gap, we propose EgoBody, a unique, large-scale egocentric dataset capturing high-quality 3D human motions during social interactions. We focus on 2-people interaction cases, and define interaction scenarios based on the social interaction categories studied in sociology . Unlike most existing datasets that only provide RGB streams, EgoBody collects egocentric multi-modal data, with accurate 3D human shape, pose and motion ground-truth for both interacting subjects, accompanied by eye gaze tracking for the camera wearer. Furthermore, EgoBody includes accurate 3D scene reconstructions, providing a holistic and consistent 3D understanding of the physical world around the camera wearer.

The egocentric data is captured with a Microsoft HoloLens2 headset , which provides rich multi-modal streams: RGB, depth, head, hand and eye gaze tracking, correlated in space and time. In particular, eye gaze carries vital information about human attention during interactions. By providing eye gaze tracking synchronized with other modalities, EgoBody opens the door to study relationships between human attention, interactions and motions. We obtain high-quality 3D human shape and motion annotations in an automated way, by leveraging a marker-less motion capture approach. Namely, we utilize a multi-camera rig consisting of multiple Azure Kinects as our motion capture system.

However, combining raw data streams from the egocentric- and the third-person-view remains highly challenging due to hardware limitations. Specifically, the Kinect-HoloLens2 calibration exhibit inaccuracies due to not perfectly accurate factory calibration and tracking drift. We address this by proposing a refinement scheme based on body keypoints. With carefully calibrated data, we further build an efficient motion capture pipeline based on to fit the SMPL-X body model to multi-view and egocentric RGB-D data, reconstructing accurate 3D full-body meshes for both the camera wearer and the interactee. In this way, we get accurate and well calibrated ground truth across all sensor coordinates, as well as the world coordinate, which is not available in most existing datasets. The setup is lightweight and easy to deploy in various environments.

With EgoBody we propose the first benchmark for 3D human pose and shape estimation (3DHPS) of the interactee, in interactions captured by the HMD. By evaluating state-of-the-art 3DHPS methods on the EgoBody’s test set, we carefully analyze and highlight the limitations of existing methods in this egocentric setup. We show the usefulness of EgoBody by fine-tuning three recent methods on its training set, obtaining significantly improved performance on our test set. Finally, in a cross-dataset evaluation we show how models fine-tuned on EgoBody also achieve a better performance on the You2Me dataset .

Contributions. In summary, we: (1) provide the first large-scale egocentric dataset, EgoBody, comprising both egocentric- and third-person-view multi-modal data, annotated with high-quality 3D ground-truth motions for both interacting people and 3D scene reconstructions; (2) extensively evaluate state-of-the-art 3DHPS methods on our test set, showing their shortcomings in this egocentric setup and providing insights for future methods in this direction; (3) show the usefulness of our training set: a simple fine-tuning on it significantly improves existing methods’ performance and robustness on both our test set and a different egocentric dataset; (4) provide the first benchmark for 3DHPS estimation of the interactee in the egocentric view during social interactions.

Related Work

Datasets for 3D human pose, motion and interactions. A large number of datasets focus on 3D human pose and motion from third-person-views . For example, Human3.6M and AMASS use optical marker-based motion capture to collect large amounts of high-quality 3D motion sequences; they are limited to constrained studio setups and images – when available – are polluted by markers. PROX performs marker-less capture of people moving in 3D scenes from monocular RGB-D, without human-human interactions. The quality of the reconstructed motion is further improved by LEMO . The Panoptic Studio datasets capture interactions between people using a multi-view camera system, annotated with body and hand 3D joints plus facial landmarks. CHI3D focuses on close human-human contacts, using a motion capture system to extract ground-truth 3D skeletons. 3DPW reconstructs the 3D shape and motion of people by fitting SMPL to IMU data and RGB images captured with a hand-held camera, without 3D environment reconstruction. None of these datasets provides egocentric data.

Among datasets for egocentric vision, a lot of attention has been put on hand-object interactions and action recognition, often without 3D ground-truth . Mo2Cap2 and xR-EgoPose provide image-3D skeleton pairs for egocentric body pose prediction of the camera wearer, without the interactee involved. HPS reconstructs the body pose and shape of the camera wearer moving in large 3D scenes; only a few frames include interactions with an interactee. You2Me provides 3D skeletons for both interacting people paired with images captured with a chest-mounted camera plus external cameras; there are no body shape or 3D scene annotations. EgoMoCap analyzes the interactee body shape and pose in outdoor social scenarios capturing only the egocentric RGB stream.

Table 1 compares EgoBody with the most related human motion datasets. EgoBody is the first motion capture dataset that collects calibrated egocentric- and third-person-view images, with various interaction scenarios, multi-modal data and rich 3D ground-truth. Additionally, EgoBody provides the camera wearer’s eye gaze to facilitate potential social interaction studies which jointly analyze human attention and motion.

3D human pose estimation. The problem of estimating 3D human pose from third-person-view RGB(D) images has been extensively studied in the literature – either from single frames , monocular videos or multi-view camera sequences . SPIN estimates SMPL parameters from single RGB images by combining deep learning with optimization frameworks. METRO reconstructs human meshes without relying on parametric body models. Most methods require “full-body” images and therefore lack robustness when parts of the body are occluded or truncated, as it is the case with the interactee in egocentric videos. EFT injects crop augmentations at training time to better reconstruct highly truncated people. PARE explicitly learns to predict body-part-guided attention masks. However, these methods exhibit a significant performance drop when applied to egocentric data. Our dataset helps fill this performance gap, as we show in Sec. 5.

The problem of egocentric pose estimation is receiving growing attention. Most methods estimate the camera wearer’s 3D skeleton, based on images, IMU data, scene cues or body-object interactions . You2Me estimates the camera wearer’s pose given the interactee’s pose as an additional cue. Liu et al. estimate 3D human pose and shape of the interactee given egocentric videos in outdoor scenes, with limited interaction diversity.

Egocentric social interaction learning. Egocentric videos provide a unique way to study social interactions. Most methods focus on social interaction recognition . Lee et al. produce a storyboard summary of the camera wearer’s day given egocentric videos. Northcutt et al. collect an egocentric communication dataset focusing on conversations. Recently Ego4D dataset collects massive egocentric videos for various tasks including hand-object and social interaction understanding, making significant advances in stimulating future research in the egocentric domain. EgoBody is unique in that we are the first egocentric dataset that provides rich 3D annotations including accurate 3D human pose and shape for all interacting subjects.

Building the EgoBody Dataset

EgoBody collects sequences capturing subjects performing diverse social interactions in various indoor scenes. For each sequence, two subjects are involved in one or more interaction scenarios (Sec. 3.1). Their performance is captured from both egocentric- and third-person-views. One subject (the camera wearer) wears a HoloLens2 headset , capturing multi-modal egocentric data (RGB, depth, head, hand and eye gaze tracking streams). Their interaction partner, i.e. interactee, does not wear any device. The camera wearer’s HoloLens2 is calibrated and synchronized with three to five Azure Kinect cameras which capture the interaction from different viewpoints (Sec. 3.2). Based on this multi-view data, we acquire rich ground-truth annotations for all frames, including 3D full-body pose and shape for both interacting subjects and the reconstructed 3D scene (Sec. 3.3). Statistics for EgoBody are reported in Sec. 4.

To guide the subjects and obtain rich, diverse body motions, we define multiple interaction scenarios within five major interaction categories in sociology studies : cooperation, social exchange, conflict, conformity and others, spanning diverse action types (Tab. 2) and body poses (Fig. 5). For each sequence, we pre-define one or more interaction scenarios and ask the two participants to interact accordingly. We allow the subjects to improvise within each interaction scenario to ensure intra-class variation. The motion diversity is further increased with various human-scene interactions by capturing in 3D scenes.

2 Data Acquisition Setup

As mentioned above, EgoBody collects egocentric- and third-person-view multi-modal data, plus 3D scene reconstructions. Fig. 2 illustrates our system setup.

Egocentric-view capture. We use a Microsoft Hololens2 headset to record egocentric data. Using the Research Mode API , we capture RGB videos (1920×\times1080) at 30 FPS, long-throw depth frames (512×\times512) at 1-5 FPS, as well as eye gaze, hand and head tracking at 60 FPS. Note that we do not record depth at a higher framerate (AHAT) due to the “depth aliasing” described in . We observe that captures exhibit typical challenges for limited power-devices, like frame drops and blurry images.

Third-person multi-view capture. We use three to five Azure Kinect cameras (denoted by Cam1∼\simCam5) to capture multi-view, synchronized RGB-D videos of interacting subjects. Having multi-view data helps our motion reconstruction pipeline for ground-truth acquisition (Sec. 3.3). The cameras are fixed during recording. They capture synchronized RGB frames (1920×\times1080) and depth frames (640×\times576) at 30 FPS.

3D scene representation. We pre-scan the environment using an iPhone12 Pro Max running the 3D Scanner app . Scene reconstructions are stored as 3D triangulated meshes, each with 105∼10610^{5}\sim 10^{6} vertices. We choose this procedure for its efficiency and reconstruction quality.

Calibration and Synchronization. For each Kinect, we extract its camera parameters via the Azure Kinect DK . For the HoloLens2, we get its camera parameters as exposed by Research Mode . We synchronize the Kinects via hardware, using audio cables. Since it is not possible to synchronize HoloLens2 and Kinect in a similar way, we use a flashlight visible to all devices as signal for the first frame. Kinect-Kinect and Kinect-HoloLens2 cameras are spatially calibrated using a checkerboard and refined by rigid alignment steps (ICP ).

The Kinect-HoloLens2 calibration is further optimized based on body keypoints (Sec. 3.3). We use Cam1 to define our world coordinate frame origin. Once we calibrate the HoloLens2 coordinate frame with Cam1’s world origin, we can track the headset position, and therefore its cameras, by relying on its built-in head tracker . We also register the 3D scene into the coordinate frame of Cam1, and reconstruct the human body in this space (see details in Supp. Mat.).

3 Ground-truth Acquisition

Data preprocessing. We use OpenPose to detect 2D body joints in all (Kinect and HoloLens2) RGB frames. OpenPose identifies people in the same image by assigning a body index to each detected person. In general, this works well, but gives false positives which we process afterwards. To extract human body point clouds from Kinect depth frames, we use Mask-RCNN and DeepLabv3 . We manually inspect the data to remove spurious detections (e.g. irrelevant people in the background, and scene objects misdetected as people). We also ensure consistent subject identification across frames and views, and manually fix inaccurate 2D joint detections, mostly due to body-body and body-scene occlusions. See Supp. Mat. for more details.

Per-frame fitting. As in , given Kinect depth and 2D joints, we first optimize the SMPL-X parameters for each subject/frame separately, minimizing an objective function similar to that defined in :

where β\bm{\beta}, γ,θ,ϕ\bm{\gamma},\bm{\theta},\bm{\phi} are optimized SMPL-X parameters.

Given the preprocessed OpenPose 2D joints JOPvJ_{OP}^{v} from nn views (v∈{1,...,n}v\in\{1,...,n\}), the multi-view joint error term EJE_{J} minimizes the sum of 2D distances between JOPvJ_{OP}^{v} and the 2D projection of SMPL-X joints onto camera view vv for all views:

where KvK_{v} denotes the intrinsics parameters of camera vv, and TvT_{v} denotes the extrinsics between Cam vv and Cam 1. The depth term EDE_{D} penalizes discrepancies between the estimated body surface and body depth point clouds for all views; EpriorE_{prior} represents body pose, shape and expression priors; EcontactE_{contact} encourages scene-body contacts; and EcollE_{coll} penalizes scene-body collisions. The λi\lambda_{i}s weight the contribution of each term. We refer the reader to for more details.

Kinect-HoloLens2 calibration refinement. The Kinect-Hololens2 calibration is represented by the extrinsics TT between Kinect Cam1, and the HoloLens2 coordinate system’s origin. For each capture session, this origin is fixed in the world ; as the HMD moves, its head tracker provides the transformation between this origin and each egocentric frame tt, denoted by TtegoT^{ego}_{t}. To address the inaccurate initial Kinect-HoloLens2 calibration TinitT_{init} caused by imperfect HoloLens2 depth factory calibration, we propose a keypoint-based scheme to refine it. For each frame tt, we project the 3D SMPL-X joints J3D,tJ_{3D,t} (obtained from per-frame fitting, in Cam1’s coordinate) onto the egocentric image. We minimize the 2D error between the projected 2D joints and the OpenPose joint detections JOP,tegoJ_{OP,t}^{ego} of the egocentric frame tt, and optimize the transformation TT:

where KegoK^{ego} denotes the HoloLens2 RGB camera intrinsic parameters, and λ\lambda weights the regularizer.

Temporally consistent fitting. Per-frame fitting gives us a set of reasonable, initial pose estimates, which however are jittery and inconsistent over time. We therefore run a second optimization stage based on LEMO priors to obtain smooth, realistic human motions. Furthermore, to improve consistency between egocentric- and third-person-view estimates, we consider also egocentric data given the refined Kinect-HoloLens2 calibration. We take OpenPose 2D joint estimations from HoloLens2 RGB frames and use them as further constraints. Still, we optimize for each subject separately. The resulting objective function minimized in the temporal fitting stage is:

where EfricE_{fric} is the contact friction term defined in to prevent body sliding, EsmoothE_{smooth} and EpriorE_{prior} denote temporal and static priors as in . EJegoE_{J_{ego}} is the 2D projection term which minimizes the error between OpenPose detections on egocentric view frames and the 2D projections of SMPL-X joints onto the egocentric view; EJegoE_{J_{ego}} is only enabled for the interactee when they are visible in the egocentric frames. The λi\lambda_{i}s weight balance the contribution of each term.

EgoBody Dataset

EgoBody collects 125 sequences from 36 subjects (18 male and 18 female) performing diverse social interactions in 15 indoor scenes. In total, there are 219,731 synchronized frames captured from Azure Kinects, from multiple third-person-views. We refer to this as the “Multi-view (MV)Set”. For each MV frame, we provide 3D human full-body pose and shape annotations (as SMPL-X parameters) for both interacting subjects together with the 3D scene mesh. Furthermore, we have 199,111 egocentric RGB frames (the “EgoSet”), captured from HoloLens2, calibrated and synchronized with Kinect frames. Given the camera wearer’s head motion, the interactee is not visible in every egocentric frame; in total, we have 175,611 frames with the interactee visible in the egocentric view (“EgoSet-interactee”). Fig. 5 shows example images. For EgoSet, we also collect the head, hand and eye tracking data, plus the depth frames from the HoloLens2. We also provide SMPL body annotations via the official transfer tool . Below we provide dataset statistics; for more detailed analysis and ground truth annotation quality please refer to the Supp. Mat.

Training/validation/test splits. We split data into training, validation and test sets such that they have no overlapping subjects. The EgoBody training set contains 116,630 MVSet frames, 105,388 EgoSet frames and 90,124 EgoSet-interactee frames. The EgoBody validation set contains 29,140 MVSet frames, 25,416 EgoSet frames and 23,332 EgoSet-interactee frames. The test set contains 73,961 MV frames, 68,307 EgoSet frames and 62,155 Ego-interactee frames.

Joint visibility. The camera wearer’s motion, the headset’s field of view and the close distance between the interacting subjects cause the interactee to be often truncated in the egocentric view. To quantify the occurrences of truncations, we project the fitted 3D body joints onto the HoloLens2 images, and deem a projected 2D joint as “visible” if it lies inside the image. As shown in Fig. 6 (2nd row, right), the lower body parts are more frequently truncated in the images. Please refer to Sec. 5 for the impact of joint visibility on 3DHPS estimation performance.

Eye gaze and attention. We can combine the HoloLens2 eye gaze tracking with our 3D reconstruction of the scene/people to estimate the 3D location the user looks at, and project it on the egocentric images (interpreted as where the user’s “attention” is focused), thereby obtaining valuable data to understand interactions. We observe that the camera wearer’s attention is highly focused on the interactee during interactions. and tends to be closer to the upper body joints (Fig. 4), which in turn results in lower visibility for the lower body parts.

Experiments

We leverage EgoBody to introduce the first benchmark for 3D human pose and shape (3DHPS) estimation from egocentric images. Given a single RGB image of a target subject, the goal of a 3DHPS method is to estimate a human body mesh and a set of camera parameters, which best explain the image data. State-of-the-art (SoTA) 3DHPS methods are mostly trained and evaluated on third-person-view data, and their performance is starting to saturate on common third-person-view datasets (see Fig. 4); yet, their capabilities to generalize to real-world scenarios (e.g. cropped or blurry images) are still limited . With EgoBody, we can test their capabilities on egocentric images.

We define a benchmark for 3DHPS methods on our EgoSet-interactee test set. Within the social interaction scenarios, the input will be an egocentric view image of the interactee. We evaluate SoTA methods and show that their performance significantly drops on our data. We expose limitations of existing methods by in-depth analysis (Sec. 5.2), given that the egocentric view brings considerable challenges that are rarely present in existing third-person-view datasets.

We also provide valuable insights to boost their performance for egocentric scenarios. In particular, we show that our EgoSet-interactee training set can help address the challenges brought by egocentric view data: using it, we fine-tune three recent methods, SPIN , METRO and EFT , achieving significantly improved accuracy and robustness on both our test set (Sec. 5.3) and over a cross-dataset evaluation on the You2Me dataset (Sec. 5.4).

We employ two common metrics: Mean Per-Joint Position Error (MPJPE) and Vertex-to-Vertex (V2V) errors. We use two types of alignments before computing the accuracy for each metric: (1) translation-only alignment (aligns the bodies at the pelvis joint ) and (2) Procrustes Alignment (“PA”, solves for scale, translation and rotation). Results are by default reported with translation-only alignment unless specified with the “PA-” prefix. MPJPE is the mean Euclidean distance between predicted and ground-truth 3D joints, evaluated on 24 SMPL body joints. V2V error is the mean Euclidean distance over all body vertices, computed between two meshes.

2 Baseline Evaluation

Tab. 3 summarizes the evaluation of SoTA 3DHPS methods from different categories: (1) fitting-based method ; and regression-based methods that (2) predict parameters of a parametric body model or (3) predict non-parametric body meshes . For each baseline method, we use the best performing model provided by the authors (trained with the optimal training data).

In Fig. 4 we plot the PA-MPJPE error of these methods on our dataset and on an existing major third-person-view benchmarkThe results on 3DPW are taken from the respective original papers., On average, the methods yield a 77% higher 3D joint error on EgoBody than on 3DPW. More importantly, while the accuracy curve drives towards saturation on 3DPW, different SoTA methods still show largely varying performance on our dataset. This suggests that current datasets are not sufficient to train models that can handle egocentric view images well. Below we discuss two key challenging factors that impact performance.

Motion blur. Motion blur is common in the egocentric view images due to the motion of the camera wearer. To study how motion blur influences 3DHPS estimation accuracy, we plot in Fig. 6 (1st row, left) the MPJPE of all methods vs. the image sharpness score. The sharpness score is defined as the variance of the Laplacian of an image , upper-thresholded at 60; higher scores mean sharper images. We observe that, surprisingly, most methods are insensitive to blurriness, except for heavily blurred cases (score <10). However, our fine-tuned models (SPIN-ft / METRO-ft / EFT-ft ) are more robust against motion blur: among all methods, they achieve the lowest standard deviation over the seven image sharpness levels; see the number next to each method in the legend of Fig. 6 (1st row, left).

Joint visibility. While most 3DHPS methods assume that the target body is (almost) fully visible in the image as in existing third-person-view datasets such as 3DPW , this is seldom the case in egocentric view images. To assess the importance of this issue, we analyze the performance of each baseline with respect to the portion of visible body joints (“visibility”, see Sec. 4) in the images from our test set. The result is summarized in Fig. 6 (row 1, right). Note that our definition of a joint’s visibility is related to, but differs from, the concept of occlusion: both measure how much pixel information is missing for a body part, but visibility focuses on how much of the body is truncated from the image. A joint that is occluded by an object can still be considered visible by our definition.

Overall, all methods yield a lower error when there is less body truncation. Two recent methods, PARE and EFT , achieve the best results. PARE is designed to be robust against occlusions by explicitly employing a body part attention mechanism, whereas EFT handles body truncation “implicitly” by aggressively cropping images as training data augmentation.

We further plot the MPJPE and the invisibility ratio of each joint group in Fig. 6 (2nd row). Overall the two are in accordance: the lesser a joint is visible, the higher the error it exhibits. An exception is on the wrist joints: despite good visibility, their error remains relatively high. As observed also in , high errors on the extremities are a common problem with existing 3DHPS models, possibly because most current models only use a single, global feature from the input image for regression. This points to potential future work that deploys local image features, which has been shown effective in recent 3DHPS models .

3 Baseline Improvement

To evaluate the effectiveness of the EgoBody training set, we use it to fine-tune three of the baseline methods: two model-based methods, SPIN and EFT , as they both use the same architecture (HMR network) that is the backbone for many other recent models ; a model-free method METRO which directly predicts the body mesh. The pre-trained EFT differs from SPIN majorly in that it is trained with extended 3D pseudo ground-truth data (from the EFT-dataset) and uses aggressive image cropping as data augmentation. We use the same hyperparameters provided by the authors and select the fine-tuned model with the best validation score.

As shown in Tab. 3, after fine-tuning, the error is largely reduced for all three methods on all metrics: SPIN-ft/METRO-ft/EFT-ft has 42%/36%/18% lower MPJPE, and 35%/33%/14% V2V than their corresponding original models.

The improvement can also be seen for all blurriness/visibility categories in Figs. 6. For the motion blur specifically, the fine-tuned models not only achieve a lower error at every image sharpness level, but also show increased robustness. This is shown by the standard deviations of each method across the sharpness levels, dropping from 27.3 to 2.4 for SPIN, from 16.7 to 3.2 for METRO, and from 11.1 to 1.3 for EFT, respectively, after fine-tuning. The results show that our training set can serve as an effective source to adapt existing 3DHPS methods to the egocentric setting.

4 Cross-dataset Evaluation on You2Me

Is the effect of our training set only specific to our capture scenario, or does it generalize to other egocentric pose estimation datasets? To verify this, we evaluate SPIN, EFT and METRO against their fine-tuned counterparts on the You2Me dataset. Here we report the PA-MPJPE for pose errors (in mm): SPIN (152.8) vs. SPIN-ft (87.9); EFT (95.8) vs. EFT-ft (85.6), and METRO (117.7) vs. METRO-ft (88.2). Again, fine-tuning on our training set improves all models’ performance; see Supp. Mat. for more details. These results suggest that our data empowers existing models with the ability to address challenges faced in the generic egocentric view setup.

Conclusion

We presented EgoBody, a dataset capturing human pose, shape and motions of interacting people in diverse environments. EgoBody collects multi-modal egocentric- and third-person-view data, accompanied by ground-truth 3D human pose and shape for all interacting subjects. With this dataset, we introduced a benchmark on egocentric-view 3D human body pose and shape (3DHPS) estimation, systematically evaluated and analyzed limitations of state-of-the-art methods on the egocentric setting, and demonstrated a significant, generalizable performance gain in them with the help of our annotations. This paper has shown EgoBody’s unique value for the 3DHPS estimation task, and we see its great potential in moving the fields towards a better understanding of egocentric human motions, behaviors, and social interactions. In the future, adding more participants and even richer data modalities (e.g. audio recordings and motion annotations by natural language descriptions) could further enrich the dataset.

Acknowledgements. This work was supported by the SNF grant 200021 204840 and Microsoft Mixed Reality & AI Zurich Lab PhD scholarship. Qianli Ma is partially funded by the Max Planck ETH Center for Learning Systems. We sincerely thank Francis Engelmann, Korrawe Karunratanakul, Theodora Kontogianni, Qi Ma, Marko Mihajlovic, Sergey Prokudin, Matias Turkulainen, Rui Wang , Shaofei Wang and Samokhvalov Vyacheslav for helping with the data capture and processing, Xucong Zhang for the discussion of data collection and Jonas Hein for the discussion of the hardware setup.

References

Appendix 0.A Details of Dataset Building

To spatially calibrate the Kinects-Kinects and Kinects-HoloLens2, we employ a checkerboard to obtain an initial calibration; we then refine the result by running ICP on the scene point clouds reconstructed from the depth sensor of the devices. Additionally, the Kinects-HoloLens2 calibration is refined by the keypoint-based optimization scheme as described in the main paper Sec. 3.3. To register camera data into the 3D scene, we first manually annotate a set of correspondence points between the scene mesh and scene point clouds given by the Kinect depth frames to obtain an initial rigid transformation, which is again refined via ICP.

A.2 SMPL-X Body Model

A.3 Data Processing

Body point cloud extraction. We use Mask-RCNN to get coarse human instance segmentation masks in Kinect RGB frames and refine the masks with DeepLabv3 . The obtained human instance segmentation masks are mapped to Kinect depth frames to segment the human body point clouds from point clouds extracted from depth.

Subject index reordering. Note that OpenPose , DeepLabv3 and Mask-RCNN all work per-frame, without temporal tracking. For each sequence, we therefore reorder subject indices from OpenPose detections and human masks by their relative position to each other in 2D, such that each subject has a consistent index across all frames and all Kinect views.

Data cleaning. 2D joint detection and human instance segmentation can fail in the presence of body-body or body-scene occlusions (Fig. S1 left/middle). Thus we manually clean the failed detections and exclude them from the reconstruction pipeline. We also manually clean inaccurate 2D joint detections due to self-occlusions (Fig. S1 right). We also leverage the depth information from Kinect cameras, to filter out 2D joints with a large difference between its depth value and the median depth value of all 2D joints of the target person.

EgoSet-interactee subset selecting. Egocentric image frames with extreme human body truncations are excluded from EgoSet-interactee subset by the following filtering procedure. We run OpenPose 2D joint detection on all egocentric image frames, and manually exclude spurious detections (irrelevant people in the background or false positives on scene objects) for each frame. As OpenPose may split the joints from the same body into several detections, we merge them into one body in such case, where the joint conflict is resolved by the confidence score. We include the frames in EgoSet-interactee subset when at least six valid joints (OpenPose BODY_25 format) of the interactee are detected. We threshold the joint confidence score by 0.2 as in . Note that five joints concentrated on the head are considered as one joint due to the close distances, as well as the four joints on each foot. The bounding box is computed with the processed results.

Appendix 0.B Ground-truth Annotation Quality

To evaluate the 2D accuracy of the reconstructed body in the egocentric view, we randomly select 1,517 frames and manually annotate their 2D joints via Amazon Mechanical Turk (AMT), using the SMPL-X skeleton definition. By projecting 3D joints estimated by our pipeline on the egocentric view, the mean 2D joint error (2D Euclidean distance between the projections and AMT annotations) is 31.08 pixels (image in 1920×10801920\times 1080 resolution). We also evaluate the 3D accuracy of our ground truth on two metrics: Scan-to-body Chamfer Distance (1.65cm), and 3D per joint error (PJE) (4.32cm). The 3D PJE is computed as follows: we leverage AMT to annotate 2D body joints in all Kinect views for 100 frames, from which the annotated 3D joints are obtained via multi-view triangulation. The 3D PJE is then measured between our ground truth SMPL-X body joints and the annotated 3D joints.

B.2 Motion Smoothness

The motion smoothness of our dataset is evaluated with the Power Spectrum KL divergence score (PSKL) as in : we measure the distance between the distribution of joint accelerations in our dataset and that in the high-quality mocap dataset AMASS . The lower the score is, the more the motions resemble the natural motions in AMASS. Besides, we compare with HPS , a recent egocentric view dataset (Tab. S1). The significantly lower PSKL score of EgoBody reflects the high quality of our ground truth motions.

Appendix 0.C More Statistics

Joint invisibility. We consider 2 types of “invisibility” measurements: frame-wise invisibility and joint-wise invisibility ratio. For each frame, the frame-wise invisibility ratio calculates the percentage of invisible joints among all body joints. As shown in Fig. S2(a), partial body invisibility occurs in most frames, with even over 60% body joints invisible in extreme cases. For each body joint, the joint-wise invisibility ratio calculates the ratio of frames when the joint is invisible among all frames (See Fig. S2(b)). The lower part of the body exhibits higher chances of invisibility (knees around 50% and feet around 80%). The upper body parts are more visible: neck, shoulder, spine, and elbow joints above all, while wrist and head joints have slightly higher invisibility (around 10%).

Eye gaze and attention. The HoloLens2 eye tracking provides the eye gaze 3D ray’s starting point and orientation, which we can intersect with our 3D reconstructions to calculate the location the user looks at. By projecting the 3D eye gaze point onto the egocentric images, we perform analysis on the distances between this projected 2D eye gaze point (attention area) and the body joints of the interactee on the egocentric view images. For all frames where the 2D gaze point lies within the image, Fig. S3(a) plots the distribution of the Euclidean distance between the 2D gaze point and its nearest body joint. For more than 75%75\% of the frames, the distance between the 2D gaze point and its nearest joint lies within 120 (pixels), indicating that the camera wearer’s attention is highly focused on the interactee during interactions. The mean distance between each joint and the 2D gaze point over all frames (Fig. S3(b)) reveals that the subjects’ attention tends to be closer to the upper body joints during interactions.

Interaction distance. EgoBody covers a large range of indoor interaction distances: the 3D Euclidean distance between the pelvis joints of the interacting subjects ranges from 0.90m to 3.48m. Fig. S4(a) shows the distribution of the interaction distances between 2 interacting subjects.

Motion blur. The image sharpness score (variance of the Laplacian of the image) quantifies the motion blur in our dataset (the distribution is shown in Fig. S4(b)). A higher score indicates sharper images.

Appendix 0.D Experiments: Details and Discussions

While MPJPE is a commonly used measurement for pose estimation accuracy, it is very sparse and does not penalize the error caused by wrong joint twisting. This motivates us to also include the V2V error as a metric in our benchmark: it is not only a straightforward measure for the shape estimation, but also measures the pose error in a denser and stricter way than MPJPE as it also penalizes erroneous longitudinal joint rotations.

By default we consider the MPJPE and V2V errors without the Procrustes Alignment (PA), as PA eliminates the discrepancy in the global orientation, a major source of errors for most methods. Deprecating PA-based metrics is becoming a recent trend .

D.2 Baseline Improvement on EgoBody

Implementation details. We fine-tune SPIN , METRO and EFT using the official codes, but with slight customization in the training as follows. For SPIN, we disable the SMPLify-in-the-loop during training since the EgoBody training set already provides direct 3D supervision from the pseudo ground truth. Note that this is in fact the default setting in SPIN when the 3D ground truth is available.

Extended results and discussions. As shown in Fig. S6, while SPIN works well on images when the full body is visible, it fails on images where the subjects are truncated. EFT, in contrast, is more robust against such truncation as the model is trained on aggressively cropped images as data augmentation during training. The effectiveness of EFT’s data augmentation is further supported by our fine-tuning experiments. After fine tuning SPIN and EFT on our training set, both models show greater robustness against motion blur and image truncation, and quantitatively achieve lower errors than the original models on all metrics, as shown in the main paper Tab. 3. Together with the experiments on the You2Me dataset (see main paper Sec. 5.4), this shows that our training set can help adapt existing 3DHPS models to egocentric view data.

D.3 Details of the Cross-dataset Evaluation on You2Me

Here we report the experimental setup of the You2Me dataset experiment. The You2Me dataset provides egocentric view images (taken with a chest-mounted GoPro camera), and the ground truth 3D joint locations of both the camera wearer and the interactee. The ground truth 3D joints are however in a world coordinate system, making it infeasible to compute the translation-only MPJPE (see main paper Sec. 5.1): since the camera calibration between the GoPro and the world coordinate is unknown, even a perfect prediction (in the camera coordinate system) may differ from the ground truth up to a rigid transformation. To account for this problem, we first perform the Procrustes Alignment, which solves for the scale, translation and global rotation, to align the predicted 3D body joints with the ground truth, and then compute the MPJPE of the aligned bodies, resulting in the PA-MPJPE errors reported in the paper.

D.4 Experiment with Motion Blur Augmentation

Data augmentation could potentially simulate blurring and truncation. Can the performance of existing methods be enhanced on EgoBody by simply fine-tuning them on the original dataset that they are trained on with extra data augmentation? EFT is trained with aggressive image cropping which to an extent simulates body truncation in our dataset. Indeed its superior performance has proven the effect of data augmentation, but such augmentation does not fully address the challenges in EgoBody: a clear performance gap can be seen between the original EFT and all models fine-tuned on our dataset (SPIN-ft, METRO-ft, EFT-ft, see Tab. 4 in main paper). Likewise, here we additionally analyze motion blur augmentation. We fine-tune the pre-trained SPIN model with additional motion blur augmentation on the datasets it’s originally trained on. For the motion blur we randomly set blur direction, angle, and kernel size to blur the training images with a probability of 0.5 during fine-tuning, and we experiment with multiple settings with different kernel sizes. No improvements are observed compared with the original SPIN model when evaluated on our egocentric test set (Tab. S2). Both observations indicate that existing data augmentation techniques cannot fully resolve the challenges in the egocentric setup, and EgoBody fills this gap.

D.5 Details for Baseline Methods

CMR firstly regresses 3D locations of SMPL body vertices via a graph convolutional network. An image-based CNN encodes the input image into a feature vector, which is attached to the graph network defined by a mesh template. A Multi-Layer Perceptron (MLP) predicts the SMPL parameters based on regressed body vertices.

METRO adopts the model-free formulation to estimate body vertices and 3D body joints directly. A transformer encoder models interactions for vertex-vertex and vertex-joint via self-attention mechanism.

SPIN integrates iterative optimization loops into the neural network training to combine advantages of both regression-based and optimization-based methods. The optimization fits the SMPL body model to 2D joints on the image to enable more robust supervision for the regressor.

EFT augments existing large-scale 2D datasets with 3D annotations by Exemplar Fine-Tuning. Starting from a pre-trained 3D pose regressor, the model weights are fine-tuned by fitting 2D joints to images. Supervised by the obtained 3D annotations, a model with the same architecture as SPIN is trained with extreme crop augmentation and auxiliary input representations.

LGD proposes to use neural networks to predict the parameter update rules in the optimization framework. A Gradient Updating Network regresses the update step for SMPL parameters in each optimization iteration.

PARE leverages the visibility information of each body part, and predicts body-part-guided attention masks to achieve robust prediction for SMPL parameters with body occlusions.

Appendix 0.E AMT Annotation Details

To evaluate the body shape and pose annotation accuracy, we collect manual annotations of 2D locations of 17 body joints on the EgoSet-interactee frames via Amazon Mechanical Turk (AMT). The user interface is illustrated in Fig. S7(a). We exclude body joints that are ambiguous for manual annotating (head, spine1, spine2, spine3, left_collar, right_collar) from the first 22 SMPL-X body joints, and add the nose joint which is easy to define for users. The user is provided with the definition of body joints (Fig. S7(b)), and an image of the target person to annotate (in case of irrelevant people in the background). Self-occluded keypoints need to be inferred, while keypoints occluded by scene objects are not required to be annotated. We downsample with a rate of 50 on the EgoSet-interactee data, which yields a total number of 2,286 frames.

For better annotation quality, each frame is annotated by five users, and joints annotated by less than three users are ignored. A small part of the users flip the left and right side, inducing non-negligible noise for the ground-truth evaluation. To address this issue, we filter out the outliers and correct the flipped annotations by the following procedure. For each annotated joint, the 2D distance from each annotation xi\mathbf{x}_{i} (i=0,1,...,4i=0,1,...,4) to the mean location xˉ\bar{\mathbf{x}} is calculated as di=∣∣xi−xˉ∣∣2d_{i}=||\mathbf{x}_{i}-\bar{\mathbf{x}}||_{2}. An annotation is considered as the outlier if di−dˉσ>1.5\frac{d_{i}-\bar{d}}{\sigma}>1.5, where dˉ\bar{d} is the average distance, and σ=1n(di−dˉ)2\sigma=\sqrt{\frac{1}{n}(d_{i}-\bar{d})^{2}} is the standard deviation of the distance. For joints that have the counterpart on the other side of body we flip the left/right side of the annotations and perform the same outlier detection, to fix cases when the users flip the left and right side.

Appendix 0.F Limitations

As there exists no solution to synchronize HoloLens2 and Kinect via hardware, we align their clocks via software, using a flashlight which is visible to all devices as signal for the first frame. Although HoloLens2 exhibits frame drops occasionally, the corresponding frames of all devices can be aligned according to the timestamps provided by HoloLens2 Research Mode API . Besides, we empirically observe a small temporal misalignment. In our case, the misalignment can be observed for fast motions (for example, for hand movements, as shown in Fig. S5). However, this issue is inevitable for the synchronization between third-person view cameras and HMDs . Despite the small misalignment, our reconstruction reaches a high accuracy as proved by the reconstruction accuracy in Sec. 4.2.