BlazePose GHUM Holistic: Real-time 3D Human Landmarks and Pose Estimation

Ivan Grishchenko, Valentin Bazarevsky, Andrei Zanfir, Eduard Gabriel Bazavan, Mihai Zanfir, Richard Yee, Karthik Raveendran, Matsvei Zhdanovich, Matthias Grundmann, Cristian Sminchisescu

Introduction

Accurate real-time inference of the human body skeleton enables a variety of applications ranging from motion capture to interactive video games. Running on a consumer’s phone or laptop without relying on dedicated sensors (e.g. IR on the Kinect ) would vastly democratize the technology for non-professional users. Over the past decade, there have been numerous advances in estimating 3D landmarks on the human body or volumetric representations . The majority of these are either too computational expensive to be run on mobile devices, require a specialized lab setup (e.g. multi-camera) or lack sufficient detail w.r.t. body topology (e.g. no fingers). To overcome these limitations, we created BlazePose GHUM Holistic, a lightweight neural network pipeline that predicts 3D landmarks and pose of the human body on-device, including hands, from a single monocular image, runs in real-time at 15 FPS on most modern mobile phones and browsers and is available to developers and creators via MediaPipe.

Our approach handles three issues that we identified with current interactive motion capture solutions. First, we address the challenge of acquiring diverse 3D ground truth of the human body. We introduce a novel approach that is based on fitting a statistical 3D human model GHUM to a diverse set of 2D annotations. To further improve accuracy, we propose to use depth ordering annotations as supervision during fitting.

Second, the standard topology for on-device body landmarks prediction does not include hands and fingers, e.g. 33 landmarks by BlazePose and PoseNet . This impedes the building of a unified motion capture system for the entire body. To overcome this limitation, we use a spatial transformer approach to crop high-res hand regions from the original image seeded by BlazePose’s palm prediction as a prior. Then we run a retrained hand tracking model to predict 21 3D hand landmarks for each hand in a single feed-forward pass.

Finally, we tackle the issue of moving beyond accurate 3D representations of the human body towards high-level semantic understanding and mapping to enable expressive use cases like 3D avatars. To this end, we present a lightweight model that predicts the full body and hand pose from 3D landmarks represented as joint rotations of a 3D human model. We employ a statistical 3D human model called GHUM that additionally acts as a pose prior constraining the state of predictions to plausible and realistic movements.

3D body landmarks

The key challenge to build the 3D part of our pose model is obtaining realistic, in-the-wild 3D data. In contrast to 2D, which can be obtained via human annotation, accurate manual 3D annotation is a uniquely challenging task. It requires either a lab setup, specialised hardware with depth sensors for 3D scans or building a synthetic dataset which has an inherent domain gap to real world data. Each of these approaches introduces additional challenges to preserve a good level of human and environment diversity as present in real-world pictures.

Our approach is based on a statistical 3D human body model GHUM, which is built using a large corpus of human shapes and motions. To obtain 3D human body pose ground truth, we fit the GHUM model to our existing 2D pose dataset, which covers various domains (yoga/fitness/dance), surroundings (indoor/outdoor), devices (mobile/laptop). In doing so, we obtain real world 3D keypoint coordinates in metric space (see fig. 3). During the fitting process the shape and the pose variables of GHUM were optimized such that the reconstructed model aligns with the underlying image evidence. This includes 2D keypoint and silhouette semantic segmentation alignment as well as shape and pose regularization terms (check HUND, THUNDR).

Due to the nature of 3D to 2D projection fitting can result in several realistic 3D body poses for the given 2D annotation (i.e. with the same X and Y but different Z). To minimize this ambiguity we asked annotators to provide depth order between pose skeleton edges where they are certain, similarly to Ordinal Depth Supervision approach. This task proved to be easy showing high consistency between annotators (98% on cross-validation) and helped to reduce the depth ordering errors during fitting from 25% to 3%.

Model

BlazePose GHUM Holistic utilizes a two-step detector-tracker approach where the tracker operates on a cropped region-of-interest containing the human within the original image. Thus the model is trained to predict 3D body pose in relative coordinates of a metric space with origin in the subject’s hips center.

Experiments

To evaluate the quality of our models against other well-performing publicly available solutions, we use yoga domain, as one of the most challenging in body poses articulations. Each image contains only a single person located 2-4 meters from the camera. To be consistent with other solutions, we perform evaluation only for 17 keypoints from COCO topology. As shown in Tab. 1, our approach outperforms that of leading commercial and academic solutions. We train a variety of different models (Heavy, Full, Lite) that provide various levels of trade-off between accuracy and on-device inference speed, see Tab. 2.

Hand landmarks for holistic human pose

Holistic human pose estimation requires accurate tracking of hands in addition to the body. To include hand landmarks in the BlazePose topology, we had to solve two main issues. First, the BlazePose model’s input resolution of 256x256 is insufficient to capture hand details. Second, we need to ensure that hands prediction is spatially invariant for left and right hands to guarantee the same level of accuracy as leading on-device hand tracking models . Using a single model would require us to a) increase the input resolution which in turn would slow down model inference and b) balance the loss for small and big body parts a single pixel error is more significant for fingers, than for example the hip landmark. Instead, we opted for an alternative approach that uses a separate hand prediction modelthat is inferred on high resolution crops obtained from ROIs provided by BlazePose (i.e. re-cropping).

The four palm landmarks produced by BlazePose give us a rough estimate of the hand region. But this is insufficient to use as input for the subsequent hand landmark model as it was trained on more accurate crops obtained from all 21 hand landmarks with only slight transformation augmentations. One solution would be to train with more aggressive transformation augmentations but this significantly reduces accuracy of the model due to its limited capacity targeting real-time on-device inference. To close this gap, we trained a re-crop model that takes raw crop from BlazePose output and refines it on a higher resolution to a level acceptable for the subsequent hand landmark model. Using BlazePose as a prior for hand locations also helps us to untangle hands and body parts of different people.

Experiments

We compare the hand landmarks quality of a standalone hand tracker with our novel re-crop approach, see Tab. 3. Our pipeline produces a better hand landmark quality. We postulate this is due to a more accurate crop on the current frame, whereas the approach of uses a one frame delay.

Body pose estimation

To enable expressive use cases like 3D avatars, we must go beyond inferring 3D coordinates and obtain a high-level semantic understanding via joint rotations that can be used to drive rigged characters. To this end, we obtain the 3d pose and shape GHUM mesh from the set of our previously 3d landmarks inferred directly from images.

Lifter.

The BlazePose GHUM Holistic network outputs 33 body landmarks, and 21 landmarks for each hand, in a root-centered 3D camera coordinate system. In order to increase the expressivity of the outputs, without sacrifing performance, we propose a sample-and-train methodology, based on a novel GHUM Lifter neural network. Our neural network takes as input the concatenated 3D body and hands landmarks and outputs GHUM mesh parameters.

MLPMixer.

At the core of our lifter lies an MLP-Mixer architecture . Our mixer takes as input a sequence of S 3D keypoints, each one projected to a desired hidden dimension C, with the same projection matrix. The mixer consists of multiple layers of identical size, and each layer consists of two MLP blocks. The first one is the token-mixing MLP: it acts on columns of the mixer input and is shared across all columns. The second one is the channel-mixing MLP: it acts on rows of the token-mixing MLP and is shared across all rows. Each MLP block contains two fully-connected layers and a nonlinearity applied independently to each row of its input data tensor.

Full model training.

Experiments.

In order to validate our GHUM lifter architecture we ran experiments on a held out test set of 10,000 in the wild images with very challenging poses (yoga, fitness, dancing) containing GHUM fits which were curated for any errors. We compared with various SOTA methods for 3D pose and shape estimation and observe significant improvements for our method even though it runs one order of magnitude faster. Results are reported in Tab. 4.

Applications and conclusion

BlazePose GHUM Holistic provides a simple and meaningful representation of the human body that can be used for multiple applications right out of the box. 3D landmarks enable us to move beyond 2D space to a real world coordinate system. Pose estimation gives us high level interpretation of 3D landmarks as well as extra hand points to add fine grain details when needed (e.g. gesture detection). It enables motion capture for avatar control, repetition counting and posture correction for fitness and sports, as well as 3D effects for AR/VR.

To showcase BlazePose GHUM Holistic we created an open-source avatar demo (see Fig. 4) in MediaPipe (https://mediapipe.dev). The demo is available for web browsers and allows to control body and hands of a standard Mixamo avatar at 15 FPS (MacBook Pro 15-inch 2017).

In the future, we plan on further improving the model that predicts 3D landmarks as well as capturing reliable facial expressions to add them to BlazePose GHUM Holistic.

References