Effective Whole-body Pose Estimation with Two-stages Distillation
Zhendong Yang, Ailing Zeng, Chun Yuan, Yu Li
Introduction
Whole-body pose estimation plays a crucial role in numerous human-centric perception, understanding, and generation tasks, including 3D whole-body mesh recovery , human-object interaction , and pose-conditioned human image and motion generation . Furthermore, capturing human poses for virtual content creation and VR/AR has gained significant popularity, relying on user-friendly algorithms like OpenPose and MediaPipe . Despite the convenience of these tools, their performance remains unsatisfactory, limiting their potential. Therefore, further advancements in human pose estimation technology are essential to fully unleash the potential of user-driven content creation. Compared with human pose estimation with body-only keypoints detection, whole-body pose estimation faces more challenges from 1) the hierarchical structures of the human body for fine-grained keypoints localization; 2) the small resolutions of hand and face; 3) the complex body parts matching for multiple persons in an image, especially for occlusion and complex hand poses; 4) data limitation, especially for diverse hand pose and head pose for the whole-body images.
Besides, before deploying a model, it is essential to compress it into a lightweight network. The basic compression tools comprise distillation , pruning , and quantization . Knowledge distillation (KD) is proposed to enhance the efficiency of a compact model without incurring extra costs during inference. This technique enables a student to inherit knowledge from a larger teacher and has found widespread application in various tasks, such as classification , detection , and segmentation .
In this paper, we explore KD for whole-body pose estimation to benefit many downstream applications, resulting in a series of real-time pose estimators with high performance and efficiency. Specifically, we propose a novel two-stages pose distillation framework DWPose, which achieves state-of-the-art performance, as shown in Fig. 1. We adopt the latest pose estimator RTMPose as the basic model, which has been trained on COCO-WholeBody .
In the first-stage distillation, we natively leverage the teacher’s (e.g., RTMPose-x) intermediate layer and final logits to guide the student model (e.g., RTMPose-l). Previous pose training distinguishes keypoints via visibility and only uses visible keypoints for supervision. Unlike that, we use the teacher’s complete outputs with both visible and invisible keypoints as final logits, which can impart reasonable and comprehensive values to facilitate the student’s learning process. Meanwhile, we employ a weight-decay strategy to enhance the efficacy, gradually reducing the distillation’s weight throughout the entire training phase. Due to a better head will determine a more precise localization, the second-stage distillation proposes a head-aware self-KD to enhance the capacity of the head. We construct two identical models and select one as the teacher and the other as the student to be updated. The student backbone is frozen, and only its head is updated through the logit-based distillation. Notably, this plug-and-play approach allows the student to achieve better results with 20% training time, whether trained from scratch with distillation or without, and can be used for any dense prediction heads.
Data volume and diversity addressing different scales of human body parts will affect the model performance. Suffering from the limited holistic annotated keypoints on existing datasets, existing estimators fail to localize well on fine-grained fingers and face landmarks. Thus, we explore the data impact by incorporating an additional UBody dataset, primarily comprising diverse face and hand keypoints captured in various real-life scenes.
Therefore, our contributions can be summarized as:
We introduce a two-stage pose knowledge distillation method, pursuing efficient and precise whole-body pose estimation.
To break the whole-body data limitation, we explore more comprehensive training data, especially on diverse and expressive hand gestures and facial expressions, making it practical for real-life applications.
Based on the latest RTMPose as our base model, our proposed distillation and data strategies can significantly improve RTMPose-l from 64.8% to 66.5% AP, even surpassing RTMPose-x teacher with 65.3% AP. We also validate the powerful effectiveness and efficiency of DWPose on the generation task.
Related work
This task targets locating expressive body, hand, feet, and face keypoints for all persons in an image simultaneously . Due to the lack of whole-body annotations, most previous models are designed for body-only , hand-only , or face-only . Openpose combines different datasets for separate body parts. MediaPipe builds a perception pipeline for easy-to-use applications, especially for whole-body landmark detection. With the emergence of whole-body data , the models for whole-body pose estimation make great progress . Specifically, ZoomNet proposes the first top-down method with a hierarchical single network to solve the scale variance of different body parts. ZoomNAS further explores a neural architecture search framework for jointly searching the model architecture and the connections between different sub-modules to promote both accuracy and efficiency. TCFormer introduces progressive clustering and merging vision tokens for various locations, sizes, and shapes in multiple stages, preserving different scale information well. Recently, RTMPose has discussed key factors in pose estimation and built a real-time model, achieving state-of-the-art results on COCO-WholeBody. However, it still suffers from redundant model designs and data limitations, especially for diverse hand and face poses.
2 Knowledge Distillation
Knowledge distillation is a way to compress the model. Hinton et al. first proposed to supervise the student with the soft labels from the teacher’s output. The method is originally designed for classification and is also called logit-based distillation. Some following works utilize teacher’s logits in different ways, transferring more knowledge from soft labels, target and non-target logits . From the logit-based distillation to feature-based distillation, the knowledge is transferred from intermediate layers and it extends the distillation to various tasks, including detection , segmentation , generation and so on.
Utilizing KD in human pose estimation has been rarely studied . Existing works either distill the heavy heatmaps for body-only pose estimation or focus on gathering separate body-part experts’ knowledge into a single deep network designed for whole-body 2D-3D pose detection . 2D whole-body pose estimation is a basic task for 3D pose estimation and is more holistic than body-only pose estimation. Our proposed DWPose is the first work to explore efficient KD strategies for this task.
Method
In the following, we provide a detailed exposition of the two-stage pose distillation (TPD). As shown in Fig. 2, it comprises two distinct stages. The first-stage distillation involves a pre-trained teacher guiding the student from scratch at both the feature and logit levels. On the other hand, the second-stage distillation can be considered a self-KD approach. The model employs its own logits to train its head without any labeled data, leading to significant performance enhancements within a concise training period.
We denote the feature from the teacher’s and student’s backbone as and , and the teacher and student’s final output logit as and . The first-stage distillation forces the student to learn the teacher’s feature and logit .
For the feature-based distillation, we force the student to mimic the teacher’s layer from the backbone directly. We utilize MSE loss to calculate the distance between the student’s feature and the teacher’s feature . To learn the knowledge from the teacher’s feature map, the distillation loss of the feature can be formulated as:
where is a 11 convolutional layer to reshape the to the same dimension as . denote the height, width and channel of the teacher’s feature.
1.2 Logit-based distillation
RTMPose predicts pose keypoints with a SimCC-based algorithm that treats keypoint localization as a classification task for horizontal and vertical coordinates. Following this design, we can also apply the logit-based knowledge method to it. To begin with, we review the original classification loss for RTMPose as follows:
where is the number of the person samples in a batch, is the number of keypoints, e.g., 133 for COCO-WholeBody , is the length of the x or y localization bins. is a target weight mask to distinguish invisible keypoints. is the label value.
For the logit-based distillation, we follow the form of the original loss . It’s worth noting that we drop the target weight mask for distillation. Different from the label value, the invisible keypoints can also be distributed a reasonable value by the teacher. So we argue such value is also helpful, and we also verify it in Sec. 5.6. The distillation loss of the logits can be formulated as:
1.3 Weight-decay strategy for distillation
With feature distillation loss and logits distillation loss , we can train the student with the total loss as:
where and are hyper-parameters to balance the loss.
Inspired by a detection distillation method TADF , we apply a weight-decay strategy for the distillation to reduce the distillation penalty gradually. This strategy helps the student to focus more on the label and achieve better performance. We utilize a time function to implement the strategy, which is as follows:
where is the current epoch and is the total epochs for training. Then the final loss for the first-stage distillation can be formulated as:
2 The Second-stage distillation
In the second distillation stage, we try to utilize the trained student model to teach itself for a better performance. In this way, it can bring improvements for the students, whether trained from scratch with distillation or not.
The pose estimator comprises the encoder (backbone) and decoder (head). Based on the trained model, we first build a student with a trained backbone and an untrained head. The teacher is the same model with a trained backbone and head. During training, we freeze the student’s backbone and update the head. Because the teacher and the student have the same architecture, we only need to extract the feature from the backbone once. Then, the feature is fed into the teacher’s trained head and the student’s untrained head to get the logits and , respectively. Following the form in Eq. 3, we train the student with for the second-stage distillation. It’s worth noting that we drop the original loss , which is calculated with label value. Using to denote the hyper-parameter for loss scale, the final loss for the second-stage distillation can be formulated as:
Different from previous self-KD methods, our proposed head-aware distillation can efficiently distill the knowledge from the head with only 20% training time and further improve the localization capability.
Experiments
Datasets. We conduct experiments on COCO and UBody . For the COCO dataset, we follow the standard splitting of train2017 and val2017, which use the 118K train images for training and 5K val images for testing, respectively. Unless specifically, we adopt a commonly used person detector provided by SimpleBaseline with 56.4% AP for the COCO val dataset. UBody consists of over 1M frames from 15 real-life scenarios. It provides the corresponding 133 2d keypoints and SMPL-X parameters. Notably, the original dataset only focuses on 3D whole-body estimation and does not validate the effectiveness of 2D annotations. we pick every frame at an interval of 10 frames from the video used for both training and testing.
Implementation details. For the first-stage distillation, we utilize two hyper-parameters and in Eq. 6 to balance the loss scale. For all the experiments, we adopt on both COCO and UBody. The second-stage distillation has one hyper-parameter to balance the loss scale in Eq. 7. For all the experiments, we adopt . The training setting, such as the optimizer, learning rate, and training epochs for the first-stage distillation, is the same as training the student without distillation . For the two-stage distillation, we only need a short training time of about 1/5 of the whole training epochs. The other training settings still remain the same. This early stopping method helps to save much time for training. We use 8 GPUs to conduct the experiments with MMPose based on Pytorch . As a top-down pose estimator following RTMPose, we use the person detection boxes with 56.4 AP on the COCO val2017 dataset and the provided ground-truth box on Ubody.
2 Main Results
For a fair comparison, we evaluate our models on the public COCO-WholeBody dataset. As shown in Tab. 1, we utilize the larger RTMPose-x and RTMPose-l as the teacher to guide DWPose-l and the other student models, respectively. With our TPD, the models with different sizes and input resolutions all achieve significant improvements. Specifically, DWPose-m gets 60.6 whole AP with 2.2 GFLOPs. The performance is 4.1% higher than the baseline, while the consumption for the inference still remains the same, making it friendly to deploy. Interestingly, DWPose-l achieves 63.1 and 66.5 whole AP under two different input resolutions, which both beat the teacher RTMPose-x with fewer parameters and flops. DWPose-l also achieves the new state-of-the-art model for human whole-body pose estimation. With the proposed distillation TPD and more data, we provide a series of effective models with competitive accuracy.
Fig. 3 shows some qualitative comparisons of how our distillation helps the students to perform better. TPD helps the model to predict more accurately, reduces false pose detection, and increases true pose detection, especially for the improvement of finger keypoint localization. We also compare our state-of-the-art model with two widely used models OpenPose and MediaPipe , as presented in Fig. 4. Our DWPose also surpasses the other two methods significantly, especially for the robustness of truncation, occlusion, and effectiveness of fine-grained localization. This enables our method to replace these popular methods to benefit corresponding downstream applications effectively.
Analysis
We explore the effects of the proposed distillation method TPD and the used UBody dataset of improving whole-body pose estimation in Tab. 2. The model achieves considerable AP improvements with an extra UBody dataset, especially for hand pose detection. For example, RTMPose-l achieves 62.1 whole-body AP and 55.1 Hand AP, which is 1.0 and 3.2 higher than the model trained just on COCO-Wholebody. Our distillation method TPD further boosts the model’s performance, helping the model to get 63.1% whole AP. The results demonstrate the effectiveness of the TPD method and UBody dataset.
2 Performance on UBody
We first evaluate our method on the COCO WholeBody dataset, as we describe above. In this subsection, we evaluate the models on the UBody dataset, as shown in Tab. 3. We compare the models under two different input resolutions and report the corresponding AP of different human parts. The extra UBody data for training and our distillation method TPD are both helpful to the students, bringing them significant improvements under both input resolutions. Different from COCO, the gains that our TPD brings on UBody mainly focus on the face and hand. As for COCO, the performance on the body, foot, and hand all get significant improvements, but the gains on the face are limited, as shown in Tab. 2. The results on UBody also demonstrate the effectiveness of our distillation method TPD.
3 Effects of First and Second Stage Distillation
We propose the two-stage pose distillation (TPD), which includes the first and second stage distillation. To evaluate the impact of each distillation stage, we conduct experiments by using RTMPose-x to distill RTMPose-l on the mixed dataset, as presented in Tab. 4. Both two distillation stages are beneficial for the students, and their combination leads to further improvements in performance. When combining the first-stage and second-stage distillation together, we achieve 63.1 whole AP, which surpasses the performance achieved by using either distillation loss alone. It’s worth noting that the second-stage distillation just needs to fine-tune the head, which helps to save much training time. Interestingly, it helps the student to surpass the teacher RTMPose-x with 63.0% AP.
4 Second-stage Distillation for Trained Models
Our second-stage distillation is available not only for the models trained with our first-stage distillation but also for those trained without distillation. So it can be applied when there lacks a better and larger teacher. We can utilize the model itself as a teacher to improve it with a short training time. As shown in Tab. 5, we pick three different models and evaluate our second-stage distillation on COCO and the combination of COCO and UBody. For all settings, models with S2 significantly improve, especially for the foot and hand. Compared with traditional distillation and self-KD, it saves much time in training the model from scratch and costs to obtain a better model.
5 Ablation Study of the First-stage Distillation
As we describe in Eq. 6, our first-stage distillation calculates the loss through the ground-truth label (GT), teacher’s feature (Fea), and teacher’s logits (Logit). Furthermore, we apply a weight-decay strategy (Decay) to further improve the student. In this subsection, we analyze the effects of every component by using RTMPose-l to distill RTMPose-m, as shown in Tab. 6. The knowledge from the feature brings the student 1.4% AP gains. When combing the distillation on the logit, the AP gains get to 1.6%. This proves that the knowledge from the feature and logit are both helpful and complementary to each other. Finally, the weight-decay strategy brings another 0.3% AP gains, helping the student to achieve 62.3% AP.
Interestingly, we try to drop the GT label and train the student just with the teacher’s logit. The student achieves 60.9% AP, which is even 0.5% higher than the model trained with the GT label. This indicates we can label the new data through a teacher model instead of annotating manually, which can save much cost in time and manual efforts, and achieve a better model through such data for training. However, when combining the feature distillation together, the performance with the teacher’s logit gets lower than that with the GT label. Thus, we adopt the GT, Fea, and Logit together for distillation.
6 Target Mask for Logit-based Distillation
In our logit-based distillation, we deliberately omit the target weight mask , which is employed to differentiate between visible and invisible keypoints, as shown in Eq. 3. We conducted an in-depth investigation into how this target mask affects the distillation process. As indicated in Tab. 7, it is evident that the presence of the target weight mask significantly hampers the distillation performance, resulting in a notable 1.1% drop in the student’s performance. These results underscore the significance of the teacher’s input for invisible keypoints, affirming its positive impact on the student’s learning process.
7 Better Pose, Better Image Generation
Recently, controllable image generation has witnessed significant advancements. For human image generation, precise skeleton information is crucial to guide the pose, particularly for whole-body skeletons. Mainstream techniques like ControlNet often rely on OpenPose due to its efficiency and user-friendly nature in generating human poses. However, OpenPose’s performance, as shown in Tab. 1, reaches only 44.2% AP, which leaves room for improvement. Consequently, we aim to replace OpenPose with our DWPose to enhance ControlNet’s image generation without the need for additional training. Utilizing a top-down approach, we first employ YOLO-X to detect all individuals and then use our pose estimator to extract keypoints from the detection results, thus boosting the overall image generation process.
In Fig. 5, we employ ControlNet to visualize and compare the generated images using both OpenPose and our DWPose, demonstrating that a more precise and expressive skeleton leads to higher-quality image generation. Additionally, we present a comparison of inference speed with OpenPose in Tab. 8. Thanks to the efficient architecture of RTMPose , DWPose requires only about one percent of the time taken by OpenPose to infer the same image. Moreover, as the number of persons in the image increases, the runtime for OpenPose significantly increases. For a single person, the inference times for OpenPose and DWPose are 5.78 s and 0.068 s, respectively. However, when the number of persons reaches nine, the inference time for OpenPose triples, whereas the inference time for DWPose is only about 1.5 times longer.
Conclusion
In this paper, we aim to obtain both an efficient and effective model for human whole-body pose estimation. To this end, we apply distillation to the latest effective RTMPose. Accordingly, we first propose a Two-stage Pose Distillation to enhance the lightweight model’s performance. Moreover, the second-stage distillation is available when a larger teacher lacks, and it only needs a short training time to obtain a better model. Then, we investigate the UBody dataset to further improve its performance, obtaining DWPose. Extensive experiments prove that our method is simple yet effective. We also explore the impact of a better pose estimator on the controllable image generation task.
Acknowledgement. This work was supported by the National Key RD Program of China (2022YFB4701400/4701402), the SZSTC project Grant (JCYJ20190809172201639, WDZC20200820200655001), Shenzhen Key Laboratory (ZDSYS20210623092001004).