Versatile Multi-Modal Pre-Training for Human-Centric Perception

Fangzhou Hong, Liang Pan, Zhongang Cai, Ziwei Liu

Introduction

As a long-standing problem, human-centric perception has been studied for decades, ranging from sparse prediction tasks, such as human action recognition , 2D keypoints detection and 3D pose estimation , to dense prediction tasks, such as human parsing and DensePose prediction . Unfortunately, to train a model with reasonable generalizability and robustness, an enormous amount of labeled real data is necessary, which is extremely expensive to collect and annotate. Therefore, it is desirable to have a versatile pre-train model that can serve as a foundation for all the aforementioned human-centric perception tasks.

With the development of sensors, the human body can be more conveniently perceived and represented in multiple modalities, such as RGB, depth, and infrared. In this work, we argue that the multi-modality nature of human-centric data can induce effective representations that transfer well to various downstream tasks, due to three major advantages: 1) Learning a modal-invariant latent space through pre-training helps efficient task-relevant mutual information extraction. 2) A single versatile pre-train model on multi-modal data facilitates multiple downstream tasks using various modalities. 3) Our multi-modal pre-train setting bridges heterogeneous human-centric datasets through their common modality, which benefits the generalizability of pre-train models.

We mainly explore two groups of modalities as shown in Fig. 1 a): dense representations (e.g. RGB, depth, infrared) and sparse representations (e.g. 2D keypoints, 3D pose). Dense representations can provide rich texture and/or 3D geometry information. But they are mostly low-level and noisy. On the contrary, sparse representations obtained by off-the-shelf tools are semantic and structured. But the sparsity results in insufficient details. We highlight that it is non-trivial to integrate these heterogeneous modalities into a unified pre-training framework for the following two main challenges: 1) learning representations suitable for dense prediction tasks in the multi-modality setting; 2) using weak priors from sparse representations effectively for pre-training.

Challenge 1: Dense Targets. Existing methods perform contrastive learning densely on pixel-level features to achieve view-invariance for dense prediction tasks. However, those methods require multiple views of a static 3D scene , which is inapplicable for human-centric applications with only single view. Furthermore, it is preferable to learn representations that are continuously and orderly distributed over the human body. In light of this, we generalize the widely used InfoNCE and propose a dense intra-sample contrastive learning objective that applies a soft pixel-level contrastive target, which can facilitate learning ordinal and continuous dense feature distributions.

Challenge 2: Sparse Priors. To employ priors in contrastive learning, previous works mainly use the supervision to generate semantically positive pairs. However, these methods only focus on the sample-level contrastive learning, which means each sample is encoded to a global embedding. It is not optimal for human dense prediction tasks. To this end, we propose a sparse structure-aware contrastive learning target, which uses semantic correspondences across samples as positive pairs to complement positive intra-sample pairs. Particularly, leveraging sparse human priors leads to an embedding space where semantically corresponding parts are aligned more closely.

To sum up, we propose HCMoCo, a Human-Centric multi-Modal Contrastive learning framework for versatile multi-modal pre-training. To fully leverage multi-modal observations, HCMoCo effectively utilizes both dense measurements and sparse priors using the following three-levels hierarchical contrastive learning objectives: 1) sample-level modality-invariant representation learning; 2) dense intra-sample contrastive learning; 3) sparse structure-aware contrastive learning. As an effort towards establishing a comprehensive multi-modal human parsing benchmark dataset, we label human segments for RGB-D images from NTU RGB+D dataset , and contribute the NTURGBD-Parsing-4K dataset. To evaluate HCMoCo, we transfer our pre-train model to four human-centric downstream tasks using different modalities, including DensePose estimation (RGB) , human parsing using RGB or depth frames, and 3D pose estimation (depth) . Under full set and data-efficient training settings, HCMoCo constantly achieves better performance than training from scratch or pre-train on ImageNet. To name a few, as shown in Fig. 1 b), we achieve 7.16% improvement in terms of GPS AP on 10%10\% training data of DensePose estimation; 12% improvement in terms of mIoU on 20%20\% training data of Human3.6M human parsing. Moreover, we evaluate the modal-invariance of the latent space learned by HCMoCo for dense prediction on NTURGBD-Parsing-4K with two settings: cross-modality supervision and missing-modality inference. Compared against conventional contrastive learning targets, our method improves the segmentation mIoU by 29% and 24% for the two settings, respectively. To the best of our knowledge, we are the first to study multi-modal pre-training for human-centric perception.

The main contributions are summarized below: 1) As the first endeavor, we provide an in-depth analysis for human-centric pre-training, which is formulated as a challenging multi-modal contrastive learning problem. 2) Together with the novel hierarchical contrastive learning objectives, a comprehensive framework HCMoCo is proposed for effective pre-training for human-centric tasks. 3) Through extensive experiments, HCMoCo achieves superior performance than existing methods, and meanwhile shows promising modal-invariance properties. 4) To benefit multi-modal human-centric perception, we contribute an RGB-D human parsing dataset, NTURGBD-Parsing-4K.

Related Work

Human-Centric Perception. Many efforts have been put into human-centric perception in decades. Lots of work in 2D keypoint detection has achieved robust and accurate performance. 3D pose estimation has long been a challenging problem and is approached from two aspects, lifting from 2D keypoints and predicting from depth map . Human parsing can be defined in two ways. The first one parses garments together with visible body parts . The second one only focuses on parsing human parts . In this work, we focus on the second setting because the depth and 2D keypoints do not contain the texture information needed for garment parsing. There are a few works about human parsing on depth maps. However, the data and annotations are too coarse or unavailable. To further push the accuracy of human-centric perception, DensePose is proposed to densely model each human body surface point. The cost of DensePose annotation is enormous. Therefore, we also explore data-efficient learning of DensePose.

Multi-Modal Contrastive Learning. Multi-modality naturally provides different views of the same sample which fits well into the contrastive learning framework. CMC proposes the first multi-view contrastive learning paradigm which takes any number of views. CLIP learns a joint latent space from large-scale paired image-language dataset. Extensive studies focus on video-audio contrastive learning. Recently, 2D-3D contrastive learning has also been studied with the development in 3D computer vision. In this work, aside from commonly used modalities, we also explore the potential of 2D keypoints in human-centric contrastive learning.

Our Approach

In this section, we first introduce the general paradigm of HCMoCo (3.1). Following the design principles (3.2), hierarchical contrastive learning targets are formally introduced (3.3). Next, an instantiation of HCMoCo is introduced (3.4). Finally, we propose two applications of HCMoCo to show the versatility (3.5).

As shown in Fig. 2, HCMoCo takes multiple modalities of perceived human body as input. The target is to learn human-centric representations, which can be transferred to downstream tasks. The input modalities can be categorized into dense and sparse representations. Dense representations Id∗I_{d}^{*} are the direct output of imaging sensors, e.g. RGB, depth, infrared. They typically contain rich information but are low-level and noisy. Sparse representations are structured abstractions of the human body, e.g. 2D keypoints, 3D pose, which can be formulated as graph Is∗=G(V,E)I_{s}^{*}=G(V,E). Different representations of the same view of a human should be spatially aligned, which means intra-sample correspondences can be obtained for dense contrastive learning. HCMoCo aims to pre-train multiple encoders Ed∗E_{d}^{*} and Es∗E_{s}^{*} that produce embeddings of dense representations and sparse representations for downstream tasks transfer.

which is analyzed and explained as follows.

2 Principles of Learning Targets Design

In this subsection, we analyze the intuitions when designing learning targets, which makes the following three principles. 1) Mutual Information Maximization: Inspired by , we propose to maximize the lower bound on mutual information, which has been proved by many previous works to be able to produce strong pre-train models. 2) Continuous and Ordinal Feature Distribution: Inspired by the property of human-centric perception, it is desirable for the feature maps of the human body to be continuous and ordinal. The human body is a structural and continuous surface. The dense predictions, e.g. human parsing , DensePose , are also continuous. Therefore, such property should also be reflected in the learned representations. Besides, for an anchor point on human surfaces, closer points have higher probabilities of sharing similar semantics with the anchor point than that of far away points. Therefore, the learned dense representations should also align with such ordinal relationship. 3) Structure-Aware Semantic Consistency: Sparse representations are abstractions of the human body, which contains valuable structural semantics about the human body. Instead of identity information, the human pose and structure understanding are the keys to our target downstream tasks. Therefore, it is reasonable to eliminate the identity information and enhance the structure information by enforcing structure-aware semantic consistency where semantically close embeddings (e.g. embeddings of left hands from different samples) are pulled close and vice versa.

3 Hierarchical Contrastive Learning Targets

Based on the above three principles, we formally define hierarchical contrastive learning targets in this subsection.

Sample-level modality-invariant representation learning aims at learning a joint latent space at the sample level using global embeddings, which fulfills the first principle. Inspired by , the learning target can be formulated as

where F∗gF_{*}^{g} is a set of global embeddings of one modality, SgS_{g} is the set of F∗gF_{*}^{g} of all modalities, f2gˉ\bar{f_{2}^{g}} is the embedding of the paired view of that of f1gf_{1}^{g}, τ\tau is the temperature. It should be noticed that f1gf_{1}^{g} can be sampled from the global embeddings of either dense or sparse representations.

where Wxymn\mathcal{W}_{xy}^{mn} is the weight, τ\tau is the temperature, (x,y),(m,n),(x′,y′)(x,y),(m,n),(x^{\prime},y^{\prime}) are coordinates on the dense representation, 1≤x,x′,m≤H1\leq x,x^{\prime},m\leq H, 1≤y,y′,n≤W1\leq y,y^{\prime},n\leq W. The above equation is a generalized version of InfoNCE . InfoNCE is a special case when Wxymn\mathcal{W}_{xy}^{mn} is set to 11 if x=mx=m and y=ny=n else . We use the normalized distances as the weights, which is formulated as

For each pair of dense representations, the above learning target is calculated between each pair of dense embeddings. Therefore, the whole learning target is defined as

where F∗dF_{*}^{d} is a set of dense embeddings of one modality, SdS_{d} is the set of all F∗dF_{*}^{d}, fd1df_{d1}^{d} and fd2df_{d2}^{d} are two paired embeddings. It should be noticed that the ‘soft’ learning target cannot guarantee an ordinal feature distribution. Instead, it serves as a computationally efficient relaxation of the requirement of ordinal distribution.

Sparse structure-aware contrastive learning takes two sparse representations f1sf_{1}^{s} and f2sf_{2}^{s} as inputs. The paired features f1jsf_{1j}^{s} and f2jsf_{2j}^{s} (i.e. features of the jj-th joint) should be pulled close while unpaired features are pushed away. The two sparse representations can be sampled from the same or different modalities, intra- or inter-sample. The intra-sample alignment satisfies the first principle. The inter-sample alignment follows the third principle. The sparse structure-aware contrastive learning target is formulated as

where F∗sF_{*}^{s} is a set of sparse embeddings of one modality, SsS_{s} is the set of F∗sF_{*}^{s}, τ\tau is the temperature, f1s,f2sf_{1}^{s},f_{2}^{s} are sampled from the union of F1sF_{1}^{s} and F2sF_{2}^{s}. To conclude, the overall learning target is formulated as Eq. 1, where λ∗\lambda_{*} are the weights to balance the targets.

4 Instantiation of HCMoCo

In this section, we introduce an instantiation of HCMoCo. As shown in Fig. 3, for dense representations, we use RGB and depth. Large-scale paired human RGB and depth data is easy to obtain with affordable sensors e.g. Kinect. These two modalities are the most commonly encountered in human-centric tasks . Moreover, a proper pre-train model for depth is highly desired. Therefore, RGB and depth are reasonable choices of human dense representations, both of which are easy to acquire and important to downstream tasks. For sparse representations, 2D keypoints are used, which provide positions of human body joints in the image coordinate. Off-the-shelf tools are available to quickly and robustly extract human 2D keypoints given RGB images. Using 2D keypoints as the sparse representation is a good balance between the amount of human prior and acquisition difficulty.

For RGB inputs Id1I_{d}^{1}, an image encoder Ed1E_{d}^{1} is applied to obtain feature maps Ed1(Id1)E_{d}^{1}(I_{d}^{1}). Similarly, for depth inputs Id2I_{d}^{2}, an image encoder or 3D encoder Ed2E_{d}^{2} can be applied to extract feature maps Ed2(Id2)E_{d}^{2}(I_{d}^{2}). 2D keypoints IsI_{s} are encoded by a GCN-based encoder EsE_{s} to produce sparse features Es(Is)E_{s}(I_{s}). Mapper networks comprise a single linear layer and a normalization operation.

As for the implementation of contrastive learning targets, we choose to use a memory pool to store all the global embeddings which are updated in a momentum way. Sparse and dense embeddings cannot all fit in memory. Therefore, for the last two types of contrastive learning targets, the negative samples are sampled within a mini-batch.

5 Versatility of HCMoCo

On top of the pre-train framework HCMoCo, we propose to further extend it on two direct applications: cross-modality supervision and missing-modality inference. The extensions are based on the key design of HCMoCo: dense intra-sample contrastive learning target. With the feature maps of different modalities aligned, it is straightforward to implement the two extensions, which are shown in Fig. 4.

Cross-Modality Supervision is a novel task where we train the network on the source modality, while test on the target modality. This is a practical scenario where people transfer the knowledge of some single modality dataset to other modalities. At training time, an additional downstream task head (e.g. segmentation head) DD is attached to the backbone of the source modality. The hierarchical contrastive learning targets L\mathcal{L} together with downstream task loss L′\mathcal{L}^{\prime} are used for end-to-end training. At inference time, DD is attached to the backbone of the target modality. The extracted feature maps of the target modality are passed to DD for prediction.

Missing-Modality Inference is another novel task where we train the network using multi-modal data and inference on single modality. Multi-modal data collection in practice would inevitably result in data with incomplete modalities, which brings the requirement of missing-modality inference. At training time, the feature maps of multiple modalities are fused using max-pooling and fed to a downstream task head DD. Similarly, hierarchical contrastive learning targets L\mathcal{L} and downstream task loss L′\mathcal{L}^{\prime} are used for co-training. At inference time, the feature map of a single modality is passed to DD for missing-modality inference.

NTURGBD-Parsing-4K Dataset

Although RGB human parsing has been well studied , human parsing on depth or RGB-D data has not been fully addressed due to the lack of labeled data. Therefore, we contribute the first RGB-D human parsing dataset: NTURGBD-Parsing-4K. The RGB and depth are uniformly sampled from NTU RGB+D (60/120) . As shown in Fig. 5, we annotate 24 human parts for paired RGB-D data. The partition protocols follow that of . The train and test set both have 19631963 samples. The whole dataset contains 39263926 samples. Hopefully, by contributing this dataset, we could promote the development of both human perception and multi-modality learning.

Experiments

Implementation Details. The default RGB and depth encoders are HRNet-W18 . The default datasets for pre-train are NTU RGB+D and MPII . The former provides paired indoor human RGB, depth, and 2D keypoints, The latter provides in-the-wild human RGB and 2D keypoints. Mixing human data from different domains helps our pre-train models adapt to a wilder domain.

Downstream Tasks. We test our pre-train models on four different human-centric downstream tasks, two on RGB images and two on depth. 1) DensePose estimation on COCO : DensePose aims at mapping pixels of the observed human body to the surface of a 3D human body, which is a highly challenging task. 2) RGB human parsing on Human3.6M . Human3.6M provides pure human part segmentation, which aligns with our objectives. We uniformly sample 2fps of the video for training and evaluation. 3) Depth human parsing on NTURGBD-Parsing-4K. 4) 3D pose estimation from depth maps on ITOP (only side view). For all the above downstream tasks, we use the pre-train backbones for end-to-end fine-tune.

Comparison Methods. Since there are few previous human-centric multi-modal pre-train methods, we propose to use general multi-modal contrastive learning methods CMC and MMV as the baselines. Although there are other multi-modal contrastive learning works, they either require the multi-view calibration or focus on multi-modal downstream tasks and therefore are not suitable for comparison. In addition, for RGB tasks, we also experiment under two settings, one initializes encoders with supervised ImageNet (IN) pre-train while the other does not.

2 Performance on Downstream Tasks

DensePose Estimation. As shown in Tab. 1, we test DensePose estimation under two settings: full and 10%10\% of the training data. The trained models are tested on the full validation set of DensePose. Firstly, if not using IN pre-train, our pre-train model significantly outperforms both ‘From Scratch’ and two baseline methods. Especially under 10%10\% of training data, 12.7% improvement in terms of GPS AP is observed. And our pre-train model even outperforms that using IN pre-train by 4.13% in terms of GPS AP. When we use IN pre-train as initialization, which is a common practice for 2D tasks, our method still outperforms all the baselines. Our method surpasses IN pre-train by 7.2% and 5.4% in terms of GPS/GPSM AP under 10%10\% setting. To further test the performance of in-domain transfer, we also pre-train models using training sets of NTU RGB+D and COCO. The performance gain under 10%10\% setting further improves to 9.7% and 7.5% in terms of GPS/GPSM AP.

RGB Human Parsing. As shown in Tab. 2, we test four settings on Human3.6M : full, 20%20\%, 10%10\% and 1%1\% training data. In all settings, our method outperforms all baselines in all metrics. On full training data, we outperform IN pre-train by 5.6% in terms of mIoU. The performance gain increases with the amount of training data decreases. It is worth noticing that with only 10%10\% of training data, our method outperforms IN pre-train with full training data.

Depth Human Parsing. As shown in Tab. 4, we test the pre-train depth backbone on our proposed Dataset NTURGBD-Parsing-4K with all training data and 20%20\% training data. We outperform all baselines on two settings. Especially, only using 20%20\% of training data, we surpass IN pre-train by 6.4% and MMV by 4.6% in terms of mIoU.

3D Pose Estimation. As shown in Tab. 5, we test the pre-train depth backbone on ITOP with six different ratios of training data. Our pre-train model outperforms all baselines on most settings. With only 10%10\% training data, the accuracy of our method outperforms that of IN pre-train with all training data. It is also worth noticing that 0.1%0.1\% of training data are 1717 samples, which makes this a few-shot learning setting. With such limited training data, IN pre-train barely produce meaningful results, while our method improves the accuracy by 48.2%.

3 Ablation Study

In this subsection, we perform a thorough ablation study on HCMoCo to justify the design choices. As shown in Tab. 3, we firstly report the results of only applying sample-level modality-invariant representation learning. Then we add dense intra-sample contrastive learning and sparse structure-aware contrastive learning in order. To further demonstrate the effect of the ‘soft’ design in dense intra-sample contrastive learning, we also report results of the ‘hard’ learning target, which takes the form of a classic InfoNCE . We report the results of the ablation study on all four downstream tasks under data-efficient settings.

For DensePose estimation, it is important to learn feature maps that are continuously and ordinally distributed, which is the expected result of soft dense intra-sample contrastive learning. The performance gain of the soft learning target over the hard counterpart justifies the observation and the learning target design. The dense intra-sample contrastive learning also shows superiority on three other downstream tasks, which shows the importance of fine-grained contrastive learning targets for dense prediction tasks.

Explicitly injecting human prior into the network through sparse structure-aware contrastive learning also proves its effectiveness by further improving the performance on DensePose. Thanks to the strong hints provided by 2D keypoints, the performance of 3D pose estimation is improved. Moreover, the sparse structure-aware contrastive learning boosts the performance of human parsing both on RGB and depth maps by 1.9% and 2.8% respectively in terms of mIoU. Although 2D keypoints are sparse priors, they still provide the rough location of each part of the human body, which facilitate the feature alignment of same body parts. To summarize, the sparse and dense learning targets both contribute to the performance of our methods, which is in line with our analysis.

4 Performance on HCMoCo Versatility

Cross-Modality Supervision. We test the cross-modality supervision pipeline on the task of human parsing on NTURGBD-Parsing-4K because it has two modalities and respective dense annotations. Two baseline methods are adopted: 1) using CMC contrastive learning target; 2) no contrastive learning target. For a fair comparison, the backbones of all methods are initialized by CMC pre-train. At training time, the target modality of training data is not available. We experiment on two settings where we supervise on RGB, test on depth (RGB →\rightarrow Depth), and vice versa (Depth →\rightarrow RGB). As shown in Tab. 6, our method outperforms both baselines under two settings. Specifically, our method improves the mIoU of both settings by 29.2% and 23.0%, respectively. Even compared to methods with direct supervision, we can achieve comparable results.

Missing-Modality Inference. For missing-modality inference, we report the experiments on the same dataset and same baselines as above. As shown in Tab. 7, with no pixel-level alignment, the two baseline methods struggle in two missing-modality settings i.e. ‘Only RGB’ and ‘Only Depth’. While our method improves the segmentation mIoU by 24.3% and 19.6% on two settings.

5 Further Analysis

Faster Convergence. One of the advantages of pre-training is the fast convergence speed when transferred to downstream tasks. Our HCMoCo also shows superiority in this feature. We log the validation mIoU of Human3.6M human parsing at different training epochs. As shown in Fig. 6, compared with IN pre-train and CMC , our pre-train model is able to converge within a few training epochs in both the full training data and data-efficient settings.

Changing Backbone. So far our experiments are all performed on HRNet-W18. To further demonstrate HCMoCo’s performance on other backbones, for the 2D backbone, we also experiment with HRNet-W32 . For the depth backbone, we choose to test with PointNet++ . For the RGB pre-train model, we experiment on the 10%10\% DensePose estimation. For the depth pre-train model, we experiment on the 20%20\% NTURGBD-Parsing-4K. As shown in Tab. 8, our method outperforms its pre-train counterparts by a reasonable margin, which is in line with our previous experimental results.

Discussion and Conclusion

In this work, we propose the first versatile multi-modal pre-training framework HCMoCo specifically designed for human-centric perception tasks. Hierarchical contrastive learning targets are designed based on the nature of human datasets and the requirements of human-centric downstream tasks. Extensive experiments on four different human downstream tasks of different modalities demonstrated the effectiveness of our pre-training framework. We contribute a new RGB-D human parsing dataset NTURGBD-Parsing-4K to support research of human perception on RGB-D data. Besides downstream task transfer, we also propose two novel applications of HCMoCo to show its versatility and ability in cross-modal reasoning.

Potential Negative Impacts & Limitations. Usage of large amounts of data and long training time might negatively impact the environment. Moreover, even though we did not collect any new human data in this work, human data collection could happen if our framework is used in other applications, which potentially raises privacy concerns. As for the limitations, due to limited resources, we could only experiment with one possible instantiation of HCMoCo. And for the same reason, even though the theoretical possibility exists, we do not have the chance to further scale up the amount of human dataset and network size.

Acknowledgments This work is supported by NTU NAP, MOE AcRF Tier 2 (T2EP20221-0033), and under the RIE2020 Industry Alignment Fund – Industry Collaboration Projects (IAF-ICP) Funding Initiative, as well as cash and in-kind contribution from the industry partner(s).

References

Implementation Details

Network Details. For the proposed instantiation of HCMoCo, we implement the sample-level modality-invariant representation learning target by maintaining a memory pool, which is adapted from an open-sourced implementation https://github.com/HobbitLong/PyContrast. The memory pool is updated in a momentum style with the momentum of 0.50.5. For global embeddings, we sample 1638416384 negative samples from the memory pool. For other hyper-parameters, we use a batch size of 224224, a learning rate of 0.030.03, a temperature of 0.070.07 for all three contrastive learning targets. For the pre-train, 4 NVIDIA V100 GPUs are used. The training process is divided in two steps. The first step only pre-train the model using sample-level modality-invariant representation learning target for 100100 epochs. The second stage adds the other two learning targets and trains for another 100100 epochs. The whole training process takes approximately 48 hours.

Mixing Heterogeneous Datasets. Since we mix several heterogeneous human datasets for pre-train, we need to mask out the missing modalities. For example, when we use NTU RGB+D and MPII for pre-train. The former dataset has all the required modalities, while the latter one misses depth maps. Therefore, for the hierarchical contrastive learning targets, we mask out the missing depth embeddings of MPII for all the positive pairs sampling. By using the masking technique, it is possible to combine multiple heterogeneous datasets into this pre-train paradigm as long as there are at least two common modalities.

Datasets for Pre-train. For NTU RGB+D, we only use the version with 60 actions . With the provided RGB-D videos, we uniformly sample one frame from every 30 frames, which makes 143648143648 samples. The RGB and depth frames are calibrated by the correspondences provided by the 2D keypoints positions on RGB and depths. For MPII and COCO, we use the full training sets for pre-train.

2 DensePose Estimation

For the DensePose estimation, we use the official open-sourced implementation https://github.com/facebookresearch/detectron2. For the full training set, we train the network for 130000130000 iterations with a batch size of 1616, a learning rate of 0.010.01 on 4 NVIDIA V100 GPUs, which takes around 8080 hours to train. For the 10%10\% training set, we train the network for 1300013000 iterations with a learning rate of 0.0050.005 and other settings the same, which takes 9 hours to train. The 10%10\% training set is uniformly sampled from the default ordered training set.

3 RGB Human Parsing

For the RGB human parsing on Human3.6M , we use the official HRNet semantic segmentation implementation https://github.com/HRNet/HRNet-Semantic-Segmentation. Different ratios of training settings are uniformly sampled from the default ordered full training set. For the full training set, we train the network for 5050 epochs with a learning rate of 0.0070.007, a batch size of 4040 on 2 NVIDIA V100 GPUs. For other data-efficient settings, we train the network for 150150 epochs with other settings the same. We use the the standard dataset split protocol, where the subjects 1,5,6,7,81,5,6,7,8 are for training and the subjects 99 and 1111 are for evaluation.

4 Depth Human Parsing

For the depth human parsing on NTURGBD-Parsing-4K, we use the same implementation as that of RGB human parsing. To use the HRNet to encode depth maps, we repeat the depth dimension for three times to fit the RGB input, which is also how HCMoCo deals with depth inputs. For all training settings, we train the network for 150150 epochs with a learning rate of 0.0070.007, a batch size of 8080 on 2 NVIDIA V100 GPUs. Even though the encoder is used to deal with depth inputs, we still initialize it using ImageNet pre-train for that it might help with the performance proved by some previous works .

5 3D Pose Estimation on Depth

For the 3D pose estimation from depth maps on ITOP , we choose to adapt the official implementation https://github.com/zhangboshen/A2J of A2J . The original implementation uses ResNet as the backbone. And we switch to HRNet. Since the original implementation only provides validation scripts, we re-implement the whole training pipeline. We change the original normalization method where a global mean and variance is counted for a global normalization. Instead, we perform an online instance normalization where we only centralize each depth pixel to zero mean but do not normalize its variance, since its a better way to prevent the over-fitting to the relatively small dataset. We train the network for 5050 epochs with a learning rate of 0.000350.00035 and a batch size of 1212 on one NVIDIA V100 GPU. As for the dataset, we use the side-view of ITOP since the depth maps in pre-train are side views. Following the official dataset split, there are 1799117991 samples for training and 48634863 for testing. Following the practice of A2J , we initialize the encoders using ImageNet pre-train.

6 Cross-Modality Supervision

To experiment with the cross-modality supervision, we choose the downstream task of human parsing on NTURGBD-Parsing-4K. The modalities to experiment with are RGB and depth. To make the experiment fair and the networks to converge faster, the backbones are initialized by CMC pre-train. The following descriptions are for the setting of ‘RGB→\rightarrowDepth’, where the source modality is RGB and the target modality is depth. To implement ‘Depth→\rightarrowRGB’, one can simply switch the source and target modalities. At training time, a randomly initialized segmentation header, which is the same one used for human parsing experiments, is attached to the dense mapper network of RGB. Then the network is trained with both the hierarchical contrastive learning targets L\mathcal{L} and a cross-entropy loss L′\mathcal{L}^{\prime} for the supervision of the segmentation. For the ‘No Contrastive’ baseline, we only train with L′\mathcal{L}^{\prime}. As for the ‘CMC’ baseline, the network is supervised by both the learning target proposed by CMC L\mathcal{L} and the segmentation loss L′\mathcal{L}^{\prime}. Note that, during the whole training time, including the CMC pre-train, the target modality of NTURGBD-Parsing-4K is not exposed to better simulate the application scenario. In order to build the connection between RGB and depth during training time, we mix the NTURGBD-Parsing-4K with NTU RGB+D which is the same one used for our pre-train. At inference time, we attach the trained segmentation head to the mapper network of depth. Since the dense embeddings of RGB and depth are aligned thanks to our hierarchical contrastive learning targets, it is reasonable for the segmentation head to be able to handle the dense embeddings of depth.

7 Missing-Modality Inference

We also use human parsing on NTURGBD-Parsing-4K to experiment with our extension of missing-modality inference. The basic setup is the same as that of the cross-modality supervision experiments. At training time, we take the dense embeddings of both RGB and depth together for a max pooling operation for a simple feature-level fusion. Then the fused dense embedding is passed to a segmentation header, which is the same one used by the human parsing experiment, to produce the segmentation prediction. The network is supervised with both the hierarchical contrastive learning targets L\mathcal{L} and a cross-entropy loss L′\mathcal{L}^{\prime} for segmentation supervision. Similarly, the ‘No contrastive’ baseline does not use any contrastive learning targets. The ‘CMC’ baseline uses the contrastive learning target proposed in CMC as L\mathcal{L}. At inference time, if RGB is missing, then the dense embedding of depth is passed to the trained segmentation header for prediction. Since the dense embeddings of RGB and depth are aligned and the segmentation header is trained with the fusion of both embeddings, missing one of them will still produce reasonable predictions.

More Quantitative Results

DensePose Estimation. Due to the page limitation, we could not report all metrics for the DensePose estimation. Therefore, we report them in this supplementary material. As shown in Tab. 9, detailed results of all settings mentioned in the main paper are listed. Specifically, for the initialization of the network, we test with the network randomly initialized (‘From Scratch’) and the network initialized by ImageNet pre-train (‘IN Pre-train’). As for the ratio of training data, we test with the full training set and 10%10\% of the training set. As for the pre-train datasets, we test with two combinations: NTU RGB+D ++ MPII and NTU RGB+D ++ COCO. As for the backbone, we test with HRNet-W18 and HRNet-W32. Compared with the baseline and two other state-of-the-art pre-train counterparts, our method outperforms them in most of the metrics. Especially, our method has advantages in GPS and GPSM, which are two critical metrics for DensePose quality. Additionally, we also report full results of the ablation study. The detailed results further validates the analysis in the main paper.

RGB Human Parsing. We further report detailed RGB human parsing results on Human3.6M that could not fit into the main paper. As shown in Tab. 10, we report the per-class IoU for all the settings reported in the main paper. Similarly, for the initialization of the network, we test with the network randomly initialized (‘From Scratch’) and the network initialized by ImageNet pre-train (‘IN Pre-train’). As for the ratio of training data, we test with the full training set, 20%20\%, 10%10\% and 1%1\% of the training set. The pre-train datasets are NTU RGB+D ++ MPII. In most classes, our method outperforms comparison methods. Moreover, we also report per-class IoU for the four settings in ablation study, which are in line with our analysis in the main paper.

Depth Human Parsing. We report detailed depth human parsing results on NTURGBD-Parsing-4K. As shown in Tab. 11, we report the per-class IoU for all the settings reported in the main paper. We initialize the networks using ImageNet pre-train. Two ratios of the training set, i.e. full and 20%20\%, are tested. We also change the backbone to PointNet++ (‘PN++’). Since it is a point-based backbone, the ‘background’ class is ignored and not included in the calculation of mIoU. The per-class IoU results also agree with the conclusion in the main paper that our method is superior than other comparison methods.

Cross-Modality Supervision. As shown in Tab. 12, we report detailed per-class IoU for the experiments of cross-modality supervision. In both ‘RGB→\rightarrowDepth’ and ‘Depth→\rightarrowRGB’ settings, our method outperforms other baseline methods in all classes. Especially, other baseline methods barely make correct predictions while ours makes a huge improvement.

Missing-Modality Inference. As shown in Tab. 12, we list detailed per-class IoU for the experiments of missing-modality inference. In both ‘Only RGB’ and ‘Only Depth’ settings, our method outperforms baseline methods in most classes. Therefore, the detailed results further validates the conclusions made in the main paper.

More Qualitative Results

More qualitative results of RGB human parsing on Human3.6M and depth human parsing on NTURGBD-Parsing-4K are shown in Fig. 7, Fig 8 and Fig. 9. We choose to visualize both the full training set and 10%10\% training set for RGB human parsing. The segmentation results produced by our pre-train model are superior than those of other comparison methods, especially in data-efficient settings. For challenging classes like hands and elbows, our method is capable of producing correct predictions constantly while other methods struggle. The depth map is a challenging modality for the dense prediction task like semantic segmentation. Our method manages to produce reasonable predictions that are better than those of other comparison methods.