Occluded Human Mesh Recovery
Rawal Khirodkar, Shashank Tripathi, Kris Kitani
Introduction
Estimating accurate 3D human meshes from single images has diverse applications in modeling human-scene interactions, understanding human behaviour, AR/VR and robotics. While recent approaches perform particularly well in images containing a single person, human mesh recovery for complex real-world scenes with multiple occluded people remains a challenging task. This can be attributed in part to simplifying assumptions made by existing methods. For instance, most top-down approaches expect a single subject in the input image, which affects robustness under in-the-wild scenarios containing severe person-person occlusion, such as crowding. In this paper, we address human mesh recovery in multi-person scenarios by mitigating the limitations of the single-person assumption of top-down approaches.
Current human mesh recovery methods can be categorized into top-down and bottom-up methods. Top-down methods reduce the problem to a simpler task of single human mesh recovery by relying on a person detector to detect individual bounding box for each person in the image. Since each bounding box is scaled to the same size, top-down methods are less sensitive to scale variations among subjects and can achieve pixel accurate mesh alignment . In contrast, bottom-up methods simultaneously predict meshes for all subjects in the input image but are limited to a fixed input resolution due to computational constraints. e.g. ROMP , a bottom-up method, recovers a limited number of human meshes from a resized input whereas SPIN , a top-down method, scales each bounding box to , retaining higher input resolution per person(see Fig. 1). This observation has also been discussed by Cheng et al. albeit in the context of 2D human pose estimation. Thus, top-down methods are currently the best performers on various multi-human benchmarks . Despite the advantages, due to the single-human assumption, when presented with multi-human inputs like crowded scenes, top-down methods are forced to select a single plausible mesh per detection bounding box. Bottom-up methods do not have this limitation and typically perform better under occlusion.
A general method should have both traits – be robust to scale variations and person-person occlusions. To this end, we rethink top-down human mesh recovery by predicting multiple meshes from the input bounding box. We condition the top-down model on image spatial-context in the form of body-center maps, refer Fig. 2. Our choice of using center maps for representing humans under occlusion is inspired by crowd-counting literature and recent works in detection . Our method, OCHMR, predicts the output mesh from the input image for the person of interest in the subject-specific local center-map. Similar to bottom-up methods, we also use information from the global center-map for understanding overall scene context, which is helpful for occlusion reasoning. With this strategy, we obtain the best of both worlds – OCHMR achieves pixel accurate mesh alignment similar to top-down methods and is robust to occlusions similar to bottom-up methods (See Fig. 1).
To design a top-down architecture capable of contextual conditioning using centermaps, we adopt the mechanism of feature normalization and propose a novel Context Normalization (CoNorm) block to process the global and local centermaps. The CoNorm blocks are used to inject contextual information at multiple depths in the deep feature backbone network. The spatial context is necessary for 3D occlusion reasoning, and the CoNorm block allows for adaptive normalization of intermediate features of the network without changing the backbone. We show that unlike early fusion (e.g. channel-wise concatenation) of centermaps with input image , CoNorm can effectively utilize the contextual information from the image. OCHMR is general and can be extended to other top-down human mesh recovery methods with minimal effort.
While the use of spatial-context allows our method to reason about occlusions, our method must also reason about the intersection of a set of 3D human meshes. To address this, following CRMH , we use an interpenetration loss to penalize intersections among reconstructed meshes and a differentiable depth-ordering loss for depth-consistent human mesh recovery. Furthermore, we make use of training-time data augmentation like scaling and cropping, which affords OCHMR the ability to predict meshes from a variety of body-center locations. We show that our proposed method is also robust to errors in estimated body-centers under severe-occlusion. Our empirical results show that OCHMR does not require precise centermaps that correspond to actual body-centers but can also work with any point in its vicinity.
Overall, OCHMR outperforms both top-down and bottom-up methods on various datasets. For challenging datasets such as 3DPW-PC , CrowdPose and OCHuman , containing a larger proportion of cluttered scenes (with multiple overlapping people), OCHMR sets a new state-of-the-art for 3D reconstruction error (PMPJPE) and 2D keypoint average precision (AP) achieving PMPJPE, AP and AP respectively on the val sets outperforming bottom-up methods (Tab. 1). Further, when evaluating using ground-truth bounding boxes, OCHMR dramatically improves SPIN by AP and AP on the OCHuman and CrowdPose dataset respectively. In summary:
OCHMR advances top-down human mesh recovery methods by addressing limitations caused by the single-human assumption. Our method leverages spatial-context in the form of centermaps to predict multiple mesh outputs from an input image.
We introduce novel Context Normalization (CoNorm) blocks to inject global and local centermap information at multiple depths of the top-down network.
Our approach achieves state-of-the-art results on the occluded 3DPW-PC, CrowdPose and OCHuman datasets. Empirically, we also show that OCHMR is resilient to noisy body center estimates and demonstrates robust 3D reasoning using multi-person losses.
Related Work
Deep learning has significantly advanced 3D human mesh recovery , facilitating the more challenging task of mesh recovery under severe multi-person occlusion , which is the main focus of this work.
Biased Human Mesh Recovery Benchmarks. Most benchmark datasets used for learning human mesh recovery focus on a single person and do not accurately represent the distribution of possible occlusions present in the real world. Human3.6M , HumanEva and TotalCapture are popular datasets collected using motion capture (mocap) systems using optical markers. While providing accurate annotations, they only have a single subject in the image with limited image complexity due to the lack of background variation. In contrast, datasets like MPI-INF-3DHP , PanopticStudio and 3DPW contain multi-person annotations but have limited person-person occlusion – less than 27% of all annotations have crowding (at IoU 0.5). Although previous methods leverage 2D keypoint annotations from datasets like COCO , MPII , LSP-Extended , the 2D datasets are also known to contain similar biases . These biases have affected critical design decisions in state-of-the-art methods which lead to poor generalization under heavy occlusion . Recently, challenging datasets such as OCHuman , CrowdPose and 3DPW-PC containing heavy occlusion have been proposed to capture these biases. OCHMR shows a significant improvement over existing works under such challenging conditions.
Top-Down Human Mesh Recovery. Top-down methods estimate 3D human mesh of a single person within a person bounding box. The bounding box is usually generated using a person detector . As the input bounding boxes are cropped and scaled to the same size, top-down methods are less sensitive to person scale variations in the image. In contrast, bottom-up methods have to deal with scale variations which compromises pixel alignment in the reconstruction results. For these reasons , most state-of-the-art 2D pose estimation methods are also top-down. However, top-down methods inherently assume a single person in the input image and often fail under occlusions in multi-person scenarios. Recent works like use 2D/3D poses as input along with bounding boxes for human mesh recovery. However, obtaining accurate 2D poses under occlusion is difficult and pose errors like joint swaps are magnified during the 3D reconstruction . CRMH handles multi-person scenarios by using RoI-aligned features of each person to predict the SMPL parameters. However, the reliance on bounding-box-level features makes it hard to effectively differentiate between two overlapping bounding boxes. OCHMR resolves these issues by conditioning the top-down model on image context in the form of body-centers – a representation which helps in resolving ambiguity under multi-person occlusion.
Bottom-Up Human Mesh Recovery. Unlike top-down methods, only few methods exist which use the bottom-up paradigm for human mesh recovery. Zanfir et al. uses intermediate 3D poses to estimate the 3D mesh of each person in a bottom-up fashion. ROMP uses a fixed resolution body-center map to disambiguate between multiple persons under occlusion. Due to the fixed input size, , ROMP is limited to predicting a small numbers of meshes. In contrast, OCHMR is top-down and can leverage input resizing of subject bounding box to a higher resolution for pixel accurate shape estimation. Being top-down, OCHMR can be applied to all detected persons in the input image.
Method
OCHMR leverages the strengths of both top-down and bottom-up methods for multi-person mesh recovery under severe person-person occlusion/crowding. In this section, we briefly describe the top-down method used as baseline architecture in our approach. Then we provide details of our contextual representations i.e. local and global centermaps and the context estimation network. Finally, we describe the proposed architectural improvements in the form of Context Normalization (CoNorm) blocks and multi-person losses used in training.
Similar to , we define a deep regression model as our baseline top-down architecture for human mesh recovery. The bounding box at training and inference is scaled to and is provided as an input to . Let denote the ground-truth SMPL and camera parameters corresponding to the human in the input image . The deep regression model transforms input to a single 3D mesh , such that . is trained to minimize the sum of various 2D/3D pose and shape losses (using 2D pose annotations and segmentations masks if available) denoted by .
We propose to modify the top-down deep regression model to predict multiple meshes as follows. Let be the number of ground-truth subjects present in the image . is set to the total number of subjects with atleast visible 2D keypoints in the image. Let be the corresponding ground-truth mesh parameters. Our modified deep regression model predicts instances, for an input . This is achieved by conditioning the network on the spatial-context individually for each subject. accepts both and as input and predicts where . We define the OCHMR’s single person loss as follows,
During inference, we vary the spatial-context to extract multiple mesh predictions from the same input image . In cases of severely overlapping people, it is hard for the baseline top-down method to estimate diverse body meshes from similar image patches . OCHMR uses spatial-context to resolve the implicit ambiguity of the bounding-box input representation in such multi-person cases.
2 Global and Local Center Map Estimation
Our top-down framework relies heavily on the representation of the spatial-context . It is crucial to define a representation which is explicit and robust to occlusion. Inspired by , we choose body centers to encode the spatial context of the image. Specifically, we represent the contextual information of instance as where is the body-center heatmap of all the instances present in the image and is the body-center heatmap of the instance (see Fig. 2). is calculated by thresholding and iterating over pixel locations in . While informs the network about the subject of interest, places the subject in the context of its neighbors, thereby helping the network disambiguate between occluding persons.
3 Context Normalization Block
A key challenge is to design an architecture that incorporates spatial-context as a conditioning input. A naïve early fusion approach would be to simply concatenate the input image with the spatial-context . Similarly, late fusion would concatenate feature maps from later layers within the network with appropriately down-sampled context . However, both of these approaches fail to improve performance.
We describe the Context Normalization (CoNorm) block that can be easily introduced in any existing feature extraction backbone to overcome this issue (see Fig. 3). The key intuition is that CoNorm allows normalization of intermediate feature maps using the conditioning input . The deep regression model uses CoNorm blocks to leverage contextual information for predicting multiple meshes from the input image . Similar to Batch Normalization , CoNorm learns to influence the output of the neural network by applying an affine transformation to the network’s intermediate features based on .
4 Multi-Person Losses
In multi-person scenarios, the regression model can often predict meshes that are intersecting and have incoherent depth ordering. Following , we adopt two multi-person losses – i) interpenetration and ii) depth-ordering loss, refer Fig. 4. We briefly describe the losses here for completeness but refer to for more details.
Interpenetration Loss. Let be the modified Signed Distance Field (SDF) over the 3D space. takes a positive value for all the points inside the 3D human mesh , proportional to the distance from the mesh surface and is everywhere else. We compute a separate distance field for each human mesh in the image . We define the pairwise interpenetration loss between mesh and mesh as follows,
is the sum of valid pairwise mesh collisions (Fig.4).
Depth-ordering Loss. We now define the depth-ordering loss . The key idea is to leverage the ground-truth instance segmentation maps available in the COCO datasets . We render all the meshes and the corresponding depth maps onto the image plane using a differentiable renderer and optimize the vertex locations based on the agreement with the ground-truth instance segmentation map of the image (Fig.4).
Finally, we train the network to minimize the loss where and are loss weights,
Experiments
OCHMR. For a fair comparison with other approaches , we use ResNet-50 as the default backbone for the mesh regression model and HRNet-W32 as the backbone for the context estimator . We insert CoNorm blocks after each of the ResNet block in the backbone. We set the CoNorm’s latent space dimensionality as for all our experiments. The input images are resized to , keeping the same aspect ratio and padding with zeros. Following , gaussians of size pixels is used to generate the local/global centermaps. The train-time data-augmentation, training schedule and all other hyper-parameters are set similar to . The loss weights are set to , , to ensure that the weighted loss items are of the same magnitude. The threshold of the local/global center heatmaps is set to .
Training Datasets. Similar to , we use MPI-INF-3DHP , COCO , MPII , LSP-Extended for training (we do not use Human3.6M due to licensing issues). Only the training sets are used, following the standard split protocols. We use ground-truth SMPL annotations from MPI-INF-3DHP and 2D annotations from COCO, MPII and LSP-Extended. The instance segmentation masks from COCO are used to compute .
Evaluation Benchmarks. 3DPW-PC is employed as the main benchmark for evaluating 3D mesh/joint error since it contains in-the-wild multi-person videos with abundant 2D/3D annotations. 3DPW-PC is the person-occluded subset of 3DPW . We also evaluate OCHMR under severe occlusion on Crowdpose and OCHuman which are crowded-in-the-wild 2D pose benchmarks. For completeness, we also benchmark our approach on the general datasets like 3DPW and COCO.
Evaluation Metrics. We report mean per joint position error (MPJPE), Procrustes-aligned MPJPE (PMPJPE) and per-vertex error (PVE) on the 3D datasets. MPJPE and PMPJPE evaluates the 3D joint rotation accuracy and PVE evaluates the 3D surface error. Also, to evaluate the pose accuracy under occlusion, we report standard metrics such as at various Object Keypoint Similarity . We also report results using bounding boxes obtained via Faster R-CNN detector.
2 Comparison to the State-of-the-Art
Occlusion benchmarks. To validate the stability under occlusion, we evaluate OCHMR on multiple occlusion benchmarks. Firstly, on the person-occluded 3DPW-PC, OCHuman and Crowdpose, results in Tab. 1 show that OCHMR significantly outperforms previous state-of-the-art methods . Additionally, in Fig. 5, we qualitatively demonstrate the robustness of OCHMR under severe occlusion in comparison to top-down SPIN and bottom-up ROMP . Further, when using ground-truth bounding boxes, the gains of OCHMR are significant in comparison to baselines. These results show that using high-resolution input images along with global/local centermaps is key for occlusion reasoning.
General benchmarks. We also compare OCHMR with other approaches on general benchmarks like 3DPW (Tab. 2) and COCO (Tab. 3). OCHMR undergoes no performance degradation on non-occlusion cases. Infact, OCHMR improves baseline SPIN’s MPJPE error by mm on 3DPW. Without using extra supervision, our method achieves comparable performance to ROMP with ResNet-50 backbone. We also outperform other methods on the COCO dataset.
3 Analysis
We perform all our analysis on the 3DPW-PC dataset with ground-truth boxes for evaluations.
CoNorm Block Architecture. We compare CoNorm blocks against early and late fusion in Tab. 4. In early fusion, we perform channel-wise concatenation of input image, global centermap and local centermap. In late fusion we concatenate the intermediate feature after the third ResNet block with the downsampled context information. We observe that injection of high-resolution context information at multiple-depths in the form of CoNorm blocks is important for accurate human mesh recovery under occlusion. Further, we vary the dimension of the latent space of the four CoNorm blocks in the OCHMR backbone. We show that increasing improves performance under occlusion in comparison to baseline SPIN, with achieving the optimal balance between the parameter overhead and human recovery performance.
Effect of Multi-Person losses. To understand the effect of multi-person losses like interpenetration loss and depth-ordering loss , we perform an ablative study using the loss weights in the OCHMR framework in Tab. 5. We achieve the best performance when using both losses, however the use of supervised loss gives better gains than self-supervised . Note, OCHMR still significantly outperforms baseline SPIN when only using .
Choice of Context. CoNorm blocks allow conditioning the network with various representations of the spatial-context . Tab. 6 shows the effect of using ground-truth and predicted (using F) Local and Local + Global Centermaps along with 2D keypoints as . We use offshelf pose-estimation network HRNet-W48 trained on COCO dataset as our . In case of 2D keypoints, is a -channel heatmap corresponding to keypoint locations. In comparison to local centermaps, the addition of global centermaps helps improve performance under occlusion. Interestingly, conditioning using ground-truth 2D keypoints outperforms all other choices. However, when ground-truth keypoints are unavailable, body centers outperform estimated 2D keypoints as estimating accurate 2D pose under occlusion is more challenging than estimating body centers.
Limitations. OCHMR is a multi-stage top-down method and is, therefore, not real-time during inference. Though OCHMR improves performance under multi-person occlusion, it is still susceptible to failure under truncation and extreme cropping due to object occlusion. Moreover, OCHMR fails to handle extreme poses and shapes due to the lack of training data, as shown in the Sup. Mat. In the future, OCHMR can be extended and incorporated with the recent progress to handle various kinds of occlusions .
Conclusion
Most top-down methods for human mesh recovery assume a single subject in the input, causing them to fail under severe person-person occlusion. In this work, we introduce OCHMR, a novel top-down method to handle multiple occluded people in crowded scenes. Our key idea is conditioning top-down models on spatial-context from the image, in the form of local and global centermaps which allows OCHMR to effectively disambiguate between overlapping humans. We propose Contextual Normalization (CoNorm) blocks, a novel architectural improvement which can be easily extended to any existing top-down method. While OCHMR draws inspiration from bottom-up methods, we retain the advantages of both bottom-up and top-down methods, resulting in a method that can handle multi-person occlusion and achieve pixel-aligned reconstruction results.