Self-Supervised 3D Human Pose Estimation via Part Guided Novel Image Synthesis
Jogendra Nath Kundu, Siddharth Seth, Varun Jampani, Mugalodi Rakesh, R. Venkatesh Babu, Anirban Chakraborty
Introduction
Analyzing humans takes a central role in computer vision systems. Automatic estimation of 3D pose and 2D part-arrangements of highly deformable humans from monocular RGB images remains an important, challenging and unsolved problem. This ill-posed classical inverse problem has diverse applications in human-robot interaction , augmented reality , gaming industry, etc.
In a fully-supervised setting , the advances in this area are mostly driven by recent deep learning architectures and the collection of large-scale annotated samples. However, unlike 2D landmark annotations, it is very difficult to manually annotate 3D human pose on 2D images. A usual way of obtaining 3D ground-truth (GT) pose annotations is through a well-calibrated in-studio multi-camera setup , which is difficult to configure in outdoor environments. This results in a limited diversity in the available 3D pose datasets, which greatly limits the generalization of supervised 3D pose estimation models.
To facilitate better generalization, several recent works leverage weakly-supervised learning techniques that reduce the need for 3D GT pose annotations. Most of these works use an auxiliary task such as multi-view 2D pose estimation to train a 3D pose estimator . Instead of using 3D pose GT for supervision, a 3D pose network is supervised with loss functions on multi-view projected 2D poses. To this end, several of these works still require considerable annotations in terms of paired 2D pose GT , multi-view images and known camera parameters . Dataset bias still remains a challenge in these techniques as they use paired image and 2D pose GT datasets which have limited diversity. Given the ever-changing human fashion and evolving culture, the visual appearance of humans keeps varying and we need to keep updating the 2D pose datasets accordingly.
In this work, we propose a differentiable and modular self-supervised learning framework for monocular 3D human pose estimation along with the discovery of 2D part segments. Specifically, our encoder network takes an image as input and outputs 3 disentangled representations: 1. view-invariant 3D human pose in canonical co-ordinate system, 2. camera parameters and 3. a latent code representing foreground (FG) human appearance. Then, a decoder network takes the above encoded representations, projects them onto 2D and synthesizes FG human image while also producing 2D part segmentation. Here, a major challenge is to disentangle the representations for 3D pose, camera, and appearance. We achieve this disentanglement by training on video frame pairs depicting the same person, but in varied poses. We self-supervise our network with consistency constraints across different network outputs and across image pairs. Compared to recent self-supervised approaches that either rely on videos with static background or work with the assumption that temporally close frames have similar background , our framework is robust enough to learn from large-scale in-the-wild videos, even in the presence of camera movements. We also leverage the prior knowledge on human skeleton and poses in the form of a single part-based 2D puppet model, human pose articulation constraints, and a set of unpaired 3D poses. Fig. 1 illustrates the overview of our self-supervised learning framework.
Self-supervised learning from in-the-wild videos is challenging due to diversity in human poses and backgrounds in a given pair of frames which may be further complicated due to missing body parts. We achieve the ability to learn on these wild video frames with a pose-anchored deformation of puppet model that bridges the representation gap between the 3D pose and the 2D part maps in a fully differentiable manner. In addition, the part-conditioned appearance decoding allows us to reconstruct only the FG human appearance resulting in robustness to changing backgrounds.
Another distinguishing factor of our technique is the use of well-established pose prior constraints. In our self-supervised framework, we explicitly model 3D rigid and non-rigid pose transformations by adopting a differentiable parent-relative local limb kinematic model, thereby reducing ambiguities in the learned representations. In addition, for the predicted poses to follow the real-world pose distribution, we make use of an unpaired 3D pose dataset. We interchangeably use predicted 3D pose representations and sampled real 3D poses during training to guide the model towards a plausible 3D pose distribution.
Our network also produces useful part segmentations. With the learned 3D pose and camera representations, we model depth-aware inter-part occlusions resulting in robust part segmentation. To further improve the segmentation beyond what is estimated with pose cues, we use a novel differentiable shape uncertainty map that enables extraction of limb shapes from the FG appearance representation.
We make the following main contributions:
We present techniques to explicitly constrain the 3D pose by modeling it at its most fundamental form of rigid and non-rigid transformations. This results in interpretable 3D pose predictions, even in the absence of any auxiliary 3D cues such as multi-view or depth.
We propose a differentiable part-based representation which enables us to selectively attend to foreground human appearance which in-turn makes it possible to learn on in-the-wild videos with changing backgrounds in a self-supervised manner.
We demonstrate generalizability of our self-supervised framework on unseen in-the-wild datasets, such as LSP and YouTube. Moreover, we achieve state-of-the-art weakly-supervised 3D pose estimation performance on both Human3.6M and MPI-INF-3DHP datasets against the existing approaches.
Related Works
Human 3D pose estimation is a well established problem in computer vision, specifically in fully supervised paradigm. Earlier approaches proposed to infer the underlying graphical model for articulated pose estimation. However, the recent CNN based approaches focus on regressing spatial keypoint heat-maps, without explicitly accounting for the underlying limb connectivity information. However, the performance of such models heavily relies on a large set of paired 2D or 3D pose annotations. As a different approach, proposed to regress latent representation of a trained 3D pose autoencoder to indirectly endorse a plausibility bound on the output predictions. Recently, several weakly supervised approaches utilize varied set of auxiliary supervision other than the direct 3D pose supervision (see Table 1). In this paper, we address a more challenging scenario where we consider access to only a set of unaligned 2D pose data to facilitate the learning of a plausible 2D pose prior (see Table 1).
In literature, while several supervised shape and appearance disentangling techniques exist, the available unsupervised pose estimation works (i.e. in the absence of multi-view or camera extrinsic supervision), are mostly limited to 2D landmark estimation for rigid or mildly deformable structures, such as facial landmark detection, constrained torso pose recovery etc. The general idea is to utilize the relative transformation between a pair of images depicting a consistent appearance with varied pose. Such image pairs are usually sampled from a video satisfying appearance invariance or synthetically generated deformations .
Beyond landmarks, object parts can infer shape alongside the pose. Part representations are best suited for 3D articulated objects as a result of its occlusion-aware property as opposed to simple landmarks. In general, the available unsupervised part learning techniques are mostly limited to segmentation based discriminative tasks. On the other hand, explicitly leverage the consistency between geometry and the semantic part segments. However, the kinematic articulation constraints are well defined in 3D rather than in 2D . Motivated by this, we aim to leverage the advantages of both non-spatial 3D pose and spatial part-based representation by proposing a novel 2D pose-anchored part deformation model.
Approach
We develop a differentiable framework for self-supervised disentanglement of 3D pose and foreground appearance from in-the-wild video frames of human activity.
Our self-supervised framework builds on the conventional encoder-decoder architecture (Sec. 3.2). Here, the encoder produces a set of local 3D vectors from an input RGB image. This is then processed through a series of 3D transformations, adhering to the 3D pose articulation constraints to obtain a set of 2D coordinates (camera projected, non-spatial 2D pose). In Sec. 3.1, we define a set of part based representations followed by carefully designed differentiable transformations required to bridge the representation gap between the non-spatial 2D pose and the spatial part maps. This serves three important purposes. First, their spatial nature facilitates compatible input pose conditioning for the fully-convolutional decoder architecture. Second, it enables the decoder to selectively synthesize only FG human appearance ignoring the variations in the background. Third, it facilitates a novel way to encode the 2D joint and part association using a single template puppet model. Finally, Sec. 3.3 describes the proposed self-supervised paradigm which makes use of the pose-aware spatial part maps for simultaneous discovery of 3D pose and part segmentation using image pairs from wild videos.
One of the major challenges in unsupervised pose or landmark detection is to map the model-discovered landmarks to the standard landmark conventions. This is essential to facilitate the subsequent task-specific pipelines, which expect the input pose to follow a certain convention. Prior works rely on paired supervision to learn this mapping. In absence of such supervision, we aim to encode this convention in a canonical part dictionary where the association of 2D joints with respect to the body parts is extracted from a single manually annotated puppet template (Fig. 2C, top panel). This can be interpreted as a 2D human puppet model, which can approximate any human pose deformation via independent spatial transformation of body parts while keeping intact the anchored joint associations.
a) shape uncertainty map as , and
b) single-channel FG-BG map as .
The above formalization bridges the representation gap between the raw joint locations, and the output spatial maps , , and , thereby facilitating them to be used as differentiable spatial maps for the subsequent self-supervised learning.
Depth-aware part segmentation. For 3D deformable objects, a reliable 2D part segmentation can be obtained with the help of following attributes, i.e. a) 2D skeletal pose, b) part-shape information, and c) knowledge of inter-part occlusion. Here, the 2D skeletal pose and the knowledge of inter-part occlusion can be extracted by accessing camera transformation of the corresponding 3D pose representation. Let, the depth of the 2D joints in with respect to the camera be denoted as and . We obtain a scalar depth value associated with each limb as, . We use these depth values to alter the strength of depth-unaware part-wise pose maps, at each spatial location, by modulating the strength of part map intensity as being inversely proportional to the depth values. This is realized in the following steps:
a) ,
b) , and
c) .
Here, indicates the spatial-channel dedicated for the background. Additionally, a non-differentiable 2D part-segmentation map (see Fig. 2E) is obtained as,
2 Self-supervised pose network
The architecture for self-supervised pose and appearance disentanglement consists of a series of pre-defined differentiable transformations facilitating discovery of a constrained latent pose representation. As opposed to imposing learning based constraints , we devise a way around where the 3D pose articulation constraints (i.e. knowledge of joint connectivity and bone-length) are directly applied via structural means, implying guaranteed constraint imposition.
As compared to spatial 2D geometry , discovering the inherent 3D human pose is a highly challenging task considering the extent of associated non-rigid deformation, and rigid camera variations . To this end, we define a canonical coordinate system , where face-vector of the skeleton is canonically aligned along the +ve X-axis, thus making it completely view-invariant. Here, the face-vector is defined as the perpendicular direction of the plane spanning the neck, left-hip and right-hip joints. As shown in Fig. 2B, in , except the pelvis, neck, left-hip and right-hip, all other joints are defined at their respective parent relative local coordinate systems (i.e. parent joint as the origin with axis directions obtained by performing Gram-Schmidt orthogonalization of the parent-limb vector and the face-vector). Accordingly, we define a recursive forward kinematic transformation to obtain the canonical 3D pose from the local limb vectors, i.e. , which accesses a constant array of limb length magnitudes .
Here, the camera extrinsics, (3 rotation angles and 3 restricted translations ensuring that the camera-view captures all the skeleton joints in ) is obtained at the encoder output, whereas a fixed perspective camera projection is applied to obtain the final 2D pose representation, i.e. . A part based deformation operation on this 2D pose (Sec. 3.1) is shown as , where and (following the depth aware operations on ). Finally, denotes the entire series of differentiable transformations, i.e. , as shown in Fig. 2A. Here, denotes composition operation.
b) Decoder network. The decoder takes a concatenated representation of the FG appearance, and pose, as input to obtain two output maps, i) a reconstructed image , and ii) a predicted part segmentation map via a bifurcated CNN decoder (see Fig. 3A). The common decoder branch, consists of a series of up-convolutional layers conditioned on the spatial pose map at each layer’s input (i.e. multi-scale pose conditioning). Whereas, and follow up-convolutional layers to their respective outputs.
3 Self-supervised training objectives
The prime design principle of our self-supervised framework is to leverage the interplay between the pose and appearance information by forming paired input images of either consistent pose or appearance.
Given a pair of source and target image, , sampled from the same video i.e. with consistent FG appearance, the shared encoder extracts their respective pose and appearance as and (see Fig. 3A). We denote the decoder outputs as while the decoder takes in pose, with appearance, (this notation is consistent in later sections). Here, is expected to depict the person in the target pose . Such a cross-pose transfer setup is essential to restrict leakage of pose information through the appearance. Whereas, the low dimensional bottleneck of followed by the series of differentiable transformations prevents leakage of appearance through pose .
To effectively operate on wild video frames (i.e. beyond the in-studio fixed camera setup ), we aim to utilize the pose-aware, spatial part representations as a means to disentangle the FG from BG. Thus, we plan to reconstruct with a constant BG color and segmented FG appearance (see Fig. 3A). Our idea stems from the concept of co-saliency detection , where the prime goal is to discover the common or salient FG from a given set of two or more images. Here, the part appearances belonging to the model predicted part regions have to be consistent across and for a successful self-supervised pose discovery.
Access to unpaired 3D/2D pose samples. We denote and as a 3D pose and its projection (via random camera), sampled from an unpaired 3D pose dataset , respectively. Such samples can be easily collected without worrying about BG or FG diversity in the corresponding camera feed (i.e. a single person activity). We use these samples to further constrain the latent space towards realizing a plausible 3D pose distribution.
a) Image reconstruction objective. Unlike , we do not have access to the corresponding ground-truth representation for the predicted , giving rise to an increased possibility of producing degenerate solutions or mode-collapse. In such cases, the model focuses on fixed background regions as the common region between the two images, specifically for in-studio datasets with limited BG variations. One way to avoid such scenarios is to select image pairs with completely diverse BG (i.e. select image pairs with high L2 distance in a sample video-clip).
To explicitly restrict the model from inculcating such BG bias, we incorporate content-based regularization. An uncertain pseudo FG mask is used to establish a one-to-one correspondence between the reconstructed image and the target image . This is realized through a spatial-mask, which highlights the salient regions (applicable for any image frame) or regions with diverse motion cues (applicable for frames captured in a fixed camera). We formulate a pseudo (uncertain) reconstruction objective as,
Here, denotes pixel-wise weighing and is a balancing hyperparameter. Note that, the final loss is computed as average over all the spatial locations, . This loss enforces a self-supervised consistency between the pose-aware part maps, and the salient common FG to facilitate a reliable pose estimate.
As a novel direction, we utilize to form a pair of image predictions following simultaneous appearance invariance and pose equivariance (see Fig. 3B). Here, a certain reconstruction objective is defined as,
b) Part-segmentation objective. Aiming to form a consistency between the true pose and the corresponding part segmentation output, we formulate,
Here, CE, and SE denote the pixel-wise cross-entropy and self-entropy respectively. Moreover, we confidently enforce segmentation loss with respect to one-hot map (Sec. 3.1) only at the certain regions, while minimizing the Shanon’s entropy for the regions associated with shape uncertainty as captured in . Here, the limb depth required to compute is obtained from (Fig. 3C).
In summary, the above self-supervised objectives form a consistency among , , and ;
a) enforces consistency between and ,
b) enforces consistency between (via ) and ,
c) enforces consistency between (via ) and .
However, the model inculcates a discrepancy between the predicted pose and the true pose distributions. It is essential to bridge this discrepancy as and rely on true pose , whereas relies on the predicted pose . Thus, we employ an adaptation strategy to guide the model towards realizing a plausible pose prediction.
c) Adaptation via energy minimization. Instead of employing an ad-hoc adversarial discriminator , we devise a simpler yet effective decoupled energy minimization strategy . We avoid a direct encoder-decoder interaction during gradient back-propagation, by updating the encoder parameters, while freezing the decoder parameters and vice-versa. However, this is performed while enforcing a reconstruction loss at the output of the secondary encoder in a cyclic auto-encoding scenario (see Fig. 3C). The two energy functions are formulated as and , where and .
The decoder parameters are updated to realize a faithful , as the frozen encoder expects to match its input distribution of real images (i.e. ) for an effective energy minimization. Here, the encoder can be perceived as a frozen energy network as used in energy-based GAN . A similar analogy applies while updating the encoder parameters with gradients from the frozen decoder. Each alternate energy minimization step is preceded by an overall optimization of the above consistency objectives, where both encoder and decoder parameters are updated simultaneously (see Algo. 1).
Experiments
We perform a thorough experimental analysis to establish the effectiveness of our proposed framework on 3D pose estimation, part segmentation and novel image synthesis tasks, across several datasets beyond the in-studio setup.
Implementation details. We employ an ImageNet trained Resnet-50 architecture as the base CNN for the encoder . We first bifurcate it into two CNN branches dedicated to pose and appearance, then the pose branch is further bifurcated into two multi-layer fully-connected networks to obtain the local pose vectors and the camera parameters . While training, we use separate AdaGrad optimizers for each loss term at alternate training iterations. We perform appearance (color-jittering) and pose augmentations (mirror flip and inplane rotation) selectively for and conceding their invariance effect on and respectively.
Datasets. We train the base-model on image pairs sampled from a mixed set of video datasets i.e. Human3.6M (H3.6M) and an in-house collection of in-the-wild YouTube videos (YTube). As opposed to the in-studio H3.6M images, the YTube dataset constitutes a substantial diversity in apparels, action categories (dance forms, parkour stunts, etc.), background variations, and camera movements. The raw video frames are pruned to form the suitable image pairs after passing them through an off-the-shelf person-detector . We utilize an unsupervised saliency detection method to obtain for the wild YTube frames, whereas for samples from H3.6M is obtained directly through the BG estimate . Further, LSP and MPI-INF-3DHP (3DHP) datasets are used to evaluate generalizability of our framework. We collect the unpaired 3D pose samples, from MADS and CMU-MoCap dataset keeping a clear domain gap with respect to the standard datasets chosen for benchmarking our performance.
Inline with the prior arts , we evaluate our 3D pose estimation performance in the standard protocol-II setting (i.e. with scaling and rigid alignment). We experimented on 4 different variants of the proposed framework with increasing degrees of supervision levels. The base model in absence of any paired supervision is regarded as Ours(unsup). In presence of multi-view information (with camera extrinsics), we finetune the model by enforcing consistent canonical pose and camera shift for multi-view image pairs, termed as Ours(multi-view-sup). Similarly, finetuning in presence of a direct supervision on the corresponding 2D pose GT is regarded as Ours(weakly-sup). Lastly, finetuning in presence of a direct 3D pose supervision on 10% of the full training set is referred to as Ours(semi-sup). Table 2 depicts our superior performance against the prior arts in their respective supervision levels.
2 Evaluation on MPI-INF-3DHP
With our framework, we also demonstrate a higher level of cross-dataset generalization, thereby minimizing the need for finetuning on novel unseen datasets. The carefully devised constraints at the intermediate pose representation are expected to restrict the model from producing implausible poses even when tested in unseen environments. To evaluate this, we directly pass samples of MPI-INF-3DHP (3DHP) test-set through Ours(weakly-sup) model trained on YTube+H3.6M and refer it as unsupervised transfer, denoted as -3DHP in Table 4. Further, we finetune the Ours(weakly-sup) on 3DHP dataset at three supervision levels, a) no supervision, b) full 2D pose supervision, and c) 10% 3D pose supervision as reported in Table 4, at rows 10, 7, and 11 respectively. The reported metrics clearly highlight our superiority against the prior arts.
3 Ablation study
We evaluate the effectiveness of the proposed local vector representation followed by the forward kinematic transformations, against a direct estimation of the 3D joints in camera coordinate system in the presence of appropriate bone length constraints . As reported in Table 3, our disentanglement of camera from the view-invariant canonical pose shows a clear superiority as a result of using the 3D pose articulation constraints in the most fundamental form. Besides this, we also perform ablations by removing or from the unsupervised training pipeline. As shown in Fig. 5B, without the model predicts implausible part arrangements even while maintaining a roughly consistent FG silhouette segmentation. However, without , the model renders a plausible pose on the BG area common between the image pairs, as a degenerate solution.
As an ablation of the semi-supervised setting, (see Fig. 4), we train the proposed framework on progressively increasing amount of 3D pose supervision alongside the unsupervised learning objective. Further, we perform the same for the Encoder network without the unsupervised objectives (thus, discarding the decoder networks) and term it as Encoder(Only-3D-sup). The plots in Fig. 4 clearly highlight our reduced dependency on the supervised data implying graceful and faster transferability.
4 Evaluation of part segmentation
For evaluation of part-segmentation, we standardize the ground-truth part conventions across both LSP and H3.6M datasets via SMPL model fitting . This convention roughly aligns with the part-puppet model used in Fig. 2C, thereby maintaining a consistent part to joint association. Note that, is supposed to account for the ambiguity between the puppet-based shape-unaware segmentation against the image dependent shape-aware segmentation. In Fig. 5A, we show the effectiveness of this design choice, where the shape-unaware segmentation is obtained at after depth-based part ordering, and the corresponding shape-aware segmentation is obtained at output. Further, quantitative comparison of part segmentation is reported in Table 5. We achieve comparable results against the prior arts, in absence of additional supervision.
5 Qualitative results
To evaluate the effectiveness of the disentangled factors beyond the intended primary task of 3D pose estimation and part segmentation, we manipulate them to analyze their effect on the decoder synthesized output image. In pose-transfer, pose obtained from an image is transferred to the appearance of another. However, in view-syn., we randomly vary the camera extrinsic values in . The results shown in Fig. 5A are obtained from Ours(unsup) model, which is trained on the mixed YTube+H3.6 dataset. This demonstrates the clear disentanglement of pose and appearance. Fig. 6 depicts qualitative results for the primary 3D pose estimation and part segmentation tasks using Ours(weakly-sup) model as introduced in Sec. 4.1. In Fig. 6B, we show results on the unseen LSP dataset, where the model has not seen this dataset even during self-supervised training. A consistent performance on such unseen dataset further establishes generalizability of the proposed framework.
Conclusion
We proposed a self-supervised 3D pose estimation method that disentangles the inherent factors of variations via part guided human image synthesis. Our framework has two prominent traits. First, effective imposition of both human 3D pose articulation and joint-part association constraint via structural means. Second, usage of depth-aware part based representation to specifically attend to the FG human resulting in robustness to changing backgrounds. However, extending such a framework for multi-person or partially visible human scenarios remains an open challenge.
Acknowledgements. This project is supported by a Indo-UK Joint Project (DST/INT/UK/P-179/2017), DST, Govt. of India and a Wipro PhD Fellowship (Jogendra).