Encoder-decoder with Multi-level Attention for 3D Human Shape and Pose Estimation

Ziniu Wan, Zhengjia Li, Maoqing Tian, Jianbo Liu, Shuai Yi, Hongsheng Li

Introduction

3D human shape and pose estimation from a single image or video is a fundamental topic in computer vision. It is difficult to directly estimate the 3D human shape and pose from monocular images without any 3D information. To tackle this problem, massive 3D labeled data and 3D parametric human body models with prior knowledge are necessary. Tremendous works based on Deep Neural Network (DNN) have been made to increase the accuracy and robustness of this task.

However, existing DNN-based methods often fail in some challenging scenarios, including cluttered background, occlusion and extreme pose. To overcome these challenges, three intrinsic relations should be jointly modeled for the video-based 3D human shape and pose estimation: a). Spatial relation: For the pose estimation task, the human joints areas and the spatial correlations among body parts are directly related to the pose prediction. It is critical to carefully utilize the spatial relation, especially in the scene of cluttered background. b). Temporal relation: Everyone has particular temporal trajectory in a given video. In occlusion cases, this temporal relation should be exploited to infer the pose of current occluded frame from surrounding frames. c). Human joint relation: In the parametric 3D body model SMPL , human joints are organized as a kinematic tree. Once pose changes, the parent joint rotates first, and then rotates the children. When the pose amplitude is large, we argue that the prior of the dependence among joints is especially helpful for accurate pose estimation. However, none of the existing methods fully utilizes the above three relations in a unified framework.

Motivated by the above observations, we propose Multi-level Attention Encoder-Decoder Network (MAED) for video-based 3D human shape and pose estimation. MAED is the first work to explore the above three relations by exploiting corresponding multi-level attentions in a unified framework. It includes Spatial-Temporal Encoder (STE) for spatial-temporal attention and Kinematic Topology Decoder (KTD) for human joint attention.

Specifically, the STE consists of several cascaded blocks, and each block uses two parallel branches to learn spatial and temporal attention respectively. We call the two branches Multi-Head Self-Attention Spatial (MSA-S) and Multi-Head Self-Attention Temporal (MSA-T), which are inspired by Multi-Head Self-Attention (MSA) mechanism in Transformer related works . Derived from MSA, MSA-S and MSA-T have Transformer-like structures, but are different in the order of input features dimensions. As illustrated in Figure 1(a), MSA-S focuses on the critical spatial positions in image, highlighting significant features for pose estimation. Meantime, MSA-T concentrates on improving the prediction of current frame by exploiting frames that are informative to current one according to the calculated temporal attention scores.

On the other hand, existing methods usually use an iterative feedback regressor to regress the SMPL parameters, in which the pose parameters of all joints are generated simultaneously. However, they ignore the human joint relation. To exploit the dependence among joints, we further propose KTD to simulate the SMPL kinematic tree for joint level attention modeling. In KTD, each joint is assigned a unique linear regressor to regress its pose parameters. As shown in Figure 1(b), these parameters are generated through a top-down hierarchical regression process. To estimate a joint, besides image feature, we also take the predicted pose parameters of its ancestors as the input of linear regressor. In this manner, the bias of the parent joint’s estimation incurs substantial negative impact on the estimation of all its children, which forces the KTD to predict more accurate results for ancestor joints. In other words, although KTD does not explicitly allocate an attention score to each joint, the top-down regression process implicitly encourages the model to pay more attention to the parent joints with more children. As a result, the proposed KTD captures the inherent relation of joints and effectively reduce the prediction error.

We summarize the contributions of our method below:

We propose Multi-level Attention Encoder-Decoder Network (MAED) for video-based 3D human shape and pose estimation. Our proposed MAED contains Spatial-Temporal Encoder (STE) and Kinematic Topology Decoder (KTD). It learns different attentions at spatial level, temporal level and human joint level in a unified framework.

Our proposed STE leverages the MSA to construct MSA-S and MSA-T to encode the spatial and temporal attention respectively in the given video.

Our proposed KTD considers hierarchical dependence among human joints and implicitly captures human joint level attention.

Related Works

Recent works have made significant advances in 3D human pose and shape estimation due to the parametric 3D human body models, such as SMPL , SMPL-X and SCAPE , which utilize the statistics of human body and provide 3D mesh based on few hyper-parameters. Later, various studies focus on estimating the hyper-parameters of 3D human model directly from image or video input.

Previous parametric 3D human body model based methods are split into two categories: optimization-based methods and regression-based methods. The optimization-based methods fit the parametric 3D human body models to pseudo labels, like 2D keypoints, silhouettes and semantic mask. SMPLify , one of the first end-to-end optimization-based methods, uses strong statistics priors to guide the optimization supervised by 2D keypoints. The work utilizes silhouettes along with 2D keypoints to supervise the optimization. On the other hand, regression-based methods train deep neural network to regress the hyper-parameters directly. HMR is trained with the supervision of re-projection keypoints loss along with adversarial learning of human shape and pose. SPIN exploits SMPLify in the training loop to provide more supervision. VIBE is a video-based method that employs adversarial learning of the motions.

2 Transformer in Computer Vision

Transformer is first proposed in NLP field. It is an encoder-decoder model, completely replacing commonly used recurrent neural networks with Multi-Head Self-Attention mechanism, and later achieves great success in various NLP tasks . Motivated by the achievements of Transformer in NLP, various works start to apply Transformer to computer vision tasks. Vision Transformer (ViT) views an image as a 16x16 patch sequence, and trains a Transformer for image classification. The work explores distillation to use smaller datasets to get more efficient ViT. Some works study various Transformer structures which are more suitable for visual classification tasks. In addition, Transformer has also achieved impressive results in many downstream computer vision tasks, including denoising , object detection , video action recognition , 3D mesh reconstruction , panoptic segmentation , etc. In this paper, we focus on using Transformer to fully exploit the spatial-temporal level attention from video for better human pose and shape estimation.

Methods

In this section, we first revisit the parametric 3D human body model (SMPL ). Secondly, we give an overview of our proposed framework. Finally, we describe the proposed STE and KTD in detail.

2 Framework Overview

Figure 2 shows the architecture of our proposed network. It takes a video clip of length TT as input, and adopts a CNN backbone to extract the basic feature for each frame. The global pooling layer at the end of the CNN is omitted, resulting in TT feature maps of size (h×w×d)(h\times w\times d), where hh/ww/dd denotes the height/width/channel size of feature map. We reshape each feature map into 1D sequence of size (hw×d)(hw\times d), and prepend a trainable embedding to each sequence (Following , we denote a token in the sequence as a patch). Thus, the CNN outputs a matrix of size (T×N×d)(T\times N\times d), where N=hw+1N=hw+1. Then our proposed Spatial-Temporal Encoder (STE) is used to perform spatial-temporal modeling on these basic features. The encoded vector corresponding to the prepended embedding serves as the output of STE. Finally, our proposed Kinematic Topology Decoder (KTD) is employed to estimate shape β⃗\vec{\beta}, pose θ⃗\vec{\theta} and camera ϕ⃗\vec{\phi} parameters from the output of STE. These predicted parameters allow us to utilize SMPL to calculate 3D joints and their 2D projection, J2d=Πϕ⃗(J3d)J_{2d}=\Pi_{\vec{\phi}}\left(J_{3d}\right), where Πϕ⃗(.)\Pi_{\vec{\phi}}(.) is the projection function.

After getting {β⃗,θ⃗,J3d,J2d}\{\vec{\beta},\vec{\theta},J_{3d},J_{2d}\}, the model is supervised by the following 4 losses:

where L2DL_{2D}/L3DL_{3D} denotes the 2D/3D keypoint loss, LSMPLL_{SMPL} denotes the SMPL parameters loss, and LNORML_{NORM} denotes the L2-Normalization loss. J3dgt,J2dgt,θ⃗gt,β⃗gtJ_{3dgt},J_{2dgt},\vec{\theta}_{gt},\vec{\beta}_{gt} represent the ground truth of J3d,J2d,θ⃗,β⃗J_{3d},J_{2d},\vec{\theta},\vec{\beta} respectively.

3 Spatial-Temporal Encoder

Transformer is able to effectively model the interaction of tokens in a sequence. Recently, applying Transformer to model the temporal attention on global pooling feature of each frame is widely used in many video-based computer vision tasks. However, the global pooling operation will inevitably lose the spatial information in a frame, which makes it difficult to estimate detailed human pose. In our method, to perform spatial and temporal modeling simultaneously, we serialize the input video clip in multiple ways, and design three variants based on Multi-Head Self-Attention (MSA) : Multi-Head Spatial Self-Attention (MSA-S), Multi-Head Temporal Self-Attention (MSA-T) and Multi-Head Self-Attention Coupling (MSA-C). Then we further design three forms of Spatial-Temporal Encoder (STE) Block as shown in Figure 3, which endows the encoder with both global spatial perception and temporal reasoning capability. Finally, we stack multiple STE Blocks to construct the STE.

MSA Variants. The standard MSA can only learn attention of one dimension, so the different order of input dimensions affects the meaning of learned attention. Our proposed three variants have similar model structure, but are different on the order of the input dimensions.

MSA-S aims at finding the key spatial information in a frame, such as joints and limbs of human body. It is shown in the blue box in Figure 3(a), where each self-attention head outputs a heatmap of size (T×N×N)(T\times N\times N) computed by scaled dot-multiplication. However, in this setting, temporal relations among frames are not captured, as a patch in one frame does not interact with any patch in other frames.

MSA-T is pretty similar to MSA-S, except that it first reshapes the input matrix from size (T×N×d)(T\times N\times d) to (N×T×d)(N\times T\times d) as shown in the green box in Figure 3(b). Each head of MSA-T outputs the heatmap of size (N×T×T)(N\times T\times T) , where each score reflects the attention of a patch to the patch in the same position in other frames. Although temporal semantics is modeled explicitly, MSA-T ignores spatial relation of patches in the same frame.

MSA-C flattens patch sequence and frame sequence together, i.e., reshape the input matrix from size (T×N×d)(T\times N\times d) to (TN×d)(TN\times d), as shown in the yellow box in Figure 3(c). In this way, the heatmap of size (TN×TN)(TN\times TN) enables each patch interacts with any other patches in the video clip.

STE Blocks. As depicted in Figure 3, we design three kinds of STE Blocks based on these MSA variants. Coupling Block consists of a MSA-C followed by a Multi-Layer Perception (MLP) layer, modeling spatial-temporal information in a coupling fashion. However, it greatly increases the complexity since the complexity of dot-multiplication is quadratic to sequence length.

Connection of MSA-T and MSA-S makes it possible to combine image and video datasets to train more robust models. When it comes to image input, we simply bypass or disconnect the MSA-T in the blocks to ignore the non-existent temporal information.

Considering the trade-off between accuracy and speed, we empirically choose Parallel Block in our STE, as the Parallel Block is able to dynamically adjust the attentive weights between spatial and temporal attention and yields the best results compared with other variants. The quantitative comparison is discussed in Section 4.4.2 in detail.

4 Kinematic Topology Decoder

As aforementioned, previous works ignore the inherent dependence among joints and regard them as equally important. Therefore, we design Kinematic Topology Decoder (KTD) to implicitly model the attention at the joint level.

Therefore, the position of a joint is affected by its own and ancestral pose parameters. The more children a joint has, the greater its impact on the accuracy of the overall joint position estimation. Despite this, currently widely used iterative feedback regressor does not pay more attention to the parent joints, especially the root of kinematic tree. As a result, it can only get sub-optimal results. However, our proposed KTD can avoid the problem. In KTD, we first decode the shape/cam parameters with a matrix Wshape\mathbf{W}_{\text{shape}}/Wcam\mathbf{W}_{\text{cam}} as shown in Eq (4).

By KTD, we establish the dependence between the parent joint and its children, which is consistent with kinematic tree structure. In traditional regressor, the error of the parent joint’s pose estimation only affects itself. While in KTD, the error will be propagated to its children as well. This encourages the model to learn an attention at the joint level and pay more attention to parent joints, so as to achieve more accurate estimation results.

Experiments

Training. Following previous works , we use mixed datasets for training, including 3D video datasets, 2D video datasets and 2D image datasets. For 3D video datasets, Human3.6M and MPI-INF-3DHP provide 3D keypoints and SMPL parameters in indoor scene. For 2D video datasets, PennAction and PoseTrack provide ground-truth 2D keypoints annotation, while InstaVariaty provides pseudo 2D keypoints annotation using a keypoint detector . For image-based datasets, COCO , MPII and LSP-Extended are adopted, providing in-the-wild 2D keypoints annotation. Meanwhile, we conduct ablation study on the 3DPW dataset.

Evaluation. We report the experiments results on Human3.6M , MPI-INF-3DHP and 3DPW evaluation set. We adopt the widely used evaluation metrics following previous works , including Procrustes-Aligned Mean Per Joint Position Error (PA-MPJPE), Mean Per Joint Position Error (MPJPE), Per Vertex Error (PVE) and ACCELeration error (ACCEL). We report the results with and without 3DPW training set for fair comparison with previous methods.

2 Training Details

Data Augmentation. Horizontal flipping, random cropping, random erasing and color jittering are employed to augment the training samples. Different frames of the same video input share consistent augmentation parameters.

Model Details. Following , we use a modified ResNet-50 as the CNN backbone to extract the basic feature of an input image. For STE, 6 STE Parallel Blocks are stacked, and each block has 12 heads. We adopt the weights from to initialize the ResNet-50 and STE.

The whole training process is divided into two stages. In the first stage, the model aims at accumulating sufficient spatial prior knowledge, and thus is trained with all image-based datasets and frames from Human3.6M and MPI-INF-3DHP. We fix the number of epochs as 100 and the mini-batch size as 512 for this stage. In the second stage, we use both video and image datasets for temporal modeling. For video datasets, we sample 16-frame clips at a interval of 8 as training instances. We train another 100 epochs with a mini-batch size of 32 for this stage. The model is optimized by Adam optimizer with an initial learning rate of 10−410^{-4} which is decreased by 10 at the 60-th and 90-th epochs. Finally, each term in the loss function has different weighting coefficients. Refer to Sup. Mat. for further details. All experiments are conducted on 16 Nvidia GTX1080ti GPUs.

3 Comparison to state-of-the-art results

In this section, we compare our method with the state-of-the-art models on 3DPW, MPI-INF-3DHP and Human3.6M, and the results are summarized in Table 1. On the 3DPW and MPI-INF-3DHP datasets, our method outperforms other competitors including image- and video-based methods by a large margin, whether or not using 3DPW training set. On Human3.6M, our method achieves results on-par with I2LMeshNet . We also observe MEVA , an two-stage method that aims at producing both smooth and accurate results, ranks best in ACCEL metric on 3DPW. However, considering all indicators, our method achieves better performance overall.

These results validate our hypothesis that the exploitation of the attentions at spatial-temporal level and human joint level greatly helps to achieve more accurate estimation. The leading performance on these three datasets (especially the in-the-wild dataset 3DPW) demonstrates the robustness and the potential to real-world applications of our method.

4 Ablation Study

The upper part of Table 2 verifies the effectiveness of our proposed STE and KTD. Compared with CNN encoder+Iterative decoder, STE and KTD brings 4.7 and 1.3 mm improvement in PA-MPJPE metric respectively. Moreover, STE and KTD together further improves the performance by 6.5 mm. This proves the attention at different levels extracted by STE and KTD effectively complement rather than conflict each other.

We can also observe that when using CNN encoder, the gain of KTD in PA-MPJPE metric is smaller than that when using CNN+STE encoder. Even there is a small decline in MPJPE metric. This is because the CNN loses too much spatial information due to the global pooling operation, and fails to provide detailed human body clue for KTD. However, with hard downsampling removed, STE not only preserves more spatial information, but also pay more attention to more informative locations, which makes KTD capture more precise attention between joints.

4.2 Influence of different encoders

In the middle part of Table 2, we compare the performance of various forms of STE. SE denotes the encoder with only MSA-S. TE denotes the encoder with only MSA-T and CNN global pooling layer kept. STEparallel{}_{\text{parallel}}v1 and STEparallel{}_{\text{parallel}}v2 denote the Parallel Block w/o and w/ attentive addition respectively. We conclude that all the variants of STE benefit the model, while STEparallel{}_{\text{parallel}}v2 yields the most significant gain. This is because the attentive weights dynamically computed in the Parallel Block effectively act as a valve which adjusts the proportion of temporal and spatial information passing through the network. When it comes to occlusion or ambiguity, the valve will allow more temporal information to pass through to complement the lack of information in current frame, and do otherwise when the current frame is clear. Surprisingly, STEcoupling{}_{\text{coupling}} yields only modest improvement over encoder with only MSA-S (49.8→\rightarrow49.3), which has no temporal modeling capability. We also observe that STEcoupling{}_{\text{coupling}} converges more slowly compared to other STE variants. We argue that flattening the spatial and temporal dimension together may harm human pose estimation mainly due to the extremely long sequence. Tremendous irrelevant patches (such as background and joints that are too far apart) overwhelm valid information, making it challenging for the current patch to allocate reasonable attention.

4.3 Influence of different decoders

We choose CNN+STE as the encoder and report the results with different decoders in the lower part of Table 2. KTDrandom{}_{\text{random}} denotes the KTD on a randomly generated kinematic tree. KTDreverse{}_{\text{reverse}} denotes the KTD on the reverse kinematic tree, namely, exchange the relationship between parent joint and its children. Decodervanilla{}_{\text{vanilla}} denotes the standard decoder in with 6 layers. It takes as input the zero sequence of length 37 (24 for pose, 10 for shape and 3 for camera) and outputs SMPL parameters. We observe that KTD outperforms Iterative by a large margin. While KTDrandom{}_{\text{random}} and KTDreverse{}_{\text{reverse}} have no obvious improvement, even are slightly worse, proving unreasonable kinematic tree is useless prior knowledge, which brings difficulties to the optimization of the network. We also observe that Decodervanilla{}_{\text{vanilla}} brings no improvement. Although it can capture the relation between different joints with the self-attention mechanism, the predictions of all joints are generated simultaneously, not in the sequential way as KTD. As a result, it can not pay more attention to the parent joints.

5 Visualization Analysis

Figure 5 includes qualitative results of MAED from two representative scenarios. For these challenging cases including extreme pose in Figure 5(a) and cluttered background and occlusion in Figure 5(b), our model predicts reasonable spatial and temporal attention maps and further produce proper estimations.

Conclusion

This paper describes MAED, an approach that utilizes multi-level attentions at spatial-temporal level and human joint level for 3D human shape and pose estimation. We design multiple variants of MSA and STE Block to construct STE to learn spatial-temporal attention from the output feature of CNN backbone. In addition, we propose KTD, which simulates the process of joint rotation based on SMPL kinematic tree to decode human pose. MAED makes significant accuracy improvement on multiple datasets but also brings non-negligible computation overhead, which we explore further in the Sup. Mat. Thus, future work could consider reducing computation overhead or extending this method to capture the relation between multiple people.

References