MHFormer: Multi-Hypothesis Transformer for 3D Human Pose Estimation

Wenhao Li, Hong Liu, Hao Tang, Pichao Wang, Luc Van Gool

Introduction

3D human pose estimation (HPE) from monocular videos is a fundamental vision task with a wide range of applications, such as action recognition , human-computer interaction , and augmented/virtual reality . This task is typically solved by dividing it into two decoupled subtasks, i.e., 2D pose detection to localize the keypoints on the image plane, followed by 2D-to-3D lifting to infer joint locations in the 3D space from 2D keypoints. Despite their impressive performance , it remains an inherently ill-posed problem because of self-occlusion and depth ambiguity in 2D representations.

To alleviate such issues, most methods focus on exploring spatial and temporal relationships. They either employ graph convolutional networks to estimate 3D poses with a spatio-temporal graph representation of human skeletons or apply a pure Transformer-based model to capture spatial and temporal information from 2D pose sequences . Yet, the 2D-to-3D lifting from monocular videos is an inverse problem where multiple feasible solutions (i.e., hypotheses) exist due to its ill-posed nature given the missing depth . Those approaches ignore this problem and only estimate a single solution, which often leads to unsatisfactory results, especially when the person is severely occluded (see Figure 1).

Recently, a couple of methods that generate multiple hypotheses have been proposed for the inverse problem. They often rely on the one-to-many mapping by adding multiple output heads to an existing architecture with a shared feature extractor, while failing to build the relationships among the features of different hypotheses. That is an important shortcoming, as such ability is vital to improve the expressiveness and the performance of the model. In view of the ambiguous inverse problem of 3D HPE, we argue that it is more reasonable to conduct a one-to-many mapping first and then a many-to-one mapping with various intermediate hypotheses, as this way can enrich the diversity of features and produce a better synthesis for the final 3D pose.

To this end, we present Multi-Hypothesis Transformer (MHFormer), a novel Transformer-based method for 3D HPE from monocular videos. The key insight is to allow the model to learn spatio-temporal representations of diverse pose hypotheses. To accomplish this, we introduce a three-stage framework that starts from generating multiple initial representations and gradually communicates across them to synthesize a more accurate prediction, as shown in Figure 2. This framework more effectively models multi-hypothesis dependencies while also building stronger relationships among hypothesis features. Specifically, in the first stage, a Multi-Hypothesis Generation (MHG) module is built to model the intrinsic structure information of human joints and generate several multi-level features in the spatial domain. Those features contain diverse semantic information in different depths from shallow to deep and hence can be regarded as initial representations of multiple hypotheses.

Next, we propose two novel modules to model temporal consistencies and enhance those coarse representations in the temporal domain, which have not been explored by the existing works that generate multiple hypotheses. In the second stage, a Self-Hypothesis Refinement (SHR) module is proposed to refine every single-hypothesis feature. The SHR consists of two new blocks. The first block is a multi-hypothesis self-attention (MH-SA) which models single-hypothesis dependencies independently to construct self-hypothesis communication, enabling message passing within each hypothesis for feature enhancement. The second block is a hypothesis-mixing multi-layer perceptron (MLP) that exchanges information across hypotheses. The multiple hypotheses are merged into a single converged representation, and then this representation is partitioned into several diverged hypotheses.

Although those hypotheses are refined by SHR, the connections across different hypotheses are not strong enough since the MH-SA in the SHR only passes intra-hypothesis information. To address this issue, in the last stage, a Cross-Hypothesis Interaction (CHI) module models interactions among multi-hypothesis features. Its key component is the multi-hypothesis cross-attention (MH-CA), which captures mutual multi-hypothesis correlations to build cross-hypothesis communication, enabling message passing among hypotheses for better interaction modeling. Subsequently, a hypothesis-mixing MLP is used to aggregate the multiple hypotheses to synthesize the final prediction.

With the proposed MHFormer, multi-hypothesis spatio-temporal feature hierarchies are explicitly incorporated into Transformer models, where the multiple hypothesis information of body joints can be independently and mutually processed in an end-to-end manner. As a result, the representation ability is potentially enhanced and the synthesized pose is much more accurate. Our contributions are summarized as follows:

We present a new Transformer-based method, called Multi-Hypothesis Transformer (MHFormer), for 3D HPE from monocular videos. MHFormer can effectively learn spatio-temporal representations of multiple pose hypotheses in an end-to-end manner.

We propose to communicate among multi-hypothesis features both independently and mutually, providing powerful self-hypothesis and cross-hypothesis message passing, and strong relationships among hypotheses.

Our MHFormer achieves state-of-the-art performance on two challenging datasets for 3D HPE, significantly outperforming PoseFormer by 3% with 1.3 mmmm error reduction on Human3.6M .

Related Work

3D Human Pose Estimation. Existing single-view 3D pose estimation methods can be divided into two mainstream types: one-stage approaches and two-stage ones. One-stage approaches directly infer 3D poses from input images without intermediate 2D pose representations , while two-stage ones first obtain 2D keypoints from pretrained 2D pose detections and then feed them into a 2D-to-3D lifting network to estimate 3D poses. Benefiting from the excellent performance of 2D human pose estimation, this 2D-to-3D pose lifting method can efficiently and accurately regress 3D poses using detected 2D keypoints. For instance, SimpleBaseline proposes a fully-connected residual network to lift 2D keypoints to 3D joint locations from a single frame. Anatomy3D decomposes the task into bone direction and bone length predictions to ensure temporal consistency over a sequence. Despite the promising results achieved by using temporal correlations from fully convolutional or graph-based architectures, these methods are less efficient in capturing global-context information across frames.

Vision Transformers. Recently, Transformer equipped with the powerfully global self-attention mechanism has received increasingly research interest in the computer vision community . For the basic image classification task, ViT is proposed to apply a standard Transformer architecture directly to sequential image patches. For the pose estimation task, PoseFormer applies a pure Transformer to capture human joint correlations and temporal dependencies. Strided Transformer introduces a Transformer-based architecture with strided convolutions to lift a long 2D pose sequence to a single 3D pose. Our work is inspired by them and similarly uses the Transformer as the basic architecture. But we do not just utilize a simple architecture with a single representation; instead, the seminal ideas of multi-hypothesis and multi-level feature hierarchies are connected within Transformers, which makes the model not only expressive but also strong. Besides, a cross-attention mechanism is introduced for effective multi-hypothesis learning.

Multi-Hypothesis Methods. Single-view 3D HPE is ill-posed and therefore assuming only a single solution might be sub-optimal. Several works generate diverse hypotheses for the inverse problem and achieve substantial performance gains . For example, Jahangiri et al. generated multiple 3D pose candidates consistent with 2D keypoints via a compositional model and anatomical constraints. Wehrbein et al. modeled the posterior distribution of 3D pose hypotheses with normalized flows. Unlike these works that focus on a one-to-many mapping, we learn a one-to-many mapping first and then a many-to-one mapping, which allows for the effective modeling of different features corresponding to the various hypotheses to improve the representation ability.

Multi-Hypothesis Transformer

The overview of the proposed MHFormer is depicted in Figure 3 (a). Given a consecutive 2D pose sequence estimated by an off-the-shelf 2D pose detector from a video, our method aims to reconstruct the 3D pose of the center frame by making full use of spatial and temporal information in the multi-hypothesis feature hierarchies. To achieve our proposed three-stage framework, MHFormer is built upon (i) three major modules: Multi-Hypothesis Generation (MHG), Self-Hypothesis Refinement (SHR), and Cross-Hypothesis Interaction (CHI), and (ii) two auxiliary modules: temporal embedding and regression head.

In this work, we adopt a Transformer-based architecture since it performs well in long-range dependency modeling. We first give a brief description of the basic components in the Transformer , including a multi-head self-attention (MSA) and a multi-layer perceptron (MLP).

MSA splits the queries, keys, and values for hh times as well as performs the attention in parallel. Then, the outputs of hh attention heads are concatenated.

MLP. The MLP consists of two linear layers, which are used for non-linearity and feature transformation:

2 Multi-Hypothesis Generation

where LN⁡(⋅)\operatorname{LN}(\cdot) is the LayerNorm layer, l∈[1,...,L1]l{\in}[1,...,L_{1}] is the index of MHG layers, X1=XˉX^{1}{=}\bar{X}, and Xm=XL1m−1X^{m}{=}X^{m-1}_{L_{1}} (m>1m{>}1). The outputs of the MHG (i.e., XL1m{X}^{m}_{L_{1}}) are multi-level features containing diverse semantic information. Therefore, those features can be regarded as initial representations of different pose hypotheses and need to be further enhanced.

3 Temporal Embedding

The MHG helps to generate initial multi-hypothesis features in the spatial domain, whereas the capabilities of such features are not strong enough. Considering this limitation, we propose to build relationships across hypothesis features and capture temporal dependencies in the temporal domain with two carefully designed modules: an SHR module followed by a CHI module (see Figure 3 (c) and (d)).

4 Self-Hypothesis Refinement

In the temporal domain, we first construct the SHR to refine single-hypothesis features. Each SHR layer consists of a multi-hypothesis self-attention (MH-SA) block and a hypothesis-mixing MLP block.

where l∈[1,...,L2]l{\in}[1,...,L_{2}] is the index of SHR layers. Therefore, the message of different hypothesis features can be passed in a self-hypothesis way for feature enhancement.

Hypothesis-Mixing MLP. The multiple hypotheses are processed independently in the MH-SA, but there is no information exchange across hypotheses. To handle this issue, we add a hypothesis-mixing MLP after the MH-SA. The features of multiple hypotheses are concatenated and fed into the hypothesis-mixing MLP to merge (i.e., converge) themselves. Then, the converged features are evenly partitioned (i.e., diverged) into non-overlapping chunks along the channel dimension, forming refined hypothesis representations. The procedure can be formulated as:

where Concat⁡(⋅)\operatorname{Concat}(\cdot) is the concatenation operation and HM-MLP⁡(⋅)\operatorname{HM-MLP}(\cdot) is the function of hypothesis-mixing MLP which shares the same format as Eq. (2). This process explores the relations among channels of different hypotheses.

5 Cross-Hypothesis Interaction

We then model interactions among multi-hypothesis features via the CHI, which contains two blocks: multi-hypothesis cross-attention (MH-CA) and hypothesis-mixing MLP.

MH-CA. The MH-SA lacks connections across hypotheses, which limits its interaction modeling. To capture multi-hypothesis correlations mutually for cross-hypothesis communication, the MH-CA composed of multiple multi-head cross-attention (MCA) elements in parallel is proposed.

The MCA measures the correlation among cross-hypothesis features and has a similar structure to MSA. The common configuration of MCA uses the same input between keys and values . However, an issue with this configuration is that it will result in more blocks (e.g., 6 MCA blocks for 3 hypotheses). Here, we adopt a more efficient strategy, which reduces the number of parameters by using different inputs (only require 3 MCA blocks), as shown in Figure 4 (Right). The multiple hypotheses ZmZ^{m} are alternately regarded as queries, keys, and values and fed into the MH-CA:

where l∈[1,...,L3]l{\in}[1,...,L_{3}] is the index of CHI layers, Z0m=Z~L2m{Z}_{0}^{m}{=}\widetilde{Z}^{m}_{L_{2}}, m1m_{1} and m2m_{2} are the other two corresponding hypotheses, and MCA⁡(Q,K,V)\operatorname{MCA}(Q,K,V) denotes the function of the MCA. Thanks to the MH-CA, the message passing can be performed in a crossing way to significantly improve modeling power.

Hypothesis-Mixing MLP. The hypothesis-mixing MLP in the CHI serves as the same function as the process in Eq. (5). The outputs of the MH-CA are fed into it:

6 Regression Head

7 Loss Function

The entire model is trained in an end-to-end manner with a Mean Squared Error (MSE) loss, which is applied to minimize the error between the estimated and ground truth poses:

where X~in\widetilde{X}_{i}^{n} and YinY_{i}^{n} represent the predicted and ground truth 3D poses of joint ii at frame nn, respectively.

Experiments

We evaluate our method on two widely-used datasets for 3D HPE: Human3.6M and MPI-INF-3DHP .

Human3.6M. The Human3.6M dataset is the largest and most representative benchmark for 3D HPE. This dataset consists of 3.6 million images captured from four synchronized cameras at 50 Hz. There are 15 daily activities performed by 11 human subjects in an indoor environment. Following previous works , we train a single model on five subjects (S1, S5, S6, S7, S8) and test it on two subjects (S9 and S11). We adopt the most commonly used evaluation protocols: Protocol 1 is the MPJPE which measures the mean Euclidean distance between the ground truth and estimated joints in millimeters; Protocol 2 is the MPJPE after aligning the predicted 3D pose with the ground truth using translation, rotation, and scale (P-MPJPE).

MPI-INF-3DHP. The MPI-INF-3DHP is a large 3D pose dataset in both indoor and outdoor environments. This dataset provides 1.3 million frames, containing more diverse motions than Human3.6M. Following the setting in , we report metrics of MPJPE, Percentage of Correct Keypoint (PCK) with the threshold of 150 mmmm, and Area Under Curve (AUC) for a range of PCK thresholds.

2 Implementation Details

In our implementation, the proposed MHFormer contains L1=4L_{1}{=}4 MHG, L2=2L_{2}{=}2 SHR, and L3=1L_{3}{=}1 CHI layers. The MHFormer model is implemented in PyTorch framework on one GeForce RTX 3090 GPU. We train our model in an end-to-end manner from scratch using Amsgrad optimizer. The initial learning rate is set to 0.001 with a shrink factor of 0.95 applied after each epoch and 0.5 after every 5 epochs. For a fair comparison, the same horizontal flip augmentation is adopted following . We perform the 2D pose detection using cascaded pyramid network (CPN) for Human3.6M following and ground truth 2D pose for MPI-INF-3DHP following .

3 Comparison with State-of-the-Art Methods

Results on Human3.6M. The proposed MHFormer is compared with the state-of-the-art methods on Human3.6M. The results of our model with a receptive field of 351 frames using 2D detected inputs are reported in Table 8 (top). Without bells and whistles, our MHFormer outperforms all previous state-of-the-art methods by a large margin under both Protocol 1 (43.0 mmmm) and Protocol 2 (34.4 mmmm, see supplemental material). Compared to the very recent Transformer-based method, i.e., PoseFormer , MHFormer noticeably surpasses it by 1.3 mmmm in MPJPE (relative 3% improvement). Figure 5 shows the qualitative comparison with the PoseFormer and the baseline model (same architecture as ViT ) on some challenging poses. To further explore the lower bound of our method, we compared our MHFormer with the state-of-the-art methods with ground truth 2D poses as inputs. The results are shown in Table 8 (bottom). It can be seen that our method achieves the best performance (30.5 mmmm in MPJPE), outperforming all other methods.

Additionally, our method is compared with previous methods of generating multiple 3D pose hypotheses. The results are shown in Table 2. It is noteworthy that these methods report metrics for the best hypothesis due to the adopted one-to-many mapping, while our method reports metrics with a specific solution by learning a deterministic mapping, which is much more practical in reality. Even though we use much fewer hypothesis numbers (3 vs. 200), our proposed method consistently outperforms previous works.

Results on MPI-INF-3DHP. To assess the generalization ability, we evaluate our method on MPI-INF-3DHP dataset. Following , we use 2D pose sequences of 9 frames as our model input due to the fewer samples and shorter sequence lengths of this dataset compared to Human3.6M. The results in Table 3 demonstrate that our method achieves the best performance on all metrics (PCK, AUC, and MPJPE). It emphasizes the effectiveness of our MHFormer in improving performance in outdoor scenes.

4 Ablation Study

To verify the impact of each component and design in the proposed model, we conduct extensive ablation experiments on Human3.6M dataset under Protocol 1 with MPJPE.

Impact of Receptive Fields. For the video-based 3D HPE task, a large receptive field is essential for estimation accuracy. Table 4 shows the results of our method with different input frames. It can be seen that our method obtains larger gains with more frames fed into the model. The error has a great decrease of 16.7% from 9-frames to 351-frames with GT input, which indicates the effectiveness of our method in capturing long-range dependencies across frames with a large receptive field. Next, ablations in the following parts are carried out using a receptive field of 27 frames to balance the computation efficiency and performance.

Impact of Parameters in MHG. In the top part of Table 5, we report the results with different numbers of MHG layers. Experiments show that stacking more layers in MHG can slightly improve the performance with few parameter increases, but the gain disappears when the layer number is larger than 4. Moreover, we investigate the influence of using different numbers of hypotheses in MHG. The results are shown in the bottom part of Table 5. Increasing the number of hypotheses can improve the result, but the performance saturates when using 3 hypothesis representations. Notably, our model equipped with 3 hypotheses shows significant gains over the single-hypothesis model with 1.7 mmmm error reduction. This demonstrates that exploiting different representations of multiple pose hypotheses is helpful to improve the performance of the model, validating our motivation.

Impact of Parameters in SHR and CHI. Table 6 reports how the different parameters of SHR and CHI impact the performance and computation complexity of our model. The results show that enlarging the embedding dimension from 256 to 512 can boost the performance, but using dimensions larger than 512 cannot bring further improvements. In addition, we observe no more gains by stacking either more SHR or CHI layers. Therefore, the optimal parameters for our model are L2=2L_{2}{=}2, L3=1L_{3}{=}1, and C=512C{=}512.

Effect of Model Components. In Table 7, we carry out experiments to quantify the influence of our proposed components. Firstly, we compare our method with the baseline model. For a fair comparison, the results of the baseline are reported at the same number of layers as MHFormer with different embedding dimensions, since our hypothesis-mixing MLP in MHFormer takes concatenated hypothesis features as inputs (the dimension is 512×3=1536512{\times}3{=}1536). The results show that the baseline model is prone to overfitting due to the excessive number of parameters, whereas our method performs well. Additionally, it can be seen that our MHFormer built upon MHG, SHR, and CHI outperforms varying variants of baseline models (1.9 mmmm improvement). Then, when we incorporate multi-hypothesis representations and SHR or CHI within the baseline, the performance has significant gains (-1.3 mmmm for MHG-SHR and -1.0 mmmm for MHG-CHI). Besides, we remove the MHG in MHFormer (SHR-CHI). At this point, the model only captures temporal information and its error heavily increases by 1.3 mmmm. These ablations indicate that learning multi-hypothesis spatio-temporal representations is significant for 3D HPE, and the different hypothesis representations should be modeled in both independent and mutual ways.

We also explore the use of multi-level features in MHG by simply building the MHG upon several parallel Transformer encoders (MHFormer ∗). As shown in the table, our MHFormer equipped with the multi-level features increases the performance, which indicates that the multi-level features can bring valuable information to the final estimation.

Qualitative Results

Although our method does not aim to produce multiple 3D pose predictions, for better observation, we add additional regression layers and finetune our model to visualize the intermediate hypotheses. The several qualitative results are shown in Figure 6. It can be seen that our method is able to generate different plausible 3D pose solutions, especially for ambiguous body parts with depth ambiguity, self-occlusion, and 2D detector uncertainty. Moreover, the final 3D pose synthesized by aggregating multi-hypothesis information is more reasonable and accurate.

Conclusion

This paper presents Multi-Hypothesis Transformer (MHFormer), a new Transformer-based three-stage framework for the ambiguous inverse problem of 3D HPE from monocular videos. MHFormer first generates initial representations of multiple pose hypotheses in the spatial domain and then communicates across them in both independent and mutual ways in the temporal domain. Extensive experiments show that the proposed MHFormer has a fundamental advantage over single-hypothesis Transformers and achieves state-of-the-art performance on two benchmark datasets. We hope that our approach will foster further research in 2D-to-3D pose lifting considering various ambiguities.

Limitation. One limitation of our method is the relatively larger computational complexity. The excellent performance of Transformers comes at a price of high computational cost.

References

Appendix A Multi-Head Cross-Attention

The common configuration of MCA uses the same input between keys and values , i.e., the inputs x≠y=zx\neq y=z. Instead, we adopt a more efficient strategy by using different inputs, i.e., the inputs x≠y≠zx\neq y\neq z.

Appendix B Additional Quantitative Results

Table 8 shows the results of our proposed MHFormer on Human3.6M under Protocol 2. The input 2D poses are estimated by CPN . Without bells and whistles, our MHFormer achieves promising results that outperform the state-of-the-art approaches.

Several methods adopt a pose refinement module, which is first proposed by ST-GCN , to further improve the estimation accuracy. Following , we adopt the refine module and the results are shown in Table 9. It can be seen that our method can use the refine module to improve the performance, achieving an error of 42.4 mmmm in MPJPE which surpasses all other approaches by a large margin.

Appendix C Additional Ablation Studies

Effect of Model Components. Here, we give more details about how to build the different variants of MHFormer in Table 7 of our main manuscript:

Baseline: The baseline model contains 3 layers for standard Transformer encoder (same architecture as ViT ).

SHR-CHI: We remove the MHG module. SHR-CHI contains L2=2L_{2}{=}2 SHR and L3=1L_{3}{=}1 CHI layers.

MHG-SHR: We replace the CHI layers in MHFormer with SHR layers. MHG-SHR contains L1=4L_{1}{=}4 MHG and L3=3L_{3}{=}3 SHR layers.

MHG-CHI: We replace the SHR layers in MHFormer with CHI layers. SHR-CHI contains L1=4L_{1}{=}4 MHG and L3=3L_{3}{=}3 CHI layers.

MHFormer ∗: The MHG in MHFormer is simply built upon several parallel Transformer encoders.

MHFormer: Our proposed method that contains L1=4L_{1}{=}4 MHG, L2=2L_{2}{=}2 SHR, and L3=1L_{3}{=}1 CHI layers. Please refer to Figure 3 in our main manuscript.

Impact of Configurations in MH-CA. As mentioned in Sec. 3.5 of our main manuscript, the common configuration of MCA uses the same input between keys and values , which will result in more blocks. We adopt a more efficient configuration by using different inputs among queries, keys, and values. The performance and computational complexity of these two configurations are given in Table 10. We can see that using the same input between keys and values in MH-CA (MH-CA ∗) requires more parameters and FLOPs but cannot bring further performance gains. It illustrates the effectiveness of our efficient strategy in MCA.

Impact of Receptive Fields. For the video-based 3D human pose estimation task, the number of receptive fields directly influences the estimation results. Figure 7 (a) shows the results of our model with different receptive fields (between 1 and 351) on Human3.6M. Increasing the receptive field can improve the result under both CPN and GT 2D pose inputs, which demonstrates the great power of our method in long-range dependency modeling with a long input sequence.

Impact of 2D Detections. To show the effectiveness of our method on different 2D pose detectors, we carry out experiments with the detections from Stack Hourglass (SH) , Detectron , and CPN . In addition, to evaluate the robustness of our method to various levels of noise, we also conduct experiments on 2D ground truth plus different levels of additive Gaussian noise. The results are shown in Figure 7 (b). It can be observed that the curve has a nearly linear relationship between MPJPE of 3D poses and two-norm errors of 2D poses. These experiments validate both the effectiveness and robustness of our proposed method.

Appendix D Additional Visualization Results

3D Reconstruction Visualization. Figure 8 and Figure 9 show qualitative results of our method on Human3.6M dataset, MPI-INF-3DHP dataset, and challenging in-the-wild videos. Moreover, Figure 10 shows the qualitative comparison with the baseline method and the previous state-of-the-art method (PoseFormer ) on some wild videos. It can be seen that our method can produce more accurate and reasonable 3D poses, especially when the human action is complex and rare.

Hypothesis Visualization. For visualization purposes, we add additional regression layers and finetune our model to output intermediate hypotheses. Figure 11 shows the visualization results of intermediate 3D pose hypotheses generated by our proposed method. We can see that our MHFormer can generate different plausible 3D pose solutions, especially for ambiguous body parts with depth ambiguity, self-occlusion, and 2D detector uncertainty.

Attention Visualization. Visualization results of the multi-head attention maps of the first layers from the Multi-Hypothesis Generation (MHG) module and Self-Hypothesis Refinement (SHR) module (351-frame model with 3 hypotheses) are shown in Figure 12 and Figure 13, respectively. It can be found that the maps of multiple hypotheses contain diverse patterns and semantics. This indicates multiple representations in our method actually learn various modal information of pose hypotheses.