Direct Multi-view Multi-person 3D Pose Estimation

Tao Wang, Jianfeng Zhang, Yujun Cai, Shuicheng Yan, Jiashi Feng

Introduction

Multi-view multi-person 3D pose estimation aims to localize 3D skeleton joints for each person instance in a scene from multi-view camera inputs. It is a fundamental task that benefits many real-world applications (such as surveillance, sportscast, gaming and mixed reality) and is mainly tackled by reconstruction-based and volumetric approaches in previous literature, as shown in Fig. 1 (a) and (b). The former first estimates 2D poses in each view independently and then aggregates them and reconstructs their 3D counterparts via triangulation or a 3D pictorial structure model. The volumetric approach builds a 3D feature volume through heatmap estimation and 2D-to-3D un-projection at first, based on which instance localization and 3D pose estimation are performed for each person instance individually. Though with notable accuracy, the above paradigms are inefficient due to highly relying on those intermediate tasks. Moreover, they estimate 3D pose for each person separately, making the computation cost grow linearly with the number of persons.

Targeted at a more simplified and efficient pipeline, we were wondering if it is possible to directly regress 3D poses from multi-view images without relying on any intermediate task? Though conceptually attractive, adopting such a direct mapping paradigm is highly non-trivial as it remains unclear how to perform skeleton joints detection and association for multiple persons within a single stage. In this work, we address these challenges by developing a novel Multi-view Pose transformer (MvP) model which significantly simplifies the multi-person 3D pose estimation. Specifically, MvP represents each skeleton joint as a learnable positional embedding, named joint query, which is fed into the model and mapped into final 3D pose estimation directly (Fig. 1 (c)), via a specifically designed attention mechanism to fuse multi-view information and globally reason over the joint predictions to assign them to the corresponding person instances. We develop a novel hierarchical query embedding scheme to represent the multi-person joint queries. It shares joint embedding across different persons and introduces person-level query embedding to help the model in learning both person-level and joint-level priors. Benefiting from exploiting the person-joint relation, the model can more accurately localize the 3D joints. Further, we propose to update the joint queries with input-dependent scene-level information (i.e., globally pooled image features from multi-view inputs) such that the learnt joint queries can adapt to the target scene with better generalization performance.

To effectively fuse the multi-view information, we propose a geometrically-guided projective attention mechanism. Instead of applying full attention to densely aggregate features across spaces and views, it projects the estimated 3D joint into 2D anchor points for different views, and then selectively fuses the multi-view local features near to these anchors to precisely refine the 3D joint location. we propose to encode the camera rays into the multi-view feature representations via a novel RayConv operation to integrate multi-view positional information into the projective attention. In this way, the strong multi-view geometrical priors can be exploited by projective attention to obtain more accurate 3D pose estimation.

Comprehensive experiments on 3D pose benchmarks Panoptic , as well as Shelf and Campus demonstrate our MvP works very well. Notably, it obtains 92.3% AP25 on the challenging Panoptic dataset, improving upon the previous best approach VoxelPose by 9.8%, while achieving nearly 2×2\times speed up. Moreover, the design ethos of our MvP can be easily extended to more complex tasks—we show that a simple body mesh branch with SMPL representation trained on top of a pre-trained MvP can achieve competitively qualitative results.

Our contributions are summarized as follows: 1) We strive for simplicity in addressing the challenging multi-view multi-person 3D pose estimation problem by casting it as a direct regression problem and accordingly develop a novel Multi-view Pose transformer (MvP) model, which achieves state-of-the-art results on the challenging Panoptic benchmark. 2) Different from query embedding designs in most transformer models, we propose a more tailored and concise hierarchical joint query embedding scheme to enable the model to effectively encode person-joint relation. Additionally, we mitigate the commonly faced generalization issue by a simple query adaptation strategy. 3) We propose a novel projective attention module along with a RayConv operation for fusing multi-view information effectively, which we believe are also inspiring for model designs in other multi-view 3D tasks.

Related Works

3D pose estimation from monocular inputs is an ill-posed problem as multiple 3D predictions may result in the same 2D projection. To alleviate such projective ambiguities, multi-view methods have been explored. Research works on single-person scenes use either multi-view geometry for feature fusion and triangulation , or pictorial structure models for fast and robust 3D pose reconstruction , achieving promising results. However, it is more challenging as we progress towards multi-person scenes. Current approaches mainly exploit a multi-stage pipeline for multi-person tasks, including reconstruction-based and volumetric paradigms. Despite their notable accuracy, these methods suffer expensive computation cost from the intermediate tasks, such as cross-view matching and heatmap back-projection. Moreover, the total computation cost grows linearly with the number of persons in the scene, making them hardly scalable for larger scenes. Different from all previous approaches that rely on a multi-stage pipeline with computation redundancy, our method views multi-person 3D pose estimation as a direct regression problem based on a novel Multi-view Pose transformer model, enables an intermediate task-free single stage solution.

Attention and Transformers

Driven by the recent success in natural language fields, there have been growing interests in exploring the Transformers for computer vision tasks, such as image recognition and generation , as well as more complicated object detection and video instance segmentation . However, multi-person 3D pose estimation has not been explored along this direction. In this study, we propose a novel Multi-view Pose Transformer architecture with a joint query embedding scheme and a projective attention module to regress 3D skeleton joints from multi-view images directly, delivering a simplified and effective pipeline.

Multi-view Pose Transformer (MvP)

To build a direct multi-person 3D pose estimation framework from multi-view images, we introduce a novel Multi-view Pose transformer (MvP). MvP takes in the multi-view feature representations, and transforms them into groups of 3D joint locations directly (Fig. 2 (a)), delivering multi-person 3D pose results, with the following carefully designed query embedding and attention schemes for detecting and grouping the skeleton joints.

Inspired by transformers , MvP represents each skeleton joint as a learnable positional embedding, which is fed into the transformer decoder and mapped into final 3D joint location by jointly attending to other joints and the multi-view information (Fig. 2 (a)). The learnt embeddings encode a prior knowledge about the skeleton joints and we name them as joint queries. MvP develops the following concise query embedding scheme.

The most straightforward way for designing joint query embeddings is to maintain a learnable query vector for each joint per person. However, we empirically find this scheme does not work well, likely because such a naive strategy cannot share the joint-level knowledge between different persons.

With such a hierarchical embedding scheme, the number of learnable query embedding parameters is reduced from NJCNJC to (N+J)C(N+J)C.

Input-dependent Query Adaptation

2 Projective Attention for Multi-view Feature Fusion

It is crucial to aggregate complementary multi-view information to transform the joint embeddings into accurate 3D joint locations. We consider the dot product attention mechanism of transformers to fuse the multi-view image features. However, naively applying such dot product attention densely over all spatial locations and camera views will incur enormous computation cost. Moreover, such dense attention is difficult to optimize and delivers poor performance empirically since it does not exploit any 3D geometric knowledge.

Therefore, we propose a geometrically-guided multi-view projective attention scheme, named projective attention. The core idea is to take the 2D projection of the estimated 3D joint location as the anchor point in each view, and only fuse the local features near those projected 2D locations from different views. Motivated by the deformable convolution , we adopt an adaptive deformable sampling strategy to gather the localized context information in each camera view, as shown in Fig. 2 (b). Other local attention operations can also be adopted as an alternative. Formally, given joint query feature q and 3D joint position y, the projective attention is defined as

The projective attention incorporates two geometrical cues, i.e., the corresponding 2D spatial locations across views from the 3D to 2D projection and the deformed neighborhood of the anchors from the learned offsets to gather view-adaptive contextual information. Unlike naive attention where the query feature densely interacts with the multi-view key features across all the spatial locations, the projective attention is more selective for the interaction between the query and each view—only the features from locations near to the projected anchors are aggregated, and thus is much more efficient.

The positional encoding is an important component of the transformer, which provides positional information of the input sequence. However, a simple per-view 2D positional encoding scheme cannot encode the multi-view geometrical information. To tackle this limitation, we propose to encode the camera ray directions that represent positional information in 3D space into the multi-view feature representations. Concretely, the camera ray direction Rv\textbf{R}_{v}, generated with the view-specific camera parameters, is concatenated channel-wisely to the corresponding image feature representation Zv\textbf{Z}_{v}. Then a standard convolution is applied to obtain the updated feature representation Z^v\hat{\textbf{Z}}_{v}, with the view-dependent geometric information:

We name the operation as RayConv. With it, the obtained feature representation Z^v\hat{\textbf{Z}}_{v} is used for the projective attention by replacing Zv{\textbf{Z}}_{v} in Eqn. (3).

Such drop-in replacement introduces negligible computation, while injecting strong multi-view geometrical prior to augment the projective attention scheme, thus helping more precisely predict the refined 3D joint position.

3 Architecture

Our overall architecture (Fig. 2 (a)) is pleasantly simple. It adopts a convolution neural network, designed for 2D pose estimation , to obtain high-resolution image features {Zv}v=1V\{\textbf{Z}_{v}\}^{V}_{v=1} from multi-view inputs {Iv}v=1V\{\textbf{I}_{v}\}^{V}_{v=1}. The features are then fed into the transformer decoder consisting of multiple decoder layers to predict the 3D joint locations. Each layer conducts a self-attention to perform pair-wise interaction between all the joints from all the persons in the scene; a projective attention to selectively gather the complementary multi-view information; and a feed-forward regression to predict the 3D joint positions and their confidence scores. Specifically, the transformer decoder applies a multi-layer progressive regression scheme, i.e., each decoder layer outputs 3D joint offsets to refine the input 3D joint positions from previous layer.

MvP learns skeleton joints feature representations and is extendable to recovering human mesh with a parametric body mesh model . Specifically, after average pooling on the joint features into per-person feature, a feed-forward network is used to predict the corresponding body mesh represented by the parametric SMPL model . Similar to the joint location prediction, the SMPL parameters follow multi-layer progressive regression scheme.

4 Training

MvP infers a fixed set of MM joint locations for NN different persons, where M=NJM=NJ. The main training challenge is how to associate the skeleton joints correctly for different person instances. Unlike the post-hoc grouping of detected skeleton joints as in bottom-up pose estimation methods , MvP learns to directly predict the multi-joint 3D human pose in a group-wise fashion as shown in Fig. 2 (a). This is achieved by a grouped matching strategy during model training.

The ground truth set Y∗\textbf{Y}^{*} of 3D poses of different person instances is smaller than the prediction set of size NN, which is padded to size NN with empty element ∅\varnothing. Then we find a bipartite matching between the prediction set and the ground truth set by searching for a permutation of σ^∈ℵN\hat{\sigma}\in\aleph_{N} that achieves the lowest matching cost:

We consider both the regressed 3D joint position and confidence score for the matching cost:

where Yn∗≠∅\textbf{Y}^{*}_{n}\neq\varnothing, and L1\mathcal{L}_{1} computes the L1L_{1} loss error. Following , we employ the Hungarian algorithm to compute the optimal assignment σ^\hat{\sigma} with the above matching cost.

Objective Function

We compute the Hungarian loss with the obtained optimal assignment σ^\hat{\sigma}:

Experiments

In this section, we aim to answer following questions. 1) Can MvP provide both efficient and accurate multi-person 3D pose estimation? 2) How does the proposed attention mechanism help multi-view multi-person skeleton joints information fusing? 3) How does each individual design choice affect model performance? To this end, we conduct extensive experiments on several benchmark datasets.

Panoptic is a large-scale benchmark with 3D skeleton joint annotations. It captures daily social activities in an indoor environment. We conduct extensive experiments on Panoptic to evaluate and analyze our approach. Following VoxelPose , we use the same data sequences except ‘160906_band3’ in the training set due to broken images. Unless otherwise stated, we use five HD cameras (3, 6, 12, 13, 23) in our experiments. All results reported in the experiments follow the same data setup. We use Average Precision (AP) and Recall , as well as Mean Per Joint Position Error (MPJPE) as evaluation metrics. Shelf and Campus are two multi-person datasets capturing indoor and outdoor environments, respectively. We split them into training and testing sets following . We report Percentage of Correct Parts (PCP) for these two datasets.

Implementation Details

Following VoxelPose , we adopt a pose estimation model build upon ResNet-50 for multi-view image features extraction. Unless otherwise stated, we use a stack of six transformer decoder layers. The model is trained for 40 epochs, with the Adam optimizer of learning rate 10−410^{-4}. During inference, a confidence threshold of 0.1 is used to filter out redundant predictions. Please refer to supplementary for more implementation details.

1 Main Results

We first evaluate our MvP model on the challenging Panoptic dataset and compare it with the state-of-the-art VoxelPose model . As shown in Table 1, Our MvP achieves 92.3 AP25, improving upon VoxelPose by 9.8%, and achieves much lower MPJPE (15.8 v.s 17.8). Moreover, MvP only requires 170ms to process a multi-view input, about 2×2\times faster than VoxelPoseWe count averaged per-sample inference time in millisecond on Panoptic test set. For all methods, the time is counted on GPU GeForce RTX 2080 Ti and CPU Intel i7-6900K @ 3.20GHz.. These results demonstrate both accuracy and efficiency advantages of MvP from estimating 3D poses of multiple persons in a direct regression paradigm. To further demonstrate efficiency of MvP, we compare its inference time with VoxelPose’s when processing different numbers of person instances. As shown in Fig. 3, the inference time of VoxelPose grows linearly with the number of persons in the scene due to the per-person regression paradigm. In contrast, MvP keeps constant inference time no matter how many instances in the scene. Notably, it takes only 185ms for MvP to process scenes even with 100 person instances (the blue line), demonstrating its great potential to handle crowded scenarios.

Shelf and Campus

We further compare our MvP with state-of-the-art approaches on the Shelf and Campus datasets. The reconstruction-based methods use 3D pictorial model or conditional random field within a multi-stage paradigm; and the volumetric approach VoxelPose highly relies on computationally intensive intermediate tasks. As shown in Table 2, our MvP achieves the best performance in all the actors on the Shelf dataset. Moreover, it obtains a comparable result on the Campus dataset as VoxelPose without relying on any intermediate task. These results further confirm the effectiveness of MvP for estimating 3D poses of multiple persons directly.

2 Visualization

We visualize some 3D pose estimations of MvP on the challenging Panoptic dataset in Fig. 4. It can be observed that MvP is robust to large pose deformation (the 1st example) and severe occlusion (the 2nd example), and can achieve geometrically plausible results w.r.t. different viewpoints (the rightmost column). Moreover, MvP is extendable to body mesh recovery and can achieve fairly good reconstruction results (the 2nd and 4th rows). All these results verify both effectiveness and extendability of MvP. Please see supplementary for more examples.

Attention Mechanism

We visualize the projective attention and the self-attention in Fig. 5. Benefiting from the 3D-to-2D projection, the projective attention can accurately locate the skeleton joint in each camera view (the green point) based on the current estimated 3D joint location. We observe it learns to gather adaptive local context information (the red points) with the deformable sampling operation. For instance, when regressing the 3D position of mid-hip (the 1st example), the projective attention selectively attends to informative joints such as the left and right hips as well as thorax, which offers sufficient contextual information for accurate estimation. We also visualize the self-attention, which learns pair-wise interaction between all the skeleton joints in the scene. From the 3D plot in Fig. 5, we can observe a certain skeleton joint mainly attends to other joints of the same person instance (more opaque). It also attends to joints from other person instances, but with less attention (more transparent). This phenomenon is reasonable as the skeleton joints of a human body are strongly correlated to each other, e.g., with certain pose priors and bone length.

3 Ablation

MvP introduces RayConv to encode multi-view geometric information, i.e., camera ray directions into image feature representations. As shown in Table 3(a), if removing RayConv, the performance drops significantly—4.8 decrease in AP25 and 1.6 increase in MPJPE. This indicates the multi-view geometrical information is important for the model to more precisely localize the skeleton joints in 3D space. Without RayConv, the transformer decoder cannot accurately capture positional information in 3D space, resulting in performance drop.

Importance of Hierarchical Query Embedding

As shown in Table 3(b), compared with the straightforward and unstructured per-joint query embedding scheme, the proposed hierarchical query embedding boosts the performance sharply—14.1 increase in AP25 and 23.4 decrease in MPJPE. Its advantageous performance clearly verifies introducing the person-level queries to collaborate with the joint-level queries can better exploit human body structural information and improve model to better localize the joints. Upon the hierarchical query embedding scheme, adding the query adaptation strategy further improves the performance significantly, reaching AP25 of 92.3 and MPJPE of 15.8. This shows the proposed approach effectively adapts the query embeddings to the target scene and such adaptation is indeed beneficial for the generalization of MvP to novel scenes.

Different Model Designs

We also examine effects of varying the following designs of the MvP model to gain better understanding on them.

Confidence Threshold During inference, a confidence threshold is used to to filter out the low-confidence and erroneous pose predictions, and obtain the final result. Adopting a higher confidence will select the predictions in a more restrictive way. As shown in Table 3(c), a higher confidence threshold brings lower MPJPE as it selects more accurate predictions; but it also filters out some true positive predictions and thus reduces the average precision.

Number of Decoder Layers Decoder layers are used for refining the pose estimation. Stacking more decoder layers thus gives better performance (Table 3(d)). For instance, the MPJPE is as high as 49.6 when using only two decoder layers, but it is significantly reduced to 22.8 when using three decoder layers. This clearly justifies the progressive refinement strategy of our MvP model is effective. However the benefit of using more decoder layers diminishes when the number of layers is large enough, implying the model has reached the ceiling of its model capacity.

Number of Camera Views Multi-view inputs provide complementary information to each other which is extremely useful when handling some challenging environment factors in 3D pose estimation like occlusions. We vary the number of camera views to examine whether MvP can effectively fuse and leverage multi-view information to continuously improve the pose estimation quality (Table 3(e)). As expected, with more camera views, the 3D pose estimation accuracy monotonically increases, demonstrating the capacity of MvP in fusing multi-view information.

Number of Deformable Sampling Points Table 3(f) shows the effect of the number of deformable sampling points KK used in the projective attention. With only one deformable point, MvP already achieves a respectable result, i.e., 88.6 in AP25 and 17.4 in MPJPE. Using more sampling points further improves the performance, demonstrating the projective attention is effective at aggregating information from the useful locations. When K=4K=4, the model gives the best result. Further increasing KK to 8, the performance starts to drop. It is likely because using too many deformable points introduces redundant information and thus makes the model more difficult to optimize.

Conclusion

We introduced a direct and efficient model, named Multi-view Pose transformer (MvP), to address the challenging multi-view multi-person 3D human pose estimation problem. Different from existing methods relying on tedious intermediate tasks, MvP substantially simplifies the pipeline into a direct regression one by carefully designing the transformer-alike model architecture with a novel hierarchical joint query embedding scheme and projective attention mechanism. We conducted extensive experiments to verify its superior performance and speed over the well-established baselines.

We empirically found MvP needs sufficient data for model training since it learns the 3D geometry implicitly. In the future, we will study how to enhance the data-efficiency of MvP by leveraging the strategy like self-supervised pre-training or exploring more advanced approaches. Similar to prior works, we also found MvP suffers from performance drop for cross-camera generalization, that is, generalizing on novel camera views. We will explore approaches like disentangling camera parameters and multi-view feature learning to improve this aspect. Besides, we will explore the large-scale applications of MvP and further extend it to other relevant tasks. Thanks to its efficiency, MvP would be scalable to handle very crowded scenes with many persons. Moreover, the framework of MvP is general and thus extensible to other 3D modeling tasks like dense mesh recovery of common objects.

References

More Implementation Details

We use PyTorch to implement the proposed Multi-view Pose transformer (MvP) model. Our MvP model is trained on 8 Nvidia RTX 2080 Ti GPUs, with a batch size of 1 per GPU and a total batch size of 8. We use the Adam optimizer with an initial learning rate of 1e-4 and decrease the learning rate by a factor of 0.1 at 20 epochs during training. The hyper-parameter λ\lambda for balancing confidence score and pose regression losses is set to 2.5. We use the image feature representations (256-d) from the de-convolution layer of the 2D pose estimator PoseResNet for multi-view inputs. Additionally, we provide the code of MvP, including the implementation of model architecture, training and inference, in the folder of “./mvp” for better understanding our method.

Architecture Details

Fig. 6 (a) illustrates our proposed hierarchical query embedding scheme. As shown in Eqn. (1), each person-level query is added individually to the same set of joint-level queries to obtain the per-person customized joint queries. This scheme shares the joint-level queries across different persons and thus reduces the number of parameters (the joint embeddings) to learn, and helps the model generalize better. The generated per-person joint query embedding is further augmented by adding the scene-level feature extracted from the input images.

Decoder Layer

The decoder of MvP transformer consists of multiple decoder layers for regressing 3D joint locations progressively. Fig. 6 (b) demonstrates the detailed architecture of a decoder layer, which contains a self-attention module to perform pair-wise interaction between all the joints from multiple persons in the scene; a projective attention module to selectively gather the complementary multi-view information; and a feed-forward network (FFN) to predict the 3D joint locations and their confidence scores.

More Ablation Studies

MvP encodes camera ray directions into the multi-view image feature representations via RayConv. We also compare with the simple positional embedding baseline that uses 2D coordinates as the positional information to embed, similar to the previous transformer-based models for vision tasks . Specifically, we replace the camera ray directions with 2D spatial coordinates of the input images in RayConv. Results are shown in Table 4. We can observe using the 2D coordinates in RayConv results in much worse performance, i.e., 83.3 in AP25 and 18.1 in MPJPE. This result demonstrates that using such view-agnostic 2D coordinates information cannot well encode multi-view geometrical information into the model; while using camera ray directions can effectively encode the positional information of each view in 3D space, thus leading to better performance.

Replacing Projective Attention with Dense Attention

We further investigate the effectiveness of the proposed projective attention by comparing it with the dense dot product attention, i.e., conducting attention densely over all spatial locations and camera views for multi-view information gathering. Results are given in Table 5. We observe MvP with the dense attention (MvP-Dense) delivers very poor performance (0.0 AP25 and 114.5 MPJPE) since it does not exploit any 3D geometries and thus is difficult to optimize. Moreover, such dense dot product attention incurs significantly higher computation cost than the proposed projective attention—MvP-Dense costs 31 G GPU memory, more than 5×\times larger than MvP with the projective attention, which only costs 6.1 G GPU memory.

More Results

We also evaluate our MvP model on the most widely used single-person dataset Human3.6M collected in an indoor environment. We follow the standard training and evaluation protocol and use MPJPE as evaluation metric. Our MvP model achieves 18.6 MPJPE which is comparable to state-of-the-art approaches (18.6 v.s 17.7 and 19.0) .

Qualitative Result

Here we present more qualitative results of MvP on Panoptic (Fig. 7), Shelf and Campus (Fig. 8) datasets. From Fig 7 we can observe that MvP can produce satisfactory 3D pose and body mesh estimations even in case of strong pose deformations (the 1st example) and large occlusion (the 2nd and 3rd examples). Moreover, the performance of MvP is robust even in the challenging crowded scenario, as shown in the 1st example in Fig. 8.