Camera-Space Hand Mesh Recovery via Semantic Aggregation and Adaptive 2D-1D Registration

Xingyu Chen, Yufeng Liu, Chongyang Ma, Jianlong Chang, Huayan Wang, Tian Chen, Xiaoyan Guo, Pengfei Wan, Wen Zheng

Introduction

Monocular 3D mesh recovery has attracted tremendous attention due to its extensive applications in AR/VR, human–machine interaction, etc. The task is to estimate 3D locations of mesh vertices from a single RGB image. It is particularly challenging owing to highly articulated structures, 2D-to-3D ambiguity, and self-occlusion. Significant efforts have been made recently for accurate 2D-to-3D reconstruction, including , to name a few.

Most of the aforementioned methods have difficulty in predicting absolute camera-space coordinates. Instead, they define a root (i.e., wrist of the hand) and estimate root-relative coordinates of the 3D mesh. In this aspect, these methods cannot be applied to many high-level tasks, e.g., hand-object interaction, that requires camera-space mesh information. To this end, we propose to jointly solve root-relative mesh recovery and root recovery by integrating these two sub-tasks into a unified framework, thereby bridging the gap between root-relative predictions and camera-space estimation.

RGB images consist of 2D patterns that are indirect cues of the underlying 3D structure. Therefore, 2D cues have long been leveraged to assist 3D tasks. For example, 2D pose and silhouette have been used to facilitate 3D pose regression .

However, the relationship between 2D cues and 3D structure remains unclear. We observe that 2D joint landmarks together with their semantic relations describe the 2D pose, while the silhouette indicates the holistic 3D-to-2D projection of the hand. They have different 2D properties and should be treated in different manners in the 3D task. Inspired by these observations, we set to explore the following aspects of the 2D-to-3D task: (1) different roles of 2D cues, (2) the reason for their different effects, and (3) how to construct more effective 2D cues.

In this paper, we propose a camera-space mesh recovery (CMR) framework to integrate the tasks of 3D hand mesh and root recovery into a unified framework. CMR consists of three phases, i.e., 2D cue extraction, 3D mesh recovery, and global mesh registration. For 2D cue extraction, we predict joint landmarks and silhouette from a single RGB image. For mesh recovery, we introduce an Inception Spiral Module for robust 3D decoding. Moreover, an aggregation method is designed for composing more effective 2D cues. Specifically, instead of implicitly learning the relations among joints, we exploit their known relations by aggregating landmark heatmaps in groups, which proves to be effective for the subsequent 3D task. Finally, camera-space root location is obtained by a global mesh registration step that aligns the generated 3D mesh with the extracted 2D landmarks and silhouette. This step is carried out via an adaptive 2D-1D registration method that achieves robustness by leveraging matching objectives in different dimensions. Our full pipeline surpasses state-of-the-art methods in the 3D mesh and root recovery tasks. While our approach is mainly described for hand mesh, it can be readily applied to full-body mesh as shown in the experiments. Figure 1 demonstrates several example results of our CMR for camera-space mesh recovery.

Our main contributions are summarized as follows:

We propose a novel aggregation method to collect effective 2D cues and exploit high-level semantic relations for root-relative mesh recovery.

We design an adaptive 2D-1D registration method to sufficiently leverage both joint landmarks and silhouette in different dimensions for robust root recovery.

We present a unified pipeline CMR for camera-space mesh recovery and demonstrate state-of-the-art performance on both mesh and root recovery tasks via extensive experiments on FreiHAND, RHD, and Human3.6M.

Related Work

According to different output property, we categorize methods for single-view RGB-based mesh recovery into three types, i.e., RGB→\rightarrowMANO/SMPL , RGB→\rightarrowVoxel , and RGB→\rightarrowCoord (coordinate) .

MANO and SMPL are parameterized 3D models of hand and human body, factorizing 3D human mesh into coefficients of shape and pose. Tremendous literature attempts to predict these coefficients for human/hand mesh recovery. For example, Zhou et al. estimated MANO coefficients based on the kinematic chain and developed an inverse kinematics network to improve prediction accuracy on pose coefficients. MANO/SMPL can reconstruct 3D mesh, but they embed 3D information into a parametrized space (e.g., PCA space), where the 3D structure is less straightforward (compared to 3D vertices).

Voxel is one type of Euclidean 3D representation, to which the canonical convolutional operator can be directly applied . Thereby, the mesh recovery task can be explored in voxels. For example, Moon et al. proposed an I2L-MeshNet by dividing voxels into three lixel spaces, where a 2.5D representation is leveraged for human mesh. The voxel/2.5D-based paradigm has impressive performance in terms of human mesh recovery because the merits of Euclidean space are fully leveraged. However, voxel/2.5D representations are not efficient enough in capturing 3D details (compared to 3D vertices).

Defferrard et al. proposed a graph convolution network (GCN) based on spectral filtering to process 3D vertices in the non-Euclidean space. Based on GCN, Kolotouros et al. developed a graph convolutional mesh regressor to directly estimate the 3D coordinates of mesh vertices. Ge et al. also developed a graph-based method for hand mesh recovery by learning from mixed real and synthetic data. Instead of spectral filtering, Lim et al. proposed spiral convolution (SpiralConv) to process mesh data in the spatial domain. Based on SpiralConv, Kulon et al. developed an encoder-decoder structure for efficient hand mesh recovery. We follow the RGB→\rightarrowCoord paradigm and explore versatile aggregation of 2D cues.

Root recovery.

In analogy to root recovery, estimation of external camera parameters has been widely studied . For instance, Zhang et al. designed an iterative regression method to simultaneously estimate external camera parameters and MANO coefficients. However, camera parameter estimation from RGB data is an ill-posed problem, leading to relatively low generalization performance. Moon et al. proposed RootNet to predict the absolute 3D human root. RootNet essentially modeled object size in images, but the pixel-level object area has a relatively low correlation with 3D root. Rogez et al. predicted both 2D and 3D pose so that the 3D root location can be obtained by aligning predicted 2D pose with projected 3D pose. We argue that this 2D-3D alignment cannot sufficiently leverage 2D information and propose an adaptive 2D-1D registration method for root recovery.

D cues in 3D shape/pose recovery.

Researchers have long exploited 2D cues in recovering 3D human shape and body parts. Pavlakos et al. utilized 2D pose to regress pose coefficients and used silhouette to estimate shape coefficients. Varol et al. first predicted 2D joint landmarks and body part segmentation, both of which were then combined to predict 3D pose. In this work, we aim to investigate how 2D cues work in 3D tasks and to leverage them effectively for hand mesh and root recovery.

Our Method

To represent a 3D mesh in camera-space, we divide it into the root-relative mesh and the camera-space root location. As shown in Figure 2, CMR includes three phases, i.e., 2D cue extraction, 3D mesh recovery, and global mesh registration. In the step of 2D cue extraction, we predict 2D pose and silhouette, which are used later for both mesh and root recovery. In the step of 3D mesh recovery, we generate a root-relative mesh which is then registered to the camera space in the final phase of the pipeline.

The first phase of our pipeline extracts 2D pose and silhouette. Both 2D pose and silhouette are represented by heatmaps. To refine the 2D cues gradually, we use a multi-stack hourglass network .

The 3D mesh is defined by its shape and pose . The silhouette is a holistic 3D-to-2D projection so it captures important shape cues. However, it can hardly describe the pose accurately. On the other hand, joint landmark locations are very informative for the pose. Given their different roles, how to better combine them becomes an interesting question. Therefore, we have an insight that to improve accuracy of the subsequent 3D tasks, it is essential not only to improve the accuracy of the 2D tasks respectively, but also to better aggregate them according to their semantic relations. Specifically, we propose to aggregate a series of 2D cues denoted as follows:

Hp\mathbf{H}_{p}: NN heatmaps of 2D poses. Each heatmap corresponds to a joint landmark.

Hs\mathbf{H}_{s}: a single heatmap of the silhouette.

cat(Hp,Hs\mathbf{H}_{p},\mathbf{H}_{s}): concatenating heatmaps of Hp\mathbf{H}_{p} and Hs\mathbf{H}_{s} to aggregate 2D pose and silhouette.

sum(Hp\mathbf{H}_{p}): combining all the joint landmarks as a single heatmap to aggregate joint locations.

group(Hp\mathbf{H}_{p}): concatenating Hp\mathbf{H}_{p} and tip-, part- , or level-grouped landmarks to aggregate joint semantics for high-level semantic relations.

2D silhouette and joint landmarks represent pixel-level locations in different aspects. A straightforward way of combining them is cat(Hp,Hs\mathbf{H}_{p},\mathbf{H}_{s}). Another simple baseline, sum(Hp\mathbf{H}_{p}), would discard semantics of individual joints by encoding their locations in a single heatmap. Thus, comparing Hp\mathbf{H}_{p} and sum(Hp\mathbf{H}_{p}) would reveal the effect of joint semantics. These simple baselines, as we will show, are inferior to more semantically meaningful ways of aggregation: group(Hp\mathbf{H}_{p}). That is, we sum the joint heatmaps in groups. As shown in Figure 3, three ways of grouping are introduced, i.e., by part, by level, and by tip grouping. Part grouping integrates joint landmarks on a finger, leg, arm, or torso, while level grouping integrates joint landmarks at the kinematic level . Tip grouping integrates pairwise part tips. As a result, group(Hp\mathbf{H}_{p}) forms sub-poses that exploit high-level semantic relation of 2D joints.

Spiral decoder.

The second phase of our pipeline generates the root-relative 3D mesh from the aggregated 2D cues using an improved spiral convolution decoder.

A 3D mesh M\mathcal{M} contains vertices V={vi=(xi,yi,zi)}i=1M\mathcal{V}=\{\mathbf{v}_{i}=(x_{i},y_{i},z_{i})\}_{i=1}^{M} and faces F\mathcal{F}. Convolution methods for M\mathcal{M} essentially process vertex features f(v)f(\mathbf{v}). SpiralConv is a graph-based convolution operator, which processes vertex features in the spatial domain. By explicitly formulating the order of aggregating neighboring vertices, SpiralConv++ presents an efficient version of SpiralConv. SpiralConv++ depends on a spiral manner of neighbor selection and adopts a fully-connected layer for feature fusion:

As shown in Figure 4(left), SpiralConv++ collects vertex neighbors (black dots) of the cell (red star) in a spiral manner. Then these neighbors are treated indiscriminately by a fully connected layer. Inspired by Inception and residual models , we design an Inception Spiral Module (ISM) to enhance the receptive field of SpiralConv. Specifically, as shown in Figure 4(right), we distinguish neighbors according to the spiral hierarchy (red, orange, and green dots) and adopt parallel layers with diverse receptive field for 3D decoding. The ISM can be described as

where [⋅][\cdot] denotes concatenating. In ISM we keep the number of parameters manageable by controlling the channel size of oo. Note that the i-diski\text{-disk} for each vertex may contain different number of elements. Similar to SpiralConv++, we truncate it to obtain a fixed-length sequence so that WiW_{i} and bib_{i} can be shared for all the vertices.

The overall architecture of our 3D mesh decoder is shown in Figure 5. Our design improves the spiral decoder in three aspects: (1) we replace SpiralConv with ISM; (2) we leverage multi-scale prediction and coarse-to-fine fusion; and (3) we introduce a self-regression mechanism by concatenating scale-level predictions with the same-scale feature. Meanwhile, we use a convolutional decoder which runs in parallel with the spiral decoder to refine the estimation of 2D pose and silhouette.

2 Root Recovery by Global Mesh Registration

The silhouette reflects holistic 3D-to-2D projection and contains strong geometric information for root recovery. Given the intrinsic matrix KK of the camera, predicted 3D vertices V\mathcal{V} can be projected into the 2D space by KVK\mathcal{V}, resulting in a 2D mesh which consists of 2D vertices V2D={vi2D=(xi2D,yi2D)}i=1M\mathcal{V}^{2D}=\{\mathbf{v}^{2D}_{i}=(x^{2D}_{i},y^{2D}_{i})\}_{i=1}^{M} with the original connectivity. Ideally, the 2D mesh should align well with the silhouette. However, they cannot be directly aligned because (1) the large number of 2D points leads to prohibitive computational cost, i.e., the silhouette contains thousands of pixels while the 2D mesh has 778 (for hand) or 6,890 (for human body) vertices; (2) 2D mesh vertices have no explicit correspondence with respect to the silhouette.

To overcome these difficulties, we design a group of 1D projections to align the silhouette with the 2D mesh. First, the silhouette is converted into contours C={c=(xic,yic)}i=1C\mathcal{C}=\{\mathbf{c}=(x^{c}_{i},y^{c}_{i})\}_{i=1}^{C} by edge detection . We define a set of 1D axes A={aj=(xja,yja)}j=1A\mathcal{A}=\{\mathbf{a}_{j}=(x^{a}_{j},y^{a}_{j})\}_{j=1}^{A}. These axes have unit length and are uniformly distributed, i.e., the angle difference between neighboring axes is π/A\pi/A. As shown in Figure 6, we project contours and 2D vertices onto aj\mathbf{a}_{j}, resulting in their 1D span, i.e., SjC={ci⋅aj}i=1C\mathcal{S}^{\mathcal{C}}_{j}=\{\mathbf{c}_{i}\cdot\mathbf{a}_{j}\}_{i=1}^{C} and SjV2D={vi2D⋅aj}i=1M\mathcal{S}^{\mathcal{V}^{2D}}_{j}=\{\mathbf{v}^{2D}_{i}\cdot\mathbf{a}_{j}\}_{i=1}^{M}.

Adaptive 2D-1D registration.

We denote the camera-space root by t=(xr,yr,zr)\mathbf{t}=(x^{r},y^{r},z^{r}) and 2D joint landmark predictions by P={pi=(xip,yip)}i=1N\mathcal{P}=\{\mathbf{p}_{i}=(x^{p}_{i},y^{p}_{i})\}_{i=1}^{N}. Given KK and joint regressor JJ defined by MANO or SMPL , the camera-space 3D vertices V+t\mathcal{V}+\mathbf{t} can be converted into 2D joints by Q=KJ(V+t)={qi=(xiq,yiq)}i=1N\mathcal{Q}=KJ\mathcal{(}\mathcal{V}+\mathbf{t})=\{\mathbf{q}_{i}=(x^{q}_{i},y^{q}_{i})\}_{i=1}^{N}. Since P\mathcal{P} and Q\mathcal{Q} have intrinsic correspondence, we define the energy function for 2D matching as

whose solution is denoted as t2D=arg⁡min⁡tE2D(t)\mathbf{t}^{2D}=\mathop{\arg\min}_{\mathbf{t}}E_{2D}(\mathbf{t}). Further, we use V+t\mathcal{V}+\mathbf{t} to implement the aforementioned 3D-2D-1D projection. Camera-space SV2D\mathcal{S}^{\mathcal{V}^{2D}} is thereby produced. Then the 1D correspondence is obtained using two endpoints of the 1D span. The energy function for alignment of 1D spans is defined as

In 1D space, t1D=arg⁡min⁡tE1D(t)\mathbf{t}^{1D}=\mathop{\arg\min}_{\mathbf{t}}E_{1D}(\mathbf{t}). Both 2D and 1D optimizations are based on quadratic programming . With t1D\mathbf{t}^{1D} and t2D\mathbf{t}^{2D}, we develop an adaptive method for the final root t∗\mathbf{t}^{*} with their distance d=∣∣t2D−t1D∣∣2d=||\mathbf{t}^{2D}-\mathbf{t}^{1D}||_{2}:

where δ1>δ2\delta_{1}>\delta_{2}, both of which are robust hyper-parameters according to 3D scale. Without wide search, we empirically use 0.06,0.020.06,0.02 for the hand and 1.0,0.51.0,0.5 for the body. This design attempts to sufficiently leverage the merits of 2D-1D registration. It is known that 2D pose is more fragile than silhouette. Hence, 2D pose is prone to be erroneous when there is a huge 2D-1D discrepancy, and silhouette is more dependable if d>δ1d>\delta_{1}. From another perspective, joint correspondence is more explicit than that of 1D projection. Thereby, the 2D process is more reliable if 2D and 1D results are similar (d<δ2d<\delta_{2}). The whole process of adaptive 2D-1D registration is presented in Figure 7, from which we can see that geometrical information is sufficiently exploited from 3D to 1D spaces for root recovery.

3 Loss Functions

We use L1 norm for loss terms of 3D mesh/pose Lmesh,Lpose3D\mathcal{L}_{mesh},\mathcal{L}_{pose3D}, and our 2D pose/silhouette losses Lpose2D,Lsil\mathcal{L}_{pose2D},\mathcal{L}_{sil} are based on binary cross entropy (BCE). We adopt normal loss Lnorm\mathcal{L}_{norm} and edge length loss Ledge\mathcal{L}_{edge} for smoother reconstruction . Formally, we have

where F,V\mathcal{F},\mathcal{V} are faces and vertices of a mesh; J,JJ,\mathcal{J} are the pose regressor and 3D joints; nk⋆\mathbf{n}_{\mathbf{k}}^{\star} indicates unit normal vector of face k\mathbf{k}; U,SU,S are heatmaps of 2D pose and silhouette; and ⋆\star denotes the ground truth. Following , U⋆U^{\star} is constructed with Gaussian distribution.

Our overall loss function is Ltotal=Lmesh+Lpose3D+λpLpose2D+λsLsil+λnLnorm+Ledge\mathcal{L}_{total}=\mathcal{L}_{mesh}+\mathcal{L}_{pose3D}+\lambda_{p}\mathcal{L}_{pose2D}+\lambda_{s}\mathcal{L}_{sil}+\lambda_{n}\mathcal{L}_{norm}+\mathcal{L}_{edge}, where λp=10,λs=0.5,λn=0.1\lambda_{p}=10,\lambda_{s}=0.5,\lambda_{n}=0.1 are used to balance different terms.

Experiments

We conduct experiments on several commonly-used benchmarks as listed below.

is a 3D hand dataset with 130,240 training images and 3,960 evaluation samples. The annotations of the evaluation set are not available, so we submit our predictions to the official server for online evaluation.

consists of 41,258 and 2,728 virtually rendered samples for training and testing on hand pose estimation, respectively.

is a large-scale 3D body pose benchmark containing 3.6 million video frames with annotations of 3D joint coordinates. SMPLify-X is used to obtain ground-truth SMPL coefficients. We follow existing methods to use subjects S1, S5, S6, S7, S8 for training and subjects S9, S11 for testing.

is a wild dataset with annotations of 2D human joints. Following previous work , we use the SMPL coefficients to produce the 3D human mesh. COCO is used for training.

We use the following metrics in quantitative evaluations.

measures the mean per joint/vertex position error in terms of Euclidean distance (mm) between the root-relative prediction and ground-truth coordinates.

is the MPJPE/MPVPE based on procrustes analysis with global variation being ignored.

measures MPJPE/MPVPE in the camera space for evaluation of the root recovery task.

is the area under the curve of PCK (percentage of correct keypoints) vs. error thresholds.

Our backbone is based on ResNet and we use the Adam optimizer to train the network with a mini-batch size of 3232. Serving as network inputs, image patches are cropped and resized to resolutions of 224×224\text{224}\times\text{224} (for the hand) or 256×256\text{256}\times\text{256} (for the body). Because of different data amounts, the total number of iterations is set as 38 and 25 epochs for tasks on hand and body. The initial learning rate is 10−410^{-4}, which is divided by 1010 at the 20th or 30th epoch. Data augmentation includes random box scaling/rotation, color jitter, etc.

2 Ablation Study

Our ablation studies are based on ResNet18. YoutubeHand serves as the baseline, but its code and models are inaccessible. Our re-implemented model obtains 9.069.06mm PA-MPVPE (see Table 1). With our spiral decoder, 8.548.54mm PA-MPVPE is achieved. Thus, our designs can strengthen the robustness of 2D-to-3D decoding.

Effects of various 2D cues.

This paper explores cues of 2D pose and silhouette for 3D mesh recovery, and a two-stack network is leveraged. At first, discarding the second stack, we only use one encoder-decoder structure for exposing details of 2D-cue effects. As shown in Table 2, Hs\mathbf{H}_{s} leads to 8.108.10mm PA-MPVPE while sum(Hp\mathbf{H}_{p}) reduces PA-MPVPE to 7.947.94mm. Both Hs\mathbf{H}_{s} and sum(Hp\mathbf{H}_{p}) are heatmaps that contain location information, so the locations of joint landmarks are more instructive than that of a holistic shape. This phenomenon is evident, because pose is more difficult to estimate than shape, and joint locations can directly provide cues on poses. Furthermore, Hp\mathbf{H}_{p} leads to an improved PA-MPVPE of 7.777.77mm by providing both joint locations and semantics. Compared to sum(Hp\mathbf{H}_{p}), it is seen that joint semantics is also important. In addition, the effect of cat(Hp,Hs\mathbf{H}_{p},\mathbf{H}_{s}) is worse than that of Hp\mathbf{H}_{p}. Therefore, there is no complementary benefit from 2D pose and silhouette.

To reveal potential reasons for different effects of 2D cues, we dissect feature representation acted by them. It is known that channel-specific feature usually embeds semantics, so we focus on typical channels in the first encoding block in the 3D mesh recovery phase. As shown in Figure 8, this block tends to describe trivial properties such as edges (see “none”). If Hs\mathbf{H}_{s} is employed, holistic 2D shapes emerge, but pose cues are ignored. With sum(Hp\mathbf{H}_{p}), holistic 2D shapes still can be learned based on joint landmark locations. Thus, this phenomenon can be the reason why there exists no complementary benefit in joint landmarks and silhouette. Moreover, holistic joint locations can also be captured with sum(Hp\mathbf{H}_{p}). As for Hp\mathbf{H}_{p}, although joint information is provided in separate heatmaps, the features after Hp\mathbf{H}_{p} have a tendency that multiple joints are simultaneously activated. Features shown in Hp\mathbf{H}_{p} of Figure 8 imply semantic relation of joints, which can have significant impact on 3D tasks. However, relation representation invited by Hp\mathbf{H}_{p} is not exhaustive enough. Specifically, the left two representations have similar patterns while both the 3rd and 4th representations mainly focus on the tips of thumb and middle finger.

With CMR-PG and CMR-SG, the same tendency emerges, which also validates the effectiveness of our designs. Note that CMR-PG induces better 2D pose (i.e., 0.7980.798 AUC) while CMR-SG invites better silhouette (i.e., 0.8320.832 mIoU), so their 2D and 1D performances are distinct. Overall, CMR-SG achieves the best performance on root recovery and camera-space mesh recovery.

We perform a comprehensive comparison on the FreiHAND dataset. As shown in Table 5, our proposed CMR outperforms other methods in terms of all the aforementioned metrics. Specifically, ResNet50-based CMR-PG achieves the best performance on root-relative mesh recovery, i.e., 7.07.0mm PA-MPVPE and 6.96.9mm PA-MPJPE.

In Table 5, we evaluate the root recovery performance of camera-space mesh on the FreiHAND dataset. Our CMR outmatches ObMan and I2L-MeshNet by 36.536.5mm and 11.511.5mm on CS-MPVPE. To the best of our knowledge, CMR achieves state-of-the-art performance on camera-space mesh recovery, i.e., 48.948.9mm CS-MPVPE and 48.848.8mm CS-MPJPE. From Figure 9, we can see that the proposed CMR outperforms all the compared methods on 3D PCK by a large margin. With the root-relative mesh from CMR, we demonstrate that our adaptive 2D-1D registration consistently outperforms RootNet as shown in Table 4.2.

Referring to CMR-PG’s PA-MPJPE and CS-MPJPE in Table 5, only an error of 6.96.9mm is incurred by relative pose with an error larger than 4040mm from global translation, rotation, and scaling. Thus, compared with root-relative information, camera-space 3D reconstruction is more essential to improve the practicability of hand mesh recovery, and we advocate studying the camera-space problem.

In Figure 9, we compare our CMR with several pose estimation methods on the RHD dataset. Following the criterion of PA-MPJPE, the predicted 3D pose are processed with procrustes analysis. The AUC of CMR-PG and CMR-SG is 0.9440.944 and 0.9490.949, respectively, surpassing all the other methods. In addition, we directly use the FreiHAND models for RHD test, inducing comparable AUC of 0.8720.872 and 0.8520.852. Hence, the cross-domain generalization ability of CMR is verified.

In Table 6, we compare our method with several state-of-the-art approaches on body mesh recovery task using the Human3.6M and COCO datasets. Pose2Mesh uses joint coordinates as the input and its performance downgrades slightly if the COCO dataset is added. As an RGB-to-voxel method, I2L-MeshNet performs better when both the Human3.6M and COCO datasets are used. By contrast, our CMR is more robust to different choices of training data and we achieve comparable or even better numbers of MPJPE and PA-MPJPE. Moreover, our approach has faster inference speed and consumes less GPU memory.

In Figure 1, we show several qualitative evaluation results on the FreiHAND and Human3.6M datasets. As demonstrated in this figure, our CMR can deal with a variety of complex situations for camera-space mesh recovery, e.g., challenging poses, object occlusion, and truncation.

In this work, we aim to recover 3D hand and human mesh in camera-centered space and we present CMR to unify tasks of root-relative mesh recovery and root recovery. We first investigate 2D cues including 2D joint landmarks and silhouette for 3D tasks. Then, an aggregation method is proposed to collect effective 2D cues. Through aggregation of joint semantics, high-level semantic relations are explicitly captured, which is instructive for root-relative mesh recovery. We also explore 2D information for root recovery and design an adaptive 2D-1D registration to sufficiently leverage 2D pose and silhouette to estimate absolute camera-space information. Our CMR achieves state-of-the-art performance on both mesh and root recovery tasks when evaluated on FreiHAND, RHD, and Human3.6M datasets.

In the future, we plan to integrate 2D semantic information together with biomechanical relationship for more robust monocular 3D representation. We are also interested in extending our CMR with a human/hand detector in a top-down manner for multi-person tasks.

Supplementary Materials

As shown in Figure 10, the 2D registration is instructive for finger alignment while the 1D process is beneficial to holistic shape alignment. Thus, joint landmarks and silhouette have different effects on the task of root recovery, and both of them can be effectively leveraged by our adaptive 2D-1D registration scheme.

Effects of our designs for spiral decoder.

We improve the spiral decoder with ISM, multi-scale mechanism, and self-regression. As shown in Table 7, all of these three design choices are beneficial to 2D-to-3D decoding, where ISM and multi-scale mechanism has relatively more significant impact.

Full result comparison on FreiHAND dataset.

As shown in Table 8. For root-relative and camera-space tasks, CMR achieves state-of-the-art performance on all the merits. Figure 11 plots PCK curves of 3D joints, which can serve as a complement of Figure 9 of our main text.

Full-feature representation after 2D cues.

Figure 13 serves as a complement of Figure 8 in the main text. hs and sum(hp) induce holistic shape and pose. In contrast, hp invites simultaneously activated joint representation that essentially implies semantic relation. However, this relation representation is not comprehensive enough. We design group(hp) for explicitly exploring known high-level semantics so that more comprehensive joint relations can be captured.

More qualitative results.

Figure 14 and 15 illustrate comprehensive qualitative results of our predicted silhouette, 2D pose, projection of mesh, side-view mesh, camera-space mesh and pose in meter. Different from most methods, 3D roots required by image-mesh alignment are provided by CMR itself rather than ground truth, so CMR can handle images in the wild.

Referring to Figure 14, FreiHAND’s challenges include hard poses, object interactions, and truncation. Overcoming these difficulties, CMR can generate accurate silhouette, 2D pose, and camera-space 3D information.

Referring to Figure 15, samples of RHD, STB, and real-world dataset released by are illustrated. We directly use the FreiHAND model for these datasets, and equally accurate predictions are obtained. Thus, CMR demonstrates superior capability of cross-domain generalization. Figure 15 also shows examples on Human3.6M and COCO. It can be seen that our CMR achieves reasonable results in the task of human body recovery.

Failure case analysis.

Figure 12 shows three typical failure cases of CMR-SG. When only a small portion of the hand is visible in the input (Figure 12(a)), CMR-SG predicts wrong silhouette and 2D pose. Consequently, the camera-space information is not accurate. For cases of occlusion (e.g., Figure 12(b), in which the forefinger is completely occluded by the middle finger), although 2D pose and silhouette prediction results are still reasonable, it is difficult to obtain accurate 3D mesh since self-occlusion is challenging for the mesh recovery stage. Referring to Figure 12(c), strong contrast and extreme illumination change in the RGB input leads to large but consistent errors in silhouette, 2D pose, and 3D mesh prediction results.