Interacting Attention Graph for Single Image Two-Hand Reconstruction

Mengcheng Li, Liang An, Hongwen Zhang, Lianpeng Wu, Feng Chen, Tao Yu, Yebin Liu

Introduction

Interacting two-hand reconstruction is one of the fundamental tasks towards manifold industrial applications such as virtual reality (VR), human-computer-interaction (HCI), robotics, holoportation, digital medicine, etc. Recently, monocular single hand pose and shape recovery has witnessed great success owing to deep neural networks and large scale datasets . However, two-hand reconstruction is more challenging and remains unsolved for two reasons. First, severe mutual occlusions and appearance similarity confuse the feature extractors, making it difficult for networks to align hand poses with image features. Second, the interaction context between two hands is difficult to be effectively formulated during network design and training.

Monocular depth-based two-hand tracking has been studied for years and promising results have been demonstrated. However, the energy demand and algorithm complexity restrict the ubiquitous application of depth-based methods. Recently, Wang et al. contributes a monocular RGB based two-hand reconstruction by tracking dense matching map. However, the tracking procedure itself is inherently sensitive to fast motion, and does not take full advantage of prior knowledge between interacting hands. Since the proposal of the large scale two-hand dataset InterHand2.6M , learning based single image two-hand reconstruction methods have emerged. Existing methods either employ 2.5D heatmaps to estimate hand joint positions , or use them as attention maps to extract sparse image features . However, such sparse local image features encoded in the heatmaps could not effectively model hand surface occlusions, and could not extract dense interaction context. In contrast, vertex-based graph convolutional network (GCN) has achieved great success in single hand reconstruction , yet it has not been demonstrated in two-hand conditions, and the previously mentioned challenges remain to be addressed.

In this paper, we propose Interacting Attention Graph Hand (IntagHand), a novel GCN based single image two-hand reconstruction method. As a basic pipeline, we initially utilize GCN to regress mesh vertices of each hand in a coarse-to-fine manner, similar to traditional GCN . However, for the two-hand task, naively using a two-stream GCN to generate two hand vertices fails to utilize the interaction context between two hands, making the network confused regarding two-hand mutually occluded parts. Moreover, without any image feature feedback, the network has difficulty aligning vertices to image features as suggested by . To address these issues, we equip the GCN with two novel attention modules. The first module is a pyramid image feature attention (PIFA) module, which uses a transformer encoder to update the latent vertex features with patched image features. Unlike projection based vertex-image alignment , PIFA benefits from the global sensing ability of the attention mechanism to help each vertex seek alignment over all image patches. Furthermore, as GCN upsamples mesh vertices in a coarse-to-fine manner, we design an encoder-decoder based image feature extraction module to extract pyramid features, forcing the high resolution mesh to leverage fine-grained features. The second module is a cross hand attention (CHA) module that encodes interaction context into hand vertex features. The CHA module allows vertices of each hand to pay dense attention to the other hand’s vertex features in order to disambiguate interhand occlusions. Benefiting from the GCN structure and the novel attention based modules, the IntagHand outperforms existing methods on InterHand2.6M by a large margin (8.8mm v.s. 13.5mm). Moreover, our method is efficient for real-time applications, producing well-aligned two-hand results on in-the-wild images and live video streams, as shown in Fig. 1 and our project page. Overall, our contributions are summarized as:

We propose the first two-hand reconstruction method using GCN based mesh regression, named IntagHand, demonstrating the effectiveness of GCN for the two-hand reconstruction task.

We propose a pyramid image feature attention (PIFA) module to distill local occlusion information with global image patch attention, producing better alignment between the hand vertices and the image features.

We propose a cross hand attention (CHA) module to implicitly model the two-hand interaction context, improving the reconstruction accuracy for closely interacting poses.

Our method achieves the new state-of-the-art results and outperforms existing solutions by a large margin on the InterHand2.6M benchmark. Furthermore, We demonstrate the generalization ability of our method on in-the-wild images.

Related Works

Since the last century, hand pose estimation and gesture recognition have attracted a substantial interest . In the deep learning era, estimating the 3D hand skeleton from a single image has achieved great success . Since the proposal of the popular parametric hand model MANO and various large scale datasets , reconstructing both hand pose and shape has become a mainstream approach. Among all of these methods, the most recent transformer–based models yield the best results, demonstrating the ability of the attention mechanism to learn the nonlocal relationship between any two vertices. This excellent performance inspires us to use the attention mechanism to improve mesh-image alignment and model mesh-mesh interaction.

2 Two-Hand Reconstruction

Although nearly all single-hand reconstruction methods could extend to two-hand reconstruction tasks, few works demonstrate a result for close interacting hands. Two-hand reconstruction is one of the key challenges for human total motion capture. Previous body and hand simultaneous reconstruction methods all treat each hand in a separate manner and thus cannot handle close hand interaction cases such as finger knots. A recent multiview tracking based method could reconstruct high-quality interactive hand motions, however, its hardware setup is expensive, and the algorithm is time-consuming. Monocular kinematic tracking based two-hand motion estimation methods, regardless of whether a depth sensor or an RGB camera is incorporated, are sensitive to fast motion and possible tracking failure. However, their dense mapping strategy, which queries correspondences between hand vertices and image pixels, inspires us to seek mesh-image alignment using dense features. In contrast, deep learning based methods such as directly reconstruct per-frame two-hand interaction. Unfortunately, all of these methods either employ 2.5D heatmaps to estimate hand joint positions , or use them as attention maps to extract sparse image features , or reconstruct each hand respectively and fine-tune later . As hands are naturally 3D surfaces, sparse local image features encoded in the heatmaps may not effectively capture hand surface occlusions and hands interaction context. Therefore, the mentioned methods usually fail to obtain two-hand reconstructions well aligned to images.

3 Convolutional Mesh Regression

Convolutional mesh regression (CMR), which directly regresses mesh vertices in a coarse-to-fine manner from image features using a graph convolutional network (GCN), has been proven successful for generating image-aligned 3D objects , faces , bodies or hands . A typical CMR pipeline passes the global image feature vector through two or more cascaded graph convolution and upsampling layers and produces per-vertex 3D coordinates of the target object. Compared with joint based or rotation parameter based methods, the CMR method has denser and more semantic model representation and thus has the ability to better align image features in a per-vertex manner. However, existing CMR methods build a single forward pass without explicit image feature feedback strategy, limiting their mesh-image alignment performance as suggested by . Some recent single hand reconstruction works also employ GCN as part of the network structure; however, they discard the coarse-to-fine nature of CMR and simply use a single GCN to enhance local sensing ability.

Formulation

2 System Overview

Interacting Attention Graph

To directly produce two-hand vertices, our IntagHand is basically built upon previous GCN by extending one hand stream to two-hand streams. However, different from vanilla GCN which transforms the latent vector FGF_{G} to a larger unshared per-vertex feature, we utilize fully-connected (FC) layer gh(⋅)g_{h}(\cdot) to map FGF_{G} to a more compact feature vector gh(FG)g_{h}(F_{G}) which is shared across vertices, and concatenate dense matching encoding (positional embedding) cic_{i} of the ithi^{th} vertex with the shared vector to form per-vertex feature FViF_{V}^{i} (Fig. 2), which can be denoted as:

where L^t\hat{L}^{t} is the scaled Laplacian matrix, TktT_{k}^{t} is the kthk^{th} term of the K-order Chebyshev polynomial, WktW_{k}^{t} is the learnable parameter and σ\sigma is a nonlinear activation function. FGCNtF_{GCN}^{t} denotes the intermediate features that are passed to the PIFA module. Inspired by ResNet , we add residual connection for every two GraphConv operations to assist gradient propagation and enhance learning ability; see Fig. 3.

2 Pyramid Image Feature Attention Module

Although Mesh Graphformer utilizes image feature attention (called ‘grid feature’) as well, they use the same low–resolution image feature (7×77\times 7) in the whole network while our image features are multi-scale (8×8→16×16→32×328\times 8\rightarrow 16\times 16\rightarrow 32\times 32). While low–resolution image features encode more compact (or global) information, high–resolution features contain more semantic (or local) knowledge as they are closer to the input and output. Therefore, the pyramid structure forces the sparse mesh to attend to the global image features while the dense mesh to the local image features, and could yield better vertex-image alignment.

To demonstrate the function of PIFA, we compute the attention map between the vertex domain and the image domain (please refer to ViT for details). By adding up the PIFA attention maps of all three blocks, we observe that our PIFA module could distinguish between left and right hands on image pixels, and we note that PIFA pays more attention to the area of close interaction. That means PIFA module learns correct vertex-image mapping as we expect.

3 Cross Hand Attention Module

It has been shown that the poses of two interacting hands are correlated ; therefore, it is important to model hands interacting context for two-hand reconstruction. Instead of simply representing interaction as one hand’s joints in the other hand’s coordinate system , we use a symmetric cross-hand attention (CHA) module to implicitly formulate this correlation between two hands. For simplicity, we ignore tt in FPIFAtF_{PIFA}^{t} and use FLF_{L} and FRF_{R} to indicate FPIFAF_{PIFA} for the left and right hands, respectively.

As shown in Fig. 3, we first perform MHSA on each individual hand to get Qh,Kh,VhQ_{h},K_{h},V_{h} (h∈L,Rh\in{L,R}) indicating the query, key and value feature of each hand. Then we use the query feature QhQ_{h} of one hand to fetch the key feature KhK_{h} and the value feature VhV_{h} of the other hand through Multi-Head Attention (MHA, see Fig. 3) as

where FR→LF_{R\rightarrow L} and FL→RF_{L\rightarrow R} are the cross-hand attention features encoding the correlation between two hands, and dd is a normalization constant. Afterwards, the cross-hand attention features are merged into the hand vertex features by a pointwise MLP layer fp(⋅)f_{p}(\cdot) as

where FL′F_{L}^{{}^{\prime}} and FR′F_{R}^{{}^{\prime}} are the output hand vertex features, which act as FVt+1F_{V}^{t+1} of both hands for the next t+1tht+1^{th} block (t<Nbt<N_{b}).

It is shown in Fig. 4 that CHA also pays more attention to the closely interacting area, especially the finger-tips. This indicates that the CHA module helps to address mutual collision between hands implicitly.

4 Loss Functions

For training the image encoder-decoder, we use smooth L1 loss to supervise the 2D dense matching encoding and mean square error (MSE) loss to supervise 2D heatmaps.

For training IntagHand, we utilize (1) vertex loss, (2) regressed joint loss and (3) mesh smooth loss.

Vertex Loss. We use L1 loss to supervise the 3D coordinates of hand vertices and MSE loss to supervise the 2D projection of vertices:

where Vh,iV_{h,i} is ithi^{th} vertex, h=L,Rh=L,R means left or right hand, and Π\Pi is the 2D projection operation, the same below. Vertex loss is applied for each submesh, which we ignore here for simplicity.

Regressed Joint Loss. By multiplying the predefined joint regression matrix J\mathcal{J}, hand joints can be regressed from the predicted hand vertices. We penalize the joint error by the following loss:

Mesh Smooth Loss. To ensure the geometric smoothness of the predicted vertices, two different smooth losses are applied. First, we regularize the normal consistency between the predicted and the ground truth mesh:

where ff is the face index of the hand mesh, ef,i(i=1,2,3)e_{f,i}(i=1,2,3) are the three edges of face ff and nfGTn_{f}^{GT} is the normal vector of this face calculated from the ground truth mesh. Second, we minimize the L1 distance of each edge length between the predicted mesh and the ground truth mesh:

Note that, both the image encoder-decoder and the IntagHand are trained simultaneously in an end-to-end manner.

Experiments

Implementation Details. Our network is implemented using PyTorch. We use ResNet50 pretrained on ImageNet as backbone to encode the image feature. Following , our image decoders utilize three simple deconvolutional layers to predict 2D joint heatmaps, 2D segmentations and dense mapping encodings.

Training Details. We train our model using the Adam optimizer on 4 NVIDIA RTX 2080Ti GPUs with the minibatch size for each GPU set as 32. The whole training takes 100 epochs across 2.5 days, with the learning rate decaying to 1×10−51\times 10^{-5} at 50th50^{th} epoch from the initial rate 1×10−41\times 10^{-4}. During training, data augmentations including scaling, rotation, random horizontal flip and color jittering are applied. Note that, we pretrain the last upsampling layer of GCN (see Fig. 2) using posed MANO meshes and fix its weights during further training.

Evaluation Metrics. To evaluate both the pose and shape accuracy of reconstructed hands, we compare the Mean Per Joint Position Error (MPJPE) and Mean Per Vertex Position Error (MPVPE) in millimeters. For fair comparison, we follow Zhang et al. to scale the length of the middle metacarpal of each hand to 9.5cm9.5cm during training and rescale it back to the ground truth bone length during evaluation. This is performed after root joint alignment of each hand. We also report the Percentage of Correct Keypoints (PCK) curve and Area Under the Curve (AUC) across linearly spanned thresholds between 0 and 50 millimeters to compare reconstruction accuracy.

2 Datasets

InterHand2.6M Dataset. As the only dataset with two-hand mesh annotation, all networks in this paper are trained on InterHand2.6M datasetWe use the v1.0_5fps version of the InterHand2.6M which is CC-BY-NC 4.0 licensed.. Because we only focus on two-hand reconstruction, we pick out the interacting two-hand (IH) data with both human and machine (H+M) annotated, and discard invalid labeling according to the hand_type_validhand\_type\_valid annotation provided by . Ultimately, 366K training samples and 261K testing samples from InterHand2.6M are utilized. At preprocessing, we crop out the hand region according to the 2D projection of hand vertices and resize it to 256×256256\times 256 resolution.

RGB2Hands and EgoHands Datasets. RGB2Hands dataset consists of 4 sequences of videos with different types of two-hand interactions, and EgoHands dataset contains 48 egocentric videos capturing complex two people interactions such as playing chess. Both datasets have no mesh annotation, therefore we only use them for qualitative evaluation.

3 Qualitative Results

Our qualitative results on InterHand2.6M are shown in Fig. 5 and Fig. 6. As shown in Fig. 5, our method could generate high quality two-hand reconstruction results under severe occlusions and various kinds of interaction context. Compared with previous state-of-the-art method , our method produces more realistic finger interactions and less mutual collisions of two hands (see Fig. 6).

Beyond existing methods which only show results in dome setting , we further demonstrate the generalization ability of our method on in-the-wild images. As shown in Fig. 7, our method performs well on our real-life data captured by a common USB camera. Besides, without additional training, our model yields excellent results on images from RGB2Hands dataset and EgoHands dataset, showing the potential to be applied in both third/egocentric viewpoint conditions. Moreover, our model runs at 30fps on single NVIDIA RTX 3090 GPU during inference, which enables future real-time applications.

4 Quantitative Comparisons

We first compare our IntagHand network with state-of-the-art single-hand reconstruction methods, as shown in Tab. 1. Within single-hand reconstruction methods, each hand is cropped from image by ground truth bounding box and processed separately. It is shown that reconstructing each hand individually works poorly due to heavy occlusion and appearance confusion.

We further compare IntagHand with recent two-hand reconstruction methods. One is Moon et al. which regresses 3D skeletons of two hands directly. Another is Zhang et al. , which predicts the pose and shape parameters of two MANO models. For a fair comparison, we run their released source code on the same subset of InterHand2.6M to that we utilize (see Sec. 5.1). Comparison results are shown in Tab. 1 and Fig. 8. It is clearly shown in Tab. 1 that our method significantly reduces MPJPE and MPVPE. We attribute this success to the dense mesh reasoning ability of GCN and our novel attention based modules, which better align the mesh with the input image. The PCK curve in Fig. 8 further demonstrates the superior performance of our method at all error threshold levels.

5 Ablation study

Baseline GCN. We train a baseline GCN model by directly modifying the GCN decoder of Ge et al. for two-hand output (see ‘GCN baseline’ in Tab. 2). Although directly leveraging GCN structure shows excellent numeric performance, inaccurate interaction reconstructions still exist without the attention modules.

Adding Attention Modules. Based on ‘GCN baseline’, we first add the CHA module to model interaction context (+CHA) and then add the PIFA module to further enhance vertex-mesh alignment (+CHA +PIFA), as shown in Tab. 2. By modeling interaction context with CHA, we achieve more than 0.6mm performance gain, proving the effectiveness of CHA for occlusion handling. By adding PIFA, our method further achieves more than 0.5mm performance improvement, affirming the ability of PIFA for vertex-image alignment. Qualitative comparison is shown in Fig. 9.

Pyramid or Not. Note that our model utilizes pyramid image features with increasing resolutions (8×8→16×16→32×328\times 8\rightarrow 16\times 16\rightarrow 32\times 32). By removing pyramid structure, we use the consistent small (8×88\times 8) or large (32×3232\times 32) image features in all the three IntagHand blocks (see ‘IFA-8’ and ‘IFA-32’ in Tab. 2). Similar to , we find that using the small image feature performs better than using the large one. More importantly, our pyramid structure further improves the reconstruction accuracy by leveraging both the global and local information for mesh regression.

Discussion

Conclusion. We present the interacting attention graph hand (IntagHand) method to reconstruct two interacting hands from a single RGB image. Specifically, we introduce a novel pyramid image feature attention (PIFA) module to formulate the attention relationship between hand meshes and image features, together with a novel cross-hand attention (CHA) module to encode the interaction context between two hands. Comprehensive experiments demonstrate the supreme performance of our network on InterHand2.6M dataset and in-the-wild images, and verify the effectiveness of our PIFA and CHA modules.

Limitation & Impact. The major limitation of our method is the absence of explicit mesh collision handling, resulting in occasional mesh intersections between hands. Note that, it is possible for our method to work with more than two hands where a preliminary detection network is necessary to extract hand regions and predict the hand number in each region. It is also possible to extend our method to other 2-way interactions (hand-object, human-human, etc.) as long as the subjects are encoded as vertex feature.

Acknowledgement: This paper is sponsored by NSFC No.62125107, NSFC No.62171255 and National Key R&D Program of China (2021ZD0113503).

References