Reconstructing Interacting Hands with Interaction Prior from Monocular Images

Binghui Zuo, Zimeng Zhao, Wenqian Sun, Wei Xie, Zhou Xue, Yangang Wang

Introduction

Reconstruction of interacting hands is significant for enhancing the behavioral realism of digital avatars in communication, thinking and working. With the advent of the RGB dataset recording two-hand interactions, numerous attempts have been implemented to reconstruct interacting hands from monocular RGB images. Inspired by the existing single-hand frameworks , pioneer works localize and identify all two-hand joints as the interacting clues. Unfortunately, this process can be seriously misguided by the regional occlusion and local similarity between hands. Subsequent improvements include optimizing re-projection errors , localizing mesh vertices from coarse to fine , and querying all-in-one visible heatmap . Nevertheless, they still rely on more accurate joint 2D estimators, more diverse marker-less training data, or more computational complexity.

To overcome this hurdle, our key idea is to first construct a comprehensive interaction prior with multimodal datasets and then sample this pre-built prior according to the interaction cues extracted from a monocular image. It is noted that existing frameworks are always trained with paired data of calibrated images and mesh annotations. This may lead to difficulty in generalizing since the well-known benchmark contains simple backgrounds and only around 8.5K interaction patterns. We break this images-paired manner and construct an interaction prior with multimodal datasets, including marker-based data, marker-less data and hands-object data. To do this, a dataset with 500K two-hand patterns is proposed, which contains physically plausible 3D hand joints and MANO parameters. This dataset is used for the unsupervised training of a prior container, which can be formulated by a VAE . As a result, each two-hand interaction pattern is mapped to an interaction code in the prior space. Since the correlation between the two hands is considered, this representation is more compact than doubling the hand joint/vertex positions or MANO parameters .

We argue that accurate joint localization is challenging for the monocular reconstruction of interacting hands. As an alternative, we sample the above pre-built interaction prior according to the interaction adjacency heatmap (IAH). This heatmap is defined as the mixture coordinate distribution of this joint and other two-hand joints within its coordinate neighborhood. Compared with the 2.5D joint heatmap , our IAH abandons the pseudo depth and concatenates more on spatial correlations of the target joint. This heatmap formulation is easier to regress because even for an invisible joint, humans can determine its identity and location according to its spatial neighborhood. Considering that the Gaussian distribution has a more ambiguous boundary, the Laplacian distribution is selected as the kernel function of each joint. This effectively reduces the aliasing of interacting adjacency information. These IAHs are further converted to be the corresponding interaction codes through the ViT module and then are regarded as conditions to sample reasonable interaction from the latent space.

∙\bullet A powerful interaction reconstruction framework that compactly represents two-hand patterns as latent codes, which are learned from multimodal datasets in an unsupervised manner.

∙\bullet An effective feature extraction strategy that utilizes interaction adjacency as clues to identify each joint, which is inspired by human perception and is more friendly for network learning.

∙\bullet A large-scale multimodal dataset that records 500K patterns of closely interacting hands, which is more conducive to the construction of our latent prior space.

Related Work

Hand monocular reconstruction has received a breakthrough after more than a decade of development. Previous studies have always focused on 3D joint estimation . introduced the first baseline to predict 3D joint position from a single image. The proposal of MANO brings a new research direction to this field. With presenting the first end-to-end solution for learning 3D hand shape and pose from a monocular RGB image, estimating model parameters from inputs has become a major trend. Some researches even directly regressed 3D vertices from the inputs. However, most of them are only applicable to a single-pose representation. In this work, we construct an interaction prior which is used to effectively estimate plausible hand poses. Our proposed unified framework can be applied to 3D joints, vertices and MANO parameters.

Interacting hand reconstruction is critical to promote the development of human-computer interaction(HCI). Due to the severe self-occlusion and the similar appearance of the entangled two hands, previous works heavily relied on depth cameras or multi-view cameras . Benefiting from the promotion of deep learning and the proposal of interacting hand dataset , previous works tried to estimate interacting hand pose from monocular color images. Most of them attempted to extract distinguishable features of each hand or decouple the interaction . Unfortunately, the traditional feature extraction schemes are unsuitable for extracting effective interaction details, and it is unreliable to decouple hands relying on features. Recently, an attention mechanism has been widely adopted to yield more interacting attention features. Among them, Zhang et al. utilized pose-aware attention and context-aware refinement module to improve the pose accuracy. Hampali et al. employed a transformer architecture to model interaction, but the joint angle representation did not perform perfectly. Li et al. fed the bundled features to the attention module, progressively regressing the two-hand vertices. Besides, further demonstrated the power of the transformer. Inspired by the above studies, we propose a novel feature extraction module that gives more attention to the interaction region. Meanwhile, a ViT-based model is used to fuse them.

Learning-based prior has sparked more attention in different domains. Both VAE and GAN are the mainstream models for constructing diverse priors. Most related researchers constructed a prior to avoid implausible situations. To model human motion, used VAE to build motion prior. Other pioneers introduced pose or shape prior to refine geometric details. Similarly, designed adversarial prior to learn plausible reconstruction. Most similar to us are , who applied the built prior to hand pose estimation. Wan et al. combined GANs and VAEs to build two separate latent spaces. To estimate hand pose from depth maps, they used a mapping function to connect these two spaces. Spurr et al. proposed a cross-modal framework where both RGB images and depth maps can be used to build the prior. In addition, committed to aligning the learned latent space with different modalities jointly. To this end, we extend the VAE framework to build a novel interaction prior. Compared with competitors, the biggest challenge of two-hand reconstruction is the lack of interaction status. Existing two-hand reconstruction methods are trained on with the paired images. The smaller number of interaction states (8.5K) limits their generalization performance. Fortunately, our core improvement is to construct prior without paired images, which means multimodal datasets can be applied. To compensate for the lack of interaction between two hands, we provide a larger two-hand dataset Two-hand 500K , which has more interaction states than .

Method

where the first term is the Laplacian kernel for the identity ȷj\bm{\jmath}_{\text{j}} with the variance σj\sigma_{\text{j}}. The other terms are the Laplacian kernels for adjacent joints ȷk∈Aj(d)\bm{\jmath}_{\text{k}}\in\mathcal{A}_{\text{j}}(d) with the variance ασj\alpha\sigma_{\text{j}}. α>1\alpha>1 is a zoom factor. As shown in Fig. 3(c)&(d), we select Laplacian instead of Gaussian as the kernel function mainly due to its clearer distribution boundary. M\mathbf{M} contains the visible parts of the left and right hand. Because τ\bm{\tau} is often regarded as a global feature , it is regressed by an extra MLP. The overall loss term of this part is:

In practice, ResNet50 (with 4 cascading residual blocks) is selected as the feature extraction backbone, and the decoder for 2D local features is designed as a symmetrical structure with 4 blocks. For MLP used for regressing τ\bm{\tau}, after passing through an adaptive pooling layer behind the high-level feature map, we connect two fully connected layers to obtain τ\bm{\tau}.

Latent utilization. Besides the above explicitly supervised features, we further utilize the low-level feature maps F\mathbf{F} from the first block of our extractor as additional visual guidance. As shown in Fig. 2 (a), it contains more dense responses and the same map size as H\mathbf{H} and M\mathbf{M}.

Implementation details. In our experiment, the variance σj\sigma_{\text{j}} of identity ȷj\bm{\jmath}_{\text{j}} is set to 2.0, the zoom factor α\alpha is set to 2.0 and adjacent region size dd is set to 2.5. To balance each loss term, we set λ1\lambda_{1}=11 and λ2\lambda_{2}=20002000. To normalize the translation, we fix the right translation to 0 and predict τ\bm{\tau} between the left to right hand. We use Pytorch to implement the feature extraction network and train it on a single NVIDIA GeForce RTX 3090. To update the network parameters, we use Adam optimizer with a fixed learning rate 1e-4. We set the batch size to 64 and total training iterations to 500K. Before training, we crop the interacting hand regions with the annotated hand 2D vertices coordinates and resize it to 256×\times256. To improve the generalization, we perform data augmentation, including random rotation, random flip and color blur.

2 Interacting State Sampling

Prior construction. Building the prior allows us to sample the reasonable interaction from the extracted features even if one of the hands is completely occluded. Therefore, the expressiveness and accuracy of the constructed prior directly affect the final reconstruction. Similar to , we deploy the VAE framework to build the interaction prior, which consists of an encoder and a decoder. The encoder implicitly maps the input x\bm{x} to p(z)p(\bm{z}) that conforms to the normal distribution, where p(z)p(\bm{z}) is the prior on the latent space. The decoder reconstructs x^\hat{\bm{x}} that is close to x\bm{x}. We represent the encoder as the conditional probability distribution q(z∣x)q\left(\bm{z}\mid\bm{x}\right) and the decoder as p(x^∣z)p\left(\hat{\bm{x}}\mid\bm{z}\right). The building process is shown in Eqn. 3.

Training procedure. Both the encoder and decoder are composed of fully connected layers that follow ReLU activations. The encoder is a four-layers feed-forward network that models the input x\bm{x} to the latent space with dimension dzd_{z}. We force the output dimension of the encoder to 2dz2d_{z}, where the first dzd_{z} is used for the regression of mean μ\mu and the second for variance σ\sigma. We strive to shape the distribution with μ\mu and σ\sigma into a standard normal distribution and encourage it using Kullback-Leibler divergence loss. Afterward, the latent variable z\bm{z} sampled by the reparameterization trick is passed to the decoder to reconstruct the expected result x^\hat{\bm{x}}, where the decoder consists of six linear layers. We use MSE as the reconstruction loss to supervise it. The total loss for interaction prior is defined as:

In the inference phase, we discard the encoder and only use the decoder with the frozen parameters to obtain the reconstruction.

Hand representation. Benefiting from the embedding ability of VAE, we conduct experiments on three different hand pose representations, including 3D joints coordinates, 3D vertices coordinates and MANO parameters . For compatibility with our framework, only pose and shape parameters are considered when embedding MANO parameters, as the relative translation has been estimated in the feature extraction module.

Implementation details. To balance quality and generalization, we set loss weight λ3\lambda_{3} to 100100 to ensure losses are within one order of magnitude . During training, we flatten inputs as a vector and map them to latent space with dzd_{z}=128. We use Adam optimizer with a base learning rate of 1e-5 and a batch size of 64.

3 Interacting Feature Fusion

Feature alignment. To ensure the consistency between the dimension of ViT output and the pre-built interaction prior, we add a feature alignment block consisting of a linear layer after the final transformer block. We treat the output of the feature alignment block as a condition and sample the expected reconstruction from the pre-built interaction prior. Only the VAE decoder with frozen parameters is employed for this process.

Training procedure. We train the ViT-based fusion and simultaneously fine-tune the feature extraction module in an end-to-end manner. In addition to the feature extraction loss defined in Eqn. 2, we also apply other special losses to supervise the reconstruction of different hand representations. For the representation of 3D joints and 3D vertices, we use the L1L_{1} distance as a loss function to ensure consistency between the prediction and ground truth. And for the representation of MANO parameters, we also adopt additional loss terms to make the hand surface smoother and physically plausible, including normal loss and penetration loss .

Implementation details. In our framework, we set the size of each patch PP to 8, patch embeddings size DD to 1024 and total use NattnN_{attn}=66 transformer blocks. Different from feature extraction and interaction prior module, both the Adam optimizer and the learning rate scheduler are used. The training process costs 150 epochs with a batch size of 64 on a single NVIDIA GeForce RTX 3090. We fix the learning rate at 1e-4 in the first 100 epochs and then adaptively reduce it with the scheduler. Besides the above, as the parameters of the feature extraction module are also updated, we use the same data augmentation as the training of the feature extraction network.

4 Interacting Modality Expansion

Motivation. To eliminate the dependence on image datasets when building the interaction prior, we propose the following measures to increase the diversity of hand interaction: (i) Capturing skeleton data according to the marker-based system simultaneously; (ii) Randomly combining the poses sampled from single-hand datasets. Fig. 4 shows our Two-hand 500K .

More diversity. To obtain more diverse interaction states, we construct Two-hand 500K in a multimodal manner. With the above data generation measures, more than 500K interaction states are captured. Besides the data captured by the marker-based MoCap system, we splice left-right hand instances sampled from single-hand datasets . It should be noted that although sufficient MoCap data could be collected, considering the cost of the MoCap system, it is still meaningful to utilize single-hand data. Fig. 5 (a) visualizes the distribution of related two hand datasets and Two-hand 500K , showing that our proposed dataset is more diverse than them.

Less penetration. For the marker-based data, we fit MANO parameters from 3D skeleton by solving inverse kinematics (IK). As the fingers are often tangled together, penetration and dysmorphism always exist between two hands. To ensure physical interaction, we use the physics engine to optimize the fitted hand pose. Similar to , we adopt a sampling-based optimization scheme to iteratively refine interaction. The same strategy is also applied to the splicing process to ensure the assemblies are plausible. Fig. 5 (b) demonstrates the interaction before and after optimization. For more details about our dataset, please refer to Sup. Mat .

Experiments

Prior data. Data from three different domains are used to construct interaction prior in our practice: marker-based data (Two-hand 500K ), marker-less data and hands-object data . Since the relative translation is estimated by the network, all two-hand data in can be used to construct interaction prior, even if the hands are not strictly interacting. We use the physics engine to process implausible interactions in the dataset.

Reconstruction data. We only use to train the procedure from image to reconstruction. Before training, we pick out interacting instances and corresponding labels annotated by human and machine (H+M), containing 366K training instances and 261K testing instances. Interacting subjects in are only employed to demonstrate qualitative performance.

Metrics. To evaluate the accuracy of hand pose, we use the mean per joint position error (MPJPE) in millimeters. For fair comparisons, we calculate the MPJPE after aligning the root joints and scaling the bone lengths of each hand . Apart from that, the percentage of correct keypoints (PCK) and the area under the curve (AUC) in the range of 0 to 50 millimeters are taken to assess the evaluations. We also report the mean per vertex position error (MPVPE) to evaluate the quality of the reconstructed hand surface.

2 Comparison with the SOTA

We use the hand pose representation of MANO parameters to report the comparison results. In fairness, the following comparisons are performed on the basis of constructing interaction prior with only the (ALL) branch of .

Qualitative results. Comprehensive comparisons show the satisfactory performance of our method. Fig. 7 demonstrates the qualitative results compared with previous SOTA methods for interacting hands reconstruction. Compared to them, our reconstructed interaction generates less penetration while ensuring fidelity, which means that the constructed interaction prior effectively addresses the problem caused by severe occlusion and homogeneous appearance. Fig. 8 further demonstrates the qualitative analysis of the reconstruction. For the completely occluded instances in the gray background, the reconstructed interaction is also consistent with the ground truth (shown on the top right). More results on and can be seen in Sup. Mat .

Quantitative results. The summarized results shown in Tab. 1 indicate the superior performance of our method. From the first four rows in Tab. 1, we present the reconstruction performance of single-hand reconstruction methods . The poor performance suggests that applying the single-hand reconstruction method directly to our task is undesirable due to heavy self-occlusion and appearance confusion. We further compare with almost all recent two-hand reconstruction methods in the community. In exception to them, is not considered because they actually estimate single hand by interacting hand de-occlusion and removal. Compared to our full model, the interaction prior corresponding to w/Inter2.6M is only trained with . It can be concluded that non-image paired multimodal training data is positive for constructing interaction prior. Fig. 6 depicts the PCK curve of our method, which is also superior to other methods.

3 Ablation Study

We use the hand pose representation of MANO parameters to ablate the effectiveness of each component, but also report the accuracy of other pose representations with the same configuration.

Effectiveness of feature extraction. When extracting interacting features from inputs, we do not overemphasize extracting the unique features for each hand, as the self-occlusion between interacting hands makes it difficult. We design a novel feature extraction module that reflects more global-local context information and report the effectiveness in Tab. 2. We first perform ablation by removing the extracted local features and only using the global features F\mathbf{F} to reflect all information. The poor performance in Row.a shows the importance of the feature extraction module, indicating that the extracted features provide more interacting clues. We further ablate the impact of each part in the feature extraction module (Row.b and Row.c) by discarding the corresponding part separately. The visualization about the effectiveness of feature extraction is shown in Fig. 9 (b).

Effectiveness of IAH. To demonstrate that IAH is more suitable for interacting hands reconstruction, we express the heatmaps with different forms and list the corresponding effects. As shown in Row.d and Row.e of Tab. 2, although the conventional heatmaps help to reduce the errors, the impact on reconstruction is more marginal than our proposed IAH. We attribute this success to the adaptability of IAH. Row.f investigates the IAH with Gaussian distribution, and the inferior performance suggests that the Laplacian distribution is more suitable for IAH. More ablations can be found in Sup. Mat .

Effectiveness of prior. Benefiting from the constructed interaction prior, unreal interaction states have been excluded. We analyze the impact of the interaction prior by replacing it with an MLP architecture . Row.g in Tab. 2 shows the result without interaction prior. The unsatisfactory performance highlights the importance of interaction prior, denoting it contributes more to improving accuracy than the feature extraction module. Besides, Fig. 9 (c) further demonstrates the significant decrease in mesh quality.

Effectiveness of ViT-based fusion. The effects of both the extracted features and pre-built prior have been analyzed. Maximizing the performance of extracted features and accurately sampling the constructed prior are critical to the final reconstruction. We compare two other powerful backbones ResNet50 and HRNet32 , each of which is applied to completely replace ViT. Comparing Row.h and Row.i in Tab. 2, we see that the ViT-based network gives more powers when fusing interaction features. That is because ViT obtains more global-local context information and effectively models the interactions between two hands. It is noted that only 6 ViT blocks are adopted in the experiments, making the parameter count comparable among different networks and ensuring the fairness of the ablations.

Influence of different representations. We further compare the interaction priors constructed separately with 3D hand joints, 3D hand vertices and MANO parameters within the same dimension. The corresponding performance is reported in Row.j and Row.k of Tab. 2. Among them, the best performance is achieved by the MANO representation, while the lowest accuracy occurs in the representation of 3D vertices. We attribute the reason to the self-restriction of MANO parameters and it is more difficult to embed discrete 3D coordinates.

Discussion on prior structure. We discuss two candidate prior structures: auto-encoder (AE) and variational auto-encoder (VAE). AE is data-dependent and can not generate data. While VAE drives the latent variable to conform to the standard normal distribution and uses reparameterization tricks to improve generativity and robustness , which is more reasonable to construct interaction prior.

Conclusion

This work treats the interacting hands as a whole, constructs interaction prior based on multimodal datasets, and utilizes joint-wise interaction adjacency to reconstruct interacting hands from monocular images. Compared to most existing works, our framework elegantly combines multimodal datasets to build interaction prior and further recasts the reconstruction as the conditional sampling from this prior. To facilitate its training, Two-hand 500K dataset is further constructed with modal diversity and physical plausibility considerations. Our framework based on cross-modal interaction prior would also bring inspiration to other multi-body reconstruction tasks.

Limitations and Future Work. Although we can obtain reasonable interaction from the constructed interaction prior, the penetration is still unavoidable for complicated entanglements. In the future, using the physics engine to guide interaction could bring more benefits to the community.

References