CHORE: Contact, Human and Object REconstruction from a single RGB image

Xianghui Xie, Bharat Lal Bhatnagar, Gerard Pons-Moll

Introduction

In order to deploy robots and intelligent systems in the real world, they must be able to perceive and understand humans interacting with the real world from visual input. While there exists a vast literature in perceiving humans in 3D from single images, the large majority of works perceive humans in isolation . The joint 3D perception of humans and objects has received much less attention and is the main focus of this work. Joint reconstruction of humans and objects is extremely challenging. The object and human occlude each other making inference hard; it is difficult to predict their relative size and spatial arrangement in the 3D world, and the visual evidence for contacts and interactions consists of only a very small region in the image.

Recent work addresses these challenges by reconstructing 3D objects and humans separately, and imposing hand-crafted rules like manually defined object contact regions and object size. However, heuristics cannot scale as it is difficult to specify all possible human-object contacts beforehand. Consider the chair in Fig. 3: one can sit, lean, move it from the handle, back or legs, which constitutes different pairs of human-object contacts that are not easy to manually enumerate. Furthermore, reasoning human and object separately can produce errors that cannot be corrected afterwards, see Fig. 4. These motivate a method that can jointly reconstruct humans and objects and learn the interactions from data, without hand-crafted rules.

In this work, we introduce CHORE, a novel method to jointly recover a 3D parametric body model and the 3D object, from a single image. CHORE is built on two high-level ideas. First, instead of reconstructing human and object separately, we argue for jointly reason them in the first place. Secondly, we learn spatial configuration priors of interacting human and object from data without using heuristics. Instead of regressing human parameters directly from image, we first obtain neural distance fields to human and object surfaces as well as a correspondence field to the SMPL body model. To obtain robust object fitting, we also predict a neural field for object rotation and translation. This allows us to formulate a more robust optimization energy to fit the SMPL model and the object template to the image. To train our network on real data, we propose a training strategy that enables us to effectively learn pixel-aligned implicit functions with perspective cameras, which leads to more accurate prediction and joint fitting. In summary, our key contributions are:

We propose CHORE, the first end-to-end learning approach that can reconstruct human, object and contacts from a single RGB image. Our correspondence and contact prediction allow us to accurately register a controllable body model and 3D object to the image.

Different from prior works that use weak perspective cameras and learn from synthetic data, we use perspective camera model, which is crucial to train on real data. To this end, we propose a new training strategy that allows effective pixel-aligned implicit learning with perspective cameras.

Through our effective training and joint reconstruction, our model achieves a remarkable performance improvement of over 50%50\% compared to prior art . Our code and models are publicly available to foster future research.

Related Work

We propose the first learning-based approach to jointly reconstruct 3D human, object and contacts from a single image. In this regard, we first review recent advances in separate human or object reconstructions. We then discuss approaches that can jointly reason about human-object interaction.

Human reconstruction from images. One common approach to reconstruct 3D humans from images is to fit a parametric model such as SMPL to the image . Learning-based approaches have also been used to directly regress model parameters such as pose and shape as well as clothing . More recently, implicit function based reconstruction are applied to capture fine-grained details including clothing and hair geometry . These methods however cannot reconstruct 3D objects, let alone human-object interaction.

Object reconstruction from images. Given a single image, 3D objects can be reconstructed as voxels , point clouds and meshes . Similar to human reconstruction, implicit functions have shown huge success for object reconstruction as well . The research in this direction has been significantly aided by large scale datasets . We urge the readers to see the recent review of image-based 3D reconstruction of objects. These approaches have shown great promise in rigid object reconstruction but articulated humans are significantly more challenging and human-object interaction even more so. Unlike these methods, our approach can jointly reconstruct humans, objects and their interactions.

Human object interactions. Modelling 3D human-object interactions is extremely challenging. Recent works have shown impressive performance in modelling hand-object interactions from 3D , 2.5D and images . Despite being quite impressive, these works have a shortcoming in that they are restricted to only hand-object interactions and cannot predict the full body. Modelling full-body interactions is even more challenging, with works like PROX and being able to fit 3D human model to satisfy the 3D scene constraints , capture interaction from mult-view or reconstruct 3D scene based on human-scene interactions . Recently, this direction has also been extended to model human-human interaction , self-contacts . Broadly speaking, these works can model human-object interactions to to satisfy scene constraints but cannot jointly reconstruct human-object contacts from single images. Few works directly reason about 3D contacts from images and video but do not reconstruct 3D human and object from a single RGB image.

Closest to our work are and PHOSA . Weng et al. predict 3D human and scene layout separately and use scene constraints to optimize human reconstruction. They use predefined contact weights on SMPL-X vertices to optimize human meshes, which is prone to errors. PHOSA first fits SMPL and object meshes separately and then use hand crafted heuristics like predefined contact pairs to reason human object interaction. These heuristics are neither scale-able nor very accurate. On the other hand, our method does not rely on heuristics and can learn joint reconstruction and contact priors directly from data. Our experiments show that our learning based method easily outperforms heuristic-based PHOSA and .

Method

We describe here the components of CHORE, a method for 3D reconstruction of the human, object and their contacts from a single RGB image. This work focuses on a single person in close interaction with one dynamic movable object. Our approach is inspired by the classical image based model fitting methods , which fit a body model parameterized by shape and pose θ,β\mathbf{\boldsymbol{\theta},\boldsymbol{\beta}} by minimizing an energy:

where EdataE_{data} is a 2D image based loss like a silhouette loss, EJ2DE_{J2D} is re-projection loss for body joints, and EregE_{reg} is a body pose and shape prior loss. Estimating the human and the object jointly however is significantly more challenging. The relative configuration of human and object needs to be accurately estimated despite occlusions, lack of depth, and hard to detect contacts. Hence, a plain extension of a hand-crafted 2D objective to incorporate the object is far too prone to local minima. Our idea is to jointly learn to predict, based on a single image, 3D neural fields (CHORE fields) for the object, human and their relative pose and contacts, to formulate a learned robust 3D data term. We then fit a parametric body model and a template object mesh to the predicted CHORE neural fields, see overview Fig. 2.

In this section, we first explain our CHORE neural fields in Sec. 3.1, and then describe our novel human-object fitting formulation using learned CHORE neural fields in Sec. 3.2. Unlike prior work on pixel-aligned neural implicit reconstruction (PiFu ) which learns from synthetic data and assume a weak camera model, we learn from real data captured using perspective cameras. This poses new challenges which we address with a simple and effective deformation of the captured real 3D surfaces to account for perspective effects (Sec. 3.3).

As representation, we use the SMPL body model H(θ,β)H(\boldsymbol{\theta},\boldsymbol{\beta}) which parameterizes the 3D human as a function of pose θ\boldsymbol{\theta} (including global translation) and shape β\boldsymbol{\beta}, and a template mesh for the object. Our idea is to leverage a stronger 3D data term coming from learned CHORE neural fields to fit SMPL and the object template to the image. In contrast to a plain joint surface reconstruction of human and object, the resulting SMPL and object mesh can be further controlled and edited. To obtain a good fit to the image, we would ideally i) minimize the distance between the optimized meshes and the corresponding ground truth human and object surfaces, and ii) enforce that human-object contacts of the meshes are consistent with input. Obviously, ground truth human and object surfaces are unknown at test time. Hence, we approximate them with CHORE fields. Specifically, a learned neural model predicts 1) the unsigned distance fields of the object and human surfaces, and 2) a part correspondence field to the SMPL model, for more robust SMPL fitting and contact modeling, and 3) an object pose field, which we use to initialize the object pose.

where fiuf^{u}_{i} can be human (fhuf^{u}_{h}) or object (fouf^{u}_{o}) UDF predictor. In practice, points p\mathbf{p} are initialized from a fixed 3D volume and projection is done iteratively.

To train neural distance function fiuf^{u}_{i}, we sample a set of points Pi\mathcal{P}_{i} near the ground truth surfaces and compute the ground truth distance UDFgt(p)\text{UDF}_{\text{gt}}(\mathbf{p}). The decoder fiuf^{u}_{i} is then trained to minimize the L1L_{1} distance between clamped prediction and ground truth distance:

where δ\delta is a small clamping value to focus on the volume nearby the surface.

Our feature encoder fencf^{enc} and neural predictors fu,fp,fR,fcf^{u},f^{p},f^{R},f^{c} are trained jointly with the objective: L=λu(Luh+Luo)+λpLp+λRLR+λcLcL=\lambda_{u}(L_{u_{h}}+L_{u_{o}})+\lambda_{p}L_{p}+\lambda_{R}L_{R}+\lambda_{c}L_{c}. See supplementary for more details on our network architecture and training.

2 3D human and object fitting with CHORE fields

With CHORE fields, we reformulate Eq. 1 to include a stronger data term EdataE_{data} in 3D, itself composed of 3 different terms for human, object, and contacts which we describe next.

3D human term. This term consists of the sum of distances from SMPL vertices to the CHORE neural UDF for the human fhuf^{u}_{h}, and a part based term leveraging the part fields fpf^{p} discussed in Sec. 3.1:

Here, Lp(⋅,⋅)L_{p}(\cdot,\cdot) is the standard categorical cross entropy loss function and lpl_{\mathbf{p}} denotes our predefined part label on SMPL vertex p\mathbf{p}. λh,λp′\lambda_{h},\lambda_{p\prime} are the loss weights. The part-based term Lp(⋅,⋅)L_{p}(\cdot,\cdot) is effective to avoid, for example, matching a person’s arm to the torso or a hand to the head. To favour convergence, we initialize the human pose estimated from the image using .

The first term computes the distance between fitted object and corresponding surface represented implicitly with the network fouf_{o}^{u}. The second term Locc-silL_{\text{occ-sil}} is an occlusion-aware silhouette loss . We additionally use LregL_{\text{reg}} computed as the distance between predicted object center and center of O′\mathbf{O}^{\prime}. This regularization prevents the object from being pushed too far away from predicted object center.

The object rotation is initialized by first averaging the rotation matrices of the object pose field near the object surface (points with fou(Fp)<4mmf^{u}_{o}(\mathbf{F}_{\mathbf{p}})<4mm): Ro=1M∑k=1MRk\mathbf{R}_{o}=\frac{1}{M}\sum^{M}_{k=1}\mathbf{R}_{k}. We further run SVD on Ro\mathbf{R}_{o} to project the matrix to SO(3)SO(3). Object translation to\mathbf{t}_{o} is similarly initialized by averaging the translation part of the pose field nearby the surface. Our experiments show that pose initialization is important for accurate object fitting, see Sec. 4.6.

Joint fitting with contacts. While Eq. 4 and Eq. 5 allow fitting human and object meshes coherently, fits will not necessarily satisfy fine-grained contact constraints. Our key idea here is to jointly reconstruct and fit humans and objects while reasoning about contacts. We minimize a joint objective as follows:

where EdatahE_{\text{data}}^{h} and EdataoE_{\text{data}}^{o} are defined in Eq. 4 and Eq. 5 respectively. The contact term EdatacE_{\text{data}}^{c} consists of the chamfer distance d(⋅,⋅)d(\cdot,\cdot) between the sets of human HjcH_{j}^{c} and object points OjcO_{j}^{c} predicted to be in contact:

Points on the j−thj-th body part of SMPL (Hj(θ,β)H_{j}(\boldsymbol{\theta},\boldsymbol{\beta})) are considered to be in contact when their UDF fou(Fp)f^{u}_{o}(\mathbf{F}_{\mathbf{p}}) to the object is smaller than a threshold: Hjc(θ,β)={p ∣ p∈Hj(θ,β) and fou(Fp)≤ϵ}H_{j}^{c}(\boldsymbol{\theta},\boldsymbol{\beta})=\{\mathbf{p}\,|\,\mathbf{p}\in H_{j}(\boldsymbol{\theta},\boldsymbol{\beta})\,\textrm{and}\,f^{u}_{o}(\mathbf{F}_{\mathbf{p}})\leq\epsilon\}. Analogously, object points are considered to be in contact with body part jj when their UDF to the body is small and they are labelled as part jj: Ojc={p ∣ p∈O′ and fhu(Fp)≤ϵ and fp(Fp)=j}O_{j}^{c}=\{\mathbf{p}\,|\,\mathbf{p}\in\mathbf{O}^{\prime}\,\textrm{and}\,f^{u}_{h}(\mathbf{F}_{\mathbf{p}})\leq\epsilon\ \textrm{and}\,f^{p}(\mathbf{F}_{\mathbf{p}})=j\}. The contact term EdatacE_{\text{data}}^{c} encourages the object to be close with the corresponding SMPL parts in contact. The loss is zero when no contacts detected. This is important to have a physically plausible and consistent reconstruction, see Sec. 4.5.

Substituting our data term into Eq. 1, we can now write our complete optimization energy as: E(θ,β,Ro,to,so)=Edatao+Edatah+λcEdatac+λJEJ2D+λrEregE(\boldsymbol{\theta},\boldsymbol{\beta},\mathbf{R_{o}},\mathbf{t_{o}},s_{o})=E_{data}^{o}+E_{data}^{h}+\lambda_{c}E_{data}^{c}+\lambda_{\text{J}}E_{J2D}+\lambda_{r}E_{\text{reg}}. See Supp. for details about different loss weights λ\lambda’s.

3 Learning pixel-aligned neural fields on real data

Our training data consists of real images paired with reference SMPL and object meshes. Here, we want to follow the pixel-aligned training paradigm, which has been shown effective for 3D shape reconstruction. Since both depth and scale determine the size of objects in images this makes learning from single images ambiguous. Hence, existing works place the objects at a fixed distance to the camera, and render them using a weak perspective or perspective camera model . This effectively allows the network to reason about only scale, instead of both scale and depth at the same time.

This strategy to generate training pairs is only possible when learning from synthetic data, which allows to move objects and re-render the images. Instead, we learn from real images paired with meshes. Re-centering meshes to a fixed depth breaks the pixel alignment with the real images. An alternative is to adopt a weak perspective camera model (assume all meshes are centered at a constant depth), but the alignment is not accurate hence leads to inaccurate learning, see Fig. 3 b) and c). Alternatively, attempting to learn a model directly from the original pixel-aligned data with the intrinsic depth-scale ambiguity leads to poor results, as we show in Fig. 3 e). Hence, our idea is to adopt a perspective camera model, and transform the data such that it is centered at a fixed distance to the camera, while preserving pixel alignment with the real images. We demonstrate that this can be effectively achieved by applying depth-dependent scale factor.

Our first observation is that scaling the mesh vertices V\mathbf{V} by ss does not change its projection to the image. Proof: Let v=(vx,vy,vz)∈V\mathbf{v}=(\mathbf{v}_{x},\mathbf{v}_{y},\mathbf{v}_{z})\in\mathbf{V} be a vertex of the original mesh, and v′=(svx,svy,svz)∈sV\mathbf{v}^{\prime}=(s\mathbf{v}_{x},s\mathbf{v}_{y},s\mathbf{v}_{z})\in s\mathbf{V} a vertex of the scaled mesh. For a camera with perspective focal length fx,fyf_{x},f_{y} and camera center cx,cyc_{x},c_{y}, the vertex v\mathbf{v} is projected to the image as π(v)=(fxvxvz+cx,fyvyvz+cy)\pi(\mathbf{v})=(f_{x}\frac{\mathbf{v}_{x}}{\mathbf{v}_{z}}+c_{x},f_{y}\frac{\mathbf{v}_{y}}{\mathbf{v}_{z}}+c_{y}). The scaled vertices will be projected to the exact same pixels, π(v′)=(fxsvxsvz+cx,fysvysvz+cy)=π(v)\pi(\mathbf{v}^{\prime})=(f_{x}\frac{s\mathbf{v}_{x}}{s\mathbf{v}_{z}}+c_{x},f_{y}\frac{s\mathbf{v}_{y}}{s\mathbf{v}_{z}}+c_{y})=\pi(\mathbf{v}) because the scale cancels out. This result might appear un-intuitive at first. What happens is that the object size gets scaled, but its center is also scaled, hence making the object bigger/smaller and pushing it further/closer to the camera at the same time. This effects cancel out to produce the same projection.

Given this result, all that remains to be done, is to find a scale factor which centers objects at a fixed distance z0z_{0} to the camera. The desired scale factor that satisfies this condition is s=z0μzs=\frac{z_{0}}{\mu_{z}}, where μz=1n∑i=1nvzi\mu_{z}=\frac{1}{n}\sum_{i=1}^{n}\mathbf{v}^{i}_{z} is the mean of the un-scaled mesh vertices depth. Proof: The proof follows immediately from the linearity property of the empirical mean. The depth of the transformed mesh center is:

This shows that the depth μz′\mu_{z}^{\prime} of the scaled mesh center is fixed at z0z_{0} and its image projection remains unchanged, as desired.

We scale all meshes in our training set as described above to center them at z0=2.2mz_{0}=2.2m and recompute GT labels. To maintain pixel alignment we use the camera intrinsic from the dataset. At test time, the images have various resolution and unknown intrinsic with person at large range of depth (see Fig. 6). We hence crop and resize the human-object patch such that the person in the resized patch appears as if they are at z0z_{0} under the camera intrinsic we use during training. We leverage the human mesh reconstructed from as a reference to determine the patch resizing factor. Please see supplementary for details.

Experiments

To verify the generalization ability, we also tested our method on COCO where each image contains at least one instance of the object categories in BEHAVE and the object should not be occluded more than 50%. Images where there is no visible contact between human and object are also excluded. There are four overlapping categories between BEHAVE and COCO (backpack, suitcase, chair and sports ball) and in total 913 images are tested.

We additionally test on the NTU-RGBD dataset, which is quite interesting since it features human-object interactions. We select images from the NTU-RGBD dataset following the same criteria for COCO and in total 1.5k images of chair or backpack interactions are tested.

We also compare our method qualitatively with PHOSA and in Fig. 4. It can be seen that PHOSA fails to reconstruct the object at the correct position relative to the human. also fails quite often as it uses predefined contact weights on SMPL vertices instead of predicting contacts from inputs. On the other hand, CHORE works well because it learns a 3D prior over human-object configurations based on the input image. See supplementary for more examples.

2 Generalisation beyond BEHAVE

To verify our model trained on the BEHAVE dataset can generalize to other datasets, we further test it on the NTU-RGBD as well as COCO dataset and compare with .

Comparison with PHOSA. We test our method and PHOSA on the selected COCO images discussed above. We then randomly select 50 images and render ours and PHOSA reconstruction results side by side. The results are collected in an anonymous user study survey and we ask 50 people in Amazon Mechanical Turk (AMT) to select which reconstruction is better. Notably, people think our reconstruction is better than PHOSA on 72% of the images. A similar user study for images from the NTU-RGBD dataset is released to AMT and on average 84% of people prefer our reconstruction. Some example images are shown in Fig. 6 and Fig. 5. See supplementary for details on the user study and more qualitative examples.

Comparison with Weng et al. . One limitation of is that they use some indoor scene layout assumptions, which makes it unsuitable to compare on ’in the wild’ COCO images. Instead, we only test it on the NTU-RGBD dataset. Similar to the comparison with PHOSA, we released a user study on AMT and ask 50 users to select which reconstruction is better. Results show that our method is preferred by 94% people. Our method clearly outperforms , see two examples in Fig. 5 and more examples in supplementary.

3 Comparison with a human reconstruction method

Our method takes SMPL pose predicted from FrankMocap as initialisation and fine tunes it together with object using network predictions. We compare the human reconstruction performance with FrankMocap and PHOSA in Tab. 2. It can be seen that our joint reasoning of human and object does not deteriorate the human reconstruction. On the other hand, it slightly improves the human reconstruction as it takes the interacting object into account.

4 Pixel-aligned learning on real data

CHORE is trained with the depth-aware scaling strategy described in Sec. 3.3 for effective pixel aligned implicit surface learning on real data. To evaluate its importance, we compare with two baseline models trained with weak perspective and direct perspective (no depth-aware scaling on GT data) camera. Qualitative comparison of the neural reconstruction is shown in Fig. 3. We further fit SMPL and object meshes using these baseline models (with object pose prediction) on BEHAVE test set and compute the errors for reconstructed meshes, see Tab. 3b) and c). It can be clearly seen that our proposed training strategy allows effective learning of implicit functions and significantly improves the reconstruction.

5 Contacts modeling

We have argued that contacts between human and the object are crucial for accurately modelling interaction. Quantitatively, the contact term EdatacE^{c}_{data} only reduces the object error slightly (Tab. 3f) and g) as contacts usually constitute a small region of the full body. Nevertheless, it makes significant difference qualitatively, as can be seen in Fig. 7. Without the contact information, the object cannot be fitted correctly to the location where contact is happening. Our contact prediction helps snap the object to the correct contact location, leading to a more physically plausible and accurate reconstruction. Figure 7: Accurate human and object reconstruction requires modeling contacts explicitly. Without contact information, the human and the objects are not accurately aligned and the resulting reconstruction is not physically plausible. Methods SMPL ↓\downarrow Object ↓\downarrow a. neural recon. 14.06 15.34 b. weak persp. 9.14 25.24 c. direct persp. 7.63 15.58 d. w/o fRf^{R} 6.58 17.54 e. w/o fcf^{c} 6.26 14.25 f. w/o contact 5.51 10.86 g. full model 5.58 10.66 Table 3: Ablation studies. We can see that neither weak perspective (b) nor direct perspective camera (c) can work well on real data. Our proposed training strategy, together with rotation and translation prediction achieves the best performance.

6 Other ablation studies

We further ablate our design choices in Tab. 3. Note that all numbers in Tab. 3 are Chamfer distances (cm) and computed on meshes after joint fitting except a) which is computed on the neural reconstructed point clouds. We predict object rotation and translation to initialize and regularize object fitting, which are important for accurate reconstruction. In Tab. 3, it can be seen that fitting results without rotation d) or without translation initialization e) leads to worse perofrmance compared to our full model g). Without initialization the object gets stuck in local minima. See supplementary for qualitative examples.

We also compute the Chamfer distance of the neural reconstructed point clouds (Eq. 2) using alignments computed from fitted meshes. It can be seen that this result is much worse than the results after joint fitting (Tab. 3g). This is because network predictions can be noisy in some regions while our joint fitting incorporate priors of human and object making the model more robust.

Limitations and Future Work

We introduced the first learning based method to reconstruct 3D human and object from images. Although the results are promising, there are some limitations. For instance, we assume known object template for the objects shown in the images. Future works can remove this assumption by selecting template from 3D CAD database or reconstructing shape directly from image. Non-rigid deformation, multiple objects, multi-person interactions are exciting avenues for future research. Please also refer to supplementary for discussions of failure cases.

Conclusion

In this work we presented CHORE, the first method to learn a joint reconstruction of human, object and contacts from single images. CHORE combines powerful neural implicit predictions for human, object and its pose with robust model-based fitting. Unlike prior works on pixel-aligned implicit surface learning, we learn from real data instead of synthetic scans, which requires a different training strategy. We proposed a simple and elegant depth-aware scaling that addresses the depth-scale ambiguity while preserving pixel alignment.

Experiments on three different datasets demonstrate the superiority of CHORE compared to the SOTA. Quantitative experiments on BEHAVE dataset evidence that joint reasoning leads to a remarkable improvement of over 50%50\% in terms of Chamfer distance compared to PHOSA. User studies on the NTU-RGBD and the challenging COCO dataset (no ground truth) show that our method is preferred over PHOSA by 84% and 72% users respectively. We also conducted extensive ablations, which reveal the effectiveness of the different components of CHORE, and the proposed depth-aware scaling for pixel-aligned learning on real data. Our code and models are released to promote further research in this emerging field of single image 3D human-object interaction capture.

We would like to thank RVH group members for their helpful discussions. Special thanks to Beiyang Li for supplementary preparation. This work is funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) - 409792180 (Emmy Noether Programme, project: Real Virtual Humans), and German Federal Ministry of Education and Research (BMBF): Tübingen AI Center, FKZ: 01IS18039A. Gerard Pons-Moll is a Professor at the University of Tübingen endowed by the Carl Zeiss Foundation, at the Department of Computer Science and a member of the Machine Learning Cluster of Excellence, EXC number 2064/1 – Project number 390727645.

References