BEHAVE: Dataset and Method for Tracking Human Object Interactions

Bharat Lal Bhatnagar, Xianghui Xie, Ilya A. Petrov, Cristian Sminchisescu, Christian Theobalt, Gerard Pons-Moll

Introduction

The last decade has seen rapid progress in modelling the appearance of humans ranging from body pose, shape , faces and even detailed clothing . With various practical use cases like virtual try-on, personalised avatar creation, and several applications in augmented and mixed reality, or human-robot collaboration, the focus on humans is justified. Beyond modelling appearance, few methods have focused on capturing and synthesizing human interactions (human-object/scene interaction). There exists work to capture humans in a static 3D scene , even without using external cameras , and work to synthesize static poses , or full body movement in a 3D scene.

These methods show growing interest in modelling human behavior, highlighting a need to capture real human interactions. Existing methods however are learned from high quality curated data captured using optical marker based motion capture systems or wearable sensors. Unfortunately, such commercial systems are expensive, drastically limit the interactions that can be captured, and often fail when tracking humans and objects under occlusion. In addition, the recording volume is spatially confined and difficult to re-locate, thus limiting the activities, scenes, and objects that can be captured. Wearable sensors are not restricted in volume, but close range interaction can not be accurately captured. Altogether, the lack of diverse 3D interaction data, and the lack of accurate and flexible capture methods both constitute barriers in modelling human behavior.

With the goal of simplifying the data capture process and hence allowing faster progress in the field, we propose BEHAVE, a method to capture diverse 3D human interactions in natural environments, using a setup comprising of portable, cheap, and easy to use RGBD cameras. Tracking human interactions from sparse consumer grade cameras is however extremely challenging. Depth data is inherently noisy and incomplete. Moreover, the person and object occlude each other frequently during interactions. Furthermore, capturing interactions requires estimating human-object contacts accurately, which is difficult because contacts represent small regions in the image, close to the observable (resolution) limit. This requires innovation that goes significantly beyond the current state of the art trackers. We propose to track the human using a parametric human model (such as SMPL ) and track objects using template meshes. Naively fitting the human model and an object 3D template to the point-cloud completely fails due to the aforementioned challenges. Our key idea is to train a neural model which jointly completes the human and object shape, represented with implicit surfaces, while predicting a correspondence field to the human, as well as an object orientation field. These rich outputs allow us to formulate a powerful human-object fitting objective which is robust to missing data, noise and occlusion.

To train and evaluate BEHAVE, we capture the largest dataset of human-object interactions in natural environments. The BEHAVE dataset contains 20 3D objects, 8 subjects (5 male, 3 female), 5 different locations and totals around 15.2k frames of recording. We provide ground truth SMPL and 3D object meshes as well as contacts. Our contributions can be summarized as follows:

We propose the first approach that can accurately 3D track humans, objects and contacts in natural environments using multi-view RGBD images.

We collect the largest dataset of multi-view RGBD sequences and corresponding human models, object and contact annotations. See Sec. 3 for details regarding its usefulness to the community.

Since there exists no publicly available code and datasets to accurately track human-object interactions in natural environments, we will release our code and data for further research in this direction.

Related Work

In this section, we first briefly review work focused on object and human reconstruction, in isolation from their environmental context. Such methods focus on modelling appearance and do not consider interactions. Next, we cover methods focused on humans in static scenes and finally discuss closer-related work to ours, for modelling dynamic human-object interactions.

Perceiving humans from monocular RGB data and under multiple views settings has been widely explored. Recent work tends to focus on reconstructing fine details like hand gestures and facial expressions , self-contacts , interactions between humans, and even clothing . These methods benefit from representing human with parametric body models , thus motivating our use of recent implicit diffused representations as backbone for our tracker. Following the success of pixel-aligned implicit function learning , recent methods can capture human performance from sparse or even a single RGB camera . However, capturing 3D humans from RGB data involves a fundamental ambiguity between depth and scale. Therefore, recent methods use RGBD or volumetric data for reliable human capture. These insights motivate us to build novel trackers based on multi-view RGBD data.

Object reconstruction

Most existing work on reconstructing 3D objects from RGB and RGBD data does so in isolation, without the human involvement or the interaction. While challenging, it is arguably more interesting to reconstruct objects in a dynamic setting under severe occlusions from the human.

2 Interaction modelling: Humans and objects with scene context

Modelling how humans act in a scene is both important and challenging. Tasks like placement of humans into static scenes , motion prediction or human pose reconstruction under scene constrains, or learning priors for human-object interactions , have been investigated extensively in recent years. These methods are relevant but restricted to modelling humans interacting with static objects. We address a more challenging problem of jointly tracking human-object interactions in dynamic environments where objects are manipulated.

Dynamic human object interactions

Recently, there has been a strong push on modeling hand-object interactions based on 3D , 2.5D and 2D data. Although powerful, these methods are currently restricted to modelling only hand-object interactions. In contrast, we are interested in full body capture. Methods for dynamic full body human object interaction approach the problem via 2D action recognition or reconstruct 3D object trajectories during interactions . Despite being impressive, such methods either lack full 3D reasoning or are limited to specific objects . More recent work reconstructs and tracks human-object interactions from RGB or RGBD streams , but does not consider contact prediction, thus missing a component necessary for accurate interaction estimates. Very relevant to our work, PHOSA reconstructs humans and objects from a single image. PHOSA uses hand crafted heuristics, instance specific optimization for fitting, and pre-defined contact regions, which limits generalization to diverse human-object interactions. Our method on the other hand learns to predict the necessary information from data, making our models more scale-able. As shown in the experiments, the accuracy of our method is significantly higher to PHOSA.

BEHAVE Dataset

We present BEHAVE dataset, the largest dataset of human-object interactions in natural environments, with 3D human, object and contact annotation, to date. See Tab. 1 for comparison with other datasets. Our dataset contains multi-view RGBD frames, with accurate pseudo-ground truth SMPL , object fits, human and object segmentation masks, and contact annotations.

We setup and calibrate 4 Kinects at 4 corners of our square recording volume where all interactions are performed by 8 subjects (5 male, 3 female). Interactions are captured at 5 disparate indoor locations with 20 commonly used, yet diverse objects: 5 different boxes, 2 chairs, 2 tables, crate, backpack, trashcan, monitor, keyboard, suitcase, basketball, exercise ball, yoga mat, stool and a toolbox. We include common interactions such as lifting, carrying, sitting, pushing and pulling with hands and feet, as well as free interactions. See our supplementary video for sample sequences. In total, our dataset contains 10.7k frames for training and 4.5k frames for testing respectively.

Human segmentation and SMPL fitting

We segment the human in our images by running DetectronV2 followed by manual correction with on the segmentation masks. These masks are then used to segment multi-view depth maps and lift human point clouds from 2D to 3D. We use FrankMocap to initialize SMPL’s pose from the images and then use instance specific optimization to fit the SMPL model to the segmented human point cloud. For more accurate fitting, we additionally obtain the SMPL shape parameters of each subject from 3D scans using . We report a chamfer error of 1.80cm between the segmented kinect point cloud and our SMPL fits.

Object segmentation and fitting

To obtain object segmentation, we pre-scan objects using a 3D scanner . We then use multi-view object keypoints, marked manually by AMT annotators in images, to optimize the 6D pose of the pre-scanned object mesh to the given frame. We obtain the chamfer error of 2.42cm between the segmented Kinect point cloud and object fit. The segmentation masks are then obtained by projecting fitted object meshes to the images.

Contact annotation

Based on the pseudo-GT SMPL and object fits as described above, we automatically detect contacts if a point on the human surface (registered SMPL) is closer than 2cm to the object surface. For every object point, we store a binary contact label (whether there is a contact or not) and correspondence to the human (contact location on the surface).

See supplementary for more details on data acquisition.

How will this dataset be useful to the community?

We devote significant effort in recording the largest, so far, dataset of natural, full body, day-to-day human interactions with common objects in different natural environments. We propose following challenges with BEHAVE dataset:

Tracking human-object interactions. Track humans and objects using multi-view RGBD data. This can further be extended to track with just multi-view RGB, no-depth, and eventually just a single camera.

Reconstruction from a single image. Joint 3D reconstruction of 3D humans and objects from a single RGB image. Currently, there is no dataset that can be used for benchmarking let alone to learn such a model.

Pose and shape estimation. Benchmarking pose and shape estimation methods in challenging natural environments where the person is heavily occluded by the interacting object.

Apart from these tasks, the research community is free to explore other applications of the BEHAVE dataset.

Method: Tracking human, object and contacts

We present BEHAVE, a method to jointly track humans, objects and their interactions (represented as surface contacts) based on multi-view RGBD input. We formulate our method as an extended per-frame registration problem: we register the human (using SMPL ) and the object (using its pre-scanned object mesh), and predict contacts as correspondences between SMPL and object meshes. See Fig. 3 for the overview of our method. Our formulation must obey three properties, (i) the SMPL model M(⋅)M(\cdot) should fit the human in the multi-view input, (ii) the object mesh Wo\mathbf{W}^{o} should fit the input object and, (iii) SMPL model and object should satisfy contacts. To facilitate joint reasoning of the human, object and contacts directly in 3D, we lift the human Sh\mathcal{S}^{h}, and object So\mathcal{S}^{o} point clouds to 3D using multi-view depth and semantic segmentation. Our joint formulation fits SMPL M(⋅)M(\cdot) and the object Wo\mathbf{W}^{o} to multi-view RGB-D data at each time step, using explicit contacts. This takes the following form

Fitting SMPL to the human point cloud Sh\mathcal{S}^{h} requires, (i) that distance between the SMPL model and the human point cloud should be minimized and (ii) the correct SMPL parts fit the corresponding body parts of the point cloud. The latter is important to avoid degenerate cases such as 180∘180^{\circ} flipped fitting, where the left hand is erroneously matched to the right side of the body or vice-versa . With these considerations, we design our SMPL fitting objective as:

If the correspondences predicted by the network ci\mathbf{c}_{i} deviate from the SMPL surface, these cannot be skinned using the SMPL model as its function is only defined on the body surface. To alleviate this issue, we use the LoopReg formulation that allows us to pose and shape off-the-surface correspondences as well. The final term Ereg=EJ2D+Eθ+EβE^{\text{reg}}=E^{\text{J2D}}+E^{\boldsymbol{\theta}}+E^{\boldsymbol{\beta}}, adds regularisation for SMPL joints, EJ2D=∑k=1K∣πkMJ(θ,β)−J2Dk∣2E^{\text{J2D}}=\sum_{k=1}^{K}|\pi_{k}M^{J}(\boldsymbol{\theta},\boldsymbol{\beta})-\mathbf{J}_{2D}^{k}|_{2}, where πk\pi_{k} is the camera projection matrix of camera kk, MJ(⋅)M^{J}(\cdot) are the 3D body joints and J2Dk\mathbf{J}_{2D}^{k} are the 2D joints detected in the kthk^{th} Kinect image. EθE^{\boldsymbol{\theta}} and EβE^{\boldsymbol{\beta}} are regularisation terms on SMPL pose and shape similar to .

2 Fitting the object mesh to the object point cloud

In order to fit the object mesh, we must ensure that distance from the input object point cloud to the object mesh is minimized. Minimizing this one-sided distance is necessary but not sufficient. Since severe occlusions are common in our interaction setting, large parts of object might be missing from the object point cloud, making fitting difficult. To alleviate this issue we must also ensure that all the vertices of the object mesh are correctly placed w.r.t. the input, even when the point cloud is incomplete. To do so, we take the object mesh vertices vjo∈Wo,j∈{1,…,L}\mathbf{v}^{o}_{j}\in\mathbf{W}^{o},j\in\{1,\dots,L\} and obtain the corresponding point feature Fj\mathbf{F}_{j}, same as Sec. 4.1, where LL is the number of object mesh vertices. We then obtain the unsigned distances to the object and human surfaces using the point feature ujo,ujh=fϕudf(Fj)u_{j}^{o},u_{j}^{h}=f_{\phi}^{\text{udf}}(\mathbf{F}_{j}). Since vjo\mathbf{v}^{o}_{j} is a vertex on the object mesh, its distance to the object surface ujou^{o}_{j} must be zero for a correct fit. This allows us to accurately fit the object vertices to the point data even when corresponding parts are missing from the object point cloud.

where, d(So,Wo)\boldsymbol{d}(\mathcal{S}^{o},\mathbf{W}^{o}) minimizes the point-to-mesh distance between the object point cloud and the object mesh, and the term ∑j=1L∣ujo∣\sum_{j=1}^{L}|u^{o}_{j}| uses implicit unsigned distance prediction to reason about missing object parts.

3 Refining human & object models using contacts

EcontactE^{\text{contact}} allows us to jointly optimise the SMPL model and the object parameters Ro,to\mathbf{R}^{o},\mathbf{t}^{o} to satisfy the contacts predicted by the network.

4 Network training

In this section we elaborate on training our networks.

We use a 3D CNN similar to IF-Net to obtain a voxel aligned multi-scale grid of features F=fϕenc(Sh,So)\mathbf{F}=f_{\phi}^{\text{enc}}(\mathcal{S}^{h},\mathcal{S}^{o}).

Unsigned distance prediction.

To train the network fϕudff_{\phi}^{\text{udf}}, we sample NN query points {p1,…,pN}\{\mathbf{p}_{1},\dots,\mathbf{p}_{N}\} in 3D. For each query point pj\mathbf{p}_{j} we obtain its point feature Fj\mathbf{F}_{j} (Sec. 4.1) and use this to predict the unsigned distance to human and object surface ujo,ujh=fϕudf(Fj)u^{o}_{j},u^{h}_{j}=f_{\phi}^{\text{udf}}(\mathbf{F}_{j}). We jointly train fϕenc,fϕudff_{\phi}^{\text{enc}},f_{\phi}^{\text{udf}} with standard L2 loss. The GT for ujo,ujhu^{o}_{j},u^{h}_{j} is easily available as our dataset contains GT SMPL and object fits allowing us to obtain GT distance of point pj\mathbf{p}_{j} from the SMPL and object mesh.

SMPL correspondence prediction.

To train fϕcorrf_{\phi}^{\text{corr}}, we use the point feature Fj\mathbf{F}_{j} of sampled query point pj\mathbf{p}_{j} to predict its correspondence to the SMPL model cj=fϕcorr(Fj)\mathbf{c}_{j}=f_{\phi}^{\text{corr}}(\mathbf{F}_{j}). We jointly train fϕenc,fϕcorrf_{\phi}^{\text{enc}},f_{\phi}^{\text{corr}} using a standard L2 loss. Since we have the GT SMPL fit in our dataset we simply find the closest SMPL surface point for the query point pj\mathbf{p}_{j} and use this as the GT correspondence.

Object orientation prediction.

Experiments

In this section we compare our approach with existing methods. Our experiments show that we clearly outperform existing baselines. Next, we ablate our design choices and highlight the importance of contact and object orientation prediction in capturing human-object interactions.

We find PHOSA , a method to reconstruct humans and objects from a single image, quite relevant to our work. Although PHOSA uses only a single image whereas we use multi-view images, thus giving our method an advantage, it is still the closest competing method. We run Procrustes alignment on PHOSA results to remove depth ambiguity. It should be noted that PHOSA depends on pre-defined fixed contact regions whereas our approach can freely predict full-body contacts and PHOSA uses hand crafted heuristics to model contacts whereas our approach learns contact modelling from data, making our method more scalable. We compare our method with PHOSA in Fig. 6 and Tab. 2, and clearly outperform it.

2 Why not fit human and object models directly to point clouds?

Since there are no existing methods that can jointly track humans, objects and the contacts from a multi-view input, we create an obvious baseline where we fit the SMPL and object meshes directly to the input point cloud. We show (Tab. 2) that direct fitting easily gets stuck in local minima. This is because the point clouds are very noisy and large parts are missing due to heavy occlusion between the person and the object during interactions. Our network, on the other hand, can implicitly reason about missing parts, thus generating more accurate results.

3 Why can’t existing human registration approaches be extended to our setting?

There are no direct baselines that can jointly track humans, objects, and contacts from multi-view input. There are works that pursue similar ideas of predicting correspondences and fitting SMPL to the human point cloud. In this subsection we explore their suitability in our setting.

IPNet takes as input a human point cloud and predicts an implicit reconstruction of the human and sparse correspondences to the SMPL model, which enables its fitting to the implicit reconstruction. This approach has three major disadvantages. First, querying occupancies for a 1283128^{3} grid to obtain implicit reconstruction is expensive. Second, it predicts occupancies which requires water-tight surfaces. And third, running traditional Marching Cubes makes occupancy prediction non-differentiable w.r.t. SMPL fitting. Our formulation in Eqs. 2 and 3 alleviates these problem as we can fit SMPL by only querying N=30kN=30k points instead of 1283(∼2M)128^{3}(\sim 2M) points. Since we use unsigned distance prediction, our method can work with non-water tight surfaces. We can also fit SMPL directly to unsigned distance predictions, thus removing the requirement for Marching Cubes. We compare our approach with IP-Net (trained on our dataset) in Tab. 2 and show that we obtain better performance than IP-Net at much lower cost (30k30k (ours) vs. ∼2M\sim 2M (IP-Net) query points and no Marching Cubes). This shows that our formulation is superior than IPNet even for human registration. We can additionally handle objects and interactions. Qualitative comparisons are given in the supplementary material.

Comparison with LoopReg [10]

LoopReg fits SMPL to the input point cloud by explicitly predicting correspondences. We find the idea interesting and use their diffused SMPL formulation in our method. LoopReg is, however, not directly applicable in our setting as it assumes a noise free and complete human point cloud. When the point cloud is incomplete due to occlusions, no correspondences are predicted for missing parts. Since LoopReg can only use surface points for fitting, this makes registration inaccurate. BEHAVE handles this case by using distances to the SMPL surface( Eq. 3) predicted for each of the sampled query points to fit the body model, thus allowing the use of non-surface points for fitting. This is important as the Kinect point cloud is noisy. We outperform LoopReg (trained on our dataset) and show (Tab. 2) that our formulation is robust to missing parts and noisy input.

4 Importance of contacts

In this experiment we show that our network predicted contacts are key for physically plausible tracking. Even though quantitative difference is not significant (Tab. 3), it can be seen in Fig. 5 that without contact information, the human and the object do not lock into the correct location. Hence, we notice unnatural results like floating objects. Using our contact prediction alleviates such issues.

We encourage the readers to see our supplementary document for detailed discussion regarding limitations and future work with BEHAVE.

Conclusions

We have presented BEHAVE, the first methodology to jointly track humans, objects and explicit contacts in natural environments. By introducing neural networks to predict correspondences to a 3D human body model along with unsigned distance fields defined over human and object surfaces, we are able to accurately model human-object contacts. We further integrate such neural predictions into a proposed joint registration method resulting in the robust 3D tracking of human-object interactions. Along with our proposed method we also provide BEHAVE, the largest dataset of RGBD sequences and annotated humans, objects, and contacts to date. BEHAVE dataset is the first benchmark for the part of the research community interested in modelling human-object interactions. We propose real-world challenges like reconstructing humans and object from a single RGB image, tracking human-object interactions from multiple and single-view RGB(D) input, pose estimation etc. Our dataset together with our code is released in order to stimulate future research in this important emerging domain.

Special thanks to RVH team members , and reviewers, their feedback helped improve the manuscript. This work is funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) - 409792180 (Emmy Noether Programme, project: Real Virtual Humans), German Federal Ministry of Education and Research (BMBF): Tübingen AI Center, FKZ: 01IS18039A and ERC Consolidator Grant 4DRepLy (770784). Gerard Pons-Moll is a member of the Machine Learning Cluster of Excellence, EXC number 2064/1 – Project number 390727645.

References