BEHAVE: Dataset and Method for Tracking Human Object Interactions
Bharat Lal Bhatnagar, Xianghui Xie, Ilya A. Petrov, Cristian Sminchisescu, Christian Theobalt, Gerard Pons-Moll
Introduction
The last decade has seen rapid progress in modelling the appearance of humans ranging from body pose, shape , faces and even detailed clothing . With various practical use cases like virtual try-on, personalised avatar creation, and several applications in augmented and mixed reality, or human-robot collaboration, the focus on humans is justified. Beyond modelling appearance, few methods have focused on capturing and synthesizing human interactions (human-object/scene interaction). There exists work to capture humans in a static 3D scene , even without using external cameras , and work to synthesize static poses , or full body movement in a 3D scene.
These methods show growing interest in modelling human behavior, highlighting a need to capture real human interactions. Existing methods however are learned from high quality curated data captured using optical marker based motion capture systems or wearable sensors. Unfortunately, such commercial systems are expensive, drastically limit the interactions that can be captured, and often fail when tracking humans and objects under occlusion. In addition, the recording volume is spatially confined and difficult to re-locate, thus limiting the activities, scenes, and objects that can be captured. Wearable sensors are not restricted in volume, but close range interaction can not be accurately captured. Altogether, the lack of diverse 3D interaction data, and the lack of accurate and flexible capture methods both constitute barriers in modelling human behavior.
With the goal of simplifying the data capture process and hence allowing faster progress in the field, we propose BEHAVE, a method to capture diverse 3D human interactions in natural environments, using a setup comprising of portable, cheap, and easy to use RGBD cameras. Tracking human interactions from sparse consumer grade cameras is however extremely challenging. Depth data is inherently noisy and incomplete. Moreover, the person and object occlude each other frequently during interactions. Furthermore, capturing interactions requires estimating human-object contacts accurately, which is difficult because contacts represent small regions in the image, close to the observable (resolution) limit. This requires innovation that goes significantly beyond the current state of the art trackers. We propose to track the human using a parametric human model (such as SMPL ) and track objects using template meshes. Naively fitting the human model and an object 3D template to the point-cloud completely fails due to the aforementioned challenges. Our key idea is to train a neural model which jointly completes the human and object shape, represented with implicit surfaces, while predicting a correspondence field to the human, as well as an object orientation field. These rich outputs allow us to formulate a powerful human-object fitting objective which is robust to missing data, noise and occlusion.
To train and evaluate BEHAVE, we capture the largest dataset of human-object interactions in natural environments. The BEHAVE dataset contains 20 3D objects, 8 subjects (5 male, 3 female), 5 different locations and totals around 15.2k frames of recording. We provide ground truth SMPL and 3D object meshes as well as contacts. Our contributions can be summarized as follows:
We propose the first approach that can accurately 3D track humans, objects and contacts in natural environments using multi-view RGBD images.
We collect the largest dataset of multi-view RGBD sequences and corresponding human models, object and contact annotations. See Sec. 3 for details regarding its usefulness to the community.
Since there exists no publicly available code and datasets to accurately track human-object interactions in natural environments, we will release our code and data for further research in this direction.
Related Work
In this section, we first briefly review work focused on object and human reconstruction, in isolation from their environmental context. Such methods focus on modelling appearance and do not consider interactions. Next, we cover methods focused on humans in static scenes and finally discuss closer-related work to ours, for modelling dynamic human-object interactions.
Perceiving humans from monocular RGB data and under multiple views settings has been widely explored. Recent work tends to focus on reconstructing fine details like hand gestures and facial expressions , self-contacts , interactions between humans, and even clothing . These methods benefit from representing human with parametric body models , thus motivating our use of recent implicit diffused representations as backbone for our tracker. Following the success of pixel-aligned implicit function learning , recent methods can capture human performance from sparse or even a single RGB camera . However, capturing 3D humans from RGB data involves a fundamental ambiguity between depth and scale. Therefore, recent methods use RGBD or volumetric data for reliable human capture. These insights motivate us to build novel trackers based on multi-view RGBD data.
Object reconstruction
Most existing work on reconstructing 3D objects from RGB and RGBD data does so in isolation, without the human involvement or the interaction. While challenging, it is arguably more interesting to reconstruct objects in a dynamic setting under severe occlusions from the human.
2 Interaction modelling: Humans and objects with scene context
Modelling how humans act in a scene is both important and challenging. Tasks like placement of humans into static scenes , motion prediction or human pose reconstruction under scene constrains, or learning priors for human-object interactions , have been investigated extensively in recent years. These methods are relevant but restricted to modelling humans interacting with static objects. We address a more challenging problem of jointly tracking human-object interactions in dynamic environments where objects are manipulated.
Dynamic human object interactions
Recently, there has been a strong push on modeling hand-object interactions based on 3D , 2.5D and 2D data. Although powerful, these methods are currently restricted to modelling only hand-object interactions. In contrast, we are interested in full body capture. Methods for dynamic full body human object interaction approach the problem via 2D action recognition or reconstruct 3D object trajectories during interactions . Despite being impressive, such methods either lack full 3D reasoning or are limited to specific objects . More recent work reconstructs and tracks human-object interactions from RGB or RGBD streams , but does not consider contact prediction, thus missing a component necessary for accurate interaction estimates. Very relevant to our work, PHOSA reconstructs humans and objects from a single image. PHOSA uses hand crafted heuristics, instance specific optimization for fitting, and pre-defined contact regions, which limits generalization to diverse human-object interactions. Our method on the other hand learns to predict the necessary information from data, making our models more scale-able. As shown in the experiments, the accuracy of our method is significantly higher to PHOSA.
BEHAVE Dataset
We present BEHAVE dataset, the largest dataset of human-object interactions in natural environments, with 3D human, object and contact annotation, to date. See Tab. 1 for comparison with other datasets. Our dataset contains multi-view RGBD frames, with accurate pseudo-ground truth SMPL , object fits, human and object segmentation masks, and contact annotations.
We setup and calibrate 4 Kinects at 4 corners of our square recording volume where all interactions are performed by 8 subjects (5 male, 3 female). Interactions are captured at 5 disparate indoor locations with 20 commonly used, yet diverse objects: 5 different boxes, 2 chairs, 2 tables, crate, backpack, trashcan, monitor, keyboard, suitcase, basketball, exercise ball, yoga mat, stool and a toolbox. We include common interactions such as lifting, carrying, sitting, pushing and pulling with hands and feet, as well as free interactions. See our supplementary video for sample sequences. In total, our dataset contains 10.7k frames for training and 4.5k frames for testing respectively.
Human segmentation and SMPL fitting
We segment the human in our images by running DetectronV2 followed by manual correction with on the segmentation masks. These masks are then used to segment multi-view depth maps and lift human point clouds from 2D to 3D. We use FrankMocap to initialize SMPL’s pose from the images and then use instance specific optimization to fit the SMPL model to the segmented human point cloud. For more accurate fitting, we additionally obtain the SMPL shape parameters of each subject from 3D scans using . We report a chamfer error of 1.80cm between the segmented kinect point cloud and our SMPL fits.
Object segmentation and fitting
To obtain object segmentation, we pre-scan objects using a 3D scanner . We then use multi-view object keypoints, marked manually by AMT annotators in images, to optimize the 6D pose of the pre-scanned object mesh to the given frame. We obtain the chamfer error of 2.42cm between the segmented Kinect point cloud and object fit. The segmentation masks are then obtained by projecting fitted object meshes to the images.
Contact annotation
Based on the pseudo-GT SMPL and object fits as described above, we automatically detect contacts if a point on the human surface (registered SMPL) is closer than 2cm to the object surface. For every object point, we store a binary contact label (whether there is a contact or not) and correspondence to the human (contact location on the surface).
See supplementary for more details on data acquisition.
How will this dataset be useful to the community?
We devote significant effort in recording the largest, so far, dataset of natural, full body, day-to-day human interactions with common objects in different natural environments. We propose following challenges with BEHAVE dataset:
Tracking human-object interactions. Track humans and objects using multi-view RGBD data. This can further be extended to track with just multi-view RGB, no-depth, and eventually just a single camera.
Reconstruction from a single image. Joint 3D reconstruction of 3D humans and objects from a single RGB image. Currently, there is no dataset that can be used for benchmarking let alone to learn such a model.
Pose and shape estimation. Benchmarking pose and shape estimation methods in challenging natural environments where the person is heavily occluded by the interacting object.
Apart from these tasks, the research community is free to explore other applications of the BEHAVE dataset.
Method: Tracking human, object and contacts
We present BEHAVE, a method to jointly track humans, objects and their interactions (represented as surface contacts) based on multi-view RGBD input. We formulate our method as an extended per-frame registration problem: we register the human (using SMPL ) and the object (using its pre-scanned object mesh), and predict contacts as correspondences between SMPL and object meshes. See Fig. 3 for the overview of our method. Our formulation must obey three properties, (i) the SMPL model should fit the human in the multi-view input, (ii) the object mesh should fit the input object and, (iii) SMPL model and object should satisfy contacts. To facilitate joint reasoning of the human, object and contacts directly in 3D, we lift the human , and object point clouds to 3D using multi-view depth and semantic segmentation. Our joint formulation fits SMPL and the object to multi-view RGB-D data at each time step, using explicit contacts. This takes the following form
Fitting SMPL to the human point cloud requires, (i) that distance between the SMPL model and the human point cloud should be minimized and (ii) the correct SMPL parts fit the corresponding body parts of the point cloud. The latter is important to avoid degenerate cases such as flipped fitting, where the left hand is erroneously matched to the right side of the body or vice-versa . With these considerations, we design our SMPL fitting objective as:
If the correspondences predicted by the network deviate from the SMPL surface, these cannot be skinned using the SMPL model as its function is only defined on the body surface. To alleviate this issue, we use the LoopReg formulation that allows us to pose and shape off-the-surface correspondences as well. The final term , adds regularisation for SMPL joints, , where is the camera projection matrix of camera , are the 3D body joints and are the 2D joints detected in the Kinect image. and are regularisation terms on SMPL pose and shape similar to .
2 Fitting the object mesh to the object point cloud
In order to fit the object mesh, we must ensure that distance from the input object point cloud to the object mesh is minimized. Minimizing this one-sided distance is necessary but not sufficient. Since severe occlusions are common in our interaction setting, large parts of object might be missing from the object point cloud, making fitting difficult. To alleviate this issue we must also ensure that all the vertices of the object mesh are correctly placed w.r.t. the input, even when the point cloud is incomplete. To do so, we take the object mesh vertices and obtain the corresponding point feature , same as Sec. 4.1, where is the number of object mesh vertices. We then obtain the unsigned distances to the object and human surfaces using the point feature . Since is a vertex on the object mesh, its distance to the object surface must be zero for a correct fit. This allows us to accurately fit the object vertices to the point data even when corresponding parts are missing from the object point cloud.
where, minimizes the point-to-mesh distance between the object point cloud and the object mesh, and the term uses implicit unsigned distance prediction to reason about missing object parts.
3 Refining human & object models using contacts
allows us to jointly optimise the SMPL model and the object parameters to satisfy the contacts predicted by the network.
4 Network training
In this section we elaborate on training our networks.
We use a 3D CNN similar to IF-Net to obtain a voxel aligned multi-scale grid of features .
Unsigned distance prediction.
To train the network , we sample query points in 3D. For each query point we obtain its point feature (Sec. 4.1) and use this to predict the unsigned distance to human and object surface . We jointly train with standard L2 loss. The GT for is easily available as our dataset contains GT SMPL and object fits allowing us to obtain GT distance of point from the SMPL and object mesh.
SMPL correspondence prediction.
To train , we use the point feature of sampled query point to predict its correspondence to the SMPL model . We jointly train using a standard L2 loss. Since we have the GT SMPL fit in our dataset we simply find the closest SMPL surface point for the query point and use this as the GT correspondence.
Object orientation prediction.
Experiments
In this section we compare our approach with existing methods. Our experiments show that we clearly outperform existing baselines. Next, we ablate our design choices and highlight the importance of contact and object orientation prediction in capturing human-object interactions.
We find PHOSA , a method to reconstruct humans and objects from a single image, quite relevant to our work. Although PHOSA uses only a single image whereas we use multi-view images, thus giving our method an advantage, it is still the closest competing method. We run Procrustes alignment on PHOSA results to remove depth ambiguity. It should be noted that PHOSA depends on pre-defined fixed contact regions whereas our approach can freely predict full-body contacts and PHOSA uses hand crafted heuristics to model contacts whereas our approach learns contact modelling from data, making our method more scalable. We compare our method with PHOSA in Fig. 6 and Tab. 2, and clearly outperform it.
2 Why not fit human and object models directly to point clouds?
Since there are no existing methods that can jointly track humans, objects and the contacts from a multi-view input, we create an obvious baseline where we fit the SMPL and object meshes directly to the input point cloud. We show (Tab. 2) that direct fitting easily gets stuck in local minima. This is because the point clouds are very noisy and large parts are missing due to heavy occlusion between the person and the object during interactions. Our network, on the other hand, can implicitly reason about missing parts, thus generating more accurate results.
3 Why can’t existing human registration approaches be extended to our setting?
There are no direct baselines that can jointly track humans, objects, and contacts from multi-view input. There are works that pursue similar ideas of predicting correspondences and fitting SMPL to the human point cloud. In this subsection we explore their suitability in our setting.
IPNet takes as input a human point cloud and predicts an implicit reconstruction of the human and sparse correspondences to the SMPL model, which enables its fitting to the implicit reconstruction. This approach has three major disadvantages. First, querying occupancies for a grid to obtain implicit reconstruction is expensive. Second, it predicts occupancies which requires water-tight surfaces. And third, running traditional Marching Cubes makes occupancy prediction non-differentiable w.r.t. SMPL fitting. Our formulation in Eqs. 2 and 3 alleviates these problem as we can fit SMPL by only querying points instead of points. Since we use unsigned distance prediction, our method can work with non-water tight surfaces. We can also fit SMPL directly to unsigned distance predictions, thus removing the requirement for Marching Cubes. We compare our approach with IP-Net (trained on our dataset) in Tab. 2 and show that we obtain better performance than IP-Net at much lower cost ( (ours) vs. (IP-Net) query points and no Marching Cubes). This shows that our formulation is superior than IPNet even for human registration. We can additionally handle objects and interactions. Qualitative comparisons are given in the supplementary material.
Comparison with LoopReg [10]
LoopReg fits SMPL to the input point cloud by explicitly predicting correspondences. We find the idea interesting and use their diffused SMPL formulation in our method. LoopReg is, however, not directly applicable in our setting as it assumes a noise free and complete human point cloud. When the point cloud is incomplete due to occlusions, no correspondences are predicted for missing parts. Since LoopReg can only use surface points for fitting, this makes registration inaccurate. BEHAVE handles this case by using distances to the SMPL surface( Eq. 3) predicted for each of the sampled query points to fit the body model, thus allowing the use of non-surface points for fitting. This is important as the Kinect point cloud is noisy. We outperform LoopReg (trained on our dataset) and show (Tab. 2) that our formulation is robust to missing parts and noisy input.
4 Importance of contacts
In this experiment we show that our network predicted contacts are key for physically plausible tracking. Even though quantitative difference is not significant (Tab. 3), it can be seen in Fig. 5 that without contact information, the human and the object do not lock into the correct location. Hence, we notice unnatural results like floating objects. Using our contact prediction alleviates such issues.
We encourage the readers to see our supplementary document for detailed discussion regarding limitations and future work with BEHAVE.
Conclusions
We have presented BEHAVE, the first methodology to jointly track humans, objects and explicit contacts in natural environments. By introducing neural networks to predict correspondences to a 3D human body model along with unsigned distance fields defined over human and object surfaces, we are able to accurately model human-object contacts. We further integrate such neural predictions into a proposed joint registration method resulting in the robust 3D tracking of human-object interactions. Along with our proposed method we also provide BEHAVE, the largest dataset of RGBD sequences and annotated humans, objects, and contacts to date. BEHAVE dataset is the first benchmark for the part of the research community interested in modelling human-object interactions. We propose real-world challenges like reconstructing humans and object from a single RGB image, tracking human-object interactions from multiple and single-view RGB(D) input, pose estimation etc. Our dataset together with our code is released in order to stimulate future research in this important emerging domain.
Special thanks to RVH team members , and reviewers, their feedback helped improve the manuscript. This work is funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) - 409792180 (Emmy Noether Programme, project: Real Virtual Humans), German Federal Ministry of Education and Research (BMBF): Tübingen AI Center, FKZ: 01IS18039A and ERC Consolidator Grant 4DRepLy (770784). Gerard Pons-Moll is a member of the Machine Learning Cluster of Excellence, EXC number 2064/1 – Project number 390727645.