Object pop-up: Can we infer 3D objects and their poses from human interactions alone?
Ilya A. Petrov, Riccardo Marin, Julian Chibane, Gerard Pons-Moll
Introduction
Complex interactions with the world are among the unique skills distinguishing humans from other living beings. Even though our perception might be imperfect (we cannot hear ultrasonic sounds or see ultraviolet light ), our cognitive representation is enriched with a functional perspective, i.e., potential ways of interacting with objects or, as introduced by Gibson and colleagues, the affordance of the objects . Several behavioural studies confirmed the centrality of this concept , which plays a fundamental role also for kids’ development . Computer Vision is well-aware that the function of an object complements its appearance , and exploited this in tasks like human and object reconstruction . Previous literature approaches the interaction analysis from an object perspective (i.e., given an object, analyze the human interaction), building object-centric priors , generating realistic grasps given the object , reconstructing hand-object interactions . Namely, objects induce functionality, so a human interaction (e.g., a mug suggests a drinking action; an handle a grasping one). For the first time, our work reverts the perspective, suggesting that analyzing human motion and behaviour is naturally a human-centric problem (i.e., given a human interaction, what kind of functionality is it suggesting, Fig. 2). Moving the first step in this new research direction, we pose a fundamental question: Can we infer 3D objects and their poses from human interactions alone?
At first sight, the problem seems particularly hard and significantly under-constrained since several geometries might fit the same action. However, the human body complements this information in several ways: physical relations, characteristic poses, or body dynamics serve as valuable proxies for the involved functionality, as suggested by Fig. 1. Such hints are so powerful that we can easily imagine the kind of object and its location, even if such an object does not exist. Furthermore, even if the given pose might fit several possible solutions, our mind naturally comes to the most natural one indicated by the observed behaviour. It also suggests that solely focusing on the contact region (an approach often preferred by previous works) is insufficient in this new viewpoint. The reification principle of Gestalt psychology highlights that "the whole" arises from "the parts" and their relation. Similarly, in Fig. 3, the hand grasps in B) pop up a binocular in our mind because we naturally consider the relationship with the other body parts.
Finally, moving to a human-centric perspective in human-object interaction is critical for human studies and daily-life applications. Modern systems for AR/VR , and digital interaction are both centered on humans, often manipulating objects that do not have a real-world counterpart. Learning to decode an object from human behaviour enables unprecedented applications.
To answer our question, we deploy a first straightforward and effective pipeline to "pop up" a rigid object from a 3D human point cloud. Starting from the input human point cloud and a class, we train an end-to-end pipeline to infer object location. In case a temporal sequence of point clouds is available, we suggest post-processing to avoid jittering and inconsistencies in the predictions, showing the relevance of this information to handle ambiguous poses. We show promising results on previously unaddressed tasks in digital and real-world scenes. Finally, our method allows us to analyze different features of human behaviour, highlight their contribution to object retrieval, and point to exciting directions for future works.
We formulate a novel problem, changing the perspective taken by previous works in the field, and open to a yet unexplored research direction;
We introduce a method capable of predicting the object starting from an input human point cloud;
We analyze different components of the human-object relationship: the contribution of different pieces of interactions (hands, body, time sequence), the point-wise saliency of input points, and the confusion produced by objects with similar functions.
Related Work
In the core of human perception of objects, functionality complements physical appearance, enhancing our perception. Gibson introduced the idea that humans use affordances of objects for perception. Affordance can be defined as: "an intrinsic property of an object, allowing an action to be performed with the object" .
From a Computer Vision perspective, object functionality supports several tasks such as scene analysis , object classification , object properties inferring , and it is also possible to learn object-specific human interaction models from 2D images. These works suggest an intimate entanglement between the action and the object itself.
2 Human-Object interaction
Modelling environment-aware humans and their interactions in 3D is one of the most recent challenges to creating virtual humans. We see two main lines of work there: hand-focused and full-body ones.
Several works tackle the problem of human interaction, focusing solely on the hands . This has been done starting from 2D , 2.5D , and 3D data . Particularly promising seems the application of well-designed priors for the motion . The class of objects involved in these works are mainly limited to graspable ones. We argue that interactions involving different body parts are common in everyday life, more attractive from an applicative perspective, and more challenging. Moreover, full-body context is crucial in reconstructing even grasp interactions since body pose contains information on an object’s properties, e.g. pounding with a hammer affects the whole posture to support the action.
Fully-Body Interaction.
On this line, several works focus on the interaction between a human and a scene . Also, in this case, priors can be used to regularize the motion . Several datasets are also available to study the interaction between a human and a single object. For example, recent BEHAVE , GRAB , and InterCap capture full-body interactions with diverse objects. Works address the task of humans interactions reconstruction from different kinds of data sources like single image , video , and multi-view capturing , and synthetization of them as well . However, in all these works, we observe a general object-centric perspective: given a scene or an object, they aim to recreate the humans interacting with them. We argue that the significantly less explored complementary one has more concrete applications in daily life, especially in VR/XR contexts where the human is central to the system.
3 Human-centric perspective
While the general trend mainly focuses on the surrounding environment and objects, there is a growing interest and availability of egocentric tools for humans also interacting with objects . They provide a subjective view and a valuable paradigm for several applications, like letting a user interact with objects in the digital world. Recent works also involve more sophisticated devices , while they are still far from applicability. In a similar direction to our work, others propose to recover objects arrangements in a room starting from the human motion , to hallucinate a coherent 2D image from a human pose , or predicting physical properties of the objects (e.g., the weight of a box) from human joints . While the principle inspires us, our study significantly differs: we focus on object pose and its spatial relation with humans, starting solely from unordered point clouds.
Method
This section describes our setting and the main components of our methodology, both at inference and training time. An overview of our pipeline can be found in Fig. 4.
Object Center.
Training a model to predict an object pose from a human point cloud poses several challenges. Such a task requires the network to understand the location of different body parts and their subtle relations while jointly developing a sense of its spatial relationship with the human. Empirically, we observed that this is only feasible by carefully deconstructing the problem and designing different features to ease the learning process. As the first step to decompose this problem, we train a PointNet++ architecture to predict the object center starting from . At training time, this is supervised with an L2 loss against the ground truth center :
Local Neighbourhood.
Object displacement.
To predict the object’s final position, we empirically observed that directly predicting a rotation and a translation is not a good solution. Inspired by recent works that suggest a point-wise offset prediction to recover 3D human shapes , we apply a similar approach to our task. Our goal is to predict a point-wise shift for the vertices to align them to the target pose. We append the features , , the one-hot encoding of the object class , and a positional encoding to the centered key points, and we pass them to a decoder. At training time, we consider the following loss:
The network is then trained end-to-end using:
The weighting coefficient is .
2 Template fitting
The point-wise offset produced by the network potentially distorts the key points structure in a non-rigid way. To recover the desired global rigid transformation, we rely on a Procrustes alignment . This procedure takes as input two point clouds and returns the rotation and the translation to minimize the L2 distances of the points:
We apply this to the template key points and their configuration obtained with our network:
Finally, we recover the desired object pose as:
Time Smoothing.
While our pipeline is designed to work with a single point cloud as input, considering the temporal evolution of interaction is often crucial, shaping the context of the individual poses. If a temporal sequence of point clouds is available, we provide a post-processing smoothing technique to take advantage of this further information. After running our method for each frame, we smooth the centering prediction across the sequence using a Gaussian kernel. Later, we will discuss a variation of our approach that also predicts the object class. In that case, we consider the most frequent class prediction over the whole set of frames to fix a class for the sequence.
Experiments
In this Section, we will describe the datasets used for training and testing our method. Then, we will present the baseline and the evaluation metrics. Finally, we will provide validation of our method as well as an analysis of its extended version that allows object class prediction.
We jointly train on the union of BEHAVE and GRAB , obtaining a set of:
subjects, first subjects from the GRAB dataset and subjects from the official training part of BEHAVE;
different classes of objects, including all objects from the BEHAVE dataset and selected objects from the GRAB dataset;
We downsample training sequences of GRAB and BEHAVE to . To evaluate our method, we select subjects and from the GRAB dataset and downsample the sequences to . For the BEHAVE dataset, we use the official test part, which includes all sequences at with subject and part of the sequences with subjects . As an input, we use point clouds with points sampled uniformly over the SMPL-H meshes. We refer to raw point clouds from the BEHAVE dataset used in our experiments as BEHAVE-Raw. We use point clouds that are fused from 4 Kinect sensors and subsample points from them.
Data augmentation.
During training, to simulate errors in the center prediction, we randomly translate and rotate the object around the ground-truth center .
Implementation details.
We implement our method using PyTorch framework and use Nvidia RTX3090 GPU for training and evaluation. The model is trained using Adam optimizer for epochs with a learning rate of , which decays times after -th and -th epochs. For the first epochs of training, we use ground-truth object center instead of predicted to select local neighborhood , to warm up the local PointNet++ encoder.
Nearest-Neighbor Baseline.
Since we are the first to tackle this task and no competitors are available, we propose a simple while informative baseline. Given the input point cloud, we recover the most similar in the training dataset in an L2 sense. Then, we recover the object handled by that subject and pose it in space in the same way. This baseline demonstrates that the task is non-trivial and the generalization to unseen poses and subjects of our method. Also, this baseline requires that the target point cloud and the ones in the training set share the same number of points. Hence, if the input point cloud is a raw scan, this baseline is not applicable. Our method, instead, does not rely on this assumption and is more general.
Object classification.
In our research, we also investigate the possibility of incorporating class prediction inside the network training. This task is significantly difficult at a single-frame level since an isolated pose often does not suggest a clear functionality. However, including this step is interesting to analyze the interaction and the nature of the network confusion. Hence, we modify our method by adding a decoder module that takes the global features and the local ones as input to predict the object class. Then, we add to the training a simple cross-entropy loss between the predicted class and the ground truth one .
1 Metrics
In qualitative experiments, we use three main metrics to evaluate our results. In the tables, we report the average error across the considered test samples.
In most cases, our resulting object and the target one share the same number of vertices. Hence, we can compute the error between our prediction and the ground truth as a point-to-point error:
When such error is computed only between the object centers, we will refer to it as .
Chamfer distance.
When we evaluate the network that also predicts the class, target objects and selected templates might not share the same number of vertices. In that case, as a metric we use bi-directional Chamfer distance:
Classification Accuracy.
In case we use our network to predict the object class, we measure our misclassification error in terms of accuracy.
2 Object pose Evaluation
In Tab. 1, we report the quantitative evaluation on the test set of the datasets, comparing our method to the baseline. Our approach significantly outperforms the baseline, even if this latter exploits the points order information. We obtain the most significant margin on the GRAB dataset, where objects are small and mainly involve hands, showing the precision of our method. The baseline cannot be applied on BEHAVE-Raw since it does not share the same number of vertices as the training set, while our method shows only a limited performance decrease, pointing to generalization also to point clouds coming from different sources. We report qualitative results of our method on GRAB (first two rows) and BEHAVE (last row) in Fig. 5, and on BEHAVE-Raw in Fig. 6. Finally, to further evaluate the generalization of our method to unseen poses and subjects, we also considered point clouds obtained by an egocentric pipeline. In this scenario, we record a user motion with an XSens system , retarget the pose to an SMPL+H model, and obtain the point cloud by sampling the resulting mesh. Outputs of our method can be observed in Fig. 1 and Fig. 7. Generalizing to unseen poses and subjects acquired with wearable systems prone to measurement errors opens several exciting applications for VR/XR contexts.
3 Human Affordance analysis
We use our method to analyze human-object interaction considering three key factors: changes in the input information, the points saliency, and confusion in classification.
Since we are curious to analyze how information is encoded in the inputs, we consider three more scenarios:
Hands: We generate the input data using the MANO hand model annotations included in GRAB dataset. Hence, the model should infer the pose of the object without relying on other body parts.
SMPL: We consider all the points of the subject, but the ones from the hands are in the rest pose provided by SMPL . In this way, we analyze how much the network captures the information of the finger pose.
SMPLH+T: We consider all the points of the subject with the finger correctly posed (as in our method), and also we apply the temporal smoothing outlined in Sec. 3. Temporal information contextualizes the interactions and regularizes the predictions throughout the sequence.
For the first two cases, we train ad-hoc networks. Results of this analysis are reported in Tab. 2. We notice that hands provide crucial information for the GRAB dataset. The result is expected since, as the name suggested, all the interactions in the dataset are focused on hands. However, full-body context is crucial for reconstructing interactions from the BEHAVE dataset, as they often involve multiple body parts.
Points saliency.
As further evidence that object interaction involves different body parts, we conducted a study to discover what input points are crucial for the network. We follow a recent protocol to find 3D point cloud saliency : we cast the input through the network, we compute a loss on the output (in our case, the one in Eq. 2), and we modify the input using the backpropagated gradient. This procedure is performed iteratively, and we refer to for the details. We report the results of this study in Fig. 8, where points modified by the procedure are highlighted in red. We find the results of this analysis fascinating. As expected, the contact region is always essential to infer the correct object location. However, feet play a crucial role in all the reported cases, since they provide information about the human position and the consequent pose of the body. Another highlighted region is the head: different orientations give clues about object location and body posture.
Confusion in classification.
As a final analysis, we explore the learning object classification during the training. We jointly train a further MLP module that takes as an input and predicts the object class, using a cross-entropy loss. When a time sequence of point clouds is available, we exploit it by selecting the class with the highest score across the frames and applying it to the whole sequence. We report results in Tab. 3. Our experiment suggests that the task is challenging while, given the number of classes (40), we still consider our results a promising first step. Also, we notice that temporal smoothing significantly helps classification accuracy. Temporal context disambiguate poses without a clear functionality. In Fig. 9, we report the confusion matrix for a subset of the classes for our method. The misclassification mainly arises from interactions of objects with similar functionality.
Conclusions
In this work, we have addressed a novel and inspiring problem that changes the perspective on object-human interaction. Our proposed model is simple, carefully designed, and inspired by behavioural studies. We collected evidence of the method’s effectiveness on a large set of object classes and empirically proved its generalization on noisy and different inputs. Finally, our analysis of human affordance is unprecedented, showing that human-object interaction can also involve body parts distant from the object and pointing to interesting relations useful for applications and subsequent works.
As the first exploration in this direction, our study enables several future possibilities. In this work, the temporal information is only used after the training procedure. Incorporating this information can create other patterns, further improving the results’ quality. Our empirical evidence suggests that class prediction requires further investigation and more sophisticated techniques, like specialized attention mechanisms. Finally, we do not consider sequences that involve long (e.g., hours-long) and complex (e.g., multi-objects) interactions, which are difficult to capture. We hope our work can foster the community to collect such datasets.
Acknowledgements
Special thanks to the RVH team and reviewers, their feedback helped improve the manuscript. We also thank Omid Taheri for the help with GRAB objects. This work is funded by the Deutsche Forschungsgemeinschaft - 409792180 (EmmyNoether Programme, project: Real Virtual Humans) and the German Federal Ministry of Education and Research (BMBF): Tübingen AI Center, FKZ: 01IS18039A. G. Pons-Moll is a member of the Machine Learning Cluster of Excellence, EXC number 2064/1 – Project number 390727645. The authors thank the International Max Planck Research School for Intelligent Systems (IMPRS-IS) for supporting I. Petrov. R. Marin is supported by an Alexander von Humboldt Foundation Research Fellowship. The project was made possible by funding from the Carl Zeiss Foundation.
References
Appendix A Ablation Study
We perform an ablation study to validate our design choices and analyze their impact on the system. We report the results in Tab. A.1.
R,t. Our method outputs a vertex-wise offset. Given that our goal is recovering a global pose of a rigid object, a more natural option is to predict a global rotation and translation directly. However, training a model this way significantly decreases performance, suggesting that our richer output representation provides the network more flexibility.
No . One of our hypotheses is that the correct object location arises from the union of human parts. In our design, we enrich the features from the whole body with others that are focused around the predicted center location of the object. By removing the latter, we observe up to impact on the performance. We discover this information is particularly relevant for small objects (i.e., GRAB), where recovering the pose requires a finer understanding. The performance for the larger objects (i.e., BEHAVE) is on par with the main model, as in this case, the whole body becomes predominant in the prediction.
Appendix B Architecture and implementation details
In this section we provide more details about the proposed architecture. We gather the used notations in Tab. A.2.
Training the model for epochs on Nvidia RTX3090 GPU takes approximately hours. Sequences from the GRAB dataset are downsampled from 120fps to 10fps for training and 30fps for evaluation. The BEHAVE dataset provides extended annotations at 30fps for some sequences, that are downsampled to 10fps for training. For evaluation on BEHAVE, the original 1fps annotations are used. We align the meshes to share the same ground plane, keeping the original global rotation.
B.2 Object pop-up with class prediction
The overview of the model with class prediction is presented in Fig. A.1. Object class is predicted from global using MLP, apart from that module other modules are similar to the main model.
This model is trained in exactly the same setting as the model without class prediction. We use cross entropy loss , to supervise class prediction.
The network is then trained using the following loss:
The weighting coefficient is .
B.3 Nearest neighbour baseline
Here we provide a detailed description of the Nearest Neighbour baseline. First of all, we consider the set of training point clouds
where each is equipped with a correctly posed object template and the relative class . Then, given a new input point cloud , we look for the closest point cloud in the training set:
When the task is to recover the object with the class given as input, we consider only the training samples of the given class, and we retrieve the object . When the method also has to predict the class, we consider the entire training dataset, and we also output the associated class .
Remark. This baseline can be applied only if the input point cloud shares the same number and order of vertices as the ones coming from the training set. While this gives an advantage to the baseline, we decided to proceed this way for a computational reason, given the large training dataset size.
Appendix C Results
This section presents more qualitative results of the proposed Object pop-up method and discusses the failure cases.
Qualitative results for Object pop-up trained on data generated from MANO hand meshes from the GRAB dataset are presented in Fig. A.2. The vast majority of actions in GRAB are done with hands, so that the model can predict the object’s position quite well. However, for other types of interactions (i.e. involving other body parts), having full-body context is crucial. In Fig. A.3 we show qualitative results for Object pop-up model trained on data generated from SMPL meshes from GRAB and BEHAVE. This data differs from the primary training data generated from SMPL-H meshes by the absence of articulated hand pose. Some easier interactions with objects through hands, like lifting a camera on the left side of Fig. A.3, are handled by the model well. However, more complex cases, like cutting with a knife in the middle of Fig. A.3, lead to erroneous prediction because the model lacks local features essential for the interaction. At the same time, interactions involving full-body, like sitting in the right of Fig. A.3, are perceived well by the model because hands are not contributing much to the local interaction context there.
GRAB and BEHAVE.
We present more qualitative results on GRAB in Fig. A.6 and BEVAHE in Fig. A.5. Our method can predict the object’s realistic location for many classes.
Generalization.
We present additional results on BEHAVE-Raw in Fig. A.4, and in Fig. A.7 we report more qualitative examples on the data recorded with IMU sensors. Our method shows generalization to both noisy point clouds of BEHAVE-Raw and unseen subjects, recorded with a wearable IMU setup.
Failure cases.
Our method sometimes predicts objects interpenetrating the human body (e.g. yoga ball on the left of Fig. A.8). The absence of an explicit surface in the input data requires the network to recover such a complex structure and the non-interpenetrating relationship. In some cases, the method struggles to predict correct object placement for objects with small handles (e.g. mug in the middle of Fig. A.8, as grabbings by different parts (e.g., by the handle, or by the central body of the object) are all plausible solutions. These two cases differ only in a slight variation of hand pose, which can be hard to grasp even with the local focus of the network. Another failure case is incorrect object placement for interactions that do not involve objects’ functionality (i.e. lifting, passing, inspecting). An example of such a case is a teapot interaction on the right side of Fig. A.8, where it is just shifted in the space. For this case, the method still predicts an object pose which is more common for the object’s functionality (Fig. A.6 presents such distinctive teapot grasp).
Appendix D Human affordance details
Here we report the details about Points saliency estimation, following the procedure from the Algorithm of .
Given an input point cloud and the associated class , first of all we compute the center of as the median of the three individual coordinates of the points:
For each point, we compute the vector that connect the center to it:
We cast and through the network, obtaining the output offsets . We use them to compute as reported in the main manuscript.
For each input point of the point cloud, we recover the gradient by backpropagation:
We construct the point-wise saliency map as:
We pick the input points ( of the point cloud) associated to the top saliency scores, and we shift the position of each of these points toward the shape median:
We substitute these values in the original point cloud, and we restart the procedure from the beginning for times.
In Fig. A.9 we report further results of this procedure (in red, the points touched by these iterations).
D.2 Confusion matrix of classes.
For the sake of completeness, in Fig. A.10 we report the full confusion matrix for all the classes on the classification task. We observe that the task is particularly challenging, and several ambiguous cases exist. We believe this is due to the presence of objects with similar functions (e.g. various boxes, two chairs and a stool, etc.), but also the ambiguity of some datasets sequence (e.g., a human inspecting object without actually using it with a clear functionality). In future, the collection of other datasets designed explicitly for this task will significantly ease the learning.