Populating 3D Scenes by Learning Human-Scene Interaction
Mohamed Hassan, Partha Ghosh, Joachim Tesch, Dimitrios Tzionas, Michael J. Black
Introduction
Humans constantly interact with the world around them. We move by walking on the ground; we sleep lying on a bed; we rest sitting on a chair; we work using touchscreens and keyboards. Our bodies have evolved to exploit the affordances of the natural environment and we design objects to better “afford” our bodies. While obvious, it is worth stating that these physical interactions involve contact. Despite the importance of such interactions, existing representations of the human body do not explicitly represent, support, or capture them.
In computer vision, human pose is typically estimated in isolation from the 3D scene, while in computer graphics 3D scenes are often scanned and reconstructed without people. Both the recovery of humans in scenes and the automated synthesis of realistic people in scenes remain challenging problems. Automation of this latter case would reduce animation costs and open up new applications in augmented reality. Here we take a step towards automating the realistic placement of 3D people in 3D scenes with realistic contact and semantic interactions (Fig. 1). We develop a novel body-centric approach that relates 3D body shape and pose to possible world interactions. Learned parametric 3D human models represent the shape and pose of people accurately. We employ the SMPL-X model, which includes the hands and face, as it supports reasoning about contact between the body and the world.
While such body models are powerful, we make three key observations. First, human models like SMPL-X do not explicitly model contact. Second, not all parts of the body surface are equally likely to be in contact with the scene. Third, the poses of our body and scene semantics are highly intertwined. Imagine a person sitting on a chair; body contact likely includes the buttocks, probably also the back, and maybe the arms. Think of someone opening a door; their feet are likely in contact with the floor, and their hand is in contact with the doorknob.
Based on these observations, we formulate a novel model, that makes human-scene interaction (HSI) an explicit and integral part of the body model. The key idea is to encode HSI in an ego-centric representation built in SMPL-X. This effectively extends the SMPL-X model to capture contact and the semantics of HSI in a body-centric representation. We call this POSA for “Pose with prOximitieS and contActs”. Specifically, for every vertex on the body and every pose, POSA defines a probabilistic feature map that encodes the probability that the vertex is in contact with the world and the distribution of semantic labels associated with that contact.
POSA is a conditional Variational Auto-Encoder (cVAE), conditioned on SMPL-X vertex positions. We train on the PROX dataset , which contains subjects, fit with SMPL-X meshes, interacting with real 3D scenes. We also train POSA using use the scene semantic annotations provided by the PROX-E dataset . Once trained, given a posed body, we can sample likely contacts and semantic labels for all vertices. We show the value of this representation with two challenging applications.
First, we focus on automatic scene population as illustrated in Fig. 1. That is, given a 3D scene and a body in a particular pose, where in the scene is this pose most likely? As demonstrated in Fig. 1 we use SMPL-X bodies fit to commercial 3D scans of people , and then, conditioned on the body, our cVAE generates a target POSA feature map. We then search over possible human placements while minimizing the discrepancy between the observed and target feature maps. We quantitatively compare our approach to PLACE , which is SOTA on a similar task, and find that POSA has higher perceptual realism.
Second, we use POSA for monocular 3D human pose estimation in a 3D scene. We build on the PROX method that hand-codes contact points, and replace these with our learned feature map, which functions as an HSI prior. This automates a heuristic process, while producing lower pose estimation errors than the original PROX method.
To summarize, POSA is a novel model that intertwines SMPL-X pose and scene semantics with contact. To the best of our knowledge, this is the first learned human body model that incorporates HSI in the model. We think this is important because such a model can be used in all the same ways that models like SMPL-X are used but now with the addition of body-scene interaction. The key novelty is posing HSI as part of the body representation itself. Like the original learned body models, POSA provides a platform that people can build on. To facilitate this, our model and code are available for research purposes at https://posa.is.tue.mpg.de.
Related Work
Humans & Scenes in Isolation: For scenes , most work focuses on their 3D shape in isolation, e.g. on rooms devoid of humans or on objects that are not grasped . For humans, there is extensive work on capturing or estimating their 3D shape and pose, but outside the context of scenes.
Human Models: Most work represents 3D humans as body skeletons . However, the 3D body surface is important for physical interactions. This is addressed by learned parametric 3D body models . For interacting humans we employ SMPL-X , which models the body with full face and hand articulation.
HSI Geometric Models: We focus on the spatial relationship between a human body and objects it interacts with. Gleicher uses contact constraints for early work on motion re-targeting. Lee et al. generate novel scenes and motion by deformably stitching “motion patches”, comprised of scene patches and the skeletal motion in them. Lin et al. generate 3D skeletons sitting on 3D chairs, by manually drawing 2D skeletons and fitting 3D skeletons that satisfy collision and balance constraints. Kim et al. automate this, by detecting sparse contacts on a 3D object mesh and fitting a 3D skeleton to contacts while avoiding penetrations. Kang et al. reason about the physical comfort and environmental support of a 3D humanoid, through force equilibrium. Leimer et al. reason about pressure, frictional forces and body torques, to generate a 3D object mesh that comfortably supports a given posed 3D body. Zheng et al. map high-level ergonomic rules to low-level contact constraints and deform an object to fit a 3D human skeleton for force equilibrium. Bar-Aviv et al. and Liu et al. use an interacting agent to describe object shapes through detected contacts , or relative distance and orientation metrics . Gupta et al. estimate human poses “afforded” in a depicted room, by predicting a 3D scene occupancy grid, and computing support and penetration of a 3D skeleton in it. Grabner et al. detect the places on a 3D scene mesh where a 3D human mesh can sit, modeling interaction likelihood with GMMs and proximity and intersection metrics. Zhu et al. use FEM simulations for a 3D humanoid, to learn to estimate forces, and reasons about sitting comfort.
Several methods focus on dynamic interactions . Ho et al. compute an “interaction mesh” per frame, through Delaunay tetrahedralization on human joints and unoccluded object vertices; minimizing their Laplacian deformation maintains spatial relationships. Others follow an object-centric approach . Al-Asqhar et al. sample fixed points on a 3D scene, using proximity and “coverage” metrics, and encode 3D human joints as transformations w.r.t. these. Pirk et al. build functional object descriptors, by placing “sensors” around and on 3D objects, that “sense” 3D flow of particles on an agent.
HSI Data-driven Models: Recent work takes a data-driven approach. Jiang et al. learn to estimate human poses and object affordances from an RGB-D scene, for 3D scene label estimation. SceneGrok learns action-specific classifiers to detect the likely scene places that “afford” a given action. Fisher et al. use SceneGrok and interaction annotations on CAD objects, to embed noisy 3D room scans to CAD mesh configurations. PiGraphs maps pairs of {verb-object} labels to “interaction snapshots”, i.e. 3D interaction layouts of objects and a human skeleton. Chen et al. map RGB images to “interaction snapshots”, using Markov Chain Monte Carlo with simulated annealing to optimize their layout. iMapper maps RGB videos to dynamic “interaction snapshots”, by learning “scenelets” on PiGraphs data and fitting them to videos. Phosa infers spatial arrangements of humans and objects from a single image. Cao et al. map an RGB scene and 2D pose history to 3D skeletal motion, by training on video-game data. Li et al. follow to collect 3D human skeletons consistent with 2D/3D scenes of , and learn to predict them from a color and/or depth image. Corona et al. use a graph attention model to predict motion for objects and a human skeleton, and their evolving spatial relationships.
Another HSI variant is Hand-Object Interaction (HOI); we discuss only recent work . Brahmbhatt et al. capture fixed 3D hand-object grasps, and learn to predict contact; features based on object-to-hand mesh distances outperform skeleton-based variants. For grasp generation, 2-stage networks are popular . Taheri et al. capture moving SMPL-X humans grasping objects, and predict MANO hand grasps for object meshes, whose 3D shape is encoded with BPS . Corona et al. generate MANO grasping given an object-only RGB image; they first predict the object shape and rough hand pose (grasp type), and then they refine the latter with contact constraints and an adversarial prior.
Closer to us, PSI and PLACE populate 3D scenes with SMPL-X humans. Zhang et al. train a cVAE to estimate humans from a depth image and scene semantics. Their model provides an implicit encoding of HSI. Zhang et al. , on the other hand, explicitly encode the scene shape and human-scene proximal relations with BPS , but do not use semantics. Our key difference to is our human-centric formulation; inherently this is more portable to new scenes. Moreover, instead of the sparse BPS distances of , we use dense body-to-scene contact, and also exploit scene semantics like .
Method
Our training data corpus is a set of pairs of 3D meshes
2 POSA Representation for HSI
We encode the relationship between the human mesh and the scene mesh in an egocentric feature map that encodes per-vertex features on the SMPL-X mesh . We define as:
where is the contact label and is the semantic label of the contact point. is the feature dimension.
For each vertex on the body, , we find its closest scene point . Then we compute the distance :
Given , we can compute whether a is in contact with the scene or not, with :
The contact threshold is chosen empirically to be cm. The semantic label of the contacted surface is a one-hot encoding of the object class:
where is the number of object classes. The sizes of , , and are , and respectively. All the features are computed once offline to speed training up. A visualization of the proposed representation is in Fig. 2.
3 Learning
Our goal is to learn a probabilistic function from body pose and shape to the feature space of contact and semantics. That is, given a body, we want to sample labelings of the vertices corresponding to likely scene contacts and their corresponding semantic label. Note that this function, once learned, only takes the body as input and not a scene – it is a body-centric representation.
To train this we use the PROX dataset, which contains bodies in 3D scenes. We also use the scene semantic annotations from the PROX-E dataset . For each body mesh , we factor out the global translation and rotation and around the and axes. The rotation around the axis is essential for the model to differentiate between, e.g., standing up and lying down.
Given pose and shape parameters in a given training frame, we compute a . This gives vertices from which we compute the feature map that encodes whether each is in contact with the scene or not, and the semantic label of the scene contact point .
where and are the reconstructed contact and semantic labels, KL denotes the Kullback Leibler divergence, and denotes the reconstruction loss. BCE and CCE are the binary and categorical cross entropies respectively. The is a hyperparameter inspired by Gaussian -VAEs , which regularizes the solution; here . encourages the reconstructed samples to resemble the input, while encourages to match a prior distribution over , which is Gaussian in our case. We set the values of and to .
Since is defined on the vertices of the body mesh , this enables us to use graph convolution as our building block for our VAE. Specifically, we use the Spiral Convolution formulation introduced in . The spiral convolution operator for node in the body mesh is defined as:
where denotes layer in a multi-layer perceptron (MLP) network, and is a concatenation operation of the features of neighboring nodes, . The spiral sequence is an ordered set of vertices around the central vertex . Our architecture is shown in Fig. 3. More implementation details are included in the Sup. Mat. For details on selecting and ordering vertices, please see .
Experiments
We perform several experiments to investigate the effectiveness and usefulness of our proposed representation and model under different use cases, namely generating HSI features, automatically placing 3D people in scenes, and improving monocular 3D human pose estimation.
We evaluate the generative power of our model by sampling different feature maps conditioned on novel poses using our trained decoder , where and is the randomly generated feature map. This is equivalent to answering the question: “In this given pose, which vertices on the body are likely to be in contact with the scene, and what object would they contact?” Randomly generated samples are shown in Fig. 4.
We observe that our model generalizes well to various poses. For example, notice that when a person is standing with one hand pointing forward, our model predicts the feet and the hand to be in contact with the scene. It also predicts the feet are in contact with the floor and hand is in contact with the wall. However this changes for the examples when a person is in a lying pose. In this case, most of the vertices from the back of the body are predicted to be in contact (blue color) with a bed (light purple) or a sofa (dark green).
These features are predicted from the body alone; there is no notion of “action” here. Pose alone is a powerful predictor of interaction. Since the model is probabilistic, we can sample many possible feature maps for a given pose.
2 Affordances: Putting People in Scenes
Given a posed 3D body and a 3D scene, can we place the body in the scene so that the pose makes sense in the context of the scene? That is does the pose match the affordances of the scene ? Specifically, given a scene, , semantic labels of objects present, and a body mesh, , our method finds where in this given pose is likely to happen. We solve this problem in two steps.
First, given the posed body, we use the decoder of our cVAE to generate a feature map by sampling as in Sec. 4.1. Second, we optimize the objective function:
where is the body translation, is the global body orientation and is the body pose. The afforance loss
and are the observed distance and semantic labels, which are computed using Eq. 2 and Eq. 4 respectively. and are the generated contact and semantic labels, and denotes dot product. and are and respectively. is a penetration penalty to discourage the body from penetrating the scene:
. is a regularizer that encourages the estimated pose to remain close to the initial pose of :
Although the body pose is given, we optimize over it, allowing the parameters to change slightly since the given pose might not be well supported by the scene. This allows for small pose adjustment that might be necessary to better fit the body into the scene. .
The input posed mesh, , can come from any source. For example, we can generate random SMPL-X meshes using VPoser which is a VAE trained on a large dataset of human poses. More interestingly, we use SMPL-X meshes fit to realistic Renderpeople scans (see Fig. 1).
We tested our method with both real (scanned) and synthetic (artist generated) scenes. Example bodies optimized to fit in a real scene from the PROX test set are shown in Fig. 5 (top); this scene was not used during training. Note that people appear to be interacting naturally with the scene; that is, their pose matches the scene context. Figure 5 (bottom) shows bodies automatically placed in an artist-designed scene (Archviz Interior Rendering Sample, Epic Games)https://docs.unrealengine.com/en-US/Resources/Showcases/ArchVisInterior/index.html. POSA goes beyond previous work to produce realistic human-scene interactions for a wide range of poses like lying down and reaching out.
While the poses look natural in the above results, the SMPL-X bodies look out of place in realistic scenes. Consequently, we would like to render realistic people instead, but models like SMPL-X do not support realistic clothing and textures. In contrast, scans from companies like Renderpeople (Renderpeople GmbH, Köln) are realistic, but have a different mesh topology for every scan. The consistent topology of a mesh like SMPL-X is critical to learn the feature model.
Clothed Humans: We address this issue by using SMPL-X fits to clothed meshes from the AGORA dataset . We then take the SMPL-X fits and minimize an energy function similar to Eq. 9 with one important change. We keep the pose, , fixed:
Since the pose does not change, we just replace the SMPL-X mesh with the original clothed mesh after the optimization converges; see Sup. Mat. for details.
Qualitative results for real scenes (Replica dataset ) are shown in Fig. 6, and for a synthetic scene in Fig. 1. More results are shown in Sup. Mat. and in our video.
We quantiatively evaluate POSA with two perceptual studies. In both, subjects are shown a pair of two rendered scenes, and must choose the one that best answers the question “Which one of the two examples has more realistic (i.e. natural and physically plausible) human-scene interaction?” We also evaluate physical plausibility and diversity.
Comparison to PROX ground truth: We follow the protocol of Zhang et al. and compare our results to randomly selected examples from PROX ground truth. We take real scenes from the PROX test set, namely MPH16, MPH1Library, N0SittingBooth and N3OpenArea. We take SMPL-X bodies from the AGORA dataset, corresponding to different 3D scans from Renderpeople. We take each of these bodies and sample one feature map for each, using our cVAE. We then automatically optimize the placement of each sample in all the scenes, one body per scene. For unclothed bodies (Tab. 1, rows 1-3), this optimization changes the pose slightly to fit the scene (Eq. 9). For clothed bodies (Tab. 1, rows 4-5), the pose is kept fixed (Eq. 13). For each variant, this optimization results in unique body-scene pairs. We render each 3D human-scene interaction from views so that subjects are able to get a good sense of the 3D relationships from the images. Using Amazon Mechanical Turk (AMT), we show these results to different subjects. This results in unique ratings. The results are shown in Tab. 1. POSA (contact only) and the state-of-the-art method PLACE are both almost indistinguishable from the PROX ground truth. However, the proposed POSA (contact + semantics) (row ) outperforms both POSA (contact only) (row ) and PLACE (row ), thus modeling scene semantics increases realism. Lastly, the rendering of high quality clothed meshes (bottom two rows) influences the perceived realism significantly.
Comparison between POSA and PLACE: We follow the same protocol as above, but this time we directly compare POSA and PLACE. The results are shown in Tab. 2. Again, we find that adding semantics improves realism. There are likely several reasons that POSA is judged more realistic than PLACE. First, POSA employs denser contact information across the whole SMPL-X body surface, compared to PLACE’s sparse distance information through its BPS representation. Second, POSA uses a human-centric formulation, as opposed to PLACE’s scene-centric one, and this can help generalize across scenes better. Third, POSA uses semantic features that help bodies do the right thing in the scene, while PLACE does not. When human generation is imperfect, inappropriate semantics may make the result seem worse. Fourth, the two methods are solving slightly different tasks. PLACE generates a posed body mesh for a given scene, while our method samples one from the AGORA dataset and places it in the scene using a generated POSA feature map. While this gives PLACE an advantage, because it can generate an appropriate pose for the scene, it also means that it could generate an unnatural pose, hurting realism. In our case, the poses are always “valid” by construction but may not be appropriate for the scene. Note that, while more realistic than prior work, the results are not always fully natural; sometimes people sit in strange places or lie where they usually would not.
Physical Plausibility: We take bodies from the AGORA dataset and place all of them in each of the test scenes of PROX, leading to a total of samples, following . Given a generated body mesh, , a scene mesh, , and a scene signed distance field (SDF) that stores distances for each voxel , we compute the following scores, defined by Zhang et al. : (1) the non-collision score for each , which is the ratio of body mesh vertices with positive SDF values divided by the total number of SMPL-X vertices, and (2) the contact score for each , which is if at least one vertex of has a non-positive value. We report the mean non-collision score and mean contact score over all samples in Tab. 3; higher values are better for both metrics. POSA and PLACE are comparable under these metrics.
Diversity Metric: Using the same samples, we compute the diversity metric from . We perform K-means () clustering of the SMPL-X parameters of all sampled poses, and report: (1) the entropy of the cluster sizes, and (2) the cluster size, i.e. the average distance between the cluster center and the samples belonging in it. See Tab. 4; higher values are better. While PLACE generates poses and POSA samples them from a database, there is little difference in diversity.
3 Monocular Pose Estimation with HSI
Traditionally, monocular pose estimation methods focus only on the body and ignore the scene. Hence, they tend to generate bodies that are inconsistent with the scene. Here, we compare directly with PROX , which adds contact and penetration constraints to the pose estimation formulation. The contact constraint snaps a fixed set of contact points on the body surface to the scene, if they are “close enough”. In PROX, however, these contact points are manually selected and are independent of pose.
We replace the hand-crafted contact points of PROX with our learned feature map. We fit SMPL-X to RGB image features such that the contacts are consistent with the 3D scene and its semantics. Similar to PROX, we build on SMPLify-X . Specifically, SMPLify-X optimizes SMPL-X parameters to minimize an objective function of multiple terms: the re-projection error of 2D joints, priors and physical constraints on the body;
where represents the pose parameters of the body, face (neck, jaw) and the two hands, , denotes the body translation, and the body shape. is a re-projection loss that minimizes the difference between 2D joints estimated from the RGB image and the 2D projection of the corresponding posed 3D joints of SMPL-X. is a prior penalizing extreme bending only for elbows and knees. The term penalizes self-penetrations. For details please see .
We turn off the PROX contact term and optimize Eq. 14 to get a pose matching the image observations and roughly obeying scene constraints. Given this rough body pose, which is not expected to change significantly, we sample features from and keep these fixed. Finally, we refine by minimizing
where represents the SMPLify-X energy term as defined in Eq. 14, are the generated contact labels, is the observed distance, and represents the body-scene penetration loss as in Eq. 11. We compare our results to standard PROX in Tab. 5. We also show the results of RGB-only baseline introduced in PROX for reference. Using our learned feature map improves accuracy over the PROX’s heuristically determined contact constraints.
Conclusions
Traditional 3D body models, like SMPL-X, model the a priori probability of possible body shapes and poses. We argue that human poses in isolation from the scene, make little sense. We introduce POSA, which effectively upgrades a 3D human body model to explicitly represent possible human-scene interactions. Our novel, body-centric, representation encodes the contact and semantic relationships between the body and the scene. We show that this is useful and supports new tasks. For example, we consider placing a 3D human into a 3D scene. Given a scan of a person with a known pose, POSA allows us to search the scene for locations where the pose is likely. This enables us to populate empty 3D scenes with higher realism than the state of the art. We also show that POSA can be used for estimating human pose from an RGB image, and that the body-centered HSI representation improves accuracy. In summary, POSA is a good step towards a richer model of human bodies that goes beyond pose to support the modeling of HSI.
Limitations: POSA requires an accurate scene SDF; a noisy scene mesh can lead to penetration between the body and scene. POSA focuses on a single body mesh only. Penetration between clothing and the scene is not handled and multiple bodies are not considered. Optimizing the placement of people in scenes is sensitive to initialization and is prone to local minima. A simple user interface would address this, letting naive users roughly place bodies, and then POSA would automatically refine the placement.
Acknowledgements: We thank M. Landry for the video voice-over, B. Pellkofer for the project website, and R. Danecek and H. Yi for proofreading. This work was partially supported by the International Max Planck Research School for Intelligent Systems (IMPRS-IS). Disclosure: MJB has received research funds from Adobe, Intel, Nvidia, Facebook, and Amazon. While MJB is a part-time employee of Amazon, his research was performed solely at, and funded solely by, Max Planck. MJB has financial interests in Amazon, Datagen Technologies, and Meshcapade GmbH.
References
Appendix A Training Details
The global orientation of the body is typically irrelevant in our body-centric representation, so we rotate the training bodies around the and axes to put them in a canonical orientation. The rotation around the axis, however, is essential to enable the model to differentiate between standing up and lying down. The semantic labels for the PROX scenes are taken from Zhang et al. , where scenes were manually labeled following the object categorization of Matterport3D , which incorporates object categories.
Our encoder-decoder architecture is similar to the one introduced in Gong et al. . The encoder consists of spiral convolution layers interleaved with pooling layers . Pool stands for a downsampling operation as in COMA , which is based on contracting vertices. FC is a fully connected layer and the number in the bracket next to it denotes the number of units in that layer. We add additional fully connected layers to predict the parameters of the latent code, with fully connected layers of units each. The input to the encoder is a body mesh where, for each vertex, , we concatenate vertex positions, and vertex features. For computational efficiency, we first downsample the input mesh by a factor of . So instead of working on the full mesh resolution of vertices, our input mesh has a resolution of vertices. The decoder architecture consists of spiral convolution layers only . We attach the latent vector to the 3D coordinates of each vertex similar to Kolotouros et al. .
We build our model using the PyTorch framework. We use the Adam optimizer , batch size of , and learning rate of without learning rate decay.
Appendix B SDF Computation
Appendix C Random Samples
We show multiple randomly sampled feature maps for the same pose in Fig. S.1. Note how POSA generate a variety of valid feature maps for the same pose. Notice for example that the feet are always correctly predicted to be in contact with the floor. Sometimes our model predicts the person is sitting on a chair (far left) or on a sofa (far right).
The predicted semantic map is not always accurate as shown in the far right of Fig. S.1. The model predicts the person to be sitting on a sofa but at the same time predicts the lower parts of the leg to be in contact with a bed which is unlikely.
Appendix D Affordance Detection
The complete pipeline of the affordance detection task is shown in Fig. S.2. Given a clothed 3D mesh that we want to put in a scene, we first need a SMPL-X fit to the mesh; here we take this from the AGORA dataset . Then we generate a feature map using the decoder of our cVAE by sampling . Next we minimize the energy function in Eq. 16.
Finally, we replace the SMPL-X mesh with the original clothed.
We show additional qualitative examples of SMPL-X meshes automatically placed in real and synthetic scenes in Fig. S.3. Qualitative examples of clothed bodies placed in real and synthetic scenes are shown in Fig. S.4. We show qualitative comparison between our results and PLACE in Fig. S.5.
Appendix E Failure Cases
We show representative failure cases in Fig. S.6. A common failure mode is residual penetrations; even with the penetration penalty the body can still penetrate the scene. This can happen due to thin surfaces that are not captured by our SDF and/or because the optimization becomes stuck in a local minimum. In other cases, the feature map might not be right. This can happen when the model does not generalize well to test poses due to the limited training data.
Appendix F Effect of Shape
Fig. S.7 shows that our model can predict plausible feature maps for a wide range of human body shapes.
Appendix G Scene population.
In Fig. S.8 we show the three main steps to populate a scene: (1) Given a scene, we create a regular grid of candidate positions (Fig. S.8.1). We place the body, in a given pose, at each candidate position and evaluate Eq. 10 once. (2) We then keep the best candidates with the lowest energy (Fig. S.8.2), and (3) iteratively optimize Eq. 10 for these; Fig. S.8.3 shows results at three positions, with the best one highlighted with green.