MIME: Human-Aware 3D Scene Generation

Hongwei Yi, Chun-Hao P. Huang, Shashank Tripathi, Lea Hering, Justus Thies, Michael J. Black

Introduction

Humans constantly interact with their environment. They walk through a room, touch objects, rest on a chair, or sleep in a bed. All these interactions contain information about the scene layout and object placement. In fact, a mime is a performer who uses our understanding of such interactions to convey a rich, imaginary, 3D world using only their body motion. Can we train a computer to take human motion and, similarly, conjure the 3D scene in which it belongs? Such a method would have many applications in synthetic data generation, architecture, games, and virtual reality. For example, there exist large datasets of 3D human motion like AMASS and such data rarely contains information about the 3D scene in which it was captured. Could we take AMASS and generate plausible 3D scenes for all the motions? If so, we could use AMASS to generate training data containing realistic human-scene interaction.

To answer such questions, we train a new method called MIME (Mining Interaction and Movement to infer 3D Environments) that generates plausible indoor 3D scenes based on 3D human motion. Why is this possible? The key intuitions are that (1) A human’s motion through free space indicates the lack of objects, effectively carving out regions of the scene that are free of furniture. And (2), when they are in contact with the scene, this constrains both the type and placement of 3D objects; e.g, a sitting human must be sitting on something, such as a chair, a sofa, a bed, etc.

To make these intuitions concrete, we develop MIME, which is a transformer-based auto-regressive 3D scene generation method that, given an empty floor plan and a human motion sequence, predicts the furniture that is in contact with the human. It also predicts plausible objects that have no contact with the human but that fit with the other objects and respect the free-space constraints induced by the human motion. To condition the 3D scene generation with human motion, we estimate possible contact poses using POSA and divide the motion in contact and non-contact snippets. The non-contact poses define free-space in the room, which we encode as 2D floor maps, by projecting the foot vertices onto the ground plane. The contact poses and corresponding 3D human body models are represented by 3D bounding boxes of the contact vertices predicted by POSA. We use this information as input to the transformer and auto-regressively predict the objects that fulfill the contact and free-space constraints.

To train MIME, we built a new dataset called 3D-FRONT Human that extends the large-scale synthetic scene dataset 3D-FRONT . Specifically, we automatically populate the 3D scenes with humans, i.e., non-contact humans (a sequence of walking motion and standing humans) as well as contact humans (sitting, touching, and lying humans). To this end, we leverage motion sequences from AMASS , as well as static contact poses from RenderPeople scans.

At inference time, MIME is generating a plausible 3D scene layout for the input motion, represented as 3D bounding boxes. Based on this layout, we select 3D models from the 3D-FUTURE dataset and refine their 3D placement based on geometric constraints between the human poses and the scene.

In comparison to pure 3D scene generation baselines like ATISS , our method generates a 3D scene that supports human contact and motion while putting plausible objects in free space. In contrast to Pose2Room which is a recent pose-conditioned generative model, our method enables the generation of objects that are not in contact with the human, thus, predicting the entire scene instead of isolated objects. We demonstrate that our method can directly be applied to real captured motion sequences such as PROXD without finetuning.

In summary, we make the following contributions:

a novel motion-conditioned generative model for 3D room scenes that auto-repressively generates objects that are in contact with the human or avoid free-space defined by the motion.

a new 3D scene dataset with interacting humans and free space humans which is constructed by populating 3D FRONT with static contact/standing poses from RenderPeople and motion data of AMASS.

Related Work

Generative Scene Synthesis (No People). Most prior work on indoor scene synthesis, ignores the human and is based on (1) procedural modeling with grammars; (2) graph neural networks ; (3) autoregressive neural networks ; or (4) transformers . Some works leverage lexical text or a sentence as input to guide the 3D scene synthesis. Fisher et al. take 3D scans as input and synthesize the corresponding 3D object arrangements. This is extended to also include functionality aspects in the reconstruction. Recently, ATISS performs scene synthesis using a transformer-based architecture. ATISS takes a floorplan as input and auto-regressively generates a 3D scene that is represented as an unordered set of objects.

All methods mentioned above do not take human motion into consideration to guide the 3D scene synthesis. In contrast, we generate 3D scenes that are compatible with the humans defined by a given input motion. Specifically, the objects in the generated scene should support the human motion (e.g., a chair or couch for sitting) and should not collide with the path of a walking human, To this end, we build upon the auto-regressive scene synthesis architecture of ATISS and incorporate contact and free-space information into the pipeline.

Human-aware Scene Reconstruction. Qi et al. propose a method that synthesizes a 3D scene based on a human’s affordance map together with a spatial And-Or graph. PiGraphs learns a probability distribution over human pose and object geometry from interactions. It does not model the lack of interaction, i.e. the free space carved out by movement. Similarly, recent methods explore how to estimate a 3D scene from human behaviors and interactions. Mura et al. predict the “3D floor plan” from a 2D human walking trajectory in a deterministic way. The approach only indicates the room layout and furniture footprints and does not model objects or contact. Nie et al. propose Pose2Room, which predicts 3D objects inside a room from 3D human pose trajectories in a probabilistic way, by learning 3D object arrangement distribution. It only predicts contacted objects and can not generate objects in free space. In addition, it cannot take floor plans as input. We find these crucial in our experiments since object arrangements are highly related to the floor plan; e.g. some furniture is designed to go against a wall.

Human-Scene Interaction Datasets. Many datasets exist for understanding humans or scenes in separation, but relatively few address humans and scenes together. Human bodies are commonly captured using optical markers , IMU sensors , and multiple RGB cameras . See for a comprehensive review. These datasets contain only humans, forgoing the 3D environments which the subjects interact with, e.g., floor plane, walls, furniture. In contrast, real 3D scene datasets such as Matterport3D , ScanNet and Replica are captured primarily through time-of-flight sensors, where humans are excluded since only static content is reconstructed. Consequently, despite having a large variety of scenes, they are not suitable for modeling human-scene interaction.

To train MIME, we need diverse scene arrangement given a set of sparse or continuously-moving bodies. While recent real datasets capture both humans and environments, they fail to provide sufficient variety because the a priori scanned scenes are static and only the subject moves. This limits the variety of scenes that can be practically captured. Hassan et al. use mocap to capture a person interacting with objects like chairs, sofas and tables. They then augment the dataset by changing the size and shape of the objects and updating the human pose using inverse kinematics. The approach does not capture full scenes. For MIME we need a dataset with more variety. Composite or synthetic datasets such as are also widely used for human mesh recovery, but the meaningful human-scene interaction in them is fairly limited. To our knowledge, Pose2Room and GTA-IM are the closest to our needs. However, they represent humans with 3D skeletons, which cannot represent realistic contact between the body surface and the scene. Also the scene arrangement is still not rich enough to train a generative model. Thus, we introduce a new dataset called 3D FRONT Human, which is generated by populating 3D scenes from 3D FRONT with humans that move and interact with the scene.

Method

Given input motion of a human and an empty or partially occupied room of a specific kind (e.g., bedroom, living room, etc.) with its floor plan, we learn a generative model that can populate the room with objects that do not collide with the input humans and also support them. To this end, we propose a human-aware autoregressive model that represents scenes as one unordered set of objects. We divide the objects into two kinds, i.e., contact objects and non-contact objects, based on the human-object interaction. Contact objects are ones that humans interact with. Non-contact objects can be placed anywhere in the free space of a room that makes semantic sense. These objects enrich the content and potential functionality of a room.

In the following, we describe our human-aware scene synthesis model, MIME, which consists of two components: (1) a generative scene synthesis method based on 3D bounding boxes with object labels, and (2) a 3D refinement method that takes 3D human-scene interactions into account to optimize the rotation and placement of the generated objects. In Sec. 4, we detail the dataset generation process to train our model.

Given humans H\mathcal{H} and a floor plan F\mathcal{F}, our goal is to generate a “habitat” X={H,F,S}\mathcal{X}=\{\mathcal{H},\mathcal{F},\mathcal{S}\} where the 3D scene S\mathcal{S} can support all human interactions and motions. In contrast to the pure 3D scene generation methods , we focus on leveraging information from human motion to guide the 3D scene generation. To this end, we extract two types of information from the input motion and the corresponding human bodies: (i) contact humans C\mathcal{C} and (ii) free-space humans. We use POSA , to take posed human meshes and automatically label which of their vertices are potentially in contact with an object. Free-space humans are those that are only in contact with the floor plane, F\mathcal{F}. These define a binary mask that we call free-space mask FS\mathcal{FS}, which is constructed by the union of all projected foot contact points on F\mathcal{F}. This free-space mask FS\mathcal{FS} defines the region of a room that is free from objects as a human can stand and walk there. Given all contact humans, we compute the bounding boxes of their contact vertices and keep only the non-overlapping boxes using non-maximum suppression; we denote these as cic_{i}. The collection of contact boxes is referred to as C={ci}i=1N\mathcal{C}=\{c_{i}\}_{i=1}^{N}. Instead of storing all contact vertices of all bodies, our features are compact and encode complementary information. The contact humans, represented by C\mathcal{C}, indicate where to locate an object. See Fig. 2 top and middle rows for an illustration.

We represent a 3D scene S\mathcal{S} as an unordered set of objects, consisting of two kinds of objects based on human-object interaction. Objects in contact with the input human are referred to as contact objects O={o}i=1N\mathcal{O}=\{o\}_{i=1}^{N}, while non-contact objects Q={q}i=1M\mathcal{Q}=\{q\}_{i=1}^{M} are without any human interaction. Formally, a 3D scene is the union of contact and non-contact objects: S=O∪Q\mathcal{S}=\mathcal{O}\cup\mathcal{Q}.

The free-space mask FS\mathcal{FS}, the floor plan F\mathcal{F}, the contact humans C\mathcal{C} as well as the already existing objects S\mathcal{S} are input to an auto-regressive transformer model. Each input is encoded with a respective encoder, detailed below.

The log-likelihood of the generation of scene S\mathcal{S} including contact objects and non-contact objects is:

To calculate the likelihood of all generated contact objects Q\mathcal{Q}, we accumulate the likelihood of every contact object:

where p(oj∣o<j,F,FS,c≥j)p\left(o_{j}\mid o_{<j},\mathcal{F},\mathcal{FS},c_{\geq j}\right) is the probability of generating the jthj_{\text{th}} object conditioned on the input floor plan, free-space humans, the rest of contact humans and the previously generated objects, and π\pi is the random permutation function for those generated contact objects in the scene. The likelihood of all non-contact objects Q\mathcal{Q} is computed by replacing the input contact humans with the corresponding generated contact objects. During the training, we remove all contact humans inside the room, thus, all contact objects O\mathcal{O} can be treated as non-contact objects Q′\mathcal{Q^{\prime}}:

We follow to use Monte Carlo sampling to approximate all different object permutations during training, to make our model invariant to the order of generated objects.

The 2D free-space mask FS\mathcal{FS} is encoded together with the 2D floor plan F\mathcal{F} using a ResNet-18 . The encoded feature provides the information to the transformer encoder about where an object can be placed.

Contact Encoder.

We represent the contact humans as 3D bounding boxes, which consist of the contact label II, the contact class category kk (sitting, touching, lying), the translation tt, the rotation rr, and the size ss. During generation of a scene, we set the contact label II of one contact human to 11 while the others are labeled . This label highlights the contribution of the specific contact human to the next generated contacted object. Note that we remove contact humans from the input set if they are already in contact with an existing object in the scene. Otherwise, we encode the jthj_{th} input contact human by applying:

where λ(⋅)\lambda\left(\cdot\right) is a learnable embedding for the contact class category kk, and p(⋅)p\left(\cdot\right) is the positional encoding for the translation tt, rotation rr and size ss.

Furniture Encoder.

The furniture encoder computes the embedding of existing objects in the room:

Note that the furniture encoder is sharing the same weight as the contact encoder. The contact labels of the objects are all zero, where j∈[1,M]j\in[1,M].

Scene Synthesis Transformer.

To decode the attribute distribution (k^,t^,r^,s^)(\hat{k},\hat{t},\hat{r},\hat{s}) of the generated object oM+1{o_{M+1}} from q^\hat{q}, we follow the same design from ATISS . Specifically, we employ an MLP for each attribute in a consecutive fashion. Given q^\hat{q}, we first predict the class category label k^\hat{k}, then we predict the t^\hat{t}, r^\hat{r} and s^\hat{s} in this specific order, where the previous attribute will be concatenated with the input q^\hat{q} for the next attribution prediction.

2 Training and Inference.

We train our model on the training set of 3D FRONT HUMAN, by maximizing the log-likelihood of each generated scene S\mathcal{S} in Eq. (1). During training, we select a human-populated scene in 3D FRONT HUMAN and add a random permutation π(⋅)\pi\left(\cdot\right) on all N\mathit{N} contact and M\mathit{M} non-contact objects. We randomly select the m\textslth+1m_{\textsl{th}}+1 as the generated object, where m∈[0,N+M]m\in[0,N+M]. Note that, m=0m=0 represents an empty scene, while m=N+Mm=\mathit{N+M} indicates the generated scene is already full and the class label of the predicted object is an extra end symbol. Our model predicts the attribute distribution of the generated object, conditioned on the floor plane F\mathcal{F}, free space FS\mathcal{FS}, previous mm objects and contact humans C\mathcal{C}; see Fig. 3. To enable our model to generate both contact objects and non-contact objects, we make a data augmentation for adding input contact humans or dropping them out in equal frequency.

During inference, we start from an empty floor plane FF with input humans including free-space humans FS\mathcal{FS}, and contact humans C\mathcal{C}. We autoregressively sample the attribute of the next generated object to put one object into a scene. By default, we set the contact label of the first contact human to 11, and the rest are . After each generation step, we remove contact humans that are already in contact, by computing the 2D IoU of the human bounding box and the generated object by projecting them on the ground plane. Specifically, if the IoU is larger than 0.50.5, we remove the contact human from the input. Once the end symbol is generated, the generated scene is finished.

3 3D Scene Refinement

The generated scene from our model is represented with 3D bounding boxes. Based on the bounding box size and class category label, we retrieve the closest mesh model from 3D FUTURE . To improve the human-scene interaction between the generated scenes and input humans, we apply the collision loss and the contact loss from MOVER to refine the object position, as can be seen in Fig. 4. We calculate a unified SDF volume and accumulate all contact vertices for all humans in the 3D space, and jointly optimize the object alignment to improve human-object contact and resolve 3D interpenetrations between humans and the scene. The MOVER contact loss weight and the collision loss weight are 1e51e5 and 1e31e3 respectively.

Dataset Generation of 3D FRONT HUMAN

To enable 3D scene generation from humans, we need a dataset that consists of large numbers of rooms with a wide variety of human interactions. Since no such dataset exists, we generate a new synthetic dataset by populating the 3D rooms in the 3D FRONT with interactive humans. We name the resulting dataset 3D FRONT HUMAN. To populate the rooms of 3D FRONT with people, we insert humans with contact and humans that stand or walk in free space, as shown in Fig. 5. We represent people with the SMPL-X model and add contact humans from RenderPeople by randomly assigning plausible interactions to different contactable objects in the room. Specifically, we allow for three types of contact interactions: touching, sitting, and lying. In Fig. 5 (bottom), we put a lying down person on a bed, and multiple humans interact with a nightstand or wardrobe. In the free space, we put a random number of static standing people and add multiple walking motion clips from AMASS with random start positions and directions to the scene, and remove humans that intersect with objects in the scene.

Experiments

We qualitatively and quantitatively evaluate our method and compare with two baselines. Specifically, we compare to the 3D scene generation method ATISS and the human-aware scene reconstruction method Pose2Room .

Our human-populated dataset 3D FRONT HUMAN contains four room types: 1) 56895689 bedrooms, 2) 29872987 living rooms, 3) 25492549 dining rooms and 4) 679679 libraries. We use 2121 object categories for the bedrooms, 2424 for the living and dining rooms, and 2525 for the libraries. We independently train our model four times on the four kinds of rooms. Following our baseline ATISS , for each kind of room, we split the data 80%, 10%, 10% into training, validation and test sets. We train and validate MIME on the training and validation sets respectively, and evaluate it on the test set. See more details in Sup. Mat. Since ATISS does not provide a pretrained model, we retrain it with the official code1{}^{\text{1}} following the same training strategy on the original 3D FRONT dataset as one of our baseline. https://github.com/nv-tlabs/ATISS/commit/6b46c11.

To evaluate the effectiveness and generalization of our method, we test MIME on a real RGB-D motion captured dataset PROX-D and compare it with Pose2Room . Pose2Room needs a sequence of human motions that are in contact with objects. Our 3D FRONT HUMAN does not provide these interactive human-object motions, so we cannot enable fine tune and evaluate Pose2Room on 3D FRONT HUMAN.

Evaluation Metrics.

We compare MIME with the baselines in two different ways: (i) the plausibility between human-scene interaction and (ii) the realism of the generated scenes only. We propose a interpenetration loss (↓\downarrow) to evaluate the collision between the generated objects and the free space, through computing the ratio of the violated free space by the 2D projection of the generated objects:

where pp denotes each pixel on the floor plane image. We calculate the 2D IoU and 3D IoU between generated objects and input contact bounding boxes to measure the human-object interaction. To evaluate the realism and diversity of generated scenes, we follow common practice and calculate the FID (at 2562256^{2} resolution) score between bird-eye view orthographic projections of generated scenes and real scenes from the test set, as well as the category KL divergence. We compute the FID score 10 times and report the mean and variance of it. All these evaluation experiments are conducted on the test split of the 3D FRONT HUMAN dataset.

1 Human-aware Scene Synthesis.

In Fig. 6, we visualize the ability of our method to generate plausible 3D scenes from input motion and floor plans for different kinds of rooms; we also show our baseline methods for comparison. See Sup. Mat. for more examples. Note that the original ATISS model generates a 3D scene only based on the floor plan, without taking the humans into account. Thus, generated scenes from ATISS violate free space constraint and are not consistent with the human contact. For a more fair comparison, we extend ATISS to take information about the human motion as input. Specifically, we adapt the 2D input floor plan to also contain the free space information of the walking and standing humans. However, ATISS with input free space still generates objects in free space, while also generating implausible object configurations such as the white closet inside the bed (Fig. 6, top). In contrast, MIME generates plausible 3D scenes that have less interpenetration with the free space and support interacting humans; e.g. a bed beneath a lying person and a chair under a sitting person.

The observations in the qualitative comparison are also confirmed by a quantitative evaluation in Tab. 1. MIME achieves significant improvements on human-scene interaction evaluation metrics compared with ATISS. Note, since our scene generation is constrained by the input motion, the diversity scores (FID, KL divergence) are lower than of ATISS, which is not human-aware. This is not a failure/limitation of MIME.

To evaluate the generalization of our method, we test it on a real dataset of human motion. We consider the PROXD dataset and the 3D bounding box annotation from . We use it without finetuning, and use the motions to generate scenes. We compare our method with Pose2Room , which predicts 3D objects from a motion sequence of 3D skeletons. Note that Pose2Room can only predict contact objects, it does not predict an entire scene which is the goal of our method. Fig. 7 presents a qualitative comparison of the methods and we report the quantitative metrics in Tab. 2. Specifically, we compute the mean average precision with 3D IoU 0.50.5 (mAP@0.5) to evaluate the 3D object detection accuracy for those contact objects only. Both methods are probabilistic generative models to predict the object attribute distribution. Following Pose2Room, we use the same 55 input motions and sample 10 scenes for each motion sequence, and report the mean value of it. Our method achieves better 3D object detection accuracy compared to Pose2Room without pretraining.

2 Ablation Study on Input Humans

In Fig. 8, we evaluate the influence of the density of free-space humans, and the number of contact humans, that we provide as input to MIME. We observe that MIME generates contact objects according to the number of contact humans and, as the density of free-space humans increases, MIME generates fewer objects in scenes. This is as expected.

Discussion

Given a sequence of human motions, MIME generates diverse and plausible scenes with which the humans interact. We assume that the generated scenes are static, and future work should explore generating moving objects by exploring the interaction between humans and moving objects, such as moving a chair, grasping a cup, opening a door, etc.

MIME, like ATISS, needs a pre-defined floor plan room layout as input. The resolution of the 2D floor plan is coarse; i.e., 1 pixel stands for around 10 centimeters, which is extracted as a 512512 dimension feature by ResNet-18. Introducing a finer floor plan representation, such as dividing one floor plan into multiple patches (cf. ViT) or simply enlarging the size of the feature dimension could improve the generated object placement, resulting in less collision between the humans and the free space. Another interesting direction is to estimate a floor plan and 3D object layout jointly from input humans only.

During inference, MIME uses a hand-crafted metric 2D IoU between the generated objects and the input contact humans to factor out which human it is in contacted with. A simple extension would be to use the network to learn this information. Our model directly estimates 3D bounding boxes as a 3D scene representation, followed by a scene refinement that places the mesh models into the scene. Learning to directly estimate the mesh models from the interacting humans is another promising direction.

Conclusion

We have introduced MIME, which generates varied furniture layouts that are consistent with input human movement and contacts. To train MIME, we built a new dataset called 3D FRONT HUMAN, by populating humans into the large-scale synthetic scene dataset . We have demonstrated that by incorporating input human motion into free space and contact boxes, our method can generate multiple realistic scenes, where the input motion can take place. MIME has many applications, particularly for generating synthetic training data at scale. MIME provides a means of taking existing human motion capture data and “upgrading” it to include plausible 3D scenes that are consistent with it.

Acknowledgments. We thank Despoina Paschalidou, Wamiq Para for useful feedback about the reimplementation of ATISS, and Yuliang Xiu, Weiyang Liu, Yandong Wen, Yao Feng for the insightful discussions, and Benjamin Pellkofer for IT support. This work was supported by the German Federal Ministry of Education and Research (BMBF): Tübingen AI Center, FKZ: 01IS18039B.

Disclosure. MJB has received research gift funds from Adobe, Intel, Nvidia, Meta/Facebook, and Amazon. MJB has financial interests in Amazon, Datagen Technologies, and Meshcapade GmbH. JT has received research gift funds from Microsoft Research.

References

Appendix A Training Details

During training, we apply the Adam optimizer with learning rate 1e−41e^{-4} and no weight decay. In Adam optimizer, we use the default PyTorch implemented parameters, i.e., β1=0.9\beta_{1}=0.9, β2=0.999\beta_{2}=0.999 and ϵ=1e−8\epsilon=1e-8. We train MIME with the batch size 128128 for 100k100k iterations. We perform random global rotation augmentation between $$ degrees on the holistic populated scene, including the floor plane, all objects, the free space and all contact humans.

Appendix B More Qualitative Examples

We present more qualitative examples for different kinds of rooms, in Fig. 9, Fig. 10, and Fig. 11. Compared with our baseline methods , our method can generate more plausible 3D scenes that input motions can interact with.