Human-Aware Object Placement for Visual Environment Reconstruction

Hongwei Yi, Chun-Hao P. Huang, Dimitrios Tzionas, Muhammed Kocabas, Mohamed Hassan, Siyu Tang, Justus Thies, Michael J. Black

Introduction

Human behavior and the interaction of humans with their environment are fundamentally about the 3D world. Hence, 3D reconstruction of both the human and scene can facilitate behavior analysis. Where and how humans interact with a scene can be used to predict future motions and interactions for human-centered AI and robots, or to synthesize these for AR/VR and other computer-graphics applications.

Tremendous progress has been made in reconstructing 3D human bodies and 3D scenes from monocular images or videos, typically in isolation from each other. In real life, though, humans always interact with scenes. Consequently, humans (partially) occlude the scene, and the scene (partially) occludes humans. Strong human-scene occlusion can cause problems for both scene and human reconstruction.

In contrast, recent work on human-scene interaction (HSI), estimates humans and scenes together . PROX demonstrates how HSI can be used to constrain 3D human pose estimation, but it requires a 3D scan of the full scene to be known a priori. This is often unrealistic and cumbersome, as it requires one to conduct offline 3D reconstruction by walking around the scene with a depth sensor to observe it from many view points.

What we need, instead, is a method that estimates the scene and humans from images of a single color camera. This is challenging, as the lack of depth information causes the scale and placement of objects to be inconsistent w.r.t. the humans interacting with them. This leads to physically implausible results, like humans penetrating objects, or lacking physical contact when walking, sitting, or lying down, causing bodies to “hover” in the air (see Fig. 2). Methods that reconstruct 3D humans from single views leverage statistical body models as priors on the body shape and pose. However, the same tools do not exist for the collective space of 3D scene layouts. This is due to the enormous space of possible object arrangements in indoor 3D scenes, the large number of different object classes, and the huge inter-class (e.g., chairs and desks) and intra-class (e.g., desk chair and club chair) shape variability.

To address the above issues, we present MOVER, which stands for “human Motion driven Object placement for Visual Environment Reconstruction”. MOVER leverages information across several HSI frames to estimate both a plausible 3D scene and a moving human that interacts with the scene. Figure Human-Aware Object Placement for Visual Environment Reconstruction provides a high-level overview. MOVER takes as input: (1) a set of color frames from a static monocular camera, (2) a 3D human mesh inferred for each frame , and (3) a 3D shape inferred for each object detected in the scene . As output, MOVER produces a refined 3D scene, comprised of repositioned input objects, so that it is consistent with the estimated 3D human; i.e., it satisfies the expected contacts on the body , while preventing interpenetration. MOVER uses a novel optimization scheme, that jointly optimizes over camera pose, ground-plane pose, and the size and position of 3D objects, while being constrained by various HSI constraints.

MOVER takes three types of HSI constraints into account: (1) humans who move in a scene are occluded or occlude objects, thus, defining the depth ordering of the objects (c.f. ), (2) humans move in free space that is not occupied by objects and do not interpenetrate objects, (3) contact between humans and objects means that the contacting parts of their surfaces occupy the same place in space. Thus, we leverage both explicit (i.e., contact) and implicit (i.e., free space, no penetrations) HSI cues. MOVER is able to use these because it employs detailed meshes for both the scene and the moving human. In contrast, the few attempts that have been made in this direction use oversimplified shapes , i.e., 3D bounding boxes for objects and skeletons for humans, work only for static humans that contact a single object , or do not integrate information across several interaction frames .

Comparisons against the state of the art on the PROX and PiGraphs datasets show, that MOVER estimates more accurate and realistic 3D scene layouts that satisfy the expected contacts, while minimizing penetrations, w.r.t. the moving humans. Interestingly, we find that MOVER’s estimated 3D scene can be used to refine the human poses, with a PROX-like method . While estimating 3D scenes and humans from a single camera is challenging, our results suggest that they are synergistic tasks that benefit each other.

Related Work

Single-view 3D Human Pose in “Isolation”: Estimating human pose from an image is a long standing problem . Typically, this is cast as estimating 2D or 3D joints of body or whole-body skeletons . Recently, there has been a significant shift in research interest towards reconstructing the 3D human body surface which, in contrast to the joints, interacts directly with objects and can be observed by commodity cameras. To this end, many non-parametric methods have been developed, that estimate either depth maps , 3D voxels , 3D distance fields , or free-form 3D meshes . While these methods can reconstruct bodies with details like hair and clothing, they do not encode body parts or provide correspondence across people and poses. In contrast, parametric statistical 3D shape models of the body or body, face, and hands provide this information and allow re-posing. Since parametric models represent the shape and pose in a low-dimensional space, they are a powerful tool to estimate the surface from incomplete data (e.g., 2D images with occlusions) through optimization , regression , or hybrid approaches .

However, all the above methods reason about the human in “isolation”, i.e. without taking the surrounding objects and scenes into account. Thus, they struggle to reconstruct details like contact with objects, and often fail due to occlusions (e.g., bodies standing behind furniture). PARE addresses this by leveraging localized features and attention, gaining robustness to occlusions. We initialize our approach with to refine the 3D scene layout.

Single-view 3D Scene in Isolation: 3D reconstruction from single views has been addressed in several recent works that leverage learned geometric priors for specific object classes or entire scenes. Shapes from single views are reconstructed using generative models for specific object classes . The methods differ in the underlying representation, which ranges from volumetric representations like occupancy fields and implicit surface functions , to explicit surface representations like triangular meshes . To reconstruct scenes, single objects can be detected and reconstructed in isolation. Mesh-RCNN detects the objects in an RGB image, and predicts geometry for each object individually. Instead of a generative mesh model, Izadinia et al. and Kuo et al. retrieve individual CAD models for the detected objects in the scene. Bansal et al. infer a normal map from the input image that is used to align a retrieved CAD model. Instead of predicting normal maps from the input image, several methods estimate depth maps , or pixel-aligned implicit functions for objects and scenes . Joint estimation of the room layout and objects with scene context information is done for isolated 3D scenes without humans in them .

Note that there are also methods that predict room layouts with 3D bounding boxes . In contrast, we reconstruct the detailed object geometry to leverage explicit contact point constraints based on the human scene interactions, while optimizing for the scene layout.

3D Human-Scene Interaction: Humans inhabit 3D scenes. Several methods model this and learn to populate a 3D scene . In contrast, our work reasons about the human and its interaction with the 3D scene from RGB observations. There are several methods that explore different kinds of HSI; these can be divided into three categories by the interaction granularity between the human and scene: (1) Hand-Object . (2) Body-Object . (3) Body-Scene .

Our proposed method focuses on reconstructing 3D scenes composed of objects and structural elements like the floor plane, using accumulated human scene interactions (body-objects and body-scene). Table 1, overviews the most related work that operates on single-view RGB images/videos. PHOSA infers humans and objects together when they are in contact. They do not consider the fact that humans do not need to contact an object to constrain its location; their movement through free space constrains object placement. Zanfir et al. only consider feet-ground contact. iMapper maps RGB videos to dynamic “interaction snapshots”, by learning “scenelets” from PiGraphs data and fitting them to videos. However, the estimated scene is not aligned with the 2D image, and consists of pre-defined CAD templates with fixed shape and size. Holistic++ takes learned 3D HOI (Human Object Interaction) into account to jointly reason about the arrangement of bodies and objects. Both and do not model geometrically detailed human-scene interaction, due to their simplified representation of the scene and bodies. Weng et al. jointly optimize the reconstructed mesh-based 3D scene and bodies, which are initialized from and . The approach only considers interpenetration between objects and the human, and does not model the explicit human-scene contact. Additionally, both do not model the coherence of human-scene interactions across frames from monocular video. In contrast to the prior work, our contribution lies in incorporating multiple human-scene interactions collectively, such that we can reconstruct a more accurate and consistent scene, with physically plausible human-scene interactions.

Method

MOVER is an optimization-based approach that reconstructs a physically plausible 3D scene that is consistent with predicted human-scene interactions over time (see Fig. 3). Specifically, our method takes an RGB video or multiple images {It}t=1T\left\{I_{t}\right\}_{t=1}^{T} as input and reconstructs the human bodies at each time step tt as well as the numerous static scene objects, all of which reside in a common 3D space and are supported by a ground plane. In our experiments, we consider indoor scenes containing large objects with which humans frequently interact, i.e., chairs, beds, sofas, and tables.

We initialize our approach using separate estimates for the 3D human poses , the 3D scene , and the ground plane. Using the estimated body poses, we predict contact vertices C\mathcal{C} for all bodies using POSA , which predicts likely contact vertices on the body conditioned on pose. We further divide these vertices into foot contacts Cfeet\mathcal{C}^{\text{feet}} and other body part contacts Cbody\mathcal{C}^{\text{body}}. The explicit foot contact points Cfeet\mathcal{C}^{\text{feet}} are used as constraints to refine the camera orientation and ground plane prediction. Based on this initialization, we optimize the alignment of the objects by minimizing an objective function based on multiple human-scene interactions (HSIs) across the entire input data.

Our method leverages multiple HSIs to refine the 3D scene. Recall that these HSIs provide the following constraints: (1) humans that move in a scene are occluded or occlude objects, thus, defining the depth ordering of the objects (depth order constraint), (2) humans move through free space and do not interpenetrate objects (collision constraint), (3) when humans and objects are in contact, the contact surfaces occupy the same place in space (contact constraint). Using these constraints, our objective is:

The occlusion between humans and objects can provide clues about the object’s depth. We assume the human’s depth is accurate. If a human occludes an object, then the far side of the person sets a limit on how close the object can be. Alternatively, if the object occludes the person, then the visible side of the person sets a maximum distance for the object. This is summarized in Fig. 4. In this way, human-object occlusion provides constraints on scene layout even when there is no human-object contact.

Directly applying the ordinal depth loss proposed by Jiang et al. for each image is inefficient, as the required memory increases with the number of images. In contrast, we accumulate all single depth ordering maps into one far depth range map D^far\hat{D}_{\text{far}} and one near depth range map D^near\hat{D}_{\text{near}} as:

where the pixel pp is in the overlapping region between the human bodies and the objects. Using these accumulated depth range maps, we constrain the depth Di(q)D_{i}(q) of a projected pixel qq from object ii to lie in the corresponding range:

where SiliSil_{i} is the rendered silhouette of the object ii, MiM_{i} is its 2D segmentation mask, and Di(q)D_{i}(q) is the depth of the object ii at the pixel qq. See more details in Sup. Mat.

To penalize all interpenetrating vertices of objects and bodies in the scene, we use the signed distance field (SDF) of all reconstructed bodies. Specifically, we calculate a signed distance field volume VjV_{j} for each body jj in a shared 3D world space, and accumulate them into a global SDF volume as V^=min⁡(V1,...,Vj,...)\hat{V}=\min(V_{1},...,V_{j},...). The SDF V^\hat{V} is stored in a volumetric grid of size 2563256^{3}, which spans a padded bounding box of all bodies. For a vertex uiu_{i} of an object OiO_{i}, we compute the voxel coordinates f(ui)=(p(ui),q(ui),k(ui))f(u_{i})=(p(u_{i}),q(u_{i}),k(u_{i})) in the global SDF volume, and retrieve the corresponding SDF value V^f(ui)\hat{V}_{f(u_{i})}.

Based on the SDF values of all vertices of all NN objects, we resolve the scene-body interpenetration by penalizing vertices with a negative SDF value:

When humans and objects are in contact, the contact surfaces occupy the same place in space. We propose a contact constraint to minimize the distance between the contacted body parts and its assigned corresponding contacted object. PHOSA proposes a loss in which they assign a whole body to only one object, whereas humans sometimes interact with multiple objects; e.g., a person sits on a chair and puts their hand on a table. In contrast, we directly assign the contacted body vertices Cibody\mathcal{C}^{body}_{i} of each body to different objects, based on the overlap between the 2D projection of the vertices and the detected object masks, and based on the 3D distances between them. We consider the vertices of sofa and chair backs and seat bottoms as contactable regions, see more details in Sup. Mat.

We minimize the distance between the contacted bodies and the contacted object parts:

2 Optimization

We optimize Eq. (1) for a specific scene w.r.t. the parameters si\mathbf{s}_{i} (scale), θi\mathcal{\theta}_{i} (rotation), ti\mathbf{t}_{i} (translation) of the objects {i=1...N}\{{i=1...N}\}, with the Adam optimizer . In the following, we detail the initialization of the 3D scene and the HPS.

Initialization of the ground and camera.

As shown in the third column of Fig. 3, the estimated ground plane and camera orientation from are inconsistent with the reconstructed bodies (e.g., people float in the air). Previous methods either fix the camera orientation and only optimize the ground plane and humans , or estimate them independently per image , which generates inconsistent camera orientations and ground planes throughout a video. However, the camera orientation and ground plane are essential for producing plausible HSIs. Thus, we jointly estimate the ground, camera and multiple humans together, by applying:

where R{R} is the camera rotation matrix calculated from pitch, and roll, and ρ\rho denotes a robust Geman-McClure error function for down-weighting outliers and σ1=0.1\sigma_{1}=0.1.

Initial Estimate of 3D Bodies.

To obtain an initial body shape and pose estimate for the input images {It}t=1T\left\{I_{t}\right\}_{t=1}^{T}, we use OpenPose and SMPLify-X . Specifically, we use a perspective camera and estimate the pose parameters θt\theta_{t} of SMPL-X for each frame with shared body shape parameters β\beta. SMPLify-X requires a good initialization and, for this, we use PARE because it is robust to occlusion and our scenes involve significant occlusion. PARE outputs SMPL, which we convert to SMPL-X , and use the resulting 3D joints to initialize SMPLify-X, see more details in Sup. Mat.

We then optimize all SMPL-X parameters to minimize an objective function EBodyE_{\text{Body}} of multiple terms, as described in SMPLify-X (see ESMPLify-XE_{\text{SMPLify-X}}) :

To reduce jitter, we add a constant-velocity motion smoothing term on 3D joints JJ and their 2D projections JProjJ^{\text{Proj}}:

Experiments

To evaluate the influence of accumulated HSIs on the optimized 3D scene layout, we use two different datasets, PiGraphs and PROX (see Sup. Mat.). In comparison to and , we achieve state-of-the-art 3D scene layout reconstruction, both quantitatively (see Sec. 4.1) and qualitatively (see Sec. 4.3). On the PROX quantitative dataset, we find that our 3D scene reconstructions lead to more accurate human shape and pose estimations than our baselines. In Sec. 4.2, we analyze the different energy terms and how they contribute to our final results.

We perform several experiments to investigate the effectiveness of our proposed method in three parts: 3D scene reconstruction, human-scene interaction (HSI) reconstruction, and human pose and shape (HPS) estimation.

Following , we compute the 3D IoU and 2D IoU of object bounding boxes to evaluate the 3D scene reconstruction and the consistency between the 3D world and 2D image on PROX and PiGraphs. However, the 3D IoU is coarse and does not capture the error in an object’s orientation, which is quite important for physically plausible HSI, e.g., a human can not sit on an armed chair with the wrong orientation. Therefore, we introduce the point2surface distance (p2sp2s) to measure the distance from a cropped object mesh to the estimated 3D object mesh. It enables 3D scene reconstruction evaluation with more geometric details including orientation and shape. Given 2D labeled or detected bounding boxes and masks, our method improves the input significantly, and outperforms on all scene-reconstruction metrics and different datasets, as shown in Tab. 2 and Tab. 3.

Furthermore, we evaluate the error of the camera orientation and ground plane penetration using the estimated foot contact vertices (see Tab. 4). We find that jointly optimizing the camera orientation and the ground plane using foot contact significantly improves accuracy compared to the initial estimate from .

Human-scene Interaction Reconstruction.

To evaluate the physical plausibility of the estimated scene, we compute the metrics used in prior work . Specifically, for each reconstructed body and 3D scene, we calculate (1) the non-collision score to measure the ratio of body mesh vertices that do not penetrate the estimated 3D scene, divided by the number of all body mesh vertices, and (2) the contact score to denote whether the body is in contact with the 3D scene or not. The contact score is 11, if at least one vertex of a body interpenetrates the 3D scene. We report the mean non-collision score and mean contact scores among all videos and all bodies. In Tab. 2, MOVER achieves the best balance between non-collision and contact.

The estimated scenes with detected 2D boxes and masks provide lower HSI scores than with 2D GT. This is mainly because of the mis-detected objects from . Since the reconstructed scenes of do not support human-scene contact well, e.g., a sitting body often floats, due to the lack of explicit human-scene contact modeling, it has a better non-collision score but a lower contact score.

Human Pose and Shape (HPS) Estimation.

2 Ablation Study

To analyze the contribution of the accumulated HSIs and the different constraints, we conducted multiple ablation studies; see Tab. 2. All three proposed HSI constraints (depth order, collision, and contact) help to improve 3D scene reconstruction in different ways. The contact constraint produces the highest human-scene contact scores, but decreases the non-collision score. The collision and depth order both contribute to the non-collision score. However, using only the depth order constraint achieves a slightly better 3D scene than our full model, but leads to worse human-scene contact scores. By applying all constraints, our method can generate a 3D scene that supports more physically plausible HSIs.

3 Qualitative Analysis

In Fig. 5, we show reconstructed 3D scenes and humans along with frames from the RGB videos, to demonstrate the effectiveness and generality of our approach on different datasets (PROX and PiGraphs ). MOVER recovers better 3D scenes and HPS compared to Total3D and HolisticMesh . See Sup. Mat. for more examples.

Discussion

Based on single-view inputs, our proposed method optimizes the 3D pose of objects in a scene. While we assume a static camera, future work should explore moving cameras and structure-from-motion techniques to better estimate the 3D scene. We also assume that the scene is static. However, humans move objects when interacting with the world, resulting in a dynamic scene layout. We believe that our proposed constraints based on HSIs will be beneficial for future work on reconstructing dynamic scenes. Besides optimizing the 3D scene layout, we do not change the initial shape estimate of an object. A more flexible and adjustable geometric object representation, e.g., an implicit representation, would be beneficial. One could then optimize over the space of object shapes in addition to object poses. While here we focus on large objects like furniture, hand-held objects are also important and are likely subject to different constraints. During HSI, bodies are often occluded, causing errors in estimated 3D human pose. These estimates could be improved by incorporating strong human motion priors .

Conclusion

We have introduced MOVER, which reconstructs a 3D scene by exploiting 3D humans interacting with it. We have demonstrated that accumulated HSIs, computed from a monocular video, can be leveraged to improve the 3D reconstruction of a scene. The reconstructed scene, in turn, can be used to improve 3D human pose estimation. In contrast to the state of the art, MOVER can reconstruct a consistent, physically plausible 3D scene layout.

Acknowledgments. We thank Yixin Chen, Yuliang Xiu for the insightful discussions, Yao Feng, Partha Ghosh and Maria Paola Forte for proof-reading, and Benjamin Pellkofer for IT support. This work was supported by the German Federal Ministry of Education and Research (BMBF): Tübingen AI Center, FKZ: 01IS18039B.

Disclosure. https://files.is.tue.mpg.de/black/CoI_CVPR_2022.txt

References

Appendix A Dataset

PiGraphs consists of 6060 RGB-D videos of 3030 scenes. The dataset is recorded with a Microsoft Kinect One, and is designed to capture human and object arrangements in different kinds of interaction. Each video recording is about 22-minute long with 55 fps. It contains labeled 3D bounding boxes of objects in the scene and human poses represented as 3D skeletons. We use this dataset to evaluate the scene reconstruction and compare with . Note that the provided human poses are noisy and not suitable for an evaluation of 3D human shape and pose estimation.

PROX Qualitative.

PROX qualitative contains 6161 RGB-D videos at 3030 fps of human motion/interaction in 1212 scanned static 3D scenes. The data has been recorded using the Microsoft Kinect One and StructureIO sensor. To enable 3D scene reconstruction evaluation on this dataset, we segment and label each object with its 3D bounding box. Since there are two scenes (i.e., “BasementSittingBooth” and “N0SittingBooth”) containing an inseparable object, we evaluate all methods on the remaining 1010 scenes (see Fig. 6) using the corresponding 5151 videos as input.

PROX Quantitative.

PROX quantitative captures a sequence of human-scene interaction RGB-D frames within a synchronized Vicon marker-based motion capturing system. In total, the dataset contains 178178 frames and provides groundtruth body meshes, which accounts for human pose and shape (HPS) evaluation. For fair evaluation on HPS, we input all images into HolisticMesh and ours to get a refined scene and use a refined scene to get refined bodies. In addition, we also label this scene for 3D scene reconstruction evaluation, see Fig. 6.

Appendix B Implementation Details

The scale term prevents object scales ss deviating far from the initial estimates sinits^{init} from Total3D :

Initial Estimate of 3D Bodies.

We use PARE to initialize the body poses and shape (shape β\beta, pose θ\theta, scale ss). Since our approach uses the SMPL-X model, we apply to convert the SMPL parameter estimated from PARE. In addition, we use perspective projection with the calibrated camera intrinsic parameters, KK provided by the datasets (PiGraph and PROX). To convert the estimations of PARE using a weak perspective camera model, we compute the corresponding translation tbodyt^{body} by:

where K0K_{0} denotes the camera intrinsic parameters of the weak perspective camera model with focal length 5000. Then we extract the resulting 3D joints to initialize EbodyE_{body}.

Contact Regions of Objects.

We automatically calculate the contact regions of objects based on the normal of the vertices. Specifically, the vertices, whose normals are along y-axis, are the bottom or top part of the objects, while the vertices with along z-axis normal are the back part of the objects. We term that sofas and chairs have two contact regions, i.e., bottom and back parts, while beds and tables only have the top part as the contact region, shown in Fig. 7.

Optimization.

We use the Adam optimizer to optimize the final energy term with a step size of 0.0020.002 and 30003000 iterations. We set λ1,λ2,λ3\lambda_{1},\lambda_{2},\lambda_{3} as 1000,0.3,10001000,0.3,1000 respectively, for 2D bounding box term, occlusion-aware term and scale term. The weights of our proposed depth order constraint, collision constraint, and contact constraint are set to λ4=8,λ5=1000\lambda_{4}=8,\lambda_{5}=1000, and λ6=1e5\lambda_{6}=1e5, respectively.

Our method takes around 3030 minutes for 30003000 iterations to optimize a 3D scene with accumulated HSIs constraints. In comparison, HolisticMesh which jointly optimizes human and a 3D scene for one single image, directly trains the parameters of the network in Total3D to regress the 3D scene, which is time-consuming and costs around 40 minutes. For the human optimization, it runs twice in 5 minutes, i.e., the first pass is a HPS initialization used to refine the scenes, and the second pass is done using the refined scenes. In total, HolisticMesh takes 45 minutes for one single image. Our method takes almost the same time for a scene (around 10 objects) regardless how many frames in the input video. The number of frames in a video only influences the time of calculating the depth map, the SDF volume and the contact information of each body. However, this can be done once and is easily processed in parallel before the optimization. In contrast, HolisticMesh processes a video sequentially, i.e., one frame after another. Therefore, the optimization time increases w.r.t. the number of frames in a video.

Appendix C Sensitivity Analysis.

Our approach uses HSIs observed in a video. A longer video potentially has more HSIs, which results in more constraints for our objective function. In Tab. 6, we analyze how different video lengths influence scene reconstruction, by reporting the 3D intersection-over-union (IoU) metric. Specifically, we use 1010 sequences of the PROX qualitative dataset (one sequence per scene) and randomly sample 1010 segments of 1010s, 2020s, 3030s length from each sequence. We observe that longer sequences result in better performance, i.e., higher IoU and lower standard deviation. We observe that the performance of 3D scene reconstruction depends on the number of HSIs and not the video length, i.e., a short video with many HSIs results in a better reconstruction than a long video with a few unique HSIs.

We also do a sensitivity study w.r.t. noise in the initialization. In Tab. 7, we add uniform noise on the initial scale, translation and orientation of objects predicted by Total3D , and report the 3D IoU. MOVER is robust to noisy orientation and translation estimates from Total3D , but sensitive to the scale variation. This is because we currently regularize the optimization to the initial scale relatively strongly; i.e., we cannot deviate much from a noisy estimate to “correct” it. Relaxing Lscale\mathcal{L}_{\text{scale}} easily resolves this.

Appendix D More Evaluation Results on PROX Quantitative Dataset.

We also evaluate 3D scene reconstruction and human-scene interaction on PROX quantitative, as shown in Tab. 8. Our method improves our input baseline significantly and outperforms the previous method with a big margin in both 3D scene reconstruction metrics and human-scene interaction metrics.

Appendix E Failure Cases

In this section, we discuss and show the failure cases of our method. Besides optimizing the 3D scene layout, we do not change the initial shape estimate of an object. Thus, wrong estimated geometry shape can still violate human’s interaction, as shown in (A) in Fig. 8. A more flexible and adjustable geometry representation, e.g., an implicit representation, would be needed. Human motion reconstruction struggles with severe occlusions in the input, that leads to wrong body poses as well as poor estimations of HSIs, and, thus, influences our 3D scene layout prediction, see (B) in Fig. 8. While not the scope of our work, the robustness and accuracy of human motion estimation can be improved by incorporating human motion priors or learning-based probabilistic human pose and estimation network. Severe occlusion can also cause missing objects in the scene, like the chair in Fig. 8(C).

In our pipeline, we currently consider the contact between detected objects and bodies. As a potential future extension of our method, one can also leverage the information from 2D learning-based human-object interaction (HOI) detection network , by using contacted bodies to discover missing objects; or learn a model that jointly regress human-object interaction and their geometry shape.

Appendix F Additional Qualitative Results

In Fig. 9 and Fig. 10, we present additional qualitative results on PROX qualitative and PiGraphs respectively. As can be seen, our method performs well on a variety of different scenes and predicts a physically plausible scene layout. We also refer to the suppl. video for results.

Appendix G Discussion of Potential Misuse

Our approach is not intended for any surveillance application. Our goal is to understand how humans interact and move in scenes from videos (e.g., from TV sitcoms, movies, etc.), to this end both the scene geometry and the human pose and shape need to be reconstructed. Our method could be misused in potential surveillance applications that curtail human rights and civil liberties, but we will restrict the usage of our method in a legal way.