Holistic 3D Scene Understanding from a Single Image with Implicit Representation
Cheng Zhang, Zhaopeng Cui, Yinda Zhang, Bing Zeng, Marc Pollefeys, Shuaicheng Liu
Introduction
3D indoor scene understanding is a long-lasting computer vision problem and has tremendous impact on several applications, e.g., robotics, virtual reality. Given a single color image, the goal is to reconstruct the room layout as well as each individual object and estimate its semantic type in the 3D space. Over decades, there are plenty of works consistently improving the performance of such a task over two focal points of the competition. One is the 3D shape representation preserving fine-grained geometry details, evolving from the 3D bounding box, 3D volume, point cloud, to the recent triangulation mesh. The other is the joint inference of multiple objects and layout in the scene leveraging contextual information, such as co-occurring or relative locations among objects of multiple categories. However, the cluttered scene is a double-blade sword, which unfortunately increases the complexity of 3D scene understanding by introducing large variations in object poses and scales, and heavy occlusion. Therefore, the overall performance is still far from satisfactory.
In this work, we propose a deep learning system for holistic 3D scene understanding, which predicts and refines object shapes, object poses, and scene layout jointly with deep implicit representation. At first, similar to previous methods, we exploit standard Convolutional Neural Networks (CNN) to learn an initial estimation of 3D object poses, scene layout as well as 3D shapes. Different from previous methods using explicit 3D representation like volume or mesh, we utilize the local structured implicit representation of shapes motivated by . Instead of taking depth images as input like , we design a new local implicit shape embedding network to learn the latent shape code directly from images, which can be further decoded to generate the implicit function for 3D shapes. Due to the power of implicit representation, the 3D shape of each object can be reconstructed with higher accuracy and finer surface details compared to other representations.
Then, we propose a novel graph-based scene context network to gather information from local objects, i.e., bottom-up features extracted from the initial predictions, and learn to refine the initial 3D pose and scene layout via scene context information with the implicit representation. Being one of the core topics studied in scene understanding, the context has been achieved in the era of deep learning mainly from two aspects - the model architecture and the loss function. From the perspective of model design, we exploit the graph-based convolutional neural network (GCN) to learn context since it has shown competitive performance to learn context . With the deep implicit representation, the learned local shape latent vectors are naturally a compact and informative feature measuring of the object geometries, which results in more effective context models compared to features extracted from other representations such as mesh.
Not only the architecture, but the deep implicit representation also benefits the context learning on the loss function. One of the most basic contextual information yet still missing in many previous works - objects should not intersect with each other, could be easily applied as supervision by penalizing the existence of 3D locations with negative predicted SDF in more than one objectsThe object interior is with negative SDF, and thus no location should be inside of two objects.. We define this constraint as a novel physical violation loss and find it particularly helpful in preventing intersecting objects and producing reasonable object layouts.
Overall, our contributions are mainly in four aspects. First, we design a two-stage single image-based holistic 3D scene understanding system that could predict object shapes, object poses, and scene layout with deep implicit representation, then optimize the later two. Second, a new image-based local implicit shape embedding network is proposed to extract latent shape information which leads to superior geometry accuracy. Third, we propose a novel GCN-based scene context network to refine the object arrangement which well exploits the latent and implicit features from the initial estimation. Last but not least, we design a physical violation loss, thanks to the implicit representation, to effectively prevent the object intersection. Extensive experiments show that our model achieves the state-of-the-art performance on the standard benchmark.
Related works
Single Image Scene Reconstruction. As a highly ill-posed problem, single image scene reconstruction sets a high bar for learning-based algorithms, especially in a cluttered scene with heavy occlusion. The problem can be divided into layout estimation, object detection and pose estimation, and 3D object reconstruction. A simple version of the first problem is to simplify the room layout as a bounding box . To detect objects and estimate poses in 3D space, Recent works try to infer 3D bounding boxes from 2D detection by exploiting relationships among objects with a graph or physical simulation. At the same time, other works further extend the idea to align a CAD model with similar style to each detected object. Still, the results are limited by the size of the CAD model database which results in an inaccurate representation of the scene. To tackle the above limitations of previous works, Total3D is proposed as an end-to-end solution to jointly estimate the layout box and object poses while reconstructing each object from the detection and utilizing the reconstruction to supervise the pose estimation learning. However, they only exploit relationships among objects with features based on appearance and 2D geometry.
Shape Representation. In the field of computer graphics, traditional shape representation methods include mesh, voxel, and point cloud. Some of the learning-based works try to encode the shape prior into a feature vector but stick to the traditional representations by decoding the vector into mesh , voxel or point cloud . Others try to learn structured representations which decompose the shape into simple shapes . Recently, implicit surface function has been widely used as a new representation method to overcome the disadvantages of traditional methods (i.e. unfriendly data structure to neural network of mesh and point cloud, low resolution and large memory consumption of voxel). Most recent works try to combine the structured and implicit representation which provides a physically meaningful feature vector while introducing significant improvement on the details of the decoded shape.
Graph Convolutional Networks. Proposed by , graph neural networks or GCNs have been widely used to learn from graph-structured data. Inspired by convolutional neural networks, convolutional operation has been introduced to graph either on spectral domain or non-spectral domain which performs convolution with a message passing neural network to gather information from the neighboring nodes. Attention mechanism has also been introduced to GCN and has been proved to be efficient on tasks like node classification , scene graph generation and feature matching . Recently, GCN has been even used on super-resolution which is usually the territory of CNN. In the 3D world which interests us most, GCN has been used on classification and segmentation on point cloud, which is usually an enemy representation to traditional neural networks. The most related application scenario of GCN with us is 3D object detection on points cloud. Recent work shows the ability of GCN to predict relationship or 3D object detections from point cloud data.
Our method
As shown in Figure 2, the proposed system consists of two stages, i.e., the initial estimation stage, and the refinement stage. In the initial estimation stage, similar to , a 2D detector is first adopted to extract the 2D bounding box from the input image, followed by an Object Detection Network (ODN) to recover the object poses as 3D bounding boxes and a new Local Implicit Embedding Network (LIEN) to extract the implicit local shape information from the image directly, which can further be decoded to infer 3D geometry. The input image is also fed into a Layout Estimation Network (LEN) to produce a 3D layout bounding box and relative camera pose. In the refinement stage, a novel Scene Graph Convolutional Network (SGCN) is designed to refine the initial predictions via the scene context information. As 2D detector, LEN, ODN has the standard architecture similar to prior works , in this section, we will describe the details of the novel SGCN and LIEN in detail. Please refer to our supplementary materials for the details of our 2D detector, LEN, ODN.
As shown in Figure 2, motivated by Graph R-CNN , we model the whole 3D scene as a graph , in which the nodes represent the objects, the scene layout, and their relationships. The graph is constructed starting from a complete graph with undirected edges between all objects and layout nodes, which allows information to flow among objects and the scene layout. Then, we add relation nodes to each pair of neighboring object/layout nodes. Considering the nature of directional relation , we add two relation nodes between each pair of neighbors in different directions.
It is well known that the input features are the key to an effective GCN . For different types of nodes, we design features carefully from different sources as follows. For each node, features from different sources are flattened and concatenated into a vector, then embedded into a node representation vector with the same length using MLP.
Layout Node. We use the feature from the image encoder of LEN, which encodes the appearance of layout, and the parameterized output of layout bounding box and camera pose from LEN, as layout node features. We also concatenate the camera intrinsic parameters normalized by the image height into the feature to add camera priors.
Object Node. We collect the appearance-relationship feature from ODN, and the parameterized output of object bounding box from ODN, along with the element centers in the world coordinate and analytic code from LIEN (which we will further describe in the next section). We also use the one-hot category label from the 2D detector to introduce semantic information to SGCN.
Relationship Node. For nodes connecting two different objects, the geometry feature of 2D object bounding boxes and the box corner coordinates of both connected objects normalized by the image height and width are used as features. The coordinates are flattened and concatenated in the order of source-destination, which differentiate the relationships of different directions. For nodes connecting objects and layouts, since the relationship is presumably different from object-object relationship, we initialize the representations with constant values, leaving the job of inferring reasonable relationship representation to SGCN.
Since the graph is modeled with different types of nodes, which makes a difference in the information needed from different sources to destinations, we define independent message passing weights for each of the source-destination types. We denote the linear transformation and the adjacent matrix from source node to destination node with type and as and , in which node types can be source object (or layout) , destination object (or layout) , and relationships . Thus, the representation of object and layout nodes can be updated as
and the relationship node representations can be updated as
After four steps of message passing, independent MLPs are used to decode object node representations into residuals for corresponding object bounding box parameters , and layout node representation into residuals for initial layout box and camera pose . Please refer to our supplementary or for the details of the definition. The shape codes can be also refined in the scene graph, while we find that it doesn’t improve empirically as much as for the layout and object poses in our pipeline because our local implicit embedding network, which will be introduced in the following, is powerful enough to learn accurate shapes.
2 Local Implicit Embedding Network
With a graph constructed for each scene, we naturally ask what features help SGCN effectively capture contextual information among objects. Intuitively, we expect features that well describe 3D object geometry and their relationship in 3D space. Motivated by Genova et al., we propose to utilize the local deep implicit representation as the features embedding object shapes due to its superior performance for single object reconstruction. In their model, the function is a combination of 32 3D elements (16 with symmetry constraints), with each element described with 10 Gaussian function parameters analytic code and 32-dim latent variables (latent code). The Gaussian parameters describe the scale constant, center point, radii, and Euler angle of every Gaussian function, which contains structured information of the 3D geometry. We use analytic code as a feature for object nodes in SGCN, which should provide information on the local object structure. Furthermore, since the centers of the Gaussian functions presumably correspond to centers of different parts of an object, we also transform them from the object coordinate system to the world coordinate system as a feature for every object node in SGCN. The transformation provides global information about the scene, which makes SGCN easier to infer relationships between objects. The above two features make up the implicit features of LIEN.
As LDIF is designed for 3D object reconstruction from one or multiple depth images, we design a new image-based Local Implicit Embedding Network (LIEN) to learn the 3D latent shape representation from the image which is obviously a more challenging problem. Our LIEN consists of a Resnet-18 as image encoder, along with a three-layer MLP to get the analytic and latent code. Additionally, in order to learn the latent features effectively, we concatenate the category code with the image feature from the encoder to introduce shape priors to the LIEN, which improves the performance greatly. Please refer to our supplementary material for the detailed architecture of the proposed LIEN.
3 Loss Function
Losses for Initialization Modules. When training LIEN along with LDIF decoder individually, we follow to use the shape element center loss with weight and point sample loss,
where and evaluates losses for near-surface samples and uniformly sampled points. When training LEN and ODN, we follow to use classification and regression loss for every output parameter of LEN and ODN,
Joint Refinement with Object Physical Violation Loss. For the refinement stage, we aim to optimize the scene layout and object poses using the scene context information by minimizing the following loss function,
Experiments
In this section, we compare our method with state-of-the-art 3D scene understanding methods in various aspects and provide an ablation study to highlight the effectiveness of major components.
Datasets. We follow to use two datasets to train each module individually and jointly. We use two datasets for training and evaluation. 1) Pix3D dataset is presented as a benchmark for shape-related tasks including reconstruction, providing 9 categories of 395 furniture models and 10,069 images with precise alignment. We use the mesh fusion pipeline from Occupancy Network to get watertight meshes for LIEN training and evaluate LIEN on original meshes. 2) SUN RGB-D dataset contains 10K RGB-D indoor images captured by four different sensors and is densely annotated with 2D segmentation, semantic labels, 3D room layout, and 3D bounding boxes with object orientations. Follow Total3D , we use the train/test split from on the Pix3D dataset and the official train/test split on the SUN RGB-D dataset. The object labels are mapped from NYU-37 to Pix3D as presented by .
Metrics. We adopt the same evaluation metrics with , including average 3D Intersection over Union (IoU) for layout estimation; mean absolute error for camera pose; average precision (AP) for object detection; and chamfer distance for single-object mesh generation from single image.
Implementation. We use the outputs of the 2D detector from Total3D as the input of our model. We also adopted the same structure of ODN and LEN from Total3D. LIEN is trained with LDIF decoder on Pix3D with watertight mesh, using Adam optimizer with a batch size of 24 and learning rate decaying from 2e-4 (scaled by 0.5 if the test loss stops decreasing for 50 epochs, 400 epochs in total) and evaluated on the original non-watertight mesh. SGCN is trained on SUN RGB-D, using Adam optimizer with a batch size of 2 and learning rate decaying from 1e-4 (scaled by 0.5 every 5 epochs after epoch 18, 30 epochs in total). We follow to train each module individually then jointly. When training SGCN individually, we use without , and put it into the full model with pre-trained weights of other modules. In joint training, we adopt the observation from that object reconstruction depends on clean mesh for supervision, to fix the weights of LIEN and LDIF decoder.
2 Comparison to State-of-the-art
In this section, we compare to the state-of-the-art methods for holistic scene understand from aspects including object reconstruction, 3D object detection, layout estimation, camera pose prediction, and scene mesh reconstruction.
3D Object Reconstruction. We first compare the performance of LIEN with previous methods, including AtlasNet , TMN , and Total3D , for the accuracy of the predicted geometry on Pix3D dataset. All the methods take as input a crop of image of the object and produce 3D geometry. To make a fair comparison, the one-hot object category code is also concatenated with the appearance feature for AtlasNet and TMN . For our method, we run a marching cube algorithm on 256 resolution to reconstruct the mesh. The quantitative comparison is shown in Table 1. Our method produces the most accurate 3D shape compared to other methods, yielding the lowest mean Chamfer Distance across all categories. Qualitative results are shown in Fig. 4. AtlasNet produces results in limited topology and thus generates many undesired surfaces. MGN mitigates the issue with the capability of topology modification, which improves the results but still leaves obvious artifacts and unsmooth surface due to the limited representation capacity of the triangular mesh. In contrast, our method produces 3D shape with correct topology, smooth surface, and fine-grained details, which clearly shows the advantage of the deep implicit representation.
3D Object Detection. We then evaluate the 3D object detection performance of our model. Follow , we use mean average precision (mAP) with the threshold of 3D bounding box IoU set at 0.15 as the evaluation metric. The quantitative comparison to state-of-the-art methods is shown in Table 2. Our method performs consistently the best over all semantic categories and significantly outperforms the state-of-the-art (i.e. improving AP by 18.83%). Fig. 5 shows some qualitative comparison. Note how our method produces object layout not only more accurate but also in reasonable context compared to Total3D, e.g. objects are parallel to wall direction.
Layout Estimation. We also compare the 3D room layout estimation with Total3D and other state-of-the-arts . The quantitative evaluation is shown in Table 3 (Layout IoU). Overall, our method outperforms all the baseline methods. This indicates that the GCN is effective in measuring the relation between layout and objects and thus benefits the layout prediction.
Camera Pose Estimation. Table 3 also shows the comparison over camera pose prediction, following the evaluation protocol of Total3D. Our method achieves 5% better camera pitch and slightly worse camera roll.
Holistic Scene Reconstruction. To our best knowledge, Total3D is the only work achieving holistic scene reconstruction from a single RGB, and thus we compare to it. Since no ground truth is presented in SUN RGB-D dataset, we mainly show qualitative comparison in Fig. 5. Compared to Total3D, our model has less intersection and estimates more reasonable object layout and direction. We consider this as a benefit from a better understanding of scene context by GCN. Our proposed physical violation loss also contributes to less intersection.
3 Ablation Study
In this section, we verify the effectiveness of the proposed components for holistic scene understanding. As shown in Table 4, we disable certain components and evaluate the model for 3D layout estimation and 3D object detection, We do not evaluate the 3D object reconstruction since it is highly related to the usage of deep implicit representation, which has been already evaluated in Section 4.2.
Does GCN Matter? To show the effectiveness of GCN, we first attach the GCN to the original Total3D to improve the object and scene layout (Table 4, Total3D+GCN). For the difference between MGN of Total3D and LIEN of ours, we replace deep implicit features with the feature from image encoder of MGN and use their proposed partial Chamfer loss instead of . Both object bounding box and scene layout are improved. We also train a version of our model without the GCN (Ours-GCN), and the performance drops significantly. Both experiments show that GCN is effective in capturing scene context.
Does Deep Implicit Feature Matter? As introduced in Section 3.2, the LDIF representation provides informative node features for the GCN. Here we demonstrate the contribution from each component of the latent representation. Particularly, we remove either element centers or analytic code from the GCN node feature (Ours-element, Ours-analytic), and find both hurts the performance. This indicates that the complete latent representation is helpful in pursuing better scene understanding performance.
Does Physical Violation Loss Matter? Additionally, we evaluate the effectiveness of the physical violation loss. We train our model without it (Ours-), and also observe performance drop for both scene layout and object 3D bounding box in Table 4. We refer to supplementary material for qualitative comparison.
Evaluating on Other Metrics. We also test our method in other aspects including supporting relation, geometry accuracy, and room layout as shown in Table 5. 1) We calculate the mean distance between the predicted bottom of on-floor objects and the ground truth floor to measure the supporting relationship. As ground truth, an object is considered to be on-floor if its bottom surface is within 15cm to the floor. While GCN significantly improves the metric, slightly hurts possibly because it tends to push objects away. Further qualitative results are shown in the supplementary material. Besides, we also measure the average volume of the collision per scene between objects (Coll Vol), and our full model effectively prevent collision. 2) We follow Total3D to evaluate the alignment between scene reconstruction and ground truth depth map with global loss , and our full model performs the best. 3) We also project the predicted layout onto the image and evaluate with image based metrics . Our full model achieves the best on both corner and pixel errors. Overall, the GCN and benefit on all these aspects.
4 Generalization to other datasets
We also show qualitative results of our method tested on the 3D detection dataset ObjectNet3D and the layout estimation dataset in without fine-tuning in Fig. 6. Our method shows good generalization capability and performs reasonably well on these unseen datasets.
Conclusion
We have presented a deep learning model for holistic scene understanding by leveraging deep implicit representation. Our model not only reconstructs accurate 3D object geometry, but also learns better scene context using GCN and a novel physical violation loss, which can deliver accurate scene and object layout. Extensive experiments show that our model improves various tasks in holistic scene understanding over existing methods. A promising future direction could be exploiting object functionalities for better 3D scene understanding.
Acknowledgement: This research was supported by National Natural Science Foundation of China (NSFC) under grants No.61872067 and No.61720106004.
References
Appendix A Architecture of Our Pipeline
We show the architecture of LEN, ODN, LIEN, and SCGN in Fig. 7.
2D Detector, LEN, ODN. Following , we use Faster RCNN trained on COCO dataset and fine-tuned on SUN RGB-D as 2D detector. The 2D detection results on SUN RGB-D are filtered and matched with the ground-truth 3D object bounding box during the data preparation procedure provided by . During the initialization stage, we use LEN and ODN architecture shown in Fig. 7 similar with .
LIEN. Our proposed LIEN consists of an image encoder followed by a three-layer MLP to embed a single image into a code. When evaluating on SUN RGB-D, the category labels are mapped to the ones used by Pix3D and concatenated to the image feature following . To construct the shape elements for LDIF decoder, we follow to reshape the 1344-dim vector into a 32x42 array, which corresponds to 42-dim (10 for analytic code and 32 for latent code) codes of the 32 shape elements.
SGCN. Before being fed into each node, the features from different sources are flattened, concatenated, and embedded into a 512-dim representation using FC layers. The weights of the embedding network for layout, object, and relationship nodes are independent of each other. After updated with four steps of message passing, the representations of layout and object nodes are decoded into parameters with the networks specially designed for each of the node types. The decoding networks follow the design of LEN and ODN, and refine the parameterized initial outputs of them.
Appendix B Implementation details
Data Processing. For the training of LIEN, watertight meshes must be used to retrieve the ground-truth values of inside-outside labels. However, the models of Pix3D are not that clean with inverted surface normals and holes occasionally, which causes failure with the traditional flood fill algorithm. To get more robust results, we utilize the mesh fusion pipeline which generates watertight meshes by fusing signed distance fields from several virtual cameras and applying the marching cube algorithm on it. Although the mesh fusion pipeline makes the model thicker and introduces noise to the ground-truth sample points, we evaluate it on the original mesh to directly compare with previous works.
Hyper Parameters. When training LIEN, we use 1024 near-surface samples and 1024 uniformly samples, and set their loss weight and . For shape element center loss, we let . Following , classification and regression loss is used for parameters of both LEN and ODN, which we denote as . Other parameters of LEN and ODN are using only regression loss. For camera parameters, we set , and . For layout box parameters, we set , and . For object box parameters, we set , and . When training with cooperative loss and object physical violation loss, we set .
Scene Mesh Reconstruction. Since our LIEN is trained on Pix3D with only 9 categories like MGN of Total3D, we suffer from the same problem with them when testing on SUN RGB-D, that our LIEN can not generalize to some of the categories. For the accuracy of the scene reconstruction, we follow to only consider certain categories of objects (i.e. cabinet, bed, chair, sofa, table, door, bookshelf, desk, shelves, dresser, refrigerator, television, box, whiteboard, nightstand). As a result, the reconstructed scene mesh has fewer objects than 3D detections.
Appendix C 3D Detection on all categories
In this section, we report the average precision of 3D object detection on all categories of SUN RGB-D for a full comparison in Table 6. Our method achieves the best performance for 27 over 33 categories and a significantly better mean average precision.
Appendix D More Qualitatively Comparison with MGN on Object Mesh Reconstruction
In this section, we show more results on the object reconstruction in Fig. 10. Compared to MGN in Total3D, our method produces more accurate geometry preserving high-quality details especially on chairs, bookshelves, and those shapes with relatively more complex topology.
Appendix E More Qualitatively Comparison on 3D Detection and Scene Reconstruction
In main paper Section 4.2, we show qualitative results of the 3D object detection and scene reconstruction. Here, we show more results in Fig. 11, Fig. 12, and Fig. 13. We can observe that compared to the state-of-the-art method , our method produces significantly more accurate object pose estimation with fewer flying objects (Fig. 11e, Fig. 12a), fewer objects intersected with each other (Fig. 11a, Fig. 12d, Fig. 13e), and more accurate object orientation estimation (Fig. 11c, Fig. 12e, Fig. 13c). We also observe fewer objects intersected with the layout box (Fig. 11d, Fig. 12e, Fig. 13a).
Appendix F Qualitative Comparison of Ablation Study
In main paper Section 4.3, we quantitatively compare the improvement of our proposed . While exhibiting a small gap from the metric, we show in qualitative results (Fig. 8) that the visual difference is relatively large. Objects are more likely to intersect with each other when trained without , which disobeys physical context severely. On the contrary, training with effectively prevents these errors in the results.
We also quantitatively compare the supporting relation, in main paper Section 4.3. Here in Fig. 9, we qualitatively compare the understanding of supporting relation in the front view of object 3D detection.
Appendix G Failure Cases
We also show some failure cases in Fig. 14. We observe that although our LIEN performs well on Pix3D and is generalized to SUN RGB-D, it still cannot make plausible reconstruction for some objects in rarely seen shapes (i.e. the desks of (a) and (b), the bookshelves of (b) and (c), the bed of (d)). For object detection, our pipeline fails to correctly estimate the pose of the bed in (e), which might result from the clustered scenes. Also, in some extreme cases, heavy occlusion might cause our pipeline to fail like in (f).