3D-Aware Scene Manipulation via Inverse Graphics
Shunyu Yao, Tzu Ming Harry Hsu, Jun-Yan Zhu, Jiajun Wu, Antonio Torralba, William T. Freeman, Joshua B. Tenenbaum
Introduction
Humans are incredible at perceiving the world, but more distinguishing is our mental ability to simulate and imagine what will happen. Given a street scene as in Fig. 1, we can effortlessly detect and recognize cars and their attributes, and more interestingly, imagine how cars may move and rotate in the 3D world. Motivated by such human abilities, in this work we seek to obtain an interpretable, expressive, and disentangled scene representation for machines, and employ the learned representation for flexible, 3D-aware scene manipulation.
Deep generative models have led to remarkable breakthroughs in learning hierarchical representations of images and decoding them back to images (Goodfellow et al., 2014). However, the obtained representation is often limited to a single object, hard to interpret, and missing the complex 3D structure behind visual input. As a result, these deep generative models cannot support manipulation tasks such as moving an object around as in Fig. 1. On the other hand, computer graphics engines use a predefined, structured, and disentangled input (e.g., graphics code), often intuitive for scene manipulation. However, it is in general intractable to infer the graphics code from an input image.
In this paper, we propose 3D scene de-rendering networks (3D-SDN) to incorporate an object-based, interpretable scene representation into a deep generative model. Our method employs an encoder-decoder architecture and has three branches: one for scene semantics, one for object geometry and 3D pose, and one for the appearance of objects and the background. As shown in Fig. 2, the semantic de-renderer aims to learn the semantic segmentation of a scene. The geometric de-renderer learns to infer the object shape and 3D pose with the help of a differentiable shape renderer. The textural de-renderer learns to encode the appearance of each object and background segment. We then employ the geometric renderer and the textural renderer to recover the input scene using the above semantic, geometric, and textural information. Disentangling 3D geometry and pose from texture enables 3D-aware scene manipulation. For example in Fig. 1, to move a car closer, we can edit its position and 3D pose, but leave its texture representation untouched.
Both quantitative and qualitative results demonstrate the effectiveness of our method on two datasets: Virtual KITTI (Gaidon et al., 2016) and Cityscapes (Cordts et al., 2016). Furthermore, we create an image editing benchmark on Virtual KITTI to evaluate our editing scheme against 2D baselines. Finally, we investigate our model design by evaluating the accuracy of obtained internal representation. Please check out our code and website for more details.
Related Work
Interpretable image representation. Our work is inspired by prior work on obtaining interpretable visual representations with neural networks (Kulkarni et al., 2015; Chen et al., 2016). To achieve this goal, DC-IGN (Kulkarni et al., 2015) freezes a subset of latent codes while feeding images that move along a specific direction on the image manifold. A recurrent model (Yang et al., 2015) learns to alter disentangled latent factors for view synthesis. InfoGAN (Chen et al., 2016) proposes to disentangle an image into independent factors without supervised data. Another line of approaches are built on intrinsic image decomposition (Barrow and Tenenbaum, 1978) and have shown promising results on faces (Shu et al., 2017) and objects (Janner et al., 2017).
While prior work focuses on a single object, we aim to obtain a holistic scene understanding. Our work most resembles the paper by Wu et al. (2017a), who propose to ‘de-render’ an image with an encoder-decoder framework that uses a neural network as the encoder and a graphics engine as the decoder. However, their method cannot back-propagate the gradients from the graphics engine or generalize to a new environment, and the results were limited to simple game environments such as Minecraft. Unlike Wu et al. (2017a), both of our encoder and decoder are differentiable, making it possible to handle more complex natural images.
Deep generative models. Deep generative models (Goodfellow et al., 2014) have been used to synthesize realistic images and learn rich internal representations. Representations learned by these methods are typically hard for humans to interpret and understand, often ignoring the 3D nature of our visual world. Many recent papers have explored the problem of 3D reconstruction from a single color image, depth map, or silhouette (Choy et al., 2016; Kar et al., 2015; Tatarchenko et al., 2016; Tulsiani et al., 2017; Wu et al., 2017b, 2016b; Yan et al., 2016b; Soltani et al., 2017). Our model builds upon and extends these approaches. We infer the 3D object geometry with neural nets and re-render the shapes into 2D with a differentiable renderer. This improves the quality of the generated results and allows 3D-aware scene manipulation.
Deep image manipulation. Learning-based methods have enabled various image editing tasks, such as style transfer (Gatys et al., 2016), image-to-image translation (Isola et al., 2017; Zhu et al., 2017a; Liu et al., 2017), automatic colorization (Zhang et al., 2016), inpainting (Pathak et al., 2016), attribute editing (Yan et al., 2016a), interactive editing (Zhu et al., 2016), and denoising (Gharbi et al., 2016). Different from prior work that operates in a 2D setting, our model allows 3D-aware image manipulation. Besides, while the above methods often require a given structured representation (e.g., label map (Wang et al., 2018)) as input, our algorithm can learn an internal representation suitable for image editing by itself. Our work is also inspired by previous semi-automatic 3D editing systems (Karsch et al., 2011; Chen et al., 2013; Kholgade et al., 2014). While these systems require human annotations of object geometry and scene layout, our method is fully automatic.
Method
We propose 3D scene de-rendering networks (3D-SDN) in an encoder-decoder framework. As shown in Fig. 2, we first de-render (encode) an image into disentangled representations for semantic, textural, and geometric information. Then, a renderer (decoder) reconstructs the image from the representation.
The semantic de-renderer learns to produce the semantic segmentation (e.g. trees, sky, road) of the input image. The 3D geometric de-renderer detects and segments objects (cars and vans) from image, and infers the geometry and 3D pose for each object with a differentiable shape renderer. After inference, the geometric renderer computes an instance map, a pose map, and normal maps for objects in the scene for the textural branch. The textural de-renderer first fuses the semantic map generated by the semantic branch and the instance map generated by the geometric branch into an instance-level semantic label map, and learns to encode the color and texture of each instance (object or background semantic class) into a texture code. Finally, the textural renderer combines the instance-wise label map (from the textural de-renderer), textural codes (from the textural de-renderer), and 3D information (instance, normal, and pose maps from the geometric branch) to reconstruct the input image.
Fig. 3 shows the 3D geometric inference module for the 3D-SDN. We first segment object instances with Mask-RCNN (He et al., 2017). For each object, we infer its 3D mesh model and other attributes from its masked image patch and bounding box.
As shown in Fig. 3, given an object’s masked image and estimated bounding box, the geometric de-renderer learns to predict the mesh by first selecting a mesh from eight candidate shapes, and then applying a Free-Form Deformation (FFD) (Sederberg and Parry, 1986) with inferred grid point coordinates . It also predicts the scale, rotation, and translation of the 3D object. Below we describe the training objective for the network.
3D attribute prediction loss. The geometric de-renderer directly predicts the values of scale and rotation . For translation , it instead predicts the object’s distance to the camera and the image-plane 2D coordinates of the object’s 3D center, denoted as . Given the intrinsic camera matrix, we can calculate from and . We parametrize in the log-space (Eigen et al., 2014). As determining from the image patch of the object is under-constrained, our model predicts a normalized distance , where is the width and height of the bounding box. This reparameterization improves results as shown in later experiments (Sec. 4.2). For , we follow the prior work (Ren et al., 2015) and predict the offset relative to the estimated bounding box center . The 3D attribute prediction loss for scale, rotation, and translation can be calculated as
Reprojection consistency loss. We also use a reprojection loss to ensure the 2D rendering of the predicted shape fits its silhouette (Yan et al., 2016b; Rezende et al., 2016; Wu et al., 2016a, 2017b). Fig. 4 and Fig. 4 show an example. Note that for mesh selection and deformation, the reprojection loss is the only training signal, as we do not have a ground truth mesh model.
3D model selection via REINFORCE. We choose the mesh from a set of eight meshes to minimize the reprojection loss. As the model selection process is non-differentiable, we formulate the model selection as a reinforcement learning problem and adopt a multi-sample REINFORCE paradigm (Williams, 1992) to address the issue. The network predicts a multinomial distribution over the mesh models. We use the negative reprojection loss as the reward. We experimented with a single mesh without FFD in Fig. 4. Fig. 4 shows a significant improvement when the geometric branch learns to select from multiple candidate meshes and allows flexible deformation.
2 Semantic and Textural Inference
The semantic branch of the 3D-SDN uses a semantic segmentation model DRN (Yu et al., 2017; Zhou et al., 2017) to obtain an semantic map of the input image. The textural branch of the 3D-SDN first obtains an instance-wise semantic label map by combining the semantic map generated by the semantic branch and the instance map generated by the geometric branch, resolving any conflict in favor of the instance map (Kirillov et al., 2018). Built on recent work on multimodal image-to-image translation (Zhu et al., 2017b; Wang et al., 2018), our textural branch encodes the texture of each instance into a low dimensional latent code, so that the textural renderer can later reconstruct the appearance of the original instance from the code. By ‘instance’ we mean a background semantic class (e.g., road, sky) or a foreground object (e.g., car, van). Later, we combine the object textural code with the estimated 3D information to better reconstruct objects.
Formally speaking, given an image and its instance label map , we want to obtain a feature embedding such that can later reconstruct . We formulate the textural branch of the 3D-SDN under a conditional adversarial learning framework with three networks : a textural de-renderer , a texture renderer and a discriminator are trained jointly with the following objectives.
where denotes the -th layer of a pre-trained VGG network (Simonyan and Zisserman, 2015) with elements. Similarly, for our our discriminator , denotes the -th layer with elements. and denote the number of layers in network and . We fix the network during our training. Finally, we use a pixel-wise image reconstruction loss as:
The final training objective is formulated as a minimax game between and :
where and control the relative importance of each term.
Decoupling geometry and texture. We observe that the textural de-renderer often learns not only texture but also object poses. To further decouple these two factors, we concatenate the inferred 3D information (i.e., pose map and normal map) from the geometric branch to the texture code map and feed both of them to the textural renderer . Also, we reduce the dimension of the texture code so that the code can focus on texture as the 3D geometry and pose are already provided. These two modifications help encode textural features that are independent of the object geometry. It also resolves ambiguity in object poses: e.g., cars share similar silhouettes when facing forward or backward. Therefore, our renderer can synthesize an object under different 3D poses. (See Fig. 5 and Fig. 7 for example).
3 Implementation Details
Semantic branch. Our semantic branch adopts Dilated Residual Networks (DRN) for semantic segmentation (Yu et al., 2017; Zhou et al., 2017). We train the network for 25 epochs.
Geometric branch. We use Mask-RCNN for object proposal generation (He et al., 2017). For object meshes, we choose eight CAD models from ShapeNet (Chang et al., 2015) including cars, vans, and buses. Given an object proposal, we predict its scale, rotation, translation, FFD grid point coefficients, and an -dimensional distribution across candidate meshes with a ResNet-18 network (He et al., 2015). The translation can be recovered using the estimated offset , the normalized distance , and the ground truth focal length of the image. They are then fed to a differentiable renderer (Kato et al., 2018) to render the instance map and normal map.
We empirically set . We first train the network with using Adam (Kingma and Ba, 2015) with a learning rate of for epochs and then fine-tune the model with and REINFORCE with a learning rate of for another epochs.
Textural branch. We first train the semantic branch and the geometric branch separately and then train the textural branch using the input from the above two branches. We use the same architecture as in Wang et al. (2018). We use two discriminators of different scales and one generator. We use the VGG network (Simonyan and Zisserman, 2015) as the feature extractor for loss (Eqn. 3). We set the dimension of the texture code as . We quantize the object’s rotation into bins with one-hot encoding and fill each rendered silhouette of the object with its rotation encoding, yielding a pose map of the input image. Then we concatenate the pose map, the predicted object normal map, the texture code map , the semantic label map, and the instance boundary map together, and feed them to the neural textural renderer to reconstruct the input image. We set and , and train the textural branch for epochs on Virtual KITTI and epochs on Cityscapes.
Results
We report our results in two parts. First, we present how the 3D-SDN enables 3D-aware image editing. For quantitative comparison, we compile a Virtual KITTI image editing benchmark to contrast 3D-SDNs and baselines without 3D knowledge. Second, we analyze our design choices and evaluate the accuracy of representations obtained by different variants. The code and full results can be found at our website.
Datasets. We conduct experiments on two street scene datasets: Virtual KITTI (Gaidon et al., 2016) and Cityscapes (Cordts et al., 2016). Virtual KITTI serves as a proxy to the KITTI dataset (Geiger et al., 2012). The dataset contains five virtual worlds, each rendered under ten different conditions, leading to a sum of images. For each world, we use either the first or the last consecutive frames for training and the rest for testing. For object-wise evaluations, we use objects with more than visible pixels, a occlusion ratio, and a truncation ratio, following the ratios defined in Gaidon et al. (2016). In our experiments, we downscale Virtual KITTI images to and Cityscapes images to .
We have also built the Virtual KITTI Image Editing Benchmark, allowing us to evaluate image editing algorithms systematically. The benchmark contains pairs of images in the test set with the camera either stationary or almost still. Fig. 1 shows an example pair. For each pair, we formulate the edit with object-wise operations. Each operation is parametrized by a starting position , an ending position (both are object’s 3D center in image plane), a zoom-in factor , and a rotation with respect to the -axis of the camera coordinate system.
The Cityscapes dataset contains training images with pixel-level semantic segmentation and instance segmentation ground truth, but with no 3D annotations, making the geometric inference more challenging. Therefore, given each image, we first predict 3D attributes with our geometric branch pre-trained on Virtual KITTI dataset; we then optimize both attributes and mesh parameters and by minimizing the reprojection loss . We use the Adam solver (Kingma and Ba, 2015) with a learning rate of for iterations.
The semantic, geometric, and textural disentanglement provides an expressive 3D image manipulation scheme. We can modify the 3D attributes of an object to translate, scale, or rotate it in the 3D world, while keeping the consistent visual appearance. We can also change the appearance of the object or the background by modifying the texture code alone.
Methods. We compare our 3D-SDNs with the following two baselines:
2D: Given the source and target positions, the naïve 2D baseline only applies the 2D translation and scaling, discarding the rotation.
2D+: The 2D+ baseline includes the 2D operations above and rotates the 2D silhouette (instead of the 3D shape) along the -axis according to the rotation in the benchmark.
Metrics. The pixel-level distance might not be a meaningful similarity metric, as two visually similar images may have a large L1/L2 distance (Isola et al., 2017). Instead, we adopt the Learned Perceptual Image Patch Similarity (LPIPS) metric (Zhang et al., 2018), which is designed to match the human perception. LPIPS ranges from 0 to 1, with 0 being the most similar. We apply LPIPS on (1) the full image, (2) all edited objects, and (3) the largest edited object.
Besides, we conduct a human study, where we show the target image as well as the edited results from two different methods: 3D-SDN vs. 2D and 3D-SDN vs. 2D+. We ask 120 human subjects on Amazon Mechnical Turk which edited result looks closer to the target. For better visualization, we highlight the largest edited object in red. We then compute, between a pair of methods, how often one method is preferred, across all test images.
Results. Fig. 5 and Fig. 6 show qualitative results on Virtual KITTI and Cityscapes, respectively. By modifying semantic, geometric, and texture codes, our editing interface enables a wide range of scene manipulation applications. Fig. 7 shows a direct comparison to a state-of-the-art 2D manipulation method pix2pixHD (Wang et al., 2018). Quantitatively, Table 1(a) shows that our 3D-SDN outperforms both baselines by a large margin regarding LPIPS. Table 1(b) shows that a majority of the human subjects perfer our results to 2D baselines.
2 Evaluation on the Geometric Representation
Methods. As described in Section 3.1, we adopt multiple strategies to improve the estimation of 3D attributes. As an ablation study, we compare the full 3D-SDN, which is first trained using then fine-tuned using , with its four variants:
w/o : we only use the 3D attribute prediction loss .
w/o normalized distance : we predict the original distance in log space rather than the normalized distance .
w/o MultiCAD and FFD: we use a single CAD model without free-form deformation (FFD).
We also compare with a 3D bounding box estimation method (Mousavian et al., 2017), which first infers the object’s 2D bounding box and pose from input and then searches for its 3D bounding box.
Results. Table 2 shows that our full model has significantly smaller 2D reprojection error than other variants. All of the proposed components contribute to the performance.
Conclusion
In this work, we have developed 3D scene de-rendering networks (3D-SDN) to obtain an interpretable and disentangled scene representation with rich semantic, 3D structural, and textural information. Though our work mainly focuses on 3D-aware scene manipulation, the learned representations could be potentially useful for various tasks such as image reasoning, captioning, and analogy-making. Future directions include better handling uncommon object appearance and pose, especially those not in the training set, and dealing with deformable shapes such as human bodies.
Acknowledgements. This work is supported by NSF #1231216, NSF #1524817, ONR MURI N00014-16-1-2007, Toyota Research Institute, and Facebook.