Pix3D: Dataset and Methods for Single-Image 3D Shape Modeling

Xingyuan Sun, Jiajun Wu, Xiuming Zhang, Zhoutong Zhang, Chengkai Zhang, Tianfan Xue, Joshua B. Tenenbaum, William T. Freeman

Introduction

The computer vision community has put major efforts in building datasets. In 3D vision, there are rich 3D CAD model repositories like ShapeNet and the Princeton Shape Benchmark , large-scale datasets associating images and shapes like Pascal 3D+ and ObjectNet3D , and benchmarks with fine-grained pose annotations for shapes in images like IKEA . Why do we need one more?

Looking into Figure 1, we realize existing datasets have limitations for the task of modeling a 3D object from a single image. ShapeNet is a large dataset for 3D models, but does not come with real images; Pascal 3D+ and ObjectNet3D have real images, but the image-shape alignment is rough because the 3D models do not match the objects in images; IKEA has high-quality image-3D alignment, but it only contains 90 3D models and 759 images.

We desire a dataset that has all three merits—a large-scale dataset of real images and ground-truth shapes with precise 2D-3D alignment. Our dataset, named Pix3D, has 395 3D shapes of nine object categories. Each shape associates with a set of real images, capturing the exact object in diverse environments. Further, the 10,069 image-shape pairs have precise 3D annotations, giving pixel-level alignment between shapes and their silhouettes in the images.

Building such a dataset, however, is highly challenging. For each object, it is difficult to simultaneously collect its high-quality geometry and in-the-wild images. We can crawl many images of real-world objects, but we do not have access to their shapes; 3D CAD repositories offer object geometry, but do not come with real images. Further, for each image-shape pair, we need a precise pose annotation that aligns the shape with its projection in the image.

We overcome these challenges by constructing Pix3D in three steps. First, we collect a large number of image-shape pairs by crawling the web and performing 3D scans ourselves. Second, we collect 2D keypoint annotations of objects in the images on Amazon Mechanical Turk, with which we optimize for 3D poses that align shapes with image silhouettes. Third, we filter out image-shape pairs with a poor alignment and, at the same time, collect attributes (i.e., truncation, occlusion) for each instance, again by crowdsourcing.

In addition to high-quality data, we need a proper metric to objectively evaluate the reconstruction results. A well-designed metric should reflect the visual appealingness of the reconstructions. In this paper, we calibrate commonly used metrics, including intersection over union, Chamfer distance, and earth mover’s distance, on how well they capture human perception of shape similarity. Based on this, we benchmark state-of-the-art algorithms for 3D object modeling on Pix3D to demonstrate their strengths and weaknesses.

With its high-quality alignment, Pix3D is also suitable for object pose estimation and shape retrieval. To demonstrate that, we propose a novel model that performs shape and pose estimation simultaneously. Given a single RGB image, our model first predicts its 2.5D sketches, and then regresses the 3D shape and the camera parameters from the estimated 2.5D sketches. Experiments show that multi-task learning helps to boost the model’s performance.

Our contributions are three-fold. First, we build a new dataset for single-image 3D object modeling; Pix3D has a diverse collection of image-shape pairs with precise 2D-3D alignment. Second, we calibrate metrics for 3D shape reconstruction based on their correlations with human perception, and benchmark state-of-the-art algorithms on 3D reconstruction, pose estimation, and shape retrieval. Third, we present a novel model that simultaneously estimates object shape and pose, achieving state-of-the-art performance on both tasks.

Related Work

For decades, researchers have been building datasets of 3D objects, either as a repository of 3D CAD models or as images of 3D shapes with pose annotations . Both directions have witnessed the rapid development of web-scale databases: ShapeNet was proposed as a large repository of more than 50K models covering 55 categories, and Xiang et al. built Pascal 3D+ and ObjectNet3D , two large-scale datasets with alignment between 2D images and the 3D shape inside. While these datasets have helped to advance the field of 3D shape modeling, they have their respective limitations: datasets like ShapeNet or Elastic2D3D do not have real images, and recent 3D reconstruction challenges using ShapeNet have to be exclusively on synthetic images ; Pascal 3D+ and ObjectNet3D have only rough alignment between images and shapes, because objects in the images are matched to a pre-defined set of CAD models, not their actual shapes. This has limited their usage as a benchmark for 3D shape reconstruction .

With depth sensors like Kinect , the community has built various RGB-D or depth-only datasets of objects and scenes. We refer readers to the review article from Firman for a comprehensive list. Among those, many object datasets are designed for benchmarking robot manipulation . These datasets often contain a relatively small set of hand-held objects in front of clean backgrounds. Tanks and Temples is an exciting new benchmark with 14 scenes, designed for high-quality, large-scale, multi-view 3D reconstruction. In comparison, our dataset, Pix3D, focuses on reconstructing a 3D object from a single image, and contains much more real-world objects and images.

Probably the dataset closest to Pix3D is the large collection of object scans from Choi et al. , which contains a rich and diverse set of shapes, each with an RGB-D video. Their dataset, however, is not ideal for single-image 3D shape modeling for two reasons. First, the object of interest may be truncated throughout the video; this is especially the case for large objects like sofas. Second, their dataset does not explore the various contexts that an object may appear in, as each shape is only associated with a single scan. In Pix3D, we address both problems by leveraging powerful web search engines and crowdsourcing.

Another closely related benchmark is IKEA , which provides accurate alignment between images of IKEA objects and 3D CAD models. This dataset is therefore particularly suitable for fine pose estimation. However, it contains only 759 images and 90 shapes, relatively small for shape modelingOnly 90 of the 219 shapes in the IKEA dataset have associated images.. In contrast, Pix3D contains 10,069 images (13.3x) and 395 shapes (4.4x) of greater variations.

Researchers have also explored constructing scene datasets with 3D annotations. Notable attempts include LabelMe-3D , NYU-D , SUN RGB-D , KITTI , and modern large-scale RGB-D scene datasets . These datasets are either synthetic or contain only 3D surfaces of real scenes. Pix3D, in contrast, offers accurate alignment between 3D object shape and 2D images in the wild.

Single-image 3D reconstruction.

The problem of recovering object shape from a single image is challenging, as it requires both powerful recognition systems and prior shape knowledge. Using deep convolutional networks, researchers have made significant progress in recent years . While most of these approaches represent objects in voxels, there have also been attempts to reconstruct objects in point clouds or octave trees . In this paper, we demonstrate that our newly proposed Pix3D serves as an ideal benchmark for evaluating these algorithms. We also propose a novel model that jointly estimates an object’s shape and its 3D pose.

Shape retrieval.

Another related research direction is retrieving similar 3D shapes given a single image, instead of reconstructing the object’s actual geometry . Pix3D contains shapes with significant inter-class and intra-class variations, and is therefore suitable for both general-purpose and fine-grained shape retrieval tasks.

D pose estimation.

Many of the aforementioned object datasets include annotations of object poses . Researchers have also proposed numerous methods on 3D pose estimation . In this paper, we show that Pix3D is also a proper benchmark for this task.

Building Pix3D

Figure 2 summarizes how we build Pix3D. We collect raw images from web search engines and shapes from 3D repositories; we also take pictures and scan shapes ourselves. Finally, we use labeled keypoints on both 2D images and 3D shapes to align them.

We obtain raw image-shape pairs in two ways. One is to crawl images of IKEA furniture from the web and align them with CAD models provided in the IKEA dataset . The other is to directly scan 3D shapes and take pictures.

The IKEA dataset contains 219 high-quality 3D models of IKEA furniture, but has only 759 images for 90 shapes. Therefore, we choose to keep the 3D shapes from IKEA dataset, but expand the set of 2D images using online image search engines and crowdsourcing.

For each 3D shape, we first search for its corresponding 2D images through Google, Bing, and Baidu, using its IKEA model name as the keyword. We obtain 104,220 images for the 219 shapes. We then use Amazon Mechanical Turk (AMT) to remove irrelevant ones. For each image, we ask three AMT workers to label whether this image matches the 3D shape or not. For images whose three responses differ, we ask three additional workers and decide whether to keep them based on majority voting. We end up with 14,600 images for the 219 IKEA shapes.

D scan.

We scan non-IKEA objects with a Structure Sensorhttps://structure.io mounted on an iPad. We choose to use the Structure Sensor because its mobility enables us to capture a wide range of shapes.

The iPad RGB camera is synchronized with the depth sensor at 30 Hz, and calibrated by the Scanner App provided by Occipital, Inc.https://occipital.com The resolution of RGB frames is 2592×\times1936, and the resolution of depth frames is 320×\times240. For each object, we take a short video and fuse the depth data to get its 3D mesh by using fusion algorithm provided by Occipital, Inc. We also take 10–20 images for each scanned object in front of various backgrounds from different viewpoints, making sure the object is neither cropped nor occluded. In total, we have scanned 209 objects and taken 2,313 images. Combining these with the IKEA shapes and images, we have 418 shapes and 16,913 images altogether.

2 Image-Shape Alignment

To align a 3D CAD model with its projection in a 2D image, we need to solve for its 3D pose (translation and rotation), and the camera parameters used to capture the image.

We use a keypoint-based method inspired by Lim et al. . Denote the keypoints’ 2D coordinates as X2D={x1,x2,⋯ ,xn}\bm{X}_{\text{2D}}=\{\mathbf{x}_{1},\mathbf{x}_{2},\cdots,\mathbf{x}_{n}\} and their corresponding 3D coordinates as X3D={X1,X2,⋯ ,Xn}\bm{X}_{\text{3D}}=\{\mathbf{X}_{1},\mathbf{X}_{2},\cdots,\mathbf{X}_{n}\}. We solve for camera parameters and 3D poses that minimize the reprojection error of the keypoints. Specifically, we want to find the projection matrix P\bm{P} that minimizes

where ProjP(⋅)\text{Proj}_{\bm{P}}(\cdot) is the projection function.

where ff is the focal length, and ww and hh are the width and height of the image. Therefore, there are altogether seven parameters to be estimated: rotations θ,ϕ,ψ\theta,\phi,\psi, translations x,y,zx,y,z, and focal length ff (Rotation matrix RR is determined by θ\theta, ϕ\phi, and ψ\psi).

To solve Equation 1, we first calculate a rough 3D pose using the Efficient PnP algorithm and then refine it using the Levenberg-Marquardt algorithm , as shown in Figure 2. Details of each step are described below.

Perspective-n-Point (PnP) is the problem of estimating the pose of a calibrated camera given paired 3D points and 2D projections. The Efficient PnP (EPnP) algorithm solves the problem using virtual control points . Because EPnP does not estimate the focal length, we enumerate the focal length ff from 300 to 2,000 with a step size of 10, solve for the 3D pose with each ff, and choose the one with the minimum projection error.

The Levenberg-Marquardt algorithm (LMA).

We take the output of EPnP with 50 random disturbances as the initial states, and run LMA on each of them. Finally, we choose the solution with the minimum projection error.

Implementation details.

For each 3D shape, we manually label its 3D keypoints. The number of keypoints ranges from 8 to 24. For each image, we ask three AMT workers to label if each keypoint is visible on the image, and if so, where it is. We only consider visible keypoints during the optimization.

The 2D keypoint annotations are noisy, which severely hurts the performance of the optimization algorithm. We try two methods to increase its robustness. The first is to use RANSAC. The second is to use only a subset of 2D keypoint annotations. For each image, denote C={c1,c2,c3}C=\{c_{1},c_{2},c_{3}\} as its three sets of human annotations. We then enumerate the seven nonempty subsets Ck⊆CC_{k}\subseteq C; for each keypoint, we compute the median of its 2D coordinates in CkC_{k}. We apply our optimization algorithm on every subset CkC_{k}, and keep the output with the minimum projection error. After that, we let three AMT workers choose, for each image, which of the two methods offers better alignment, or neither performs well. At the same time, we also collect attributes (i.e., truncation, occlusion) for each image. Finally, we fine-tune the annotations ourselves using the GUI offered in ObjectNet3D . Altogether there are 395 3D shapes and 10,069 images. Sample 2D-3D pairs are shown in Figure 3.

Exploring Pix3D

We now present some statistics of Pix3D, and contrast it with its predecessors.

Figures 4 and 5 show the category distributions of 2D images and 3D shapes in Pix3D; Figure 6 shows the distribution of the number of images each model has. Our dataset covers a large variety of shapes, each of which has a large number of in-the-wild images. Chairs cover the significant part of Pix3D, because they are common, highly diverse, and well-studied by recent literature .

Quantitative evaluation.

As a quantitative comparison on the quality of Pix3D and other datasets, we randomly select 25 chair and 25 sofa images from PASCAL 3D+ , ObjectNet3D , IKEA , and Pix3D. For each image, we render the projected 2D silhouette of the shape using its pose annotation provided by the dataset. We then manually annotate the ground truth object masks in these images, and calculate Intersection over Union (IoU) between the projections and the ground truth. For each image-shape pair, we also ask 50 AMT workers whether they think the image is picturing the 3D ground truth shape provided by the dataset.

From Table 1, we see that Pix3D has much higher IoUs than PASCAL 3D+ and ObjectNet3D, and slightly higher IoUs compared with the IKEA dataset. Humans also feel IKEA and Pix3D have matched images and shapes, but not PASCAL 3D+ or ObjectNet3D. In addition, we observe that many CAD models in the IKEA dataset are of an incorrect scale, making it challenging to align the shapes with images. For example, there are only 15 unoccluded and untruncated images of sofas in IKEA, while Pix3D has 1,092.

Metrics

Designing a good evaluation metric is important to encourage researchers to design algorithms that reconstruct high-quality 3D geometry, rather than low-quality 3D reconstruction that overfits to a certain metric.

Many 3D reconstruction papers use Intersection over Union (IoU) to evaluate the similarity between ground truth and reconstructed 3D voxels, which may significantly deviate from human perception. In contrast, metrics like shortest distance and geodesic distance are more commonly used than IoU for matching meshes in graphics . Here, we conduct behavioral studies to calibrate IoU, Chamfer distance (CD) , and Earth Mover’s distance (EMD) on how well they reflect human perception.

The definition of IoU is straightforward. For Chamfer distance (CD) and Earth Mover’s distance (EMD), we first convert voxels to point clouds, and then compute CD and EMD between pairs of point clouds.

We first extract the isosurface of each predicted voxel using the Lewiner marching cubes algorithm. In practice, we use 0.1 as a universal surface value for extraction. We then uniformly sample points on the surface meshes and create the densely sampled point clouds. Finally, we randomly sample 1,024 points from each point cloud and normalize them into a unit cube for distance calculation.

Chamfer distance (CD).

For each point in each cloud, CD finds the nearest point in the other point set, and sums the distances up. CD has been used in recent shape retrieval challenges .

Earth Mover’s distance (EMD).

where ϕ:S1→S2\phi:S_{1}\to S_{2} is a bijection. We divide EMD by the size of the point cloud for normalization. In practice, calculating the exact EMD value is computationally expensive; we instead use a (1+ϵ)(1+\epsilon) approximation algorithm .

2 Experiments

We then conduct two user studies to compare these metrics and benchmark how they capture human perception.

We run three shape reconstructions algorithms (3D-R2N2 , DRC , and 3D-VAE-GAN ) on 200 randomly selected images of chairs. We then, for each image and every pair of its three constructions, ask three AMT workers to choose the one that looks closer to the object in the image. We also compute how each pair of objects rank in each metric. Finally, we calculate the Spearman’s rank correlation coefficients between different metrics (i.e., IoU, EMD, CD, and human perception). Table 2 suggests that EMD and CD correlate better with human ratings.

How good is it?

We randomly select 400 images, and show each of them to 15 AMT workers, together with the voxel prediction by DRC and the ground truth shape. We then ask them to rate the reconstruction, on a scale of 1 to 7, based on how similar it is to the ground truth. The scatter plot in Figure 7 suggests that CD and EMD have higher Pearson’s coefficients with human responses.

Approach

Pix3D serves as a benchmark for shape modeling tasks including reconstruction, retrieval, and pose estimation. Here, we design a new model that simultaneously performs shape reconstruction and pose estimation, and evaluate it on Pix3D.

Our model is an extension of MarrNet , both of which use 2.5D sketches (the object’s depth, surface normals, and silhouette) as an intermediate representation. It contains four modules: (1) a 2.5D sketch estimator that predicts the depth, surface normals, and silhouette of the object; (2) a 2.5D sketch encoder that encodes the 2.5D sketches into a low-dimensional latent vector; (3) a 3D shape decoder and (4) a view estimator that decodes a latent vector into a 3D shape and camera parameters, respectively. Different from MarrNet , our model has an additional branch for pose estimation. We briefly describe them below, and please refer to the supplementary material for more details.

The first module takes an RGB image as input and predicts the object’s 2.5D sketches (its depth, surface normals, and silhouette). We use an encoder-decoder network. The encoder is based on a ResNet-18 and turns a 256×\times256 image into 384 feature maps of size 16×\times16; the decoder has three branches for depth, surface normals, and silhouette, respectively, each consisting of four sets of 5×\times5 transposed convolutional, batch normalization and ReLU layers, followed by one 5×\times5 convolutional layer. All output sketches are of size 256×\times256.

5D sketch encoder.

We use a modified ResNet-18 that takes a four-channel image (three for surface normals and one for depth). Each channel is masked by the predicted silhouette. A final linear layer outputs a 200-D latent vector.

D shape decoder.

Our 3D shape decoder has five sets of 4×\times4×\times4 transposed convolutional, batch-norm, and ReLU layers, followed by a 4×\times4×\times4 transposed convolutional layer. It outputs a voxelized shape of size 128×\times128×\times128 in the object’s canonical view.

View estimator.

The view estimator contains three sets of linear, batch normalization, and ReLU layers, followed by two parallel linear and softmax layers that predict the shape’s azimuth and elevation, respectively. Here, we treat pose estimation as a classification problem, where the 360-degree azimuth angle is divided into 24 bins and the 180-degree elevation angle is divided into 12 bins.

Training paradigm.

For training, we use Mitsuba to render each chair in ShapeNet from 20 random views using three types of backgrounds: 1/3 on a white background, 1/3 on high-dynamic-range backgrounds with illumination channels, and 1/3 on backgrounds randomly sampled from the SUN database . We augment our training data by random color and light jittering.

We first train the 2.5D sketch estimator. We then train the 2.5D sketch encoder and the 3D shape decoder (and the view estimator if we’re predicting the pose) jointly. We finally concatenate them for prediction.

Experiments

We now evaluate our model and state-of-the-art algorithms on single-image 3D shape reconstruction, retrieval, and pose estimation, all using Pix3D. For all experiments, we use the 2,894 untruncated and unoccluded chair images.

We compare our model, with and without the pose estimation branch, with the state-of-the-art systems, including 3D-VAE-GAN , 3D-R2N2 , DRC , and MarrNet . We use pre-trained models offered by the authors and we crop the input images as required by each algorithm. The results are shown in Table 3 and Figure 8. Our model outperforms the state-of-the-arts in all metrics. Our full model gets better results compared with the variant without the view estimator, suggesting multi-task learning helps to boost its performance. Also note the discrepancy among metrics: MarrNet has a lower IoU than DRC, but according to EMD and CD, it performs better.

Image-based, fine-grained shape retrieval.

For shape retrieval, we compare our model with 3D-VAE-GAN and MarrNet . We use the latent vector from each algorithm as its embedding of the input image, and use L2 distance for image retrieval. For each test image, we retrieve its K nearest neighbors from the test set, and use Recall@K to compute how many retrieved images are actually depicting the same shape. Here we do not consider images whose shape is not captured by any other images in the test set. The results are shown in Table 4 and Figure 9. Our model (without the pose estimation module) achieves the highest numbers; our model (with the pose estimation module) does not perform as well, because it sometimes retrieves images of objects with the same pose, but not exactly the same shape.

D pose estimation.

We compare our method with Render for CNN . We calculate the classification accuracy for both azimuth and elevation, where the azimuth is divided into 24 bins and the elevation into 12 bins. Table 5 suggests that our model outperforms Render for CNN in pose estimation. Qualitative results are included in Figure 10.

Conclusion

We have presented Pix3D, a large-scale dataset of well-aligned 2D images and 3D shapes. We have also explored how three commonly used metrics correspond to human perception through two behavioral studies and proposed a new model that simultaneously performs shape reconstruction and pose estimation. Experiments showed that our model achieved state-of-the-art performance on 3D reconstruction, shape retrieval, and pose estimation. We hope our paper will inspire future research in single-image 3D shape modeling.

This work is supported by NSF #1212849 and #1447476, ONR MURI N00014-16-1-2007, the Center for Brain, Minds and Machines (NSF STC award CCF-1231216), the Toyota Research Institute, and Shell Research. J. Wu is supported by a Facebook fellowship.

References

A Network Parameters

As mentioned in Section 6 in the main text, we proposed a new model that simultaneously performs 3D shape reconstruction and camera view estimation. Here we provide more details about the network structure.

As shown in Figure 11, our model consists of four components: (1) a 2.5D sketch estimator, which estimates 2.5D sketches from an RGB image, (2) a 2.5D sketch encoder, which encodes 2.5D sketches into a 200-dimensional latent vector, (3) a 3D shape decoder, which decodes a latent vector into voxels and (4) a view estimator, which estimates the camera view from a latent vector.

Table 6 shows the network configuration summary of the 2.5D sketch estimator. We use an encoder-decoder network. The first four rows in Table 6 shows the encoder’s structure and the other rows describe the decoder. The encoder takes in an input RGB image of size 256×\times256 and encodes it into 384 16×\times16 feature maps. The decoder takes in 384 16×\times16 feature maps and decodes them into the object’s surface normals, depth, and silhouette of size 256×\times256.

For the encoder, we use a truncated ResNet-18 with last two layers (average pooling and fully connected) removed. The truncated ResNet-18 is followed by a transposed convolutional layer, a batch normalization layer, and a ReLU layer. For the decoder, we use four sets of 5×\times5 transposed convolutional layers, batch normalization layers and ReLU layers, followed by one 5×\times5 convolutional layer. We do not share weights of layers between three sketches (i.e., surface normal, depth, silhouette).

5D sketch encoder.

The 2.5D sketch encoder is modified from a ResNet-18. It takes in a four-channel image with size 256×\times256 obtained by stacking the three-channel surface normal image and single-channel depth image, both of which are masked by the silhouette. It then encodes them into a 200-dimensional latent vector.

For the first layer of ResNet-18, we change the number of input channels from 3 to 4. We also change the average pooling layer into an adaptive average pooling layer. For the last fully connected layer, we change the output dimensional to 200.

D shape decoder.

Table 7 shows the network architecture of the 3D shape decoder. It takes in a 200-dimensional latent vector and decodes it into a voxel grid of size 128×\times128×\times128. We use five sets of 4×\times4×\times4 3D transposed convolutional layers, 3D batch normalization layers and ReLU layers, followed by one 4×\times4×\times4 transposed convolutional layer.

View estimator.

Table 8 shows the network configuration summary of the view estimator. We use three sets of fully connected, batch normalization, and ReLU layers, followed by two parallel fully connected and softmax layers that predict azimuth and elevation, respectively.

B Training Paradigms

As mentioned in Section 7 in the main text, we train our proposed method and test it on three different tasks. Here we provide more details about training.

We first train the 2.5D sketch estimator. We then train the 2.5D sketch encoder and the 3D shape decoder (and the view estimator if we’re predicting the pose) jointly.

The loss function is defined as the sum of mean squared error between predicted sketches and ground truth sketches (with size average). Specifically,

where MSE is mean square error with size average, pred stands for prediction, and gt stands for ground truth.

The batch size is 4. We use Adam as the optimizer and set the learning rate to 2×10−42\times 10^{-4}. The model is trained for 270 epochs, each with 6,000 batches. We choose to use the one with the minimum validation loss.

Shape and view estimation.

The loss function is defined as the weighted sum of the 3D reconstruction loss and the pose estimation loss. The loss function for 3D reconstruction is

where BCEL\text{BCE}_{L} is the binary cross-entropy between the target and the output logits (no sigmoid applied) with size average, pred stands for prediction, and gt stands for ground truth. The loss function for pose estimation is

where BCE is the binary cross-entropy between the target and the output with size average, pred stands for prediction, and gt stands for ground truth. Note that we have already applied softmax to azimuth and elevation predictions in our model. The global loss function is thus

We set α\alpha to 0.6. The batch size is 4. We use stochastic gradient descent with a momentum of 0.9 as the optimizer and set the learning rate to 0.1. The model is trained for 300 epochs, each with 6,000 batches. We choose to use the one with the minimum validation loss.

C Evaluation Metrics

Here, we explain in detail our evaluation protocol for single-image 3D shape reconstruction. As different voxelization methods may result in objects of different scales in the voxel grid, for a fair comparison, we preprocess all voxels and point clouds before calculating IoU, CD and EMD.

For IoU, we first find the bounding box of the object with a threshold of 0.1, pad the bounding box into a cube, and then use trilinear interpolation to resample to the desired resolution (323\text{32}^{\text{3}}). Some algorithms reconstruct shapes at a resolution of 1283\text{128}^{\text{3}}. In this case, we first, apply a 4×\times max pooling before trilinear interpolation; without the max pooling, the sampling grid can be too sparse and some thin structure can be left out. After the resampling of both the output voxel and the ground truth voxel, we search for an optimal threshold that maximizes the average IoU score over all objects, from 0.01 to 0.50 with a step size of 0.01.

For CD and EMD, we first sample a point cloud from the voxelized reconstructions. For each shape, we compute its isosurface with a threshold of 0.1, and then sample 1,024 points from the surface. All point clouds are then translated and scaled such that the bounding box of the point cloud is centered at the origin with its longest side being 1. We then compute CD and EMD for each pair of point clouds.

D Nearest Neighbors of 3D Shapes

In Section 5 in the main text, we have compared three different metrics from two different perspectives. We here compare them in another way: for a 3D shape, we retrieve three nearest neighbors from Pix3D according to IoU, EMD and CD, respectively. Results are shown in Figure 12. EMD and CD perform slightly better than IoU.

E Sample Data Points in Pix3D

We supply more sample data points in Figures 13, 14, and 15. Figures 13 and 14 show the diversity of 3D shapes and the quality of 2D-3D alignment in Pix3D. Figure 15 shows that each shape in Pix3D is matched with a rich set of 2D images.