Pixel2Mesh: Generating 3D Mesh Models from Single RGB Images
Nanyang Wang, Yinda Zhang, Zhuwen Li, Yanwei Fu, Wei Liu, Yu-Gang Jiang
Implementation Details
The perceptual feature pooling layer projects a 3D vertex onto the image plane and pools features from the image feature pathway. For a 3D vertex with coordinate , its 2D projection in image is:
where are focal lengths along horizontal and vertical image axis, and is the center of projection of the camera. These parameters can be readily obtained by intrinsic camera calibration once for a camera.
2 Bilinear feature pooling
Suppose the coordinate of the projection is . The feature is pooled using bilinear interpolation [bilinear]:
where , , and are the integral coordinates of the pixel square where the projection reside in.
Comparison with Octree based Voxel Generation Method
This section provides a comparison with a Octree based voxel Reconstruction method [TatarchenkoDB17], which produces higher resolution voxels. Following Tartarchenko et al.[TatarchenkoDB17], we train on Shapenet-car and compare to it as shown in Tab. 1. Our model consistently outperform the octree based approach in all the evaluation metrics.
Hausdorff Distance for the Ablation Study
Hausdorff distance [haufdist] between meshes measures how far two meshes are from each other, we evaluate the ablation study with Hausdorff distance in Tab. 2. The full model outperforms all except the one without Laplacian. Again, we think Laplacian term regulates deformation and thus benefits the surface smoothness and continuity, which is hard to be reflected in Hausdorff distance.
Other Views of the Generated Mesh
We provide mesh visualizations from other viewpoints, please see Fig. 1 for two examples. As can be seen, our model also recovers invisible parts of the 3D mesh.
Sensitivity to the Initial Meshes
We test our approach with different initial meshes. Specifically, we have tried sphere, noise ellipsoid (EllipsoidN) and ellipsoid in vertical (EllipsoidV) in addition to the original ellipsoid in horizontal (EllipsoidH) as shown in Fig. 3. With each different shape, we fine-tune the model with 30,000 iterations and report quantitative and qualitative results in the Fig. 3 and Tab. 3. As can be seen, our method is not sensitive to the initial shape.
More Qualitative Results
We show more qualitative results on rendered images from the ShapeNet dataset in Fig. 2. The first column shows the color images; the 2nd column shows the volume results from Choy et al.[ChoyXGCS16] and the mesh converted using Marching Cube [LorensenC87]; the 3rd column shows the point cloud from Fan et al.[FanSG16] and the mesh converted using Ball Pivoting [BernardiniMRST99]; the 4th column shows the results of Neural 3D Mesh Renderer [KatoUH2018] and the last column shows our results. Notice how our method produces both smooth surface and the sharp details compared to other methods.
2 Real-world images
We show more results on real-world images in Fig. 4. As can be seen, our method generalizes well to real-world images.