PU-Net: Point Cloud Upsampling Network
Lequan Yu, Xianzhi Li, Chi-Wing Fu, Daniel Cohen-Or, Pheng-Ann Heng
Introduction
Point cloud is a fundamental 3D representation that has drawn increasing attention due to the popularity of various depth scanning devices. Recently, pioneering works qi2016pointnet; qi2017pointnet++; klokov2017escape began to explore the possibility of reasoning point clouds by means of deep networks for understanding geometry and recognizing 3D structures. In these works, the deep networks directly extract features from the raw 3D point coordinates without using traditional features, e.g., normal and curvature. These works present impressive results for 3D object classification and semantic scene segmentation.
In this work we are interested in an upsampling problem: given a set of points, generate a denser set of points to describe the underlying geometry by learning the geometry of a training dataset. This upsampling problem is similar in spirit to the image super-resolution problem shi2016real; ledig2016photo; however, dealing with 3D points rather than a 2D grid of pixels poses new challenges. First, unlike the image space, which is represented by a regular grid, point clouds do not have any spatial order and regular structure. Second, the generated points should describe the underlying geometry of a latent target object, meaning that they should roughly lie on the target object surface. Third, the generated points should be informative and should not clutter together. Having said that, the generated output point set should be more uniform on the target object surface. Thus, simple interpolation between input points cannot produce satisfactory results.
To meet the above challenges, we present a data-driven point cloud upsampling network. Our network is applied at a patch-level, with a joint loss function that encourages the upsampled points to remain on the underlying surface with a uniform distribution. The key idea is to learn multi-level features per point, and then expand the point set via a multi-branch convolution unit implicitly in feature space. The expanded feature is then split to a multitude of features, which are then reconstructed to an upsampled point set.
Our network, namely PU-Net, learns geometry semantics of point-based patches from 3D models, and applies the learned knowledge to upsample a given point cloud. It should be noted that unlike previous network-based methods designed for 3D point sets qi2016pointnet; qi2017pointnet++; klokov2017escape, the number of input and output points in our network are not the same.
We formulate two metrics, distribution uniformity and distance deviation from underlying surfaces, to quantitatively evaluate the upsampled point set, and test our method using variety of synthetic and real-scanned data. We also evaluate the performance of our method, and compare it to baseline and state-of-the-art optimization-based methods. Results show that our upsampled points have better uniformity, and are located closer to the underlying surfaces.
An early work by Alexa et al. alexa2003computing upsamples a point set by interpolating points at vertices of a Voronoi diagram in the local tangent space. Lipman et al. lipman2007parameterization present a locally optimal projection (LOP) operator for points resampling and surface reconstruction based on an median. The operator works well even when the input point set contains noise and outliers. Successively, Huang et al. huang2009consolidation propose an improved weighted LOP to address the point set density problem.
Although these works have demonstrated good results, they make a strong assumption that the underlying surface is smooth, thus restricting the method’s scope. Then, Huang et al. huang2013edge introduce an edge-aware point set resampling method by first resampling away from edges and then progressively approaching edges and corners. However, the quality of their results heavily relies on the accuracy of the normals at given points and careful parameter tuning. It is worth mentioning that Wu et al. wu2015deep propose a deep points representation method to fuse consolidation and completion in one coherent step. Since its main focus is on filling large holes, global smoothness is, however, not enforced, so the method is sensitive to large noise. Overall, the above methods are not data-driven, thus heavily relying on priors.
Related work: deep-learning-based methods.
Points in a point cloud do not have any specific order nor follow any regular grid structure, so only a few recent works adopt a deep learning model to directly process point clouds. Most existing works convert a point cloud into some other 3D representations such as the volumetric grids maturana2015voxnet; wu20153d; riegler2016octnet; dai2017shape and geometric graphs bruna2013spectral; masci2015geodesic for processing. Qi et al. qi2016pointnet; qi2017pointnet++ firstly introduced a deep learning network for point cloud classification and segmentation; in particular, the PointNet++ uses a hierarchical feature learning architecture to capture both local and global geometry context. Subsequently, many other networks were proposed for high-level analysis problems with point clouds klokov2017escape; hua2017pointwise; li2018pointcnn; wang2017sgpn; qi2017frustum. However, they all focus on global or mid-level attributes of point clouds. In another work, Guerrero et al. guerrero2017pcpnet developed a network to estimate the local shape properties in point clouds, including normal and curvature. Other relevant networks focus on 3D reconstruction from 2D images fan2016point; lin2017learning; Groueix2018AtlasNet. To the best of our knowledge, there are no prior works focusing on point cloud upsampling.
Network Architecture
Given a 3D point cloud with point coordinates in nonuniform distributions, our network aims to output a denser point cloud that follows the underlying surface of the target object while being uniform in distribution. Our network architecture (see Fig. 1) has four components: patch extraction, point feature embedding, feature expansion, and coordinate reconstruction. First, we extract patches of points in varying scales and distributions from a given set of prior 3D models (Sec. 2.1). Then, the point feature embedding component maps the raw 3D coordinates to a feature space by hierarchical feature learning and multi-level feature aggregation (Sec. 2.2). After that, we expand the number of features using the feature expansion component (Sec. 2.3) and reconstruct the 3D coordinates of the output point cloud via a series of fully connected layers in the coordinate reconstruction component (Sec. 2.4).
We collect a set of 3D objects as prior information for training. These objects cover a rich variety of shapes, from smooth surface to shapes with sharp edges and corners. Essentially, for our network to upsample a point cloud, it should learn local geometry patterns from the objects. This motivates us to take a patch-based approach to train the network and to learn the geometry semantics.
In detail, we randomly select points on the surface of these objects. From each selected point, we grow a surface patch on the object, such that any point on the patch is within a certain geodesic distance () from the selected point over the surface. Then, we use Poisson disk sampling to randomly generate points on each patch as the referenced ground truth point distribution on the patch. In our upsampling task, both local and global context contribute to a smooth and uniform output. Hence, we set with varying sizes, so that we can extract patches of points on the prior objects with varying scale and density.
2 Point Feature Embedding
To learn both local and global geometry context from the patches, we consider the following two feature learning strategies, whose benefits complement each other:
Hierarchical feature learning. Progressively capturing features of growing scales in a hierarchy has been proved to be an effective strategy for extracting local and global features. Hence, we adopt the recently proposed hierarchical feature learning mechanism in PointNet++ qi2017pointnet++ as the very frontal part in our network. To adopt hierarchical feature learning for point cloud upsampling, we specifically use a relatively small grouping radius in each level, since generating new points usually involves more of the local context than the high-level recognition tasks in qi2017pointnet++.
Multi-level feature aggregation. Lower layers in a network generally correspond to local features in smaller scales, and vice versa. For better upsampling results, we should optimally aggregate features in different levels. Some previous works adopt skip-connections for cascaded multi-level feature aggregation long2015fully; ronneberger2015u; qi2017pointnet++. However, we found by experiments that such top-down propagation is not very efficient for aggregating features in our upsampling problem. Therefore, we propose to directly combine features from different levels and let the network learn the importance of each level hariharan2015hypercolumns; xie2015holistically; hou2016deeply.
3 Feature Expansion
We therefore propose an efficient feature expansion operation based on the sub-pixel convolution layer shi2016real. This operation can be represented as:
It is worth mentioning that the feature sets generated from the first convolution in each set have a high correlation, and this would cause the final reconstructed 3D points to be located too close to one another. Hence, we further add another convolution (with separate weights) for each feature set. Since we train the network to learn the different convolutions for the feature sets, these new features can include more diverse information, thus reducing their correlations. This feature expansion operation can be implemented by applying separated convolutions to the feature sets; see Fig. 1. It can also be implemented by more computation efficient grouped convolution krizhevsky2012imagenet; xie2016aggregated; zhang2017shufflenet.
4 Coordinate Reconstruction
End-to-End Network Training
Point cloud upsampling is an ill-posed problem due to the uncertainty or ambiguity of upsampled point clouds. Given a sparse input point cloud, there are many feasible output point distributions. Therefore, we do not have the notion of “correct pairs” of input and ground truth. To alleviate this problem, we propose an on-the-fly input generation scheme. Specifically, the referenced ground truth point distribution of a training patch is fixed, whereas the input points are randomly sampled from the ground truth point set with a downsampling rate of at each training epoch. Intuitively, this scheme is equivalent to simulating many feasible output point distributions for a given sparse input point distribution. Additionally, this scheme can further enlarge the training dataset, allowing us to depend on a relatively small dataset for training.
2 Joint Loss Function
We propose a novel joint loss function to train the network in an end-to-end fashion. As we mentioned earlier, the function should encourage the generated points to be located on the underlying object surfaces in a more uniform distribution. Therefore, we design a joint loss function that combines the reconstruction loss and repulsion loss.
where indicates the bijection mapping.
Actually, Chamfer Distance (CD) is another candidate for evaluating the similarity between two point sets. However, compared with CD, EMD can better capture the shape (see fan2016point for more details) to encourage the output points to be located close to the underlying object surfaces. Hence, we choose to use EMD in our reconstruction loss.
Repulsion loss. Although training with the reconstruction loss can generate points on the underlying object surfaces, the generated points tend to be located near the original points. To distribute the generated points more uniformly, we design the repulsion loss, which is represented as:
where is the number of output points, is the index set of the -nearest neighbors of point , and is the L2-norm. is called the repulsion term, which is a decreasing function to penalize if is located too close to other points in . To penalize only when it is too close to its neighboring points, we add two restrictions: (i) we only consider points in the -nearest neighborhood of ; and (ii) we use the fast-decaying weight function in the repulsion loss; that is, we follow lipman2007parameterization; huang2009consolidation to set in our experiments.
Altogether, we train the network in an end-to-end manner by minimizing the following joint loss function:
where indicates the parameters in our network, balances the reconstruction loss and repulsion loss, and denotes the multiplier of weight decay. For simplicity, we ignore the index of each training sample.
Experiments
Since there are no public benchmarks for point cloud upsampling, we collect a dataset of 60 different models from the Visionair repository timmurphy.org, ranging from smooth non-rigid objects (e.g., Bunny) to steep rigid objects (e.g., Chair). Among them, we randomly select 40 for training, and use the rest for testing The complete object list can be found in the supplementary material.. We crop 100 patches for each training object, and we use patches to train the network in total. For testing objects, we use Monte-Carlo random sampling approach to sample 5000 points on each object as input. To further demonstrate the generalization ability of our network, we directly test our well-trained network on the SHREC15 Lian2015 dataset, which contains 1200 shapes from 50 categories. In detail, we randomly choose one model from each category for testing, considering that each category contains 24 similar objects in various poses. As for ModelNet40 wu20153d and ShapeNet chang2015shapenet, we found it hard to extract patches from those objects due to the low mesh quality (e.g., holes, self-intersection, etc.). Therefore, we use them for testing; see the supplementary material for the results.
2 Implementation Details
The default point number of each patch is 4096, and the upsampling rate is 4. Therefore, each input patch has 1024 points. To avoid overfitting, we augment the data by randomly rotating, shifting and scaling the data. We use 4 levels with grouping radii 0.05, 0.1, 0.2 and 0.3 in the point feature embedding component, and the dimension of the restored feature is 64. For details on other network architecture parameters, please see our supplementary material. Parameters and in repulsion loss are set as 5 and 0.03, respectively. The balancing weights and are set as 0.01 and = , respectively. The implementation is based on TensorFlow https://github.com/yulequan/PU-Net. For the optimization, we train the network for 120 epoch using the Adam kingma2014adam algorithm with a minibatch size of 28 and a learning rate of 0.001. Generally, the training took about 4.5h on the NVIDIA TITAN Xp GPU.
3 Evaluation Metric
To quantitatively evaluate the quality of the output point sets, we formulate two metrics to measure the deviation between the output points and the ground truth meshes, as well as the distribution uniformity of the output points. For surface deviation, we find the closest point on the mesh for each predicted point , and calculate the distance between them. Then we compute the mean and standard deviation over all the points as one of our metrics.
As for the uniformity metric, we randomly put equal-size disks on the object surface ( in our experiments) and calculate the standard deviation of the number of points inside the disks. We further normalize the density of each object and then compute the overall uniformity of the point sets over all the objects in the testing dataset. Therefore, we define the normalized uniformity coefficient (NUC) with disk area percentage as:
where is the number of points within the -th disk of the -th object, is the total number of points on the -th object, is the total number of test objects, and is the percentage of the disk area over the total object surface area. Note that we use geodesic distance rather than Euclidean distance to form the disks. Fig. 2 shows three different point distributions with their corresponding NUC values. As we can see, the proposed NUC metric can effectively reveal the point set uniformity: the lower the UNC value, the more uniform the point set distribution is.
4 Comparisons with Other Methods
Comparison with an optimization-based method. We compare our method with the Edge Aware Resampling (EAR) method huang2013edge, which is a state-of-art method for point cloud upsampling. The results are shown in Fig. 3, where the Chair is from our collected testing dataset and the Spider is from SHREC15. We color-code the point clouds to show the deviation from the ground truth meshes. There are 1024 points in the input and we do a 4X upsampling. Since EAR relies on the normal information, to be fair, we calculate the normal according to the ground truth mesh. We show two results of EAR with increasing radius, while setting other parameters to their default values. As we can see, the radius parameter has a great influence on EAR’s performance. For relatively small radius, the output has low surface deviation but the added points are not uniform, while more outliers are introduced if the radius is large. In contrast, our method can better balance the deviation and uniformity without the need to carefully tune the parameters.
Comparison with deep learning-based methods. As far as we know, we are not aware of any deep learning-based method for point cloud upsampling, so we design some baseline methods for comparison. Since PointNet qi2016pointnet and PointNet++ qi2017pointnet++ are pioneers for 3D point cloud reasoning with deep learning techniques, we design the baselines based on them. Specifically, we adopt the semantic segmentation network architecture for point feature embedding and use one set of convolutions for feature expansion. Note that we consider two versions of PointNet++: basic PointNet++ and PointNet++ with multi-scale grouping (MSG) for handling non-uniform sampling density; hence, we have three baselines in total, and we train them only with the reconstruction loss. Please refer to the supplemental material for details of the baseline network architectures. Although we modify the PointNet, PointNet++, and PointNet++(MSG) architectures for the upsampling problem, for convenience, we still call the baselines by their original names.
Tables 1 and 2 list the quantitative comparison results on our collected dataset and the SHREC15 dataset, respectively. Note that measuring NUC with small shows local distribution uniformity in small regions, while measuring NUC with large shows more global uniformity. Among the baselines, PointNet performs the worst, since it cannot capture local structure information. Compared with PointNet++, PointNet++(MSG) can slightly improve the uniformity due to the explicit multi-scale information grouping. However, it involves more parameters, thus significantly prolonging the training and testing time. Overall, our PU-Net achieves the best performance with the lowest deviation from surface and the best distribution uniformity compared to the baselines, especially on the local uniformity.
Fig. 4 shows results for visual comparison, where points are colored by their distance deviations from surface. As we can see, the point clouds predicted by our method better match the underlying surface with lower deviations.
5 Architecture Design Analysis
Analyzing the feature expansion. We compare our proposed feature expansion scheme with two interpolation-like schemes. The first one is similar to a naive point interpolation (denoted as interp1). After extracting the point feature of each point, we combine its own features and the features from the nearest neighboring points to generate the upsampled features. The second one introduces more randomness (denoted as interp2). Instead of using the features from the nearest neighbors, we use a radius based ball query to find the neighborhood and combine the features from these points to generate the upsampled features. We train these two networks with the reconstruction loss (also named as the EMD loss) and the results are listed in the top two rows of Table 3. For fair comparison, we also train our network only with the EMD loss. The results are shown in the fourth row. Comparing with these two interpolation-like schemes, our proposed scheme can generate more uniform outputs with comparable surface distance deviation.
Comparing different loss functions. As mentioned above, the EMD can better capture the object shape than CD. Comparing their performance in Table 3, we can see that the EMD loss improves the output uniformity with low surface distance deviation when comparing with the CD loss, meaning that the EMD loss can better encourage the output points to lie on the underlying surface. Furthermore, by comparing the last two rows in Table 3, we can see that the repulsion loss can further improve the uniformity of the output.
6 More Experiments
Surface reconstruction from upsampled point sets. An important application of point cloud upsampling is to improve the surface reconstruction quality. Hence, we compare the reconstruction results of different methods with the direct Poisson surface reconstruction method poisson2006 provided in MeshLab LocalChapterEvents:ItalChap:ItalianChapConf2008:129-136; see Fig. 5. We can observe that the reconstruction result from our method is the closest to the ground truth, while other methods either miss certain structures (e.g., the leg of the Horse) or overfill the hole.
Results of iterative upsampling. To study the ability of our network to handle varying number of input points, we design an iterative upsampling experiment, which takes the output of the previous iteration as the input of the next iteration. Fig. 6 shows the results. The initial input point cloud has 1024 points and we increase fourfold in each iteration. From the results, we can see that our network can produce reasonable results for different number of input points. Furthermore, this iterative upsampling experiment also shows the anti-noise ability of our network to resist the accumulated errors introduced in the iterative upsampling.
Results from noisy input point sets. Fig. 7 shows the surface reconstruction results from noisy point clouds (Gaussian noise of 0.5% and 1% of object bounding box diagonal), which demonstrate that our network facilitates the production of better surfaces even with noisy inputs.
Results on real-scanned point clouds. Lastly, we evaluated the ability of our network to upsample real-scanned point clouds, which were downloaded from Visionair timmurphy.org. In Fig. 8, the left-most column presents the real-scanned point clouds. Even though each real-scanned point cloud contains millions of points, the phenomenon of inhomogeneity still exists. For better visualization, we cut small patches from the original point clouds and show the patches in the middle column. We can observe that the real-scanned points tend to have line structures, while our network still has the ability to uniformly add points in the sparse regions.
Conclusion
In this paper, we present a deep network for point cloud upsampling, with the goal of generating a denser and uniform set of points from a sparser set of points. Our network is trained at a patch-level using a multi-level feature aggregation manner, thus capturing both local and global information. The design of our network bypasses the need for a prescribed order among the points, by operating on individual features that contain non-local geometry to allow a context-aware upsampling. Our experiments demonstrate the effectiveness of our method. As the first attempt using deep networks, our method still has a number of limitations. Firstly, it is not designed for completion, so our network can not fill large holes and missing parts. Besides, our network may not be able to add meaningful points for tiny structures that are severely undersampled.
In the future, we would like to investigate and develop more means to handle irregular and sparse data, both for regression purposes and for synthesis. One immediate step is to develop a downsampling method. Although, downsampling seems like a simpler problem, there is room to devise proper losses and architecture that maximize the preservation of information in the decimated point set. We believe that in general, the development of deep learning methods for irregular structures is a viable research direction.
Acknowledgments. We thank anonymous reviewers for the comments and suggestions. The work is supported in part by the National Basic Program of China, the 973 Program (Project No. 2015CB351706), the Research Grants Council of the Hong Kong Special Administrative Region (Project no. CUHK 14225616), the Shenzhen Science and Technology Program (No. JCYJ20170413162617606), and the CUHK strategic recruitment fund.
References
Appendix A Overview
In this supplementary material, we first provide more details about our collected dataset in Section B. Then, we show the details of our network architecture as well as the baseline networks employed in the experiments in Section C.
Appendix B Details of our Collected Dataset
We collect 60 different 3D models to form our training and testing datasets. The specific name of each model is shown in Table 4. We also present the shapes of some training and testing 3D models in our dataset in Fig. 9 and Fig. 10, respectively. As we can see, our collected datasets have a large variation in geometry shapes, containing 3D models with smooth surface regions (first row) and 3D models with sharp corners and edges (second row). There is also a large variation between training and testing 3D models, indicating a good generalization ability of our proposed method.
Appendix C Details of Network Architectures
The details of our network architecture are listed as follows.
In the hierarchical feature learning component, we use four levels to extract local features. Following the notations in PointNet++, we use (, , []) to represent a level with local regions of ball radius , and [] the fully connected layers with width (). Therefore, the parameters we use are , , and .
In the coordinate reconstruction component, we use two fully connected layers with 64 and 3 output channels, respectively.
The details of the baseline architectures are illustrated in Fig. 11, Fig. 12 and Fig. 13.
All the convolution layers and fully connected layers in the above networks are followed by the ReLU operator, except for the last coordinate regression layer.