BiPointNet: Binary Neural Network for Point Clouds
Haotong Qin, Zhongang Cai, Mingyuan Zhang, Yifu Ding, Haiyu Zhao, Shuai Yi, Xianglong Liu, Hao Su
Introduction
With the advent of deep neural networks that directly process raw point clouds (PointNet (Qi et al., 2017a) as the pioneering work), great success has been achieved in learning on point clouds (Qi et al., 2017b; Li et al., 2018; Wang et al., 2019a; Wu et al., 2019; Thomas et al., 2019; Liu et al., 2019b; Zhang et al., 2019b). Point cloud applications, such as autonomous driving and augmented reality, often require real-time interaction and fast response. However, computation for such applications is usually deployed on resource-constrained edge devices. To address the challenge, novel algorithms, such as Grid-GCN (Xu et al., 2020b), RandLA-Net (Hu et al., 2020), and PointVoxel (Liu et al., 2019d), have been proposed to accelerate those point cloud processing networks. While significant speedup and memory footprint reduction have been achieved, these works still rely on expensive floating-point operations, leaving room for further optimization of the performance from the model quantization perspective. Model binarization (Rastegari et al., 2016; Bulat & Tzimiropoulos, 2019; Hubara et al., 2016; Wang et al., 2020; Zhu et al., 2019; Xu et al., 2019) emerged as one of the most promising approaches to optimize neural networks for better computational and memory usage efficiency. Binary Neural Networks (BNNs) leverage 1) compact binarized parameters that take small memory space, and 2) highly efficient bitwise operations which are far less costly compared to the floating-point counterparts.
Despite that in 2D vision tasks (Krizhevsky et al., 2012; Simonyan & Zisserman, 2014; Szegedy et al., 2015; Girshick et al., 2014; Girshick, 2015; Russakovsky et al., 2015; Wang et al., 2019b; Zhang et al., 2021) has been studied extensively by the model binarization community, the methods developed are not readily transferable for 3D point cloud networks due to the fundamental differences between 2D images and 3D point clouds. First, to gain efficiency in processing unordered 3D points, many point cloud learning methods rely heavily on pooling layers with large receptive field to aggregate point-wise features. As shown in PointNet (Qi et al., 2017b), global pooling provides a strong recognition capability. However, this practice poses challenges for binarization. Our analyses show that the degradation of feature diversity, a persistent problem with binarization (Liu et al., 2019a; Qin et al., 2020b; Xie et al., 2017), is significantly amplified by the global aggregation function (Figure 2), leading to homogenization of global features with limited discriminability. Second, the binarization causes immense scale distortion at the point-wise feature extraction stage, which is detrimental to model performance in two ways: the saturation of forward-propagated features and backward-propagated gradients hinders optimization, and the disruption of the scale-sensitive structures (Figure 3) results in the invalidation of their designated functionality.
In this paper, we provide theoretical formulations of the above-mentioned phenomenons and obtain insights through in-depth analysis. Such understanding allows us to propose a method that turns full-precision point cloud networks into extremely efficient yet strong binarized models (see the overview in Figure 1). To tackle the homogenization of the binarized features after passing the aggregation function, we study the correlation between the information entropy of binarization features and the performance of point cloud aggregation functions. We thus propose Entropy-Maximizing Aggregation (EMA) that shifts the feature distribution towards the statistical optimum, effectively improving expression capability of the global features. Moreover, given maximized information entropy, we further develop Layer-wise Scale Recovery (LSR) to efficiently restore the output scale that enhances optimization, which allows scale-sensitive structures to function properly. LSR uses only one learnable parameter per layer, leading to negligible storage increment and computation overhead.
Our BiPointNet is the first binarization approaches to deep learning on point clouds, and it outperforms existing binarization algorithms for 2D vision by convincing margins. It is even almost on par (within 1-2%) with the full-precision counterpart. Although we conduct most analysis on the PointNet baseline, we show that our methods are generic and can be readily extendable to other popular backbones, such as PointNet++ (Qi et al., 2017b), PointCNN (Li et al., 2018), DGCNN (Wang et al., 2019a), and PointConv (Wu et al., 2019), which are the representatives of mainstream categories of point cloud feature extractors. Moreover, extensive experiments on multiple fundamental tasks on the point cloud, such as classification, part segmentation, and semantic segmentation, highlight that our BiPointNet is task-agnostic. Besides, we highlight that our EMA and LSR are efficient and easy to implement in practice: in the actual test on popular edge devices, BiPointNet achieves speedup and storage savings compared to the full-precision PointNet. Our code is released at https://github.com/htqin/BiPointNet.
Related Work
Network Binarization. Recently, various quantization methods for neural networks have emerged, such as uniform quantization (Gong et al., 2019; Zhu et al., 2020), mixed-precision quantization (Wu et al., 2018; Yu et al., 2020), and binarization. Among these methods, binarization enjoys compact binarized parameters and highly efficient bitwise operations for extreme compression and acceleration (Rastegari et al., 2016; Qin et al., 2020a). In general, the forward and backward propagation of binarized models in the training process can be formulated as:
where denotes the element in floating-point weights and activations, denotes the element in binarized weights and activations . , and donate the gradient and , respectively, where is the cost function for the minibatch. In forward propagation, function is directly applied to obtain the binary parameters. In backward propagation, the Straight-Through Estimator (STE) (Bengio et al., 2013) is used to obtain the derivative of the function, avoiding getting all zero gradients. The existing binarization methods are designed to obtain accurate binarized networks by minimizing the quantization error (Rastegari et al., 2016; Zhou et al., 2016; Lin et al., 2017), improving loss function (Ding et al., 2019; Hou et al., 2017), reducing the gradient error (Liu et al., 2018; 2020), and designing novel structures and pipelines (Martinez et al., 2020). Unfortunately, we show in Sec 3 that these methods, designed for 2D vision tasks, are not readily transferable to 3D point clouds.
Deep Learning on Point Clouds. PointNet (Qi et al., 2017a) is the first deep learning model that processes raw point clouds directly. The basic building blocks proposed by PointNet such as MLP for point-wise feature extraction and max pooling for global aggregation (Guo et al., 2020) have become the popular design choices for various categories of newer backbones: 1) the pointwise MLP-based such as PointNet++ (Qi et al., 2017b); 2) the graph-based such as DGCNN (Xu et al., 2020b); 3) the convolution-based such as PointCNN (Li et al., 2018), PointConv (Wu et al., 2019) RS-CNN (Liu et al., 2019c) and KP-Conv (Thomas et al., 2019). Recently, methods are proposed for efficient deep learning on point clouds through novel data structuring (Xu et al., 2020b), faster sampling (Hu et al., 2020), adaptive filters (Xu et al., 2020a), efficient representation (Liu et al., 2019d) or convolution operation (Zhang et al., 2019b) . However, they still use expensive floating-point parameters and operations, which can be improved by binarization.
Methods
Binarized models operate on efficient binary parameters, but often suffer large performance drop. Moreover, the unique characteristics of point clouds pose even more challenges. We observe there are two main problems: first, aggregation of a large number of points leads to a severe loss of feature diversity; second, binarization induces an immense scale distortion, that undermines the functionality of scale-sensitive structures. In this section, we discuss our observations, and propose our BiPointNet with theoretical justifications.
We first give a brief introduction to our framework that binarizes a floating-point network. For example, deep learning models on point clouds typically contain multi-layer perceptrons (MLPs) for feature extraction. In contrast, the binarized models contain binary MLPs (BiMLPs), which are composed of binarized linear (bi-linear) layers. Bi-linear layers perform the extremely efficient bitwise operations (XNOR and Bitcount) on the lightweight binary weight/activation. Specifically, the activation of the bi-linear layer is binarized to , and is computed with the binarized weight to obtain the output :
where denotes the inner product for vectors with bitwise operations XNOR and Bitcount. When and denote the random variables in and , we represent their probability mass function as , and .
Similarly, when , and denote the random variables sampled from , and , we represent their probability mass function as , and .
2 Entropy-maximizing Aggregation
Unlike images pixels that are arranged in regular lattices, point clouds are sets of points without any specific order. Hence, features are usually processed in a point-wise manner and aggregated explicitly through pooling layers. Our study shows that the aggregation function is a performance bottleneck of the binarized model, due to severe homogenization as shown in Figure 2.
We apply information theory (Section 3.2.1) to quantify the effect of the loss of feature diversity, and find that global feature aggregation leads to a catastrophic loss of information entropy. In Section 3.2.2, we propose the concept of Entropy-Maximizing Aggregation (EMA) that gives the statistically maximum information entropy to effectively tackle the feature homogenization problem.
Ideally, the binarized tensor should reflect the information in the original tensor as much as possible. From the perspective of information, maximizing mutual information can maximize the information flow from the full-precision to the binarized parameters. Hence, our goal is equivalent to maximizing the mutual information of the random variables and :
where is the information entropy, and is the conditional entropy of given . as we use the deterministic sign function as the quantizer in binarization (see Section A.1 for details). Hence, the original objective function Eq. (4) is equivalent to:
where is the set of possible values of . We then study the information properties of max pooling, which is a common aggregation function used in popular point cloud learning models such as PointNet. Let the max pooling be the last layer of the multi-layer stacked , and the input of is defined as . The data flow of Eq. (3) can be further expressed as , and the information entropy of binarized feature can be expressed as
where is the number of elements aggregated by the max pooling, and is the random variable sampled from . The brief derivation of Eq (6) is shown in Appendix A.2. Theorem 1 shows the information properties of max pooling with the normal distribution input on the binarized network architecture.
For input of max pooling with arbitrary distribution, the information entropy of the binarized output goes to zero as goes to infinity, i.e., . And there exists a constant , for any and , if , we have , where is the number of elements to be aggregated.
The proof of Theorem 1 is included in Appendix A.2, which explains the severe feature homogenization after global feature pooling layers. As the number of points is typically large (e.g. 1024 points by convention in ModelNet40 classification task), it significantly reduces the information entropy of binarized feature , i.e., the information of is hardly retained in , leading to highly similar output features regardless of the input features to pooling layer as shown in Figure 2.
Furthermore, Theorem 1 provides a theoretical justification for the poor performance of existing binarization methods, transferred from 2D vision tasks to point cloud applications. In 2D vision, the aggregation functions are often used to gather local features with a small kernel size (e.g. in ResNet (He et al., 2016; Liu et al., 2018) and VGG-Net (Simonyan & Zisserman, 2014) which use pooling kernels). Hence, the feature homogenization problem on images is not as significant as that on point clouds.
2.2 EMA for Maximum Information Entropy
Therefore, we need a class of aggregation functions that maximize the information entropy of to avoid the aggregation-induced feature homogenization.
We study the correlation between the information entropy of binary random variable and the distribution of the full-precision random variable . We notice that the function used in binarization has a fixed threshold and decision levels, so we get Proposition 1 about information entropy of and the distribution of .
When the distribution of the random variable satisfies , the information entropy is maximized.
The proof of Proposition 1 is shown in Appendix A.3. Therefore, theoretically, there is a distribution of that can maximize the mutual information of and by maximizing the information entropy of the binary tensor , so as to maximally retain the information of in .
To maximize the information entropy , we propose the EMA for feature aggregation in BiPointNet. The EMA is not one, but a class of binarization-friendly aggregation layers. Modifying the aggregation function in the full-precision neural network to a EMA keeps the entropy maximized by input transformation. The definition of EMA is
where denotes the aggregation function (e.g. max pooling and average pooling) and denotes the transformation unit. Note that a standard normal distribution is assumed for because batch normalization layers are placed prior to the pooling layers by convention. can take many forms; we discover that a simple constant offset is already effective. The offset shifts the input so that the output distribution satisfies , to maximize the information entropy of binary feature . The transformation unit in our BiPointNet can be defined as .
When max pooling is applied as , we obtain the distribution offset for the input that maximizes the information entropy by solving the objective function
In a nutshell, we provide two possible variants of : first, we show that a simple shift is sufficient to turn a max pooling layer into an EMA (EMA-max); second, average pooling can be directly used (EMA-avg) without modification as a large number of points does not undermine its information entropy, making it adaptive to the dynamically changing number of point input. Note that modifying existing aggregation functions is only one way to achieve EMA; the theory also instructs the development of new binarization-friendly aggregation functions in the future.
3 One-scale-fits-all: Layer-wise Scale Recovery
In this section, we show that binarization leads to feature scale distortion and study its cause. We conclude that the distortion is directly related to the number of feature channels. More importantly, we discuss the detriments of scale distortion from the perspectives of the functionality of scale-sensitive structures and the optimization.
To address the severe scale distortion in feature due to binarization, we propose the Layer-wise Scale Recovery (LSR). In LSR, each bi-linear layer is added only one learnable scaling factor to recover the original scales of all binarized parameters, with negligible additional computational overhead and memory usage.
The scale of parameters is defined as the standard deviation of their distribution. As we mentioned in Section 3.2, balanced binarized weights are used in the bi-linear layer aiming to maximize the entropy of the output after binarization, i.e., and .
When we let and in bi-linear layer to maximize the mutual information, for the binarized weight and activation , the probability mass function for the distribution of output can be represented as . The output has approximately a normal distribution .
The proof of Theorem 2 is found in Appendix A.4. Theorem 2 shows that given the maximized information entropy, the scale of the output features is directly related to the number of feature channels. Hence, scale distortion is pervasive as a large number of channels is the design norm of deep learning neural networks for effective feature extraction.
We discuss two major impacts of the scale distortion on the performance of binarized point cloud learning models. First, the scale distortion invalidates structures designed for 3D deep learning that are sensitive to the scale of values. For example, the T-Net in PointNet is designed to predict an orthogonal transformation matrix for canonicalization of input and intermediate features (Qi et al., 2017a). The predicted matrix is regularized by minimizing the loss term . However, this regularization is ineffective for the with huge variance as shown in Figure 3.
Second, the scale distortion leads to a saturation of forward-propagated activations and backward-propagated gradients (Ding et al., 2019). In the binary neural networks, some modules (such as and ) rely on the Straight-Through Estimator (STE) Bengio et al. (2013) for feature binarization or feature balancing. When the scale of their input is amplified, the gradient is truncated instead of increased proportionally. Such saturation, as shown in Fig 4(c), hinders learning and even leads to divergence.
3.2 LSR for Output Scale Recovery
To recover the scale and adjustment ability of output, we propose the LSR for bi-linear layers in our BiPointNet. We design a learnable layer-wise scaling factor in our LSR. is initialized by the ratio of the standard deviations between the output of bi-linear and full-precision counterpart:
where denotes as the standard deviation. And the is learnable during the training process. The calculation and derivative process of the bi-linear layer with our LSR are as follows:
where and denotes the gradient and , respectively. By applying the LSR in BiPointNet, we mitigate the scale distortion of output caused by binarization.
Compared to existing methods, the advantages of LSR is summarized in two folds. First, LSR is efficient. It not only abandons the adjustment of input activations to avoid expensive inference time computation, but also recovers the scale of all weights parameters in a layer collectively instead of expensive restoration in a channel-wise manner (Rastegari et al., 2016). Second, LSR serves the purpose of scale recovery that we show is more effective than other adaptation such as minimizing quantization errors (Qin et al., 2020b; Liu et al., 2018).
Experiments
In this section, we conduct extensive experiments to validate the effectiveness of our proposed BiPointNet for efficient learning on point clouds. We first ablate our method and demonstrate the contributions of EMA and LSR on three most fundamental tasks: classification on ModelNet40 (Wu et al., 2015), part segmentation on ShapeNet (Chang et al., 2015), and semantic segmentation on S3DIS (Armeni et al., 2016). Moreover, we compare BiPointNet with existing binarization methods where our designs stand out. Besides, BiPointNet is put to the test on real-world devices with limited computational power and achieve extremely high speedup () and storage saving (). The details of the datasets and the implementations are included in the Appendix E.
As shown in Table 1, the binarization model baseline suffers a catastrophic performance drop in the classification task. EMA and LSR improve performance considerably when used alone, and they further close the gap between the binarized model and the full-precision counterpart when used together.
In Figure 4, we further validate the effectiveness of EMA and LSR. We show that BiPointNet with EMA has its information entropy maximized during training, whereas the vanilla binarized network with max pooling gives limited and highly fluctuating results. Also, we make use of the regularization loss for the feature transformation matrix of T-Net in PointNet as an indicator, the of the BiPointNet with LSR is much smaller than the vanilla binarized network, demonstrating LSR’s ability to reduce the scale distortion caused by binarization, allowing proper prediction of orthogonal transformation matrices.
Moreover, we also include the results of two challenging tasks, part segmentation, and semantic segmentation, in Table 1. As we follow the original PointNet design for segmentation, which concatenates pointwise features with max pooled global feature, segmentation suffers from the information loss caused by the aggregation function. EMA and LSR are proven to be effective: BiPointNet is approaching the full precision counterpart with only mIoU difference on part segmentation and mIoU gap on semantic segmentation. The full results of segmentation are presented in Appendix E.6.
2 Comparative Experiments
In Table 3, we show that our BiPointNet outperforms other binarization methods such as BNN (Hubara et al., 2016), XNOR (Rastegari et al., 2016), Bi-Real (Liu et al., 2018), ABC-Net (Lin et al., 2017), XNOR++ (Bulat & Tzimiropoulos, 2019), and IR-Net (Qin et al., 2020b). Although these methods have been proven effective in 2D vision, they are not readily transferable to point clouds due to aggregation-induced feature homogenization.
Even if we equip these methods with our EMA to mitigate information loss, our BiPointNet still performs better. We argue that existing approaches, albeit having many scaling factors, focus on minimizing quantization errors instead of recovering feature scales, which is critical to effective learning on point clouds. Hence, BiPointNet stands out with a negligible increase of parameters that are designed to restore feature scales. The detailed analysis of the performance of XNOR is found in Appendix C. Moreover, we highlight that our EMA and LSR are generic, and Table 3 shows improvements across several mainstream categories of point cloud deep learning models, including PointNet (Qi et al., 2017a), PointNet++ (Qi et al., 2017b), PointCNN (Li et al., 2018), DGCNN (Wang et al., 2019a), and PointConv (Wu et al., 2019).
3 Deployment Efficiency on Real-world Devices
To further validate the efficiency of BiPointNet when deployed into the real-world edge devices, we further implement our BiPointNet on Raspberry Pi 4B with 1.5 GHz 64-bit quad-core ARM CPU Cortex-A72 and Raspberry Pi 3B with 1.2 GHz 64-bit quad-core ARM CPU Cortex-A53.
We compare our BiPointNet with the PointNet in Figure 5(a) and Figure 5(b). We highlight that BiPointNet achieves inference speed increase and storage reduction over PointNet, which is recognized as a fast and lightweight model itself. Moreover, we implement various binarization methods over PointNet architecture and report their real speed performance on ARM A72 CPU device. As Figure 5(c), our BiPointNet surpasses all existing binarization methods in both speed and accuracy. Note that all binarization methods adopt our EMA and report their best accuracy, which is the important premise that they can be reasonably applied to binarize the PointNet.
Conclusion
We propose BiPointNet, the first binarization approach for efficient learning on point clouds. We build a theoretical foundation to study the impact of binarization on point cloud learning models, and proposed EMA and LSR in BiPointNet to improve the performance. BiPointNet outperforms existing binarization methods, and it is easily extendable to a wide range of tasks and backbones, giving an impressive speedup and storage saving on resource-constrained devices. Our work demonstrates the great potential of binarization. We hope our work can provide directions for future research.
Acknowledgement This work was supported by National Natural Science Foundation of China (62022009, 61872021), Beijing Nova Program of Science and Technology (Z191100001119050), and State Key Lab of Software Development Environment (SKLSDE-2020ZX-06).
References
Appendix A Main Proofs and Discussion
In our BiPointNet, we hope that the binarized tensor reflects the information in the original tensor as much as possible. From the perspective of information, our goal is equivalent to maximizing the mutual information of the random variables and :
where and , are the joint and marginal probability mass functions of these discrete variables. is the information entropy, and is the conditional entropy of given . According to the Eq. (15) and Eq. (18), the conditional entropy can be expressed as
Hence, the original objective function is equivalent to maximizing the information entropy :
A.2 Proofs of Theorem 1
For input of max pooling with arbitrary distribution, the information entropy of the binarized output to zero as to infinity, i.e., . And there is a constant , for any and , if , we have , where is the number of aggregated elements.
Proof. We obtain the correlation between the probability mass function of input and output of max pooling, intuitively, all values are negative to give a negative maximum value:
Since the function is applied as the quantizer, the of binarized feature can be expressed as Eq. (6).
(1) When obeys a arbitrary distribution, the probability mass function must satisfies . According to Eq. (6), let , we have
(2) For any , we can obtain the representation of the information entropy :
Let p_{n}=\Big{(}\sum_{x_{\phi}<0}p_{X_{\phi}}(x_{\phi})\Big{)}^{n}, the can be expressed as
and the derivative of is
the is maximized when takes 0.5, and is positive correlation with when since the when .
Therefore, when the constant satisfies p_{c}=\Big{(}\sum_{x_{\phi}<0}p_{X_{\phi}}(x_{\phi})\Big{)}^{c}\geq 0.5, given the , we have , and .
A.3 Proofs of Proposition 1
When the distribution of the random variable satisfies , the information entropy is maximized.
Then we can get the derivative of with respect to
When we let to maximize the , we have . Sine the deterministic function with the zero threshold is applied as the quantizer, the probability mass function of is represented as
and when the information entropy is maximized, we have
A.4 Discussion and Proofs of Theorem 2
The bi-linear layers are widely used in our BiPointNet to model each point independently, and each linear layer outputs an intermediate feature. The calculation of the bi-linear layer is represented as Eq. (2). Since the random variable is sampled from or obeying Bernoulli distribution, the probability mass function of can be represented as
where is the probability of taking the value +1. The distribution of output can be represented by the probability mass function of and .
In bi-linear layer, for the binarized weight and activation with probability mass function and , the probability mass function for the distribution of output can be represented as .
Proof. To simplify the notation in the following statements, we define and . Then, for each element in output , we have
Observe that is independent to and the value of both variables are either or . Therefore, the discrete probability distribution of can be defined as
Notice that can be parameterized as a binomial distribution. Then we have
Observe that obeys the same distribution as . Finally, we have
Proposition 2 shows that the output distribution of the bi-linear layer depends on the probability mass functions of binarized weight and activation. Then we present the proofs of Theorem 2.
When we let and in bi-linear layer to maximize the mutual information, for the binarized weight and activation , the probability mass function for the distribution of output can be represented as . The distribution of output is approximate normal distribution .
Proof. First, we prove that the distribution of can be approximated as a normal distribution. For bi-linear layers in our BiPointNet, all weights and activations are binarized, which can be represented as and , respectively. And the value of an element in can be expressed as
and the value of the element can be expressed as
The only can take from two values and its value can be considered as the result of one Bernoulli trial. Thus for the random variable sampled from the output tensor , the probability mass function, can be expressed as
where denotes the probability that the element takes . Note that the Eq. (45) is completely equivalent to the representation in Proposition 2. According to the De Moivre–Laplace theorem, the normal distribution can be used as an approximation of the binomial distribution under certain conditions, and the can be approximated as
and then, we can get the mean and variance of the approximated distribution with the help of equivalent representation of in Proposition 2. Now we give proof of this below.
According to Proposition 2, when , we can rewrite the equation as
Then we move to calculate the mean and standard variation of this distribution. The mean of this distribution is defined as
By the virtue of binomial coefficient, we have
Besides, when is an even number, we have . These equations prove the symmetry of function . Finally, we have
The standard variation of is defined as
To calculate the standard variation of , we use Binomial Theorem and have several identical equations:
These identical equations help simplify Eq. (57):
Now we proved that, the distribution of output is approximate normal distribution .
A.5 Discussion of the Optimal δ𝛿\delta for EMA-max
A.6 Discussion of the Optimal δ𝛿\delta for EMA-avg
Appendix B Implementation of BiPointNet on ARM Devices
We further implement our BiPointNet on Raspberry Pi 4B with 1.5 GHz 64-bit quad-core ARM Cortex-A72 and Raspberry Pi 3B with 1.2 GHz 64-bit quad-core ARM Cortex-A53, and test the real speed that one can obtain in practice. Although the PointNet is a recognized high-efficiency model, the inference speed of BiPointNet is much faster. Compared to PointNet, BiPointNet enjoys up to speedup and storage saving.
We utilize the SIMD instruction SSHL on ARM NEON to make inference framework daBNN (Zhang et al., 2019a) compatible with our BiPointNet and further optimize the implementation for more efficient inference.
B.2 Implementation Details
Figure 6 shows the detailed structures of six PointNet implementations. In Full-Precision version (a), BN is merged into the later fully connected layer for speedup, which is widely chosen for deployment in real-world applications. In Binarization version (b)(c)(d)(e), we have to keep BN unmerged due to the binarization of later layers. Instead, we merge the scaling factor of LSR into BN layers. The HardTanh function is removed because it does not affect the binarized value of input for the later layers. We test the quantization for the first layer and last layer in the variants (b)(c)(d)(e). In the last variant(f), we drop the BN layers during training. The scaling factor is ignored during deployment because it does not change the sign of the output.
B.3 Ablation analysis of time cost and quantization sensitivity
Table 4 shows the detailed configuration including overall accuracy, storage usage, and time cost of the above-mentioned six implementations. The result shows that binarization of the middle fully connected layers can extremely speed up the original model. We achieve storage saving, speedup on A72, and speed on A53. The quantization of the last layer further helps save storage consumption and improves the speed with a slight performance drop. However, the quantization of the first layer causes a drastic drop in accuracy without discernible computational cost reduction. The variant (f) without BN achieves comparable performance with variant (b). It suggests that our LSR method could be an ideal alternative to the original normalization layers to achieve a fully quantized model except for the first layer.
Appendix C Comparison between Layer-Wise Scale Recovery and other Methods
In this section, we will analyze the difference between the LSR method with other model binarization methods. Theorem 2 shows the significance of recovering scale in point cloud learning. However, IRNet and BiReal only consider the scale of weight but ignore the scale of input features. Therefore, these two methods cannot recover the scale of output due to scale distortion on the input feature. A major difference between these two methods is that LSR opts for layer-wise scaling factor while XNOR opts for point-wise one. Point-wise scale recovery needs dynamical computation during inference while our proposed LSR only has a layer-wise global scaling factor, which is independent of the input. As a result, our method can achieve higher speed in practice.
Table 3 shows that XNOR can alleviate the aggregation-induced feature homogenization. The point-wise scaling factor helps the model to achieve comparable adjustment capacity as full-precision linear layers. Therefore, although XNOR suffers from feature homogenization at the beginning of the training process, it can alleviate this problem with the progress of training and achieve acceptable performance, as shown in Figure 7.
Appendix D Comparison with Other Efficient Learning Methods
We compare our computation speedup and storage savings with several recently proposed methods to accelerate deep learning models on point clouds. Note that the comparison is for reference only; tests are conducted on different hardware, and for different tasks. Hence, direct comparison cannot give any meaningful conclusion. In Table 5, we show that BiPointNet achieves the most impressive acceleration.
Appendix E Experiments
ModelNet40: ModelNet40 (Wu et al., 2015) for part segmentation. The ModelNet40 dataset is the most frequently used dataset for shape classification. ModelNet is a popular benchmark for point cloud classification. It contains 12,311 CAD models from 40 representative classes of objects.
ShapeNet Parts: ShapeNet Parts (Chang et al., 2015) for part segmentation. ShapeNet contains 16,881 shapes from 16 categories, 2,048 points are sampled from each training shape. Each shape is split into two to five parts depending on the category, making up to 50 parts in total.
S3DIS: S3DIS for semantic segmentation (Armeni et al., 2016). S3DIS includes 3D scan point clouds for 6 indoor areas including 272 rooms in total, each point belongs to one of 13 semantic categories. We follow the official code (Qi et al., 2017a) for training and testing.
E.2 Implementation Details of BiPointNet
We follow the popular PyTorch implementation of PointNet and the recent geometric deep learning codebase (Fey & Lenssen, 2019) for the implementation of PointNet baselines. Our BiPointNet is built by binarizing the full-precision PointNet. All linear layers in PointNet except the first and last one are binarized to bi-linear layer, and we select as our activation function instead of ReLU when we binarize the activation before the bi-linear layer. For the part segmentation task, we follow the convention (Wu et al., 2014; Yi et al., 2016) to train a model for each of the 16 classes. We also provide our PointNet baseline under this setting.
Following previous works, we train 200 epochs, 250 epochs, 128 epochs on point cloud classification, part segmentation, semantic segmentation respectively. To stably train the binarized models, we use a learning rate of 0.001 with Adam and Cosine Annealing learning rate decay for all binarized models on all three tasks.
E.3 More Backbones
We also propose four other models: BiPointCNN, BiPointNet++, BiDGCCN, and BiPointConv, which are binarized versions of PointCNN (Li et al., 2018), PointNet++ (Qi et al., 2017b), DGCNN (Wang et al., 2019a), and PointConv (Wu et al., 2019), respectively. This is attributed to the fact that all these variants have characteristics in common, such as linear layers for point-wise feature extraction and global pooling layers for feature aggregation (except PointConv, which does not have explicit aggregators). In PointNet++, DGCNN, and PointConv, we keep the first layer and the last layer full-precision and binarize all the other layers. In PointCNN, we keep every first layer of XConv full precision and keep the last layer of the classifier full precision.
E.4 Binarization Methods
For comparison, we implement various representative binarization methods for 2D vision, including BNN (Hubara et al., 2016), XNOR-Net (Rastegari et al., 2016), Bi-Real Net (Liu et al., 2018), XNOR++ (Bulat & Tzimiropoulos, 2019), ABC-Net (Lin et al., 2017), and IR-Net (Qin et al., 2020b), to be applied on 3D point clouds. Note that the Case 1 version of XNOR++ is used in our experiments for a fair comparison, which applies layerwise learnable scaling factors to minimize the quantization error. These methods are implemented according to their open-source code or the description in their papers, and we take reference of their 3x3 convolution design when implementing the corresponding bi-linear layers. We follow their training process and hyperparameter settings, but note that the specific shortcut structure in Bi-Real and IR-Net is ignored since it only applies to the ResNet architecture.
E.5 Training Details
Our BiPointNet is trained from scratch (random initialization) without leveraging any pre-trained model. Amongst the experiments, we apply Adam as our optimizer and use the cosine annealing learning rate scheduler to stably optimize the networks. To evaluate our BiPointNet on various network architectures, we mostly follow the hyper-parameter settings of the original papers (Qi et al., 2017a; Li et al., 2018; Qi et al., 2017b; Wang et al., 2019a).
E.6 Detailed Results of Segmentation
We present the detailed results of part segmentation on ShapeNet Part in Table 6 and semantic segmentation on S3DIS in Table 7. The detailed results further prove the conclusion of Section 4.1 as EMA and LSR improve performance considerably in most of the categories (instead of huge performance in only a few categories). This validates the effectiveness and robustness of our method.