PointGrow: Autoregressively Learned Point Cloud Generation with Self-Attention

Yongbin Sun, Yue Wang, Ziwei Liu, Joshua E. Siegel, Sanjay E. Sarma

Introduction

Recently 3D generative model has attracted enormous research interests because it directly promotes the development of emerging applications, such as virtual/augmented reality and self-driving cars. For example, it is capable of completing the LIDAR scans that might suffer from occlusion issues . Therefore, intelligent systems that are able to automatically generate realistic and diverse 3D shapes are highly desired.

3D shapes are usually represented as triangle meshes or point clouds due to their light-weight nature and simple form. Such shape representations are flexible for rendering, but impose a problem when applying computer vision techniques for processing, because most standard operations are designed based on regular grid-based formats (e.g. image), while meshes and point clouds are fundamentally irregular: vertex or point positions are continuously distributed in the space, and any permutation of their face or point ordering does not change the spatial distribution. Therefore, one major line of 3D research discretizes a continuous shape representation onto a 3D grid . But such volume-based representations consume significant memory and introduce quantization artifacts, making it difficult to generate high-res 3D shapes and retain fine-grained surface details. Therefore, it is more desirable to design a framework specific to raw shape representations, rather than relying upon intermediate representations. In this paper, we choose to represent shapes using point clouds, because they comprise the output of most existing 3D sensing technologies, showing more application values but with less complexity.

The vast majority of existing learning-based works for 3D point clouds generation rely on two types of distance metric between point sets, Chamfer Distance (CD) and Earth Mover’s Distance (EMD) , to handle the irregularity problem of point clouds. Serving as the loss function, these distance metric penalizes the dissimilarity between the generated and ground truth point sets, and help deep generative models learn to produce visually-plausible 3D shapes for various types of inputs . However, the generative process is hard to interpret due to the intrinsic limitation of inter-set distance metric.

In this work, we seek to explore a different approach to better understand and interpret the point cloud generative process. We begin by observing that the points constituting a 3D shape have correlations. For example, most man-made shapes are symmetric, such as the four legs of a table; also, points are distributed in a correlated way to form different parts of a shape (e.g. to produce an airplane, some points form its wings, while others have to form its body structure). To explicitly learn and utilize such inter-point correlations for shape generation, we adopt a probabilistic approach to jointly model the spatial distribution of all the points of a shape in an nn dimensional space, where nn is the number of points in a point cloud. Underpinned by a joint point distribution reflecting the underlying inter-point correlation in the data, a high joint probability value should correspond to the point set distribution of a plausible 3D shape, while a low value should indicate an implausible one. In this way, the point cloud generation process can be cast as a sampling process in an nn dimensional point space.

Furthermore, since a joint probability can be decomposed by chain rule as the product of a series of conditional probabilities, the joint point set probability can be expressed by a series of conditional point probabilities, where each point is conditioned on its previously generated ones. This property naturally enables us to visualize the shape generative process and interpret inter-point correlations in a point-by-point manner during the point sampling process, as shown in Figure 1.

To this end, we propose an autoregressive framework dubbed PointGrow to generate every point recurrently. Specifically, PointGrow estimates a conditional distribution of the point under consideration given all its preceding points. However, the irregularity of point clouds imposes difficulties when aggregating meaningful information from a given point set, especially when such information is contained by distant points. Therefore, we further propose two point cloud-based self-attention modules to dynamically aggregate long-range dependencies from available points. Our experiments show that those two modules can improve information flow between points, and successfully capture meaningful semantic information.

The contributions of our work are summarized as below:

We propose a novel autogressive model, PointGrow, for point cloud generation, which models the joint 3D spatial distribution in a point-by-point manner. PointGrow has two appealing properties: 1) it is capable of generating diverse and realistic 3D clouds, 2) it constitutes an interpretable 3D shape generative process.

Two self-attention modules are carefully designed to capture long-range dependencies and semantic correlations between points, facilitating the generation of plausible part configurations within 3D objects.

Besides point cloud generation, our framework also enables several important applications, such as diverse shape completions, unsupervised feature learning and shape arithmetic operations.

Related Work

While 3D data processing and generation has a long history, here, we only discuss the directly related work of using deep networks to analyze 3D shapes, autoregressive networks, and self-attention.

Volumetric Methods. 3D shape recognition and generation has been studied using 3D voxel grids . Voxelization often produces a sparsely-occupied 3D grid, which limits resolution and introduces quantization artifacts. Recent frameworks have been proposed to reduce spatial complexity of volumetric shape representations, , though these generally suffer from high computation costs.

Mesh-Based Methods. 3D meshes are a lightweight approach for geometric modelling via a set of vertices and triangular or quad primitives. Recent work has extended standard convolutions to mesh surfaces for aggregating and propagating local features . Relevant work reconstructing 3D shapes as 3D meshes is found in .

Point Cloud-Based Methods. PointNet is the pioneering work in applying deep neural nets to point sets, using a symmetric function to aggregate feature vectors for all points in a permutation-invariant manner. PointNet’s successors explore ways to accumulate local information in the spatial and embedded feature domains to achieve high performance. To address point cloud generative tasks, introduced two symmetric distance metrics, CD and EMD, to measure the distance between two point sets. These metrics are order-invariant, which makes them suitable as loss function operated directly on point clouds. By taking advantage of these metrics, models have been proposed to address point cloud synthesis problems under different settings . However, existing generative approaches focus on measuring inter-set dissimilarity, and the inter-point relationship within a point set is not well understood.

2 Autoregressive Networks

Autoregressive networks model current values as a function of their own previous values, and have been adopted to model the joint distribution of image pixels and audio samples . The joint distribution is cast as a product of conditional distributions, and each condition distribution is modeled using a deep neural network that takes as input previously generated values and outputs a distribution for the value currently under consideration. But it is not trivial to adapt autoregressive frame to point cloud due to the irregularity problem of point clouds.

3 Self-Attention

Attention is a flexible mechanism to capture information in a self-adaptive manner such that accumulated important information is weighted highly. It improves performance in tasks including image recognition and natural language processing . Recently, self-attention has been adopted into generative tasks, such as image generation . In our experiments, we demonstrate that self-attention modules can be extended to process unordered point sets and capture inter-point correlations.

PointGrow

This section introduces the formulation and implementation of PointGrow (and its conditional version), a point-by-point generative model for 3D point clouds.

Unconditional PointGrow. A point cloud, S, that consists of nn points is defined as S={s1,s2,...,sn}\textbf{S}=\{\textbf{s}_{1},\textbf{s}_{2},...,\textbf{s}_{n}\}, with its ithi^{th} point si={xi,yi,zi}\textbf{s}_{i}=\{x_{i},y_{i},z_{i}\} in 3D space. Our goal is to assign a probability p(S)p(\textbf{S}) to each point cloud. We do so by factorizing the joint probability of S as a product of conditional probabilities over all its points:

The value p(si∣s≤i−1)p(\textbf{s}_{i}|\textbf{s}_{\leq i-1}) is the conditional probability of the ithi^{th} point si\textbf{s}_{i} given all its previously generated points, and computed as a joint probability over its coordinates:

where each coordinate is conditioned on available coordinates. To facilitate the generation process, we sort training points according to their zz coordinates to encourage a shape to be generated mainly along its primary axis during testing (like 3D printing), hoping that semantic information can be better captured to produce consistent shapes (e.g. a rear car body to be generated has to match an existing front). But since our model is sampling-based, the generated coordinates are not strictly larger than their previous ones and processing modules (introduced later) should be invariant to point permutation. Here, we model the conditional probability distribution of each coordinate using a deep neural network. Prior art shows that a softmax discrete distribution is more flexible than a continuous one to model any arbitrary distribution. So we discrete point coordinates by scaling them into the range and quantizing them to dd uniformly distributed values. Note that different from many existing voxel-based methods generating tensors of size d3d^{3} and operating on 3D volumes, our model outputs tensors of 3×n×d3\times n\times d (usually much smaller than d3d^{3} considering the resolution of shape volumes) and operates directly on sparse point representation. We set n=1024n=1024 and d=200d=200 as a trade-off between generative performance and quantization artifacts. Larger nn and dd can be used to achieve better visual results but at the cost of slower performance.

Context Awareness Operation. Context awareness improves model inference. For example, in and , a global feature is obtained by applying max pooling along each feature dimension, and then used to provide context information for solving semantic segmentation tasks. Similarly, we obtain context-aware features for all sets of generated available points in the point cloud generation process, as illustrated in Figure 2 (bottom left). Each row of resultant context-aware features aggregates the context information of all the previously generated points dynamically by fetching and averaging. This Context Awareness (CA) operation is implemented as a plug-in module in our model, and mean pooling is used in our experiments.

Self-Attention Context Awareness Operation. The CA operation accumulates point features in a fixed manner via pooling. Improving this, we propose two learning-based operations to determine the weights for aggregating point features self-adaptively. We define these as Self-Attention Context Awareness (SACA) operations, and the weights as self-attention weights.

Figure 2 shows the first SACA operation, SACA-A. To allow each input point feature to understand its importance in context and later determine its weight, we associate it with its context-aware feature obtained after a CA module. The combined feature vector is then passed into a Multi-Layer Perception (MLP) to learn self-attention weights. Given an n×fn\times f point feature matrix, F, with its ithi^{th} row, fi\textbf{f}_{i}, representing the feature vector of the ithi^{th} point, we compute the ithi^{th} self-attention weight vector, wi\textbf{w}_{i}, as below:

where Mean{⋅}Mean\{\cdot\} is mean pooling, ⊕\oplus is concatenation, and MLP(⋅)MLP(\cdot) is a sequence of fully connected layers. The self-attention weights encode information about context changes due to each newly generated point, and are unique to that point. Next, we conduct element-wise multiplication between input point features and self-attention weights to obtain weighted features, which are then accumulated sequentially to generate corresponding context features. The process to calculate the ithi^{th} context feature, ci\textbf{c}_{i}, is summarized as:

where ⊗\otimes is element-wise multiplication. Finally, we shift context features downward by one row, because for the ithi^{th} point, si\textbf{s}_{i}, only its previous points, s≤i−1\textbf{s}_{\leq i-1}, are available. A zero vector of the same size is attached to the beginning as the initial context feature, indicating no a-priori context knowledge is available.

Figure 2 also shows the other SACA operation, SACA-B. SACA-B differs from SACA-A in the way to compute and apply self-attention weights. In SACA-B, the ithi^{th} context-aware feature after CA operation is shared by all the first ii point features to obtain self-attention weights, which are used to compute ci\textbf{c}_{i}. This process is described as:

Compared to SACA-A, SACA-B self-attention weights encode the importance of each point feature under a common context. Their differences are highlighted in Eq. (4) and (5).

In Figure 3, we visualize the attention fields during generative processes by visualizing Euclidean distances between the context feature of a query point and the point features before the SACA operation of its previously generated points.

Model Architecture. Figure 2 top shows the proposed network to output conditional coordinate distributions. The top, middle and bottom branches model p(zi∣s≤i−1)p(z_{i}|\textbf{s}_{\leq i-1}), p(yi∣s≤i−1,zi)p(y_{i}|\textbf{s}_{\leq i-1},z_{i}) and p(xi∣s≤i−1,zi,yi)p(x_{i}|\textbf{s}_{\leq i-1},z_{i},y_{i}), respectively. Note that the input points in the latter two cases are masked so that the network cannot see information not-yet-generated. During training, points are available to compute all the context features, thus coordinate distributions can be estimated in parallel. During the generative phase, the point coordinates are sampled according to the estimated softmax probability distributions. This occurs sequentially, since each sampled coordinate needs to be fed as input back into the network, as shown in Figure 1. Our proposed autoregressive architecture models point coordinate distribution categorically, thus a simple cross-entropy loss is sufficient to handle the shape learning process efficiently without requiring the use of a computationally-intensive set-to-set distance loss function.

Conditional PointGrow. Given a condition or embedding vector, h, we intend to generate a shape satisfying the latent meaning of h. To achieve this, Eq. (1) and (2) are adapted to Eq. (6) and (7), respectively, as below:

The additional condition, h, affects the coordinate distributions by adding biases and potential constrains in the generative process. We implement this by changing the operation between adjacent fully-connected layers from xi+1=f(Wxi)\textbf{x}^{i+1}=f(\textbf{W}\textbf{x}^{i}) to xi+1=f(Wxi+Hh)\textbf{x}^{i+1}=f(\textbf{W}\textbf{x}^{i}+\textbf{H}\textbf{h}), where xi+1\textbf{x}^{i+1} and xi\textbf{x}^{i} are feature vectors in the i+1thi+1^{th} and ithi^{th} layer, respectively, W is a weight matrix, H is a matrix that transforms h into a vector with the same dimension as Wxi\textbf{W}\textbf{x}^{i}, and f(⋅)f(\cdot) is a nonlinear activation function. In this paper, we experimented with h as an one-hot categorical vector which adds class dependent bias, and an high-dimensional embedding vector of a 2D image which imposes geometric constraints.

Experiments

Datasets. We evaluated our framework on the ShapeNet CAD dataset. We used a subset consisting of 17,687 models across 7 categories: airplanes, cars, tables, chairs, benches, cabinets and lamps. To generate corresponding point clouds, we sampled 10,000 points uniformly from each mesh file, and then used farthest point sampling to select 1,024 points representing the shape. Each category follows a split ratio 0.9/0.1 to separate training from testing sets. ModelNet40 and PASCAL3D+ are used for additional analysis and demonstration.

We first evaluate unconditional PointGrow by addressing the following questions: (1) Can the model generate visually plausible 3D shapes in an interpretable manner? (2) Can the model generate diverse shapes? (3) Is the proposed self-attention module important for shape generation? (4) Does the model learn meaningful feature representation?

(1) Generative Process Interpretability and Shape Quality. We start evaluation by showing qualitative results of generated point clouds of unconditional PointGrow, shown in Figure 3. Fine-detailed structures can be observed from the generated shapes, such as the jet engine of airplanes and legs of furniture (e.g. table and chair). In the bottom part of Figure 3, we additionally show attention fields when generating corresponding query points in the process. It can be observed that the model focuses on different regions when generating points for different parts, and the focused areas are usually semantically related no matter their spatial distances. For example, when generating airplane wing points, the model aggregates structural knowledge from available wing part points; when generating points close to table legs, other previously generated leg points contribute most; when generating torchiere shade points for the lamp, the model considers both existing torchiere shade and base areas for a proper structural match. We also show a sampled generative process for an airplane in Figure 1. Note that the model produces different conditional distribution for points at different parts (e.g. in the second row of Figure 1, the network model outputs a roughly symmetric distribution along the XX axis, describing the airplane’s wings). From the above investigation, the interpretability of the proposed generative process is demonstrated. In this experiment, we train our model for each category separately, because an unconditional model lacks knowledge about the target shape when sampling from scratch. In later experiments, we show that when a categorical condition is given, it is possible to train the model across multiple shape categories and generate plausible shapes.

Next, we quantitatively evaluate the quality of generated shapes. The negative log-likelihood is commonly used to evaluate autoregressive models for image and audio generation . However, we observed inconsistency between this value and the visual quality of generated 3D shapes. This is validated by comparing two baseline models: CA-Mean and CA-Max, where the SACA operation is replaced with the CA operation implemented by mean and max pooling, respectively. In Figure 5, we report negative log-likelihoods in bits per coordinate on ShapeNet testing sets of airplane and car categories, and visualize their representative results. Despite CA-Max shows lower negative log-likelihood values, it gives less visually plausible results (i.e. airplanes lose wings and cars lose rear ends).

To actually evaluate the generated shape quality, we argue that if generated 3D shapes contain consistent semantic features as real shapes, a classification model trained on real shapes should perform well on generated ones, and vice versa. Therefore, after training on ShapeNet sets, we generate 300 point clouds per category (2,100 in total for 7 categories), and conduct two classification tasks: one training on original ShapeNet training sets and testing on generated shapes, the other training on generated shapes and testing on original ShapeNet testing sets. Here, PointNet , a widely-uesd model, is used as the point cloud shape classifier. We implement another two GAN-based competing methods and report classification results in Table 1, together with model complexity using number of model parameters. In the first classification task, our SACA-A model outperforms existing models by a relatively large margin, while in the second task, SACA-A and SACA-B models show similar performance.

(2) Shape Diversity. To demonstrate PointGrow can generate diverse shapes, we conducted a shape completion task. Given an initial set of points, our model is capable of completing shapes in multiple ways. Figure 4 visualizes examples. The input points are sampled from ShapeNet testing sets, which are not seen during the training process. The shapes generated by our model are different from the original ground truth point clouds, but still look plausible. A current limitation of our model is that it works only when the input point set is given as the beginning part of a shape along its primary axis, and in future work we will investigate how to complete shapes when partial point clouds are given from any directions.

(3) Ablation Study on Self-Attention Module. We conduct an ablation study to investigate the importance of the Self-Attention module in shape generation process. To quantitatively evaluate generated shapes and inspired by Fréchet Inception Distance (FID) , a metric widely used to measure the visual quality of generated images, we provide a similar metric, “PointNet Distance” (PND), to measure the geometric quality of generated point clouds. PND follows the same assumption and formula as FID, except that the features from a point cloud is obtained from the global feature of a PointNet model pre-trained on ModelNet40 with a global feature dimension of 128. Three model architectures are considered: without the second CA module (No CA), CA-Mean baseline, and SACA-A. We generate 200 point clouds per category for each model architecture. The PND is computed for each category and reported in Table 2. We observe a large performance improvement by gathering context information with self-attention supported.

(4) Unsupervised Feature Learning. To prove the model actually learns meaningful representations, we extract learned features and use them for classification tasks. We obtain the feature vector of a shape by applying different types of “symmetric” functions as illustrated in (i.e. min, max and mean pooling) on features of each layer before the SACA operation, and concatenate them all. Following , we pre-train our model on 7 categories from the ShapeNet dataset, and then use this model to extract feature vectors for both training and testing shapes from the ModelNet40 dataset. We experimented with both linear SVM and single layer classifiers, following the same settings as and , respectively. We report our best results in Table 3. The SACA-A model achieves the best performance using SVM classifier, and performs slightly worse than MTN using single layer classifier.

2 Conditional Point Cloud Generation

To evaluate conditional PointGrow, we answer the following questions: (1) Can the model be trained across multiple categories? (2) Can the model generate plausible shapes for image conditions?

(1) Conditioned on Category Label. We first experiment with category-conditional modelling of point clouds, given an one-hot vector h with its nonzero element hih_{i} indicating the ithi^{th} shape category. The one-hot condition provides categorical knowledge to guide the shape generation process, enabling the model to be trained across multiple categories. Figure 6 shows generated shape examples. Failure cases are also observed: generated shapes present interwoven geometric properties from other shape types. For example, the airplane misses wings and generates a car-like body; the lamp and the car develop chair leg structures.

(2) Conditioned on 2D Image. Next, we experiment with image conditions for point cloud generation. Image conditions add additional constrains to the point cloud generation process such that the geometric structures of sampled shapes match their 2D projections. In our experiments, we obtain an image condition vector through an image encoder, and optimize it together with the rest model components from scratch. The model is trained on synthetic ShapeNet dataset, and one out of 24 views of a shape (provided by ) is selected as the image condition input. The trained model is also tested on foreground objects of real images from the PASCAL3D+ dataset to prove its generalizability. The PASCAL3D+ dataset is challenging because the images are captured in real environments. Testing examples are shown on Figure 7 upper left.

We quantitatively evaluate the conditional generation results in terms of mean Intersection-over-Union (mIoU) and point-wise 3D Euclidean distance . Here we obtain ground truth points as voxel centers of 3D volumes from , and only consider shapes containing more than 500 occupied voxels, with 500 of them uniformly sampled to describe the shape. To compensate for the sampling randomness of PointGrow, we align generated points to their nearest voxels within a neighborhood of 2-voxel radius. When calculating surface-to-surface 3D Euclidean distance metric, we further remove interior points with 26 non-empty neighbors for fair comparison when selecting ground truth points. As shown in Table 4 and Table 5, PointGrow achieves above-par performance on conditional 3D shape generation.

Further, we demonstrate that intermediate 3D shapes can be generated from linearly interpolated embedding vectors of image pairs (Figure 7 bottom), and compositive shapes can be generated by applying arithmetic on embedded image condition vectors (Figure 7 upper right).

Discussion and Conclusion

This work studies the problem of point cloud generation. Unlike previous work, which minimizes set-to-set distances for generative learning, our model builds upon an autoregressive architecture and exploits point-to-point relations during generation, allowing the generative process to be better understood and interpreted. To further capture long-range dependencies in a self-adaptive manner and address the irregularity of point cloud data, two self-attention models are integrated within our framework. Extensive experiments validates the efficacy of this approach across a wide range of tasks.

Though our model generates visually-plausible 3D shapes, it faces two potential limitations. Firstly, due to the iterative property intrinsic to autoregressive models, the model scales poorly when generating large point sets. Recent work has accelerated autoregression-based audio generation. Similar techniques are also applicable here (e.g. generating clouds in a hierarchical rather than sequential manner). Secondly, our model only generates a point cloud along its primary axis as determined in training. This does not hinder generation performance if the shape is sampled from scratch, but limits its applicability to applications like shape completion. Generating clouds more flexibly will be an important topic for further research.

References