3DViewGraph: Learning Global Features for 3D Shapes from A Graph of Unordered Views with Attention

Zhizhong Han, Xiyang Wang, Chi-Man Vong, Yu-Shen Liu, Matthias Zwicker, C. L. Philip Chen

Introduction

Global features of 3D shapes can be learned from raw 3D representations, such as meshes, voxels, and point clouds. As an alternative, a number of works in 3D shape analysis employed multiple views Su and others (2015); Han et al. (2019b) as raw 3D representation, exploiting the advantage that multiple views can facilitate understanding of both manifold and non-manifold 3D shapes via computer vision techniques. Therefore, effectively and efficiently aggregating comprehensive information over multiple views, is critical for the discriminability of learned features, especially in deep learning models.

Pooling was designed as a procedure for information abstraction in deep learning models. In order to describe a 3D shape by considering features from multiple views, view aggregation is usually performed by max or mean pooling, where pooling only employs the max or mean value of each dimension across all view features Su and others (2015). Although pooling is able to eliminate the rotation effect of 3D shapes, both the content information within views and the spatial relationship among views cannot be fully preserved. As a consequence, this limits the discriminability of learned features. In this work, we address the challenge to learn 3D features in a deep learning model by more effectively aggregating the content information within individual views, and the spatial relationship among multiple unordered views.

To tackle this issue, we propose a novel deep learning model called 3D View Graph (3DViewGraph), which learns 3D global features from multiple unordered views. By taking multiple views around a 3D shape on a unit sphere, we represent the shape as a view graph formed by the views, where each view denotes a node, and the nodes are fully connected with each other by edges. 3DViewGraph learns highly discriminative global 3D shape features by simultaneously encoding both the content information within the view nodes, and the spatial relationship among the view nodes.

We propose a novel deep learning model called 3DViewGraph for 3D global feature learning by effectively aggregating multiple unordered views. It not only encodes the content information within all views, but also preserves the spatial relationship among the views.

We propose an approach to learn a low-dimensional latent semantic embedding of the views by directly capturing the similarities between each view and a set of latent semantic patterns. As an advantage, 3DViewGraph avoids mining the latent semantic patterns across the whole training set explicitly.

We perform view aggregation by integrating a novel spatial pattern correlation, which encodes the content information and the spatial relationship in each pair of views.

We propose a novel attention mechanism to increase the discriminability of learned features by highlighting the unordered view nodes with distinctive characteristics and depressing the ones with appearance ambiguities.

Related work

Deep learning models have made a big progress on learning 3D shape features from different raw representations, such as meshes Han and others (2018), voxels Wu and others (2016), point clouds Qi and others (2017) and views Su and others (2015). Because of page limit, we focus on reviewing view-based deep learning models to highlight the novelty of our view aggregation.

View-based methods. View-based methods represent a 3D shape as a set of rendered views Kanezaki et al. (2018) or panorama views Sfikas and others (2017). Besides direct set-to-set comparison Bai and others (2017), pooling is the widely used way of aggregating multiple views in deep learning models Su and others (2015). In addition to global feature learning, pooling can also be used to learn local features Huang et al. (2017); Yu et al. (2018) for segmentation or correspondence by aggregating local patches.

Although pooling can aggregate views on the fly in the models, it can not encode all the content information within views and the spatial relationship among views. Thus, the strategies of concatenation Savva and others (2016), view pair weighting Johns et al. (2016), cluster specified pooling Wang and others (2017), RNN Han and others (2019), were employed to resolve this issue. However, these methods can not learn from unordered views or fully capture the spatial information among unordered views.

To resolve the aforementioned issues, 3DViewGraph aggregates unordered views more effectively by simultaneously encoding their content information and spatial relationship.

Graph-based methods. To handle the irregular structure of graphs, various methods have been proposed Hamilton and others (2017). Although we formulate the multiple views from a 3D shape as a view graph, existing methods proposed for graphs cannot be directly used for learning the 3D feature in our scenario. The reasons are two-fold. First, these methods mainly focus on how to locally learn meaningful representation for each node in a graph from its raw attributes rather than globally learning the feature of the whole graph. Second, these methods mainly learns how to process the nodes in a graph with firm order, while the order of views involved in 3DViewGraph are always ambiguous because of the rotation of 3D shapes.

Moreover, some methods have employed graphs to retrieve 3D shapes from multiple views Anan et al. (2015); An-An et al. (2016). Different from these methods, 3DViewGraph employs a more efficient way of view aggregation in deep learning models, which makes the learned features useful for both classification and retrieval.

3DViewGraph

We first take a set of unordered views vi={vji∣j∈[1,V]}\bm{v}^{i}=\{v_{j}^{i}|j\in[1,V]\} on a unit sphere centered at mim^{i}, as shown in Fig. 1(a). Here, we use “unordered views” to indicate that the views cannot be organized in a sequential way. The views vjiv_{j}^{i} are regarded as view nodes DjiD_{j}^{i} (briefly shown by symbols) of an undirected graph GiG^{i}, where each DjiD_{j}^{i} is fully connected with other view nodes Dj′iD_{j^{\prime}}^{i} by edges Ej,j′iE_{j,j^{\prime}}^{i}, such that Gi=({Dji},{Ej,j′i})G^{i}=(\{D_{j}^{i}\},\{E_{j,j^{\prime}}^{i}\}).

To resolve the effect of rotation, 3DViewGraph encodes the content and spatial information of GiG^{i} by exhaustively computing our novel spatial pattern correlation between each pair of view nodes. As illustrated in Fig. 1(d), we compute the pattern correlation cj,j′i\bm{c}_{j,j^{\prime}}^{i} between DjiD_{j}^{i} and each other node Dj′iD_{j^{\prime}}^{i}, and we weight it with their spatial similarity sj,j′is_{j,j^{\prime}}^{i}. In addition, for each node DjiD_{j}^{i}, we compute its cumulative correlation Cji\bm{C}_{j}^{i} to summarize all spatial pattern correlations as the characteristics of the 3D shape from the jj-th view node DjiD_{j}^{i}.

Finally, we obtain the global feature Fi\bm{F}^{i} of shape mim^{i} by integrating all cumulative correlations Cji\bm{C}_{j}^{i} with our novel attention weights αi\bm{\alpha}^{i}, as shown in Fig. 1(e) and (f). αi\bm{\alpha}^{i} aims to highlight the view nodes with distinctive characteristics while depressing the ones with appearance ambiguity.

Latent semantic mapping learning. To learn global features from unordered views, 3DViewGraph encodes the content information within all views and the spatial relationship among views in a pairwise way. 3DViewGraph relies on the intuition that correlations between pairs of views can effectively represent discriminative characteristics of a 3D shape, especially considering the relative spatial position of the views. To implement this intuition, each view should be encoded in terms of a small set of common elements across all views in the training set. Unfortunately, the low-level features fji\bm{f}_{j}^{i} are too high dimensional and not suitable as a representation of the views in terms of a set of common elements.

where the similarity K(fji,ϕn)K(\bm{f}_{j}^{i},\bm{\phi}_{n}) is inversely proportional to the distance between fji\bm{f}_{j}^{i} and ϕn\bm{\phi}_{n} through exp()exp(), and gets normalized across the similarities between fji\bm{f}_{j}^{i} and all ϕn\bm{\phi}_{n}. Parameter β\beta controls the decay of the response with the distance. This equation can be further simplified by cancelling the norm of fji\bm{f}_{j}^{i} from the numerator and the denominator as follows,

Based on the last line in Eq. 2, we implement the latent semantic mapping Φ\Phi as a row-wise convolution with each pair of {ωn}\{\bm{\omega}_{n}\} and {εn}\{\varepsilon_{n}\} corresponding to a filter and a row-wise softmax normalization, as shown in Fig. 2.

Spatial pattern correlation. The pattern correlation cj,j′i\bm{c}_{j,j^{\prime}}^{i} aims to encode the content of view nodes DjiD_{j}^{i} and Dj′iD_{j^{\prime}}^{i}. cj,j′i\bm{c}_{j,j^{\prime}}^{i} makes the semantic patterns that co-occur in both views more prominent while the non-co-occurring ones more subtle. More precisely, we use the latent semantic embeddings dji\bm{d}_{j}^{i} and dj′i\bm{d}_{j^{\prime}}^{i} to compute cj,j′i\bm{c}_{j,j^{\prime}}^{i} as follows,

where cj,j′i\bm{c}_{j,j^{\prime}}^{i} is a N×NN\times N dimensional matrix whose entry cj,j′i(n,n′)\bm{c}_{j,j^{\prime}}^{i}(n,n^{\prime}) measures the correlation between the semantic pattern ϕn\bm{\phi}_{n} contributing to dji\bm{d}_{j}^{i} and ϕn′\bm{\phi}_{n^{\prime}} contributing to dj′i\bm{d}_{j^{\prime}}^{i}.

We further enhance the pattern correlation cj,j′i\bm{c}_{j,j^{\prime}}^{i} between the view nodes DjiD_{j}^{i} and Dj′iD_{j^{\prime}}^{i} by their spatial similarity sj,j′is_{j,j^{\prime}}^{i}, which forms the spatial pattern correlation sj,j′icj,j′is_{j,j^{\prime}}^{i}\bm{c}_{j,j^{\prime}}^{i}.

Fig. 3 visualizes how we compute the spatial similarity sj,j′is_{j,j^{\prime}}^{i}. In Fig. 3(a), we show all edges Ej,j′iE_{j,j^{\prime}}^{i} connecting DjiD_{j}^{i} to all other view nodes Dj′iD_{j^{\prime}}^{i} in different colors, where DjiD_{j}^{i} is briefly shown by symbols. The length of Ej,j′iE_{j,j^{\prime}}^{i} is measured by the length of the shortest arc connecting the two view nodes DjiD_{j}^{i} and Dj′iD_{j^{\prime}}^{i} on the unit sphere. Thus, Ej,j′i=2π×1×(θ/2π)=θE_{j,j^{\prime}}^{i}=2\pi\times 1\times(\theta/2\pi)=\theta as illustrated in Fig. 3(b), where θ\theta is the central angle of the arc and the factor 11 corresponds to the radius of the unit sphere. To reduce the high variance of {Ej,j′i}\{E_{j,j^{\prime}}^{i}\}, we employ Ej,j′i=0.5(1−cos⁡θ)E_{j,j^{\prime}}^{i}=0.5(1-\cos\theta) instead of Ej,j′i=θE_{j,j^{\prime}}^{i}=\theta, which normalizes Ej,j′iE_{j,j^{\prime}}^{i} into the range of $.Finally,. Finally,s_{j,j^{\prime}}^{i}isinverselyproportionaltois inversely proportional toE_{j,j^{\prime}}^{i}$ as follows,

where σ\sigma is a parameter to control the decay of the response with the edge length. In Fig. 3(c), we visualize sj,j′is_{j,j^{\prime}}^{i} by mapping the value of sj,j′is_{j,j^{\prime}}^{i} to the width of edges Ej,j′iE^{i}_{j,j^{\prime}}.

To represent the characteristics of 3D shape mim^{i} from the jj-th view node DjiD_{j}^{i} on GiG^{i}, we finally introduce the cumulative correlation Cji\bm{C}_{j}^{i}, which encodes all spatial pattern correlations starting from DjiD_{j}^{i} as follows,

Attentioned correlation aggregation. Intuitively, more views will provide more information to any deep learning model, which should allow it to produce more discriminative 3D features. However, additional views may also introduce appearance ambiguities that negatively affect the discriminability of learned features, as shown in Fig. 4.

To resolve this issue, 3DViewGraph employs a novel attention mechanism in the aggregation of the 3D shape characteristics from all unordered view nodes of a shape, as illustrated in Fig. 1(e). 3DViewGraph learns attention weights αi={αji∣j∈[1,V]}\bm{\alpha}^{i}=\{\alpha_{j}^{i}|j\in[1,V]\} for all view nodes DjiD_{j}^{i} on GiG^{i}, where αji\alpha_{j}^{i} would be a large value (the second row in Fig. 4) if the view vjiv_{j}^{i} has distinctive characteristics, while αji\alpha_{j}^{i} would be a small value (the first row in Fig. 4) if vjiv_{j}^{i} exhibits appearance ambiguity with views from other shapes. Note that ∑j=1Vαji=1\sum_{j=1}^{V}\alpha_{j}^{i}=1.

Our novel attention mechanism evaluates how distinctive each view is to the views that 3DViewGraph has processed. To comprehensively represent the characteristics of the views that 3DViewGraph has processed, the attention mechanism employs the fully connected weights WF\bm{W}_{F} in the final softmax classifier which accumulates the information of all views, as shown in Fig. 1(f). The attention mechanism projects the characteristics Cji\bm{C}_{j}^{i} of 3D shape mim^{i} from the jj-th view node DjiD^{i}_{j} and the characteristics WF\bm{W}_{F} of the views that 3DViewGraph has processed into a common space to calculate the distinctiveness of view vjiv_{j}^{i}, as defined below,

Based on αi\bm{\alpha}^{i}, the characteristics Cji\bm{C}_{j}^{i} of 3D shape mim^{i} from all view nodes are aggregated with weighting αi\bm{\alpha}^{i} into attentioned correlation aggregation Ci\bm{C}^{i}, as defined below,

where Ci\bm{C}^{i} represents 3D shape mim^{i} as a N×NN\times N matrix, as shown in Fig. 1(e). Finally, the global feature Fi\bm{F}^{i} of 3D shape mim^{i} is learned by a fully connected layer with attentioned correlation aggregation Ci\bm{C}^{i} as input, as shown in Fig. 1(f), where the fully connected layer is followed by a sigmoid function.

Using Fi\bm{F}^{i}, the final softmax classifier computes the probabilities Pi\bm{P}^{i} to classify the 3D shape mim^{i} into one of LL shape classes as

Learning inference. The parameters involved in 3DViewGraph are optimized by minimizing the log-likelihood OO over MM 3D shapes in the training set, where Qi\bm{Q}^{i} is the truth label,

The parameter optimization is conducted by back propagation of classification errors of 3D shapes. Noteworthy, WF\bm{W}_{F} is updated by two elements with the learning rate ε\varepsilon as follows,

The advantage of Eq. (10) is that WF\bm{W}_{F} can be learned more flexibly for optimization convergence. WF\bm{W}_{F} also enables αi\bm{\alpha}^{i} to simultaneously observe the characteristics of shape mim^{i} from each view node DjiD_{j}^{i} and take all views that have been processed from different shapes as reference.

Results and analysis

We evaluate 3DViewGraph by comparing it with the state-of-the-art methods in shape classification and retrieval under ModelNet40 Wu and others (2015), ModelNet10 and ShapeNetCore55 Savva and others (2017). We also show ablation studies to justify the effectiveness of novel elements.

Parameters. We first explore how the important parameters FF, NN and σ\sigma affect the performance of 3DViewGraph under ModelNet40. The comparison in Table. 1, 2, and 3 shows that their effects are slight in a proper range.

Classification. As compared under ModelNet in Table 4, 3DViewGraph outperforms all the other methods under the same conditionWe use the same modality of views from the same camera system for the comparison, where the results of RotationNet are from Fig.4 (d) and (e) in https://arxiv.org/pdf/1603.06208.pdf. Moreover, the benchmarks are with the standard training and test split.. In addition, we show the single view classification accuracy in VGG fine-tuning (“VGG(ModelNet)”). To highlight the contribution of VGG fine-tuning, spatial similarity, and attention, we remove fine-tuning (“Ours(No finetune)”) or set all spatial similarity (“Ours(No spatiality)”) and attention (“Ours(No attention)”) to 1. The degenerated results indicate these elements are important for 3DViewGraph to achieve high accuracy. Similar phenomena is observed when we justify the effect of Cji\bm{C}^{i}_{j} and WF\bm{W}_{F} in Eq. 6 by setting them to 1 (“Ours(No attention-)”), respectively. We also justify the latent semantic embedding and spatial pattern correlation by replacing them by single view features (“Ours(No latent)”) and summation (“Ours(No correlation)”), the degenerated results also show that they are important elements. Finally, we compare our proposed view aggregation with mean (“Ours(MeanPool)”) and max pooling (“Ours(MaxPool)”) by directly pooling all single view features together. Due to the loss of content information in each view and spatial information among multiple views, pooling performs worse.

3DViewGraph also achieves the best under the more challenging benchmark ShapeNetCore55, based on the fine-tuned VGG (“VGG(ShapeNetCore55)”), as shown in Table 5. We also find that different parameters do not significantly affect the performance, such as NN and σ\sigma.

Attention visualization. We visualize the attention learned by 3DViewGraph under ModelNet40, which demonstrates how 3DViewGraph understands 3D shapes by analyzing views on a view graph. In Fig. 5, attention weights αi\bm{\alpha}^{i} on view nodes DjiD_{j}^{i} of GiG^{i} are visualized as a vector which is represented by scattered black nodes, where the corresponding views are also shown nearby, such as the views of a toilet in Fig. 5(a), a table in Fig. 5(b) and a cone in Fig. 5(c). The coordinates of black nodes along the y-axis indicate how much attention 3DViewGraph pays to the corresponding view nodes. In addition, the views that is paid the most and least attention to are highlighted by the red upward and blue downward arrow, respectively.

Fig. 5 demonstrates that 3DViewGraph is able to understand each view, since the view with the most ambiguous appearance in a view graph is depressed while the view with the most distinctive appearance is highlighted. For example, the most ambiguous views of toilet, table and cone merely show some basic shapes that provide little useful information for classification, such as the rectangles of the toilet and table, and the circle of the cone. In contrast, the most distinctive views of toilet, table and cone exhibit more unique and distinctive characteristics.

Retrieval. We evaluate the retrieval performance of 3DViewGraph under ModelNet in Table 7. We outperform the state-of-the-art methods, where the retrieval range is also shown. We further detail the precision and recall curves of these results in Fig. 7. In addition, 3DViewGraph also achieve the best results under ShapeNetCore55 in Table 6. We compare 10 state-of-the-art methods under testing set in the SHREC2017 retrieval contest Savva and others (2017) and Taco Cohen et al. (2018), where we summarize all the 10 methods (“All”) by presenting the best result of each metric due to page limit. Finally, we demonstrate that 3DViewGraph is also superior to other graph-based multi-view learning methods Anan et al. (2015); An-An et al. (2016) under Princeton Shape Benchmark (PSB) in Fig. 6.

Conclusion

In view-based deep learning models for 3D shape analysis, view aggregation via widely used pooling, leads to information loss about content and spatial relationship of views. We propose 3DViewGraph to address this issue for 3D global feature learning by more effectively aggregating unordered views with attention. By organizing unordered views taken around a 3D shape into a view graph, 3DViewGraph learns global features of the 3D shape by simultaneously encoding both the content information within view nodes and the spatial relationship among the view nodes. Through a novel latent semantic mapping, low-level view features are projected into a meaningful, lower-dimensional latent semantic embedding using a learned kernel function, which directly captures the similarities between low-level view features and latent semantic patterns. The latent semantic mapping successfully facilitates 3DViewGraph to encode the content information and the spatial relationship in each pair of view nodes using a novel spatial pattern correlation. Further, our novel attention mechanism effectively increases the discriminability of learned features by efficiently highlighting the unordered view nodes with distinctive characteristics and depressing the ones with appearance ambiguity. Our results in classification and retrieval under three large-scale benchmarks show that 3DViewGraph can learn better global features than the state-of-the-art methods due to its more effective view aggregation.

Acknowledgments

This work was supported by National Key R&D Program of China (2018YFB0505400), NSF (1813583), University of Macau (MYRG2018-00138-FST), and FDCT (273/2017/A). We thank all anonymous reviewers for their constructive comments.

References