When Does Self-Supervision Help Graph Convolutional Networks?

Yuning You, Tianlong Chen, Zhangyang Wang, Yang Shen

Introduction

Graph convolutional networks (GCNs) (Kipf & Welling, 2016) generalize convolutional neural networks (CNNs) (LeCun et al., 1995) to graph-structured data and exploit the properties of graphs. They have outperformed traditional approaches in numerous graph-based tasks such as node or link classification (Kipf & Welling, 2016; Veličković et al., 2017; Qu et al., 2019; Verma et al., 2019; Karimi et al., 2019; You et al., 2020), link prediction (Zhang & Chen, 2018), and graph classification (Ying et al., 2018; Xu et al., 2018), many of which are semi-supervised learning tasks. In this paper, we mainly focus our discussion on transductive semi-supervised node classification, as a representative testbed for GCNs, where there are abundant unlabeled nodes and a small number of labeled nodes in the graph, with the target to predict the labels of remaining unlabeled nodes.

In a parallel note, self-supervision has raised a surge of interest in the computer vision domain (Goyal et al., 2019; Kolesnikov et al., 2019; Mohseni et al., 2020) to make use of rich unlabeled data. It aims to assist the model to learn more transferable and generalized representation from unlabeled data via pretext tasks, through pretraining (followed by finetuning), or multi-task learning. The pretext tasks shall be carefully designed in order to facilitate the network to learn downstream-related semantics features (Su et al., 2019). A number of pretext tasks have been proposed for CNNs, including rotation (Gidaris et al., 2018), exemplar (Dosovitskiy et al., 2014), jigsaw (Noroozi & Favaro, 2016) and relative patch location prediction (Doersch et al., 2015). Lately, Hendrycks et al. (2019) demonstrated the promise of self-supervised learning as auxiliary regularizations for improving robustness and uncertainty estimation. Chen et al. (2020) introduced adversarial training into self-supervision, to provide the first general-purpose robust pretraining.

In short, GCN tasks usually admit transductive semi-supervised settings, with tremendous unlabeled nodes; meanwhile, self-supervision plays an increasing role in utilizing unlabeled data in CNNs. In view of the two facts, we are naturally motivated to ask the following interesting, yet rarely explored question:

Can self-supervised learning play a similar role in GCNs to improve their generalizability and robustness?

Contributions. This paper presents the first systematic study on how to incorporate self-supervision in GCNs, unfolded by addressing three concrete questions:

Could GCNs benefit from self-supervised learning in their classification performance? If yes, how to incorporate it in GCNs to maximize the gain?

Does the design of pretext tasks matter? What are the useful self-supervised pretext tasks for GCNs?

Would self-supervision also affect the adversarial robustness of GCNs? If yes, how to design pretext tasks?

Directly addressing the above questions, our contributions are summarized as follows:

We demonstrate the effectiveness of incorporating self-supervised learning in GCNs through multi-task learning, i.e. as a regularization term in GCN training. It is compared favorably against self-supervision as pretraining, or via self-training (Sun et al., 2019).

We investigate three self-supervised tasks based on graph properties. Besides the node clustering task previously mentioned in (Sun et al., 2019), we propose two new types of tasks: graph partitioning and completion. We further illustrate that different models and datasets seem to prefer different self-supervised tasks.

We further generalize the above findings into the adversarial training setting. We provide extensive results to show that self-supervision also improves robustness of GCN under various attacks, without requiring larger models nor additional data.

Related Work

Graph-based semi-supervised learning. Semi-supervised graph-based learning works with the crucial assumption that the nodes connected with edges of larger weights are more likely to have the same label (Zhu & Goldberg, 2009). There are abundance of work on graph-based methods, e.g. (randomized) mincuts (Blum & Chawla, 2001; Blum et al., 2004), Boltzmann machines (Getz et al., 2006; Zhu & Ghahramani, 2002) and graph random walks (Azran, 2007; Szummer & Jaakkola, 2002). Lately, graph convolutional network (GCN) (Kipf & Welling, 2016) and its variants (Veličković et al., 2017; Qu et al., 2019; Verma et al., 2019) have gained their popularity by extending the assumption from a hand-crafted one to a data-driven fashion. A detailed review could be referred to (Wu et al., 2019b).

Self-supervised learning. Self-supervision is a promising direction for neural networks to learn more transferable, generalized and robust features in computer vision domain (Goyal et al., 2019; Kolesnikov et al., 2019; Hendrycks et al., 2019). So far, the usage of self-supervision in CNNs mainly falls under two categories: pretraining & finetuning, or multi-task learning. In pretraining & finetuning. the CNN is first pretrained with self-supervised pretext tasks, and then finetuned with the target task supervised by labels (Trinh et al., 2019; Noroozi & Favaro, 2016; Gidaris et al., 2018), while in multi-task learning the network is trained simultaneously with a joint objective of the target supervised task and the self-supervised task(s). (Doersch & Zisserman, 2017; Ren & Jae Lee, 2018).

To our best knowledge, there has been only one recent work pursuing self-supervision in GCNs (Sun et al., 2019), where a node clustering task is adopted through self-training. However, self-training suffers from limitations including performance “saturation” and degrading (to be detailed in Sections 3.2 and 4.1 for theoretical rationales and empirical results). It also restricts the types of self-supervision tasks that can be incorporated.

Adversarial attack and defense on graphs. Similarly to CNNs, the wide applicability and vulnerability of GCNs raise an urgent demand for improving their robustness. Several algorithms are proposed to attack and defense on graph (Dai et al., 2018; Zügner et al., 2018; Wang et al., 2019a; Wu et al., 2019a; Wang et al., 2019b).

Dai et al. (2018) developed attacking methods by dropping edges, based on gradient descent, genetic algorithms and reinforcement learning. Zügner et al. (2018) proposed an FSGM-based approach to attack the edges and features. Lately, more diverse defense approaches emerge. Dai et al. (2018) defended the adversarial attacks by directly training on perturbed graphs. Wu et al. (2019a) gained robustness by learning graphs from the continuous function. Wang et al. (2019a) used graph refining and adversarial contrasting learning to boost the model robustness. Wang et al. (2019b) proposed to involve unlabeled data with pseudo labels that enhances scalability to large graphs.

Method

In this section, we first elaborate three candidate schemes to incorporate self-supervision with GCNs. We then design novel self-supervised tasks, each with its own rationale explained. Lastly we generalize self-supervised to GCN adversarial defense.

2 Three Schemes: Self-Supervision Meets GCNs

Pretraining & finetuning. In the pretraining process, the network is trained with the self-supervised task as following:

Despite improving performance in previous few-shot experiments, M3S shows performance gain “saturation” in Table 2 as the label rate grows higher, echoing literature (Zhu & Goldberg, 2009; Li et al., 2018). Further, we will show and rationalize their limited performance boost in Section 4.1.

Multi-task learning. Considering a target task and a self-supervised task for a GCN with (2), the output and the training process can be formulated as:

In the problem (4), we regard the self-supervised task as a regularization term throughout the network training. The regularization term is traditionally and widely used in graph signal processing, and a famous one is graph Laplacian regularizer (GLR) (Shuman et al., 2013; Bertrand & Moonen, 2013; Milanfar, 2012; Sandryhaila & Moura, 2014; Wu et al., 2016) which penalizes incoherent (i.e. nonsmooth) signals across adjacent nodes (Chen & Liu, 2017). Although the effectiveness of GLR has been shown in graph signal processing, the regularizer is manually set simply following the smoothness prior without the involvement of data, whereas the self-supervised task acts as the regularizer learned from unlabeled data under the minor guidance of human prior. Therefore, a properly designed task would introduce data-driven prior knowledge that improves the model generalizability, as show in Table 1.

In total, multi-task learning is the most general framework among the three. Acting as the data-driven regularizer during training, it makes no assumption on the self-supervised task type. It is also experimentally verified to be the most effective among all the three (Section 4).

3 GCN-Specific Self-Supervised Tasks

While Section 3.2 discusses the “mechanisms” by which GCNs could be trained with self-supervision, here we expand a “toolkit” of self-supervised tasks for GCNs. We show that, by utilizing the rich node and edge information in a graph, a variety of GCN-specific self-supervised tasks (as summarized in Table 3) could be defined and will be further shown to benefit various types of supervised/downstream tasks. They will assign different pseudo-labels to unlabeled nodes and solve formulation in (4).

With the clusters of node sets, we assign cluster indices as self-supervised labels to all the nodes:

Graph partitioning. Clustering-related algorithms are node feature-based, with the rationale of grouping nodes with similar attributes. Another rationale to group nodes can be based on topology in graph data. In particular two nodes connected by a “strong” edge (with a large weight) are highly likely of the same label class (Zhu & Goldberg, 2009). Therefore, we propose a topology-based self-supervision using graph partitioning.

Different from node clustering based on node features, graph partitioning provides the prior regularization based on graph topology, which is similar to graph Laplacian regularizer (GLR) (Shuman et al., 2013; Bertrand & Moonen, 2013; Milanfar, 2012; Sandryhaila & Moura, 2014; Wu et al., 2016) that also adopts the idea of “connection-prompting similarity”. However, GLR, which is already injected into the GCNs architecture, locally smooths all nodes with their neighbor nodes. In contrast, graph partitioning considers global smoothness by utilizing all connections to group nodes with heavier connection densities.

Graph completion. Motivated by image inpainting a.k.a. completion (Yu et al., 2018) in computer vision (which aims to fill missing pixels of an image), we propose graph completion, a novel regression task, as a self-supervised task. As an analogy to image completion and illustrated in Figure 2, our graph completion first masks target nodes by removing their features. It then aims at recovering/predicting masked node features by feeding to GCNs unmasked node features (currently restricted to second-order neighbors of each target node for 2-layer GCNs).

We design such a self-supervised task for the following reasons: 1) the completion labels are free to obtain, which is the node feature itself; and 2) we consider graph completion can aid the network for better feature representation, which teaches the network to extract feature from the context.

4 Self-Supervision in Graph Adversarial Defense

With the three self-supervised tasks introduced for GCNs to gain generalizability toward better-performing supervised learning (for instance, node classification), we proceed to examine their possible roles in gaining robustness against various graph adversarial attacks.

Adversarial attacks. We focus on single-node direct evasion attacks: a node-specific attack type on the attributes/links of the target node vnv_{n} under certain constraints following (Zügner et al., 2018), whereas the trained model (i.e. the model parameters (θ∗,Θ∗)(\theta^{*},\boldsymbol{\Theta}^{*})) remains unchanged during/after the attack. The attacker gg generates perturbed feature and adjacency matrices, X′\boldsymbol{X}^{\prime} and A′\boldsymbol{A}^{\prime}, as:

with (attribute, links and label of) the target node and the model parameters as inputs. The attack can be on links, (node) features or links & features.

Adversarial training for graph data can then be formulated as both supervised learning for labeled nodes and recovering pseudo labels for unlabeled nodes (attacked and clean):

Adversarial defense with self-supervision. With self-supervision working in GCNs formulated as in (4) and adversarial training in (6), we formulate adversarial training with self-supervision as:

Experiments

In this section, we extensively assess, analyze, and rationalize the impact of self-supervision on transductive semi-supervised node classification following (Kipf & Welling, 2016) on the aspects of: 1) the standard performances of GCN (Kipf & Welling, 2016) with different self-supervision schemes; 2) the standard performances of multi-task self-supervision on three popular GNN architectures — GCN, graph attention network (GAT) (Veličković et al., 2017), and graph isomorphism network (GIN) (Xu et al., 2018); as well as those on two SOTA models for semi-supervised node classification — graph Markov neural network (GMNN) (Qu et al., 2019) that introduces statistical relational learning (Koller & Pfeffer, 1998; Friedman et al., 1999) into its architecture to facilitate training and GraphMix (Verma et al., 2019) that uses the Mixup trick; and 3) the performance of GCN with multi-task self-supervision in adversarial defense. Implementation details can be found in Appendix A.

Self-supervision incorporated into GCNs through various schemes. We first examine three schemes (Section 3.2) to incorporate self-supervision into GCN training: pretraining & finetuning, self-training (i.e. M3S (Sun et al., 2019)) and multi-task learning. The hyper-parameters of M3S are set at default values reported in (Sun et al., 2019). The differential effects of the three schemes combined with various self-supervised tasks are summarized for three datasets in Table 5, using the target performances (accuracy in node classification). Each combination of self-supervised scheme and task is run 50 times for each dataset with different random seeds so that the mean and the standard deviation of its performance can be reported.

Through the remaining two schemes, GCNs with self-supervision incorporated could see more significant improvements in the target task (node classification) compared to GCN without self-supervision. In contrast to pretraining and finetuning that switches the objective function after self-supervision in (3) and solves a new optimization problem in (2), both self-training and multi-task learning incorporate self-supervision into GCNs through one optimization problem and both essentially introduce an additional self-supervision loss to the original formulation in (2).

Their difference lies in what pseudo-labels are used and how they are generated for unlabeled nodes. In the case of self-training, the pseudo-labels are the same as the target-task labels and such “virtual” labels are assigned to unlabeled nodes based on their proximity to labeled nodes in graph embedding. In the case of multi-task learning, the pseudo-labels are no longer restricted to the target-task labels and can be assigned to all unlabeled nodes by exploiting graph structure and node features without labeled data. And the target supervision and the self-supervision in multi-task learning are still coupled through common graph embedding. So compared to self-training, multi-task learning can be more general (in pseudo-labels) and can exploit more in graph data (through regularization).

Multi-task self-supervision on SOTAs. Does multi-task self-supervision help SOTA GCNs? Now that we have established multi-task learning as an effective mechanism to incorporate self-supervision into GCNs, we set out to explore the added benefits of various self-supervision tasks to SOTAs through multi-task learning. Table 6 shows that different self-supervised tasks could benefit different network architectures on different datasets to different extents.

When does multi-task self-supervision help SOTAs and why? We note that graph partitioning is generally beneficial to all three SOTAs (network architectures) on all the three datasets, whereas node clustering do not benefit SOTAs on PubMed. As discussed in Section 3.2 and above, multi-task learning introduce self-supervision tasks into the optimization problem in (4) as the data-driven regularization and these tasks represent various priors (see Section 3.3).

(1) Feature-based node clustering assumes that feature similarity implies target-label similarity and can group distant nodes with similar features together. When the dataset is large and the feature dimension is relatively low (such as PubMed), feature-based clustering could be challenged in providing informative pseudo-labels.

(2) Topology-based graph partitioning assumes that connections in topology implies similarity in labels, which is safe for the three datasets that are all citation networks. In addition, graph partitioning as a classification task does not impose the assumption overly strong. Therefore, the prior represented by graph partitioning can be general and effective to benefit GCNs (at least for the types of the target task and datasets considered).

(3) Topology and feature-based graph completion assumes the feature similarity or smoothness in small neighborhoods of graphs. Such a context-based feature representation can greatly improve target performance, especially when the neighborhoods are small (such as Citeseer with the smallest average degree among all three datasets). However, the regression task can be challenged facing denser graphs with larger neighborhoods and more difficult completion tasks (such as the larger and denser PubMed with continuous features to complete). That being said, the potentially informative prior from graph completion can greatly benefit other tasks, which is validated later (Section 4.2).

Does GNN architecture affect multi-task self-supervision? For every GNN architecture/model, all three self-supervised tasks improve its performance for some datasets (except for GMNN on PubMed). The improvements are more significant for GCN, GAT, and GIN. We conjecture that data-regularization through various priors could benefit these three architectures (especially GCN) with weak priors to begin with. In contrast, GMNN sees little improvement with graph completion. GMNN introduces statistical relational learning (SRL) into the architecture to model the dependency between vertices and their neighbors. Considering that graph completion aids context-based representation and acts a somewhat similar role as SRL, the self-supervised and the architecture priors can be similar and their combination may not help. Similarly GraphMix introduces a data augmentation method Mixup into the architecture to refine feature embedding, which again mitigates the power of graph completion with overlapping aims.

We also report in Appendix B the results in inductive fully-supervised node classification. Self-supervision leads to modest performance improvements in this case, appearing to be more beneficial in semi-supervised or few-shot learning.

2 Self-Supervision Boosts Adversarial Robustness

What additional benefits could multi-task self-supervision bring to GCNs, besides improving the generalizability of graph embedding (Section 4.1)? We additionally perform adversarial experiments on GCN with multi-task self-supervision against Nettack (Zügner et al., 2018), to examine its potential benefit on robustness.

What self-supervision task helps defend which types of graph attacks and why? In Tables 7 and 8 we find that introducing self-supervision into adversarial training improves GCN’s adversarial defense. (1) Node clustering and graph partitioning are more effective against feature attacks and links attacks, respectively. During adversarial training, node clustering provides the perturbed feature prior while graph partitioning does perturbed link prior for GCN, contributing to GCN’s resistance against feature attacks and link attacks, respectively. (2) Strikingly, graph completion boosts the adversarial accuracy by around 4.5 (%) against link attacks and over 8.0 (%) against the link & feature attacks on Cora. It is also among the best self-supervision tasks for link attacks and link & feature attacks on Citeseer, albeit with a smaller improvement margin (around 1%). In agreement with our earlier conjecture in Section 4.1, the topology- and feature-based graph completion constructs (joint) perturbation prior on links and features, which benefits GCN in its resistance against link or link & feature attacks.

3 Result Summary

We briefly summarize the results as follows.

First, among three schemes to incorporate self-supervision into GCNs, multi-task learning works as the regularizer and consistently benefits GCNs in generalizable standard performances with proper self-supervised tasks. Pretraining & finetuning switches the objective function from self-supervision to target supervision loss, which easily “overwrites” shallow GCNs and gets limited performance gain. Self-training is restricted in what pseudo-labels are assigned and what data are used to assign pseudo-labels. And its performance gain is more visible in few-shot learning and can be diminishing with slightly increasing labeling rates.

Second, through multi-task learning, self-supervised tasks provide informative priors that can benefit GCN in generalizable target performance. Node clustering and graph partitioning provide priors on node features and graph structures, respectively; whereas graph completion with (joint) priors on both help GCN in context-based feature representation. Whether a self-supervision task helps a SOTA GCN in the standard target performance depends on whether the dataset allows for quality pseudo-labels corresponding to the task and whether self-supervised priors complement existing architecture-posed priors.

Last, multi-task self-supervision in adversarial training improves GCN’s robustness against various graph attacks. Node clustering and graph partitioning provides priors on features and links, and thus defends better against feature attacks and link attacks, respectively. Graph completion, with (joint) perturbation priors on both features and links, boost the robustness consistently and sometimes drastically for the most damaging feature & link attacks.

Conclusion

In this paper, we present a systematic study on the standard and adversarial performances of incorporating self-supervision into graph convolutional networks (GCNs). We first elaborate three mechanisms by which self-supervision is incorporated into GCNs and rationalize their impacts on the standard performance from the perspective of optimization. Then we focus on multi-task learning and design three novel self-supervised learning tasks. And we rationalize their benefits in generalizable standard performances on various datasets from the perspective or data-driven regularization. Lastly, we integrate multi-task self-supervision into graph adversarial training and show their improving robustness of GCNs against adversarial attacks. Our results show that, with properly designed task forms and incorporation mechanisms, self-supervision benefits GCNs in gaining both generalizability and robustness. Our results also provide rational perspectives toward designing such task forms and incorporation tasks given data characteristics, target tasks and neural network architectures.

Acknowledgements

We thank anonymous reviewers for useful comments that help improve the paper during revision. This study was in part supported by the National Institute of General Medical Sciences of the National Institutes of Health [R35GM124952 to Y.S.], and a US Army Research Office Young Investigator Award [W911NF2010240 to Z.W.].

References