Exploiting Sentiment and Common Sense for Zero-shot Stance Detection

Yun Luo, Zihan Liu, Yuefeng Shi, Stan Z Li, Yue Zhang

Introduction

Stance detection aims to identify the authors’ attitudes or positions (Pro (support), Con (oppose), Neu (neutral)) towards a specific target such as an entity, a topic. Mohammad et al. (2017, 2016); Walker et al. (2012); Qiu et al. (2015); Zhang et al. (2017). It is crucial for understanding opinions and analyzing how opinions are presented in texts regarding specific issues, and much work has been done building stance detection models Wei et al. (2016); Dias and Becker (2016); Allaway and Mckeown (2020). There are two salient challenges to the task. First, obtaining rich annotated data in stance detection is time-consuming and labor-intensive. To address this issue, Allaway and Mckeown (2020) propose the dataset VAST containing various topics for few-shot and zero-shot stance detection tasks, requiring the model to classify the stance of topics unseen in the training set. Second, the topic is often not explicitly mentioned in the document, resulting in difficulty. Considering Figure 1 Example 1, the document does not explicitly contain the topic ‘Olympics’, but ‘Games’ and ‘Athlete’ implicitly refer to the topic.

Existing work incorporates external knowledge to solve the challenges Liu et al. (2021); Jayaram and Allaway (2021). For example, CKE-Net achieves the state-of-the-art results for zero-shot stance detection, which uses pre-trained model BERT and commonsense knowledge graph on ConceptNet Liu et al. (2021). However, such a method only considers the knowledge relations between documents and topics (i.e., the commonsense knowledge in two-hop directed paths on the ConceptNet from documents to topics), limiting the generalization of adding other types of related knowledge. In Figure 1 Example 1, the word ‘games’ can also represent the computer programs in a different document. Such knowledge cannot be used for that document if no relation between ‘game’ and ‘computer program’ can be learned from the relations between documents and topics in the dataset.

We consider incorporating two types of general knowledge, including common sense and sentiment. First, we incorporate commonsense knowledge into the stance detection model using a graph autoencoder module. We take a pre-training method to train the graph autoencoder, separately to the stance detection module. Second, stance detection is significantly influenced by the sentiment information Li and Caragea (2019); Sobhani et al. (2016); Hardalov et al. (2022) (case study can be seen in Appendix). In Figure 1 Example 2, the document contains many positive words like ‘good’, and ‘better’ regarding the topic ‘nuclear power’, which implies a Pro stance. However, little existing work has considered sentiment knowledge for zero-shot stance detection. We use the sentiment-aware BERT (SentiBERT henceforth) to extract the sentiment information, assisting in classifying the stances of topics.

Existing work on injecting knowledge into NLP models can be broadly classified into two categories. One uses a graph encoder to integrate structural knowledge into a neural encoder Li et al. (2019); Ghosal et al. (2020); baietal and the other injects knowledge by using training losses to tune model parameters Jayaram and Allaway (2021); Peters et al. (2019); Logan et al. (2019); Liu et al. (2019). In our work, we consider the former for commonsense knowledge and the latter for sentiment due to the sources of information. In the component of knowledge graph encoding, a graph autoencoder consisting of relational graph convolutional network (RGCN) encoders Schlichtkrull et al. (2018) and a DisMult decoder Yang et al. (2014) is trained using negative sampling to obtain the relations of concepts on the commonsense knowledge graph. We inject sentiment knowledge encoded by SentiBERT into BERT using a cross attention module and tuning the fusing process by the training loss of the stance detection.

Our model achieves the state-of-the-art performance on the benchmark dataset VAST Allaway and Mckeown (2020) in both zero-shot and few-shot stance detection, improving the performance on many challenging linguistic phenomena such as sarcasm and quotations. We analyze the performance of our model with respect to different sentiment and common sense features, finding that the data with the corresponding sentiment and stance pairs (i.e., (Pos, Pro) and (Neg, Con)) are the easiest part for models to classify; in addition, increased commonsense knowledge leads to improved performance of the stance detection model. To our knowledge, we are the first to incorporate both sentiment and common sense into zero-shot stance detection model. The code has been released https://github.com/LuoXiaoHeics/StanceCS.

Related Work

Stance detection, also known as stance classification Walker et al. (2012), stance identification Zhang et al. (2017), stance prediction Qiu et al. (2015), debate-side classification Anand et al. (2011), and debate stance classification Hasan and Ng (2013), aims to identify the stance of the text author towards a target (an entity, event, idea, opinion, claim, topic, etc.) either explicitly mentioned or implied within the text. For the initial task of stance detection, models are trained an individual classifier for each topic Lin et al. (2006); Beigman Klebanov et al. (2010); Sridhar et al. (2015); Hasan and Ng (2013, 2014); Li et al. (2018) or only a small number of topics are both in training and evaluation sets Faulkner (2014); Du et al. (2017); Hardalov et al. (2021).

However, given rich and varying topics, data annotation can be time-consuming and labor-intensive. Researchers attempt to solve the task in a cross-target setting Augenstein et al. (2016); Xu et al. (2018a), training the model in a topic and testing it on another one, and propose several weakly supervised approaches using unlabeled data related to the test topics Zarrella and Marsh (2016); Wei et al. (2016); Dias and Becker (2016). Other studies propose the tasks of zero-shot and few-shot stance detection, which requires training the model in data of several topics and testing it on some unseen topics Allaway and Mckeown (2020).

Allaway and Mckeown (2020) propose to solve the task using a topic-grouped attention net, which uses the relation between the training and evaluation topics in an unsupervised way, and they also analyze the relationship between sentiment and stance from the perspective of the model by corrupting sentences with replacing sentiment words. Jayaram and Allaway (2021) use human rationales as attribution priors to provide faithful explanations of models. Liu et al. (2021) propose to incorporate commonsense knowledge to learn the relations between different topics utilizing a CompGCN (a variant of graph convolution networks). However, it limits the content of knowledge (only knowledge from documents to stances in the training data). Our model differs from such a method in that our model adopts the related concepts of both documents and topics and uses a pre-trained graph autoencoder to obtain commonsense information. Adversarial learning is also applied to solve the zero-shot task by using unlabeled raw data Allaway et al. (2021). Unlike the above work, we consider integrating external knowledge for zero-shot stance detection, including sentiment and commonsense information that are rarely considered. To our knowledge, we are the first to systematically incorporate sentiment and commonsense knowledge into the stance detection model and analyze the relationship between them (in Section 4.5 and 4.6).

Method

The architecture of our model is illustrated in Figure 2, which contains two components: (1) knowledge graph encoding, which integrates commonsense knowledge from ConceptNet (Section 3.1); (2) stance detection with sentiment and commonsense knowledge (Section 3.2).

Formally, the ConceptNet is represented as a directed labeled graph G={V,E,R}\mathcal{G=\{V,E,R}\}, with concepts vi∈Vv_{i}\in\mathcal{V} and labeled edges (vi,r,vj)∈E(v_{i},r,v_{j})\in\mathcal{E}, where r∈Rr\in\mathcal{R} is the relation type of edge between viv_{i} and vjv_{j}. The concepts in ConceptNet are unigram words or n-gram phrases in the triplet format. For example, one such triplet from ConceptNet is (teacher, RelatedTo, job).

ConceptNet has a large size of approximately 14 million edges. We extract a subset of edges related to the VAST dataset for our task. From the training documents in VAST, we first extract the set of all unique nouns, adjectives, and adverbs. These words are treated as the seeds that we use to filter the ConceptNet to a sub-graph. We extract all the triplets with a one-edge distance to any of those seed concepts, resulting in a sub-graph G′={V′,E′,R′}\mathcal{G^{\prime}=\{V^{\prime},E^{\prime},R^{\prime}}\} with 310k concepts and 750k edges. The top 5 relations include ‘RelatedTo’, ‘HasContext’, ‘IsA’, ‘Synonym’ and ‘DerivedFrom’. The sub-graph G′\mathcal{G^{\prime}} contains all the concepts related to stance targets in the VAST dataset.

Following Schlichtkrull et al. (2018), we construct a graph autoencoder to compute the representations of concepts in the sub-graph G′\mathcal{G}^{\prime}. The autoencoder takes an incomplete set (randomly sampled with 50% probability in our model) of edges E^′\hat{\mathcal{E}}^{\prime} from E′\mathcal{E}^{\prime} in G′\mathcal{G}^{\prime} as input. E^′\hat{\mathcal{E}}^{\prime} is negative sampled to the overall set of samples denoted T\mathcal{T} (details in Training). Then we assign the possible edges (vi,r,vj)∈T(v_{i},r,v_{j})\in\mathcal{T} with scores to determine the probability these edges are in E′\mathcal{E}^{\prime}. Our graph autoencoder consists of a relational concept network (RGCN) Schlichtkrull et al. (2018) encoder to obtain the latent feature representations of concepts and a DistMult scoring decoder Yang et al. (2014) to recover the missing facts of triplets.

where ff is the encoder network (requiring inputs of feature vector xix_{i} and the rank of the layer ll), NirN^{r}_{i} denotes the neighbouring concepts ii with the relation r∈Rr\in\mathcal{R}; vi,rv_{i,r} is a normalization constant, which can be set in advance vi,r=∣Nir∣v_{i,r}=|N^{r}_{i}| or learned by network learning; σ\sigma is the activation function like ReLU and Wr(1/2),W0(1/2)W_{r}^{(1/2)},W_{0}^{(1/2)} are learnable parameters though training.

Training. We use DistMult factorization as the decoder to assign scores. For a given triplet (vi,r,vj)(v_{i},r,v_{j}), the score can be obtain as follows:

Our graph autoencoder module is trained using negative sampling Schlichtkrull et al. (2018). We randomly corrupt the positive triplets, i.e., triplets in E^′\hat{\mathcal{E}}^{\prime}, to create an equal number of negative samples. The corruption is performed by modifying either of the connected concepts or relations randomly, creating the overall set of samples denoted by T\mathcal{T}. The training objective is a binary classification between positive/negative (denoted as uu) triplets with a cross entropy loss function:

2 Stance Detection Module

Sentiment Feature Encoding. To learn sentiment knowledge, we follow Zhou et al. (2020) to continually train BERT with sentiment masking. We mask the sentiment-related tokens such as sentiment lexicons, emoticons, and ratings with higher probability than general tokens. The model is trained to reconstruct the masked sentiment tokens and predict the rating of the sentences. The corrupted text x^\hat{x} is fed into BERT to obtain each word representation hi\textbf{h}_{i} and the sentence representation hCLS\textbf{h}^{CLS}. Softmax layers are used on hi\textbf{h}_{i} to predict each word’s probability, the sentiment of words, and emoticon probability, respectively. A softmax layer on hCLS\textbf{h}^{CLS} is also used to predict the rating of the text x^\hat{x}. The tasks are trained using cross-entropy loss. Following Zhou et al. (2020) , the SentiBERT are trained on Amazon review dataset Ni et al. (2019) and Yelp 2020https://www.yelp.com/dataset challenge dataset.

After pre-training the SentiBERT, given a document dd and a topic tt, we concatenate dd and tt as our model input xx in the following format: x=[CLS] d [SEP] t [SEP]x=[CLS]\ d\ [SEP]\ t\ [SEP], SentiBERT to obtain its hidden states:

where the parameters of SentiBERT are fixed in our model to keep sentiment information stabilized.

Commonsense Feature Encoding. After training the graph autoencoder, in order to extract the document-specific commonsense graph feature for the document dd and the topic tt, the unique nouns, adjectives, and adverbs in the document dd and the topic tt are extracted at first, which we denote as SS. Then we extract a sub-graph GS′\mathcal{G}^{\prime}_{S} from G′\mathcal{G}^{\prime}, which contains all the triplets either of whose concepts are in SS or within the vicinity of radius 1 from any of the concepts in SS. Next, we make a forward pass of GS′\mathcal{G}^{\prime}_{S} through the encoder of graph autoencoder to obtain the feature vectors hj\textbf{h}_{j} for all unique concepts jj in GS′\mathcal{G}^{\prime}_{S} . The average of feature vectors hj\textbf{h}_{j} for all unique concepts in GS′\mathcal{G}^{\prime}_{S} is regarded as the commonsense graph feature vector hKG\textbf{h}^{KG} for the document dd. The commonsense graph feature vector hKG\textbf{h}^{KG} is feed into a encoder layer to obtain hidden states hK\textbf{h}^{K}:

where WkW_{k} and bkb_{k} are the trained parameters of the linear layer.

Stance Classification. The input xx is first fed into BERT to obtain its hidden states:

Then the hidden states of hBERT,hsentfix{\textbf{h}_{BERT}},{\textbf{h}^{fix}_{sent}} are concatenated and fed into a cross attention module to fuse the information of BERT and SentiBERT:

where hCLS\textbf{h}^{CLS} is the hidden states of [CLS][CLS] token in BERT. The hidden states vectors of hK\textbf{h}^{K} and hCLS\textbf{h}^{CLS} are concatenated to for classification:

where WW and bb are the parameters and pp is the probability distribution on the three stance labels.

Training. Given the input and its golden label (xi,yi)(x_{i},y_{i}), the loss function Lcls\mathcal{L}_{cls} for classifying stance is cross entropy:

where ∣N∣|N| is the number of data samples. To further ensure stronger topic invariance constraints of hKG\textbf{h}_{KG}, we add a shared decoder layer DreconD_{recon} with a reconstruction loss:

Experiments

We verify the effectiveness of sentiment and common sense influence for zero-shot and few-shot stance detection. We also prove the significance of each module in our model in Section 4.4 and analyze the relationship between sentiment (common sense) and stance in Section 4.5 (4.6).

Dataset: We adopt the dataset for zero-shot and few-shot stance detection task–VAried Stance Topics (VAST) Allaway and Mckeown (2020), which is practical and useful for real-world applications. The dataset consists of thousands of topics, and the statistics are summarized in Table 1. The zero-shot topics only appear in the test set, and the few-shot topics only contain a few training examples.

Training Details We perform experiments using the official pre-trained BERT model provided by Huggingfacehttps://huggingface.co/. For the pre-trained model with sentiment information, we adopt the model provided by Zhou et al. (2020), which is a continually trained BERT on sentiment datasets. We train our model on 1 GPU (Nvidia GTX2080Ti) using the Adam optimizer Kingma and Ba (2014). For training the graph autoencoder, the initial learning rate is 1e-2. For the stance detection training process, the initial learning rate is 1.5e-5, the max sequence length for BERT and SentiBERT is 256, the batch size for training is 4, and the model is trained for three epochs.

Baselines We compare our model with several state-of-the-art baselines: (1) BiCond Augenstein et al. (2016), a model for cross-domain target stance detection task which uses one BiLSTM to encoding the topic and another BiLSTM to encoded the text; (2) CrossNet Xu et al. (2018b), a model based on the BiCond adding an aspect-specific attention layer for cross-target setting; (3) SENT Zhang et al. (2020), a model using the semantic-emotion heterogeneous graph to enhance BiLSTM for cross-traget stance detection; (4) BERT-sep, a model that encodes the text and topic separately, using BERT, and then classification with a two-layer feed-forward neural network; (5) BERT-joint Allaway and Mckeown (2020), a model with contextual conditional encoding followed by a two-layer feed-forward neural network; (6) TGA-Net Allaway and Mckeown (2020), a model using contextual conditional encoding and topic-grouped attention. In addition, we also consider the models BERT-joint-ft and TGA-Net-ft where the BERT module is fine-tuned; (7) Prior-Bin:gold Jayaram and Allaway (2021), a model applying human rationales as attributions to assist the stance detection; (8) BERT-GCN Liu et al. (2021), a model applying the conventional GCN Kipf and Welling (2016), which considers node information aggregation; (9) CKE-Net Liu et al. (2021), a model based on BERT, using the CompGCN Vashishth et al. (2019) to obtain the commonsense information.

2 Results

The results are shown in Table 2. Compared with previous models, our model achieves the state-of-the-art performance in zero-shot, few-shot, and all the topics of VAST. In particular, the macro F1 scores are 72.6%, 70.2%, and 71.3%, which are 2.4%, 0.1%, and 1.2% higher than CKE-Net model, respectively. The results of B-RGCN (our model without SentiBERT module) are 71.2% and 69.9%, with a higher macro F1 score on zero-shot topics but a similar result on few-shot topics compared with CKE-Net. The performances of both our model and B-RGCN increase largely on the zero-shot topics but less on few-shot topics, which implies that our graph autoencoder module can achieve a similar effect compared with the GCN module of CKE-Net in the few-shot topics but can improve the effectiveness in extracting relation information in zero-shot topics. This verifies the intuition that only considering the relations between documents and topics limits the transferability of CKE-Net for the zero-shot task. Compared with Prior-Bin:gold, the macro F1 scores of our model are 3.4%, 1.0%, and 2.2% higher on zero-shot, few-shot, and all the topics sets, respectively. It implies that commonsense knowledge and sentiment information are more effective than the set of specific human rationales by Prior-Bin:gold as attributions.

Our model achieves better performance on Con labels (67.4%, 66.5%, 66.9%) compared with Pro labels (60.8%, 60.0%, 60.4%), which is similar to most of the previous models (BERT-GCN, TGA-Net, and so on). The phenomenon also appears in B-RGCN and BS, which are our models without SentiBERT and without BERT, respectively (the analysis of the ablation study is explained in Section 4.4 in detail). The results suggest that the use of SentiBERT does not cause the imbalanced performance on different stances and the detection difficulty is mainly on Pro labels. In addition, the results of Neu stance labels are the highest (89.5%, 83.9%, 86.6%) than those of other labels. It indicates that it is easier for models to classify the Neu, where the topics are mostly unrelated to documents.

3 Breakdown Evaluation

We also test our model on five special phenomena of the test set on VAST following Allaway and Mckeown (2020): (1) Imp: non-neutral stance examples where the topics are not explicit in the documents; (2) mlT: documents having multiple stance topics with different topics; (3) mlS: documents having multiple stance topics with different and non-neutral labels; (4) Qte: documents with quotations; (5) Sarc: documents with sarcasm.

The results are shown in Table 3. Our model achieves the state-of-the-art performance on mlT, mlS, Qte, and Sarc with 64.7%, 55.6%, 70.1%, and 71.7%, respectively. In particular, the improvement of our model on mlS implies that different types of knowledge features help models extract stance topics-related information. The most challenging task is mlS, with a macro F1 score of 55.6% by our model. The results demonstrate that it is highly challenging to classify the topics with different stances since the stance information extracted in the model is more related to the whole sentence but more minor to the topics. The macro F1 score of Sarc increases the most, 3.5% higher than that of CKE-NET, implying that the sentiment information helps boost the model performance in understanding sarcasm, which is a sentiment-related linguistic phenomenon. The accuracy of our model on Imp is the second-highest (slightly lower than that of CKE-Net), which indicates that introducing commonsense graph knowledge can help improve the model performance on the zero-shot task.

4 Ablation Study

We conduct ablation studies of BS, S-RGCN, and B-RGCN to understand the significance of the graph autoencoder, BERT, and SentiBERT modules, respectively. The results are shown in Table 2. First, BS fuses BERT and SentiBERT feature vectors using Eq(5-6) and classifies the stance using hCLS\textbf{h}^{CLS} with a linear layer. It achieves macro F1 scores of 71.7%, 69.9%, and 70.6% on the zero-shot, few-shot, and all the topics, which are 3.2%, 1.5%, and 2.2% higher than those of BERT-joint-ft, respectively, which proves that sentiment information can help boost the performance of stance detection task.

Second, B-RGCN and S-RGCN are models without fusing the BERT and SentiBERT feature vectors. The feature vectors of [CLS][CLS] tokens from BERT or SentiBERT (the parameters of SentiBERT are not fixed) are directly concatenated with knowledge graph feature vectors to classify the stance. The macro F1 scores of S-RGCN are 69.9% and 66.5% on the zero-shot topics and the few-shot topics, 1.3%, and 3.4% lower than those of B-RGCN, respectively. It indicates that it is not sufficient to use a sentiment-specific model to do stance classification. The macro F1 score of B-RGCN on the zero-shot set is 71.2%, 1.0% higher than that of CKE-Net, which shows that our graph autoencoder module can achieve better performance for zero-shot stance detection than CompGCN. However, BS, B-RGCN, and S-RGCN do not outperform BS-RGCN in the zero-shot topics and all the topics set, which shows that the graph autoencoder, BERT, and SentiBERT are all useful for the stance detection task.

5 Sentiment and Stance

Allaway and Mckeown (2020) indicate that models of BERT-Joint are reliant on sentiment cues, and the models learn the strong association between the Neg (negative) sentiment and the Con stance, yet weak association between Pos (positive) sentiment and Pro stance. Their analysis is based on experiments where the documents are corrupted by replacing the text’s sentiment words. Here we take a different perspective and carry out experiments with respect to different stances and sentiment pairs on both B-RGCN and BS-RGCN. We use opinion lexicon Hu and Liu (2004) to classify the sentiment of document, (i.e, if a document contains more positive/negative words, we treat it as a document with the Pos (positive)/Neg (negative) sentiment; otherwise, we treat it as a document with the Neu (neutral) sentiment).

The results are shown in Figure 3 (the model trained on all the topics is tested in this experiment). For BS-RGCN, the accuracy on the corresponding stance and sentiment (Neg, Con) is 78.9%, higher than 71.4% of (Pos, Con) and 70.1% of (Neu, Con). Similarly, the accuracy on (Pos, Pro) is 56.6%, higher than 47.5% of (Neg, Pro) and 43.6% of (Neu, Pro). This suggests that data samples with corresponding sentiment and stance pairs ((Pos, Pro), (Neg, Con)) are easier to classify by our model. The performance of B-RGCN is similar to BS-RGCN, with an accuracy of 76.4% for (Neg, Pro), a little higher than those of (Pos, Con) (75.9%) and (Neu, Con) (76.0%). The same model achieves an accuracy of 50% of (Pos Pro), 10% higher than that of (Neg, Pro), and 8.1% higher than that of (Neu, Pro). The model without the sentiment module can also predict corresponding sentiment and stance pairs with higher accuracy, demonstrating that sentiment information can help stance detection models. The accuracies for B-RGCN and BS-RGCN are both significantly higher on data with Con stances than those with Pro stances. The phenomenon indicates that it is difficult for models to predict Pro stance in the VAST dataset, and the difference in performance is not caused by the difference of associations between data of (Pos, Pro) and (Neg, Con). For the data of Neu stance, the performance is less related to sentiments. The models can achieve much better results on Neu stance data, where the topics may be not related to the documents, 91.4% on (Neg, Neu), 86.9% on (Pos, Neu), 83.1% on (Neu, Neu) for BS-RGCN , and 86.7% on (Neg, Neu), 85.8% on (Pos, Neu), 85.7% on (Neu, Neu) for B-RGCN. The phenomenon demonstrates that it is easy for the model to judge whether the topic is related to the documents.

6 Common Sense and Stance

We show the relationship between common sense and stance by pre-training the graph autoencoder w.r.t different percentages of extracted concepts (Section 3.1). Using the commonsense feature with the pre-trained autoencoder, we show the results of the stance detection models B-GCN, S-GCN, and BS-RGCN on the zero-shot task. The results are given in Figure 4. As observed, the performance of the three models increases with increasing coverage of commonsense knowledge. It indicates that commonsense knowledge is directly useful for stance detection models.

7 Case Study

We also show some cases from the test data using the model trained on all the topics. In the first case, sentiment words such as ‘tanking’ or ‘standstill’ imply the negative sentiments towards the influence of the Olympics on the economy of Brazil, which further expresses an opposing stance towards ‘Olympics’. Our model outputs the correct label towards the target thanks to the sentiment information. In the second case, no explicit expression of the target ‘College’ is contained in the document. Only some implications, including ‘foreign language programs’, have relation to the ‘College’, and with the commonsense knowledge encoding, our model outputs the correct stance. The third case proves that both common sense and sentiment information can benefit the stance detection model, that ‘inhumane’ expresses a negative sentiment, and the topic ‘nail removal’ is implicitly involved by the word ‘declawing’. Our model can also give the correct stance for case III.

Conclusion

We proposed a stance detection model incorporating commonsense knowledge and sentiment information, achieving state-of-the-art zero-shot and few-shot stance detection results on the standard dataset. The ablation study showed the significance of each module, such as knowledge graph autoencoder, SentiBERT, and BERT. We also analyzed the relation between sentiment/common sense and stance, which indicate the effectiveness of this external knowledge.

Acknowledgements

Yue Zhang is the corresponding author. We would also like to thank the anonymous reviewers for the detailed and thoughtful reviews. The work is funded by the Zhejiang Province Key Project 2022SDXHDX0003.

References

Appendix A Appendix

Human Labeling for Sentiment and Stance Detection

In this part, we manually label some samples (randomly selected) from VAST dataset to prove the relation between sentiments and stances. Opinion lexicon Hu and Liu (2004) is adopted as the sentiment vocabulary. As shown in Table 5, there are many samples (7 in 10) that sentiment knowledge plays a significant role for stance detection, and few samples have a conflicting relation.