A Paragraph-level Multi-task Learning Model for Scientific Fact-Verification

Xiangci Li, Gully Burns, Nanyun Peng

Introduction

Many seemingly convincing rumors such as “Most humans only use 10 percent of their brain” are widely spread, but ordinary people are not able to rigorously verify them by searching for scientific literature. In fact, it is not a trivial task to verify a scientific claim by providing supporting or refuting evidence rationales, even for domain experts. The situation worsens as misinformation is proliferated on social media or news websites, manually or programmatically, at every moment. As a result, an automatic fact-verification tool becomes more and more crucial for combating the spread of misinformation.

The existing fact-verification tasks usually consist of three sub-tasks: document retrieval, rationale sentence extraction, and fact-verification. However, due to the nature of scientific literature that requires domain knowledge, it is challenging to collect a large scale scientific fact-verification dataset, and further, to perform fact-verification under a low-resource setting with limited training data. Wadden et al. (2020) collected a scientific claim-verification dataset, SciFact, and proposed a scientific claim-verification task: given a scientific claim, find evidence sentences that support or refute the claim in a corpus of scientific paper abstracts. Wadden et al. (2020) also proposed a simple, pipeline-based, sentence-level model, VeriSci, as a baseline solution based on DeYoung et al. (2019).

VeriSci is a pipeline model that runs modules for abstract retrieval, rationale sentence selection, and stance prediction sequentially, and thus the error generated from an upstream module may propagate to the downstream modules. To overcome this drawback, we hypothesize that a module jointly optimized on multiple sub-tasks may mitigate the error-propagation problem to improve the overall performance. In addition, we observe that a complete set of rationale sentences usually contains multiple inter-related sentences from the same paragraph. Therefore, we propose a novel, paragraph-level, multi-task learning model for the SciFact task.

In this work, we employ compact paragraph encoding, a novel strategy of computing sentence representations using BERT-family models. We directly feed an entire paragraph as a single sequence to BERT, so that the encoded sentence representations are already contextualized on the neighbor sentences by taking advantage of the attention mechanisms in BERT. In addition, we jointly train the modules for rationale selection and stance prediction as multi-task learning (Caruana 1997) by leveraging the confidence score of rationale selection as the attention weight of the stance prediction module. Furthermore, we compare two methods of transfer learning that mitigate the low-resource issue: pre-training and domain adaptation (Peng and Dredze 2017). Our experiments show that:

The compact paragraph encoding method is beneficial over separately computing sentence embeddings.

With negative sampling, the joint training of rationale selection and stance prediction is beneficial over the pipeline solution.

SciFact Task Formulation

Given a scientific claim cc and a corpus of scientific paper abstracts AA, the SciFact (Wadden et al. 2020) task retrieves all abstracts E^(c)\hat{E}(c) that either Supports or Refutes cc. Specifically, the stance prediction (a.k.a. label prediction) task classifies each abstract a∈Aa\in A into y(c,a)∈{Support,Refutes,NoInfo}y(c,a)\in\{\text{{Support}},\text{{Refutes}},\text{{NoInfo}}\} with respect to each claim cc; the rationale selection (a.k.a. sentence selection) task retrieves all rationale sentences S^(c,a)={s1^(c,a),...,sl^(c,a)}\hat{S}(c,a)=\{\hat{s_{1}}(c,a),...,\hat{s_{l}}(c,a)\} of each aa that Supports or Refutes cc. The performance of both tasks are evaluated with F1F1 measure at both abstract-level and sentence-level, as defined by Wadden et al. (2020), where {Supports,Refutes}\{\text{{Supports}},\text{{Refutes}}\} are considered as the positive labels and NoInfo is the negative label for stance prediction.

Approach

We formulate the SciFact task (Wadden et al. 2020) as a sentence-level sequence-tagging problem. We first apply an abstract retrieval module to filter out negative candidate abstracts that do not contain sufficient information with respect to each given claim. Then we propose a novel model for joint rationale selection and stance prediction using multi-task learning (Caruana 1997).

In contrast to the TF-IDF similarity used by Wadden et al. (2020), we leverage BioSentVec (Chen, Peng, and Lu 2019) embedding, which is the biomedical version of Sent2Vec (Pagliardini, Gupta, and Jaggi 2018), for a fast and scalable sentence-level similarity computation. We first compute the BioSentVec (Chen, Peng, and Lu 2019) embedding of each abstract in the corpus by treating the concatenation of each title and abstract as a single sentence. Then for each given claim, we compute the cosine similarities of the claim embedding against the pre-computed abstract embeddings, and choose the top kretrievalk_{retrieval} similar abstracts as the candidate abstracts for the next module.

2 Joint Rationale Selection and Stance Prediction Model

A major usage of BERT-family models (Devlin et al. 2018; Liu et al. 2019) for sentence-level sequence tagging computes each sentence embedding in a paragraph with batches. Since each batch is independent, such method leaves the contextualization of the sentences to the subsequent modules. Instead, we propose a novel method of encoding paragraphs by directly feeding the concatenation of the claim cc and the whole paragraph PP to a BERT model BERTBERT as a single sequence SeqSeq. By separating each sentence ss using the BERT model’s [SEP][SEP] token, we fully leverage the multi-head attention (Vaswani et al. 2017) within the BERT model to compute the contextualized word representations hSeqh_{Seq} with respect to the claim sentence and the whole paragraph.

Sentence Representations via Word-level Attention

Next, we apply a weighted sum to the contextualized word representations of each sentence hsenth_{sent} to compute the sentence representations hsih_{s_{i}}. The weights are obtained by applying a self-attention SelfAttnwordSelfAttn_{word} with a two-layer multi-layer perceptron on the word representations in the scope of each sentence, as separated by the [SEP][SEP] tokens.

Dynamic Rationale Representations

We use a two-layer multi-layer perceptron MLPrationaleMLP_{rationale} to compute the rationale score and use the softmaxsoftmax function to compute the probability of each candidate sentence being a rationale sentence prp^{r} or not pnot_rp^{not\_r} with respect to the claim sentence cc. Then we only feed rationale sentences rr into the next stance prediction module.

Stance Prediction

We use two variants for stance prediction: a simple sentence-level attention and the Kernel Graph Attention Network (KGAT) (Liu et al. 2020).

Simple Attention. We apply another weighted summation on the predicted rationale sentence representations hrih_{r_{i}} to compute the whole paragraph’s rationale representation, where the attention weights are obtained by applying another self-attention SelfAttnsentenceSelfAttn_{sentence} on the rationale sentence representations hrh_{r}. Finally, we apply another two-layer multi-layer perceptron MLPstanceMLP_{stance} and the softmaxsoftmax function to compute the probability of the paragraph serving the role of {Supports, Refutes, NoInfo} with respect to the claim cc.

Kernel Graph Attention Network. Liu et al. (2020) proposed KGAT as a stance prediction module for their pipeline solution on the FEVER (Thorne et al. 2018) task. In addition to the Graph Attention Network (Veličković et al. 2017), which applies attention mechanisms on each word pair and sentence pair in the input paragraph, KGAT applies a kernel pooling mechanism (Xiong et al. 2017) to extract better features for stance prediction. We integrate KGAT (Liu et al. 2020) into our multi-task learning model for stance prediction on SciFact (Wadden et al. 2020). The KGAT module KGATKGAT takes the word representation of the claim hch_{c} and the predicted rationale sentence representations hRh_{R} as inputs, and outputs the probability of the paragraph serving the role of {Supports, Refutes, NoInfo} with respect to the claim cc.

3 Model Training

We train our model on rationale selection and stance prediction using multi-task learning approach (Caruana 1997). We use cross-entropy loss as the training objective for both tasks. We introduce a coefficient γ\gamma to adjust the proportion of two loss values LrationaleL_{rationale} and LstanceL_{stance} in the joint loss LL.

Scheduled Sampling

Because the stance prediction module takes the predicted rationale sentences as the input, errors in rationale selection may propagate to the stance prediction module, especially during the early stage of training. To mitigate this issue, we apply scheduled sampling (Bengio et al. 2015), which starts by feeding the ground truth rationale sentences to the stance prediction module, and gradually increasing the proportion of the predicted rationale sentences, until eventually all input sentences are the predicted rationale sentences. We use a sinsin function to compute the probability of sampling predicted rationale sentences psamplep_{sample} as a function of the progress of the training:

Negative Sampling and Down-sampling

Although the abstract retrieval module filters out the majority of the negative candidate abstracts, the false-positive rate is still inevitably high, in order to ensure the retrieval of most of the positive abstracts. As a result, the input to the joint prediction model is highly biased towards negative samples. Therefore, in addition to the positive samples from the SciFact dataset (Wadden et al. 2020), we perform negative sampling (Mikolov et al. 2013) to sample the top ktraink_{train} similar negative abstracts using our abstract retrieval module as an augmented dataset for training and validation to increase the downstream model’s tolerance to false positive abstracts. Furthermore, in order to increase the diversity of the dataset, we augment the dataset by down-sampling sentences within each paragraph.

FEVER Pre-training

As Wadden et al. (2020) proposed, due to the similar task structure of FEVER (Thorne et al. 2018) and SciFact (Wadden et al. 2020), we first pre-train our model on the FEVER dataset, then fine-tune on the SciFact dataset by partially re-initializing the rationale selection and stance prediction attention modules.

Domain Adaptation

Instead of pre-training, we also explore domain adaptation (Peng and Dredze 2017) from FEVER (Thorne et al. 2018) to SciFact (Wadden et al. 2020). We use shared representations for the compact paragraph encoding and word-level attention, while using domain-specific representations for the rationale selection and stance prediction modules.

4 Implementation Details

We follow Wadden et al. (2020) in using Roberta-large (Liu et al. 2019) as our BERT-family model.

We dynamically feed only the predicted rationale sentence representations to the stance prediction module. To address the special case when an abstract contains no rationale sentences, we append a fixed dummy sentence (e.g.“@”) whose rationale label is always at the beginning of each of the paragraph. When the stance prediction module has no actual rationale sentence to take as input, we feed it with the representation of the dummy sentence and expect the module to predict NoInfo.

To prevent inconsistency between the outputs of rationale selection and stance prediction, we enforce the predicted stance to be NoInfo if no rationale sentence is proposed.

Table 1 lists the hyper-parameters used for training the Joint-Paragraph model in Table 4 https://github.com/jacklxc/ParagraphJointModel, where kFEVERk_{FEVER} refers to the number of negative samples retrieved from FEVER (Thorne et al. 2018) for model pre-training.

Experiments

SciFact (Wadden et al. 2020) is a small dataset, whose corpus contains 5183 abstracts. There are 1409 claims, including 809 in the training set, 300 in the development set and 300 in the test set.

2 Abstract Retrieval Performance

Table 2 compares the performance of abstract retrieval modules using using TF-IDF and BioSentVec (Chen, Peng, and Lu 2019). As Table 2 indicates, the overall difference between these two methods is small. Wadden et al. (2020) chose kretrieval=3k_{retrieval}=3 to maximize the F1F_{1} score of the abstract retrieval module, while we choose a larger kretrievalk_{retrieval} to pursue a larger recall score, in order to retrieve more positive abstracts for the down-stream models.

3 Baseline Models

Along with the SciFact task and dataset, Wadden et al. (2020) proposed VeriSci, a sentence-level, pipeline-based solution. After retrieving the top similar abstracts for each claim with TF-IDF vectorization method, they applied a sentence-level “BERT to BERT” model DeYoung et al. (2019) to extract rationales, sentence by sentence, with a BERT model, and they predict the stance with another BERT model using the concatenation of the extracted rationale sentences. Wadden et al. (2020) used Roberta-large (Liu et al. 2019) as their BERT model and pre-trained their stance prediction module on the FEVER dataset (Thorne et al. 2018).

Very recently, Pradeep et al. (2020) proposed a strong model VerT5erini, based on T5 (Raffel et al. 2019). They applied T5 for all three steps of the SciFact task in a sentence-level, pipeline fashion. Because of the known significant performance gap between Roberta-large (Liu et al. 2019) that we use and T5 (Raffel et al. 2019; Pradeep et al. 2020), we only use VerT5erini as a reference (marked with *).

4 Model Performances and Ablation Studies

We experiment on the oracle task, which performs rationale selection and stance prediction given the oracle abstracts (Table 3), and the open task, which performs the full task of abstract retrieval, rationale selection, and stance prediction (Table 4). We tune our models based on the sentence-level, final development set performance (Selection+Label). The test labels are not released by Wadden et al. (2020). Unless explicitly stated, all models are pre-trained on FEVER (Thorne et al. 2018).

We compare our paragraph-level pipeline model against VeriSci (Wadden et al. 2020), which is a sentence-level solution on the oracle task. As Table 3 shows, our paragraph-level pipeline model (Paragraph-Pipeline) outperforms VeriSci, particularly on rationale selection. This suggests the benefit of computing the contextualized sentence representations using the compact paragraph encoding over individual sentence representations.

Although our joint model does not show benefits over the pipeline model on the oracle task (Table 3), the benefit emerges on the open task. Along with negative sampling, which greatly increases the tolerance of models to false positive abstracts, the Paragraph-Joint model shows its benefit over the Paragraph-Pipeline model. The small difference between the Paragraph-Joint model and the same model except with TF-IDF abstract retrieval (Paragraph-Joint TF-IDF) shows that the performance improvement is mainly attributed to the joint training, instead of replacing TF-IDF similarity with BioSentVec embedding similarity in abstract retrieval.

We also compare two methods of transfer learning from FEVER (Thorne et al. 2018) to SciFact (Wadden et al. 2020). Table 4 shows that the effect of pre-training (Paragraph-Joint) or domain adaptation (Peng and Dredze 2017) (Paragraph-Joint DA) is similar. Both of them are effective as transfer learning, as they significantly outperform the same model that is only trained on SciFact (Paragraph-Joint SciFact-only).

We expected a significant performance improvement by applying the strong stance prediction model KGAT (Liu et al. 2020), but the actual improvement is limited. This is likely due to the strong regularization of KGAT that under-fits the training data.

By the time this paper is updated, our Paragraph-Joint model trained on the combination of SciFact training set and development set achieved the first place on the SciFact leaderboard https://scifact.apps.allenai.org/leaderboard, as of January 23, 2021.. We obtain test sentence-level F1F_{1} score (Selection+Label) of 60.9%60.9\% and test abstract-level F1F_{1} score (Label+Rationale) of 67.2%67.2\%.

Related Work

Fact-verification has been widely studied. There are many datasets available on various domains (Vlachos and Riedel 2014; Ferreira and Vlachos 2016; Popat et al. 2017; Wang 2017; Derczynski et al. 2017; Popat et al. 2017; Atanasova 2018; Baly et al. 2018; Chen et al. 2019; Hanselowski et al. 2019), among which the most influential one is FEVER shared task (Thorne et al. 2018), which aims to develop systems to check the veracity of human-generated claims by extracting evidences from Wikipedia. Most existing systems (Nie, Chen, and Bansal 2019) leverages a three-step pipeline approach by building modules for each of the step: document retrieval, rationale selection and fact verification. Many of them focus on the claim verification step (Zhou et al. 2019; Liu et al. 2020), such as KGAT (Liu et al. 2020), one of the top models on FEVER leader board. On the other hand, there are some attempts on jointly optimizing rationale selection and stance prediction. TwoWingOS (Yin and Roth 2018) leverages attentive CNN (Yin and Schütze 2018) to inter-wire two modules, while Hidey et al. (2020) used a single pointer network (Vinyals, Fortunato, and Jaitly 2015) for both sub-tasks. We propose another variation that directly links two modules by a dynamic attention mechanism.

Because SciFact (Wadden et al. 2020) is a scientific version of FEVER (Thorne et al. 2018), systems designed for FEVER can be applied to SciFact in principle. However, as a fact-verification task in scientific domain, SciFact task has inherited the common issue of lacking sufficient data, which can be mitigated with transfer learning by leveraging language models and introducing external dataset. The baseline model by Wadden et al. (2020) leverages Roberta-large (Liu et al. 2019) fine-tuned on FEVER dataset (Thorne et al. 2018), while VerT5erini (Pradeep et al. 2020) leverages T5 (Raffel et al. 2019) and fine-tuned on MS MARCO dataset (Bajaj et al. 2016). In this work, in addition to fine-tuning Roberta-large on FEVER, we also explore domain adaptation (Peng and Dredze 2017) to mitigate the low resource issue.

Conclusion

In this work, we propose a novel paragraph-level multi-task learning model for SciFact task. Experiments show that (1) The compact paragraph encoding method is beneficial over separately computing sentence embeddings. (2) With negative sampling, the joint training of rationale selection and stance prediction is beneficial over the pipeline solution.

Acknowledgement

We thank the anonymous reviewers for their useful comments, and Dr. Jessica Ouyang for her feedback. This work is supported by a National Institutes of Health (NIH) R01 grant (LM012592). The views and conclusions of this paper are those of the authors and do not reflect the official policy or position of NIH.

References