Human Parity on CommonsenseQA: Augmenting Self-Attention with External Attention
Yichong Xu, Chenguang Zhu, Shuohang Wang, Siqi Sun, Hao Cheng, Xiaodong Liu, Jianfeng Gao, Pengcheng He, Michael Zeng, Xuedong Huang
Introduction
Transformers (Vaswani et al., 2017) have revolutionized many areas of AI with state-of-the-art performance in a wide range of tasks (Devlin et al., 2018; Dosovitskiy et al., 2020). The most notable and effective component in a Transformer model is the self-attention mechanism, which enables the model to dynamically leverage different parts of the input for computation, with no information loss for even the most distant parts of the input. With the success of pre-trained models (Devlin et al., 2018; Liu et al., 2019), the Transformer and its self-attention mechanism have been widely adopted as the cornerstone of foundation models trained on huge amounts of data (Bommasani et al., 2021).
One phenomenon found during the development of Transformer models is that models with larger sizes tend to have better learning abilities, especially when combined with large-scale data. This has prompted the recent boom of super large Transformer models, ranging from BERT (Devlin et al., 2018) with 110 million parameters, to GPT-3 (Brown et al., 2020) with 175 billion parameters. Nevertheless, numerous studies have shown that the corresponding understanding and generation capabilities of these huge models are still behind humans (Bommasani et al., 2021). Furthermore, the sheer size of these models already poses serious practical challenges in utilization, deployment, interpretation, and environmental impact (Patterson et al., 2021). Thus, the recent “scaling-up” approach to Transformer-based NLP modeling is unsustainable and has been questioned in recent studies (Bommasani et al., 2021).
In this paper, we take a step back and examine the mechanism of current Transformer-based models. Self-attention was designed to allow the model to better analyze the inner structure of input data, and the model is trained to have its parameters grasp and memorize all the content and patterns of the training data. When the model is given a novel input , the implicitly stored knowledge in the parameters about related information is activated to facilitate the analysis of . This could partly explain why larger models pre-trained with more data have an advantage in performance.
While Transformer models process input by looking inward via self-attention, we propose to make the model look outward by providing it with related context and knowledge from various sources. We then let the model conduct self-attention on the input while also computing external attention to the knowledge (Figure 1). As the context and knowledge can usually be stored in a non-parametric and symbolic way (e.g., plain text, knowledge graph and dictionary entries), even moderately-sized Transformer models can perform exceptionally well on NLP tasks. This approach allows one to shrink the size of Transformer-based foundation models, which is critical to the accessibility and democratization of AI technology. This approach is also analogous to the way humans conduct intelligence; we often resort to search engines, dictionaries, or information from other people to navigate the world.
Another benefit of external attention is that, as the related knowledge is stored outside of the model, practitioners can easily update the knowledge source to change the behavior of their models. For example, one could add or delete entries from a knowledge graph or rewrite certain paragraphs in Wikipedia. By explicitly representing knowledge, the decision process of the model becomes much more transparent and explainable.
In this paper, we use the commonsense reasoning task CommonsenseQA (Talmor et al., 2019) as a case study in leveraging external attention to obtain and integrate information related to the input. Given a commonsense question and a choice, we retrieve knowledge from three external sources: a knowledge graph (ConceptNet), a dictionary (Wiktionary), and labeled training data (CommonsenseQA and 16 related QA datasets). The retrieved knowledge is directly appended to the input and sent to the language model with no revision to the underlying architecture. We show that with the proposed external attention, the accuracy of commonsense reasoning using a DeBERTa-xxlarge model (He et al., 2020) can be significantly boosted from 83.8% to 90.8% on the dev set, while fine-tuned large-scale models like GPT-3 can only achieve 73.0%. The ensembled version of our model, Knowledgeable External Attention for commonsense Reasoning (KEAR), reaches an accuracy of 93.4% on the dev set and 89.4% on the test set, surpassing human performance (88.9%) for the first time (Talmor et al., 2019).
The benefits of our approach extend beyond commonsense reasoning. First, the external attention dramatically reduces our system’s dependence on large-scale models, i.e., achieving human parity with models up to 1.5B parameters. Second, the external information is obtained via computationally efficient methods, such as information retrieval and word matching, adding little computational cost to the main model. Third, the text-level concatenation of input and knowledge leads no change to the Transformer model, enabling existing systems to easily adopt this new external attention mechanism.
Method
We first describe our external attention framework in Sec 2.1. Next, we describe our external knowledge sources in Sec 2.2. Last, we present additional modeling techniques for improving commonsense reasoning in Sec 2.3.
We focus on the multiple-choice question answering task in this paper, where the goal is to select the correct answer from a given list for a commonsense question . The output of the model is a distribution on .
1 External Attention
The majority of recent language models are based on the Transformer architecture (Vaswani et al., 2017). One of the most important components in Transformer is the self-attention mechanism, which can be formulated as
External Attention.
For commonsense question answering, the required information needed to answer the question is usually absent from the input. Thus, we need to integrate external knowledge into the model. In this work, we denote the extra knowledge in text format as . There are many ways to integrate the external knowledge into the model, such as using graph neural networks (Lin et al., 2019). In this paper we simply concatenate the knowledge to the input text: . The advantage of this input-level integration is that the existing model architecture does not need to be modified. Then, applying self-attention on can make the model freely reason between the knowledge text and the question/choices, therefore equipping the model with enhanced reasoning capacity.
2 External Knowledge Sources
The knowledge to append to the input for external attention is crucial for getting the correct prediction. For commonsense reasoning, we collect three external knowledge sources to complement the input questions and choices.
Knowledge graphs (KG) contain curated facts that can help with commonsense reasoning which might not appear in a general corpus. We follow KCRhttps://github.com/jessionlin/csqa to retrieve relevant relation triples in the ConceptNet graph (Speer et al., 2017). Suppose the question entity is and the choice contains entity In CommonsenseQA dataset, both and are provided. Otherwise, we use entity linking to find related knowledge graph nodes to the input text (see “Training Data” part later in this section).. If there is a direct edge from to in ConceptNet, we choose this triple . Otherwise, we retrieve all the triples originating from . We score each triple by the product of its confidence (provided by ConceptNet) and the defined relation type weight : , where is the relation type of , is the total number of triples originating from , is the number of triples with relation among these triples. We then choose the triple with highest weight. Finally, if the selected triple is , we denote the knowledge from the KG as .
Dictionary.
Although pre-trained language models are exposed to large-scale text data, the long tail distribution of words means that the quality of a word’s representation is highly dependent on that word’s frequency in the pre-training corpus. Dictionaries, on the other hand, can provide accurate semantic explanation of words regardless of their frequency in datasets. To help understand key concepts in the question and answer, we follow DEKCOR (Xu et al., 2021) to use the Wiktionary definitions of the question and answer concepts as external knowledge. For every concept, we fetch the first (most frequent) definition from Wiktionary using its closest lexical match. Let be the definition text for and be the definition text for , we denote the dictionary knowledge as .
Training Data.
Although recent language models are giant in terms of the number of parameters, recent studies show that they cannot perfectly memorize all the details of their training data (Wang et al., 2022).
To tackle this challenge, we propose to retrieve relevant questions and answers from the training data as additional knowledge. We use BM25 (Schütze et al., 2008) to retrieve top relevant questions and answers from the training data. For each question , we index the concatenation of the question text, the ground-truth choice , ConceptNet triples and Wiktionary definitions: . For a new question and a potential choice , we similarly build a query to retrieve most similar questions from the training set. For each retrieved question from the training data, we drop the knowledge part and employ the retrieved question and the ground-truth answer as external knowledge. Suppose the retrieved questions and (correct) answers are , we denote the knowledge from training data as . During training, for each question we filter itself from the retrieved results to avoid data leakage.
Different from Wang et al. (2022) where the retrieval questions are only obtained from the same dataset, we experiment with three sources of training data for retrieval: i) CSQA training data, ii) CSQA+OBQA+RiddleSense, a small collection of datasets focusing on ConceptNet knowledge, and iii) a pool of 17 datasets focusing on commonsense reasoning (we describe details of these 17 datasets in the Appendix). Since most datasets do not provide the question and choice entity for every question-choice pair, we use entity linking to find all entities appearing in the question and choice text respectively. We select the entity with the maximum length in and as the question and choice entity for Wiktionary definitions. For ConceptNet triples, we find edges between and and choose the one with the maximum total length.
Finally, we concatenate the retrieved knowledge from our three sources to form a final knowledge input: . In practice, the semicolon is replaced by the separator token (e.g., [SEP]). We name our knowledge retrieval and integration technology as Knowledgeable External Attention for commonsense Reasoning (KEAR), shown in Figure 1.
3 General Methods to Improve Commonsense Reasoning
Prior works have proposed other methods to improve general NLU performance, and it is therefore natural to wonder if these methods also work for commonsense reasoning. Here, we explore two general methods for improving commonsense reasoning performance: i) using different text encoders and ii) virtual adversarial learning.
Previous methods for natural language understanding (NLU) have tried using BERT (Devlin et al., 2018), RoBERTa (Liu et al., 2019), ALBERT (Lan et al., 2019), T5 (Raffel et al., 2019), ELECTRA (Clark et al., 2020) and DeBERTa (He et al., 2020) as the text encoder, achieving state-of-the-art performance on the GLUE benchmark (Wang et al., 2019). Thus, we evaluate these models as encoders for the commonsense reasoning task.
Virtual Adversarial Training (VAT).
Previous works show that virtual adversarial training (VAT) can improve the performance for general NLU and question answering tasks (Jiang et al., 2020). In the multiple-choice commonsense reasoning task, the goal is to minimize the cross-entropy loss:
where produces the model prediction (distribution on the choices), represents the model parameters, is the one-hot ground-truth answer vector, CE is cross-entropy, and is the empirical data distribution. VAT first finds the update that leads to the largest change in the predicted distribution, subject to a -norm constraint. Then, a consistency regularization loss term is added to minimize the difference in the function’s output when compared to the input variation :
where and are hyperparameters.
Experiments
We present our empirical results in this section Our source code is released at https://github.com/microsoft/KEAR. . Each of our three external knowledge sources can boost the commonsense reasoning performance, and combining all the three techniques helps us reach the human parity on the CommonsenseQA benchmark.
We focus on the CommonsenseQA (CSQA, Talmor et al., 2019) benchmark. CommonsenseQA is a widely used multiple-choice question answering dataset that requires commonsense knowledge. It contains 12k questions created using ConceptNet (Speer et al., 2017). For an edge (subject, relation, object) in ConceptNet, Talmor et al. (2019) retrieve other object concepts with the same subject and relation as distractors for a question. A human worker is then asked to i) write a question containing the subject and with the object as the correct answer, ii) pick the most distractive answer from the retrieved concepts, and iii) write another distractor for the question. The final question contains five choices, with one correct choice, two random retrieved concepts, one human-picked concept, and one human-curated answer.
We present details of the 17 datasets that we use for training data retrieval in Table 1. All the datasets are multiple-choice or classification datasets related to commonsense reasoning, and we include dataset details in the appendix.
Model Setup.
Implementation Details.
We finetune the model using the AdamW optimizer (Loshchilov and Hutter, 2017). The batch size is set to 48 or smaller to fit the batch onto a single GPU. We train the model for 10 epochs and take the best result on the dev set. We choose the weight decay in . The learning rate are chosen from for all encoders except for DeBERTa; following the DeBERTa paper (He et al., 2020) we use a smaller learning rate, chosen from . We use the DeBERTa v2 model and choose from the pretrained model or model finetuned on MNLI. We also try out the recent DeBERTa V3 model (He et al., 2021) which combines DeBERTa with adversarial pretraining. For VAT, we choose the weight multiplier and set input variation norm (see Eqn. 3). For retrieving from training data, we choose the data source with the best validation set performance from the three retrieval source datasets. We set number of retrieved questions . We run each experiment with 3 different seeds and present results from the best run.
2 Effects of Individual Components
As shown in Table 2, there is a positive correlation between general performance on NLI tasks and commonsense reasoning abilities on CommonsenseQA. Notice that the fine-tuned GPT-3 model with 175 billion parameters could only achieve 73.0% on the dev set of CommonsenseQA. Based on these results, we choose ELECTRA-large and DeBERTa variants (He et al., 2020, 2021) as the encoders for subsequent experimentation.
For virtual adversarial training, we find that VAT can improve commonsense reasoning accuracy for ELECTRA, improving the result from 81.3% to 82.1%. It does not show much improvement for DeBERTa models in our experiment. Therefore, we apply VAT to ELECTRA in subsequent experiments.
Effect of External Attention.
As shown in Table 3, all of the proposed knowledge sources bring gains in commonsense reasoning accuracy across all base encoder models. The dictionary, knowledge graph and training data bring 0.5%, 2.1%, and 2.5% improvement, respectively, when DeBERTaV3-large (He et al., 2021) is the base encoder model.
We find that the best training data retrieval source depends on the exact encoders and the techniques applied, and we show a comparison in Table 4. In general, the 17-dataset pool achieves the best performance for DeBERTa, but for ELECTRA retrieving from the CSQA training set alone can get the best performance. Table 3 and 4 demonstrate the effectiveness of our proposed knowledge retrieval and concatenation methods.
3 Combining the Techniques
Table 5 shows the results of KEAR, which combines the best techniques in previous experiments, i.e., best encoders and external attention to all knowledge sources, to further boost the performance. The best single model (DeBERTaV3-large + KEAR) achieves 91.2% accuracy on the dev set. To get the best performance, we train KEAR models with ELECTRA large, DeBERTa xlarge (900M), xxlarge (1.5B) and V3 large as encoders with 12 different seeds, resulting in 48 models in total. We rank the models by their dev set performance as . The ensemble prediction uses a majority vote on individual predictions. We picked the first models such that the dev set performance of ensembling is the best. We ended up with 39 models with 12 ELECTRA models, 12 DeBERTaV3 models, 11 DeBERTa-xxlarge models and 4 DeBERTa-xlarge models. Our ensemble model reaches 93.4% accuracy on the dev set. Table 6 shows the official leaderboard result on the hidden test set. Our ensemble model exceeds the previously best DEKCOR model by over 6% and exceeds the human performance (88.9%) by 0.5%.
4 Case Study
We present two examples from CSQA in Table 7 to illustrate how the model can reason between all retrieved knowledge sources to get the correct answer. For the first question, the knowledge graph helps rule out the wrong answer triangle since it does not have a surface. The dictionary and training data attention further confirms that a tetrahedron has four sides/faces, which is the correct answer. For the second question, again knowledge graph rules out “salad” since a dog does not desire salads. The dictionary attention results suggest that bones are important for a dog, and training data attention suggests that bones are good food for a dog. This leads to the correct answer (bone). This suggests that all three knowledge sources are critical for getting the correct answer. Having access to all three knowledge sources makes it easier for model reasoning to get the correct answer.
Related work
Many previous works have proposed ways of incorporating external knowledge sources into Transformer architectures. For commonsense question answering, specialized knowledge graphs like ConceptNet (Speer et al., 2017) and ATOMIC (Sap et al., 2019) are the most popular choices for external knowledge (Chang et al., 2021; Yao et al., 2022; Song et al., 2021). Lin et al. (2019) construct a scheme graph from concepts in the question and choices and uses an LSTM to reason on paths between question and choice concepts. Yasunaga et al. (2021) construct a joint graph containing the QA context and KG, then use graph neural networks to reason over the two knowledge sources. Liang et al. (2021) proposes a KG-Transformer for using knowledge graphs in generative question answering.
Another line of work explores less structured knowledge such as Wikipedia and dictionaries for commonsense reasoning (Xu et al., 2021; Lv et al., 2020). Bhakthavatsalam et al. (2020) combine the knowledge from ConceptNet, WordNet, and other corpora to form 3.5M generic statements and show that this knowledge can help boost accuracy and explanation quality. Mitra et al. (2020) compares several ways of incorporating external knowledge from a relevant corpus for commensense question answering.
Recently, there are approaches to generate facts from pretrained language models to complement missing facts in the external knowledge source. Bosselut et al. (2019) finetune a pretrained model on ATOMIC for commonsense knowledge graph completion. Liu et al. (2021) directly prompt the GPT-3 model (Brown et al., 2020) to get knowledge for reasoning.
Beyond commonsense reasoning, external knowledge can also help boost performance on other language processing tasks like open domain question answering (Yu et al., 2021), relation classification (Yu et al., 2020a) dialog response generation (Ghazvininejad et al., 2018), conversational QA (Qin et al., 2019), multilingual NLU (Fang et al., 2021) and text generation (Yu et al., 2020b). Compared with prior work that uses extra modules (e.g., GNNs) or extra models (e.g., GPT-3), our external attention framework is extremely lightweight. It operates via a combination of non-parametric retrieval and text concatenation, which we show is highly effective, able to surpass human parity on the CommonsenseQA task.
Conclusion
We propose external attention as a lightweight framework for retrieving and integrating external knowledge for language understanding. Compared with self-attention which benefits from ever-increasing model sizes, external attention can bring related information from external sources to supplement the input. We demonstrate that this strategy can lead to considerable gains in performance with little additional computational cost. By leveraging knowledge from knowledge graphs, dictionaries, and training data, we show that our technology, KEAR, achieves human parity on the CommonsenseQA benchmark for the first time. For future work, we will apply the technique to other NLP tasks to improve language model performance with external knowledge.
Acknowledgement
We thank the anonymous reviewers for their comments on our paper. We thank Reid Pryzant for proof-reading the paper.