DAGN: Discourse-Aware Graph Network for Logical Reasoning
Yinya Huang, Meng Fang, Yu Cao, Liwei Wang, Xiaodan Liang
Introduction
A variety of QA datasets have promoted the development of reading comprehensions, for instance, SQuAD Rajpurkar et al. (2016), HotpotQA Yang et al. (2018), DROP Dua et al. (2019), and so on. Recently, QA datasets with more complicated reasoning types, i.e., logical reasoning, are also introduced, such as ReClor Yu et al. (2020) and LogiQA Liu et al. (2020). The logical questions are taken from standardized exams such as GMAT and LSAT, and require QA models to read complicated argument passages and identify logical relationships therein. For example, selecting a correct assumption that supports an argument, or finding out a claim that weakens an argument in a passage. Such logical reasoning is beyond the capability of most of the previous QA models which focus on reasoning with entities or numerical keywords.
A main challenge for the QA models is to uncover the logical structures under passages, such as identifying claims or hypotheses, or pointing out flaws in arguments. To achieve this, the QA models should first be aware of logical units, which can be sentences or clauses or other meaningful text spans, then identify the logical relationships between the units. However, the logical structures are usually hidden and difficult to be extracted, and most datasets do not provide such logical structure annotations.
An intuitive idea for unwrapping such logical information is using discourse relations. For instance, as a conjunction, “because” indicates a causal relationship, whereas “if” indicates a hypothetical relationship. However, such discourse-based information is seldom considered in logical reasoning tasks. Modeling logical structures is still lacking in logical reasoning tasks, while current opened methods use contextual pre-trained models Yu et al. (2020). Besides, previous graph-based methods Ran et al. (2019); Chen et al. (2020a) that construct entity-based graphs are not suitable for logical reasoning tasks because of different reasoning units.
In this paper, we propose a new approach to solve logical reasoning QA tasks by incorporating discourse-based information. First, we construct discourse structures. We use discourse relations from the Penn Discourse TreeBank 2.0 (PDTB 2.0) Prasad et al. (2008) as delimiters to split texts into elementary discourse units (EDUs). A logic graph is constructed in which EDUs are nodes and discourse relations are edges. Then, we propose a Discourse-Aware Graph Network (DAGN) for learning high-level discourse features to represent passages.The discourse features are incorporated with the contextual token features from pre-trained language models. With the enhanced features, DAGN predicts answers to logical questions. Our experiments show that DAGN surpasses current opened methods on two recent logical reasoning QA datasets, ReClor and LogiQA.
We propose to construct logic graphs from texts by using discourse relations as edges and elementary discourse units as nodes.
We obtain discourse features via graph neural networks to facilitate logical reasoning in QA models.
We show the effectiveness of using logic graph and feature enhancement by noticeable improvements on two datasets, ReClor and LogiQA.
Method
Our intuition is to explicitly use discourse-based information to mimic the human reasoning process for logical reasoning questions. The questions are in multiple choices format, which means given a triplet (context, question, answer options), models answer the question by selecting the correct answer option. Our framework is shown in Figure 1. We first construct a discourse-based logic graph from the raw text. Then we conduct reasoning via graph networks to learn and update the discourse-based features, which are incorporated with the contextual token embeddings for downstream answer prediction.
Our discourse-based logic graph is constructed via two steps: delimiting text into elementary discourse units (EDUs) and forming the graph using their relations as edges, as illustrated in Figure 1(1).
It is studied that clause-like text spans delimited by discourse relations can be discourse units that reveal the rhetorical structure of texts Mann and Thompson (1988); Prasad et al. (2008). We further observe that such discourse units are essential units in logical reasoning, such as being assumptions or opinions. As the example shown in Figure 1, the “while” in the context indicates a comparison between the attributes of “pure analog system” and that of “digital systems”. The “because” in the option provides evidence “error cannot occur in the emission of digital signals” to the claim “digital systems are the best information systems”.
We use PDTB 2.0 Prasad et al. (2008) to help drawing discourse relations. PDTB 2.0 contains discourse relations that are manually annotated on the 1 million Wall Street Journal (WSJ) corpus and are broadly characterized into “Explicit” and “Implicit” connectives. The former apparently presents in sentences such as discourse adverbial “instead” or subordinating conjunction “because”, whereas the latter are inferred by annotators between successive pairs of text spans split by punctuation marks such as “.” or “;”. We simply take all the “Explicit” connectives as well as common punctuation marks to form our discourse delimiter library (details are given in Appendix A), with which we delimit the texts into EDUs. For each data sample, we segment the context and options, ignoring the question since the question usually does not carry logical content.
Discourse Graph Construction
We define the discourse-based graphs with EDUs as nodes, the “Explicit” connectives as well as the punctuation marks as two types of edges. We assume that each connective or punctuation mark connects the EDUs before and after it. For example, the option sentence in Figure 1 is delimited into two EDUs, “digital systems are the best information systems” and “error cannot occur in the emission of digital signals” by the connective “because”. Then the returned triplets are and . For each data sample with the context and multiple answer options, we separately construct graphs corresponding to each option, with EDUs in the same context and every single option. The graph for the single option is denoted by .
2 Discourse-Aware Graph Network
We present the Discourse-Aware Graph Network (DAGN) that uses the constructed graph to exploit discourse-based information for answering logical questions. It consists of three main components: an EDU encoding module, a graph reasoning module, and an answer prediction module. The former two are demonstrated in Figure 1(2), whereas the final component is in Figure 1(3).
An EDU span embedding is obtained from its token embeddings. There are two steps. First, similar to previous works Yu et al. (2020); Liu et al. (2020), we encode such input sequence “ context question || option ” into contextual token embeddings with pre-trained language models, where and are the special tokens for RoBERTa Liu et al. (2019) model, and || denotes concatenation. Second, given the token embedding sequence , the -th EDU embedding is obtained by , where is the set of token indices belonging to -th EDU.
Graph Reasoning
After EDU encoding, DAGN performs reasoning over the discourse graph. Inspired by previous graph-based models Ran et al. (2019); Chen et al. (2020a), we also learn graph node representations to obtain higher-level features. However, we consider different graph construction and encoding. Specifically, let denote a graph corresponding to the -th option in answer choices. For each node , the node embedding is initialized with the corresponding EDU embedding . indicates the neighbors of node . is the adjacency matrix for one of the two edge types, where indicates graph edges corresponding to the explicit connectives, and indicates graph edges corresponding to punctuation marks.
The model first calculates weight for each node with a linear transformation and a sigmoid function , then conducts message propagation with the weights:
After the message propagation, the node representations are updated with the initial node embeddings and the message representations by
where and are weight and bias respectively. The updated node representations will be used to enhance the contextual token embedding via summation in corresponding positions. Thus , where and is the corresponding token indices set for -th EDU.
Answer Prediction
The probabilities of options are obtained by feeding the discourse-enhanced token embeddings into the answer prediction module. The model is end-to-end trained using cross entropy loss. Specifically, the embedding sequence first goes through a layer normalization Ba et al. (2016), then a bidirectional GRU Cho et al. (2014). The output embeddings are then added to the input ones as the residual structure He et al. (2016). We finally obtain the encoded sequence after another layer normalization on the added embeddings.
We then merge the high-level discourse features and the low-level token features. Specifically, the variant-length encoded context sequence, question-and-option sequence are pooled via weighted summation wherein the weights are softmax results of a linear transformation of the sequence, resulting in single feature vectors separately. We concatenate them with “” embedding from the backbone pre-trained model, and feed the new vector into a two-layer perceptron with a GELU activation Hendrycks and Gimpel (2016) to get the output features for classification.
Experiments
We evaluate the performance of DAGN on two logical reasoning datasets, ReClor Yu et al. (2020) and LogiQA Liu et al. (2020), and conduct ablation study on graph construction and graph network. The implementation details are shown in Appendix B.
ReClor contains 6,138 questions modified from standardized tests such as GMAT and LSAT, which are split into train / dev / test sets with 4,638 / 500 / 1,000 samples respectively. The training set and the development set are available. The test set is blind and hold-out, and split into an EASY subset and a HARD subset according to the performance of BERT-base model Devlin et al. (2019). The test results are obtained by submitting the test predictions to the leaderboard. LogiQA consists of 8,678 questions that are collected from National Civil Servants Examinations of China and manually translated into English by professionals. The dataset is randomly split into train / dev / test sets with 7,376 / 651 / 651 samples respectively. Both datasets contain multiple logical reasoning types.
2 Results
The experimental results are shown in Tables 1 and 2. Since there is no public method for both datasets, we compare DAGN with the baseline models. As for DAGN, we fine-tune RoBERTa-Large as the backbone. DAGN (Aug) is a variant that augments the graph features.
DAGN reaches 58.20% of test accuracy on ReClor. DAGN (Aug) reaches 58.30%, therein 75.91% on EASY subset, and 44.46% on HARD subset. Compared with RoBERTa-Large, the improvement on the HARD subset is remarkably 4.46%. This indicates that the incorporated discourse-based information supplements the shortcoming of the baseline model, and that the discourse features are beneficial for such logical reasoning. Besides, DAGN and DAGN (Aug) also outperform the baseline models on LogiQA, especially showing 4.01% improvement over RoBERTa-Large on the test set.
3 Ablation Study
We conduct ablation study on graph construction details as well as the graph reasoning module. The results are reported in Table 3.
We first use clauses or sentences in substitution for EDUs as graph nodes. For clause nodes, we simply remove “Explicit” connectives during discourse unit delimitation. So that the texts are just delimited by punctuation marks. For sentence nodes, we further reduce the delimiter library to solely period (“.”). Using the modified graphs with clause nodes or coarser sentence nodes, the accuracy of DAGN drops to 64.40%. This indicates that clause or sentence nodes carry less discourse information and act poorly as logical reasoning units.
Varied Graph Edges
We make two changes of the edges: (1) modifying the edge type, (2) modifying the edge linking. For edge type, all edges are regarded as a single type. For edge linking, we ignore discourse relations and connect every pair of nodes, turning the graph into fully-connected. The resulting accuracies drop to 64.80% and 61.60% respectively. It is proved that in the graph we built, edges link EDUs in reasonable manners, which properly indicates the logical relations.
Ablation on Graph Reasoning
We remove the graph module from DAGN and give a comparison. This model solely contains an extra prediction module than the baseline. The performance on ReClor dev set is between the baseline model and DAGN. Therefore, despite the prediction module benefits the accuracy, the lack of graph reasoning leads to the absence of discourse features and degenerates the performance. It demonstrates the necessity of discourse-based structure in logical reasoning.
Related Works
Recent datasets for reading comprehension tend to be more complicated and require models’ capability of reasoning. For instance, HotpotQA Yang et al. (2018), WikiHop Welbl et al. (2018), OpenBookQA Mihaylov et al. (2018), and MultiRC Khashabi et al. (2018) require the models to have multi-hop reasoning. DROP Dua et al. (2019) and MA-TACO Zhou et al. (2019) need the models to have numerical reasoning. WIQA Tandon et al. (2019) and CosmosQA Huang et al. (2019) require causal reasoning that the models can understand the counterfactual hypothesis or find out the cause-effect relationships in events. However, the logical reasoning datasets Yu et al. (2020); Liu et al. (2020) require the models to have the logical reasoning capability of uncovering the inner logic of texts.
Deep neural networks are used for reasoning-driven RC. Evidence-based methods Madaan et al. (2020); Huang et al. (2020); Rajagopal et al. (2020) generate explainable evidence from a given context as the backup of reasoning. Graph-based methods Qiu et al. (2019); De Cao et al. (2019); Cao et al. (2019); Ran et al. (2019); Chen et al. (2020b); Xu et al. (2020b); Zhang et al. (2020) explicitly model the reasoning process with constructed graphs, then learn and update features through message passing based on graphs. There are also other methods such as neuro-symbolic models Saha et al. (2021) and adversarial training Pereira et al. (2020). Our paper uses a graph-based model. However, for uncovering logical relations, graph nodes and edges are customized with discourse information.
Discourse information provides a high-level understanding of texts and hence is beneficial for many of the natural language tasks, for instance, text summarization Cohan et al. (2018); Joty et al. (2019); Xu et al. (2020a); Feng et al. (2020), neural machine translation Voita et al. (2018), and coherent text generation Wang et al. (2020); Bosselut et al. (2018). There are also discourse-based applications for reading comprehension. DISCERN Gao et al. (2020) segments texts into EDUs and learns interactive EDU features. Mihaylov and Frank (2019) provide additional discourse-based annotations and encodes them with discourse-aware self-attention models. Unlike previous works, DAGN first uses discourse relations as graph edges connecting EDUs for texts, then learns the discourse features via message passing with graph neural networks.
Conclusion
In this paper, we introduce a Discourse-Aware Graph Network (DAGN) to addressing logical reasoning QA tasks. We first treat elementary discourse units (EDUs) that are split by discourse relations as basic reasoning units. We then build discourse-based logic graphs with EDUs as nodes and discourse relations as edges. DAGN then learns the discourse-based features and enhances them with contextual token embeddings. DAGN reaches competitive performances on two recent logical reasoning datasets ReClor and LogiQA.
Acknowledgements
The authors would like to thank Wenge Liu, Jianheng Tang, Guanlin Li and Wei Wang for their support and useful discussions. This work was supported in part by National Natural Science Foundation of China (NSFC) under Grant No.U19A2073 and No.61976233, Guangdong Province Basic and Applied Basic Research (Regional Joint Fund-Key) Grant No.2019B1515120039, Shenzhen Basic Research Project (Project No. JCYJ20190807154211365), Zhijiang Lab’s Open Fund (No. 2020AA3AB14) and CSIG Young Fellow Support Fund.
References
Appendix A Discourse Delimiter Library
Our discourse delimiter library consists of two parts, the “Explicit” connectives annotated in Penn Discourse TreeBank 2.0 (DPTB 2.0) Prasad et al. (2008), as well as a set of punctuation marks. The overall discourse delimiters used in our method are presented in Table 4.
Appendix B Implementation Details
We fine-tune RoBERTa-Large Liu et al. (2019) as the backbone pre-trained language model for DGAN, which contains 24 hidden layers with hidden size 1024. The overall model is end-to-end trained and updated by Adam Kingma and Ba (2015) optimizer with an overall learning rate of 5e-6 and a weight decay of 0.01. The overall dropout rate is 0.1. The maximum sequence length is 256. We tune the model on the dev set to obtain the best iteration steps of graph reasoning, which is 2 for ReClor data, and 3 for LogiQA data. The model is trained for 10 epochs with a batch size of 16 on Nvidia Tesla V100 GPU.
For the answer prediction module, the hidden size of GRU is the same as the token embeddings in the pre-trained language model, which is 1024. The two-layer perceptron first projects the concatenated vectors with a hidden size of 1024 3 to 1024, then project 1024 to 1.