Discern: Discourse-Aware Entailment Reasoning Network for Conversational Machine Reading

Yifan Gao, Chien-Sheng Wu, Jingjing Li, Shafiq Joty, Steven C. H. Hoi, Caiming Xiong, Irwin King, Michael R. Lyu

Introduction

Conversational Machine Reading (CMR) is challenging because the rule text may not contain the literal answer, but provide a procedure to derive it through interactions Saeidi et al. (2018). In this case, the machine needs to read the rule text, interpret the user scenario, clarify the unknown user’s background by asking questions, and derive the final answer. Taking Figure 1 as an example, to answer the user whether he is suitable for the loan program, the machine needs to interpret the rule text to know what are the requirements, understand he meets “American small business” from the user scenario, ask follow-up clarification questions about “for-profit business” and “not get financing from other resources”, and finally it concludes the answer “Yes” to the user’s initial question.

Existing approaches Zhong and Zettlemoyer (2019); Sharma et al. (2019); Gao et al. (2020) decompose this problem into two sub-tasks. Given the rule text, user question, user scenario, and dialog history (if any), the first sub-task is to make a decision among “Yes”, “No”, “Inquire” and “Irrelevant”. The “Yes/No” directly answers the user question and “Irrelevant” means the user question is unanswerable by the rule text. If the user-provided information (user scenario, previous dialogs) are not enough to determine his fulfillment or eligibility, an “Inquire” decision is made and the second sub-task is activated. The second sub-task is to capture the underspecified condition from the rule text and generate a follow-up question to clarify it. Zhong and Zettlemoyer (2019) adopt BERT Devlin et al. (2019) to reason out the decision, and propose an entailment-driven extracting and editing framework to extract a span from the rule text and edit it into the follow-up question. The current state-of-the-art model EMT Gao et al. (2020) uses a Recurrent Entity Network Henaff et al. (2017) with explicit memory to track the fulfillment of rules at each dialog turn for decision making and question generation.

In this problem, document interpretation requires identification of conditions and determination of logical structures because rules can appear in the format of bullet points, in-line conditions, conjunctions, disjunctions, etc. Hence, correctly interpreting rules is the first step towards decision making. Another challenge is dialog understanding. The model needs to evaluate the user’s fulfillment over the conditions, and jointly consider the fulfillment states and the logical structure of rules for decision making. For example, disjunctions and conjunctions of conditions have completely different requirements over the user’s fulfillment states. However, existing methods have not considered condition-level understanding and reasoning.

In this work, we propose Discern: Discourse-Aware Entailment Reasoning Network . To better understand the logical structure of a rule text and to extract conditions from it, we first segment the rule text into clause-like elementary discourse units (EDUs) using a pre-trained discourse segmentation model Li et al. (2018). Each EDU is treated as a condition of the rule text, and our model estimates its entailment confidence scores over three states: Entailment, Contradiction or Neutral by reading the user scenario description and existing dialog. Then we map the scores to an entailment vector for each condition, and reason out the decision based on the entailment vectors and the logical structure of rules. Compared to previous methods that do little entailment reasoning Zhong and Zettlemoyer (2019) or use it as multi-task learning Gao et al. (2020), Discern is the first method to explicitly build the dependency between entailment states and decisions at each dialog turn.

Discern achieves new state-of-the-art results on the blind, held out test set of ShARC Saeidi et al. (2018). In particular, Discern outperforms the previous best model EMT Gao et al. (2020) by 3.8% in micro-averaged decision accuracy and 3.5% in macro-averaged decision accuracy. Specifically, Discern performs well on simple in-line conditions and conjunctions of rules while still needing improvements on understanding disjunctions. Finally, we conduct comprehensive analyses to unveil the limitation of Discern and current challenges for the ShARC benchmark. We find one of the biggest bottlenecks is the user scenario interpretation, in which various types of reasoning are required.

Discern Model

Discern answers the user question through a three-step process shown in Figure 2:

First, Discern segments the rule text into individual conditions using discourse segmentation.

Taking the user-provided information including the user question, user scenario and dialog history as inputs, Discern predicts the entailment state and maps it to an entailment vector for each segmented condition. Then it reasons out the decision by considering the logical structure of the rule text and the fulfillment of each condition.

Finally, if the decision is “Inquire”, Discern generates a follow-up question to clarify the underspecified condition in the rule text.

The goal of rule segmentation is to understand the logical structure of the rule text and parse it into individual conditions for the ease of entailment reasoning. Ideally, each segmented unit should contain at most one condition. Otherwise, it will be ambiguous to determine the entailment state for that unit. Determining conditions is easy when they appear as bullet points, but in most cases (65% samples in the ShARC dataset), one rule sentence may contain several in-line conditions as exemplified in Figure 2. To extract these in-line conditions, we find discourse segmentation in discourse parsing to be useful. In the Rhetorical Structure Theory or RST Mann and Thompson (1988) of discourse parsing, texts are first split into a sequence of clause-like units called elementary discourse units (EDUs). We utilize an off-the-shelf discourse segmenter Li et al. (2018) to break the rule text into a sequence of EDUs. The segmenter uses a pointer network and achieves 92.2% F-score with Glove vectors and 95.55% F-score with ELMo embeddings on the standard RST benchmark testset, which is close to human agreement of 98.3% F-score Joty et al. (2015); Lin et al. (2019b). As exemplified in Figure 2 Step \raisebox{-0.9pt}{1}⃝, the rule sentence is broken into three EDUs, in which two conditions (“If a worker has taken more leave than they’re entitled to”, “unless it’s been agreed beforehand in writing”) and the outcome (“their employer must not take money from their final pay”) are split out precisely. For rule texts which contain bullet points, we directly treat these bullet points as conditions.

2 Decision Making via Entailment Reasoning

As shown in Figure 2 Step \raisebox{-0.9pt}{2}⃝, inputs to Discern include the segmented conditions (EDUs) in the rule text, user question, user scenario, and follow-up question-answer pairs in dialog history, each of which is a sequence of tokens. In order to get the sentence-level representations for all individual sequences, we insert an external [CLS] symbol at the start of each sequence, and add a [SEP] symbol at the end of every type of inputs. Then, Discern concatenates all sequences together, and uses RoBERTa Liu et al. (2019) to encode the concatenated sequence. The encoded [CLS] token represents the sequence that follows it. In this way, we extract sentence-level representations of conditions (EDUs) as e1,e2,...,eN\mathbf{e}_{1},\mathbf{e}_{2},...,\mathbf{e}_{N}, and also the representations of the user question uQ\mathbf{u}_{Q}, user scenario uS\mathbf{u}_{S}, and MM turns of dialog history u1,...,uM\mathbf{u}_{1},...,\mathbf{u}_{M}. All these vectorized representations are of dd dimensions (768 for RoBERTa-base).

Entailment Prediction.

In order to reason out the correct decision for the user question, it is necessary to figure out the fulfillment of conditions in the rule text. We propose to formulate the fulfillment prediction of conditions into a multi-sentence entailment task. Given a sequence of conditions (premises) and a sequence of user-provided information (hypotheses), a system should output Entailment, Contradiction or Neutral for each condition listed in the rule text. In this context, Neutral indicates that the condition has not been mentioned from the user information.

We utilize an inter-sentence transformer encoder Vaswani et al. (2017) to predict the entailment states for all conditions simultaneously. Taking all sentence-level representations [e1\mathbf{e}_{1}; e2\mathbf{e}_{2}; …; eN\mathbf{e}_{N}; uQ\mathbf{u}_{Q}; uS\mathbf{u}_{S}; u1\mathbf{u}_{1}; …; uM\mathbf{u}_{M}] as inputs, the LL-layer transformer encoder makes each condition attend to all the user-provided information to predict whether the condition is entailed or not. We also allow all conditions can attend to each other to understand the logical structure of the rule text.

where ci=[cE,i,cC,i,cN,i]∈3\mathbf{c}_{i}=[c_{\text{E},i},c_{\text{C},i},c_{\text{N},i}]\in 3 contains confidence scores of three entailment states Entailment, Contradiction, Neutral for the ii-th condition in the rule text.

Since there are no ground truth entailment labels for individual conditions, we adopt a heuristic approach similar to Gao et al. (2020) to get the noisy supervision signals. Given the rule text, we first collect all associated follow-up questions in the dataset. Each follow-up question is matched to a segmented condition (EDU) in the rule text which has the minimum edit distance. For conditions in the rule text which are mentioned by follow-up questions in the dialogue history, we label the entailment state of a condition as Entailment if the answer for its mentioned follow-up question is Yes, and label the state of this condition as Contradiction if the answer is No. The remaining conditions not covered by any follow-up question are labeled as Neutral. Let rr indicate the correct entailment state. The entailment prediction is weakly supervised by the following cross entropy loss, normalized by total number of KK conditions in a batch:

Decision Making.

After knowing the entailment state for each condition in the rule text, the remaining challenge for decision making is to perform logical reasoning over different rule types such as disjunction, conjunction, and conjunction of disjunctions. To achieve this, we first design three dd-dimension entailment vectors VE\mathbf{V}_{\text{E}} (Entailment), VC\mathbf{V}_{\text{C}} (Contradiction), VN\mathbf{V}_{\text{N}} (Neutral), and map the predicted entailment confidence scores of each condition to its vectorized entailment representation:

The overall loss for the Step \raisebox{-0.9pt}{2}⃝ decision making is the weighted-sum of decision loss and entailment prediction loss:

3 Follow-up Question Generation

If the predicted decision is “Inquire”, the follow-up question generation model is activated, as shown in Step \raisebox{-0.9pt}{3}⃝ of Figure 2. It extracts an underspecified span from the rule text which is uncovered from the user’s feedback, and rephrases it into a well-formed question. Existing approaches put huge efforts in extracting the underspecified span, such as entailment-driven extracting and ranking Zhong and Zettlemoyer (2019) or coarse-to-fine reasoning Gao et al. (2020). However, we find that such sophisticated modelings may not be necessary, and we propose a simple but effective approach here.

We split the rule text into sentences and concatenate the rule sentences and user-provided information into a sequence. Then we use RoBERTa to encode them into vectors grounded to tokens, as here we want to predict the position of a span within the rule text. Let [t1,1\mathbf{t}_{1,1}, …, t1,s1\mathbf{t}_{1,s_{1}}; t2,1\mathbf{t}_{2,1}, …, t2,s2\mathbf{t}_{2,s_{2}}; …; tN,1\mathbf{t}_{N,1}, …, tN,sN\mathbf{t}_{N,s_{N}}] be the encoded vectors for tokens from NN rule sentences, we follow the BERTQA approach Devlin et al. (2019) to learn a start vector ws∈d\mathbf{w}_{s}\in d and an end vector we∈d\mathbf{w}_{e}\in d to locate the start and end positions, under the restriction that the start and end positions must belong to the same rule sentence:

where i,ji,j denote the start and end positions of the selected span, and kk is the sentence which the span belongs to. The training objective is the sum of the log-likelihoods of the correct start and end positions. To supervise the span extraction process, the noisy supervision of spans are generated by selecting the span which has the minimum edit distance with the to-be-asked question. Lastly, following Gao et al. (2020), we concatenate the rule text and span as the input sequence, and finetune UniLM Dong et al. (2019), a pre-trained language model to rephrase it into a question.

Experiments

ShARC Saeidi et al. (2018) dataset is the current benchmark to test entailment reasoning in conversational machine reading Leaderboard: https://sharc-data.github.io/leaderboard.html. The dataset contains 948 rule texts clawed from 10 government websites, in which 65% of them are plain text with in-line conditions while the rest 35% contain bullet-point conditions. Each rule text is associated with a dialog tree (follow-up QAs) that considers all possible fulfillment combinations of conditions. In the data annotation stage, parts of the dialogs are paraphrased into the user scenario. These parts of dialogs are marked as evidence which should be extracted (entailed) from the user scenario, and are not provided as inputs for evaluation. The inputs to the system are the rule text, user question, user scenario, and dialog history (if any). The output is the answer among Yes, No, Irrelevant, or a follow-up question. The train, development, and test dataset sizes are 21890, 2270, and 8276, respectively.

Evaluation Metrics.

The decision making sub-task uses macro- and micro- accuracy of four classes “Yes”, “No”, “Irrelevant”, “Inquire” as metrics. For the question generation sub-task, we evaluate models under both the official end-to-end setting Saeidi et al. (2018) and the recently proposed oracle setting Gao et al. (2020). In the official setting, the BLEU score Papineni et al. (2002) is calculated only when both the ground truth decision and the predicted decision are “Inquire”, which makes the score dependent on the model’s “Inquire” predictions. For the oracle question generation setting, models are asked to generate a question when the ground truth decision is “Inquire”.

Implementation Details.

For the decision making sub-task, we finetune RoBERTa-base model Wolf et al. (2019) with Adam Kingma and Ba (2015) optimizer for 5 epochs with a learning rate of 5e-5, a warm-up rate of 0.1, a batch size of 16, and a dropout rate of 0.35. The number of inter-sentence transformer layers LL and the loss weight λ\lambda for entailment prediction are hyperparameters. We try 1,2,3 for LL and 1.0, 2.0, 3.0, 4.0, 5.0 for λ\lambda, and find the best combination is L=2,λ=3.0L=2,\lambda=3.0, based on the development set results. For the question generation sub-task, we train a RoBERTa-base model to extract spans under the same training scheme above, and finetune UniLM Dong et al. (2019) 20 epochs for question rephrasing with a batch size of 16, a learning rate of 2e-5, and a beam size 10 for decoding in the inference stage. We repeat 5 times with different random seeds for all experiments on the development set and report the average results along with their standard deviations. It takes two hours for training on a 4-core server with an Nvidia GeForce GTX Titan X GPU.

2 Results

The decision making results in macro- and micro- accuracy on the blind, held out test set of ShARC are shown in Table 1. Discern outperforms the previous best model EMT Gao et al. (2020) by 3.8% in micro-averaged accuracy and 3.5% in macro-averaged accuracy. We further analyze the class-wise decision prediction accuracy on the development set of ShARC in Table 2, and find that Discern have far better predictions than all existing approaches whenever a decision on the user’s fulfillment is needed (“Yes”, “No”, “Inquire”). It is because the predicted decisions from Discern are made upon the predicted entailment states while previous approaches do not build the connection between them.

Question Generation Sub-task.

Discern outperforms existing methods under both the official end-to-end setting (Table 1) and the recently proposed oracle setting (Table 3). Because the comparison among models is only fair under the oracle question generation setting Gao et al. (2020), we compare Discern with E3 Zhong and Zettlemoyer (2019), E3+UniLM Gao et al. (2020), EMT Gao et al. (2020), and our ablation Discern (BERT) in Table 3. Interestingly, we find that, in this oracle setting, our proposed simple approach is even better than previous sophisticated models such as E3 and EMT which jointly learn question generation and decision making via multi-task learning. From our results and investigations, we believe the decision making sub-task and the follow-up question generation sub-task do not share too many commonalities so the results are not improved for each task in their multi-task training. On the other hand, our question generation model is easy to optimize because this model is separately trained from the decision making one, which means there is no need to balance the performance between these two sub-tasks. Besides, RoBERTa backbone performs comparably with its BERT counterpart.

In our detailed analyses, we find Discern can locate the next questionable sentence with 77.2% accuracy, which means Discern utilizes the user scenario and dialog history well to locate the next underspecified condition. We try to add entailment prediction supervision to help Discern to locate the unfulfilled condition but it does not help. We also try to simplify our approach by directly finetuning UniLM to learn the mapping between concatenated input sequences and the follow-up clarification questions. However, the poor result (around 40 for BLEU1) suggests this direction still remains further investigations.

3 Ablation Study

Table 4 shows an ablation study of Discern for the decision making sub-task on the development set of ShARC, and we have the following observations:

Discern (BERT) replaces the RoBERTa backbone with BERT while other modules remain the same. The better performance of RoBERTa backbone matches findings from Talmor et al. (2019), which indicate that RoBERTa can capture negations and handle conjunctions of facts better than BERT.

Discourse Segmentation vs. Sentence Splitting.

Discern (w/o EDU) replaces the discourse segmentation based rule parsing with simple sentence splitting, and we observe there is a 1.63% drop on the micro-accuracy. This is intuitive because we observe 65% of the rule texts in the training set contains in-line conditions. To better understand the effect of discourse segmentation, we also evaluate Discern and Discern (w/o EDU) on just that portion of examples that contains multiple EDUs. The micro-accuracy of decision making is 75.75 for Discern while it is 70.98 for Discern (w/o EDU). The significant gap shows that discourse segmentation is extremely helpful.

Are Inter-sentence Transformer Layers Necessary?

We investigate the necessity of inter-sentence transformer layers because RoBERTa-base already has 12 transformer layers, in which the sentence-level [CLS] representations can also interact with each other via multi-head self-attention. Therefore, we remove the inter-sentence transformer layers and use the RoBERTa encoded [CLS] representations for entailment prediction and decision making. The results show that removing the inter-sentence transformer layers (Discern w/o Trans) hurts the performance, which suggests that the inter-sentence self-attention is essential.

Both Condition Representations and Entailment Vectors Facilitate Decisions.

4 Analysis of Logical Structure of Rules

To see how Discern understands the logical structure of rules, we evaluate the decision making accuracy according to the logical types of rule texts. Here we define four logical types: “Simple”, “Conjunction”, “Disjunction”, “Other”, which are inferred from the associated dialog trees. “Simple” means there is only one requirement in the rule text while “Other” denotes the rule text have complex logical structures, for example, a conjunction of disjunctions or a disjunction of conjunctions. Table 5 shows decision prediction results categorized by different logical structures of rules. Discern achieves the best performance on the “Simple” logical type which only needs to determine the single condition is satisfied or not. On the other hand, Discern does not perform well on rules in the format of disjunctions. We conduct further analysis on this category and find that the error comes from user scenario interpretation: the user has already provided his fulfillment in the user scenario but Discern fails to extract it. Detailed analyses are further conducted in the following section.

5 How Far Has the Problem Been Solved?

In order to figure out the limitations of Discern, and the current challenges of ShARC CMR, we disentangle the challenges of scenario interpretation and dialog understanding in ShARC by selecting different subsets, and evaluate decision making and entailment prediction accuracy on them.

Because the classification for unanswerable questions (“irrelevant” class) is nearly solved (99.3% in Table 2), we create the baseline subset by removing all unanswerable examples from the development set. Results for this baseline are shown in ShARC (Answerable) of Table 6.

Dialog History Subset.

We first want to see how Discern understands dialog histories (follow-up QAs) without the influence of user scenarios. Hence, we create a subset of ShARC (Answerable) in which all samples have an empty user scenario. The performance over 224 such samples is shown in “Dialog History Subset” of Table 6. Surprisingly, the results on this portion of samples are much better than the overall results, especially for the entailment prediction (92.41% micro-accuracy).

Scenario Subset.

With the curiosity to see what is the bottleneck of our model, we test the model ability on scenario interpretation. Similarly, we create a “Scenario Subset” from ShARC (Answerable) in which all samples have an empty dialog history. Results in Table 6 (“Scenario Subset”) show that interpreting scenarios to extract the entailment information within is exactly the current bottleneck of Discern. We analyze 100 error cases on this subset and find that various types of reasoning are required for scenario interpretation, including numerical reasoning (15%), temporal reasoning (12%), and implication over common sense and external knowledge (46%). Besides, Discern still fails to extract user’s fulfillment when the scenarios paraphrase the rule texts (27%). Examples for each type of error are shown in Figure 3.5. Among three classes of entailment states, we find that Discern fails to predict Entailment or Contradiction precisely – it predicts Neutral in most cases for scenario interpretation, resulting in high micro-accuracy in entailment prediction but the macro-accuracy is poor. The decision accuracy is subsequently hurt by the entailment results.

ShARC (Evidence).

Based on the above observation, we replace the user scenario in the ShARC (Answerable) by its evidence and re-evaluate the overall performance on these answerable questions. As described in Section 3.1 Dataset, the evidence is the part of dialogs that should be entailed from the user scenario. Table 6 shows that the model improves 11.38% in decision making micro-accuracy if no scenario interpretation is required, which validates our above observation.

Understanding entailments (or implications) of text is essential in dialog and question answering systems. ROPES Lin et al. (2019a) requires reading descriptions of causes and effects and applying them to situated questions, while ShARC Saeidi et al. (2018), the focus of Discern, requires to understand rules and apply them to questions asked by users in a conversational manner. Most existing methods simply use BERT to classify the answer without considering the structures of rule texts Zhong and Zettlemoyer (2019); Sharma et al. (2019); Lawrence et al. (2019). Gao et al. (2020) propose Explicit Memory Tracker (EMT), which firstly addresses entailment-oriented reasoning. At each dialog turn, EMT recurrently tracks whether conditions listed in the rule text have already been satisfied to make a decision.

In this paper, we also explicitly model entailment reasoning for decision making, but there are three key differences between our Discern and EMT: (1) we apply discourse segmentation to parse the rule text, which is extremely helpful because there are many in-line conditions in rules; (2) Our stacked inter-sentence transformer layers extract better features for entailment prediction, which could be seen as a generalization of their recurrent explicit memory tracker. (3) Different from their utilization of entailment prediction which is treated as multi-task learning for decision making, we directly build the dependency between entailment prediction states and the predicted decisions.

Discourse Applications.

Discourse analysis uncovers text-level linguistic structures (e.g., topic, coherence, co-reference), which can be useful for many downstream applications, such as coherent text generation Bosselut et al. (2018) and text summarization Joty et al. (2019); Cohan et al. (2018); Xu et al. (2020). Recently, discourse information has also been introduced in neural reading comprehension. Mihaylov and Frank (2019) design a discourse-aware semantic self-attention mechanism to supervise different heads of the transformer by discourse relations and coreferring mentions. Different from their use of discourse information, we use it as a parser to segment surface-level in-line conditions for entailment reasoning.

In this paper, we present Discern, a system that does discourse-aware entailment reasoning for conversational machine reading. Discern explicitly builds the connection between entailment states of conditions and the final decisions. Results on the ShARC benchmark shows that Discern outperforms existing methods by a large margin. We also conduct comprehensive analyses to unveil the limitations of Discern and challenges for ShARC. In future, we plan to explore how to incorporate discourse parsing into the current decision making model for end-to-end learning. One possibility would be to frame them as multi-task learning with a common (shared) encoder. Another direction is leveraging current methods in question generation Gao et al. (2019); Li et al. (2019) to improve the follow-up question generation sub-task since Discern is on par with the previous best model EMT.

We thank Max Bartolo and Patrick Lewis for evaluating our submitted models on the hidden test set. The work described in this paper was partially supported by following projects from the Research Grants Council of the Hong Kong Special Administrative Region, China: CUHK 2300174 (Collaborative Research Fund, No. C5026-18GF); CUHK 14210717 (RGC General Research Fund).