Coreference Resolution as Query-based Span Prediction
Wei Wu, Fei Wang, Arianna Yuan, Fei Wu, Jiwei Li
Introduction
Recent coreference resolution systems (Lee et al., 2017, 2018; Zhang et al., 2018a; Kantor and Globerson, 2019) consider all text spans in a document as potential mentions and learn to find an antecedent for each possible mention. There are two key issues with this paradigm, in terms of task formalization and the algorithm.
At the task formalization level, mentions left out at the mention proposal stage can never be recovered since the downstream module only operates on the proposed mentions. Existing models often suffer from mention proposal (Zhang et al., 2018a). The coreference datasets can only provide a weak signal for spans that correspond to entity mentions because singleton mentions are not explicitly labeled. Due to the inferiority of the mention proposal model, it would be favorable if a coreference framework had a mechanism to retrieve left-out mentions.
At the algorithm level, existing end-to-end methods (Lee et al., 2017, 2018; Zhang et al., 2018a) score each pair of mentions only based on mention representations from the output layer of a contextualization model. This means that the model lacks the connection between mentions and their contexts. Semantic matching operations between two mentions (and their contexts) are performed only at the output layer and are relatively superficial. Therefore it is hard for their models to capture all the lexical, semantic and syntactic cues in the context.
To alleviate these issues, we propose CorefQA, a new approach that formulates the coreference resolution problem as a span prediction task, akin to the question answering setting. A query is generated for each candidate mention using its surrounding context, and a span prediction module is further employed to extract the text spans of the coreferences within the document using the generated query. Some concrete examples are shown in Figure 1. This is an illustration of the question formulation. The actual operation is described in Section 3.4.
This formulation provides benefits at both the task formulation level and the algorithm level. At the task formulation level, since left-out mentions can still be retrieved at the span prediction stage, the negative effect of undetected mentions is significantly alleviated. At the algorithm level, by generating a query for each candidate mention using its surrounding context, the CorefQA model explicitly considers the surrounding context of the target mentions, the influence of which will later be propagated to each input word using the self-attention mechanism. Additionally, unlike existing end-to-end methods (Lee et al., 2017, 2018; Zhang et al., 2018a), where the interactions between two mentions are only superficially modeled at the output layer of contextualization, span prediction requires a more thorough and deeper examination of the lexical, semantic and syntactic cues within the context, which will potentially lead to better performance.
Moreover, the proposed question answering formulation allows us to take advantage of existing question answering datasets. Coreference annotation is expensive, cumbersome and often requires linguistic expertise from annotators. Under the proposed formulation, the coreference resolution has the same format as the existing question answering datasets (Rajpurkar et al., 2016a, 2018; Dasigi et al., 2019a). Those datasets can thus readily be used for data augmentation. We show that pre-training on existing question answering datasets improves the model’s generalization and transferability, leading to additional performance boost.
Experiments show that the proposed framework significantly outperforms previous models on two widely-used datasets. Specifically, we achieve new state-of-the-art scores of 83.1 (+3.5) on the CoNLL-2012 benchmark and 87.5 (+2.5) on the GAP benchmark.
Related Work
Coreference resolution is a fundamental problem in natural language processing and is considered as a good test of machine intelligence (Morgenstern et al., 2016). Neural network models have shown promising results over the years. Earlier neural-based models (Wiseman et al., 2016; Clark and Manning, 2015, 2016) rely on parsers and hand-engineered mention proposal algorithms. Recent work (Lee et al., 2017, 2018; Kantor and Globerson, 2019) tackled the problem in an end-to-end fashion by jointly detecting mentions and predicting coreferences. Based on how entity-level information is incorporated, they can be further categorized as (1) entity-level models Björkelund and Kuhn (2014); Clark and Manning (2015, 2016); Wiseman et al. (2016) that directly model the representation of real-world entities and (2) mention-ranking models (Durrett and Klein, 2013; Wiseman et al., 2015; Lee et al., 2017) that learn to select the antecedent of each anaphoric mention. Our CorefQA model is essentially a mention-ranking model, but we identify coreference using question answering.
2 Formalizing NLP Tasks as question answering
Machine reading comprehension is a general and extensible task form. Many tasks in natural language processing can be framed as reading comprehension while abstracting away the task-specific modeling constraints.
McCann et al. (2018) introduced the decaNLP challenge, which converts a set of 10 core tasks in NLP to reading comprehension. He et al. (2015) showed that semantic role labeling annotations could be solicited by using question-answer pairs to represent the predicate-argument structure. Levy et al. (2017) reduced relation extraction to answering simple reading comprehension questions, yielding models that generalize better in the zero-shot setting. Li et al. (2019a, b) cast the tasks of named entity extraction and relation extraction as a reading comprehension problem. In parallel to our work, Aralikatte et al. (2019) converted coreference and ellipsis resolution in a question answering format, and showed the benefits of training joint models for these tasks. Their models are built under the assumption that gold mentions are provided at inference time, whereas our model does not need that assumption – it jointly trains the mention proposal model and the coreference resolution model in an end-to-end manner. Chada (2019) proposed an extractive QA model for the resolution of ambiguous pronouns, and showed better results on the GAP dataset by only fine-tuning the pre-trained BERT model.
3 Data Augmentation
Data augmentation is a strategy that enables practitioners to significantly increase the diversity of data available for training models. Data augmentation techniques have been explored in various fields such as question answering (Talmor and Berant, 2019), text classification (Kobayashi, 2018) and dialogue language understanding (Hou et al., 2018). In coreference resolution, Zhao et al. (2018); Emami et al. (2019); Zhao et al. (2019) focused on debiasing the gender bias problem; Aralikatte et al. (2019) explored the effectiveness of joint modeling of ellipsis and coreference resolution. To the best of our knowledge, we are the first to use existing question answering datasets as data augmentation for coreference resolution.
Model
In this section, we describe our CorefQA model in detail. The overall architecture is illustrated in Figure 2.
Given a sequence of input tokens in a document, where denotes the length of the document. denotes the number of all possible text spans in . Let denotes the -th span representation , with the start index first(i) and the end index last(i). . The task of coreference resolution is to determine the antecedents for all possible spans. If a candidate span does not represent an entity mention or is not coreferent with any other mentions, a dummy token is assigned as its antecedent. The linking between all possible spans defines the final clustering.
2 Input Representations
We use the SpanBERT model https://github.com/facebookresearch/SpanBERT to obtain input representations following Joshi et al. (2019a). Each token is associated with a SpanBERT representation . Since the speaker information is indispensable for coreference resolution, previous methods Wiseman et al. (2016); Lee et al. (2017); Joshi et al. (2019a) usually convert the speaker information into binary features indicating whether two mentions are from the same speaker. However, we use a straightforward strategy that directly concatenates the speaker’s name with the corresponding utterance. This strategy is inspired by recent research in personalized dialogue modeling that use persona information to represent speakers (Li et al., 2016; Zhang et al., 2018b; Mazaré et al., 2018). In subsection 5.2, we will empirically demonstrate its superiority over the feature-based method in Lee et al. (2017).
To fit long documents into SpanBERT, we use a sliding-window approach that creates a -sized segment after every /2 tokens. Segments are then passed to the SpanBERT encoder independently. The final token representations are derived by taking the token representations with maximum context.
3 Mention Proposal
Similar to Lee et al. (2017), our model considers all spans up to a maximum length as potential mentions. To improve computational efficiency, we further prune the candidate spans greedily during both training and evaluation. To do so, the mention score of each candidate span consists of three parts: (1) is the start of a span; (2) is the end of a span; and (3) and form a valid span. The third part (i.e., (3)) is necessary because each sentence can contain multiple spans. The first part is computed by feeding into a feed-forward layer:
Similarly, the first part is computed by feeding into a feed-forward layer:
The third part computed by feeding the concatenation of and into into a feed-forward layer:
) denotes the feed-forward neural network that computes a nonlinear mapping from the input vector to the mention score. The three involved ) use separate sets of parameters. The overall score for span being a mention is the average of the three parts:
We only keep up to (where is the document length) spans with the highest mention scores.
It is crucial that the mention proposal model is pretrained. Otherwise, most of the proposed mentions that are fed to the linking stage are invalid mentions. The mention proposal model is pretrained by jointly training three binary classification models: (1) whether is the start of a span; (2) whether is the end of a span; and (3) whether and should be combined. This leads to the objective of the mention proposal model as follows:
4 Mention Linking as Span Prediction
Given a mention proposed by the mention proposal network, the role of the mention linking network is to give a score for any text span , indicating whether and are coreferent. We propose to use the question answering framework as the backbone to compute . It operates on the triplet {context (X), query (q), answers (a)}. The context is the input document. The query is constructed as follows: given , we use the sentence that resides in as the query, with the minor modification that we encapsulates with special tokens . The answers are the coreferent mentions of . A query is considered unanswerable in the following scenarios: (1) the candidate span does not represent an entity mention or (2) the candidate span represents an entity mention but is not coreferent with any other mentions in .
Following Devlin et al. (2019), we represent the input query and the context as a single packed sequence. The for any span , we first compute the score of being the answer for query , denoted by . Let and respectively denote the representations for first(j) and last(j) from BERT, where is used as query concatenated to the context. is computed by feeding the first and the last of its constituent token representations (i.e., and ) into a feed-forward layer:
denotes the feed-forward neural network that computes a nonlinear mapping from the input vector to the mention score. Comparing Eq.8 with Eq.3, we can observe their relatedness and difference: both of the equations compute scores for a span. But for Eq.8, the query is additionally used to check whether span is the answer for .
A closer look at Eq.8 reveals that it only models the uni-directional coreference relation from to , i.e., is the answer for query . This is suboptimal since if is a coreference mention of , then should also be the coreference mention . We thus need to optimize the bi-directional relation between and .This bidirectional relationship is actually referred to as mutual dependency and has shown to benefit a wide range of NLP tasks such as machine translation Hassan et al. (2018) or dialogue generation Li et al. (2015). The final score is thus given as follows:
can be computed in the same way as , in which is used as the query:
where and respectively denote the representations for first(i) and last(i) from BERT, where is used as query concatenated to the context.
For a pair of text span and , the premises for them being coreferent mentions are (1) they are mentions and (2) they are coreferent. This makes the overall score for and the combination of Eq.3 and Eq.7:
is the hyperparameter to control the tradeoff between mention proposal and mention linking.
5 Antecedent Pruning
Given a document with length and the number of spans , the computation of Eq.9 for all mention pairs is intractable with the complexity of . Given an extracted mention , the computation of Eq.9 for regarding all is still extremely intensive since the computation of the backward span prediction score requires running question answering models on all query . A further pruning procedure is thus needed: For each query , we collect span candidates only based on the scores, and then use Eq. 9 to compute the overall scores.
6 Training
For each mention proposed by the mention proposal network, it is associated with potential spans proposed by the mention linking network based on , we aim to optimize the marginal log-likelihood of all correct antecedents implied by the gold clustering. Following Lee et al. (2017), we append a dummy token to the candidates. The model will output it if none of the span candidates is coreferent with . For each mention , the model learns a distribution over all possible antecedent spans based on the global score from Eq. 9:
The mention proposal module and the mention linking module are jointly trained in an end-to-end fashion using training signals from Eq.10, with the SpanBERT parameters shared.
7 Inference
Given an input document, we can obtain an undirected graph using the overall score, each node of which represents a candidate mention from either the mention proposal module or the mention linking module. We prune the graph by keeping the edge whose weight is the largest for each node based on Eq.10. Nodes whose closest neighbor is the dummy token are abandoned. Therefore, the mention clusters can be decoded from the graph.
8 Data Augmentation using Question Answering Datasets
We hypothesize that the reasoning (such as synonymy, world knowledge, syntactic variation, and multiple sentence reasoning) required to answer the questions are also indispensable for coreference resolution. Annotated question answering datasets are usually significantly larger than the coreference datasets due to the high linguistic expertise required for the latter. Under the proposed QA formulation, coreference resolution has the same format as the existing question answering datasets (Rajpurkar et al., 2016a, 2018; Dasigi et al., 2019a). In this way, they can readily be used for data augmentation. We thus propose to pre-train the mention linking network on the Quoref dataset Dasigi et al. (2019b), and the SQuAD dataset Rajpurkar et al. (2016b).
9 Summary and Discussion
Comparing with existing models Lee et al. (2017, 2018); Joshi et al. (2019b), the proposed question answering formalization has the flexibility of retrieving mentions left out at the mention proposal stage. However, since we still have the mention proposal model, we need to know in which situation missed mentions could be retrieved and in which situation they cannot. We use the example in Figure 1 as an illustration, in which {many people, They, themselves} are coreferent mentions: If partial mentions are missed by the mention proposal model, e.g., many people and They, they can still be retrieved in the mention linking stage when the not-missed mention (i.e., themselves) is used as query. But, if all the mentions within the cluster are missed, none of them can be used for query construction, which means they all will be irreversibly left out. Given the fact that the proposal mention network proposes a significant number of mentions, the chance that mentions within a mention cluster are all missed is relatively low (which exponentially decreases as the number of entities increases). This explains the superiority (though far from perfect) of the proposed model. However, how to completely remove the mention proposal network remains a problem in the field of coreference resolution.
Experiments
The special tokens used to denote the speaker’s name () and the special tokens used to denote the queried mentions () are initialized by randomly taking the unused tokens from the SpanBERT vocabulary. The sliding window size = 512, and the mention keep ratio = 0.2. The maximum length for mention proposal = 10 and the maximum number of antecedents kept for each mention = 50. The SpanBERT parameters are updated by the Adam optimizer (Kingma and Ba, 2015) with initial learning rate and the task parameters are updated by the Range optimizer https://github.com/lessw2020/Ranger-Deep-Learning-Optimizer with initial learning rate .
2 Baselines
We compare the CorefQA model with previous neural models that are trained end-to-end:
e2e-coref (Lee et al., 2017) is the first end-to-end coreference system that learns which spans are entity mentions and how to best cluster them jointly. Their token representations are built upon the GLoVe (Pennington et al., 2014) and Turian (Turian et al., 2010) embeddings.
c2f-coref + ELMo (Lee et al., 2018) extends Lee et al. (2017) by combining a coarse-to-fine pruning with a higher-order inference mechanism. Their representations are built upon ELMo embeddings (Peters et al., 2018).
c2f-coref + BERT-large(Joshi et al., 2019b) builds the c2f-coref system on top of BERT (Devlin et al., 2019) token representations.
EE + BERT-large (Kantor and Globerson, 2019) represents each mention in a cluster via an approximation of the sum of all mentions in the cluster.
c2f-coref + SpanBERT-large (Joshi et al., 2019a) focuses on pre-training span representations to better represent and predict spans of text.
3 Results on CoNLL-2012 Shared Task
The English data of CoNLL-2012 shared task (Pradhan et al., 2012) contains 2,802/343/348 train/development/test documents in 7 different genres. The main evaluation is the average of three metrics – MUC (Vilain et al., 1995), (Bagga and Baldwin, 1998), and (Luo, 2005) on the test set according to the official CoNLL-2012 evaluation scripts http://conll.cemantix.org/2012/software.html.
We compare the CorefQA model with several baseline models in Table 1. Our CorefQA system achieves a huge performance boost over existing systems: With SpanBERT-base, it achieves an F1 score of 79.9, which already outperforms the previous SOTA model using SpanBERT-large by 0.3. With SpanBERT-large, it achieves an F1 score of 83.1, with a 3.5 performance boost over the previous SOTA system.
4 Results on GAP
The GAP dataset (Webster et al., 2018) is a gender-balanced dataset that targets the challenges of resolving naturally occurring ambiguous pronouns. It comprises 8,908 coreference-labeled pairs of (ambiguous pronoun, antecedent name) sampled from Wikipedia.
We follow the protocols in Webster et al. (2018); Joshi et al. (2019b) and use the off-the-shelf resolver trained on the CoNLL-2012 dataset to get the performance of the GAP dataset. Table 2 presents the results. We can see that the proposed CorefQA model achieves state-of-the-art performance on all metrics on the GAP dataset.
Ablation Study and Analysis
We perform comprehensive ablation studies and analyses on the CoNLL-2012 development dataset. Results are shown in Table 3.
Replacing SpanBERT with vanilla BERT leads to a 3.5 F1 degradation. This verifies the importance of span-level pre-training for coreference resolution and is consistent with previous findings (Joshi et al., 2019a).
Effect of Pre-training Mention Proposal Network
Skipping the pre-training of the mention proposal network using golden mentions results in a 7.2 F1 degradation, which is in line with our expectation. A randomly initialized mention proposal model implies that mentions are randomly selected. Randomly selected mentions will mostly be transformed to unanswerable queries. This makes it hard for the question answering model to learn at the initial training stage, leading to inferior performance.
Effect of QA pre-training on the augmented datasets
One of the most valuable strengths of converting anaphora resolution to question answering is that existing QA datasets can be readily used for data augmentation purposes. We see a contribution of 0.7 F1 from pre-training on the Quoref dataset (Dasigi et al., 2019a) and a contribution of 0.3 F1 from pre-training on the SQuAD dataset (Rajpurkar et al., 2016a).
Effect of Question Answering
We aim to study the pure performance gain of the paradigm shift from mention-pair scoring to query-based span prediction. For this purpose, we replace the mention linking module with the mention-pair scoring module described in Lee et al. (2018), while others remain unchanged. We observe an 8.1 F1 degradation in performance, demonstrating the significant superiority of the proposed question answering framework over the mention-pair scoring framework.
2 Analyses on speaker modeling strategies
We compare our speaker modeling strategy (denoted by Speaker as input), which directly concatenates the speaker’s name with the corresponding utterance, with the strategy in Wiseman et al. (2016); Lee et al. (2017); Joshi et al. (2019a) (denoted by Speaker as feature), which converts speaker information into binary features indicating whether two mentions are from the same speaker. We show the average F1 scores breakdown by documents according to the number of their constituent speakers in Figure 3.
Results show that the proposed strategy performs significantly better on documents with a larger number of speakers. Compared with the coarse modeling of whether two utterances are from the same speaker, a speaker’s name can be thought of as speaker ID in persona dialogue learning Li et al. (2016); Zhang et al. (2018b); Mazaré et al. (2018). Representations learned for names have the potential to better generalize the global information of the speakers in the multi-party dialogue situation, leading to better context modeling and thus better results.
3 Analysis on the Overall Mention Recall
Since the proposed framework has the potential to retrieve mentions missed at the mention proposal stage, we expect it to have higher overall mention recall rate than previous models (Lee et al., 2017, 2018; Zhang et al., 2018a; Kantor and Globerson, 2019).
We examine the proportion of gold mentions covered in the development set as we increase the hyperparameter (the number of spans kept per word) in Figure 4. Our model consistently outperforms the baseline model with various values of . Notably, our model is less sensitive to smaller values of . This is because missed mentions can still be retrieved at the mention linking stage.
4 Qualitative Analysis
We provide qualitative analyses to highlight the strengths of our model in Table 4.
Shown in Example 1, by explicitly formulating the anaphora identification of the company as a query, our model uses more information from a local context, and successfully identifies Freddie Mac as the answer from a longer distance.
The model can also efficiently harness the speaker information in a conversational setting. In Example 3, it would be difficult to identify that [Thelma Gutierrez] is the correct antecedent of mention [I] without knowing that Thelma Gutierrez is the speaker of the second utterance. However, our model successfully identifies it by directly feeding the speaker’s name at the input level.
Conclusion
In this paper, we present CorefQA, a coreference resolution model that casts anaphora identification as the task of query-based span prediction in question answering. We showed that the proposed formalization can successfully retrieve mentions left out at the mention proposal stage. It also makes data augmentation using a plethora of existing question answering datasets possible. Furthermore, a new speaker modeling strategy can also boost the performance in dialogue settings. Empirical results on two widely-used coreference datasets demonstrate the effectiveness of our model. In future work, we will explore novel approaches to generate the questions based on each mention, and evaluate the influence of different question generation methods on the coreference resolution task.
Acknowledgement
We thank all anonymous reviewers for their comments and suggestions. The work is supported by the National Natural Science Foundation of China (NSFC No. 61625107 and 61751209).