Revealing the Myth of Higher-Order Inference in Coreference Resolution

Liyan Xu, Jinho D. Choi

Introduction

Coreference resolution has always been considered one of the unsolved NLP tasks due to its challenging aspect of document-level understanding Wiseman et al. (2015, 2016); Clark and Manning (2015, 2016); Lee et al. (2017). Nonetheless, it has made a tremendous progress in recent years by adapting contextualized embedding encoders such as ELMo Lee et al. (2018); Fei et al. (2019) and BERT Kantor and Globerson (2019); Joshi et al. (2019, 2020). The latest state-of-the-art model shows the improvement of 12.4% over the model introduced 2.5 years ago, where the major portion of the improvement is derived by representation learning (Figure 1).

Most of these previous models have also adapted higher-order inference (HOI) for the global optimization of coreference links, although HOI clearly has not been the focus of those works, for the fact that gains from HOI have been reported marginal. This has inspired us to analyze the impact of HOI on modern coreference resolution models in order to envision the future direction of this research.

To make thorough ablation studies among different approaches, we implement an end-to-end coreference system in PyTorch (Sec 3.1), and two HOI approaches proposed by previous work, attended antecedent and entity equalization (Sec 3.2), along with two of our original approaches, span clustering and cluster merging (Sec 3.3). These approaches are experimented with two Transformer encoders, BERT and SpanBERT, to assess how effective HOI is even when coupled with those high-performing encoders (Sec 4). To the best of our knowledge,this is the first work to make a comprehensive analysis on multiple HOI approaches side-by-side for the task of coreference resolution.Source codes and models are available at https://github.com/lxucs/coref-hoi.

Related Work

Most neural network-based coreference resolution models have adapted antecedent-ranking Wiseman et al. (2015); Clark and Manning (2015); Lee et al. (2017, 2018); Joshi et al. (2019, 2020), which relies on the local decisions between each mention and its antecedents. To achieve deeper global optimization,Wiseman et al. (2016); Clark and Manning (2016); Yu et al. (2020) built entity representations in the ranking process, whereas Lee et al. (2018); Kantor and Globerson (2019) refined the mention representation by aggregating its antecedents’ information.

It is no secret that the integration of contextualized embeddings has played the most critical role in this task. While the following are based on the same end-to-end coreference model Lee et al. (2017), Lee et al. (2018); Fei et al. (2019) reported 3.3% improvement by adapting ELMo in the encoders Peters et al. (2018). Kantor and Globerson (2019); Joshi et al. (2019) gained additional 3.3% by adapting BERT Devlin et al. (2019). Joshi et al. (2020) introduced SpanBERT that gave another 2.7% improvement over Joshi et al. (2019).

Most recently, Wu et al. (2020) proposes a new model that adapts question-answering framework on coreference resolution, and achieves state-of-the-art result of 83.1 on the CoNLL’12 shared task.

Approach

We reimplement the end-to-end c2f-coref model introduced by Lee et al. (2018) that has been adapted by every coreference resolution model since then. It detects mention candidates through span enumeration and aggressive pruning. For each candidate span xx, the model learns the distribution over its antecedents y∈Y(x)y\in\mathcal{Y}(x):

where s(x,y)s(x,y) is the local score involving two parts: how likely the spans xx and yy are valid mentions, and how likely they refer to the same entity:

gx,gyg_{x},g_{y} are the span embeddings of xx and yy, ϕ(x,y)\phi(x,y) is the meta-information (e.g., speakers, distance), and wm,wcw_{m},w_{c} are the mention and coreference scores, respectively (FFNN: feedforward neural network).

We use different Transformers-based encoders, and follow the “independent” setup for long documents as suggested by Joshi et al. (2019).

2 Span Refinement

Two HOI methods presented by recent coreference work are based on span refinement that aggregates non-local features to enrich the span representation with more “global” information. The updated span representation gx′g_{x}^{\prime} can be derived as in Eq. 3, where gx′g_{x}^{\prime} is the interpolation between the current and refined representation gxg_{x} and axa_{x}, and WfW_{f} is the gate parameter. gx′g_{x}^{\prime} is used to perform another round of antecedent-ranking in replacement of gxg_{x}.

The following two methods share the same updating process for gx′g_{x}^{\prime}, but with different ways to obtain the refined span representation axa_{x}.

takes the antecedent information to enrich gx′g_{x}^{\prime} (Lee et al., 2018; Fei et al., 2019; Joshi et al., 2019, 2020). The refined span axa_{x} is the attended antecedent representation over the current antecedent distribution P(y)P(y), where gy∈Y(x)g_{y\in\mathcal{Y}(x)} is the antecedent representation:

Entity Equalization (EE)

takes the clustering relaxation as in Eq. 3.2 to model the entity distribution (Kantor and Globerson, 2019), where Q(x∈Ey′)Q(x\in E_{y^{\prime}}) is the probability of the span xx referring to an entity Ey′E_{y^{\prime}} in which the span y′y^{\prime} is the first mention. P(y)P(y) is the current antecedent distribution.

The refined span axa_{x} is the attended entity representation, where ey(x)e_{y}^{(x)} is the entity representation to which the span yy belongs till the span xx:

3 HOI with Clustering

This section introduces two new HOI methods for a more extensive study in HOI.

is also based on span refinement, and it constructs the actual clusters and obtains the “true” predicted entities using P(y)P(y) instead of modeling the “soft” entity clusters through the relaxation as in EE (Section 3.2). This way, although we lose the differentiable property, the obtaining of true entities with the same empirical inference time as EE has made SC desirable.

The entity representation eie_{i} for an entity cluster CiC_{i} is given by the attended spans in this cluster:

The entity clusters CiC_{i} are constructed in the same way as in the final cluster prediction. The refined span axa_{x} is then equal to the representation of entity eie_{i} to which it belongs (gx∈Cig_{x}\in C_{i}).

Cluster Merging (CM)

performs sequential antecedent ranking combining both antecedent and entity information to gradually build up the entity clusters, which is distinguished from span refinement methods that simply re-rank antecedents. Algorithm 1 describes the ranking process for CM. gig_{i} is the ii’th span, Y(i)\mathcal{Y}(i) is the indices of gig_{i}’s antecedents, and CiC_{i} is the cluster that gig_{i} belongs to. The ranking score sx(y)s_{x}(y) consists of both antecedent score faf_{a} (see Eq. 2) and cluster score fcf_{c}. To avoid overlapping between faf_{a} and fcf_{c}, we set fcf_{c} as if the cluster is the initial cluster (L6). Thus, fcf_{c} becomes the consultation such that when fc>0f_{c}>0, the span gxg_{x} is likely to match the cluster CyC_{y}, and vice versa. fcf_{c} is computed by FFNN similar to faf_{a}, and ϕ(Cy)\phi(C_{y}) is the meta-feature such as the cluster size.

Two simple configurations can be tuned for CM. We can have the sequential left-to-right ranking order or the easy-first order (L3) whose sequence is ordered by each span’s max antecedent score, building the most confident clusters first (Ng and Cardie, 2002; Clark and Manning, 2016). There can be element-wise mean or max-reduction among the spans in the two merging clusters (L10).

Distinguished from Wiseman et al. (2016), clusters in CM are searched and merged in training without the use of oracle clusters, closing the gap between training and test time.

Experiments

For our experiments, the CoNLL 2012 English shared task dataset is used (Pradhan et al., 2012). Given the end-to-end coreference system in Section 3.1, six models are developed as follows:Appdendix A.1 provides details of our experimental settings.

BERT: BERT Devlin et al. (2019) as the encoder

SpanBERT: SpanBERT Joshi et al. (2020) as the encoder

+AA: SpanBERT with attended antecedent (§3.2)

+EE: SpanBERT with entity equalization (§3.2)

+SC: SpanBERT with span clustering (§3.3)

+CM: SpanBERT with cluster merging (§3.3)

Note that BERT and SpanBERT completely rely on only local decisions without any HOI. Particularly, +AA is equivalent to Joshi et al. (2020).

Table 1 shows the best results in comparison to previous state-of-the-art systems. We also report the mean scores and standard deviations from 5 repeated developments, which we could not find from the previous works.

The impact of SpanBERT over BERT is clear, showing 2.4% improvement on average. However, none of the HOI models shows a clear advantage over SpanBERT which adapts no HOI. In fact, all HOI models except for CM show negative impact. The best result is achieved by CM with the Avg-F1 of 80.2, surpassing the previous best result of 79.6 based on c2f-coref reported by Joshi et al. (2020).

2 Impact Analysis of HOI

Three HOI methods based on span refinement, AA, EE, and SC, show negative impact upon local decisions. We suspect that error propagation from antecedent-ranking may downgrade the quality of refinement. On the other hand, CM shows marginal improvement, suggesting that maintaining entity clusters can be superior to span refinement, at the cost of more inference time from the sequential ranking process. To analyze the direct impact of HOI, we take the trained models of each HOI method and evaluate them on the test set while turning off HOI, making it compatible to SpanBERT.

The averaged performance drop w.r.t Avg-F1 after turning off HOI is less than 0.2 for all methods(Appendix A.3), implying that none of the HOI method has a significantly direct impact to the final performance of the model using SpanBERT.

In further investigation, we examine the change of coreferent links w.r.t their correctness. Specifically, Table 2 shows the four types of link changes before and after HOI. It demonstrates that the benefits from HOI is diminished because the effects are two-sided: there are roughly same amounts of links (about 1%) becoming correct or wrong after HOI, therefore neither HOI method leads to much improvement overall.

It is worth mentioning that the impact of HOI is not limited to only global decisions. HOI implicitlyserves as a way of regularization that impacts local decisions as well, since HOI and local ranking are mutually dependent during training. Such indirect influence of HOI makes it difficult to assess its true impact, which we will explore more in the future.

3 Analysis of Pronoun Resolution

For the error analysis, we examine the direct inference between two personal pronouns.Ambiguous pronouns such as “you” are excluded in direct inference analysis, and included in indirect inference analysis. SP/PS in Table 3 shows the numbers of links that one pronoun incorrectly selects another pronoun with different plurality as its antecedent. We find that adapting HOI shows slightly higher impact than switching to a more advanced encoder. AA can reinforce the pronoun representation to bias towards singularity and lead to lower SP error and higher PS error, while the difference between BERT and SpanBERT is trivial on SP/PS.

We also look at the general types of coreferent errors involving two pronouns. False Link (FL) falsely links a non-anaphoric pronoun to another pronoun as antecedent; Wrong Link (WL) links an anaphoric pronoun to another wrong pronoun as antecedent. Table 3 shows that EE and CM reduce FL errors by 4+%, suggesting that the aggregation of non-local features indeed leads to more conservative linking decisions. However, adapting an advanced encoder shows higher impact on WL errors, as SpanBERT reduces almost 10% compared to BERT, implying that representation learning is still more important for semantic matching.

Indirect Inference

The plurality of ambiguous pronouns such as you depends on the context. Two indirect links of (he, you) and (you, they) can be common to induce incorrect clusters that contain both singular and plural pronouns (Wiseman et al., 2016; Lee et al., 2018). Table 3 shows the numbers of these erroneous clusters in prediction. Surprisingly, very few of these clusters contain ambiguous pronouns in either approach. This observation moderates the long-standing movitation of HOI.

Additionally, the change of representation from BERT to SpanBERT has far more impact that reduces 10% of these erroneous clusters, while the four HOI methods fail to show significant difference compared to SpanBERT.

Conclusion

We implement the end-to-end coreference resolution model and investigate four higher-order inference methods, including two of our own methods. Our best model shows the new result of 80.2 on the CoNLL 2012 dataset. We thoroughly analyze the empirical effectiveness of HOI and demonstrate why it fails to boost performance on the CoNLL 2012 dataset compared to the improvement from encoders. We show that current HOI does not meet up with the original motivation, suggesting that a new perspective of HOI is needed for this task in the era of deep learning-based NLP.

Acknowledgments

We gratefully acknowledge the support of the AWS Machine Learning Research Awards (MLRA). Any contents in this material are those of the authors and do not necessarily reflect the views of AWS.

References

Appendix A Appendices

We implement the experimented models using PyTorch. BERTLarge and SpanBERTLarge are used as encoders. For each experiment, the best performed model on the development set is selected and evaluated on the test set.

Similar to Joshi et al. (2019, 2020), documents are split into independent segments with maximum 384 word pieces for BERTLarge and 512 for SpanBERTLarge. In our final setting, BERT-parameters and task-parameters have separate learning rates (1×10−51\times 10^{-5} and 3×10−43\times 10^{-4} respectively), separate linear decay schedule, and separate weight decay rates (10−210^{-2} and respectively). Models are trained 24 epochs with dropout rate 0.3.

The implementation of EE is based on the Tensorflow implementation from Kantor and Globerson (2019) which requires O(k2)\mathcal{O}(k^{2}) memory with kk being the number of extracted spans, while other HOI approaches only requires O(k)\mathcal{O}(k) memory The maximum number of antecedents for all models is set to 50 which is constant.. To keep the GPU memory usage within 32GB, we limit the maximum number of span candidates for EE to be 300, which may have a negative impact on the performance.

Experiments are conducted on Nvidia Tesla V100 GPUs with 32GB memory. The average training time is around 7 hours for BERT and SpanBERT without HOI, and ranges from 9 - 15 hours with HOI methods.

A.2 Results

Table 4 reports the macro-average F1 scores out of 5 repeated developments of each approach. CM still has the best performance with 79.9 averaged F1 score. Span refinement-based HOI approaches, AA, EE, and SC, still have lower F1 scores than the local-only SpanBERT.

We do not find different configurations for CM make any huge impact to the performance. The final configuration for CM is sequential order and max reduction (Algorithm 1).

A.3 Analysis

Table 5 shows the averaged performance drop and its standard deviations w.r.t Avg-F1 after turning off the corresponding HOI in trained models, to see the direct performance impact of HOI over local decisions.

In our analysis, the following personal pronouns are regarded as ambiguous pronouns: “you”, “your”, “yours”.