Contextualizing Citations for Scientific Summarization using Word Embeddings and Domain Knowledge

Arman Cohan, Nazli Goharian

Introduction

In scientific literature, related work is often referenced along with a short textual description regarding that work which we call citation text. Citation texts usually highlight certain contributions of the referenced paper and a set of citation texts to a reference paper can provide useful information about that paper. Therefore, citation texts have been previously used to enhance many downstream tasks in IR/NLP such as search and summarization (e.g. (Ritchie et al., 2008; Qazvinian and Radev, 2008; Cohan and Goharian, 2015)).

While useful, citation texts might lack the appropriate context from the reference article (Teufel and Moens, 2002; de Waard and Maat, 2012; Cohan et al., 2015). For example, details of the methods, assumptions or conditions for the obtained results are often not mentioned. Furthermore, in many cases the citing author might misunderstand or misquote the referenced paper and ascribe contributions to it that are not intended in that form. Hence, sometimes the citation text is not sufficiently informative or in other cases, even inaccurate (Sándor and De Waard, 2012). This problem is more serious in life sciences where accurate dissemination of knowledge has direct impact on human lives.

We present an approach for addressing such concerns by adding the appropriate context from the reference article to the citation texts. Enriching the citation texts with relevant context from the reference paper helps the reader to better understand the context for the ideas, methods or findings stated in the citation text.

A challenge in citation contextualization is the discourse and terminology variations between the citing and the referenced authors. Hence, traditional IR models that rely on term matching for finding the relevant information are ineffective.

We propose to address this challenge by a model that utilizes word embeddings and domain specific knowledge. Specifically, our approach is a retrieval model for finding the appropriate context of citations, aimed at capturing terminology variations and paraphrasing between the citation text and its relevant reference context.

We perform two sets of experiments to evaluate the performance of our system. First, we evaluate the relevance of extracted contexts intrinsically. Then we evaluate the effect of citation contextualization on the application of scientific summarization. Experimental results on TAC 2014 benchmark show that our approach significantly outperforms several strong baselines in extracting the relevant contexts. We furthermore, demonstrate that our contextualization models can enhance summarizing scientific articles.

Contextualizing citations

Given a citation text, our goal is to extract the most relevant context to it in the reference article. These contexts are essentially certain textual spans within the reference article. Throughout, colloquially, we refer to the citation text as query and reference spans in the reference article as documents. Our approach extends Language Models for IR (LM) by incorporating word embeddings and domain ontology to address shortcomings of LM for this research purpose. The goal in LM is to rank a document dd according to the conditional probability p(d∣q)∝p(q∣d)=∏qi∈qp(qi∣d)p(d|q)\propto p(q|d)=\prod_{q_{i}\in q}p(q_{i}|d) where qiq_{i} shows the tokens in the query qq. Estimating p(qi∣d)p(q_{i}|d) is often achieved by maximum likelihood estimate from term frequencies with some sort of smoothing. Using Dirichlet smoothing (Zhai and Lafferty, 2004), we have:

where f(qi,d)f(q_{i},d) shows the frequency of term qiq_{i} in document dd, C\mathcal{C} is the entire collection, V\mathcal{V} is the vocabulary and μ\mu the Dirichlet smoothing parameter. In the citation contextualization problem, (i) the target reference sentences are short documents and (ii) there exist terminology variations between the citing author and the referenced author. Hence, the citation terms usually do not appear in the documents and relying only on the frequencies of citation terms in the documents (f(qi,d)f(q_{i},d)) for estimating p(qi∣d)p(q_{i}|d) yields an almost uniform smoothed distribution that is unable to decisively distinguish between the documents.

where fsemf_{sem} is a function that measures semantic relatedness of the query term qiq_{i} to the document dd and is defined as: fsem(qi,d)=∑dj∈ds(qi,dj)f_{sem}(q_{i},d)=\sum_{d_{j}\in d}s(q_{i},d_{j}); where djd_{j}’s are document terms and s(qi,dj)s(q_{i},d_{j}) is the relatedness between the the query term and document term which is calculated by applying a similarity function to the distributed representations of qiq_{i} and djd_{j}. We use a transformation (ϕ\phi) of dot products between the unit vectors e(qi)e(q_{i}) and e(dj)e(d_{j}) corresponding to the embeddings of the terms qiq_{i} and djd_{j} for the similarity function: s(qi,dj)={ϕ(e(qi).e(dj));if e(qi).e(dj)>τ0;otherwise\small s(q_{i},d_{j})=\begin{cases}\phi(e(q_{i}).e(d_{j}));&\text{if }e(q_{i}).e(d_{j})>\tau\\ 0;&\text{otherwise}\end{cases}

We first explain the role of τ\tau and then the reason for considering the function ϕ\phi instead of raw dot product. τ\tau is a parameter that controls the noise introduced by less similar words. Many unrelated word vectors have non-zero similarity scores and adding them up introduces noise to the model and reduces the performance. τ\tau’s function is to set the similarity between unrelated words to zero instead of a positive number. To identify an appropriate value for τ\tau, we select a random set of words from the embedding model and calculate the average and standard deviation of pointwise absolute value of similarities between terms from these two samples. We then select the threshold τ\tau to be two standard deviations larger than the average to only consider very high similarity values (this choice was empirically justified).

Examining term similarity values between words shows that there are many terms with high similarities associated with each term and these values are not highly discriminative. We apply a transfer function ϕ\phi to the dot product e(qi).e(dj)e(q_{i}).e(d_{j}) to dampen the effect of less similar words. In other words, we only want highly related words to have high similarity values and similarity should quickly drop as we move to less related words. We use the logit function for ϕ\phi to achieve this dampening effect:

Figure 1 shows this effect. The purple line is the normalized dot product of a sample word with the most similar words in the model. As illustrated, the similarity score differences among top words is not very discriminative. However, applying the logit function (green line) causes the less similar words to have lower similarity values to the target word.

Successful word embedding methods have previously shown to be effective in capturing syntactic and semantic relatedness between terms. These co-occurrence based models are data driven. On the other hand, domain ontologies and lexicons that are built by experts include some information that might not be captured by embedding methods (Hill et al., 2015). Therefore, using domain knowledge can further help the embedding based retrieval model; we incorporate it in our model in the following ways:

1) Retrofitting: Faruqui et al. (2015) proposed a model that uses the constraints on WordNet lexicon to modify the word vectors and pull synonymous words closer to each other. To inject the domain knowledge in the embeddings, we apply this model on two domain specific ontologies, namely, Mesh and Protein Ontologies (PO)https://www.nlm.nih.gov/mesh/; http://pir.georgetown.edu/pro/. We chose these two biomedical domain ontologies because they are in the same domain as the articles in the TAC dataset. Mesh is a broad ontology that consists of biomedical terms and PO is a more focused ontology related to biology of proteins and genes.

2) Interpolating in the LM: We also directly incorporate the domain knowledge in the retrieval model; we modify the LM into the following interpolated LM with parameter λ\lambda:

where p1p_{1} is estimated using Eq. 2 and p2p_{2} is similar to p1p_{1} except that we replace fsemf_{sem} with the function fontf_{ont} which considers domain ontology in calculating similarities:

where γ∈\gamma\in is a parameter and qi≈djq_{i}\approx d_{j} shows that there is an is-synonym relation in ontology between qiq_{i} and djd_{j}The values of the parameters γ\gamma and λ\lambda were selected empirically by grid search.

Experiments

Data. We use the TAC 2014 Biomedical Summarization benchmarkhttp://www.nist.gov/tac/2014/BiomedSumm/. This dataset contains 220 scientific biomedical journal articles and 313 total citation texts where the relevant contexts for each citation text are annotated by 4 experts.

Baselines. To our knowledge, the only published results on TAC 2014 is (Cohan et al., 2015), where the authors utilized query reformulation (QR) based on UMLS ontology. In addition to (Cohan et al., 2015), we also implement several other strong baselines to better evaluate the effectiveness of our model: 1) BM25; 2) VSM: Vector Space Model that was used in (Cohan et al., 2015); 3) DESM: Dual Embedding Space Model which is a recent embedding based retrieval model (Mitra et al., 2016); and 4) LMD-LDA: Language modeling with LDA smoothing which is a recent extension of the LMD to also account for the latent topics (Jian et al., 2016). All the baseline parameters are tuned for the best performance, and the same preprocessing is applied to all the baselines and our methods.

First, we analyze the effectiveness of our proposed approach for contextualization intrinsically. That is, we evaluate the quality of the extracted citation contexts using our contextualization methods in terms of how accurate they are with respect to human annotations.

Evaluation. We consider the following evaluation metrics for assessing the quality of the retrieved contexts for each citation from multiple aspects: (i) Character offset overlaps of the retrieved contexts with human annotations in terms of precision (c-P), recall (c-R) and F-score (c-F). These are the recommended metrics for the task per TAChttps://tac.nist.gov/2014/BiomedSumm/guidelines.html. (ii) nDCG: we treat any partial overlaps with the gold standard as a correct context and then calculate the nDCG scores. (iii) Rouge-N scores: To also consider the content similarity of the retrieved contexts with the gold standard, we calculate the Rouge scores between them. (iv) Character precision at KK (c-P@K): Since we are usually interested in the top retrieved spans, we consider character offset precision only for the top KK spans and we denote it with “c-P@K”.

Results. The results of intrinsic evaluation of contextualization are presented in Table 1. Our models (last 4 rows of table 1) achieve significant improvements over the baselines consistently across most of the metrics. This shows the effectiveness of our models viewed from different aspects in comparison with the baselines. The best baseline performance is the query reformulation (QR) method by (Cohan et al., 2015) which improves over other baselines.

To analyze the performance of our system more closely, we took the context identified by 1 annotator as the candidate and the other 3 as gold standard and evaluated the precision to obtain an estimate of human performance on each citation. We then divided the citations based on human performance to 4 groups by quartiles. Table 3 shows our system’s performance on each of these groups. We observe that, when human precision is higher (upper quartiles in the table), our system also performs better and with more confidence (lower std). Therefore, the system errors correlate well with human disagreement on the correct context for the citations. Averaged over the 4 annotators for each citation, the mean precision was 56.7% (note that this translates to our c-P@1 metric). In Table 1, we observe that our best method (c-P@1 of 56.1%) is comparable with average human precision score (c-P@1 of 56.7%) which further demonstrates the effectiveness of our model.

2. External evaluation

Citation-based summarization can effectively capture various contributions and aspects of the paper by utilizing citation texts (Qazvinian and Radev, 2008). However; as argued in section 1, citation texts do not always accurately reflect the original paper. We show how adding context from the original paper can address this concern, while keeping the benefits of citation-based summarization. Specifically, we compare how using no contextualization, versus various proposed contextualization approaches affect the quality of summarization. We apply the following well-known summarization algorithms on the set of citation texts, and the retrieved citation-contexts: LexRank, LSA-based, SumBasic, and KL-Divergence (For space constraints, we will not explain these approaches here; refer to (Nenkova and McKeown, 2012) for details). We then compare the effect of our proposed contextualization methods using the standard Rouge-N summarization evaluation metrics.

Related work

Related work has mostly focused on extracting the citation text in the citing article (e.g. (Abu-Jbara and Radev, 2012)). In this work, given the citation texts, we focus on extracting its relevant context from the reference paper. Related work have also shown that citation texts can be used in different applications such as summarization (Qazvinian and Radev, 2008; Mei and Zhai, 2008; Wan et al., 2009; Cohan and Goharian, 2015; Jaidka et al., 2016; Cohan and Goharian, 2017). Our proposed model utilizes word embeddings and the domain knowledge. Embeddings have been recently used in general information retrieval models. Vulić and Moens (2015) proposed an architecture for learning word embeddings in multilingual settings and used them in document and query representation. Mitra et al. (2016) proposed dual embedded space model that predicts document aboutness by comparing the centroid of word vectors to query terms. Ganguly et al. (2015) used embeddings to transform term weights in a translation model for retrieval. Their model uses embeddings to expand documents and use co-occurrences for estimation. Unlike these works, we directly use embeddings in estimating the likelihood of query given documents; we furthermore incorporate ways to utilize domain specific knowledge in our model. The most relevant prior work to ours is (Cohan et al., 2015) where the authors approached the problem using a vector space model similarity ranking and query reformulations.

Conclusions

Citation texts are textual spans in a citing article that explain certain contributions of a reference paper. We presented an effective model for contextualizing citation texts (associating them with the appropriate context from the reference paper). We obtained statistically significant improvements in multiple evaluation metrics over several strong baseline, and we matched the human annotators precision. We showed that incorporating embeddings and domain knowledge in the language modeling based retrieval is effective for situations where there are high terminology variations between the source and the target (such as citations and their reference context). Citation contextualization not only can help the readers to better understand the citation texts but also as we demonstrated, they can improve other downstream applications such as scientific document summarization. Overall, our results show that citation contextualization enables us to take advantage of the benefits of citation texts, while ensuring accurate dissemination of the claims, ideas and findings of the original referenced paper.

Acknowledgements

We thank the three anonymous reviewers for their helpful comments and suggestions. This work was partially supported by National Science Foundation (NSF) through grant CNS-1204347.

References