The CL-SciSumm Shared Task 2018: Results and Key Insights
Kokil Jaidka, Michihiro Yasunaga, Muthu Kumar Chandrasekaran, Dragomir Radev, Min-Yen Kan
Introduction
CL-SciSumm explores summarization of scientific research in the domain of computational linguistics. The Shared Task dataset comprises the set of citation sentences (i.e., “citances”) that reference a specific paper as a (community-created) summary of a topic or paper . Citances for a reference paper are considered a synopses of its key points and also its key contributions and importance within an academic community . The advantage of using citances is that they are embedded with meta-commentary and offer a contextual, interpretative layer to the cited text. Citances offer a view of the cited paper which could complement the reader’s context, possibly as a scholar or a writer of a literature review .
The CL-SciSumm Shared Task is aimed at bringing together the summarization community to address challenges in scientific communication summarization. It encourages the incorporation of new kinds of information in automatic scientific paper summarization, such as the facets of research information being summarized in the research paper, and the use of new resources, such as the mini-summaries written in other papers by other scholars, and concept taxonomies developed for computational linguistics. Over time, we anticipate that the Shared Task will spur the creation of other new resources, tools, methods and evaluation frameworks.
CL-SciSumm task was first conducted at TAC 2014 as part of the larger BioMedSumm Taskhttp://www.nist.gov/tac/2014. It was organized in 2016 and 2017 as a part of the Joint Workshop on Bibliometric-enhanced Information Retrieval and Natural Language Processing for Digital Libraries (BIRNDL) workshop at the Joint Conference on Digital Libraries (JCDLhttp://www.jcdl2016.org/) in 2016, and the annual ACM Conference on Research and Development in Information Retrieval (SIGIRhttp://sigir.org/sigir2017/) in 2017 . This paper provides the results for the CL-SciSumm 2018 Task being held as part of the BIRNDL 2018 workshop at SIGIR 2018 in Ann Arbor, Michigan.
Task
CL-SciSumm defines two serially dependent tasks that participants could attempt, given a canonical training and testing set of papers.
Given: A topic consisting of a Reference Paper (RP) and ten or more Citing Papers (CPs) that all contain citations to the RP. In each CP, the text spans (i.e., citances) have been identified that pertain to a particular citation to the RP. Additionally, the dataset provides three types of summaries for each RP:
the abstract, written by the authors of the research paper.
the community summary, collated from the reference spans of its citances.
a human-written summary, written by the annotators of the CL-SciSumm annotation effort.
Task 1A: For each citance, identify the spans of text (cited text spans) in the RP that most accurately reflect the citance. These are of the granularity of a sentence fragment, a full sentence, or several consecutive sentences (no more than 5).
Task 1B: For each cited text span, identify what facet of the paper it belongs to, from a predefined set of facets.
Task 2: Finally, generate a structured summary of the RP from the cited text spans of the RP. The length of the summary should not exceed 250 words. This was an optional bonus task.
Development
The CL-SciSumm 2018 corpus comprises a training set that is randomly sampled research papers (Reference papers, RPs) from the ACL Anthology corpus and the citing papers (CPs) for those RPs which had at least ten citations. The prepared dataset then comprised annotated citing sentences for a research paper, mapped to the sentences in the RP which they referenced. Summaries of the RP were also included. The CL-SciSumm 2018 corpus included a refined version of the CL-SciSumm 2017 corpus of 40 RPs as a training set, in order to encourage teams from the previous edition to participate. For details of the general procedure followed to construct and annotate the CL-SciSumm corpus, the changes made to the procedure in CL-SciSumm-2016 and the refinement of the training set in 2017, please see .
The test set was an additional corpus of 20 RPs which was picked out of the ACL Anthology Network corpus (AAN), which automatically identifies and connects the citing papers and citances for each of thousands of highly-cited RPs. Therefore, we expect that that characteristics of the test set could be somewhat different from the training set. For this year’s test corpus, every RP and its citing papers were annotated three times by three independent annotators, and three sets of human summaries were also created.
The annotation scheme was unchanged from what was followed in previous editions of the task and the original BiomedSumm task developed by Cohen et. alhttp://www.nist.gov/tac/2014: Given each RP and its associated CPs, the annotation group was instructed to find citations to the RP in each CP. Specifically, the citation text, citation marker, reference text, and discourse facet were identified for each citation of the RP found in the CP.
Overview of Approaches
Ten systems participated in Task 1 and a subset of three also participated in Task 2. The following paragraphs discuss the approaches followed by the participating systems, in lexicographic order by team name.
System 2: The team from the Beijing University of Posts and Telecommunications’ Center for Intelligence Science and Technology developed models based on their 2017 system. For Task 1A, they adopted Word Mover’s Distance (WMD) and improve LDA model to calculate sentence similarity for citation linkage. For Task 1B they presented both rule-based systems, and supervised machine learning algorithms such as: Decision Trees and K-nearest Neighbor. For Task 2, in order to improve the performance of summarization, they also added WMD sentence similarity to construct new kernel matrix used in Determinantal Point Processes (DPPs).
System 4: The team from Thomson Reuters, Center for Cognitive Computing participated in Task 1A and B. For Task 1A, they treated the citation linkage prediction as a binary classification problem and utilized various similarity-based features, positional features and frequency-based features. For Task 1B, they treated the discourse facet prediction as a multi-label classification task using the same set of features.
System 6: The National University of Defense Technology team participated in Task 1A and B. For Task 1A, they used a random forest model using multiple features. Additionally, they integrated random forest model with BM25 and VSM model and applied a voting strategy to select the most related text spans. Lastly, they explored the language model with word embeddings and integrated it into the voting system to improve the performance. For task 1B, they used a multi-features random forest classifier.
System 7: The Nanjing University of Science and Technology team (NJUST) participated in all of the tasks (Tasks 1A, 1B and 2). For Task 1A, they used a weighted voting-based ensemble of classifiers (linear support vector machine (SVM), SVM using a radial basis function kernel, Decision Tree and Logistic Regression) to identify the reference span. For Task 1B, they used a dictionary for each discourse facet, a supervised topic model, and XGBOOST. For Task 2, they grouped sentences into three clusters (motivation, approach and conclusion) and then extracted sentences from each cluster to combine into a summary.
System 8: The International Institute of Information Technology team participated in Task 1A and B. They treated Task 1A as a text-matching problem, where they constructed a matching matrix whose entries represent the similarities between words, and used convolutional neural networks (CNN) on top to capture rich matching patterns. For Task 1B, they used SVM with tf-idf and naive bayes features.
System 9: The Klick Labs team participated in Task 1A and B. For Task 1A, they explored word embedding-based similarity measures to identify reference spans. They also studied several variations such as reference span cutoff optimization, normalized embeddings, and average embeddings. They treated Task 2B as a multi-class classification problem, where they constructed the feature vector for each sentence as the average of word embeddings of the terms in the sentence.
System 10: The University of Houston team adopted sentence similarity methods using Siamese Deep learning Networks and Positional Language Model approach for Task 1A. They tackled Task 1B using a rule-based method augmented by WordNet expansion, similarly to last year.
System 11: The LaSTUS/TALN+INCO team participated in all of the tasks (Tasks 1A, 1B and 2). For Task 1A, B, they proposed models that use Jaccard similarity, BabelNet synset embeddings cosine similarity, or convolutional neural network over word embeddings. For Task 2, they generated a summary by selecting the sentences from the RP that are most relevant to the CPs using various features. They used CNN to learn the relation between a sentence and a scoring value indicating its relevance.
System 12: The NLP-NITMZ team participated in all of the tasks (Tasks 1A, 1B and 2). For task 1A and 1B they extracted each citing papers (CP) text span that contains citations to the reference paper (RP). They used cosine similarity and Jaccard Similarity to measure the sentence similarity between CPs and RP, and picked the reference spans most similar to the citing sentence (Task 1A). For Task 1B, they applied rule based methods to extract the facets. For Task 2, they built a summary generation system using the OpenNMT tool.
System 20: Team Magma treated Task 1A as a binary classification problem and explored several classifiers with different feature sets. They found that Logistic regression with content-based features derived on topic and word similarities, in the ACL reference corpus, performed the best.
Evaluation
An automatic evaluation script was used to measure system performance for Task 1A, in terms of the sentence ID overlaps between the sentences identified in system output, versus the gold standard created by human annotators. The raw number of overlapping sentences were used to calculate the precision, recall and score for each system.
We followed the approach in most SemEval tasks in reporting the overall system performance as its micro-averaged performance over all topics in the blind test set. Additionally, we calculated lexical overlaps in terms of the ROUGE-2 and ROUGE-SU4 scores between the system output and the human annotated gold standard reference spans. It should be noted that this year, the average performance on every task was obtained by calculating the average performance on each of three independent sets of annotations for Task 1, and the performance on the human summary was also an average of performances on three human summaries.
ROUGE scoring was used for Tasks 1a and Task 2. Recall-Oriented Understudy for Gisting Evaluation (ROUGE) is a set of metrics used to automatically evaluate summarization systems by measuring the overlap between computer-generated summaries and multiple human written reference summaries. ROUGE–2 measures the bigram overlap between the candidate computer-generated summary and the reference summaries. More generally, ROUGE–N measures the -gram overlap. ROUGE-SU uses skip-bigram plus unigram overlaps. Similar to CL-SciSumm 2017, CL-SciSumm 2018 also uses ROUGE-2 and ROUGE-SU4 for its evaluation.
Task 1B was evaluated as a proportion of the correctly classified discourse facets by the system, contingent on the expected response of Task 1A. As it is a multi-label classification, this task was also scored based on the precision, recall and scores.
Task 2 was optional, and also evaluated using the ROUGE–2 and ROUGE–SU4 scores between the system output and three types of gold standard summaries of the research paper: the reference paper’s abstract, a community summary, and a human summary.
The evaluation scripts have been provided at the CL-SciSumm Github repositorygithub.com/WING-NUS/scisumm-corpus where the participants may run their own evaluation and report the results.
Results
This section compares the participating systems in terms of their performance. Three of the ten systems that did Task 1 also did the bonus Task 2. The results are provided in Table 1 and Table 2. The detailed implementation of the individual runs are described in the system papers included in this proceedings volume.
For Task 1A, on using sentence overlap (F1 score) as the metric, the best performance was by four runs from NUDT (system 6) . Their performance was closely followed by three runs from CIST (system 2) . The third best system was UPF-TALN (system 11) . When ROUGE-based F1 is used as a metric, the best performance is by Klick Labs (system 9) followed by NUDT (system 6) and then NLP-NITMZ (system 12) .
The best performance in Task 1B was by several runs submitted by CIST (system 2) followed by NJUST (system 7) . Klick Labs (system 9) was the second runner-up.
For Task 2, TALN-UPF (system 11) had the best performance against the abstract and human summaries, and the second-best performance against community summaries. NLP-NITMZ (system 12) had the best performance against the community summaries and were the second runners-up in the evaluation against human summaries. CIST (system 2) summaries had the second best performance against human summaries and finished as second runners-up against abstract and community summaries.
Conclusion and Future Work
Ten teams participated in this year’s shared task, on a corpus that was 33% larger than the 2017 corpus. In follow-up work, we plan to release a detailed comparison of the annotations as well as a micro-level error analysis to identify possible gaps in document or annotation quality. We will also aim to expand the part of the corpus with multiple annotations, in the coming few months. Furthermore, we expect to release other resources complementary to the CL-scientific summarization task, such as semantic concepts from the ACL Anthology Network .
We believe that the large improvements in Task 1A this year are a sign of forthcoming breakthroughs in information retrieval and summarization methods, and we hope that the community will not give up on the challenging task of generating scientific summaries for computational linguistics. Based on the experience of running this task for four years, we believe that lexical methods would work well with the structural and semantic characteristics that are unique to scientific documents, and perhaps will be complemented with domain-specific word embeddings in a deep learning framework. The Shared Task has demonstrated potential as a transfer learning task and is also expected to allow the generalization of its methods to other areas of scientific summarization.