TalkSumm: A Dataset and Scalable Annotation Method for Scientific Paper Summarization Based on Conference Talks

Guy Lev, Michal Shmueli-Scheuer, Jonathan Herzig, Achiya Jerbi, David Konopnicki

Introduction

The rate of publications of scientific papers is increasing and it is almost impossible for researchers to keep up with relevant research. Automatic text summarization could help mitigate this problem. In general, there are two common approaches to summarizing scientific papers: citations-based, based on a set of citation sentences Nakov et al. (2004); Abu-Jbara and Radev (2011); Yasunaga et al. (2019), and content-based, based on the paper itself Collins et al. (2017); Nikola Nikolov and Hahnloser (2018). Automatic summarization is studied exhaustively for the news domain Cheng and Lapata (2016); See et al. (2017), while summarization of scientific papers is less studied, mainly due to the lack of large-scale training data. The papers’ length and complexity require substantial summarization effort from experts. Several methods were suggested to reduce these efforts Yasunaga et al. (2019); Collins et al. (2017), still they are not scalable as they require human annotations.

Recently, academic conferences started publishing videos of talks (e.g., ACLvimeo.com/aclweb, EMNLPfootnote 1, ICMLicml.cc/Conferences/2017/Videos, and more). In such talks, the presenter (usually a co-author) must describe their paper coherently and concisely (since there is a time limit), providing a good basis for generating summaries. Based on this idea, in this paper, we propose a new method, named TalkSumm (acronym for Talk-based Summarization), to automatically generate extractive content-based summaries for scientific papers based on video talks. Our approach utilizes the transcripts of video content of conference talks, and treat them as spoken summaries of papers. Then, using unsupervised alignment algorithms, we map the transcripts to the corresponding papers’ text, and create extractive summaries. Table 1 gives an example of an alignment between a paper and its talk transcript (see Table 3 in the appendix for a complete example).

Summaries generated with our approach can then be used to train more complex and data-demanding summarization models. Although our summaries may be noisy (as they are created automatically from transcripts), our dataset can easily grow in size as more conference videos are aggregated. Moreover, our approach can generate summaries of various lengths.

Our main contributions are as follows: (1) we propose a new approach to automatically generate summaries for scientific papers based on video talks; (2) we create a new dataset, that contains 17161716 summaries for papers from several computer science conferences, that can be used as training data; (3) we show both automatic and human evaluations for our approach. We make our dataset and related code publicly availablehttps://github.com/levguy/talksumm. To our knowledge, this is the first approach to automatically create extractive summaries for scientific papers by utilizing the videos of conference talks.

Related Work

Several works focused on generating training data for scientific paper summarization Yasunaga et al. (2019); Jaidka et al. (2018); Collins et al. (2017); Cohan and Goharian (2018). Most prominently, the CL-SciSumm shared tasks Jaidka et al. (2016, 2018) provide a total of 40 human generated summaries; there, a citations-based approach is used, where experts first read citation sentences (citances) that reference the paper being summarized, and then read the whole paper. Then, they create a summary of 150 words on average.

Recently, to mitigate annotation cost, Yasunaga et al. (2019) proposed a method, in which human annotators only read the abstract in addition to citances (not reading the full paper). Using this approach, they generated 10001000 summaries, costing 600+ person-hours. Conversely, we generate summaries, given transcripts of conference talks, in a fully automatic manner, and, thus, our approach is much more scalable. Collins et al. (2017) also aimed at generating labeled data for scientific paper summarization, based on “highlight statements” that authors can provide in some publication venues.

Using external data to create summaries was also proposed in the news domain. Wei and Gao (2014, 2015) utilized tweets to decide which sentences to extract from news article.

Finally, alignment between different modalities (e.g., presentation, videos) and text was studied in different domains. Both Kan (2007) and Bahrani and Kan (2013) studied the problem of document to presentation alignment for scholarly documents. Kan (2007) focused on the the discovery and crawling of document-presentation pairs, and a model to align between documents to corresponding presentations. In Bahrani and Kan (2013) they extended previous model to include also visual components of the slides. Aligning video and text was studied mainly in the setting of enriching videos with textual information Bojanowski et al. (2015); Malmaud et al. (2015); Zhu et al. (2015). Malmaud et al. (2015) used HMM to align ASR transcripts of cooking videos and recipes text for enriching videos with instructions. Zhu et al. (2015) utilized books to enrich videos with descriptive explanations. Bojanowski et al. (2015) proposed to align video and text by providing a time stamp for every sentence. The main difference between these works and ours is in the alignment being used to generate textual training data in our case, rather than to enrich videos.

The TalkSumm Dataset

Recently, many computer science academic associations including ACL, ACM, IMLS and more, have started recording talks in different conferences, e.g., ACL, NAACL, EMNLP, and other co-located workshops. A similar trend occurs in other domains such as Physicswww.cleoconference.org, Biologyigem.org/Videos/Lecture_Videos, etc.

In a conference, each speaker (usually a co-author) presents their paper given a timeframe of 15-20 minutes. Thus, the talk must be coherent and concentrate on the most important aspects of a paper. Hence, the talk can be considered as a summary of the paper, as viewed by its authors, and is much more comprehensive than the abstract, which is written by the authors as well.

In this work, we focused on NLP and ML conferences, and analyzed 17161716 video talks from ACL, NAACL, EMNLP, SIGDIAL (2015-2018), and ICML (2017-2018). We downloaded the videos and extracted the speech data. Then, via a publicly available ASR servicewww.ibm.com/watson/services/speech-to-text/, we extracted transcripts of the speech, and based on the video metadata (e.g., title), we retrieved the corresponding paper (in PDF format). We used Science-Parsegithub.com/allenai/science-parse to extract the text of the paper, and applied a simple processing in order to filter-out some noise (e.g. lines starting with the word “Copyright”). At the end of this process, the text of each paper is associated with the transcript of the corresponding talk.

2 Dataset Generation

The transcript itself cannot serve as a good summary for the corresponding paper, as it constitutes only one modality of the talk (which also consists of slides, for example), and hence cannot stand by itself and form a coherent written text. Thus, to create an extractive paper summary based on the transcript, we model the alignment between spoken words and sentences in the paper, assuming the following generative process: During the talk, the speaker generates words for describing verbally sentences from the paper, one word at each time step. Thus, at each time step, the speaker has a single sentence from the paper in mind, and produces a word that constitutes a part of its verbal description. Then, at the next time-step, the speaker either stays with the same sentence, or moves on to describing another sentence, and so on. Thus, given the transcript, we aim to retrieve those “source” sentences and use them as the summary. The number of words uttered to describe each sentence can serve as importance score, indicating the amount of time the speaker spent describing the sentence. This enables to control the summary length by considering only the most important sentences up to some threshold.

We use an HMM to model the assumed generative process. The sequence of spoken words is the output sequence. Each hidden state of the HMM corresponds to a single paper sentence. We heuristically define the HMM’s probabilities as follows.

Denote by Y(1:T)Y(1:T) the spoken words, and by S(t)∈{1,...,K}S(t)\in\{1,...,K\} the paper sentence index at time-step t∈{1,...,T}t\in\{1,...,T\}. Similarly to Malmaud et al. (2015), we define the emission probabilities to be:

where words(k)words(k) is the set of words in the kk’th sentence, and simsim is a semantic-similarity measure between words, based on word-vector distance. We use pre-trained GloVe Pennington et al. (2014) as the semantic vector representations for words.

As for the transition probabilities, we must model the speaker’s behavior and the transitions between any two sentences in the paper. This is unlike the simpler setting in Malmaud et al. (2015), where transition is allowed between consecutive sentences only. To do so, denote the entries of the transition matrix by T(k,l)=p(S(t+1)=l∣S(t)=k)T(k,l)=p(S(t+1)=l|S(t)=k). We rely on the following assumptions: (1) T(k,k)T(k,k) (the probability of staying in the same sentence at the next time-step) is relatively high. (2) There is an inverse relation between T(k,l)T(k,l) and ∣l−k∣|l-k|, i.e., it is more probable to move to a nearby sentence than jumping to a farther sentence. (3) S(t+1)>S(t)S(t+1)>S(t) is more probable than the opposite (i.e., transition to a later sentence is more probable than to an earlier one). Although these assumptions do not perfectly reflect reality, they are a reasonable approximation in practice.

Following these assumptions, we define the HMM’s transition probability matrix. First, define the stay-probability as α=max⁡(δ(1−KT),ϵ)\alpha=\max(\delta(1-\frac{K}{T}),\epsilon), where δ,ϵ∈(0,1)\delta,\epsilon\in(0,1). This choice of stay-probability is inspired by Malmaud et al. (2015), using δ\delta to fit it to our case where transitions between any two sentences are allowed, and ϵ\epsilon to handle rare cases where KK is close to, or even larger than TT. Then, for each sentence index k∈{1,...,K}k\in\{1,...,K\}, we define:

where λ,γ,βk∈(0,1)\lambda,\gamma,\beta_{k}\in(0,1), λ\lambda and γ\gamma are factors reflecting assumptions (2) and (3) respectively, and for all kk, βk\beta_{k} is normalized s.t. ∑l=1KT(k,l)=1\sum_{l=1}^{K}T(k,l)=1. The values of λ\lambda, γ\gamma, δ\delta and ϵ\epsilon were fixed throughout our experiments at λ=0.75\lambda=0.75, γ=0.5\gamma=0.5, δ=0.33\delta=0.33 and ϵ=0.1\epsilon=0.1. The average value of α\alpha, across all papers, was around 0.30.3. The values of these parameters were determined based on evaluation over manually-labeled alignments between the transcripts and the sentences of a small set of papers.

Finally, we define the start-probabilities assuming that the first spoken word must be conditioned on a sentence from the Introduction section, hence p(S(1))p(S(1)) is defined as a uniform distribution over the Introduction section’s sentences.

Note that sentences which appear in the Abstract, Related Work, and Acknowledgments sections of each paper are excluded from the HMM’s hidden states, as we observed that presenters seldom refer to them.

To estimate the MAP sequence of sentences, we apply the Viterbi algorithm. The sentences in the obtained sequence are the candidates for the paper’s summary. For each sentence ss appearing in this sequence, denote by count(s)count(s) the number of time-steps in which this sentence appears. Thus, count(s)count(s) models the number of words generated by the speaker conditioned on ss, and, hence, can be used as an importance score. Given a desired summary length, one can draw a subset of top-ranked sentences up to this length.

Experiments

We evaluate the quality of our dataset generation method by training an extractive summarization model, and evaluating this model on a human-generated dataset of scientific paper summaries. For this, we choose the CL-SciSumm shared task Jaidka et al. (2016, 2018), as this is the most established benchmark for scientific paper summarization. In this dataset, experts wrote summaries of 150 words length on average, after reading the whole paper. The evaluation is on the same test data used by Yasunaga et al. (2019), namely 10 examples from CL-SciSumm 2016, and 20 examples from CL-SciSumm 2018 as validation data.

Training Data

Using the HMM importance scores, we create four training sets, two with fixed-length summaries (150 and 250 words), and two with fixed ratio between summary and paper lengths (0.3 and 0.4). We train models on each training set, and select the model yielding the best performance on the validation set (evaluation is always done with generating a 150-words summary).

Summarization Model

We train an extractive summarization model on our TalkSumm dataset, using the extractive variant of Chen and Bansal (2018). We test two summary generation approaches, similarly to Yasunaga et al. (2019). First, for TalkSumm-only, we generate a 150-words summary out of the top-ranked sentences extracted by our trained model (sentences from the Acknowledgments section are omitted, in case the model extracts any). In the second approach, a 150-words summary is created by augmenting the abstract with non-redundant sentences extracted by our model, similarly to the “Hybrid 2” approach of Yasunaga et al. (2019). We perform early-stopping and hyper-parameters tuning using the validation set.

Baselines

We compare our results to ScisummNet Yasunaga et al. (2019) trained on 1000 scientific papers summarized by human annotators. As we use the same test set as in Yasunaga et al. (2019), we directly compare their reported model performance to ours, including their Abstract baseline which takes the abstract to be the paper’s summary.

2 Results

Table 2 summarizes the results: both GCN Cited text spans and TalkSumm-only models, are not able to obtain better performance than AbstractWhile the abstract was input to GCN Cited text spans, it was excluded from TalkSumm-only.. However, for the Hybrid approach, where the abstract is augmented with sentences from the summaries emitted by the models, our TalkSumm-Hybrid outperforms both GCN Hybrid 2 and Abstract. Importantly, our model, trained on automatically-generated summaries, performs on par with models trained over ScisummNet, in which training data was created manually.

Human Evaluation

We conduct a human evaluation of our approach with support from authors who presented their papers in conferences. As our goal is to test more comprehensive summaries, we generated summaries composed of 3030 sentences (approximately 15%15\% of a long paper). We randomly selected 1515 presenters from our corpus and asked them to perform two tasks, given the generated summary of their paper: (1) for each sentence in the summary, we asked them to indicate whether they considered it when preparing the talk (yes/no question); (2) we asked them to globally evaluate the quality of the summary (1-5 scale, ranging from very bad to excellent, 3 means good). For the sentence-level task (1), 73%73\% of the sentences were considered while preparing the talk. As for the global task (2), the quality of the summaries was 3.733.73 on average, with standard deviation of 0.7250.725. These results validate the quality of our generation method.

Conclusion

We propose a novel automatic method to generate training data for scientific papers summarization, based on conference talks given by authors. We show that the a model trained on our dataset achieves competitive results compared to models trained on human generated summaries, and that the dataset quality satisfies human experts. In the future, we plan to study the effect of other video modalities on the alignment algorithm. We hope our method and dataset will unlock new opportunities for scientific paper summarization.

References

Appendix A A Detailed Example

This section elaborates on the example presented in Table 1. Table 3 extends Table 1 by showing the manually-labeled alignment between the complete text of the paper’s Introduction section, and the corresponding transcript. Table 4 shows the alignment obtained using the HMM. Each row in this table corresponds to an interval of consecutive time-steps (i.e., a sub-sequence of the transcript) in which the same paper sentence was selected by the Viterbi algorithm. The first column (Paper Sentence) shows the selected sentences; The second column (ASR transcript) shows the transcript obtained by the ASR system; The third column (Human transcript) shows the manually corrected transcript, which is provided for readability (our model predicted the alignment based on the raw ASR output); Finally, the forth column shows whether our model has correctly aligned a paper sentence with a sub-sequence of the transcript. Rows with no values in this column correspond to transcript sub-sequences which were not associated with any paper sentence in the manually-labeled alignment.