Citation Classification for Behavioral Analysis of a Scientific Field
David Jurgens, Srijan Kumar, Raine Hoover, Dan McFarland, Dan Jurafsky
Introduction
Citations play a key role in scientific development. The citations from a paper reinforce its arguments and connect it to an intellectual lineage [Latour, 1987], while the citations to a paper enable communities to evaluate its intellectual contribution and quality [Cole and Cole, 1971, Lindsey, 1980, Hirsch, 2005]. At the same time, authors employ citations in multiple ways (Figure 1) so as to build a strong and multi-faceted argument [Latour, 1987]. While research on citation function is longstanding [Swales, 1986, White, 2004, Ding et al., 2014], there is still no field-scale dataset for citation function, and so we still don’t understand the way that a field’s authors frame their citations and how this framing influences uptake by readers and future citers.
We perform the first field-scale study of citation usage by developing accurate methods for automatically classifying citation purpose, and applying these methods to an entire field’s literature.
We unify core aspects of prior citation annotation schemes [White, 2004, Ding et al., 2014, Hernández-Alvarez and Gomez, 2016] into a highly-operationalizable coarse-grained classification that captures the function the citation plays towards furthering an argument as well as whether the citation is essential for understanding a contribution or serves rather to position the paper within the scientific context. Using this scheme, we annotate a corpus of citations and use it to train a high-accuracy method for automatically labeling a corpus. We use this method to label the field of NLP, over 134,127 citations in over 20,000 papers from nearly forty years of work.
We then investigate a number of key questions concerning how authors frame their citations (how citations reflect the discourse structure of a paper, how they are influenced by venue) how readers take them up (which citations do online readers follow, how does citation framing affect future citations to a paper), and how these processes reflect the maturation and growth of the field of NLP. As further contributions, we publicly release our dataset and code.
A Corpus for Citation Function
Citations play a key role in supporting authors’ contributions throughout a scientific paper.For notational clarity, we use the term reference for the work that is cited and citation for the mention of it in the text. Multiple schemes have been proposed on how to classify these different roles, ranging from a handful of classes [Nanba and Okumura, 1999, Pham and Hoffmann, 2003] to twenty or more [Garfield, 1979, Garzone and Mercer, 2000]. While suitable for expert manual analysis, such schemes don’t fit our goal of field-scale automatic classification. Because they tend to be unidimensional, they can’t label both centrality and function [Swales, 1986], they often include fine-grained distinctions too rare to reliably identify and subjective classifications that require detailed field or author knowledge [Ziman, 1968, Swales, 1990, Harwood, 2009]. Motivated by the desire to examine large-scale trends in scholarly behavior, we address these issues by unifying the common aspects of multiple approaches in a two-dimensional model.
Our classification builds on two themes for citation role: (1) citation centrality, which reflects whether the citation is used to position the contribution within a broader context or is essential for understanding the contribution itself [Moravcsik and Murugesan, 1975, Chubin and Moitra, 1975, Swales, 1986, Latour, 1987, Valenzuela et al., 2015] and (2) citation function, which reflects the particular purpose a citation is serving in the discourse, e.g., providing background or serving as contrast [Oppenheim and Renn, 1978, Spiegel-Rüsing, 1977, Teufel et al., 2006a, Garfield, 1979, Garzone and Mercer, 2000]. These themes capture complementary information and, as we will see, reveal meaningful patterns in author behavior.A third potential theme is citation sentiment [Athar, 2014, Kumar, 2016], but we omit this because researchers have shown that negative sentiment is rare in practice [Chubin and Moitra, 1975, Vinkler, 1998, Case and Higgins, 2000] and quite subjective to classify due to textual mixtures of praise and criticism [Peritz, 1983, Swales, 1986, Brooks, 1986, Teufel, 2000].
Centrality In citing, an author attempts to persuade the reader of a paper’s merit [Gilbert, 1977] by incorporating other work to further their own approach, or by contextualizing their work within the broader literature. Both kinds of citations are necessary; for example as ?) and others have argued, positioning citations secure the inductive gaps between an author’s arguments, and allow the author to identify with a lineage of work as motivation for their own position. The first theme in our classification, citation centrality, thus specifies whether a citation is (a) Essential to understanding the contributions of the citing paper or (b) Positioning by situating the citing paper within a broader context.
Function Citation function reflects the specific purpose a citation plays with respect to the current paper’s contributions. We unify the functional roles common in several classifications, e.g., [Spiegel-Rüsing, 1977, Garfield, 1979, Peritz, 1983, Teufel et al., 2006a, Harwood, 2009, Dong and Schäfer, 2011], into the seven classes shown in Table 1.
2 Annotation Process and Dataset
Annotation guidelines were created using a pilot study of 10 papers sampled from the ACL Anthology Reference Corpus (ARC) [Bird et al., 2008]. Annotators completed two rounds of pre-annotation to discuss their process and design guidelines. All citations were then doubly-annotated by two trained annotators using the Brat tool [Stenetorp et al., 2012] and then fully adjudicated to ensure quality.
The citation scheme was applied to a random sample of 52 papers drawn from the ARC. Each paper was processed using ParsCit [Councill et al., 2008] to extract citations and their references. As expected from prior studies [Teufel et al., 2006a, Dong and Schäfer, 2011], some citation functions were infrequent. We therefore attempted to oversample the infrequent classes Future, Continuation, and Motivation, by using keywords biased toward extracting citing sentences of a particular class (such as the word “future” for the Future class). The resulting citing sentences were then annotated and could potentially be assigned to any class. In total, 1436 contexts were annotated for the fully-labeled 52 papers (mean 27.6 citations/paper) and 533 supplemental contexts were added by sampling, bringing the total number of instances to 1969.
Table 2 shows the final dataset. Consistent with prior work, the majority of citations are Background [Moravcsik and Murugesan, 1975, Spiegel-Rüsing, 1977, Teufel et al., 2006b]. While some citation functions are strongly associated with one centrality type, e.g., Background as Positioning; citation function does not wholly predict centrality, highlighting the need for the two complementary classifications.
Automatically Classifying Citations
The structure of a scientific article provides multiple cues for a citation’s purpose. Our work draws on multiple approaches [Hernández-Alvarez and Gomez, 2016] to develop a classifier based on (1) structural features describing where the citation is located, (2) lexical and grammatical features for how the citation is described, (3) field features that take into account venue or other external information, and (4) usage features on how the reference is cited throughout the paper. Table 3 shows our features, including those drawn from state of the art systems [Teufel, 2000, Teufel et al., 2006b, Dong and Schäfer, 2011, Wan and Liu, 2014, Valenzuela et al., 2015] as well as ten novel feature classes.
Following, we describe in detail the three main categories of novel features.
Pattern-based Features Patterns provide a powerful mechanisms for capturing regularity in citation usage [Dong and Schäfer, 2011]. Our patterns are a sequence of cue phrases, parts of speech, or lexical categories, like positive-sentiment words or specific categories that allow generalizations across phrases like “we extend” and “we build upon.” We began with the largest publicly-available list of citation patterns [Teufel, 2000], extending it with 132 new patterns and 13 new lexical categories based on a manual analysis of the corpus.
We then used bootstrapping to automatically identify new patterns. Each annotated context was converted into fixed-length patterns using (a) our 42 lexical categories, (b) part of speech wild cards, or (c) the tokens directly. To avoid semantic drift [Riloff and Jones, 1999], a bootstrapped pattern was only included as a feature if the majority of its occurrences were with a single citation function.For computational efficiency, patterns were restricted to having between 3 and 8 tokens and at most two part of speech wild cards. Due to its high frequency, patterns for Background were required to occur in at least 100 contexts. Table 4 shows examples of these bootstrapped patterns.
Previous patterns primarily use cues from the same sentence as the citation [Teufel, 2000]. However, authors often use multiple sentences to indicate a citation’s purpose [Abu-Jbara and Radev, 2012, Ritchie et al., 2008, He et al., 2011, Kataria et al., 2011]. For example, authors may first introduce a work positively, only to contrast with it in later sentences [Peritz, 1983, Brooks, 1986, Mercer et al., 2004]. Indeed the average text pertaining to a citation spans 1.6 sentences in the ARC [Small, 2011].
We therefore induce bootstrapped patterns specific to the citation sentence as well as the preceding and following sentences. Ultimately, 805 new bootstrapped patterns were added for the citing sentence, 669 for the preceding context, and 1159 for the following, over four times the number of manually curated patterns.
Topic-based Features A context’s thematic framing can point to the purpose of a citation even in the absence of explicit cues. For example, a context describing system performances and results is likely to be a Compare or Contrast, whereas one describing methodology is more likely to be Uses. We quantify this thematic framing by using features based on topic models, computed over the sentence containing the citation and also over an extended -1/+3 context around the citing sentence. For each type of context, a topic model is trained with 100 topics over 321,129 respective contexts from the ARC. Table 5 shows example topics.
Prototypical Argument Features We also explored richer grammatical features, drawing on selectional preferences reflecting expectations for predicate arguments. We construct a prototype for each citation function by identifying the frequent arguments seen in different syntactic positions. For example, Continuation citations occur frequently as objects of verbs such as “follow” and “use,” whereas Uses citations have techniques or artifact words as dependents; Table 6 shows more examples. Each class’s selectional preferences are represented using a vector for the argument at each relation type, constructed by summing the vectors of all words appearing in it. Each function is represented as a separate feature whose value is the average similarity of an instance’s arguments with the class’s preferences for all observed syntactic relationships (i.e., how similar are the syntactically-related words to the function’s preferences).
2 Experimental Setup
Models All models were trained using a Random Forest classifier, which is robust to overfitting even with large numbers of features [Fernández-Delgado et al., 2014]. After limited grid search, we set the number of random trees to 500, initialized trees using a balanced subsample of the data, and required each leaf to match 7 instances. The classifier is implemented using SciKit [Pedregosa et al., 2011] and syntactic processing was done using CoreNLP [Manning et al., 2014]. Selectional preferences used pretrained 300-dimensional vectors from the 840B token Common Crawl [Pennington et al., 2014].
Data Annotated data is crucial for developing high accuracy for rare citation classes. Therefore, we integrate portions of the dataset of ?), which has fine-grained citation function labeled for ACL-related documents using the annotation scheme of ?). We map their 12 function classes into six of ours. When combining the two datasets, we omit the data labeled with their Background-equivalent class to reduce the effects of a large majority class and because instances of the Future class are merged into Background according to their scheme. The resulting citation function dataset contains 3083 instances. As there is no precise mapping from function to centrality (cf. Table 2), centrality classifiers use only our annotated data.
Evaluation Evaluation is performed using cross-validation where each fold leaves out all citations of a single paper. Stratifying by paper instead of instance is critical: since multiple citations may appear in the same sentence, instance-based stratification would leak information between training and test. We report micro- and macro-averaged F1 scores across the seven function classes and for centrality, report precision and recall on the binary classification where Essential is the positive class.
Comparison Systems The proposed citation classifiers are compared against three systems. For state of the art, we compare against ?) which is the most similar model to ours that is experimentally reproducible; the original implementation used a custom syntactic tool (for e.g., verb tense), which we replaced with CoreNLP. Two baselines are used for comparison: a Random baseline that selects labels at chance and a Single-class baseline that labels all instances with the most frequent citation function Background or in the binary centrality classification, with the positive class Essential.
3 Results and Discussion
Our methods substantially outperformed the closest state of the art and both baselines for both classification tasks, as shown in Table 7. All improvements over comparison systems are statistically significant (McNemar’s, p0.01).
An ablation test suggests that each of our novel features contributed to the final performance. Notably, we observe that selectional preference and topic features had the largest impact on performance. While both we and multiple prior works have focused on patterns to recognize function, our results suggest that machine learned features (topics or word vectors) are superior. Indeed, examining the feature weighting in the random forest shows that features for structure (e.g., section number), topic, and selectional preference comprised most of the 100 highest-weighted features (76%).
The use of conjunctive features was critical for performance, with all other classifiers we tried (SVM, Naive Bayes, -nearest neighbor, and Decision Trees) providing significantly worse results.Indeed, replacing the -nearest neighbors classifier used in ?) with a random forest improves citation function classification by 0.123 (Macro F1) and 0.175 (Micro F1).
The resulting classifier performance is sufficient to apply it to the entire ARC dataset for the analyses in the next four sections. Nonetheless, errors remain. Our error analysis revealed that a main challenge is incorporating information external to the citing sentence. Consider the following example:
BilderNetle is our new data set of German noun-to-ImageNet synset mappings. ImageNet is a large-scale and widely used image database, built on top of WordNet, which maps words into groups of images, called synsets (Deng et al., 2009).
Here the citing sentence appears much like a Background citation when read in isolation; however, the preceding sentence reveals that the citing work’s data is based on the citation, making its function Uses though no explicit cues suggest this in the citing sentence. Automated methods need to do a better job of inferring (a) which context inside and outside the citing sentence relates to the citation and (b) how the cited paper’s content relates to this context; together these require richer textual understanding than is present in most current methods.
Narrative Structure of Citation Function
In the next four sections, we apply the classifier trained on our combined dataset (3083 citation function instances, 1969 instances for citation centrality) to the ACL Anthology to study what citation roles can tell us about scientific uptake and direction.
As a qualitative demonstration, we first examine the narrative structure of citation function across section. Scientific papers commonly follow a structured narrative according to section: Introduction, Methodology, Results, and Discussion [Skelton, 1994, Nwogu, 1997]. Each part in the narrative adopts argumentative moves designed to convince the reader of the work’s claims [Swales, 1986, Swales, 1990]. We expect that this narrative is mirrored in how authors use their citations in sections, with the section serving to frame the meaning of a citation [Goffman, 1974, Gumperz, 1982].
To test this hypothesis, the function classifier was applied to all 21,474 papers of the latest 2016 release of the ACL Anthology. This yielded a dataset of 134,127 citations between papers in the ARC. The resulting distributions of citation function (Figure 2), show that authors’ citation usage indeed parallels the expected rhetorical framing: (1) establishing framing via Background citations in the Introduction, Motivation, and Related Work sections to (2) introducing methodology with Uses citations in the Methodology and Evaluation sections, (3) a large increase in Comparison or Contrast for related literature in the Results and Discussions, and finally (4) closing comparisons and pointers to future directions. These trends also mirror the thematic structure identified in full-paper textual analyses [Skelton, 1994, Nwogu, 1997]. By showing that a section contains citations serving a variety of functions, our findings further point to a new direction for citation placement studies [Hu et al., 2013, Ding et al., 2013, Bertin et al., 2016], which have largely treated all citations within a section as equivalent.
Venues and Citation Patterns
Does the venue in which a paper appears affect the way it cites? We used the same experimental setup as the previous section. Figure 3 shows citation function by venue for the 134,127 citations.
We find that similar venues have similar citation functions. Journals have the highest percentage of Background citations, suggesting that their extra space and wider temporal scope lends itself more to positioning. Conference venues devote proportionally more space to contrast and comparison, presumably because new NLP work is first presented at conferences, hence requiring demonstrating the proposed technique is better than existing ones. Workshops, by contrast, have relatively little comparison and instead use more Background; the experimental nature of workshop papers presumably results in fewer potential prospects for comparison.
Citation Network Navigation
The narrative in a scholarly work highlights salient aspects about cited works for the reader. As a result, a citation’s framing may motivate or discourage readers from seeking it out. Understanding readers’ citation-following behavior is an important part of understanding scholars’ information gathering goals and strategies. Yet while human navigation on hyperlinked networks such as Wikipedia have been explored extensively (e.g. ?)), we know very little about how scholars browse citation networks [Rouse et al., 1982, Bergström and Whitehead Jr, 2006]; and no studies have looked at data from users’ real-world behavior. Do readers generally follow links to methodologies? Are they more likely to follow links to essential papers?
The ACL Anthology offers an opportunity to examine these questions in detail, since the server logs from the Anthology allow us to estimate when readers follow citation links from one paper to another. We use this data to understand how the way a reference is cited is predictive of it being read.
Experimental Setup Paper requests were gathered from the complete Apache logs of the ACL Anthology website from November 2014 to September 2015. After removing requests by bots and webspiders, the logs contain roughly 3.6M requests for 34,821 papers, with between 1 and 9588 requests per paper (mean 97.5). The access logs do not directly record unique user identifiers for tracking behavior over extended periods of time (e.g., as compared to session cookies); however, each request includes a string describing the user agent, which reports specific details on the operating system, browser, and version numbers. Because of the high number of unique combinations of user agent features and relatively low traffic to the website, these user agent strings are sufficient for identifying access patterns of individual users during short time intervals. Activity traces are constructed by identifying sequential requests made by the same user agent string for two or more papers, with no more than 60 minutes between requests. The vast majority of traces consist of requests for two (80%) or three papers (9.6%).To guard against cases where a user agent string might be used by multiple individuals in parallel, we remove all traces with 50 or more requests, though this is rare in practice (0.3%).
To compare with the observed human behavior, we construct a null model over the expected frequencies with which references of each type are visited. For each trace beginning with paper and continuing to one of its references , the null model randomly selects a reference in with equal probability. Visiting the same initial papers seen in the observed data controls for differences in the number of references. The expected frequency distributions for reading behavior was created by generating visit counts from 500 simulations per trace with the null model and calculating the number of standard deviations away the observed frequency was from the expected frequency, i.e., the z-score. A positive z-score indicates that the reader accesses the paper more frequently than expected by chance.
The effects of how a reference is cited is measured by calculating the z-score in visits with respect to references by class. We treat a reference as Essential if any of its citations are labeled as such. Because a reference may have multiple functions (e.g., both Background and Uses), we use a fractional count proportional to the number of functions.
Results Centrality was strongly predictive of whether a reader would subsequently access the reference, shown in Table 8 (top). Confirming the hypothesis of ?), we find that Essential references are significantly more likely to be read than at chance (p 0.01), while Positioning are less likely (p 0.01). However, the effect of citation function (Table 8, bottom) reveals a more complex story of reader behavior.
Readers are more likely to follow the temporal thread of an idea, reading the papers from which the current work originates (by others Extends or themselves Continuation) or how the current work could be applied in the future (Future). However, readers are significantly less likely to follow Uses citations, despite these nearly always being Essential. We speculate that when authors extend their work or the work of others, they are more likely to omit details of the original paper; hence, a reader must go back to the original paper to obtain a complete picture. And because works within a research area often share methodologies and data (Uses citations), a reader may be familiar with these references and not seek them out again.
By using the browsing patterns of an entire field, we have demonstrated the strategies that scholars use to explore research. Our findings more broadly motivate the need for scholarly search engines like Semantic Scholar (www.semanticscholar.org/) to support browsing preferences by surfacing temporally-connected references instead of methodological citations.
Predicting Future Impact
The scholarly narrative told through citations provides the reader with support for its claims and technical competence [Latour, 1987, p. 34]. This framing could affect the perception of the work and, ultimately, how it is received and cited within the community [Shi et al., 2010]. Does the narrative a paper constructs through citation roles (the way it compares to related work, or motivates, or points to the future) affect its reception?
Experimental Setup To quantify the impact of citation usage, we test its ability to predict the cumulative number of citations a paper received within the first five years after publication, known to be highly representative of the eventual citation count [Wang et al., 2013, Stern, 2014]. Our baseline model consists of state of the art features for predicting citation count [Yan et al., 2011, Yan et al., 2012, Chakraborty et al., 2014, Dong et al., 2016]; we then ask whether prediction is improved by adding the number of citations per function and centrality.
All papers with at least five years of publication history in the anthology were considered, yielding a set of 10,434 papers. We used negative binomial models, which are more appropriate than linear regression as citation counts are non-negative discrete counts, and compared them using Akaike Information Criterion (AIC).
Five types of baseline features were included. To model the amount of attention received by different research areas, each paper is associated with its distribution over 100 topics, built using LDA over the ARC; to capture diversity, we include the entropy of the topic distribution. We model the different publication venues via a categorical feature for the ACL venue in which the paper was published. We include the publication year since the size of the field changes over time. Multi-author papers are known to receive higher citation counts [Gazni and Didegah, 2011], partially due to the effects of self-citation [Fowler and Aksnes, 2007], and therefore we include the number of authors on the paper. Finally, we include the number of references.To control for collinearity between citation-related predictors, we regress out the number of citations from the citation function counts and then in turn, regress out the citation function counts from citation centrality counts [Kutner et al., 2004, O’brien, 2007]. The resulting model has a variance inflation factor of 10 for all variables.
Results Knowledge of how a paper contextualizes itself indeed helps predict its future impact; citation function and centrality each improve AIC model fit (likelihood ratio test, p 0.01 for each), and adding them together improves AIC again (p 0.01).
Four main insights can been seen through which types of citations are significantly predictive of higher impact (p0.05), shown in Table 9. First, papers that closely integrate more works as essential to their contribution are cited more than those using more citations to position their work. This trend parallels the effects seen with Uses citations, which suggests that papers that leverage existing methodologies and data have higher impact. Second, citations to motivating works have the most impact on future citation. We speculate that such citations provide the reader with an explicit signal of external support for a line of research which inspires future work in the area as a result. Third, ?, p. 54) has suggested that authors may deflect criticism of their work (improving its perception) by claiming it as an extension, rather than comparing it with prior work. However, we find that additional comparison increases impact, while no significant effect was seen for Extends. We speculate that Extends and Continuation citations could lead to the work being perceived as incremental, thereby reducing its novelty and anticipated impact. Fourth, citations indicating a temporal connection had no significant effect on citation frequency (p0.05)—despite our earlier results showing that these citations are more likely to be read.
The Growth of Rapid Discovery Science
As scientific fields evolve, new subfields are often initially based around method technology, which starts a community discussion of the best way to use and improve the technology [Moody, 2004]. NLP has witnessed the emergence of several such subfields from the early grammar based approaches in the 1950s-1970s to the statistical revolution in the 1990s to the recent deep learning models [Spärck Jones, 2001, Anderson et al., 2012]. ?) proposed that a field can undergo a particular shift, to what he calls rapid discovery science, when the field (a) reaches high consensus on research topics and methodologies, and (b) develops genealogies of method technologies that build upon one another. As a result of increased consensus, a field emphasizes moving to new research problems rather than contesting the results of prior approaches. Collins claims this shift characterizes natural sciences, but not many social sciences, which are instead more likely to emphasize this process of contesting.
We propose to measure the two facets of rapid discovery science by quantifying the degrees to which authors (a) defend their work through positioning and comparison, with lower rates suggesting more consensus, and (b) focus on the same works for methodology and comparison, with increased rates signaling the development of methodological lineages. We propose that the increased use of shared evaluations, and the statistical methodology borrowed originally from electrical engineering [Hall et al., 2008] has led NLP to undergo a shift towards rapid discovery science.
Experimental Setup The expected frequencies of citation functions and centrality were determined using the fully-classified ACL Anthology network. Due to variance in the number of papers per year, we report bootstrapped confidence intervals in the expected percentage of each at 68% and 95%.
Results The NLP field shows a significant increase in consensus consistent with the rise in rapid discovery science, evidenced through two main trends.
First, NLP authors use a decreasing number of Positioning references (Pearson’s = -0.446, p 0.01)—Figure 4(a)—with an even sharper decrease in the number of works used for comparison and contrast (= -0.899, p 0.01)—Figure 4(b). Instead of comparing to others, it seems that authors simply acknowledge prior work as Background. Despite an increase in Background citations, the total percentage of comparative and background citations (the two main Positioning functions; see Table 2) still declines (= -0.663, p 0.01), with authors instead increasingly including more Uses citations. ?, p. 50) argues that Positioning references are critical to an author’s defense of an idea. We therefore interpret the observed decrease in Positioning references as signaling a reduced need for authors to defend aspects of their work. Authors are able to compare against fewer papers due to the fields growing consensus on the validity of the problem and methodological contribution.
Note that there is a small but significant increase in the number of Positioning references between 2009 and 2011. This transition corresponds to the date at which ACL venues began allowing unlimited references (2010 for ACL, 2011 for NAACL, etc.). Unlimited extra space for citations acted to distort the framing of citations; given unlimited space, authors chose to include proportionally more Positioning citations.Note that this change acts against the general decrease in Positioning; considering only 1980-2009, the decrease in Positioning is even larger (= -0.568, p 0.01 ).
In the second trend, authors are more likely to use and compare against the same papers, as shown in Figure 4(c) by the rise in expected incoming citations to those works compared against (= 0.734, p 0.01) and used (=0.889, p 0.01). For example, in 1991, authors compared with a diffuse group of parsing papers, e.g., [Shieber, 1988, Pereira and Warren, 1983, Haas, 1989], with such papers receiving at most three citations that year; whereas in 2000, most comparisons were to a core set of parsing papers, e.g., [Collins, 2003, Buchholz et al., 1999, Collins, 1997], with a much sharper (lower entropy) distribution of citations. These trends show the increased incorporation of prior work to form a lineage of method technologies as well as show increased consensus on which works are sufficient for comparing against in order to establish a claim. These results also empirically confirm the observation of ?) that a major trend in NLP in the 1990s was an increase in reusable technologies and evaluations, like the BNC [Leech, 1992] and the Penn Treebank [Marcus et al., 1993].
More broadly, our work points to the future of NLP as a quickly moving field of high consensus and suggests that artifacts that facilitate consensus such as shared tasks and open source research software will be necessary to continue this trend.
Conclusion
Not all citations are created equal. Using a new corpus annotated with citation function and centrality, we develop a state-of-the-art classifier, demonstrating the importance of novel unsupervised features related to topic models and argument structure, and label all the citations for an entire field (with all data and materials released at https://github.com/anon).
We then show that citation function and centrality reveal salient behaviors of writers, readers, and the field as a whole: (1) authors are sensitive to discourse structure and venue when citing, (2) readers follow temporal threads of ideas in their browsing of citation networks, rather than methodological threads, (3) the way in which an author relates their work to the literature is predictive of the number of citations its receives, with the community favoring well-motivated works that integrate and compare with other works, and (4) the NLP field as a whole has seen increased consensus in what constitutes valid work—with a reduced need for positioning and excessive comparison—demonstrating its shift towards rapid discovery science.