Language Model Evaluation Beyond Perplexity
Clara Meister, Ryan Cotterell
Introduction
Neural language modelsIn this work, we do not use the term language model to refer to cloze language models such as BERT Devlin et al. (2019), which do not give us a distribution over strings. have become shockingly good at modeling natural language data in recent years Merity et al. (2017); Conneau and Lample (2019); Radford et al. (2019). Thus, to test just how well neural language models capture language NLP researchers have started to look beyond standard evaluation metrics such as perplexity, endeavoring to understand which underlying attributes of human language these models are learning. To this end, a nascent literature has emerged that focuses on probing language models Belinkov and Glass (2019), i.e., determining whether models encode linguistic phenomena. For the most part, these works have been limited to analyses of sentence-level phenomenon, such as subject–verb agreement Gulordava et al. (2018) and garden path effects van Schijndel and Linzen (2018) among a myriad of other properties (Blevins et al., 2018; Chowdhury and Zamparelli, 2018, inter alia).
In this work, we attempt to understand which macro-level phenomena of human language today’s language models reflect. That is, we pose the question: Do neural language models exhibit the statistical tendencies of human language? Phenomena that can be measured at this level provide an alternate view of a model’s comprehension; for example, rather than exploring whether morphological agreement is captured, we look at whether our models learn the trends across a corpus as a whole, e.g., the token rank–frequency (Zipf’s) relationship. In comparison to standard probing techniques, this framework does not require we know a priori how linguistic phenomena should manifest themselves. That is, when there is no law stating the theoretical tendencies of an attribute of natural language or we have reason to believe our language domain does not follow such a law, we can use the statistical tendencies present in empirical data as our baseline. This characteristic both allows us to assess a model’s fit to highly corpus-dependent distributions—like the length distribution—and mitigates the biases introduced by our own preconceptions regarding properties of natural language.Such biases are naturally introduced by many probing techniques that e.g., draw conclusions from hand-constructed challenge tasks.
More concretely, our paper describes an experimental design and accompanying hypothesis tests to determine precisely whether text generated from language models follows the same empirical trends as human language. Our experiments reveal that adherence to natural language tendencies varies widely with both model architecture and generation strategy, e.g., Fig. 1 shows varying degrees of adherence to the empirical type–token relationship, an artifact that perplexity alone could not reveal. Our findings suggest this framework is a valuable tool for gaining a deeper understanding of where today’s language models are succeeding and failing at capturing human language.
Language Models
Language models are probability distributions over natural language sentences. We define the support of a language model with parameters as
where is the model’s vocabulary and tokens eos and bos demarcate the beginning and end of a string, respectively, and is the Kleene closure of . In this paper, we term vocabularies consisting of words closed and those consisting of BPE tokens Sennrich et al. (2016) open.
In the case when is locally normalized, which is the predominant case for language models, is defined as the product of probability distributions:
Statistical Tendencies of Language
Human languages are thought to exhibit statistical tendencies, several of which are explicitly quantified by laws Altmann and Gerlach (2016). In this section, we review a subset of these distributions–both with and without well-established forms—over which we subsequently perform analyses.
where is the normalizing constant of our probability mass function (pmf). The adherence of language to Zipf’s law has been widely studied and is considered one of the canonical laws of quantitative linguistics Baroni (2009); Li et al. (2010); Moreno-Sánchez et al. (2016).
Estimating from an observed set of rank–frequency pairs can be done using standard estimation techniques. Here we use the maximum-likelihood estimateDerivation in App. A. We may also estimate using, e.g., least squares over the original or log–log transform of our distribution. However, it has been empirically observed that least-squares estimates under this paradigm are not reliable Clauset et al. (2009) and further, directly incorporate assumptions that contradict power law behavior Schluter (2020). (MLE), employing numerical optimization to solve for since the MLE of the discrete power law lacks a closed form solution.
Type–Token.
Heaps’ law Herdan (1960), also known as the type–token relationship, states that the number of additional unique tokens (i.e., number of types) in a document diminishes as its length increases. Formally, we can express the expected number of types as a function of the length of the string via the relationship where is a free parameter. Types may be, e.g., unigrams or bigrams.
The above formulation of Heaps’ law lacks an obvious probabilistic interpretation. However, if we frame Heaps’ law as modeling the expected value of the number of types for any given length document, then we can model the relation as a Poisson process, where the marginal distribution over document length follows Heaps’ proposed power law. Specifically, we model the number of types for a document of a given length as a non-homogeneous Poisson process (NHPP; Ross, 1996) where our rate parameter is Heaps’ power law relation. The probability that there are types in a document of length is then
for . Similarly to Eq. 4, we can fit parameters using MLE (see App. A).
2 Other Tendencies
Natural language has other quantifiable distributions, e.g., over document length or unigrams. While there may not exist well-established laws for the behavior of these (often highly corpus-dependent) distributions, we can observe their empirical distributions w.r.t. a corpus. We review a few here and leave the exploration of others to future work.
Using notation from earlier, we estimate the pmf of the distribution over the length of documents in a corpus as
Unigram.
Notably, the rank–frequency law of § 3.1 leaves the categorical distribution over words unspecified, i.e., it defines the frequency for the ranked word without specifying the word itself. In order to make explicit comparisons, we define the unigram distribution w.r.t. corpus as
Stopwords and Symbols.
Certain percentages of words in a string consist of either symbols, i.e., numbers and punctuation, or stopwords, i.e., common words such as “that” or “so” that primarily serve a syntactic function. We can model this percentage as a (continuous) random variable and estimate its probability density function (pdf) as
Statistical Distances
In this work, we aim to quantify the degree to which the linguistic distributions of text generated from language models match—or differ from—those of natural language. To this end, we propose the use of several probability metrics Mostafaei and Kordnourie (2011); Rachev et al. (2013) as our notion of statistical distance.Some of these metrics are formally pseudo-distances, as they are not necessarily symmetric. For each of these metrics, we present nonparametric statistical significance tests, i.e.,author=clara,color=orange,size=,fancyline,caption=,]not a precise def of nonparametric tests, should we tests that may be used when the underlying distribution of observed data is not known.
Perhaps the simplest method for measuring the distance between two random variables is through differences in expectations, e.g., means or variances. (Semi-)distances of this nature are formally called primary metrics. To estimate this distance, we can use observations from random samples and , e.g., .
Observing a value of on its own is not enough to confirm a difference between and ; we need to assess whether the observed distance is significantly above or below . Formally, our null and alternative hypotheses are:
In a nutshell, a permutation test provides a simple method for constructing the sampling distribution of a test statistic through empirical observations. The method uses the value of over all possible rearrangements of the observed data points to represent the distribution of the test statistic under the null hypothesis. Using this distribution, we can determine the probability of observing a value of the test statistic (or a more extreme value), which if low, may give us reason to reject a specific null hypothesis. In this work, we only consider statistics over two samples. We provide pseudocode for this case in App. B.When the number of possible permutations of the data is computationally prohibitive, we may instead use a MC sampling approach, where we sample from the set of possible permutations Good (2000).
2 Simple Metrics
Primary metrics provide only a weak measure of the sameness of random variables as they are completely dependent on a single statistic of a distribution.author=clara,color=orange,size=,fancyline,caption=,]check or reword On the other hand, we know a random variable can be completely described by its distribution function. As such, we turn to simple metrics of distance between random variables.
Given cumulative distribution functions (cdfs) and over one-dimensional random variables, the Kolmogorov–Smirnov (KS) metric is
where and indicates the distributions are identical. However, not all random variables can be described in terms of a cdf. For categorical distributions where the support of our random variable is not ordinal, the natural counterpart to the KS metric is the Chi-square distance. This metric has a number of drawbacks (discussed in App. C)—primarily that its value can be hard to interpret and so we instead turn to the total variation distance (tvd)—a widely used metric of distance between probability distributions.
Given two pmfs and , we define tvd as
where similarly to the KS metric, tvd is bounded above by 1 and a value of indicates identical distributions. In our setting, we consider two use cases for the KS metric and tvd: as distance metrics between an empirical and theoretical distribution (one-sample) and between two empirical distributions (two-sample). The corresponding hypotheses that we can test with these metrics are:
O ne-Sample Case: (12) T wo-Sample Case: (13) where in the two-sample case, the exact form of does not need to be known. These hypotheses require the following tests.
The KS test Smirnov (1948) is a nonparametric goodness-of-fit test originally designed to assess the fit of a continuous cdf to empirically-observed data; the two-sample version tests whether two samples come from the same distribution. The method has since been extended to discrete distributions and is regarded as one of the most widely applicable nonparametric goodness-of-fit tests for comparing two distributions Horn (1977); Moreno-Sánchez et al. (2016). The test uses the KS metric as its test statistic; under our null hypothesis, converges to 0 almost surely in the limit as our number of samples by the Glivenko–Cantelli theorem.Also known as the “fundamental theorem of statistics.” We may reject the null hypothesis if our test statistic is greater than the critical value, which is computed based off of our sample size and a desired significance level.Under the null hypothesis, our text statistic follows a Kolmogorov distribution. In the two sample case, the critical value is dependent on the size of both samples.
A Test for tvd.
Unlike the KS metric, we do not have a (theoretical) limiting distribution for tvd between samples from the same distribution that holds for all density functions Devroye and Győrfi (1990). However, we can construct this distribution using resampling techniques. Formally, when and are drawn from the same distribution —where need not be known—then the test statistic tvd() follows the sampling distribution , i.e., . The distribution of can be computed using permutations of our samples, in the same manner as defined in § 4.1.
Experiments
We use the above framework to assess the degree to which language models learn various distributions of natural language, i.e., we report metrics outlined in § 4 measured over the distributions and quantities defined in § 3. We compare samples generated from language models to a reserved test set taken from the same corpus as the model’s training data. Each set contains 1 million samples.Due to our large sample sizes, we should anticipate that our results will almost always be significant, even when effect sizes are trivially small. As such, we will almost assuredly reject our null hypotheses that model-generated samples come from the same distribution as natural language ones. While in this light, the presentation of hypothesis tests in § 4 may seem pointless, we provide them for cases where generating many samples for each model setting is computationally prohibitive. We tokenize all samples using the Moses decoder toolkit Koehn et al. (2007). All text is lower-cased and only complete unigrams are considered, i.e., when BPE is used, only the detokenized unigram is considered. Length of a string is computed as the number of tokens separated by whitespace. Note that when reporting the KS metric (), we always report the metric between (a) an empirical cdfauthor=clara,color=orange,size=,fancyline,caption=,]do we need to clarify how this is done? computed over the respective model-generated samples and (b) a reference cdf, where indicates direct comparison with empirical cdf of the test set. and indicate comparison with cdfs of a parametric distribution, whose parameters are estimated on the model and test set, respectively.
Simulating Corpora from Language Models.
Given the distribution , we may exactly compute statistics and distributions for language models over the entire set , weighting examples by the probability assigned to each string; however, doing so is infeasible due to the size of the output space and non-Markovian structure of most neural models. Rather, we turn to sampling to create a representative set from . We explore three sampling schemes: ancestral random sampling (Random), nucleus sampling (Nucleus), and beam sampling (Beam).The latter two sampling designs do not result in samples drawn according to our original . As such, the schemes lead to two “new” distributions, and , respectively.
In ancestral random sampling, are constructed iteratively according to the distribution
where . Under the local normalization scheme of Eq. 2, sampling according to Eq. 14 is equivalent to sampling directly from . In nucleus sampling, our distribution is truncated to the most probable items covering portion of the probability mass. Formally, we now sample
where is the smallest subset such that and . Beam sampling uses Eq. 14 as the sampling distribution, but extends a “beam” of sequences at each sampling iteration. I.e., extensions are sampled from and the most probable of the sampled items remain on the beam; note that unlike standard beam search, this is a stochastic procedure.Note that this is the default sampling scheme for language generation in the fairseq library. We use a beam size of 5 in all experiments.
Models.
We perform our tests on neural models with three different architectures: a transformer Vaswani et al. (2017); Baevski and Auli (2019) (only decoder portion), LSTM Hochreiter and Schmidhuber (1997), and Convolutional Neural Network Dauphin et al. (2017). All models are implemented and trained using fairseq.github.com/pytorch/fairseq/ We train models on corpora processed both with and without BPE. We include details for each model in Tab. 1. We additionally estimate a trigram model on the training data; formally, we build a model where the probability of observing token at position of the text is estimatedauthor=ryan,color=violet!40,size=,fancyline,caption=,]Are the indices here in the right order?author=clara,color=orange,size=,fancyline,caption=,]yes as
where denotes the function counting occurrences of a sequence in some implicit . Note that we do not employ smoothing techniques in this model, thus, perplexity over a held-out dataset may diverge and so is not reported in Tab. 1. Vocabulary statistics for each sample are shown in Fig. 2. We provide samples of model-generated text in App. E.
1 Rank–Frequency
To understand the rank–frequency relationship implicitly learned by language models—and how it relates to the rank–frequency distribution present in natural language—we compute the three KS metrics previously described: , , and . Specifically, for the first two values, we use the cdf of a Zipfian distribution parameterized by as our reference—where is estimated using model generated samples or the test set, respectively. is known to vary with the corpus size Powers (1998), however is the same for all sets, so this should not affect our analysis. These metrics give us a sense of how well the rank–frequency distribution under our language models match a Zipfian distribution. Since the power-law behavior of the token rank–frequency distribution is known to fall off at higher ranks Piantadosi (2014); Moreno-Sánchez et al. (2016), we consider solely the first 10,000 ranks in each sample, including when computing . We report these values in Tab. 2. Values of estimates of and plots of rank–frequency are shown in App. D.
Our results indicate that our models’ empirical rank–frequency distributions do not adhere very closely to a standard Zipfian distribution (as shown by and ),author=ryan,color=violet!40,size=,fancyline,caption=,]what is meant by ?author=clara,color=orange,size=,fancyline,caption=,]Im trying to say far from zero. Do you have a suggestion? despite appearing to at a superficial level (see App. D). However, the same is true for our test (), which suggests that our models fit a Zipfian distribution perhaps no more poorly than natural language does. Rather, the model produces qualitatively worst text (see App. E)—a trigram model under the beam sampling generation strategy—follows a power law trend the most closely of any of our samples. On the other hand, the small values of suggest our models learn the empirical rank–frequency trends of human text quite well, something that would not be evident by simply looking at adherence to a Zipfian distribution. The combination of these results suggest the limitation of using adherence to Zipf’s law as a gauge for a model’s consistency with natural language.
2 Type–Token
Fig. 3 shows the type–token trend for all corpora and generation schemes. While most models appear not to follow the same trend as the natural language distribution (as depicted by our test set), we observe that transformers under the nucleus sampling generation scheme match it most closely. Indeed, both models based on the transformer architecture exhibit remarkably similar trends in these experiments, despite having different vocabulary sizes and hyperparameters: both in their generally close fit to the natural language type–token distribution and in their visible fall-off for longer length sequences. The latter observation reveals a deficiency that is seemingly specific to the transformer architecture—one that may be linked to observations in natural language generation tasks. More specifically, we take this as quantitative evidence for recent qualitative observations that when left to generate lots of text, neural language models based on the transformer architecture tend to babble repetitively Holtzman et al. (2020); Cohen and Beck (2019); Eikema and Aziz (2020).
To provide a more mathematically rigorous analysis, we compute KS metrics, § 3.1 provides motivation for comparing distributions at individual time steps rather than collectively over time; analyzing Eq. 5 for all document lengths simultaneously would not give us a sense of how the power-law fit changes as a function of document length. again presenting three values: , , and . In Fig. 4, we can see that model-generated text follows a NHPP parameterized by Heaps’ law moderately well (); there are larger divergences at the tails of document length. However, most do not follow an NHPP with the same parameters as our test set (). Further, in contrast to rank–frequency, the type–token distribution is more disparate from the empirical natural language distribution than our parameterized ones, as shown by high values of . While both transformers exhibit the closest fit for all document lengths, which is in-line with our observations in Fig. 3, statistical distance from the natural language distribution for all models and in all settings increases with document length.
3 Unigram Distribution
Because we do not have a well-established law dictating the form of the natural language unigram distribution, we compare only empirical pmfs from model-generated samples and the test set directly. Further, as the distribution over unigrams is categorical, we employ tvd following § 4.2. Our results in Tab. 2 indicate that language models generally capture the unigram distribution quite well. The transformer (AS), which has a closed vocabulary, consistently performs poorly in comparison to other models. While we might speculate this outcome is a result of disparate tails between empirical cdfs—i.e., the part of the distribution over infrequent words, which may have been omitted from the closed vocabulary but could still be generated using BPE—the tvd metric in this setting should generally be robust to tail probabilities.We observe this empirically; calculating tvd between distributions truncated to the (union of the) first 1000 ranked unigrams lead to almost the exact same result. This suggests that BPE (or similar) vocabulary schemes may lead to models that can better fit this natural language distribution.
4 Length, Stopwords and Symbols
Similarly to the unigram distribution, for length, stopwords and symbols, we compare solely empirical cdfs. We use the set of English stopwords defined by NLTK Bird et al. (2009). We define the set of symbols as tokens consisting solely of punctuation and numerical values. Our results in Tab. 3 demonstrate that our language models—at least when using random and nucleus sampling—mimic these natural language distributions quite well. Notably, text generated from an LSTM using random sampling follows all three distributions the closest of any model, suggesting LSTMs may have an inductive bias that is helpful for capturing these distributions. On the other hand, using beam sampling leads to strong divergence from natural language distributions across the board. Results for differences in distribution means in the permutation testing framework can be found in App. D.
With respect to the length distribution, these results are perhaps surprising: the local-normalization scheme used by the majority of language generation models (and by those in these experiments) has been claimed to result in models that favor shorter than typical sequences Sountsov and Sarawagi (2016); Murray and Chiang (2018). The results in 3 and 5 suggest otherwise. Specifically, we see that our models fit the natural language length distribution of our corpus quite closely, in terms of both overall distributions and means (see App. D). Rather, it appears that the generation strategy may be the cause of prior observations. This finding raises further questions: since models capture the length distribution well, is a language model more likely to produce degenerate text (e.g., repetitions) than the EOS token if only long documents are used in training? We posit that corpus preprocessing should perhaps be more carefully considered in light of these results.
5 Consistent Trends
Across results, we observe that text generated using the nucleus sampling decoding scheme often aligns with natural language more closely than text produced using other generation strategies. This suggests that nucleus sampling performs a helpful alteration to a standard distribution learned via MLE, which may in turn provide motivation for recent efforts to employ truncated or sparse probability distributions directly at training time, e.g., truncated loss Kang and Hashimoto (2020) or -entmax loss Peters et al. (2019).
We additionally observe large discrepancies in both § 5.1 and § 5.2 between the results when using empirical natural language cdfs vs. parametric ones. We take this as a warning that assumptions about the forms of linguistic distributions—such as the ones employed by challenge tasks in probing—can have significant effects on results.
Related Work
author=clara,color=orange,size=,fancyline,caption=,]incorporate kuncoro-etal-2019-scalable and mueller-etal-2020-cross? or GLTR and HUSE? In the last few years, a number of works have extended language model analysis beyond simple evaluation metrics—like perplexity—in order to understand what attributes of human language these models are learning. Some use task-based approaches, i.e., they design a set of tasks that require a specific subset of linguistic knowledge then evaluate model performance on these tasks (Linzen et al., 2016; Gulordava et al., 2018; Jiang et al., 2020, inter alia). Others use model-based approaches, where a separate model is trained to perform some auxiliary task on representations learned by the model under test (Blevins et al., 2018; Giulianelli et al., 2018; Sorodoc et al., 2020, inter alia). We direct readers to Belinkov and Glass (2019) for a full survey of probing methods.
These approaches have drawbacks; for example, introducing a secondary model to determine what the original model has learned presents confounding factors Hewitt and Liang (2019). The designing of auxiliary tasks for assessing linguistic knowledge requires large manual effort and lends itself to implicit bias about how linguistic phenomena should manifest. In contrast, our work allows us to take a hands-offauthor=clara,color=orange,size=,fancyline,caption=,]better word? approach to analyzing language models. We see the benefit of this in § 5, where our results without an assumed model of statistical tendencies give us a much different sense of which empirical properties of human-generated text our models have learned.
Our work is closest to that of Takahashi and Tanaka-Ishii (2017, 2019) who use model generated text to visually analyze whether language models reflect well-established statistical tendencies. In contrast, our work provides a quantitative framework, along with appropriate significance tests,In this respect, our work is similar to Dror et al. (2018), whom also present statistical tests for use in NLP. for evaluating distribution fits. We additionally assess the fit of language models to our test set directly, rather than solely to established laws. Further, our analysis includes different generation strategies, multiple neural architectures, and a wider variety of empirical language distributions.
Conclusion and Future Directions
In this work, we present a framework for determining the linguistic properties learned by language models through analysis of statistical trends in generated text. We find that neural language models accurately capture only a subset of natural language distributions and that this subset is highly dependent on both model architecture and generation strategy; no one configuration stands out as capturing all linguistic distributions. Ultimately, we see this analysis framework as a means for a more fine-grained evaluation of language models than perplexity alone can provide. Uncovering which linguistic properties language models have learned—and which they have not—should help us to understand both the inductive biases of various models and via which avenues they can still be improved.
There are a number of important axes of variation that this work does not explore: perhaps most importantly, our results are limited to a single corpora in the English language. A cross-linguistic analysis may reveal whether different model architectures exhibit inductive biases compatible with different languages; observing how these metrics change as a function of corpus size would have implications about the effects of data availability. An exploration of the correlation of these metrics with other quantifications of model performance, such as perplexity or a model’s ability to capture sentence level phenomenon, may help us understand how comprehensive other evaluation metrics are. We leave these analyses as future work.
Acknowledgements
We thank Adhi Kuncoro for helpful discussion and feedback in the middle stages of our work and Tiago Pimentel, Jason Wei, and our anonymous reviewers for insightful feedback on the manuscript. We additionally thank B. Bou for his concern.
References
Appendix A Maximum Likelihood Estimates
The log-likelihood of our observed rank–frequency data under the power law paradigm is
denotes the function counting occurrences of in and denotes the total word count of . See Appendix B of Clauset et al. (2009) for full proof of correctness. The log-likelihood of our observed unique vs. total tokens under Eq. 5 is:
Appendix B Permutation Test Pseudocode
Appendix C Chi-square Goodness-of-fit Test
The Chi-square test uses the theoretical observation that (a function of) the difference between the expected frequencies and the observed frequencies in one or more categories follows a distribution. Observing a large value of suggests that two samples do not come from the same distribution. The Chi-square test has a major drawbacks: it is extremely sensitive to sample size. Specifically, when the sample size is large (), almost any small difference will appear statistically significant. Additionally, the statistic itself is not easy to interpret as it is not bounded above by any value.
Appendix D Additional Results
author=clara,color=orange,size=,fancyline,caption=,]check bigrams