Do Long-Range Language Models Actually Use Long-Range Context?

Simeng Sun, Kalpesh Krishna, Andrew Mattarella-Micke, Mohit Iyyer

Introduction

Understanding long documents requires modeling various discourse-level phenomena, including anaphora (Hobbs 1979; Grosz et al. 1995), argument structure (Grimshaw 1990), narrative scripts and trajectories (Schank and Abelson 1977; Labov and Waletzky 1997), and causal links between concepts (Mooney and DeJong 1985). Unfortunately, most language models (LMs) are trained to predict the next word given only a small window of local context, which prevents them from using long-range discourse structure to improve their predictions. Many research efforts over the years have attempted to address this issue: for example, Rosenfeld 1996 incorporated statistics from distant tokens to improve nn-gram models, while Ji et al. 2015 added document-level context to neural LMs.

More recently, the Transformer LM (Vaswani et al. 2017), which forms the backbone of state-of-the-art NLP systems (Devlin et al. 2019; Brown et al. 2020), has become the focus of numerous efforts to process longer input sequences. Transformer LMs are constrained by the inefficiency of the self-attention mechanism, whose complexity scales quadratically with the input sequence length. As such, more efficient methods based on sparse attention (Correia et al. 2019) and cached memory Rae et al. 2020 have been proposed to increase the maximum sequence length, which has progressed from GPT-2’s 1024 tokens (Radford et al. 2019) to 4096 tokens Zaheer et al. 2020 and finally 8192 tokens (Roy et al. 2021, Routing Transformer). When evaluated on the PG-19 benchmark dataset (Rae et al. 2020), which contains long documents in the public domain, these “long-range” Transformer LMs also reach lower perplexities than baseline models on held-out data.

How do “long-range” Transformer LMs make use of the long-range context? Do they actually encode important discourse information to improve their predictions? In this paper, we conduct a series of fine-grained analysis experiments to answer these questions, several of which are inspired by the context analysis of LSTM LMs conducted by Khandelwal et al. 2018. We focus specifically on analyzing the behavior of the state-of-the-art Routing Transformer and a simpler baseline model in the presence of various perturbations (e.g., word shuffling, random document replacement), and look closely at how different types of tokens are affected. Our results show that:

Providing long-range context (i.e., further than 2K tokens away) to these models has negligible impact on the perplexity of tokens near the end of a sequence in aggregate. However, a fine-grained analysis reveals that it does help a small set of tokens (tokens within subword clusters and those that can only be copied from the distant context), as well as particular types of books (fictional and continuous).

Despite the aforementioned improvements on a fraction of tokens, significantly perturbing the long-term context with word shuffling and random replacement has no notable impact on perplexity overall, suggesting that the evaluated models encode long-range context superficially at best.

Long-range context is not used for sequence-level prediction tasks that move outside the teacher-forced setting of the previous experiments.

While modern LMs can process much longer input sequences than those of the past, we conclude that they do not exploit the information available in the long-range context. We recommend that future research on long-range LMs includes analysis experiments such as those in our work to shed light on how and when they are using the distant context.

Background & Setup

In this section, we first provide an overview of the long-range language models analyzed in this work (Local Transformer, Routing Transformer). Next, we describe the experimental setup used in the remainder of the paper to measure the impact of the long-term context.

Given preceding context tokens w<iw_{<i} (the prefix), a language model (LM) computes the probability distribution of the next token p(wi∣w<i)p(w_{i}\mid w_{<i}). LMs are commonly evaluated using perplexity, which is the exponentiated negative log likelihood of a held-out corpus:

Modern LMs are most often implemented with Transformers (Vaswani et al. 2017), which compute vector representations for every token in the prefix at multiple layers and combine them together using self-attention. Since self-attention requires scoring every token in the sequence against every other token, it scales quadratically in both compute time and memory usage, which limits its application to very long sequences.

A simple way to improve the efficiency of self-attention blocks is to constrain the attention at each layer to a local window of the previous kk tokens. Such Transformers, which we refer to as Local Transformers, can be feasibly scaled up to large input sequence lengths. The receptive field of a Local Transformer scales linearly with the number of layers, as the lthl^{th} layer of this model can access the previous k×lk\times l tokens Luong et al. 2015; Child et al. 2019; Sukhbaatar et al. 2019.

Routing Transformer

The Routing Transformer (Roy et al. 2021, RT) takes a more intelligent approach to scaling self-attention. Specifically, the RT assigns keys and queries in self-attention to kk clusters, the centroids of which are learned during training. A routing attention strategy computes attention AA only over the queries QiQ_{i} and keys KjK_{j} that belong to the same cluster μ(Qi)\mu(Q_{i}) (i.e., those whose centroid is closest to query QiQ_{i}).

In contrast to the position-based local attention, this clustering-based attention takes the content of the token representations into account. This sparse self-attention mechanism reduces the complexity from O(N2)O(N^{2}) to O(N1.5)O(N^{1.5}) and has led to state-of-the-art results on tasks such as long-form question answering (Krishna et al. 2021).

2 Experimental setup

While the previously-described models can be trained with longer inputs than standard Transformers, it remains unclear how they make use of the additional context tokens to improve their predictions of the next word. To shed light on the behavior of long-range Transformer LMs, we perform a series of experiments in which we manipulate the input sequence (both length and content). For token-level experiments (§ 3, § 4), we only evaluate the perplexity of kk tokens near the end An artifact exhibited by RT causes tokens at the very end of a sequence to have much higher losses than the others; this phenomena does not exist with LT. After correspondence with the RT authors, we decide to exclude the very last 40 tokens (whose losses are affected) from all of our evaluations. More details about this artifact can be found in the Appendix. of an NN token-long input sequence (k≪Nk\ll N) to focus on the effect of long-range context. Language models are normally evaluated on non-overlapping sequences, with the begin and end sequence tokens receiving different amount of context. In our setting, all target tokens have roughly the same amount of context.

We conduct all of our analyses on the validation set of the PG-19 dataset (Rae et al. 2020). This dataset contains ∼\sim29K books from Project Gutenberg repository published before 1919 and was constructed specifically to evaluate long-range LMs (average document length ∼\sim69K tokens). The validation set contains 50 books We observe a significant gap of ∼\sim10 perplexity between the PG-19 test and validation sets and discover that this gap is largely due to the presence of an annotated edition of The Canterbury Tales and Other Poems. This book intertwines line-by-line annotations with the main text, which causes the preprocessed version in the dataset to be unreadable. We remove this book in all of our experiments, which decreases the test/validation perplexity gap to ∼\sim3. and ∼\sim3 million tokens in total. Evaluating every token in the validation set with large prefix sizes (e.g. 8K tokens) is computationally infeasible. On one RTX8000 GPU, it takes around 104h to evaluate the entire PG-19 validation set with sequence length 8K and target sequence 10. Thus, we set the number of target tokens per context k=10k=10 and sample a subset of 220K validation tokens for our experiments, which is the same data size used in the LM analysis experiments of Khandelwal et al. 2018. Evaluating RT in this way yields slightly better perplexity on the validation set of PG-19 than using the evaluation setting in the original RT paper (35.2 vs. 36.3). We ensure that the number of tokens sampled from each validation book is proportional to the length of that book.

Models:

As training long-range Transformer LMs is also infeasible without immense computational resources, we use publicly-available pretrained checkpoints for all of our experiments. The Routing Transformer (RT) checkpoint contains 490M parameters and processes sequences up to 8192 tokens long, achieving 33.2 perplexity on the PG-19 test set. The released checkpoint, which was trained on 128 TPUv3 cores for several weeks, has a subword vocabulary of ∼\sim 98K types, We follow the RT paper Roy et al. 2021 by scaling the loss by 1.248 before computing the perplexity in order to match the word-level perplexity reported by Rae et al. 2020. along with 8 attention heads in each of its 22 layers. The top two RT layers include content-based clustering attention while the remaining are composed entirely of local attention.

Our Local Transformer (LT) is derived from the same checkpoint as RT (and thus has identical model size), except that all clustering heads are replaced with local attention heads. It achieves slightly better perplexity on the PG-19 test set (38.3 vs. 39.3) compared to the LT model trained from scratch by Roy et al. 2021, possibly because the local attention heads learn a better representation of the weight matrices by using the information from the clustering heads. The original fully trained LT checkpoint was not made publicly available before the EMNLP submission deadline. The RT authors released a new LT checkpoint during the review period of this paper. We also evaluate this newly-released LT checkpoint and include the results in the Appendix F. Both the RT and the LT checkpoints can be found at https://github.com/google-research/google-research/tree/master/routing_transformer The attention heads in this model attend to the previous 256 tokens, which results in an effective receptive field of ∼\sim 5K tokens. A preliminary experiment verified that the clustering heads in the RT do attend to the long-range context, beyond 5K tokens, demonstrating that it is at least theoretically incorporating more context than the LT. While we would have also liked to analyze other long-range LMs such as the Compressive Transformer Rae et al. 2020 and Longformer Beltagy et al. 2020, these models do not have publicly-available PG-19 checkpoints; additionally, they differ from RT and LT in model size, which makes it hard to perform controlled experiments.

The effect of longer context

How does the size of the prefix affect the perplexity of long-range Transformer LMs? In this section, we evaluate our RT and LT checkpoints on the PG-19 validation set with varied prefix length. We discover that although these models are theoretically able to encode long sequences, increasing the prefix length beyond 2K tokens does not bring discernible improvements in aggregate. However, we do identify small subsets of tokens that benefit from long-range context. Additionally, we find that these models take advantage of long-range context to different degrees on different types of books (e.g., continuous fictional narratives vs. discontinuous magazine articles).

As shown in Figure 1, RT perplexity plateaus when evaluated with prefixes longer than 2K. As a point of comparison, the perplexity of the much smaller LSTM language models evaluated by Khandelwal et al. 2018 plateaus after 200 words. Additionally, Press et al. 2020 discover that the perplexity flattens after 1K for a smaller standard Transformer LM. In contrast, relative to RT, the perplexity curve for the more primitive LT starts flattening earlier at around 1K tokens (note that its effective context size is only 5K). We conclude that while RT’s clustering heads take better advantage of global context than LT, the long-range context beyond 2K tokens is not helping overall. Surprisingly, the perplexity gap between RT and LT is relatively consistent regardless of the prefix length, which indicates that much of RT’s gains do not come from its increased ability to leverage long-range context but rather from better modeling of local context.

Infrequent tokens can benefit from increased prefix length:

While the overall perplexity does not improve with increasing prefix length, we do observe different behavior when filtering the target tokens by frequency, as shown in Figure 2. We define frequent tokens to be the top 10% more frequently-occurring tokens in the subword vocabulary of PG-19 while the rest of tokens are classified as infrequent. Around 20K tokens in our target token set are classified as infrequent, which amounts to 9% of all target tokens. While adding long-range context does not improve either model’s predictions of frequent tokens, the RT’s perplexity of infrequent tokens decreases from ∼\sim 1200 with a prefix length of 2K to 11801180 with prefix length of 5K. However, we do observe that infrequent token perplexity increases back to 1200 as the input is further extended, suggesting that the additional context perhaps confounds the model. Meanwhile, the LT is significantly worse than RT on infrequent token prediction, and its perplexity increases as the prefix length is increased. This is likely an artifact due to the elimination of routing attention heads from the RT checkpoint, since the LT trained from scratch does not exhibit such behavior. More details are included in Appendix F.

Tokens inside a subword cluster benefit from longer contexts:

One issue with the previous experiment is that the frequency categorization was computed at a subword level and so may not exactly correspond to word frequency, especially for infrequent words (e.g., entities) that are split into multiple subwords. We therefore do a follow-up experiment by isolating all words that are split into multiple subwords, and then examining perplexity of these tokens as a function of their position in the subword cluster. For example, the word “Trocadero” is separated into three subword tokens “Tro”, “cade”, and “ro”. We specifically distinguish between the first subword in the cluster (“Tro”) from the rest of the subwords (“cade” and “ro”) in the plots shown in Figure 3. The perplexities are computed over 4.1K first and 5.1K rest subword tokens. The first subword category exhibits the same curve shape as those for infrequent tokens for both models, although the magnitude of the perplexities is far higher. The rest of the subwords are far easier for both models to predict, but the RT perplexity curve shows some positive impact from the long-range context until a prefix size of 5K tokens.

Routing Transformers are able to copy tokens that occur in the long-range context:

Target subword tokens that can be copied from somewhere in the prefix form another interesting group to analyze. Note that there is some overlap between the token categories we have analyzed so far. We verify in the Appendix C Table 2 that the overlap between these categories is not significant enough to confound the results. While this is commonplace for frequent words (e.g., determiners, pronouns), it also occurs for entities and rare words (e.g., character names in a novel); sometimes, a word can occur several thousand tokens after its last occurrence. We focus on the latter category of tokens, specifically using a prefix length of 2K tokens as a cutoff to distinguish local and long-range context. Perplexities are computed over 22k tokens which occur last time more than 2K tokens away, and 36K tokens that never appear in the prefix. In particular, the left plot in Figure 4 shows the perplexity of tokens that cannot be found in the previous 2K tokens, but occur somewhere in the long-range context (2K to 8K tokens away). While the LT curve for such tokens plateaus after 2K tokens, indicating that LT cannot take advantage of repeated words in the distant context, the RT curve steadily decreases until 8K tokens. The right plot, which shows the subset of target tokens which do not occur anywhere in the short or long-term context, decreases until about 5K tokens and then plateaus. Overall, these results show that long-range context is helpful for tokens that appear even several thousands tokens away.

Following patterns in the long-range context:

Besides the token categories examined above, we also qualitatively look at some examples that are too infrequent to analyze at scale. Interestingly, we observe some simple patterns (slightly more complex than copying) that the RT model picks up on. Specifically, it learns to increment chapter numbers even if the previous chapter title appears more than 2K tokens away: for example, when predicting “Chapter V” in the validation book Keith of the Border, modifying the previous chapter title “Chapter IV”, which occurs 2300 tokens away, to “Chapter V” causes the loss of the predicted token “V” to increase by over 10.

The impact of book type on the benefits of long-range context:

PG-19 contains a diverse array of topics, genres, and formats, not all of which equally benefit from long-range context modeling. For example, while continuous narratives (e.g., novels) certainly build up many high-level discourse structures over a long sequence of tokens, discontinuous text like magazines, textbooks, or short story collections may require primarily local modeling. To better understand the effect of the type of book on long-range LM perplexity, we annotate every book in PG-19’s validation set as either fiction or non-fiction and continuous We consider books with related but distinct sections (such as textbooks) to be discontinuous in our annotation. or discontinuous. We also annotate whether the work has been written by the same author or various authors, which is presented in the Appendix. Out of 49 books we annotated, 30 are non-fiction, Some magazines contain short stories or poems interspersed with news articles and essays; we count these as non-fiction in our analysis. 31 are discontinuous, and 25 books are both non-fiction and discontinuous.

We observe in Figure 5 that the RT model takes better advantage of long-range context for fictional and continuous books, as the perplexity for these books plateaus at around 5K tokens. Figure 6 shows fictional and continuous books exploit better the long-range context while predicting tokens within subword clusters. Overall, we find the improvement stems largely from continuous and fictional books; more details are included in Appendix B.

The perturbation of long-range context

The experiments in the previous section show that incorporating long-range context (further than 2K tokens away from the target) yields only marginal improvements to the overall perplexities of RT and LT. However, the long-range context does have a notable positive impact on a subset of tokens and book types. Do these improvements persist in the presence of severe perturbations to the distant context? If so, this would indicate that they are not encoding any complex discourse structure in the context but rather relying on surface information (e.g., token presence) to make better predictions. In this section, we perform a perturbation analysis to quantitatively measure the robustness of the state-of-the-art RT model. Figure 20 in the Appendix shows that the Local Transformer never uses context beyond 3K. Due to this limitation, we only present results on RT for in this section.

Formally, assume we are given a prefix sequence P=(w0,w1,…,wn)P=(w_{0},w_{1},\dots,w_{n}) with which we want to predict target sequence (wn+1,wn+2,…,wn+k)(w_{n+1},w_{n+2},\dots,w_{n+k}). We apply a perturbation ρ\rho to the first mm tokens of the prefix (w0:mw_{0:m}) to obtain the perturbed prefix

We define the following three perturbation operations for ρ\rho and report results averaged over five runs for each of them.

Sequence shuffling: Tokens within the perturbed window w0:mw_{0:m} are shuffled across the entire window (i.e., sentence boundaries are not respected).

Random sequence replacement: w0:mw_{0:m} is replaced with a random sequence from another validation book that is mm tokens long.

Specific token drop: Specific tokens within w0:mw_{0:m} (e.g., those that occur in the target) are dropped and replaced with the padding token.

We first apply sequence-level shuffling and random replacement to the distant context. Both operations have minimal impact on the perplexity of all tokens (Figure 7) as well as frequent/infrequent tokens (Figure 9) provided at least 2K tokens are left unperturbed. However, these perturbations do have increasing impact as the perturbations come closer to the target, especially for infrequent tokens. Zooming in on the long-range context, we find that random replacement consistently results in higher perplexity than shuffling, but also that shuffling distant context actually achieves slightly lower perplexity than when the model is given completely unperturbed prefixes. Overall, these results demonstrate that RT is insensitive to the word order of the long-range context.

Tokens inside subword clusters and tokens repeated in the distant context depend on word order:

Similar to the analysis in Section 3, the experiments above may hide impacts on small subsets of tokens, which motivates us to do a more fine-grained analysis. We find that tokens inside subword clusters (Figure 10) and those that can only be copied from long-range context (Figure 11) are sensitive to both the order and the content of the remote context. While random replacement is more harmful than shuffling for tokens that can be copied in the distant context (172 shuffled vs 174 random replacement perplexity when perturbing 6K tokens), shuffling is more detrimental for tokens inside subword clusters (3.8 vs 3.7 perplexity when perturbing 6K tokens).

Routing Transformer encodes token identity in the long-range context:

While the previous perturbations affected entire contiguous blocks of the prefix, we move now to more targeted perturbations. An interesting question to ask given the observation that RT perplexity decreases on copied tokens as sequence length increases (§ 3) is how much that perplexity decrease depends on word order and surrounding content. In response, we drop tokens in the distant context whose next appearance is in the target sequence. As a control experiment, we drop the same number of random tokens for each perturbation length.

As shown in the right plot of Figure 12, dropping the previous long-range occurrences of target tokens increases the perplexity of those target tokens, which shows that RT indeed memorizes token identity in the long-range context to some extent. The left plot shows that dropping long-range duplicate tokens does not affect tokens that also occur within the local context (i.e., the prior 2K tokens). The flat curve before 6K indicates the model relies only on the most recent occurrences for prediction.

Sequence-level analysis

All of the previous experiments have focused on token-level perplexity, which is the standard way in which LMs are evaluated. However, the prefixes in these evaluations consist solely of ground-truth text, mirroring the “teacher-forcing” setup that LMs are trained with. When these models are deployed practically to generate text, they have to rely on their previous predictions instead of ground-truth text, and several prior works have noted different behavior in this setting Wang and Sennrich 2020; Holtzman et al. 2020; Welleck et al. 2020. In this section, we shift from token-level tasks to analyzing RT and LT performance on sequence-level tasks. In particular, we first look at how well the models can memorize an exact sequence in the distant context, as opposed to a single token as we did previously. Next, we examine the models’ ability to identify which of six 128-token suffixes follows a given prefix, which examines their behavior outside the standard teacher-forced setting.

As a sequence-level analogue to the token-copying analysis in the previous section, we examine both RT and LT’s ability to memorize a sequence that occurs in the distant context. To test this ability, we copy the target sequence and paste it into different positions of the prefix. The left plot in Figure 13 shows that both models give a very low perplexity to the target sequence if its duplicate appears within the previous 512 tokens. However, both models lose their ability to take advantage of the copied sequence if it occurs more than 2K tokens away. This confirms our previous discovery that sequence order is in general not encoded in the long-range context.

Suffix identification:

To move beyond token-level experiments, we adopt a similar setting as the multiple choice task in SWAG (Zellers et al. 2018). Specifically, a prefix is paired with the ground-truth next 128 tokens (or suffix) as well as five randomly sampled sequences of length 128 that come from the same book and do not occur in the prefix or gold suffix. We constrain the prefix to end at a full stop and each candidate suffix to start from a new sentence so that the difference in perplexity is not due to ungrammaticality. An example is shown in Table 1. We construct 7K examples and compute the accuracy of both models at correctly choosing the correct suffix. The model makes a correct prediction when the gold suffix has lower perplexity than all other suffixes. As shown in the right plot of Figure 13, increasing prefix length beyond 2K does not improve suffix identification accuracy. Surprisingly, the LT and RT model have almost identical (and poor) performance on this task. While evaluating on the newly released LT checkpoint, the performance of LT is slightly worse, but the trend is similar. Adding context beyond 2K tokens does not keep improving the suffix identification accuracy. We direct reader to Appendix F for more details. While RT is a significantly better LM in terms of token-level perplexity, it does not appear to be superior in terms of using long-range context to improve sequence prediction. Overall, both models often predict obviously wrong negative suffixes: the full version of Table 1 together with RT’s perplexity score for each suffix is included in Appendix E.

Combined with our previous token-level analysis, we conclude that the distant context helps a subset of tokens in superficial ways; however, distant context is currently not helpful for sequence-level prediction tasks.

Related work

Our work examines recent advances in efficient Transformer variants Sukhbaatar et al. 2019; Kitaev et al. 2020; Choromanski et al. 2021; Tay et al. 2021; Katharopoulos et al. 2020; Wang et al. 2020; Wu et al. 2020 that accept longer sequences than prior approaches. Longer effective context size is often achieved by sparse attention Child et al. 2019, recurrence Dai et al. 2019, and cached memory Weston et al. 2015; Rae et al. 2020. Our work is also related to methods that incorporate long context Wang and Cho 2016 as well as document-level tasks that inherently require modeling long-range context Zhang et al. 2018; Hofstätter et al. 2020; Zhang et al. 2020.

This work is also similar to other analysis of language models, especially for long-range context. Khandelwal et al. 2018 analyze the usage of long-term context of smaller LSTM LMs. Sharan et al. 2018 prove long-term context is not needed for HMM LM due to teacher forcing. Rae and Razavi 2020 conduct an analysis exclusively for the Transformer-XL Dai et al. 2019 model. Rae et al. 2020 show that Compressive Transformer improves the performance of infrequent tokens. Our work also relates to that of Lai et al. 2020, who investigate the impact of context for pretrained masked LMs. More recently, Press et al. 2020 also observe negligible benefits of long-term context; we step further in this direction by exploring larger models with more fine-grained analysis.

Conclusion

We perform a fine-grained analysis of the impact of long-range context to both token- and sequence-level improvements on two long-range Transformer language models, using the PG-19 dataset as a testbed. Our results suggest these models rarely take advantage of the long-term context, and when they do it is mostly in superficial ways (e.g, by copying rare tokens from far away). With the proliferation of research in increasing the input size of Transformer LMs, we hope that our research will lead to more meaningful progress on integrating discourse information into these models.

Ethical concerns

The two large language models we evaluated in this work share common ethical concerns with works on language models and language generation. These pre-trained LMs can be used maliciously to generate unfaithful, hallucinated, and biased output. Our reported results do not include any kind of generation.

Energy costs

We conduct all our analysis experiments on RTX8000 GPUs. Although our work does not include training large language models, the energy costs of evaluating large pre-trained LMs, such as the Routing Transformer, should not be ignored. Each example of 8K tokens long takes around 1.3s ∼\sim 1.4s to run one forward pass with the RT model. We hope our analysis can shed light on more efficient and effective method to encode long-term context.

Acknowledgements

We are grateful to Aurko Roy for releasing the code and checkpoints and for discussing the Routing Transformer results with us. We thank Nader Akoury, Brendan O’Connor and the rest of UMass NLP group for the great advice on the draft of this paper. We thank the anonymous reviewers for their thoughtful comments on the paper. This project was partially funded by a grant from Intuit AI and also by award IIS-1955567 from the National Science Foundation (NSF).

References

Appendix A Routing Transformer and end-sequence degradation

In our analysis, instead of picking the last 10 tokens in a sequence, we chose the last 50 to last 40 tokens due to an artifact introduced by the clustering heads in the RT model. We find that in general the last 20 tokens in a sequence tend to have increasing perplexity as we evaluate on longer and longer sequence lengths. As shown in Figure 14, this phenomenon is only native to the RT model and disappears when the clustering heads are replaced with local attentions. Therefore, to make the RT and the LT comparable, we select the tokens from the range that is not affected by the end-sequence issue. Although it’s not the last 10 tokens, this short target chunk is still located near the end of a sequence, preceded by enough long context.

Appendix B Effect of longer context

In section 3 we discussed that books that are fictional and continuous benefit more from the long-range context. We also annotated the validation set by the authorship (i.e., whether a book is written by single author or various authors). Out of 49 books, 11 are written by various authors, 10 of which are non-fictions. Due to this overlap, we only show results of fic/non-fic in the main text.

In this section , we also further break down all targets to the three types of tokens we examined in section 3, and display the results by book types. Perplexity of infrequent tokens (Figure 15), tokens inside subword clusters (Figure 16), and tokens whose last occurrence is more than 2K tokens away (Figure 17). In general, for the small set of tokens whose perplexity keep decreasing as adding in more context, the major source of improvements are from the continuous and fictional books.

Appendix C Token overlaps

We have shown in section 3 that infrequent tokens, tokens inside subword clusters, and tokens that can only be copied from distant context benefit from context longer than 2K tokens. It is possible that these improvements come from the same set of tokens shared across these three types of tokens. To verify if there are significant overlaps among those three types of tokens, we compute the overlapped ratio in Table 2. Except for in-subword and infrequent tokens, the overlapped ratios are all below 0.1.

Appendix D Perturbation

Perplexity with perturbed distant prefix when evaluated with Local Transformer is shown in Figure 20. Perplexity hardly changes when perturbing up to around 6K tokens. Because LT doesn’t properly use long-range context, we only present the results of Routing Transformer in the main text.

Appendix E Suffix Identification

In the main text, we present the suffix identification results with 128-token long suffix. Here, we provide results when evaluate the accuracy with suffix of different length. Interestingly, the accuracy of distinguishing a gold suffix first increases with the suffix length, reaching the best when the suffix length is around 10 to 20, then decreases as the suffix becomes longer. This is likely because the sequence becomes more probable as incorporating more local context (part of the suffix).

In Table 3 and Table 4, we present a complete example of the suffix identification task. The prefix contains 1024 tokens, and each of the suffixes contains 128 tokens. Lower perplexity of obviously wrong suffix (e.g. negative 1,3,4) indicates current RT model is not properly taking advantage of distant context to make sequence-level predictions.

Appendix F Local Transformer checkpoint results

In the main text, we analyzed both the RT checkpoint and an LT model derived from the same RT checkpoint by replacing the clustering heads with local attention heads. After the submission deadline and before the camera ready deadline, the author of the Routing Transformer released a new LT checkpoint, which has 24 layers in total Both the RT model and the former LT model have 22 layers.. To make sure the behavior of our former LT is the same with an LT trained from scratch, we also conducted all our analysis again on this new checkpoint. Overall, we find the new LT checkpoint performs slightly better on token-level tasks but is inferior to the previous LT checkpoint on suffix identification, a sequence-level task. Since the trends are the same, no conclusion needs to be changed.

In Figure 18 are the perplexities in aggregate over all target tokens of both the RT and the new LT checkpoints. The target tokens are the same ones we used in the main text. The new LT has better perplexity than the one presented in the main text (36.5 vs. 40 as the prefix length extends to 8K).

Perturbation of long-range context

Similar to the former LT (Figure 20), the newly released LT checkpoint(Figure 19) is not sensitive at all to the perturbation further than local 2K tokens. Both models are impacted by local random replacement more than shuffling, however, the new LT checkpoint has overall better perplexity than the RT-derived LT.

Sequence-level tasks

Figure 22 shows the performance of both RT and the released LT checkpoint on sequence-level tasks. Compared to the results in the main text, the new LT checkpoint performs better at sequence-copying task, however, the trend remains the same – the order of tokens beyond 2K tokens is not properly encoded. On the other hand, the new LT checkpoint is slightly worse in suffix identification while the former RT-derived LT has almost identical performance as the RT. This implies even though the clustering heads are removed from the previous LT, useful information is preserved by the local attention heads.

Overall, we find the released LT checkpoint has better token-level performance but performs worse on suffix identification. Each plot shares the same trend with its corresponding one in the main text, thus no conclusion needs to be modified.