Time Masking for Temporal Language Models
Guy D. Rosin, Ido Guy, Kira Radinsky
Introduction
The ability to learn from context is crucial for language understanding and modeling, as proved by the rise of contextual language models in recent years (Devlin et al., 2019; Liu et al., 2019). But what is context? Context has always been textual, i.e., neighboring words (Mikolov et al., 2013; Devlin et al., 2019). We argue there are other types of context that are valuable for language modeling, and propose to use time as context.
Our language is continuously evolving, especially for content on the web. The rate of change of the web, language, and culture have compressed the time in which critical changes happen (Adar et al., 2009). Many important tasks in Natural Language Processing and Information Retrieval, such as web search, must scale not only to the number of documents, but also temporally (Kanhabua and Anand, 2016; Huang and Paul, 2019; Röttger and Pierrehumbert, 2021; Savov et al., 2021). Language models (LMs) are usually pretrained on corpora derived from a snapshot of the web crawled at a specific moment in time (Devlin et al., 2019; Liu et al., 2019). This “static” nature of training prevents language models from adapting to time and generalizing temporally (Röttger and Pierrehumbert, 2021; Lazaridou et al., 2021; Hombaiah et al., 2021; Dhingra et al., 2021).
At the heart of the masked language modeling (MLM) approach is the task of predicting a masked subset of the text given the remaining, unmasked text. In this work, we adapt this approach and propose a simple way to encode time into a language model. We explicitly model temporality by adding time information directly into texts. This modification only yields changes in the input representations, and is thus model agnostic. Most existing temporal models train separate models specialized to different time periods (Hamilton et al., 2016). In contrast, we create a single model that is inherently temporal. In our approach, sequences in the training data are concatenated with a special token of the time period. The concatenation step enables the learning of word representations that are specific to each time period. For example, “<2021> Joe Biden is the President of the United States”, where <2021> is a special time token added to the model’s vocabulary by our framework to represent the year of 2021. By applying this modification contextual embedding methods can utilize the temporal information attached to the input text. We present several masking techniques to exploit this modification when pretraining LMs, as part of (or in addition to) the standard language modeling objective of BERT (Devlin et al., 2019). We refer to this process as time masking, and to this modification applied to BERT models as TempoBERT (Section 3). Figure 1 depicts the training process of TempoBERT, along with different ways to use it for inference.
It was previously argued that the embeddings generated by BERT (Devlin et al., 2019) are by definition dependent on the time-specific context and therefore do not require diachronic fine-tuning (Martinc et al., 2020b). In this work, we challenge this hypothesis by training the language model in a temporal fashion, using time masking. We experiment two tasks—semantic change detection and sentence time prediction.
Semantic change detection is the task of identifying which words undergo semantic changes and to what extent. Most existing contextual methods detect changes by first embedding the target words in each time period, and then either aggregating them to create a time-specific embedding (Martinc et al., 2020b), or computing a cluster of the embeddings for each time (Giulianelli et al., 2020; Martinc et al., 2020a; Montariol et al., 2021; Laicher et al., 2021). The embeddings or clusters are compared to estimate the degree of change between different time periods. These approaches are similar to the traditional static word embedding models, where each word from the predefined vocabulary is presented as a unique vector, and to find the degree of change, the vectors of a word at different times are compared (Hamilton et al., 2016; Schlechtweg et al., 2019). In this work, we propose a new method to detect changes, exploiting time masking (Section 4.1). Through time masking, it is no longer needed to focus on the target word embedding, but rather learn the context of the time token. TempoBERT detects changes by analyzing the distribution of the times predicted by the transformers. We experiment with several diverse datasets in terms of time, size, genre, and language. Our empirical results show that TempoBERT outperforms state-of-the-art methods (Del Tredici et al., 2019; Schlechtweg et al., 2019; Martinc et al., 2020b; Montariol et al., 2021).
The second task is sentence time prediction. A textual unit has two main temporal dimensions: its timestamp, or creation time, and its focus time (Campos et al., 2014). Both dimensions are important to meaningfully answer temporal queries in Information Retrieval systems. Knowledge of the creation time of texts is essential for many important tasks, such as temporal search, summarization, event extraction, or temporally-focused information extraction. While existing approaches for these tasks assume accurate knowledge of the creation time, it is not always available, especially for arbitrary documents from the Web. For most of the documents on the Web, the time-stamp metadata is either missing or cannot be trusted (Kanhabua and Nørvåg, 2008). Thus, predicting the time of textual units (e.g., documents, posts, or sentences) based on their content is an important task (de Jong et al., 2005; Kumar et al., 2011; Vashishth et al., 2018; Savov et al., 2021). We tackle this task as a multiclass classification problem, with classes defined as time intervals (e.g., years or decades). In this work, we focus on dating sentences, which is a harder task than dating documents, and is a building block of the document dating task (Savov et al., 2021). We use time masking to directly predict the writing time of sentences (Section 4.2). We experiment with two datasets of different granularities (i.e., years and decades) and find that TempoBERT outperforms classifiers based on GloVe (Pennington et al., 2014) or BERT (Devlin et al., 2019) embeddings.
Overall, our work offers the following contributions: (1) We propose a simple modification to pretraining that facilitates the acquisition of temporal knowledge through time masking; (2) We conduct evaluations on the tasks of semantic change detection and sentence time prediction that demonstrate the impact of temporally-scoped pretraining; (3) On semantic change detection, we reach state-of-the-art performance on various datasets; (4) We perform a quantitative analysis of time masking for the two evaluation tasks. We also publish our code and data.https://github.com/guyrosin/tempobert
Related Work
Language models are usually pretrained on corpora derived from a snapshot of the web crawled at a specific moment in time (Devlin et al., 2019; Liu et al., 2019). This “static” nature of training prevents LMs from adapting to time and generalizing temporally (Röttger and Pierrehumbert, 2021; Lazaridou et al., 2021; Hombaiah et al., 2021; Dhingra et al., 2021). Hombaiah et al. (2021) performed incremental training to better handle continuously evolving web content. Dhingra et al. (2021) experimented with temporal language models for question answering. The authors focused on probing LMs for factual knowledge that changes over time, and showed that conditioning the temporal LM on the temporal context of the text data improves memorization of facts. Others focused on document classification by using word-level temporal embeddings (Huang and Paul, 2019) and adapting pretrained BERT models to domain and time (Röttger and Pierrehumbert, 2021).
In this work, we focus on the novel concept of temporal contextual representation. We leverage our proposed model for the tasks of semantic change detection and sentence time prediction, which were not studied in the above papers. Furthermore, we learn temporal contexts for word embeddings, introduce the concept of time masking, and demonstrate its benefits for temporal LMs.
Semantic change detection is the task of determining whether and to what extent the meaning of a set of target words has changed over time, with the help of time-annotated corpora (Kutuzov et al., 2018; Tahmasebi et al., 2018). This task is often addressed using distributional semantic models: time sensitive word representations are learned and then compared between different time periods (Hamilton et al., 2016; Bamler and Mandt, 2017; Dubossarsky et al., 2019). Gonen et al. (2020) used a simple nearest-neighbors-based approach to detect semantically-changed words. Del Tredici et al. (2019) measured changes of word meaning in online Reddit communities by employing an incremental fine-tuning approach. Montariol et al. (2021) used statistical tools to detect semantic change in a scalable way using clusters. Others attempted to simultaneously learn time-aware embeddings over all time periods and resolve the alignment problem using regularization (Yao et al., 2018), modeling word usage as a function of time Rosenfeld and Erk (2018), Bayesian skip-gram model with latent time series (Bamler and Mandt, 2017), and exponential family embeddings (Rudolph and Blei, 2018).
In all of these methods, each word has a single representation for each time period, which limits their sensitivity and interpretability. Recent contextualized architectures allow for overcoming this limitation by taking sentential context into account when inferring word token representations. Indeed, such architectures were applied to diachronic semantic change detection (Hu et al., 2019; Giulianelli et al., 2020; Martinc et al., 2020b; Martinc et al., 2020a; Laicher et al., 2021). While all these studies used language models by aggregating the information from the set of the token embeddings, we directly exploit their contextual properties by integrating time directly into their training process.
In the last decade, the exploitation of temporal information to improve document search and exploration received a notable attention from the Information Retrieval community. Specifically, in this context, Temporal Information Extraction is crucial to support retrieval systems (Campos et al., 2014; Kanhabua and Anand, 2016). Predicting the time of textual units (e.g., documents, posts, or sentences) based on their content is an important task (de Jong et al., 2005; Kumar et al., 2011; Vashishth et al., 2018; Savov et al., 2021) with numerous approaches. Most existing work uses documents as textual units: One of the first studies to model temporal information for the automatic dating of documents is de Jong et al. (2005), who used a statistical LM based on word usage statistics over time. Several papers extended this model with temporal entropy (Kanhabua and Nørvåg, 2008), KL divergence (Kumar et al., 2011), and n-grams (Jatowt and Campos, 2017). Niculae et al. (2014) tackled this task as a ranking problem. As for deep neural models for document dating, Vashishth et al. (2018) used a Graph Convolution Network together with the syntactic and temporal structure of the document, and Ray et al. (2018) used an attention-based neural model. Recently, Savov et al. (2021) used Ordinal Classification to predict the time of each sentence individually and then aggregated the scores to document level.
In this work, we focus on dating sentences. Sentences are often significantly shorter than documents, making this task potentially harder than dating documents, due to the limited context. In addition, sentence dating can be used as a building block of document dating, as was done in (Savov et al., 2021). We use time masking to directly predict the writing time of sentences.
Time Masking
In this work, we use bidirectional language models (also called “contextual language models”). These models represent words as a function of the entire context of a unit of text (e.g., sentence or paragraph), and not only conditioned on previous words as in traditional unidirectional models. Formally, given an input sequence of tokens and a position , bidirectional LMs estimate using the left and right context of that word. To estimate this probability, BERT (Devlin et al., 2019) samples positions in the input sequence randomly and learns to fill the word at the masked position. To this end, a Transformer architecture is employed.
To adapt language models to time, we propose to directly encode temporality information in the training process. We design a modification to the training process designed specifically for contextual representations: rather than including time information for each word as in existing approaches (e.g., temporal referencing (Dubossarsky et al., 2019)), we concatenate a special time token to each textual sequence. For example, “¡2021¿ Joe Biden is the President of the USA”, where “¡2021¿” is a special time token added to the model’s vocabulary by our framework to represent the year of 2021. By performing this modification at the sequence level instead of at the token level, we avoid adding time-specific variants of each token to the vocabulary, thus increasing its size significantly. In addition, contextual embedding methods can utilize this temporal information attached to the input text. The time token denotes the time at which the sequence was written or a point in time at which its assertion is valid.
In addition to encoding time into the model and thus making the model temporal, time masking enables predicting the writing time of texts. The same way prompts (Petroni et al., 2019; Jiang et al., 2020) exploit the masked language model to fill mask tokens for the task of question answering, we can fill mask time tokens to predict the time of the sequence. We exploit this feature in both tasks we experiment with in this work: semantic change detection and sentence time prediction. For example, to predict the time of the sentence “Joe Biden is the President of the USA”, we prefix it with a mask time token: “[MASK] Joe Biden is the President of the USA” and feed it into the model. The model is likely to fill the mask token with a time token, since it was pretrained on such sentences. Thus, the expected token is the relevant time token to Biden’s presidency, e.g., “<2021>”.
Formally, given a sequence and a time , we denote for clarity and prepend the time token to . Now the inputs to the model are a sequence and a position . Let us analyze the two options for : (1) If , then is one of the original tokens in the sentence. In this case, , i.e., is computed based on both the textual context and the temporal context. (2) If , then is . In this case, , i.e., we predict the time by the whole input text.
As for the granularity of , different values can be used, according to the use case. In this work, for the sentence time prediction task (Section 4.2), we experiment with granularity of a year and of a decade. For the semantic change detection task (Section 4.1), the corpora are divided into two time buckets, so we use two time tokens, one for each bucket (i.e., <1> and <2>).
Throughout this paper, we use pretrained BERT-base (Devlin et al., 2019) (with 12 layers, 768 hidden dimensions, and 110M parameters)We use Hugging Face’s implementation of BERT: https://github.com/huggingface/transformers., and then post-pretrain it on the temporal corpora with time masking, as described above. We refer to the resulting model as TempoBERT.
Finally, time masking can be performed in two ways: The straightforward one is to use the standard MLM, where time tokens are treated like any other token for the masking process. But we can also use a separate masking process for time tokens. In this case, we choose time tokens with a custom probability ; a chosen time token is then either masked (i.e., replaced with the mask token [MASK]) 80% of the time, replaced by a random time token 10% of the time, or kept unchanged 10% of the time (we follow Devlin et al. (2019) for the 80-10-10 probabilities). See Section 7.1.2 for an analysis of the different settings.
Applications for Temporal Contextual Representations
Words can change their meaning, gain or lose senses over time. Knowing whether, and to what degree, a word has changed is key to various tasks, e.g., aiding in understanding historical documents, searching for relevant content, or historical sentiment analysis. In semantic change detection (Kutuzov et al., 2018; Tahmasebi et al., 2018), the task is to rank a set of target words according to their degree of semantic change between two time periods and .
We propose two scoring functions that utilize TempoBERT for this task (see Section 7.1.1 for an analysis of them): given a target word , they approximate the semantic change it underwent between and .
We generate time specific representations of words and compare them to detect semantic changes. This method was used by Martinc et al. (2020b). It is model agnostic and does not explicitly utilize time masking. In detail, we sample sentences containing the target word from each time period and feed them into the model. For each of the sentences, we create a sequence embedding by extracting their last hidden layers and averaging them. We then extract the embedding of the specific target word. This is the contextual token embedding. Note that these representations vary according to the context in which the token appears, meaning that the same word has a different representation in each specific context (sequence). Following, the resulting embeddings are aggregated at the token level and averaged, in order to get a time specific representation for for each time , denoted by . Finally, we estimate the semantic change of by measuring the cosine distance (cos_dist) between two time specific representations of the same token:
This method is tailored to TempoBERT. We use our model’s ability to predict the time of sentences and identify semantic change by looking at the distribution of the predicted times. Specifically, we sample sentences containing the target word from each time period and add them all to a set of sentences for , denoted by . For each sentence , we prepend a mask token and feed it into the model for the task of MLM. The model fills the mask token with the most appropriate token, which in our case is a time token (since during our post-pretraining, the model was trained on sentences starting with a time token). This produces a predicted time for the sentence. Let be the probability of time to be predicted for . Consider the distribution of these predicted probabilities (for each of the time periods). Intuitively, a uniform distribution means those sentences could have been written at any time period, which means there was no semantic change. A non-uniform distribution implies there was a change that caused the sentences to be written in a specific time period. In our case, since there are just two time periods, we simply use the absolute difference of the probabilities , so the semantic change score of the target word is set to the average difference of the predicted time probabilities:
2. Sentence Time Prediction
As another evaluation of TempoBERT’s generalization potential, we examine its use for the task of sentence time prediction; it is a multiclass classification problem, where the goal is to predict the time of a sentence (de Jong et al., 2005; Kumar et al., 2011; Vashishth et al., 2018; Savov et al., 2021). This task is a building block of the document dating task (Savov et al., 2021). Tackling this task with TempoBERT is straightforward: given an input sentence, we prepend a mask token and predict it. In more detail, we feed the prepended sentence into the model for the task of MLM. The model fills the mask token with the most appropriate token, which in our case is a time token (since during our post-pretraining, the model was trained on sentences starting with a time token). This produces a predicted time for the sentence.
Datasets
For semantic change detection, we experiment with several diachronic data sources, covering a variety of genres, time periods, languages, and sizes. We use the LiverpoolFC corpus (Del Tredici et al., 2019) and the SemEval-2020 Task 1 corpora (Schlechtweg et al., 2020) (English and Latin). Each dataset is split into two time periods; see Table 1 for the datasets’ statistics.
The LiverpoolFC corpus contains 8 years of Reddit posts from the Liverpool Football Club subreddit. It was created for the task of short-term meaning shift analysis in online communities. The corpus is in English and split into two time periods from 2011-2013 and 2017. We follow Martinc et al. (2020b) and only conduct light text preprocessing on the LiverpoolFC corpus, where we remove URLs.
Semantic change evaluation dataset: Del Tredici et al. (2019) published a set of 97 words from the LiverpoolFC corpus, manually annotated with semantic shift labels by the members of the LiverpoolFC subreddit. Each word was assigned a graded label (between 0 and 1) according to its degree of semantic change. The average of these judgements is used as a gold standard semantic shift index.
The SemEval-2020 Task 1 deals with the semantic change detection task. It introduced datasets in four different languages. In this work, we use the English and Latin datasets. They are all long-term, in contrast to the LiverpoolFC dataset: the English dataset spans two centuries, and the Latin dataset spans more than 2000 years.
Semantic change evaluation dataset: The task contains two subtasks: binary classification of whether a word sense has been gained or lost (subtask 1) and ranking a list of words according to their degree of semantic change (subtask 2). We use the data for subtask 2, which is a set of target words and their graded labels. The target words are balanced for part of speech (POS) and frequency. For the English dataset, we remove POS tags from both the corpus and the evaluation set (as done in previous work (Montariol et al., 2021)).
2. Sentence Time Prediction Datasets
We use the New York Times archive with 40 years of news articles (from 1981 to 2020, 11 GB of text in total). We use two variants of this corpus (Table 2); one with a granularity of decades (NYT-decades) and the other with a granularity of individual years (NYT-years). For both of them, we downsample the corpus to 10k sentences per year (1.4M tokens in average, around 1.4 MB of text) to ease computation. The year-granularity dataset contains 40 years, so we have 40 classes for our classification task. To create the decade-granularity dataset, we merge each decade. The resulting dataset contains 4 decades (i.e., four classes), each is composed of 14 MB of text. Each row in the two datasets is composed of a sentence and its associated time (year or decade). For the evaluation, we randomly split both datasets to train and test sets using an 80-20 ratio.
Experimental Setup
We use five baseline methods for semantic change detection: (1) Del Tredici et al. (2019) was the first paper to address the notion of short-term semantic change detection. They use a Skip-gram with Negative Sampling model (SGNS) (Mikolov et al., 2013) and employ an incremental model fine-tuning approach; the model is first trained on a large Reddit corpus of around 900M tokens (a random crawl from Reddit in 2013) for the initialization of the word vectors, and then fine-tuned on the LiverpoolFC corpus. (2) Schlechtweg et al. (2019) is the state-of-the-art semantic change detection method employing non-contextual word embeddings: the Skip-gram with Negative Sampling model is trained on two periods independently and aligned using Orthogonal Procrustes. Cosine distance is used to compute the semantic change. (3) Gonen et al. (2020), who use SGNS embeddings along with a simple detection method. For each period, a word is represented by its top nearest neighbors according to cosine distance. Semantic change is then measured as the size of the intersection between the nearest neighbors lists of two periods. (4) Martinc et al. (2020b) were one of the first to use BERT for semantic change detection, by creating time-specific embeddings of words, and comparing between different times by calculating the cosine distance between the averaged embeddings. Notice, unlike our work they consider time-specific context rather than capture temporal context. They focused on short-term semantic change as well. Their model performs slightly worse than Del Tredici et al. (2019), but it is more general and scalable. (5) Montariol et al. (2021) generate a set of contextual embeddings using BERT for each word. These representations are clustered, and the derived cluster distributions are compared across time slices. Finally, words are ranked according to the distance measure, assuming that the ranking resembles a relative degree of usage shift. We use their best performing method as reported in (Montariol et al., 2021), which uses affinity propagation for clustering word embeddings and Wasserstein distance as a distance measure between clusters. For the LiverpoolFC dataset, we experiment with all the variants proposed in (Montariol et al., 2021) and choose the best performing one, which uses k-means with for clustering embeddings, and Wasserstein distance as a distance measure between clusters.
Finally, existing methods for this task that concatenate time information to individual words (i.e., Temporal Referencing (Dubossarsky et al., 2019) and Word Injection (Schlechtweg et al., 2019)) were found (Schlechtweg et al., 2019, 2020) to be outperformed by Schlechtweg et al. (2019), so we leave them out of this evaluation.
For each dataset, we train a TempoBERT model using a pretrained case-insensitive BERT-base model.For English: bert-base-uncased from the Hugging Face library, and for Latin: latin-bert from https://github.com/dbamman/latin-bert. Since the vocabulary of this model does not contain all the target words in the evaluation datasets, we add all target words to TempoBERT’s vocabulary (to avoid the tokenizer splitting any occurrences of the target words to subwords).
Semantic change detection performance is measured by the correlation between the semantic shift index (i.e., the ground truth) and the model’s semantic shift assessment for each word in the evaluation set. Previous work used different correlation measures for different datasets; for the LiverpoolFC dataset, Pearson’s was used (Del Tredici et al., 2019; Martinc et al., 2020b), while for the SemEval-2020 datasets it was Spearman’s rank-order correlation coefficient (Schlechtweg et al., 2020; Montariol et al., 2021). The difference between them is that Spearman’s only considers the order of the words, while Pearson’s considers the actual predicted values. In our evaluation, we make an effort to evaluate our methods and the baselines using both Pearson and Spearman correlation, to make the evaluation as comprehensive as possible. There were some cases where we could not reproduce the original authors’ results; in such cases we opted to report only the original result.
2. Sentence Time Prediction Setup
We use the following baselines: (1) CNN (Kim, 2014): a convolutional neural network loaded with pretrained GloVe (Pennington et al., 2014) vectors. (2) Bi-LSTM + Attention (Zhou et al., 2016): an attention-based bidirectional Long Short-Term Memory Network loaded with pretrained GloVe vectors. (3) Linear GloVe: a linear model with one hidden layer of 10 units and a softmax, loaded with pretrained GloVe vectors. (4) Boosting methods based on BERT embeddings (extracted using BERT’s last four hidden layers): XGBoost and CatBoost.https://github.com/dmlc/xgboost, https://github.com/catboost/catboost
For each dataset, we train a TempoBERT model using a pretrained uncased BERT-base model with time masking probability .
As sentence time prediction is a multiclass classification task, we report accuracy and macro-F1 scores on the test sets.
3. Implementation Details
Due to limited computational resources, we train our models with a maximum input sequence length of 128 tokens. To train TempoBERT models, we tune the following hyperparameters for each dataset: learning rate and number of epochs . To use TempoBERT for semantic change detection, we use a maximum number of sentences , to ease computation (previous work used all available sentences (Martinc et al., 2020b)). Specifically for the temporal-cosine method, we tune the number of last hidden layers to use for embedding extraction . All experiments were run using four NVIDIA GeForce RTX 2080 Ti GPUs. Runtime varied by experiment; training our best TempoBERT models on each of the datasets took between 30 minutes and 4 hours.
For fine-tuning BERT on labelled data, we train for three epochs with a batch size of 32, which corresponds to the default settings recommended by Devlin et al. (2019).
Results
Table 3 shows the results for semantic change detection on the LiverpoolFC and SemEval datasets. As can be seen, TempoBERT outperforms all the baselines using both Pearson and Spearman correlation with significant correlations (), except for the SemEval-English dataset, where it is outperformed by (Montariol et al., 2021) by 0.028, although TempoBERT got a higher Spearman correlation by 0.03. For the LiverpoolFC dataset, we observe a strong correlation (Pearson’s and Spearman’s ) along with a very large performance gap compared to the baselines (which achieve around for both measures, at max). For the SemEval corpora, we observe a moderate correlation (around 0.47–0.54) and a relatively smaller (but still considerable) gap compared to the baselines. The difference originates from the fact that the LiverpoolFC corpus is short-term, i.e., it contains texts from just a few years, in contrast to the SemEval corpora, which span centuries. We further analyze this phenomenon in Section 7.1.1.
Figure 2 shows the Pearson correlation between TempoBERT’s scores and the ground truth scores for the LiverpoolFC dataset. The correlation is strong (), so the number of false positives and false negatives is considered small; there are around 10 false positive words (top-left corner), while only one false negative (bottom-right corner). Thus, while we successfully detect changed words, future work should focus on better detecting unchanged words.
As shown in Table 3, TempoBERT is particularly effective on short-term corpora. TempoBERT encodes time information at the sentence level. As a result, its strength lies in distinguishing between sentences of different times (see the evaluation on sentence time prediction in Section 4.2). For long-term corpora, the writing style between different times (e.g., centuries) is very different; this makes the time information encoded in TempoBERT less valuable.
In long-term periods, there are major difference in writing style, which makes the prediction relatively easier to make. In short time spans, the temporal information encoded into the texts by TempoBERT helps it predict the time accurately and without direct supervision. To demonstrate the difference between short and long term corpora, we include in our evaluation two recent methods that were originally evaluated only on long-term corpora (Schlechtweg et al., 2019; Montariol et al., 2021). As showed in Table 3, these two methods achieved lower results compared to both TempoBERT and the two existing methods for short-range semantic change detection (Del Tredici et al., 2019; Martinc et al., 2020b).
Finally, Table 4 shows the comparison results of the two flavors of TempoBERT for semantic change detection, i.e., time-diff and temporal-cosine, for all the corpora, measured by both Pearson and Spearman correlation. We can clearly see that time-diff outperforms temporal-cosine on the short-term LiverpoolFC corpus, while it is outperformed on the long-term SemEval corpora. This further strengthens our hypothesis, as time-diff is based on sentence time prediction, while temporal-cosine is focused on the target word and is thus less affected by context differences (e.g., writing style).
1.2. Time masking experiments
As described in Section 3, time masking can be performed in two ways: (1) As part of the standard MLM (i.e., the time tokens are treated like any other token); (2) As a separate masking process, where time tokens are masked with a custom masking probability .
In this section, we experiment with these time masking techniques: we train TempoBERT models on the LiverpoolFC corpus, each with a different method, and compare the results for the task of semantic change detection. We train all models using a 3e-4 learning rate for just 5 epochs (39K steps), due to computational resource limits.
Figure 3 shows the comparison results, where the points show the Pearson and Spearman correlation for different time masking probabilities (), and the horizontal lines show the correlations for time masking performed as part of the standard MLM. We observe a positive trend from to . In particular, for , TempoBERT benefits from custom time masking compared to standard MLM. In addition, when training the model without time masking at all (i.e., ), performance is much lower, as expected, with correlation around –.
1.3. Comparison of BERT model sizes
TempoBERT is based on a pretrained BERT model with 768 hidden dimension size and 12 transformer layers. This is the most commonly used setting of BERT and is called BERT-base. In this section, we train TempoBERT-tiny, a variant of TempoBERT that is based on a much smaller pretrained variant of BERT, called BERT-tiny. BERT-tiny has just of the parameters of BERT-base: its hidden dimension is 128, and it contains only 2 transformer layers.
Table 5 shows the comparison results, where we compare TempoBERT with TempoBERT-tiny, the best baseline for each corpus (i.e., Del Tredici et al. (2019) for LiverpoolFC, and Montariol et al. (2021) for SemEval-English), and a tiny variant of the best BERT-based baselines (i.e., Martinc et al. (2020b) for LiverpoolFC, and Montariol et al. (2021) for SemEval-English). We evaluate only on the English datasets because a pretrained BERT-tiny model is currently available only for English. As expected, TempoBERT-tiny achieves significant () but lower results compared to the standard TempoBERT, but the interesting result is that even TempoBERT-tiny outperforms the best baseline on the LiverpoolFC dataset. On SemEval-English it is outperformed but by a small margin (0.09). In addition, we observe relatively worse performance by the two tiny baselines: Martinc et al. (2020b) deteriorated from 0.473 to 0.300, and Montariol et al. (2021) declined from 0.437 to 0.376. Compared to them, the performance decline from TempoBERT to TempoBERT-tiny is much smaller (0.637 to 0.561, and 0.467 to 0.427). To conclude, the tiny version of TempoBERT shows comparable or better results to the best baselines. We hypothesize that to understand time there is no need to use extremely large models. This is encouraging for both researchers possessing limited computational resources and for Earth. In addition, the tiny version of TempoBERT is relatively stronger compared to tiny versions of the baselines.
2. Sentence Time Prediction Results
Table 6 shows the results for sentence time prediction on the two NYT datasets. The two corpora show similar results: TempoBERT outperforms all the baselines. On NYT-decades, the accuracy gap between TempoBERT and the baselines is –, while on NYT-years the gap is around –. The difference between the two corpora is reasonable, since NYT-years is a much harder task (with 40 classes compared to 4 classes of NYT-decades): we observe x5–x8 higher accuracy for the evaluated methods on NYT-decades compared to NYT-years.
In this section, we experiment with various ways to perform time masking (Section 3): we train TempoBERT models on the NYT-decades corpus, each with a different time masking method, and compare the results for the task of sentence time prediction. We trained all models for 10 epochs (16K steps), using a 2e-4 learning rate.
Figure 4 shows the comparison results, where the points show the accuracy for different time masking probabilities (), and the horizontal line shows the accuracy for time masking performed as part of the standard MLM. First, we observe a positive trend, i.e., TempoBERT improves for sentence time prediction as it performs more time masking. This is reasonable, as the time masking itself is some kind of supervision for this task. Second, when training the model without time masking at all (i.e., ), performance is much lower, as expected, with 32.38% accuracy. This is reasonable, as the training process in this case does not involve predicting time tokens at all, i.e., the model does not get any supervision for this task (i.e., zero-shot). Finally, when performing time masking as part of the standard MLM, the model achieves 49.50% accuracy. This is worse than performing time masking as a custom masking process with .
2.2. Comparison to Fine-tuned BERT
The sentence time prediction task is a supervised task. As such, our masking approach can be compared to fine-tuning. It has been argued that various types of masking can be equivalent to fine-tuning in terms of performance for supervised tasks, while being more efficient (Zhao et al., 2020). We wish to examine the different parameters of fine-tuning on the performance of the sentence time prediction task to shed more light on when masking should be applied compared to fine-tuning. Specifically, we compare TempoBERT to a BERT model fine-tuned for the task of sentence time prediction. Note that the supervision TempoBERT receives in its training process is much limited compared to fine-tuned BERT; fine-tuned BERT is trained directly on the sentence time prediction task, whereas TempoBERT is a general temporal language model that is trained on token time prediction as a small part of its pretraining process.
We experiment with different fine-tuned BERT classifiers, trained on different sizes of data. Table 7 shows the comparison results. TempoBERT reaches similar performance as the fine-tuned BERT at 50% of the data. This shows most of the information of the time can be obtained from learning on only half the corpus but continues improving as data is available. When reaching 100% of the data as compared to 50% of it, we see a relatively smaller improvement.
Conclusion
In this paper, we presented a temporal contextual language model called TempoBERT which uses time as an additional context of texts. Our technique is based on modifying texts with time information and performing time masking—specific masking for the added time information. Using time masking, we model time as part of the context of texts. This enables us to both fill mask tokens while conditioning on time, and to predict the writing time of sentences. We experimented with two tasks: semantic change detection and sentence time prediction, and showed that both benefit from time masking. For semantic change detection, we proposed a method exploiting time masking that is especially strong for short-term corpora and outperformed existing state-of-the-art methods on diverse datasets, in terms of time, size, genre, and language. In addition, we experimented with small-sized TempoBERT models and showed their relatively strong performance on this task. For sentence time prediction, time masking allows us to directly predict the time of sentences without additional fine-tuning. We showed strong performance compared to various baselines and studied different time masking settings to demonstrate the contribution of time masking to this task.
Acknowledgements
We thank Omer Levy and Yonatan Belinkov for their useful feedback.