TaCL: Improving BERT Pre-training with Token-aware Contrastive Learning

Yixuan Su, Fangyu Liu, Zaiqiao Meng, Tian Lan, Lei Shu, Ehsan Shareghi, Nigel Collier

Introduction

Since the rising of BERT (Devlin et al., 2019), masked language models (MLMs) have become the de facto backbone for almost all natural language understanding (NLU) tasks. Despite their clear success, many existing language models pre-trained with MLM objective suffer from the anisotropic problem Ethayarajh (2019). That is, their token representations reside in a narrow subset of the representation space, therefore being less discriminative and less powerful in capturing the semantic differences of distinct tokens.

Recently, great advancement has been made in continually training MLMs with unsupervised sentence-level contrastive learning, aiming at creating more discriminative sentence-level representations (Giorgi et al., 2021; Carlsson et al., 2021; Yan et al., 2021; Kim et al., 2021; Liu et al., 2021b; Gao et al., 2021). However, such representations are only evaluated as sentence embeddings and there is no evidence that they will benefit other well-established NLU tasks. We show that these approaches hardly bring any benefit to challenging tasks like SQuAD Rajpurkar et al. (2016, 2018).

In this paper, we argue that the key of obtaining more discriminative and transferrable representations lies in learning contrastive and isotropic token-level representations. To this end, we propose TaCL (Token-aware Contrastive Learning), a new continual pre-training approach that encourages BERT to learn discriminative token representations. Specifically, our approach involves two models (a student and a teacher) that are both initialized from the same pre-trained BERT. During the learning stage, we freeze the parameters of the teacher and continually optimize the student model with (1) the original BERT pre-training objectives (masked language modelling and next sentence prediction) and (2) a newly proposed TaCL objective. The TaCL loss is obtained by contrasting the student representations of masked tokens against the “reference” representations produced by the teacher without masking the input tokens. In Figure 1, we provide an overview of our approach.

We extensively test our approach on a wide range of English and Chinese benchmarks and illustrate that TaCL brings notable performance improvements on most evaluated datasets (section 3.1.1). These results validate that more discriminative and isotropic token representations lead to better model performances. Additionally, we highlight the benefits of using our token-level method compared to current state-of-the-art sentence-level contrastive learning techniques on NLU tasks (section 3.2.1). We further analyze the inner workings of TaCL and its impact on the token representation space (section 3.2.2).

Our work, to the best of our knowledge, is the first effort on applying contrastive learning to improve token representations of Transformer models. We hope the findings of this work could facilitate further development of methods on the intersection of contrastive learning and representation learning at a more fine-grained granularities.

Token-aware Contrastive Learning

Note that the learning of the student is fully unsupervised and can be realized using the original pre-training corpus. After the learning is completed, we fine-tune the student model on downstream tasks.

Experiment

We test our approach on a wide range of benchmarks in two languages. For English benchmarks, we evaluate the BERTbase\textup{BERT}_{\textup{base}} and BERTlarge\textup{BERT}_{\textup{large}} models. For Chinese benchmarks, we test the BERTbase\textup{BERT}_{\textup{base}} model.All models are officially released by Devlin et al. (2019). After initializing the student and teacher, we continually pre-train the student on the same Wikipedia corpus as in Devlin et al. (2019) for 150k steps. The training samples are truncated with a maximum length of 256 and the batch size is set as 256. The temperature τ\tau in Eq. (1) is set as 0.01. Same as Devlin et al. (2019), we optimize the model with Adam optimizer Kingma and Ba (2015) with weighted decay, and an initial learning rate of 1e-4 (with warm-up ratio of 10%).

For English benchmarks, we use the GLUE dataset Wang et al. (2019) which contains a variety of sentence-level classification tasks covering textual entailment (RTE and MNLI), question-answer entailment (QNLI), paraphrase (MRPC), question paraphrase (QQP), textual similarity (STS-B), sentiment (SST-2), and linguistic acceptability (CoLA). Our evaluation metrics are Spearman correlation for STS-B, Matthews correlation for CoLA, and accuracy for the other tasks; the macro average score is also reported. Additionally, we conduct experiments on SQuAD 1.1 Rajpurkar et al. (2016) and 2.0 Rajpurkar et al. (2018) datasets that evaluate the model’s performance on the token-level answer-extraction task. The dev set results of Exact-Match (EM) and F1 scores are reported.

For Chinese benchmarks, we evaluate our model on two token-level labelling tasks, including name entity recognition (NER) and Chinese word segmentation (CWS). For NER, we use the Ontonotes Weischedel et al. (2011), MSRA Levow (2006), Resume Zhang and Yang (2018), and Weibo He and Sun (2017) datasets. For CWS, we use the PKU, CityU, and AS datasets from SIGHAN 2005 Emerson (2005) for evaluation. The standard F1 score is used for evaluation.

Baselines: We compare against two baselines: (1) the original BERT used to initialize the student and teacher; (2) BERT+MT (BERT with more training) which is acquired by continually pre-training the original BERT on Wikipedia for 150k stepsThe number of steps is set the same as our TaCL training. using the original BERT pre-training objectives.

Table 1 reports the results on English and Chinese benchmarks.For all tasks, the average results over five runs are reported. We observe that, on most sequence-level classification tasks in GLUE, TaCL outperforms BERT and BERT+MT. Additionally, on all token-level benchmarks (SQuAD, NER, and CWS), TaCL consistently and notably surpasses other baselines. These results indicate that the learning of an isotropic token representation space is beneficial for the model’s performance, especially on the token-centric tasks.

2 Analysis

In this section, we present further comparisons and in-depth analysis of the proposed approach.

We compare TaCL against existing sentence-level contrastive learning methods, including DeCLUTR Giorgi et al. (2021), SimCSE Gao et al. (2021), and MirrorBERT Liu et al. (2021b). We also include two ablated models to study the effect of different combinations of pre-training objectives. Specifically, the ablated model-1 is initialized with BERT and trained with the original BERT objectives (LMLM\mathcal{L}_{\textup{MLM}} and LNSP\mathcal{L}_{\textup{NSP}}) plus the sentence-level contrastive objective as proposed in Liu et al. (2021b). The ablated model-2 is initialized with BERT and trained only with the proposed token-aware contrastive objective of Eq. (1). Note that all compared models have the same size as the BERTbase\textup{BERT}_{\textup{base}} model.

Table 2 shows the performance of different models on SQuAD. We observe decreased performance of existing sentence-level contrastive methods compared with the original BERT. This could be attributed to the fact that such methods only focus on learning sentence-level representations while ignoring the learning of individual tokens. This behaviour is undesired for tasks like SQuAD that demands informative token representations. Nonetheless, the ablated model-1 shows that the original BERT pre-training objective (LMLM\mathcal{L}_{\textup{MLM}} and LNSP\mathcal{L}_{\textup{NSP}}) remedies, to some extent, the performance degeneration caused by the sentence-level contrastive methods. On the other hand, the ablated model-2 demonstrates that our token-aware contrastive objective helps the model to achieve improved results by learning better token representations.

2.2 Token Representation Self-similarity

To analyze the token representations learnt by TaCL and BERT, we follow Ethayarajh (2019) and define the averaged self-similarity of the token representations within one sequence x=[x1,...,xn]x=[x_{1},...,x_{n}] as,

where hih_{i} and hjh_{j} are the token representations of xix_{i} and xjx_{j} produced by the model. Intuitively, a lower s(x)s(x) indicates that the representations of tokens within the sequence xx are less similar to each other, therefore being more discriminative.

We sample 50k sentences from both Chinese and English Wikipedia and compute the self-similarity of representations over different layers. Figure 2 plots the results of TaCLbase\textup{TaCL}_{\textup{base}} and BERTbase\textup{BERT}_{\textup{base}} averaged over all sentences. We see that, in the intermediate layers, the self-similarity of TaCL is higher than BERT’s. In contrast, at the top layer (layer 12), TaCL’s self-similarity becomes notably lower than BERT’s, demonstrating that the final output token representations of TaCL are more discriminative.

We sample one sentence from Wikipedia and visualize the self-similarity matrix MM (where Mi,j=cosine(hi,hj)M_{i,j}=\textup{cosine}(h_{i},h_{j})) produced by BERTbase\textup{BERT}_{\textup{base}} and TaCLbase\textup{TaCL}_{\textup{base}}. The results are shown in Figure 3, where a darker color denotes a higher self-similarity score.The entries Mi,iM_{i,i} in the diagonal have a 1.01.0 self-similarity by definition, as cosine(hi,hi)=1.0\textup{cosine}(h_{i},h_{i})=1.0. We see that, as compared with BERT (Fig. 3(a)), the self-similarities of TaCL (Fig. 3(b)) are much lower in the off-diagonal entries. This further highlights that the individual token representations of TaCL are more discriminative, which in return leads to improved model performances as demonstrated (section 3.1.1, section 3.2.1).

Conclusion

In this work, we proposed TaCL, a novel approach that applies token-aware contrastive learning for the continual pre-training of BERT. Extensive experiments are conducted on a wide range of English and Chinese benchmarks. The results show that our approach leads to notable performance improvement across all evaluated benchmarks. We then delve into the inner-working of TaCL and demonstrate that our performance gain comes from a more discriminative distribution of token representations.

Acknowledgments

The first author would like to thank Jialu Xu and Piji Li for the insightful discussions and support. Many thanks to our anonymous reviewers and area chairs for their suggestions and comments.

Ethical Statement

We honor and support the ACL code of Ethics. Language model pre-training aims to improve the system’s performance on downstream NLU tasks. All pre-training corpora and pre-trained models used in this work are publically available. In addition, all evaluated datasets are from previously published works, and in our view, do not have any attached privacy or ethical issues.

References

Appendix A Statistics of Evaluated Benchmarks

A.2 Chinese Benchmarks

Appendix B Related Work

Since the introduction of BERT Devlin et al. (2019), the research community has witnessed remarkable progress in the field of language model pre-training on a large amount of free text. Such advancements have led to significant progresses in a wide range of natural language understanding (NLU) tasks Liu et al. (2019); Yang et al. (2019); Clark et al. (2020); Lan et al. (2021) and text generation tasks Radford et al. (2019); Lewis et al. (2020); Raffel et al. (2020); Su et al. (2021a, e, g, d, f, c); Zhong et al. (2021)

Generally, contrastive learning methods distinguish observed data points from fictitious negative samples. They have been widely applied to various computer vision areas, including image Chopra et al. (2005); Oord et al. (2018) and video Wang and Gupta (2015); Sermanet et al. (2018). Recently, Chen et al. (2020) proposed a simple framework for contrastive learning of visual representations (SimCLR) based on multi-class N-pair loss. Radford et al. (2021); Jia et al. (2021) applied the contrastive learning approach for language-image pretraining. Xu et al. (2021); Yang et al. (2021) proposed a contrastive pre-training approach for video-text alignment.

In the field of NLP, numerous approaches have been proposed to learn better sentence-level Reimers and Gurevych (2019); Wu et al. (2020); Meng et al. (2021a); Liu et al. (2021b); Gao et al. (2021); Su et al. (2021b) and lexical-level (Liu et al., 2021a; Vulić et al., 2021; Liu et al., 2021c; Wang et al., 2021) representations using contrastive learning. Different from our work, none of these studies specifically investigates how to utilize contrastive learning for improving general-purpose token-level representations. Beyond representation learning, contrastive learning has also been applied to other NLP applications such as NER (Das et al., 2021) and summarisation (Liu and Liu, 2021), knowledge probing for pre-trained language models (Meng et al., 2021b), and open-ended text generation (Su et al., 2022).

Many researchers Xu et al. (2019); Gururangan et al. (2020); Pan et al. (2021) have investigated how to continually pre-train the model to alleviate the task- and domain-discrepancy between the pre-trained models and the specific target task. In contrast, our proposed approach studies how to apply continual pre-training to directly improve the quality of model representations which is transferable and beneficial to a wide range of benchmark tasks.

Appendix C More Self-similarity Visualizations

In Figure 4, 5, and 6, we provide three more comparisons between the self-similarity matrix produced by TaCL and BERT (the example sentences are randomly sampled from Wikipedia).All results are generated by models with base size. From the figures, we can draw the same conclusion as in section section 3.2.2, that the token representations of BERT follow an anisotropic distribution and are less discriminative. On the other hand, the token representations of TaCL better follow an isotropic distribution, therefore different tokens become more distinguishable with respect to each other.