Contextual Representation Learning beyond Masked Language Modeling

Zhiyi Fu, Wangchunshu Zhou, Jingjing Xu, Hao Zhou, Lei Li

Introduction

In the age of deep learning, the basis of representation learning is to learn distributional semantics. The target of distributional semantics can be summed up in the so-called distributional hypothesis (Harris, 1954): Linguistic items with similar distributions have similar meanings. To model similar meanings, traditional representation approaches (Mikolov et al., 2013; Pennington et al., 2014) (e.g., Word2Vec) model distributional semantics by defining tokens using context-independent (CI) dense vectors, i.e., word embeddings, and directly aligning the representations of tokens in the same context. Nowadays, pre-trained language models (PTMs) (Devlin et al., 2019; Radford et al., 2018; Qiu et al., 2020) expand static embeddings into contextualized representations where each token has two kinds of representations: context-independent embedding, and context-dependent (CD) dense representation that stems from its embedding and contains context information. Although language modeling and representation learning have distinct targets, masked language modeling is still the prime choice to learn token representations with access to large scale of raw texts (Peters et al., 2018; Devlin et al., 2019; Raffel et al., 2020; Brown et al., 2020).

It naturally raises a question: How do masked language models learn contextual representations? Following the widely-accepted understanding (Wang and Isola, 2020), MLM optimizes two properties, the alignment of contextualized representations with the static embeddings of masked tokens, and the uniformity of static embeddings in the representation space. In the alignment property, sampled embeddings of masked tokens play as an anchor to align contextualized representations. We find that although such local anchor is essential to model local dependencies, the lack of global anchors brings several limitations. First, experiments show that the learning of contextual representations is sensitive to embedding quality, which harms the efficiency of MLM at the early stage of training. Second, MLM typically masks multiple target words in a sentence, resulting in multiple embedding anchors in the same context. This pushes contextualized representations into different clusters and thus harms modeling global dependencies.

To address these challenges, we propose a novel Token-Alignment Contrastive Objective (TACO) to directly build global anchors. By combing local anchors and global anchors together, TACO achieves better performance and faster convergence than MLM. Motivated by the widely-accepted belief that contextualized representation of a token should be the mapping of its static embedding on the contextual space given global information, we propose to directly align global information hidden in contextualized representations at all positions of a natural sentence to encourage models to attend same global semantics when generating contextualized representations. Concerning possible relationships between context-dependent and context-independent representations, we adopt the simplest probing method to extract global information via the gap between context-dependent and context-independent representations of a token for simplification, as shown in Figure 1. To be specific, we define tokens in the same context (text span) as positive pairs and tokens in different contexts as negative pairs, to encourage the global information among tokens within the same context to be more similar compared to that from different contexts.

We evaluate TACO on GLUE benchmark. Experiment results show that TACO outperforms MLM with average 1.2 point improvement and 5x speedup (in terms of sample efficiency) on BERT-small, and with average 0.9 point improvement and 2x speedup on BERT-base.

The contributions of this paper are as follows.

We analyze the limitation of MLM and propose a simple yet efficient method TACO to directly model global semantics.

Experiments show that TACO outperforms MLM with up to 1.2 point improvement and up to 5x speedup on GLUE benchmark.

Understanding Language Modeling

The key idea of MLM is to randomly replace a few tokens in a sentence with the special token [MASK] and ask a neural network to recover the original tokens. Formally, we define a corrupted sentence as x1\bm{x}_{1}, x2\bm{x}_{2}, ⋯\cdots, xL\bm{x}_{L}, and feed it into a Transformers encoder (Vaswani et al., 2017), the hidden states from the final layer are denoted as h1\bm{h}_{1}, h2\bm{h}_{2}, ⋯\cdots, hL\bm{h}_{L}. We denote the embeddings of the corresponding original tokens as e1\bm{e}_{1}, e2\bm{e}_{2}, ⋯\cdots, eL\bm{e}_{L}. The MLM objective can be formulated as:

where M\mathcal{M} denotes the set of masked tokens and ∣V∣|\mathcal{V}| is the size of vocabulary V\mathcal{V}. mi\bm{m}_{i} is hidden state of the last layer at the masked position, and can be regarded as a fusion of contextualized representations of surrounding tokens. Following the widely-accepted understanding (Wang and Isola, 2020), Eq.1 optimizes: (1) the alignment between contextualized representations of surrounding tokens and the context-independent embedding of the target token and (2) the uniformity of representations in the representation space.

In the alignment part, MLM relies on sampled contextual-independent embeddings of masked tokens as anchors to align contextualized representations in contexts, as shown in Figure 2. Local anchor is the key feature of MLM. Therefore, the learning of contextualized representations heavily relies on embedding quality. In addition, multiple local anchors in a sentence tend to pushing contextualized representations of surrounding tokens closer to different clusters, encouraging models to attend local dependencies where global semantics are neglected.

2 Empirical Analysis

To verify our understanding, we conduct comprehensive experiments to investigate: How does embedding anchor affect the learning dynamics of MLM? We re-train a BERT-small (Devlin et al., 2019) model with the MLM objective solely and analyze the changes in its semantic space during pre-training. The training details are described in Appendix A.

In general, if contextualized representations are well learned, the contextualized representations in a same context will have higher similarity than that of in different contexts. Naturally, we use the gap between intra-sentence similarity and inter-sentence similarity to evaluate contextual information in contextualized representations. We call this gap as contextual score. The similarity can be evaluated via probing methods like L2 distance, cosine similarity, etc. We observe similar findings on different probing methods and only report cosine similarity here for simplification. Figure 3(b) shows how contextual score changes during training. Other statistical results are listed in Appendix A.

Embedding similarity evaluation.

To observe how sampled embeddings affect contextualized representation learning, we evaluate the embedding similarity between co-occurrent tokens. Motivated by the target that co-occurrent tokens should have similar representations, we use the similarity score calculated by cosine similarity between co-occurrent words labeled by humans (sampled from the WordSim353 dataset (Agirre et al., 2009)) as the evaluation metric. Figure 3(a) shows how embedding similarity between co-occurrent tokens changes during training.

The learning of contextualized representations heavily relies on embeddings similarity. As we can see from Figure 3(a), the embedding similarity between co-occurrent tokens first decreases during the earliest stage of pre-training. It is because all embeddings are randomly initialized with the same distribution and the uniformity feature in MLM pushes tokens far away from each other, thus resulting in the decrease of embedding similarity. Meanwhile, the contextual score, i.e., the gap between intra-context similarity and inter-context similarity in Figure 3(b), does not increase at the earliest stage of training. It shows that random embeddings provide little help to learn contextual semantics. During 5K-10K iterations, only when embeddings become closer, contextualized representations in the same context begin to have similar features. At this stage, the randomly sampled embeddings from the same sentence, i.e., the same context, usually have similar representations and thus MLM can push contextualized tokens closer to each other.

We further verify the effects of embedding quality in Figure 4. To this end, we train two BERT models whose embedding matrices are frozen and initialized with the ones from different pre-training stage. We can see the model initialized with random embedding fails to teach contextualized representations to attend sentence meanings and representations from different contexts have almost the same similarity. However, the variant with well-trained but frozen embeddings learns to distinguish different contexts early at around 4k steps. These statistical observations verify that embedding anchors bring the efficiency and effectiveness problem.

Figure 3(a) shows that embedding similarity begins to drop after 8k steps. It shows that the model learns the specific meanings of co-occurrent tokens and begins to push them a little bit far away. Since MLM adopts local anchors, these local embeddings push contextualized representations into different clusters. The contextual score begins to decrease too. This phenomenon proves the embedding bias problem where the learning of contextualized representations is decided by the selected embeddings where the global contextual semantics are neglected.

Proposed Approach: TACO

To address the challenges of MLM, we propose a new method TACO to combine global anchors and local anchors. We first introduce TC, a token-alignment contrastive loss which explicitly models global semantics in Section 3.1, and combine TC with MLM to get the overall objective for training our TACO model in Section 3.2.

To model global semantics, the objective is expected to be capable of explicitly capturing information shared between contextualized representation of tokens within the same context. Therefore, a natural solution is to maximize the mutual information of contextual information hidden in contextualized representations in the same context. To extract shared contextual information, we first define a rule to generate contextual representations of tokens by combining embeddings and global information. Formally,

where ff is a probing algorithm and ei\bm{e}_{i} is the embedding and g\bm{g} is the global bias of a concrete context. In this paper, we adopt a straightforward probing method to get global information hidden in contextualized representations, where

Given contextualized representations of an token x\bm{x} and its nearby tokens c\bm{c} in the same context, we use gx\bm{g}_{x} and gc\bm{g}_{c} to represent global semantics hidden in these representations. The mutual information between the two global bias gx\bm{g}_{x} and gc\bm{g}_{c} is

According to van den Oord et al. 2019, the InfoNCE loss serves as an estimator of mutual information of x\bm{x} and c\bm{c}:

where L(gx,gc)\mathcal{L}(\bm{g}_{x},\bm{g}_{c}) is defined as:

where ck−\bm{c}_{k}^{-} is the kk-th negative sample of x\bm{x} and KK is the size of negative samples. Hence minimizing the objective L(gx,gc)\mathcal{L}(\bm{g}_{x},\bm{g}_{c}) is equivalent to maximizing the lower bound on the mutual information I(gx,gc)I(\bm{g}_{x},\bm{g}_{c}). This objective contains two parts: positive pairs f(gx,gc)f(\bm{g}_{x},\bm{g}_{c}) and negative pairs f(gx,gck−)f(\bm{g}_{x},\bm{g}_{c_{k}^{-}}).

Previous study (Chen et al., 2020) has shown that cosine similarity with temperature performs well as the score function ff in InfoNCE loss. Following them, we take

Positive pairs: Given each token x\bm{x}, we randomly sample a positive sample c\bm{c} from nearby tokens in the same context (sequence) within a window span where WW is the window size.

Negative pairs: Given each token x\bm{x}, we randomly sample KK tokens from other sequences in this batch as negative samples ck−\bm{c}_{k}^{-}.

To sum up, the Token-alignment Contrastive (TC) loss is applied to every token in a batch as:

where NN is the number of sequences of this batch; si\bm{s}_{i} is the ii-th sequence; jj and jcjc are tokens in si\bm{s}_{i} where jc≠jjc\neq j; gi\bm{g}^{i} is the global semantics hidden in contextualized representation of token si\bm{s}_{i}. gji\bm{g}_{j}^{i} and gjci\bm{g}_{j_{c}}^{i} are generated via:

where hji\bm{h}_{j}^{i} and eji\bm{e}_{j}^{i} are the contextualized representation and static embedding of the anchor token, respectively. hjci\bm{h}_{j_{c}}^{i} and ejci\bm{e}_{j_{c}}^{i} are the contextualized representation and static embedding of the sampled positive token in the same context.

2 Training Objective

As described before, the token-alignment contrastive loss LTC\mathcal{L}_{\text{TC}} is designed to model global dependencies while MLM is able to capture local dependencies. Therefore, we can better model contextualized representations by combining the token-alignment contrastive loss LTC\mathcal{L}_{\text{TC}} and the MLM loss to get our overall objective LTACO\mathcal{L}_{\text{TACO}}:

We implement it in a multi-task learning manner where all objectives are calculated within one forward propagation, which only introduces negligible extra computations.

Experiments

Following BERT (Devlin et al., 2019), we select the BooksCorpus (800M words after WordPiece tokenization) (Zhu et al., 2015) and English Wikipedia (4B words) as pre-training corpus. We pre-train two variants of BERT models: BERT-small and BERT-base. All models are equipped with the vocabulary of size 30,522, trained with 15% masked positions for MLM. The maximum sequence length is 256 and batch size is 1,280. We adopt optimizer AdamW (Loshchilov and Hutter, 2019) with learning rate 1e-4. All models are trained until convergence. To be specific, the small model is trained up to 250k steps with a warm-up of 2.5k steps. The base model is trained up to 500k steps with a warm-up of 10k steps. For TACO, we set the positive sample window size WW to 5, the negative sample number KK to 50, and the temperature parameter τ\tau to 0.07 after a slight grid-search via preliminary experiments. More pre-training details can be found in Appendix A.

During fine-tuning models, we conduct a grid search over batch sizes of {16, 32, 64, 128}, learning rates of {1e-5, 2e-5, 3e-5, 5e-5}, and training epochs of {4, 6} with an Adam optimizer (Kingma and Ba, 2015). We use the open-source packages for implementation, including HuggingFace Datasetshttps://github.com/huggingface/datasets and Transformershttps://github.com/huggingface/transformers. All the experiments are conducted on 16 GPU chips (32 GB V100).

Evaluation

We evaluate methods on the GLUE benchmark Wang et al. (2019). Specifically, we test on Microsoft Research Paraphrase Matching (MRPC) (Dolan and Brockett, 2005), Quora Question Pairs (QQP)https://www.quora.com/q/quoradata/First-Quora-Dataset-Release-Question-Pairs and STS-B (Conneau and Kiela, 2018) for Paraphrase Similarity Matching; Stanford Sentiment Treebank (SST-2) (Socher et al., 2013) for Sentiment Classification; Multi-Genre Natural Language Inference Matched (MNLI-m), Multi-Genre Natural Language Inference Mismatched (MNLI-mm) (Williams et al., 2018), Question Natural Language Inference (QNLI) (Rajpurkar et al., 2016) and Recognizing Textual Entailment (RTE) (Wang et al., 2019) for the Natural Language Inference (NLI) task; The Corpus of Linguistic Acceptability (CoLA) (Warstadt et al., 2019) for Linguistic Acceptability.

Following Devlin et al. (2019), we exclude WNLI (Levesque, 2011). We report F1 scores for QQP and MRPC, Spearman correlations for STS-B, and accuracy scores for the other tasks. For evaluation results on validation sets, we report the average score of 4 fine-tunings with different random seeds. For results on test sets, we select the best model on the validation set to evaluate.

Baselines

We mainly compare TACO with MLM on BERT-small and BERT-base models. In addition, we also compare TACO with related contrastive methods: a sentence-level contrastive method BERT-NCE and a span-based contrastive learning method INFOWORD, both from Kong et al. (2020). We directly compare TACO with the results reported in their paper.

2 Results on BERT-Small

Table 1 and Figure 5 show the results of TACO on BERT-small. As we can see, compared with MLM with 250k training steps ( convergence steps), TACO achieves comparable performance with only 1/5 computation budget. By modeling global dependencies, TACO can significantly improve the efficiency of contextualized representation learning. In addition, when pre-trained with the same steps, TACO outperforms MLM with 1.2 average score improvement on the validation set.

In addition to convergence, we also compare TACO and MLM on fewer training data. The results are shown in Table 2. We sample 4 tasks with the largest amount of training data for evaluation. As we can see, TACO trained on 25% data can achieve competitive results with MLM trained on full data. These results also verify the data efficiency of our method, TACO.

3 Results on BERT-Base

We also compare TACO with MLM on base-sized models, which are the most commonly used models according to the download data from Huggingfacehttps://huggingface.co/models (Wolf et al., 2020). First, from Table 3, we can see that TACO consistently outperforms MLM under all pre-training computation budgets. Notably, TACO-250kk achieves comparable performance with MLM-500kk, which saves 2x computations. Similar results are observed on TACO-100kk and BERT-250kk. These results demonstrate that TACO can achieve better acceleration over MLM. It is also a significant improvement compared to previous methods (Gong et al., 2019) focusing on accelerating BERT but only with slight speedups. In addition, as shown in Table 4, TACO achieves competitive results compared to BERT-NCE and INFOWORD, two similar contrastive methods.

Discussion

To better understand how TACO works, we conduct a quantitative comparison on the learning dynamic for BERT and TACO. Similar to Section 2.2, we plot the Cosine similarity among contextualized representations of tokens in the same context (intra-context) and different contexts (inter-context) in Figure 6. We find that the learning dynamic of TACO significantly differs from that of MLM. Specifically, for TACO, the intra-context representation similarity remains high and the gap between intra-context similarity and inter-context similarity remains large at the later stage of training. This confirms that TACO can better fulfill global semantics, which may contribute to the superior downstream performance.

2 Ablation Study

TACO is implemented as a token-level contrastive (TC) loss along with the MLM loss. Therefore, the improvement of TACO might come from two aspects, including 1) denser supervision signals from the all-token objective and 2) the benefits of the contrastive loss to strengthen global dependencies. It is helpful to figure out which factor is more important. To this end, we design two variants for ablation. One is a concentrated TACO, where the contrastive loss is built on the 15% masked positions only, keeping the same density of supervision signal with MLM. The other is an extended MLM, where not only 15% masked positions are asked to predict the original token, so do the rest 85% unmasked positions. The extended MLM has the same dense supervision with TACO but loses the benefits of modeling the global dependencies. The results on small models are shown in Figure 6.

As we can see, the performance of TACO decreases if we sample a part of token positions to implement TC objectives. It shows that more supervision signals benefit the final performance of TACO. However, simply adding more supervision signals by predicting unmasked tokens does not help MLM too much. Even equipped with the extra 85% token prediction (TP) loss, MLM+TP does not show significant improvements and it is noticeable that the performance of MLM+TP starts to drop after 150k steps. This further confirms the effectiveness of TC loss by strengthening global dependencies.

Related Work

Classic language representation learning methods (Mikolov et al., 2013; Pennington et al., 2014) aims to learn context-independent representation of words, i.e., word embeddings. They generally follow the distributional hypothesis (Harris, 1954). Recently, the pre-training then fine-tuning paradigm has become a common practice in NLP because of the success of pre-trained language models like BERT (Devlin et al., 2019). Context-dependent (or contextualized) representations are the basic characteristic of these methods. Many existing contextualized models are based on the masked language modeling objective, which randomly masks a portion of tokens in a text sequence and trains the model to recover the masked tokens. Many previous studies prove that pre-training with the MLM objective helps the models learn syntactical and semantic knowledge (Clark et al., 2019). There have been numerous extensions to MLM. For example, XLNet (Yang et al., 2019) introduced the permutated language modeling objective, which predicts the words one by one in a permutated order. BART Lewis et al. (2020) and T5 Raffel et al. (2020) investigated several denoising objectives and pre-trained an encoder-decoder architecture with the mask span infilling objective. In this work, we focus on the key MLM objective and aim to explore how MLM objective helps learn contextualized representation.

2 Contrastive-based SSL

Apart from denoising-based objectives, contrastive learning is another promising way to obtain self-supervision. In contrastive-based self-supervised learning, the models are asked to distinguish the positive samples from the negative ones for a given anchor. Contrastive-based SSL method was first introduced in NLP for efficient learning of word representations by negative sampling, i.e., SGNS (Word2Vec (Mikolov et al., 2013)). Later, similar ideas were brought into CV field for learning image representation and got prevalent, such as MoCo He et al. (2020), SimCLR Chen et al. (2020), BYOL Caron et al. (2020), etc.

In the recent two years, there have been many studies targeting at reviving contrastive learning for contextual representation learning in NLP. For instance, CERT (Fang et al., 2020) utilized back-translation to generate positive pairs. CAPT (Luo et al., 2020) applied masks to the original sentence and considered the masked sentence and its original version as the positive pair. DeCLUTR (Giorgi et al., 2020) samples nearby even overlapping spans as positive pairs. INFOWORD (Kong et al., 2020) treated two complementary parts of a sentence as the positive pair. However, the aforementioned methods mainly focus on sentence-level or span-level contrast and may not provide dense self-supervision to improve efficiency. Unlike these approaches, TACO regards the global semantics hidden in contextualized token representations as the positive pair. The token-level contrastive loss can be built on all input tokens, which provides a dense self-supervised signal.

Another related work is ELECTRA (Clark et al., 2020). ELECTRA samples machine-generated tokens from a separate generator and trains the main model to discriminate between machine-generated tokens and original tokens. ELECTRA implicitly treats the fake tokens as negative samples of the context, and the unchanged tokens as positive samples. Unlike this method, TACO does not require architectural modifications and can serve as a plug-and-play auxiliary objective, largely improving pre-training efficiency.

Conclusion

In this paper, we propose a simple yet effective objective to learn contextualized representation. Taking MLM as an example, we investigate whether and how current language model pre-training objectives learn contextualized representation. We find that the MLM objective mainly focuses on local anchors to align contextualized representations, which harms global dependencies modeling due to an “embedding bias” problem. Motivated by these problems, we propose TACO to directly model global semantics. It can be easily combined with existing LM objectives. By combining local and global anchors, TACO achieves up to 5×\times speedups and up to 1.2 improvements on GLUE score. This demonstrates the potential of TACO to serve as a plug-and-play approach to improve contextualized representation learning.

Acknowledgement

We thank the anonymous reviewers for their helpful feedback. We also thank the colleagues from ByteDance AI Lab for their suggestions on our experiment designing and paper writing.

References

Appendix A Experiment Details

All pre-training approaches involved in experiments use the same pre-training hyper-parameters but do not include BERT-NCE and INFOWORD. Results of BERT-NCE and INFOWORD are directly cited from the original paper (Kong et al., 2020). Following Liu et al. (2019), we do not use the next sentence prediction (NSP) objective and use dynamic masking for MLM with a 15% mask ratio, where the masked positions are decided on the fly.

TACO introduces three extra hyper-parameters, including negative sample size KK, positive sample window size WW and temperature τ\tau. We set the temperature τ\tau as a small value, 0.07, following Fang et al. (2020). By searching for the best KK out of {10, 50} and WW out of {3, 5, 10, 50} on the small TACO model, we found that TACO with KK=50 and WW=5 performs best, so we also apply these hyper-parameter choices for base-sized TACO. The full set of pre-training hyper-parameters are listed in Table 5. Actually, TACO outperforms MLM under most cases in our preliminary experiments. However, we still also find some extreme cases which might harm the effectiveness of TACO. If the size of negative samples KK is too small, e.g., smaller than 10, the performance of TACO degenerates nearly to the level of BERT baseline. Similar conclusions are also mentioned in related works He et al. (2020); Chen et al. (2020). Also, if the positive window size WW is too large, e.g., bigger than 50, the performance of TACO degrades, too. We suspect the over-large positive window brings more false-positive samples, which makes the sequence meaning ambiguous, thus harms the performance.

A.2 Fine-tuning Details

For small-sized models, we fine-tune all saved checkpoints (5k, 10k, 20k, 30k, 40k, 50k, 100k, 150k, 200k, 250k-step) of different pre-trained models (TACO and its ablations) with the same hyper-parameters on each task. Considering the large amount of pre-training checkpoints, we just adopt the default fine-tuning hyper-parameters and repeat fine-tuning 4 times with different random seeds. Then the best performed fine-tuned models on validation sets are used for testing. This setting helps make a fair comparison among models and avoids a large amount of grid-search runs. The task-specific hyper-parameters for small-sized models are listed in Table 7. The general fine-tuning hyper-parameters are listed in Table 6.

For base-sized models, we save checkpoints at 100k, 250k, and 500k steps, respectively. During fine-tuning, we also conduct multiple fine-tuning runs with different task-specific hyper-parameter combinations as shown in Table 8. Concretely, we randomly sample 6 different hyper-parameter combinations and report the average score for validation results. Then we select the best-performing run of 500k-step checkpoints (converged) for testing.

A.3 Statistic Details

We calculate cosine similarity of 20 randomly sampled pairs of frequently co-occurrent words from the WordSim353 dataset (Agirre et al., 2009) labeled by human annotators to plot the average similarity curve in Figure 3(b). Corresponding embeddings are obtained from the embedding layer of the BERT model and variant models mentioned in Section 2.2.

Intra-/Inter-context Similarity

For every token wiw_{i} in the corpus, we randomly sample a positive token wj≠iw_{j\neq i} within the same context (sentence) and another token wkw_{k} from other sentences.

As mentioned in Section 2.2, we take BERT Devlin et al. (2019) as our encoder to get contextualized representations through the last hidden states h\bm{h}. We mainly adopt the cosine similarity as the measurement and calculate the average intra-context similarity (between hi\bm{h}_{i} and hj\bm{h}_{j}) and the average inter-context similarity (between hi\bm{h}_{i} and hk\bm{h}_{k}) over all tokens in the corpus. It is worth noticing that we do use any masks here when generating a token’s contextualized representation for statistics.

Other Measurements

We observe the same findings for MLM under other measurements, though the statistics before are mainly based on cosine similarities. We tried other similarities or distances, e.g., L1 distance, L2 distance and L10 distance, to evaluate the discrepancy between contextualized representations from the same context and different contexts. Specifically, we make intra-context and inter-context statistics under specific measurement at different pre-training checkpoints, then calculate the ratio of intra-context measurement over the inter-context one. Table 9 shows the statistical results. As we can see, when the ratio of L1 distance decreases, the ratio of cosine similarity and the dot-production similarity increase, vice versa.

Appendix B Extra Experiments

In the standard implementation of BERT, the parameters of input embeddings are shared with output embeddings. All experiments and analyses in this paper are based on this assumption. To further confirm the effectiveness of TACO, we conduct the extra experiments without embedding sharing on BERT-small. The results are showed in Table 10. It is unexpected that the variants without embedding sharing perform worse compared their counterparts due to lack of regularization of weight sharing. From the results, we can see that the TACO without embedding sharing performs slightly worse than TACO with embedding sharing. However, compared to the MLM, it is still better than MLM than 0.9 average GLUE score when convergence. These results prove the effectiveness of TACO even when embeddings are not sharing.