Neural Mask Generator: Learning to Generate Adaptive Word Maskings for Language Model Adaptation
Minki Kang, Moonsu Han, Sung Ju Hwang
Introduction
The recent success of the language model pre-training approaches (Devlin et al., 2019; Peters et al., 2018; Radford et al., 2019; Raffel et al., 2019; Yang et al., 2019), which train language models on diverse text corpora with self-supervised or multi-task learning, have brought up huge performance improvements on several natural language understanding (NLU) tasks (Wang et al., 2019; Rajpurkar et al., 2016). The key to this success is their ability to learn generalizable text embeddings that achieve near optimal performance on diverse tasks with only a few additional steps of fine-tuning on each downstream task.
Most of the existing works on language model aim to obtain a universal language model that can address nearly the entire set of available natural language tasks on heterogeneous domains. Although this train-once and use-anywhere approach has been shown to be helpful for various natural language tasks (Devlin et al., 2019; Radford et al., 2019; Dong et al., 2019; Raffel et al., 2019), there have been considerable needs on adapting the learned language models to domain-specific corpora (e.g. healthcare or legal). Such domains may contain new entities that are not included in the common text corpora, and may contain only a small amount of labeled data as obtaining annotation on them may require expert knowledge. Some recent works (Sun et al., 2019a; Lee et al., 2019; Beltagy et al., 2019; Gururangan et al., 2020) suggest to further pre-train the language model with self-supervised tasks on the domain-specific text corpus for adaptation, and show that it yields improved performance on tasks from the target domain.
Masked Language Models (MLMs) objective in BERT (Devlin et al., 2019) has shown to be effective for the language model to learn the knowledge of the language in a bi-directional manner (Vaswani et al., 2017). In general, masks in MLMs are sampled at random (Devlin et al., 2019; Liu et al., 2019c), which seems reasonable for learning a generic language model pre-trained from scratch, since it needs to learn about as many words in the vocabulary as possible in diverse contexts.
However, in the case of further pre-training of the already pre-trained language model, such a conventional selection method may lead a domain adaptation in an inefficient way, since not all words will be equally important for the target task. Repeatedly learning for uninformative instances thus will be wasteful. Instead, as done with instance selection Ngiam et al. (2018); Jiang et al. (2018); Yoon et al. (2019); Zhu et al. (2019), it will be more effective if the masks focus on the most important words for the target domain, and for the specific NLU task at hands. How can we then obtain such a masking strategy to train the MLMs?
Several works (Joshi et al., 2019; Sun et al., 2019b, c; Glass et al., 2019) propose rule-based masking strategies which work better than random masking (Devlin et al., 2019) when applied to language model pre-training from scratch. Based on those works, we assume that adaptation of the pre-trained language model can be improved via a learned masking policy which selects the words to mask. Yet, existing models are inevitably suboptimal since they do not consider the target domain and the task. To overcome this limitation, in this work, we propose to adaptively generate mask by learning the optimal masking policy for the given task, for the task-adaptive pre-training (Gururangan et al., 2020) of the language model.
As described in Figure 1, we want to further pre-train the language model on a specific task with a task-dependent masking policy, such that it directs the solution to the set of parameters that can better adapt to the target domain, while task-agnostic random policy leads the model to an arbitrary solution. To tackle this problem, we pose the given learning problem as a meta-learning problem where we learn the task-adaptive mask-generating policy, such that the model learned with the masking strategy obtains high accuracy on the target task. We refer to this meta-learner as the Neural Mask Generator (NMG). Specifically, we formulate mask learning as a bi-level problem where we pre-train and fine-tune a target language model in the inner loop, and learn the NMG at the outer loop, and solve it using renforcement learning. We validate our method on diverse NLU tasks, including question answering and text classification. The results show that the models trained using our NMG outperforms the models pre-trained using rule-based masking strategies, as well as finds a proper adaptive masking strategy for each domain and task.
We propose to learn the mask generating policy for further pre-training of masked language models, to obtain optimal maskings that focus on the most important words for the given text domain and the NLU task.
We formulate the problem of learning the task-adaptive mask generating policy as a bi-level meta-learning framework which learns the LM in the inner loop, and the mask generator at the outer loop using reinforcement learning.
We validate our mask generator on diverse tasks across various domains, and show that it outperforms heuristic masking strategies by learning an optimal task-adaptive masking for each LM and domain. We also perform empirical studies on various heuristic masking strategies on the language model adaptation.
Related Work
Ever since Howard and Ruder (2018) suggested language model pre-training with multi-task learning, inspired by the success of fine-tuning on ImageNet pre-trained models on computer vision tasks (Liu et al., 2019b), research on the representation learning for natural language understanding tasks have focused on obtaining a global language model that can generalize to any NLU tasks. A popular approach is to use self-supervised pre-training tasks for learning the contextualized embedding from large unannotated text corpora using auto-regressive (Peters et al., 2018; Yang et al., 2019) or auto-encoding (Devlin et al., 2019; Liu et al., 2019c) language modeling. Following the success of the Masked Language Model (MLM) from (Devlin et al., 2019; Liu et al., 2019c), several works have proposed different model architecture (Lan et al., 2019; Raffel et al., 2019; Clark et al., 2020) and pre-training objectives (Sun et al., 2019b, c; Liu et al., 2019c; Joshi et al., 2019; Glass et al., 2019; Dong et al., 2019), to improve upon its performance. Some works have also proposed alternative masking policies for the MLM pre-training over random sampling, such as SpanBERT Joshi et al. (2019) and ERNIE Sun et al. (2019b). Yet, none of the existing approaches have tried to learn the task-adaptive mask in a context-dependent manner which is the problem we target in this work.
Language Model Adaptation
Pre-training the language model on the target domain, then fine-tuning on downstream tasks, is the most simple yet successful approach for adapting the language model to a specific task. Some studies (Lee et al., 2019; Beltagy et al., 2019; Whang et al., 2019) have shown the advantage of further pre-training the language model on a large unlabeled text corpus collected from a specific domain. Moreover, Sun et al. (2019a) and Han and Eisenstein (2019) investigate the effectiveness of further pre-training of the language model on small domain-related text corpora. Recently, Gururangan et al. (2020) integrates prior works and defines domain-adaptive pre-training and task-adaptive pre-training, showing that domain adaptation of the language model can be done with additional pre-training with the MLM objective on a domain-related text corpus, as well as a smaller but directly task-relevant text corpus.
Meta-Learning
Meta-learning Thrun and Pratt (1998) aims to train the model to generalize over a distribution of tasks, such that it can generalize to an unseen task. There exist large number of different approaches to train the meta-learner (Santoro et al., 2016; Vinyals et al., 2016; Ravi and Larochelle, 2017; Finn et al., 2017; Liu et al., 2019a). However, existing meta-learning approaches do not scale well to the training of large models such as masked language models. Thus, instead of the existing meta-learning method such a gradient based approach, we formulate the problem as a bi-level problem Franceschi et al. (2018) of learning the language model in the inner loop and the mask at the outer loop, and solve it using reinforcement learning. Such optimization of the outer objective using RL is similar to the formulation used in previous works on neural architecture search Zoph and Le (2017); Zoph et al. (2018).
Problem Statement
We now describe how to formulate the problem of learning to generate masks for the Masked Language Model (MLM) as a bi-level optimization problem. Then, we describe how we can reformulate it as a reinforcement learning problem.
For pre-training of the language models, we need an unannotated text corpus . Here the is the context, whose element is a single word or token. To formulate the meta-learning objective, we assume each context corpus as the part of the task consisting of the context and its corresponding task dataset. For the MLM, we need to generate a noisy version of , which we denote as . Let be the indicator of masking -th word of . If , -th word is replaced with the [MASK] token. The objective of the MLM then is to predict the original token of each [MASK] token. Therefore, we can formulate this problem as follows:
where is the parameter of the language model and is the original tokens of each [MASK] token, and is the set of masked contexts. Following the formulation from (Yang et al., 2019), we can approximate the MLM objective with as follows:
where indicates the contextualized representation of from the language model layer (e.g. Pre-trained Transformer layers in (Devlin et al., 2019)), and denotes the embedding of word from the last prediction layer. Our objective then is to learn the optimal policy for determining each mask indicator , which we will describe in detail in the next subsection.
2 Bi-level formulation
We now describe how to formulate this learning problem as a bi-level problem consisting of inner and outer level objectives. Consider that can be represented using an arbitrary function parameterized with :
where is the probability of masking i-th word from , and indicates the list of word indices to be masked and is the mask indicator parameterized by the parameter . The details of corresponding equations will be described in section 3.3. Therefore, the MLM objective has been slightly changed from its original form, into the following objective:
where the masked context is parameterized by the parameter . Now assume that we have found the optimal parameter for language model from equation 5. Then, we need to fine-tune the language model on the downstream task. Although linear heads used in both pre-training and fine-tuning are different, we describe the parameter of both models as for simplicity. Following the bi-level framework notation described in (Franceschi et al., 2018), the inner objective function for fine-tuning can be written as follows:
where is a training dataset, is the loss function of the supervised learning, and is the function representation of downstream task solver model. In case of question answering task, each consists of a context and corresponding question and is the corresponding answer spans.
Assume that we find optimal parameter from supervised task fine-tuning. Then, the final outer-level objective can be described as follows:
This will allow us to obtain the optimal parameter which minimizes the task objective function on a test dataset .
Although the outer objective is differentiable, we formulate the optimization problem of the outer objective as a reinforcement learning problem to avoid excessive computation cost caused by the two constraint terms.
In this paragraph, we explain why we use the Reinforcement Learning (RL) instead of the differentiable method to train the parameter . As indicated in the equation 7, our inner loop includes consecutive two steps of language model training. The NMG model is addressed to the pre-training step rather than the task fine-tuning step. Therefore, the direct differentiation of the outer objective contains two second-derivative terms for both the MLM loss and the task train loss . With a single-step approximation, the derivative of the outer objective is approximated as follows:
Instead, we can address the first-order approximation () to the derivative of the outer objective to avoid second-order derivative computation as follows:
where is approximated to . Such approximation trivially results in a meaningless optimization since it ignores the pre-training step induced by the parameter , which decides the masking policy (Liu et al., 2019a). Therefore, we approach solving this optimization problem with RL instead of the differential method to avoid such an issue. In the next section, we introduce how we formulate this problem as the RL with the outer objective as a non-differentiable reward.
3 Reinforcement learning formulation
We now propose a reinforcement learning (RL) framework, which given the context as the state, decides on the actions where is the number of masked tokens in the given context, each is the token index that indicates a decision on masking the token following equation 4, and . The objective of the RL agent then is to find an optimal masking policy that minimizes from the section 3.2. In addition, minimizing on can be seen as maximizing its performance. Therefore, the objective is the maximization of the performance of the model on . We can induce it by setting the reward as the accuracy improvement on . We will describe the detail of the reward design in section 4.
In general RL formulation (Sutton and Barto, 2018) following Markov Decision Process (MDP), state transition probability can be described as . The probability of masking tokens is formularized as , where consists of number of [MASK] tokens. Although the representation of and are slightly different because of the addition of [MASK] token, we can approximate them as following the approximation in equation 1, which inductively approximates representations after each word masking as a representation of original context.
Therefore, we can approximate the probability of masking tokens from the MDP problem to the problem without state-transition formulation as follows:
where denotes the index of masked word included in the context and is the masking policy of the agent. By this approximation, we do not need to consider the trajectory along the temporal horizon, which makes the problem much easier. Instead, we approximate the problem as the task of selecting multiple discrete actions simultaneously for the same state. We approximate the policy with neural network parameterized by as . As in equations 3 and 4, the mask is determined by actions generated from the neural policy . In the next section, we will describe how to train the Neural Mask Generator to generate the neural policy maximizing the reward using RL in Section 4.
Neural Mask Generator
In this section, we describe our model, Neural Mask Generator (NMG), which learns the masking policy to generate an optimal mask on an unseen context using deep RL with the detailed descriptions of the framework setup. The overview of the meta mask generator framework is shown in Figure 2. For detailed descriptions of the approaches, the procedures of both training and test phase, and algorithm, please see Appendix A.
The probability of selecting -th word of as -th action can be described as , where . Instead of in equation 2, the neural policy from the NMG is given as follows:
where is a deep neural network parameterized by , which has a self-attention layer (Vaswani et al., 2017) followed by linear layers with gelu activation Hendrycks and Gimpel (2016), and is the contextualized representation of of context from the frozen language model layer . Note that is shared with the target language model and not trained during the NMG training. Further, outputs the scalar logit of the , and the final probabilistic policy is computed by the softmax function.
Training Objective
We train the NMG model using the Advantage Actor-Critic method (Mnih et al., 2016) with the value estimator. Furthermore, we resort to off-policy learning (Degris et al., 2012) such that the agent can explore the optimal policy based on its past experiences. To this end, we leverage a prioritized experience replay, in which we store every state, action, reward, and old-policy pairs Mnih et al. (2013), and sample them based on their absolute value of the advantage Schaul et al. (2016). We use importance sampling to estimate the value of the target policy with the samples from the behavior policy , which are sampled from the replay buffer (Degris et al., 2012; Schulman et al., 2015, 2017; Wang et al., 2017). To sum up, the objectives for both policy and value network are as follows:
where the is set of sampled replays from the replay buffer, is an entropy function , and is a hyperparameter for entropy regularization. is an estimated value of the state . The value network consists of linear layers with activation function after mean pooling .
To summarize, the outer-level objective for updating the NMG parameter can be written as follows:
At each episode of meta-training, we update the NMG parameter by optimizing the above objective function.
Reward Design and Self-Play
As in Section 3.3, the reward function is considered as the accuracy improvement on the test set . Therefore, the pre-training step (equation 5) and the fine-tuning then evaluation step (equation 6, 7) should be done to get the reward in every episodes.
Since using the full size of dataset in the inner loop is generally not feasible, we randomly sample smaller sub-task from at every episode. For the evaluation, we randomly split the training set to generate a hold-out validation set and replace in equation 7 to while meta-training where is unobservable. We use a sufficiently large hold-out validation set to prevent the masking policy from overfitting to . We assume that the meta-learner NMG also performs well on if it is trained on diverse sub-tasks where .
The problem to be considered for using diverse sub-tasks is that the NMG model encounters different sub-task at every new episode. Since determines the state distribution and the data to be trained, it results in the reward scale problem that the expectation of the validation accuracy on varies depending on the composition of then makes it harder to evaluate the performance increment of the neural policy across episodes.
To address this problem, we introduce the random policy as an opponent policy to evaluate the neural policy relative to it. Therefore, the reward R is defined as , where is the sign function, and are the accuracy score on from neural and random policy respectively. However, the random policy may be too weak as the opponent. To overcome this limitation, we add another neural policy as an additional opponent to induce the zero-sum game of two learning agent by the concept of the self-play algorithm (Wei et al., 2017; Grau-Moya et al., 2018).
Then, three distinct policies are compared with each other during episodes and two neural policies are individually trained. Furthermore, in each neural agent training, only actions corresponding to disjoint comparing with others are stored in a global replay buffer for a more accurate reward assignment of each action.
Continual Adaptive Learning
For fair comparison at every episode, the same language model should be used to evaluate the policy. Initializing the language model before the start of each inner loop can be the simplest choice to handle this. However, since we pre-train the language model for only few steps during each episode of meta-training, the model is always evaluated around the original language model domain (see Figure 1). To avoid this, at each episode, the language model which is pre-trained by the NMG model of former step is continually loaded instead of the fixed checkpoint except for the first episode. By this, we intend that our agent learns the optimal policy of various environments. Furthermore, the agent can learn dynamic policy based on the learned degree of the target language model.
Experiment
We now experimentally validate our Neural Mask Generator (NMG) model on multiple NLU tasks, including question answering and text classification tasks using two different language models, and analyze its behaviors. In Section 5.1, we evaluate the NMG with several baselines. Then, we evaluate the effect of the specific design choices made for our model through ablation studies in Section 5.2. Finally, we analyze how the policy learned by the NMG model works in Section 5.3.
For question answering, we use three datasets, namely SQuAD v1.1 Rajpurkar et al. (2016), NewsQA Trischler et al. (2017), and emrQA Pampari et al. (2018) to validate our model. We use the MRQAhttps://mrqa.github.io/ version for both SQuAD and NewsQA to sustain a coherency. We also preprocess emrQA to fit the format of other datasets. We use a standard evaluation metric named Exact-Matchs (EM) and F1 score for question answering task. For text classification, we use IMDb Maas et al. (2011) and ChemProt Kringelum et al. (2016), following the experimental settings of (Gururangan et al., 2020).
Baselines
According to Devlin et al. (2019) and Joshi et al. (2019), training the language models with difficult objectives is much more beneficial when pre-training from scratch. To test whether it is also the case for task-adaptive pre-training, we experiment with two heuristic masking strategies, which we refer to as whole-random and span-random. In addition, we also tested the named entity masking proposed by Sun et al. (2019b) (entity-random). Below is the complete list of the heuristic baselines we compare against in the experiments.
1) No-PT A baseline without any further pre-training of the language model.
2) Random A random masking strategy introduced in BERT Devlin et al. (2019).
3) Whole-Random A random masking strategy which masks the entire word instead of the token (sub-word). This method is introduced by the authors of BERT Devlin et al. (2019)https://github.com/google-research/bert.
4) Span-Random A random masking strategy which selects multiple consecutive tokens.
5) Entity-Random A random masking strategy which selects named entities with highest priorities, then randomly selects other tokens.
6) Punctuation-Random A random masking strategy which selects punctuation tokens first, then randomly selects other tokens.
Implementation Details
For the language model , we use the same hyperparameters and architecture with DistilBERT Sanh et al. (2019) model (66M params) and BERTBASE Devlin et al. (2019) model (110M params). Our implementation is based on the huggingface’s Pytorch implementation version (Wolf et al., 2019; Paszke et al., 2019). We load the pre-trained parameters from the checkpoint of each language model in meta-testing and the first episode of meta-training. As for the text corpus to pre-train the language model, we use the collection of contexts from the given NLU task. We only use the Masked Language Model (MLM) objective for further pre-training. In the initial stage of meta-training, the NMG randomly selects actions for exploration. In meta-testing, it takes maximum probability indices as actions. We describe the details of language model training in Appendix B since we use different settings for each task and experiments. As for reinforcement learning (RL), we use the off-policy actor-critic method described in Section 4. For more details of the reinforcement learning framework, please see Appendix B.
1 Results
First of all, we need to discuss how the MLM works on the language model pre-training. In practice, the Cloze task Taylor (1953) benefits when words that need to learn are masked. However, for the MLM, we observed that the masking prevents learning the masked words in the language model. Rather the MLM learns representations and relations between non-masked words by predicting the masked words using them. In the case of adaptive pre-training, learning the domain-specific vocabulary is crucial for the domain adaptation. Therefore, masking out trivial words (e.g. Punctuations) may be more beneficial than masking out unique words (e.g. Named Entities) for the language model to learn the knowledge of a new domain. To see how it works, we experiment for punctuation-random, which masks only the punctuations, which are clearly useless for the domain adaptation.
In Table 1 and 2, we report the performance of baselines and our model on both question answering and text classification tasks. From the baseline results of Table 1, we speculate on the important aspects of a masking strategy for better adaptation. The whole-word, span and entity maskings often lead to better results since it makes the MLM objective more difficult and meaningful (Joshi et al., 2019; Sun et al., 2019b). For instance, in emrQA, most words are tokenized to sub-words since contexts include a lot of unique words such as medical terminologies. Therefore, whole-word masking could be most suitable for domain adaptation on emrQA. In contrast, such maskings sometimes lead to worse results than random masking on some domains such as in NewsQA. Especially, in the case of DistilBERT, punctuation masking performs better than others. These result suggests that masking complicated words is rather disturbing for the adaptation of small language model.
On the other hand, our NMG learns the optimal masking policy for a given task and domain in an adaptive manner on any language model. In Table 1 and 2, we can see that this adaptive characteristic of our model makes the neural masking results in better or at least comparable performances to the baselines for all tasks. We further analyze the learned masking strategy in Section 5.3.
2 Ablation study
We further investigate the effectiveness of self-play by comparing it with the NMG model without self-play, where the model only competes with the random agent. We validate this on the NewsQA dataset. The result in Table 3 shows that the NMG model with self-play obtains better performance than its counterpart without self-play. This result verifies that competing with the opponent neural agent while learning helps the NMG model to learn better policy.
Continual Adaptation
We also perform an ablation study of the continual adaptation learning. The result in Table 3 shows that the continual-adaptive masking strategy is significantly effective for the language model adaptation. The result suggests that helpful words for the language model to learn depends on the adaptation degree of it.
3 Analysis
To analyze how our model performs, we measure the difference between which kind of word token is masked by both the random and neural policy on the pre-trained checkpoint. For qualitative analysis, we provide examples of masked tokens on the context in Figure 3. As shown in Figure 3, NMG tends to mask highly informative words such as seminary or islam, which are parts of the answer spans. Furthermore, we analyze the masking behavior of our NMG by performing Part-of-Speech (POS) tagging on the masked words using spaCyhttps://spacy.io. Figure 4 shows the six most frequent tags for the words masked out by the random and neural policy. Figure 4 shows that the neural policy masks more words in noun, verb, and proper noun tags than the random policy, suggesting that our NMG model learns that masking such informative words is beneficial to adapt on the NewsQA task with the BERTBASE model as a language model.
Learning Curves
As already known, the RL-based methods often suffer from the instability problem. Therefore, we further analyze the learning curves of the NMG training in this section. In Figure 5, we plot three kinds of learning curves to show the detailed training process. Cumulative Regret indicates how many times the neural agent is defeated against the random agent until certain episodes. The grey plot indicates the worst case that the random agent always defeats against the neural agent. Entropy indicates the average entropy of policy for states given in a certain episode. Lower entropy means that the policy has a high probability of a few significant actions. Loss indicates the RL loss described in the equation 10. We ignore outliers in the loss plot for brevity.
From the entropy and loss plots, we can notice that the policy converges as learning proceeds. However, it seems that such convergence is not continually sustained. From the cumulative regret plot, we can observe that the neural policy still often loses against the random policy, although it is trained for a while. Such instability may come from the difficulty of the exact credit assignment on each action. Otherwise, continuous change of state distribution from the continual adaptive learning may hinder the neural policy’s convergence.
Even if the NMG shows the notable results, there is room for improvement on RL in terms of efficiency and stability. We leave it as the future work.
Conclusion
We proposed a novel framework which automatically generates an adaptive masking for masked language models based on the given context, for language model adaptation to low-resource domains. To this end, we proposed the Neural Mask Generator (NMG), which is trained with reinforcement learning to mask out words that are helpful for domain adaptation. We performed an empirical study of various rule-based masking strategies on multiple datasets for question answering and text classification tasks, which shows that the optimal masking strategy depends on both the language model and the domain. We then validated NMG against rule-based masking strategies, and the results show that it either outperforms, or obtains comparable performance to the best heuristic. Further qualitative analysis suggests that such good performance comes from its ability to adaptively mask meaningful words for the given task.
Acknowledgments
This work was supported by Samsung Advanced Institute of Technology (SAIT), the Engineering Research Center Program through the National Research Foundation of Korea (NRF) funded by the Korea Government MSIT (NRF2018R1A5A1059921), Institute for Information & communications Technology Promotion(IITP) grant funded by the Korea Government (MSIT) (No.2016-0-00563, Research on Adaptive Machine Learning Technology Development for Intelligent Autonomous Digital Companion, and No.2019-0-00075, Artificial Intelligence Graduate School Program (KAIST)), and a study on the “HPC Support” Project supported by the ‘Ministry of Science and ICT’ and NIPA.
References
Appendix A Algorithms
We provide the pseudocode of algorithm for meta-training of the Neural Mask Generator (NMG).
In the case of meta-testing, InnerLoop with the full task , pre-trained language model checkpoint , and trained policy as inputs can be considered as the meta-testing algorithm.
Appendix B Hyperparameters
We describe detailed hyperparameters in Table 4 for reinforcement learning (RL). We use the prioritized experience replay Schaul et al. (2016) with exponent value as 1. In addition, we address the concept of susampling (Mikolov et al., 2013) on replay sampling. Specifically, we divide the priority of each replay by the square root of word frequency within the corresponding context. For optimization, we use Adam Kingma and Ba (2015) optimizer to train the NMG model and its opponent (Self-Play).
B.2 LM Training in Meta-Train (Inner Loop)
We describe detailed hyperparameters in Table 5 for language model (LM) training in the meta-training. In the case of pre-processing of pre-training dataset, we use the context of triplets in Question Answering and sentences in Text Classification. Especially for emrQA (Pampari et al., 2018), we preprocess it to be same as other QA’s formats by removing yes-no and multiple answer type questions. Furthermore, we arbitrarily split a train dataset to a train and a validation set by 8 to 2 since Pampari et al. (2018) do not provide a separate validation set. For optimization, we used AdamW optimizer (Loshchilov and Hutter, 2019), with a linear learning rate scheduler. Each meta-training is done on the two Titan XP or RTX 2080 Ti GPUs and it costs maximum of 2 days in the case of BERT. The batch size is adequately selected according to the size of the data and the model.
B.3 LM Training in Meta-Test
We describe detailed hyperparameters in Table 6 for LM training meta-testing. In the meta-testing, we use same setting described in Section B.2 for pre-processing and optimization. The batch size is also adequately selected according to the size of the data and the model.
Regarding the pre-training epoch and the masking probability in meta-testing, we use two distinct settings for the baselines and NMG model. For the baselines, we train the LM for 1 epoch with following a conventional setting. However, we observed that the fewer masking with a more pre-training epoch is much more beneficial in the meta-training on the task with long contexts since it makes the NMG evaluate actions more precisely. Therefore, for the NMG model, we train the LM for 3 epochs with , following the setting of the meta-training.
B.4 Neural Mask Generator Architecture
The policy network of the NMG model consists of the single self-attention layer and two linear layers. The self-attention layer follows the configuration of the transformer layer of BERTBASE. We omit the linear layers of the original transformer Vaswani et al. (2017) implementation in our architecture. The hidden size of linear layers is 128 and gelu Hendrycks and Gimpel (2016) is used as an activation function. The value network also consists of the same linear layers after mean-pooling of word representations. The total number of parameters of the NMG model is approximately 2.5M, which is far smaller than the conventional language model.
B.5 Hyperparameter Searching
For searching proper hyperparameter, we use a manual tuning which tries conventional hyperparameters for the reinforcement learning (RL) and the language model (LM) training. A selection criterion depends on the result of the meta-testing. Especially, we set the criterion to F1 score and accuracy for question answering and classification task respectively.
Appendix C More Examples
To show the masking strategy from our NMG model, we additionally append additional examples from various datasets used in our experiments.