Exploring the Benefits of Training Expert Language Models over Instruction Tuning

Joel Jang, Seungone Kim, Seonghyeon Ye, Doyoung Kim, Lajanugen Logeswaran, Moontae Lee, Kyungjae Lee, Minjoon Seo

Introduction

Recent works show pretrained Language Models (LMs) that have been fine-tuned on multiple tasks with instructions (prompted instances), also known as multitask-prompted fine-tuned LMs and referred to as MT LMs in this work, can generalize to unseen tasks without task-specific fine-tuning (Wei et al., 2021; Sanh et al., 2021; Chung et al., 2022; Ye et al., 2022b; Ouyang et al., 2022; Wang et al., 2022a; Muennighoff et al., 2022). This paper raises some questions regarding the current paradigm of training MT LMs and is mainly divided into two parts. In Part 1, we report an unexpected finding regarding expert LMs (trained only on a single task) compared to MT LMs. In Part 2, we leverage the finding to highlight some of the benefits of expert LMs over MT LMs.

Previously, the key component to enhancing the unseen task generalization performance of MT LMs was thought to be scaling the total number of tasks used in training (Wei et al., 2021; Chung et al., 2022; Wang et al., 2022a). However, in this work, we show that training a single expert LM on oneTraining task: cosmos_qa, Prompt Name: no_prompt_text from Bach et al. (2022). out of the 300+ tasks used to train an MT LM (T0-3B (Sanh et al., 2021)) can outperform the MT LM by a non-trivial margin on 24 unseen tasks on mean accuracy.

Specifically, following the same experimental setup (training and evaluation) as T0-3B (Sanh et al., 2021), one of the most widely used MT LM, we first train expert LMs for each given training task (296) by freezing the underlying LM and updating adapters (Houlsby et al., 2019). We report a finding that shows 7 out of the 296 experts surpass T0-3B on the capability to generalize to unseen tasks on mean accuracy (shown in Figure 1). Using the top performing expert for all of the unseen task evaluation tasks surpasses T0-3B by a mean accuracy of 3.20% and 1.29% on 11 unseen datasets and 13 datasets of the BIG-Bench benchmark, respectively. We also show that applying a simple mechanism to retrieve relevant experts for each individual unseen task results in comparable performance to T0-3B. Considering the significant room for improvement when retrieving the best-performing expert for each unseen task (+11.94% compared to T0-3B), these results imply that choosing the right expert rather than naïvely utilizing a single MT LM for all of the unseen tasks can be a more efficient and effective approach.

Leveraging the finding of expert LMs showing improved unseen task generalization capability, we highlight three other advantages of training multiple expert LMs for each task and retrieving the relevant expert during inference (shown in Figure 2) compared to training MT LMs.

#1. MT LMs do not show the optimal performance for seen tasks because of negative task transfer, where learning multiple tasks at once hinders the learning of some specific tasks (Aghajanyan et al., 2021; Asai et al., 2022a; Zhang et al., 2022). Expert LMs, on the other hand, are not subject to negative task transfer (Levine et al., 2022) since each task is learned independently; We show our approach of selecting relevant experts during inference results in a +10.4% mean accuracy improvement on validation datasets of the 36 training tasks compared to T0-3B.

#2. MT LMs are susceptible to catastrophic forgetting (McCloskey & Cohen, 1989) of previous tasks when learning new tasks and require re-training on previous tasks to mitigate forgetting (Chakrabarty et al., 2022). Results show our distributed (training individual tasks in an independent manner) approach results in absolutely no degradation of seen tasks, even when adding the 8 new experts to the Expert Library, without re-training on previous tasks when learning 8 new generative tasks.

#3. We show that MT LMs show poor ability in performing composition of previously learned tasks given via concatenation of the corresponding instructions as a single compositional instruction. On the other hand, we show that merging the two experts trained on the individual tasks with mT5-3B (Xue et al., 2021) as the underlying pre-trained LM results in an expert that can outperform its MT LM counterpart, mT0-3B (Muennighoff et al., 2022), by a mean ROUGE-L score of +2.71 on 5 novel compositional tasks (summarization & translation). Details of the merging mechanism are provided in Section 3.3.

Related Work

Several studies have demonstrated that multitask fine-tuning moderately sized LMs with instructions, also referred to as instruction tuning, enables zero-shot task generalization. Specifically, Sanh et al. (2021); Wang et al. (2022a) have shown that scaling the number of training tasks, the number of prompts per task, and the size of the LM helps boost zero-shot task generalization performance. In addition to scaling these aspects, Chung et al. (2022) include Chain-of-Thought (Wei et al., 2022) tasks during instruction tuning, reaching state-of-the-art performance on zero-shot and few-shot settings with PaLM 540B (Chowdhery et al., 2022) as the underlying LM. Lin et al. (2022) improve MT LMs by adapting MT LMs on subsets of the training data retrieved given a few unlabeled examples of the unseen task. Ouyang et al. (2022) adapt MT LMs to align with human preferences through reinforcement learning. Muennighoff et al. (2022) include multilingual tasks to show cross-lingual generalization capability. Ye et al. (2022b) flip the instruction and label space to enhance generalization capability to novel unseen labels. Asai et al. (2022b) utilize instruction tuning to construct a general-purpose retrieval system. Similarly, Su et al. (2022) utilize instruction tuning to construct a general-purpose embedding model that can be used to perform different unseen tasks requiring text embeddings.

While previous literature has mostly asserted that the primary key component of MT LMs is scaling the total number of training tasks, in this paper, we propose an alternative perspective and instead show experimental results and findings that the feature of the tasks may be a more critical factor (analysis provided in Section 5); Similar findings are shown in the setting of few-shot adaptation (Chan et al., 2022) as well.

2 Retrieving task-specific embeddings

Retrieving task-specific parameters has the advantage of rapid target task adaptation, especially for low-resource scenarios (Vu et al., 2022; Asai et al., 2022a; Ye et al., 2022a; Qin & Eisner, 2021; Wang et al., 2022b; Bari et al., 2022). Vu et al. (2022) show that retrieving an optimal source soft prompt leads to better initialization for adapting to the target task. Asai et al. (2022a) also focus on retrieval of soft prompts for initialization for the target task but utilize the idea of attention weights to effectively interpolate between multiple training soft prompts. Similarly, Ye et al. (2022a) extend this idea of retrieving soft prompts, but utilize an MT LM as the underlying LM and do not fine-tune the LM to the target task, performing the target task in a zero-shot manner. Our work is motivated by Ye et al. (2022a), but proposes to replace the instruction tuning stage altogether, using vanilla pretrained LMs as the underlying LM instead of MT LMs. We accomplish this by training experts whereas previous work trained soft prompts on top of MT LMs.

3 Distributed Training of Language Models

Recent work has shown the possibilities and benefits of distributed training of LMs. Li et al. (2022) have shown that it is possible to merge individual LMs pretrained on different subsets (domains) of the training corpora to construct a single LM that shows lower overall perplexity compared to an LM trained on all of the corpora at once. Another line of work that explores merging individually fine-tuned LMs is Wortsman et al. (2022b), where they merge LMs fine-tuned on the same task with different configurations to boost performance. Similarly, Wortsman et al. (2022a) merge LMs fine-tuned on the same task, but with subsets of the training data for efficiency. Don-Yehiya et al. (2022) explore merging LMs fine-tuned on different tasks to make a multitask fine-tuned LM in a distributed manner, which has many benefits including federated learning (McMahan et al., 2017).

Other interesting extensions of distributed LM training include performing task arithmetic with task vectors (Ilharco et al., 2022), training and performing inference of several billion parameter LMs on distributed compute (Borzunov et al., 2022), introducing language-specific modules for growing the total capacity of multilingual LMs (Pfeiffer et al., 2022), finding theoretical guarantees of why merging works (Frankle et al., 2020; Ainsworth et al., 2022) and proposing novel methods of merging model weights (Matena & Raffel, 2021). In our work, we also show the benefits of distributed LM training by showing that the capability of expert LMs can be further amplified through merging individual experts.

Expert Language Models

In this section, we describe the framework of our proposed method. We train each expert by training adapters for each training task (Section 3.1). During inference, we retrieve the relevant experts from the Expert Library (Section 3.2). We additionally explore the effect of merging experts to observe the benefits of distributed training (Section 3.3).

For training the experts, we mainly explore parameter-efficient fine-tuning via adapters while freezing the underlying LM to train individual experts. We train experts for each task with the corresponding prompts and denote the resulting experts as Prompt Experts (PE).Each prompt (instruction) is referred to as tasks, following Chung et al. (2022). We also explore training experts for each dataset, which consists of multiple training prompts, referred to as Dataset Experts (DE). For training DE, instead of utilizing a parameter-efficient fine-tuning approach (adapters), we instead simply train the entire LM to observe the merging capability of expert LMs.Experimental results show that merging adapter experts does not lead to improved positive task transfer on mean accuracy (shown in Section 5). Figure 3 shows the hierarchy of the training datasets and the level at which PE and DE are trained on.

We apply a parameter-efficient method of representing experts by training additional adapters while freezing the original parameters (Houlsby et al., 2019). Specifically, given a standard Transformer LM with ll layers, input sequence XX containing TT tokens, the output for a single layer h1:Tl\textbf{h}^{l}_{1:T} is calculated by

where htl\textbf{h}^{l}_{t} is the hidden state of t-token after the l-th layer, Self-Att(·) is the self-attention module, and \textscFFNd\textsc{FFN}_{d}(·) is the feed-forward network with hidden dimensions d. When fine-tuning the LM with an adapter expert, each layer before the self-attention layer (Equation 1) changes into the following format:

where e represents the hidden dimension of the adapter feed-forward network. When using adapters to represent experts, parameters of \textscFFNe\textsc{FFN}_{e} are the only trainable parameters and the rest of the parameters in the LM are frozen.

2 Retrieval-of-Experts (RoE)

After independent (distributed) training of individual experts, we retrieve one of the experts to use during inference (Ye et al., 2022a). To this end, we construct an Expert Library and use dense retrieval to retrieve a relevant expert from the library to use during inference.

We first construct the Expert Library. This library contains keys that are each embedding representations of a single instance from the training tasks and values that are unique ids of the corresponding trained experts. For each unique expert, S training instances are randomly sampled and stored in the library. This results in [S×# of expertsS\times\textit{\# of experts}] entries in the Expert Library. To get the embedding representation of the training instances, we employ a simple Sentence Transformer (Reimers & Gurevych, 2019) as the dense retriever.We explore other text embedding models for the retriever such as Sentence-T5, SimCSE, INSTRUCTOR, etc., in Appendix B. Sentence Transformer shows the best performance among the embedding models. For the text format of the training instance that is given to the embedding model as input, we simply concatenate the answer choice (e.g. Yes∣|No, A∣|B∣|C∣|D) to the Prompted Input. The answer choice for generative tasks is given as ‘None’. We report ablation results of varying the text format given as the input to the embedding model in Appendix B.

Following the approach of Lin et al. (2022); Ye et al. (2022a), given a target task during inference, we first randomly select QQ instances from the target taskWe assume a scenario where we can perform batch-inference.. Next, we use the same text format (concatenation of Prompted Input and Answer Choice) and the same embedding model used to construct the Expert Library to obtain embedding representations of each of the QQ target queries. We then use MIPS (maximum inner product search) on our Expert Library to identify the most similar training instance (key) for each query instance, resulting in a total of QQ corresponding experts (value). We select the most frequently retrieved expert as the expert for solving the given target task.

3 Merging of Experts

Previous work has shown the possibility of distributed multitask fine-tuning by merging individually fine-tuned LMs (Don-Yehiya et al., 2022). Along with selecting the most retrieved expert, we observe how merging fully fine-tuned LMs (DE) affects the generalization performance on the unseen tasks.

A fully fine-tuned LM can be represented in the form of a vector τd=θd−θpre\tau_{d}=\theta_{d}-\theta_{pre} where θpre\theta_{pre} represent the full parameters of the vanilla pretrained LM and θd\theta_{d} represents the full parameters of the LM fine-tuned on the training dataset dd (Ilharco et al., 2022). The formula for merging of N experts can be denoted as follows:

where λi=1N\lambda_{i}=\frac{1}{N} as default if not stated otherwise. Note that when λi=1N\lambda_{i}=\frac{1}{N}, it results in merging experts uniformly. In some cases, however, performance was optimal when ∑iλi>1\sum_{i}\lambda_{i}>1 and each λi\lambda_{i} (representing the importance to place on τi\tau_{i}) and was set to a different value determined using a held-out validation dataset following Ilharco et al. (2022). A concrete example is provided in Appendix C.

Experimental Setup

Following the setting of Sanh et al. (2021), we use a total of 36 training datasets of T0 for training our experts.The original T0 (Sanh et al., 2021) paper includes 38 training datasets. However, we could not load 4 datasets from the Huggingface Dataset library: adversarial_qa/dbidaf, adversarial_qa/dbert, adversarial_qa/droberta, and duorc/SelfRC. Instead, we utilize the adversarial_qa/adversarialQA dataset and also additionally train on commonsense_qa dataset which is a variant of the cos_e dataset, resulting in a total of 36 training datasets. For each dataset, we use all of the prompts used to train T0 from the Promptsource Library (Bach et al., 2022) which results in a total of 296 prompts to train the corresponding experts (∼\sim8 prompts per training dataset). This results in 36 Dataset Experts (DE) represented via fully fine-tuned LMs, and 296 Prompt Experts (PE) via adapter training. For each individual fine-tuning, we randomly sample K=50,000K=50,000 training instances for each classification task and K=10,000K=10,000 for each generative task.We train with less number of instances for the generative tasks because the training generative tasks required longer max token length, and thus longer training time. We use the LM-adapted T5 model (Lester et al., 2021) checkpoint as our base model, and train for 5 epochs with a constant learning rate of 1e-4 for both adapter fine-tuning and full LM fine-tuning. For the construction of the Expert Library, much smaller S=100S=100 training instances are randomly sampled for each expert following Ye et al. (2022a).

We evaluate the baseline MT LMs (T0-3B, T0-11B) and our proposed method (T5-3B + DE/PE) on the same evaluation setting as the original T0 paper (Sanh et al., 2021): 11 unseen datasets that can be categorized into 4 task categories and on 13 datasets from BIG-Bench benchmark (Srivastava et al., 2022), which are diverse and challenging tasks that are not encountered during training.We exclude Novel Concepts task from the original T0 evaluation setting because it is a multi-label classification task. Multiple prompts are evaluated for each evaluation dataset. We further evaluate the models on 8 new generative tasksThe dataset details of the 8 new generative tasks are provided in Appendix A. that were not included in the original T0 paper evaluation setting. We use a rank classification evaluation by selecting the label option with higher log-likelihood following Brown et al. (2020); Sanh et al. (2021) for the classification tasks. For the generative tasks, we use the ROUGE-L score as the default metric if not stated otherwise. The details of each training and evaluation dataset are provided in Appendix A.

During inference, we set Q=32Q=32 for applying our Retrieval-of-Expert (RoE) mechanism. We do not separately perform ablations of SS and QQ, simply following the optimal setting of Ye et al. (2022a).

Expert LMs Can Generalize to Unseen Tasks

In this section, we show experimental results of expert LMs and show their potential for becoming a new paradigm over instruction tuning. Since this is a fairly novel approach of endowing LMs the capability to generalize to unseen tasks, we focus on providing proof-of-concept of some core research questions instead of making head-to-head comparisons with all of the baselines. We leave other extensive comparisons and exhaustive ablations for future work.

Table 1 shows the evaluation results on the 11 unseen datasets, Table 2 shows the results on the 13 unseen BIG-Bench tasks, and Table 3 shows the results on the 8 unseen generative tasks. Results from the three tables show that (1) a single PE significantly outperforms T0-3B, (2) the RoE (Orc.) outperforms other baselines by a non-trivial margin, and (3) our simple RoE approach outperforms T0-3B on the classification tasks, but not on generative tasks. Details of each finding are provided in the following paragraphs.

#1. In Table 1, surprisingly, T5(3B) + Cos PE, which is a Prompt Expert (PE) that is only trained on a single prompt (‘no_prompt_text’ prompt of Cosmos-qa dataset), outperforms its MT LM counterpart (T0-3B) on 8 out of 11 evaluation datasets and +3.20% on mean accuracy. Prior work shows that scaling the total number of training tasks during instruction tuning leads to better generalization; in our case, training an expert on a single task outperforms an LM trained on 300+ tasks (T0-3B). This finding is bolstered in Table 2 where the same Cos PE that shows the highest mean accuracy for the 11 unseen tasks outperforms T0-3B by +1.29% on the mean accuracy performance on 13 datasets of BIG-Bench Benchmark and in Table 3 where T5(3B) + Sam PE, which is a PE trained on (‘Given the above dialogue write a summary’ prompt of Samsum dataset), outperforms T0-3B by +6.83 mean score on the 8 generative tasks.

#2. In Table 1, we can see that T5(3B) + PE w/ RoE (Orc.), which is the upper-bound performance of choosing the best-performing expert based on the accuracy for each unseen task, outperforms T0-3B, much larger GPT-3(175B) and T0-11B by +11.94%, +4.37% and +2.61%, respectively, on the mean accuracy. T5(3B) + PE w/ RoE (Orc.) also outperforms T0-3B by +13.69 mean score on the 8 unseen generative tasks shown in Table 3. This means that RoE has a potential for strong unseen task generalization when the proper expert is chosen.

#3. T5(3B) + PE w/ RoE, which is a simple method of retrieving an expert for each unseen task leveraging an off-the-shelf retriever (Sentence Transformer (Reimers & Gurevych, 2019)), outperforms T0-3B on 8 out of 11 evaluation datasets and by +2.05% on mean accuracy. However, T5(3B) + PE w/ RoE underperforms T0-3B by -5.37 mean score on the 8 unseen generative tasks (Table 3). Considering that T5(3B) + PE w/ RoE still shows a significant performance gap compared to retrieving the best-performing expert (T5(3B) + PE w/ RoE (Orc.)), there is much room for improvement on the retriever side. One way to close the gap is to train a supervised retrieval model, which we leave for future work.

Table 4 shows the merging capability of expert LMs. The first three rows show the merging results of PE which are represented in the form of adapters. While Cos&Soc PE (Mer.), which is an expert constructed by performing uniform merging with Cos PE and Soc PESoc PE is a PE that was trained on Social_i_qa with prompt ‘no_prompt_text’ that showed the second highest mean accuracy on the 11 unseen tasks other than PE trained with Cosmos-qa. shows positive task transfers for some evaluation datasets (Copa & Story Cloze), not all of the results are the best or second best (RTE, Hellaswag, & Winogrande). This means that there was a negative task transfer when merging the adapter experts.

Thus, in order to further explore the merging capability of expert LMs, we train DE via full LM fine-tuning, known to be effective in previous literature (Ilharco et al., 2022), and merge them as shown in the last three rows in Table 4. Cos De (Cosmos-qa) and Soc De (Social-i-qa) are the two highest performing DE based on the mean accuracy performance on the 11 unseen tasks. While Cos&Soc DE (Mer.) shows only a +0.20% enhancement compared to Cos DE on mean accuracy, it still shows either the best or second best performance compared to the individual Cos and Soc DE. This implies that merging the two experts results in a composition of abilities. This opens up new possibilities of leveraging the merging of experts to unlock new capabilities which are further explored in Section 6 with the composition of instructions.

Overall, Table 4 shows that merging with adapters does not always result in positive task transfer while merging with full parameters seems to. Thus, future work should explore developing more parameter-efficient methods of merging expert LMs since always training and utilizing the entire LM weights is computationally demanding.

Figure 1 shows the mean accuracies of all the PE and DE results on the 11 unseen datasets. We highlight three main analyses from the figure and from the tables.

First, among the 8 training task categories, Multiple-Choice Question Answering (MCQA) training tasks generally show the strongest generalization capability. We hypothesize this to be the case because all of the 11 evaluation datasets are classification tasks and require some form of question answering via instructions. This extends the findings of Khashabi et al. (2020) that Multiple-Choice Question Answering (MCQA) generalizes well to not only different format QA tasks, but also different types of tasks such as natural language inference, story completion, coreference resolution, and word sense disambiguation as well.

Second, among the 36 training datasets, 3 datasets consistently ensure high performance for both PE and DE: Cosmos-qa (Huang et al., 2019), Social-i-qa (Sap et al., 2019), and Dream (Sun et al., 2019). All three datasets are commonsense reasoning datasets, which have been considered to be crucial for generalization to unseen tasks (Lourie et al., 2021). We provide the full ranking of the PE and DE for the 11 unseen tasks shown in Figure 1 in Appendix D.

Lastly, T5(3B) + Sam PE which is a PE trained on Samsum (Gliwa et al., 2019), a dataset with abstractive dialogue summaries, shows the best mean score on the 8 unseen generative tasks in Table 3, outperforming T0-3B by +6.83 mean score. However, the same PE shows one of the lowest ranks for the 11 unseen (classification) tasks (shown in Appendix D) underperforming T0-3B by -9.15% mean accuracy. This shows that there is no free lunch: a PE that shows high mean performance for unseen generative tasks do not show high mean performance for unseen classification tasks. This also implies that it is more-so important to retrieve the correct expert dynamically depending on the given context (target task).

Benefits of Expert LMs over MT LMs

In this section, we highlight the 3 main benefits of expert LMs and RoE over MT LMs.

First, we show that expert LMs are less susceptible to negative task transfer by comparing the performance of T5(3B) + PE w/ RoE on the validation datasets of the 36 training datasets with two MT LMs, T0-3B and T0-11B. As shown in Table 5, our distributed approach outperforms T0-3B and T0-11B by +10.40% and +7.70% on mean accuracy, respectively.

This is because since evaluation is done with seen instructions, our simple retrieval mechanism is highly likely to select the best-performing expert from the Expert Library, showing comparable performance to T5(3B) + PE w/ RoE (Orc.). In fact, T5(3B) + PE w/ RoE retrieves the PE from the same training dataset on 280 out of 296 seen tasks, and the PE trained with both the same prompt and dataset (oracle) on 185 out of 296 seen tasks.

In some scenarios when we want to additionally fine-tune LMs on additional datasets after model deployment, making finetuned LMs continual learners is important (Chakrabarty et al., 2022). This is because performing instruction tuning on the entire set of original and additional tasks in each update would lead to heavy computation. Previous work mitigates this issue through a rehearsal-based method, continually training the instruction-tuned LM on samples of the original and additional tasks (Chakrabarty et al., 2022). However, this approach (1) assumes that we have access to the original datasets and (2) still leads to additional computational overhead, especially when scaling the total number of seen tasks during instruction tuning.

We show that we can accomplish the same feat through distributed training of experts without any access to original, seen datasets by training separate experts for each additional task and simply adding them to the Expert Library. Specifically, we show the comparison between continually training an MT LM (T0-3B) which is referred to as CT0-3B through a rehearsal-based approach, and our distributed approach on 8 new generative tasks in Table 6. The 8 generative tasks for continual learning were chosen following the previous work (Chakrabarty et al., 2022).

The table shows that our distributed approach results in absolutely no degradation of performance for the seen task, a minor (-0.15%) degradation for unseen tasks, and superior mean performance (+1.08) for the 8 target tasks compared to the MT LM counterpart, outperforming CT0-3B on 7 out of the 8 target tasks. This shows that without any access to original datasets or heavy computational cost, our distributed approach is mostly able to retain its original ability (seen & unseen) as well as outperform CT0-3B on the target tasks. We leave scaling the number of new target tasks and how our distributed approach performs against its instruction-tuned counterpart for future work.

Prior work has shown the need for performing compositional instructions (Logeswaran et al., 2021; Corona et al., 2021; Khot et al., 2022). For example, we can give the following instruction to the LM: “Write a summary of the following English text and translate the sentence into Korean.” where “Write a summary of the following English text.” and “Translate the sentence into Korean.” are two separate instructions seen during training. To test this compositional capability, especially in a multi-lingual setting, we utilize the mT0-3B (Muennighoff et al., 2022) as our MT LM and evaluate the composition of performing 5 novel compositional tasks of summarization and translation. To explore the benefits of merging experts for performing compositional instructions, we perform 6 full fine-tuning with mT5-3B (Xue et al., 2021) as the underlying vanilla pretrained multilingual LLM: We use xsum to train one English Summarization expert and use five translation pairs in tatoeba (en→\rightarrowes, en→\rightarrowfr, en→\rightarrowja, en→\rightarrowzh, en→\rightarrowko) to train the corresponding five translation experts. During inference, we merge the summarization expert with each of the five translation expertsWe provide the specific configurations used for merging such as the λi\lambda_{i} values for each task vector τi\tau_{i} and the training and validation stats in Appendix C. Note that both xsum and tatoeba are part of the training tasks used during instruction tuning of mT0-3B.

Evaluation results on the five compositional tasks are shown in Table 7. Our distributed approach, mT5-3B + Mer. Ex., outperforms its MT LM counterpart, mT0-3B on 4 out of the 5 tasks and by a mean ROUGLE-L score of +2.71; This is due to a significant performance gap for the tasks involving low-resource languages (Korean and Japanese) because the low-resource languages are protected from negative transfer when doing distributed training. Cherry-picked output examples of the MT LM and the merged experts are provided in Table 8.

Limitations and Discussions

While we highlight some of the major drawbacks of instruction tuning and propose an alternative approach of instead training and retrieving experts in this paper, we do not perform experimental results over MT LMS that have more than >>11B parameters. For example, MT LMs with >>11B parameters may be less susceptible to negative task transfer because of increased model capacity. Also, during the inference of unseen tasks, our retrieval mechanism assumes batch inference (i.e. having access to 32 samples of the target tasks without labels). Finally, when showing the compositional instruction experiments, we assume the two optimal experts could be retrieved from the compositional instruction (concatenation of the two seen instructions) given as the input along with the evaluation instance. This might not necessarily be the case with more complex, compositional instructions, which might require a separate decomposition stage. We instead focus on showing the possibility merging experts can bring and leave developing novel methods of retrieving the optimal experts during inference for future work.

Conclusion

In this work, we provide an interesting finding that expert LMs trained on single tasks show strong generalization capability to unseen tasks, even surpassing MT LMs trained on multiple tasks (300+) by a non-trivial margin. We leverage this capability and show three main benefits of training and retrieving experts for inference over MT LMs, demonstrating that our proposed distributed approach is more robust against negative task transfer, more adapt at learning new tasks, and can perform compositional instructions. To this end, we urge the research community to further explore distributed and collaborative training of experts which may have other future benefits including efficiency, privacy, and personalization not explicitly explored in this paper.

We thank Colin Raffel, Sungdong Kim, Sejune Joo, Miyoung Ko, Eunbi Choi, Hyunji Lee, Dongkeun Yoon, Yoonjoo Lee, and Yujin Kim for the useful discussion and feedback on the paper.

References

Appendix A Details of Training and Evaluation Datasets

Following Sanh et al. (2021), we use 36 training datasets from the 8 task categories for training our experts. We provide the official names given in Huggingface Datasets: Sentiment Classification (Senti.) imdb (Maas et al., 2011), amazon_polarity (McAuley & Leskovec, 2013), rotten_tomatoes (Pang & Lee, 2005), yelp_review_full (Zhang et al., 2015b), and app_reviews. Paraphrase Identification (Para.) glue/qqp (Wang et al., 2018), glue/mrpc (Wang et al., 2018), and paws/labeled_final (Zhang et al., 2019). Topic Classification (Topic C. ag_news (Zhang et al., 2015a), dbpedia_14 (Lehmann et al., 2015), and trec (Li & Roth, 2002). Summarization (Summ.) gigaword (Graff et al., 2003), multi_news (Fabbri et al., 2019), samsum (Gliwa et al., 2019), xsum (Narayan et al., 2018), and cnn_dailymail/3.0.0 (See et al., 2017). Structure-To-Text (STS) common_gen (Lin et al., 2020) and wiki_bio (Lebret et al., 2016). Multiple-Choice Question Answering (MCQA) commonsense_qa (Talmor et al., 2019), dream (Sun et al., 2019), quail (Rogers et al., 2020a), qasc (Khot et al., 2020), quarel (Tafjord et al., 2019), cos_e/v1.11 (Rajani et al., 2019), quail (Rogers et al., 2020b), social_i_qa (Sap et al., 2019), wiqa (Tandon et al., 2019), cosmos_qa (Huang et al., 2019), sciq (Welbl et al., 2017), and wiki_hop/original (Welbl et al., 2018) Extractive Question Answering (EQA) adversarial_qa/adversarial_qa (Bartolo et al., 2020b), quoref (Bartolo et al., 2020a), ropes (Lin et al., 2019), and duorc/Paraphrase IdentificationRC (Saha et al., 2018) Closed Book Question Answering (CBQA) kilt_tasks/hotpotqa (Petroni et al., 2021) and wiki_qa (Yang et al., 2015).

Following Sanh et al. (2021), we include 11 evaluation datasets as follows: RTE (Dagan et al., 2005), CB (De Marneffe et al., 2019), ANLI (Nie et al., 2020) for natural language inference task, COPA (Roemmele et al., 2011), Hellaswag (Zellers et al., 2019), Storycloze (Mostafazadeh et al., 2016) for sentence completion task, Winogrande (Sakaguchi et al., 2021), WSC (Levesque et al., 2012) for coreference resolution task, and WiC (Pilehvar & Camacho-Collados, 2019) for word sense disambiguation task.

For BIG-bench tasks, we evaluate on 13 tasks, following Sanh et al. (2021): Known Unknown, Logic Grid, StrategyQA, Hindu Knowledge, Movie Dialog, Code Description, Conceptual, Language ID, Vitamin C, Syllogisms, Misconceptions, Logical Deduction, and Winowhy.

For the generative evaluation tasks, we follow Chakrabarty et al. (2022) and utilize 8 tasks: Text Simplification (WikiAuto) (Jiang et al., 2020), Headline Generation with constraint (HGen) (Yamada et al., 2021), Haiku Generation (Haiku), Covid QA (Möller et al., 2020), Inquisitive Question Generation (ELI5) (Fan et al., 2019), Empathetic Dialogue Generation (EmDg) (Rashkin et al., 2019), Explanation Generation (eSNLI) (Camburu et al., 2018), and Twitter Stylometry (Twitter)

Appendix B Varying the Embedding Model and Text Format for Retrieval of Experts

While Ye et al. (2022a) used T0 (Sanh et al., 2021) as the base embedding model to retrieve prompt embeddings, we explore 13 different sentence embedding models to waive the need of using instruction tuned models for retrieval of expert LMs.

More specifically, we list of embedding models we use are as follows: (a) 4 different variants of Sentence Transformer model (Reimers & Gurevych, 2019): all-MiniLM-L6-v2, all-MiniLM-L12-v2, all-mpnet-base-v2, nli-mpnet-base-v2, (b) 2 different variants of SimCSE model (Gao et al., 2021): sup-simcse-roberta-large, unsup-simcse-roberta-large, (c) 2 different variants of Instructor model (Su et al., 2022): hkunlp/instructor-large, hkunlp/instructor-xl, (d) 2 different variants of GTR model (Ni et al., 2021): gtr-t5-large, gtr-tr-xl, (e) 2 different variants of SentenceT5 model (Ni et al., 2022): sentence-t5-large, sentence-t5-xl, and (f) DiffCSE model (Chuang et al., 2022): voidism/diffcse-bert-base-uncased-sts which are all available on HuggingFace. Note that we try different embedding models in an unsupervised manner, i.e., not requiring any supervision to train the embedding model, but using it off-the-shelf. The results are shown in Table 9.

We also try different variants of text format given to the embedding model. Using Promptsource (Bach et al., 2022), we compare including the instance, label list, answer choice in 2 different formats. Specifically, the full list of text formats are as follows: (a) ‘Instance: {instance}’, (b) ‘Answer Choices: {label list}’, (c) ‘Answer Choices: {answer choice}’, (d) ‘Answer Choices: {label list}, Instance: {instance}’, (e) ‘Answer Choices: {answer choice}, Instance: {instance}’, (f) ‘{instance}’, (g) ‘{label list}’, (h) ‘{answer choice}’, (i) ‘{label list}<</s>>{instance}’, (j) ‘{answer choice}<</s>>{instance}’. Label list and answer choice differ in that while label list uses the actual label options (e.g., [‘swim’,‘fly’,‘walk’,‘run’]), answer choice organizes them with a ‘—’ deliminator in the middle (e.g. A∣|B∣|C∣|D). The results are shown in Table 10.

While we tried different variants, the oldest, yet most chosen model all-minilm-l6-v2 outperforms other options. We conjecture that this is because most of the model variants we tested were trained as sentence embedding models, not for embedding prompted instances. Prompted instances are some how structural and formatted compared to natural language sentences used for training sentence embedding models. In terms of text format, using both the prompted instance and the answer choice showed the best results. These results show that for the dense retriever to map instances, it should rely on both components, which are orthogonally important. Also, using the actual label option harms performance compared to using the answer choice, which indicates that the output format itself is important to retrieve well-matched expert LMs.

Appendix C Details of Performing Compositional Instructions

Our compositional instruction setting consists of a total of 400 instances for each task (300 instances for the validation set, and 100 instances for the test set.) per language that was obtained using google translate to change the input of the XL-Sum (Hasan et al., 2021) dataset. We thus use the ground truth label in the specified language and the input is the machine-translated version. The reason for this is that we measure the λi\lambda_{i} values (the importance to place on each task vector τi\tau_{i}) by performing evaluation on the validation datasets. Empirically, setting 1.0 for each λi\lambda_{i} value resulted in the best performance. Thus, as mentioned in the method section, the total ∑λi\sum\lambda_{i} results in 2.0, greater than 1.0.

We also vary the decoding strategies to check the performance of merging two experts finetuned from mT5-3B compared with naive mT0-3B on XL-Sum dataset. The detailed optimal setting we found is as follows:

Here are the actual inputs for the LM generated & ground truth output examples shown in Table 8. The compositional instruction portion is shown in bold.

English →\rightarrow Spanish: “Write a summary of the following English text and translate the sentence into Spanish: The French police arrested four members of the child’s family for their alleged involvement on Tuesday. Police sources told local media that the child refused to do his homework and that he was beaten with the stick of a broom. The 20 -year -old sister, his older brother and his girlfriend were present at the time of the incident and were arrested. The three called emergency services, which could not save the child. The alleged crime occurred on September 17 at the family’s home in the town of Mulhouse, in the east of the country, and of just over 100,000 inhabitants. Although the child’s mother was not at home because she was on a trip for work reasons, she was also arrested. The authorities say it will be questioned to confirm whether it encouraged the punishment. The four family members remain in police custody and must appear before the Mulhouse Prosecutor’s Office for a judicial investigation. Prosecutor Edwige Roux-Morizot will investigate the case. Moretones after the death of the child, victim of cardiac arrest, several neighbors celebrated a vigil in their honor and met with the child’s parents to offer them comfort. However, the results of the autopsy motivated the police to carry out an investigation into what happened. The child’s body presented several bruises, especially at his feet, according to AFP. Despite the confirmation of cardiac arrest, pathologists said the cause of death was probably the blows he had suffered. A police source said the child was beaten with blunt objects. Although the main suspect of the murder is the older brother, the French authorities hope that the investigation will shed light on what happened. France is one of the 13 countries of the European Union where corporal punishment is legal. A legal practice The National Assembly of France is considering approveing a law to prohibit corporal punishments for children. There are two new law proposals that would grant children a violence -free education, venting parents to use ”forms of humiliation such as physical or verbal…”

English →\rightarrow French: “Write a summary of the following English text and translate the sentence into French: The former Minister of Justice of Malawi, Ralph Kasambara, was arrested on November 8, 2013. Mr. Kasambara was found guilty of conspiracy in the assassination in September 2013 of the former budget director at the Ministry of Finance, Paul Paul MPHWIYO. The murder of Mr. Mphwiyo had led to the discovery of the scandal of ”cashgate”, the systematic looting of public resources, during the administration of President Joyce Banda. Nearly 250 million had been fraudulently paid to businessmen for services who have never been rendered. A few days before the tragedy, a subordinate official would have been found with gold bars belonging to the cash, the equivalent of more than $ 300 million, in the trunk of his car. Money was also confiscated at the home of certain officials and in chests from their vehicles. Immediately after his conviction last month, Kasambara had suggested that he would not appeal the court verdict.”

English →\rightarrow Japanese: “Write a summary of the following English text and translate the sentence into Japanese: Vice Chairman Meng Ship, the highest financial manager (CFO), was the daughter of the founder arrested in Vancouver, Canada last December, and Vice President Meng Teng was sanctioned at Vancouver Airport last December. He was arrested for violating and associated scams and was charged at the end of January this year. The United States authorities are seeking to hand over the vice chairman, but they deny the charges. Defendant Meng filed an administrative lawsuit for the Canadian government, the immigration bureau, and the police for ”significantly infringing” their citizenship. China has accused the defendant’s arrest and delivery procedure as a ”political project.” ¡Related article¿ Introduction is ”illegal” and ”Dandridy” British Columbia Senior Court on the 1st, and Meng is the Canadian government and the Royal Canadian equestrian police (RCMP), and the Canadian Immigration Bureau (CBSA). He is complaining of civil rights infringement. Before the arrest of RCMP, CBSA complained that he had detained himself on unfair claims, investigated and interrogated his belongings. The vice chairman was bail and was at Vancouver’s home, and the authorities arrested Vice Chairman Meng on the spot. He complained that it infringed on the rights based on the Canadian Characters of Human Rights. In addition, Vice -Chairman’s detention was ”illegal” and ”arbitrary”, and authorities pointed out that ”the reason for detention, the right to call lawyers, or the right to be paid to be silent.” What is the reaction of each country? The relationship between China, Canada and the United States has deteriorated over the arrest of Vice Chairman Meng. In January, the U.S. Department was charged with 23 cases of Huawei and Vice Chairman Meng. In addition to bank fraud, communication fraud, judicial obstruction, a major US telecommunications equipment T -mobile has been charged with trying to steal technology. China accused these movements as ”abuse of the handover agreement” between the United States and Canada, and stated that they…”

Englsih →\rightarrow Chinese: “Write a summary of the following English text and translate the sentence into Chinese: Dr. Craig Spencer, who is infected with Ebola virus, is currently being hospitalized at the New York Metropolitan Hospital. Caisyex said that the isolation experience is very scary and may also make other medical workers reluctant to go to West Africa to help curb the Ebola epidemic. Following New York and New Jersey, Illinois has also adopted a strict isolation policy. New measures means that those who have come into contact with any Ebola patient in West Africa will be forced to isolate for 21 days. U.S. President Obama Obama said in a weekly radio speech on September (October 25) that Americans should believe in the facts rather than being dominated. He also reiterated that he can infect the virus only with direct body fluids with Ebola patients. Higkos, who was ”criminals”, who was an isolated person, said that she had witnessed ”confusion, panic, and the most terrifying isolation” when she returned from Sierra Leone on Friday (24th). Hekox wrote a newspaper in the United States: ”I don’t know how many medical workers who fought with Ebola virus in the West African epidemic area will have the same encounter.” She said, ”Will they feel like criminals like criminals like criminals? She also said that she was isolated for seven hours at the airport terminal, but she only got a grain rod to fill her hunger. She denied that she had had a fever and said that she was just blushing at the time because she was not satisfied with the treatment at the airport. Even though Hiccoks was negative in Ebola virus testing, she was still being isolated for three weeks and was monitored by medical officials. Frontline medical staff was deeply influenced by the Ebola outbreak. After being diagnosed with Ebola patients, a doctor of New York, who had worked in Guinea last week, was diagnosed with Ebola patients, New York State and New Jersey have strengthened their isolation measures. Spencer is currently receiving isolation treatment in a hospital in New York. Mali has also recently appeared in Ebola, and President Ibrahim… ”

English →\rightarrow Korean: “Write a summary of the following English text and translate the sentence into Korean: According to the Korea Meteorological Administration, January this year was the warmest winter since 1973, when the weather observation began in the Korean peninsula. The average temperature in January last month was 2.8 degrees. This is 3.8 degrees higher than the average of minus 1.0 degrees in January, 1981 2010. The previous average temperature record was 1.6 degrees in 1979. Except for the first day of the new year, the average temperature in the country was higher than normal. Due to the high temperature, the snowfall was the lowest. The Korea Meteorological Administration cited the introduction of warm southwestern air flow into the Siberian region, and the fact that the ’pole whirl’, which traps cold air in the Arctic, was strong as an abnormal temperature. It also analyzed that the warm south wind flow was introduced to the Korean peninsula due to the high sea level temperature of the Western Pacific. Nationwide weather data in January, the average temperature in the coldest January of the year has continued to rise in recent years. According to the weather data released by the Korea Meteorological Administration in January 1973-2020, the average temperature in January in Korea is steadily rising. Choi Jung -hee, the Korea Meteorological Agency, said that the warming of winter is ”global warming impact,” and ”most of the monthly weather data tends to be similar.” Detection of the ecosystem change is detected throughout the ecosystem. The first spawning season of ’Bukbangsan Guri’, a climate change indicator, has been faster. Mudeungsan National Park Eastern Office said on the 24th of last month that the first spawning of the North Bangsan Gogi, a species designated by the Ministry of Environment, was observed. It was first observed. It is 27 days earlier than February 19 last year. This is the first time that spawning has been observed in January since 2010, when the survey began. Researchers at the Park Industrial Complex believed that the spawning day was advanced due to the exceptionally warm… ”

Appendix D Full List of PE and DE ranked on the 11 unseen datasets

Table 11 shows the full list of DE and Table 12 shows the full list of PE, both lists sorted in descending order with regards to the mean accuracy on 11 unseen tasks.