OLMoE: Open Mixture-of-Experts Language Models
Niklas Muennighoff, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Jacob Morrison, Sewon Min, Weijia Shi, Pete Walsh, Oyvind Tafjord, Nathan Lambert, Yuling Gu, Shane Arora, Akshita Bhagia, Dustin Schwenk, David Wadden, Alexander Wettig, Binyuan Hui, Tim Dettmers, Douwe Kiela, Ali Farhadi, Noah A. Smith, Pang Wei Koh, Amanpreet Singh, Hannaneh Hajishirzi
Introduction
Despite significant advances in Large Language Models (LMs) on various tasks, there remains a clear trade-off between performance and cost in both training and inference. High-performing LMs are inaccessible for many academics and open-source developers as they are prohibitively expensive to build and deploy.For example, even with 16 H100 GPUs and several optimizations, Llama 3 405B only achieves a decoding throughput of around 100 tokens per second . One approach to improve the cost-performance trade-off lies in using sparsely-activated Mixture-of-Experts (MoEs) . MoEs have several experts in each layer, only a subset of which is activated at a time (see Figure 2). This makes MoEs significantly more efficient than dense models with a similar number of total parameters, which activate all parameters for every input . For this reason, industry frontier models use MoEs including Gemini-1.5 and reportedly GPT-4 .
Most MoE models, however, are closed-source: while some have publicly released model weights , they offer limited to no information about their training data, code, or recipes (see Figure 1). While there have been prior efforts in making language modeling research fully accessible , they have been largely limited to dense LMs. This comes despite MoEs requiring more openness as they add complex new design questions to LMs, such as how many total versus active parameters to use, whether to use many small or few large experts, if experts should be shared, and what routing algorithm to use. The lack of open resources and findings about these details prevents the field from building cost-efficient open MoEs that approach the capabilities of closed-source frontier models.
To address these issues, we introduce OLMoE, a fully open Mixture-of-Experts language model with state-of-the-art performance among similarly-sized models. In particular, we pretrain OLMoE-1B-7B for 5.1 trillion tokens with 6.9B total parameters, of which only 1.3B are activated for each input token. This leads to a similar inference cost as using dense models with around 1B parameters, such as OLMo 1B or TinyLlama 1B , but requires more GPU memory to store its 7B total parameters. Our experiments show that MoEs train 2 faster than dense LMs with equivalent active parameters. In Figure 1, we show that OLMoE-1B-7B significantly outperforms all open 1B models and displays competitive performance to dense models with significantly higher inference costs and memory storage (e.g., similar MMLU scores to Llama2-13B, which is 10 more costly). Via instruction- and preference tuning, we create OLMoE-1B-7B-Instruct, which we find exceeds various larger instruct models including Llama2-13B-Chat , OLMo-7B-Instruct (0724), and DeepSeekMoE-16B on common benchmarks (MMLU, GSM8k, HumanEval, etc.).
Our comprehensive set of controlled experiments highlights key design choices for MoEs (see Table 1) and LMs in general. One critical design decision for making MoEs performant is the use of fine-grained routing with granular experts : we employ 64 small experts in each layer with 8 being activated. The choice of routing algorithm is also important: we find dropless token-based routing outperforms expert-based routing . Our findings also include those that challenge prior work, such as the ineffectiveness of shared experts and the limited benefits of sparsely upcycling a pretrained dense LM into an MoE unless under small compute budgets. Finally, we analyze the routing behavior in OLMoE-1B-7B, finding that routing saturates early in pretraining, experts are rarely co-activated, and experts exhibit domain and vocabulary specialization.
We hope our fully open MoE facilitates more research and analysis to improve our understanding of these models. We release training code, intermediate checkpoints (every 5000 steps), training logs, and training data under open-source licenses (Apache 2.0 http://www.apache.org/licenses/LICENSE-2.0 or ODC-By 1.0 https://opendatacommons.org/licenses/by/1-0/).
Pretraining and Adaptation
OLMoE is a decoder-only LM consisting of transformer layers. The feedforward network (FFN) in dense models like OLMo , is replaced with an MoE module consisting of smaller FFN modules called experts, of which a subset of experts are activated for each processed input token (also see Figure 2):
where , called the router, is a learned linear layer mapping from the input logits to the chosen experts. A softmax is applied to the router outputs to compute routing probabilities for all experts. Each selected expert processes the input , the output of which is then multiplied with its respective routing probability. The results are then summed across all chosen Top- experts to constitute the output of the MoE module for a single layer of the model out of its total layers. Key decisions in designing an MoE model include determining the number of activated and total parameters, the design of the experts (e.g., granularity, whether or not to include shared experts), and the choice of the routing algorithm. Moreover, training an MoE model can involve initializing from a dense model (sparse upcycling) and changing the training objective, such as including auxiliary load balancing and router z-losses. Experiments related to these design choices are in §4.1; Table 1 shows our final decisions.
In summary, we use 1.3B active parameters out of a total of 6.9B, with 8 activated experts out of 64 per layer. We use dropless token choice routing : for each input token, the learned router network determines 8 experts to process it. We train OLMoE-1B-7B from scratch with two auxiliary losses: load balancing loss () and router z-loss () , which we define and experiment with in §4.1.6 and §4.1.7, respectively. We multiply them with respective loss weights, and , and sum them linearly with the cross entropy loss () to arrive at our final training loss:
Our full pretraining configuration for OLMoE-1B-7B is in Appendix B.
We use a mix of data from DCLM and Dolma 1.7 , which includes the following: (1) a quality-filtered subset of Common Crawl, referred to as DCLM-Baseline, (2) StarCoder, Algebraic Stack and arXiv, used in both DCLM and Dolma 1.7, and (3) peS2o and Wikipedia from Dolma 1.7. We refer to our pretraining dataset as OLMoE-Mix.
To all sources above, we apply a filter that removes all documents with a sequence of 32 or more repeated n-grams, where an n-gram is any span of 1 to 13 tokens. For the StarCoder subset, we also remove any document that is either from a repository with fewer than 2 stars on GitHub, or whose most frequent word constitutes over 30% of the document, or whose top-2 most frequent words constitute over 50% of the document.
We shuffle all samples randomly at the beginning of each epoch and train for a total of 5.133T tokens (1.3 epochs following Muennighoff et al. ). During our annealing phase (final 100B tokens) we first reshuffle the entire dataset and then linearly decay the learning rate to 0, following prior work . Our pretraining data statistics are in Table 2.
We create OLMoE-1B-7B-Instruct by following a standard adaptation recipe split into instruction tuning followed by preference tuning building on prior open models . In our instruction tuning dataset, we add more code and math data to boost performance on downstream coding and math applications. Other models, such as GPT-4 and Llama 3 , instead include samples from math datasets like GSM8k or MATH during pretraining. We also include No Robots and a subset of Daring Anteater as they are of high quality and add diversity, two key factors for successful adaptation . We describe our adaptation datasets in Table 3 and hyperparameters in Appendix B.
Results
Our evaluation procedure consists of three parts: During pretraining, After pretraining, and After adaptation. We detail the setup for each in Appendix C.
In Figure 3 we benchmark the performance of OLMoE-1B-7B during pretraining with the current best OLMo models on commonly used downstream tasks. We find that across all tasks OLMoE-1B-7B reaches better performance with less compute (FLOPs) than the dense OLMo models. OLMoE-1B-7B matches or outperforms OLMo-7B at the end of training despite OLMoE-1B-7B having used less than half as many FLOPs for training and using only 1B active parameters. This is likely a result of the dataset and modeling changes we make to the OLMo setup including MoE-related changes, stability, and performance improvements, outlined in Appendix B. Appendix E contains training and validation loss plots showing very smooth loss curves without major loss spikes during the 5T tokens of our pretraining.
In Table 4 we benchmark OLMoE-1B-7B on common downstream tasks. We find that OLMoE-1B-7B performs best among models that use less than 2B active parameters, making it the most economical option for many use cases of LMs. For larger budgets, Qwen1.5-3B-14B has stronger performance but has more than double the active and total parameters than OLMoE-1B-7B. We find that despite requiring 6–7 less compute per forward pass, OLMoE-1B-7B outperforms some dense LMs with 7B parameters such as Llama2-7B , but falls short of others like Llama3.1-8B . Figure 1 compares MMLU performance with active parameters, a proxy for the value of a model given its cost, of OLMoE-1B-7B and other LMs. OLMoE-1B-7B is the state of the art in its cost regime.
In Table 5, we benchmark our instruction (SFT) and preference (DPO) tuning of OLMoE-1B-7B. SFT improves our model on all tasks measured. We observe a 10 gain on GSM8k, likely due to our inclusion of additional math data to account for the relatively small amounts of math data during pretraining (§2). DPO helps on most tasks, especially AlpacaEval which aligns with findings from prior work . Our DPO model, which we refer to as OLMoE-1B-7B-Instruct, has the highest average among all models benchmarked. We find it to outperform the chat version of Qwen1.5-3B-14B despite Qwen having 2 more parameters and its pretrained model outperforming OLMoE-1B-7B in Table 4. The 84% score on AlpacaEval also outperforms much larger dense models on the leaderboard,https://tatsu-lab.github.io/alpaca_eval/ such as Llama2-13B-Chat .
Experimenting with Alternative Design Choices
In this section, we present pretraining and adaptation experiments that have led to OLMoE-1B-7B. We group them into experiments on settings specific to Mixture-of-Experts (§4.1), experiments on settings applicable to both dense LMs and MoEs (§4.2), and adaptation experiments (§4.3). In pretraining experiments, we often use MMLU Var, a version of MMLU with varying few-shots and a different format that provides signal earlier during training. We describe our full evaluation setup in Appendix C and provide additional experiments in Appendix F. Each experiment links to a Weights & Biases report with more validation and downstream results, and the full configurations of the runs. To isolate the impact of changes and minimize confounders, we vary only one hyperparameter for each experiment. Nevertheless, due to the large number of hyperparameters, some results may change under different configurations and we cannot guarantee the correctness of each of our hyperparameter choices. Models are not comparable across different experiments, as we vary the base model to incorporate successful findings.
Prior work reports various speed-ups of MoEs over dense models: Artetxe et al. report that MoEs require 2–4 less compute to match dense models, MoMa exhibits 2.6 FLOP savings for language tasks, Arctic yields 4 FLOP savings but for very different dense and MoE configurations, and Switch Transformers train 2-7 faster with MoEs but for encoder-decoder models while the other works study decoder-only LMs .
In Figure 4, we compare MoEs and dense models in a controlled setup. We find that our MoE reaches the performance of the dense model with 3 fewer tokens equivalent to 3 less compute measured in FLOPs. However, due to the additional memory overhead of training the MoE with its 7B total parameters, it processes fewer tokens per second than the dense model (23,600 tokens per second per GPU for the MoE vs. 37,500 for dense). Thus, in terms of training time, it reaches the performance of the dense model only 2 faster. There are likely optimizations possible that would bring the speed-up closer to the 3 token speed-up, which we leave to future work. Based on these results, we select an MoE configuration with 6.9B total and 1.3B active parameters matching OLMo-7B in total and OLMo-1B in active parameter count, respectively.
1.2 Expert Granularity
Dai et al. propose to use small fine-grained experts to allow more combinations of experts and thus make the model more flexible. For example, the Mixtral model uses the common configuration of 8 experts per layer, 2 of which are activated. This allows for combinations per layer. By halving the size of each expert and therefore doubling the number of experts to maintain the same compute and parameter budget, we can increase the possible combinations to . Krajewski et al. investigate compute-optimal granularity configurations finding that higher compute budgets warrant more granular experts.
In Figure 5, we observe that more granular experts improve training loss, validation loss, and downstream performance. The 8-expert configuration uses 1 active expert, which yields combinations. By quartering the size of each expert but increasing the number to 32 with 4 active ones ( combinations), we observe an improvement of around 10% on HellaSwag and MMLU at around 130 billion tokens. However, we find that there are diminishing returns to granularity. The additional increase to 64 experts with 8 active ones ( combinations) improves downstream metrics by a smaller amount of 1–2%. For our OLMoE-1B-7B compute budgetApproximated via , where are active parameters (1B) and are training tokens (5T). of , Krajewski et al. predict an optimal number of experts of 256 ( in their paper). However, their predictions are for compute-optimal models , while we train for 5T tokens, which is orders of magnitude beyond what would be conventionally considered optimal for our model size. Thus, their predictions may not extend to our setup, and we stick with 64 experts for OLMoE-1B-7B, also due to the diminishing returns in Figure 5.
1.3 Shared Experts
Dai et al. propose training with a shared/fixed expert that is always used in addition to the routed experts. The intuition is to encourage the shared expert to learn common information and allow the other routed experts to learn more specialized knowledge. This should reduce redundancy among experts and thus lead to a better model as it can store more total information.
In Figure 6, we benchmark having a single shared and a single routed expert versus two routed experts. While both settings lead to similar performance, sharing an expert performs slightly worse. Sharing an expert removes flexibility from the model and thus goes against the findings in §4.1.2 suggesting that allowing for more expert combinations improves performance. Specifically, the two models in Figure 6 have and possible combinations per layer. Thus, removing one of the routed experts and turning it into a shared one eliminates almost 90% of possible combinations. This likely acts as a counterforce to the potential benefits of isolating common knowledge in a shared expert. Based on these results, we do not use shared experts in OLMoE-1B-7B, but we do think that there is merit to the idea of experts that are activated more often or even always. However, rather than enforcing this behavior via a shared expert, we believe that it should be learned by the model. This is difficult with current setups due to the necessity of a load balancing loss (§4.1.6) penalizing the model if tokens are not distributed equally among experts. Potential future work can explore removing the load balancing loss to allow for more flexible usage of experts.
1.4 Expert Choice vs. Token Choice
The MoE router determines which experts process each input token (§2). There are two common types : expert choice (EC) and token choice (TC) . For EC, each expert selects a fixed number of tokens from the incoming sequence. By design, this leads to each expert processing the same number of tokens. This is the main benefit of EC as it ensures perfect load balance, which improves training throughput and removes the need for a load balancing loss. The main downside of EC is that it is not easily usable for autoregressive generation where a single token is processed at each step rather than the entire sequence in one . Another potential downside is that EC can lead to token dropping, where some tokens are not selected by any expert, which can hurt performance . At the same time, it can lead to some tokens being processed by multiple experts, which could also be beneficial as it allows the model to allocate more compute to some tokens . For TC, each token selects a fixed number of experts. This can lead to many tokens choosing the same expert, hurting training efficiency. Therefore it is common to use TC with a load balancing loss to encourage equal distribution.
In Figure 7, we benchmark EC and TC. We find that TC outperforms EC for the same token budget for all tasks depicted as well as other tasks like PIQA, SciQ, etc. which we report at https://wandb.ai/ai2-llm/olmoe/reports/Plot-EC-vs-TC--Vmlldzo4MzkzMDM3. While Zhou et al. find EC to be better, our configuration slightly differs in that we use dropless MoEs with a load balancing loss. Thus, our TC variant is expected to perform better than the TC variant in Zhou et al. . We confirm findings that EC runs around 20% faster at 29,400 tokens per second per device versus 24,400 for TC . EC may be more beneficial in a multimodal setup as dropping noisy image tokens is likely less harmful than text tokens. Thus, while we stick with TC for this release of OLMoE, we may revisit EC for future multimodal models.
1.5 Sparse Upcycling
Komatsuzaki et al. propose turning a dense model into a Mixture-of-Experts model via sparse upcycling: (1) The dense MLP is cloned for each desired expert to constitute MoE layers. (2) A newly initialized router is added in front of each MoE layer. (3) Pretraining continues with the new model so that the cloned MLPs can gradually specialize in different things and the router can be learned. They find that the upcycling approach maintains a performance advantage over a language model trained from scratch for up to 120% of the compute budget of the original dense checkpoint that the sparse model was upcycled from. For example, if sparsely upcycling a 1.3B parameter model at 2 trillion tokens then only at 2.4 trillion tokens should an MoE trained from scratch catch up with the upcycled model. That is, the sparsely upcycled model would have been trained for another 400 billion tokens, thereby saving the equivalent of up to 2T tokens of compute. Other works such as MiniCPM , Qwen2 and reportedly Mixtral have adopted sparse upcycling but only share limited information about their configuration.
In Figure 8, we compare sparse upcycling OLMo-1B (0724) with training an MoE from scratch. We find that after 500B tokens, an otherwise equivalent MoE trained from scratch already catches up with the upcycled model, both on the metrics in Figure 8 and our additional metrics at https://wandb.ai/ai2-llm/olmoe/reports/Plot-Scratch-vs-Upcycle--Vmlldzo4NDIyOTc4. At around 600B tokens, the MoE from scratch starts outperforming the upcycled MoE. Thus, it only requires 25% of the compute budget of the original dense model to catch up as opposed to the 120% reported in Komatsuzaki et al. . However, they use expert choice routing and study encoder-decoder models . Meanwhile, we use token choice routing (§4.1.4) and decoder-only models (§2). Further, we upcycle a model that has already been significantly overtrained , i.e., a 1B model trained for 2T tokens. Its parameters are likely already in a very optimal range for a dense model, which may limit the amount of additional exploration possible after upcycling. This motivates us to experiment with adding noise to the upcycled weights outlined in Appendix F, but we do not find it to lead to better performance. A large disadvantage of upcycling is that the upcycled MoE is constrained by some hyperparameters of the dense model. Specifically, OLMo-1B (0724) was trained without QK-Norm and normal initialization, both of which hurt stability in our experiments (§4.2.5, §4.2.2). While it may be possible to simply add new QK-Norms and train them from scratch similar to the new router layer trained from scratch, it is impossible to change the initialization of the original dense model when upcycling it. Thus, as we want to change these hyperparameters and also train OLMoE-1B-7B for around 250% of the compute budget of the dense model (5T vs. 2T tokens), we do not use upcycling.
1.6 Load Balancing Loss
Shazeer et al. propose the load balancing loss to penalize the model if it is unbalanced, i.e., if it routes all tokens to only a few experts. This is based on the observation that without such penalty, models tend to update only a select few experts in each layer . To compute the load balancing loss () we multiply the fraction of tokens routed to one expert with the total routing probability allocated to for one batch and sum it across the number of experts :
The loss is further scaled by and a loss weight (see Equation 2), which is an optional weight to determine the magnitude of the loss commonly set to 0.01 . We do not experiment with changing the weight of 0.01.
In Figure 9 we investigate the performance impact of using the auxiliary load balancing loss. We find that across training loss and validation losses, using the load balancing loss leads to better performance even after only a few billion tokens. We still measure the load balancing loss even when it is not used (“No LBL”) and find that while it spikes initially, it slowly decreases over the next few billion tokens. This behavior is also visible in Figure 10 (left), where initially all tokens in the first layer are assigned to the 6th expert (pink). Eventually, the model also starts assigning some tokens to the 1st expert (yellow). However, all other experts remain largely flat and are thus “dead weights” that take up GPU memory but are not used. Given these results, we use the auxiliary load balancing loss with a weight of 0.01 following prior work . However, getting rid of the load balancing loss is an important direction for future research as it constrains the flexibility of the model by forcing it to use all experts approximately equally. This could prevent the experts from specializing in certain data domains and may be a reason prior work has failed to find strong evidence of expert specialization .
1.7 Router Z-loss
Zoph et al. propose the router z-loss to improve both the stability and quality of MoE models. This auxiliary loss penalizes large logits coming into the gating network. Such large logits can lead to numeric overflows in the large matrix multiplications happening in the MoE layer. It is computed by exponentiating the logits right before the router layer summed across the number of experts and averaged across the batch , thereby making larger logits lead to a larger loss:
The loss is further multiplied with an optional loss weight, (see Equation 2), to determine the magnitude of the loss commonly set to 0.001 . We do not experiment with changing the weight of 0.001.
In Figure 11, we confirm that across training loss, validation loss, and downstream performance adding the router z-loss improves stability (less spikes) and quality (lower loss and higher downstream performance). Thus, despite it reducing throughput by 2% we use the router z-loss for OLMoE-1B-7B with a weight of 0.001 as in Zoph et al. .
2 General Pretraining Settings
Li et al. release the DCLM-Baseline dataset and establish that it leads to better language models than Dolma 1.7 and other datasets as measured on common benchmarks like MMLU . This motivates us to mix their DCLM dataset with some components from Dolma 1.7 that we deem to be high-quality; see §2. In Figure 12, we compare our mix, OLMoE-Mix, with Dolma 1.7 in a controlled setup. We find that OLMoE-Mix leads to clear gains on all three downstream metrics, especially MMLU. DCLM-Baseline has been created through a series of dataset ablations targeting MMLU and other downstream metrics, which explains these results. We also compare adding Reddit and FLAN to our mix as detailed in Appendix F, but do not find consistent performance gains. We do not have a strong intuition for why adding these datasets does not help and a more automatic approach to dataset mixing may be desirable for future iterations . We pretrain using our mix of DCLM-Baseline and Dolma 1.7 dubbed OLMoE-Mix.
2.2 Initialization
Few prior works on Mixture-of-Experts share their initialization strategy. Even the most open MoEs prior to this work, JetMoE and OpenMoE , do not mention their initialization scheme. For DeepSeekMoE and DeepSeekV2 , the authors share that they use a normal initialization with a standard deviation (std) of 0.006. For dense language models, a normal initialization with an std of 0.02 has been commonly used as popularized by Shoeybi et al. .
In Figure 13, we find a truncated normal initialization leads to more stable training and better performance than a regular normal initialization. The difference between the two initializations only becomes clear at around 450 billion tokens, where the model with the normal initialization starts to diverge. This is despite both models using the same configuration except for the difference in weight initialization. Having to train for hundreds of billions of tokens until an experiment provides a clear signal is one of the key challenges of pretraining ablations. We use the truncated normal initialization for OLMoE-1B-7B.
2.3 RMSNorm
OLMo uses non-parametric layer normalization , mainly as it is significantly faster than the commonly used RMSNorm . This is an unusual choice as most LMs use RMSNorm, such as the Llama , Gemma , and Qwen model families.
In Figure 14, we observe that replacing the non-parametric layer normalization in OLMo with a parametric RMSNorm leads to better performance. This is likely because the non-parametric layer normalization leads to a large number of spikes in the gradients as seen in Figure 16. We clip gradients at 1.0, which prevents these spikes from leading to very large and potentially disruptive parameter updates. However, the clipped gradients may still harm the performance of the model as they are no longer the true gradients. Thus, despite RMSNorm lowering our training throughput by 15%, we train our final model with RMSNorm. We include the RMSNorm parameters in weight decay as we find that it performs slightly better (Figure 15) even though it is common practice to exclude them.https://github.com/karpathy/minGPT/pull/24#issuecomment-679316025
2.4 Decaying Embedding Parameters
Similar to the RMSNorm parameters (§4.2.3), embedding parameters are commonly excluded from weight decay.https://github.com/karpathy/minGPT/pull/24#issuecomment-679316025 In Figure 17 we find that whether or not they are decayed has only a minor impact on performance, with decaying being slightly better. Thus for simplicity, we weight decay all parameters in OLMoE-1B-7B including embedding and RMSNorm.
2.5 QK-Norm
Some works have reported stability improvements from adding layer normalization after the query and key projections (“QK-Norm”) . QK-Norm can prevent the subsequent attention operation from leading to very large logits that may lead to numeric overflows and destabilize the network, especially when training in low precision. Like layer normalization at other places in the model, the QK-Norm could be non-parametric or use the parametric RMSNorm (§4.2.3).
In Figure 18, we compare using QK-Norm with no normalization after the query and key projections. We find that QK-Norm leads to some stability and performance improvements. We perform this experiment with non-parametric layer normalization as used in OLMo , while we used parametric RMS layer normalization for OLMoE-1B-7B (§4.2.3). To ensure the benefit of QK-Norm is not an artifact of comparing with non-parametric layer normalization, we run another experiment with RMS layer normalization and still find QK-Norm to lead to slightly better training loss and to prevent a large grad norm spike.https://wandb.ai/ai2-llm/olmoe/reports/Plot-QKNorm-revisited--Vmlldzo4NTc2NTIz Thus, we use QK-Norm for OLMoE-1B-7B despite it reducing throughput by almost 10%.
2.6 AdamW Epsilon
Groeneveld et al. use an epsilon (“eps”) value of 1E-05 in the AdamW optimizer for training OLMo. A larger eps value leads to smaller steps of the optimizer but can be more stable .
In Figure 19, we find that decreasing eps to the recommended default of 1E-08 significantly improves performance while the run remains stable. Thus, we set eps to 1E-08 for our final run.
3 Adaptation Settings
We experiment with small design choices for adaptation using our evaluation setup described in Appendix C. (1) Auxiliary losses: Zoph et al. find that using the auxiliary load balancing loss (§4.1.6) during regular finetuning leads to small performance gains. For instruction tuning, however, Shen et al. do not find conclusive evidence in favor of using the load balancing or router z-loss with only small differences in performance, both in support of and against the auxiliary losses. In Table 7 we display experiments with the load balancing loss during adaptation and find that not using it leads to better performance (54.0 vs. 52.8 after instruction tuning (SFT) and 57.7 vs. 57.1 after preference tuning (DPO)). One potential problem of deactivating the load balancing loss is that it may harm balance among experts and turn some into dead weights as observed during pretraining in §4.1.6. However, when measuring the load balancing loss in Table 6 on our SFT data (§2), we find that the loss actually decreases slightly during SFT (12.16 vs. 12.22). This is likely because which experts certain tokens get routed to is determined early during pretraining, as we find later in the analysis section (§5.1). We also visualize the activation patterns of experts of the model after pretraining, and the models after SFT and DPO trained without load balancing in Appendix G (Figure 33) finding that the distribution remains around the same. Thus, as our models adapted without load balancing perform better and we find it not to impact routing substantially, we do not use load balancing during adaptation. (2) Annealing checkpoint: We also experiment with using the checkpoint pre-annealing (§2) for adaptation and find the checkpoint post-annealing leads to better performance (53.8 vs. 54.0 after SFT and 56.3 vs 57.7 after DPO), thus we use the post-annealing checkpoint. (3) Preference algorithm: Since the release of DPO (Direct Preference Optimization) , a variety of preference algorithms have been proposed . We experiment with KTO and find that it matches DPO in Table 7 for our setup (Appendix B). While we release both models, we use DPO for our final OLMoE-1B-7B-Instruct model, as it scores higher on AlpacaEval, which has a smaller chance of data contamination than our other benchmarks .
MoE Analysis
By advancing open and cost-efficient models (§1), OLMoE-1B-7B enables new research into LMs and MoEs. Making use of our released intermediate checkpoints, data, and code, we define and analyze four properties specific to MoEs: Router saturation (§5.1), Expert co-activation (§5.2), Domain specialization (§5.3), and Vocabulary specialization (§5.4).
We define router saturation as the proportion of expert activations at some intermediary checkpoint at time that matches the expert IDs activated at some final checkpoint over the same dataset:
: The total number of tokens in the dataset.
: The number of top- experts activated per input token. While we train with (§2), we also analyze by only looking at the expert with the highest routing probability.
: The set of experts activated for the th token at the th checkpoint.
: The set of experts activated for the th token at the final checkpoint .
: The number of common experts activated for the th token between the th and final checkpoints.
Router saturation thus corresponds to whether the router weights are still learning which expert will process certain data. A value of 100% indicates that the router at the intermediate checkpoint will route to the same experts as the final checkpoint router. However, even at 100% saturation the router weight can still change and adapt the exact router probability for each expert. These probabilities are used to scale the output of the respective expert in the model. For OLMoE-1B-7B with its 64 experts, random routing equals a saturation of for and for .
In Figure 20 we find that after 1% of pretraining (5000 steps or 20B tokens), up to 60% of routing to the top-8 activated experts has already saturated (right). Thus the model already uses the same 8 experts for given input data as it will at the end of pretraining. This early saturation aligns with prior work . At 40% of pretraining, saturation reaches up to 80%. However, which top-1 expert has the highest routing probability saturates slower (left). We find that routing in later layers saturates earlier during pretraining. Layer 0 is an outlier saturating significantly more slowly than other layers. Dai et al. do not use an MoE in the first layer as they find that load balancing converges more slowly for the first layer. This is likely linked to our findings on saturation. Because routing in the first layer saturates slower, the experts that certain input data get routed to frequently change. These changes may lead to one expert suddenly getting significantly more data than others thereby impairing load balancing. We are excited about future work further investigating what happens in the first layer by building on our open release.
2 Expert Co-activation
We define expert co-activation as the proportion of times two specific experts, and , are simultaneously activated out of the total number of activations of one of those experts:
: The number of times experts and are activated together.
: The total number of times expert is activated.
A co-activation of 100% indicates that if is activated, is also always activated. A value of 0% indicates that the experts never co-occur. If multiple expert pairs have high co-activation, it may suggest that these experts could be merged, benefiting less from keeping them separate. In a distributed setup, we could place highly co-activated experts on the same device to reduce communication costs during model inference.
In Figure 21, we find that there is no strong co-activation among experts in one layer, with only few exceptions. This may indicate that there is little redundancy across different experts. Overall, layers 7 and 15 show similar co-activation patterns with several groups of 3 or 2 experts that tend to get activated together. We investigate tokens that activate these experts in §5.4. Further, in Appendix G (Figure 35), we investigate whether experts across layers, rather than within one layer, tend to process tokens together.
3 Domain Specialization
We define domain specialization as the proportion of tokens from a particular domain that get routed to a particular expert :
: The domain from which the data originates.
: The number of experts considered (e.g., means considering the top 8 experts with the highest routing probabilities).
: The number of tokens from domain for which is among the top- selected experts.
: The total number of tokens from domain processed by the MoE.
Domain specialization thus refers to the specialization of expert to domain . A value of 100% indicates that all data from that domain is routed to , whereas 0% indicates the expert is never used for that domain and can be removed from the model without affecting performance in that domain.
In Figure 22 (top) we find many examples of experts that are activated significantly above or below random chance for specific domains. E.g., for arXiv, which has a very specific distribution with lots of scientific text, the first expert in layer 0 is nearly 100% specialized. This suggests that there is little redundancy in the knowledge of the experts in OLMoE-1B-7B, as they specialize in different kinds of data. GitHub and arXiv are often activated together in layer 7, which we explore further in §5.4. For generic domains, such as C4 , which is a web crawl containing various kinds of data, expert activations in OLMoE-1B-7B are much more balanced. This highlights that the load balancing (§4.1.6) works as intended and the model makes proper use of all experts for generic data. Mixtral-8x7B in Figure 22 (bottom), however, exhibits little domain specialization across both unique and generic domains. Experts are activated close to the uniform routing baseline for all layers and domains. Thus, there may be more redundancy across experts in Mixtral, as they likely contain similar knowledge. We hypothesize that this is due to Mixtral being upcycled from Mistral . The initialization from a dense model may limit the amount of possible specialization in the experts as they all start from the same local optimum. This is likely why training from scratch eventually outperforms upcycling in our pretraining experiments (§4.1.5).
4 Vocabulary Specialization
We define vocabulary specialization as the proportion of tokens with a token ID (also called vocabulary element) that are routed to one particular expert out of all experts in that layer:
: The number of experts considered (e.g., means considering the top 8 experts with the highest routing probabilities).
: The number of times input data is routed to for .
: The total number of times input data is routed across all experts for .
Vocabulary specialization thus refers to how specialized a particular expert is on some vocabulary item. We distinguish input and output variants of this specialization, where is either the input token ID or the next output token ID (either the ground-truth next token ID or the token ID predicted by the model). A value of 100% indicates that for all occurrences of that vocabulary element, input data is routed to , whereas 0% indicates an expert that is fully irrelevant for that vocabulary element and can be effectively removed from the model without affecting performance whenever the token ID appears.
In Figure 23 we find that vocabulary specialization is higher in later layers, similar to how later layers saturate earlier (§5.1). Later layers also specialize more on predicted output token IDs rather than input token IDs, i.e., the routing is decided more by the token the model is about to predict rather than the original input token. This is intuitive as in earlier layers there is more uncertainty about which token the model will predict. At 90%, expert 27 specializes the most, which we find in Table 8 to activate for many non-alphabetic tokens, such as Cyrillic and Devanagari letters. Expert 43 shows specialization on geographic terms in both input and output tokens. Experts 48 and 23 both focus on connector words, such as Then and Therefore. This is likely because they commonly process tokens together with a high co-activation of 60% in Figure 21 (middle). Based on our findings in §5.3 that for GitHub and arXiv often the same experts in layer 7 activate, we display one such expert (expert ID 4) in Table 8. It seems to specialize in measurements, such as sq, YR (year), and GHz. These are common terms in scientific papers corresponding to the arXiv domain and likely also in GitHub code for computations related to measurements. They are less likely to appear in books, which explains the low activation of expert ID 4 in layer 7 for book data in Figure 22. Expert 3 is among the three most active experts of layer 7 for book data in Figure 22 (fourth yellow bar for layer 7). This resonates when looking at its specialization on family terms in Table 8, which are far more common in books than scientific papers or code. Overall, domain specialization and vocabulary specialization are closely linked to one another, as domains are usually characterized by their distinct word distribution. In Appendix G (Figure 32), we link them more closely by comparing the extent of vocabulary specialization across domains and expert IDs. In Appendix G (Figure 30, Figure 31) we also find that OLMoE-1B-7B exhibits stronger vocabulary specialization than Mixtral-8x7B.
Related Work
Current LMs still largely follow the transformer architecture with only few architectural changes that have been widely adopted, such as decoder-only training , SwiGLU activations , RoPE , MQA/GQA and RMSNorm . Model sparsity via Mixture-of-Experts is one modification still under active exploration with some early adoption but most LMs, including Llama 3 , still rely on a dense architecture. There has been a lot of progress in improving the sparsely-gated MoE layer since its introduction : New routing techniques , fine-grained expert segmentation , stability and efficiency improvements. In this work, we perform many experiments to provide insights into training Mixture-of-Experts LMs. Subsequently, we train OLMoE-1B-7B for 5T tokens. No prior MoE has been overtrained to this extent to our knowledge making OLMoE-1B-7B the best testbed to research performance saturation of MoEs vs. dense models. With OLMoE we hope to facilitate such and other research to help the field uncover whether MoEs should make it into all future LMs and with what precise configuration.
A variety of model families have been proposed under varying degrees of openness commonly categorized based on whether model weights are available. Closed-weight models include GPT , Gemini , PaLM , Reka , and open-weight ones include Llama , Mistral , Gemma , Falcon , MPT , Qwen , GLM , Yi , DeepSeek , Nemotron , InternLM , Baichuan , Phi , StableLM , OPT . However, besides model weights, training data and code are key to enabling scientific research of these models and distributing their benefits broadly . There have been few releases also including data and code in addition to model weights which we refer to as “fully open-source”: BLOOM , GPT-NeoX , StarCoder , Pythia , OLMo , LLM360 , Cerebras-GPT , DCLM , MAP-Neo , RWKV , and SmolLM . For Mixture-of-Experts only OpenMoE aims to be fully open-source, however, its poor performance limits its usefulness. We release OLMoE-1B-7B as the first state-of-the-art Mixture-of-Experts LM that is fully open-source: model weights, data, code, and logs.
Conclusion
We open-source OLMoE-1B-7B and OLMoE-1B-7B-Instruct including model, data, code, and logs. At 1B active and 7B total parameters, our models yield state-of-the-art performance among models with a similar amount of active parameters even outperforming larger models including DeepSeekMoE-16B and Llama2-13B-Chat. We share various training experiments and define and analyze router saturation, expert co-activation, domain and vocabulary specialization of our model. Through our fully open release, we seek to help the field build better MoEs. We are excited about new iterations of OLMoE to close the gap between frontier models and fully open models.
Author Contributions
Niklas Muennighoff proposed and led the project. He ran the pretraining experiments, pretrained the model, helped run adaptation and analysis, and wrote most of the paper. Luca Soldaini created the pretraining dataset and advised on pretraining. Dirk Groeneveld advised on pretraining, especially stability and throughput improvements. Kyle Lo helped with pretraining dataset creation, analyzed data experiments, and advised on data and framing, and helped edit the paper. Jacob Morrison co-created the adaptation dataset, ran most adaptation experiments, and helped edit the paper. Sewon Min analyzed router saturation, expert correlation, and vocabulary specialization, and helped frame and edit the paper. Weijia Shi analyzed domain and vocabulary specialization, advised at various project stages, and helped edit the paper. Pete Walsh advised on pretraining, especially stability and throughput improvements. Oyvind Tafjord ran OLMES evaluations. Nathan Lambert co-created the adaptation dataset, advised on adaptation, and helped edit the paper. Yuling Gu ran OLMES evaluations and helped edit the paper. Shane Arora uploaded the models and helped with code review. Akshita Bhagia supported stability investigations and helped with DCLM evaluations. Dustin Schwenk supported stability investigations. David Wadden ran DCLM evaluations and helped with Weights & Biases reports. Alexander Wettig analyzed load balancing, routing, and domain specialization, and helped edit the paper. Binyuan Hui advised on pretraining. Tim Dettmers advised on analysis and inference experiments. Douwe Kiela advised on framing. Ali Farhadi advised on pretraining and framing. Noah A. Smith advised on pretraining, and helped frame and edit the paper. Pang Wei Koh advised on analysis, and helped frame and edit the paper. Amanpreet Singh advised on pretraining, framing and helped edit the paper. Hannaneh Hajishirzi was responsible for direction and advising of the overall effort and helped frame and edit the paper.
Acknowledgements
OLMoE would not be possible without the support of many individuals and institutions. We thank our teammates at the Allen Institute for AI, Contextual AI, and the University of Washington for their support, especially Aditya Kusupati, Ananya Harsh Jha, Caitlin Wittlif, Carissa Schoenick, Costa Huang, Crystal Nam, David Atkinson, Emma Strubell, Faeze Brahman, Hamish Ivison, Karel D’Oosterlinck, Matt Latzke, Ian Magnusson, Jack Merullo, Jay Chen, Jennifer Dumas, Jiacheng Liu, Johann Dahm, Luke Zettlemoyer, Michael Schmitz, Michael Wilson, Pradeep Dasigi, Sahil Verma, Sam Skjonsberg, Sophie Lebrecht, Stas Bekman, Taira Anderson, Valentina Pyatkin, Yanai Elazar, Yizhong Wang, and Yoganand Chandrasekhar. We also thank Armen Aghajanyan, Akshat Shrivastava, Colin Raffel, Haokun Liu, Ludwig Schmidt, and Shayne Longpre. PWK is supported by the Singapore National Research Foundation and the National AI Group in the Singapore Ministry of Digital Development and Innovation under the AI Visiting Professorship Programme (award number AIVP-2024-001).
References
Appendix A Artifacts
Appendix B Training Configuration
We display the pretraining hyperparameter configuration of OLMoE-1B-7B in Appendix B comparing with other relevant models. We follow Groeneveld et al. using the AdamW optimizer with ZeRO via PyTorch FSDP and mixed-precision training . Our main model settings differing from Groeneveld et al. are: (1) MoE-related changes: OLMoE-1B-7B is a sparsely activated decoder-only transformer using dropless Mixture-of-Experts . Unlike most prior MoEs, we use a high granularity with 64 small experts with an FFN dimension of just 1,024 rather than a few large experts. We further use two auxiliary losses: router z-loss and load balancing loss . (2) Stability improvements: (a) We use a truncated normal initialization with a standard deviation of 0.02 and a minimum (maximum) cut-off of -0.06 (0.06) corresponding to three standard deviations. (b) We use QK normalization . (c) We use RMSNorm instead of the non-parametric LayerNorm used in Groeneveld et al. . (3) Performance improvements: Besides some of the stability improvements which also impact performance, we also reduce the AdamW epsilon to 1.0E-08 from the 1.0E-05 used in Groeneveld et al. to speed up convergence. Finally, we train OLMoE-1B-7B for significantly longer than all prior OLMo models amounting to 5T tokens and thus more than one epoch (1.2) following Muennighoff et al. . We shuffle the pretraining dataset before starting the second epoch. For the final 100B tokens, we decay the learning rate linearly from 5.0E-04 to 0 (“annealing”). We experiment with many of these settings in §4.
For finetuning we use Open Instruct .Code: https://github.com/allenai/open-instruct We filter all SFT samples to a length of fewer than 4096 tokens to match the sequence length of the model. Following Muennighoff et al. , we aggregate loss at the token level during SFT to improve performance on long generative tasks, such as AlpacaEval. We finetune in BF16 with a global batch size of 128 (4 H100 nodes with 8 GPUs each, a per device batch size of 2, and 2 gradient accumulation steps). We train for 2 epochs with a constant learning rate of 2.0E-5. For DPO , we reduce the global batch size to 32 (4 H100 nodes with 8 GPUs each and a per device batch size of 1). We train for 3 epochs with a learning rate of 5.0E-7 and a DPO beta of 0.1. Our adapted models are built on top of our annealed checkpoint, and we include the load balancing loss during both SFT and DPO based on our experiments in §4.3. Our preference tuning recipe is heavily optimized for DPO based on extensive experiments by Ivison et al. , thus for KTO we experiment with a few settings in Appendix F. Our final KTO adaptation uses the same hyperparameters as DPO, except that we use the RMSProp optimizer instead of Adam, which we use for SFT and DPO, and that we reduce the training duration to 1.3 epochs (5,000 steps) for KTO instead of the 3 epochs used for DPO.
We pretrain OLMoE-1B-7B on 256 H100 GPUs for approximately 10 days with NV-link interconnect across GPUs and InfiniBand interconnect across nodes. We also use H100 GPUs for all our experiments but some use a cluster with GCP TCPx interconnect across nodes instead. For adaptation, we use 32 H100 GPUs for 33 hours to instruction tune and for another 14 hours to preference tune via DPO. For KTO adaptation we use 8 H100 GPUs for 30 hours instead.
Appendix C Evaluation Setup
We evaluate using a similar in-loop evaluation setup as Groeneveld et al. , with the addition of more tasks such as CommonsenseQA, PIQA, and different implementations of MMLU. Following Groeneveld et al. , for the majority of the tasks, we perform 0-shot evaluation using the Completion/Cloze formulation (CF), ranking each answer string using language model probabilities. In terms of probability normalization, there is either no normalization (none) or normalization by the number of tokens in the answer (token) when ranking solely based on probability may heavily favor shorter answers . For MMLU, the in-loop evaluation also includes a setup where we increase the total number of instances by including a range of 0-shot to 5-shot setups together as we found this provides smoother trends as the training proceeds (“MMLU Var”). We also include the Multiple-choice formulation (MCF) version of MMLU, scoring prediction of answer labels like A/B/C/D, which generally starts to rise only later in training as models only gain the multiple-choice capability later (at around 1T tokens for OLMoE-1B-7B in Figure 25). We also evaluate perplexity on selected validation sets from Paloma . All code used for evaluation during pretraining is at https://github.com/allenai/OLMo/tree/61ac104d616ec5435db225796e5c7532c9abd95a/olmo/eval.
We perform evaluations following the OLMES evaluation standard , with the suite of tasks in the original paper. OLMES (Open Language Model Evaluation Standard) is a standard for reproducible LM evaluations that is open, practical, and documented, providing recommendations guided by experiments and results from the literature . It is designed to support comparisons between smaller base models that require the Cloze formulation of multiple-choice questions against larger models that can utilize the Multiple-choice formulation. To make our evaluations reproducible, we follow OLMES in prompt formatting, choice of in-context examples, probability normalization, task formulation, as well as all other details. We summarize this setup in Table 4 and refer to Gu et al. for more details.
For results on the DCLM tasks in Table 13, we precisely follow their setup using the evaluation code released by the authors at https://github.com/mlfoundations/dclm. “Core” results are the low variance tasks in their evaluation code, while “Extended” corresponds to the heavy tasks.
After supervised finetuning and direct preference optimization, we evaluate models using a subset of the evaluations and the same overall setup used in Ivison et al. and Wang et al. . We cover a wide range of model capabilities in our evaluation suite including coding (HumanEval ), general and mathematical reasoning (Big Bench Hard , GSM8k ), world knowledge (MMLU), general instruction following (AlpacaEval 1.0 , not the length-controlled variant ), precise instruction following (IFEval ) and safety (XSTest ). We refer to Wang et al. for more details on each benchmark.
Appendix D Openness of Models
We list the openness of various models summarized in Figure 1. We exclude Switch Transformers , as it was published over three years ago and is very different from more recent MoE models (MLM objective, Encoder-decoder, etc.).
Model: Their model is licensed under the open-source Apache 2.0 license.
Model: Their model is licensed under the open-source Apache 2.0 license.
Model: The model is licensed under a custom non-open-source licensehttps://www.databricks.com/legal/open-model-license with additional use-case restrictions.https://www.databricks.com/legal/acceptable-use-policy-open-model
Code: They use closed-source custom adaptations of their public libraries LLM-foundry, composer, and megablocks.https://github.com/databricks/dbrx
Model: The model is licensed under a custom non-open-source license.https://github.com/SkyworkAI/Skywork/blob/main/Skywork%20Community%20License.pdf
Model: The models are licensed under custom non-open-source licenses.https://github.com/deepseek-ai/DeepSeek-MoE/blob/main/LICENSE-MODEL and https://github.com/deepseek-ai/DeepSeek-V2/blob/main/LICENSE-MODEL
Model: The model is licensed under the open-source Apache 2.0 license.
Data: They describe their mixture but do not release it.https://medium.com/snowflake/snowflake-arctic-cookbook-series-arctics-approach-to-data-b81a8a0958bd
Model: The model is licensed under the open-source Apache 2.0 license.
Model: The model is licensed under the open-source Apache 2.0 license.
Model: The model is licensed under a custom non-open-source license.https://hf.co/Qwen/Qwen1.5-MoE-A2.7B/blob/main/LICENSE
Model: The model is licensed under the open-source Apache 2.0 license.
Data: They describe their mixture but do not release it.
Code: They make their fork of megablocks publicly available,https://github.com/yikangshen/megablocks however, their Megatron-LM training code is not available.https://hf.co/jetmoe/jetmoe-8b/discussions/5#661ee52c03251697a0b155cc
Model: The model is licensed under the open-source Apache 2.0 license.
Data: They make scripts for recreating their data available.
Code: They make their code available.https://github.com/XueFuzhao/OpenMoE/tree/main?tab=readme-ov-file#training-with-tpugpu
Model: The model is licensed under the open-source Apache 2.0 license.
Data: The data is licensed under the open-source ODC-By 1.0 license.
Code: The code is licensed under the open-source Apache 2.0 license.
Logs: Logs are available with the same open-source license as the code (Apache 2.0).
Appendix E Additional Evaluation
Appendix F Additional Experiments
In Figure 26 we benchmark adding the Reddit or FLAN subsets of Dolma 1.7 to our pretraining data mix (§2). Overall, we do not find either one to lead to consistent gains, thus we do not use them in our final data mix.
Fedus et al. selectively perform operations related to routing in full precision (FP32) to improve stability. In Figure 27, we test whether computing the load balancing loss in full precision improves stability, but do not find it to reduce spikes. Thus, we stick with bfloat16 (BF16).
For the creation of Qwen2-MoE , the authors add 50% of gaussian noise to feedforward networks before continuing training in an upcycled setup . Komatsuzaki et al. also report that they experimented with adding noise but did not find it beneficial. In Figure 28, we experiment with regular upcycling versus adding noise by randomly replacing 50% of each MLP with numbers drawn from a normal distribution with a standard deviation of 0.02 following. We find that after 700 billion tokens, the no noise variant still performs slightly better but both appear to converge to the same performance. If training further, it is possible that the noise variant eventually outperforms the no noise variant, but at that point, it may make more sense to just train the MoE from scratch (§4.1.5).
Some work has investigated Mixture-of-Experts with weights shared across layers in the context of Universal Transformers . We test whether layer-shared Mixture-of-Experts can beat non-shared dense models in Figure 29. The layer-shared MoE uses a load balancing loss that is applied at the model level rather than at the layer level. This gives the model more flexibility by allowing it to completely deactivate certain experts for some layers and even emulate a dense model by always activating one separate expert for each layer. This makes it a generalization of the dense model which motivated our hypothesis that it may perform better than the dense model. However, in practice, we find that both perform similarly with the regular dense models even maintaining a small advantage on validation loss and HellaSwag. One possible advantage of layer-shared MoEs is that they can allow for better load balancing at inference. If prompts come in continuously, then newly incoming prompts can be batched with previous prompts that have already passed through several layers and sent through the MoE module together, as the MoE module is the same regardless of whether it is the first or last layer. Sharing also reduces throughput by around 20% during training, which further motivates our decision not to use it for OLMoE-1B-7B.
In Table 14 we experiment with the number of steps (5,000 vs. 10,000) and the optimizer (Adam vs. RMS) used for KTO . Based on these experiments we use the RMS optimizer and the checkpoint at 5,000 steps in §4.3.
Appendix G Additional Analysis
Appendix H Limitations and Future Work
We highlight four key limitations with this release of OLMoE-1B-7B. We look forward to addressing these issues in future iterations of OLMoE.
OLMoE-1B-7B has 7B total parameters out of which 1B are activated for each input token. This small size makes OLMoE-1B-7B very cheap to use, yet we demonstrate in this work that it outperforms much more expensive models (Figure 1). However, using only 1B parameters for each input token also limits the capabilities of OLMoE-1B-7B as seen by its performance compared to models that use 7 more parameters, such as Llama3.1-8B in §3. While it may be possible that more parameters are not needed to match 8B models and beyond , in the short-term adding parameters is an easy way to improve the performance of OLMoE, at least allowing the model to utilize more than 1B parameters per input, possibly via recursion or agentic workflows . Relatedly, changing the allocation of parameters to e.g. vocabulary versus non-vocabulary parameters is another avenue for improvement .
We train OLMoE-1B-7B for 5 trillion tokens, however, some recent dense models train significantly longer, such as Llama 3 with 15 trillion tokens . To the best of our knowledge, there has been no large MoE that has been overtrained as much as OLMoE-1B-7B. Specifically, taking the active parameters of OLMoE-1B-7B, our token multiplier is around 5,000 (5T / 1B). There are likely benefits to training even longer, but to what degree overtraining is effective for MoEs and how it differs from dense models still requires more research .
OLMoE-1B-7B is a text-only large language model, thus it cannot take inputs or produce outputs in other modalities like images or audio. This limits its utility for the large variety of multimodal use cases of such models . There has been early work on open multimodal MoEs and we look forward to making future versions of OLMoE a part of that.
We pretrain OLMoE-1B-7B on a predominantly English corpus and exclusively evaluate on English tasks. This may severely limit the usefulness of our model for research on non-English language models . While there has been work on training language-specific LMs , it is more likely that as we add more data to build better future iterations of OLMoE we will mix in more non-English data due to data constraints . This may make future OLMoE models perform better in non-English languages.