Scaling Laws for Fine-Grained Mixture of Experts
Jakub Krajewski, Jan Ludziejewski, Kamil Adamczewski, Maciej Pióro, Michał Krutul, Szymon Antoniak, Kamil Ciebiera, Krystian Król, Tomasz Odrzygóźdź, Piotr Sankowski, Marek Cygan, Sebastian Jaszczur
Introduction
In recent years, we have witnessed Large Language Models (LLMs) achieve exceptional performance in tasks across numerous domains (Chowdhery et al., 2022; Yin et al., 2023; Agostinelli et al., 2023). However, training those massive models incurs high computational costs, measured in millions of GPU-hours (Touvron et al., 2023b), enabled only by enormous budgets (Scao et al., 2023) and leading to non-negligible carbon footprints (Faiz et al., 2024). To combat these obstacles, the research community has been striving to increase the efficiency of LLMs. One promising approach that has lately been gaining visibility is the use of Mixture of Experts (MoE) methods. Models such as Switch (Fedus et al., 2022) and Mixtral (Jiang et al., 2024) have already demonstrated that it is possible to achieve comparable effectiveness with significantly lower computational costs.
In the context of the current trend of increasing budgets for training language models, a question arises: will MoE models continue to be attractive in the future? This is an important issue, as other studies have stated that the gap in efficiency between MoE and standard Transformers narrows at scale (Artetxe et al., 2022) or even that traditional dense models may outperform MoE as the size of the models increases (Clark et al., 2022).
In this paper, we argue that previous claims lose their validity when we relax certain implicit assumptions regarding the training process, present in previous research. In particular, we refer to the fixed training duration and the constant size of experts in MoE models.
Our results suggest that a compute-optimal MoE model trained with a budget of FLOPs will achieve the same quality as a dense Transformer trained with a greater computing budget, with the compute savings rising steadily, exceeding when budget of FLOPs is surpassed (see Figure 1). Importantly, we show that the standard practice of fixing the size of experts in MoE to be the same as feed-forward layer is almost never optimal.
Introducing a new hyperparameter - granularity. Adjusting this parameter allows us to determine the optimal size of experts in MoE models, which translates into increased efficiency.
Deriving new scaling laws for MoE models that incorporate variable training duration, the number of parameters, and granularity. Such scaling laws allow us to calculate optimal training hyperparameters for MoE models.
Demonstrating that, with optimal settings, MoE models can always outperform traditional Transformers at any computing budget. This is a conclusion contrary to the results from Clark et al. (2022).
The code used to produce the results described in this work is open-sourced at github.com/llm-random/llm-random.
Related Work
Mixture of Experts. In the context of language modeling, MoE was first introduced by Shazeer et al. (2017) as a sparsely gated layer between stacked blocks of LSTM (Hochreiter & Schmidhuber, 1997). A similar technique was proposed in the context of Transformers by Shazeer et al. (2018) and Lepikhin et al. (2020). Fedus et al. (2022) proposed to route each input to only a single expert and designed a modified initialization scheme to reduce training instability. Numerous studies have proposed to modify the original routing method. Lewis et al. (2021) used a linear assignment algorithm to postprocess token-expert mappings and ensure even expert selections. Roller et al. (2021) suggested another approach involving deterministic hash functions. Zhou et al. (2022) proposed expert choice routing, eliminating the need for additional load balancing losses. Puigcerver et al. (2023) designed a fully-differentiable Soft MoE architecture.
Concurrently to our work, Dai et al. (2024) proposed to modify the MoE layer by segmenting experts into smaller ones and adding shared experts to the architecture. Independently, Liu et al. (2023) suggested a unified view of sparse feed-forward layers, considering, in particular, varying the size of memory blocks. Both approaches can be interpreted as modifying granularity. However, we offer a comprehensive comparison of the relationship between training hyperparameters and derive principled selection criteria, which they lack.
Scaling laws. Scaling laws are empirically derived equations relating the loss of a model with variables such as the number of parameters, training samples, or the computational budget. In the case of dense Transformers, scaling laws were first studied by Kaplan et al. (2020), who observed power law relationships between the final model perplexity and model and dataset size. This work was extended by Hoffmann et al. (2022) by considering variable cosine cycle lengths and formulating a modified functional form of the scaling equation.
Scaling laws have also been proposed for other architectures and training scenarios. Henighan et al. (2020) studied autoregressive modeling across various modalities, while Ghorbani et al. (2021) considered machine translation. Frantar et al. (2023) explored the impact of pruning on vision and language Transformers, deriving optimal sparsity for a given compute budget. Clark et al. (2022) studied the scaling of MoE when changing model size and number of experts on a fixed dataset, concluding that routed models are more efficient only until a certain model size. In this work, we challenge that claim by considering a variable, optimal dataset size for both model families (see Section 6.3).
Background
A standard decoder-only Transformer (Radford et al., 2018a; b; Kaplan et al., 2020; Brown et al., 2020) consists of an embedding layer, a stack of alternating attention and feed-forward layers, and an unembedding layer. In the model, each input token is converted by the embedding layer into a vector of size , the dimension maintained across all the layers in the residual stream.
The feed-forward component consists of two linear transformations and a nonlinearity in between. It can be described as , with mapping from to , and back to the original . It is standard (Radford et al., 2018a; Rae et al., 2022; Touvron et al., 2023a; Jiang et al., 2023) to set the hidden dimension as .
Feed-forward layers contain the majority of Transformer parameters and require the biggest computational budget counted in terms of FLOPs. Subsequently, they are the main focus of the Mixture of Experts models considered in this work.
Mixture of Experts.
The core idea behind MoE in Transformers is to replace the feed-forward layer with a set of experts. The size of each expert is typically (Fedus et al., 2022; Zhou et al., 2022; 2023; Jiang et al., 2024) set to mirror the original dimensions of the layer, with the hidden expert dimension equal to Therefore, the total number of parameters in MoE scales linearly with the number of experts. However, the computational cost remains approximately constant as each input is routed and then processed by a subset of experts.
2 Scaling Laws
Dense Transformers. Large Transformer-based models are known to approximately obey the power-law relationship between final loss , model size and number of training tokens This relationship is often called Chinchilla scaling laws described by Hoffmann et al. (2022) as
The power-law formula is composed of three distinct terms that characterize the intrinsic entropy of data, constraints of the model, and limitations in the training data. The term represents the minimum possible error intrinsic to the data. The remaining two terms are suboptimality terms, which address the limitations in function representation owing to the size of the model and in data signified by the number of tokens. In the limit, with infinite data and model size, the loss is reduced to .
Mixture of Experts. For MoE Transformer-based models, Clark et al. (2022) formulated the final loss for a constant dataset size of 130B tokens, allowing for variations in the expansion rate , as:
However, this result has a notable limitation as it can be applied only to the original dataset size. The scalability and effectiveness are constrained in this scenario because it is crucial to align the number of training samples with the available computational resources for optimal use. As per Kaplan et al. (2020) and Hoffmann et al. (2022), maintaining a constant dataset size while scaling up the neural network size leads to undertraining, resulting in a model that does not perform to its full potential.
Granularity
As described in Section 3, in the standard setting, the inner dimension of each expert network, , is equal to , which is the same size as the feed-forward layer of the base model.
In this work, we suggest an alternative approach where the hidden dimension of the expert is not necessarily set to mirror that of the standard feed-forward layer. Instead, it can be adjusted to a value that is the most effective. This approach allows the configuration of MoE to be articulated in terms of two key hyperparameters: granularity () and expansion rate (). In the following parts of this work, we will also use the term active parameters to refer to the non-embedding parameters used to produce output for a single token, except routing. The number of active parameters is denoted as .
Let be the hidden dimension of a single expert. Granularity is defined as
In other words, granularity denotes the multiplier factor for the change in the size of an expert from the original standard model, defined as . In this work, we investigate where experts are smaller than in the standard layer.
Note that increasing granularity does not affect the number of active parameters. As increases, the number of experts that process the token grows proportionally to . In other words, for granularity , a token is routed to fine-grained experts, thereby keeping the number of active parameters constant. See Fig. 2 for visualization.
We then define the expansion rate, which describes the increase in the number of parameters from a standard transformer layer to a MoE layer. Given that, and denote the total number of parameters in a MoE layer excluding routing and the standard feed-forward layer, respectively. The expansion rate is then defined as
Expansion rate can also be seen as the total number of parameters in a MoE layer compared to its active parameters.
The concept of the expansion rate is intricately linked to the number of experts through the idea of granularity. Indeed, the definitions of both granularity and expansion rate extend and refine our understanding of the number of experts, symbolized as .
For non-granular models, where , the expansion rate is equal to the number of experts.
Intuitively, increasing granularity for a given expansion rate gives the model more flexibility in mapping datapoints to experts, potentially improving performance. We incorporate the notion of granularity into our scaling laws in Section 5. The discussion about practical tradeoffs in changing this parameter is given in Section 6.
Scaling Laws
Granularity determines changes in the architecture of MoE. In this section, we answer a central question of this work: whether the granular MoE models follow scaling laws and, if so, how granularity affects them. Thus, we aim to derive a parametric scaling law for predicting the final loss value based on granularity , total number of non-embedding parameters , and number of training tokens .
We run over 100 experiments on the decoder-only Transformer architecture, with each feed-forward component replaced by a Mixture of Experts layer. Those experiments involve training models with sizes ranging from 129M to 3.7B parameters across different training durations, from 16B to 130B tokens. We consider logarithmically spaced values of granularity between 1 and 16. To constrain the search space, is fixed, following the recommendations of Clark et al. (2022). In addition, we also run experiments with dense Transformers to compare their performance with MoE. The details of all architectures, the training procedure, and hyperparameter choices are described in detail in Appendix A.
In the subsequent part of this paper, we will use the notation to describe a MoE model with active parameters and expansion rate
We first answer the question of whether granular models follow the scaling laws. In Figure 4(a), it can be seen that increasing granularity results in a lower loss. The returns follow approximately an exponential pattern, converging to a positive constant. The empirical relationship given by Figure 3(a) suggests the following power-law dependence of loss on a varying granularity for given and and constants and that may be dependent on them,
2 Scaling the Model and Dataset Size
As outlined in Section 3.2, the power-law given by Eq. 1 consists of three terms that describe inherent data entropy and limitations in function representation and data. This derivation is independent of the architecture. In particular, the Eq. 1 also holds for constant granularity. Empirically, we observe a power law relationship in and analogous to that in dense models as depicted in Figure 3(b) for a fixed value of granularity (see also Fig. 1, Kaplan et al. (2020)). Furthermore, the validity of this functional form is verified by fit in Section 5.4.
Since we know that separate scaling laws are valid for given granularities, in the general form, the parameters in Eq. 1 can be dependent on the model’s granularity:
3 The Form of the Joint Scaling Law
Following the above observation that models with constant granularity obey Chinchilla scaling laws given by Eq. 1, the key question arises as to how the general notion of granularity can be incorporated into the joint scaling law. Moreover, the scaling law formula from Eq. 5 for constant and has to be representable by Eq. 4. This is because the former is a more general equation, encompassing shared hyper-parameters across all , , and . It is anticipated to align with the latter, consisting of distinct power laws, each with specific parameters for different and values. Consequently, the objective is to identify a function that fulfills these criteria.
In the subsequent sections, we aim to determine which of these parameters remain independent of and identify their functional form. Furthermore, we present some rationale for the structure of our formula.
Lower Bound. Consider the limit of Eq. 5 for and growing to infinity:
with the constant term dependent on granularity. This is contradictory to the fact that it captures the inherent entropy of the dataset. Lower bound of the achievable loss when training bigger models on more samples should not depend on the architecture, therefore parameter is constant for all granularities.
Granularity and Number of Tokens . As seen in Figure 3(c), the benefit of training a model on a larger dataset is almost the same for each granularity value. This suggests that there is no interaction between and . Therefore, we can assume that
Granularity and Model Size . We consider to be a constant that describes how the function scales with . In this work, we assume polynomial functional forms that rule out the potential dependency of on given the form of Eq. 4. Therefore, the only element dependent on is :
Finally, one could consider omitting the constant in the equation above, and it would still reduce to 4 for constant and . However, this would mean that a model with infinite granularity and a small number of active parameters can achieve the perfect perplexity of the lower bound. We assume that a sparse MoE (Mixture of Experts) model is unlikely to surpass the performance of an equivalent dense model that has a matching total number of parameters, all of which are active. This means that constant can act as a marginal improvement due to granularity.
Subsequently, we fit parameters in Eq. 9 to describe the scaling of MoE. For comparison, we also perform fitting for dense transformer given by Eq. 1. Similarly to Hoffmann et al. (2022), we use Huber loss (Huber, 1964), with . The optimization is performed using the BFGS algorithm. We include a weight decay of to enhance generalization. We start with fitting parameters in Eq. 9 and then find architecture-dependent coefficients and in Eq. 1. We observe a good fit, with . The values are presented in Table 1. We depict the results in Figure 4.
4 Fitting the Parametric Scaling Law
We validate the stability of the fit by excluding the top of models with the lowest perplexity and finding the coefficients based on the remaining experiments. We observe that the formula remains almost unchanged in this scenario (see Table 5 in Appendix B). The validation RMSE is 0.019. Results are depicted in Figure 5 (a).
5 MoE Scaling Properties
Comparing the part of the formula that approximates underfitting (that is, dependent on training tokens) in MoE () and Transformer (), we can infer that MoE models need longer training to perform competitively but scale better after reaching that point. Nonetheless, this moment may still precede the compute optimal for both models. On the other hand, we can see that the exponent on dense models scales better with a total number of parameters than the MoE counterpart . This should not be surprising since dense models use all parameters on each token contrary to MoE, which gains a computational advantage by activating only a subset of them. Therefore, the fair comparison of the performance has to take into account FLOPs used by each model type. In the next section, we find compute-optimal granularity for a given FLOP budget.
Optimal Allocation of Computational Budget
In Section 5, we show that higher granularity leads to lower loss for the same number of training steps. This is not always the case if we consider the wall-clock time. As depicted in Figure 5 (b), in practice for too high values of (relative to ), training can be bottlenecked by the routing cost. Practical modeling of this situation is possible by measuring FLOPs in routing. In this section we find optimal for a given computational budget by solving the following optimization problem,
It is important to acknowledge that increasing granularity can lead to some challenges in training the model, namely higher computational and communication costs and a larger memory footprint. The main component responsible for higher costs is the increase in routing operations due to a larger pool of granular experts. This increase is proportional to the value of For standard, non-granular MoE models (), the routing overhead still exists, although it has been considered negligible.
Taking into account the routing operation overhead, the number of used FLOPs is described by the following formula:
given expansion rate , granularity , and constants that denote FLOPs per active parameter ratio, respectively, within routing () and within the rest of the network (). The term is the number of active parameters within a transformer block, while is the number of active parameters within a routing network. The in-depth analysis of constants and can be found in Appendix E. We exclude embedding and unembedding from the FLOPs calculations, following Hoffmann et al. (2022).
Observe that, in contrast to scenarios where routing operations are omitted, the FLOPs calculation that incorporates routing overhead relies on both and . Consequently, an additional condition is required to determine the scaling of and in relation to an increase in , the number of parameters. It is noted that minor variations in the depth-to-width ratio are not significant (Kaplan et al., 2020). Following this analysis, we opt to adopt the assumption that .
The total number of parameters in the feed-forward layer, excluding the routing matrix, is , and in attention (key, query, value, and output projection). This results in the following formula for the total number of parameters, .
2 Compute Optimal Formula
Taking into consideration we need to solve the following optimization problem, given ,
All these constraints are reducible to a one-dimensional optimization problem, which is, however, hard to solve analytically. Therefore we approximate the solution using Brent’s method (Brent, 1971). The results of this optimization for varying FLOPs budgets are plotted in Figure 1 while the optimal configurations of parameters for selected model sizes are presented in Table 2. To validate the uncertainty of these predictions, we follow Hoffmann et al. (2022) and calculate the 10th and 90th percentiles estimated via bootstrapping data (see Appendix C for the detailed results).
3 MoE is Always More Efficient
Contrary to the results from Clark et al. (2022), in Figure 1 we can see, that Mixture-of-Experts can be always considered more efficient than dense Transformers, regardless of the model size. According to our previous observations from Section 5.5, MoE models scale better with optimal training. However, for short training schedules, they may under-perform dense models. This means that for constant training time and increasing model size, there exists a point where both models will become very under-trained, in which scenario dense models surpass MoE. This shows why in Clark et al. (2022), where varying the number of training tokens has not been considered, MoE was predicted to be under-performing for models bigger than . However, when all training hyper-parameters are properly selected to be compute-optimal for each model, the gap between dense and sparse models only increases as we scale.
Discussion
In Section 5, we argue that model performance improves with increasing granularity. This postulate largely aligns with the empirical findings of our study. Nonetheless, at exceedingly high granularity levels, such as in models characterized by and , there is an observable decline in performance. This phenomenon is particularly evident in scenarios where the number of parameters in the routing mechanism exceeds active parameters in actual experts. Additionally, as described in Section 6, the utility of such high granularity is predominantly restricted to models of substantial size. In alignment with the principles outlined by Hoffmann et al. (2022), this research focuses more on findings that can be broadly applied rather than delving into the specific details of these corner-case situations. However, it is hypothesized that the efficiency of models with significantly high granularity could be potentially enhanced through careful expert initialization or modifications to the routing algorithm. These ideas are set aside to be investigated in future studies.
Varying Expansion Rate.
In this study, due to computational resources constraint, we focus on as recommended by Clark et al. (2022). This value of was also used for the largest models in other works (Du et al., 2022; Zhou et al., 2022) and the best-performing configuration in Fedus et al. (2022). Nonetheless, we acknowledge the importance of considering different expansion rates, as different levels of may be chosen based on factors like the target size of the model in memory. Therefore, in Appendix D, we present the results of the study for and show that the main findings of this work are still valid in such cases.
Including E𝐸E in the formula.
Another possible advancement would be to unify all of the factors and in one formula. While this would open the possibility of studying the relationships between coefficients in more detail, it would also be hard to practically recommend the optimal configuration in such a scenario using only FLOPs. This is because larger values of typically lead to better performance but also incur additional memory requirements. Therefore, the choice of expansion rate may be heavily dependent on the available hardware configuration. We leave a detailed study of these factors for future work.
Modeling the cost of granularity.
It is important to note that the exact estimation of the training cost of MoE models is dependent on the training setup, hardware, and implementation. Specifically, increasing G can lead to higher transfer costs, depending on the adopted model of distributed training. Therefore, the precise selection of hyperparameters should be made considering these factors. In this work, we model the cost of operations using FLOPs, which is common in the Scaling Laws literature (Kaplan et al., 2020; Hoffmann et al., 2022; Frantar et al., 2023). Additionally, we would like to note that in our setup, we observe significant gains of fine-grained MoE measured as wall-clock time needed to achieve given perplexity (see Fig. 5 (b) for an example).
Conclusions
This study introduces a novel hyperparameter, granularity (), and underscores the significance of adjusting it for optimizing the efficiency of experts within MoE models. A central finding of this research is that a standard granularity of is suboptimal across a broad range of FLOPs, leading to the recommendation of using higher granularity values to enhance MoE model performance and efficiency. Simultaneously, this work emphasizes the importance of varying training duration for compute-optimal settings. Consequently, both granularity and variable training length are incorporated into new scaling laws. These laws confidently demonstrate that MoE models consistently outperform dense transformers in terms of efficiency and scaling. This work not only sheds new light on the scaling laws applicable to MoE models but also provides practical guidance for improving computational efficiency in large language models. The insights are critical for the development and optimization of large-scale language models, marking a significant advancement in the field.
Reproducibility
The code used to produce the results described in this work is open-sourced and can be found at github.com/llm-random/llm-random.
Acknowledgments
We would like to express sincere gratitude to Piotr Miłoś and Tomasz Trzciński for valuable feedback and to Aleksandra Weglarz for her help with graphic design.
This work was funded by IDEAS NCBR, which also provided significant computational resources a supportive research environment and direction. The research was supported by PL-Grid infrastructure (grant PLG/2023/016148). We also benefited from the Entropy cluster (hosted at the Faculty of Mathematics, Informatics and Mechanics of the University of Warsaw) funded by NVIDIA, Intel, the Polish National Science Center grant 2022/45/N/ST6/02222, and ERC Starting Grant TOTAL. Marek Cygan was partially supported by an NCBiR grant POIR.01.01.01-00-0392/17-00.
References
Appendix A Architecture and Training Setup
In MoE, we use the Expert Choice routing algorithm, as it guarantees a balanced expert load without tuning additional hyperparameters. To maintain compatibility with autoregressive language modeling, we apply the recipe described in Zhou et al. (2022): tokens are grouped by position across different sequences. The group size is always set to We match the number of FLOPs for MoE and dense models with the same (meaning we activate an average of parameters per token in each MoE layer). In the router, softmax is performed over the expert dimension, while we choose tokens over the token dimension, as this leads to the best performance (as opposed to performing softmax over the token dimension). We put an additional layer normalization before the output of MoE layer. This gives a small improvement for standard MoE, but is crucial for the performance of models with
Table 3 and Table 4 list the considered architecture and training variants for dense and MoE models, respectively.
Appendix B Validation of the Scaling Law
In this section, we provide coefficients of the scaling law fitted with 20 of datapoints with the lowest perplexity excluded for the purpose of validation.
Appendix C Reliability of Compute Optimal Formula
In this section, we assess the stability of our predictions presented in Section 6.1. Similarly to Hoffmann et al. (2022) we calculate the 10 and 90 percentiles estimated via bootstrapping data ( of the data is sampled times). See Table 6 for the details.
Appendix D Varying Expansion Rate
In this section, we provide results for The training procedure is the same as described in App. A. The models considered in this part are listed in Table 7.
We fit Eq. 9 using the same procedure as described in Section 5.4. The results are detailed in Table 8.
Using the coefficients and FLOPs calculation formulas, we can derive the compute optimal training parameters. The results are presented in Table 9.
We can observe that similarly to the case when larger compute budgets imply larger optimal values of Note that the values for and percentiles form larger intervals in this case, as in this part we run a smaller number of experiments and keep shorter training durations. However, we believe that this preliminary study forms a valuable addition to the results in the main part.
Appendix E FLOPs Constants
The number of FLOPs used in Transformer training, considering the routing operation overhead in MoE, can be described by the following formula:
Following Hoffmann et al. (2022), we assume to be . This is interpreted as 6 FLOPs for each pair of an active parameter (in linear projection) and a processed token. The breakdown of operations is as follows:
During the forward pass, 2 operations (single multiplication and single addition) are used to compute the matrix multiplication of an input and linear projection.
During the backward pass, 2 operations are used to compute gradients wrt. the input.
During the backward pass, 2 operations are used to compute gradients wrt. the weights of linear projection.
In our work, we have assumed the routing constant, , to be 14, with the breakdown presented below. The exact number of operations may depend on the implementation of routing, but it will be between 6 and 20. However, our main conclusions of the paper are resistant to different assumptions of this constant.
During the forward pass, 2 operations are used to compute the expert logits based on an input and “routing linear projection”.
During the backward pass, 2 operations are used to compute gradients for “routing linear projection” wrt. the input.
During the backward pass, 2 operations are used to compute gradients for “routing linear projection” wrt. the weights of linear projection.
During the forward pass, 2 operations are used to route input tokens to chosen experts.
During the forward pass, 2 operations are used to route expert outputs to chosen tokens and multiply those outputs by the routing score.
During the backward pass, 2 operations are used to route gradients from output tokens to experts.
During the backward pass, 2 operations are used to route gradients from experts to input tokens.
Similarly to the calculation of FLOPs for , FLOPs come in pairs as each multiplication is followed by an addition (used to accumulate outputs or gradients).