When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method

Biao Zhang, Zhongtao Liu, Colin Cherry, Orhan Firat

Introduction

Leveraging and transferring the knowledge encoded in large-scale pretrained models for downstream applications has become the standard paradigm underlying the recent success achieved in various domains (Devlin et al., 2019; Lewis et al., 2020; Raffel et al., 2020; Dosovitskiy et al., 2021; Baevski et al., 2020), with the remarkable milestone set by large language models (LLMs) that have yielded ground-breaking performance across language tasks (Brown et al., 2020; Zhang et al., 2022b; Scao et al., 2022; Touvron et al., 2023). Advanced LLMs, such as GPT-4 (OpenAI, 2023) and PaLM 2 (Anil et al., 2023), often show emergent capabilities and allow for in-context learning that could use just a few demonstration examples to perform complex reasoning and generation tasks (Wei et al., 2022; Zhang et al., 2023; Fu et al., 2023; Shen et al., 2023). Still, LLM finetuning is required and widely adopted to unlock new and robust capabilities for creative tasks, get the most for focused downstream tasks, and align its value with human preferences (Ouyang et al., 2022; Yang et al., 2023; Gong et al., 2023; Schick et al., 2023). This becomes more significant in traditional industrial applications due to the existence of large-scale annotated task-specific data accumulated over years.

There are many potential factors affecting the performance of LLM finetuning, including but not limited to 1) pretraining conditions, such as LLM model size and pretraining data size; and 2) finetuning conditions, such as downstream task, finetuning data size and finetuning methods. Intuitively, the pretraining controls the quality of the learned representation and knowledge in pretrained LLMs, and the finetuning affects the degree of transfer to the donwstream task. While previous studies have well explored the scaling for LLM pretraining or training from scratch (Kaplan et al., 2020; Hoffmann et al., 2022) and the development of advanced efficient finetuning methods (Hu et al., 2021; He et al., 2022), the question of whether and how LLM finetuning scales with the above factors unfortunately receives very little attention (Hernandez et al., 2021), which is the focus of our study. Note, apart from improving finetuning performance, studying the scaling for LLM finetuning could help us to understand the impact of different pretraining factors from the perspective of finetuning, which may offer insights for LLM pretraining.

In this paper, we address the above question by systematically studying the scaling for two popular ways of LLM finetuning: full-model tuning (FMT) that updates all LLM parameters and parameter-efficient tuning (PET) that only optimizes a small amount of (newly added) parameters, such as prompt tuning (Lester et al., 2021, Prompt) and low-rank adaptation (Hu et al., 2021, LoRA). We first examine finetuning data scaling (Hernandez et al., 2021), on top of which we further explore its scaling relationship with other scaling factors, including LLM model size, pretraining data size, and PET parameter size. We focus on the data-limited regime, where the finetuning data is much smaller than the LLM model, better reflecting the situation in the era of LLM. For experiments, we pretrained two sets of bilingual LLMs (English&German, English&Chinese) with model size ranging from 1B to 16B, and performed large-scale study on WMT machine translation (English-German, English-Chinese) and multilingual summarization (English, German, French and Spanish) tasks with up to 20M finetuning examples. Our main findings are summarized below:

We propose the following multiplicative joint scaling law for LLM finetuning:

where {A,E,α,β}\{A,E,\alpha,\beta\} are data-specific parameters to be fitted, DfD_{f} denotes finetuning data size, and XX refer to each of the other scaling factors. We show empirical evidence that this joint law generalizes to different settings.

Scaling LLM model benefits LLM finetuning more than scaling pretraining data.

Increasing PET parameters doesn’t scale well for LoRA and Prompt, although LoRA shows better training stability.

The scaling property for LLM finetuning is highly task- and data-dependent, making the selection of optimal finetuning method for a downstream task non-trivial.

LLM-based finetuning could encourage zero-shot generalization to relevant tasks, and PET performs much better than FMT.

Setup

We consider machine translation and multilingual summarization as the downstream tasks for the finetuning, because 1) these tasks require resolving cross-lingual understanding and generation, which represent high complexity and are challenging; and 2) they are well established in NLP with rich amount of available finetuning corpora. Specially, we adopt WMT14 English-German (En-De) and WMT19 English-Chinese (En-Zh) (Kocmi et al., 2022) for translation. We combine the De, Spanish (Es) and French (Fr) portion of the multilingual summarization dataset (Scialom et al., 2020) with CNN/Daily-Mail (Hermann et al., 2015, En) for summarization and denote it as MLSum. Details about each task are listed in Table 1(a). Note for MLSum, we directly concatenate the datasets of different languages for training and evaluation, where each article is prepended a prompt indicating its language “Summarize the following document in {lang}:”.

LLMs and Preraining

We adopt the exact setup as in Garcia et al. (2023) for LLM pretraining. The model is a decoder-only Transformer with multi-query attention (Chowdhery et al., 2022) and trained with the modified UL2 objective (Tay et al., 2022). Considering the focused downstream tasks and also to ensure the generalization of our study, we pretrained two sets of bilingual LLMs, i.e. En-De LLM and En-Zh LLM. The pretraining data is a mix of monolingual data from two languages: we use En/De (En/Zh) data with about 280B (206B) tokens to pretrain the En-De (En-Zh) LLM. We train LLMs with parameter sizes from 1B to 16B by varying model configurations as in Table 3 and keep all other settings intact. All LLMs are optimized using Adafactor (Shazeer & Stern, 2018) for one training epoch under a cosine learning rate decay schedule (from 0.01 to 0.001). We refer the readers to (Garcia et al., 2023) for more details about the pretraining.

Finetuning Settings

We mainly study the scaling for the following three finetuning methods:

Full-Model Tuning (FMT): This is the vanilla way of finetuning which simply optimizes all LLM parameters;

We explore 4 different factors for the scaling, which are summarized in Table 1(b). Except LLM model scaling, all experiments are based on the corresponding 1B LLM. For pretraining data scaling, we adopt intermediate pretrained checkpoints as the proxy due to computational budget constraint while acknowledge its sub-optimality. Details for optimization are given in Appendix.

Evaluation

We use the best checkpoint based on token-level perplexity (PPL) on the dev set for evaluation. For scaling laws, we report PPL on test sets; for general generation, we use greedy decoding, and report BLEURT (Sellam et al., 2020) and RougeL (Lin, 2004) for translation and summarization, respectively. For zero-shot evaluation, we adopt Flores200 (NLLB Team, 2022) and evaluate on {Fr, De, Hindi (Hi), Turkish (Tr), Polish (Po)→\rightarrowZh} and {Fr, Zh, Hi, Tr, Po→\rightarrowDe} for En-Zh and En-De translation respectively. For scaling law evaluation, we split empirical data points into two sets, empirical fitting and held-out set, where the former is used for fitting scaling parameters and the latter is used for evaluation. We report mean absolute deviation. To reduce noise, we perform three runs, each with a different random subset of the finetuning data, and report average performance. When sampling for MLSum, we keep the mixing ratio over different languages fixed.

Why Multiplicative Joint Scaling Law?

We consider 4 scaling factors in this study but jointly modeling all of them is time and resource consuming. Instead, we treat finetuning data as the pivoting factor and perform joint scaling analysis between it and every other factor separately. Below, we start with finetuning experiments for FMT, Prompt and LoRA on WMT14 En-De, and then explore the formulation for the joint scaling.

We first examine the scaling over finetuning data size for each LLM model size independently, with a single variable formulation: L^(Df)=\nicefracADfβ+E\hat{\mathcal{L}}(D_{f})=\nicefrac{{A}}{{D_{f}^{\beta}}}+E. Following Hoffmann et al. (2022), we estimate {A,β,E}\{A,\beta,E\} using the Huber loss (δ=0.001\delta=0.001) and the L-BFGS algorithm, and select the best fit from a grid of initializations. Figure 1 shows that the above formulation well describes LLM finetuning data scaling with small predictive errors across model sizes and methods, echoing with the findings of Hernandez et al. (2021). Such scaling trend also implies that, while finetuning with small amount of examples could achieve decent results (Zhou et al., 2023; Gao et al., 2023), larger scale finetuning data still contributes to improved downstream performance, especially when the downstream application is well defined.

Additive or multiplicative joint scaling law for LLM finetuning?

Figure 1 also shows some scaling pattern over LLM model sizes, suggesting the existence of a joint scaling law. We explore two formulations: multiplicative as in Eq. (1) and additive: L^(X,Df)=\nicefracAXα+\nicefracBDfβ+E\hat{\mathcal{L}}(X,D_{f})=\nicefrac{{A}}{{X^{\alpha}}}+\nicefrac{{B}}{{D_{f}^{\beta}}}+E (Hoffmann et al., 2022), and compare them via empirical experiments.For LLM model scaling, we omitted the newly added parameters in PET because 1) the added parameters only take a very tiny proportion, and 2) the proportion across LLM model sizes is similar. Take the 1B LLM as example. ∣P∣=100|P|=100 in Prompt adds 0.017% parameters; r=4r=4 in LoRA adds 0.19% parameters. We also explored different formulations for the new parameters for PET, which don’t make a substantial difference.

In both formulations, α\alpha and β\beta reflect the impact of factor XX and finetuning data size on the performance, respectively, which are factor-specific. EE is a model- and task-dependent term, describing irreducible loss (Ghorbani et al., 2021). We notice that the meaning for β\beta and EE generalizes over different factors XX, and thus propose to estimate them first based on results for both LLM model and pretraining data scaling.We didn’t consider PET parameter scaling when estimating β\beta and EE because this scaling is pretty weak and ineffective, as shown in Section 4. Such joint fitting could also reduce overfitting and improve extrapolation ability. We apply the following joint fitting loss:

where we set AX=eaX,BX=ebX,E=eeA_{X}=e^{a_{X}},B_{X}=e^{b_{X}},E=e^{e}, and XX refers to LLM model size or pretraining data size. Note bXb_{X} is only valid in the additive formulation. We then fix β\beta and EE and refit other parameters for each factor, separately.

Table 2 (and Table 6 in Appendix) shows that both joint laws perform similarly while the multiplicative one achieves slightly lower extrapolation error on average. Therefore, we adopt Eq. (1) for follow-up analysis.

Scaling Results for LLM Finetuning

Here, we show the empirical results for LLM model, pretraining data and PET parameter scaling on WMT14 En-De, WMT19 En-Zh and MLSum in Figures 2, 3 and 4, respectively. Results for BLEURT/RougeL are given in Appendix (Figures 7, 8 and 9), which shows high correlation with the PPL scores in general (see Table 7). Fitted scaling parameters are summarized in Table 4.

In each group of experiments, we leave several data points along each scaling dimension as the held-out set. We report the mean absolute derivation on the empirical fitting (Δe\Delta_{e}) and held-out (Δh\Delta_{h}) sets to show the fitting and predictive ability, respectively. In general, we observe that Eq. (1) captures the scaling trend of different factors under finetuning data scaling with small fitting and extrapolation errors. Note there are some mismatched cases, where the empirical data points themselves could be noisy mostly caused by unstable optimization and dev-set overfitting, challenging issues when tuning on small datasets. We observe high mismatch when extrapolating to 16B, particularly for LoRA and Prompt on WMT19 En-Zh in Figure 2. We ascribe this to 1) the insufficiency of empirical data over LLM model sizes (i.e. only 4 points) – the prediction by the fitted scaling law makes sense intuitively based on 1B-8B results, and 2) the inferior of the 16B En-Zh LLM due to pretraining instability, where its pretraining performance is not well predicted by even single-variable scaling laws as in Figure 10, Appendix.

LLM finetuning benefits more from LLM model scaling than pretraining data scaling across tasks and methods.

While LLM model size and pretraining data size show similar impact on the pretraining scaling following the optimal scaling under a computational budget constraint (Hoffmann et al., 2022; Muennighoff et al., 2023), they show slightly different roles in finetuning scaling. Intuitively, finetuning heavily relies on the knowledge encoded in the LLM, where LLM model size and pretraining data size both matter. However, results in Figures 2, 3 and Table 4 show that the scaling exponent for LLM model size αm\alpha_{m} often outnumbers that for pretraining data size αp\alpha_{p} across finetuning methods and tasks, i.e. αm>αp\alpha_{m}>\alpha_{p}. This suggests that using a larger LLM model is preferred over pretraining on a larger dataset, but we also notice that the difference in scaling is highly task-dependent. Our selection of closed generation tasks, i.e. translation and summarization, might deliver biased observations and for more creative generation tasks, larger and diverse pretraining data could be more crucial.

Scaling PET parameters is ineffective, delivering limited gains for both LoRA and Prompt.

The amount of newly added trainable parameters often forms a bottleneck for the expressivity of PET, controlled by the length ∣P∣|P| and rank rr in Prompt and LoRA, respectively. However, Figure 4 and Table 4 show that increasing PET parameter sizes (i.e. enlarging ∣P∣|P| and rr) affects finetuning performance marginally as demonstrated by the small scaling exponents, ∣αt∣≪1e−2|\alpha_{t}|\ll 1e-2, and even results in inverse scaling in some settings, e.g. LoRA on En-De. Besides, we observe that scaling Prompt length suffers from training instability as optimizing larger prompt embedding becomes non-trivial, which has also been seen in previous studies (Lester et al., 2021; Hu et al., 2021). We expect that carefully optimizing finetuning hyperparameters and prompt initialization may alleviate it to some extent. In this respect, LoRA is more stable and reliable.

Finetuning data have more pronounced influence on FMT than PET, where LoRA scales better than Prompt.

Different finetuning methods show different degrees of finetuning data scaling. Table 4 shows that the scaling exponent β\beta for FMT is often significantly higher than that for PET across settings, indicating that FMT is more data-hungry and also benefits more from increasing finetuning data. While the scaling exponents are quite similar across PET, β\beta for LoRA often slightly surpasses that for Prompt. As shown in Figures 2, 3 and 4, LoRA often achieves better finetuning performance with more finetuning data than Prompt while Prompt behaves better with only few thousands of finetuning examples.

PET depends more on LLM model and pretraining data scaling than finetuning data scaling across settings.

Since the majority of LLM parameters is frozen during finetuning, PET relies heavily on the encoded knowledge in pretrained LLMs when adapting them to downstream tasks. This is reflected by Table 4 that αm\alpha_{m} and αp\alpha_{p} are clearly larger than β\beta in PET. Figure 2 and 3 further support the scaling of LLM model, where the performance gap between FMT and PET is substantially narrowed with larger LLMs.

Discussion

Unfortunately, there is no universal answer! Intuitively, there exists a critical point for finetuning data size beyond which one finetuning method performs better than another. However, the high non-linearity of the joint scaling law hinders us from identifying such points analytically, although the finetuning data size follows a power law when the performance difference between two methods is fixed (see Appendix). We thus resort to empirical methods by extrapolating the fitted scaling law. Figure 5 shows the critical points as a function of LLM model size and pretraining data size over different tasks.

The scaling trend and actual value are highly dependent on the downstream task: critical points for one task can hardly generalize to other tasks. Still, the existence of such points suggests that the selection of finetuning methods should be based on the availability of finetuning examples. When only few thousands of finetuning examples are available, PET should be considered first, either Prompt or LoRA. With sightly larger datasets, LoRA would be preferred due to its stability and slightly better finetuning data scalability. For million-scale datasets, FMT would be good.

How does finetuning affect the generalization capability of the base LLM?

While finetuning on task-specific data improves task-specific performance, it may specialize the base LLM towards the task and hurt the models’ generalization. We examine this for different finetuning methods by performing zero-shot translation for LLMs finetuned on WMT14 En-De and WMT19 En-Zh (Few-shot results are in Appendix). We focus on generalization to related tasks, where the target language is shared, i.e. De and Zh, and generalization should be relatively easier (Johnson et al., 2017). We report average performance for translation from a diverse set of source languages other than English.

Figure 6 shows the results. While specializing on a downstream task, finetuning could still elicit and improve the generalization for closely related tasks, although the overall zero-shot translation quality is inferior. Note whether finetuning benefits generalization is method- and task-dependent. Overall, Prompt and LoRA achieve relatively better results than FMT particularly when the base LLM is large, mostly because LLM parameters are frozen and the learned knowledge get inherited. This also suggests that when generalization capability is a big concern, PET should be considered.

Related Work

With the significant increase of model size, updating all LLM parameters becomes computationally inefficient and unaffordable. Researchers thus resort to parameter efficient tuning methods that target achieving the best performance with minimal tunable parameters. Efforts in this direction mainly focus on developing efficient tunable modules for LLMs, such as adapters that insert small feed-forward layers (Houlsby et al., 2019; Bapna et al., 2019), prefix and prompt tuning that appends tunable embeddings to the input (Li & Liang, 2021; Lester et al., 2021), LoRA and compacter that adopts low-rank decomposition (Hu et al., 2021; Mahabadi et al., 2021), Bitfit that adds tunable bias vectors (Zaken et al., 2021), IA3 that scales model activations (Liu et al., 2022) and QLoRA that leverages quantization (Dettmers et al., 2023), to name a few. While previous studies reported encouraging performance with PET, e.g. reaching and even surpassing FMT across various domains (He et al., 2022; Ding et al., 2022; Liu et al., 2022; Dettmers et al., 2023), they mainly focus on one or few experimental setups, leaving the question of how scaling affects the performance of different finetuning methods under-explored.

Scaling Laws

Recent research has shown that the performance of neural models can be predicted by a power-law of model and/or data sizes (Hestness et al., 2017; Kaplan et al., 2020). Such pattern widely exists across different domains and model architectures, such as computer vision (Zhai et al., 2021), autoregressive generative modeling (Henighan et al., 2020), neural machine translation (Gordon et al., 2021; Ghorbani et al., 2021; Bansal et al., 2022; Zhang et al., 2022a), multilingual translation (Fernandes et al., 2023), multi-modal modeling (Aghajanyan et al., 2023) and sparse neural architectures (Frantar et al., 2023). These laws provide a valuable tool for guiding training decisions (Hoffmann et al., 2022) and model development by understanding how model performance evolves with scale, which greatly facilitates the development of LLMs (OpenAI, 2023). Unfortunately, the study of scaling for LLM finetuning lags behind badly, and our study fills this gap.

The most closely related work to ours is (Hernandez et al., 2021) which explored the scaling for knowledge transfer by comparing finetuning with training from scratch. Our study is orthogonal to theirs with significant difference as our key focus is understanding the scaling of different factors for LLM finetuning, rather than the transfer.

Conclusion and Future Work

In this paper, we systematically studied the scaling for LLM finetuning, considering different factors including LLM model size, pretraining data size, finetuning data size, PET parameter size and diverse finetuning methods. To ensure the generality, we worked on two sets of LLMs, three different downstream tasks (translation and summarization), and three finetuning methods (FMT, Prompt and LoRA). We proposed a multiplicative joint scaling law that could describe the scaling relationship between finetuning data size and each other scaling factor. Extensive results show that increasing LLM model size has a higher impact on finetuning than pretraining data scaling, and that scaling PET parameter is ineffective. In addition, finetuning scaling is highly task- and data-dependent, making the selection of best finetuning method for a downstream task less conclusive.

We acknowledge that our work suffers from some limitations. The proposed joint scaling law is mostly based on empirical results on closed generation tasks without theoretical groundings. Whether it could generalize to different finetuning scenarios requires more experimentation, which however is beyond our current computing budget. Besides, we understand the imperfection of the optimization and evaluation for Prompt and LoRA in some setups. In the future, we would like to extend our study to multi-modal LLMs, explore the impact of finetuning data quality and consider open and creative generation tasks as well as multi-task setup for finetuning.

Acknowledgements

We thank the reviewers for their insightful comments. We thank Yamini Bansal for providing valuable feedback on the scaling laws, Xavier Garcia for reviewing this work with constructive comments, Frederick Liu for helpful discussion on PET optimization, and Quoc Le, Apu Shah and Google Translate team for supporting this research.

We also thank the colleagues building the training infrastructure used in this paper: Brian Lester, Rami Al-Rfou and Noah Constant for prompt tuning, Chu-Cheng Lin for LoRA, Xavier Garcia and the T5X team (Roberts et al., 2023) for the training framework.

References

Appendix A Appendix

For optimization, we continue the pretraining from the given pretrained checkpoint on finetuning data but with the standard conditional log-likelihood loss. More specifically, for each finetuning example, we concatenate the input and target into a single sequence and compute the log-likelihood on the target alone. Adafactor and cosine learning rate schedule are reused. Note En-De and En-Zh LLM are pretrained for 135K and 98K steps, respectively. All LLMs are further finetuned for up to 200K steps (except for WMT En-Zh (FMT) which is 300K steps) or 100 epochs, whichever comes first. To get the best performance, we optimize the initial learning rate and batch size for different finetuning methods based on the 1B LLM via grid search. Finally, we set the learning rate to 3e−1,1e−23e^{-1},1e^{-2} and 1e−31e^{-3} for Prompt, LoRA and FMT, respectively, and set the batch size to 16 and 128 for PET and FMT, respectively.

While Eq. (1) hinders us from computing DfcD_{f}^{c} directly, it still allows for theoretical analysis between two finetuning methods when their performance gap is a constant:

, which follows another power-law. Intuitively, the exponent γ\gamma captures the transferability difference of the two methods to the downstream task as scaling factor XX. We summarize the coefficients for different tasks in Table 5, where the value differs greatly over tasks and there are no clear patterns across settings.

How does finetuning affect the few-shot capability of the base LLM?

Apart from zero-shot translation, we also explore LLM’s few-shot capability after finetuning. Few-shot generation not only offers a way to inspect LLM’s capability but also is of interest to downstream applications as it provides an effective way to adapt models over domains. Figures 11, 12, 13 and 14 shows the impact of finetuning on few-shot generation.

We note that FMT degenerates LLM’s few-shot capability in most cases, where adding more finetuning data often reduces the few-shot performance. By contrast, PET behaves more robustly which retains most of LLM’s few-shot capability regardless of model size and pretraining data size.