VeRA: Vector-based Random Matrix Adaptation
Dawid J. Kopiczko, Tijmen Blankevoort, Yuki M. Asano
Introduction
In the era of increasingly large and complex language models, the challenge of efficient adaptation for specific tasks has become more important than ever. While these models provide powerful capabilities, their extensive memory requirements pose a significant bottleneck, particularly when adapting them for personalized use. Consider, for example, a cloud-based operating system assistant that continuously learns from and adapts to individual user behaviors and feedback. The need to store multiple checkpoints of finetuned models for each user rapidly escalates the required storage, even more so when multiple tasks come into play.
The situation is further exacerbated when we look at the state-of-the-art models like GPT-4 (OpenAI, 2023). Finetuning techniques like LoRA (Hu et al., 2022), while effective, still introduce considerable memory overhead. As an illustrative example, applying LoRA with a rank of 16 to the query and value layers of GPT-3 (Brown et al., 2020) would demand at least 288MB of memory, if stored in singe-precision – at a million finetuned weights, e.g., one per user, that would amount to 275TB.
Given the recent proliferation of language models and their deployment in personalized assistants, edge devices, and similar applications, efficient adaptation methods are paramount. We believe there is untapped potential for even more efficient approaches. Previous work (Aghajanyan et al., 2021) pointed out the low intrinsic dimensionality of pretrained models’ features. These studies reported numbers much lower than the trainable parameters used in LoRA, suggesting there is room for improvement.
In parallel to this, recent research has shown the surprising effectiveness of models utilizing random weights and projections (Peng et al., 2021; Ramanujan et al., 2020; Lu et al., 2022; Schrimpf et al., 2021; Frankle et al., 2021). Such models serve as the basis of our proposed solution, Vector-based Random Matrix Adaptation (VeRA), which minimizes the number of trainable parameters introduced during finetuning by reparametrizing the weights matrices. Specifically, we employ “scaling vectors” to adapt a pair of frozen random matrices shared between layers. With this approach, many more versions of the model can reside in the limited memory of a single GPU.
In summary, our main contributions are as follows:
We introduce a novel finetuning method with no additional inference time cost. Our method further reduces the number of trainable parameters compared to the state-of-the-art LoRA method, while yielding comparable results.
We compare our approach with LoRA and other parameter-efficient adaptation methods on the natural language understanding (GLUE) and natural language generation (E2E) benchmarks, and compare against LoRA on instruction-following and image classification tasks.
We perform an ablation study to better understand the individual components of our method and their effects on performance.
Related Work
LoRA offers an innovative solution to the computational challenges posed by the finetuning of large pretrained language models. Introduced by Hu et al. (2022), the method employs low-rank matrices to approximate the weight changes during finetuning, effectively reducing the number of parameters that need to be trained. Among its advantages, LoRA significantly lowers the hardware barrier for finetuning by reducing the need for gradient calculation and optimizer state maintenance for most parameters. It can also work with quantized model weights (Dettmers et al., 2023), reducing the requirements even further. Furthermore, LoRA modules are easily swappable, making task-switching efficient and less resource-intensive. Importantly, and different to adapter-based finetuning approaches (Houlsby et al., 2019; Lin et al., 2020; Pfeiffer et al., 2021; Rücklé et al., 2021), LoRA incurs no additional inference time cost when deployed, as the trainable matrices can be merged with the frozen weights.
Based on this, AdaLoRA (Zhang et al., 2023b) extends the LoRA method, introducing dynamic rank adjustment for the low-rank matrices during finetuning. The core idea is to optimally distribute the parameter budget by selectively pruning less important components of the matrices based on an importance metric.
Parameter Efficiency in Existing Methods
While methods such as LoRA have shown significant improvements in finetuning performance, they still require a considerable amount of trainable parameters. According to Aghajanyan et al. (2021), the upper bound for intrinsic dimensions is much smaller than what is typically utilized in such methods. For instance, the The smallest dimension that provides a satisfactory solution, which is 90% of the full training metric, as defined by Li et al. (2018). for RoBERTabase is reported to be , whereas authors of the LoRA paper reported using M trainable parameters for this model, suggesting that the parameter count could be reduced further.
Although AdaLoRA takes steps in this direction by dynamically allocating parameters to more critical layers, we posit that a different approach could achieve substantial parameter reduction, while tolerating a marginal performance degradation. This sets the stage for the method we introduce in the following section.
Random Models and Projections.
The concept of using random matrices and projections for model efficiency is supported by multiple strands of research. Frankle & Carbin (2019) identified that randomly-initialized neural networks contain subnetworks that are capable of reaching high performance when trained. Meanwhile, Ramanujan et al. (2020) revealed that there exist subnetworks that can achieve impressive results even in the absence of training. Aghajanyan et al. (2021) showed that training only a small number of parameters, randomly projected back into the full space, could achieve 90% of the full-parameter model performance. Ruiz et al. (2023) introduced a parameter-efficient finetuning method for personalization of text-to-image models, utilising random frozen matrices inside LoRA. Other works (Lu et al., 2022; Schrimpf et al., 2021; Frankle et al., 2021) have shown that frozen, randomly initialized models, with small sections finetuned, can perform surprisingly well.
Collectively, these works create a compelling case for the utilization of frozen random matrices in finetuning methods, providing both a theoretical and an empirical foundation for the approach taken in this paper.
Method
In this section, we introduce Vector-based Random Matrix Adaptation, a novel parameter-efficient finetuning method that builds upon and extends the state-of-the-art method, LoRA. The central innovation in VeRA lies in the reparameterization of the low-rank matrices. Specifically, we freeze a single pair of randomly initialized matrices, shared across all adapted layers, and introduce trainable scaling vectors that allow for layer-wise adaptation, as shown in Figure 1. Similarly to LoRA, trained scaling vectors along with low-rank matrices can be merged into original weights, eliminating additional inference latency.
where we undeline the parameters updated via gradient descent. This approximation enables the model to keep the original weight frozen while optimizing only the new low-rank matrices and . These matrices are much smaller in size than the original matrix due to their rank-reduced nature. has shape and has shape , where serves as the bottleneck dimension. In contrast, our VeRA method is expressed as:
2 Parameter Count
We use to denote the number of finetuned layers and to represent the dimension of these layers. The number of trainable parameters in VeRA is then governed by , contrasting with LoRA’s . Specifically, for the lowest rank (i.e., ), VeRA requires approximately half the trainable parameters of LoRA. Moreover, as the rank increases, VeRA’s parameter count increases by for each increment, a substantial saving compared to LoRA’s . This parameter efficiency becomes notably significant in the context of extremely deep and wide models, such as GPT-3 (Brown et al., 2020), which has 96 attention layers and a hidden size of 12288.
Building on this efficiency, the main advantage of VeRA is its minimal memory footprint for storing the trained weight adjustments. Because the random frozen matrices can be regenerated from a random number generator (RNG) seed, these do not need to be stored in memory. This substantially reduces the memory requirement, which is now limited to the bytes needed for the trained and vectors and a single RNG seed. The memory efficiency in comparison to LoRA is shown in Table 1.
3 Initialization Strategies
Shared Matrices: In our method, we employ Kaiming initialization (He et al., 2015) for the frozen low-rank matrices and . By scaling the values based on matrix dimensions, it ensures that a matrix product of and maintains a consistent variance for all ranks, eliminating the need to finetune the learning rate for each rank.
Scaling Vectors: The scaling vector is initialized to zeros, which aligns with the initialization of matrix in LoRA and ensures that the weight matrix is unaffected during the first forward pass. The scaling vector is initialized with a single non-zero value across all its elements, thereby introducing a new hyperparameter that may be tuned for better performance.
Figure 1 illustrates example initializations for the low-rank matrices and scaling vectors in VeRA. Specifically, the low-rank matrices are initialized using a normal distribution, and the vector is initialized with ones. Note that alternative initializations, such as uniform distribution for and , and other non-zero constants for , are also explored in our experiments.
Experiments
In this section, we conduct a series of experiments to evaluate our finetuning method. We start by comparing our approach to LoRA and other baselines on the GLUE and E2E benchmarks. Following this, we turn our attention to instruction-tuning of Llama models, and image classification with Vision Transformers. Next, we select one task and vary the rank for both methods, LoRA and VeRA, to examine how performance scales with the number of trainable parameters. Lastly, an ablation study sheds light on the importance of each component in our method, including the influence of different initializations.
We compare VeRA to the following baselines:
Full finetuning - the model is initialized with pretrained weights and all parameters are being trained.
Bitfit - this baseline involves the sole finetuning of bias vectors, keeping all other parameters fixed. This technique has been investigated in depth by Zaken et al. (2022).
Adapter tuning - initially introduced by Houlsby et al. (2019), involves the integration of adapter layers between the self-attention and MLP modules, followed by a residual connection. This setup includes two fully connected layers and a nonlinearity and is denoted as . A variation by Lin et al. (2020), , employs the adapter layer solely after the MLP module and subsequent to a LayerNorm. This closely resembles an alternative design suggested by Pfeiffer et al. (2021), referred to as . Another baseline, termed AdapterDrop by Rücklé et al. (2021), enhances efficiency by omitting certain adapter layers and is represented as .
LoRA (Hu et al., 2022) - as introduced in the earlier section.
1 GLUE Benchmark
We evaluate our approach on the General Language Understanding Evaluation (GLUE) benchmark (Wang et al., 2019), employing the RoBERTabase and RoBERTalarge models (Liu et al., 2019). For RoBERTabase we use a rank of 1024, and for RoBERTalarge a rank of 256. The shared matrices are initialized using the uniform version of Kaiming initialization as implemented in PyTorch (Paszke et al., 2019), with an initial value of 0.1 for the vector.
Our experimental setup generally aligns with that of Hu et al. (2022), applying our method to the query and value projection matrices in each self-attention module and fully training the classification head. Unlike Hu et al. (2022), who used an additional hyperparameter to adjust gradients for the adapted layers, we introduce separate learning rates for the classification head and the adapted layers. We determine the learning rates and the number of training epochs through hyperparameter tuning; for detailed settings, refer to the Table 8 in Appendix A. The batch size is set to 64 for RoBERTabase and 32 for RoBERTalarge, with maximum sequence lengths of 512 and 128 respectively.
Due to time constraints and budget limitations, we omit the time-intensive MNLI and QQP tasks, thus forgoing the use of the MNLI trickFor the RoBERTabase model and MRPC, RTE and STS-B tasks, Hu et al. (2022) initialized the model with the best weights finetuned on the MNLI task. for tasks MRPC, RTE, and STS-B. In line with Hu et al. (2022), we report the number of trainable parameters attributable to the finetuned layers, explicitly excluding the classification head, which is trained in a standard way. We perform 5 runs with different random seeds, recording the best epoch’s outcome for each run, and report the median of these results.
Table 2 reveals that VeRA performs competitively with LoRA across both models, yet achieves these results with an order of magnitude fewer parameters.
2 E2E Benchmark
For the E2E benchmark (Novikova et al., 2017), we follow the experimental setup from Hu et al. (2022) and finetune the GPT-2 (Radford et al., 2019) Medium and Large models. For LoRA we use the implementation and set of hyperparameters provided in Hu et al. (2022), while for VeRA we change the rank and learning rate, both of which are tuned. Table with all hyperparameters used can be found in Appendix A.
We report results from the last epoch. Table 3 shows that VeRA outperforms LoRA with 3 and 4 times less trainable parameters, for GPT2 Medium and Large respectively.
3 Instruction tuning
Instruction tuning is a process by which language models are finetuned to follow specific instructions more effectively (Ouyang et al., 2022). We demonstrate the efficacy of VeRA in enabling Llama (Touvron et al., 2023a) and Llama2 (Touvron et al., 2023b) models to follow instructions using only M and M trainable parameters, for 7B and 13B variants respectively, in contrast to M and M trainable parameters when employing LoRA with a rank of 64 as proposed by Dettmers et al. (2023).
We perform finetuning using both LoRA and VeRA, by applying both methods on all linear layers except the top one, similarly to Dettmers et al. (2023). Additionally, we leverage the quantization techniques from Dettmers et al. (2023) to train the model on a single GPU.
For our experiment, we employ the Alpaca dataset (Taori et al., 2023), specifically its cleaned versionhttps://huggingface.co/datasets/yahma/alpaca-cleaned. This dataset comprises 51K instructions and demonstrations and is suitable for instruction-tuning. The cleaned version corrects multiple issues such as hallucinations, merged instructions, and empty outputs. We train for one epoch, preceded by a learning rate sweep.
We evaluate finetuned models on MT-Bench (Zheng et al., 2023), by generating model responses to a pre-defined set of 80 multi-turn questions and subsequently evaluating these using GPT-4 (OpenAI, 2023). GPT-4 reviews the answers and assigns a quantitative score on a scale of 10 to each response. We present the average scores alongside the number of trainable parameters in Table 4.
We find that despite the 100x reduction in the number of trainable parameters, our method closely matches the performance of LoRA-based finetuning.
4 Image Classification
To evaluate the method on the image classification task, we adapt Vision Transformer (ViT) (Dosovitskiy et al., 2021), Base and Large variants, on datasets - CIFAR100 (Krizhevsky, 2009), Food101 (Bossard et al., 2014), Flowers102 (Nilsback & Zisserman, 2008), and RESISC45 (Cheng et al., 2017). For each dataset we train on a subset of 10 samples per class, and evaluate on the full test set (CIFAR100, Food101, Flowers102) or on all the remaining samples (RESISC45). We use weights of ViT models pretrained on the ImageNet-21k (Deng et al., 2009) dataset.
We evaluated LoRA and VeRA methods applied on the query and value layers of ViT, along with two baselines - fully-finetuned model (referred to as Full), and training the classification head only (referred to as Head). Similarly to the GLUE benchmark, we use rank 8 for LoRA, and rank 256 for VeRA. We tuned learning rates for all methods and reported results after 10 epochs in Table 5. The reported parameter count excludes the classification head, which has to be trained in all methods.
We find that VeRA approaches performance of LoRA on the Base model for three datasets and outperforms it for Flowers102, despite using over 10x fewer trainable parameters. For ViT-Large, it outperforms LoRA for three datasets: CIFAR100, Flowers102 and RESISC45.
5 Scaling the Number of Trainable Parameters
Finally, we investigate the trade-offs involved in parameter scalability for both LoRA and our method using the RoBERTalarge model on the RTE task from the GLUE benchmark. We use a set of ranks for VeRA and for LoRA, and observe the trade-off between trainable parameters and the accuracy. We replicate each configuration five times for different random seeds, and report the median of results. For LoRA, we employ the HuggingFace PEFT (Mangrulkar et al., 2022) implementation, adhering to the hyperparameters specified in Hu et al. (2022). Our own method uses the same hyperparameters as employed in the RTE experiments from the previous subsection. The results, depicted in Figure 3, reveal that our method is significantly more parameter-efficient. Notably, when the higher-rank VeRA has the same number of parameters as standard LoRA, it outperforms LoRA by 4 accuracy percentage points.
6 Ablation Study
In this section, we conduct an ablation study to examine the impact of individual components of our method. All subsequent experiments focus on the MRPC and RTE tasks and utilize the RoBERTalarge model. We adhere to the hyperparameters used in previous experiments, modifying only the component under investigation for each test. Each experiment is run with 5 random seeds, and we report the mean and standard deviation of the results.
We first investigate the necessity of both the and scaling vectors in our method. We create two ablation setups: one that excludes (termed as only ) and another that omits (termed as only ). In the only setup, is initialized with zeros. As shown in Table 6, omitting either scaling vector compromises performance. The only configuration performs slightly better than its only counterpart. This disparity in performance underscores the higher expressiveness of the scaling vector over the vector. Specifically, modulates the rows of both low-rank matrices, thereby influencing a broader aspect of the final constructed matrix. In contrast, only scales the rows of the final matrix resulting from the product of the low-rank matrices.
Initialization of Shared Matrices
We examine three different initialization schemes for the shared matrices: Kaiming normal, Kaiming uniform, and uniform initialization within the range . As per the results in Table 6, both Kaiming initializations outperform the uniform range initialization, with uniform variant having slightly better results than the normal one.
Initialization of Scaling Vector
We further explore the impact of the initialization values for the vector. Experiments are conducted with set at , , and . The results in Table 6 show that the choice of significantly influences the method’s performance; in the settings we examined, values and outperformed , potentially offering more flexibility in the optimization process through early sign changes in selected rows of the frozen matrices.
Magnitude of Adaptation
In Figure 3 we provide a visualisation of the magnitude of the changes of the vectors after finetuning on RTE task. Because the low-rank frozen matrices remain the same for each layer, we can directly compare the length of the vector across layers to account for its relative adaptation. Overall, we find that the largest adaptation happens for query matrices compared to the value ones, indicating a larger need or ease for finetuning a model there. Furthermore, similar to previous efficient adaptation methods’ findings (Zhang et al., 2023b; Liu et al., 2021) we also observe a higher adaptation for the later layers compared to earlier ones.
Sharing Random Matrices
We conduct experiments on RTE, MRPC, CoLA, and STS-B tasks to assess the impact of sharing random matrices on the performance. We evaluate two setups - one with random matrices shared across all adapted layers, and another with uniquely generated ones. Results in Table 7 show that the mean performance is identical in case of tasks RTE and STS-B, and there is a slight improvement for MRPC and CoLA when using unique matrices.
Conclusion
In this work, we introduce a finetuning method that significantly reduces the number of trainable parameters compared to LoRA, yielding similar or better results on downstream tasks. Specifically, it achieved ten-fold reduction in parameters yielding the same performance on the GLUE benchmark for RoBERTalarge, ten-fold reduction on image classification tasks, and three-fold reduction on the E2E benchmark. This method is particularly well-suited for scenarios that require frequent swapping of numerous finetuned models, such as cloud-based AI services personalized for individual users. Due to the minimal size of the scaling vectors, many versions can reside in the limited memory of a single GPU, thus substantially improving serving efficiency and removing the bottleneck of loading specific models into memory.
While the current study focuses on language and vision models with Transformer architecture, the applicability of the method across different architectures and domains remains an area for future research. Moreover, the performance of the method may benefit from additional refinements, such as dynamic parameter budget allocation, or different initialization and regularization techniques.
This work is financially supported by Qualcomm Technologies Inc., the University of Amsterdam and the allowance Top consortia for Knowledge and Innovation (TKIs) from the Netherlands Ministry of Economic Affairs and Climate Policy. We also acknowledge the use of the National Supercomputer Snellius and Distributed ASCI Supercomputer 6 (Bal et al., 2016) for essential computational tasks.
References
Appendix A Hyperparameters
In Table 8, we provide the hyperparameters used for the GLUE benchmark in the main paper. Note that due to our academic compute we were not able to run full grid searches on any hyperparameters. We only evaluated different learning rates and number of epochs and even relied on existing configurations of LoRA (Optimizer, Warmup ratio, LR schedule).
Appendix B Relative performance gain.
Figure 4 quantifies the efficiency of each method in terms of performance gains per 1K trainable parameters. For a focused comparison, we select the RTE task and RoBERTalarge model.
To establish a baseline, we conduct auxiliary experiments where only the classification head is trained while the remainder of the model is frozen. This baseline is constructed using the same hyperparameters as in our VeRA method. We then evaluate the performance gain attributable to each method, normalized by the additional trainable parameters introduced, relative to the baseline. The results clearly show that VeRA yields the highest performance gain per 1K trainable parameters.
Appendix C Impact on training time and memory usage
To evaluate the training time and GPU memory benefits of our method, we conducted a comparison between LoRA and VeRA while fine-tuning LLaMA 7B with the same rank (64) on instruction tuning dataset, introduced earlier in this work. The results are summarized in Table 12:
While VeRA includes more operations than LoRA because of the additional vector multiplies in the forward pass, we find that it only results in a modest 1.8% increase in training time. For the GPU memory, we observe a 7.4% reduction in memory usage with VeRA, as it does not require storing optimizer states and gradients for shared random matrices.
Appendix D Similarities of trained weights
We compared the weights trained with LoRA and VeRA at a single rank of 64 across all query layers. For each method and adapted layer, we constructed a weight difference. In LoRA’s case, this involved the multiplication of two low-rank matrices, while for VeRA, it also included multiplication by scaling vectors. We then calculated the cosine similarity of these flattened weights. Additionally, we compared the similarity between trained LoRA weights and randomly initialized matrices as a baseline: We find that similarities of VeRA to LoRA are on average 2e-3 while LoRA to random matrices is -8e-5.
In Figure 5 we can see a notable increase in similarity between the trained weights, particularly in the latter layers. This observation aligns with our earlier findings (Figure 3) that the highest adaptation occurs in these layers. These results support the notion that VeRA can approximate the weights trained with LoRA.
Appendix E Expressivity of VeRA
We conducted an experiment on the expressivity of LoRA and VeRA on the task of fitting random square 10x10 matrices, with results seen in Figure 6. For given number of trainable parameters, both methods perform equally well, with VeRA providing more flexibility, e.g. by allowing for much lower parametrization - below LoRA’s rank 1.
Appendix F Instruction-tuning with Vicuna Eval
Results and samples from evaluation of instruction tuned Llama 7B model with Vicuna Eval (Chiang et al., 2023), predecessor of MT-Bench. The model has been finetuned on a 10K subset of cleaned Alpaca dataset.