Tied-Lora: Enhancing parameter efficiency of LoRA with weight tying

Adithya Renduchintala, Tugrul Konuk, Oleksii Kuchaiev

Introduction

Large language models (LLMs) are increasingly employed as the fundamental component in a wide array of applications that require proficiency in Natural Language Processing (NLP) tasks. A key contributor to this widespread adoption is the capacity to fine-tune pretrained LLMs for downstream tasks in a parameter-efficient manner. This fine-tuning procedure enables the development of specialized language models that excel on particular domains and tasks. Despite smaller training data sizes compared to pretraining, finetuning still demands considerable computational resources, particularly for large models containing billions of parameters. Furthermore, it is often desirable to maintain several different customization for each combination of user ×\times task of the service, many of which are often served simultaneously. As the number of users and tasks per user grow, so do the costs associated with customization. Thus, making efficient use of customizable parameters paramount.

Low-rank Adaptation (LoRA) (Hu et al., 2021) has emerged as a popular parameter-efficient finetuning (PEFT) method because of its straightforward implementation and the ability to merge LoRA weights into the base model. However, despite its advantages, LoRA training can still be expensive, especially as the base models become increasingly larger. While prior work has attempted to make LoRA more parameter efficient, they concentrated on appropriate low-rank selection. However, we introduce a novel approach, Instead of controlling the number of parameters by the rank, we employ simple weight tying coupled with selective training. By integrating these two core ideas, we propose a range of Tied-LoRA configurations and study the performance of each configuration on five diverse customization tasks.

We propose a range of Tied-LoRA configurations that use simple weight tying in LoRA along with selective training to boost the parameter efficiency of LoRA.

We study this spectrum of possible Tied-LoRA configurations on diverse tasks that resemble real-world customization problems.

Based on the results of our study, we propose the specific vB\faChainuA\faChain\mathbf{v}\mathbf{B}_{{}_{\text{\faChain}}}\mathbf{u}\mathbf{A}_{{}_{\text{\faChain}}} configuration as the best option for maintaining performance while reducing parameters by 87%~{}87\%.

Related Work

Recent work on PEFT of pretrained language models has shown competitive capabilities, often matching full fine-tuning performance for task-specific model customization while utilizing significantly fewer trainable parameters (Houlsby et al., 2019; Lin et al., 2020; Pfeiffer et al., 2021; Rücklé et al., 2021; Liu et al., 2022).

Low-Rank adaptation (LoRA).

One of the most popular PEFT techniques is LoRA, introduced by Hu et al. (2021). LoRA employs low-rank matrix approximations of full weights’ gradient-descent (GD) update to significantly reduce the number of trainable parameters. Importantly, LoRA can incorporate the low-rank updates into the frozen base weights after the fine-tuning process, avoiding any inference speed penalties or model architecture changes. In summary, LoRA paves the way for efficient fine-tuning for task-specific customization of large models with minimal computational overhead and no changes to the model’s architecture.

Extensions to LoRA.

Since its arrival, there have been several efforts to improve the LoRA method. These methods mostly concentrated around reducing the trainable parameters and memory footprint while increasing the performance of the method on downstream tasks. AdaLoRA (Zhang et al., 2023) introduces dynamic rank adjustment for the low-rank matrices during the fine-tuning process. The fundamental premise of this extension is to optimally distribute the parameter budget over model layers. Chavan et al. (2023) combined the adapter tuning with LoRA to derive a generalized framework that utilized both methods for increased flexibility and capability across a wide variety of tasks and datasets. Kopiczko et al. (2023) proposes the VeRA method the freezes randomly initialized projection matrices and introduces trainable scaling vectors that vary across layers. This method shows similar performance to the \color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{v}\mathbf{B}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{u}\mathbf{A}(LoRA) method while dramatically reducing the number of trainable parameters. Our work draws significant inspiration from the principles of the VeRA method.

Tangential to the efforts that aim to reduce trainable parameters, QLoRA (Dettmers et al., 2023), significantly reduces the memory usage of LoRA using a 4-bit or 8-bit quantized base language model during training. The method provides algorithms and custom kernels to backpropagate gradients through the frozen, quantized base model to update low-rank matrices during training, resulting in considerable reduction in memory usage. Combining quantization and reduction in the number of trainable parameters is a direction of future work.

Weight tying.

Weight tying (Press and Wolf, 2017) is a common approach that reduces the number of parameters by using the same set of weights for both the input word embedding layer and the output word embedding layer (sometimes referred to as the language model head). In this study, we apply weight tying to the low-rank weight matrices used in LoRA, and share them across the layers of the base language model. This simple procedure leads to efficient training methods where the number of trainable parameters are either unaffected by, or only increases marginally with the number of hidden layers. As models get deeper this approach naturally provides greater parameter reduction over original LoRA method.

Method

In this section, we introduce tied \color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{v}\mathbf{B}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{u}\mathbf{A}(LoRA) , a generalized paradigm for parameter-efficient fine-tuning of large language models through low-rank weight-update approximations. Our framework offers a range of training strategies through a series of design choices over selective parameter training and weight tying, including some of the existing PEFT methodologies available in the literature. Specifically, we use weight tying alongside pairs of projection matrices and scaling vectors that can be selectively either trained or frozen. As the low-rank computation path does not introduce any non-linearity, all Tied-LoRA configurations can be merged into the base model weights to preventing additional latency during inference. Table 1 provides an overview of the scenarios we study.

The overall structure of the tied LoRA framework can be seen in Figure 1. Note that the original LoRA (Hu et al., 2021) uses a dedicated pair of low-rank projections for each of the Q,K,VQ,K,V matrices. However, in our formulation, WW is a d×3dd\times 3d matrix that jointly projects Q,KQ,K, and VV attention matrices, where dd is the hidden size of the base language model. Therefore, our down projection AA is a d×rd\times r shaped matrix and up projection matrix BB has shape r×3dr\times 3d, where rr is the low-rank bottleneck dimension. Essentially, the down projection AA is shared by Q,KQ,K, and VV, leading to fewer trainable parameters (4dr4dr) than the original LoRA (6dr6dr).

For a linear layer with a frozen pretrained weight matrix WW, we define the layer output as

where ΔW\Delta W is the full-rank update matrix, α\alpha is a scaling factor, AA and BB are low-rank projection matrices, and Λu\Lambda_{u} and Λv\Lambda_{v} are diagonal matrices with diagonal elements given by uu and vv, respectively. Herein, ΛvBΛuAx\Lambda_{v}B\Lambda_{u}Ax is the low-rank approximation to the parameter update matrix ΔW\Delta W. Unlike the original LoRA, where α\alpha is a hyper-parameter that can be manually set, we simply set α=r\alpha=r, effectively removing its scaling effect.

Equation 1 is a generalized formulation for methods that utilize low-rank approximations to estimate parameter updates. Particular settings of parameter updates and weight tying reduces this equation to some of the existing formulations in the literature. Setting and freezing Λu=Λv=I\Lambda_{u}=\Lambda_{v}=I and untying AA and BB results in LoRA:

Similarly, randomly initializing AA and BB matrices and tying them across all layer leads the the VeRA formulation (Kopiczko et al., 2023):

2 Weight Tying

The third column of Table 1 presents representations for number of trainable parameters each Tied-Lora configuration requires. As is apparent from the table, weight tying is a critical ingredient of our proposed approach which drastically reduces the number of trainable parameters. For example, \color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{v}\mathbf{B}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{u}\mathbf{A}(LoRA) training using the 7B LLaMA-2 (Touvron et al., 2023) language model with a typical low rank setting of 88 requires ∼4.2\sim 4.2M trainable parameters. By merely introducing weight tying across the 3232 layers of this model reduces the trainable parameters to ∼131\sim 131K, which is a 96.875%96.875\% reduction. In comparison, the Vera method results in a reduction of 90.6%90.6\%.

3 Selective Training

Through the flexible framework that equation 1 offers, we are given the opportunity to investigate a range training configurations. By selectively updating the components A,B,uA,B,u, and vv during the training process, we can generate a variety of methodological variations. These variations not only exhibit differences in parameter count, but they also demonstrate distinct capabilities across a variety of tasks. This exploration allows us to investigate the intriguing regime of extremely low-parameter and low-rank PEFT models. This is a key step towards the customization of models, enabling them to excel at specific tasks while maintaining a minimal parameter count. Our ultimate goal here is to harness the power of this methodology to create highly efficient, task-specific models that achieve high performance with reduced complexity.

Experiments

We now turn to evaluating the different configurations possible within our Tied-LoRA paradigm. While \color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{v}\mathbf{B}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{u}\mathbf{A}(LoRA) and PEFT methods can be used to train models for general instruction following (Sun et al., 2023; Lermen et al., 2023; Sun et al., 2023), we focus our evaluations in a “task customization” perspective, where each model is trained on a specific task and is evaluated on a test set from the same task.

To evaluate the performance of each Tied-LoRA configuration across diverse data settings, we utilized the following types of tasks:

is a common task where the model is expected to “read” some relevant text (the context) and answer questions. The answers are usually exact sub-strings from the provided context. We use SQuADv1 dataset (Rajpurkar et al., 2016) in our experiments. Since the official test split of this dataset does not contain ground-truth answers, we use the validation set as our test set. We create a validation set comprising of a random sample of 48004800 examples extracted from the training set.

Summarization

is a central problem in NLP and several variations of summarization datasets have been proposed. We employ the DialogSum dataset (Chen et al., 2021) to study our models’ performance on this task. DialogSum includes summaries of real-word conversations on a diverse set of topics and scenarios. This dataset was an attractive option as the length of the conversations and summarizes are within the context lengths (40964096 tokens) of the base language models.

Commonsense Natural Language Inference (NLI)

is a task designed to probe the ability of language models to apply “commonsense reasoning” to choose a possible ending for a given situation described in natural language. These tasks are typically trivial for humans but language models can still struggle. We use the HellaSwag dataset (Zellers et al., 2019) to study the performance of our proposed models on this type of task. As HellaSwag contains multiple-choice questions, it can be viewed as a classification problem.

Translation

Machine translation is a natural language generation task which is widely used in research and industry. Translation is inherently multilingual and thus offers a challenging domain to study our Tied-LoRA paradigm. There are several large scale translation datasets but we focus on a moderately sized IWSLT 2017 German-to-English translation dataset (Cettolo et al., 2017). The dataset contains translation of spoken language into various other natural languages. With over 206k206k training examples this is the largest dataset that we study.

Mathematical Reasoning

is a challenging domain where large language models still lag behind human performance. Using PEFT methods on such tasks further amplifies these challenges as there are very few trainable parameters. In our experiments, we use the GSM8K benchmark (Cobbe et al., 2021) which contains 8.58.5K high-quality, grade-school level math word problems. Each example in the GSM8K benchmark contains a question and an answer. The answers are provided with natural language solutions which contain explanations of each step used to obtain the final answer. The final numerical answer is demarcated from the rest of the natural language solution. We evaluate our models by comparing these final numerical answers.

[] \color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{v}\mathbf{B}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{u}\mathbf{A}(LoRA) \mathbf{v}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{B}_{{}_{\text{\faChain}}}\mathbf{u}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{A}_{{}_{\text{\faChain}}}(Vera) \color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{v}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{B}_{{}_{\text{\faChain}}}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{u}\mathbf{A}_{{}_{\text{\faChain}}} \mathbf{v}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{B}_{{}_{\text{\faChain}}}\mathbf{u}\mathbf{A}_{{}_{\text{\faChain}}} \color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{v}\mathbf{B}_{{}_{\text{\faChain}}}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{u}\mathbf{A}_{{}_{\text{\faChain}}} vB\faChainuA\faChain\mathbf{v}\mathbf{B}_{{}_{\text{\faChain}}}\mathbf{u}\mathbf{A}_{{}_{\text{\faChain}}} \color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{v}\mathbf{B}_{{}_{\text{\faChain}}}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{u}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{A}_{{}_{\text{\faChain}}} \color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{v}\mathbf{B}_{{}_{\text{\faChain}}}\mathbf{u}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{A}_{{}_{\text{\faChain}}}

2 Base Language Models

Although PEFT enables the base language model to perform new tasks, the final performance heavily depends on the inherent abilities learned during pretraining. This necessitates investigating the performance of Tied-LoRA on multiple base models with different inherent capabilities. Therefore, we use a relatively small two billion parameter, GPT-2B-001 modelhttps://huggingface.co/nvidia/GPT-2B-001 released by NVIDIA and the moderately large 77B LLaMA 2 model (Touvron et al., 2023) released by Meta.

In addition to the size differences, these models also differ in the amount of pretraining data used. The GPT-2B-001 model was trained on 1.11.1 trillion tokens of text from publicly available multilingual text spaning 5353 languages. The LLaMA2 77B model was trained on 22 trillion tokens of predominately English text. Both models are auto-regressive language models with a context size of 40964096 tokens.

3 Implementation Details

We use the open-source NeMo Framework to implement all the algorithms presented in this paper. Our implementation is publicly available through the NeMo GitHub repository.https://github.com/NVIDIA/NeMo/tree/adithyare/vera All training routines were run for 2k2k max steps, but training was terminated sooner using early stopping with a patience of 1010 to prevent overfitting. We trained all configurations using AdamW optimizer (Loshchilov and Hutter, 2017) with a weight decay of 0.010.01 and a cosine learning rate schedule with 5050 warm-up steps. For each Tied-Lora method we tried two learning rates, a high rate of 1e−41e-4 and a low learning rate of 1e−51e-5. While the “typical” range of the low-rank dimension rr is 4−164-16 we find that some complex tasks benefit from higher rr so we trained all our models with a wide range of r∈{2,8,32,128}r\in\{2,8,32,128\}. Each task was trained with a global batch size of 256256 and a validation check interval of 3030 steps. The only exception was the IWSLT translation dataset for which we set global batch size and validation check interval of 10241024 and 6060 respectively. No extensive hyper-parameter search was conducted.

During inference, we used greedy-decoding to generate the models’ predictions with a limit of 500500 tokens.

Results

Table 2 shows average scores attained by each Tied-Lora configuration over the 55 tasks, per low-rank value. We can immediately see that \color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{v}\mathbf{B}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{u}\mathbf{A}(LoRA) is the best performing model for both the 2B and 7B base language models. This is hardly surprising as \color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{v}\mathbf{B}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{u}\mathbf{A}(LoRA) is the most expensive method which does not use tied weights. With this in mind we see that vB\faChainuA\faChain\mathbf{v}\mathbf{B}_{{}_{\text{\faChain}}}\mathbf{u}\mathbf{A}_{{}_{\text{\faChain}}} is a consistently the next best performing method with average scores comparable to \color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{v}\mathbf{B}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{u}\mathbf{A}(LoRA) , demonstrating the efficacy of weight tying. \color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{v}\mathbf{B}_{{}_{\text{\faChain}}}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{u}\mathbf{A}_{{}_{\text{\faChain}}} however does not perform as well suggesting that the scaling vectors u\mathbf{u} and v\mathbf{v} provide an additional boost in performance especially as the rank rr is increased to 128128 (at the cost of more trainable parameters).

Next best Tied-Lora configuration is \mathbf{v}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{B}_{{}_{\text{\faChain}}}\mathbf{u}\mathbf{A}_{{}_{\text{\faChain}}} which obtains third place for 66 out of the 88 settings shown in Table 2. Note that \mathbf{v}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{B}_{{}_{\text{\faChain}}}\mathbf{u}\mathbf{A}_{{}_{\text{\faChain}}} beats other Tied-LoRA methods which use more parameters. Interestingly, \mathbf{v}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{B}_{{}_{\text{\faChain}}}\mathbf{u}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{A}_{{}_{\text{\faChain}}}(Vera) which uses fewer parameters than \color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{v}\mathbf{B}_{{}_{\text{\faChain}}}\mathbf{u}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{A}_{{}_{\text{\faChain}}} and \color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{v}\mathbf{B}_{{}_{\text{\faChain}}}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{u}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{A}_{{}_{\text{\faChain}}} has better performance. \color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{v}\mathbf{B}_{{}_{\text{\faChain}}}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{u}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{A}_{{}_{\text{\faChain}}} and \color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{v}\mathbf{B}_{{}_{\text{\faChain}}}\mathbf{u}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{A}_{{}_{\text{\faChain}}} does the worst in most cases, especially when rr is increased.

Figure 2 shows the performance for each task individually. We see that for tasks like HellaSwag and SQuAD Tied-LoRA methods (\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{v}\mathbf{B}_{{}_{\text{\faChain}}}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{u}\mathbf{A}_{{}_{\text{\faChain}}} and vB\faChainuA\faChain\mathbf{v}\mathbf{B}_{{}_{\text{\faChain}}}\mathbf{u}\mathbf{A}_{{}_{\text{\faChain}}} specifically) are virtually the same as \color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{v}\mathbf{B}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{u}\mathbf{A}(LoRA) over the entire range of ranks, fewer parameters. \color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{v}\mathbf{B}_{{}_{\text{\faChain}}}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{u}\mathbf{A}_{{}_{\text{\faChain}}} for example, only uses 4.2%4.2\% and 3.1%3.1\% of parameters that \color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{v}\mathbf{B}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{u}\mathbf{A}(LoRA) uses for the 2B and 7B models, respectively.

On the flip side tasks like GSM8K seem to benefit from the additional parameters provided by \color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{v}\mathbf{B}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{u}\mathbf{A}(LoRA) . A similar gap between \color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{v}\mathbf{B}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{u}\mathbf{A}(LoRA) and Tied-LoRA methods can be seen for the translation task as well especially in the 2B model. We hypothesize that tasks in which the base language model already performs well can be easily enhanced by Tied-Lora, while tasks that are not “natutal” to the base model (like mathematical reasoning) requires more parameters.

Again, we can see that in Tied-LoRA methods the addition of untied parameters uu and vv are most helpful as the rr is increased. This suggests that the untied parameters act as a per layer “adjustment” in the Tied-LoRA paradigm. We also see that it is best to either train both AA and BB or just freeze BB and train AA (with untied weights uu and vv when applicable). Lastly, we see that in the specific cases of \color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{v}\mathbf{B}_{{}_{\text{\faChain}}}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{u}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{A}_{{}_{\text{\faChain}}} and \color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{v}\mathbf{B}_{{}_{\text{\faChain}}}\mathbf{u}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{A}_{{}_{\text{\faChain}}} there is extreme instability when rr is increased. This pattern is consistent across all the tasks we studied.

Conclusion & Future Work

We have presented our Tied-Lora paradigm of extending the parameter efficiency of Lora by using simple technique of weight-tying and selective training of low-rank matrices. Our study suggests that for several tasks vB\faChainuA\faChain\mathbf{v}\mathbf{B}_{{}_{\text{\faChain}}}\mathbf{u}\mathbf{A}_{{}_{\text{\faChain}}} configuration can perform as well as Lora (over a range of low-rank dimensions) with just 13%13\% of the parameters of Lora when rr is within the typical setting of 88. Increasing to larger rr result is more aggressive reduction of trainable parameters compared to \color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{v}\mathbf{B}\color[rgb]{0.26171875,0.578125,0.765625}\definecolor[named]{pgfstrokecolor}{rgb}{0.26171875,0.578125,0.765625}\mathbf{u}\mathbf{A}(LoRA) . This is especially true for tasks which the base language model already has some abilities, such as commonsense NLI, extractive QA and summarization. Given that the baseline abilities of LLMs are consistently improving with each iteration of LLMs, we hope that our best Tied-LoRA configuration can be used as a replacement for LoRA for more tasks in the future.

References