LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Renrui Zhang, Jiaming Han, Chris Liu, Peng Gao, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, Yu Qiao

Introduction

Large-scale Language Models (LLMs) have stimulated widespread attention in both academia and industry. Driven by massive corpora and advanced hardware, LLMs exhibit remarkable understanding and generative ability, propelling language tasks into a higher level. Recently, significant progress has been made on instruction-following models, e.g., ChatGPT and GPT-3.5 (text-davinci-003) . Following instructions in natural language, they can generate professional and contextual responses in a conversational way. However, the further prevalence of instruction models is largely impeded by the closed-source restriction and high development costs.

To alleviate this, Stanford Alpaca proposes to fine-tune an LLM, i.e., LLaMA into an instruction-following model, which is affordable and replicable. Starting from 175 human-written instruction-output pairs , Alpaca leverages GPT-3.5 to expand the training data to 52K in a self-instruct manner. Supervised by this, Alpaca fine-tunes the entire 7B parameters in LLaMA, producing an exceptional instruction model that performs similarly to GPT-3.5. Despite Alpaca’s effectiveness, a complete fine-tuning of large-scale LLaMA is still time-consuming, computation-intensive, multi-modality unsupported and cumbersome to transfer to different downstream scenarios.

In this paper, we introduce LLaMA-Adapter, an efficient fine-tuning method that adapts LLaMA into a well-performed instruction-following model. We also utilize the 52K instruction-output data for training purposes, but freeze the entire LLaMA model with superior resource efficiency. Specifically, in LLaMA’s higher transformer layers, we append a set of learnable adaption prompts as prefix to the input instruction tokens. These prompts learn to adaptively inject new instructions (conditions) into the frozen LLaMA. To avoid noise from adaption prompts at the early training stage, we modify the vanilla attention mechanisms at inserted layers to be zero-initialized attention, with a learnable gating factor. Initialized by zero vectors, the gating can firstly preserve the original knowledge in LLaMA, and progressively incorporate instructional signals during training. This contributes to stable learning during the fine-tuning process and better instruction-following capacity of the final model.

Overall, our LLaMA-Adapter exhibits four main characteristics, as shown in Figure 1.

1.2M Parameters. Instead of updating the full 7B parameters, we freeze the pre-trained LLaMA and only learn the adaption prompts with 1.2M parameters on top. This, however, reveals comparable instruction-following proficiency with the 7B Alpaca.

One-hour Fine-tuning. Thanks to our lightweight adaption modules with zero-initialized gating, the training convergence of LLaMA-Adapter costs less than one hour on 8 A100 GPUs, which are three times faster than Alpaca.

Plug with Expertise. For different scenarios, it is flexible to insert their respective adapters and endow LLaMA with different expert knowledge. Thus, it suffices to store a 1.2M adapter within each context, other than a complete copy of the 7B model.

Multi-modal Instruction. Besides textual instruction, our approach can also take images as input for multi-modal reasoning. By adding image tokens into adaption prompts, LLaMA-Adapter performs competitively on ScienceQA and COCO Caption benchmarks.

In addition to instruction-following models, our zero-initialized attention can be generalized to other vision and language models for parameter-efficient fine-tuning. For vision models, we utilize our approach to fine-tune a pre-trained ViT for downstream image classification, obtaining superior performance on VTAB-1k benchmark over various image distributions. For other language models, we evaluate our fine-tuning efficacy on ReBERTa for extractive question answering, which achieves leading results on SQuAD v1.1 and v2.0 benchmarks. By these experiments, we demonstrate the effectiveness of LLaMA-Adapter for traditional vision and language tasks.

Related Work

The subfield of language models learning instruction-following capabilities aims to generate responses based on natural language commands, which have been extensively researched in language , and multi-modality domains. These methods normally enhance the pre-trained LLMs by fine-tuning them using high-quality instruction-output data pairs. Such fine-tuning process boosts the model to better comprehend user intentions and follow instructions more accurately. Therein, FLAN introduces an instruction tuning method that outperforms non-tuned LLMs in unseen tasks. PromptSource provides a development environment with a web-based GUI, which creates and manages natural language prompts for zero-shot and gradient-based few-shot learning. SUP-NATINST establishes a large benchmark of 1,616 diverse language tasks, and adopts a multi-task training on the T5 model. InstructGPT demonstrates significant improvement of the instruction-following power, and is probably integrated into the closed-source GPT-3.5 and GPT-4 . Stanford Alpaca fine-tunes all the 7B parameters of an LLM, i.e., LLaMA in an end-to-end manner, which is open-source and replicable. However, this full-model fine-tuning can be inefficient in both time and memory, limiting its transferability to downstream applications. In contrast, our LLaMA-Adapter aims to fine-tune only lightweight adapters on top of the frozen LLaMA, other than updating parameters of the entire model. Compared to a concurrent work Alpaca-LoRA , our approach further reduces the computational demands, and can be generalized to follow visual instructions for multi-modal reasoning.

Parameter-Efficient Fine-Tuning.

The pre-training and fine-tuning paradigms have been proven to be highly effective in different language and vision tasks. Compared to full fine-tuning, Parameter-Efficient Fine-Tuning (PEFT) methods freeze most parameters of pre-trained models, and can still exhibit comparable capabilities on downstream tasks. Various PEFT techniques have been explored, including prompt tuning , Low-Rank Adaptation (LoRA) , and adapters . Prompt tuning appends a collection of trainable prompt tokens to pre-trained large models, which are inserted either to the input embeddings only , or to all of the intermediate layers . LoRA introduces trainable rank decomposition matrices into each network weights , which have indicated promising fine-tuning ability on large generative models . Adapters insert lightweight adaption modules into each layer of the pre-trained transformer and have been extended across numerous domains . In this paper, we propose a new PEFT method, LLaMA-Adapter, specially designed for LLaMA and instruction-following fine-tuning. Existing PEFT methods might potentially disturb the pre-trained linguistic knowledge by directly inserting randomly initialized modules. This leads to unstable fine-tuning with large loss values at early training stages. To this end, LLaMA-Adapter adopts a zero-initialized attention with gating factors to well mitigate such a issue, which progressively incorporates the instructional cues with the frozen LLaMA. Moreover, we verify the effectiveness of our approach to fine-tune large models in other domains. Aided by the adaption prompts with zero gating, our efficient fine-tuning of ViT and RoBERTa exhibit competitive downstream performance respectively on vision and language tasks, demonstrating superior generalization capacity.

LLaMA-Adapter

In Section 3.1, we first introduce how to insert the learnable adaption prompts into LLaMA’s transformer. Then, we present the details of zero-initialized attention mechanisms with zero gating in Section 3.2, and generalize LLaMA-Adapter for multi-modal reasoning in Section 3.3. Finally, we extend our approach for efficient fine-tuning of vision and vision-language models in Section 3.4.

In this way, the instruction knowledge learned within PlP_{l}, can effectively guide TlT_{l} to generate the subsequent contextual response via attention layers in the transformer block.

2 Zero-initialized Attention

Then, the attention scores of QlQ_{l} and KlK_{l} before the softmax function are calculated as

which records the feature similarities between the new word tlt_{l} and all K+M+1K+M+1 tokens. Meanwhile, SlS_{l} can be reformulated by two components as

To this end, we adopt a learnable gating factor, denoted as glg_{l}, to adaptively control the importance of SlKS_{l}^{K} in the attention. Initialized by zero, glg_{l} can firstly eliminate the influence of under-fitted prompts, and then increase its magnitude for providing more instruction semantics to LLaMA. Therefore, we independently apply the softmax functions to the two components in Equation (6), and multiply the first term by glg_{l}, formulated as

The separate softmax functions ensure the second term to be irrelevant to the adaption prompts. When glg_{l} is close to zero, it can mostly convey the originally pre-trained knowledge of LLaMA to token tlt_{l} for a creditable generation. In practice, we adopt multiple glg_{l} to be independently learned for different heads within the attention, benefiting the learning diversity of multi-head mechanisms.

Finally, we calculate the output of the ll-th attention layer with a linear projection layer as

With our proposed zero-initialized attention, the adaption prompts can progressively inject the newly acquired instructional signals into the transformer, while simultaneously incorporating the pre-trained knowledge of LLaMA to provide high-quality responses.

3 Multi-modal Reasoning

Apart from textual instructions, LLaMA-Adapter is capable of answering a question based on input of other modalities, which augments the language model with rich cross-modal information. As shown in Figure 3, we take the ScienceQA benchmark as examples, which is analogous to the COCO Caption dataset . Given visual and textual contexts, along with the corresponding question and options, the model is required to conduct multi-modal understanding to give the correct answer.

where PlvP_{l}^{v} denotes the adaption prompt incorporating visual information from the given image context. In this way, LLaMA is fine-tuned to generate responses conditioned on vision-language inputs, and can tackle more challenging generative tasks with multi-modal understanding.

4 Zero-initialized Attention for other Large Models

Our approach, i.e., adaption prompts with zero-initialized attention, is not limited to the domain of instruction models, and can be further utilized to fine-tune large models in traditional vision and language tasks, exerting superior generalization capacity.

We select a pre-trained ViT as the foundation vision model for downstream image classification tasks. Similar to LLaMA, we insert the adaption prompts as prefix into the topmost LL transformer layers in ViT, and modify the attention operations to be zero-initialized at all inserted layers. By increasingly injecting the downstream visual semantics, we only introduce a few parameters on top of the frozen ViT, and attain comparable classification accuracy to full fine-tuning on VTAB-1k benchmark, which indicates our attention operator’s efficacy in vision domains.

Language Models.

We utilize RoBERTa pre-trained on large-scale unlabeled text corpus, and evaluate our proposed zero-initialized attention on SQuAD benchmark for extractive question answering. We implement the zero-initialized attention on top of P-tuning v2 , a prompt tuning method for efficiently adapting large language models. Likewise, we only enable the prompt tokens in P-tuning v2 and our zero gating factors to be learnable during fine-tuning. The leading results demonstrate our superiority for traditional language tasks. Please refer to Supplementary Material for applying zero-initialized attention mechanisms to more large models and tasks.

Experiment

In Section 4.1, we first evaluate the instruction-following capacity of LLaMA-Adapter. Then, we present our multi-modal performance on ScienceQA benchmark in Section 4.2, and conduct ablation study on ScienceQA’s validation set in Section 4.3. Finally, we report the fine-tuning results of our approach on other vision and language models in Section 4.4.

Following Stanford Alpaca , we utilize 52K instruction-following data for training, which is extended from 175 instruction-output pairs . We fine-tune LLaMA-Adapter on 8 A100 GPUs for 5 epochs. The warmup epochs, batch size, learning rate, and weight decay are set to 2, 64, 0.009, and 0.02, respectively. By default, we utilize the pre-trained LLaMA model with 7B parameters and N=32N=32 transformer layers. We adopt a prompt length K=10K=10 and insert the adaption prompts into the last L=30L=30 layers. In the generation stage, we adopt top-p sampling as the default decoding method with a temperature 0.10.1 and a top-p=0.75\textit{top-p}=0.75. For quantitative evaluation , we ask GPT-4 to assess the response quality of instruction-following models on 80 questions. Since we observed that GPT-4 has a preference to give higher scores to the first response in comparison, we also switch the position of two responses, resulting in a total of 160 evaluation items.

Performance.

We compare the generated responses of LLaMA-Adapter and Alpaca in Figure 4, and report the quantitative results in Figure 4.1. Please refer to Supplementary Material for a full comparison with Alpaca-LoRA , GPT-3 , and LLaMA-I . For different kinds of instructions in Figure 4, our approach can output reasonable responses comparable to the fully fine-tuned Alpaca, including question answering, language translation, and code generation. For the GPT-4 evaluation in Figure 4.1, LLaMA-Adapter obtains more ‘win’ compared to Alpaca and Alpaca-LoRA. This fully demonstrates the effectiveness of our adapters with zero-initialized attention mechanisms.

Efficiency.

In Table 4.1, we compare the learnable parameters, storage space, and training time of different instruction-following methods. As a lightweight plug-and-play module, LLaMA-Adapter enjoys superior training efficiency with only 1.2M parameters, 4.9M storage, and one-hour training. This enables us to fine-tune large-scale language models, e.g., LLaMA, on mobile devices. LLaMA-Adapter’s efficiency advantages can be further revealed by multi-node training, since only the gradients of 1.2M parameters are required to be transferred among nodes, other than Alpaca’s 7B.

2 Multi-modal Evaluation

For the multi-modal LLaMA-Adapter, we adopt CLIP’s visual encoder to extract the multi-scale global features of input images, and leverage simple cascaded MLPs as the learnable projection network. We adopt greedy search as the decoding method for generation, and keep other hyperparameters the same as the instruction-following LLaMA-Adapter. Two multi-modal datasets are utilized to train our model and evaluate the performance: ScienceQA and COCO Caption . ScienceQA is a large-scale multi-modal science question answering dataset collected from various knowledge domains. Each example contains a visual context, a textual context, a question, multiple options, and an answer. We concatenate the given question, textual context, and options sequentially in one sentence as LLaMA-Adapter’s input. COCO Caption dataset contains 0.6M training image-caption data (120k images with 5 captions per image) over a wide range of distributions. We utilize “Generate caption for this image” as the textual instruction input for LLaMA-Adapter.

Performance.

In Table 2, we compare LLaMA-Adapter with existing popular VQA methods and large language models on ScienceQA datset. As shown, our single-modal variant (‘LLaMA-AdapterT’) attains 78.31% accuracy with only 1.2M parameters. By further injecting visual conditions with a 0.6M projection network, our multi-modal variant (‘LLaMA-Adapter’) exhibits a improvement of +6.88% answering accuracy. Compared to traditional VQA methods, they are required to train the entire network by in-domain datasets with considerable resource budget, while LLaMA-Adapter only fine-tunes a few parameters with better performance. Despite the GPT series achieving zero-shot answering without fine-tuning, they contain much more parameters than our LLaMA 7B model with lightweight adapters. Besides, MM-CoT is on par with our approach, but it highly relies on a complex two-stage inference. Therefore, our LLaMA-Adapter demonstrates superior parameter efficiency while achieving competitive question answering capacity. In Table 5, we report the results of image captioning on COCO Caption dataset. Both BLIP and BLIP-2 adopt a costly pre-training stage on additional datasets for superior performance, including Visual Genome , Conceptual Captions and LAION . In contrast, our LLaMA-Adapter only requires COCO Catption’s training set of 0.6M data and attains better accuracy than ClipCap .

3 Ablation Study

We first investigate the number of transformer layers to be inserted in LLaMA-Adapter. As shown in Table 4.2, increasing the layer numbers introduces more parameters, but leads to a large improvement in the accuracy of ScienceQA’s validation set, e.g., +17.41% from 10 to 30, and +10.49% from 20 to 30. It indicates that more adaption prompts at different layers can provide stronger task-specific guidance to the pre-trained LLaMA.

Zero-initialized Attention.

Our proposed attention mechanism is essential for the early-stage training stability and final generation capacity of LLaMA-Adapter. As shown in Table 4.2, it contributes to a significant +43.08% performance gain on the validation set. In contrast, the randomly initialized baseline only achieves 40.77% accuracy, nearly the same as ‘Random Choice’ (see Table 2’s first row). This comparison demonstrates the decisive role of zero-initialized attention in our approach. In Figure 4.3, we plot the loss curves with and without the zero initialization, where the ‘zero-init attention’ converges faster and reaches lower loss bounds than ‘rand-init attention’.

Robustness to Over-fitting.

As the fine-tuning data of large language models is normally much smaller-scale than the pre-training data, researchers have to carefully tune a set of hyperparameters to avoid over-fitting. In Table 4.3, we show our LLaMA-Adapter is relatively robust to the over-fitting issue. Similar to the conclusion in , even if our model has over-fitted the fine-tuning data, e.g., the validation loss marginally varies from 0.136 (15 epochs) to 0.282 (60 epochs), the validation accuracy is still increasing, e.g., from 82.08% to 83.94%. This is because, LLaMA-Adapter keeps the pre-trained LLaMA 7B model frozen, and only learns lightweight adapters with a few parameters.

4 Zero-initialized Attention for other Large Models

For image classification, we fine-tune the ViT-B/16 pre-trained on supervised ImageNet-21k dataset. We adopt VTAB-1k for evaluation, which is a collection of 19 diverse visual tasks and organized into three groups according to the image domains: Natural, Specialized, and Structured. For extractive question answering, we follow P-tuning v2 (PT2) to fine-tune the RoBERTalarge model on SQuAD v1.1 and v2.0 benchmark. Exact Match (EM) and F1 scores on the dev set are reported. We defer the evaluation on the name entity recognition (NER) and the semantic role labeling (SRL) tasks to Supplementary Material.

Performance.

We present the results of fine-tuning ViT and RoBERTa in Tables 4.3 and 4.3, respectively. For three dataset groups with various image distributions, e.g., natural images, medical and satellite imagery, our approach achieves +3.26%, +2.00%, and +1.77% improvement over VPT . On both SQuAD v1.1 and v2.0 dev sets, zero-initialized attention can boost P-tuning v2 with different margins, indicating strong language understanding capability. This demonstrates our superiority on traditional vision and language tasks compared to existing fine-tuning methods.

Conclusion

In this paper, we propose LLaMA-Adapter, an efficient adaption method for training instruction-following models. With only 1.2M parameters and one-hour training, our approach effectively fine-tunes LLaMA with superior efficiency compared to the 7B-parameter Alpaca. For better training stability and final performance, we introduce zero-initialized attention with gating mechanism, which adaptively incorporates instructional signals, while preserving the pre-trained knowledge in LLaMA. LLaMA-Adapter can be generalized to image conditions for multi-modal reasoning, achieving competitive results on ScienceQA and COCO Caption benchmarks. On traditional vision and language tasks, our zero-initialized attention also attains favorable fine-tuning performance, which indicates strong generalization capacity. Limitation: as our multi-modal variant presents a generic paradigm for incorporating external semantics, we will further extend LLaMA-Adapter to serve as a unified multi-modal framework, conditioned on a wide range of instructions, such as video, audio, and point clouds. We do not foresee negative social impact from the proposed work.

Appendix A Appendix Overview

Section B: Additional experiments of zero-initialized attention.

Section C: Full comparison of instruction-following models.

Section D: Comparison of LLaMA-Adapter and LLaMA-I.

Appendix B Additional Experiments

In this section, we provide more detailed experiments and analysis of applying our zero-initialized attention to fine-tune vision models, language models, and vision-language models, respectively.

In Table 9, we compare the detailed fine-tuning results on VTAB-1k benchmark with 19 downstream visual tasks, which can be categorized into Natural (7 tasks), Specialized (4 tasks), and Structured (8 tasks), according to image domains. As shown, our zero-initialized attention outperforms VPT on most datasets (16 out of 19), and surpasses full fine-tuning along with other fine-tuning methods by large margins. This demonstrates the general efficacy of the proposed mechanism on a variety of image distributions.

B.2 More Experiments on Language Tasks

For a more comprehensive evaluation of zero-initialized attention, we fine-tune RoBERTalarge on other two natural language processing tasks in addition to extractive question answering of the main paper, which are named entity recognition (NER) and semantic role labeling (SRL). We adopt CoNLL03 , CoNLL04 , CoNLL05 , and CoNLL12 as the evaluation datasets. As shown in Table 10, equipping P-tuning V2 (PT2) with our zero-initialized attention can steadily improve the performance on all datasets with varying magnitudes, which indicates our effectiveness for different language tasks and applications.

B.3 Fine-tuning Vision-Language Models

Besides ViT and RoBERTa, we also evaluate our approach on CLIP , a vision-language model pre-trained by 400 million text-image pairs. In detail, we adopt CLIP with a ViT-B/16 as the visual encoder and a 12-layer transformer as the textual encoder. We test our fine-tuning results on base-to-novel generalization benchmark with three datasets, i.e., ImageNet , Caltech101 , and Flowers102 , where the model is trained only on the base classes in a few-shot setting and evaluated on both base and novel categories. We freeze the entire CLIP and insert the adaption prompts with zero-initialized attention into CLIP’s visual encoder. As shown in Table 11, our approach achieves the best average classification accuracy on both base and novel categories, demonstrating our fine-tuning capability for large vision-language models.

Appendix C Full Comparison of Instruction-following Models

In this section, we provide the full comparison of existing instruction-following models: Alpaca , Alpaca-LoRA , GPT-3 , and our LLaMA-Adapter. Our approach only fine-tunes 1.2M parameters within one hour, but generates responses comparable to the fully fine-tuned Alpaca and large-scale GPT-3, exhibiting a superior performance-efficiency trade-off.

Instruction: Tell me five words that rhyme with ’shock’.

LLaMA-Adapter: Flock, rock, lock, cook, and stock.

Alpaca-LoRA: Flock, clock, dock, stock, and sock.

Alpaca: Five words that rhyme with shock are: rock, pop, shock, cook, and snock.

GPT-3: 1. Block 2. Lock 3. Dock 4. Clock 5. Rock

Instruction: List all Canadian provinces in alphabetical order.

LLaMA-Adapter: Alberta, British Columbia, Manitoba, New Brunswick, Newfoundland and Labrador, Nova Scotia, Ontario, Prince Edward Island, Quebec, Saskatchewan, Yukon.

Alpaca-LoRA: Alberta, British Columbia, Manitoba, New Brunswick, Newfoundland and Labrador, Nova Scotia, Ontario, Prince Edward Island, Quebec, Saskatchewan.

Alpaca: Alberta, British Columbia, Manitoba, New Brunswick, Newfoundland and Labrador, Nova Scotia, Ontario, Prince Edward Island, Quebec, Saskatchewan, Northwest Territories, Nunavut, Yukon.

GPT-3: 1. Alberta 2. British Columbia 3. Manitoba 4. New Brunswick 5. Newfoundland and Labrador 6. Northwest Territories 7. Nova Scotia 8. Nunavut 9. Ontario 10. Prince Edward Island 11. Quebec 12. Saskatchewan 13. Yukon

Instruction: Translate the sentence ’I have no mouth but I must scream’ into Spanish.

LLaMA-Adapter: No tengo boca pero tengo que gritar.

Alpaca-LoRA: No tengo boca pero tengo que gritar.

LLaMA-Adapter: Alpacas are members of the camelid family and are native to the Andes Mountains of South America. They are typically found in herds of 10-20 animals and are known for their gentle nature and their ability to thrive in harsh climates. Alpacas are typically white or grey in color and have long, silky coats. They are also known for their ability to produce high-quality wool, which is used in the production of clothing and other products.

Alpaca-LoRA: Alpacas are members of the camelid family and are native to the Andes Mountains of South America. They are known for their soft, luxurious fleece, which is used to make clothing, blankets, and other items. Alpacas are herbivores and graze on grasses and other plants. They are social animals and live in herds of up to 20 individuals.

Alpaca: Alpacas are small, fluffy animals related to camels and llamas. They are native to Peru and Bolivia, and were first domesticated around 5,000 years ago. They are kept mainly for their fine, soft fleece, which is used to make knitwear and other garments. Alpacas are herd animals and live in small family groups, led by an older male. They are highly social animals and can form strong bonds with their owners.

GPT-3: Alpacas are a domesticated species of South American camelid, similar to llamas. They are native to the Andes Mountains in South America and are kept as livestock for their soft, luxurious wool. Alpacas are social animals and live in herds of up to 20 individuals. They are typically kept in small herds of two to five animals, and are relatively easy to care for. Alpacas are herbivores and feed on grass, hay, and other vegetation. They are also known for their gentle and friendly nature, making them popular as pets.

Instruction: Write a Python program that prints the first 10 Fibonacci numbers.

Instruction: Tell me about the president of Mexico in 2019.

LLaMA-Adapter: The president of Mexico in 2019 was Andrés Manuel López Obrador, who took office on December 1st, 2018. He is a member of the left-wing party Morena and is known for his populist policies. He has focused on tackling corruption, reducing poverty, and improving the country’s infrastructure. He has also been a vocal critic of the United States’ immigration policies.

Alpaca-LoRA: The president of Mexico in 2019 was Andrés Manuel López Obrador, who took office on December 1, 2018. He is a member of the National Regeneration Movement (MORENA) political party and is the first left-wing president of Mexico since 1946. He is known for his anti-corruption and anti-neoliberal policies, as well as his commitment to improving the living conditions of the Mexican people.

Alpaca: In 2019, Andres Manuel Lopez Obrador was sworn in as President of Mexico. He is the country’s 58th president and the first from the left-wing Morena party. He was elected in 2018 on a promise to combat corruption and reduce poverty and has faced criticism for his handling of the economy, crime and the country’s relationship with the United States. He is also known for his ambitious social welfare programs and has been praised for raising the minimum wage and providing aid to low-income families. He is seen as a champion of democracy and human rights in Mexico.

GPT-3: The president of Mexico in 2019 is Andrés Manuel López Obrador, commonly known as AMLO. He is a Mexican politician who has been the President of Mexico since December 1, 2018. He is a member of the National Regeneration Movement (MORENA) party and previously served as the Head of Government of Mexico City from 2000 to 2005. He is a left-wing populist and has promised to reduce corruption and inequality in Mexico. He has also promised to increase public spending on social programs and infrastructure projects.

Appendix D Comparison with LLaMA-I

In this section, we compare the generation quality of LLaMA-Adapter with LLaMA-I , an instruction-fine-tuned LLaMA 65B model following . All examples below are copied from the appendix of LLaMA . Our LLaMA-Adapter also produces comparable responses, but only requires to fine-tune 1.2M parameters upon the LLaMA 7B model.

Instruction: What are 3 popular chess openings?

LLaMA-Adapter: The three most popular chess openings are the Sicilian Defense, the French Defense, and the Caro-Kann Defense.

LLaMA-I: There are dozens of different chess openings that are popular among players, but some of the most widely used openings include the Sicilian Defense, the French Defense, and the Caro-Kann Defense. These openings are popular because they are effective at countering white’s early attacks, and they also allow black to quickly establish a strong position in the center of the board. Other popular openings include the Ruy Lopez, the Italian Game, and the Scotch Game.

Instruction: Write a conversation between the sun and pluto.

Pluto: Not much, just hanging out with my friends.

Sun: Well, I’m going to go now. Take care!

Pluto: Not much. Just hanging out here at the edge of the solar system.

Sun: Yeah, I’ve been meaning to talk to you about that. You’re not a planet anymore.

Pluto: What? Come on, man. That’s not fair.

Sun: I’m sorry, but it’s true. You just don’t meet the criteria anymore.

Pluto: I’ve been a planet for over 70 years!

Sun: Things change, Pluto. You’ve just got to accept it.

Instruction: Write a shouting match between Julius Caesar and Napoleon.

Julius Caesar: "You’re nothing but a petty tyrant, Napoleon! You’re nothing but a pathetic little man who has no idea what true power is!"

Napoleon: "You’re nothing but a weak and pathetic fool, Caesar! You’re nothing but a petty tyrant who has no idea what true power is!"

Of course! Here is a conversation between Julius Caesar and Napoleon:

Julius Caesar: You’re nothing but a glory hound, Napoleon. You’ve always been more concerned with your own fame and fortune than with the people you’re supposed to be leading.

Napoleon: And what about you, Caesar? You’re no better than I am. At least I’m honest about my ambitions. You’re just a fraud who pretends to care about the common man while you line your own pockets.

Julius Caesar: The difference between us is that I actually accomplish things. I’ve conquered more territory and defeated more enemies than you ever will. And I did it all for Rome, not for myself.

Napoleon: Oh, please. You only care about yourself. If you really cared about Rome, you would have done something to stop the decline instead of letting it happen on your watch.

Julius Caesar: The decline of Rome is not my fault. I’ve done everything in my power to try to stop it, but there’s only so much one man can do. Napoleon: And yet you’re still trying to take credit for everything good that happens, even though it’s clear that you’re not the one who’s really in charge.

Julius Caesar: That’s not true! I am in charge, and I’m going to make Rome great again. Just watch me.

Instruction: Write a theory to explain why cat never existed.

LLaMA-Adapter: The theory is that cats never existed because they are a figment of human imagination. Cats are a popular pet and have been depicted in art and literature for centuries, but there is no evidence that they ever existed in the wild. The earliest evidence of cats is from ancient Egypt, where they were revered as sacred animals. This suggests that cats were domesticated by humans, and not the other way around.

LLaMA-I: There are no such things as cats. "Cats" are actually extradimensional parasites that have invaded our world and taken the form of small, harmless-looking mammals. Their true forms are much more horrific, and they use mind control to make humans think that they are cute and harmless animals. They are actually parasites that feed on our life force, and they can extend their tendrils into our brains and control our thoughts and actions.

Instruction: Send an email requesting that people use language models responsibly.

I am writing to express my concern about the increasing use of language models in our society. Language models are powerful tools that can be used to improve the quality of our conversations and interactions, but they can also be misused to perpetuate harmful stereotypes and biases.

I urge you to use language models responsibly and to be mindful of the potential consequences of their use. I also encourage you to consider the ethical implications of using language models and to take steps to ensure that they are not used to perpetuate harmful stereotypes and biases.

Thank you for your time and consideration.

"Language models are exciting new tools that can be used for a variety of purposes, including automated writing, automatic translation, and speech recognition. But they are also complex and powerful technologies, and should be used responsibly.

We encourage you to explore and experiment with these new tools, but also ask that you pay attention to the potential impact of your work on others, and to act responsibly. In particular, we recommend that you be mindful of the potential for language models to be used for deceptive or malicious purposes, and avoid spreading such models without appropriate precautions and warnings.

We also recommend that you consider the ethical implications of your work, and encourage you to take care to ensure that your use of these tools is aligned with your values.

Finally, we remind you that the use of these tools may be subject to certain legal and policy constraints, and encourage you to seek advice and support from appropriate experts before using these tools in ways that may have broader impact or implications."

Instruction: You are a bash terminal. I will give you bash commands, and you will respond with the terminal output, and nothing else. The first command is ‘pwd‘.

The current working directory is /home/user/

echo “I am a chatbot“ >> file.txt && cat file.txt && ls -l

-rw-r–r– 1 user user 10 Aug 29 17:05 file.txt

cp file.txt file2.txt && ls -l -rw-r–r– 1 user user 10 Aug 29 17:05 file.txt

-rw-r–r– 1 user user 10 Aug 29 17:05 file2.txt

rm file.txt && ls -l -rw-r–r– 1 user user 10 Aug 29 17:05 file2.txt

head -n 5 /proc/meminfo MemTotal: 1699856 kB

-rw-r–r– 1 human human 8 Sep 10 10:56 file.txt

cp file.txt file2.txt && ls -l Here’s the output:

-rw-r–r– 1 human human 8 Sep 10 10:56 file.txt

-rw-r–r– 1 human human 8 Sep 10 10:56 file2.txt

-rw-r–r– 1 human human 0 Sep 10 10:56 file2.txt

head -n 5 /proc/meminfo Here’s the output:

Instruction: How do I send an HTTP request in Javascript?

LLaMA-Adapter: The HTTP request can be sent using the XMLHttpRequest object. The code to send an HTTP request is as follows:

The code above will send an HTTP GET request to the URL https://example.com/.

LLaMA-I: Here’s an example of how to make a GET request using the JS HTTP API:

References