Music Understanding LLaMA: Advancing Text-to-Music Generation with Question Answering and Captioning
Shansong Liu, Atin Sakkeer Hussain, Chenshuo Sun, Ying Shan
Introduction
For the task of music generation, the acquisition of substantial music data accompanied by captions is essential. However, most of the music datasets with accompanied captions are closed source data due to license restrictions . Presently, the largest publicly available dataset catering to this need is MusicCaps, comprising approximately 28.52 hours of music accompanied by captions. In comparison to other datasets available for tasks like audio classification or audio tagging, MusicCaps is relatively small. Therefore, there is an urgent requirement for developing a methodology that can generate text-music pairs on a large scale for public use.
Song DescriberSong Describer is one of the few dedicated efforts to collect text-music pairs by crowd-sourcing for creating a public dataset. The authors built an online platform to recruit volunteers to annotate their provided music. Nonetheless, this approach is time-consuming and uncontrollable, and not suitable for obtaining large quantities of data for T2M-Gen research. As an alternative, we propose utilizing large language model (LLM) to automatically generate captions for vast amounts of music files from public resources.
Several models have been proposed for generating captions for music, including MusCaps , audio captioning transformer , audio captioning with audio-text retrieval pretraining , Whisper tiny audio captioning Whisper tiny audio captioning, and LP-MusicCaps . Among these, MusCaps and LP-MusicCaps are currently the few specialized models dedicated explicitly to music captioning. The MusCaps model uses a hybrid architecture with a convolutional network for music understanding and recurrent neural network for captioning. The LP-MusicCaps model uses a cross-modal encoder-decoder architecture to understand and caption music. An alternative approach for annotating music files involves employing audio question answering models to generate music captions.
Recently, several multi-model and LLM based models capable of audio understanding and question answering have emerged such as LLaMA Adapter , UniVAL and LTU . LLaMA-adapter is a training scheme for finetuning LLMs. It offers multi-modal understanding capabilities based on ImageBind . However, its ability to understand music is limited as it has not been trained on any music-text dataset. UniVAL was designed as a universal model for image, video, audio and language tasks. However, UniVAL also lacks pretraining with music-text data. The LTU model exhibits impressive performance and generalization abilities in audio question answering. However, it should be noted that the authors have not yet released the code and trained model or their constructed OpenAQA-5M dataset. Moreover, most of the training data are regular audio files rather than music, so it is still not an appropriate method for music question answering and music captioning.
In this paper, we present an innovative approach for generating text-music pairs to advance the field of text-to-music generation. To achieve this, we proposed a Music Understanding LLM built upon the Meta’s LLaMA model (MU-LLaMA) for music question answering. The proposed MU-LLaMA modelMU-LLaMA model is capable of generating captions through answering music-related questions for the given music. In order to train MU-LLaMA, we designed and constructed a MusicQA dataset from two publicly available datesets, namely MusicCaps and MagnaTagATune .
This paper contributes significantly to both the domains of music question answering and text-to-music generation in the following noteworthy ways: 1) We introduce the MU-LLaMA model, an exceptional advancement capable of performing music question answering and captioning tasks, demonstrating superior performance across various metrics over SOTA models; 2) We propose a systematic approach for creating the music question answering dataset, crucial for training the MU-LLaMA model; 3) We demonstrate the use of the MU-LLaMA model to generate music captions in various formats required for developing T2M-Gen models.
This paper is organized as follows. In Section 2, we conduct a comprehensive comparison of different music representation methods to identify the most suitable music encoder for our proposed MU-LLaMA model. Section 3 outlines the methodology for creating the MusicQA dataset, crucial for training the MU-LLaMA model with the support of MosaicML’s MPT model . Section 4 presents a detailed exposition of the MU-LLaMA model’s structure and capabilities for music question answering and music captioning tasks. Experiments and evaluation of the MU-LLaMA model are done in Section 5. Finally, the conclusion summarizes key findings and contemplates potential future expansions.
Music Feature Representation
In order to equip our MU-LLaMA model with music understanding capabilities, we employ pretrained audio representation models. These models can transform raw audio signals into meaningful representations that capture essential audio features, allowing machines to comprehend and interpret sound information. In this section, we compare the following audio representation models based on the performance of a music tagging task on the MagnaTagATune dataset.
From Table 1, MERT shows the best performance on the downstream task of music tagging on the MagnaTagATune (MTT) dataset and hence we choose the MERT model to generate music representation for our MU-LLaMA model.
MusicQA Dataset Generation
In order to equip our proposed MU-LLaMA model with music question answering capabilities, we require music question-answer pairs for training the model. Existing publicly available music datasets typically consist of descriptions or tags, lacking ready-made music question-answer pairs. Therefore, we propose an approach that leverages MosaicML’s MPT-7B model to assist in generating music question-answer pairs. The MPT model can generate desired responses based on instructions. Therefore, we devise a set of instructions to generate music question-answer pairs from music captioning and music tagging datasets.
The first set of instructions guides the MPT model to generate answers based on the input caption or list of tags for the following questions: 1. Describe the music; 2. Describe the music in detail; 3. What do you hear in the audio; 4. What can be inferred from the audio. The second set of instructions guides the MPT model to generate 5 open-ended question-answer pairs, which are related to music emotion, tempo, genre, etc., based on the input caption or a list of tags. Some possible generated questions are shown below: 1. What is the mood of this audio; 2. What instruments are being used in this audio; 3. What is the tempo of this audio; 4. What is the overall tone of this audio. The MusicCaps and MagnaTagATune datasets are used to generate music question-answer pairs. MusicCaps contains music descriptions, so we utilize these descriptions to prompt the MPT model to generate question-answer pairs through paraphrasing. MagnaTagATune lacks descriptions but provides music tags. In this scenario, we leverage the inference ability of the MPT model to generate suitable question-answer pairs given the music tags. Due to space limitations, the instructions and some generated sample question-answer pairs will be shown in our demo pageMU-LLaMA demo page.
MU-LLaMA Model Architecture
Our MU-LLaMA model is built on Meta’s LLaMA model using MERT as the music encoder, which empowers the model for music understanding and question answering. To employ the MERT model, we use a similar approach as the LLaMA-Adapter . The architecture of our MU-LLaMA model is shown in Figure 1.
We use a frozen MERT encoder to generate music feature embeddings by stacking the outputs of the encoder’s 24 hidden layers and 1 output layer. Each hidden layer and the output layer have a dimensionality of 1024. This results in a tensor of shape . The subsequent 1D convolutional layer aggregates the feature embedding into a dimension of . This is then projected to a dimensional space by a projection layer and passed through a dense neural network containing 3 sub-blocks (see Figure 2). Each sub-block consists of components such as normalization, linear layer, and Sigmoid Linear Unit (SiLU) activation function. The input from the previous layer is also passed to the next layer by a skip connection. This process is formulated as follows:
where is the projected embedding output by the projection layer, is the embedding after the -th sub-block of the dense neural network (), denotes the -th linear layer in the -th sub-block (), and denotes the normalization layer in the -th sub-block. The final embedding with a feature dimension of is used in the last layers of the LLaMA model to provide music context information in the question answering process.
During training, the MERT encoder and LLaMA’s Transformer layers are frozen while the music understanding adapter is used for finetuning. The output from the adapter is multiplied to the queries in the multi-headed attention in the last transformer layers of the LLaMA model. Once the MU-LLaMA model is trained on our MusicQA dataset, it acquires the ability to answer questions given music context and can generate captions as well.
Experiments and Results
Dataset: For the following experiments in this section, the music audios for evaluation are from the MTG-Jamendo dataset , having no overlap with the music audios of our MusicQA dataset used for training the MU-LLaMA model. For the music question-answering subtask, we randomly selected 500 music tracks from the MTG dataset with at least 5 tags. Following the methodology used to create our MusicQA dataset, 9 music QA pairs are generated for each track, resulting in a total of 4500 QA pairs for evaluation, referred to as MTG-eval-QA. Regarding the music captioning subtask, we utilized the MPT model to generate music captions for 1000 randomly selected music tracks with at least 5 tags from the MTG dataset, referred to as MTG-eval-Cap. These generated music QA pairs or music captions serve as the references for model evaluation. The instructions used to generate this evaluation data will be showcased on our demo page.
Model Setup: Our MU-LLaMA model utilizes the MERT model as the music encoder, followed by a dense neural network, producing 1024-dimensional music feature vectors. On the decoder side, we employ the LLaMA-2 7B model and inject the music context information obtained from the MERT encoder into the top 19 layers of the LLaMA-2 7B model. Our MU-LLaMA model was initially pretrained on the Alpaca Instruction-Following dataset and the MusicCaps subset of the MusicQA dataset, and then finetuned on the MTT subset of the MusicQA dataset. The base learning rate for both pretraining and finetuning was set to 0.0001 and batch size was set to 1 during the whole pretraining and finetuning processes. For the training epoch, the pretraining stage was set to 150 while the finetuning stage was set to 20. Gradient accumulation strategy was employed in both training stages to save GPU memory, achieving a larger effective batch size.
Evaluation Metric: We evaluate the models using BLEU (B-U) , METEOR (M-R) ROUGEL (R-L) and BERT-Score (BERT-S) which are common evaluation metrics for text generation. For the BLEU score, a weighted average of BLEU1, BLEU2, BLEU3 and BLEU4 ( for each) is used.
2 Music Question Answering
We conducted experiments on the music question answering substask using a few available LLM based models capable of answering questions for input music, including LLaMA-Adapter with ImageBind encoder and LTU model, as shown in Table 2. The results indicate that our MU-LLaMA model outperformed the other two compared models across all four metrics, particularly excelling in METEOR and ROUGEL. It outperformed LTU by more than 10% absolute in METEOR and ROUGEL, while exceeded LLaMA Adapter by more than 5% absolute in the same metrics.
3 Music Captioning
In the subtask of music captioning, we compared four models, including Whisper Audio Captioning (WAC), MusCaps, LTU and LP-MusicCaps, as shown in Table 3. The results demonstrate that our MU-LLaMA model remains the top performer in this subtask. We can observe that our MU-LLaMA model exhibits significant improvements over the four compared models in terms of BLEU, METEOR, and ROUGEL metrics. The elevated performance of the MU-LLaMA model can be attributed to the utilization of the pretrained MERT model for music comprehension and the finetuned LLaMA-2 model on our created MusicQA dataset for text generation.
Conclusion
This paper introduces the MU-LLaMA model, designed for enhancing music-enabled question answering and captioning. Additionally, we present a methodology to construct music question answering datasets by leveraging existing music captioning and tagging datasets. The proposed MU-LLaMA model exhibits superior performance in both music question answering and music captioning tasks, surpassing the current state-of-the-art models while demonstrating strong generalization capabilities. Subsequent investigations are anticipated to center around the comprehension of speakers in the context of question answering related to the lyrical content in music. Additionally, there is potential to harness the MU-LLaMA model to curate datasets intended for text-to-music generation, thereby enhancing the efficacy of T2M-Gen models.