AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo Mandic, Wenwu Wang, Mark D. Plumbley

Introduction

Generating sound effects, music, or speech according to personalized requirements is important for applications such as augmented and virtual reality, game development, and video editing. Traditionally, audio generation has been achieved through signal processing techniques (Andresen, 1979; Karplus & Strong, 1983). In recent years, generative models (Oord et al., 2016; Ho et al., 2020; Song et al., 2021; Tan et al., 2022), either unconditional or conditioned on other modalities (Kreuk et al., 2022; Żelaszczyk & Mańdziuk, 2022), have revolutionized this task. Previous studies primarily worked on the label-to-sound setting with a small set of labels (Liu et al., 2021b; Pascual et al., 2022) such as the ten sound classes in the UrbanSound8K dataset (Salamon et al., 2014). In comparison, natural language is considerably more flexible than labels as they can include fine-grained descriptions of audio signals, such as pitch, acoustic environment, and temporal order. The task of generating audio prompted with natural language descriptions is known as text-to-audio (TTA) generation.

TTA systems are designed to generate a wide range of high-dimensional audio signals. To efficiently model the data, we adopt a similar approach as DiffSound (Yang et al., 2022) by employing a learned discrete representation to efficiently model high-dimensional audio signals. We also draw inspiration from the recent advancements in autoregressive modelling of discrete representation learnt on the waveform, such as AudioGen (Kreuk et al., 2022), which has surpassed the capabilities of DiffSound. Building on the success of StableDiffusion (Rombach et al., 2022), which uses latent diffusion models (LDMs) for high-quality image generation, we extend previous TTA approaches to continuous latent representations, instead of learning discrete representations. Additionally, as audio manipulations, such as style transfer (Engel et al., 2020; Pascual et al., 2022), are desired for some applications such as games, we explore and achieve various zero-shot text-guided audio manipulations with LDMs, which have not been demonstrated before.

For previous TTA works, a potential limitation for generation quality is the requirement of large-scale high-quality audio-text data pairs, which are usually not readily available, and where they are available, are of limited quality and quantity (Liu et al., 2022f). To better utilize the low-quality data, several methods for text preprocessing have been proposed (Kreuk et al., 2022; Yang et al., 2022). However, these preprocessing steps limit generation performances by overlooking the relations of sound events (e.g., a dog is barking at the bark is transformed into dog bark park). By comparison, our proposed method only requires audio data for generative model training, circumvents the challenge of text preprocessing, and performs better than using audio-text paired data, as we will discuss later.

In this work, we present a TTA system, AudioLDM, which achieves high generation quality with continuous LDMs, with good computational efficiency and enables text-conditional audio manipulations. The overview of AudioLDM design for TTA generation and text-guided audio manipulation is shown in Figure 1. Specifically, AudioLDM learns to generate the representation in a latent space encoded by a mel-spectrogram-based variational auto-encoder (VAE). An LDM conditioned on a contrastive language-audio pretraining (CLAP) embedding is developed for VAE latent generation. By leveraging the audio-text-aligned embedding space in CLAP, we remove the requirement for paired audio-text data during training LDM, as the condition for VAE latent generation can directly come from the audio itself. We demonstrate that training an LDM with audio only can be even better than training with audio-text data pairs. The proposed AudioLDM achieves leading TTA performance on the AudioCaps dataset with a Freshet distance (FD) of 23.3123.31, outperforming the DiffSound baseline (FD of 47.6847.68) by a large margin. Our system also enables zero-shot audio manipulations in the sampling process. In summary, our contributions are as follows:

We demonstrate the first attempt to develop a continuous LDM for TTA generation. Our AudioLDM method outperforms existing methods in both subjective evaluation and objective metrics.

We utilize CLAP embeddings to enable TTA generation without using language-audio pairs to train LDMs.

We experimentally show that using audio only data in LDM training can obtain a high-quality and computationally efficient TTA system.

We show that our proposed TTA system can perform text-guided audio styles manipulation, such as audio style transfer, super-resolution, and inpainting, without fine-tuning the model on a specific task.

Related Work

Text-to-Audio Generation has gained a lot of attention recently. Two works (Yang et al., 2022; Kreuk et al., 2022) explore how to learn audio representations in a discrete space given a natural language description, and then decode the representations to the audio waveform. Since both works require audio-text paired data for training the latent generation model, they have both proposed methods to address the issues of low quality and scarcity of paired data.

DiffSound (Yang et al., 2022) consists of a text encoder, a decoder, a vector-quantized variational autoencoder (VQ-VAE), and a vocoder. To alleviate the scarcity of audio-text paired data, they propose a mask-based text generation strategy (MBTG) for generating text descriptions from audio labels. For example, the label dog bark, a man speaking will be represented as [M] [M] dog bark [M] man speaking [M], where [M] represent the mask token. However, the text generated by MBTG still only includes the label information, which might potentially limit model performance.

AudioGen (Kreuk et al., 2022) uses a Transformer-based decoder to learn to generate the target discrete tokens that are directly compressed from the waveform. AudioGen is trained on 1010 datasets and proposes data augmentation methods to enhance the diversity of training data. When creating the language-audio pairs, they pre-process the language descriptions to labels to better match the class-label annotation distribution and simplify the task. For example, the text description a dog is barking at the park is transformed to dog bark park. For data augmentation, they mix audio samples according to various signal-to-noise ratios and concatenate the transformed language descriptions. This means that the detailed text descriptions showing the spatial and temporal relationships are discarded.

Diffusion Models (Ho et al., 2020; Song et al., 2021) have achieved state-of-the-art sample quality in tasks such as image generation (Dhariwal & Nichol, 2021; Ramesh et al., 2022; Saharia et al., 2022), image restoration (Saharia et al., 2021), speech generation (Chen et al., 2021; Kong et al., 2021b; Leng et al., 2022), and video generation (Singer et al., 2022; Ho et al., 2022). For speech or audio synthesis, diffusion models have been studied for both mel-spectrogram generation (Popov et al., 2021; Chen et al., 2022c) and waveform generation (Lam et al., 2022; Lee et al., 2022; Chen et al., 2022b).

A major concern with diffusion models is that the iterative generation process in a high-dimensional data space will result in a low inference speed. One of the solutions is to employ diffusion models in a small latent space, an approach used, for example, in image generation (Vahdat et al., 2021; Sinha et al., 2021; Rombach et al., 2022). For TTA generation, the audio waveform has redundant information (Liu et al., 2022e, c) that increases modeling complexity and decreases inference speed. To overcome this, DiffSound (Yang et al., 2022) uses text-conditional discrete diffusion models to generate discrete tokens as a compressed representation of mel-spectrograms. However, the quality of the sound generated by their method is limited. In addition, audio manipulation methods are not explored.

Text-Conditional Audio Generation

Text-to-image generation models have shown stunning sample quality by utilizing Contrastive Language-Image Pretraining (CLIP) (Radford et al., 2021) for generating the image prior. Inspired by this, we leverage Contrastive Language-Audio Pretraining (CLAP) (Wu et al., 2022) to facilitate TTA generation.

After training the CLAP model, an audio sample xx can be transformed into an embedding Ex\boldsymbol{E}^{x} within an aligned audio and text embedding space. The generalization ability of CLAP model has been demonstrated by various downstream tasks such as the zero-shot audio classification (Wu et al., 2022). Then, for unseen language or audio samples, CLAP embeddings also provide cross-modal information.

2 Conditional Latent Diffusion Models

Diffusion models (Ho et al., 2020; Song et al., 2021) consist of two processes: i) a forward process to transform the data distribution into a standard Gaussian distribution with a predefined noise schedule 0<β1<⋯<βn<…βN<10<\beta_{1}<\dots<\beta_{n}<\dots\beta_{N}<1, and ii) a reverse process to gradually generate data samples from the noise according to an inference noise schedule.

In the forward process, at each time step n∈[1,…,N]n\in[1,\dots,N], the transition probability is given by:

where ϵ∼N(0,I)\boldsymbol{\epsilon}\sim\mathcal{N}(\boldsymbol{0},\boldsymbol{I}) denotes injected noise, αn\alpha_{n} is a reparameterization of 1−βn1-\beta_{n} and αˉn:=∏s=1nαs\bar{\alpha}_{n}:=\prod_{s=1}^{n}\alpha_{s} represents the noise level at each step. At the final time step NN, zN∼N(0,I)\boldsymbol{z}_{N}\sim\mathcal{N}(\boldsymbol{0},\boldsymbol{I}) has a standard isotropic Gaussian distribution. For model optimization, we employ the reweighted noise estimation training objective (Ho et al., 2020; Kong et al., 2021b; Rombach et al., 2022):

where Ex\boldsymbol{E}^{x} is the embedding of the audio waveform xx produced by the pretrained audio encoder faudio(⋅)f_{\text{audio}}(\cdot) in CLAP. In the reverse process, starting from Gaussian noise distribution p(zN)∼N(0,I)p(\boldsymbol{z}_{N})\sim\mathcal{N}(\boldsymbol{0},\boldsymbol{I}) and the text embedding Ey\boldsymbol{E}^{y}, a denoising process conditioned on Ey\boldsymbol{E}^{y} gradually generates the audio prior z0\boldsymbol{z}_{0} by the following process:

The mean and variance are parameterized as (Ho et al., 2020):

where ϵθ(zn,n,Ey)\boldsymbol{\epsilon}_{\theta}(\boldsymbol{z}_{n},n,\boldsymbol{E}^{y}) is the predicted generation noise, and σ12=β1\boldsymbol{\sigma}_{1}^{2}=\beta_{1}. In the training stage, we learn the generation of an audio prior z0\boldsymbol{z}_{0} given the cross-modal representation Ex\boldsymbol{E}^{x} of an audio sample xx. Then, in TTA generation, we provide the text embeddings Ey\boldsymbol{E}^{y} to predict the noise ϵθ(zn,n,Ey)\boldsymbol{\epsilon}_{\theta}(\boldsymbol{z}_{n},n,\boldsymbol{E}^{y}). Built on the CLAP embeddings, our LDM realizes TTA generation without text supervision in the training stage. We provide the details of network architecture in Appendix B.

3 Conditioning Augmentation

In text-to-image generation, diffusion-based models have demonstrated an ability to capture the fine-grained details between objects and backgrounds (Ramesh et al., 2022; Saharia et al., 2022; Liu et al., 2022d). One of the reasons for this success is the large-scale language-image training pairs, such as 400400 million image-text pairs in the LAION dataset (Schuhmann et al., 2021). For TTA generation, it is also desired to generate compositional audio signals whose relationships are consistent with natural language descriptions. However, the scale of available language-audio datasets is not comparable to that of language-image datasets. For data augmentation, AudioGen (Kreuk et al., 2022) use a mixup strategy which mixes pairs of audio samples and concatenates their respective processed text captions to form new paired data. In our work, as shown in Equation 3, we provide the audio only embedding Ex\boldsymbol{E}^{x} as conditioning information when training LDMs, we can implement data augmentation on audio only signals instead of needing to augment language-audio pairs. Specifically, we perform mixup augmentation on audio x1x_{1} and x2x_{2} by:

where λ\lambda is a scaling factor varying between [0,1]\left[0,1\right] sampled from a Beta distribution B(5,5)\mathcal{B}(5,5) (Gong et al., 2021). Here we do not need to consider the corresponding text description y1,2y_{1,2}, since text information is not needed during LDM training. By mixing audio pairs, we increase the number of training data pairs (z0,Ex)(\boldsymbol{z}_{0},\boldsymbol{E}^{x}) for LDMs, which makes LDMs robust to CLAP embeddings. In the sampling process, given the text embedding Ey\boldsymbol{E}^{y} from unseen language descriptions, LDMs are expected to generate the corresponding audio prior z0\boldsymbol{z}_{0}.

4 Classifier-free Guidance

For diffusion models, controllable generation can be achieved by introducing guidance at each sampling step. After classifier guidance (Song et al., 2021; Nichol & Dhariwal, 2021), classifier-free guidance (Ho & Salimans, 2021; Nichol et al., 2021) (CFG) has been the state-of-the-art technique for guiding diffusion models. During training, we randomly discard our condition Ex\boldsymbol{E}^{x} with a fixed probability, e.g., 10%10\% to train both the conditional LDMs ϵθ(zn,n,Ex)\boldsymbol{\epsilon}_{\theta}(\boldsymbol{z}_{n},n,\boldsymbol{E}^{x}) and the unconditional LDMs ϵθ(zn,n)\boldsymbol{\epsilon}_{\theta}(\boldsymbol{z}_{n},n). In generation, we use text embedding Ey\boldsymbol{E}^{y} as condition and perform sampling with a modified noise estimation ϵ^θ(zn,n,Ey)\hat{\boldsymbol{\epsilon}}_{\theta}(\boldsymbol{z}_{n},n,\boldsymbol{E}^{y}):

where ww determines the guidance scale. Compared with AudioGen (Kreuk et al., 2022), we have two differences. First, they leverage CFG on a transformer-based auto-regressive model, while our LDMs retain the theoretical formulation behind the CFG (Ho & Salimans, 2021). Second, our text embedding Ey\boldsymbol{E}^{y} is extracted from unprocessed natural language and therefore enables CFG to make use of the detailed text descriptions as guidance for audio generation. However, AudioGen removed the text details showing spatial or temporal relationships with text preprocessing methods.

5 Decoder

Text-Guided Audio Manipulation

Style Transfer Given a source audio sample xsrcx^{src}, we can calculate its noisy latent representation zn0\boldsymbol{z}_{n_{0}} with a predefined time step n0≤Nn_{0}\leq N according to the forward process shown in Equation 2. By utilizing zn0\boldsymbol{z}_{n_{0}} as the starting point of the reverse process of a pretrained AudioLDM model, we enable the manipulation of audio xsrcx^{src} with text input yy with a shallow reverse process pθ(z0:n0∣Ey)p_{\theta}(\boldsymbol{z}_{0:n_{0}}|\boldsymbol{E}^{y}):

where n0n_{0} controls the manipulation results. If we define a n0≈Nn_{0}\approx N, the information provided by source audio will not be retained and the manipulation would be similar to TTA generation. We show the effect of n0n_{0} in Figure 3, where larger manipulations can be seen in the setting of n0=3N/4n_{0}=3N/4.

Inpainting and Super-Resolution Both audio inpainting and audio super-resolution refer to generating the missing part given the observed part xobx^{ob}. We explore these tasks by incorporating the observed part in latent representation zob\boldsymbol{z}^{ob} into the generated latent representation z\boldsymbol{z}. Specifically, in reverse process, starting from p(zN)∼N(0,I)p(\boldsymbol{z}_{N})\sim\mathcal{N}(\boldsymbol{0},\boldsymbol{I}), after each inference step shown in Equation 5, we modify the generated zn−1\boldsymbol{z}_{n-1} with:

The values of observation mask m\boldsymbol{m} depend on the observed part of a mel-spectrogram X\boldsymbol{X}. As we adopt a convolutional structure in VAE to learn the latent representation z\boldsymbol{z}, we can roughly retain the spatial correspondency in mel-spectrogram, as it is shown in Figure 7 in Appendix C. Therefore, if a time-frequency bin Xt,f\boldsymbol{X}_{t,f} is observed, we set the observation mask mtr,fr\boldsymbol{m}_{\frac{t}{r},\frac{f}{r}} in latent space as 11. By using m\boldsymbol{m} to denote the generation part and observation part in z\boldsymbol{z}, according to Equation 11, we can generate the missing information conditioned on the text prompt with TTA models, while retaining the ground-truth observation zob\boldsymbol{z}^{ob}.

Experiments

Training dataset The datasets we used in this paper includes AudioSet (AS) (Gemmeke et al., 2017), AudioCaps (AC) (Kim et al., 2019), Freesound (FS)https://freesound.org/, and BBC Sound Effect library (SFX)https://sound-effects.bbcrewind.co.uk/search. AS is currently the largest audio dataset, with 527527 labels and over 5,0005,000 hours of audio data. AC is a much smaller dataset with around 49,00049,000 audio clips and text descriptions. Most of the data in AudioSet and AudioCaps are in-the-wild audio from YouTube, so the quality of the audio is not guaranteed. To expand the dataset, especially with high-quality audio data, we crawl the data from the FreeSound and BBC SFX datasets, which have a wide range of categories such as music, speech, and sound effects. We show our detailed data processing methods and training configuration in Appendix E.

Evaluation dataset We evaluate the model on both AC and AS. Each audio clip in AC has 55 text captions. We generate the evaluation set by randomly selecting one of them as text condition. Because the authors of AC intentionally remove the audio with the label related to music (Kim et al., 2019), to evaluate model performance with a wider range of sound, we randomly select 10%10\% audio samples from AS as another evaluation set. Since AS does not contain text descriptions, we use the concatenation of labels as text descriptions, such as Speech, hip hop music, and crowd cheering.

Evaluation methods We perform both objective evaluation and human subjective evaluation. The main metrics we use for objective evaluation include frechet distance (FD), inception score (IS), and kullback–leibler (KL) divergence. Similar to the frechet inception distance in image generation, the FD in audio indicates the similarity between generated samples and target samples. IS is effective in evaluating both sample quality and diversity. KL is measured at a paired sample level and averaged as the final result. All of these three metrics are built upon a state-of-the-art audio classifier PANNs (Kong et al., 2020b). To compare with (Kreuk et al., 2022), we also adopt the frechet audio distance (FAD) (Kilgour et al., 2019). FAD has a similar idea to FD but it uses VGGish (Hershey et al., 2017) as a classifier which may have inferior performance than PANNs. To better measure the generation quality, we choose FD as the main evaluation metric in this paper. For subjective evaluation, we recruit six audio professionals to carry on a rating process following (Kreuk et al., 2022; Yang et al., 2022). Specifically, the generated samples are rated based on i) overall quality (OVL); and ii) relevance to the input text (REL) between a scale of 11 to 100100. We include the details of human evaluation in Appendix E. We open-source our evaluation pipeline to facilitate reproducibilityhttps://github.com/haoheliu/audioldm_eval.

Models We employ two recently proposed TTA systems, DiffSound (Yang et al., 2022) and AudioGen (Kreuk et al., 2022) as our baseline models. DiffSound is trained on AS and AC datasets with around 400400M parameters. AudioGen is trained on AS, AC, and eight other datasets with around 285285M parameters. Since AudioGen has not released publicly available implementation, we reuse the KL and FAD results reported in their paper. We train two AudioLDM models. One is a small model named AudioLDM-S, which has 181181M parameters, and the other is a large model named AudioLDM-L with 739739M parameters. We describe the details of UNet architecture in Appendix B. To demonstrate the advantage of our method, we simply train these two models only with the AC dataset. Moreover, to explore the effect of the scale of training data, we develop an AudioLDM-L-Full model which is trained on AC, AS, FreeSound, and BBC SFX datasets.

We show the main evaluation results on the AC test set in Table 1. Given the single training dataset AC, AudioLDM-S can achieve better generation results than the baseline models on both objective and subjective evaluations, even with smaller model size. By expanding model capacity with AudioLDM-L, we further improve the overall results. Then, by incorporating AS and the two other datasets into training, our model AudioLDM-L-Full achieves the best quality, with an FD of 23.3123.31. Although RoBERTa and CLAP have the same text encoder structure, CLAP has an advantage in that it decouples audio-text relationship learning from generative model training. This decoupling is intuitive as CLAP has already modelled the relationship between audio and text by aligning their embedding spaces. On the other hand, AudioLDM-S-Full-RoBERTa, in which the text encoder only represents textual information, requires the model to learn the text-audio relationships while simultaneously learning the audio generation process. Additionally, our CLAP-based method allows for model training using audio-only data. Therefore, using Roberta without pretraining with CLAP may increase the difficulty of training.

Our human evaluation shows a similar trend as other evaluation metrics. Our proposed methods have OVL and REL of around 6464, outperforming DiffSound with OVL of 45.0045.00 and REL of 43.8343.83 by a large margin. On the AudioLDM model size, we notice that the larger model is advantageous for the overall audio qualities. After scaling up the training data, both OVL and REL show significant improvements. Figure 4 shows the score statistic of different models averaged between all the raters. We notice our model is more concentrated on the higher scores compared with DiffSound. Our spam cases, which are randomly selected real recordings, show high scores, indicating the rating result is reliable.

To perform the evaluation on audio data that could include music, we further evaluate our model on the AS evaluation set. We compare our method with DiffSound and show the results in Table 2. Our three AudioLDM models show a similar trend as they perform on the AC test set. We can outperform the DiffSound baseline by a large margin on all the metrics.

Conditioning Information As we train LDMs conditioned on the audio embedding Ex\boldsymbol{E}^{x} but provide the text embedding Ey\boldsymbol{E}^{y} to LDMs in TTA generation, a natural concern is that if stronger results could be achieved by directly using the text embedding as training condition. We conduct experiments and show the results in Table 3. For a fair comparison, we also conduct data augmentation and we adopt the strategy from AudioGen. Specifically, we use the same mixing method for audio pairs shown in Section 3.3, and concatenate two text captions as conditioning information. Table 3 shows by training LDMs on Ex\boldsymbol{E}^{x}, we can achieve better results than training with Ey\boldsymbol{E}^{y}.

We believe the primary reason for the result in Table 3 is that text embedding cannot represent the generation target as good as audio embedding. Firstly, due to the ambiguity and complexity of sound, the text caption is difficult to be accurate and comprehensive. Different human annotators may have different perceptions and descriptions over the same audio, which make training with text-audio pair less stable than with audio only. Moreover, some of the captions are at a highly-abstracted level and cannot correctly describe the audio content. For example, there is an audio in the BBC SFX dataset with caption Boats: Battleships-5.25 conveyor space, which is even difficult for humans to imagine how it sounds. This quality of language-audio pairs may hinder model optimization. By comparison, if we use Ex\boldsymbol{E}^{x} from CLAP latents as a condition, it is extracted directly from the audio signal and is aligned with ideally the best text caption, which enables us to provide strong conditioning information to LDMs without considering the noisy labeled text description. Figure 5 shows sample quality as a function of training progress. We notice that i) training with audio embedding can lead to significantly better results than text embedding throughout the entire training process; and ii) larger models may converge more slowly but can achieve better final performance.

Compression Level We study the effect of compression level rr on generation quality. Table 4 shows the performance comparison with r=4,8,16r\texttt{=}4,8,16. We observe a decreasing trend with the increase of compression levels. Nevertheless, in the setting of r=16r\texttt{=}16 where we compress the 6464-band mel-spectrogram into only 44 dimensions in the frequency axis, our performance is still on par with AudioGen on KL, and better than DiffSound on all the metrics.

If we set the compression level as r=1r\texttt{=}1, which means we directly generate mel-spectrogram from CLAP embeddings, the training process is difficult to implement on a single RTX 30903090 GPU. Similar results happen on r=2r\texttt{=}2. Moreover, the inference speed will be low with r=1,2r\texttt{=}1,2. In our studies, r=4r\texttt{=}4 achieves high generation quality while reducing the computational load to a reasonable level. Hence, we use it as the default setting in our experiments.

Text-Guided Audio Manipulation We show the performance of our text-guided audio manipulation methods on two tasks: super-resolution and inpainting. Specifically, for super-resolution, we upsample the audio signal from 88 kHz to 1616 kHz sampling rate. For the inpainting task, we remove the audio signal between 2.52.5 and 7.57.5 seconds and refill this part by inpainting. Since most studies on audio super-resolution work on speech signal (Liu et al., 2021a, 2022a), we demonstrate our results on both AudioCaps, and a speech dataset VCTK (Yamagishi et al., 2019), which is a multi-speaker speech dataset. For super-resolution, we employ two models AudioUNet (Kuleshov et al., 2017) and NVSR (Liu et al., 2022a) as baseline models, and employ log-spectral distance (LSD) (Wang & Wang, 2021) as the evaluation metric for comparison. For the inpainting task, we use FAD as a metric and establish a baseline for this task.

Table 5 shows that AudioLDM can outperform the strong AudioUNet baseline, but the result is not as good as NVSR (Liu et al., 2022a). Recall that AudioLDM is a model trained on a diverse set of audio signals, including those with heavy background noise. This can lead to the presence of white noise or other non-speech sound events in the output of our super-resolution process, potentially reducing performance. Nevertheless, our contribution could open the door to achieving text-guided audio manipulation with the TTA system in a zero-shot way. Further improvements could be expected based on our benchmark results. We provide several samples of our results in Appendix I.

2 Ablation Study

Table 6 shows the result of our ablation study on AudioLDM-S. By simplifying the attention mechanism in UNet into a one-layer multi-head self-attention (w. Simple attn), the performance in each metric will have a notable decrease, which indicates complex attention mechanism is preferred. Also, we notice the widely used balanced sampling strategy (Gong et al., 2021; Liu et al., 2022b) in audio classification does not show improvement in TTA (w. Balance samp). Conditional augmentation (see Section 3.3) shows improvement in the subjective evaluation, but it does not show improvement in the objective evaluation metrics (w. Cond aug). The reason could be that conditioning augmentation generates training data that is not representative of the AudioCaps dataset, resulting in model outputs that are not well-aligned with the evaluation data, ultimately leading to lower metric scores. Nevertheless, conditioning augmentation can improve two subjective metrics and we still recommend using it as a data augmentation technique.

DDIM Sampling Step The number of inference steps in the reverse process of DDPMs can directly affect the generation quality (Ho et al., 2020; Song et al., 2021). Generally, the sample quality can be improved with an increase in the number of sampling steps and computational load at the same time. We explore the effect of the DDIM (Song et al., 2020) sampling steps on our latent diffusion model. Table 7 shows that more sampling steps lead to better quality. With enough sampling steps such as 100100, the gain of adding sampling steps becomes less significant. The result of 200200 steps is only slightly better than that of 100100 steps.

Guidance Scale represents a trade-off between conditional generation quality and sample diversity. A suitable guidance scale can improve the consistency between generated samples and conditioning information at an acceptable cost of generation diversity. We show the effect of guidance scale ww on TTA in Figure 6. When w=3w=3, we achieve the best results in both FD and KL, but not in FAD. We suppose the reason is the audio classifier in FAD is not as good as FD, as mentioned in Section 5. In this case, the improvement in the adherence to detailed language description may become misleading information to the classifier in FAD. Considering previous studies report FAD results instead of FD, we set w=2w=2 for comparison, but also provide detailed effects of ww on FAD, FD, IS, and KL, respectively.

Case Study We conduct case study and show the generated results in Appendix I, including style transfer (see Figure 9-11), super-resolution (see Figure 12), inpainting (see Figure 13-14), and text-to-audio generation (see Figure 15-22). Specifically, for text-to-audio, we demonstrate the controllability of AudioLDM, including the control of the acoustic environment, material, sound event, pitch, musical genres, and temporal orders.

Conclusions

We have presented a new method AudioLDM for text-to-audio (TTA) generation, with contrastive language-audio pretraining (CLAP) models and latent diffusion models (LDMs). Our method is advantageous in generation quality, computational efficiency, and audio manipulations. With a single training dataset AudioCaps and a single GPU, AudioLDM achieves SOTA generation quality evaluated by both subjective and objective metrics. Moreover, AudioLDM enables zero-shot text-guided audio style transfer, super-resolution, and inpainting.

Acknowledgement

We would like to thank James King and Jinhua Liang for the useful discussion on the latent diffusion model. This research was partly supported by the British Broadcasting Corporation Research and Development (BBC R&D), Engineering and Physical Sciences Research Council (EPSRC) Grant EP/T019751/1 “AI for Sound”, and a PhD scholarship from the Centre for Vision, Speech and Signal Processing (CVSSP), Faculty of Engineering and Physical Science (FEPS), University of Surrey. For the purpose of open access, the authors have applied a Creative Commons Attribution (CC BY) license to any Author Accepted Manuscript version arising.

References

Appendix

We follow the pipeline of the contrastive language-audio pretraining (CLAP) models proposed by (Wu et al., 2022) to capture the similarity between text and audio, and project them into joint latent space. The training dataset includes the currently largest public dataset LAION-Audio-630630K, the AudioSet dataset whose text caption is augmented with keyword-to-captionhttps://github.com/gagan3012/keytotext by T55 model (Raffel et al., 2020), the AudioCaps dataset and the Clotho dataset (Drossos et al., 2020). The LAION-Audio-630630K dataset contains 633,526633,526 language-audio pairs and 4325.394325.39 hours of audio samples. The AudioSet dataset contains 1,912,0241,912,024 pairs and 463.48463.48 hours of audio samples. The AudioCaps dataset contains 49,27449,274 pairs and 136.87136.87 hours of audio samples. The Clotho dataset contains 3,8393,839 pairs and 23.9923.99 hours of audio samples. These datasets contain various natural sounds, audio effects, music and human activity.

where τ\tau is a learnable temperature parameter and DD is the batch size.

B Latent Diffusion Model

We adopt the UNet backbone of StableDiffusion (Rombach et al., 2022) as the basic architecture of LDM for AudioLDM. As shown in Equation 5, the UNet model is conditioned on both the time step tt and the CLAP embedding E\boldsymbol{E}. We map the time step into a one-dimensional embedding and then concatenate it with E\boldsymbol{E} as conditioning information. Since our condition vector is only one-dimensional, we do not use the cross-attention mechanism in StableDiffusion for conditioning. Instead, we directly use the feature-wise linear modulation layer (Perez et al., 2018) to merge conditioning information with the feature map of the UNet convolution block. The UNet backbone we use has four encoder blocks, a middle block, and four decoder blocks. With a basic channel number of cuc_{u}, the channel dimensions of encoder blocks are [cu,2cu,3cu,5cu][c_{u},2c_{u},3c_{u},5c_{u}]. The channel dimensions of decoder blocks are the reverse of encoder blocks, and the channel of the middle block has 5cu5c_{u} dimensions. We add an attention block in the last three encoder blocks and the first three decoder blocks. Specifically, we add two multi-head self-attention layers with a fully-connected layer in the middle as the attention block. The number of heads is determined by dividing the embedding dimension of the attention block with a parameter chc_{h}. We set AudioLDM-S and AudioLDM-L with cu=128,ch=32c_{u}\texttt{=}128,c_{h}\texttt{=}32, and cu=256,ch=64c_{u}\texttt{=}256,c_{h}\texttt{=}64, respectively. In the forward process, we use N=1000N=1000 steps. A linear noise schedule from β1=0.0015\beta_{1}=0.0015 to βN=0.0195\beta_{N}=0.0195 is used. In sampling, we employ the DDIM (Song et al., 2020) sampler with 200200 sampling steps. For classifier-free guidance, a guidance scale ww of 2.02.0 is used in Equation 9.

C Variational Autoencoder

We train our VAE using the Adam optimizer (Kingma & Ba, 2014) with a learning rate of 4.5×10−64.5\times 10^{-6} and a batch size of six. The audio data we use includes AudioSet, AudioCaps, Freesound, and BBC SFX. We perform experiments with three compression-level settings r=4,8,16r\texttt{=}4,8,16, for which the latent channels are C=8,16,32C=8,16,32, respectively. VAEs in all three settings are trained with at least 1.51.5M steps on a single NVIDIA RTX 3090 GPU. To stabilize training, we do not apply the adversarial loss in the first 5050K training steps. We apply the mixup (Kong et al., 2020b) strategy for data augmentation.

Table 8 shows the reconstruction performance of our VAE model with different values of rr. All three settings achieve comparable metrics score with the GT Mel + Vocoder setting, indicating the autoencoder can perform reliable mel-spectrogram encoding and decoding.

D Vocoder

In this work, we employ HiFi-GAN (Kong et al., 2020a) as a vocoder, which is widely used for speech waveform generation. It contains two sets of discriminators, a multi-period discriminator, and a multi-scale discriminator, to enhance the perceptual quality. To synthesize the audio waveform, we train it on the AudioSet dataset. For the input samples at the sampling rate of 16,00016,000Hz, we extract 6464 bands mel-spectrogram. Then we follow the default settings of HiFi-GAN V11. The window, FFT, and hop size are set to 10241024, 10241024, and 160160. The fminf_{\text{min}} and fmaxf_{\text{max}} are set as and 80008000. We use the AdamW optimizer with 0.80.8 and 0.990.99. The learning rate starts from 2×10−42\times 10^{-4} and a learning rate decay of 0.9990.999 is used. We use a batch size of 9696 and train the model with 66 NVIDIA 30903090 GPUs. We release this pretrained vocoder in our open-source implementation.

E Experiment Details

Data Processing The duration of the audio samples in AudioSet and AudioCaps is 1010 seconds, while it is much longer in FreeSound and BBC SFX datasets. To avoid overusing the data from long audio, which usually have repeated sound, we only use the first thirty seconds of the audio in both the FreeSound and BBC SFX datasets and segment them into ten-second long audio files. Finally, we have in total 3,302,5533,302,553 ten-seconds audio samples for model training. It should be noted that even if some datasets, e.g., AudioCaps and BBC SFX, have text captions for the audio, we do not utilize them during the training of LDMs. We only use the audio samples for training. We resample all the datasets into 1616kHz sampling rate and mono format, and all samples are padded to 1010 seconds.

Configuration For each LDM model, we use the compression level r=4r\texttt{=}4 as the default setting. Then, we train AudioLDM-S and AudioLDM-L for 0.60.6M steps on a single GPU, NVIDIA RTX 30903090, with the batch size of 55 and 88, respectively. The learning rate is set as 3×10−53\times 10^{-5}. The AudioLDM-L-Full is trained for 1.51.5M steps on one NVIDIA A100100 with a batch size of 88. The learning rate is 10−510^{-5}. For better performance on AudioCaps, we further fine-tune AudioLDM-L-Full on AudioCaps for 0.250.25M steps before evaluation. It should be noted that we limit our batch size because of the scarcity of GPU. However, this potentially restricts the performance of AudioLDM models. In comparison, DiffSound uses 3232 NVIDIA V100100 GPUs for model training with a batch size of 1616 on each GPU. AudioGen utilizes 6464 A100100 GPUs with a batch size of 256256.

Human evaluation We construct the dataset for human subjective evaluation with 7070 randomly selected samples where 3030 audios are from AudioCaps, 3030 audios are from AudioSet, and 1010 randomly selected real recordings, which we will refer to as spam cases. Therefore, each model should generate 6060 audio samples given the corresponding text descriptions. We gather the output from models in one folder and anonymize them with random identifiers. An example questionnaire is shown in Table 9. The participant will need to fill in the last two columns for each audio file given the text description. Our final result shows that all the human raters have an average score above 9090 on the spam cases. Hence, their evaluation result is considered reliable.

F The Effect of Finetuning

Table 10 compares the results obtained with and without fine-tuning on the evaluation set. We observe an improvement in various evaluation metrics, which is expected since the training set of AudioCaps has a distribution that is similar to the evaluation set. However, it is important to note that higher performance on the limited distribution of the evaluation set may not necessarily indicate better performance overall. A model that can generate broader distributions of audio may perform worse on the evaluation set, even though it may have better generalization capabilities. Future work in audio generation can focus on building an evaluation protocal that is more aligned with human perceptions.

G Computation Efficiency Comparison

As shown in Figure 8(a), AudioLDM-S can generate eight ten-second-long audios within ten seconds without classifier-free guidance. With classifier-free guidance, AudioLDM-Small can generate eight ten-second-long audios with 150150 DDIM steps. Figure 8(c) shows our model is faster than the DiffSound on different batch sizes. Our model can generate eight ten-second-long audios with 2020 seconds while DiffSound needs more than 4040 seconds. Since AudioGen has not been open-sourced yet, we did not perform a speed comparison with AudioGen.

H Limitations

There are several limitations to our study that warrant further investigation in future work. For example, the sampling rate of our model is still insufficient, especially for the generation of music. Exploring higher-fidelity sampling rates such as 32 kHz or 48 kHz could improve the quality of the generated audio. Also, all the modules in AudioLDM are trained separately, which may result in misalignment between different modules. For instance, the latent space learned by VAE may not be optimal for the latent diffusion model. Future work can explore approaches to better align the different modules, such as end-to-end fine-tuning.

The possible negative impact of our method might be the abuse of our technology or released models, e.g., generating fake audio effects to provide misleading information. Moreover, sensitive text content should be restricted in future work to prevent the creation of harmful audio content.

I Demos

We show three examples of zero-shot audio style transfer with AudioLDM-S, using the developed shallow reverse process (see Equation 10). In Figure 9, we show the transfer from drum beats to ambient music. From left to right, we show the source audio sample drum beats, and the six generated samples guided by text prompt ambient music with different starting points n0n_{0}. Given a smaller n0n_{0} (i.e., the left part of the figure), the generated sample is similar to drum beats, while when we set n0=0.8×Nn_{0}=0.8\times N for the last sample, the generated sample will be aligned with the text input ambient music. Similarly, we show the source audio trumpt, and the seven generated samples guided by text prompt children singing in Figure 10. We show the source audio sheep vocalization, and the five generated samples guided by text prompt narration, monologue in Figure 11.

In Figure 12, we show four cases of zero-shot audio super-resolution with AudioLDM-S: 11) violin, 22) sneezing sound from a woman, 33) baby crying, and 44) female speech. The sampling rate of input samples (left) is 88 kHz, and that of generated samples (middle) and ground-truth samples (right) is 1616 kHz. Our visualization shows we can retain the ground-truth observation in the low-frequency part (below 88 kHz), while generating the high-frequency missing part (from 88 kHz to 1616 kHz) with pretrained AudioLDM-S. The generated high-frequency information is consistent with the low-frequency observation.

In Figure 13, we show four samples of zero-shot audio inpainting with AudioLDM-S. The time length of each audio sample is 1010 seconds. In the unprocessed part, we remove the content between 2.52.5 and 7.57.5 seconds from the ground-truth sample as the input of inpainting. In the inpainting result part, we show the generated samples guided by the same text prompt of the ground-truth sample. In the ground truth part, we show the ground-truth sample for comparison.

In Figure 14, we use one sample to demonstrate the audio inpainting guided by different text prompts. Given the observed audio signal shown in the top row, we guide the inpainting process with four different text prompts: 11) ambient music; 22) a man is speaking with bird calls in the background; 33) a cat is meowing; 44) raining with wind blowing. As can be seen, the observed audio signal is preserved in each generated sample, while the generated content can be controlled by text input.

In Figure 15, we demonstrate that AudioLDM can control the acoustic environment of generated samples with a text description. The four samples are generated with the same random seed, but with different text prompts. Their common text information is “A man is speaking in”, while the specific text information describes the acoustic environment as “a small room”, “a huge room”, “a huge room without background noise”, and “a studio”. These samples show the ability of AudioLDM to capture the fine-grained text description about the acoustic environment, and control the corresponding effects on audio samples, such as reverberation or background noise.

In Figure 16, we show the generated music samples when we control the music characteristics with text input. The first sample is generated by “Theme music with bass drum”. Then, we add specific text information “flute”, “fast, flute”, or “flute in the background”, to change the text input. The corresponding variations can be seen in generated mel-spectrograms. We use these samples to demonstrate the ability of AudioLDM to add new musical instruments to music samples, tune the speed of music, and control the foreground-background relations.

In Figure 17, we show the ability of AudioLDM to control the pitch of generated samples. Pitch is an important characteristic of sound effects, music and speech. Here, we set the common text information as “Sine wave with ⋯\cdots pitch”, and input the specific text information “low”, “medium”, and “high”. The text-controlled pitch variation can be seen from the three generated samples.

In Figure 18, we show the ability of AudioLDM to control the materials which generate audio samples. We show four samples generated by the common action “hit” between different materials, e.g., wooden object and wooden environment, or metal object and wooden environment.

In Figure 19, we show the ability of AudioLDM to control the temporal order between generated compositional audio signals. When the text description includes multiple sound effects, AudioLDM can generate the audio signals, and the temporal order between them is consistent with the text input.

In Figure 20, we show four text-to-audio generation results with AudioLDM-S. They include sound effects in natural environment, human speech, human activity, and sound from objects interaction.

In Figure 21, we show four novel audio samples generated with AudioLDM-S. Their text description is rarely seen, e.g., “A wolf is singing a beautiful song.”. We use them to exhibit the generalization ability of AudioLDM.

In Figure 22, we show four music samples generated with AudioLDM-S. Here, we are using the labels of AudioSet as text description for music generation, and we are able to specify the music genres of generated samples such as Classical music.