Symbolic Music Generation with Diffusion Models

Gautam Mittal, Jesse Engel, Curtis Hawthorne, Ian Simon

Introduction

Denoising diffusion probabilistic models (DDPMs) are a promising new class of generative models that can synthesize comparably high-quality samples by learning to invert a diffusion process from data to Gaussian noise. Unlike many existing deep generative models, DDPMs sample through an iterative refinement process inspired by Langevin dynamics , which enables post-hoc conditioning of models trained unconditionally for creative applications.

Despite these exciting advances, DDPMs have not yet been applied to symbolic music generation because their iterative refinement sampling process is confined to continuous domains such as images and audio . Similarly, DDPMs cannot take advantage of the recent advances in modeling long-term structure that use a two-stage process of modeling discrete tokens extracted by a separate low-level autoencoder.

In this paper, we demonstrate that it is possible to overcome these limitations by training DDPMs on the continuous latents of a low-level variational autoencoder (VAE) to generate long-form discrete symbolic music. Our key findings include:

High-quality unconditional sampling of discrete melodic sequences (1024 tokens) with DDPMs through iterative refinement of lower-level VAE latents.

DDPMs outperforming strong autoregressive baselines (TransformerMDN) in hierarchical modeling of continuous latents, partly due to a lack of teacher forcing and exposure bias during training.

Post-hoc conditional infilling of melodic sequences for creative applications.

Background

DDPMs are a class of generative models that define latents x1,...,xNx_{1},...,x_{N} of the same dimensionality as the data x0∼q(x0)x_{0}\sim q(x_{0}). Diffusion models are comprised of a forward process and a reverse process. The forward process starts from the data x0x_{0} and iteratively adds Gaussian noise according to a fixed noise schedule for NN diffusion steps:

where β1,β2,...,βN\beta_{1},\beta_{2},...,\beta_{N} is a noise schedule that converts the data distribution x0x_{0} into latent xNx_{N}. The choice of noise schedule has been shown to have important effects on sampling efficiency and quality .

The reverse process is defined by a Markov chain parameterized by θ\theta that iteratively refines latent point xN∼N(0,I)x_{N}\sim\mathcal{N}(0,I) into data point x0x_{0}. The learned transition probabilities are defined as,

where the objective is to gradually denoise samples at each reverse diffusion step tt. In practice, σθ\sigma_{\theta} is set to an untrained time-dependent constant based on the noise schedule, and found σθ(xt,t)=σt=1−αˉt−11−αˉtβt\sigma_{\theta}(x_{t},t)=\sigma_{t}=\frac{1-\bar{\alpha}_{t-1}}{1-\bar{\alpha}_{t}}\beta_{t} to have reasonable practical results, where αt=1−βt\alpha_{t}=1-\beta_{t}, and αˉt=∏i=1tαi\bar{\alpha}_{t}=\prod_{i=1}^{t}\alpha_{i}.

The training objective is to maximize the log likelihood of pθ(x0)=∫pθ(x0,...,xN)dx1:Np_{\theta}(x_{0})=\int p_{\theta}(x_{0},...,x_{N})dx_{1:N}, but the intractability of this marginalization leads to the following evidence lower bound (ELBO):

Additionally, observed that the forward process can be computed for any step tt such that q(xt∣x0)=N(xt;αˉtx0,(1−αˉt)I)q(x_{t}|x_{0})=\mathcal{N}(x_{t};\sqrt{\bar{\alpha}_{t}}x_{0},(1-\bar{\alpha}_{t})I), which can be viewed as a stochastic encoder. To simplify the above variational bound, propose training on pairs of (xt,x0)(x_{t},x_{0}) to learn to parameterize this process with a simple squared L2 loss. The following objective is simpler to train, resembles denoising score matching and was found to yield higher-quality samples:

We refer readers to Algorithms 2 and 3 based on in our supplementary materialSupplementary material including additional implementation details, audio, and tables is available at https://goo.gl/magenta/symbolic-music-diffusion-examples for the full training and sampling procedure.

2 Variational Autoencoders

Variational autoencoders are generative models that define p(y,z)=p(y∣z)p(z)p(y,z)=p(y|z)p(z) where zz is a learned latent code for data point yy. Additionally, the latent code is constrained such that zz is distributed according to prior p(z)p(z) where the prior is usually an isotropic Gaussian. The VAE is comprised of an encoder qγ(z∣y)q_{\gamma}(z|y) which models the approximate posterior p(z∣y)p(z|y) and a decoder pθ(y∣z)p_{\theta}(y|z) which models the conditional distribution of data yy given latent code zz.

The training objective is to maximize the log likelihood of pθ(x)=∫pθ(y∣z)p(z)dzp_{\theta}(x)=\int p_{\theta}(y|z)p(z)dz, but this marginalization is intractable, and we use the following variational bound maximized with qγ(z∣y)q_{\gamma}(z|y) as the approximate posterior:

The flexible implementation of variational autoencoders allows them to learn representations over a wide variety of domains. Of particular interest to us are sequential autoencoders which use long short-term memory cells to model temporal context in sequential data distributions.

In practice, there is a trade-off between the quality of reconstructions and the distance between the approximate posterior qγ(z∣y)q_{\gamma}(z|y) and the Gaussian prior p(z)p(z). This makes sampling more difficult for VAEs with better reconstructions due to latent “holes" in the approximate posterior and is one of the primary shortcomings of these models.

Model

A diagram and description of our multi-stage diffusion model is shown in Figure 1. We refer readers to the supplement for full implementation details.

Our model learns to generate discrete sequences of notes (known as MIDI) by first training a VAE with parameters γ\gamma on the sequences and then training a diffusion model to capture the temporal relationships among the kk VAE latents. Sequence VAEs such as MusicVAE are difficult to train on long sequences , which we overcome by pairing the short 2-bar MusicVAE model with a diffusion model capable of modeling dependencies between k=32k=32 latents, thus modeling 64 bars in total.

MusicVAE embeddings: Each musical phrase is a sequence of one-hot vectors with 16 quantized steps per measure and the vocabulary contains 90 possible tokens (1 note on + 1 note off + 88 pitches). We then parameterize each 2-bar phrase using the pre-trained 2-bar melody MusicVAE and generate a sequence of continuous latent embeddings z1,...,zkz_{1},...,z_{k} to parameterize an entire sequence. The MusicVAE model employs bidirectional recurrent neural networks as an encoder and autoregressive decoding as shown in Figure 4 in the supplement. As we use the pretrained model from the original work, full model details can be found in . After encoding each 2-measure phrase into a latent zz embedding, we perform linear feature scaling such that the domain of each embedding is $.ThisensuresconsistentlyscaledinputsstartingfromtheisotropicGaussianlatent. This ensures consistently scaled inputs starting from the isotropic Gaussian latentx_{N}$ for the diffusion model.

This positional encoding e1,e2,...,eke_{1},e_{2},...,e_{k} is added to inputs xtx_{t} before being fed through the transformer encoder layers allowing the model to capture the temporal context of the continuous inputs.

Noise schedule and conditioning: As described in both the original diffusion model framework and in , we use an additional sinusoidal encoding to condition the diffusion model on a continuous noise level during training and sampling. This noise encoding is identical to the positional encoding described above but with the frequency of each sinusoid scaled by 5000 to account for the updated domain. We use feature-wise linear modulation to generate γ\gamma (scale) and ξ\xi (shift) parameters given a noise encoding and apply the transformation γϕ+ξ\gamma\phi+\xi to the output ϕ\phi of each layer normalization block in each residual layer, allowing for effective conditioning of the diffusion model. Our model uses a linear noise schedule with N=1000N=1000 steps and β1=10−6\beta_{1}=10^{-6} and βN=0.01\beta_{N}=0.01.

2 Unconditional Generation

In the unconditional generation task, the goal is to produce samples that exhibit long-term structure. Because of our multi-stage approach, this works even in the scenario where the KL divergence between the marginal posterior qγ(z)q_{\gamma}(z) and the Gaussian prior is quite large because the diffusion model accurately captures the structure of the latent space therefore improving the sample quality. Additionally, we extend the underlying VAE to samples longer than what it was trained to model by using the diffusion model to predict sequences of latent embeddings and attempt to generate unconditional samples with coherent patterns across a large number of measures.

3 Infilling

One of the benefits of using a sampling process that iteratively refines noise into data samples is that the trajectory of the reverse process can be steered and arbitrarily conditioned without the need for retraining the diffusion model. In creative domains, this post-hoc conditioning is especially useful for artists without the computational resources to modify or re-train deep models for new tasks. We demonstrate the power of diffusion modeling applied to music with conditional infilling of latent embeddings using an unconditionally trained diffusion model.

The infilling procedure extends the sampling procedure described in Algorithm 3 by incorporating information from a partially occluded sample ss. At each step of sampling, we diffuse the fixed regions of ss with the forward process q(st∣s)=N(st;αˉts,(1−αˉt)I)q(s_{t}|s)=\mathcal{N}(s_{t};\sqrt{\bar{\alpha}_{t}}s,(1-\bar{\alpha}_{t})I) and use a mask mm to add the diffused fixed regions to the updated sample xt−1x_{t-1}. The final output x0x_{0} will be a version of ss with the occluded regions inpainted by the reverse process.

We refer readers to Algorithm 1 for the modified sampling procedure that allows for post-hoc conditional infilling.

Methods

We use the Lakh MIDI Dataset (LMD) for all experiments. The dataset contains over 170,000 MIDI files with 99% of those files used for training and the remaining used for validation. We extracted 988,893 64-bar monophonic sequences for training and 11,295 for validation from the provided MIDI files. Each sequence was encoded into 32 continuous latent embeddings using MusicVAE. We set the softmax temperature for MusicVAE to 0.001 for decoding generated embeddings in all experiments. A diagram of the MusicVAE architecture used is shown in the supplementary material.

2 Autoregressive Baseline

We compare our model to an autoregressive transformer with a mixture density output layer and train on the same dataset as the diffusion model. To ensure a fair comparison, we use the same architecture as our diffusion model with L=6L=6, H=8H=8, and K=2K=2 with 2048 neurons for each fully-connected layer. The mixture density layer outputs a mixture of 100 Gaussians to ensure sufficient mode coverage. In total, our baseline model has 38M trainable parameters. While the model is the same as the diffusion model (25.58M trainable parameters) up until the output layer, the autoregressive model has more parameters due to the much larger output layer. We refer to this model as TransformerMDN.

3 Training

All models were trained using Adam with default parameters. We trained our diffusion model for 500K steps on a single NVIDIA Tesla V100 GPU for 6.5 hours using a learning rate of 10−310^{-3} annealed with a decay rate of 0.98 every 4000 steps and batch size 64. Unlike the diffusion model, which is non-autoregressive, we train TransformerMDN with teacher forcing. We use a batch size of 128, learning rate 3×10−43\times 10^{-4}, and train for 250K steps on a single NVIDIA Tesla V100 GPU for 6.5 hours.

We used the open-source implementation of MusicVAE written in TensorFlow and the publicly available 2-bar melody checkpoints trained on LMD. We trained our diffusion and baseline modelsOur implementation is available at https://github.com/magenta/symbolic-music-diffusion with JAX and Flax .

4 Framewise Self-similarity Metric

We clip Consistency and Variance such that samples with μOA\mu_{OA} or σOA2\sigma^{2}_{OA} with greater than 100% percent error from the ground truth are considered to have zero relative similarity.

5 Latent Space Evaluation

We evaluate the similarity of latent embeddings generated by each of our models using the Fréchet distance (FD) and Maximum Mean Discrepancy (MMD) with a polynomial kernel, which are popular evaluation metrics in the generative modeling literature. These metrics measure the distance between the models’ continuous output distributions and the original data distribution in latent space. It is important to note that this metric does not measure long-term temporal consistency or quality of produced sequences and only measures the quality of the intermediate continuous representation before the final sequence is generated using the MusicVAE decoder.

Results

To evaluate unconditional sample quality, we compare batches of 1000 samples (32 latents each) generated by our proposed diffusion model with random draws from the training and test sets. We compare against a set of baseline generators including TransformerMDN, the independent N(0,I)\mathcal{N}(0,I) MusicVAE prior, and spherical interpolation between two MusicVAE embeddings at the start and end of an example from the test set.

As seen in Table 1, the diffusion model quantitatively produces samples most similar to the training data according to the relative framewise overlapping area metrics for note pitch and duration. The diffusion model outperforms TransformerMDN, which is challenged by modelling the relatively high-dimensional continuous latents autoregressively, even with a mixture of Gaussians output. The autoregressive models are also trained with teacher forcing that results in exposure bias, leading to divergence from the data distribution during sampling. This is reflected in the lower pitch and duration consistency and higher variance in the absolute overlapping area numbers seen in Supplemental Table 2. Additionally, the diffusion model is able to capture the joint dependencies of the sequences better because it learns to model all latents simultaneously as opposed to autoregressively. Note that the Gaussian prior also suffers from low consistency and high variance, due to lack of temporal dependencies, while the interpolated samples conversely suffer from low variance and too much consistency due to high repetition.

Table 3 in our supplement presents latent space evaluations of our generated samples. The TransformerMDN outperformed all other models, likely due to the Gaussian mixture prior on its output layer whereas the diffusion model must learn the output distribution from scratch. Furthermore, the latent space metrics are limited by assumptions about the latent manifold distribution and are unable to fully capture the detail of the entire space, further highlighting the necessity of our quantitative framewise self-similarity metrics and qualitative evaluations.

Figure 3 helps us to further understand the iterative refinement process by showing the improvement in sample quality as a function of iterative refinement time for both latent space and framewise self-similarity metrics. Interestingly, latent metrics improve steadily, while consistency similarity starts fairly high, and variance similarity only emerges at the end of refinement. We refer the reader to the supplementary material for extended visual and audio samples of the generated sequences from each modelhttps://goo.gl/magenta/symbolic-music-diffusion-examples.

2 Infilling

To probe the diffusion model’s ability to perform post-hoc conditional generation, we remove the middle 32 measures (16 embeddings) and generate new embeddings following Algorithm 1 by conditioning on the first and last 8 embeddings. As in the unconditional setting, we compare to interpolation and independent samples from the prior.

Figure 2 visualizes the task by plotting the resulting note sequences for two different examples. Even qualitatively, it is visually apparent that the diffusion model produces notes with a consistency and variance similar to the original data. Latent interpolation is very consistent but unrealistically repetitive, and sampling independently from the prior produces a sequence with extremely large variance that is inconsistent in both pitch and duration.

The quantitative evaluations in Table 1 back up these observations. Similar to the unconditional generation task, the diffusion model outperforms the baselines in both consistency and variance similarity. We do not include the autoregressive baseline because it is unable to condition on the final 8 embeddings.

Related Work

Multi-stage learning: Several models have achieved long-term structure by first modeling some intermediate representation and then using that to guide the final generative process. Wave2Midi2Wave uses a transformer to generate MIDI-like symbolic data and then a WaveNet to synthesize that symbolic data into audio. Jukebox and DALL-E use similar approaches for text-conditioned generative music audio and image models. There has also been work investigating the extension of a single-measure VAE to multiple measures by using an autoregressive LSTM with a mixture density output layer , similar to TransformerMDN.

Iterative refinement: use an orderless NADE with blocked Gibbs sampling to iteratively rewrite musical harmonies based on surrounding context. use a gradient-based sampler combined with a restricted Boltzmann machine to generate polyphonic piano music. Similarly, investigated the use of a score-based generative model to generate Bach chorales with annealed Langevin dynamics. A key distinction between our method and prior work is the use of a VAE to parameterize the discrete space of musical notes for improved generation with a DDPM while previous methods have performed iterative refinement in the discrete space directly.

Conditional sampling from unconditional models: We build on top of the breadth of material that investigate steering generation in the latent space of an unconditionally trained generative model . Most similar to our work is which train an additional model on top of a pre-trained variational autoencoder to steer generation in the latent space of that autoencoder. Our approach builds on top of this by not only improving sampling and generation of the underlying autoencoder but also extending generation to sequences much longer than those used to train the VAE. We also use a diffusion model to refine generation and provide conditional infilling while the works mentioned primarily used conditional GANs and VAEs to extend the underlying autoencoder for a single latent embedding as opposed to a sequence of latent embeddings.

Conclusion

We have proposed and demonstrated a multi-stage generative model comprised of a low-level variational autoencoder with continuous latents modeled by a higher-level diffusion model. This approach enables using diffusion models on discrete data, and as priors for modeling long-term structure in multi-stage systems. We demonstrate that this model is useful for symbolic music generation, both in unconditional generation and conditional infilling. Future work will include extending this approach to other discrete data such as text, and exploring a greater array of approaches for post-hoc conditioning in creative applications.

Acknowledgements

We thank Anna Huang, Carrie Cai, Sander Dieleman, Doug Eck and the rest of the Magenta team for helpful discussions, feedback, and encouragement.

References

Appendix A DDPM Training and Sampling

To train and sample from our diffusion model, we use the algorithms as described in .

Appendix B MusicVAE Architecture

Appendix C Trimming Latents

During training VAEs typically learn to only utilize a fraction of their latent dimensions. As shown in Figure 5, by examining the standard deviation per dimension of the posterior q(z∣y)q(z|y) averaged across the entire training set, we are able to identify underutilized dimensions where the average embedding standard deviation is close to the prior of 1. The VAE loss encourages the marginal posterior to match to the prior , but to encode information, dimensions must have smaller variance per an example.

In all experiments, we remove all dimensions except for the 42 dimensions with standard deviations below 1.0, before training the diffusion model on the input data. We find this latent trimming to be essential for training as it helps to avoid modeling unnecessary high-dimensional noise and is very similar to the distance penalty described in . We also tried reducing the dimensionality of embeddings with principal component analysis (PCA) but found that the lower dimensional representation captured too many of the noisy dimensions and not those with high utilization.

Appendix D Tables

In Tables 2 and 3, we present the unnormalized framewise self-similarity results as well as the latent space evaluation of each model.

Appendix E Additional Samples

In Figure 6 we provide piano rolls of sequences drawn from the test set and in Figures 7, 8, 9, and 10 we present additional samples unconditionally generated by our diffusion model, TransformerMDN, spherical interpolation, and through independent sampling from the MusicVAE prior, respectively. Additional piano roll visualizations from infilling experiments are provided in Figure 11.

For extended visual and audio samples of the generated sequences from each model, we refer the reader to the online supplement available at https://goo.gl/magenta/symbolic-music-diffusion-examples.