Unsupervised Cross-Domain Singing Voice Conversion
Adam Polyak, Lior Wolf, Yossi Adi, Yaniv Taigman
Introduction
We tackle the audio synthesis task of singing voice conversion. In this task, a given template song is reproduced by another singer’s voice. The conversion retains the content and musical expression of the template song. Singing voice conversion can aid in improving the vocal qualities of a given singing segment, create mimicry effects and even enable a single amateur singer to record an entire chorus in their home studio.
Our method is inspired by recent work on speech voice conversion and music synthesis . We combine melody extracted features with acoustic speech features and a powerful neural audio generation framework to create a singing voice synthesizer. All features are extracted directly from the raw audio, therefore, our method does not require supervision in the form of a labelled dataset with lyrics and notes or a parallel dataset with singers singing the same songs.
From a technical perspective, we present multiple contributions: (i) introducing an audio generation framework which employs task-related perceptual losses to improve the quality of the generated audio, (ii) the first singing voice conversion method, as far as we know, conditioned on sine-excitation of the song’s melody, together with intermediate features of a speech recognition network, and (iii) presenting a speaker invariant singing voice conversion method trained on voices from either speech or singing datasets (i.e., mimicking in singing either speaking or singing voices), with either single or multiple identities.
Related work
Neural Audio Generation Recent advancements in neural audio generation enabled computers to generate natural sounding speech and music. Autoregressive models, such as WaveNet and SampleRNN generate high-quality audio in the waveform domain, one sample at a time, resulting in slow inference. WaveRNN enabled fast synthesis by training a compact neural network further optimized via sparsification. Feed-forward networks were suggested to further speed up inference, via knowledge distillation from an autoregressive teacher-model . WaveGlow trained a feed-forward network without knowledge distillation. Recently, generative adversarial networks (GANs) were able to match the quality of autoregressive and large feed-forward models. ParallelWaveGAN trained a mel-inversion network, by combining adversarial networks with a multi-scale spectral loss. Other methods, employed multi-scale spectral loss to train convolutional neural networks for spectrogram inversion , to directly predict the parameters of a differentiable synthesizer and for speech synthesis with sine excitation as input .
Singing Synthesis and Conversion Previous singing synthesis methods applied unit selection methods or HMM based parametric methods . Neural networks were used to synthesize singing , by training a WaveNet-like network conditioned on notes and lyrics to generate vocoder features. This method was later extended to synthesize new singers based on a few minutes of them singing . Mellotron , presented a sequence to sequence architecture conditioned on text and pitch to synthesize singing, without training on a singing dataset. Recently, non-autoregressive models were used for singing synthesis. Feed-forward transformers removed the constraint of time-aligned phonemes, by employing a duration model. Generative adversarial networks , conditioned on pitch and time-aligned phonemes, predicted a singing spectrogram in a single pass instead of a frame-by-frame prediction.
Initial methods for the task of singing voice conversion , relied on parallel datasets composed of paired samples of singers singing the same piece. The method of Unsupervised Singing Voice Conversion learned to convert between a fixed set of singers without relying on a parallel-dataset. This was done by learning singer-agnostic features via a domain confusion term on the singer identity. PitchNet further improved the method, by applying an additional domain confusion term on the pitch, but was only applied on conversion between a fixed group of male singers.
Very recently, variational autoencoders were used to convert between singers and their vocal techniques , learned from non-parallel corpora. However, their method was limited to mel-spectrograms, thereby bounding the audio quality. A method employing an automatic speech recognition (ASR) engine was used for singing voice conversion. The ASR extracts phonemes probabilities, which the model then converts to the acoustic features of the target singer. The method was demonstrated on a many-to-one scenario and was only trained on a singing dataset, while our method is applicable both to many-to-many conversion and can also be trained on either speech or singing domains. Similar to us, used WaveRNN conditioned on phoneme probabilities, pitch and a speaker i-vector to generate singing audio in the waveform-domain. Our method differs by (i) using a non-autoregressive model for real-time generation, and (ii) the usage of perceptual losses which, as we demonstrate, greatly boost the performance of our method.
Method
The proposed model is based on a Generative Adversarial Network, with a generator network and a discriminator network . The model is conditioned on both speech and musical features, while in the multi-singer generation case, it is additionally conditioned on a learned singer identity vector. Each feature set is forwarded via a separate context-stack, similar to , before feeding it to . The generator is a non-autoregressive WaveNet, which generates the audio waveform directly from a random noise vector. Fig. 1 depicts the architecture.
Recent works demonstrated the need of both linguistic and musical features in the context of singing generation. As a result, we extract both the loudness measure and the fundamental frequency (F0) as the musical features. Loudness is represented by the log-scaled A-Weighting of the power spectrum, , while F0 is extracted using CREPE , denoted by , similarly to . We experimented with different representations of F0, such as: octave, note, etc., and achieved similar performance.
In preliminary experiments, we observed that using F0 as input produces inconsistent shakes in the pitch of the generated samples. Therefore, we turn into conditioning on a synthesized melody generated from the F0 instead. The melody is synthesized via a single sinusoid sine-excitation, denoted by .
For speech features, we follow and utilize an intermediate representation from a pre-trained acoustic model optimized for the task of Automatic Speech Recognition (ASR) as an additional input to the model. Specifically, we use the public implementation of Wav2Letter , denoted by . Since an ASR network is speaker-agnostic by design , our method does not require any disentanglement terms or domain-specific (speaker) confusion terms.
Finally, we concatenate and upsample the features in the temporal domain to match the audio frequency.
Single Singer Objective Function We follow the least-squares GAN setup where the discriminator and generator would like to minimize the following terms,
accordingly. is the set of samples, is the audio sample synthesized from a random noise vector sampled from a uniform distribution , and the concatenated features .
In addition, we include a reconstruction loss to further improve optimization stability. We note that two audio samples might be perceptually similar while being the exact opposite in the waveform representation, e.g., in the case of a simple phase inversion (multiply by -1). To mitigate this, we include a spectral amplitude distance loss in multiple FFT resolutions . The spectral amplitude distance loss, for a given FFT size , is defined as follows:
where and denotes the Forbenius and the norms, and denotes the Short-time Fourier transform magnitudes of the original and synthesized samples respectively, and the number of elements. The first term of the expression emphasizes spectral peaks, while the second penalizes silent sections of the audio. The multi-resolution loss is defined as the sum of the above loss for multiple scales:
Lastly, to further improve the generation quality, we add perceptual losses on top of the generator output. Specifically, we compute the -distance between intermediate activations of the ASR network, and the CREPE network, as follows:
Overall, the optimization loss for the generator, , is defined as:
where are weight factors to balance the contribution of each loss term.
Multi-singer Training Losses In the multi-singer regime, we include a speaker embedding as an additional input to . The speaker embeddings are learned during training and stored in a Look Up Table. Then, the reconstruction of a sample, , pronounced by speaker is updated to be .
Moreover, we introduce two additional training schemes. The first one is performed by converting a sample from singer to singer , while omitting the reconstruction loss. The second one introduces novel virtual training samples by creating parallel samples using back-translation and mixup . This was previously shown by to improve singing voice conversion.
These additional objective functions for unaligned samples are defined as follows,
where . Note the reconstruction loss is omitted, since we do not have the target sample of singer singing sample .
To virtually simulate unseen singers we follow the mixup scheme while generating a conversion to a new virtual singer. Specifically, we use a convex combination of two different singers embeddings, and , as:
where is drawn from the uniform distribution. Sample is then generated from , and translated back to singer as follows: . The produced set of artificial examples is denoted as . Finally, the discriminator and generator are optimized by minimizing the loss over supervised, unaligned and virtual samples:
Note that in the mixup setting, the .
Architecture Generator , is based on a non-causal WaveNet architecture . It receives as input a random noise vector sampled from a uniform distribution, and the input features described above. Each input feature () is passed through a separate convolutional stack , which is composed of two blocks of eight non-causal convolutional layers. The layers in each block have an exponentially increasing dilation rate. Each layer employs 128 filters and a kernel-size of 3. The features are then upsampled by a series of interleaved nearest neighbor upsampling and convolutional layers. Once temporally-aligned to the audio, the features are concatenated to form the conditioning signal.
The generator is composed of a series of 30 non-causal layers ordered in three blocks. The dilation rate in a single block of layers is exponentially increasing. Thus, the model has a receptive field of 3,072 samples, which means each sample is generated based on a window of in future and past directions. Each layer has 128 residual-channels, 128 skip-channels and a kernel-size of 3. Our model achieves inference speed of 10.14 times faster than real-time, on a single Tesla V100 GPU.
The discriminator is composed of ten layers of 1-D convolutions followed by a leaky-ReLU activation with a leakiness of 0.2, with a linearly increasing dilation rate. Each convolution layer consists of 128 filters with a kernel-size of 3. The discriminator outputs a prediction for each time-step in the input audio signal. The loss is computed by averaging across the time-domain. Both the discriminator and generator apply weight normalization on all layers.
Experiments
We perform a series of experiments to evaluate the proposed method against several baselines. We experimented with generating singing using speech-only, singing-only and mixed datasets. We explore both many-to-one conversion on a large single identity corpus and many-to-many conversion using learned identities from a corpus with a variable amount of audio per identity. Moreover, we perform an extensive ablation study to better understand the contribution of each component. Audio samples are available online at https://singing-conversion.github.io/, as well as in the supplementary material.
Datasets We report results on several datasets. LJ is a large single speaker speech corpus with approximately 24 hours of audio recording. LCSING is a studio recordings of a single professional singer , which was filtered using an off-the-shelf voice activity detector and contains 3 hours and 40 minutes of very expressive and high dynamic range recordings, including some melodic singing without lyrics. For the multi-speaker experiments, we use the speech corpus, VCTK , to learn 109 singers with 44 hours of audio recordings. Finally, the NUS-48E dataset, which includes six male singers and six female singers, has both singing and reading of four songs per voice, resulting in a total of 15 minutes per singer. All audio was down-sampled to 16kHz with a single channel. All datasets were randomly split according to a 80%/10%/10% of train/val/test partitions.
Hyperparameters We train our models for 800K steps, with a batch-size of 8 one second long audio segments. We use RADAM optimizer with a learning rate of 0.0001. The learning rate is halved every 200K steps. The discriminator joins the training process after 100K steps and the perceptual losses after 50K steps. For the CREPE perceptual loss, we use the intermediate-activation before the final sigmoid activation. For the ASR network loss, we use the output of the tenth convolutional block. We use for the weight factors in Eq. 5. The model is trained with a mixup batch every 3 steps after 100K steps of training had passed.
Singing conversion All experiments were performed by converting singing recordings of the NUS-48E to the target identities learned from the datasets described above. Therefore, experiments conducted on the LJ, LCSING and VCTK evaluate the methods’ invariance to the input voice identity, while experiments conducted on the NUS-48E dataset, evaluate the methods’ performance while converting across a fixed set of singers.
Evaluation metrics are based on subjective and objective success metrics: (i) Mean Opinion Scores (MOS), human raters rate the naturalness of the audio samples on a scale of 1–5. Each experiment, included 40 randomly selected samples rated by 20 raters. (ii) ABX testing for similarity, in which we present each rater with two audio samples A and B. The examples originate from two different singers. These two samples are followed by a third utterance X randomly selected to be from the same identity as A or B. Next, the rater must decide whether X has the same identity as A or B. We report the success rate across all raters. (iii) Automatic identification metric by training a multi-class classifier on the ground-truth training partitions of all datasets and reporting the success rate of the classifier, similar to . (iv) Voicing Decision Error (VDE) , which measures the portion of frames with voicing decision error, (v) F0 Frame Error (FFE) , measures the percentage of frames that contain a deviation of more than 20% in pitch value or have a voicing decision error.
Mellotron showed good results on singing generation by training on speech datasets. Therefore, we use it as the baseline for the experiments on speech datasets. Given a template singing sample, Mellotron extracts the rhythm, the alignment between text and spectral features, which is then used to generate the conversion. For emotive samples, like in the NUS-48E dataset, the rhythm extraction sometime failed. Therefore, we used a forced-aligner to create a synthetic alignment map and replace the rhythm extracted by Mellotron with the synthetic one before generating the conversion.
For a fair comparison we do not report Mellotron results on LCSING dataset due to instability in the training caused by the following reasons. In the LCSING dataset there are only 40 minutes of transcribed segments and many of the recordings contain non-lexical vocables.
Our experiments on NUS-48E dataset involved baselines which showed convincing results: WGANSing and Unsupervised Singing Voice Conversion (USVC) . Both methods require multi-singer voice dataset for training. Therefore, we do not apply them on LJ and VCTK. For the single singer dataset, we use the same architecture as USVC but without the confusion term.
Table 1 presents the results for all of the above models. On speech datasets, LJ and VCTK, our models outperform Mellotron, despite the latter utilizing the underlying text as input. In addition to subjective quality, our method predicts both voicing decision and pitch accuracy better than Mellotron. Results on LCSING, show that our method is better than the baseline and is able to generate recognizable samples. On a multi-singer dataset, NUS-48E, our method generates subjectively higher quality samples, which are more identifiable than the baselines.
Ablation We perform ablation for the suggested single singer training framework. Table 2 summarizes the results. The comparison between the F0 condition and the Melody condition, shows that providing melody as input to the model reduces the overall error with regard to pitch generation. The addition of the ASR perceptual loss slightly improves the model MOS scores at the cost of slightly reducing the pitch metrics. Adding the CREPE perceptual loss adds a significant gain to the model performance across all metrics.
Conclusion
We present an unsupervised method that can convert a singing voice to a voice that is sampled either as speaking or singing. The method employs multiple pre-trained encoders and perceptual losses and achieves state of the art results on both objective and subjective measures. Conditioning the generator on a sine-excitation was shown to be beneficial while further improving the results. As future work, we would like to focus on temporal modification of the input singing to further match the style of the target singer.