HiddenSinger: High-Quality Singing Voice Synthesis via Neural Audio Codec and Latent Diffusion Models

Ji-Sang Hwang, Sang-Hoon Lee, Seong-Whan Lee

I Introduction

Singing voice synthesis (SVS) systems aim to generate high-quality expressive singing voices from musical scores. Recent advancements in generative models have led to rapid development in deep-learning-based SVS systems, resulting in high performance. Most SVS systems first synthesize an intermediate acoustic representation, such as Mel-spectrogram, from a musical score using an acoustic model . Subsequently, separately trained vocoders convert the generated representation into audio, as shown in Fig 1(a).

However, conventional two-stage SVS systems face certain limitations. These systems depend on pre-defined intermediate representation, making it difficult to apply latent learning to improve audio generation. Moreover, a training-inference mismatch problem occurs because the predicted intermediate representation differs from the ground-truth intermediate representation. To resolve these issues, an end-to-end SVS system, VISinger , directly synthesizes audio by employing variational inference.

Although existing systems can improve the audio quality, several challenges remain: 1) SVS systems require high-dimensional audio or a linear-spectrogram to synthesize high-fidelity audio, resulting in high computational costs in high-dimensional space. 2) The training-inference mismatch problem persists in end-to-end systems. A gap between the posterior distribution from the audio and the prior distribution from the musical score exists, which results in inaccurate pitch and mispronunciations in the generated singing voice. Moreover, systems based on Normalizing Flows are trained in a backward direction but perform inference in a forward direction . 3) SVS systems require audio-musical score corpora for training, wherein it is time-consuming to obtain high-quality paired datasets.

To address the aforementioned problems, we propose HiddenSinger, an advanced high-quality SVS system utilizing a neural audio codec and latent diffusion models. Our approach involves multiple components to enhance the synthesis process: First, we introduce an audio autoencoder that can efficiently encode audio into a compressed latent representation, resulting in a lower-dimensional representational space. We also adopt residual vector quantization in the audio autoencoder to regularize the arbitrarily high-variance latent space. Subsequently, we employ the powerful generative ability of latent diffusion models to generate a latent representation conditioned on a musical score, which is converted into audio through the audio autoencoder. Moreover, we propose an unsupervised singing voice learning framework that leverages unpaired singing voice data containing only audio. The experimental results demonstrate that HiddenSinger outperforms previous SVS models in terms of audio quality. Furthermore, our model can synthesize high-quality singing voices, even for speakers who are represented in unpaired data, by using the proposed unsupervised singing voice learning framework (HiddenSinger-U).

Our study makes the following contributions:

We introduce HiddenSinger, which utilizes a neural audio codec and latent diffusion models to synthesize high-quality singing voices. The latent generator generates a latent representation conditioned on a musical score. Subsequently, the audio autoencoder synthesizes high-quality singing voice audio from the generated latent representation.

We extend our proposed model to HiddenSinger-U, an unsupervised singing voice learning framework that performs training with both paired and unpaired datasets using acoustic features from audio. HiddenSinger-U can synthesize a high-quality singing voice of a speaker without a musical score during training.

The proposed model is demonstrated to outperform previous SVS models. Audio samples are available at https://jisang93.github.io/hiddensinger-demo/

II Related Studies

Singing voice synthesis (SVS) systems are designed to generate a singing voice based on a musical score. Since singing voices comprise significant pitch variability and an extended duration of vowel, SVS systems require additional input data, such as note pitch, note duration, and lyrics. Conventional SVS systems follow a two-stage manner comprising an acoustic model and a vocoder to synthesize a realistic singing voice. Although previous SVS systems improved the singing voice quality, the two-stage pipeline has inherent limitations that prevent it from surpassing the upper bound of the vocoder performance.

To address these limitations, researchers have proposed end-to-end SVS systems that use a well-learned latent representation to enhance the quality of singing voices and simplify the training procedure. However, the end-to-end method still faces problems, including a training-inference mismatch problem. Specifically, the gap between the posterior and prior distributions leads to degraded audio reconstruction performance. In this study, we leverage a well-learned latent representation, which is converted into an audio codec, to improve the quality of the reconstructed audio.

II-B Neural Audio Synthesis

To generate natural audio, neural vocoders are generally used to convert signal processing components, such as the Mel-spectrogram, into raw waveform audio. For high-quality audio generation, a generative adversarial network (GAN)-based neural vocoder adopts a multi-scale discriminator and a multi-period discriminator to capture the specific characteristics of the waveform audio. Although a diffusion-based neural vocoder has been presented, several limitations persist in the audio quality and inference speed in tasks concerning waveform audio generation.

In recent developments, neural audio codecs have emerged in conjunction with neural vocoders. These audio codecs efficiently compress the audio in an autoencoder architecture. For improved compression, approaches such as SoundStream introduce residual vector quantization, leading to enhanced coding efficiency. Encodec also represents the audio as discrete units with the residual vector quantization and incorporates a multi-scale short-time Fourier transform (STFT)-based discriminator to reduce artifacts in the reconstructed audio. Drawing inspiration from these studies, we adopt neural audio codecs to achieve high-fidelity audio generation and computational efficiency.

II-C Diffusion Probabilistic Model

Diffusion probabilistic models (also known as diffusion models) are a class of generative models that have achieved remarkable results in various domains, such as image , audio and video generation. Particularly, in the audio domain, previous studies have mainly used diffusion models to generate acoustic features.

For the acoustic feature generation, Grad-TTS , DiffSinger , and DDDM-VC utilize the diffusion-based decoder to generate a high-quality Mel-spectrogram. Each model uses a conditional-diffusion decoder to condition text distribution for a text-to-speech system , the musical score for a SVS system, and speaker information for a voice conversion system. To improve the generation efficiency by compressing the Mel-spectrogram into discrete latent space, DiffSound introduces a discrete diffusion-based token decoder in a non-autoregressive manner. Make-an-audio adopts the latent diffusion models to generate a continuous latent representation that converts into Mel-spectrogram. For waveform generation, Diffwave and WaveGrad generate high-fidelity speech waveform from the Mel-spectrogram. In contrast to the above approaches, WaveGrad 2 and FastDiff adopt an end-to-end manner that generates the audio without any intermediate features (e.g., Mel-spectrogram). Inspired by the success of diffusion-based generation, we adopt latent diffusion models to generate a latent representation conditioned on a musical score.

III Preliminary

Diffusion models comprise two processes: a forward process (diffusion process) and reverse process (denoising process). In the forward process, the data X0X_{0} are gradually corrupted with a tiny Gaussian noise through a TT-step Markov chain. The reverse process, which follows the reverse trajectory of the forward process, aims to generate the data X0X_{0} from the Gaussian noise XTX_{T}.

In , a stochastic differential equation (SDE) was used to approximate the trajectory between X0X_{0} and XTX_{T}. In the speech domain, Grad-TTS and Guided-TTS applied an SDE to the text-to-speech task. Following , forward process that perturbs the data X0X_{0} into the noise XTX_{T} is defined with the pre-defined noise schedule βt=β0+(βT−β0)t\beta_{t}=\beta_{0}+(\beta_{T}-\beta_{0})t:

where WtW_{t} represents the standard Brownian motion and tt denotes a continuous timestep t∈[0,T]t\in[0,T].

The reverse process is defined as a reverse-time SDE that formulates the trajectory from Gaussian noise XTX_{T} to the data X0X_{0} as follows:

A neural network sθs_{\theta} learns to estimate the score, which is parameterized by θ\theta, to model the data distribution pt(Xt)p_{t}(X_{t}). By solving Eq. 2, X0∼p0(X)X_{0}\sim p_{0}(X) can be obtained by starting from the noisy data XTX_{T} and iteratively removing the noise using the score estimation networks sθs_{\theta}.

IV HiddenSinger

In this paper, we propose a SVS system using neural audio codecs and latent diffusion for high-quality singing voice audio. We introduce an audio autoencoder using residual vector quantization to achieve high-fidelity audio generation and computational efficiency. Additionally, we adopt latent diffusion models in a latent generator to generate a latent representation conditioned on a musical score, which is converted into audio by the audio autoencoder. Furthermore, we extend HiddenSinger to HiddenSinger-U, which can train the model without musical scores. In the following subsection, we describe the details of HiddenSinger and an unsupervised singing voice learning framework (HiddenSinger-U).

For efficient coding and high-quality audio generation, we introduce the audio autoencoder to compress the audio into an audio codec, which provides a low-dimensional representation. The audio autoencoder comprises three modules: an encoder, residual vector quantization (RVQ) blocks, and a decoder, as illustrated in Fig. 2 (a).

The encoder takes a high-dimensional linear-spectrogram as the input and extracts a low-dimensional continuous latent representation z0z_{0} from the audio yy. Inspired by , the latent space is regularized through vector quantization (VQ) to avoid an arbitrarily high-variance of the latent space. A previous study in which sampling was performed using latent diffusion models demonstrated that a model trained on the VQ-regularized latent space achieved better quality than the Kullback-Leibler (KL)-regularized latent space. In our preliminary experiments, we observed that the KL-regularized latent space achieved sub-optimal performance when the diffusion models restored the latent representation. However, conventional VQ is insufficient for high-fidelity audio reconstruction because a quantized vector should represent multiple features of a raw waveform. Therefore, we apply RVQ to the continuous latent representation z0z_{0} for efficient audio compression.

IV-A2 Residual Vector Quantization Blocks

As indicated in Fig. 2 (b), the first vector quantizer discretizes the continuous latent representation z0z_{0} into the closest entry in a codebook. Subsequently, the residual is computed. The next quantizer is used with the second codebook, with this process repeated as many times as the number of quantizers CC. The number of quantizers is related to the trade-off between the computational cost and coding efficiency. We follow the training procedure described in to train the codebook for each quantizer. Furthermore, we apply the commitment loss to stabilize the codebook training. We found that the low-weighted commitment loss helps to converge the RVQ blocks during training:

where z0,cz_{0,c} represents the residual vector of the cc-th quantizer and qc(z0,c)q_{c}(z_{0,c}) denotes the closest entry in the cc-th codebook.

IV-A3 Decoder

The decoder generates a raw waveform from the audio codec y^=G(zq)\hat{y}=G{\left(z_{q}\right)}. We calculate a reconstruction loss Lrecon\mathcal{L}_{recon} between the generated x^mel\hat{x}_{mel} and ground-truth Mel-spectrograms xmelx_{mel} to improve the training efficiency of the decoder. The reconstruction loss is defined as

Moreover, we adopt adversarial learning to improve the quality of the generated audio. We use a multi-scale STFT-based (MS-STFT) discriminator , which expands a multi-resolution spectrogram discriminator . The MS-STFT discriminator operates on a multi-scale complex-valued STFT that contains both real and imaginary parts. Similar to the work of , we observed that the MS-STFT discriminator trains the decoder efficiently and facilitates the synthesis of audio with better quality than the combination of a multi-period discriminator and multi-scale discriminator . Furthermore, we adopt the feature matching loss Lfm\mathcal{L}_{fm} , which is a perceptual loss for GAN training:

where zqz_{q} denotes the quantized latent representation, LL is the total number of layers in discriminator DD, NlN_{l} represents the number of features, and DlD_{l} extracts the feature map in the ll-th layer of the discriminator.

IV-A4 Auxiliary Multi-task Learning

We introduce auxiliary tasks based on a lyrics predictor and note-pitch predictor to improve the capability of the linguistic and acoustic information in the audio codec. Each predictor takes the compressed latent representation zqz_{q} to predict a frame-level target feature. We calculate the connectionist temporal classification (CTC) loss between the predicted and target feature. We only apply the CTC loss to paired datasets that contain a musical score.

IV-A5 Final Loss

The final loss term for the audio autoencoder is defined as:

where λ∗\lambda_{*} is the loss weight, Llyrics\mathcal{L}_{lyrics} represents the CTC loss between the predicted and ground-truth lyrics, and Lnote\mathcal{L}_{note} denotes the CTC loss between the predicted and ground-truth pitch IDs according to the musical instrument digital interface (MIDI) standard.

IV-B Condition Encoder

We present a condition encoder to guide the diffusion models. The condition encoder comprises a lyrics encoder, a melody encoder, an enhanced condition encoder, and a prior estimator.

The lyrics encoder takes a phoneme-level lyrics sequence with positional embedding as the input, and then extracts a lyrics representation. We use a grapheme-to-phoneme tool to convert the lyrics sequence into a phoneme-level lyrics sequence before feeding it into the lyrics encoder.

IV-B2 Melody Encoder

We introduce the melody encoder to generate a singing voice with an adequate melody from a musical score. Before using the musical score, we divide the notes into a phoneme-level note sequence. A Korean syllable generally comprises an onset, nucleus, and coda. Following the previous Korean SVS systems , we assign onset and coda to a maximum of three frames with the remainder considered as the nucleus.

Subsequently, the melody encoder extracts a melody representation from the concatenation of a note pitch, note duration, and note tempo embedding sequence with positional embedding. The note pitch sequence is transformed into the note pitch embedding. The note duration embedding sequence is represented by a fixed set of duration tokens, among which the resolution is represented by a specific note duration (e.g., the 64th note). The note tempo is calculated in beats per minute and encoded into the tempo embedding.

IV-B3 Enhanced Condition Encoder

The enhanced condition encoder encodes the summation of the outputs of the lyrics and melody encoders to provide a more informative condition representation hcondh_{cond}. Before summing the two representations, they are expanded into the frame-level based on the note duration. In our preliminary experiments, we observed that the enhanced condition encoder effectively stabilized the pronunciation of synthesized singing voices, similar to the result in .

IV-C Latent Generator

We adopt the latent diffusion models in the latent generator to generate the latent representation of the audio autoencoder. The latent representation z^0\hat{z}_{0} is sampled using the latent diffusion models, following which the generated latent representation z^0\hat{z}_{0} is converted into the audio codec in the audio autoencoder. Furthermore, the latent representation is normalized to ease the sampling.

We use data-driven priors in the latent diffusion models to improve their generation abilities. Previous studies have demonstrated that the use of data-driven priors helps approximate the trajectories between the complex data and known priors. Following , we design the diffusion models to start denoising from noise close to the target z0′z^{\prime}_{0}, which is easier than denoising from standard Gaussian noise. We predict μ^\hat{\mu} from the condition representation hcondh_{cond} using the prior estimator of the condition encoder. We apply the negative log-likelihood loss Lprior\mathcal{L}_{prior} between the normalized latent z0′z^{\prime}_{0} and the predicted μ^\hat{\mu} to consider μ^\hat{\mu} as a mean-shifted Gaussian distribution N(μ^,I)\mathcal{N}(\hat{\mu},I).

IV-C2 Latent Diffusion Models

The diffusion process is defined using a forward stochastic differential equation (SDE) with the data-driven priors given a time horizon t∈[0,1]t\in\left[0,1\right]. The forward SDE converts the normalized latent representations z0′z^{\prime}_{0} into Gaussian noise:

where WtW_{t} is the standard Brownian motion and βt\beta_{t} is the non-negative pre-defined noise schedule. Its solution is expressed as:

According to the properties of Ito^\hat{\text{o}}’s integral, the transition density pt(zt′∣z0′)p_{t}{\left(z^{\prime}_{t}|z^{\prime}_{0}\right)} is the Gaussian distribution pt(zt′∣z0′)∼N(zt′;ρt,λt)p_{t}{\left(z^{\prime}_{t}|z^{\prime}_{0}\right)}\sim\mathcal{N}{\left(z^{\prime}_{t};\rho_{t},\lambda_{t}\right)}, as follows:

We define the reverse process as an SDE solver to obtain the normalized latent representations z0′∼p0(z′)z^{\prime}_{0}\sim p_{0}{(z^{\prime})}. We use a score estimation network sθs_{\theta} to approximate the intractable score:

Following , we compute the expected value of the estimated gradients of the log-density of the noisy latent zt′z^{\prime}_{t}:

where ∇zt′log⁡pt(zt′∣z0′)=−λt−1ϵt\nabla_{z^{\prime}_{t}}\log{p_{t}{(z^{\prime}_{t}|z^{\prime}_{0}})}=-\lambda_{t}^{-1}\epsilon_{t} and ϵt∈N(0,I)\epsilon_{t}\in\mathcal{N}{(0,I)}. Furthermore, we adopt a temperature parameter τ\tau for the data-driven prior distribution N(μ^,τ−1I)\mathcal{N}{\left(\hat{\mu},\tau^{-1}I\right)} during sampling, which helps the latent generator to maintain the quality when τ>1\tau>1, similar to the approach in .

We jointly optimize the latent generator and condition encoder based on the following objective:

where λprior\lambda_{prior} is the loss weight for the prior loss Lprior\mathcal{L}_{prior}.

IV-D Unsupervised Singing Voice Learning Framework

Conventional SVS models require paired data (audio-musical score corpora) for training. Furthermore, these models cannot synthesize the singing voice of an untrained speaker without special techniques such as zero-shot adaptation. We extended our proposed model to HiddenSinger-U, an unsupervised singing voice learning framework, to mitigate the difficulty of collecting paired datasets. This framework enables the model to use unlabeled data during training. We introduce two additional encoders into the condition encoder to model the unsupervised lyrics and melody representation, as shown in Fig. 2 (c): an unsupervised lyrics encoder (lyrics-U encoder) and an unsupervised melody encoder (melody-U encoder). Furthermore, we employ contrastive learning in the proposed framework.

We use a self-supervised speech representation method for the linguistic information. Previous works have demonstrated that the speech representation from the middle layer of a self-supervised model contains phonetic information. Therefore, the phonetic information can be leveraged by extracting the self-supervised representation from the target audio. We perform information perturbation before extracting the self-supervised representation to mitigate speaker information in the target audio. The information perturbation causes the self-supervised model to focus on extracting only phonetic information. Subsequently, the lyrics-U encoder encodes the self-supervised representation into a frame-level unsupervised lyrics representation.

IV-D2 Melody-U Encoder

SVS models still require melody information of the target audio to synthesize singing voices. We first extract the fundamental frequency (F0F0) from the audio to extract melody information. Thereafter, we quantize the F0F0 and encode it into a pitch embedding to obscure speaker information in the target audio. Subsequently, the melody-U encoder takes the pitch embedding to extract a frame-level unsupervised melody representation.

IV-D3 Contrastive Learning

We observed that it is insufficient to only use the objective Llg\mathcal{L}_{lg} to optimize HiddenSinger-U owing to the gap between the paired representations (e.g., the lyrics and unsupervised lyrics representation). To maximize the agreement and penalize the dissimilarity between the paired representations, we introduce the contrastive loss for the paired data as follows:

where cos⁡(⋅,⋅)\cos(\cdot,\cdot) calculates the cosine similarity between the pairs, τcont\tau_{cont} denotes the temperature, and ξ[k≠t]\xi_{[k\neq t]} represents a set of random time indices as negative samples. Following , we randomly select several unmatched frames within each paired representation for negative samples. We apply the contrastive loss for each type of representation h∗∈[hlyrics,hmelody]h_{*}\in\left[h_{lyrics},h_{melody}\right]. The gap between the paired representations can be reduced by adopting the contrastive terms Lcont∗\mathcal{L}_{cont_{*}} in the objective Llg\mathcal{L}_{lg}.

V Experiment and Results

We trained HiddenSinger on the Guide vocal datasethttps://bit.ly/3GbEUIX to synthesize the singing voice The Guide vocal dataset contains approximately 157.39157.39 hours of audio for 4,0004,000 paired Korean songs. We divided the audio into segments of two-bar segments to facilitate the model training, resulting in 93,12793,127 samples. Subsequently, we divided our dataset into three subsets: 89,18689,186 samples for training, 1,9751,975 samples for validation, and 1,9661,966 samples for testing.

We trained HiddenSinger-U using the Guide vocal dataset and an internal singing voice dataset containing approximately 3.303.30 hours of audio for 316316 Korean songs that do not have musical scores to evaluate the unsupervised singing voice learning framework. The internal dataset was divided into three subsets: 1,1301,130 samples for training, 9999 samples for validation, and 9797 samples for testing. Moreover, we considered specific speakers in the Guide vocal dataset as unlabeled data during training. To train the audio autoencoder, we used the aforementioned dataset, a multi-speaker singing datasethttps://bit.ly/3Q9rOkn, and children singing dataset , which contain a total of 285.1285.1 hours of audio for 8,7818,781 K-pop songs.

V-A2 Pre-processing

We downsampled the audio at 24,000 Hz for training. We transformed the audio into a linear-spectrogram with 1,025 bins to train the audio autoencoder. For the reconstruction loss, we used the Mel-spectrogram with 128 bins. We grouped words into phrases and separated the phrases with the 16th rest in a text sequence for the lyrics encoder input. Subsequently, we converted the text sequence into a phoneme sequence using the grapheme-to-phoneme toolhttps://github.com/Kyubyong/g2p. We used a 64th note resolution for the note duration tokens. We used the range $$ for the tempo values of the tempo tokens. We extracted the self-supervised representation from the middle of XLS-R , pre-trained wav2vec 2.0 with 128 language dataset including Korean, as inputs for the lyrics-U encoder. Prior to the extraction, we resampled the audio at 16,000 Hz and perturbed it. We interpolated the extracted representation back to 24,000 Hz sampling rate.

V-A3 Training

We trained the audio autoencoder using the AdamW optimizer with a learning rate of 2×10−42\times 10^{-4}, β1=0.8\beta_{1}=0.8, β2=0.99\beta_{2}=0.99, and a weight decay of λ=0.01\lambda=0.01. We adopted a windowed generator training for efficiency. We randomly extracted segments of the raw waveform with a window size of 128 frames as the input for the encoder to capture the linguistic features. Furthermore, the decoder took a randomly sliced segment of the quantized latent representation zqz_{q} with a window size of 32 frames. We used the corresponding audio segment from the ground-truth audio as the training target. Four NVIDIA RTX A6000 GPUs were used for the training. The batch size was set to 32 per GPU and the model was trained for up to 1M steps.

We jointly trained the condition encoder and latent generator using the AdamW optimizer with a learning rate of 5×10−55\times 10^{-5}, β1=0.8\beta_{1}=0.8, β2=0.99\beta_{2}=0.99, and a weight decay of λ=0.01\lambda=0.01. We randomly extracted segments of the latent representations z0z_{0} with a window size of 128 frames for efficient training. We used two NVIDIA RTX A6000 GPUs for training and set the batch size to 32 per GPU. The model was trained for up to 2M steps.

V-B Implementation Details

The encoder comprises non-causal WaveNet residual blocks, as proposed by . The decoder uses a HiFi-GAN V1 generator . We implemented 30 quantizers with codebook sizes of 1,024 entries and 128 dimensions for the residual vector quantizer blocks.

V-B2 Condition Encoder

The lyrics, melody, and enhanced condition encoders comprise four feed-forward Transformer (FFT) blocks with relative-position encoding following Glow-TTS . In each FFT block, we set the number of attention heads to 2, the hidden size to 192, and kernel size to 9. The prior estimator is a single linear layer.

V-B3 Latent Generator

As illustrated in Fig. 3, a non-causal WaveNet-based denoiser architecture is used for the score estimation network sθs_{\theta}, similar to the architecture in . We set the number of dilated convolution layers to 20, the residual channels to 256, and kernel size to 3 for the score estimation network. We set the dilation to 1 in each layer. We set β0=0.05\beta_{0}=0.05, β1=20\beta_{1}=20 and T=1T=1 to train the latent generator and τ=1.5\tau=1.5 to sample the latent representation during inference.

V-B4 Unsupervised Learning Module

The lyrics-U and melody-U encoders have the same architecture as the lyrics and melody encoders, respectively, which consist of four FFT blocks with relative-position encoding. We used the 12th layer of the pre-trained XLS-R to extract the self-supervised representation. We quantized F0F0 into 128 intervals to mitigate the speaker information.

V-C Subjective Metrics

We conducted a five-scale naturalness mean opinion score (nMOS) listening test on the test dataset to evaluate the naturalness of the audio. Each audio was evaluated by 15 native Korean speakers. The subjective metrics are reported with 95% confidence intervals in this paper.

V-D Objective Metrics

We calculated the objective metrics to evaluate various types of distance between the ground-truth and synthesized audio. We considered four metrics to evaluate the SVS quality: 1) spectrogram mean absolute error (MAE); 2) pitch error; 3) periodicity error; and 4) F1 score of voiced/unvoiced classification (V/UV F1). We used the implementation of CARGAN to evaluate the pitch, periodicity, and V/UV F1. Moreover, we provided additional objective metrics for the reconstruction quality, namely the perceptual evaluation of speech quality (PESQ) , in Subsection V-F.

where sis_{i} and si′s^{\prime}_{i} denote the ii-th spectrogram frame from the ground-truth and synthesized waveform, respectively. TT represents the frame lengths of the spectrogram.

V-D2 Pitch error

where pip_{i} and pi′p^{\prime}_{i} represent the ii-th extracted pitch representations from the ground-truth and synthesized waveform by using torchcrepehttps://github.com/maxrmorrison/torchcrepe, respectively. As following CARGAN, we only measure the pitch error on voiced parts in a waveform.

V-D3 Periodicity error

where ϕi\phi_{i} and ϕi′\phi^{\prime}_{i} are the ii-th extracted phase features from the ground-truth and synthesized waveform by using torchcrepe, respectively.

Note that the length of the synthesized and target singing voices are the same, because of the musical score that informs the duration of each note. Therefore, we do not consider time alignment, such as dynamic time warping , to calculate objective evaluations.

V-E Singing Voice Synthesis

We compared the audio generated by our proposed models, HiddenSinger and HiddenSinger-U, to the outputs of the following systems: 1) GT, Ground-truth audio; 2) HiFi-GAN , in which we reconstructed the audio from the ground-truth Mel-spectrogram using HiFi-GAN; 3) FastSpeech 2 + HiFi-GAN, in which we added a melody encoder for SVS; 4) DiffSinger + HiFi-GAN; and 5) VISinger , which is an end-to-end SVS system. We trained HiddenSinger-U on the same SVS dataset, of which 10% was defined as unlabeled data. Moreover, for fair comparisons, we trained the HiFi-GAN using the same datasets and training steps that were used to train the audio autoencoder.

As indicated in Table I, according to the subjective audio evaluation, HiddenSinger and HiddenSinger-U outperformed the other SVS models in terms of naturalness. Moreover, our proposed models reduced the pitch error without variation predictions, such as pitch or energy prediction. These results indicate that HiddenSinger can learn accurate pitch information.

However, VISinger achieved better performance in terms of the MAE, periodicity error, and V/UV F1 score. As our proposed models generate the latent representation through stochastic iterations, the stochasticity of the models may increase the distance between the ground-truth and synthesized audio. We computed the F0F0 contour from the synthesized audio of HiddenSinger using Parselmouthhttps://github.com/YannickJadoul/Parselmouth to demonstrate the stochasticity of the models. As indicated in Fig. 4 (a), we performed inference five times for a speaker with the same musical score. It can be observed that HiddenSinger synthesized singing voices that contained appropriate tunes based on the musical score and variations such as intonation. As indicated in Fig. 4 (b), we synthesized singing voices using five different speakers and the same musical score. It can be observed that HiddenSinger generated various styles of singing voices from different speakers.

Furthermore, we visualized the Mel-spectrograms of the synthesized audio to compare the models. Although the shapes of the harmonics that were synthesized by HiddenSinger differed slightly from those of the ground-truth Mel-spectrogram, the harmonics in the high-frequency band of HiddenSinger were more fine-grained than those of the other systems, as illustrated in Fig. 5. These results demonstrate that HiddenSinger generates high-fidelity and natural singing voices using the denoising process that can inject several variations.

V-F Audio Autoencoder

To demonstrate the performance of the audio autoencoder, we evaluated the quality of the reconstructed audio. We reconstructed the singing voice dataset used to train the VISinger for a fair comparison. As each decoder of our audio autoencoders leverages the HiFi-GAN V1 generator , they achieved similar performance to HiFi-GAN in terms of the objective evaluation metrics in Table II. However, in terms of naturalness, our audio autoencoders achieved slightly better performance than HiFi-GAN. Moreover, the reconstruction results of the VISinger exhibited the worst performance in terms of the subjective and objective evaluation measures. These observations suggest that the end-to-end training may reduce the quality of the reconstructed audio, resulting in the upper bound of the audio generation being degraded.

We evaluated the effectiveness of the different combinations of our audio autoencoder and the latent generator. We trained the latent generator separately using different regularized latent spaces. As indicated in Table III, the latent generator with the RVQ-regularized autoencoder outperformed the other combinations. Furthermore, it was difficult for the latent space without regularization to generate the latent representation with the latent diffusion models. These results indicate that the RVQ-regularized latent space is more suitable for sampling targets than the KL-regularized latent space in our setting, similar to the results reported in .

V-G Unsupervised Singing Voice Learning Framework

We compared the changes in the evaluation metrics according to the ratio of unlabeled data in the training dataset to verify the effectiveness of the unsupervised singing voice learning framework. We pre-defined certain speakers as unlabeled data that consisted of only audio for verification. We conducted the nMOS test to evaluate the naturalness of the audio. Moreover, we conducted a four-scale similarity MOS (sMOS) test to evaluate the voice similarity between the ground-truth and generated audio. We evaluated samples of pre-defined speakers that were considered unlabeled data in every setting, except for the 0% and 2% ratio settings in both MOS tests. The 0% ratio setting represents HiddenSinger, which has been trained without the unsupervised singing voice learning framework.

It can be observed from Table IV that the nMOS results were statistically insignificant in most of the settings. This suggests that the unsupervised singing voice learning framework helps the model learn to synthesize a natural singing voice, regardless of changes in the unlabeled ratio. Moreover, the objective evaluations demonstrate that the proposed framework can be trained stably in every setting.

However, as shown in Table IV, the similarity of the synthesized singing voice decreased with an increase in the unlabeled ratio. As there were differences between the note pitch of a musical score and the F0F0 of a human speaker’s singing voice, the models were trained with slightly different speaker identities due to the difference. Therefore, the human listener differentiated between the ground-truth and synthesized audio according to the difference. Although the difference degrades the similarity according to the increasing unlabeled ratio, the proposed framework is effective in synthesizing a natural singing voice with proper linguistic information and a perceptually similar speaker identity. Moreover, the contrastive terms Lcont∗\mathcal{L}_{cont_{*}} and information perturbation aid in stabilizing the training. In particular, it is difficult to synthesize an appropriate singing voice when training is performed without the contrastive terms.

V-H Ablation Study

We conducted an ablation study to verify the effectiveness of each module in the proposed system. The results are presented in Table V. It can be observed that the subjective and objective evaluations significantly degraded with the removal of the enhanced condition encoder. Furthermore, the pronunciation of the synthesized audio was highly inaccurate without the enhanced condition encoder. Therefore, the enhanced condition encoder is necessary for the appropriate functioning of the proposed model.

We performed training on the latent generator with the audio codec zqz_{q} as the target of the latent diffusion models. Table V indicates that the generation of z0z_{0} could provide more natural audio than the generation of zqz_{q} in the latent generator. As the RVQ blocks may refine the sampled latent representation z^0\hat{z}_{0} with residual operations, the generation of z0z_{0} is superior in terms of naturalness.

Furthermore, we considered a standard Gaussian as the priors following the original denoising diffusion probabilistic models . However, the data-driven priors outperformed the standard Gaussian-based priors. This indicates that the trajectory between the data space and data-driven priors can be more stably approximate than the trajectory between the data space and standard Gaussian.

VI Conclusions

We have introduced HiddenSinger, a novel approach that enables the synthesis of high-quality and high-diversity singing voice audio through the integration of a neural audio codec and latent diffusion models. Our study demonstrated the efficacy of the audio autoencoder in reconstructing high-fidelity audio using low-dimensional audio codecs. Furthermore, we successfully generated latent representations conditioned on a musical score using latent diffusion models. The audio was successfully reconstructed from the generated latent representation by the audio autoencoder. We extended our model to an unsupervised singing voice learning framework that can be trained without lyrics and note information using self-supervised representation. Our latent diffusion models could be used in any speech domain, including text-to-speech and voice conversion systems. However, our model still has limitations regarding novel singing style adaptation, not voice. In future works, we will attempt to implement a zero-shot singing style transfer by adopting style-generalized generative models.

VII Discussion

Recently, neural audio codecs have been used in various tasks . As following concurrent works , our proposed model can be extended to a text-to-speech system. Moreover, we can address the data scarcity problem by applying our unsupervised learning framework to a low-resource language.

VII-B Social Negative Impact

Although HiddenSinger may have practical applications such as podcasts or music generation, there is an increased risk of potential misuse of such technologies. In particular, unauthorized usage of data from web crawlers in SVS can give rise to concerns related to copyright infringement and voice spoofing. We want to emphasize that we strongly discourage the utilization of our work for any illicit or unethical purposes.

VII-C Limitation

Although we adopt the latent diffusion models for high-efficient latent generation, the diffusion models require a number of iterative processes to generate the representations. In the future, we will introduce the consistency models to distill the teacher diffusion models for a single-step generation.

References