Speech Resynthesis from Discrete Disentangled Self-Supervised Representations

Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, Emmanuel Dupoux

Introduction

Learning unsupervised speech representations, both continuous and discrete, has seen a significant leap in performance following the recent success of Self-Supervised Learning (SSL) methods . In the self-supervised setting, unlabeled inputs define an auxiliary task that can generate pseudo-labeled training data. This data can then be used to train a model using supervised techniques. The learned representations are often used for downstream tasks with a minimal amount of supervised data. For example, learning speech recognition with only 10 minutes of transcribed speech and 53K hours of untranscribed speech as in wav2vec 2.0 and HuBERT . The learned self-supervised discrete representations also showed impressive performance on the conditional and unconditional spoken generative language modeling (GSLM) task .

Despite its success, most studies on SSL for speech are focused on generating and evaluating the quality of the learned representations in the context of Automatic Speech Recognition (ASR). It remains unclear how suitable these representations are for speech synthesis. Moreover, in the context of expressive and controllable generation, it is unknown to what extent the speaker identity and F0 information are encoded in the learned representations.

Traditionally, speech synthesis and text-to-speech models produce Mel-spectrogram autoregressively given textual features as input. Next, a vocoder is applied to reconstruct the phase from the Mel-spectrogram (e.g., Griffin-Lim , WaveNet , WaveGlow , or HiFi-GAN ). In this study, we suggest using the learned speech units as an input to a vocoder module with no spectrogram estimation. We additionally augment the learned units with quantized F0 representation and a global speaker embedding. Figure 1 presents the overall proposed method. This allows the evaluation of the learned units with respect to speech content, speaker identity, and F0 information, as well as better control the audio synthesis. We experiment with signal reconstruction, voice conversion, and F0 manipulation using several datasets and encoder models. Finally, equipped with our previous findings, we demonstrate how the learned units can function as an ultra-lightweight speech codec. Following the proposed method, we can reach an encoding rate of 365 bits per second (bps) while being superior to the baseline methods by a significant margin, including lightweight and heavyweight codecs.

Our Contribution: (i) We demonstrate the usage of discrete speech units, learned in a self-supervised manner, for high-quality synthesis purposes (no Mel-spectrogram estimation); (ii) We provide an extensive evaluation of the SSL speech units from a synthesis point of view, i.e., signal reconstruction, voice conversion, and F0 manipulation; (iii) We build an ultra-lightweight speech codec from the obtained speech units.

Related Work

Unsupervised Speech Representation Learning Studies on unsupervised speech representation learning can roughly be divided into reconstruction and self-supervised learning methods. Auto-encoding is the common approach for signal reconstruction, where speech is first encoded into a low-dimensional latent representation, and then decoded back to speech. Various constraints can be imposed on the encoded space, such as temporal smoothness , discreteness , and hierarchy .

SSL methods have shown remarkable results for ASR , phoneme segmentation , and GSLM . The authors in suggested training a convolutional neural network to distinguish true future samples from random distractor samples using a Contrastive Predictive Coding (CPC) loss function. Similar to CPC, the authors in use an encoder and a predictor, which is trained contrastively to distinguish positive and negative samples. However, unlike it discretizes and masks segments of the encoder’s output. In HuBERT , the model is trained with a masked prediction task similar to BERT but with masked continuous audio signals.

Speech Resynthesis Recent advancements in neural-based vocoders enabled generating natural-sounding speech and music . These are often conditioned on the log Mel-spectrogram for the generation process. Recently, unsupervised learning of low bitrate speech representations was explored under the Zero-Resource Challenge . Vector-Quantized Variational Auto-Encoder (VQ-VAE) employs a learned fixed-sized codebook, decoded by a WaveNet model for speech synthesis. In the authors proposed a VQ-VAE model followed by an FFTNet vocoder model . The authors in suggested using transformer together with a VQ-VAE model for unsupervised unit discovery, and in they combine vector quantization with contrastive predictive coding for acoustic unit discovery. The study by is the closest to our work. In which, the authors suggest training a VQ-VAE with two encoders, one for the waveform and the other for F0. The authors demonstrated how such modeling improves overall generation quality. In contrast, we study SSL-based speech encoders and empirically show these representations are better disentangled, and apply them as an ultra-low bitrate speech codec. Another line of work suggests using intermediate representations obtained from an ASR acoustic model. These representations are being used together with the identity and prosodic information for voice conversion . Unlike all of the above, we suggest synthesizing speech directly from the discrete units. Moreover, the resynthesis process sheds light on the encoded information in each of the evaluated representations.

Speech Codec Speech codecs typically employ a carefully hand-engineered pipeline combining an encoder and a decoder which is influenced by speech production physics to remove redundancies in the data and yield a compact bitstream. Low bitrate parametric speech codecs have long been studied , but their quality has been severely limited. Despite some advances , modeling the excitation signal has remained a challenge. Neural speech codecs have been recently proposed and demonstrated promising results . In an LPCNet vocoder was conditioned on hand-crafted features and a uniform quantizer. In a WaveNet model was conditioned on discrete units obtained from a VQ-VAE model, while in the Opus codec was fed to WaveNet.

Method

The proposed architecture is comprised of three pre-trained and fixed encoders, namely: (i) content encoder; (ii) F0 encoder; (iii) speaker identity encoder; and a decoder network. The first two encoders extract discrete representation from the raw audio while the latter extracts a single global representation. The overall architecture is depicted in Figure 1.

Content Encoder The input to a content encoder network, EcE_{c}, is a speech utterance, x\bm{x}, and the output is a sequence of spectral representations sampled at a low frequency as follows Ec(x)=(v1,…,vT′)E_{c}(\bm{x})=(\bm{v}_{1},\dots,\bm{v}_{T^{\prime}}). We evaluated three state-of-the-art unsupervised representation learning functions as EcE_{c}. Specifically, we experimented with: (i) CPC which attempts to predict the future states of the encoder based on the past and optimizes a contrastive loss comparing the actual future from that of random sequences; (ii) HuBERT which was trained with a masked prediction task similar to BERT on masked continuous audio signals as inputs. The targets are obtained through clustering of raw speech features or learned features from earlier iterations; (iii) and VQ-VAE which performs similarly to a Variational Auto Encoder where the encoder’s output is discrete rather than continuous.

Since the representations learned by CPC and HuBERT are continuous, a k-means algorithm is applied over the models’ outputs to generate discrete units, denoted as zc=(z1,…,zL)\bm{z}_{c}=(z_{1},\ldots,z_{L}). Each element ziz_{i} in zc\bm{z}_{c} is a positive integer, zi∈{0,1,..,K}z_{i}\in\{0,1,..,K\} for 1≤i≤L1\leq i\leq L, where KK is the number of discrete units. We did not follow the same approach for VQ-VAE as its representations are already quantized.

F0 Encoder To generate low frequency discrete F0 representation, zF0=(z1,…,zL′)\bm{z}_{F_{0}}=(z_{1},\dots,z_{L^{\prime}}), a separate encoder, EF0E_{F_{0}}, is applied over the F0 extracted from the input signal. Each element in zF0\bm{z}_{F_{0}} is an integer zs∈{0,1,..,K′}z_{s}\in\{0,1,..,K^{\prime}\}, where K′K^{\prime} is the encoder dictionary size. The YAAPT algorithm is used to extract the F0 from the input signal, x\bm{x}, generating p=(p1,…,pT′)\bm{p}=(p_{1},\dots,p_{T^{\prime}}).

To generate zF0\bm{z}_{F0}, we use the indices of the mapped latent vectors rather than the vectors.

2 Decoder

A neural vocoder is employed to decode the speech signal from the discrete representation. This study considers the decoder to be a modified version of the HiFi-GAN neural vocoder.

The HiFi-GAN architecture is comprised of a generator, GG, and a set of discriminators, DD. The generator is built from a set of look-up tables (LUT) that embed the discrete representation and a series of blocks composed of transposed convolution and a residual block with dilated layers. The transposed convolutions upsample the encoded representation to match the input sample rate, while the dilated layers increase the receptive field.

As an input, the generator receives the encoded representation (zc,zF0,zspk)(\bm{z}_{c},\bm{z}_{F_{0}},\bm{z}_{spk}). The discrete content sequence, zc\bm{z}_{c}, and the discrete pitch sequence, zF0\bm{z}_{F_{0}}, are converted to a continuous representation via LUTcLUT_{c} and LUTF0LUT_{F_{0}} accordingly. The sequences are up-sampled and concatenated together. The speaker embedding, zspk\bm{z}_{spk}, is concatenated to each frame in the up-sampled sequence.

The discriminator comprises two networks, a Multi-Period Discriminator (MPD) and a Multi-Scale Discriminator (MSD). The MPD consists of multiple sub-discriminators operating on equally spaced samples from the input signal. The period sub-discriminators differ from each other based on the space between the samples. Similar to , the MPD employs a total of five-period discriminators with a period hops of $.Multi−scalediscriminator(MSD)employsmultiplesub−discriminatorsoperatingatdifferentscalesoftheinputsignal.Specifically,weusethreescales:theoriginalinputscale,. Multi-scale discriminator (MSD) employs multiple sub-discriminators operating at different scales of the input signal. Specifically, we use three scales: the original input scale,\times 2downsampledscale,anddownsampled scale, and\times 4downsampledscale.Overalleachsub−discriminatordownsampled scale. Overall each sub-discriminatorD_{j}$ is tasked with minimizing the following,

where x^=G(LUTc(zc),LUTF0(zF0),zspk)\hat{\bm{x}}=G(LUT_{c}(\bm{z}_{c}),LUT_{F_{0}}(\bm{z}_{F_{0}}),\bm{z}_{spk}), is the resynthesized signal from the encoded representation.

Additionally, two terms are added to the loss function. The first one is a reconstruction term computed between the Mel-spectrogram of the input signal and the generated signal,

where ϕ\phi is a spectral operator computing Mel-spectrogram. The second term is a feature-matching loss which measures the distance between discriminator activations of the real signal and those of the resynthesized signal,

where ψi\psi_{i} is an operator which extracts the activations of the discriminator ii-th layer, MiM_{i} is the number of features in layer ii, and RR is the total number of layers in DjD_{j}.

The final loss with respect to the sub-discriminators composing discriminator DD and generator GG is:

where we set λfm=2\lambda_{fm}=2 and λr=45\lambda_{r}=45.

Results

Our results cover three different settings: (i) speech reconstruction experiments; (ii) speaker conversion and F0 manipulation; (iii) bitrate analysis with subjective tests for speech codec evaluation. We employ two datasets: LJ single speaker dataset and VCTK multi-speaker dataset. All datasets were resampled to a 16kHz sample rate.

Implementation Details We follow the same setup as in . For CPC, we used the model from , which was trained on a “clean” 6k hour sub-sample of the LibriLight dataset . We extract a downsampled representation from an intermediate layer with a 256-dimensional embedding and a hop size of 160 audio samples. For HuBERT we used a Base 12 transformer-layer model trained for two iterations on 960 hours of LibriSpeech corpus . This model downsamples the raw audio ×320\times 320 into a sequence of 768-dimensional vectors. Similarly to , activations were extracted from the sixth layer.

For CPC and HuBERT, the k-means algorithm is trained on LibriSpeech clean-100h dataset to convert continuous frames to discrete codes. We quantize both learned representations with K=100K=100 centroids. Leading to a bitrate of 700bps for CPC and 350bps for HuBERT.

Similarly to CPC models, we trained the VQ-VAE content encoder model on the “clean” 6K hours subset from the LibriLight dataset. We use an encoder operating on the raw signal to extract discrete units, similar to . In addition, “random restarts” were performed when the mean usage of a codebook vector fell below a predetermined threshold. Finally, we used HiFiGAN (architecture and objective) as the decoder instead of a simple convolutional decoder, as it improved the overall audio quality. This model encodes the raw audio into a sequence of discrete tokens from 256 possible tokens with a hop size of 160 raw audio samples. The VQ-VAE discrete code operates at a bitrate of 800bps. We additionally experimented with 100 discrete units for VQ-VAE, however results were the best for 256. This finding is consistent with .

The speaker verification network uses the architecture proposed in . It was trained on the VoxCeleb2 dataset, achieving a 7.4% Equal Error Rate (EER) for speaker verification on the test split of the VoxCeleb1 dataset.

Only a single F0 representation is considered across all evaluated models, trained on the VCTK dataset. The F0 is extracted from the raw audio using a window size of 20ms and a 5ms hop. As a result, the F0 sequence is sampled at 200Hz. The quantization described at Sec. 3, is applied using an F0 codebook of K′=20K^{\prime}=20 tokens and an encoder that downsamples the signal by ×16\times 16. Hence, the discrete F0 representation is sampled at 12.5Hz, leading to a bitrate of 65bps. The final bitrate of the evaluated codecs is the sum of the pitch code bitrate with the content code bitrate.

Evaluation Metrics We consider both subjective and objective evaluation metrics. For subjective tests, we report the Mean Opinion Scores (MOS). In which human evaluators rate the naturalness of audio samples on a scale of 1–5. Each experiment, included 50 randomly selected samples rated by 30 raters. For objective evaluation, we consider: (i) Equal Error Rate (EER) as an automatic speaker verification metric obtained using a pre-trained speaker verification network. We report EER between test utterances and enrolled speakers; (ii) Voicing Decision Error (VDE) , which measures the portion of frames with voicing decision error; (iii) F0 Frame Error (FFE) , measures the percentage of frames that contain a deviation of more than 20% in pitch value or have a voicing decision error; (iv) Word Error Rate (WER) and Phoneme Error Rate (PER), proxy metrics to the intelligibility of the generated audio. We used a pre-trained ASR network on both reconstructed and converted samples to calculate both metrics.

Reconstruction & Conversion We start by reporting the reconstruction performance. Results are summarized in Table 1. When considering the intelligibility of the reconstructed signal HuBERT reaches the lowest PER and WER scores across all models, where both CPC and HuBERT are superior to VQ-VAE. However, when considering F0 reconstruction VQ-VAE outperforms both HuBERT and CPC by a significant margin. This results are somewhat intuitive, bearing in mind VQ-VAE objective is to fully reconstruct the input signal. In terms of subjective evaluation, all models reach similar MOS scores, with one exception of CPC on LJ.

To better evaluate the disentanglement properties of each method with respect to speaker identity and F0, we conducted an additional set of experiments aiming at speaker conversion and F0 manipulation. For voice conversion, we converted each test utterance into five random target speakers. Next, we employed a speaker verification network, which extracts d-vector representation to evaluate speaker-converted utterances’ similarity to real speaker utterances (low error-rate indicates good conversion), providing measurement to the speaker identity’s disentanglement from the evaluated coding method. The error-rate is reported between converted test utterances and enrolled speakers. For the LJ speech single speaker dataset, we converted samples from the VCTK dataset to the single speaker and enrolled all VCTK speakers together with the single speaker. Results are summarized in Table 2 (left). Unlike resynthesis results, on voice conversion CPC and HuBERT outperform VQ-VAE on both LJ and VCTK datasets, indicating VQ-VAE contains more information about the speaker in the encoded units, hence producing more artifacts. Notice, this also affects WER, PER, and the overall subjective quality (MOS).

Next, to evaluate the presence of F0 in the discrete units, we flattened the F0 units before synthesizing the signal and calculated VDE and FFE with respect to the original F0 values. F0 flattening was done by setting the speakers’ mean F0 value across all voiced frames. In this experiment, we expected units that contain F0 information to be better at F0 reconstruction over disentangled units. Results are summarized in Table 2 (right). Notice VQ-VAE can still reconstruct the F0 almost at the same level as when using the original F0 as conditioning (5.2 vs 7.03, and 5.59 vs 7.8), in contrast to CPC and HuBERT.

Speech Codec Our final experiment evaluates the obtained speech units as a low bitrate speech codec. We use a subjective MUSHRA-type listening test to measure the perceived quality of the proposed speech codec with regard to its bitrate constraints. In MUSHRA evaluations, listeners are presented with a labeled uncompressed signal for reference, a set of test samples to rate, a copy of the uncompressed reference, and a low-quality anchor. Listeners are asked to rate each test utterance and the copy of the uncompressed reference with respect to the labeled reference in a scale of 1-100.

The experiment is performed on the VCTK dataset . For evaluation, we used 20 utterances from 5 speakers. The set of speakers in the test data is disjoint with those in the training data. For this experiment, HuBERT models with 50, 100, and 200 units were trained as described in Sec. 4. For comparison, we included other speech codecs in our evaluation: Opus wideband at 9 kbps VBR, Codec2 at 2.4 kbps and LPCNet operating at 1.6 kbps. The LPCNet model was trained from scratch on the VCTK dataset following the experimental setup in . The VQ-VAE model employs the HiFiGAN decoder trained on the LibriLight dataset to match the amount of data reported in . We compressed the anchor sample with Speex at 4 kbps as a low anchor. Fig. 2 depicts the results. HuBERT with 50 units reaches the best MUSHRA score while its bitrate is only 365bps, which is significantly lower than the baseline methods.

Conclusion

We applied self-supervised discrete representations for the task of speech resynthesis. Furthermore, we demonstrated the efficiency of disentangled representations for signal reconstruction, voice conversion, and F0 manipulation. Our evaluations shed light on the properties encoded by each method in the context of speech synthesis. Finally, we adapt the HuBERT speech representation as an ultra lightweight speech codec, providing superior subjective results than the baselines with lower bitrate.

References