A Multi-Discriminator CycleGAN for Unsupervised Non-Parallel Speech Domain Adaptation

Ehsan Hosseini-Asl, Yingbo Zhou, Caiming Xiong, Richard Socher

Introduction

Neural-based acoustic models have shown promising improvements in building automatic speech recognition (ASR) systems . However, it tends to perform poorly when evaluated on out-of-domain data, because of mismatch between the training and testing distribution (Table 1).

Domain mismatch is mainly due to variation in non-linguistic features, such as different speaker identity, unseen environmental noise, large accent variations, etc. Therefore, training a robust ASR system is highly dependent on factorizing linguistic features (text) from non-related variations, or adapting the inter-domain variations of source and target.

Voice conversion (VC) has been widely used to adapt the non-linguistic variations, such as statistical methods , and Neural-based models . However, traditional VC methods require parallel data of source and target that is difficult to obtain in practice. In addition, the requirement of parallel data prevent these methods from using more abundant unsupervised data. Therefore, an unsupervised domain adaptation is desirable for building a robust ASR system.

In this paper, we propose a new generative model based on CycleGAN for unsupervised non-parallel domain adaptation. Since differences in magnitude of frequency is the main variation across domains for spectrogram representations, it is imperative that CycleGAN correctly catch the spectro-temporal variations between different frequency bands across domains during training. This will allow the generator to learn the mapping function which can convert spectrogram from source to target domain. In this paper, we show that the original CycleGAN model is failing to learn the correct mapping function between domains, and the generator collapses into learning an identity mapping function, which results in generating a noisy and unnatural-sounding audio.

To accommodate generative adversarial network for training on non-parallel spectrogram domains, the generator should be back-propagated with multiple gradient signals (from different discriminators), that each represents the variations between source and target domains at a specific frequency band. To achieve this goal, we propose to use multiple and independent discriminators for each domain, similar to generative multi adversarial network (GMAN) . We show that the proposed Multi-Discriminator CycleGAN, without pretraining the discriminators, outperforms CycleGAN with pretrained discriminator, for spectrogram adaptation. Furthermore, we show that the multi discriminator architecture can overcome the checkerboard artifacts problem caused by deconvolution layer in generator , and generates natural clean audio. To evaluate the performance of the proposed model, gender-based domains are selected as domain adaptations.

Generative Adversarial Network (GAN) is a family of non-parametric density estimation models which learn to model the data generating distribution using adversarial training. Conditional GANs (CGAN) was first proposed for supervised (parallel) domain adaptation, where the goal is to convert source distribution to match the target. CGAN has been used in various data domains, especially image domains, both for parallel and non-parallel domain adaptation .

Recently, CGAN is used for speech enhancement on parallel datasets . Speech denoising is achieved by conditioning the generator on noisy speech to learn the de-noised version . Donahue et al. proposed a GAN model on audio (WaveGAN) and spectrogram (SpecGAN), which is actually trained CGAN on parallel domains. Kaneko et al. proposed a cycle-consistent adversarial network (CycleGAN) with gated convolutional neural network (CNN) as the generator part, where the model is trained on Mel-cepstral coefficients (MCEPs) features. Hsu et al. proposed a combination of variational inference network, using variational autoencoder (VAE) , and adversarial network, using Wasserstein GAN (WGAN) . In , the goal is to disentangle the linguistic from nuisance latent variables via VAE using spectra (SP for short), aperiodicity (AP), and pitch contours (F0) features, followed by adversarial training to learn the target distribution from the inferred linguistic latent distribution. A recurrent VAE is also proposed to capture the temporal relationships in the disentangled representation of sequence data, using Mel-scale filter bank (FBank).

Contributions of the proposed generative model are, (1) It is a robust GAN model developed for non-parallel unsupervised domains, compared to parallel-based SpecGAN and WaveGAN , (2) The choice of multiple discriminator is adjustable to the spectro-temporal structure of the intended domains, compared to domain-specific model design of , (3) Proposed GAN model training is robust and invariant to the choice of adversarial objective, i. e. binary cross-entropy or least square (LS-GAN ), while the CycleGAN in is only stable using least square loss, with additional using of identity mapping loss in generator, (4) Source and target domains in are sampled from same speakers, both including male and female, only uttering different sentences, while our approach is more natural as source and targets distribution is strongly diverged due to different speaker, gender, and uttered sentences. (5) Compared to FHVAE , our models improves ASR performance on TIMIT female set by 2.067%2.067\% PER (Table 3), when only trained on male.

Proposed Model

In this section, the proposed model is explained. We first describe the generative model based on adversarial network. Generative Model based on adversarial training (GAN) has been proposed by Goodfellow et al. to model the data generating distribution. Training GAN is based on minimizing Jensen-Shannon divergence between data generating distribution pdata(x)p_{data}(x) and model data distribution pz(z)p_{z}(z). Learning is through minimization of the adversarial loss between generator network G(z)G(z), which learns a mapping function G:Z→XG:Z\rightarrow X, and discriminator network D(x)D(x). The generator is learning to model the data distribution pdata(x)p_{data}(x) by generating indistinguishable samples x^=G(z)\hat{x}=G(z) from xx, using a source noise signal zz to minimize (1), whereas discriminator is learning to discriminate between real data xx and generated x^\hat{x} by maximizing the adversarial loss,

For domain adaptation between parallel domains XX and YY, Conditional GAN (CGAN) is proposed, using a generator that directly learns the mapping function G:X→YG:X\rightarrow Y, by minimizing parallel conditional adversarial loss LP−CGAN\mathcal{L}_{P-CGAN},

where DD is discriminating between pair of real parallel data (x,y)(x,y) and generated pair (x,G(z,x))(x,G(z,x)). To apply CGAN for adaptation between non-parallel domains XX and YY, a conditional GAN using cycle consistent adversarial loss (CycleGAN) has been proposed . In CycleGAN , there are two conditional generators, i. e., GX:X→YG_{X}:X\rightarrow Y and GY:Y→XG_{Y}:Y\rightarrow X, each trained in adversarial setting with DYD_{Y} and DXD_{X}, respectively. In other words, there are two pairs of Non-parallel conditional adversarial loss LNP−CGAN(GX,DY)\mathcal{L}_{NP-CGAN}(G_{X},D_{Y}) and LNP−CGAN(GY,DX)\mathcal{L}_{NP-CGAN}(G_{Y},D_{X}), where,

Therefore, CycleGAN learns unsupervised mapping functions between XX and YY domains by combining (LABEL:eq:cgan-unparallel) and (4), to maximize the adversarial loss LCycleGAN\mathcal{L}_{CycleGAN}, where,

2 Multi-Discriminator CycleGAN (MD-CycleGAN)

In this section, we propose a multiple discriminator generative model based on cycle consistency loss (5). The model is based on generative multi adversarial network (GMAN) . In this paper, XX and YY represents spectrogram feature datasets of different speech domains. Spectrogram feature represents the frequency variation of audio data through time dimension. In order to allow CycleGAN to learn the mapping function of spectrogram between different speech domains, the generators {GX,GY}\left\{G_{X},G_{Y}\right\} should be able to learn the variations in each frequency band for each aligned time window, across domains.

In order to learn the frequency-dependent mapping functions {GX,GY}\left\{G_{X},G_{Y}\right\} that catch the variation per each frequency bands, we define multiple frequency-dependent discriminators {DXfj∈n,DYfi∈m}\left\{D_{X}^{f_{j\in n}},D_{Y}^{f_{i\in m}}\right\}, where fj∈nf_{j\in n} represents the ithi^{th} frequency band of domain XX with nn frequency bands, and fi∈mf_{i\in m} represents jj-th frequency band od domain YY, respectively. The frequency band definition in each domain can share a portion of frequency spectrum, or be exclusive, based on the domain spectrogram distribution. We are also using the non-saturating version of GAN, NS-GAN, where the generator GG is learned through maximizing the probability of predicting generated samples x^\hat{x} as drawn from data generating distribution pdata(x)p_{data}(x). Accordingly, the adversarial loss for each pair of generator and discriminator {(GX,DYfi∈m),(GY,DXfj∈n)}\left\{\left(G_{X},D^{f_{i\in m}}_{Y}\right),\left(G_{Y},D^{f_{j\in n}}_{X}\right)\right\} in (LABEL:eq:cgan-unparallel) and (5) is

The Multi-Discriminator CycleGAN (MD-CycleGAN) is training by maximizing LMD−CycleGAN\mathcal{L}_{MD-CycleGAN}, where,

A natural extension to the proposed MD-CycleGAN is to use multiple frequency-dependent generators jointly with discriminators as well. This can follow in two configurations. In one-one setting, each generator is trained on a specific frequency band with the corresponding discriminator, i. e., set of {(GXfi,DYfi):i∈m}\left\{\left(G_{X}^{f_{i}},D_{Y}^{f_{i}}\right):i\in m\right\}. Additionally, in one-many setting, each frequency-dependent generator is trained with all frequency-dependent discriminators, i. e., set of {(GXfj,DYfi∈m):j∈n}\left\{\left(G_{X}^{f_{j}},D_{Y}^{f_{i\in m}}\right):j\in n\right\} trained in adversarial setting.

Experiment

We used TIMIT and Wall Street Journal (WSJ) corporas to evaluate the performance of proposed model on domain adaptation. TIMIT dataset contains broadband 1616kHz recordings of phonetically-balanced read speech of 63006300 utterances (5.45.4 hours). Male/Female ratio of speakers across train/validation/test sets are approximately 7070% to 3030%. WSJ contains ≈80\approx 80 hours of standard si284/dev93/eval92 for train/validation/test sets, with equally distributed genders.

The spectrogram representation of audio is used for training the CycleGAN and ASR models, which is computed with a 2020ms window and 1010ms step size. Each spectrogram is normalized to have zero mean and unit variance. To implement MD-CycleGAN, three non-overlapping frequency bands are defined, i. e. m=n=3m=n=3 with $bandwidth,formaleandfemaledomains.Wedenotethesizeoftheconvolutionlayerbythetuplebandwidth, for male and female domains. We denote the size of the convolution layer by the tuple(\text{C, F, T, SF, ST}),whereC,F,T,SF,andSTdenotenumberofchannels,filtersizeinfrequencydimension,filtersizeintimedimension,strideinfrequencydimensionandstrideintimedimensionrespectively.CycleGANmodelarchitectureisbasedonwithsomemodifications.ThegeneratorisbasedonU−netarchitecturewith, where C, F, T, SF, and ST denote number of channels, filter size in frequency dimension, filter size in time dimension, stride in frequency dimension and stride in time dimension respectively. CycleGAN model architecture is based on with some modifications. The generator is based on U-net architecture with4convolutionallayersofsizes(8,3,3,1,1),(16,3,3,1,1),(32,3,3,2,2),(64,3,3,2,2)withcorrespondingdeconvolutionlayers.Wenoticedthatthediscriminatorinoutputsavectorwithdimensionequaltothechannelsizeoffinalconvolutionlayer,insteadofoutputtingascalar.Itwasobservedthatthiscausesinstabilityinabalancedtrainingbetweengeneratoranddiscriminator.Wemodifiedthisbyaddingafullyconnectedlayerasfinallayer,tomatchthediscriminatorin.Discriminatorhasconvolutional layers of sizes (8,3,3,1,1), (16,3,3,1,1), (32,3,3,2,2), (64,3,3,2,2) with corresponding deconvolution layers. We noticed that the discriminator in outputs a vector with dimension equal to the channel size of final convolution layer, instead of outputting a scalar . It was observed that this causes instability in a balanced training between generator and discriminator. We modified this by adding a fully connected layer as final layer, to match the discriminator in . Discriminator has4convolutionlayersofsizes(8,4,4,2,2),(16,4,4,2,2),(32,4,4,2,2),(64,4,4,2,2),asdefaultkernelandstridesizesin.WeusedGriffin−limalgorithmforaudioreconstruction,toassessitsquality.ASRmodelisbasedon,trainedwithmaximumlikelihood,andnopolicygradient.Themodelhasoneconvolutionallayerofsize(32,41,11,2,2),andfiveresidualconvolutionblocksofsize(32,7,3,1,1),(32,5,3,1,1),(32,3,3,1,1),(64,3,3,2,1),(64,3,3,1,1)respectively.Followingtheconvolutionallayersareconvolution layers of sizes (8,4,4,2,2), (16,4,4,2,2), (32,4,4,2,2), (64,4,4,2,2), as default kernel and stride sizes in . We used Griffin-lim algorithm for audio reconstruction, to assess its quality. ASR model is based on , trained with maximum likelihood, and no policy gradient. The model has one convolutional layer of size (32,41,11,2,2), and five residual convolution blocks of size (32,7,3,1,1), (32,5,3,1,1), (32,3,3,1,1), (64,3,3,2,1), (64,3,3,1,1) respectively. Following the convolutional layers are4layersofbidirectionalGRURNNswithlayers of bidirectional GRU RNNs with1024hiddenunitsperdirectionperlayer,onefully−connectedhiddenlayerofsizehidden units per direction per layer, one fully-connected hidden layer of size1024$ and final output layer.

In this section, ASR model is employed to evaluate the performance of proposed model, where domains are different genders. First, gender generators {GM→F,GF→M}\left\{G_{M\rightarrow F},G_{F\rightarrow M}\right\} GM→F:Male→FemaleG_{M\rightarrow F}:Male\rightarrow Female, GF→M:Female→MaleG_{F\rightarrow M}:Female\rightarrow Male are trained on gender-separated train set. These generators are then evaluated for train→\rightarrowtest and test→\rightarrowtrain adaptation using ASR model. In former, ASR model is retrained on the adapted train set, while in latter, a more applicable case, ASR model is fixed and evaluated on the new adapted test sets.

Results on adapting TIMIT train set are shown in Table 2 and 3. As ablation study to CycleGAN-VC , performance is significantly improved with three discriminator compared to single one. Compared to FHVAE , phoneme error rate is improved by 2.067%2.067\% in Table 3. To evaluate the generalization of the generators, we used them on WSJ dataset without retraining. As shown in Table 4, ASR performance is significantly improved by reducing the gap to the corresponding male and female baselines. For a fair comparison, ASR performance trained on WSJ train set is 5.55%5.55\% WER. It is worth mentioning that relatively lower performance on TIMIT is due to smaller size of dataset.

1.2 Test→→\rightarrowTrain Adaptation

Test set adaptation of TIMIT and WSJ are shown in Table 5 and 6. It is clear that using the proposed model, ASR performance is significantly improved by adapting test→\rightarrowtrain , compared to original CycleGAN. Qualitative assessment of ASR predictions are shown in Tables 1 and Appendix A.

2 Qualitative Evaluation

In this section, the quality of generated spectrogram for male↔\leftrightarrowfemale adaptation is assessed. The characteristic difference between male and female spectrograms is the variation rate of frequency for a fixed time window. As shown in Figure 1, top row depicts the original spectrograms, where male is characterized by smooth frequency variation, opposed to peaky and high-rate variations of female. Well-trained generators should catch these inter-domain variations. As ablation study, we are also showing the generated spectrogram by CycleGAN (One-D CycleGAN), in middle row, comparing with Three-D CycleGAN in bottom row. One-D CycleGAN learns to convert the spectrogram only using a pretrained discriminator. It is noticeable that the converted spectrogram in One-D CycleGAN fails to match the target domain characteristics, at some frequency bands, and simply copied the source spectrogram. However, with no pretraining of Three-D CycleGAN, it learns a better mapping function, by either suitably smoothing the spectrogram (female→\rightarrowmale), or generating peaky variations (male→\rightarrowfemale). The checkerboard artifacts is a common problem in deconvolution-based generators. This problem is visible in One-D CycleGAN, with discontinuous artifacts through time and frequency dimensions, which results in a noisy and unnatural-sounding audio. This problem is mitigated in Three-D CycleGAN, by learning the target domain characteristics using multiple independent discriminators.

Conclusion and Future Directions

In this paper, a new cyclic consistent generative adversarial network based on multiple discriminators is proposed (MD-CycleGAN) for unsupervised non-parallel speech domain adaptation. Based on the frequency variation of spectrogram between domains, the multiple discriminators enabled MD-CycleGAN to learn an appropriate mapping functions that catch the frequency variations between domains. The performance of MD-CycleGAN is measured by ASR prediction, when train and test set are sampled from different genders. It was shown that MD-CycleGAN can improve the ASR performance on unseen domains. As future extension, this model will be evaluated on datasets adaptation, e. g. TIMIT↔\leftrightarrowWSJ, and accent, e. g. American↔\leftrightarrowIndian adaptations.

References

Appendix A ASR Prediction Assessment

In this section, qualitative assessment of ASR predictions are shown, when test set is sampled from the opposite gender, and when MD-CycleGAN model is used to adapt the test set distribution to the train set. Tables 7, 8 show the selected ASR predictions when adapting female→\rightarrowmale, whereas Tables 9, 10 show adaptation for male→\rightarrowfemale. Note that MD-CycleGAN is trained on TIMIT, and used for adaptation on WSJ dataset.