Towards Generalized Speech Enhancement with Generative Adversarial Networks
Santiago Pascual, Joan Serrà, Antonio Bonafonte
Introduction
Speech enhancement comprises the improvement of intelligibility and quality of speech contaminated by noise, typically in an additive form . It is thus broadly applicable to scenarios that need to palliate masking artifacts over the speech, like communication systems, hearing aids, or cochlear implants, where enhancing the signal prior to amplification significantly reduces discomfort and increases intelligibility . Deep learning has been broadly applied to this field, either in the form of regression of waveform samples or clean spectral components, as well as prediction of masking elements over spectral bins . A recent trend in speech enhancement is also the use of generative adversarial networks (GANs) , with its first application being the speech enhancement GAN (SEGAN) . Other GAN variations have been used to avoid using aligned clean/noisy corpora or using other adversarial losses and/or domains . SEGAN has also been applied to speech regeneration as a whispered-to-voiced alaryngeal speech conversion , thus extending the previous denoising approach.
Athough GANs were designed as an unsupervised learning strategy, they are proven to be more stable and effective when coupled with labels that are either injected as input conditionings or that help classify some traits about the data being generated . Also the successful speech synthesis model parallel WaveNet proved that a multi-task aggregation on top of the speech generative model is beneficial to improve speaker identity, prosodic, and content traits.
In this work, we first extend SEGAN towards a more generalized speech enhancement such that we recover cleaner speech out of severely distorted utterances. This generalization towards different speech distortions is interesting and directly applicable to modern communication technologies, where connections are interrupted (i.e. losing voice packets), and voice processing pipelines can distort our signals through amplifiers, compression of transmission encoders, etc. In this work we emulate these effects by introducing four applied distortions in this work, with different severity levels per distortion. Then, we propose the use of SEGAN to palliate them with three different schemes. First, a plain adversarially trained SEGAN version serves us as a baseline for our recovery experiments. Then, we propose a new regression component on top of the discriminator that serves as an additional self-supervised task, boosting the generated recovered speech quality. This new task is specially effective when we use the second modification, a two-stage adversarial training schedule as a warm up and fine-tunning sequence. Objective and subjective results show the effectiveness of both proposed mechanisms when they are jointly applied.
Speech enhancement GAN Review
Generalized Speech Enhancement
In this work, we depart from the classical denoising task and consider a more general class of enhancement problems. More specifically, we consider the problem of reconstructing speech that has been degraded by a number of (different) signal manipulations, each of them being potentially highly perceptually harmful. We propose to mix up such manipulations together with several speaker identities at training time, with the objective of being able to train generative algorithms to recover multiple (and diverse) speech components simultaneously. As a first approximation towards this generalized speech enhancement task, we consider the following signal manipulations:
Whispered speech: This is the distortion is similar to the one introduced in the WSEGAN work . However, it differs in the fact that instead of using magnetic sensors, we synthesize whispering speech by encoding the clean speech with a vocoder, removing the log-F0 (i.e. making all frames unvoiced), and recovering the signal into a version that whispers, hence artificially removing voicing. The vocoder used is Ahocoder with the default parameters .
Bandwith reduction: We downsample the audio signals with different severity factors, ranging from to , thus reducing the bandwidth. Then the generalized enhancement model must reconstruct entire frequency bands, according to the clean signals seen in training.
Chunk removal: For the parts of the waveform that contain speech, a random number of chunks are subtracted by inserting silences. Thus, we cancel the signal at random temporal sections that contain speech. The length of a silence is sampled from one of two distributions , and (numerical values in seconds).
Clipping: The waveform is clipped globally by different severity factors relative to the maximum absolute peak of the whole utterance (e.g., 30%). The regenerated signal thus has to re-condition the signal inside a proper, non-distorted range of amplitudes.
Acoustic Mapping Discriminator
To better deal with the added difficulty that the previous signal manipulations can introduce to the speech signal, we propose two modifications to the existing SEGAN pipeline: the addition of two acoustic losses and the introduction of a two-stage training schedule for applying such losses.
In general, can be understood as a learnable loss function, where the realistic features we want to generate are implicit in the back-propagation, and gets better because of ’s gradient flows . This builds a need for to be a competitive feature extractor, so that the better the features extracted by , the better can capture reality. The use of auxiliary classifying labels in is generally helpful (Sec. 1). Additionally, multi-task setups can boost generative modeling performance, where multiple factors of the signal are predicted at the output of the generative model to enforce modeling better perceptual qualities of the signal . This is applicable to the SEGAN setup where the discriminator only learns features that lead to a real or fake decision, but is not concerned about specific factors of the signal that are important to remark the reconstructed speaker identity, prosody, or contents. Note that we can also have some additional regularization loss in the output of too, aggregated to the adversarial loss coming from . For instance the first version of SEGAN used an regularization term that made output a zero-centered signal , and the posterior WSEGAN implementation substituted it by a less restrictive power loss , enforcing proper energy allocation in the frequency bands.
2 Adversarial Pre-Training
Whereas SEGAN with only an adversarial loss is stable and learns in a steady equilibrium (eq. 1), the addition of the acoustic losses induces a particular unbalance effect during training. The addition of these terms makes learn quicker and converge faster in both losses, conversely making converge slower. Importantly, both and should maintain an equilibrium learning from each other, so making one of them quickly better discourages the other to perform properly . We hypothesize that scheduling the learning of the discriminator as a two-stage process makes the addition of these acoustic losses more effective for it first gets high-level representations classifying, and then focuses on specific speech properties by doing regressions. Hence we first do an adversarial warm up in with Eq. 1 formulation, and then add the acoustic losses.
Experimental Setup
We employ the VCTK Corpus for our experiments . We select 80 speakers for training and 14 for test. We trim large silence regions to 100 ms with the help of a voice activity detector. After this, we get roughly 20 h of training speech and 3 h of test speech. We incorporate the aforementioned distortions (Sec. 3) with an online process jointly working with training. When we construct a training minibatch, we (1) load the clean utterance, (2) get a random 1 s chunk (16384 samples) inside the utterance, and (3) apply a series of transforms that get activated independently under a probability . This way one or two distortions coincide often, and zero (thus purely autoencoder) or more than two coincide less. Table 1 shows this activation distribution with up to the 4 mentioned combinations. In addition to being active or not, each transform has a certain severity level or factor, as mentioned before (sec. 3). The only transform with just one level of severity is the whispering transform. Hence every time a transformation is activated in the pipeline, one of the possible factors is randomly selected. These possibilities are: clip factors of 30%, 40% and 50%, signal resampling factors of 2, 4, and 8 and, in the case of chunk removal, we allow the system to zero out up to chunks from speech regions. Then, inside each region, the length of the chunk is sampled from one of the two mentioned Gaussian distributions (sec. 3).
For the test split, we have two different setups, one for an objective evaluation and another one for a subjective one. For the objective test set, we proceed as with the training set. For the subjective test, we generate a special split with two subsets in order to account for two different effects present in the system outcomes. The first subset is designed to focus on generated speaker identity (SpkID pool). To do so, we just use the most severe version of bandwidth reduction (8) and the whispering manipulations, so the ones that affect the most this feature. Then, for each test speaker we degrade four utterances, two with bandwidth reduction and two with whispering. The second subset focuses on the generated speech naturalness (Nat pool). To evaluate this aspect, we pick 4 test speakers (two male and two female) and apply clipping, resampling, and whispering to 6 utterances in total (two utterances per distortion). Both clipping and resampling have severity factors 30% and 8, respectively.
2 Experiments
We experiment with three models applied over the aforementioned data. Firstly, we use SEGAN in a plain adversarial setup with the LSGAN loss as in Eq. 1, where the only signal comes from real or fake decisions. This is the baseline, which is trained for 400 epochs with two-timescale update rule (TTUR) learning rates and . Secondly, we apply the acoustic losses presented in Sec. 4.1. This system trains also for 400 epochs, but the learning rates are balanced equal , because the discriminator already has an advantage with the extra signals and TTUR involved a noisier result on . We name this approach SEGAN-Aco. Finally, we include our improved proposal, for which we pre-train the system over 100 epochs the same way we do with the baseline, and then we activate the acoustic losses for the remaining 300 epochs, also lowering the learning rates to . We name this approach SEGAN-PTAco.
3 Evaluation
To assess each model result we perform two evaluations. First, we run a number of objective distortion metrics, often applied in speech synthesis problems: Mel cepstral distortion (MCD; in dB) , F0 root mean squared error (RMSE; in Hz), and the voiced/unvoiced frame prediction error (UV) . These errors give us a first clue on how close is each system to the clean original signal in terms of content, identity, and intonation.
Given the importance of perceptual scores, we also conduct a subjective evaluation with 26 listeners. This has two stages, with two tasks to be evaluated by the listeners: speaker identification and naturalness rating. For speaker identification, listeners are asked to determine how close a reconstructed signal is towards the original speaker identity. They are presented with 4 randomly-selected utterances out of the SpkID pool (Sec. 5.1). For each utterance, the clean reference is shown, as well as the four systems to be rated: (1) the degraded input signal to , (2) the SEGAN baseline, (3) SEGAN-Aco, and (4) SEGAN-PTAco. The rating is as simple as ordering them by preference, such that position 1 is for closest system towards the reference and position 4 is the furthest one. For naturalness rating, 6 utterances taken randomly out of the Nat pool (Sec. 5.1) are shown to each listener. For each utterance, the four different aforementioned systems are shown and asked to be ranked from most natural (1) to least natural (4). There is no reference shown for this case as there is no similarity trait like a speaker identity, it is only comparing the perceptual quality of the synthesized utterances. In all subjective evaluations, systems are shown in random order at every utterance.
4 SEGAN Setup
We use the same kernel widths, strides, and feature map configurations as in the WSEGAN work . These are kernels of width 31 in all convolutional layers for both and . The feature maps are incremental in the encoder and decremental in the decoder, having in and in . The discriminator is then the one diverging from its predecessors in earlier works . Firstly, it has a multi-layer perceptron (MLP) with 16384 inputs (102416 feature map, unrolled), 256 hidden PReLU units, and the single output for real and fake predictions. Secondly, we find the acoustic branch for SEGAN-Aco and SEGAN-PTAco models at the fourth convolutional layer, which outputs 51264 feature maps for an input of approximately 1 s at 16 kHz sampling rate. These corresponds to 64 frames, each one injected into the acoustic prediction MLP of 128 hidden PReLU units and 277 linear outputs. These outputs predict 257 log-power spectral bins, 16 MFCCs, 1 log-F0 value of the frame, 1 voiced/unvoiced frame flag, 1 frame energy coefficient, and 1 frame zero crossing rate. Mini-batches of 150 samples are used for all models. Additionally, implements spectral normalization to avoid sudden exploding gradients leading to training collapse and phase shift of to reduce high-frequency artifacts in the output of . All the experiments were developed with the public SEGAN framework at https://github.com/santi-pdp/segan_pytorch.
Results
Objective results are shown in Table 3. We can see that the distorted signals themselves have the noisiest behaviors, with large variances in all error metrics. This is expectable, as we have a wide range of distortion conditions. We can see how, in all metrics, SEGAN-PTAco is the best system and the least noisy one. A lower MCD can be assumed to correlate with better content preservation, identity reconstruction, and naturalness generation . The lower values obtained for F0-RMSE and UV error also typically denote better intonation schemes matching the input identity with prosodic contents.
Subjective results are shown in Table 3. In general, they also highlight the importance of both the acoustic losses and the two-stage adversarial training schedule. We can see that the baseline and SEGAN-Aco are clustered together with the distorted signals for the Speaker ID test. This implies that, for both bandwidth extension and dewhispering problems, neither systems reconstruct a proper speaker identity from the input signal. A qualitative listening in this case tells us that the baseline imposes an arbitrary identity to the reconstruction, and SEGAN-Aco sounds robotic and muffled. In contrast, SEGAN-PTAco is consistently ranked above the distorted signal, and we can clearly hear a competitive identity consistency with the conditioning. In the naturalness task, the baseline and SEGAN-Aco systems do better than the distorted signals. This was expected, as they involve an enhancement process that removes very noticeable artifacts and recover speech nuances. Nevertheless, SEGAN-PTAco is still the best system by a large margin. Some audio samples are available at http://veu.talp.cat/gsegan/.
Conclusion
In this work, we propose the application of speech enhancement GANs in a more generalized signal recovery framework. We first introduce four aggressive distortions applied in our experiments emulating aggressive distortions to the speech signals. Next, we introduce a new acoustic regression component in the SEGAN and, in addition, we propose a two-stage adversarial training method to make the new acoustic regression component reach an appropriate performance. Objective and subjective results are provided, showing how both the regression losses and and the two-stage schedule outperform the regression losses alone and the purely adversarial approach.
Acknowledgements
This research was partially supported by the project TEC2015-69266-P (MINECO/FEDER, UE). We deeply thank the participants of the subjective evaluation.