Improving GANs for Speech Enhancement

Huy Phan, Ian V. McLoughlin, Lam Pham, Oliver Y. Chén, Philipp Koch, Maarten De Vos, Alfred Mertins

I Introduction

The goal of speech enhancement is to improve the quality and intelligibility of speech which are degraded by background noise . Speech enhancement can serve as a front-end to improve performance of an automatic speech recognition system . It also plays an important role in applications like communication systems, hearing aids, and cochlear implants in which contaminated speech needs to be enhanced prior to signal amplification to reduce discomfort . Significant progress on this research topic has been made with the involvement of deep learning paradigms. Deep neural networks (DNNs) , convolutional neural networks (CNNs) , and recurrent neural networks (RNNs) have been exploited either to produce the enhanced signal directly via a regression form or to estimate the contaminating noise, which is subtracted from the noisy signal to obtain the enhanced signal . Significant improvements on speech enhancement performance have been reported by these deep-learning based methods over more conventional ones, such as Wiener filtering , spectral subtraction or minimum mean square error (MMSE) estimation .

There exists a class of generative methods relying on GANs , which have been demonstrated to be efficient for speech enhancement . When GANs are used for this task, the enhancement mapping is accomplished by the generator GG whereas the discriminator DD, by discriminating between real and fake signals, transmits information to GG so that GG can learn to produce output that resembles the realistic distribution of the clean signals. Using GANs, speech enhancement has been done using either magnitude spectrum input or raw waveform input .

Existing speech enhancement GAN (SEGAN) systems share a common feature – the enhancement mapping is accomplished via a single stage by a single generator GG , which may not be optimal. Here, we aim to divide the enhancement process into multiple stages and accomplish it via multiple enhancement mappings, one at each stage. Each of the mappings is realized by a generator, and the generators are chained to enhance a noisy input signal gradually, step by step, to yield an enhanced signal. By doing so, a generator is tasked to refine or correct the output produced by its predecessor. We hypothesize that it would be better to carry out multi-stage enhancement mapping rather than a single-stage one as in prior works . We then propose two new SEGAN frameworks, namely iterated SEGAN (ISEGAN) and deep SEGAN (DSEGAN) as illustrated in Fig. 1, to study two scenarios: (1) using a common mapping for all the enhancement stages and (2) using independent mappings at different enchancement stages. In the former the generators’ parameters are tied and parameter sharing constrains ISEGAN’s generators to learn a common mapping (i.e. the generators apply the same mapping iteratively). The latter’s generators have independent parameters, allowing them to learn different enhacement mappings flexibly. Note that, due to parameter sharing, ISEGAN’s footprint is expected to be smaller than that of DSEGAN.

We will demonstrate that the proposed method obtains more favorable results than the SEGAN baseline on both objective and subjective evaluation metrics and that learning independent mappings with DSEGAN leads to better performance than learning a common one with ISEGAN.

II SEGAN

To improve the stability, SEGAN further employs least-squares GAN (LSGAN) to replace the discriminator DD’s cross-entropy loss by the least-square loss. The least-squares objective functions of DD and GG are explicitly written as

III Iterated SEGAN and Deep SEGAN

Quan et al. showed that using an additional generator chained to the generator of a GAN leads to better image-reconstruction performance. In light of this, instead of using the single-stage enhancement mapping with one generator as in SEGAN, we propose to learn multiple mappings with a chain of NN generators G ⁣= ⁣G1 ⁣ ⁣→ ⁣ ⁣G2 ⁣ ⁣→ ⁣ ⁣… ⁣ ⁣→ ⁣ ⁣GN\mathfrak{G}\!=\!G_{1}\!\!\rightarrow\!\!G_{2}\!\!\rightarrow\!\!\ldots\!\!\rightarrow\!\!G_{N} with N ⁣> ⁣1N\!>\!1 to perform multi-stage enhancement. We study both the cases when a common mapping is learned and shared by all the stages (i.e. ISEGAN) and when independent mappings are learned at different stages (i.e. DSEGAN). In ISEGAN, the generators share their parameters (i.e. they are realized by a common generator GG) and can be viewed as an iterated generator with the number of iterations of NN. In constrast, DSEGAN’s generators are independent and can be viewed as a deep generator with the depth of NN. ISEGAN and DSEGAN with N ⁣= ⁣2N\!=\!2 are illustrated alongside SEGAN in Fig. 1. Both ISEGAN and DSEGAN reduce to SEGAN when N ⁣= ⁣1N\!=\!1.

At the enhancement stage nn, 1≤n≤N1\leq n\leq N, the generator GnG_{n} receives the output x^n−1\mathbf{\hat{x}}_{n-1} of its predecessor Gn−1G_{n-1} together with the latent representation zn\mathbf{z}_{n} and is expected to produce a better enhanced signal x^n\mathbf{\hat{x}}_{n}:

IV Network architecture

IV-B Discriminator D

The discriminator DD has similar architecture to the encoder part of the generators described in Section IV-A, except that it has two-channel input and uses virtual batch-norm before LeakyReLU activation with α=0.3\alpha=0.3. In addition, DD is topped up with a one-dimensional convolutional layer with one filter of width one (i.e. 1×11\times 1 convolution) to reduce the last convolutional output size from 8×10248\times 1024 to 88 features before classification takes place with a softmax layer.

V Experiments

To assess the performance of the proposed ISEGAN and DSEGAN and demonstrate their advantages over SEGAN, we conducted experiments on the database in which was used to evaluate SEGAN in . The database is originated from the Voice Bank corpus and consists of data from 30 speakers. Following the database’s original split, data from 28 speakers was used for training and data from two remaining speakers was used for testing.

A total of 40 noisy conditions was made in the training data by combining ten types of noises (two artificial and eight stemmed from the Demand database ) with four signal-to-noise ratios (SNRs) each: 15, 10, 5, and 0 dB. For the test data, 20 noisy conditions were created, combining five types of noise from the Demand database with four SNRs each: 17.5, 12.5, 7.5, and 2.5 dB. There are about 10 and 20 utterances for each noisy condition per speaker in the training and test set, respectively. All utterances were downsampled to 16 kHz.

V-B Baseline system

SEGAN was used as a baseline for comparison. We repeated training SEGAN to ensure a similar experimental setting across systems. In addition, to shed some light on how generative models like ISEGAN and DSEGAN perform on the speech enhancement task in relation to discriminative models, we also compared the proposed method to two discriminative deep learning methods: (1) the popular DNN proposed in and (2) the two-stage network (TSN) recently proposed in . The DNN baseline was implemented based on , but with three main modifications: (a) wideband operation (16 kHz, leading to doubling of the feature dimension), (b) smaller frame size and shift (25 ms and 10 ms, respectively), and (c) use of the Adam optimizer [Kingma2015] and simplified training (i.e. without unsupervised pre-training). In addition, early stopping was carried out during training via a leave-out validation set (10% of the training data). While these modifications may lead to a better baseline, they also allow a fair comparison with the SEGAN-based systems. The TSN baseline was configured based on , except for the use of wideband speech. For both the baselines, the features (log-power spectra) were normalized at utterance level to zero mean and unit standard deviation. De-normalization was then performed before waveform reconstruction. The utterance-based mean and standard deviation computed from the input noisy features were used for both normalization and de-normalization.

V-C Network parameters

The implementation was based on Tensorflow framework . The networks were trained for 100 epochs with RMSprop optimizer and a learning rate of 0.00020.0002. The SEGAN baseline was trained with a minibatch size of 100 while it was reduced to 50 to train ISEGAN and DSEGAN to cope with their larger memory footprints. We experimented with different values for N={2,3,4}N=\{2,3,4\} to investigate the influence of the number of iterations of ISEGAN and the depth of DSEGAN.

As in , during training, raw speech segments of length 16384 samples were extracted from the training utterances with 50% overlap. A high-frequency preemphasis filter of coefficient 0.950.95 was applied to each signal segment before presenting to the networks. During testing, raw speech segments were extracted from a test utterance without overlap. They were processed by a trained network, deemphasized, and eventually concatenated to produce the enhanced utterance.

V-D Objective evaluation

We quantified the quality of the enhanced signals based on five objective signal-quality metrics, including PESQ, CSIG, CBAK, COVL, and SSNR, as suggested in and the speech-intelligibility measure STOI . The tool used for computing the first five metrics is based on the implementation in . This is also the one used in . The metrics were computed for each system by averaging over all 824 files of the test set. Since we found that the performance may vary with different network checkpoints, the mean and standard deviation of each metric over the 5 latest network checkpoints are reported.

The objective evaluation results are shown in Table I. As expected, SEGAN enhances the noisy signals to result in speech signals with better quality and intelligibility, evidenced by its better results across the objective metrics compared to those measured from the noisy signals. In comparision to SEGAN, on the one hand, ISEGAN performs comparably in terms of speech-quality metrics, slightly surpassing the baseline in PESQ, CBAK, and SSNR (i.e. with N ⁣= ⁣2N\!=\!2 and N ⁣= ⁣4N\!=\!4) but marginally underperforming in CSIG and COVL. On the other hand, DSEGAN obtains the best results, consistently outperforming both SEGAN and ISEGAN across all the speech quality metrics. For example, with N=2N=2, DSEGAN leads to relative improvements of 7.3%7.3\%, 4.7%4.7\%, 6.9%6.9\%, 6.2%6.2\%, and 18.2%18.2\% over the baseline on PESQ, CSIG, CBAK, COVL, and SSNR, respectively. In terms of speech intelligibility, ISEGAN and DSEGAN obtain similar STOI results and both of them outperform SEGAN on this metric. The results in the table also suggest marginal impact of ISEGAN’s number of iterations and DSEGAN’s depth larger than N ⁣= ⁣2N\!=\!2 since no significant performance improvements are seen.

Interestingly, quite opposite results are seen between the discriminative baselines (DNN and TSN) and the generative models (ISEGAN and DSEGAN). In terms of speech quality, the discriminative models outperform the generative counterparts on PESQ, CSIG, COVL but underperform on CBAK and especially on SSNR. In addition, both DNN and TSN perform poorly on speech intelligibility. Degradation on STOI metric is even seen by DNN while TSN brings up modest improvement. On the contrary, both ISEGAN and DSEGAN obtain far better results on speech intelligibility. These results suggest that the discriminative models may alter the noisy input more aggressively than the generative ones and, as a result, introduce more artifacts to the enhanced signals.

To shed light on how the perfomance evolves during the enhancement process of DSEGAN and ISEGAN, we extracted and evaluated the output signals after each of their generators. The results are shown in Fig. 4. One can observe diverging patterns between DSEGAN and ISEGAN. With DSEGAN, overall, the enhancement performance is gradually improved when the signal is passed though the generators one after another. On the contrary, ISEGAN exposes a downward trend on most of the metrics with further enhancement iterations, except for SSNR. The rationale behind the SSNR improvement is that this measure best reflects the least-squares loss that was used to train the network. However, the improved SSNR does not properly reflect other metrics such as human perception and intelligibility represented by PESQ and STOI, which rely on frame-wise weighted frequency domain. This result tends to agree with the finding in psychoacoustics . We speculate that parameter independency/sharing is the key. With independent parameters, each DSEGAN’s generators is tasked for enhacement with one condition of noise and has full freedom to adapt to it. On the other hand, parameter sharing forces the common generator of ISEGAN to deal with all conditions of noise, which is hard to achieve. Of note, instead of using all generators as a whole (i.e. the results in Table I), output of any generators can be used for inferencing. For ISEGAN, using the outputs of earlier generators for this purpose is apparently reasonable as suggested in Fig. 4.

V-E Subjective evaluation

To validate the objective evaluation, we conducted a small-scale subjective evaluation of four conditions: noisy signals, SEGAN, ISEGAN and DSEGAN signals (with N ⁣= ⁣2N\!=\!2). Twenty volunteers aged 18–52 (F=6, M=14), with self-reported normal hearing, were asked to provide forced binary quality assessments between pairs of 20 randomly presented sentences, balanced in terms of speakers and noise types, i.e. each comparison varied only in the type of system. Following a familiarization session, tests were run individually using MATLAB, with listeners wearing Philips SHM1900 headphones in a low-noise environment. For each pair of utterances, the selected higher quality one was rewarded 1.01.0 while the lower quality received no reward. A preference score was obtained for each system by dividing its accumulated reward by the count of its occurrences in the test. Due to the small sample size, we assessed statistical significance of results using tt-test. Results confirm that the three SEGAN signals are perceived as higher quality than the noisy signals (0.550.55 to 0.450.45, with p< ⁣0.05p<\!0.05). DSEGAN and ISEGAN together significantly outperform SEGAN (0.670.67 to 0.330.33, p ⁣< ⁣0.001p\!<\!0.001). However, DSEGAN and ISEGAN qualities were not significantly different (0.480.48 to 0.520.52) in this small test. Results support the detailed objective evaluation in which DSEGAN performs much better than either SEGAN or noise, however we find that ISEGAN also performs well in subjective tests.

VI Conclusions

This paper presented a GAN method with multiple generators to tackle speech enhancement. Using multiple chained generators, the method aims to learn multiple enhancement mappings, each corresponding to a generator in the chain, to accomplish a multi-stage enhancement process. Two new architectures, ISEGAN and DSEGAN, were proposed. ISEGAN’s generators share their parameters and, as a result, are constrained to learn a common mapping for all the enhancement stages. DSEGAN, in contrast, has independent generators that allow them to learn different mappings at different stages. Objective tests demonstrated that the proposed ISEGAN and DSEGAN perform comparably and are better than SEGAN on speech-quality metrics and that learning independent mappings leads to better performance than a common mapping. In addition, both the proposed systems achieve more favourable results than SEGAN on the speech-intelligibility metric as well as the subjective perceptual test.

References