Exploring Speech Enhancement with Generative Adversarial Networks for Robust Speech Recognition
Chris Donahue, Bo Li, Rohit Prabhavalkar
Introduction
Speech enhancement techniques aim to improve the quality of speech by reducing noise. They are crucial components, either explicitly or implicitly , in ASR systems for noise robustness. Even with state-of-the-art deep learning-based ASR models, noise reduction techniques can still be beneficial . Besides the conventional enhancement techniques , deep neural networks have been widely adopted to either directly reconstruct clean speech or estimate masks from the noisy signals. Different types of networks have also been investigated in the literature for enhancement, such as denoising autoencoders , convolution networks and recurrent networks .
In their limited history, GANs have attracted attention for their ability to synthesize convincing images when trained on corpora of natural images. Refinements to network architecture have improved the fidelity of the synthetic images . Isola et al. demonstrate the effectiveness of GANs for image “translation” tasks, mapping images in one domain to related images in another. In spite of the success of GANs for image synthesis, exploration on audio has been limited. Pascual et al. demonstrate promising performance of GANs for speech enhancement in the presence of additive noise, posing enhancement as a translation task from noisy signals to clean ones. Their method, speech enhancement GAN (SEGAN), yields improvements to perceptual speech quality metrics over the noisy data and traditional enhancement baselines. Their investigation seeks to improve speech quality for telephony rather than ASR.
In this work, we study the benefit of GAN-based speech enhancement for ASR. In order to limit the confounding factors in our study, we use an existing ASR model trained on clean speech data to measure the effectiveness of GAN-based enhancement. To gauge performance under more-realistic ASR conditions, we consider reverberation in addition to additive noise. We first train a SEGAN model to map simulated noisy speech to the original clean speech in the time domain. Then, we measure the performance of the ASR model on noisy speech before and after enhancement by SEGAN. Our experiment indicates that SEGAN does not improve ASR performance under these noise conditions.
To address this, we refine the SEGAN method to operate on a time-frequency representation, specifically, log-Mel filterbank spectra. With this spectral feature mapping (SFM) approach, we can pass the output of our enhancement model directly to the ASR model (Figure 1). While deep learning has previously been applied to SFM for ASR , our work is the first to use GANs for this task. Michelsanti et al. employ GANs for SFM, but target speaker verification rather than ASR. Our frequency-domain approach improves ASR performance dramatically, though performance is comparable to the same enhancement model trained with an L1 reconstruction loss. Anecdotally speaking, the GAN-enhanced spectra appear more realistic than the L1-enhanced spectra when visualized (Figure 3), suggesting that ASR models may not benefit from the fine-grained details that GAN enhancement produces.
State-of-the-art ASR systems use MTR to achieve robustness to noise at inference time. While this strategy is known to be effective, a resultant model may still benefit from enhancement as a preprocessing stage. To measure this effect, we also use an existing ASR model trained with MTR and compare its performance on noisy speech with and without enhancement. We find that GAN-based enhancement degrades performance of this model, even with retraining. However, retraining the MTR model with both noisy and enhanced features in its input representation improves performance.
Generative Adversarial Networks
Generative adversarial networks (GANs) are unsupervised generative models that learn to produce realistic samples of a given dataset from low-dimensional, random latent vectors . GANs consist of two models (usually neural networks), a generator and a discriminator. The generator maps latent vectors drawn from some known prior to samples: , where . The discriminator is tasked with determining if a given sample is real (, a sample from the real dataset) or fake (, where is the implicit distribution of the generator when ). The two models are pitted against each other in an adversarial framework.
Real-world datasets often contain additional information associated with each example, e.g. the type of object depicted in an image. Conditional GANs (cGANs) use this information by providing it as input to the generator, typically in a one-hot representation: . After training, we can sample from the generator’s implicit posterior by fixing and sampling . To accomplish this, is trained to minimize the following objective, while is trained to maximize it:
Recently, researchers have used full-resolution images as conditioning information. Isola et al. propose a cGAN approach to address image-to-image “translation” tasks, where appropriate datasets consist of matched pairs of images in two different domains. Their approach, dubbed pix2pix, uses a convolutional generator that receives as input an image , a latent vector and produces : an image of identical resolution to . A convolutional discriminator is shown pairs of images stacked along the channel axis and is trained to determine if the pair is real or fake .
For conditional image synthesis, prior work demonstrates the effectiveness of combining the GAN objective with an unstructured loss. Noting this, Isola et al. use a hybrid objective to optimize their generator, penalizing it for L1 reconstruction error in addition to the adversarial objective:
Method
We describe our approach to spectral feature mapping using GANs, beginning by outlining the related time-domain SEGAN approach.
The generator’s encoder consists of layers of stride- convolution with increasing depth, resulting in a feature map at the bottle-neck of timesteps with depth . Here, the authors append a latent noise vector of the same dimensionality along the channel axis. The resultant matrix is input to an -layer upsampling decoder, with skip connections from corresponding input feature maps. As a departure from pix2pix, the authors remove batch normalization from the generator. Furthermore, they use 1D filters of width instead of 2D filters of size . They also substitute the traditional GAN loss function with the least squares GAN objective .
In agreement with observations from , we found that the SEGAN generator learned to ignore . We hypothesize that latent vectors may be unnecessary given the presence of noise in the input. We removed the latent vector from the generator altogether; the resultant deterministic model demonstrated improved performance in our experiments.
2 FSEGAN
It is common practice in ASR to preprocess time-domain speech data into time-frequency spectral representations. Phase information of the signal is typically discarded; hence, an enhancement model only needs to reconstruct the magnitude information of the time-frequency representation. We hypothesize that performing GAN enhancement in the frequency domain will be more effective than in the time domain. With this motivation, we propose a frequency-domain SEGAN (FSEGAN), which performs spectral feature mapping using an approach similar to pix2pix . FSEGAN ingests time-windowed spectra of noisy speech and is trained to output clean speech spectra.
The fully-convolutional FSEGAN generator contains encoder and decoder layers ( filters with stride and increasing depth), and features skip connections across the bottleneck between corresponding layers. The final decoder layer has linear activation and outputs a single channel. As with SEGAN, we exclude both batch normalization and latent codes from our generator, resulting in a deterministic model. The discriminator contains convolutional layers with filters and a stride of . A final layer (stride with sigmoid nonlinearity) aggregates the activations from frequency bands into a single decision for each of timesteps. We train FSEGAN with the objective in Equation 2. Other architectural details are identical to pix2pix .We modify the following open-source implementation: https://github.com/affinelayer/pix2pix-tensorflow The FSEGAN approach is depicted in Figure 2.
Experiments
As a source of reverberation for MTR, we use a room simulator as described in . The simulator randomizes the positions of the speech and noise sources, the position of a virtual stereo microphone, the T of the reverberation, and the room geometry. Through this process, our monaural speech data becomes stereo. Room configurations for training and testing are drawn from distinct sets; they are randomized during training and fixed during testing.
2 ASR Model
We train a monaural listen, attend and spell (LAS) model on the clean WSJ training data as described in Section 4.1, performing early stopping by the WER of the model on the validation set. To compare the effectiveness of GAN-based enhancement to MTR, we also train the same model using MTR as described in Section 4.1, using only one channel of the noisy speech. We refer to the clean-trained model as ASR-Clean and the MTR-trained model as ASR-MTR.
To process these features, our LAS encoder contains two convolutional layers with filter sizes: 1) , and 2) . The activations of the second layer are passed to a bidirectional, convolutional LSTM layer , followed by three bidirectional LSTM layers. The decoder contains a unidirectional LSTM with additive attention whose outputs are fed to a softmax over characters.
3 GAN
Results
We compute the WER of ASR-Clean and ASR-MTR on both the clean and MTR test sets. We also compute the WER of both models on the MTR test set enhanced by SEGAN and FSEGAN. Results are shown in Table 1. While the WER of ASR-Clean () is not state-of-the-art, we focus more on relative performance with enhancement. Our previous work has shown that LAS can approach state-of-the-art performance when trained on larger amounts of data.
The SEGAN method degrades performance of ASR-Clean on the MTR test set by relative. To verify the accuracy of this result, we also ran an experiment to remove only additive noise with SEGAN: the conditions in the original paper. Under that condition, we found that SEGAN improved performance of ASR-Clean by relative, indicating that SEGAN struggles to suppress reverberation.
In contrast, our FSEGAN method improves the performance of ASR-Clean by relative. While this is a dramatic improvement, it does not exceed the performance achieved with MTR training ( vs. WER). Furthermore, FSEGAN degraded performance for ASR-MTR, consistent with observations in .
We show a visualization of FSEGAN enhancement in Figure 3. The procedure appears to reduce both the presence of additive noise and reverberant smearing. Despite this, the procedure degrades performance of ASR-MTR. We hypothesize that the enhancement process may be introducing hitherto-unseen distortions that compromise performance.
Hoping to improve performance beyond that of MTR training alone, we retrain ASR-MTR using FSEGAN-enhanced features. To examine the effectiveness of the adversarial component of FSEGAN, we also experiment with training the same enhancement model using only the L1 portion of the hybrid loss function ( from Equation 2).
Considering that the model may benefit from knowledge of both the enhanced and noisy features, we also train a model to ingest these two representations stacked along the channel axis. We initialize this new hybrid model from the existing ASR-MTR checkpoint, setting the additional parameters to zero to ensure identical performance at the start of training. To ensure that the hybrid model is not strictly benefiting from increased parametrization, we train an LAS model from scratch with stereo MTR input. Results for these experiments appear in Table 2.
Retraining ASR-MTR with FSEGAN-enhanced features improves performance by relative to naively feeding them, but still falls short of MTR training. Hybrid retraining with both the original noisy and enhanced features improves performance further, exceeding the performance of stereo MTR training alone by relative. Our results indicate that training the same enhancer with the L1 objective achieves better ASR performance than an adversarial approach, suggesting limited usefulness of GANs in this context.
Conclusions
We have introduced FSEGAN, a GAN-based method for performing speech enhancement in the frequency domain, and demonstrated improvements in ASR performance over a prior time-domain approach. We provide evidence that, with retraining, FSEGAN can improve the performance of existing MTR-trained ASR systems. Our experiments indicate that, for ASR, simpler regression approaches may be preferable to GAN-based enhancement. FSEGAN appears to produce plausible spectra and may be more useful for telephonic applications if paired with an invertible feature representation.
Acknowledgements
The authors would like to thank Arun Narayanan, Chanwoo Kim, Kevin Wilson, and Rif A. Saurous for helpful comments and suggestions on this work.