iMetricGAN: Intelligibility Enhancement for Speech-in-Noise using Generative Adversarial Network-based Metric Learning
Haoyu Li, Szu-Wei Fu, Yu Tsao, Junichi Yamagishi
Introduction
Speech is the main media used by humans to communicate in daily life. However, in speech application systems, the intelligibility of speech messages is inevitably degraded due to background noise and reverberation. Several modification algorithms have been studied to enhance the intelligibility by preprocessing a signal before it is played out [Chermaz2019]. This task is usually termed near-end listening enhancement (NELE). There are various strategies for NELE tasks, such as spectral tilt flattening , formant shifting , and dynamic range compression . A common idea of the above algorithms is to reallocate the speech energy in the time-frequency (T-F) domain in such a way as to boost the acoustic cues that are perceptually crucial.
In this paper, we utilize deep neural networks (DNNs) to reallocate the speech energy. The output of a DNN acts as scale factor , which is then multiplied to the T-F bin of the speech. The T-F bin energy is boosted with and suppressed otherwise. This framework is very similar to the masking-based speech enhancement approach , where spectral mask is predicted by NN and applied to the T-F bin as well. Although NELE task shares a similar solution to DNN-based speech enhancement, very few related works have been done so far. One of the biggest challenges with the DNN-based NELE approach is that there is no ground truth label that can be provided for supervised training. Specifically, given an unmodified plain speech, there is no standard that explicitly defines what the perfectly intelligible speech should be, and thus no ground truth label can be prepared. In contrast, in a speech enhancement task, clean speech without noise mixed can be easily prepared and regarded as the training label of corresponding noisy speech.
Recently, MetricGAN was proposed and shown to be effective in optimizing the evaluation metrics in the field of speech enhancement. Inspired by its success, we adapt it to a modified iMetricGAN that fits in the intelligibility enhancement task. iMetricGAN is a generative adversarial network (GAN) system that consists of a generator to enhance the speech signal as the intelligibility enhancement module and a discriminator to learn to predict the intelligibility scores of modified speech. Instead of discriminating fake from real, the discriminator aims to closely approximate the intelligibility metrics as a learned surrogate, and then the generator can be trained properly with the guidance of this surrogate. From another point of view, with the framework of iMetricGAN, the ground truth speech label can be implicitly defined as a modified speech that achieves the maximum value of the intelligibility scores predicted by the discriminator. Consequently, the proposed iMetricGAN can effectively optimize the intelligibility of speech even though no ground truth label is provided. Furthermore, iMetricGAN is a flexible language-independent framework that can be easily extended to optimize multiple metrics simultaneously.
Problem Formulation and Assumptions
where denotes a convolution operation, is the room impulse response (RIR)Loudspeaker response is integrated into RIR for simplicity., and is the additive background noise. An assumption is made that the RIR and the noise can be estimated with an acoustic echo cancellation (AEC) technique . Therefore, the general NELE task is formulated as finding an algorithm to modify the natural speech to improve its intelligibility in known noise and room conditions.
In this work, for simplicity, we take only the noise signal into account and disregard , the influence of reverberation. Meanwhile, we do not change the RMS level and the duration of speech before and after modifications. With these assumptions and constraints, the problem can thus be formulated to design a DNN-based mapping function , as
Proposed iMetricGAN Model
In the field of speech enhancement, MetricGAN has shown a powerful ability to optimize complex and even non-differentiable speech quality metrics, such as PESQ . The proposed iMetricGAN adapts and revises the original MetricGAN for the intelligibility enhancement task, where the target metric to be optimized is the speech intelligibility measure.
2 Model description and training process
The model framework is depicted in Fig. 2. It consists of a generator (G) network and a discriminator (D) network. G receives speech and noise and then generates the enhanced speech. An energy normalization layer is inserted to guarantee the energy is maintained after modification. The final processed speech is notated as . The cascading D is utilized to predict the intelligibility score of the enhanced speech , given and . The output of D is notated as and expected to be close to the true intelligibility score calculated by a specific measure. We introduce the function to represent the intelligibility measures to be modeled, i.e., SIIB and ESTOI. With the above notations, the training target of D, shown in Fig. 2 (a), can be represented to minimize the following loss function:
Moreover, we introduce , the signal example that is enhanced by reference modification algorithms such as SSDRC , in the D training process. The loss function is thus extended to Equation (4).
The motivation for introducing is to improve the generalization of D. By feeding it with not only but also the signals modified by various other algorithms, D is encouraged to predict the intelligibility scores in a more accurate way. Thus Equation (4) can be seen as the loss function with auxiliary knowledge, while Equation (3) is the loss function with zero knowledge. Note that should not be regarded as the ground truth or the training label. In fact, the experimental results in Section 5 demonstrate that iMetricGAN still works well even without introducing .
For the G training process shown in Fig. 2 (b), D’s parameters are fixed and G is trained to reach intelligibility scores as high as possible. To achieve this, the target score in Equation (5) is assigned to the maximum value of the intelligibility measure.
G and D are iteratively trained until convergence. G acts as an enhancement module and is trained to cheat D in order to achieve a higher intelligibility score. On the other hand, D tries to not be cheated and to accurately evaluate the score of the modified speech. This minimax game finally makes both G and D effective. Consequently, the input speech can be enhanced to a more intelligible level by G.
Experimental Setup
The data used in our experiments were provided by the Hurricane Challenge 2 (https://hurricane-challenge.inf.ed.ac.uk/). Three languages are considered (English, German, Spanish) and 290 matrix sentences are available (100 each for German and Spanish; 90 for English). Sentences were uttered by native speakers under 3 reverberation conditions (Near, Mid, and Far). For each reverberation, sentences were presented with 3 different SNRs corresponding to three intelligibility levels: approximately 25%, 50%, and 75% correctly understood words. The detailed configuration is provided in Table 1. Since we ignore the influence of reverberation, there are actually 9 SNRs for each sentence per language. All signals are sampled at 44.1 kHz and the masker signal is Cafeteria noise.
We chose German and Spanish speech as the training set, with a total of 900 sentences (100 sentences 9 SNRs) for each language. Also, for data augmentation, we extracted 1,720 external sentences: 720 Spanish from the Sharvard corpus and 1000 German from the EMIME corpus . Each of them was resampled to 44.1 kHZ and mixed with masker signals at six different SNR levels (randomly selected from 9 corresponding SNRs) in order to form 10,320 extra sentences. In total, 12,120 German and Spanish sentences were used for training. We did not include English speech as training data since we wanted to investigate a language-mismatched condition as well as language-matched conditions. The 810 English sentences (90 sentences 9 SNRs) were used for testing only. To prepare the enhanced signal example introduced in Section 3.2, we selected three reference algorithms: (1) OptSII , a linear filter to maximize the Speech Intelligibility Index (SII), (2) OptMI , a linear filter to optimally redistribute energy based on mutual information criterion, and (3) SSDRC , a method to integrate spectral shaping and dynamic range compression. Each training sentence was randomly processed by one of these three algorithms to obtain its enhanced example.
2 Model architecture
All input signals are first transformed to magnitude spectrograms by short-time Fourier transform (STFT). A 1024-point Hanning window with 512-point hop size is applied and results in 513 frequency bins. Power-law compression with parameter is followed to compress spectrograms. We concatenate the spectrograms of input speech and masker noise to form 1026-dimensional (5132) input features for iMetricGAN. After modification, the processed spectrogram is converted into a time-domain waveform using inverse STFT with the phase of the input speech.
The generator model G is composed of two BLSTM layers, each with 400 hidden nodes, and two fully connected layers, each with 600 nodes. The activation function for the first fully connected layer is LeakyReLU with slope , and the last output layer is set as follows:
where is the result of the previous layer. This output serves as scale factors, which are point-wise multiplied with the input spectrogram (unmodified speech) to produce an enhanced spectrogram. Scale factors modify the input speech by redistributing its energy: the T-F bin is boosted or declined with the corresponding scale value. The scale range of Equation (6), which is approximately to , is empirically chosen. We expect such a wide range will facilitate the processing ability of G. Once processed by G, the enhanced spectrogram is normalized by the energy normalization layer, where the total squared energy of the output spectrogram is normalized to be the same as that of the input. The final processed spectrogram is sequentially passed on to the network D.
As shown in Fig. 2, the input features for D are 3-channel spectrograms, i.e., (processed, unprocessed, noise). D consists of five layers of 2-D CNN with the following number of filters and kernel size: [8 (5, 5)], [16 (7, 7)], [32 (10, 10], [48 (15, 15)], and [64 (20, 20)], each with LeakyReLU activation. Global average pooling is followed by the last CNN layer to produce a fixed 64-dimensional feature. Two fully connected layers are successively added, each with 64 and 10 nodes with LeakyReLU. The last layer of D is also fully connected and its output represents the scores of the intelligibility metrics. Therefore, the number of nodes in the last layer is equal to that of the intelligibility metrics we consider. For example, if we have D predict SIIB and ESTOI scores simultaneously, it should be set to . We normalize the SIIB score so that it ranges from 0 to 1, which is consistent with the range of the ESTOI score. Since both metrics of interest are bounded in , the sigmoid activation function is used in the last layer. Similar to , all the layers in D are constrained to be 1-Lipschitz continuous by spectral normalization to stabilize the training processSource codes of this work are available at https://github. com/nii-yamagishilab/intelligibility-MetricGAN.
Results
As described in Section 3.1, we have different options for learning metrics (SIIB, ESTOI, or both). In addition, two loss functions, Equations (3) and (4), can be chosen in the training process, depending on the use of enhanced examples. Therefore, we built and compared the iMetricGAN model with three different variationsAudio samples of the tested systems are available at https:// nii-yamagishilab.github.io/samples-iMetricGAN. To explain them, we use the following notations.
SiibGAN-zs: Learning target is SIIB, with Equation (3) as the loss function. Since there is no enhanced example provided in this loss function, the model is trained in a zero-short (zs) manner.
SiibGAN: Learning target is SIIB, with Equation (4) as the loss function.
MultiGAN: Learning target includes multiple metrics, SIIB and ESTOI, with Equation (4) as the loss function.
2 Objective evaluations
In addition to different iMetricGAN variants, three reference algorithms (OptMI, OptSII, SSDRC) used to produce enhanced examples were evaluated. Experimental results (SIIB and ESTOI) are presented in Table 2.
As shown, with modifications, the intelligibility scores of unmodified plain speech were effectively improved. Among all testing algorithms, MultiGAN had the best overall performance considering both ESTOI and SIIB scores, as expected. It surpassed the state-of-the-art SSDRC approach as well as OptMI and OptSII. By optimizing multiple metrics, it significantly outperformed another two iMetricGAN variants in terms of ESTOI, with only a slightly lower SIIB score compared to SiibGAN. For the SiibGAN-zs approach, the performance was degraded because the D network cannot be well trained using a zero-shot approach. Even so, it brought a significant intelligibility gain to plain speech, which further demonstrates the effectiveness of our proposed iMetricGAN model.
3 Subjective evaluations
Formal listening tests were conducted using the framework of the Hurricane Challenge 2. We submitted the entry of the MultiGAN algorithm, since it achieved the best performance in the objective evaluations. These listening tests took into account the sentences of three languages (English, German, Spanish) combined with different SNR and reverberation conditions. The SNRs used in the final tests were slightly different from those listed in Table 1 due to technical reasonsHurricane Challenge organizers adjusted the SNRs to allow for sufficient headroom in which the entries could show their performance.. Note that all provided languages and reverberations were considered in the subjective tests, while only English sentences without reverberation were used for the objective evaluation described in the previous section. We returned only the modified (enhanced) speech to the challenge organizers. These modified speech signals were remixed with noise and RIR by organizers, and then evaluated by native listeners. The masker signal was still Cafeteria noise but not a sample-by-sample equivalent of the signals we used in the training phase. Specifically, more than 180 listeners (60 for each language) were asked to type in what they heard in the experiments. Word accuracy rate, i.e., the percentage of correct words in a transcription, was then calculated as the performance measure of intelligibility.
Figure 3 shows the results for different listening conditions. Our proposed iMetricGAN achieved significant intelligibility gains in each language under all SNR and reverberation conditions. Note that we disregarded the influence of RIR when building the iMetricGAN model, but it still worked quite well in reverberant environments.
4 Discussions
In Fig. 3 (a), we can see that the word accuracy was improved for the English test set even though English sentences were not included in the training set. This demonstrates the good generalization capability of the iMetricGAN model, which can generalize well to mismatched language, speaker, and SNR level conditions. Future work will include the real-time implementation of iMetricGAN. To achieve this, we should change the original BLSTM to the uni-directional version and use noise power spectral density (PSD), which can be feasibly estimated as input noise information, instead of a raw noise signal. In the real-time inference stage, no global energy constraint can be guaranteed since the future signal values are not available, while the volume of speech can be adaptively maintained in a proper way by utilizing the automatic gain control (AGC) technique . Another future direction will involve introducing more advanced intelligibility metrics such as HASPI and HEGP to the model training. We also plan to investigate ways of integrating speech quality metrics such as PESQ to enhance the quality of the modified speech.
Conclusion
In this paper, we proposed the iMetricGAN model to enhance the intelligibility of speech-in-noise. Objective results show that our approach outperforms the state-of-the-art SSDRC method in terms of SIIB and ESTOI scores. Large-scale formal listening tests further show its effectiveness in intelligibility enhancement across different languages and background environment conditions.
Acknowledgments This work was partially supported by a JST CREST Grant (JPMJCR18A6, VoicePersonae project), Japan, and by MEXT KAKENHI Grants (16H06302, 17H04687, 18H04120, 18H04112, 18KT0051, 19K24372), Japan. The numerical calculations were carried out on the TSUBAME 3.0 supercomputer at the Tokyo Institute of Technology.