Multi-Objective Learning and Mask-Based Post-Processing for Deep Neural Network Based Speech Enhancement

Yong Xu, Jun Du, Zhen Huang, Li-Rong Dai, Chin-Hui Lee

Introduction

Classical speech enhancement (SE) approaches, such as spectral subtraction , MMSE-based spectral amplitude estimator and optimally modified log-MMSE estimator , are considered as unsupervised techniques having been studied extensively for several decades. Based on key assumptions for the interactions between speech and noise, the tremendous progress has been made for those techniques in the past. However some issues, such as fast changing noise (e.g., machine gun ) and negative spectrum estimation, still need to be addressed.

On the other hand, supervised machine learning approaches have also been developed in recent years. They were shown to generate enhanced speech with good qualities . Non-negative matrix factorization (NMF) based speech enhancement was one notable example in which speech and noise basis models were learned separately from training speech and noise databases. Then the clean speech could be decomposed given the noisy speech. However, speech and noise are assumed uncorrelated and it limited the quality of the enhanced speech signals. Following recent successes in deep learning based speech processing we have recently proposed a deep neural network (DNN) based speech enhancement framework in which DNN was regarded as a regression model to predict the clean log-power spectra (LPS) features from noisy LPS features. DNN also acts as a mapping function to learn the relationship between clean and noisy speech features without imposing any assumption. Similar DNN-based speech denoising methods were also proposed in . In , DNN-based method was demonstrated to be better than the NMF-based methods in speech separation. In DNN-based speech enhancement, the minimum mean square error (MMSE) between the target features and the predicted features was always used as the objective function. It is difficult to design a better cost function to directly optimize the DNN model, especially for features that are correlated. In it was shown that other cost functions, such as the Kullback Leibler divergence or the Itakura-Saito divergence , all performed worse than the MMSE.

In this paper, a multi-objective learning framework is proposed to optimize a joint objective function, encompassing errors not only for the primary clean LPS features but also errors in secondary targets for continuous features, such as MFCC, and for categorical information, such as ideal binary mask (IBM) . This joint optimization of different but related targets can potentially improve the DNN prediction performance of the primary target LPS which is then used to reconstruct the enhanced waveform. In the LPS domain, the target values of different frequency bins were predicted independently without any correlation constraint, and some knowledge in auditory perception is not easily utilized. Nonetheless in the MFCC domain, mel-filtering is first applied and the correlation of each channel is represented in the MFCC coefficients. Furthermore, IBM is the most important concept in the computational auditory scene analysis (CASA) . IBM which represents the noise-dominant or speech-dominant meta information can also improve DNN training and the estimated IBM could further be used for post-processing. Finally, MFCC and IBM can be combined together to help predict the target clean LPS features.

In our SE experiments, we find that learning MFCC and/or IBM as secondary tasks provides improvements to DNN-based speech enhancement. Furthermore, IBM-based post-processing also gives an additional 1.5 dB improvement of segmental signal-to-noise ratio (SSNR) .

Multi-objective Learning for DNN-based Speech Enhancement

In , DNN is adopted as a mapping function to predict the clean LPS features from the noisy LPS features. The relationship between the clean and noisy speech features can be well learned because nearly no assumptions were imposed during the training process. However, other DNN-based methods, such as binary or soft mask based speech enhancement, assume that speech and noise are independent at each time-frequency (T-F) unit.

Normalized MMSE is used to update the DNN weights,

where ErEr is the normalized mean squared error and it can also be treated as the reciprocal of signal-to-noise ratio (SNR). This normalized squared error always reduces the distribution diversity of the clean training data and makes DNN training more stable. It should be noted that all the input and output features are normalized with a global mean and variance of the noisy training data. Hence, X^n\hat{\textbf{X}}_{n} and Xn\textbf{X}_{n} denote the estimated and clean normalized LPS at sample index nn, respectively, with NN representing the mini-batch size, Yn±τ\textbf{Y}{{}_{n\pm\tau}} being the noisy LPS feature vector where the window size of the context is 2∗τ+12*\tau+1, with (W,b)(\textbf{W},\textbf{b}) denoting the weight and bias parameters to be learned.

In this study, multi-objective learning is proposed to jointly predict the primary LPS features together with other secondary continuous features, such as MFCC, or/and some discrete category information, such as IBM, to enhance DNN learning as follows,

where X^cont\hat{\textbf{X}}^{\text{cont}} and Xcont\textbf{X}^{\text{cont}} denote the estimated and clean continuous features (also normalized), respectively. Ycont\textbf{Y}{{}^{\text{cont}}} represents the second noisy continuous feature. X^cate\hat{\textbf{X}}^{\text{cate}} and Xcate\textbf{X}^{\text{cate}} denote the estimated and target meta category information, respectively. α\alpha and β\beta are the weighting coefficients of this two other error parts, respectively. Unlike linear continuous features, meta information just has binary values, which makes the normalization not necessary for squared error related with the category part. Fig. 1 presented the structure of the proposed multi-objective learning. In fact, it was similar to the multi-task learning , but different from the multi-task learning in recent DNN-based speech recognition with only one input feature type. The prediction for the secondary continuous feature should be complementary with the prediction for the primary LPS using the shared DNN. The learning for the category information with linear activation function should also promote the prediction of clean LPS. Overall, multi-objective learning can improve the generalization capability of DNN for the clean LPS estimation.

MFCC is one of the most popular speech features used in speech recognition , speaker recognition and music modeling [logan2000mel]. Mel-filtering is applied to make it consistent with human auditory perception. However there is so far no prior auditory knowledge adopted in the LPS domain except for the log-compression. We believe the clean LPS features would be better predicted with a MFCC constraint imposed at the output layer. Furthermore, the discrete cosine transformation (DCT) operation in MFCC can incorporate the correlation information of different channels into each MFCC coefficient. We therefore expect correlated and consistent distortion across different frequency bins can be learned when predicting the clean LPS. Noted that DCT here is not performing dimension reduction which means the same dimensional MFCC features as the Mel-filter bank features are extracted.

One similar work in showed that the concatenation of different input features could improve the performance of DNN-based speech separation. However the motivation of our work is multi-objective learning with a novel architecture in both input and output layers, which is totally different from the motivation of feature fusion in . It is expected that the enhancement of MFCC would be complimentary to the enhancement of LPS.

2 Joint Prediction of LPS with IBM

IBM is one type of category information often used to represent the noise-dominant or speech-dominant nature at a certain T-F bin . If the local SNR of a T-F bin is greater than a threshold, the IBM is set to one otherwise it is set to zero. Just like MFCC, IBM is also used as a constraint term in the joint objective function. IBM explicitly offers the additional speech presence information at T-F units. With this discriminative information, the speech components would be emphasized while reducing more noise components.

In addition, the joint prediction of clean LPS with clean MFCC and IBM can be combined together. The noisy MFCC augmented in the input with the noisy LPS can also improve the IBM-based post-processing performance with an accurate IBM estimation to be discussed in the next section.

3 IBM-based Post-processing

The direct prediction of the clean LPS using DNN may lead to an overestimate or underestimate problem at some T-F units. The estimated IBM can be used for post-processing to simultaneously control the noise reduction level and speech distortion as follows,

where IBM^n(d)\widehat{\text{IBM}}_{n}(d) denotes the estimated IBM at time frame nn and frequency bin dd. Noted that the estimated IBM is close to the range $$. If the estimated IBM value is very large indicating that it has very high SNR at certain T-F unit, it is not necessary to perform noise reduction which can potentially result in the speech distortion. This is also the mask concept in . If the estimated IBM has a medium value, the average value between the noisy LPS and the estimated LPS was used. Otherwise, the DNN predicted LPS was adopted. The proposed IBM post-processing scheme in Eq. (3) is therefore different from where the estimated soft mask was used as a Wiener gain to perform speech enhancement. In contrast to adopting DNN to learn the mask there is no independence assumption between speech and noise in our DNN based mapping strategy.

Experimental Results and Analysis

In , all experiments were conducted on waveforms with 8kHz sample rate, in this work we extended it to 16kHz sample rate. 104 noise types were used in , however, in this study 115 noise types including some musical noises were adopted to further improve the generalization capacity of DNN. These 115 noise types include 100 noise types recorded by G. Hu and 15 home-made noise typesThe 115 noise types for training are N1-N17: Crowd noise; N18-N29: Machine noise; N30-N43: Alarm and siren; N44-N46: Traffic and car noise; N47-N55: Animal sound; N56-N69: Water sound; N70-N78: Wind; N79-N82: Bell; N83-N85: Cough; N86: Clap; N87: Snore; N88: Click; N88-N90: Laugh; N91-N92: Yawn; N93: Cry; N94: Shower; N95: Tooth brushing; N96-N97: Footsteps; N98: Door moving; N99-N100: Phone dialing; N101: AWGN; N102: Babble; N103-N105: Car; N106-N115: musical instruments. And all of them can be downloaded at http://home.ustc.edu.cn/~xuyong62/demo/115noises.html. And the clean speech data is derived from the TIMIT corpus . All 4620 utterances from the training set of the TIMIT database were corrupted with the abovementioned 115 noise types at six levels of SNR, i.e., 20dB, 15dB, 10dB, 5dB, 0dB, and -5dB, to build 80 hours multi-condition training set, consisting of pairs of clean and noisy speech utterances. The 192 utterances from the core test set of TIMIT database were used to construct the test set for each combination of noise types and SNR levels. As we only conduct the evaluation of unseen noise types in this paper, three other noise types, namely Buccaneer1, Destroyer engine and HF channel were adopted for testing. All of them are collected from the NOISEX-92 corpus . An improved version of OM-LSA , denoted as LogMMSE, was used for performance comparison with our DNN approach.

A short-time Fourier analysis was used to compute the DFT of each overlapping windowed frame. Then 257 dimensions LPS features were used to train DNNs. Segmental SNR (SSNR in dB) , perceptual evaluation of speech quality (PESQ) , and short-time objective intelligibility (STOI) were used to assess the quality and intelligibility of the enhanced speech. Frequency-dependent log-spectral distortion, defined as subtracting estimated LPS from clean LPS at each frequency bin, was also proposed to analyze the consistency of distortion across frequencies. Rectified linear units (ReLU) was used as the activation function of DNN, and the DNN was initialized with random weights. Dropout and static noise aware training as in were used to improve its generalization capacity for unseen noise environments. Mean and variance normalization was applied to the input and target feature vectors of the DNN. All DNN configurations were fixed at L=3L=3 hidden layers, 2500 units at each hidden layer, and 7-frame input. The MFCC used in Section 2.1 had 40 dimensions of static feature and one energy dimension using 40 Mel-filters. The empirical value of α\alpha and β\beta in Eq. (2) are set to 0.1 and 0.002, respectively. The empirical value of γ\gamma and ε\varepsilon in Eq. (3) are set to 0.9 and 0.6, respectively.

In Table 1, average PESQ and SSNR comparison on the test set at different SNRs of the three unseen noise environments among: DNN baseline, MFCC augmented in the output (denoted as MFCC-o) and MFCC augmented in both the input and output (denoted as MFCC), were given. MFCC-o system consistently outperformed the DNN baseline in PESQ and SSNR which indicated that the simultaneous prediction of MFCC was beneficial for the estimation of clean LPS. Furthermore, the noisy MFCC was complementary with the noisy LPS in the input to improve the prediction of clean LPS. And the MFCC system got the best performance, such as the average PESQ improved from 2.508 to 2.664. The multi-task of MFCC enhancement and LPS enhancement shared the DNN weights and promoted each other.

The frequency-dependent log-spectral distortion between the DNN baseline and MFCC systems calculated from 192 testing utterances at SNR=0dB corrupted by the Buccaneer1 noise was also given in Fig. 2. The overall shape of this log-spectral distortion is determined by the noise type, such as here the Buccaneer1 noise has two continual and high energy parts at frequencies shown in the ellipses. But with the constraint of MFCC, the speech distortion at low frequencies where the most of speech info located was largely reduced and more consistent. This was because MFCC emphasized the info at low frequencies with the Mel-filtering.

2 Joint Prediction of LPS and IBM with Post-processing

Table 1 also presented the average PESQ and SSNR comparison for joint prediction of LPS and IBM on the test set at different SNRs of the three unseen noise environments. With the IBM constraint in the output, better average PESQ and SSNR performance could be obtained compared with the DNN baseline, especially in SSNR which improved from -0.084 dB to 0.251 dB at SNR=0dB. Moreover, the IBM-based post-processing can obtain large PESQ and SSNR improvements, especially at high SNRs, e.g., SSNR improved from 7.641 dB to 11.455 dB at SNR=20dB which implies that the baseline DNN might hurt the speech components due to under-estimation, especially at the T-F units with high SNRs. Hence, IBM-based post-processing is crucial in achieving less speech distortion. This also conformed the mask concept in that it was not necessary to reduce noise when the speech energy is much larger than the noise energy at the certain T-F unit. In addition, IBM could be combined with MFCC. Compared with the performance of MFCC system, the combined system (MFCC+IBM in Table 1) gave slightly better results at all SNR levels. For example SSNR was improved from -1.433 dB to -1.086 dB at SNR=-5dB. Finally, the average SSNR of the best MFCC+IBM+PP system was improved from 3.664 dB to 5.194 dB.

3 Overall Performance Comparison

PESQ and STOI are often adopted to represent the objective quality and intelligibility of the enhanced speech, respectively. And STOI is often more meaningful at lower SNRs. An overall PESQ and STOI comparison of different SE techniques discussed in this study on the test set at different SNRs of the three unseen noise environments is displayed in Table 2. Compared with the noisy speech results, LogMMSE could yield 0.418 PESQ improvement while only 0.011 STOI improvement on average. The DNN baseline improved the LogMMSE with an average STOI from 0.801 to 0.845 across six SNRs. Our proposed MFCC+IBM+PP system overwhelms LogMMSE at all SNRs, especially at low SNRs, e.g., 0.163 STOI improvement and 0.626 PESQ improvement at SNR=-5dB.

Fig. 3 presented spectrograms of an utterance. The non-stationary noise was successfully reduced in the DNN-enhanced spectrum, while LogMMSE could not well track the non-stationary Buccaneer1 noise (its spectrogram can be seen at the demo websitehttp://home.ustc.edu.cn/~xuyong62/demo/IS15.html). Compared with the baseline DNN-enhanced spectrogram, the improved DNN can enhance the speech with less speech distortion shown in the three dashed arrow areas, especially at the consonant portions which are similar to noise. Furthermore the improved DNN can also reduce noise shown in the rectangle highlight segments. More enhanced waveforms of real-world noisy speech can also refer to the website.

Conclusion

In this paper, multi-objective learning is proposed to improve DNN training for speech enhancement. Adding constraints from features like MFCC or IBM in the objective function is shown to obtain more accurate estimation of clean LPS. MFCC can make the log-spectral distortion more consistent across low frequencies; IBM can explicitly represent the speech presence information at T-F units, so higher SSNR could be obtained. Furthermore, the estimated IBM can be adopted to do post-processing to alleviate the over-estimate or under-estimate problems in regression-based DNN. And IBM-based post-processing was crucial to reduce speech distortion, especially at high SNR T-F units. Compared with DNN baseline, about 0.2 PESQ and 0.03 STOI improvements were obtained on average. In the future, other continuous features and meta information will be further explored.

Acknowledgement

This work was partially supported by the National Nature Science Foundation of China (Grant Nos. 61273264 & 61305002).

References