DNSMOS: A Non-Intrusive Perceptual Objective Speech Quality metric to evaluate Noise Suppressors

Chandan K A Reddy, Vishak Gopal, Ross Cutler

Introduction

Subjective evaluation of speech quality is the most reliable way to evaluate Speech Enhancement (SE) methods . However, subjective tests are not scalable as they require a considerable number of listeners, the process is laborious, time-consuming, and expensive. Conventional objective speech quality metrics such as Perceptual Evaluation of Speech Quality (PESQ) , Perceptual Objective Listening Quality Analysis (POLQA) and Signal to Distortion Ratio (SDR) are widely used to evaluate Speech Enhancement (SE) algorithms optimized for human perception. Some of these metrics are designed to predict the subjective Mean Opinion Score (MOS) obtained using the Telecommunication Standardization Sector of the International Telecommunication Union (ITU-T) Recommendation P.800 . There are also the intrusive metrics requiring a reference clean speech to compute the quality score. However, they are shown to correlate poorly with human rating when used for SE tasks that involve perceptually invariant transformations . Also, intrusive metrics cannot be used to evaluate real recordings when clean reference is unavailable in realistic scenarios.

ITU-T Recommendation P.563 is a non-intrusive technique and can directly operate on the degraded signal . However, it was developed for narrow-band applications and works well on limited impairment types. Recently, Deep Neural Networks (DNN) based approaches have been proposed to estimate the speech quality scores . Some of these learning-based approaches use other objective metrics as the ground truth to train their speech quality predictor. A handful of published methods use MOS obtained using P.800 as the ground truth to train their models. In , the authors trained the model to identify the Just Noticeable Difference (JND). MOS predictors trained on actual human ratings are more reliable than the ones trained to predict other objective metrics like PESQ or POLQA. The accuracy and robustness of the learned models depend on the quality of the human labels and also the quantity and diversity of the audio clips. There is no large scale human-labeled data set publicly available to train a model that can robustly generalize to a variety of audio impairments, especially for SE task. Though the learning-based approaches have a higher correlation to human ratings when tested on a limited and small test set, none of these speech quality metrics are as widely used as some of the intrusive objective metrics. Most of the publications in the area of noise suppression and SE test their methods on synthetic noisy speech test sets using intrusive speech quality metrics such as PESQ, POLQA and SDR.

In this paper, we have developed a robust objective perceptual speech quality metric called DNSMOS. It can be used to stack rank different Deep Noise Suppression (DNS) methods based on MOS estimates with great accuracy and hence the name DNSMOS. A simple Convolutional Neural Network (CNN) based model is trained using the ground truth human ratings obtained using ITU-T P.808 . P.808 is an online subjective testing framework that is highly reproducible and shown to stack rank noise suppression models with high accuracy when each model is tested as an average over a statistically significant number of clips. However, P.808 ratings per clip can be noisy due to various factors. The absolute MOS per clip might vary for the same clip in different P.808 runs due to the inherent noise in human rating . The number of ratings per clip varies due to spam removal. Also, all the spammers might not be filtered due to the limitations of spam filtering logic in P.808 implementation. We train a multi-stage self-teaching model inspired by continual life long learning to learn in the presence of noisy labels. We show that the multi-stage training significantly improves the correlation and generalizes remarkably well to other impairment types that are very different from what was used in the training set. We found the DNSMOS very useful in our audio/speech research as it generalizes well to a variety of speech impairments. Hence, we are providing the DNSMOS as an Azure service for other researchers to use. The details of the API are at www.microsoft.com/en-us/research/dns-challenge/dnsmos.

Data and subjective ratings

We used 600 noisy speech test clips comprised of a combination of synthetic and real recordings. The synthetic test set is a combination of both non-reverberant and reverberant clips. The real recordings were captured in a variety of noise types and Signal to Noise Ratio (SNR) and target levels. The test set is comprised of over 100 noise types and speakers. More details about the creation of these test sets can be found in . These 600 clips were processed by over 200 different noise suppression methods. Some of these methods improved the overall speech quality, but they also introduced a variety of artifacts as a consequence of over or under suppression of noise, resulting in speech and noise distortions. The speech quality ratings of the processed clips varied from very poor (MOS=1) to excellent (MOS=5). The distribution of the MOS scores in the training data is shown in Figure 1a. The scores are highly skewed with most ratings populated in the range 3

The subjective human ratings are obtained in several P.808 runs conducted over several months. Multiple noise suppression methods are compared in each P.808 run. Each P.808 run included the best performing noise suppressor, original noisy speech, and a couple of methods with intermediate perceptual quality from previous runs as anchors. Hence, some of the clips were rated multiple times. In total, we have about 120,000 audio clips with associated MOS scores as ground truth. The average length of each audio clip was about 9 seconds. In , we show that P.808 is highly reproducible with high SRCC between runs when we average the ratings across clips processed by a given model. However, the absolute values of the scores vary between P.808 runs as the raters differ and humans tend to be inconsistent in perceptual tasks due to various biases.

DNSMOS

The variation in the absolute values between P.808 runs might pose a challenge in training models that learn to predict MOS scores when the training data includes the same audio clip with different ratings. We know from the literature that DNN models can generalize well in spite of the noise for classification tasks . The variation in the absolute scores from different runs can be treated as noise in the labels. Also, the clips with fewer ratings can be treated as noise due to higher standard deviation. There are other unknown human biases that play a role in the variation of ratings. We initially took a direct approach where we trained a deep model by including all the available data without any modification. The model tends to learn the weighted average of the ratings across runs and also tries to average out other biases. This plain approach gave a good correlation with human ratings when tested on the data that was very similar to the training set. However, the correlation dropped when tested on a more challenging test set that was very different from the training data.

We improved the generalization and accuracy of the predictor by training a few stages of self-teaching models. In , a sequence of self-teachers were used to improve the generalizability of a classification task with binary cross-entropy loss in noisy and weakly labeled conditions. We use a similar approach for the regression problem at hand with Mean Squared Error (MSE) as the loss function. Unlike knowledge distillation where a larger teacher model is used to produce soft labels to train a smaller student model, we use the same model architecture for both teacher and the student, and hence the name self-teacher. The student model uses the weighted average of the predictions from one or multiple teacher models and the original human MOS ratings. The primary teacher model M0\textbf{{M}}_{0} is trained until the loss saturates using the original ground truth human MOS ratings r. The student model Ms\textbf{{M}}_{s} at stage s, is trained using the new target given by,

where r^i\hat{\textbf{{r}}}_{i} is the prediction of model Mi\textbf{{M}}_{i} and ∑i=0sαi=1\sum_{i=0}^{s}\alpha_{i}=1. In this work, we restrict s to 2. With some assumptions to the noise distribution, it can be theoretically proven that the Ms\textbf{{M}}_{s} is better than M0\textbf{{M}}_{0} for s>0\textit{s}>0. Since the distribution of noise in our case is complex, we show through empirical experiments that the predicted MOS of Mi\textbf{{M}}_{i} is of higher accuracy than M0\textbf{{M}}_{0} for appropriately chosen values of α\alpha.

2 Features

Recently, researchers have seen success in learning features within the model for tasks such as SE , speech, and music synthesis and to learn acoustic models . They show that using the time-domain waveform requires a larger model trained on a larger and diverse data set to ensure generalization. The ground truth MOS scores are obtained for audio clips with an average length of 9 secs sampled at 16 kHz. This leads to a very large input dimension if we are treating it as a vector and the model requires many layers to compress and extract input features. We used Log power Mel spectrogram as input feature as it correlates well with human perception and is proven to work very well for analyzing speech quality . For spectral features, we used a frame size of 20 ms with a hop length of 10 ms and 120 Mel frequency bands. The input features are then converted to dB scale.

3 Prediction model

For predicting the MOS scores, we explored different configurations of CNN based models. The architecture for the best performing model is shown in Table 1, which is the architecture for all M. The input to the model is log power Mel Spectrogram with 120 Mel bands computed over a clip of length 9 secs sampled at 16 kHz with a frame size of 20 ms and hop length of 10 ms. This results in an input dimension of 900 x 120. The model was trained with a batch size of 32 using the Adam optimizer and MSE loss function until the loss saturated. We experimented by adding batch normalization layers after every Conv layer in Table 1. However, adding batch normalization reduces the prediction accuracy of low volume clips. Humans tend to give lower ratings to the clips with low amplitudes. We want the model to capture the variations in the target levels of the data. Hence, we avoid any kind of feature normalization. We also explored different network architectures including CNN followed by LSTM. The model in Table 1 generalized the best and was of least complexity.

Experimental Results

Pearson Correlation Coefficient (PCC) or MSE between the predictions of the developed objective metric and the ground truth human ratings is commonly used to measure the accuracy of the model . From the earlier discussion, we know that P.808 correlation is highly repeatable between runs when averaged across a set of clips per condition, which can be formed by grouping clips enhanced by a particular SE model or based on other criteria like SNR or reverb RT60 times. The PCC computed on the average of ratings per group across different runs is >0.9. We also found that PCC computed on the same clips but from two different P.808 runs is only about 0.5 due to the high rating noise per clip.

Hence, for stack ranking different noise suppressors we evaluate by computing the average of ratings across the entire test set for each model. Therefore, we compute SRCC and PCC between averaged human ratings and averaged DNSMOS per model. SRCC gives us the stack ranking accuracy of various SE models.

2 Results

Table 2 shows the PCC and SRCC between human ratings and widely used PESQ, POLQA, SDR and DNSMOS (M0\textbf{{M}}_{0}) computed on the Interspeech DNS challenge blind test results . The results show that SDR correlates poorly with MOS. PESQ and POLQA reasonable PCC and SRCC. But the primary DNSMOS model (M0\textbf{{M}}_{0}) correlates significantly better than other objective metrics and is highly reliable in stack ranking the SE models. The scatter plots in Figure. 2 shows that DNSMOS aligns with MOS better than SDR, PESQ, and POLQA. Both PESQ and POLQA have a cluster on the left top of the graph indicating that they tend to penalize certain artifacts more than we humans do.

2.2 Generalizability of DNSMOS

We initially analyzed the effect of α0\alpha_{0} on M1\textbf{{M}}_{1} to understand the weightage given to the teacher model M0\textbf{{M}}_{0} vs human ratings. Figure. 3 shows the SRCC between MOS and M1\textbf{{M}}_{1} for different values of α0\alpha_{0} computed on one of the P.808 runs with 12 closely stack ranked models. We see in Figure. 3 that the SRCC is higher for α0=0.8\alpha_{0}=0.8.

The noisy and the processed clips comprised of a similar distribution as that of the training set. In order to test the generalization capability of training multiple stages of self-teaching models, we evaluated on the ICASSP 2021 DNS Challenge 2 final submissions , which comprised of the blind test set processed by 18 different noise suppressors. The blind set included emotional speech, singing, utterances in English and non-English languages including a few tonal languages. The DNSMOS models were trained on the data which was predominantly in English with very little or no emotional content such as laughter, crying, yelling, anger, etc. The training set did not include singing and other non-English languages as well. Table 3 shows the SRCC between the DNSMOS models at various stages and P.808 MOS for each category in the blind set. The SRCC of the primary model M0\textbf{{M}}_{0} is quite low as it does not generalize well for categories other than English. M1\textbf{{M}}_{1} significantly improved SRCC from 0.65 to 0.9 for α0=0.8\alpha_{0}=0.8. The optimal values of α0\alpha_{0} were different for each category in table 3. We also experimented with M2\textbf{{M}}_{2} for a couple of settings of (α0\alpha_{0},α1\alpha_{1},α2\alpha_{2}). For M2\textbf{{M}}_{2}(-, 0.5, 0.5) with M1\textbf{{M}}_{1}(0.2, 0.8, -) for stage 2 gave the best results. The other settings did not show much improvement over M1\textbf{{M}}_{1}.

Conclusion and Future work

DNSMOS is a robust speech quality metric designed to stack rank noise suppressors with great accuracy. The multi-stage self-teaching approach significantly improved accuracy. In the future, we would like to deepen our theoretical understanding of how the choice of αi\alpha_{i} influences the accuracy of Mi\textbf{{M}}_{i} for different datasets. This will help in training DNSMOS on other impairment types such as network distortions, codec artifacts, and reverberation without the need for large scale data with human labels.

References