Non-intrusive speech quality assessment using neural networks

Anderson R. Avila, Hannes Gamper, Chandan Reddy, Ross Cutler, Ivan Tashev, Johannes Gehrke

Introduction

In speech communication systems, the audio signal can be affected by background noise, reverberation, enhancement algorithms as well as by network impairments. In such scenarios, as providers strive to guarantee optimal and reliable services to their customers, estimating the perceived quality of the audio signal has become crucial. For instance, speech quality prediction can be useful during network design and development as well as for monitoring and improving customers’ quality of experience (QoE) .

The subjective listening test is the most accurate method for evaluating perceived speech signal quality . In this approach, the estimated quality is the average of users’ judgment, usually in a scale ranging from 1 to 5. The average of all participants’ scores over a specific condition is referred to as the mean opinion score (MOS) and represents the perceived speech quality after leveling out individual factors . Such subjective measurements are not always feasible as they: (1) require a considerable number of listeners; (2) can be laborious and time-consuming; (3) can be expensive; and (4) perhaps, more importantly, cannot be done in real-time .

As an alternative, several objective instrumental quality measures have been proposed and standardized. The ITU-T Recommendation P.862, referred to as Perceptual Evaluation of Speech Quality (PESQ) , is one of the most widely used measures for audio quality assessment, followed by its improved version ITU-T Recommendation P.863, also known as Perceptual Objective Listening Quality Assessment (POLQA) . These models, however, were developed specifically for distortions introduced by speech compression (i.e., codecs) and show low performance when the audio signal is corrupted by noise, reverberation, or processed by new enhancement algorithms .

In addition, many of these approaches are intrusive as they require the reference clean speech signal to estimate the MOS. This limits their application to use with a synthetic dataset. There are numerous algorithms and standards which are non-intrusive ; they use only the corrupted speech signal for quality assessment. Normally, these algorithms are expected to be less accurate in estimating the perceptual sound quality.

Despite the breakthroughs of neural networks in so many areas, to date, only a handful of neural network-based models have been proposed . To the best of our knowledge, even the most recent methods to predict MOS present serious limitations. First, most of them are developed to measure intelligibility , which is just one aspect of audio quality . Second, these neural network-based models are trained on a limited number of conditions, usually with no interaction between different impairments, which is quite unrealistic and rarely happens in everyday scenarios. Finally, the ground truth to train these models frequently are not the subjective scores (MOS), but the scores of another model, such as PESQ , which leaves out most of the relevant human factors.

To tackle these limitations, we generated a realistic dataset, and we labeled it using a crowd-based QoE estimation . We explore three neural network-based architectures to predict MOS. In the first approach, we use a psychoacoustic inspired feature, namely the constant Q transform , which has been successfully adopted to distinguish natural and unnatural speech as well as to perform music analysis. These features are used as input to a convolutional neural network (CNN). In the second approach, we explore the low-dimensional total variability (TV) space . The features projected in the TV space, also referred to as i-vectors, are then used as input to a fully connected deep neural network (DNN). The third approach is based on the Mel-frequency features, combined with a DNN. The performance of the proposed approaches are compared with three instrumental measures: PESQ, the ITU-T Recommendation P.563, and the speech-to-reverberation energy ratio (SRMR).

The remainder of this document is organized as follows. In Section 22, we present the data generation process. The features adopted in this work are described in Section 33. Section 44 presents the neural network models evaluated in this paper. Our experimental setup is described in Section 55, and results are discussed in Section 66. Section 77 concludes the paper.

The Audio Quality Evaluation Dataset

In everyday environments, it is expected that an audio signal is subject to a variety of acoustic background distortions. To create realistic scenarios for our listening quality test, we created a dataset of 10,000 samples, representing the conditions to be assessed. We first generated 2,0102,010 clean speech files, equally distributed by gender: 670670 males, 670670 females and 670670 children. Each speech file is approximately 2020 seconds long, starting with 44 seconds of silence, followed by 33 utterances, separated by 22 seconds of silence. All samples are normalized to −23-23 dB FS, and are sampled with 1616 kHz. The human voice levels were modeled with a mean of 6565 dB SPL at 11 meter, and deviation of 88 dB. Then the clean signal is convolved with a randomly selected room impulse response (RIR) from a library of 120 RIRs. They were measured at distance between the source (speakers) and target (microphones) varying between 0.50.5 and 33 meters, in rooms with RT60RT_{60} ranging from 300300 to 500500 ms. Anechoic and close-talk microphone conditions are also included. Next, noise is added to the convolved audio signal, with a mean level of 4545 dB SPL and deviation of 1515 dB. Three types of noise were considered: offices (8080%), homes (1010%) and others (1010%). The ratio of office noise is higher as it is a more prominent noise type in our use case. The resulting SNRs were limited to $$ dB, as depicted in Fig. 1a. Finally, half of the corrupted samples were processed with an audio processing pipeline, consisting of a noise suppressor and automatic gain control (AGC). This allowed us to investigate how processed and unprocessed speech are perceived by human users.

As listening room and equipment may influence the outcome of the experiment, listening quality tests are commonly performed with each participant using the same listening conditions, usually a quiet chamber of controlled dimensions . This, however, does not represent realistic scenarios encountered in real life. Thus, we used crowd-sourcing to label our data. In this type of experiment, online workers are assigned to the task, which can be undertaken in a variety of ambiance and listening devices. Before initiating the experiment, participants were submitted to a training phase where they listened at least once, but if necessary as many times as they wanted, to samples of each impairment. This was meant to familiarize the participants to the most uncommon distortions and evaluation scale. After the training, the labelers had to pass a mandatory qualification step. They were asked to rate gold-standard samples. Only participants who successfully passed the qualification were considered for the experiment. Fig. 2 summarizes the dataset generation and labeling procedure. The perceptual audio quality of each audio sample was rated by ten judges, and the MOS is computed by averaging the scores. Fig. 1-b and -c provide the histograms of the computed PESQ and averaged MOS. It is well visible that PESQ gives lower scores than human judges.

Proposed Features

This section describes the three features adopted as input to our neural network models. We first present the constant Q spectrum, then the low-dimensional total variability space, and finalize with Mel-frequency features.

The short-time Fourier transform (STFT) is the most popular time-frequency representation of an audio signal. To extract it, one must choose a short window function that will be multiplied along the audio signal. The length of the window function is fixed, and commonly set to values between 1010 and 3030 ms. The quality factor QcQ_{c} for the center of the frequency band fcf_{c} is defined as:

where δf\delta_{f} is the frequency bandwidth. Note that for fixed width the quality factor increases with increasing center frequency. This is not aligned well with human perception, which is known to have a constant Q factor between 500Hz500Hz and 20kHz20kHz . Perceptually motivated, the constant Q transform (CQT) was introduced in and later refined in . Applying the CQT allows better time-frequency resolution as described in . Inspired by this, we include the constant Q spectral in the set of features to be evaluated. The feature dimension is 240x220. In the case of short duration utterances, the last frame was replicated the number of times necessary to attain 220 frames. Also, exceeding frames were removed from long utterances. This procedure was performed to assure that the CNN had always the same input size. This configuration was chosen empirically based on the average duration of the speech files.

where MM is the dependent supervector (extracted from a specific utterance) and mm is the independent supervector (extracted from the UBM), TT corresponds to a rectangular low-rank total variability matrix and ww is a random vector with a normal distribution, the so-called i-vector. In our experiments, a 400-dimensional i-vector was adopted.

To extract Mel features, the audio signal is processed in frames of 512512 samples, with a step size of 160160 samples, at a sampling rate of 1616 kHz. For each frame, 2626 Mel-frequency cepstral coefficients (MFCCs) are computed. The MFCCs are combined with pitch estimate, the output of a voice activity detector (VAD), and the log-power energy of the frame, as well as their first derivatives estimated using the preceding frame. The input to the neural network consists of the features computed for each speech-active frame, as determined by the VAD, plus the 12 frames preceding and succeeding that frame, resulting in an input feature vector of size 1×14501\times 1450.

Proposed Neural Network Models

CNN architectures have been successfully applied on a 2D image arrays . It consists of two typical operations: convolution and pooling. Convolutional layers are responsible for mapping, into their units, detected features from receptive fields in previous layers. This is referred to as a feature map and is the result of a weighted sum of the input features passed through a non-linearity such as ReLU . A pooling layer will typically take the maximum or average of a set of neighboring feature maps, reducing dimensionality by merging semantically similar features. The CNN model proposed here has two convolutional layers with 3232 filters each. In the first layer 25×3025\times 30 filters are used, followed by a 2×22\times 2 max pooling. A dropout (0.20.2) is used as a regularizer before the next two convolutional layers, which has 6464 3×33\times 3 filters each. After the second layer, a 2×22\times 2 max pooling is applied prior to another dropout (0.20.2). A fully-connected layer of dimensionality 6464 is then used prior to the output unit. We adopted ReLU as an activation function within the hidden units and a learning rate of 0.0001.

As the second architecture, a multilayer perceptron (MLP) is adopted. Such a DNN learns a better feature representation by mapping the input features into a linearly separable feature space . This is achieved by successive linear combinations of the input variables, zi=wixi+biz_{i}=w_{i}x_{i}+b_{i}, where wiw_{i} and bib_{i} are weights and biases, followed by a non-linear activation function. Our first DNN architecture has 400400 input units, followed, respectively, by 200200 and 100100 units in the first and second hidden layers. We used the same activation function, dropout and learning rate adopted for the CNN. The second proposed DNN model receives a feature vector of size 1×14501\times 1450 and has four fully connected layers with 10241024 hidden units each. We adopted, respectively, 0.50.5 and 0.00040.0004 as dropout and learning rate. Adam is used as an optimization algorithm for both architectures.

Such neural network models require a fixed length of the feature vectors, while the duration of the evaluated audio signal varies. This problem can be addressed either by computing statistics of the features before sending them to the neural network (e.g. i-vectors), or by feeding the neural network with a fixed length of extracted vectors multiple times until the audio file ends, while computing statistics across the timeline. The mean or the mode is typically used, but it is also possible to have an additional classifier, such as the extreme learning machine (ELM) , adopted in this work.

Experimental Setup

We compared the performance of our proposed methods to three speech quality metrics. PESQ is adopted as a benchmark as it is one of the most widely used instrumental quality measures. We also included two non-intrusive measures as a benchmark: the speech-to-reverberation energy ratio (SRMR) , and the ITU-T Recommendation P.563 . The SRMR has shown to be a good candidate for estimating speech quality and intelligibility, outperforming PESQ and ITU-T P.563, especially in reverberant and dereverberated speech.

The performance of the tested algorithms are compared using two criteria: Pearson’s correlation (ρ\rho), and mean squared error (MSE). The data is divided in 7070% for training, 1515% for validation and hyperparameters optimization, and the other 1515% for testing. All presented results are based on estimations of the MOS from the test set and the respective MOS attained from subjective scores.

Experimental Results

The results are presented in Table 1. The first three lines are the benchmarks, followed by the CNN trained on constant Q spectral. The DNN using i-vector as a feature set is followed by the results from DNN using Mel-frequency features. The best performance for both evaluation parameters is achieved by the Mel-frequency+DNN+ELM algorithm. Its Pearson’s correlation of 0.870.87 and MSE of 0.150.15 far exceed the non-intrusive standard P.563 with 0.550.55 and 0.310.31 respectively. PESQ was also surpassed by the DNN+ELM approach. Overall, all the proposed models outperformed the benchmark ones. This is due, in great part, by the fact that the proposed DNN models were able to capture human factors potentially neglected by the standard methods as it can be observed in Fig 1, where the PESQ distribution seems to be more aligned with the SNR rather than to the human perception, represented by the MOS distribution.

Conclusion

In this work, we developed a realistic audio quality dataset based on crowd-sourcing labelling. We also propose three neural network-based approaches for estimating MOS. All models are non-intrusive and their performances are compared to three instrumental measures: PESQ, ITU-T P.563, and SRMR. Results show that all of the proposed approaches outperform the other instrumental measures. The fully connected model using Mel-frequency features as input provided the best correlations and lowest mean squared errors, followed by the DNN model combined with i-vector and the CNN model combined with the constant Q spectrum. As future work, we will evaluate the proposed methods on an extended dataset with network impairments. We will also consider training a DNN model using the raw signal.

References