An instrumental intelligibility metric based on information theory
Steven Van Kuyk, W. Bastiaan Kleijn, Richard C. Hendriks
I Introduction
Intelligibility is defined as the proportion of words correctly identified by a listener and is a natural measure for quantifying the effectiveness of speech-based communication systems . Although listening tests can provide valid data, such tests are time-consuming to conduct. For this reason, instrumental intelligibility metrics that are correlated with intelligibility and quick to compute are often preferred.
We can distinguish two types of instrumental intelligibility metrics: intrusive, and non-intrusive. Intrusive intelligibility metrics require knowledge of the clean speech and either the communication channel or degraded speech, whereas non-intrusive intelligibility metrics require only the degraded speech. In this paper we develop a new intrusive intelligibility metric based on information theory .
Existing intrusive intelligibility metrics include the speech intelligibility index (SII) , the speech transmission index (STI) , the coherence SII (CSII) , the extended SII (ESII) , the normalized covariance measure (NCM) , the hearing-aid speech perception index (HASPI) , the short-time objective intelligibility measure (STOI), the extended STOI (ESTOI), the speech-based envelope power spectrum model (sEPSM) , and the glimpse proportion metric (GP) . As a group, the above algorithms have been successful at predicting speech intelligibility in a wide-range of conditions including additive noise, filtering, reverberation, and non-linear enhancement. However, each intelligibility metric tends to perform well for only a narrow subset of conditions. This is because the above algorithms were heuristically motivated and were often designed with a specific type of distortion or data set in mind.
Information theory provides a mathematical framework for modelling communication systems. Information theoretical concepts have previously been used in the analysis of linguistics , speech production , and human hearing . Additionally, state-of-the-art speech enhancement algorithms and intelligibility metrics that are based on information theory have been developed.
Existing information theoretic intelligibility metrics, such as the mutual information k-nearest neighbour metric (MIKNN) , assume that speech can be described by a memoryless stochastic process and that the energy of a speech signal at one time-frequency location is statistically independent to the energy at all other time-frequency locations. In reality neither of these assumptions are valid, which leads to an over-estimate of the information shared between a talker and a listener.
In this paper we propose a conceptually simple intelligibility metric called SIIB. SIIB is a function of a clean acoustic signal produced by a talker and a degraded signal that is received by a listener. As described in Section II and Section III, the acoustic signals are converted to a representation of speech based on a crude model of the human auditory system. A non-parametric estimate of the mutual information rate of the signals is then computed. Unlike existing metrics, SIIB partially accounts for time-frequency dependencies in the speech signals using the Karhunen-Loève transform (KLT) and incorporates the theory developed in to account for the effect that talker-variability has on the information rate. In Section IV and Section V, SIIB is evaluated by comparing its performance to STOI , ESTOI , and MIKNN for speech degraded by noise and processed by enhancement algorithms.
II Model of Speech Communication
In this section we present a theoretical model of speech communication similar to that described in , and . The model considers the transmission of a message from a talker to a listener. Stochastic processes are denoted by , random variables are denoted by bold font, and their realisations are denoted by regular font.
The speech signal is transmitted to a listener through a communication channel that may distort the signal. Examples of distortion include noise, reverberation, speech coding algorithms, and speech enhancement algorithms. Overall, the communication process is described by a Markov chain:
We call the speech production channel and the environmental channel.
II-B Information Rate of the Communication Channel
The proposed intelligibility metric is based on the hypothesis that intelligibility is a function of the mutual information rate between the message and the degraded speech. Let , where denotes the transpose, be a vector obtained by stacking consecutive message vectors and similarly for . The mutual information rate is defined by
where is the mutual information between and given by
To estimate (3), realisations of and are needed. Estimating a realisation of requires a chorus of speech signals (see ). In typical applications of intelligibility prediction, such a chorus is not available, so instead we use an upper bound on (3). By applying the data processing inequality twice we have
In the case of a distortionless environmental channel, is unbounded from above and saturates at the information rate of the speech production channel . This maximum information rate is determined by the variability in pronunciation between different talkers. The following subsections describe how and can be calculated.
II-C Information Rate of the Environmental Channel
The mutual information rate of the environmental channel is given by
Estimating the mutual information between vectors of high dimensionality is a challenging task particularly when the vector elements have strong statistical dependencies . For this reason we introduce an invertible transform that aims to remove the dependencies between the vector elements.
where denotes the element index in the vector.
Finding an invertible that simultaneously removes the dependencies in both and is difficult. Early speech recognition systems used the discrete cosine transform (DCT), which results in Mel-frequency cepstral coefficients . It can be shown that the DCT approximates the Karhunen-Loève transform (KLT) for stationary signals . The KLT is the transformation that we use here and it is given by:
II-D Information Rate of the Speech Production Channel
Approximating and as Gaussian, the information rate of the speech production channel is
III Proposed Intelligibility Metric
The proposed intelligibility metric combines (7), (10), and (5) to give an estimate of the amount of information shared between and in bits per second. It is given by
A gammatone filterbank that includes filters linearly spaced on the ERB-rate scale between 100 Hz and 6500 Hz is used to obtain and according to (2). A sequence of stacked vectors for the clean speech is then formed by stacking consecutive vectors:
IV Evaluation Procedures
This section describes the procedures used to evaluate SIIB. The evaluation considered four intelligibility data sets and used
two performance measures to quantify the strength of the relationship between SIIB and intelligibility.
The first data set consists of speech subjected to single channel noise reduction. In phrases from the Dutch version of the Hagerman test were degraded by speech-shaped noise (SSN) at SNRs of and dB and processed by three noise reduction algorithms. The three algorithms compute a minimum mean-squared error estimate of the clean speech by multiplying the short-time spectral magnitude of the degraded speech with a gain function. In total there are 5 SNRs (3 algorithms + 1 unprocessed) = 20 conditions. The stimuli were presented to 13 normal-hearing subjects for identification.
IV-A2 KleijnPRE
The second data set consists of speech subjected to pre-processing enhancement and degraded by noise. In phrases from the Dutch version of the Hagerman test were subjected to three pre-processing enhancement algorithms and then degraded either by SSN at SNRs of and dB, or car noise at SNRs of and dB. The three enhancement algorithms optimally redistribute the energy of the clean speech according to a distortion criterion. In total there are 2 noise types 4 SNRs (3 algorithms + 1 unprocessed) = 32 conditions. The stimuli were presented to nine normal-hearing listeners for identification.
IV-A3 CookePRE
The third data set also consists of speech subjected to pre-processing enhancement. In Harvard sentences were processed by 19 pre-processing enhancement algorithms and degraded either by SSN at SNRs of and dB, or by speech from a competing talker at SNRs of and dB. The stimuli were presented to 175 normal-hearing listeners for identification. For this paper, a subset of the data in was considered because the entire data set was not available. Ten of the Harvard sentences and nine of the enhancement algorithms were used. The algorithms are referred to in as AdaptDRC, F0-shift, IWFEMD, on/offset, OptimalSII, RESSYSMOD, SBM, SEO, and SSS. In total there are 2 noise types 3 SNRs (9 algorithms + 1 unprocessed) = 60 conditions.
IV-A4 KjemsITFS
The fourth data set consists of speech subjected to ideal time-frequency segregation processing (ITFS). In phrases from the Dantale II corpus were degraded by four types of noise: SSN, cafeteria noise, noise from a bottling factory, and car noise. For each noise type, the degraded signals were processed by two types of ITFS called an ideal binary mask and a target binary mask. Three SNRs were used ( dB, and SNRs corresponding to 20% and 50% intelligibility) and eight variants of each ITFS algorithm were considered. In total there are 168 conditions. The stimuli were presented to 15 normal-hearing subjects for identification.
IV-B Performance Measures
The most important characteristic of an intelligibility metric is that it has a strong monotonic increasing relationship with intelligibility. This paper uses two performance measures to quantify the strength of the relationship: Kendall’s tau coefficient , and Pearson’s correlation coefficient . To use effectively, the relationship between the metric, , and intelligibility, , must be linear. For this reason, a monotonic function is applied to to linearise the relationship:
where are free parameters that are fit to each data set to minimise the mean squared error between and over all conditions. These free parameters are affected by the speech corpus, apparatus, and experimental procedures used during the listening test. Pearson’s correlation coefficient between and is then computed.
V Results
The performance of SIIB is compared to three state-of-the-art intelligibility metrics: STOI , ESTOI , and MIKNN . Fig. 1 shows scatter plots for each data set and each intelligibility metric. The vertical axis shows the intelligibility and the horizontal axis shows the score computed by an intelligibility metric . Each point represents a different condition in the data set. The function in (13) that is used to linearise the relationship is also shown. Table I displays for each data set and metric and, similarly, Table II displays .
The row of scatter plots corresponding to KleijnPRE shows that all of the reference metrics struggle to predict the effect that optimal energy redistribution has on intelligibility. In contrast SIIB is strongly correlated with intelligibility for this data set ( and ).
For CookePRE all of the metrics have reasonable performance except for STOI. This is in agreement with which showed that STOI performs poorly for speech degraded by modulated noise sources such as interfering talkers. An assumption sometimes made by the speech processing community is that in order to predict intelligibility for modulated noise sources, statistics have to be averaged over short-time segments to capture the affect of ‘listening for glimpses of clean speech’ . It is then surprising that SIIB performs well on this data set ( and ) because SIIB is based on global statistics only.
Compared to the reference metrics SIIB has excellent performance for JensenSCNR, KleijnPRE, and CookePRE, but poorer performance for KjemsITFS ( and ). In seventeen intelligibility metrics were evaluated using KjemsITFS and only five metrics achieved . SIIB may not perform as well on KjemsITFS because ITFS processing generates some stimuli with distortions that are not normally encountered in nature. For these stimuli it is plausible that humans are poor decoders. SIIB may correctly estimate the mutual information rate, but humans may be unable to efficiently use all of the information. This hypothesis could be tested by extensively training listeners to decode ITFS processed speech before conducting a listening test.
Notice that for maximum intelligibility, SIIB estimates an information rate of about b/s. This is higher than estimates based on linguistic models of speech communication where the information rate is 50-100 b/s . This overestimate is likely the consequence of approximating as Gaussian. Since is only approximately Gaussian, the KLT does not remove all statistical dependencies. Accounting for the remaining dependencies would give a lower information rate.
VI Conclusion
In this paper we proposed an intrusive instrumental intelligibility metric called SIIB. SIIB is based on the hypothesis that intelligibility is related to the amount of information shared between a clean and degraded speech signal in bits per second. Compared to existing metrics, SIIB is conceptually simple, theoretically motivated, and has high performance. According to Occam’s razor, these properties suggest that SIIB might generalise well to new data sets. A MATLAB implementation is available at https://stevenvankuyk.com/matlab_code/