Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation

Yi Luo, Nima Mesgarani

I Introduction

Robust speech processing in real-world acoustic environments often requires automatic speech separation. Because of the importance of this research topic for speech processing technologies, numerous methods have been proposed for solving this problem. However, the accuracy of speech separation, particularly for new speakers, remains inadequate.

Most previous speech separation approaches have been formulated in the time-frequency (T-F, or spectrogram) representation of the mixture signal, which is estimated from the waveform using the short-time Fourier transform (STFT) . Speech separation methods in the T-F domain aim to approximate the clean spectrogram of the individual sources from the mixture spectrogram. This process can be performed by directly approximating the spectrogram representation of each source from the mixture using nonlinear regression techniques, where the clean source spectrograms are used as the training target . Alternatively, a weighting function (mask) can be estimated for each source to multiply each T-F bin in the mixture spectrogram to recover the individual sources. In recent years, deep learning has greatly advanced the performance of time-frequency masking methods by increasing the accuracy of the mask estimation . In both the direct method and the mask estimation method, the waveform of each source is calculated using the inverse short-time Fourier transform (iSTFT) of the estimated magnitude spectrogram of each source together with either the original or the modified phase of the mixture sound.

While time-frequency masking remains the most commonly used method for speech separation, this method has several shortcomings. First, STFT is a generic signal transformation that is not necessarily optimal for speech separation. Second, accurate reconstruction of the phase of the clean sources is a nontrivial problem, and the erroneous estimation of the phase introduces an upper bound on the accuracy of the reconstructed audio. This issue is evident by the imperfect reconstruction accuracy of the sources even when the ideal clean magnitude spectrograms are applied to the mixture. Although methods for phase reconstruction can be applied to alleviate this issue , the performance of the method remains suboptimal. Third, successful separation from the time-frequency representation requires a high-resolution frequency decomposition of the mixture signal, which requires a long temporal window for the calculation of STFT. This requirement increases the minimum latency of the system, which limits its applicability in real-time, low-latency applications such as in telecommunication and hearable devices. For example, the window length of STFT in most speech separation systems is at least 32 ms and is even greater in music separation applications, which require an even higher resolution spectrogram (higher than 90 ms) .

Because these issues arise from formulating the separation problem in the time-frequency domain, a logical approach is to avoid decoupling the magnitude and the phase of the sound by directly formulating the separation in the time domain. Previous studies have explored the feasibility of time-domain speech separation through methods such as independent component analysis (ICA) and time-domain non-negative matrix factorization (NMF) . However, the performance of these systems has not been comparable with the performance of time-frequency approaches, particularly in terms of their ability to scale and generalize to large data. On the other hand, a few recent studies have explored deep learning for time-domain audio separation . The shared idea in all these systems is to replace the STFT step for feature extraction with a data-driven representation that is jointly optimized with an end-to-end training paradigm. These representations and their inverse transforms can be explicitly designed to replace STFT and iSTFT. Alternatively, feature extraction together with separation can be implicitly incorporated into the network architecture, for example by using an end-to-end convolutional neural network (CNN) . These methods are different in how they extract features from the waveform and in terms of the design of the separation module. In , a convolutional encoder motivated by discrete cosine transform (DCT) is used as the front-end. The separation is then performed by passing the encoder features to a multilayer perceptron (MLP). The reconstruction of the waveforms is achieved by inverting the encoder operation. In , the separation is incorporated into a U-Net 1-D CNN architecture without explicitly transforming the input into a spectrogram-like representation. However, the performance of these methods on a large speech corpus such as the benchmark introduced in has not been tested. Another such method is the time-domain audio separation network (TasNet) . In TasNet, the mixture waveform is modeled with a convolutional encoder-decoder architecture, which consists of an encoder with a non-negativity constraint on its output and a linear decoder for inverting the encoder output back to the sound waveform. This framework is similar to the ICA method when a non-negative mixing matrix is used and to the semi-nonnegative matrix factorization method (semi-NMF) , where the basis signals are the parameters of the decoder. The separation step in TasNet is done by finding a weighting function for each source (similar to time-frequency masking) for the encoder output at each time step. It has been shown that TasNet has achieved better or comparable performance with various previous T-F domain systems, showing its effectiveness and potential.

While TasNet outperformed previous time-frequency speech separation methods in both causal and non-causal implementations, the use of a deep long short-term memory (LSTM) network as the separation module in the original TasNet significantly limited its applicability. First, choosing smaller kernel size (i.e. length of the waveform segments) in the encoder increases the length of the encoder output, which makes the training of the LSTMs unmanageable. Second, the large number of parameters in deep LSTM network significantly increases its computational cost and limits its applicability to low-resource, low-power platforms such as wearable hearing devices. The third problem which we will illustrate in this paper is caused by the long temporal dependencies of LSTM networks which often results in inconsistent separation accuracy, for example, when changing the starting point of the mixture. To alleviate the limitations of the previous TasNet, we propose the fully-convolutional TasNet (Conv-TasNet) that uses only convolutional layers in all stages of processing. Motivated by the success of temporal convolutional network (TCN) models , Conv-TasNet uses stacked dilated 1-D convolutional blocks to replace the deep LSTM networks for the separation step. The use of convolution allows parallel processing on consecutive frames or segments to greatly speed up the separation process and also significantly reduces the model size. To further decrease the number of parameters and the computational cost, we substitute the original convolution operation with depthwise separable convolution . We show that with these modifications, Conv-TasNet significantly increases the separation accuracy over the previous LSTM-TasNet in both causal and non-causal implementations. Moreover, the separation accuracy of Conv-TasNet surpasses the performance of ideal time-frequency magnitude masks, including the ideal binary mask (IBM ), ideal ratio mask (IRM ), and Winener filter-like mask (WFM ) in both signal-to-distortion ratio (SDR) and subjective (mean opinion score, MOS) measures.

The rest of the paper is organized as follows. We introduce the proposed Conv-TasNet in section II, describe the experimental procedures in section III, and show the experimental results and analysis in section IV.

II Convolutional Time-domain Audio Separation Network

The fully-convolutional time-domain audio separation network (Conv-TasNet) consists of three processing stages, as shown in figure 1 (A): encoder, separation, and decoder. First, an encoder module is used to transform short segments of the mixture waveform into their corresponding representations in an intermediate feature space. This representation is then used to estimate a multiplicative function (mask) for each source at each time step. The source waveforms are then reconstructed by transforming the masked encoder features using a decoder module. We describe the details of each stage in this section.

In time-domain audio separation, we aim to directly estimate si(t),i=1,…,Cs_{i}(t),i=1,\ldots,C, from x(t)x(t).

II-B Convolutional encoder-decoder

Although we reformulate the encoder/decoder operations as matrix multiplication, the term ”convolutional autoencoder” is used because in actual model implementation, convolutional and transposed convolutional layers can more easily handle the overlap between segments and thus enable faster training and better convergence. With our Pytorch implementation, this is possibly due to the different autograd mechanisms in fully-connected layer and 1-D (transposed) convolutional layers.

II-C Estimating the separation masks

II-D Convolutional separation module

Motivated by the temporal convolutional network (TCN) , we propose a fully-convolutional separation module that consists of stacked 1-D dilated convolutional blocks, as shown in figure 1 (B). TCN was proposed as a replacement for RNNs in various sequence modeling tasks. Each layer in a TCN consists of 1-D convolutional blocks with increasing dilation factors. The dilation factors increase exponentially to ensure a sufficiently large temporal context window to take advantage of the long-range dependencies of the speech signal, as denoted with different colors in figure 1 (B). In Conv-TasNet, MM convolutional blocks with dilation factors 1,2,4,…,2M−11,2,4,\ldots,2^{M-1} are repeated RR times. The input to each block is zero padded accordingly to ensure the output length is the same as the input. The output of the TCN is passed to a convolutional block with kernel size 11 (1×1−conv1\times 1{-}conv block, also known as pointwise convolution) for mask estimation. The 1×1−conv1\times 1{-}conv block together with a nonlinear activation function estimates CC mask vectors for the CC target sources.

Figure 1 (C) shows the design of each 1-D convolutional block. The design of the 1-D convolutional blocks follows , where a residual path and a skip-connection path are applied: the residual path of a block serves as the input to the next block, and the skip-connection paths for all blocks are summed up and used as the output of the TCN. To further decrease the number of parameters, depthwise separable convolution (S-conv(⋅)S\text{-}conv(\cdot)) is used to replace standard convolution in each convolutional block. Depthwise separable convolution (also referred to as separable convolution) has proven effective in image processing tasks and neural machine translation tasks . The depthwise separable convolution operator decouples the standard convolution operation into two consecutive operations, a depthwise convolution (D-conv(⋅)D\text{-}conv(\cdot)) followed by pointwise convolution (1×1−conv(⋅)1\times 1{-}conv(\cdot)):

A nonlinear activation function and a normalization operation are added after both the first 1×1-conv1\times 1\text{-}conv and D-convD\text{-}conv blocks respectively. The nonlinear activation function is the parametric rectified linear unit (PReLU) :

III Experimental procedures

We evaluated our system on two-speaker and three-speaker speech separation problems using the WSJ0-2mix and WSJ0-3mix datasets . 30 hours of training and 10 hours of validation data are generated from speakers in si_tr_s from the datasets. The speech mixtures are generated by randomly selecting utterances from different speakers in the Wall Street Journal dataset (WSJ0) and mixing them at random signal-to-noise ratios (SNR) between -5 dB and 5 dB. 5 hours of evaluation set is generated in the same way using utterances from 16 unseen speakers in si_dt_05 and si_et_05. The scripts for creating the dataset can be found at . All the waveforms are resampled at 8 kHz.

III-B Experiment configurations

The networks are trained for 100 epochs on 4-second long segments. The initial learning rate is set to 1e−31e^{-3}. The learning rate is halved if the accuracy of validation set is not improved in 3 consecutive epochs. Adam is used as the optimizer. A 50% stride size is used in the convolutional autoencoder (i.e. 50% overlap between consecutive frames). Gradient clipping with maximum L2L_{2}-norm of 5 is applied during training. The hyperparameters of the network are shown in table I. A Pytorch implementation of the Conv-TasNet model can be found at https://github.com/naplab/Conv-TasNet.

III-C Training objective

The objective of training the end-to-end system is maximizing the scale-invariant source-to-noise ratio (SI-SNR), which has commonly been used as the evaluation metric for source separation replacing the standard source-to-distortion ratio (SDR) . SI-SNR is defined as:

III-D Evaluation metrics

We report the scale-invariant signal-to-noise ratio improvement (SI-SNRi) and signal-to-distortion ratio improvement (SDRi) as objective measures of separation accuracy. SI-SNR is defined in equation 15. The reported improvements in tables III to V indicate the additive values over the original mixture. In addition to the distortion metrics, we also evaluated the quality of the separated mixtures using both the perceptual evaluation of subjective quality (PESQ, ) and the mean opinion score (MOS) by asking 40 normal hearing subjects to rate the quality of the separated mixtures. All human testing procedures were approved by the local institutional review board (IRB) at Columbia University in the City of New York.

III-E Comparison with ideal time-frequency masks

Following the common configurations in , the ideal time-frequency masks were calculated using STFT with a 32 ms window size and 8 ms hop size with a Hanning window. The ideal masks include the ideal binary mask (IBM), ideal ratio mask (IRM), and Wiener filter-like mask (WFM), which are defined for source ii as:

IV Results

Figure 2 visualizes all the internal variables of Conv-TasNet for one example mixture sound with two overlapping speakers (denoted by red and blue). The encoder and decoder basis functions are sorted by the similarity of the Euclidean distance of the basis functions found using the unweighted pair group method with arithmetic mean (UPGMA) method . The basis functions show a diversity of frequency and phase tuning. The representation of the encoder is colored according to the power of each speaker at the corresponding basis output at each time point, demonstrating the sparsity of the encoder representation. As can be seen in figure 2, the estimated masks for the two speakers highly resemble their encoder representations, which allows for the suppression of the encoder outputs that correspond to the interfering speaker and the extraction of the target speaker in each mask. The separated waveforms for the two speakers are estimated by the linear decoder, whose basis functions are shown in figure 2. The separated waveforms are shown on the right.

Separation accuracy of different configurations in table III shows that pseudo-inverse autoencoder leads to the worst performance, indicating that an explicit autoencoder configuration does not necessarily improve the separation score in this framework. The performance of all other configurations is comparable. Because linear encoder and decoder with Sigmoid function achieves a slightly better accuracy over other methods, we used this configuration in all the following experiments.

IV-B Optimizing the network parameters

We evaluate the performance of Conv-TasNet on two speaker separation tasks as a function of different network parameters. Table II shows the performance of the systems with different parameters, from which we can conclude the following statements:

Encoder/decoder: Increasing the number of basis signals in the encoder/decoder increases the overcompleteness of the basis signals and improves the performance.

Hyperparameters in the 1-D convolutional blocks: A possible configuration consists of a small bottleneck size BB and a large number of channels in the convolutional blocks HH. This matches the observation in , where the ratio between the convolutional block and the bottleneck H/BH/B was found to be best around 5. Increasing the number of channels in the skip-connection block improves the performance while greatly increases the model size. Therefore, we selected a small skip-connection block as a trade-off between performance and model size.

Number of 1-D convolutional blocks: When the receptive field is the same, deeper networks lead to better performance, possibly due to the increased model capacity.

Size of receptive field: Increasing the size of receptive field leads to better performance, which shows the importance of modeling the temporal dependencies in the speech signal.

Length of each segment: Shorter segment length consistently improves performance. Note that the best system uses a filter length of only 2 ms (Lfs=168000=0.002s\frac{L}{fs}=\frac{16}{8000}=0.002s), which makes it very difficult to train a deep LSTM network with the same LL due to the large number of time steps in the encoder output.

Causality: Using a causal configuration leads to a significant drop in the performance. This drop could be due to the causal convolution and/or the layer normalization operations.

IV-C Comparison of Conv-TasNet with previous methods

We compared the separation accuracy of Conv-TasNet with previous methods using SDRi and SI-SNRi. Table IV compares the performance of Conv-TasNet with other state-of-the-art methods on the same WSJ0-2mix dataset. For all systems, we list the best results that have been reported in the literature. The numbers of parameters in different methods are based on our implementations, except for which is provided by the authors. The missing values in the table are either because the numbers were not reported in the study or because the results were calculated with a different STFT configuration. The previous TasNet in is denoted by the (B)LSTM-TasNet. While the BLSTM-TasNet already outperformed IRM and IBM, the non-causal Conv-TasNet significantly surpasses the performance of all three ideal T-F masks in SI-SNRi and SDRi metrics with a significantly smaller model size comparing with all previous methods.

Table V compares the performance of Conv-TasNet with those of other systems on a three-speaker speech separation task involving the WSJ0-3mix dataset. The non-causal Conv-TasNet system significantly outperforms all previous STFT-based systems in SDRi. While there is no prior result on a causal algorithm for three-speaker separation, the causal Conv-TasNet significantly outperforms even the other two non-causal STFT-based systems . Examples of separated audio for two and three speaker mixtures from both causal and non-causal implementations of Conv-TasNet are available online .

IV-D Subjective and objective quality evaluation of Conv-TasNet

In addition to SDRi and SI-SNRi, we evaluated the subjective and objective quality of the separated speech and compared with three ideal time-frequency magnitude masks. Table VI shows the PESQ score for Conv-TasNet and IRM, IBM, and WFM, where IRM has the highest score for both WSJ0-2mix and WSJ0-3mix dataset. However, since PESQ aims to predict the subjective quality of speech, human quality evaluation can be considered as the ground truth. Therefore, we conducted a psychophysics experiment in which we asked 40 normal hearing subjects to listen and rate the quality of the separated speech sounds. Because of the practical limitations of human psychophysics experiments, we restricted the subjective comparison of Conv-TasNet to the ideal ratio mask (IRM) which has the highest PESQ score among the three ideal masks (table VI). We randomly chose 25 two-speaker mixture sounds from the two-speaker test set (WSJ0-2mix). We avoided a possible selection bias by ensuring that the average PESQ scores for the IRM and Conv-TasNet separated sounds for the selected 25 samples were equal to the average PESQ scores over the entire test set (comparison of tables VI and VII). The length of each utterance was constrained to be within 0.5 standard deviation of the mean of the entire test set. The subjects were asked to rate the quality of the clean utterances, the IRM-separated utterances, and the Conv-TasNet separated utterances on the scale of 1 to 5 (1: bad, 2: poor, 3: fair, 4: good, 5: excellent). A clean utterance was first given as the reference for the highest possible score (i.e. 5). Then the clean, IRM, and Conv-TasNet samples were presented to the subjects in random order. The mean opinion score (MOS) of each of the 25 utterances was then averaged over the 40 subjects.

Figure 3 and table VII show the result of the human subjective quality test, where the MOS for Conv-TasNet is significantly higher than the MOS for the IRM (p<1e−16p<1e-16, t-test). In addition, the superior subjective quality of Conv-TasNet over IRM is consistent across most of the 25 test utterances as shown in figure 3 (C). This observation shows that PESQ consistently underestimates MOS for Conv-TasNet separated utterances, which may be due to the dependence of PESQ on the magnitude spectrogram of speech which could produce lower scores for time-domain approaches.

IV-E Processing speed comparison

Table VIII compares the processing speed of LSTM-TasNet and causal Conv-TasNet. The speed is evaluated as the average processing time for the systems to separate each frame in the mixtures, which we refer to as time per frame (TPF). TPF determines whether a system can be implemented in real time, which requires a TPF that is smaller than the frame length.

For the CPU configuration, we tested the system with one processor on an Intel Core i7-5820K CPU. For the GPU configuration, we preloaded both the systems and the data to a Nvidia Titan Xp GPU. LSTM-TasNet with CPU configuration has a TPF close to its frame length (5 ms), which is only marginally acceptable in applications where only a slower CPU is available. Moreover, the processing in LSTM-TasNet is done sequentially, which means that the processing of each time frame must wait for the completion of the previous time frame, further increasing the total processing time of the entire utterance. Since Conv-TasNet decouples the processing of consecutive frames, the processing of subsequent frames does not have to wait until the completion of the current frame and allows the possibility of parallel computing. This process leads to a TPF that is 5 times smaller than the frame length (2 ms) in our CPU configuration. Therefore, even with slower CPUs, Conv-TasNet can still perform real-time separation.

IV-F Sensitivity of LSTM-TasNet to the mixture starting point

Unlike language processing tasks where sentences have determined starting words, it is difficult to define a general starting sample or frame for speech separation and enhancement tasks. A robust audio processing system should therefore be insensitive to the starting point of the mixture. However, we empirically found that the performance of the causal LSTM-TasNet is very sensitive to the exact starting point of the mixture, which means that shifting the input mixture by several samples may adversely affect the separation accuracy. We systematically examined the robustness of LSTM-TasNet and causal Conv-TasNet to the starting point of the mixture by evaluating the separation accuracy for each mixture in the WSJ0-2mix test set with different sample shifts of the input. A shift of ss samples corresponds to starting the separation at sample ss instead of the first sample. Figure 4 (A) shows the performance of both systems on the same example mixture with different values of input shift. We observe that, unlike LSTM-TasNet, the causal Conv-TasNet performs consistently well for all shift values of the input mixture. We further tested the overall robustness for the entire test set by calculating the standard deviation of SDRi in each mixture with shifted mixture inputs similar to figure 4 (A). The box plots of all the mixtures in the WSJ0-2mix test set in figure 4 (B) show that causal Conv-TasNet performs consistently better across the entire test set, which confirms the robustness of Conv-TasNet to variations in the starting point of the mixture. One explanation for this inconsistency may be due to the sequential processing constraint in LSTM-TasNet which means that failures in previous frames can accumulate and affect the separation performance in all following frames, while the decoupled processing of consecutive frames in Conv-TasNet alleviates the effect of occasional error.

IV-G Properties of the basis functions

V Discussion

In this paper, we introduced the fully-convolutional time-domain audio separation network (Conv-TasNet), a deep learning framework for time-domain speech separation. This framework addresses the shortcomings of speech separation in the STFT domain, including the decoupling of phase and magnitude, the suboptimal representation of the mixture audio for separation, and the high latency of calculating the STFT. The improvements are accomplished by replacing the STFT with a convolutional encoder-decoder architecture. The separation in Conv-TasNet is done using a temporal convolutional network (TCN) architecture together with a depthwise separable convolution operation to address the challenges of deep LSTM networks. Our evaluations showed that Conv-TasNet significantly outperforms STFT speech separation systems even when the ideal time-frequency masks for the target speakers are used. In addition, Conv-TasNet has a smaller model size and a shorter minimum latency, which makes it suitable for low-resource, low latency applications.

Unlike STFT which has a well-defined inverse transform that can perfectly reconstruct the input, best performance in the proposed model is achieved by an overcomplete linear convolutional encoder-decoder framework without guaranteeing the perfect reconstruction of the input. This observation motivates rethinking of autoencoder and overcompleteness in the source separation problem which may share similarities to the studies of overcomplete dictionary and sparse coding . Moreover, the analysis of the encoder/decoder basis functions in section IV-G revealed two interesting properties. First, most of the filters are tuned to low acoustic frequencies (more than 60% tuned to frequencies below 1 kHz). This pattern of frequency representation, which we found using a data-driven method, roughly resembles the well-known mel-frequency scale as well as the tonotopic organization of the frequencies in the mammalian auditory system . In addition, the overexpression of lower frequencies may indicate the importance of accurate pitch tracking in speech separation, similar to what has been reported in human multitalker perception studies . In addition, we found that filters with the same frequency tuning explicitly express various phase information. In contrast, this information is implicit in the STFT operations, where the real and imaginary parts only represent symmetric (cosine) and asymmetric (sine) phases, respectively. This explicit encoding of signal phase values may be the key reason for the superior performance of TasNet over the STFT-based separation methods.

The combination of high accuracy, short latency, and small model size makes Conv-TasNet a suitable choice for both offline and real-time, low-latency speech processing applications such as embedded systems and wearable hearing and telecommunication devices. Conv-TasNet can also serve as a front-end module for tandem systems in other audio processing tasks, such as multitalker speech recognition and speaker identification . On the other hand, several limitations of Conv-TasNet must be addressed before it can be actualized, including the long-term tracking of speakers and generalization to noisy and reverberant environments. Because Conv-TasNet uses a fixed temporal context length, the long-term tracking of an individual speaker may fail, particularly when there is a long pause in the mixture audio. In addition, the generalization of Conv-TasNet to noisy and reverberant conditions must be further tested , as time-domain approaches are more prone to temporal distortions which are particularly severe in reverberant acoustic environments. In such conditions, extending the Conv-TasNet framework to incorporate multiple input audio channels may prove advantageous when more than one microphone is available. Previous studies have shown the benefit of extending speech separation to multichannel inputs , particularly in adverse acoustic conditions and when the number of interfering speakers is large (e.g., more than 3).

In summary, Conv-TasNet represents a significant step toward the realization of speech separation algorithms and opens many future research directions that would further improve its accuracy, speed, and computational cost, which could eventually make automatic speech separation a common and necessary feature of every speech processing technology designed for real-world applications.

VI Acknowledgments

This work was funded by a grant from the National Institute of Health, NIDCD, DC014279; a National Science Foundation CAREER Award; and the Pew Charitable Trusts.

References