Speaker-independent Speech Separation with Deep Attractor Network

Yi Luo, Zhuo Chen, Nima Mesgarani

I Introduction

Listening to an individual in crowded situations often takes place in the presence of interfering speakers. Such situations require the ability to separate the voice of a particular speaker from the mixed audio signal of others. Several proposed systems have shown significant performance improvement on the separation task when prior information of speakers in a mixture is given . This however is still challenging when no prior information about the speakers is available, a problem known as speaker-independent speech separation. Humans are particularly adept at this task, even in the absence of any spatial separation between speakers . This effortless task for humans, however, has proven difficult to model and emulate algorithmically. Nevertheless, it is a challenge that must be solved in order to achieve robust performance in speech processing tasks. For example, while the performance of current automatic speech recognition (ASR) systems has reached that of humans in clean conditions , these systems are still unable to perform well in noisy and crowded environments, lacking robustness when interfering speakers are present. This becomes even more challenging when separating all sources in a mixture is required, such as in meeting transcription and music separation. When signals from multiple microphones are available, beamforming algorithms can be used to improve the target-to-masker ratio ; when only one microphone is available, however, the general problem of audio separation remains largely unresolved.

Prior to the emergence of deep learning, three main categories of algorithms were proposed to solve the speech separation problem: statistical methods, clustering methods, and factorization methods, with focus on different target tasks. In statistical methods, the target speech signal is modeled with probability distributions such as complex Gaussian or methods such as independent component analysis (ICA) , where the interference signal is assumed to be statistically independent from the target speech. Maximum likelihood estimation method is typically applied based on the known statistical distributions of the target. In clustering methods, the characteristics of the target speaker, such as pitch and signal continuity are estimated from the observation and used to separate the target signal from other sources in the mixture. Methods such as computational auditory scene analysis (CASA) and spectral clustering fall into this category . Factorization models, such as non-negative matrix factorization (NMF) , formulate the separation problem as a matrix factorization problem in which the time-frequency (T-F) representation of the mixture is factorized into a combination of basis signals and activations. The activations learned for each basis signal are then used to reconstruct the target sources.

In recent years, deep learning has made important progress in audio source separation. Specifically, neural networks have been successfully applied in speech enhancement and separation and music separation with significantly better performance than that of traditional methods. A typical paradigm for neural networks is to directly estimate T-F masks of the sources given the T-F representation of the audio mixture (such as noisy speech or multiple speakers) . This formulates the separation as a supervised single-class or multi-class regression problem. Different types of masks and objective functions have been proposed. For instance, phase-aware masks for enhancement and separation have been studied in .

Limitations of the previous neural networks become evident when one considers the problem of separating two simultaneous speakers with no prior knowledge of the speakers (speaker-independent scenario). Two main challenges in this situation are the so-called permutation problem and the output dimension mismatch problem. Permutation problem refers to the observation that the order of the speakers in the target may not be the same as the order of the speakers in the output of the network. For example, when designing the target output by separating speakers S1S_{1} and S2S_{2}, both (S1,S2)(S_{1},S_{2}) and (S2,S1)(S_{2},S_{1}) are acceptable permutations of the speakers. Once the target permutation is fixed, however, the output of the network must follow the permutation of the target. In situations where the separation is successful but the outputs have incorrect permutation compared with the targets, the output error will be large, causing the network to diverge from the correct solution. Aside from this issue, the output dimension mismatch problem is also a notable problem. Since the number of speakers in a mixture can vary, a neural network with a fixed number of output targets does not have the flexibility to separate the mixtures where the number of speakers is not equal to the number of output targets.

Two deep learning methods, deep clustering (DPCL) and permutation invariant training (PIT) , have been proposed recently to resolve these problems. In deep clustering, a network is trained to generate a discriminative embedding for each T-F bin so that the embeddings of the T-F bins that belong to the same speaker are closer to each other. Because deep clustering uses the Frobenius norm between the affinity matrix of embeddings and the affinity matrix of ideal speaker assignment (e.g. ideal binary mask) as the training objective, it solves the permutation problem due to the permutation-invariant property of affinity matrices. The mask estimation process is done by applying clustering algorithms such as K-means or spectral clustering to the embeddings, while the assignment of the embeddings to the clusters forms the final estimation. Hence, the number of outputs is only determined by the number of target clusters. While DPCL is able to solve both the permutation and output dimension mismatch problems and produce a state-of-the-art performance, it is unable to use reconstruction error as the target for optimization. This is because the mask generation is done through a post-clustering step on the embeddings, which is done separately from the network network. In more recent DPCL work, minimizing the separation error is processed with an unfolded soft clustering subsystem for direct mask generation, and an additional mask enhancement network is followed for better performance . The PIT algorithm solves the permutation problem by first calculating the training objective loss for all possible permutations for CC mixing sources (C!C! permutations), and then using the permutation with the lowest error to update the network. It solves the output dimension problem by assuming a maximum number of sources in the mixture and using null output targets (very low energy Gaussian noises) as auxiliary targets when the actual number of sources in the mixture is smaller than the number of outputs in the network. PIT was proposed in in a frame-wise fashion and was later shown to have comparable performance to DPCL with a deep LSTM network structure.

We address the general source separation problem with a novel deep learning framework which we call the ‘attractor network’. The term “attractor” refers to the well-studied effects in human speech perception which suggest that biological neural networks create perceptual attractors (magnets). These attractors warp the acoustic feature space to draw in the sounds that are close to them, a phenomenon that is called the Perceptual Magnet Effect . Our proposed model works on the same principle as DPCL by first generating a high-dimensional embedding for each T-F bin. We then form a reference point (attractor) for each source in the embedding space that pulls all the T-F bins belonging to that source toward itself. This results in the separation of sources in the embedding space. A mask is estimated for each source in the mixture using the similarity between the embeddings and each attractor. Since the correct permutation of the masks is directly related to the permutation of the attractors, our network can potentially be extended to an arbitrary number of sources without the permutation problem once the order of attractors is established. Moreover, with a set of auxiliary points in the embedding space, known as the anchor points, our framework can directly estimate the masks for each source without needing a post-clustering step as in or a clustering subnetwork as in . This aspect creates a system that directly generates the reconstructed spectrograms of the sources in both training and test phases.

The rest of the paper is organized as follows. In Section II, we introduce the general problem of source separation and the embedding learning method. In Section III, we describe the original deep attractor network proposed in . In Section IV, we propose several variants and extensions to the original deep attractor network for better performance. In Section V, we evaluate the performance of the attractor network and analyze the properties of the embedding space.

II Source separation and embedding learning

We start with defining the general problem of single-channel speech separation, and describe how the method of embedding learning can be used to solve this problem.

The problem of single-channel speech separation is defined as estimating all the CC speaker sources s1(t),…,sc(t)s_{1}(t),\ldots,s_{c}(t) given the mixture waveform signal x(t)x(t)

In time-frequency (T-F) domain, the complex short-time Fourier transform (STFT) spectrogram of the mixture, X(f,t)\mathcal{X}(f,t) equals to the sum of the complex STFT spectrograms of all the sources

II-B Source separation in embedding space

High-dimensional embedding is a powerful and commonly used method in many tasks such as natural language processing and manifold learning . This technique maps the signal into a high-dimensional space where the resulting representation has desired properties. For example, word embedding is currently one of the standard tools to extract the relationship and connection between different words, and serves as a front-end to more complex tasks such as machine translation and dialogue systems . In the problem of source separation in T-F domain, a high-dimensional embedding for each T-F bin is found and speech separation is formulated as a source segmentation problem in the embedding space .

The embeddings can be either knowledge-based or data-driven. CASA is a popular frameworks for using specially designed features to represent the sources , where different types of acoustic features are concatenated to produce a high-dimensional embedding which represents the sound source in different time-frequency coordinates. An example of the data driven embedding approach for speech separation is the recently proposed deep clustering method (DPCL) . DPCL uses a neural network model to learn embeddings of T-F bins such that to minimize the in-class similarity, while at the same time maximizes the between-class similarity. For the embeddings whose corresponding T-F bins belong to the same speaker, the similarity between them should be large and vice versa. This method therefore creates a high-dimensional representation of the mixture audio that results in better segmentation and separation of speakers.

III Deep Attractor Network

In this section, we introduce the deep attractor network (DANet) and compare it to DPCL and PIT . We then discuss three methods for estimating the parameters of the network during test phase and discuss their pros and cons. We further address the limitation of DANet by presenting a new framework called ADANet. Figure 1 shows the flowchart of the overall system. Note that in this section we use the same notations as in Section II.

Following the main concept of embedding learning with neural network, DANet generates a KK-dimensional embedding vector for each of the T-F bins in the mixture magnitude spectrogram

The attractors are then estimated as follows:

After the generation of the attractors, DANet calculates the similarity between the embedding of each T-F bin and each attractor:

The neural network is then trained by minimizing a standard L2L^{2} reconstruction error as the objective function

III-B Relation to DPCL and PIT

Since DANet shares the same network structure as DPCL , it is important to illustrate the difference between DANet and DPCL. In contrast to DPCL, DANet directly optimizes the reconstruction error with a computationally simpler objective function rather than the calculation of affinity matrices in DPCL. Moreover, the direct mask estimation allows it to use flexible similarity measurements and target masks, including phase-aware and phase-sensitive mask .

On the other hand, when the attractors are considered as the trained weights of the network instead of dynamically formed by the embeddings (Eqn. 7 & 9), DANet reduces to a classification network and equation 10 is equivalent to a linear fully-connected layer. In this case, permutation invariant training becomes necessary since the masks are no longer linked to a speaker. However, the dynamic formation of the attractors in DANet allows utterance-level flexibility in the generation of the embeddings. Moreover, DANet does not assume a fixed number of outputs, since the number of attractors is decided by the size of the speaker assignment function YY during training.

IV Estimation of the attractor points

As described in equations 7 & 9, the actual speaker assignment (e.g. using the IBM or IRM methods, Eqn. 5) is necessary to form the attractors during the training phase. However, this information is not available during the test phase, causing a mismatch between training and test phases. In this section, we propose several methods for estimating the location of the attractors during the test phase. The first two methods were introduced in . Here we discuss their limitation and propose a new method for estimating the attractor points called Anchored DANet (ADANet), an extension that enables direct attractor estimation and mask generation for both training and test phases.

The simplest method to form the attractors in the embedding space is to use an unsupervised clustering algorithm such as K-means on the embeddings to determine the speaker assignment (DANet-Kmeans in Fig. 4 and Section V). This method is similar to . In this case, the centers of the clusters are treated as the attractors for mask generation. Figure 2 shows an example of this method, where the crosses represent the centers of the two clusters, which are also the estimated attractors.

IV-B Fixed attractor points

While there is no direct constraint on the location of the attractors in the embedding space, we have found empirically that the location of the attractor points are relatively constant across different mixtures. Figure 5 shows the location of attractors for 10000 different mixtures, with each opposite pair of dots corresponding to the attractors for the two speakers in a given mixture. Two pairs of attractors (marked as A1 and A2) are automatically discovered by the network. Based on this observation, we propose to first estimate all the attractors in the training phase for different mixtures, and subsequently use the mean of those attractors as the fixed attractors during the test phase (DANet-Fixed in Fig. 4 and Section V). The advantage of using fixed attractors is that it removes the need for the clustering step, allowing the system to directly estimate the mask for each time frame and enabling real-time implementation.

IV-C Anchored DANet (ADANet)

While both clustering and fixed attractor method can be used during the test phase, these approaches have several limitations. For clustering-based estimation, the K-means step increases the computational cost of the system and therefore increases the run-time delay. Additionally, the centers of the clusters are not guaranteed to match the true locations of the attractors as the (weighted) averages of the corresponding embeddings. This potential difference between the true and estimated attractors causes a mismatch in the mask formation between the training and test phases. One such instance is shown in figure 3, where the embedding space is visualized using its first two principle components. The locations of true and estimated attractors are plotted in yellow and black, and the distance between the two shows the mismatch between the true and estimated attractor locations. As suggested by equation 11, this mismatch changes the mask estimated for each source and thus reduces the accuracy of the separation. We refer to this problem as the center mismatch problem, which is caused by the unknown speaker assignment during the test phase. Using fixed attractors on the other hand relies on the assumption that the location of the attractors in training and test phases are similar. This, however, may not be the case if the test condition is significantly different from the training condition. Another drawback of fixing the number and location of attractors is the lack of flexibility when dealing with mixtures with variable number of speakers.

The entire process of ADANet can be represented as a generalized Expectation-Maximization (EM) framework, where the speaker assignment calculation is the “Expectation step” and the following attractor formation step is the “Maximization step.” Hence, the separation procedure of ADANet can be viewed as a single EM iteration. Note that this method is similar to the unfolded soft-clustering subnetwork proposed in but the parameters (i.e. statistics) for the clusters here are dynamically defined by the embeddings. Thus, they do not increase the total number of parameters in the network. Moreover, this approach also allows utterance-level flexibility in the estimation of the clusters.

IV-D Attractor formation summary

Compared to the clustering and fixed attractor methods, ADANet eliminates the need for true speaker assignment during both training and test phases. The center mismatch problem no longer exists because the attractors are determined by the anchor points, which are trained in the training phase and fixed in the test phase. The mask generation process is therefore matched during both training and test phases. Figure 4 shows the difference between various ways of attractor calculation. The fixed attractor method enables real-time processing and generates the masks without the extra clustering step, but is sensitive to the mismatch between training and test conditions. ADANet framework solves the center mismatch problem and allows direct mask generation in both training and test phase, however it increases the computational complexity since all (NC)\binom{N}{C} subsets of anchors require one EM iteration. Moreover, since the correct permutation of the anchors is unknown, permutation invariant training in is required for training the ADANet.

V Experiments and analysis

We evaluate our proposed model on the task of single-channel two and three speaker separation. Example sounds can be found here .

We use the WSJ0-2mix and WSJ0-3mix datasets introduced in which contains two 30 h training sets and a 10 h validation sets for the two tasks. These tasks are generated by randomly selecting utterances from different speakers in the Wall Street Journal (WSJ0) training set si_tr_s and mixing them at various signal-to-noise ratios (SNR) randomly chosen between 0 dB and 5 dB. Two 5 h evaluation sets are generated in the same way, using utterances from 16 unseen speakers from si_dt_05 and si_et_05 in the WSJ0 dataset. All data are resampled to 8 kHz to reduce computational and memory costs. The log magnitude spectrogram serves as the input feature, computed using short-time Fourier transform (STFT) with 32 ms window length (256 samples), 8 ms hop size (64 samples), and the square root of Hanning window.

V-B Evaluation metrics

We evaluated the separation performance on the test sets using three metrics: signal-to-distortion ratio (SDR) for comparing with PIT models in , scale-invariant signal-to-noise ratio (SI-SNR) to compare with DPCL models in , and PESQ score for evaluation of the speech quality. The SI-SNR, proposed in , is defined as:

V-C Network architecture

The network contains 4 Bi-directional LSTM layers with 600 hidden units in each layer. The embedding dimension is set to 20 according to , resulting in a fully-connected feed-forward layer of 2580 hidden units (20 ×\times 129 as K ×\times F) after the BLSTM layers. Adam algorithm is used for training, with the learning rate starting at 1e−31e^{-3} and then halved if no best validation model is found in 3 epochs. The total number of epochs is set to 100, and we used the cost function in equation 13 on the validation set for early stopping. The criterion for early stopping is no decrease in the loss function on validation set for 10 epochs. WFM (Eqn. 5) is used as the training target.

We split the input features into non-overlapping chunks of 100-frame and 400-frame length as the input to the network with a curriculum training strategy . We first train the network with 100-frame length input until converged, and then continue training the network with 400-frame long segments. The initial learning rate for 400-frame length training is set to be 1e−41e^{-4} with the same learning rate adjustment and early stopping strategies. For comparison, we trained the networks with and without curriculum training. We also studied the effect of dropout during the training of the network by which was added with probability of 0.5 to the input to each of the BLSTM layers. During test phase, the number of speakers is provided to both DANet and ADANet in all the experiments except for the mixed number of speaker experiment shown in Table V.

V-D Results

The results in table I show that using the %90 threshold leads to better performance compared to no threshold, which indicates the importance of accurate estimation of attractors for mask generation. Moreover, we observe that although Softmax with ideal speaker assignment leads to a higher performance than Sigmoid, the performance with K-means for Softmax networks are worse than Sigmoid networks. Our results also show that the K-means performance in the Softmax network is highly dependent on how the network is optimized (i.e. different network trained with different initial values may have very different performance) with a similar level of validation error. This observation is supported by our assumption that Softmax is less sensitive to the distance when the embedding dimension NN is high. As a result, adding a training regularization such as dropout does not guarantee improved performance for K-means clustering, which is also a disadvantage of a network architecture that does not directly optimize the mask generation.

Table II shows the effect of changing the number of anchors in ADANet with Softmax for mask generation. All the networks hereafter apply a 90% threshold for estimation of attractors. As the number of anchors increases, the performance of the network consistently improves. It confirms that the increased flexibility in the choice of the anchors help the final separation performance. Unlike Softmax DANet with K-means, adding dropout in BLSTM layers here always increases the performance. This shows the advantage of the framework without using the post-clustering step.

Tables III & IV compare our method with other speaker-independent techniques in separation of two and three speaker mixtures. DPCL is the original deep clustering model with K-means clustering for mask generation, and DPCL++ is an extension of deep clustering that directly generates the masks with soft-clustering layers, and further improve the masks with a second stage enhancement network. uPIT-BLSTM is the utterance level PIT model that applies PIT on the entire utterance using deep BLSTM network, and uPIT-BLSTM-ST contains a tandem second-stage mask enhancement network for performance improvement. The DPCL and DPCL++ methods also applied a similar curriculum training strategy, while the uPIT-BLSTM and uPIT-BLSTM-ST methods only used standard training. In both tables, we categorize the methods into single-stage and two-stage systems for better comparison, where the two-stage systems have a second-stage enhancement network. The ADANet system in this comparison utilizes six anchors.

In WSJ0-2mix test set (Table III), the K-means DANet outperforms all the previous one-stage systems and the two-stage uPIT-BLSTM-ST system. The ADANet with six anchors has the best performance among all one-stage systems, and is only slightly worse than the two-stage DPCL++ system. In WSJ0-3mix test set (Table IV), both K-means DANet and six-anchor ADANet outperform all previous systems with either one-stage or two-stage configuration, and ADANet also performs significantly better than K-means DANet. The PESQ scores for DANet and ADANet in WSJ0-2mix show significant improvement upon mixture, while in WSJ0-3mix the gaps between WFM and the models are larger. This is expected since the speaker with lowest energy is harder to separate than the two speaker cases, which may lead to a lower PESQ score.

Table V shows the results in the speaker separation task where the same network is used to separate both two and three speaker mixtures. The networks are first trained on three-speaker mixtures until convergence, and then continued training using both two- and three-speaker mixtures. For K-means DANet, the information about the number of speakers is given during both training and test phases, and for ADANet, we use a training strategy similar to , where the network always uses three anchors to generate three masks, and an auxiliary output mask of all zero entries is concatenated to the target masks in the two-speaker cases. During test phase, no information about the number of speakers is provided, and a simple energy-based detector is used to detect the number of speakers. If the power in an output is 20 dB less than the other outputs, it is subsequently discarded.

Table V shows that K-means DANet performs well on the two-speaker separation task, but has worse performance on three-speaker separation. This may be due to the less separated embeddings in two-speaker cases for Softmax function discussed in Table I. However, with the network configuration in ADANet, the performance is significantly better than both one-stage and two-stage PIT systems. Moreover, ADANet successfully detects the correct number of outputs in all the 3000 utterances in the WSJ0-2mix test set, showing that appending a zero-mask to the target masks of two speaker mixtures during training enables the network to learn a silent mask under two-speaker cases without sacrificing the performance in three-speaker separation. This also indicates that ADANet can automatically select proper anchors for speaker assignment estimation in mixtures with different number of sources (Fig. 6).

VI Conclusion

In this paper, we introduce the deep attractor network (DANet) for single-microphone speech separation. We discussed its advantages and drawbacks, and proposed an extension, Anchored DANet (ADANet) for better and more stable performance. DANet extends the deep clustering framework by creating attractor points in the embedding space which pull together the embeddings that belong to a specific speaker. Each attractor in this space is used to form a time-frequency mask for each speaker in the mixture. Directly minimizing the reconstruction error allows better optimization of the embeddings. We explored the mismatch problem of DANet during training and test phase, which is resolved in ADANet to allow direct mask generation in both training and test phases. This removes the post-clustering step for mask estimation. The ADANet approach provides a flexible solution and a generalized Expectation-Maximization strategy to deterministically locate the attractors from the estimated speaker assignment. This approach maintains utterance-level flexibility in attractor formation, and can also generalize to changing signal conditions. Experimental results showed that compared with previous state-of-art systems, both DANet and ADANet have comparable or better performance in both two- and three-speaker separation tasks.

The future works include exploring ways to incorporate speaker information in the estimation of anchor points and in the generation of the embeddings. These aspects can result in both speaker-dependent and speaker-independent separation by the same network. Separating more sources beyond speech signals by creating a larger and more informative embedding space is also an interesting and important topic, which prompts the possibility of designing a universal acoustic source separation framework.

VII Acknowledgement

The authors would like to thank Drs. John Hershey and Jonathan Le Roux of Mitsubishi Electric Research Lab, and Dr. Dong Yu of Tencent AI Lab for constructive discussions. Yi Luo and Zhuo Chen contributed equally to this work. This work was funded by a grant from National Institute of Health, NIDCD, DC014279, National Science Foundation CAREER Award, and the Pew Charitable Trusts.

References