Phasebook and Friends: Leveraging Discrete Representations for Source Separation

Jonathan Le Roux, Gordon Wichern, Shinji Watanabe, Andy Sarroff, John R. Hershey

I Introduction

The field of speech separation and speech enhancement has witnessed dramatic improvements in performance with the recent advent of deep learning-based techniques . Most of these algorithms rely on the estimation of some sort of time-frequency (T-F) mask to be applied to the time-frequency representation of an input mixture signal, the estimated signal then being resynthesized using some inverse transform. Let us denote by X=(xt,f)\bm{{X}}=(x_{t,f}), S=(st,f)\bm{{S}}=(s_{t,f}), and N=(nt,f)\bm{{N}}=(n_{t,f}) the complex-valued time-frequency representations of a mixture signal, a target source signal, and an interference signal, respectively, where tt denotes the time frame index and ff the frequency bin index. We also denote by θt,f=∠(st,f/xt,f)\theta_{t,f}=\angle(s_{t,f}/x_{t,f}) the phase difference between the mixture and the target source. The time-frequency representation is here typically taken to be the short-term Fourier transform (STFT), such that xt,f=st,f+nt,fx_{t,f}=s_{t,f}+n_{t,f}. The goal of speech enhancement or separation can be formulated as that of recovering an estimate S^=(s^t,f)\hat{\bm{{S}}}=(\hat{s}_{t,f}) of S\bm{{S}} from X\bm{{X}}, and we’re interested in particular in algorithms that do so by estimating a mask C=(ct,f)\bm{{C}}=(c_{t,f}) such that s^t,f=ct,fxt,f\hat{s}_{t,f}=c_{t,f}x_{t,f}. Note that the interference signal itself could also be a separate target, such as in the case of speaker separation.

In most cases, these time-frequency masks are real-valued, which means that they only modify the magnitude of the mixture in order to recover the target signal. Their values are also typically constrained to lie between 0 and 1, both for simplicity and because this was found to work well under the assumption that only the magnitude is modified, retaining the mixture phase for resynthesis.

Several reasons can be cited for focusing on modifying only the magnitude: the noisy phase is actually the minimum mean-squared error (MMSE) estimate under some simplistic statistical independence assumptions (which typically do not hold in practice); combining the noisy phase with a good estimate of the magnitude is straightforward and gives somewhat satisfactory results; until recently, getting a good estimate of the magnitude was already difficult enough such that optimizing the phase estimate was not a priority, or to put it in other words, phase was not the limiting factor in performance; estimating the phase of the target signal is believed to be a hard problem.

With the advent of recent deep learning algorithms, the quality of the magnitude estimates has improved significantly, to the point that the noisy phase has now become a limiting factor to the overall performance. Because the noisy phase is typically inconsistent with the estimated magnitude , the reconstructed time-domain signal has a different magnitude spectrogram from the intended, estimated one. As an added drawback, further improving the magnitude estimate by making it closer to the true target magnitude may actually lead to worse results when pairing it with the noisy phase, in terms of performance measures such as signal to noise ratio (SNR). Indeed, if the noisy phase is incorrect and for example opposite to the true phase, using as the estimate for the magnitude is a “better” choice than using the correct magnitude value, which may point far away in the wrong direction. Using the noisy phase is thus not only sub-optimal as a phase estimate, it likely also forces the magnitude estimation algorithms to limit their accuracy with respect to the true magnitude.

We have already started exploring this scheme for the magnitude, with the introduction of a convex softmax activation function which interpolates between the values 0,1,20,1,2 to obtain a continuous representation of the interval $$ as the target interval for the magnitude mask . We showed that this activation function led to significantly better performance when optimizing for best reconstruction after a phase reconstruction algorithm. This intuitively makes sense, because the reconstructed phase used to obtain the final time-domain signal is likely to better exploit a magnitude estimate more faithful to the clean magnitude, in particular at time-frequency bins where the clean magnitude is larger than the mixture magnitude due to cancelling interference.

We propose here a generalization of this idea of relying on discrete values to build representations for the masks. We extend the concept of convex softmax activation for the magnitude to the combination of a magnitude codebook, or magbook, with a softmax layer to build various magnitude representations, either discrete or continuous. Similarly, we propose to combine a phase codebook, or phasebook, with a softmax layer to build various phase representations, again either discrete or continuous. Finally, we propose an alternate representation which foregoes the factorization between magnitude and phase and combines a complex codebook, or combook, with a softmax layer to build various complex mask representations. These representations are flexible and can be incorporated within optimization frameworks that are regression-based, classification-based, or a combination of both.

Related works: This paper’s contributions are at the intersection of multiple directions of research: classification-based separation, discrete phase representations, complex mask estimation, phase-difference modelling, and phase reconstruction. The idea of considering separation as a classification problem was explored first using shallow methods, in particular support vector machines , and later deep neural networks , and was arguably at the onset of the deep learning revolution in this field. A few works have proposed to consider discrete representations of the phase for source separation, such as and , in both cases within a generative model based on mixtures of Gaussians. Some works have attempted to incorporate phase modeling for deep-learning-based source separation, in particular with the so-called complex ratio mask , which does consider ranges of values that are not limited to $$. While the complex ratio mask used a continuous real-imaginary representation, we here focus mainly on discrete representations involving a magnitude-phase factorization or a direct modelling of the complex value (with the real and imaginary parts considered jointly). We also model not the clean phase but a phase mask, that is, a phase difference between the mixture and the clean source, or in other words a correction to be applied to the mixture phase to get closer to the clean phase. Estimating the phase difference was recently considered within an audio-visual separation framework in , where it is reconstructed using a convolutional network that takes the estimated magnitude and the noisy phase as input. Another, potentially complementary, way to improve the phase is to use phase reconstruction. Recent works from our team applied phase reconstruction at the output of a good magnitude estimation network , then trained through an unfolded iterative phase reconstruction algorithm . We finally trained the time-frequency representations used within the phase reconstruction algorithm themselves , which is the current state-of-the-art in methods relying on time-frequency representations.

As we were finalizing this article, two related works worth mentioning were published. First, a deep-learning-based source separation algorithm, referred to as PhaseNet , attempts to estimate discretized values of the target source phase; the discretized values are fixed to a uniform quantization along the unit circle, and the network is trained using cross-entropy. As it will become clear in this article, apart from the fact that PhaseNet attempts to estimate the target phase instead of the phase difference, the representation used corresponds to a particular setup of our framework, with a fixed uniform phasebook, cross-entropy training, and argmax based inference. Our framework allows for much more variety in both training and inference schemes, in particular allowing fully end-to-end training which is cumbersome with argmax inference. Second, an updated version of the TasNet algorithm just established a new state-of-the-art on the wsj0-2mix dataset, surpassing our previous numbers as well as those presented in this article. The TasNet article introduced several interesting techniques that could be adopted in our framework, such as the use of convolution layers instead of recurrent ones, layer normalization schemes, and the use of SI-SDR as the objective instead of the L1L^{1} waveform approximation loss that we consider. It is unclear how much these techniques would influence the performance of TasNet’s competing methods, and we shall consider incorporating them in our framework as future work.

II Designing masks based on discrete representations

We propose to rely on discrete values to build representations for a complex ratio mask, either via its factorization into magnitude and phase components or directly as a complex value. In particular, we propose to model the magnitude mask using a combination of a magnitude codebook, or magbook, with a softmax layer, and to model the phase mask (i.e., the correction term between mixture phase and clean phase) using a combination of a phase codebook, or phasebook, with a softmax layer. Alternatively, we consider modelling the complex ratio mask directly using a combination of a complex codebook, or combook, with a softmax layer; magnitude and phase are then modelled jointly.

We consider scalar codebooks MM={m(1),…,m(M)}\mathcal{M}_{M}=\{m^{(1)},\dots,m^{(M)}\} for the magnitude mask, FP={θ(1),…,θ(P)}\mathcal{F}_{P}=\{\theta^{(1)},\dots,\theta^{(P)}\} for the phase mask, and CC={c(1),…,c(C)}\mathcal{C}_{C}=\{c^{(1)},\dots,c^{(C)}\} for the complex mask. At each time-frequency bin t,ft,f, a network can estimate softmax probability vectors for the magnitude mask, the phase mask, or the complex mask, denoted by

sample from the softmax distribution (“sampling”):

compute the expected value over the distribution (“interpolation”):

Note that the interpolation for the phase in Eq. (11) is performed in the complex domain and that taking the angle implies a renormalization step; this interpolation is illustrated in Fig. 3. Further note that the interpolation scheme for the magnitude is an extension of the classical sigmoid activation function for the case of a fixed magbook of size 22 with elements {0,1}\{0,1\} (referred to here as uniform magbook 2), and an extension of the convex softmax considered in for the case of a fixed magbook of size 33 with elements {0,1,2}\{0,1,2\} (referred to here as uniform magbook 3).

In the following, we shall call “phasebook layer” a layer computing phase values based on the outputs of a softmax layer and a phasebook via a method such as those above, and similarly for a “magbook layer” and a “combook layer”.

III Phasebook with argmax

To get an idea of the potential benefits of a better phase modeling, we first consider the argmax scheme for the phase mask, in which the system attempts to select the best codebook value at each T-F bin.

Given a phasebook FP={θ(1),…,θ(P)}\mathcal{F}_{P}=\{\theta^{(1)},\dots,\theta^{(P)}\}, the goal of our system is to estimate at each T-F bin (t,f)(t,f) the codebook index jt,fj_{t,f} such that:

where mt,fm_{t,f} is some estimate for the magnitude of the mask. The estimation is in fact independent of the magnitude mask value:

An important question is how to best design the phasebook. An obvious and easy choice is to use regularly spaced values. But ideally, one would like to optimize them for best performance on some training data. This can be done independently of the classification system, or together with it, optimizing both the phasebook and the classification system jointly in an end-to-end fashion. We first consider how to optimize the codebook offline in a pre-training step, for optimal performance given a magnitude estimate. That magnitude estimate may be obtained either with a pre-trained magnitude estimation network, or with an oracle mask.

The objective function for the phasebook training is:

It can be optimized using an EM-like algorithm. In the E-step, the optimal codebook assignments are computed for each T-F bin according to Eq. 14. In the M-step, we update the phasebook to further decrease the objective function by solving

which can easily be shown to be equivalent to

leading to the following update equation:

Note that a magbook could be similarly (and jointly) optimized under an argmax scheme, at each step looping in order over the updates of the magbook values, the magbook assignments, the phasebook assignments, and the phasebook values, the latter two as described above. Finally, optimization of a combook under an argmax scheme can be simply obtained via the k-means algorithm.

In our experiments, we optimize the codebooks on a speech separation task using 50 randomly selected utterances from the wsj0-2mix training dataset . Note that we noticed similar behaviors in terms of optimized codebook configurations and separation performance on a speech enhancement task with data from the CHiME2 training set . The initial codebooks can be randomly sampled from the data, or set manually. In the latter case, the phasebooks are initialized using uniform codebooks with values that partition the unit circle into equal angular intervals, making sure that is one of the elements of the codebook: FPuniform={0,…,2pπP,…,2(P−1)πP}\mathcal{F}_{P}^{\text{uniform}}=\{0,\dots,\frac{2p\pi}{P},\dots,\frac{2(P-1)\pi}{P}\}. We run the optimization algorithm for 4040 epochs, which was enough to ensure convergence. It is likely that the output of the optimization is only a local optimum, and even better codebooks could potentially be obtained by running multiple optimizations with different initializations, but we did not consider this here.

Figure 4 shows the optimized phasebooks for P=2,…,10P=2,\dots,10 and a magnitude obtained using an oracle IAM magnitude mask, together with the uniform phasebooks they were initialized from.

III-B Oracle performance

Fig. 6 shows results with uniform and optimized phasebooks for truncated ideal amplitude masks. Optimizing the codebooks leads in all cases to significant improvements, with typical gains around 2 to 3 dB.

IV Objective functions

We consider the above representations as layers within a deep learning model for source separation, and we need to optimize the parameters ϕ\bm{\phi} of the model under some objective function. We note that the magbook MM\mathcal{M}_{M}, phasebook FP\mathcal{F}_{P}, and combook CC\mathcal{C}_{C} themselves can be considered fixed (to uniform or pre-trained values as described in the previous section), or optimized jointly with the rest of the network, with the codebook values considered as part of the network parameters.

We present multiple objective functions for the magnitude and phase components as well as for the complex mask; in practice, these objective functions can be combined with each other within a multi-task learning framework. Note also that, for simplicity, we define here the objective functions on a single source-estimate pair, but the definitions can be straightforwardly extended to the permutation-free training scheme commonly used in speech separation .

We can now define an objective function based on the cross-entropy against the oracle codebook assignments for the softmax layer outputs of the magbook, phasebook, and combook layers respectively as:

If cross-entropy is used for the magnitude, the phase mask used to compute the reference magnitude can either be fixed (to , to a reference computed offline given some phasebook values, or to an initial estimate obtained by an initial phasebook network), or updated throughout training (using the reference phase mask obtained with the current phasebook if it is being optimized as well, or with the current estimate of the phase mask obtained by the network).

IV-B Magnitude objectives in the T-F domain: MA, MSA, PSA

All the classical objectives used to train mask inference networks that modify the magnitude can be used here, such as mask approximation (MA), magnitude spectrum approximation (MSA), and phase-sensitive spectrum approximation (PSA). Any norm can be considered to define these objective functions, with L1L^{1} and (squared) L2L^{2} being most commonly used. Using L1L^{1} as an example, we can define:

IV-C Complex objectives in the T-F domain: CMA, CSA

We can extend the classical MA and MSA objective functions to complex versions involving the estimated complex mask cout\bm{{c}}^{\text{out}}, either obtained directly using a combook representation, or obtained by combining magnitude and phase estimates at each T-F bin as

Again using L1L^{1} as an example, we can define a complex mask approximation (CMA) objective using the distance between the reconstructed complex ratio mask ct,foutc_{t,f}^{\text{out}} and a reference complex ratio mask ct,frefc_{t,f}^{\text{ref}} (e.g., ct,fref=st,f/xt,fc_{t,f}^{\text{ref}}=s_{t,f}/x_{t,f}):

We can also define a complex spectrum approximation (CSA) objective using the distance between the reconstructed (i.e., masked) T-F representation and the target T-F representation:

IV-D Time-domain objectives: WA, WA-MISI

Recently, we introduced a waveform approximation (WA) objective defined on the time-domain signal s^[l]\hat{s}[l] reconstructed by inverse STFT from the masked mixture . We also proposed training through an unfolded phase reconstruction algorithm such as multiple input spectrogram inversion (MISI) , using the WA objective on the reconstructed time-domain signal s^(K)[l]\hat{s}^{(K)}[l] after KK iterations.

Denoting by s[l]s[l] the reference time-domain signal, and again using L1L^{1} as an example, we define:

In the same way as we did for magnitude-only mask inference networks , we can train a network that estimates both a magnitude mask and a phase mask, or alternatively a complex mask, end-to-end using the above time-domain objective functions.

IV-E Inference considerations and expected loss

When using the cross-entropy training objectives, there is no inference scheme to be explicitly selected at training time, as the optimization is performed solely on the softmax outputs. While any of the inference schemes could be used at test time, either argmax or sampling inference seem most appropriate given the discrete characteristics of the cross-entropy objective.

For the objectives defined on the value of the estimated mask, the reconstructed time-frequency representation, or the reconstructed time-domain signal, one does need to select at training time an inference scheme used to obtain the masks, and a natural choice at test time is to use the same inference scheme as the one used during training. The interpolation scheme is by far the most convenient, because it ensures that the objective function is differentiable with respect to all parameters, and the gradients can be easily computed using straightforward back-propagation. The sampling and argmax schemes may also be considered, and would be particularly relevant if we were to introduce conditional-probability relationships between T-F bins. However, these schemes raise significant difficulties for the optimization, as both sampling-based and argmax-based selection operations break the differentiation chain.

In order to keep a discrete selection step in the training pipeline for a given loss function, one possibility is to define a corresponding expected loss function which considers all possible choices of values in the codebooks in turn to compute a loss term, weighted by their softmax probability. This corresponds to what one would obtain by sampling many times from the softmax outputs and averaging the loss obtained with the corresponding output. For example, the expected loss version of the CSA loss for the magbook-phasebook case can be defined as

While the sum can in this case be computed exactly by marginalizing over all T-F bins independently, computing such an expectation becomes much trickier for objective functions that include a coupling between the T-F bins. Such is the case for the WA objective, which is defined in the time domain on the inverted T-F representation: the final output depends on all T-F bins, which thus cannot be marginalized over independently, leading to a combinatorial explosion. We could consider approximating the expected loss as the sum of the WA losses for a given number of T-F representations obtained by sampling all T-F bins. Back-propagation could then be performed using the policy gradient technique in the REINFORCE algorithm , similarly to what was done for automatic speech recognition in . Another option would be to rely on the Gumbel-Softmax trick .

Given preliminary results described below on the CSA objective under-performing the WA objective, the significant complexity involved in implementing an expected loss for the WA objective, and the fact that relying on discrete selection instead of interpolation is expected to mostly become relevant when conditional-probability relationships between T-F bins are considered, we leave this line of research for future works.

V Experimental validation

We validate the proposed algorithms on the publicly available wsj0-2mix corpus , which is widely used in speaker-independent speech separation works. It contains 20,000, 5,000 and 3,000 instantaneous two-speaker mixtures in its 30 h training, 10 h validation, and 5 h test sets, respectively. The speakers in the validation set are seen during training, while the speakers in the test set are completely unseen. The sampling rate is 8 kHz.

For our neural networks, we follow the same basic architecture as in , containing four BLSTM layers, each with 600 units in each direction, followed by output layers. A dropout of 0.30.3 is applied on the output of each BLSTM layer except the last one. The networks are trained on 400-frame segments using the Adam algorithm. The window length is 32 ms and the hop size is 8 ms. The square root Hann window is employed as the analysis window and the synthesis window is designed accordingly to achieve perfect reconstruction after overlap-add. A 256-point DFT is performed to extract 129-dimensional log magnitude input features. All systems are implemented using the Chainer deep learning toolkit .

V-B Chimera++ network with phasebook-magbook mask inference head

We build our system based on the state-of-the-art chimera++ network , which combines within a multi-task learning framework a deep clustering head outputting a DD-dimensional embedding for each T-F bin (D=20D=20 here), and a mask-inference head with convex softmax output which predicts a magnitude mask with values in $$. The chimera++ objective function is

where LMI\mathcal{L}_{\text{MI}} can be any of the objective functions described in Section IV, and the weight α\alpha is typically set to a high value, e.g., 0.975. The loss used on the deep clustering head is the whitened k-means loss

As we explained above, the mask-inference head with convex softmax output predicting a magnitude mask can be generalized to a magbook layer. We now add a phasebook layer, similar to the magbook layer, as a new head at the output of the final BLSTM layer, as illustrated in Fig. 7. The final complex mask is obtained by combining the outputs of the magbook and phasebook layers as

and then multiplied with the complex mixture to obtain a complex T-F representation s^t,f\hat{s}_{t,f} of the target estimate:

We still refer to the branch of the network used in computing the final output as the mask-inference (MI) head, which now predicts a complex mask.

V-C Training and inference schemes for phasebook

In this experiment, we start by pre-training chimera++ networks with magbook mask-inference head, where for now we use the fixed convex softmax of for the magbook layer, referred to here as uniform magbook 3. For each of the MSA, PSA, and WA losses as MI objective function, we train such a network from scratch within the multi-task learning setting involving the deep clustering and MI objectives, then discard the deep clustering head and fine-tune the MI head only.

We now add a phasebook mask-inference head to these networks as described in Section V-B, where we assume a fixed uniform codebook with values 2pπP,p=0,…,P−1\frac{2p\pi}{P},p=0,\dots,P-1, referred to as uniform phasebook PP, and we consider: (1) training the phasebook layer by itself while keeping the rest of the network fixed, with the cross-entropy loss LCE-phase\mathcal{L}_{\text{CE-phase}}, and using the argmax scheme in Eq. 5 at inference time; (2) training the phasebook layer by itself while keeping the rest of the network fixed, assuming the interpolation scheme in Eq. (11) is used to obtain the final phase mask value, and either the CSA loss LCSA,L1\mathcal{L}_{\text{CSA},L^{1}} or the WA loss LWA,L1\mathcal{L}_{\text{WA},L^{1}} is used as the training objective; and (3) training the whole network with either the CSA loss LCSA,L1\mathcal{L}_{\text{CSA},L^{1}} or the WA loss LWA,L1\mathcal{L}_{\text{WA},L^{1}}, again assuming the interpolation scheme for the phase.

For this experiment, we consider a uniform phasebook with P=8P=8 elements. Results are shown in Table I in terms of scale-invariant SDR (dB) on the wsj0-2mix test set. From Table I, we see that the CE objective only provides SI-SDR improvements for networks pre-trained with the phase-unaware MSA objective. This intuitively makes sense, as the MSA-based magnitude estimates are likely to be closer to the true magnitude than those obtained with PSA and WA, which try to compensate for the errors in the noisy phase; once the phasebook layer fixes these errors, which it learns to do without considering the interaction with the magnitude in the CE case, the compensation performed by the magnitude estimate may become extraneous or even detrimental. The CSA objective is consistently outperformed by the WA objective both with and without joint training of the magnitude, demonstrating the importance of training through the overlap-add process. Without joint magnitude training when learning the phasebook layer, the CSA training leads to no difference in SI-SDR, while with the WA objective the largest improvement is again observed for MSA. Finally, when allowing joint training of the magbook layer, all pre-training objectives obtain their best performance, with the exception of the CSA objective with WA pre-training, where removing the overlap-add process during fine-tuning leads to a performance degradation. Pre-training with PSA and WA obtains slightly larger values than MSA, and overall, the WA objective with the interpolation scheme appears the most robust, both for pretraining and for training networks involving magbook and phasebook layers. We thus focus on this configuration going forward.

V-D Influence of the phasebook size

Figure 8 shows SI-SDR improvements for various phasebook sizes, where phasebook values are either uniform, pre-trained offline assuming an oracle IAM magnitude, or jointly trained together with the rest of the network. In each case, both the magnitude mask and phase mask layers in the inference head are jointly fine-tuned using the WA loss function, after pre-training of a chimera++ network with WA loss on the MI head. From Fig. 8, we see that all phasebooks improve on the noisy phase SI-SDR of 11.7 dB. We also note that phasebooks of size 8 appear to perform best, and the uniform phasebooks perform comparably to those with learned values. Note that, since we are interpolating over phasebook values, we can theoretically achieve the desired phase difference from any codebook, assuming it is dense enough, so the difference is mainly in the ease for the network to produce softmax outputs that are able to produce a correct estimate. We may see a different trend if we were to pick the argmax or to sample instead.

Figure 9 shows a comparison of the uniform phasebook with the pre-trained and jointly trained values for various codebook sizes. We see that both the pre-trained and jointly trained values tend to place more weight between −π/2-\pi/2 and +π/2+\pi/2 as a majority of the learned values cluster in this range; this matches the empirical distribution shown in Fig. 2. We also note that the jointly trained phasebooks appear to be quite redundant, especially for P=12P=12.

V-E Magbook

We showed in that a convex softmax interpolation of fixed values {0,1,2}\{0,1,2\} for the magnitude mask leads to state-of-the-art performance when combined with an unfolded phase reconstruction algorithm. This corresponds to a uniform magbook 3 in our proposed framework. We here consider an extension of this case using the magbook formulation, where we further train end-to-end the values to be interpolated jointly with the softmax layer under a waveform approximation objective.

Figure 10 shows examples of such learned magbooks. Interestingly, in the linear case, the network finds it best to use one or more negative magnitude elements: it is intuitive in the case of the noisy phase, where the network has an incentive to use its freedom to take negative values in order to fix the noisy phase in regions where a phase inversion is warranted; it is maybe slightly less intuitive when a phasebook layer is involved, as one may think that the phasebook layer should take care of phase inversions where they are needed instead of relying on negative magnitude mask values, but there is in fact no specific incentive in the objective function to favor a positive magnitude value mm associated with some phase θ\theta versus the opposite magnitude value −m-m with phase θ+π\theta+\pi, assuming both these phase values can be equally well generated by the phasebook layer. In the ReLU case, the network can no longer use negative magnitudes, and tends to place multiple points close to . To our surprise, it appears that magbooks obtained with the noisy phase featured slightly larger maximum values than those obtained with a phasebook layer, whereas we argued earlier that using the noisy phase should encourage the network to under-estimate the magnitude mask value. We plan to further investigate the behavior of the estimated masks in these cases by analyzing the estimated softmax probabilities and interpolated values.

V-F Combook

We have so far considered factorized representations of the complex mask as a product of a magnitude mask and a phase mask. We now consider a similar use of a discrete representation to model the complex mask, but directly using a codebook of complex values. We train Chimera++ networks where the magnitude mask estimation layer is replaced by a complex mask estimation layer consisting of a softmax layer used to interpolate values of a combook, as illustrated in Fig. 11. The networks are trained from scratch with both deep clustering and WA objectives, then fine-tuned with WA objective only.

Examples of learned combooks are shown in Fig. 12 for C∈{8,12,17}C\in\{8,12,17\}. We note that the combook size CC should not be directly compared to the phasebook size PP and magbook size MM of the previous sections, since the phasebook and magbook combine to lead to complex values: in the argmax scheme, setting aside one magbook (and combook) value which will most likely be at 0, we have PP phasebook values for each of the remaining M−1M-1 magbook values. A given combination of magbook and phasebook values can thus be considered similar to a set of combook values of size C=1+P(M−1)C=1+P(M-1), e.g., (M,P)=(3,8)(M,P)=(3,8) is akin to C=17C=17. Interestingly, for small sizes such as C=8C=8, the combook layer does not take advantage of non-real values, focusing first on covering negative values (for phase inversion), , and positive values. This is similar to what we observe with some of the linear magbooks in Fig. 10 that learn to allocate magnitude values for phase inversion. Only with C=12C=12 in Fig. 12 do we start seeing non-real values. We note however that the network does not appear to be very efficient in its usage of the available values, learning seemingly redundant values, such as the cluster of points near −3+0j-3+0j in the middle and far right plots of Fig. 12.

Table III compares SI-SDR results for combooks of various sizes, in addition to the best performing magbook and phasebook configurations. It appears that, in the current setup, the ability of the combook layer to estimate a complex mask via a single network layer works slightly better than trying to estimate magnitude and phase via separate layers. Table III also shows that performance does not improve for combook sizes greater than C=12C=12, which we use going forward.

V-G Training through unfolded MISI

Following , we now consider adding an unfolded MISI network with KK iterations at the output of the MI head, as illustrated in Fig. 13, and training the full network using the WA-MISI-K loss function. In the figure, the masks C^i\hat{\bf C}_{i} shall be considered as complex when phasebook or combook layers are involved, and as real otherwise.

Results are shown in Fig. 14 for various numbers of unfolded MISI iterations, and three different types of networks: the original chimera++ network using the noisy phase with a uniform magbook 3 layer with fixed elements {0,1,2}\{0,1,2\}, as a (state-of-the-art) baseline; a chimera++ network with the same architecture and an additional phasebook layer with P=8P=8 uniformly distributed elements; a chimera++ network with a combook layer as MI head whose C=12C=12 elements are learned end-to-end together with the rest of the network parameters. We observe that the combook network improves significantly over the noisy phase baseline and obtains the best performance among all methods for direct iSTFT reconstruction (i.e., MISI iterations), but its performance does not further improve. The phasebook network also improves significantly over the baseline, and converges to an SI-SDR value similar to that of the combook in K=2K=2 iterations. Both combook and phasebook enable a better phase estimate which can match state-of-the-art performance without the need for unfolded phase reconstruction required when using the noisy phase.

Table IV shows a comparison of the best proposed systems with three recently proposed approaches: the original Chimera++ network using noisy phase and MISI phase reconstruction as a post-processing only ; a Chimera++ network trained through unfolded MISI phase reconstruction , which is equivalent in our framework to a uniform magbook 3 with noisy phase as the initial phase; and a Chimera++ network with unfolded phase reconstruction in which the STFT and iSTFT transforms are replaced by separate (or “untied”) transforms at each layer, learned together with the rest of the network . The Jointly trained combook 12 system obtains the best performance when no MISI iteration is performed, at 12.6 dB, beating the previous state-of-the-art 12.2 dB which involves further learning a transform replacing the final iSTFT . If we allow ourselves 5 MISI iterations, all proposed systems reach 12.6 dB, but they are slightly outperformed by the system which learns replacements for the STFT/iSTFT transforms, with 12.8 dB. We shall leave it to future work to combine such transform learning with our proposed systems.

VI Conclusion and future works

According to the above experiments, both a combook layer and a combination of magbook and phasebook layers can significantly improve the performance of single-channel multi-speaker speech separation, especially reducing the need for further phase reconstruction. We have here focused mostly on end-to-end training using the waveform approximation objective, because it has led to the best results both here and in recent work : the most convenient way to use this objective for magbook, phasebook, and combook layers was to rely on the interpolation scheme, where our losses are computed on the expected outputs over the codebooks. We could also investigate training through the argmax scheme by considering expected loss functions that compute the expectation of the loss over each possible value in the codebook. This would in particular allow us to use the discrete nature of the representation to introduce conditional probability relationships between T-F bins. However, as mentioned in Section IV-E, such expectations are typically intractable, necessitating methods such as the Gumbel-Softmax or the policy gradient technique . Finally, while we here considered estimating the difference between the noisy and clean phase, we can consider also estimating the clean phase directly, and train the network to merge the two estimates based on the context.

References