Monaural Singing Voice Separation with Skip-Filtering Connections and Recurrent Inference of Time-Frequency Mask
Stylianos Ioannis Mimilakis, Konstantinos Drossos, João F. Santos, Gerald Schuller, Tuomas Virtanen, Yoshua Bengio
I. Introduction
where, and denote the entry-wise absolute and exponentiation operators respectively, and is an exponent chosen based on the assumed distributions that the sources follow. Finding (and thus an optimal for the source estimation process ) is an open optimization problem .
Deep learning methods for music source separation are trained using synthetically created mixtures (adding signals together, i.e., knowing the target decomposition). They can be divided into two categories. In the first category, the methods try to predict the mask directly from the mixture magnitude spectrum (i.e. ). This requires that an optimal is given (e.g. all the non-linear mixing parameters of the target source are known) during training as a target. However, such information for the is unknown, and an approximation of is computed from the training data using Eq. (2) and empirically chosen values, under the hypothesis that the source magnitude spectra are additive, which is not true for realistic audio signals . This implies that such models are optimized to predict non-optimal masks. The methods in the second category try to estimate all sources from the mixture(i.e. ) and then use these estimates to compute a mask . This approach is widely adopted, since it is straightforward by employing denoising autoencoders , with noise corresponding to the addition of other sources. However, the masks are dependent on the initial -power magnitude estimates of the sources (), and the mask computation is not a learned function. Instead, the mask computation uses a deterministic function which takes as inputs the outcomes () of deep neural networks, e.g. as in .
An exception to the above are the works presented in and , where these methods jointly learned and optimized the masking processes described by Eq. (1) and (2). In , highway networks were shown to be able to approximate a masking process for monaural solo source separation and in , a more robust alternative to is presented. The approach in uses a recurrent encoder-decoder with skip-filtering connections, which allow a source-dependent mask generation process, applicable to monaural singing voice separation. However, the generated masks are not robust against interferences from other music sources, thus requiring a post-processing step using the generalized Wiener filtering .
In this work we present a method for source separation that learns to generate a source-dependent mask which does not require the generalized Wiener filtering as a post-processing step. To do so, we introduce a novel recurrent inference algorithm inspired by and a sparsifying transform for generating the mask . The recurrent inference allows the proposed method to have a stochastic depth of RNNs during the mask generation process, computing hidden, latent representations which are presumably better for generating the mask. The sparsifying transform is used to approximate the mask using the output of the recurrent inference. In this method the mask prediction is not based on the above mentioned assumptions about the additivity of the magnitude spectrogram of the sources, is part of an optimization process, and is not based on a deterministic function. Additionally, the method incorporates RNNs instead of feed-forward or convolutional layers for the mask prediction. This allows the method to exploit the memory of the RNNs (compared to CNNs) and their efficiency for modeling longer time dependencies of the input data. The rest of the paper is organized as follows: Section II presents the proposed method, followed by Section III which provides information about the followed experimental procedure. Section IV presents the obtained results from the experimental procedure and Section V concludes this work.
II. Proposed Method
Our proposed method takes as an input the time domain samples of the mixture, and outputs time domain samples of the targeted source. The model consists of four parts. The first part implements the analysis and pre-processing of the input. The second part generates and applies a mask, thus creating the first estimate of the magnitude spectrogram of the targeted source. The third part enhances this estimate by learning and applying a denoising filter, and the fourth part constructs the time domain samples of the target source. We call the second part the “Masker” and the third the “Denoiser”. We differentiate between the Masker and the Denoiser because the Masker is optimized to predict a time-frequency mask, whereas the Denoiser enhances the result obtained by time-frequency masking. We implement the Masker using a single layer bi-directional RNN encoder (), a single layer RNN decoder (), a feed-forward layer (FFN), and skip-filtering connections between the magnitude spectrogram of the mixture and the output of the FFN. We implement the Denoiser using one FFN encoder (), one FFN decoder (), and skip-filtering connections between the input to the Denoiser and the output of the . We jointly train the Masker and the Denoiser using two criteria based on the generalized Kullback-Leibler divergence (), as it is shown in to be a robust criterion for matching magnitude spectrograms. All RNNs are gated recurrent units (GRU). The proposed method is illustrated in Figure 1.
ii. The Masker
Recurrent inference and mask prediction Inspired by recent optimization methods employing stochastic depth , we propose a recurrent inference algorithm that processes the latent variables of the which affect the mask generation. We use this algorithm in order to employ a stochastic depth for the network parts responsible for predicting the mask, increasing the performance of our method. The recurrent inference is an iterative process and consists in reevaluating the latent variables , produced by the , until a convergence criterion is reached, thus avoiding the need to specify a fixed number of applications of the . The stopping criterion is a threshold on the mean-squared-error () between the consecutive estimates of , with a threshold . A maximum number of iterations () is used to avoid having infinite iterations for convergence between the above mentioned consecutive estimates. is used only for the singing voice, i.e. . Let be the source-dependent and trainable function of the . The recurrent inference is performed using Algorithm 1.
is then given to the FFN layer with shared weights through time frames, in order to approximate the -th source-dependent mask as:
iii. The Denoiser
The output of the Masker is likely to contain interferences from other sources . The Denoiser aims to learn a denoising filter for enhancing the magnitude spectrogram estimated by this masking procedure. This denoising filter is implemented by an encoder-decoder architecture with the and of Fig. 1. and have shared weights through time frames. The final enhanced magnitude spectrogram estimate of the target source is computed using
iv. Training details and post-processing
We train our method to minimize the objective consisting of a reconstruction and a regularization part as:
where and are hyper-parameters penalizing the mask generation process, allowing a collaborative minimization of the overall objective. The usage of will ensure that can be used to initially estimate the target source, which is then improved by employing the Denoiser and . The penalization of the elements in the main diagonal of will ensure that the generated mask is not something trivial (e.g. a voice activity detector), while the reconstruction losses using the will ensure that a source-dependent mask is generated, that minimizes the aforementioned distance. The squared matrix norm is employed to improve the generalization of the model.
III. Experimental Procedure
We use the development subset of Demixing Secret Dataset (DSD) http://www.sisec17.audiolabs-erlangen.de and the non-bleeding/non-instrumental stems of MedleydB for the training and validation of the proposed method. The evaluation subset of DSD is used for testing the objective performance of our method. For each multi-track contained in the audio corpus, a monaural version of each of the four sources is generated by averaging the two available channels. For training, the true source is the outcome of the ideal ratio masking process , element-wise multiplied by a factor of . This is performed to avoid the inconsistencies in time delays and mixing gains between the mixture signal and the singing voice (apparent in MedleydB dataset). The length of the sequences is set to , modeling approximately seconds, and . The thresholds for the minimization of Eq.(9) are and and the corresponding scalars are , and . The hidden to hidden matrices of all RNNs were initialized using orthogonal initialization and all other matrices using Glorot normal . All parameters are jointly optimized using the Adam algorithm , with a learning rate of , over batches of , an based gradient norm clipping equal to and a total number of epochs. All of the reported parameters were chosen experimentally with two random audio files drawn from the development subset of DSD. The implementation is based on PyTorch http://pytorch.org/.
We compared our method with other state-of-the-art approaches dealing with monaural singing voice separation, following the standard metrics, namely signal to noise ratio (SIR) and signal to distortion ratio (SDR) expressed in dB, and the rules proposed in the music source separation evaluation campaign (e.g. using the proposed toolbox for SIR and SDR calculation). The compared methods are: i) GRA: Deep FFNs for predicting both binary and soft masks which are then combined to provide source estimates, ii) CHA: A convolutional encoder-decoder for magnitude source estimation, without a trainable mask approximation iii) MIM-HW: Deep highway networks for music source separation approximating the filtering process of Eq.(1), retrained using the development subset of DSD, and iv) MIM-DWF, MIM-DWF+: The two GRU encoder-decoder models combined with generalized Wiener filtering , trained on the development subset of DSD (MIM-GRUDWF) and the additional stems of MedleydB (MIM-DWF+). The methods denoted as MIM-HW, MIM-DWF, and MIM-DWF+ were re-implemented for the purposes of this work. For the rest of the methods we used their reported evaluation results obtained from . Our proposed methods are denoted as GRU-NRI, which does not include the recurrent inference algorithm, and two methods using different hyper-parameters for the recurrent inference algorithm: GRU-RIS, parametrized using a maximum number of iterations , and , and GRU-RIS parametrized using a maximum number of iterations and , which where selected according to their performance in minimizing Eq. (9).
IV. Results & Discussion
Table 1 summarizes the results of the objective evaluation for the aforementioned methods by showing the median values obtained from the SDR and SIR metrics. The proposed method based on recurrent inference and sparsifying transform is able to provide state-of-the-art results for monaural singing voice separation, without the necessity of post-processing steps such as generalized Wiener filtering, and/or additionally trained deep neural networks. Compared to methods that approximate the masking processes (GRA, MIM-HW, MIM-DWF, and MIM-DWF+) there are significant improvements in overall median performance of both the SDR and SIR metrics, especially when the masks are not a learned function, such as in the case of CHA.
Using the proposed method, a gain of dB for the SDR is observed between MIM-DWF+ and GRU-RIS and dB for the SIR between the MIM-DWF and GRU-RIS. Finally, by allowing a larger number of iterations during the recursive inference the mask generation performance and using skip-filtering connections we see an increase in SDR which outperforms the previous methods MIM-DWF and MIM-DWF+, but at the cost of a loss in SIR. A demo for the proposed method is available at https://js-mim.github.io/mss_pytorch/.
V. Conclusion
In this work we presented an approach for singing voice separation that does not require post-processing using generalized Wiener filtering. We introduced to the skip-filtering connections a sparsifying transform yielding comparable results to approaches that rely on generalized Wiener filtering. Furthermore, the introduced recurrent inference algorithm was shown to provide state-of-the-art results in monaural singing voice separation. Experimental results show that these extensions outperform previous deep learning based approaches for singing voice separation.
VI. Acknowledgements
The research leading to these results has received funding from the European Union’s H2020 Framework Programme (H2020-MSCA-ITN-2014) under grant agreement no 642685 MacSeNet. Part of the computations leading to these results were performed on a TITAN-X GPU donated to K. Drossos from NVIDIA. The authors would like to thank Paul Magron for the precious feedback.