Speech Denoising with Deep Feature Losses
Francois G. Germain, Qifeng Chen, Vladlen Koltun
I Introduction
Speech denoising (or enhancement) refers to the removal of background content from speech signals . Due to the ubiquity of this audio degradation, denoising has a key role in improving human-to-human (e.g., hearing aids) and human-to-machine (e.g., automatic speech recognition) communications. A particularly challenging but common form of the problem is the under-determined case of single-channel speech denoising, due to the complexity of speech processes and the unknown nature of the non-speech material. The complexity is further compounded by the nature of the data, since audio material contains a high density of data samples (e.g., 16,000 samples per second). Challenges also arise in mediated human-to-human communication, as perception mechanisms can make small errors still noticeable by the average user .
In this work, we present an end-to-end deep learning approach to speech denoising. Our approach trains a fully-convolutional denoising network using a deep feature loss. To compute the loss between two waveforms, we apply a pretrained audio classification network to each waveform and compare the internal activation patterns induced in the network by the two signals. This compares a multitude of features at different scales in the two waveforms. We perform extensive experiments that compare the presented approach to recent state-of-the-art end-to-end deep learning techniques for denoising. Our approach outperforms them in both objective speech quality metrics and large-scale perceptual experiments with human listeners, which indicate that our approach is more effective than the baselines. The advantages of the presented approach are particularly pronounced for the hardest, noisiest inputs, for which denoising is most challenging.
Before the popularization of deep networks, denoising systems relied on spectrogram-domain statistical signal processing methods , followed more recently by spectrogram factorization-based methods . Current denoising pipelines instead rely on deep networks for state-of-the-art performance. However, most pipelines still operate in the spectrogram domain . As such, signal artifacts then arise due to time aliasing when using the inverse short-time Fourier transform to produce the time-domain enhanced signal. This particular issue can be somewhat alleviated, but with increased computational cost and system complexity .
Recently, there has been growing interest in the design of performant denoising pipelines that are optimized end-to-end and directly operate on the raw waveform. Such approaches aim at fully leveraging the expressive power of deep networks while avoiding expensive time-frequency transformations or loss of phase information . Some of these approaches typically use simple regression loss functions for training the network (e.g., loss on the raw waveform), while ones with more advanced loss functions have shown limited gains in mismatched conditions .
For our loss function, we are inspired by computer vision research, where activations in pretrained classification networks were found to yield effective loss functions for image stylization and synthesis . To compute the loss between two images, these approaches apply a pretrained image classification network to both. Each image induces a pattern of internal activations in the network to be compared, and the loss is defined in terms of their dissimilarity. Such complex training losses have been shown to yield state-of-the-art algorithms without the need for prior expert knowledge or added complexity for the processing network itself. Furthermore, increased performance can be achieved even without task-specific loss networks . Our work develops this idea in the context of speech processing.
II Method
Let be an audio signal corresponding to speech that is corrupted by an additive background signal so that . Our goal is to find a denoising operator such that . We use a fully-convolutional network architecture based on context aggregation networks . The output signal is synthesized sample by sample as we slide the network along the input. Context aggregation networks have been previously used in the WaveNet architecture for speech synthesis . Our architecture is simpler than WaveNet – no skip connections across layers, no conditioning, no gated activations – while our loss function is more advanced, as described in Section II-B.
Adaptive normalization
II-B Feature loss
In our experiments, simple training losses (e.g., ) led to noticeably degraded output quality at lower signal-to-noise ratios (SNRs). The network seemed to improperly process low-energy speech information of perceptual importance. Instead, we train the denoising network using a deep feature loss that penalizes differences in the internal activations of a pretrained deep network that is applied to the signals being compared. By the nature of layered networks, feature activations at different depths in the loss network correspond to different time scales in the signal. Penalizing differences in these activations thus compares many features at different audio scales.
In computer vision, there are standard classification networks such as VGG-19 , pretrained on standard classification datasets such as ImageNet . Such standard classification networks do not exist in the audio processing field yet, so we design and train our own feature loss network.
We design a simple audio classification network inspired by the VGG architecture in computer vision , since it is known as a particularly effective feature loss architecture . The network consists of 15 convolutional layers with kernels, batch normalization, LReLU units, and zero padding. Each layer is decimated by 2, halving the length of the subsequent layer compared to the preceding one. The number of channels is doubled every 5 layers, with 32 channels in the first intermediate layer. Each channel in the last feature layer is average-pooled to yield the output feature vector. The receptive field is samples. We train the network using backpropagation by feeding its output vector as features to one or more logistic classifiers with a cross-entropy loss for one or more classification tasks.
Denoising loss function
Let be the -th feature layer of the feature loss network, with layers at different depths corresponding to features with various time resolutions. The feature loss function is defined as a weighted loss on the difference between the feature activations induced in different layers of the network by the clean reference signal and the output of the denoising network being trained:
where are the parameters of the denoising network. The weights are set to balance the contribution of each layer to the loss. They are set to the inverse of the relative values of after 10 training epochs. (For these first 10 epochs, the weights are set to 1.)
III Training
To generate a general-purpose feature loss network, we train it jointly on multiple audio classification tasks (only the logistic classifier parameters are trained as task-dependent). We use two tasks from the DCASE 2016 challenge : the acoustic scene classification task and the domestic audio tagging task. In the first task, we are provided with audio files featuring various scenes (e.g., beach); the goal is to determine the scene type for each file. In the second task, we are given audio files featuring events of interest (e.g., child speaking); the goal is to determine which events took place in each file (with possibly multiple events in one file).
Data
For the scene classification task, the training set consists of 30-second-long audio files sampled at 44.1kHz, split among 15 different scenes (i.e., classes). As we need to develop a feature loss for the reduced sampling frequency of 16kHz, we resample the data. The audio files are stereo, so we split them into two mono files. The training set contains files. For the tagging task, the training set CHiME-Home-refine consists of 4-second-long mono audio files sampled at 16kHz, with 7 different tags (i.e., labels). The training set contains files.
Training
Network weights are initialized with Xavier initialization . We use the Adam optimizer with a learning rate of . The model is trained for epochs. In each epoch, we iterate over the training data for each task, alternating between files from each task. The order of the files is randomized independently for each epoch. The dataset for the first task is larger than the one for the second task, so we present some of the files in the second dataset (chosen at random) a second time to preserve strict alternation between tasks. 1 epoch consists of iterations (1 file per iteration). As a data augmentation procedure, we do not present entire clips, but present a continuous section of minimal duration samples that is culled at random for each iteration.
III-B Speech Denoising
We use the noisy dataset made available in . To our knowledge, this is the largest available dataset for denoising that provides pre-mixed data with a clearly documented mixing procedure. It also has the benefit of being the dataset used in two recent works that we use as baselines. All details concerning the data can be found in . The training set is generated from the speech data of 28 speakers (14 male/14 female) and the background data of 10 unique background types. Each noise segment is used to generate four files with 0, 5, 10, and 15dB SNR. The published files are sampled at 48kHz and normalized so that the clean speech files have a maximum absolute amplitude of 0.5. We resample them to 16kHz. The complete dataset comprises files.
Training
IV Experimental Setup
As baselines, we use a Wiener filtering pipeline with a priori noise SNR estimation (as implemented in ), and two recent state-of-the-art methods that use deep networks to perform end-to-end denoising directly on the raw waveform: the Speech Enhancement Generative Adversarial Network (SEGAN) and a WaveNet-based network . This last one is designed around minor modifications to the architecture in . It uses stacked context aggregation modules with gated activation units, skip connections, and a conditioning mechanism. The modifications include training with a regression loss ( on the raw waveform) rather than a classification loss. The number of layers is larger than in our network (30), while the receptive field is smaller ( samples), capturing contextual information on more limited time scales. The network architecture is also distinctly more complex than ours. For both deep learning baselines, we use the code and models published by their respective authors. These models are optimized by their authors on the exact same training dataset, allowing fair comparison.
IV-B Data
IV-C Quantitative measures
To evaluate each system, we compare its output to the ground-truth speech signal (i.e., the clean speech alone). The common metrics to measure speech quality given ground-truth are compared in . We use here the composite scores from that were found to be best correlated with human listener ratings. These consist of the overall (OVL), the signal (SIG), and the background (BAK) scores, each on a scale from 1.0 to 5.0, and corresponding respectively to the measure of overall signal quality, the measure of quality when considering speech signal degradation alone, and the measure of quality when considering background signal intrusiveness alone . We also report the SNR , as a raw measure of the relative energies of the residual background and the speech in a given signal, quantified in decibel (dB). We use the implementations in . For all metrics, higher scores denote better performance.
The test dataset is divided into 4 mixing SNR subgroups (see Section IV-B). We argue that the dataset should be rather considered as a continuous distribution of degradation, since SNR correlates poorly with human perception of the degradation level . The continuum of degradation levels is better represented in the distribution of the background intrusiveness BAK score. (The SIG score is less informative since the undistorted speech signal is added.) To evaluate performance as a function of input degradation magnitude, we partition the test set into 8 tranches of equal size, corresponding to the 8 octiles of the BAK score distribution as shown in Figure 1, with tranches representing a different denoising difficulty.
Results
Table I reports these metrics for our approach and the baselines, evaluated over the test set. Our method outperforms all the baselines according to all measures by a comfortable margin. The plots in Figure 2 further show that our network yields the best quality for all levels of background intrusiveness separated in tranches, with a particularly significant margin according to perceptually-motivated composite measures. Table II shows the benefit of using a feature loss compared to training the same denoising network, by the same procedure on the same data, using an or an loss. Training with a feature loss outperforms networks trained with other losses. In particular, while an loss achieves a similar SNR score as our feature loss, the feature loss shows definite improvement for the BAK and OVL metrics. It also scores well for the SIG metric, especially in the noisier tranches, demonstrating the ability to capture meaningful features when important cues are hidden in the noise.
IV-D Perceptual Experiments
Objective metrics are known to only partially correlate with human audio quality ratings . Hence, we also conduct carefully designed perceptual experiments with human listeners. The procedure is based on A/B tests deployed at scale on the Amazon Mechanical Turk platform. The A/B tests are grouped into Human Intelligence Tasks (HITs). Each HIT consists of 100 “ours vs baseline” pairwise comparisons. Each comparison presents two audio clips that can be played in any order by the worker, any number of times. One of the clips is the output of our approach and one is the output of one of the baselines, for the same input from the test set. The files are presented in random order (both within each pair and among pairs), so the worker is given no information as to the provenance of the clips. The worker is asked to select, within each pair, the clip with the cleaner speech. Each HIT includes 10 additional ‘sentry’ comparisons in which the right answer is obvious to guard against negligent or inattentive workers. These sentry pairs are mixed into the HIT in random order. If a worker gives an incorrect answer to two or more sentry pairs, the entire HIT is discarded. Each HIT then contains a total of 110 pairwise comparisons. A worker is given 1 hour to complete a HIT. Each HIT is completed by 10 distinct workers.
Results
The results are summarized in Table III. This table presents the fraction of blind pairwise A/B comparisons in which the listener rated a clip denoised by our network as cleaner than the clip denoised by a baseline. The preference rates are presented versus each baseline across 4 tranches. The most notable results are for the hardest tranche, where the output of our approach was rated cleaner than the output of recent state-of-the-art deep networks in more than 83% of the comparisons. All results are statistically significant with . This demonstrates that our algorithm is more robust in this regime, in which degradation from the background signal is much more noticeable, and for which denoising is particularly useful. For easier tranches, with lower levels of degradation in the input, both our method and the baselines generally perform satisfactorily and listeners can experience more difficulty distinguishing between the different processed files, but the preference rate for our approach remains well above chance (50%), at statistically significant levels, for all baselines across all tranches.
V Conclusion
We presented an end-to-end speech denoising pipeline that uses a fully-convolutional network, using a deep feature loss network pretrained on several relevant audio classification tasks for training. This approach allows the denoising system to capture speech structure at various scales and achieve better denoising performance without added complexity in the system itself or expert knowledge in the loss design. Experiments demonstrate that our approach significantly outperforms recent state-of-the-art baselines according to objective speech quality measures as well as large-scale perceptual experiments with human listeners. In particular, the presented approach is shown to perform much better in the noisiest conditions where speech denoising is most challenging. Our paper validates the combined use of convolutional context aggregation networks and feature losses to achieve state-of-the-art performance.
References
-A Denoising Network
We denote the (consecutive) network layers by . and are 1-dimensional tensors of dimensionality and correspond to the degraded input signal and the enhanced output signal, respectively. The number of samples is not given in advance. Each intermediate layer is a 2-dimensional tensor of dimensionality , where is the width of (i.e., the number of feature maps in) each layer. For , the content of each intermediate layer is computed from the previous layer via the operation
where is the -th feature map of layer , is the -th feature map of layer , is a learned convolutional kernel, is the adaptive normalization operator and is a pointwise nonlinearity. Because of the presence of adaptive normalization, no bias term is used for these layers. The operator is a dilated convolution , i.e.,
The dilation factor for the -th layer is set at for . Between layer and , we do not use dilation (i.e., ). For the output layer , we use a linear transformation ( convolution with no nonlinearity) in order to synthesize the sample of the output signal so that
where is a learned bias term. The receptive field of the network is samples.
Nonlinear units
For the pointwise nonlinearity , we use the leaky rectified linear unit (LReLU) :
Adaptive normalization
corresponds to the adaptive normalization operation described in Section II-A. For , the operator adaptively combines batch normalization and identity mapping as
Zero padding
Our algorithm uses zero-padding at each layer so that the “effective” length of each layer tensor is constant and identical to .
Training loss
The network is trained through backpropagation using our deep feature loss as described in Section II-B (see in particular Equation 1). The feature loss classification network is further detailed in the next section.
-B Feature Loss Network
As mentioned in Section II-B, the network is inspired by the VGG architecture from computer vision. We denote its 15 (consecutive) layers by . The first layer is a 1-dimensional tensor of dimensionality and corresponds to the input signal. The number of samples is not given in advance. Each intermediate layer is a 2-dimensional tensor of dimensionality , where is the width of each layer, set to (i.e., the number of features is doubled every 5 layers). The content of each intermediate layer is computed from the previous layer through the following operation:
Classification layer
where is the logistic nonlinearity associated with the type of multi-label classification for the -th task (i.e., vector softmax nonlinearity if the task asks for a unique label for each audio file, pointwise sigmoid if the task allows for any number of labels for each audio file). is of dimension and its elements are in the range $$.
Training loss
Training is done through backpropagation using a cross-entropy loss between the vector associated with the current file (for task ) and its corresponding ground truth classification vector (i.e., the vector of dimension in which the -th element is 1 if the -th classification label is associated with the file, 0 otherwise).