Audio Super Resolution using Neural Networks

Volodymyr Kuleshov, S. Zayd Enam, Stefano Ermon

Introduction

The generative modeling of audio signals is a fundamental problem at the intersection of signal processing and machine learning; recent learning-based algorithms have enabled advances in speech recognition (Hinton et al., 2012), audio synthesis (van den Oord et al., 2016; Mehri et al., 2016), music recommendation systems (Coviello et al., 2012; Wang & Wang, 2014; Liang et al., 2015), and in many other areas (Acevedo et al., 2009). Audio processing also raises basic research questions pertaining to time series and generative modeling (Haykin & Chen, 2005; Bilmes, 2004).

One of the most significant recent advances in machine learning-based audio processing has been the ability to directly model raw signals in the time domain using neural networks (van den Oord et al., 2016; Mehri et al., 2016). Although this affords us the maximum modeling flexibility, it is also computationally expensive, requiring us to handle >10,000>10,000 audio samples at every second.

In this paper, we explore new lightweight modeling algorithms for audio. In particular, we focus on a specific audio generation problem called bandwidth extension, in which the task is to reconstruct high-quality audio from a low-quality, down-sampled input containing only a small fraction (15-50%) of the original samples. We introduce a new neural network-based technique for this problem that is inspired image super-resolution algorithms (Dong et al., 2016), which use machine learning techniques to interpolate a low-resolution image into a higher-resolution one. Learning-based methods often perform better in this context than general-purpose interpolation schemes such as splines because they leverage sophisticated domain-specific models of the appearance of natural signals.

As in image super-resolution, our model is trained on pairs of low and high-quality samples; at test-time, it predicts the missing samples of a low-resolution input signal. Unlike recent neural networks for generating raw audio, our model is fully feedforward and can be run in real-time. In addition to having multiple practical applications, our method also suggests new ways to improve existing generative models of audio.

From a practical perspective, our technique has applications in telephony, compression, text-to-speech generation, forensic analysis, and in other domains. It outperforms baselines at 2×2\times, 4×4\times, and 6×6\times upscaling ratios, while also being significantly simpler than previous methods. Whereas most existing audio enhancement methods make substantial use of signal processing theory, our approach is conceptually very simple and requires no specialized knowledge to implement. Our neural networks are simply trained to map one audio time series into another. Our approach is also among the first to use convolutional architectures for bandwidth extension; as a result, it scales better with dataset size and computational resources relative to current alternatives.

From a generative modeling perspective, our work demonstrates that purely feedforward architectures operating in a non-discretized output space can achieve good performance on an important audio generation task. This hints at the possibility of designing improved generative models for audio that combine both feedforward and recurrent components.

Setup and background

In this work, we interpret RR as the resolution of xx; our goal is to increase the resolution of audio samples by predicting xx from a fraction of its samples taken at {1R,2R,...,RTR}\{\frac{1}{R},\frac{2}{R},...,\frac{RT}{R}\}. Note that by basic signal processing theory, this is equivalent to predicting the higher frequencies of xx.

Bandwidth extension.

Audio upsampling has been studied in the audio processing community under the name bandwidth extension (Ekstrand, 2002; Larsen & Aarts, 2005). Several learning-based approaches have been proposed, including Gaussian mixture models (Cheng et al., 1994; Park & Kim, 2000) and neural networks (Li et al., 2015). These methods typically involve hand-crafted features and use relatively simple models (e.g., neural networks with at most 2-3 densely connected layers) that are often part of a larger, more complex systems. In comparison, our method is conceptually simple (operating directly on the raw audio signal), scalable (our neural networks are fully convolutional and fully feed-forward), more accurate, and is also among the few to have been tested on non-speech audio.

Method

Given a low resolution signal x={x1/R1,...xR1T1/R1}x=\{x_{1/R_{1}},...x_{R_{1}T_{1}/R_{1}}\} sampled at a rate R1R_{1}, our goal is to reconstruct a high-resolution version y={y1/R2,...yR2T2/R2}y=\{y_{1/R_{2}},...y_{R_{2}T_{2}/R_{2}}\} of xx that has a sampling rate R2>R1R_{2}>R_{1}. For example, xx may be a voice signal transmitted via a standard telephone connection at 4 KHz; yy may be a high-resolution 16 KHz reconstruction of the orignal. We use r=R2/R1r=R_{2}/R_{1} to denote the upsampling ratio of the two signals, which in our work equals r=2,4,6r=2,4,6. We thus expect that yrt/R2≈xt/R1y_{rt/R_{2}}\approx x_{t/R_{1}} for t=1,2,...,T1R1t=1,2,...,T_{1}R_{1}.

To recover the under-defined signal, we learn a model p(y∣x)p(y|x) of the higher-resolution yy, conditioned on its low-resolution instantiation xx. We assume that the relationship between the time series x,yx,y follows the equation y=fθ(x)+ϵ,y=f_{\theta}(x)+\epsilon, where ϵ∼N(0,1)\epsilon\sim\mathcal{N}(0,1) is Gaussian noise and fθf_{\theta} is a model parametrized by θ\theta. Our framework also extends to more complex noise models which the user may provide as a prior or that may be themselves parametrized by the model (similarly to how one parametrizes the normal distribution in a variational autoencoder).

The above formulation naturally leads to a mean squared error (MSE) objective

for determining the parameters θ\theta based on a dataset D={xi,yi}i=1n\mathcal{D}=\{x_{i},y_{i}\}_{i=1}^{n} of source/target time series pairs. Since our model is fully convolutional, we may take the xi,yix_{i},y_{i} to be small patches sampled from the full time series.

2 Model Architecture

We parametrize the function ff with a deep convolutional neural network with residual connections; our neural network architecture is based on ideas from Shi et al. (2016), Dong et al. (2016), and Isola et al. (2016), and is shown in Figure 1. We highlight its main features below.

Our model contains BB successive downsampling and upsampling blocks: each performs a convolution, batch normalization, and applies a ReLU non-linearity. Downsampling block b=1,2,...,Bb=1,2,...,B contains max⁡(26+b,512)\max(2^{6+b},512) convolutional filters of length min⁡(27−b+1,9)\min(2^{7-b}+1,9) and a stride of 22. Upsampling block bb has max⁡(27+(B−b+1),512)\max(2^{7+(B-b+1)},512) filters of length min⁡(27−(B−b+1)+1,9)\min(2^{7-(B-b+1)}+1,9).

Thus, at a downsampling step, we halve the spatial dimension and double the filter size; during upsampling, this is reversed. This bottleneck architecture is inspired by auto-encoders, and is known to encourage the model to learn a hierarchy of features. For example, on an audio task, bottom layers may extract wavelet-style features, while higher ones may correspond to phonemes Aytar et al. (2016). Note that the model is fully convolutional, and may run on input sequences of arbitrary length.

Skip connections.

When the source series xx is similar to the target yy, downsampling features will be also be useful for upsampling (Isola et al., 2016). We thus add additional skip connections which stack the tensor of bb-th downsampling features with the (B−b+1)(B-b+1)-th tensor of upsampling features. We also add an additive residual connection from the input to the final output: the model thus only needs to learn y−xy-x, which in practice speeds up training.

Subpixel shuffling layer.

In order to increase the time dimension during upscaling, we have implemented a one-dimensional version of the Subpixel layer of Shi et al. (2016), which has been shown to be less prone to produce artifacts (Odena et al., 2016).

An upscaling block’s convolution maps an input tensor of dimension F×dF\times d into one of size F/2×dF/2\times d. The subpixel layer reshuffles this F/2×dF/2\times d tensor into another one of size F/4×2dF/4\times 2d (while preserving the tensor entries intact); these are concatenated with F/4F/4 features from the downsampling stage, for a final output of size F/2×2dF/2\times 2d. Thus, we have halved the number of filters and doubled the spatial dimension.

Experiments

We use the VCTK dataset (Yamagishi, ) — which contains 44 hours of data from 108 different speakers — and the Piano dataset of Mehri et al. (2016) (10 hours of Beethoven sonatas). We generate low-resolution audio signal from the 16 KHz originals by applying an order 8 Chebyshev type I low-pass filter before subsampling the signal by the desired scaling ratio.

We evaluate our method in three regimes. The SingleSpeaker task trains the model on the first 223 recordings of VCTK Speaker 1 (about 30 mins) and tests on the last 8 recordings. The MultiSpeaker task assesses our ability to generalize to new speakers. We train on the first 99 VCTK speakers and test on the 8 remaining ones; our recordings feature different voices and accents (Scottish, Indian, etc.) Lastly, the Piano task extends audio-super resolution to non-vocal data; we use the standard 88%-6%-6% data split.

Methods.

We compare our method relative to two baselines: a cubic B-spline — which corresponds to the bicubic upsampling baseline used in image super-resolution — and the recent neural network-based technique of Li et al. (2015),

The latter approach takes as input the short-time Fourier transform (STFT) of the input and predicts directly the phase and the magnitudes of the high frequency components using a dense neural network with three hidden layers of size 2048 and ReLU nonlinearities. Li et al. (2015) have shown that this method is preferred over Gaussian Mixture Models in 84% of cases in a user study. This model requires that the scaling ratio be a power of 22, hence it is not applicable when r=6r=6.

We instantiate our model with B=4B=4 blocks and train it for 400 epochs on patches of length 6000 (in the high-resolution space) using the ADAM optimizer with a learning rate of 10−410^{-4}. To ensure source/target series are of the same length, the source input is pre-processed with cubic upscaling. We do not compare against previously-proposed matrix factorization techniques (Bansal et al., 2005; Liang et al., 2013), as they are typically trained on << 10 input examples (Sun & Mazumder, 2013) (due to the cost of jointly factorizing a large number of matrices), and do not scale to the size of our datasets.

Metrics

Given a reference signal yy and an approximation xx, the Signal to Noise Ratio (SNR) is defined as

The SNR is a standard metric used in the signal processing literature. The Log-spectral distance (LSD) (Gray & Markel, 1976) measures the reconstruction quality of individual frequencies as follows:

Evaluation

The results of our experiments are summarized in Table 2. Our objective metrics show an improvement of 1-5 dB over the baselines, with the strongest improvements at higher upscaling factors. Although, the spline baseline achieves a high SNR, its signal often lacks higher frequencies; the LSD metric is better at identifying this problem. Our technique also improves over the DNN baseline; our convolutional architecture appears to use our modeling capacity more efficiently than a dense neural network, and we expect such architectures will soon be more widely used in audio generation tasks.

Next, we confirmed our objective experiments with a study in which human raters were asked to assess the quality of super-resolution using a MUSHRA (MUltiple Stimuli with Hidden Reference and Anchor) test. For each trial an audio sample was upscaled using different techniquesWe have posted a our set of samples to: https://kuleshov.github.io/audio-super-res/.. We collected four VCTK speaker recordings audio samples from the MultiSpeaker testing set. For each recording, we collected the original utterance, a downsampled version at r=4r=4, as well as signals super-resolved using Splines, DNNs, and our model (six versions in total). We recruited 10 subjects and used an online survey to ask each of them to rate each sample on a scale of 0 (extremely bad) to 100 (excellent) reconstruction. The results from the experiment are summarized in Table 1. Our method ranked as being the best out of the three upscaling techniques.

Domain adaptation.

We tested the sensitivity of our method to out-of-distribution input via an audio super-resolution experiment in which the training set did not use a low-pass filter, while the test set did, and vice-versa. We focused on the Piano task and r=2r=2. The output from the model was noisier than expected, indicating that generalization is an important practical concern. We suspect this behavior may be common in super-resolution algorithms, but has not been widely documented. A potential solution would be to train on data that has been generated using multiple techniques.

In addition, we examined the ability of our model to generalize from speech to music and vice versa. We found that switching domains produced noisy output, again highlighting the specialization of the model.

Architectural analysis.

Computational performance.

Our model is computationally efficient and can be run in real time. On the Piano task (where all input signals are 12s in length), our method processed a single second of audio in 0.11s on average on a Titan X GPU. Training our models, however, required about 2 days for the MultiSpeaker task. Unlike sequence-to-sequence architectures our model does not require the complete input sequence in order to begin generating an output sequence.

1 Limitations

Finally, to explore the limits of our approach, we evaluated our method on the MagnaTagATune dataset, which consists of about 200 hours of music from 188 different genres. This dataset is larger and much more diverse that the ones we considered so far. We found that our model underfit the dataset, with very little reduction in the training error, and no improvement over the spline baseline. Other learning-based baselines fared similarly. However, we expect improved results with a larger model and more computational resources.

Previous Work and Discussion

In the machine learning literature, time series signals have most often been modeled with auto-regressive models, of which variants of recurrent networks are a special case (Gers et al., 2001; Maas et al., 2012; Mehri et al., 2016). Our approach instead generalizes conditional modeling ideas used in computer vision for tasks such as image super-resolution (Dong et al., 2016; Ledig et al., 2016) or colorization (Zhang et al., 2016).

We identify a broad class of conditional time series modeling problems that arise in signal processing, biomedicine, and other fields and that are characterized by a natural alignment among source/target series pairs and differences that are well-represented by local transformations. We propose a general architecture for such problems and show that it works well in different domains.

Bandwidth extension.

Existing learning-based approaches include Gaussian mixture models (Cheng et al., 1994; Park & Kim, 2000; Pulakka et al., 2011), linear predictive coding (Bradbury, 2000), and neural networks (Li et al., 2015). Our work proposes the first convolutional architecture, which we find to scale better with dataset size and outperform recent, specialized methods. Moreover, while existing techniques involve many hand-crafted features (see e.g., Pulakka et al. (2011)); our approach is fully domain-agnostic.

Audio applications.

In telephony, commercial efforts are underway to transmit voice at higher rates (typically 16 Khz) in specific handsets; audio-super resolution is a step towards recreating this experience in software. Similar applications could be found in compression, text-to-speech generation, and forensic analysis. More generally, our work demonstrates the effectiveness of feedforward convolutional architectures on an audio generation task.

Conclusion

Machine learning techniques based on deep neural networks have been successful at solving under-defined problems in signal processing such as image super-resolution, colorization, in-painting, and many others. Learning-based methods often perform better in this context than general-purpose algorithms because they leverage sophisticated domain-specific models of the appearance of natural signals.

In this work, we proposed new techniques that use this insight to upsample audio signals. Our technique extends previous work on image super-resolution to the audio domain; it outperforms previous bandwidth extension approaches on both speech and non-vocal music. Our approach is fast and simple to implement, and has applications in telephony, compression, and text-to-speech generation. It also demonstrates the effectiveness of feedforward architectures on an important audio generation task, suggesting new directions for generative audio modeling.

References