Real-time Speech Frequency Bandwidth Extension

Yunpeng Li, Marco Tagliasacchi, Oleg Rybakov, Victor Ungureanu, Dominik Roblek

Introduction

Frequency bandwidth reduction is one of the artifacts produced by lossy speech codecs used throughout telecommunication networks, leading to narrowband speech with frequencies up to 4kHz. This gives the received speech its familiar dull, muffled character. More recent coding standards, such as for example AMR-WB (Adaptive Multi-Rate Wideband), are explicitly designed to overcome this issue, operating in native wideband mode — also known as HD Voice — retaining frequencies up to 7kHz. However, to establish an HD Voice call, several requirements need to be met simultaneously: the mobile handsets, the underlying telephone infrastructure and the mobile cells all need to support wideband speech, and in some cases both the transmitter and the receiver need to be served by the same operator. As a consequence, HD voice calls are not yet ubiquitous, thus creating an opportunity for low-latency solutions deployed on the receiver device that aim at enhancing the quality of speech, reconstructing the high frequency content that was lost during transmission.

In this paper we propose such a solution, devising a wave-to-wave model for frequency bandwidth extension, which is fully causal and sufficiently lightweight for real-time applications on mobile devices. Our model is trained using a combination of adversarial and reconstruction objectives, defined by a discriminator learned concurrently with the generator, with a design similar to the audio-only version of SEANet (Sound Enhancement Network) that we introduced in our previous work .

In our experiments we mainly target the case of increasing the sampling frequency from 8kHz to 16kHz, thus extending the available bandwidth from 4kHz to 8kHz, respectively. We evaluate the output of the model with both objective evaluation metrics and subjective ratings, and we observe that the quality is indistinguishable from the original samples at 16kHz. We further investigate the scenario where the input bandwidth at test time does not exactly match what was used during training, where catastrophic failure can occur even for relatively small mismatches. We show that the model can be made robust against such mismatches by using a range of bandwidths during learning. Moreover, and somewhat surprisingly, the robustness induced by such a learning scheme persists even when the input bandwidth at test time is outside of the range seen during training.

In addition, one of our main contributions is to explicitly address the problem of deploying SEANet on a mobile device, aiming at low latency. This is inspired by previous work on streaming architectures for keyword spotting , that we extended to support the operations needed to deploy a UNet generator. With this solution, we are able to process each 16ms audio frame in ∼\sim1.5ms on the CPU of a mobile device, so that the total latency is ∼\sim17.5 ms.

The design of SEANet is similar to , but we adopt the losses proposed in , in which the reconstruction loss is computed in the feature space of the discriminator, at different scales and at different layers. In addition, we explicitly address the problem of low-delay inference, which was not addressed in previous work in this area.

Method

The generator is a 1D convolutional U-Net that adopts the same overall architecture as the audio-only SEANet proposed in . The model is a symmetric encoder-decoder network with skip-connections and residual units. The encoder and the decoder each have four convolution blocks, which are sandwiched between two plain convolution layers. The encoder follows a down-sampling scheme of (2, 2, 8, 8) while the decoder up-samples in the reverse order. The number of channels is doubled whenever down-sampling and halved whenever up-sampling. Each encoder block consists of three residual units each containing convolutions with dilation rates of 1, 3, and 9, respectively, followed by a down-sampling layer in the form of a strided convolution. The decoder block mirrors the encoder block, and consists of a transposed convolution for up-sampling followed by the same three residual units. A skip-connection is added between each encoder block and its mirrored decoder block. The out-most skip-connection directly connects the input waveform to the output waveform. We refer interested readers to for details of the SEANet architecture.

To be able to process live audio streams in real-time and potentially on low-power mobile devices, our model differs from the original SEANet version in several ways:

All convolutions are causal. This means that padding is only applied to the past but not the future in training and offline inference, whereas no padding is used in streaming inference.

The model is much more lightweight. We reduce the number of channels at each layer 4-fold, which results in a model approximately 1/16 the size and computational cost of the original model.

We do not apply any normalization. While weight normalization is beneficial in the original model, it did not help in the smaller model likely due to having much fewer weights in the convolution kernels.

We call this new model the Streaming SEANet to distinguish it from the original version, and we will describe in Section 2.3 how real-time processing can be realized by means of streaming convolutions.

As in the original SEANet model, we use the same multi-resolution convolutional discriminator proposed in . Three structurally identical discriminators are applied to the input audio at different resolutions: original, 2-times down-sampled, and 4-times down-sampled. Each discriminator consists of an initial plain convolution followed by four grouped convolutions , each of which has a group size of 4, a down-sampling factor of 4, and a channel multiplier of 4 up to a maximum of 1024 output channels. They are followed by two more plain convolution layers to produce the final output, i.e., the logits. Since the discriminator is fully convolutional, the number of logits in the output is proportional to the length of the input audio. See for more architectural details of the discriminator.

We use ELU activation in the generator and Leaky ReLU activation with the default α=0.2\alpha=0.2 in the discriminator. All residual and skip connections in the generator are pre-activation . Layer normalization is used in the discriminator whereas the generator does not use any normalization.

2 Learning

We use the same adversarial and reconstruction objective as , albeit without the auxiliary accelerometer signal. Given a time-aligned input-target audio pair (x,y)(x,y), let G(x)G(x) be the output audio of the generator. The adversarial loss is a hinge loss over the logits of the discriminator, averaged over multiple resolutions and over time. More formally, let k∈{1,…,K}k\in\{1,\dots,K\} index over the individual discriminators for different resolutions (K=3K=3 in our case), and tt index over the length of the output, i.e., the number of logits TkT_{k}, of the discriminator at scale kk. The discriminator is trained to classify clean vs. generated audio, by minimizing

while the adversarial loss for the generator is

The reconstruction loss is the “feature” loss proposed in , namely the average absolute difference between the discriminator’s internal layer outputs for the generated audio and those for the corresponding target audio.

where LL is the number of internal layers, Dl,t(l)D_{l,t}^{(l)} (l∈{1,…,L}l\in\{1,\dots,L\}) is the tt-th output of layer ll of discriminator kk, and Tk,lT_{k,l} denotes the length of the layer in the time dimension. The overall generator loss is a weighted sum of the adversarial loss LGadv\mathcal{L}_{G}^{\text{adv}} and the reconstruction loss LGrec\mathcal{L}_{G}^{\text{rec}}.

In all our experiments, we train for 1 million steps with a batch size of 16 using the same optimizer parameters and a weighting factor of 100 for LGrec\mathcal{L}_{G}^{\text{rec}} as in .

3 Real-time Processing of Streaming Input

Unlike offline inference where the entire input audio is available at once, in streaming applications the input is fed continuously. For efficient computation with low delay, convolutions must be performed in a streaming fashion where rolling buffers are used to maintain past intermediate outputs. A key challenge is to provide automated conversion between regular convolutions for training and streaming convolutions for inference under a common interface, as this allows the decoupling of high-level model architecture design from low-level considerations for streaming inference.

To this end, we extend the Streaming-aware Neural Network to further support three types of operations: strided convolutions, transposed convolutions, and convolutions with shortcut connections. While the original framework already supports plain convolutions, this extension enables streaming convolutions for U-Net architectures. We omit the implementation details here in the interest of space, as the extended framework is open-source and freely available at .

Latency: Since all convolutions in our model are causal, the architectural latency results entirely from the presence of strided and transposed convolutions: Given that the innermost layers have a temporal resolution 2×2×8×8=2562\times 2\times 8\times 8=256-times coarser than the audio, the decoder can only produce output audio samples in chunks of 256. This translates to a latency of 16ms at 16KHz sampling rate. Profiling the model on a single CPU core of a Pixel 4 mobile phone indicates a processing time of 1.5ms for each 16ms chunk of audio, giving a total latency of 17.5ms.

Experiments

We focus on extension of frequency bandwidth of sub-4KHz up to 8KHz (at a sampling rate of 16KHz), since the quality gain of speech audio beyond 8KHz is marginal in comparison with the 4KHz–8KHz range. We evaluate our model on the widely-used VCTK dataset , using the default training/testing split. All the audio is resampled to 16KHz.

Two objective metrics are used to measure the reconstruction quality of the reconstructed audio with respect to the ground truth:

Scale-invariant signal-to-distortion ratio (SI-SDR) over the audio waveform, which measures the per-sample fidelity up to a uniform scaling factor.

VGG distance, the L2-distance between the ground truth and the reconstructed audio embeddings computed by a pre-trained VGGish network .

Since the VGGish network operates on mel-spectrograms and is trained on a wide range of audio classification tasks, the VGG distance is expected to be less sensitive to per-sample alignment than SI-SDR and more reflective of perceived audio quality. For each model configuration, training is repeated 5 times so as to compute the mean and its standard error.

To assess the robustness of frequency extension, we evaluate on three slightly different input frequency bands: 100–3800Hz (“wide”), 200–3600Hz (“medium”), and 300–3400Hz (“narrow”). Two types of models are trained based on frequency bands of their inputs during training: fixed “medium” band (200–3600Hz) and “variable” band, where the low and high frequency cutoffs are sampled uniformly in the ranges of [0, 300Hz] and [3400Hz, 4000Hz] respectively. The input is produced from the ground truth audio by means of band-pass filtering. For the variable-band model, frequency band sampling and band-pass filtering are performed on-the-fly during training akin to typical data augmentation approaches.

Table 1 shows the average SI-SDR and VGG distance respectively, of applying each model variant on each type of test input. We can see that while the model trained on fixed-band inputs performs slightly better on test inputs of the exact same type, it is highly susceptible to frequency band mismatch. In contrast, the variable-band model is able to maintain good performance over a range of test inputs. This shows that training with variable-band inputs makes the model much more robust.

Surprisingly, the robustness generalizes even to inputs with unseen frequency ranges. To illustrate this, we force the models to “enhance” the ground truth itself. As expected, the fixed-band model fails completely (with SI-SDR below -9dB and VGG distance over 3.3). The variable-band model, however, is able to avoid drastic degradation (maintaining SI-SDR around 15dB and VGG distance around 1.3) despite having never encountered this kind of input during training.

For the rest of the paper, we will assume models to be trained on variable-band inputs and tested on “medium” unless noted otherwise.

2 Effects of Discriminator-based Learning Objectives

3 Comparison with Existing Baselines

We further compare our Streaming SEANet with two baselines, the original offline SEANet and the model by Germain et al. , both of which are more than an order of magnitude more expensive in computational cost, with architectural latency in the hundreds of milliseconds. All models are trained with the same loss functions. Table 2 compares the performance of these models. The results show the streaming model achieves comparable performance to these baselines, despite being much more lightweight and low-latency.

To assess perceptual quality, we perform subjective evaluation using the MUSHRA methodology with 10 VCTK test audio clips (2–5 seconds each). We compare the ground truth and the input (200–3600Hz band-passed), together with Streaming SEANet and the two baselines described above. This results in 10 groups of 5 audio clips, where clips within a group have the same speech content, but vary in quality. To calibrate our 10 raters, we first present them two band-passed examples (score = 50) with their corresponding ground truth (score = 100). We then ask the raters to assign scores between 0 and 100 to each of the 10×510\times 5 clipshttps://google-research.github.io/seanet/freqext/examples/vctk.html. We summarize the results in Figure 2 and confirm the quality ranking Ground truth >> Offline SEANet >> Streaming SEANet >> Germain et al. >> Input (band-passed) is statistically significant, using the Wilcoxon signed-rank test (p-value =0.0038=0.0038 between Offline and Streaming SEANet, and <10−6<10^{-6} everywhere else). We note this ranking is consistent with the quantitative metrics in Table 2, and that there exists a relatively small gap between Streaming SEANet and the offline SEANet.

Conclusion

We presented a wave-to-wave model for frequency bandwidth extension, and showed that it can be made robust to unexpected out-of-distribution inputs by varying input bandwidth during training. Our model is lightweight and fully causal, and yet attains quality levels comparable to those of much larger offline models. With our extended support for streaming convolution, the model is capable of processing live audio in real time on ordinary mobile devices.

References