HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec
Dongchao Yang, Songxiang Liu, Rongjie Huang, Jinchuan Tian, Chao Weng, Yuexian Zou
Introduction
The purpose of an audio codec is to reduce the amount of data needed to store or transmit an audio signal without significantly degrading its quality. The basic principle behind audio codecs is to remove redundant or irrelevant information from the audio signal. There are many different types of audio codecs, each with its own strengths and weaknesses. Some codecs are designed for use in real-time applications, such as telephony or streaming audio, and prioritize low latency and low bitrates. Others are designed for high-quality audio production and prioritize fidelity and accuracy. In addition to compression, many audio codecs also include features such as error correction, noise reduction, and dynamic range compression. These features can help to improve the quality and reliability of the audio signal, especially in challenging environments such as noisy or low-bandwidth networks.
In this study, we focus on using audio codec models to help solve generation problems in audio-related works, such as Text-to-Speech (TTS), music generation , audio generation . The Encodec and SoundStream are the most related to our work. The core technique in Encodec and SoundStream is residual vector quantization (RVQ), which uses multiple VQ codebooks to represent the intermediate features. Encodec and SoundStream both adopt an encoder-decoder framework. The encoder first compresses the waveform into compact deep representations, then RVQ is used to quantize the intermediate features. Lastly, the decoder is used to recover waveform from the quantized representations.
In our experiments, we found that the most of information is saved in the first codebook when RVQ is used, e.g. for speech data, text information, and timbre information can be recovered by only using one codebook. The following codebooks save some details information, which may influence the audio quality. However, such information is sparse and scattered in the hidden space. Thus, such a quantization style needs many codebooks to realize good reconstruction performance. e.g. Encodec needs 12 codebooks to realize high-quality reconstruction performance. In audio generation fields, using multiple codebooks will bring burden to the generation model, e.g. long sequence is hard to model by the transformer.
Related works
The usage of self-supervised learning (SSL) has got great success in audio-related fields, such as auto-speech recognition (ASR) and audio compression . Inspired by vector quantization (VQ) techniques , SoundStream presents a residual vector quantization (RVQ) architecture for high-level representations that carry semantic information. Similarly, many works try to use VQ-VAE model to compress the time-frequency spectrogram (e.g. Mel-spectrogram) into high-level representations.
2 Speech and Audio Generation with Audio Codec
Recently, many works propose to model speech and audio in the discrete latent space with the help of Audio Codec. The core idea is that using an audio codec model compresses the speech or sound into a group of discrete tokens, then uses a generation model to generate these tokens. e.g. AudioLM utilize the audio codec model encodes the waveform into discrete tokens, and then uses Language Model (LM) to model the generation process. Similarly, InstructTTS uses discrete diffusion models to generate discrete tokens, then uses an audio codec model to recover the waveform.
3 Audio Codec
The study of low-bitrate parametric audio codecs dates back to , but their quality is often limited. Recently, researchers have proposed several neural network-based audio codecs that show promising results . These methods typically use an encoder to extract deep features in a latent space, which is then quantized before being fed to the decoder. The most relevant related works to ours are the SoundStream and Encodec models, where these methods propose to use a fully convolutional encoder-decoder architecture with a Residual Vector Quantization (RVQ) layers. These models were optimized using both reconstruction loss and adversarial perceptual losses. To accelerate the process of compression and decompression, Encodec proposes a language model method to predict tokens. MQ-TTS also proposes to use multiple codebooks to quantize intermediate features, but they assume speaker information can be explicit from the speaker labels and do not use residual VQ to preserve more audio information. Although previous works have got great success in terms of reconstruction performance and compression rate, these works may not be suitable for generation tasks due to many codebooks being needed to maintain good reconstruction performance.
Proposed Method
In this section, we will introduce the details of HiFi-Codec. We first introduce the overview of HiFi-Codec model, then we discuss the details of each part in HiFi-Codec.
We consider a single-channel audio signal with duration , represented as a sequence , where and is the audio sample rate. The HiFi-Codec model comprises three main components: (1) an encoder network that takes the input audio and generates a latent feature representation ; (2) a group-residual quantization layer that produces a compressed representation ; and (3) a decoder that reconstructs the audio signal from the compressed latent representation . The model is trained end-to-end, optimizing a reconstruction loss applied over both time and frequency domains, along with a perceptual loss in the form of discriminators operating at different resolutions. A visual description of the proposed method can be seen in Figure 1.
2 Encoder and Decoder
Our model’s encoder and decoder architecture draws inspiration from the designs of Encodec and SoundStream . The architecture is based on a convolutional framework with sequential modeling applied over the latent representation on both the encoder and decoder sides. The encoder model comprises a 1D convolution with channels and a kernel size of 7, followed by convolution blocks. Each convolution block features a single residual unit, which is followed by a down-sampling layer consisting of a strided convolution with a kernel size of twice the stride . The residual unit consists of two convolutions with a kernel size of 3 and a skip connection. The number of channels is doubled whenever down-sampling occurs. The convolution blocks are then followed by a two-layer LSTM for sequence modeling and a final 1D convolution layer with a kernel size of 7 and output channels. In our study, we explore different settings for , , and , such as , , and or (2, 4, 5, 6) or (2, 2, 2, 4). The decoder mirrors the encoder and uses transposed convolutions instead of stride convolutions, with the strides in reverse order as in the encoder. The decoder outputs the final audio signal.
3 Group-residual Vector Quantization (GRVQ)
In this study, we want to design an audio codec model that contains fewer quantizers while enjoying good reconstruction performance. We think that one of the drawbacks of the RVQ is that the first layer of codebooks in RVQ will save the most of information, but the remaining codebooks only save a little information. Thus we propose to add more codebooks in the first layer. Specifically, for any latent feature representation , we first split into several groups averagely (in our study, we split into two groups, ), and using multiple RVQ to quantize each group features. Lastly, we combine multiple group RVQ output to obtain the final quantization results. The whole process can be summarized as Algorithm 1.
4 Discriminator
In this study, we use three discriminators: A multi-scale STFT-based (MS-STFT) discriminator, which is used on Encodec ; a multi-period discriminator (MPD) and a multi-scale discriminator (MSD) from HiFi-GAN vocoder . For the MS-STFT discriminator, which consists in identically structured networks operating on multi-scaled complex-valued STFT with the real and imaginary parts concatenated. We adopt the same configuration as Encodec for each sub-network, which consists of a 2D convolutional layer followed by 2D convolutions with increasing dilation rates in the time dimension (1, 2, and 4) and a stride of 2 over the frequency axis. A final 2D convolution with a kernel size of 3 x 3 and stride (1, 1) provides the final prediction. We use five different scales with STFT window lengths of . For the multi-period discriminator and multi-scale discriminator, we maintain the same structure as HiFi-GAN, but reduce the channel number to make the discriminator have similar parameters to MS-STFT.
5 Training Loss
Our approach is based on a GAN objective, in which we optimize both the generator and the discriminators. Specifically, we jointly optimize a reconstruction loss term, a perceptual loss term (via discriminators), and the GRVQ commitment loss for the generator. The training objective of the generator comprises several loss terms, including a time domain term, a frequency domain term, three discriminator losses, and the corresponding feature loss terms acting as a perceptual loss and the GRVQ commitment loss. The discriminator loss is based on the adversarial hinge-loss function.
Our reconstruction loss comprises two aspects: (1) time domain loss and (2) time-frequency loss. For the time domain loss, we directly use the L1 distance loss to optimize and . For the time-frequency loss, we follow a similar approach to Encodec and apply a loss term on the mel-spectrogram with several time scales.
5.2 Discriminator loss
The adversarial loss is used to promote perceptual quality. We use three types of discriminator, where MS-STFT discriminator try to make the spectrogram-level reconstruction results as similar as the original one. MPD and MSD discriminators try to make the waveform-level reconstruction results as similar as the original one. To train the discriminator, we can optimize the following objective function:
where denotes the number of discriminators. Furthermore, we can define the adversarial loss as a hinge loss over the logits of these discriminators:
Furthermore, the feature loss is computed by taking the average absolute difference between the discriminator’s internal layer outputs for the generated audio and those for the corresponding real audio.
5.3 GRVQ Commitment Loss
For the i-th group c-th residual quantizer, we can calculate the commitment loss based on following formula:
Based on previous discussion, we can use following formula to train the generator.
where denotes the adversarial loss. denotes the feature loss. denotes the reconstruction loss. , , and are the hyper-parameters to control the training objective function. In our experiments, we try to balance each loss terms by scale these hyper-parameters.
Experiments
In this study, we evaluate the audio codec model’s performance by measuring the gap between reconstruction audio and the target one. We adopt the metrics from speech enhancement fields, such as the PESQ and STOI to evaluate the performance.
2 Dataset
We use TTS dataset to train audio codec models. Our training data comes from public datasets, such as LibriTTS, VCTK, AISHELL, which mainly includes English and Chinese speech.
3 Experimental results
Table 1 shows the experimental results. We can see that our proposed HiFi-Codec realizes good reconstruction performance while only using 4 codebooks. The best performance can be abtained when we set downsample times as 240, and the number of codebooks as 8 (we set the each layer includes 4 codebooks, and two residual layers are used). Furthermore, our reproduced models (Encodec and SoundStream) also get comparable performance with Encodec . We strongly recommend readers to use the HiFi-Codec model with 4 codebooks when readers try to train a generation model.
Conclusion
In this study, we present group-residual vector quantization method, and build a novel audio codec model: HiFi-Codec, which is specially designed for generation tasks. HiFi-Codec can bring better reconstruction performance than Encodec even using 4 codebooks. Furthermore, we also release the training process of Encodec and SoundStream models, which can help readers to train their own codec models. In the future, we will continue to optimize the HiFi-Codec models, and try to train better Encodec and SoundStream models. We expect this project can facilitate the research in audio generation tasks.
Limitations
Although HiFi-Codec models realize good construction performance than Encodec model, the limitations still exist. (1) We do not use large-scale dataset to train a universal audio codec, the generalization cannot be validated very well. (2) We find that the objective evaluation metrics may not very accurate to assess the reconstruction performance. Subjective evaluation is always the best choice, but this part is missed in this study. (3) HiFi-Codec aims to help generation tasks, but we donot provide enough down-stream tasks to evaluate the performance. We take this direction to our future works.