LMCodec: A Low Bitrate Speech Codec With Causal Transformer Models

Teerapat Jenrungrot, Michael Chinen, W. Bastiaan Kleijn, Jan Skoglund, Zalán Borsos, Neil Zeghidour, Marco Tagliasacchi

Introduction

Speech coding, which consists of compressing speech signals to a limited number of bits with minimal distortion, is at the core of communication technologies such as mobile telephony or Voice over IP (VoIP). Opus and EVS are state-of-the-art speech coding techniques that combine traditional coding tools, such as Linear Predictive Coding (LPC), Code Excited Linear Prediction (CELP), and Modified Discrete Cosine Transformation (MDCT) to achieve high coding efficiency over different content types and bitrates. These waveform and parametric codecs rely on psychoacoustics expertise to design signal processing pipelines with maximal coding efficiency. Yet, while fast and interpretable, such handcrafted pipelines only represent a fraction of the potential models for a speech codec.

This has motivated data-driven approaches to train neural networks to perform speech coding. These networks leverage large amounts of training data while relaxing the assumptions made on the type of transformations applied by the system . In particular, the SoundStream neural codec combines a causal convolutional architecture with a residual vector quantizer. This quantization method produces a hierarchy of coarse-to-fine codes, and allows for efficient compression while providing bitrate scalability. As a result, SoundStream at 3  3\text{\,}~{}kbps matches the quality Opus at 12  12\text{\,}~{}kbps. However, the quality of most codecs, be they handcrafted or trained, degrades significantly at bitrates lower than 3  3\text{\,}~{}kbps.

In this work, we introduce LMCodec, a low bitrate speech codec that combines recent advances in neural audio coding and audio generative modeling. LMCodec uses autoregressive Transformers on SoundStream tokens to (i) model the entropy of the distribution of coarse tokens and (ii) predict fine tokens from the coarse ones. At inference, LMCodec extracts the codes of a SoundStream model from the input waveform. However, instead of sending all codes to the receiver like a SoundStream codec would do, LMCodec only transmits entropy-coded coarse tokens. On the receiver side, a generative language model is used to predict fine tokens from the coarse ones, and a SoundStream decoder then reconstructs audio from the complete token sequence.

LMCodec takes inspiration from the AudioLM generative model, which also predicts fine SoundStream tokens from coarse ones. However, unlike AudioLM, LMCodec does low bitrate compression rather than generative modeling, and to do so leverages AudioLM both as a generative model and an entropy model. Other Transformer-based models for low bitrate coding have been proposed . The codec in enriches SoundStream with embeddings extracted from a self-supervised speech representation model and achieves speech compression at a rate of 600  600\text{\,}~{}bps. synthesizes speech from a combination of phonetic, pitch and speaker representations to achieve 365 bps. Unlike these models, LMCodec is a fully causal model, which is thus amenable to online encoding and decoding. Our primary contribution is the design of a new neural speech codec, which achieves state-of-the-art results outperforming many previous codecs operating at three to four times the rates according to subjective human evaluation metrics.

Subjective evaluations demonstrate how LMCodec allows for low bitrate speech coding with minimal distortion, with LMCodec at approximately 1 1\text{\,}-1.5  1.5\text{\,}~{}kbps matching the performance of Opus at 12  12\text{\,}~{}kbps. We furthermore analyze the failure modes of our system, as well as the discrepancies in bit allocations between speech and non-speech sections of an audio signal.

Proposed Model

In this section, we describe our proposed speech codec consisting of four components: an encoder, a residual quantizer, an AudioLM block, and a decoder. The encoder, residual quantizer, and decoder follow similar structures from SoundStream. At the very high level, the encoder takes raw speech in the time domain as an input and extracts low-rate features that contain sufficient information to reconstruct the speech. The residual quantizer finds discrete representations of the inherently continuous encoded features. AudioLM poses the modeling of the quantized discrete representation as a language modeling problem and estimates the probability distribution of the next discrete audio token given previous audio tokens. Finally, the decoder reconstructs the input speech signal from the discrete encoded features.

We now briefly describe the SoundStream model that we used for creating high-quality audio tokens.

1.2 Residual Vector Quantizer (RVQ)

1.3 Decoder

2 AudioLM

For the rest of this paper, let NCN_{{\mathcal{C}}} and NFN_{{\mathcal{F}}} denote the number of quantizers for the coarse-level and fine-level AudioLMs, respectively. Figure 1 shows the overall architecture of our proposed model, in which we use NC=4N_{{\mathcal{C}}}=4 and NF=8N_{{\mathcal{F}}}=8. In our experiment, we use various combination of (NC,NF)(N_{\mathcal{C}},N_{\mathcal{F}}) ranging from NC+NF=3N_{\mathcal{C}}+N_{\mathcal{F}}=3 to NC+NF=24N_{\mathcal{C}}+N_{\mathcal{F}}=24. Additionally, let ck(n)c^{(n)}_{k} denote the SoundStream token at frame nn and VQ layer kk.

The goal of the coarse-level AudioLM is to model the distribution of the next coarse SoundStream tokens. Specifically, we are interested in modeling the conditional distribution of the next SoundStream tokens given the past information

Given the distribution of the future SoundStream tokens, we build a codec by using lossless Entropy Coding (Section 2.3). More specifically, the discrete probability distribution of SoundStream tokens can be estimated both at the sender and the receiver sides, and we use this to drive an entropy codec. Note that in our proposed method, we only need to transmit NCN_{{\mathcal{C}}} tokens per single audio frame. The remaining NFN_{{\mathcal{F}}} tokens are generated at the receiver side only as described in the next section.

2.2 Fine-level AudioLM

Similar to the coarse-level AudioLM, the fine-level AudioLM predicts the top VQ layers given the information about bottom VQ layers in addition to the past information. Specifically, we are interested in modeling the distribution of the fine-level SoundStream tokens conditioned on the coarse-level tokens and the past information:

for k∈{NC+1,…,NC+NF}k\in\{N_{{\mathcal{C}}}+1,\ldots,N_{{\mathcal{C}}}+N_{{\mathcal{F}}}\}. Note that our model is causal, in contrast to AudioLM.

Since we only transmit the coarse-level tokens, we model the distribution of the fine-level tokens by assuming that we have access to ground-truth coarse-level SoundStream tokens. We note that, while also proposes a similar fine-level AudioLM stage, our contribution here is the causal formulation of the task, which makes our approach more suitable and amenable to online decoding.

3 Entropy Coding (EC)

Given the distribution of coarse-level SoundStream tokens, we transmit data by using entropy coding, a lossless data compression technique. In this work, we provide experimental results using Huffman coding, in addition to the estimated entropy rate. We treat each code from the residual VQs separately and do not perform any grouping to reduce the upper bound on the bitrate.

We first note that our proposed codec requires only sending coarse-level SoundStream tokens using entropy coding. Specifically, given raw audio, LMCodec first encodes audio into SoundStream tokens and models the probability distribution of the next SoundStream tokens, driving the entropy codec. Note that the discrete probability distribution of SoundStream tokens can be estimated both at the sender and the receiver sides, so the receiver can losslessly reconstruct the coarse tokens. To generate audio output from only coarse-level tokens, we use a fine-level AudioLM to synthesize fine-level tokens from the transmitted coarse-level tokens and then generate audio from both coarse-level and fine-level tokens using SoundStream decoder.

4 Training Strategy

We adopt a 2-stage training paradigm. First, we train only the encoder, quantizer, and decoder. Then, we freeze the weights of these components and train only the AudioLM components. We train the coarse-level and fine-level AudioLM models separately.

We trained the SoundStream model using the standard adversarial loss, feature matching loss, reconstruction loss, and quantization loss according to . In training AudioLM models, we use the standard cross-entropy loss for language modeling over the vocabulary space.

4.2 Training configurations

To create our codec modules, we adapted the architectures of the encoder, quantizer, generator, and discriminators used in SoundStream and AudioLM from T5X. Both AudioLM models are the decoder-only models based on the base model of t5.1.1 (with approximately 250 million parameters).

We trained multiple coarse-level and fine-level AudioLM models to achieve varieties of bitrates. The bitrates are calculated based on the entropy coding of codes from coarse-level AudioLM.

Evaluation

To demonstrate the performance of our proposed method, we evaluate LMCodec using both objective and subjective evaluations. For objective evaluation, we report the accuracy of LMCodec future token prediction and objective metrics including ViSQOL , WARP-Q , SSL-MOS , WER, and CER together with bitrate based on the test split from the clean LibriSpeech dataset .

For subjective evaluation, we perform two MUSHRA-like subjective tests to compare the audio quality with standard state-of-the-art speech codecs at medium bitrate (i.e., 1  1\text{\,}~{}kbps to 12  12\text{\,}~{}kbps) and low rate (i.e., 0.5  0.5\text{\,}~{}kbps to 1.5  1.5\text{\,}~{}kbps). The tests were conducted respectively on 91 and 94 crowd-sourced raters using headphones over 32 clean utterances from VCTK dataset . Raters who did not score the reference above 80 at least 80% of the time were discarded, as were raters who rated more than 75% of non-reference samples 80 or above. 40 raters for the medium rate test and 33 raters for the low rate test met this requirement.

As shown in Figure 2, the raters found that LMCodec-4/6 with 4 quantizers at 1.1  1.1\text{\,}~{}kbps perform significantly better than 12  12\text{\,}~{}kbps Opus. LMCodec-8/12 with 8 quantizers at 2.6  2.6\text{\,}~{}kbps has comparable performancce to SoundStream at 6  6\text{\,}~{}kbps. The low-rate MUSHRA test compares recent transformer neural codecs and lower bitrate SoundStream models. The raters preferred LMCodec to the transformer models from and SoundStream at the same rate.

Table 1 shows the accuracy of the future token prediction and the bitrate performance of LMCodec from the test split of the clean LibriSpeech . For accuracy, we note that perfect accuracy means the model knows perfectly what the next tokens are. In the context of fine-level AudioLM, this suggests that the model does not necessarily need to synthesize the correct code to produce reasonable audio output. The bitrates are computed based on the future token’s distributions obtained from LMCodec. For Huffman coding, we use the ground truth tokens encoded with the Huffman algorithm. Additionally, we note that the distributions of future tokens are updated every timestep based on the model, different from how other entropy codecs that may have fixed distributions operate. So, the Huffman bitrate may sometimes be lower than the bitrate derived from the entropy.

In this section, we additionally discuss some of the interesting audio effects from LMCodec. We suggest that readers listen to some of the audio samples from our model. In particular, our model with only one quantizer is able to produce reasonable human voice with some babbling effects. The amount of babbling is reduced as the number of quantizers used in the codec increases. This suggests that there are some underlying hierarchical structure in SoundStream tokens, and the proposed codec can potentially be operating at very low bitrate, given that the coarse-to-fine prediction is accurate.

In Figure 4, we visualize the distribution of code prediction from the AudioLM model when the input is at the middle of a phoneme and between phonemes. We also found that the model is very confident if the audio input is the middle of the phonemes, as the language model network is able to learn underlying linguistic behavior of the utterances. On the other hand, the model has lower confidence in predicting the next token when reaching silence sections, suggesting that our proposed causal model is unable to predict future word really well. This confirms the babbling effect that we observed in the audio output from our proposed codec, which increases as we restrict the amount of information to describe each frame (e.g., by transmitting fewer codes or dropping frames).

Figure 2 shows the comparison of LMCodec with low-rate and medium-rate audio codecs. In particular, we find that LMCodec-4/6 performs better than SoundStream with 3 quantizers at 1.5  1.5\text{\,}~{}kbps but slightly worse than SoundStream with 12 quantizers at 6  6\text{\,}~{}kbps which is on par with LMCodec-8/12. We note that LMCodec-4/6 and LMCodec-8/12 are based on SoundStream with 6 and 12 quantizers respectively. Our results suggest that LMCodec effectively takes advantages from entropy coding and synthesizing reasonable fine-level codes from coarse-level codes. When comparing with SoundStream at similar rate, LMCodec essentially outperforms.

2 Voice Activity Detection (VAD)

Table 2 shows the bitrate of LMCodec on two scenarios: (i) transmitting only voices and (ii) transmitting entire speech signals but using zero bits for non-voices. We report the bitrate derived from the entropy and the bitrate based on Huffman coding. We note the first scenario has slightly lower bitrates as compared to bitrates from Table 1 because the entropy for non-speech signals is usually higher than the entropy for speech signals. Additionally, the second scenario provides the lower bound estimate of bitrates when transmitting very low bits for non-voice signals similar to Opus with variable bitrate scheme.

3 Objective Evaluation

We present an objective evaluation on the audio examples from VCTK dataset in Figure 3. First, we demonstrate that the word error rate (WER) and character error rate (CER) are decreasing as the number of quantizers used in the LMCodec increases until around 4-6 quantizers, suggesting that the semantic content is stored in the coarse tokens. To evaluate WER and CER, we use two ASR models from AWS Transcribe service and Conformer model trained on LibriSpeech . Second, ViSQOL and WARP-Q , metrics designed for neural speech codecs, increases and decreases respectively, implying that the fine tokens are responsible for fine-grained acoustic details. Third, SSL-MOS shows that the overall speech quality improves by increasing the number of quantizers.

Despite neural speech codecs metrics ViSQOL and WARP-Q indicating worse performance at about 4-6 quantizers, our listening test shows very high quality audio results with small number of quantizers. This suggests that the language model of LMCodec is able to model the distribution of the fine tokens given the coarse tokens reasonably well even if the synthesized fine tokens are different from the ground truth ones. This drives metrics like ViSQOL and WARP-Q down as they primarily rely on the comparison between synthesized audio and its corresponding ground truth reference audio.

When comparing LMCodec with different total number of quantizers, we first note that the upper bound performance of LMCodec with 6 quantizers is lower than the upper bound performance of LMCodec with 12 or 24 quantizers. However, LMCodec with a lower total number of quantizers reaches better performance faster than LMCodec with a higher total number of quantizers.

Conclusion

Our experiments show that the proposed codec significantly outperforms the original neural speech codec with respect to the quality of synthesized speech when operating in the ultra-low bitrate regime. In addition, the subjective experiments indicate comparable to or better perceptual speech quality compared to conventional codecs operating at higher rates.

References