Attentive Temporal Pooling for Conformer-based Streaming Language Identification in Long-form Speech

Quan Wang, Yang Yu, Jason Pelecanos, Yiling Huang, Ignacio Lopez Moreno

Introduction

Language Identification (LangID / LID) is the task of automatically identifying the spoken language of a digitized speech utterance , which has been widely used in various modern speech processing systems. Common applications include automatic call routing, multilingual speech transcription, multilingual speech translation, and content-based audio retrieval . Particularly, a commercially viable application in recent years is the interactive voice assistive system, where LangID can be used to automatically select the output from several automatic speech recognition (ASR) models that run in parallel, such that the user is able to interact with the system with multiple different languages . Additionally, there is an increasing research interest in multilingual ASR systems that are suitable for code-switching applications. For such applications, LangID can either be jointly trained with the ASR model , or provide auxiliary input to the ASR model .

In this paper, we are particularly interested in applications where language identification is used for analyzing streaming long-form audio, such as podcast, telephony speech, and video streaming. Language identification results produced at runtime can be used for various purposes, such as triggering ASR captioning with the right language, context-based indexing and searching, triaging the content for human auditors, recommending the content to the most relevant audience, and delivering advertisement in the right language. Such applications usually have these requirements:

Low latency. In streaming applications, it is critical to produce the language prediction signal accurately in a timely manner to be consumed by downstream components such as ASR and natural language understanding (NLU). Thus some non-causal model topologies, such as bi-directional LSTM, can be detrimental.

No recency bias. Long-form speech often allows the model to predict the language based on a relatively long context. However, recurrent neural networks such as LSTM often suffer from a bias towards short-term context .

Noise robustness. Long-form utterances are often interleaved with silence, non-speech, as well as speech segments from various sources, including different speakers, difference devices, different reverberant conditions, and different signal-to-noise ratio (SNR). The non-speech segments and noisy speech segments may cause degradation of the overall language prediction accuracy if we simply average the results from all segments.

Parallelization. As deep learning accelerators become available in more and more cloud computing services as well as consumer devices , neural networks that can perform inference in parallel will better benefit from the computational power of such hardware. In many scenarios, such as high-load services and on-device applications, parallelizable models are largely preferred over non-parallelizable models.

Based on these requirements, we introduce a language identification model based on conformer layers and an attentive temporal pooling mechanismAn open source implementation of the attentive temporal pooling module based on Lingvo is provided at: https://github.com/google/speaker-id/tree/master/lingvo. The inference of this model runs in a streaming fashion, and the conformers can be easily parallelized on accelerator hardware to minimize the latency. The attentive temporal pooling mechanism computes an attention weight for each temporal step of the conformer model, and uses a weighted moving average to produce an aggregated embedding at each step. This mechanism helps the model to better attend to speech segments that contain more information related to the spoken language.

At the same time, different application domains may have (i) different prior distributions of languages and (ii) different data properties. Here we study two simple domain adaptation approaches that can improve the empirical performance of the LangID model in various application domains without retraining the neural network. This significantly reduces the development cost of LangID models while achieving improved performance for each application.

The rest of this paper is organized as follows. In Section 2, we will briefly review existing works that are related to our LangID system. In Section 3, we will introduce our LangID system in detail, including the feature frontend in Section 3.1, the data augmentation strategy in Section 3.2, the model topology in Section 3.3, the attentive temporal pooling mechanism in Section 3.4, and the domain adaptation methods in Section 3.5. We describe our experiments and results in Section 4, and present our conclusions in Section 5.

Related Work

LangID is also often referred to as Spoken Language Recognition (SLR) to avoid confusion with text-based language identification systems . While multiple systems have been proposed to exploit the acoustic, phonetic, morphologic and semantic level representations of the utterances , most common LangID systems today directly take acoustic features of the utterance as input, such as Mel-frequency cepstral coefficients (MFCC) or log Mel-filterbank energies (LFBE).

Since the proposal of one of the earliest HMM-based LangID systems in 1977 , various models and approaches have been explored. Unlike speech recognition and speech synthesis, where the task can be viewed as a “sequence transduction” problem, LangID is usually considered as a “sequence summation” problem. Thus many LangID systems are largely inspired by speaker recognition systems. For example, i-vector based LangID systems had been very popular in early 2010s .

As deep learning has gained in popularity, most modern LangID systems are based on neural networks. Deep feed-forward neural network (DNN) based LangID models had been proven significantly more accurate than i-vector based models . Convolutional neural networks (CNN) and recurrent neural networks (RNN) such as LSTM had also been explored to reduce model sizes as well as improve LangID accuracy .

Conformers are convolution-augmented transformers that are used as building blocks in many different deep learning systems . It was originally proposed for speech recognition, but has found application in many other tasks such as speech enhancement , speech separation , speaker diarization , and sound event detection . In , a conformer-based feature extractor is used for joint ASR and LangID training. In , conformer-based models are first pretrained for ASR tasks, then used for transfer learning on various language recognition tasks . We chose conformers as the basic building block of our LangID system due to its organic combination of convolution and multi-head self-attention, the demonstrated accuracy improvements, and its ability to parallelize on accelerator hardware.

Various works have examined statistics pooling and attention modeling for speaker recognition, where the focus is mostly on the utterance level analysis. However, for real-time systems where latency is critical, there is a need to provide intermediate inference outputs in an online fashion. In this paper, we share a straightforward approach for generating these statistics using either segment level or per-frame recursion analysis for LangID.

System Description

In the feature frontend of our LangID system, we first apply automatic gain control to the input audio, then extract 32ms Hanning-windowed frames with a step of 10ms. For each frame, 128-dimensional log Mel-filterbank energies (LFBE) are computed in the range between 125Hz and 7500Hz. These filterbank energies are then stacked by 4 frames and subsampled by 3 frames, resulting in final features of 512 dimensions with a frame rate of 30ms.

2 Data augmentation

To make sure our LangID model is robust against various acoustic environments and noise conditions, we apply data augmentation during the training of the LangID model. We randomly apply multi-style training (MTR) to part of the training utterances with an SNR ranging from 5dB to 25dB. The noise source consists of ambient noises recorded in cafes, kitchens, vehicles, and quiet environments, as well as audio clips of music and sound effects downloaded from the YouTube Audio Libraryhttps://youtube.com/audiolibrary and Getty Imageshttps://www.gettyimages.com/about-music. The room configurations consist of 24 million convolutional room impulse responses generated by a room simulator .

We also apply SpecAugment to the rest of the training utterances where MTR did not apply. We found this practice pretty useful, as applying both MTR and SpecAugment to the same utterance can easily lead to over-augmentation.

3 Model topology

The diagram of our conformer-based LangID model topology is shown in Fig. 1. After performing data augmentation followed by feature extraction, we add absolute positional encodings to these features, and feed them to the conformer encoder, which has a stack of 12 conformer layers. Similar to the original conformer paper , we explore three different model sizes, where the dimensionality of each layer is 144 (small), 256 (medium), and 512 (large), respectively. Each layer has a multi-head self attention with 8 heads. The 1-D depth-wise convolutional components span 32 elements. Additionally, we perform a stack-by-2 then subsample-by-2 operation on the third conformer layer output to reduce inference cost, and insert a non-linear projection layer with output dimension of 144/256/512 (for S/M/L, respectively) after the fourth conformer layer. This means each inference step would be of size 2 (every 2 frames), covering a receptive field of about 0.06 seconds of the input. This conformer encoder generates a 144/256/512 (for S/M/L, respectively) dimensional output for each input frame. The forward propagation calls of individual layers within the conformer network for different temporal positions are independent of each other. Thus they can be batched to run in parallel on accelerator hardware during both training and inference time.

The outputs of these conformer layers will be sent to the attentive temporal pooling layer to produce a weighted mean vector, which will be optionally concatenated with a weighted standard deviation vector. Then this weighted vector will be fed into a feed-forward network, consisting of a 256-dim layer with ReLU activation, and a linear projection whose output dimension equals the number of candidate languages. Finally, a softmax layer is applied to produce the probability distribution of different language candidates.

4 Attentive temporal pooling

We introduce a weighted moving average implementation to allow the attentive temporal pooling mechanism to run in an online fashion. Assume that at time step tt, the conformer layers produce an output embedding ht\mathbf{h}_{t}. The attentive temporal pooling module will compute an attention weight wtw_{t} based on the current embedding:

The function fatt(ht)f_{\text{att}}(\mathbf{h}_{t}) is calculated as the sigmoid activation of a linear transform of ht\mathbf{h}_{t}. A small value (ϵ=0.0001\epsilon=0.0001) is added to ensure numerical stability.

The following sufficient statistics (counts, sums and sums of squares) are tracked at time tt:

Note here and elsewhere that (⋅)2(\cdot)^{2} is used to denote the element-wise square.

Then the weighted mean μt\boldsymbol{\mu}_{t} and weighted standard deviation σt\boldsymbol{\sigma}_{t} can be calculated only using the sufficient statistics at time tt:

To build the attentive temporal pooling as part of our neural network inference graph, and to handle streaming inference, we utilize the following recurrent form of the sufficient statistics in our implementation:

Note that they can be initialized as η0=0\eta_{0}=0, A0=0\boldsymbol{A}_{0}=\boldsymbol{0} and Q0=0\boldsymbol{Q}_{0}=\boldsymbol{0}. This representation allows us to incrementally calculate {μt,σt}\{\boldsymbol{\mu}_{t},\boldsymbol{\sigma}_{t}\} from the previous sufficient statistics (state variables in the inference graph) {ηt−1,At−1,Qt−1}\{\eta_{t-1},\boldsymbol{A}_{t-1},\boldsymbol{Q}_{t-1}\} and the current intermediate network output frame ht\mathbf{h}_{t}. The streaming model in this paper uses this recurrent formulation interpretation.

This attentive temporal pooling mechanism can be easily extended to a multi-head version, where each output embedding ht\mathbf{h}_{t} produces multiple weights, and we concatenate multiple weighted mean and weighted standard deviation vectors as the input to the feed-forward layers. However, in our LangID experiments, multi-head attentive temporal pooling has very similar performance with single-head attentive temporal pooling.

5 Domain adaptation

Training a neural network model for the LangID system described above can be very expensive, especially when the model is trained on massively multilingual datasets. For simplicity, when we deploy the LangID system to different applications, it is usually desired that the same neural network model is deployed to all applications, while each application has the ability to adjust the output probabilities to make it consistent with the application specific class prior probabilities and perhaps data differences.

For example, assume we want to deploy a LangID model to identify the language in videos for two different video hosting websites. It is possible that most videos on website A are in English, while most videos on website B are in Spanish. At training time, we could include datasets from both websites to make our model robust. However, when deploying the model, we want the system to incorporate such prior information and additional data related differences.

One approach is the use of a Gaussian backend classifier applied to i-vectors or x-vectors . In this work, we explore basic approaches of what can be done if only the model output probabilities are available. For example, the results are generated by a cloud service. Here we describe two domain adaptation approaches: prior replacement and a discriminatively trained output transform. They do not require updating the LangID neural network itself, thus making it possible to deploy a unified inference backend with one LangID model to handle requests from different domains, as shown in Fig. 2.

In the first approach, we re-estimate the posterior probabilities of the languages for a specific application, using the new class prior probabilities given the new target domain data. Suppose we have a set of randomly drawn samples from the target application data. We then count the number of samples drawn for each of the KK languages {Li}i=1,…,K\{L_{i}\}_{i=1,\ldots,K}. Let cic_{i} be the number of samples related to language LiL_{i}. A smoothed estimate of the class prior probabilities may be given as :

where RR is the prior data relevance count. (Empirically, when the sample size is N=10000N=10000, we found R=4R=4 to work well.)

We treat the softmax outputs from the LangID neural network as the probabilities of the languages given the input utterance. The LangID neural network was trained with equal class prior probabilities. Given the new application specific prior probabilities, and a neural network trained on uniform class prior probabilities, the posterior probability of a language given the new application target and utterance XX is computed as :

Here, DoldD_{old} represents the information regarding class priors built into the existing model trained on the old data. In this work, the language class priors are uniform and this gives a simplified version of the result in because the common values cancel. DnewD_{new} represents information regarding the class priors for the new application. The term P(Li∣Dnew)P(L_{i}|D_{new}) is the new estimate of the application specific class prior probabilities, and P(Li∣X,Dold)P(L_{i}|X,D_{old}) is the probability generated by the softmax of the neural network based on the old equal class prior assumptions.

5.2 Output transform

In the second approach, we assume a dev-set from the new target domain is available, and we optimize a transform on this dev-set.

Assume that the original output of the LangID model is denoted as a probability distribution over KK languages p∈K\mathbf{p}\in^{K}. For the new domain, let a\mathbf{a} and b\mathbf{b} be two KK-dimensional vectors. We transform p\mathbf{p} into a new probability distribution:

where ⊙\odot denotes element-wise multiplication.

Given a dev-set of NN samples from this specific domain (N=10000N=10000 in our experiments), we optimize {a,b}\{\mathbf{a},\mathbf{b}\} to minimize a regularized cross entropy between p~\widetilde{\mathbf{p}} and the ground truth language y\mathbf{y} on this dev-set:

where LcentL_{\text{cent}} denotes cross entropy loss and wregw_{\text{reg}} is the weight for the regularization term.

Once we have optimized {a,b}\{\mathbf{a},\mathbf{b}\} on the dev-set, we can deploy Eq.12 as a post-processing step together with the LangID model without modifying the model itself as shown in Fig. 2. At the same time, while the LangID model is trained on batched short segments for efficiency, Eq. 13 is based on long-form inference outputs. This can better compensate for the duration gap between training and inference.

Although Eq. 12 can be replaced by more complicated forms, such as using a full K×KK\times K transform matrix instead of the vector a\mathbf{a}, or even using a tiny neural work, such approaches can easily lead to overfitting on the dev-set, especially that the dev-set is usually much smaller than the training set. We found that adaptation with only 2K2K parameters (i.e. a\mathbf{a} and b\mathbf{b}) effectively improves in-domain performance without overfitting the dev-set.

Interestingly, if a\mathbf{a} is fixed as a vector of ones, the optimization result is closely related to the prior replacement method in Eq. 11. The key differences are that parameters are trained discriminatively instead of directly using the class priors from the dev-set, and the regularization/constraints are different.

Experiments

Our model is trained to distinguish between 65 different languagesThe list of the 65 languages: Afrikaans, Amharic, Arabic, Azeri, Belarusian, Bulgarian, Bengali, Catalan, Chinese, Czech, Danish, German, Greek, English, Spanish, Basque, Farsi, Finnish, Filipino, French, Galician, Gujarati, Hebrew, Hindi, Hungarian, Armenian, Indonesian, Icelandic, Italian, Japanese, Javanese, Georgian, Khmer, Kannada, Korean, Lao, Lithuanian, Latvian, Malayalam, Marathi, Malay, Burmese, Norwegian, Nepali, Dutch, Polish, Portuguese, Romanian, Russian, Sinhala, Slovak, Slovenian, Serbian, Sundanese, Swedish, Swahili, Tamil, Telugu, Thai, Turkish, Ukrainian, Urdu, Vietnamese, Cantonese, Zulu.. The training and evaluation utterances comprise anonymized voice queries from Google Assistant, and long-form utterances extracted from YouTube videosThe YouTube dataset only covers 61 languages, with Chinese, Filipino, Cantonese and Zulu missing. transcribed by human annotators. The average length of a voice query is about 3.3 seconds with a standard deviation of 1.5 seconds, while the average length of a long-form utterance is about 20.7 minutes with a standard deviation of 10.6 minutes. The size of the training set for each language varies from 1M to 20M utterances (including both voice queries and long-form); while the size of the evaluation set for each language is around 20K utterances for voice queries, and 20K utterances for long-form. We use a pretrained Voice Activity Detector (VAD) to remove the non-speech parts from the utterances for both training and evaluation. MTR and SpecAug are applied to the training utterances as described in Section 3.2. In this paper, we are interested in unconstrained language identification and use the softmax cross entropy loss with Adam optimization in trainingIn applications where each request is constrained to a subset of candidate languages, the TupleMax loss is preferred..

For each evaluation, we report the average accuracy of 65 languages as the model performance metric. For each language, the accuracy is defined by the percentage of the utterances whose ground truth language has the highest score in the predicted probability distribution. From a binary classification perspective, if we only accept the language when this language has the highest score in the predicted probability distribution, then this accuracy number is equivalent to the recall rate of this language.

2 Comparison of encoder models

In our first experiment, we compare the performance of the LSTM, transformer and conformer architectures, which are the commonly used encoder models for speech recognition. As mentioned in Section 3.3, we experiment with three different model size constraints: small, medium and large, where the number of parameters are around 7M, 30M, and 120M, respectively. For this experiment, no temporal pooling layers are included for all the models — we directly take the last-frame embedding as the final embedding for the softmax layer. Configurations for each model architecture under each constraint are detailed as below.

LSTM model: The LSTM models we used in this experiment have similar topologies to the ones used in . The model has a stack of LSTM layers, with each LSTM layer, except the last one, followed by a projection layer . The LSTM layers have a pyramid-like shape, where the bottom layer is the largest one and the following layer dimensions decrease linearly. Experimentally, we find that both adding projection layers and reducing the size of layers further up in the network significantly speed up training and inference without hurting performance.

Transformer model: The transformer model is implemented based on the work of Transformer-XL . The transformer models for all sizes have 14 transformer layers as we found that the model depth is important for model performance. The layer dimensions of small, medium, and large models are 144, 256, and 1024, respectively.

Conformer model: As described in Section 3.3, the model has 12 conformer layers. The layer dimensions of small, medium, and large models are 144, 256, and 512, respectively, following the experiment setting from the original conformer paper .

The experimental results are shown in Table 1. As we can see, under all different model size constraints, the conformer model shows the best performance, followed by the transformer model. This is consistent with observations from speech recognition experiments . By counting total_float_ops with TensorFlow Profilerhttps://github.com/tensorflow/profiler, we also report the number of floating point operations needed to process 1 second of audio (FLOP/s) for each model in the table. Conformer models are relatively computationally efficient across the three model sizes. Given that conformer models are also more accelerator-friendly, they are the preferred choice from both performance and efficiency perspectives.

3 Attentive temporal pooling

To evaluate the impact of the attentive temporal pooling described in Section 3.4, we trained four additional medium-size conformer models:

Naive mean pooling, with equal weight for each frame;

Naive mean and standard deviation pooling, with equal weight for each frame;

Weighted mean (Eq. 5) and standard deviation (Eq. 6) pooling.

The experimental results are shown in Table 2. As we can see from the table, attentive temporal pooling shows improved language identification accuracy compared with no temporal pooling as well as the non-attentive naive pooling approaches.

4 Domain adaptation

As mentioned in Section 4.1, our training and evaluation data comprise both anonymized voice queries and long-form speech from YouTube. These are two domains that are very different in many respects, including: (1) the textual content; (2) the length of the speech; and (3) the prior language distribution. At training time, we always train the LangID model on the joint dataset to increase model robustness and reduce development cost. But to achieve improved in-domain performance, we adapt our model to these two different domains with a held out dev-set from each domain, as described in Section 3.5. The models with and without domain adaptation are then evaluated for each domain.

The evaluation results for domain adaptation are shown in Table 3. Note that for this experiment, we report the “total accuracy” for the entire evaluation dataset, instead of the “average accuracy” over all languages. The latter asserts a uniform class prior distribution for all languages and may not reflect specifics of the application. As we can see, the two domain adaptation methods — prior replacement and the discriminatively trained output transform — improve the language identification total accuracy on both domains.

The prior replacement approach shows a larger improvement (over no adaptation) for voice queries compared with long-form data. This observation aligns with the perplexity values shown in the table. The perplexity (PPPP) of the distribution of languages is smaller for voice queries (PP=33.8PP=33.8) than for long-form data (PP=56.2PP=56.2). If the languages were balanced (equal class priors), the perplexity would be PP=65PP=65 or the number of languages supported by the system.

The output transform approach has a larger improvement than prior replacement, especially for the long-form domain. This is expected because: (1) the output transform makes use of the LangID model outputs and reference labels on the dev-set, while prior replacement only makes use of the reference labels (related to class priors) of the dev-set; (2) the output transform better compensates for the duration gap between training and inference.

At the same time, it is also worth noting that adaptation with out-of-domain parameters (i.e. the shaded area in Table 3) will hurt the performance for specific domains. This can be mitigated by using a larger smoothing parameter RR in Eq. 10 for the prior replacement approach, or using a larger regularization weight wregw_{\text{reg}} in Eq. 13 for the output transform approach.

Conclusion

In this paper, we described a novel language identification system based on conformer layers and attentive temporal pooling. This model can be parallelized on accelerator hardware, and perform inference in a streaming fashion. Our experiments confirm that conformer based models significantly outperform LSTM or transformer based models under different model size constraints, and are relatively computationally efficient. We also show that attentive temporal pooling further improves performance. We studied two different domain adaptation approaches, namely prior replacement and a discriminatively trained output transform. They allow for the deployment of the same LangID model to different application domains where the prior language distributions and/or data are different, and effectively improve the domain performance.

Acknowledgment

The authors thank Benjamin Lee and Jonathan Shen for their help with Lingvo integration, and Pedro Moreno Mengibar, Di Li, Bhaskar Gurram, Sunha Ahn, Pavan Desikan and the anonymous Odyssey reviewers for the helpful discussions.

References