Personal VAD: Speaker-Conditioned Voice Activity Detection

Shaojin Ding, Quan Wang, Shuo-yiin Chang, Li Wan, Ignacio Lopez Moreno

Introduction

In modern speech processing systems, voice activity detection (VAD) usually lives in the upstream of other speech components such as speech recognition and speaker recognition. As a gating module, VAD not only improves the performance of downstream components by discarding non-speech signals, but also significantly reduces the overall computational cost due to its relatively small size.

A typical VAD system uses a frame-level classifier on acoustic features to make speech/non-speech decisions for each audio frame (e.g. with 25ms width and 10ms step). Poor VAD systems could either mistakenly accept background noise as speech or falsely reject speech. False accepting non-speech as speech largely slows down the downstream automatic speech recognition (ASR) processing. It is also computationally expensive as ASR models are normally much larger than VAD models. On the other hand, false rejecting speech leads to deletion errors in ASR transcriptions (a few milliseconds of missed audio could remove an entire word). A good VAD model needs to work accurately in challenging environments, including noisy conditions, reverberant environments and environments with competing speech. Significant research has been devoted to finding the optimal VAD features and models . In the literature, LSTM-based VAD is a popular architecture for sequential modeling of the VAD task, showing state-of-the-art performance .

In many scenarios, especially on-device speech recognition , the computational resources such as CPU, memory, and battery are typically limited. In such cases, we wish to run the computationally intensive components such as speech recognition only when the target user is talking to the device. False triggering such components in the background while only speech signals from other talkers or TV noises are present would cause battery drain and bad user experience. Although such concerns usually can be easily addressed by introducing a keyword detection (a.k.a. wake word detection ) model, in many applications, the users would largely prefer a more seamless and natural interaction with the voice assistant without having to speak a predefined keyword. Thus, having a tiny model that only passes through speech signals from the target user is very necessary, which is our motivation of developing the personal VAD system.

Although standard speaker recognition and speaker diarization techniques can be directly used for the same task, we argue that the personal VAD system is largely preferred here for a couple of reasons:

To minimize the latency of the whole system, an accept/reject decision is needed upon the arrival of each frame immediately, which prefers frame-level inference of the model. However, many state-of-the-art speaker recognition and diarization systems usually require window-based or segment-based inference, or even offline full-sequence inference.

To minimize battery consumption on the device, the model must be very small, while most speaker recognition and diarization models are pretty big (typically millions of parameters).

Unlike speaker recognition or diarization, in personal VAD, it is unnecessary to distinguish between different non-target speakers, as we only trigger downstream components for the target speaker.

In fact, we implemented a baseline system by directly combining a standard speaker verification model and a standard VAD model for the personal VAD task, as described in Section 2.2.1, and found that its performance is worse than a dedicated personal VAD model. To the best of our knowledge, this work is the first lightweight solution that aims at directly detecting the voice activity of a target speaker in real time.

The proposed personal VAD is a VAD-alike neural network, conditioned on the target speaker embedding or the speaker verification score. Instead of determining whether a frame is speech or non-speech in standard VAD, personal VAD extends the determination to three classes: non-speech, target speaker speech, and non-target speaker speech.

The rest of the paper is organized as following. In Section 2.1, we first briefly describe our speaker verification system, which will be used during the training of personal VAD. Then in Section 2.2, we propose four different architectures to achieve personal VAD. In the training of personal VAD, we first treat it as a three-class classification problem and use cross entropy loss to optimize the model. In addition, we noticed that the discriminitivity between non-speech and non-target speaker speech is relatively less important than between target speaker speech and the other two classes in personal VAD. Therefore, we further propose a weighted pairwise loss to enforce the model to learn these differences, as introduced in Section 2.3. We evaluate the model on an augmented version of the LibriSpeech dataset , with experimental setup described in Section 3.2, model configuration described in Section 3.3, metrics explained in Section 3.4, and results presented in Section 3.5. Conclusions are drawn in Section 4.

Approach

Personal VAD relies on a pre-trained text-independent speaker recognition/verification model to encode the speaker identity into embedding vectors. In this work, we use the “d-vector” model introduced in , which has been successfully applied to various applications including speaker diarization , speech synthesis , source separation , speech translation , and audio voice preservation tests . We retrained the 3-layer LSTM speaker verification model using data from 8 languages for language robustness and better performance. During inference, the model produces embeddings on sliding windows, and a final aggregated embedding named “d-vector” is used to represent the voice characteristics of this utterance, as illustrated in Fig. 1. The cosine similarity between two d-vector embeddings can be used to measure the similarity of two voices.

In a real application, users are required to follow an enrollment process before enabling speaker verification or personal VAD. During enrollment, d-vector embeddings are computed from the target user’s recordings, and stored on the device. Since the enrollment is a one-off experience and can happen on server-side, we can assume that the embeddings of the target speakers are available at runtime with no cost.

2 System architecture

A personal VAD system should produce frame-level class labels for three categories: non-speech (ns), target speaker speech (tss), and non-target speaker speech (ntss). We implemented four different architectures to achieve personal VAD, as illustrated by Fig. 2. All four architectures rely on the embedding of the target speaker, which is acquired via the enrollment process.

Our first approach to implement personal VAD is to simply combine a standard pre-trained speaker verification system and a standard VAD system, as shown in Fig. 2(a). We use this implementation as a baseline for other approaches, since it does not require training any new model.

where zt=[zts,ztns]\mathbf{z}_{t}=[z_{t}^{\tt s},z_{t}^{\tt ns}]. The speaker verification model produces an embedding et\mathbf{e}_{t} at each frame:

To transform the standard VAD probability ztsz_{t}^{\tt s} to personal VAD probabilities zttssz_{t}^{\tt tss} and ztntssz_{t}^{\tt ntss}, we combined it with the resulting speaker verification cosine similarity score sts_{t}, such that:

There are two major disadvantages of this architecture. First, it is running a window-based speaker verification model at a frame level without any adaptation, and such inconsistency could cause significant performance degradation. However, training frame-level speaker verification models is often unscalable due to the difficulties to batch utterances of different length. Second, this architecture requires running a speaker verification system at runtime, which can be expensive since speaker verification models are usually much bigger than VAD models.

2.2 Score conditioned training (ST)

As shown in Fig. 2(b), our second approach uses the speaker verification model to produce a cosine similarity score sts_{t} for each frame, as explained in Eq. (3), then concatenates this cosine similarity score to the acoustic features:

The concatenated feature vector x^t\hat{\mathbf{x}}_{t} is 41-dimensional, as xt\mathbf{x}_{t} represents the 40-dimensional log Mel-filterbank energies. We train a new personal VAD network that takes the concatenated features as input, and outputs the probabilities of the three class labels for each frame:

where zt=[zttss,ztntss,ztns]\mathbf{z}_{t}=[z_{t}^{\tt tss},z_{t}^{\tt ntss},z_{t}^{\tt ns}].

This approach still requires running the speaker verification model at runtime. However, since it retrains the personal VAD model based on the speaker verification scores, it is expected to perform better than simply combining the scores of two individually trained systems.

2.3 Embedding conditioned training (ET)

As shown in Fig. 2(c), the third approach directly concatenates the target speaker embedding (acquired in the enrollment process) with the acoustic features:

Since our embedding is 256-dimensional, the concatenated feature vector here is 296-dimensional. Then we train a new personal VAD network, which outputs the probabilities of three class at the frame level similar to Eq. (6).

This approach is similar to a knowledge distillation process. The large speaker verification model was pre-trained on a large-scale dataset individually. Following this, when we train the personal VAD model, we use the speaker embeddings of the target speaker to “distill the knowledge” from the large speaker verification model to the small personal VAD model. As a result, it does not require running the large speaker verification model at runtime, which becomes the most lightweight solution among all architectures.

2.4 Score and embedding conditioned training (SET)

As shown in Fig. 2(d), this approach concatenates both the frame-level speaker verification score and the target speaker embedding to the acoustic features to train a new personal VAD model:

The concatenated feature vector in this approach is 297-dimensional. This approach makes use of the most information from the speaker verification system. However, it still requires running the speaker verification model at runtime, so it’s not a lightweight solution.

3 Weighted pairwise loss

However, in personal VAD, our goal is to detect the voice activity from only the target speaker. Audio frames that are classified into class ns and ntss will be discarded similarly by downstream components. As a result, confusion errors between have less impact to the system performance than errors between and . Inspired by Tuplemax loss , here we propose a weighted pairwise loss to model the different tolerance to each class pair. Given z\mathbf{z} and yy, we define weighted pairwise loss as:

where w<k,y>w_{<k,y>} is the weight between class kk and class yy. By setting lower weight to errors than and errors, we can enforce the model to be more tolerant to the confusion between and to focus on distinguishing tss from ns and ntss.

Experiments

An ideal dataset to train and evaluate personal VAD would be a dataset such that: (1) each utterance in it contains natural speaker turns; and (2) it contains enrollment utterances for each individual speaker. Unfortunately, to the best of our knowledge, no public dataset in the community really satisfies both requirements. Although some datasets for speaker diarization have natural speaker turns, they do not provide enrollment utterances for individual speakers. Alternatively, datasets containing enrollment utterances for individual speakers usually do not have natural speaker turns.

To address this limitation, we conducted experiments on an augmented version of the LibriSpeech dataset . To simulate speaker turns, we concatenate single-speaker utterances from different speakers into multi-speaker utterances (see Section 3.2.1). We also noisify the concatenated utterances with reverberant room simulators to mitigate the concatenation artifacts (see Section 3.2.2).

In the LibriSpeech dataset, the training set contains 960 hours of speech, where 460 hours of them are “clean” speech and the other 500 hours are “noisy” speech. The testing set also consists of both “clean” and “noisy” speech. In all the experiments, we use the concatenated LibriSpeech training set to train the models. We use both the original LibriSpeech testing set and the concatenated LibriSpeech testing set for evaluation, as described in the following sections. For all the datasets, to produce the frame-level ground truth personal VAD labels used in training and evaluation, we run forced alignment with a pretrained speech recognition model.

2 Experimental settings

In the training corpora of standard VAD, each utterance usually only contains the speech from one single speaker. However, personal VAD aims to find the voice activity of a target speaker in a conversation where multiple speakers could be engaged. Therefore, we cannot directly use the standard VAD training corpora to train personal VAD. To simulate the conversational speech, we concatenate utterances from multiple speakers into a longer utterance, and then we randomly select one of the speakers as the target speaker in the concatenated utterance.

To generate a concatenated utterance, we draw a random number nn indicating the number of utterances used for concatenation from a uniform distribution:

where aa and bb are the minimal and maximal numbers of utterances used for concatenation. The waveforms from the nn randomly selected utterances are concatenated, and one of the speakers is assumed as the target speaker of the concatenated utterance. At the same time, we modify the VAD ground truth label of each frame according to the target speaker: “non-speech” frames remain the same, while “speech” frames are modifed to either “target speaker speech” or “non-target speaker speech” according to whether the source utterance is from the target speaker.

In our experiments, we generated 300,000300,000 concatenated utterances for training set and 5,0005,000 concatenated utterances for testing sets. We use a=1a=1 and b=3b=3 for both sets, to cover both single-speaker and multi-speaker scenarios.

2.2 Multistyle training

For both training and evaluations, we apply a data augmentation technique named “multistyle training” (MTR) on our datasets to avoid domain overfitting and mitigate concatenation artifacts. During MTR, the original (concatenated) source utterance is noisified with multiple randomly selected noise sources, using a randomly selected room configuration. Our noise sources include:

827 audios of ambient noises recorded in cafes;

786 audios recoreded in silent environments;

6433 YouTube segments containing background music or noise.

We generated 3 million room configurations using a room simulator to cover different reverberation conditions. The distribution of the signal-to-noise ratio (SNR) of our MTR is shown in Fig. 3.

3 Model configuration

The acoustic features are 40-dimensional log Mel-filterbank energies, extracted on frames with 25ms width and 10ms step. For both standard VAD model and personal VAD model, we used a 2-layer LSTM network with 64 cells, followed by a fully-connected layer with 64 neurons. We also tried larger networks but did not see performance improvements, possibly due to the limited variety in training data. We used TensorFlow for training and inference. During training, we used Adam optimizer with a learning rate of 5×10−55\times 10^{-5}. For the models with weighted pairwise loss, we set w<tss,ns>=w<tss,ntss>=1w_{\tt<tss,ns>}=w_{\tt<tss,ntss>}=1 and explored different values for w<ns,ntss>∈{0.01,0.05,0.1,0.5,1.0}w_{\tt<ns,ntss>}\in\{0.01,0.05,0.1,0.5,1.0\}.

To reduce the model size and accelerate the runtime inference, we quantized the parameters of the model to 8-bit integer values following . With this quantization, our model using the ET architecture, which has only around 130 thousand parameters and is the smallest among all architectures (see Table 1), will be only 130 KB in size.

4 Metrics

To evaluate the performance of the proposed method, we computed the Average Precision (AP) for each class and the mean Average Precision (mAP) over all the classes. AP and mAP are most common metrics for multi-class classification problems. AP summarizes a precision-recall curve as the weighted mean of precisions achieved at each threshold, with the increase in recall from the previous threshold used as the weight. AP can be computed as:

where RnR_{n} and PnP_{n} are the recall and precision at the nn-th threshold, respectively. We adopted the micro-meanhttps://scikit-learn.org/stable/modules/generated/sklearn.metrics.average_precision_score.html over all the classes when computing mAP to take class imbalance into account, which averages APs over all the samples.

5 Results

We conducted three groups of experiments to evaluate the proposed method. First, we compared the four architectures for personal VAD. Following this, we examined the effectiveness of weighted pairwise loss and compared it against conventional cross entropy loss. Finally, we evaluated personal VAD on a standard VAD task, to see if personal VAD can replace standard VAD without performance degradation.

In the first group of experiments, we compared the performance of four personal VAD architectures described in Fig. 2. We evaluated these systems on the concatenated LibriSpeech testing set. Additionally, to explore the performance of personal VAD on noisy speech, we also applied data augmentation technique (MTR) on the testing set. In personal VAD tasks, the most important metric is the AP for class tss, as downstream processes will only be applied to the speech produced by the target speaker.

We reported the evaluation results on the testing set with and without MTR, as shown in Table 1. Results show that ST, ET, and SET significantly outperform the baseline SC system in all cases. When applying MTR to the testing set, we observed an even larger performance gain between the proposed methods and the baseline. Among the proposed systems, SET achieved the highest AP for tss, and ST slightly outperforms ET. However, both ST and SET require to run speaker verification model (4.88 million parameters) to compute the cosine similarity score during inference time, which would largely increase both the number of parameters in the system and inference computational cost. By contrast, ET obtained 0.932 (without MTR) / 0.878 (with MTR) AP for class tss on the testing set with a model of only 0.13 million parameters (∼\sim 40 times smaller), which is more appropriate for on-device applications.

5.2 Loss function comparisons

In the second group of experiments, we compared the proposed weighted pairwise loss against the conventional cross entropy loss. Here we only consider the ET architecture, as it is much more lightweight while achieving reasonably good performance. Similarly, we evaluated the systems on the concatenated LibriSpeech testing set with and without MTR.

In Fig. 4, we plot the AP for tss against different values of w<ns,ntss>w_{\tt<ns,ntss>} in weighted pairwise loss. From the results, we observed that using a smaller value of w<ns,ntss>w_{\tt<ns,ntss>} than w<tss,ns>w_{\tt<tss,ns>} and w<tss,ntss>w_{\tt<tss,ntss>} will improve the performance, which demonstrates that confusion errors between have less impact to the system performance than errors between and .

However, when w<ns,ntss>w_{\tt<ns,ntss>} becomes too small (e.g. 0.050.05 or 0.010.01), we found performance degradations from the curve. This result shows that completely ignoring the difference between ntss and ns is harmful to the system performance as well. In another word, it is insufficient to simply treat personal VAD task as a binary classification problem (target speaker speech v.s. other). The best performance is reached when setting w<ns,ntss>=0.1w_{\tt<ns,ntss>}=0.1, with detailed results listed in Table 1.

5.3 Personal VAD on standard VAD tasks

If we want to replace a standard VAD component with personal VAD, we also need to guarantee that the performance degradation on a standard speech/non-speech task is minimal. Finally, we conducted an experiment for personal VAD on standard VAD tasks. We evaluated two personal VAD models (ET architecture with cross entropy loss, and ET architecture with weighted pairwise loss) on the non-concatenated LibriSpeech testing data (so each utterance only has the target speaker). For comparison purpose, we also implemented a standard VAD model with the same network structure (2-layer LSTM network with 64 cells, followed by a fully-connected layer with 64 neurons).

The results are shown in Table 2. We can see that the AP for class speech (s) is very close between personal VAD models and the standard VAD model, which justifies replacing standard VAD by personal VAD. Additionally, the architectures of personal VAD models and the standard VAD model are the same in this experiment, so replacing standard VAD by personal VAD will not increase the model size or computational cost at inference time.

Conclusions

In this paper, we proposed four different architectures to implement personal VAD, a system that detects the voice activity of a target user in real time. Among the different architectures, using a single small network that takes acoustic features and enrolled target speaker embedding as inputs achieves near-optimal performance with smallest runtime computational cost. To model the tolerance to different types of errors, we proposed a new loss function, the weighted pairwise loss, which proves to have better performance than a conventional cross entropy loss. Our experiments also show that personal VAD and standard VAD perform equally well on a standard VAD task. In summary, our findings suggest that, by focusing only on the desired target speaker, a personal VAD can reduce the overall computational cost of speech recognition systems operating in noisy environments.

References