Speaker Diarization with Region Proposal Network

Zili Huang, Shinji Watanabe, Yusuke Fujita, Paola Garcia, Yiwen Shao, Daniel Povey, Sanjeev Khudanpur

Introduction

Speaker diarization, the process of partitioning an input audio stream into homogeneous segments according to the speaker identity (often referred as “who spoke when”), is an important pre-processing step for many speech applications.

As shown in Figure 1 left, a standard diarization system consists of four steps. (1) Segmentation: this step removes the non-speech portion of the audio with speech activity detection (SAD), and the speech regions are further cut into short segments. (2) Embedding extraction: in this step, a speaker embedding is extracted for each short segment. Typical speaker embeddings include i-vector and deep speaker embeddings. (3) Clustering: after the speaker embedding is extracted for each short segment, the segments are grouped into different clusters. Each cluster corresponds to one speaker identity. (4) Re-segmentation: this is an optional step that further refines the diarization prediction. Among the re-segmentation methods, VB re-segmentation is the most famous one.

Despite the successful applications in many scenarios, standard diarization systems have two major problems. (1) Many individual modules: to build a diarization system, you need a SAD model, a speaker embedding extractor, a clustering module and a re-segmentation module, all of which are optimized individually. (2) Overlap: the standard diarization system cannot handle the overlapped speech. To deal with the overlapped speech, some new modules are needed to detect and classify the overlaps, which makes the procedure even more complicated. The overlapped speech will also hurt the performance of clustering, which is the main reason standard diarization systems cannot perform well in highly overlapped scenarios .

Inspired by Faster R-CNN, one of the best-known frameworks in object detection, we propose Region Proposal Network based Speaker Diarization (RPNSD). As shown in Figure 1 right, in this method, we combine the segmentation, embedding extraction and re-segmentation into one stage. The segment boundaries and speaker embeddings are jointly optimized in one neural network. After the speech segments and corresponding speaker embeddings are extracted, we only need to cluster the segments and apply non-maximum suppression (NMS) to get the diarization prediction, which is much more convenient than the standard diarization system. In addition to that, since the speech segment proposals overlap with each other, our framework solves the overlap problem in a natural and elegant way.

The experimental results on Switchboard, CALLHOME and simulated mixtures reveal that our framework achieves significant and consistent improvements over the state-of-the-art x-vector baseline, and a great portion of the improvements come from successfully detecting the overlapped speech regions. Our code is available at https://github.com/HuangZiliAndy/RPNSD.

Methodology

In this section, we will introduce our framework in details. Our framework aims to solve the speaker diarization problem and it consists of two steps. (1) Joint speech segment proposal and speaker embedding extraction. (2) Post-processing. In the first step, we predict the boundary of speech segments and extract speaker embeddings with one neural network. In the second step, we perform clustering and apply NMS to get diarization predictions.

The overall procedure of the first step is shown in Figure 2. Given an audio input, we first extract acoustic featuresWe experiment on 8kHz telephone data and we choose the STFT feature with frame size 512 and frame shift 80. During training we segment the audios into 10s chunks, so the feature shape of each chunk is (257, 1000). and feed them into convolution layers to obtain the feature maps. Then a Region Proposal Network (RPN) will generate many overlapped speech segment proposals and predict their confidence scores. After that, the deep features corresponding to the speech segment proposals are pooled into fixed-size representations. Finally, we perform speaker classification and boundary refinement on the top of the representations.

The RPN is the key component of our framework. It takes the feature maps as the input and outputs the region proposals. The original RPN generates 2-d region proposals while our RPN generates 1-d speech segment proposals. In our framework, the RPN takes the feature maps as the inputThe size of the feature maps is (1024, 16, 63). There are 63 timesteps and each timestep corresponds 16 frames of speech. and predicts speech segment proposals. Similar to brute-force search, the RPN will consider every timestep as a possible center and expand several anchors with pre-defined sizes from it. In our system, we use 9 anchors with the size of {1,2,4,8,16,24,32,48,64}\{1,2,4,8,16,24,32,48,64\}, which covers the speech segments from 16 to 1024 frames. Meanwhile, the RPN will also predict scores and refine boundaries for each speech segment proposal with convolution layers. Among the 63×9=56763\times 9=567 (63 timesteps and 9 anchors per timestep) speech segment proposals, we first filter out the speech segment proposals with low confidence scores and then further remove highly overlapped segments with NMS. In the end, we keep 100 high-quality speech segment proposals after NMS during training and 50 during evaluation.

1.2 RoIAlign

After the RPN predicts the speech segment proposals, we extract corresponding regions from the feature maps as the deep features for each segment. Since the sizes of speech segment proposals vary a lot, we need RoIAlign to pool the features into fixed dimension. Suppose we want to pool the D×TD\times T speech segment proposal (DD is the feature dimension and TT is the unfixed timestep) into a fixed representation, the proposed region is first divided into N×NN\times N (N=7N=7) RoI bins. Then we uniformly sample four locations in each RoI bin and use bilinear interpolation to compute the values of them. The result is aggregated using average pooling. With the pooled feature of fixed dimension, we can perform speaker classification and boundary refinement for each speech segment proposal.

1.3 Loss Function

The training loss consists of five parts and is formulated as

where ti\mathbf{t}_{i} and ti∗\mathbf{t}_{i}^{\ast} are the coordinates of predicted segments and ground truth segments respectively, and RR is the smooth L1 loss function in . The coordinates ti\mathbf{t}_{i} and ti∗\mathbf{t}_{i}^{\ast} are defined as follows.

2 Post-processing

In RPNSD, the input of the first step is an audio and the output includes: (1) the speech segment proposals, (2) the probability of fg/bg and (3) the speaker embedding for each segment proposal. In the second step, we perform post-processing to get the diarization prediction. The whole process contains three steps.

Remove the speech segment proposals whose fg probability is lower than a threshold γ\gamma. (γ=0.5\gamma=0.5 in our experiment)

Clustering: Group the remaining speech segment proposals into clusters. (We use K-means in our experiment)

Apply NMS (NMS threshold = 0.3) for segments in the same cluster to remove the highly overlapped segment proposals.

Experiments

We train our systems on two datasets (Mixer 6 + SRE + SWBD and Simulated TRAIN) and evaluate on three datasets (Switchboard, CALLHOME and Simulated DEV) to verify the effectiveness of our framework. The dataset statistics are shown in Table 1. The overlap ratio is defined as overlap  ratio=tspk≥2tspk≥1overlap{\;}ratio=\frac{t_{spk\geq 2}}{t_{spk\geq 1}}, where tspk≥nt_{spk\geq n} denotes the total time of speech regions with more than nn speakers. Since end-to-end systems are usually data hungry and require massive training data to generalize better, we come up with two methods to create huge amount of diarization data. (1) Use public telephone conversation datasets (Mixer 6 + SRE + SWBD). (2) Use speech data of different speakers to create synthetic diarization datasets (Simulated TRAIN). Detailed introductions for each dataset are as follows.

The Mixer 6 + SRE + SWBD dataset includes Mixer 6, SRE04-10, Switchboard-2 Phase I-III and Switchboard Cellular Part 1, 2, and the majority of the dataset are 8kHz telephone conversations. For speaker recognition, we usually use single channel audios that contain only one person. While in our experiment, we sum up both channels to create a large diarization dataset. The ground truth diarization label is generated by applying SAD on single channels.In all experiments of this paper, we use the TDNN SAD model (http://kaldi-asr.org/models/m4) trained on the Fisher corpus. We also used the same data augmentation technique as and the train sets are augmented with music, noise and reverberation from the MUSAN and the RIR dataset. The augmented train set contains 10,574 hours of speech.

SWBD DEV and SWBD TEST are sampled from the SWBD dataset (We exclude these audios from Mixer6 + SRE + SWBD). They contain around 100 5-minute audios and share no common speaker with the train set. Like Mixer 6 + SRE + SWBD, the overlap ratio of SWBD DEV and SWBD TEST is quite low. We create these two datasets to evaluate the system performance on similar data.

The CALLHOME dataset is one of the best-known benchmarks for speaker diarization. As one part of 2000 NIST Speaker Recognition Evaluation (LDC2001S97), the CALLHOME dataset contains 500 audios in 6 languages including Arabic, English, German, Japanese, Mandarin, and Spanish. The number of speakers in each audio ranges from 2 to 7.

We also use synthetic datasets (same as ) to evaluate RPNSD’s performance on highly overlapped speech. The simulated mixtures are made by placing two speakers’ speech segments in a single audio file. The human voices are taken from SRE and SWBD, and we use the same data augmentation technique as . The parameter β\beta is the average length of silence intervals between segments of a single speaker, and a larger β\beta results in less overlap. In our experiment, we generate a large dataset with β=2\beta=2 for training and three datasets with β=2,3,5\beta=2,3,5 for evaluation. The training set and test set share no common speaker.

1.2 Evaluation Metrics

We evaluate different systems with Diarization Error Rate (DER). The DER includes Miss Error (speech predicted as non-speech or two speaker mixture predicted as one speaker etc.), False Alarm Error (non-speech predicted as speech or single speaker speech predicted as multiple speaker etc.) and Confusion Error (one speaker predicted as another). Many previous studies ignore the overlapped regions and use 0.25s collar for evaluation. While in our study, we score the overlapped regions and report the DER with different collars.

2 Baseline

We follow Kaldi’s CALLHOME diarization V2 recipe to build baselines. The recipe uses oracle SAD labels which are not available in real situations, so we first use a TDNN SAD model to detect the speech segments. Then the speech segments are cut into 1.5s chunks with 0.75s overlap, and x-vectors are extracted for each segment. After that, we apply Agglomerative Hierarchical Clustering (AHC) to group segments into different clusters, and the similarity matrix is based on PLDA scoring. We also apply VB re-segmentation for CALLHOME experiments.

3 Experimental Settings

We use ResNet-101 as the network architecture and Stochastic Gradient Descent (SGD) as the optimizer We refer the PyTorch implementation of Faster R-CNN in .. We start training with a learning rate of 0.010.01 and it decays twice to 0.00010.0001. The batch size is set as 8 and we train our model on NVidia GTX 1080 Ti for around 4 days. The scaling factor α\alpha in equation 1 is set to 1.01.0 for training. During adaptation, we use a learning rate of 4⋅10−54\cdot 10^{-5} and α\alpha is set to 0.10.1. The speaker embedding dimension is 128.

4 Experimental Results

In this experiment, we train RPNSD on Mixer 6 + SRE + SWBD and use Kaldi’s x-vector model for CALLHOME as the baseline.We also train a x-vector model on single channel data of Mixer 6 + SRE + SWBD as a fair comparison but the performance is slightly worse than Kaldi’s diarization model (http://kaldi-asr.org/models/m6). As shown in Table 2, RPNSD significantly reduces the DER from 15.39%15.39\% to 9.18%9.18\% on SWBD DEV and 15.08%15.08\% to 9.09%9.09\% on SWBD TEST. On SWBD TEST, the DER composition of the x-vector baseline is 8.9%8.9\% (Miss) + 1.1%1.1\% (False Alarm) + 5.0%5.0\% (Speaker Confusion) = 15.08%15.08\% (with 2.8%2.8\% Miss and 0.9%0.9\% False Alarm for SAD). For RPNSD, the DER composition is 4.0%4.0\% (Miss) + 4.8%4.8\% (False Alarm) + 0.3%0.3\% (Speaker Confusion) = 9.09%9.09\% (with 2.0%2.0\% Miss and 2.2%2.2\% False Alarm for SAD).

Since RPNSD can handle the overlapped speech, the miss error decreases from 8.9%8.9\% to 4.0%4.0\%. As a cost, the false alarm error increases from 1.1%1.1\% to 4.8%4.8\%. Surprisingly, the speaker confusion decreases largely from 5.0%5.0\% to 0.3%0.3\%. There might be two reasons for this. (1) Instead of making decisions on short segments, RPNSD makes use of longer context and extracts more discriminative speaker embeddings. (2) The training and testing condition are more matched for RPNSD. Instead of training on single speaker data, we are training on “diarization data” and testing on “diarization data”.

4.2 Experiments on CALLHOME

The CALLHOME corpus is one of the best-known benchmarks for speaker diarization. Since the CALLHOME corpus is quite small (with 17 hours of speech) and doesn’t specify dev/test splits, we follow the “pre-train and adapt” procedure and perform a 5-fold cross validation on this dataset. We use the model in section 3.4.1 as the pre-trained model, adapt it on 4/54/5 of CALLHOME data and evaluate on the rest 1/51/5. Since our model does not use any segment boundary information, it is unfair to compare it with x-vector systems using the oracle SAD label. Therefore we compare it with x-vector systems using TDNN SAD. As shown in Table 3, our system achieves better results than x-vector systems with and w/o VB re-segmentation. It largely reduces the DER from 32.30%32.30\% (or 29.54%29.54\% after VB re-segmentation) to 25.46%25.46\%. The detailed DER breakdown is shown in Table 4. Due to the ability to handle overlapped speech, RPNSD largely reduces the Miss Error from 18.6%18.6\% to 12.8%12.8\%. As a cost, the False Alarm Error increases from 5.1%5.1\% to 7.5%7.5\%. The Confusion Error of RPNSD is also lower than x-vector and x-vector (+VB).

The DER result of RPNSD (25.46%25.46\%) is even close to the x-vector system using the oracle SAD label (24.13%24.13\%). If the oracle SAD label is used, the DER of RPNSD system must be lower than 25.46−3.2=22.26%25.46-3.2=22.26\%This is because we can easily remove the False Alarm SAD error by labeling them as silence. It is more difficult to handle the Miss SAD error in this framework, but we can further reduce the DER for sure., which is better than the x-vector system (24.13%24.13\%) and quite close to x-vector (+VB) (22.12%22.12\%).

4.3 Experiments on Simulated Mixtures

According to our experience, standard diarization systems fail to perform well on highly overlapped speech. Therefore we design experiments on simulated mixtures to evaluate the system performance on overlapped scenarios. As shown in Table 5, RPNSD achieves much lower DER than i-vector and x-vector systems. Compared with permutation-free loss based end-to-end systems, the performance of RPNSD is better than BLSTM-EEND but worse than SA-EEND. However, unlike these two systems, RPNSD does not have any constraint on the number of speakers.

Conclusion

In this paper, we propose a novel speaker diarization system RPNSD. Taken an audio as the input, the model predicts speech segment proposals and speaker embeddings at the same time. With some simple post-processing (clustering and NMS), we can get the diarization prediction, which is much more convenient than the standard process. In addition to that, the RPNSD system solves the overlapping problem in an elegant way. Our experimental results on Switchboard, CALLHOME and synthetic mixtures reveal that the improvements of the RPNSD system are obvious and consistent.

References