End-to-End Speaker Diarization as Post-Processing
Shota Horiguchi, Paola Garcia, Yusuke Fujita, Shinji Watanabe, Kenji Nagamatsu
Introduction
Speaker diarization, which is sometimes referred to as “who spoke when”, has important roles in many speech-related applications. It is sometimes used to enrich transcriptions by adding speaker attributes , and at the other times, it is used to improve the performance of speech separation and recognition .
Speaker diarization methods can be classified roughly into two: clustering-based methods and end-to-end methods. Typical clustering-based methods i) first classify frames into speech and non-speech, ii) then extract an embedding which describes speaker characteristics from each speech frame, and iii) finally apply clustering to the extracted embeddings. Most methods employ hard clustering such as agglomerative hierarchical clustering (AHC) and k-means clustering; as a result, each frame belongs either to one of the speaker clusters or to the non-speech cluster. The assumption that underlies these clustering-based methods is that each frame contains at most one speaker, i.e., they treat speaker diarization as a set partitioning problem. Thus, they fundamentally cannot deal with overlapping speech. Despite the assumption, they are still strong baselines over end-to-end methods on datasets of a large number of speakers, e.g., DIHARD II dataset . This is because they handle multiple speaker problems based on unsupervised clustering without using any speech mixtures as training data. Thus, the methods do not suffer from overtraining due to the lack of the overlap speech especially for a large number of speakers.
On the other hand, some end-to-end methods called EEND treat speaker diarization as a multi-label classification problem. They predict whether each speaker is active or not at each frame; thus, they can deal with speaker overlap. Evaluation of the early models fixed the number of speakers to two . Some extensions are proposed recently to handle unknown number of speaker cases, e.g., encoder-decoder-based attractor calculation and one-by-one prediction using speaker-conditioned model . However, these methods still perform poorly when the number of speakers is large. One reason is the training datasets. Mixtures of a large number of speakers are often rare in various datasets; thus, end-to-end models cannot produce diarization results for large number of speakers because they are overtrained on mixtures of a few number of speakers. Even if the issue on the number of mixtures is solved, the EEND depends on the permutation invariant training so that it is still hard to train the model on a large number of mixtures in terms of the calculation cost. For these reasons, how to handle mixtures that contain overlapping speech of a large number of speakers is still an open problem for both clustering and end-to-end diarization methods.
In this paper, we propose to combine both clustering-based and end-to-end methods effectively to deal with overlapping speech regardless of the number of speakers. We first obtain the initial diarization result using x-vector clustering, which does not produce overlapping results in most cases. We then apply the following steps iteratively: i) frame selection to contain only two speakers and silence and ii) overlap estimation using a two-speaker EEND model. The frame selection is also used to adapt the EEND model to a dataset which contains mixtures of more than two speakers. We evaluate our method using various datasets including CALLHOME, AMI, and DIHARD II datasets.
Related work
While some methods provide supervised clustering of speaker embeddings , the most common approach is an x-vector clustering in an unsupervised manner (See the systems submitted to DIHARD II Challenge, e.g., ). Since naive x-vector clustering results in poor performance, various techniques to improve the performance have been proposed, e.g., probabilistic linear discriminate analysis (PLDA) rescoring and Variational Bayes (VB) hidden Markov model (HMM) resegmentation . In terms of overlap processing, most methods first detect overlapped frames and then assign the second speaker for the detected frames based on heuristics or the results of VB resegmentation .
Another direction is based on clustering of overlapped segments . It first extracts overlapped segments using a region proposal network, and then applies clustering for embeddings extracted from each of them. It fundamentally solved the issue of embedding extraction using a sliding window, but its accuracies are not comparable to end-to-end methods described in Section 2.2.
2 End-to-end diarization for overlapping speech
One end-to-end approach is called EEND . They calculate multiple speaker activities, each corresponding to a single speaker. Recent models can output a flexible number of speakers’ activities by using encoder-decoder-based attractor calculation modules (EDA) or speaker-conditional EEND (SC-EEND) . Another approach is called RSAN, which are based on residual masks in the time-frequency domain to extract speakers one by one .
While EEND and RSAN take only acoustic features as input, a variant of these methods also accepts a speaker embedding as input to determine the target-speaker and output his/her speech activities. For example, target-speaker voice activity detection (TS-VAD) uses i-vectors to output the corresponding speakers’ voice activities , but the number of speakers is fixed by the model architecture. Personal VAD and VoiceFilter-Lite , which are based on d-vectors, have not such a limitation, but they assume that each speaker’s d-vector is stored in the database in advance; thus they are not suited for speaker-independent diarization.
Proposed method
Given acoustic features \{\mbox{\boldmathx}_{t}\}_{t=1}^{T}, where denotes a frame index, diarization is a problem to predict a set of active frames for each speaker . is the estimated number of speakers. For simplicity, we use \mathcal{X}_{\mathcal{T}}\coloneqq\{\mbox{\boldmathx}_{t}\mid t\in\mathcal{T}\} to denote the features of selected frames .
Clustering-based methods assume that input recordings do not contain speaker overlap. It formulates diarization as a set partitioning problem, i.e., for are predicted to be disjoint, i.e., . In EEND, on the other hand, diarization is formulated as a multi-label classification to handle overlapping speech; thus, they do not have to be disjoint. The formulation of EEND is appropriate for real conversations in which speakers sometimes utter simultaneously. However, it makes the problem too difficult to be solved; when is large (e.g. 10), it rarely happens that speakers speak together. Therefore, we assume that at most speakers speak simultaneously, and refine the clustering-based results using an end-to-end model that is trained to process at most speakers. In this study, we set . The detailed algorithm is explained in Section 3.2.
2 Algorithm
To apply the iterative refinement to each pair of speakers, the processing order influences the accuracy of final diarization results. This is because we cannot select frames to include only two speakers based on estimated diarization results because they include diarization errors. For example, if we select frames not containing Speaker 1 in Figure 1, the fourth frame contains Speaker 1 according to the final results. If the ratio of such impurities among the selected frames is high, the refinement using EEND may not perform well. We found that this problem is simply solved by processing the pairs of speakers in decreasing order of the number of selected frames (Figure 1(i)). For each speaker pair , we first select a set of frames not containing speakers other than and as follows:
We then apply the refinement described below for each speaker pair in descending order of as in Figure 1 (ii-a)–(ii-c).
2.2 Iterative update of diarization results
To update the diarization results of speakers and , we first reselect a set of frames using (missing) 1 . This is because the diarization results are updated at each refinement step so that we cannot reuse the one that is calculated to decide the processing order. Then the corresponding features are input to the EEND model to obtain posteriors of two speakers and by
where and denote posteriors of the first and second speakers at frame index , respectively, and denotes the matrix transpose. We simply apply the threshold value of 0.5 to obtain the indexes of active frames of the two speakers as
Note that we have speaker permutation ambiguity between – and –, and we solve permutation to find the optimal correspondence between and as follows:
where is a function to calculate similarity between speech and non-speech activities described by two sets and defined as
Finally, we update the diarization results of speakers and . To confirm that the new results and are calculated for speaker and , we check whether they satisfy the following conditions:
where is a lower limit of the ratio of the intersection between the new results (or ) and the previous results (or ). In this study, we set . Only if the conditions in (missing) 6 are satisfied, we update the results of speakers and . When , we simply update the results with the new ones as
On the other hand, when , such fully-update strategy causes a performance drop due to impurities in the selected frames. Thus, we use the following instead of (missing) 7 to update only overlapped frames:
For the end-to-end model , we use the self-attentive EEND model with an encoder-decoder attractor calculation module (SA-EEND-EDA) . It consists of a four-layer-stacked Transformer encoder to extract embeddings for each frame and the EDA module to calculate attractors from the extracted embeddings. The EDA includes long short-term memories but we shuffled the order of embeddings just before they are fed into the EDA, which improves the diarization performance. Thus, we can consider that all the components of are independent of the order of embeddings and therefore the model can treat input features of selected frames even if they are not continuous in time.
3 Training strategy of the SA-EEND-EDA model
In the original EEND and its derived methods used matched dataset for model adaptation, i.e., only two-speaker subset of the original dataset (e.g. CALLHOMEhttps://catalog.ldc.upenn.edu/LDC2001S97) was used to finetune the models in two-speaker evaluations. This strategy cannot be used to finetune two-speaker models when the dataset does not contain two-speaker mixtures (e.g. AMI ). Even if two-speaker mixtures are included in the dataset, it does not make full use of the datasets, which may cause performance degradation.
To cope with this situation, we adopt the frame-selection technique used in (missing) 1 for model adaptation. If the input chunk contains more than two speakers, we first choose two dominant speakers and then eliminate frames in which the other speakers are active as in (missing) 1. The model is trained only using the selected frames to output speech activities of the two speakers. This makes it possible to finetune two-speaker models from any kind of multi-speaker datasets without mixture-wise selection.
Experiment
Table 1 shows the datasets used for our evaluation. The model was pretrained using simulated two-speaker mixtures Sim2spk for 100 epochs. Each mixture was simulated from two single-speaker audios derived from Switchboard-2 (Phase I & II & III), Switchboard Cellular (Part 1 & 2), or NIST Speaker Recognition Evaluation (2004 & 2005 & 2006 & 2008). Noise sources from MUSAN corpus and simulated room impulse responses are also used to simulate noisy and reverberant environments. The detailed simulation protocol is in our previous paper . For the pretraining, Adam optimizer with the learning rate scheduler proposed in was used. The number of warm-up steps was set to 100,000 following .
After the pretraining, the model was adapted on CALLHOME, AMI and DIHARD II datasets for another 100 epochs, respectively. Adam optimizer was also used in the adaptations but its learning rate was fixed to following .
We used diarization error rate (DER) and Jaccard error rate (JER) for evaluation. While some studies excluded overlapped regions from evaluation , this study scored overlapped region. We also note that our evaluations are based on estimated speech activity detection (SAD), while some studies used oracle segments or only reported confusion errors .
2 Preliminary evaluation of the training using frame selection
3 Results
We first evaluated the proposed method on CALLHOME dataset, which is composed of telephone conversations. As a clustering-based baseline, x-vectors with AHC and PLDAhttps://github.com/kaldi-asr/kaldi/tree/master/egs/callhome_diarization/v2 was used with TDNN-based speech activity detectionhttps://github.com/kaldi-asr/kaldi/tree/master/egs/aspire/s5. We also prepared the results for which VB-HMM resegmentation was applied. All the components were implemented in Kaldi recipe.
3.2 AMI
Second, we evaluated our method on AMI dataset, consisting of meeting recordings. While it includes various types of recordings, we used Headset mix recordings for this experiment. We chose the system developed during JSALT 2019 as a baseline. It is based on x-vector clustering followed by VB resegmentation and overlap detection and assignment for the second speaker candidate .
3.3 DIHARD II
Finally, we evaluated the proposed method on DIHARD II dataset, which includes recordings from 10 different domains. We used the official baseline system and the BUT system , which is the winning system of the second DIHARD Challenge, to obtain initial diarization results. Both are based on the x-vector clustering, but the BUT system is more polished in that it extracts x-vectors in shorter intervals and uses VB resegmentation and overlap detection and assignment based on heuristics.
Conclusion
In this paper, we proposed a post-processing method for clustering-based diarization using an end-to-end diarization model. We iteratively selected two speakers, picked up frames that contain the two speakers, and process the frames by the end-to-end model to update diarization results. Evaluations on CALLHOME, AMI, and DIHARD II datasets showed that our proposed method improves various types of clustering-based diarization results.
Acknowledgment
We would like to thank Federico Landini for providing the results of the winning system of the second DIHARD Challenge.