Multi-class Spectral Clustering with Overlaps for Speaker Diarization

Desh Raj, Zili Huang, Sanjeev Khudanpur

Introduction

Speaker diarization (or “who spoke when?”) refers to the task of segmenting speech into homogeneous speaker-specific regions . Conventional diarization systems consist of four major components. First, a speech activity detection module removes the non-speech segments. Next, the speech regions of the recording are divided into small (often overlapping) segments and a pretrained speaker embedding extractor is used to obtain fixed-dimensional embeddings, such as i-vectors, or neural embeddings , for each segment. The embeddings are scored pairwise using a cosine or probabilistic linear discriminant analysis (PLDA) similarity metric, and clustering (agglomerative or spectral) is performed on the resulting affinity matrix until a stopping criterion is reached or until the desired number of speaker clusters is obtained. Finally, a resegmentation module may be used for frame-level refinement of the clustering output.

Although this approach has proved to be effective through use of deep neural network based speaker embeddings, it does not handle overlapping speaker segments, since the clustering process assigns each segment to exactly one speaker. Existing approaches to solve the overlap problem fall into two categories. In the first framework, an externally trained overlap detection module identifies frames in the recording which contain overlapping speech. This “overlap detection” may be performed using hidden Markov models (HMMs) or neural networks . Once overlaps are detected, an “overlap assignment” stage assigns additional speaker labels to the overlapping frames. Recently, proposed overlap-aware resegmentation, which leverages the variational Bayes (VB)-HMM method used originally for diarization in , and applied to resegmentation in . In the second framework, end-to-end systems are used to perform overlapping diarization in a supervised setting.

In this paper, we focus on the former approach for overlap-aware speaker diarization. Specifically, we train an external overlap detector, and use its classification decision during clustering of the segment-level embeddings. Our method relies on the two-step clustering formulation proposed in . In the first step, the NP-hard discrete clustering problem is relaxed into a continuous version by ignoring the discrete constraints on the solution. The continuous problem thus obtained has a solution set generated through orthonormal transformations of eigenvectors of the normalized Laplacian. The second step involves “optimal discretization”, which finds a discrete solution under the constraint that it is close (in Frobenius norm) to any of the relaxed solutions from the solution set obtained previously. We introduce overlap awareness in the discretization stage by modifying the “sum-to-one” constraint in this subproblem. This modification makes it possible to perform overlap-aware spectral clustering at no extra computational cost beyond computing the overlap decisions. Furthermore, we use the recently proposed pp-binarization and normalized maximum eigengap (NME) techniques to self-tune the clustering process, thus requiring no hyperparameter tuning for estimating the number of speakers.

The remainder of this paper is organized as follows. We start by giving a detailed description of our method in Section 2, where we discuss pp-binarization and NME for estimating the number of speakers followed by the mathematical formulation of multi-class spectral clustering. We then introduce our modification of the method to perform overlap-aware diarization. We describe our HMM-DNN based overlap detector in Section 3. This is followed by a description of our experimental setup and results in Sections 4 and 5, respectively. We present results on the AMI meeting corpus and the LibriCSS dataset, with detailed analysis on the performance of the method on different overlap conditions. In Section 6, we summarize previous work on spectral clustering for speaker diarization. We conclude with a discussion of future work in Section 7. In the interest of reproducible research, we discuss our implementation in detail in Section 4.3, and our code has been made publicly available at: https://desh2608.github.io/pages/overlap-aware-sc.

Methodology

Our diarization follows the conventional clustering method studied extensively in previous work . We first obtain speech regions from the recording using an oracle speech activity detector (although any SAD can be used for this purpose). These are divided into small (overlapping) segments using a sliding window method (we used 1.5s segments with a stride of 0.75s in our experiments), and embeddings are extracted for each segment using an x-vector extractor . Subsequently, these embeddings are clustered to obtain speaker groupings, and each cluster is labeled as a speaker.

We first use the method described in to estimate the number of speakers K^\widehat{K} in the recording using an eigengap heuristic. Here, we describe this method briefly.

Given UU, we compute the affinity matrix A∈N×N\mathbf{A}\in^{N\times N} of raw cosine similarity values. Then, pp-binarization is performed on this matrix by replacing the pp highest similarity values in each row with 1, and the rest with 0, followed by a symmetrization operation,

We compute the unnormalized Laplacian for this matrix,

The properties of the unnormalized Laplacian of the affinity matrix have been studied extensively , and it is known that Lp\mathbf{L}_{p} has NN non-negative, real eigenvalues 0=λ1≤λ2≤…≤λN0=\lambda_{1}\leq\lambda_{2}\leq\ldots\leq\lambda_{N}. Furthermore, an implication of the Davis-Kahan perturbation theory proposes an eigengap heuristic for the optimal number of clusters. In , the authors used this heuristic to estimate the optimal pp value for the binarization (described earlier). Specifically, let ep\mathbf{e}_{p} denote the vector of differences in consecutive eigenvalues (in increasing order). We compute the quantities

Then, the optimal pp, i.e. p^\hat{p} is the one that minimizes r(p)r(p), and subsequently, the optimal number of clusters is given as

2 Multi-class spectral clustering

Bipartite graph partitioning using the affinity matrix Laplacian L\mathbf{L} is solved by node assignment based on the underlying Fiedler vector (eigenvector corresponding to the second smallest eigenvalue) . The Ng-Jordan-Weiss algorithm is a popular extension of this principle for multi-way partitioning of the graph. It applies K-means clustering on the first KK eigenvectors of LL, i.e., in the KK-eigenspace of the Laplacian. It is known that if the original samples are separable into KK groups using some transformation, then their projection on the KK-eigenspace can be easily grouped using K-means clustering. Recent work on speaker diarization through spectral clustering of x-vectors, as in and , has employed this algorithm. However, there are two major limitations of this approach. First, the K-means clustering process may get stuck in bad local optima, particularly when the affinity matrix is noisy. Second, and particularly relevant for our case, it is difficult to extend this method to handle overlaps. To remedy these issues, we use an alternative formulation of spectral clustering, proposed in .

Given A\mathbf{A} and D\mathbf{D} as defined earlier, the clustering problem requires estimating the assignment matrix XX. In graph partitioning terms, this can be represented as

Intuitively, the objective function ϵ(X)\epsilon(X) seeks to maximize the average “link-ratio”, i.e., the fraction of all link weights in a group that stay within the group. The constraint X1K=1NX\boldsymbol{1}_{K}=\boldsymbol{1}_{N} enforces the condition that each sample can belong to exactly 1 cluster. We will see later (cf. Section 2.3) how this constraint can be modified for our overlap-aware scenario.

The optimization problem in (6) is NP-complete due to the discrete constraints on XX. Instead of solving this original problem, we solve a relaxed version of this problem which ignores the constraints. Let

It is easy to verify that ZTDZ=IKZ^{T}\mathbf{D}Z=I_{K}. We can rewrite the above problem (6), by ignoring the constraints, as

Since Z has been relaxed into the continuous domain, the new optimization problem becomes tractable. Let

This implies that the global optimum is not unique; rather, it is a subspace spanned by the first KK eigenvectors of PP through orthonormal matrices.

The matrix ZZ is a continuous solution to our clustering problem. To obtain a discrete solution, we solve for a discrete approximation for ZZ. First, we note from equation (7) that

Using this transformation, we can characterize the solution obtained in equation (10) as

It is difficult to minimize ϕ(X,R)\phi(X,R) jointly in XX and RR, so we optimize it alternately in XX and RR. Suppose we are given some R∗R^{*}, then the problem (13) reduces to

The solution to this problem is given by non-maximal suppression, i.e.,

Intuitively, we set the largest entry in each row as 1 and zero out all the others. This ensures that each sample belongs to exactly 1 cluster. In the next section, we will see how to reformulate problem (14) for the case when some samples can belong to more than one clusters.

Next, we fix X∗X^{*} and solve the following problem for R∗R^{*}:

We solve the two problems (14) and (16) iteratively until convergence, and finally return X∗X^{*} as the output of the clustering procedure.

3 Extension to overlap-aware clustering

From equation (1), suppose the output of our overlap detector is given by vOL\mathbf{v}_{OL}. Then, we can reformulate problem (14) to include the overlap constraint as

Intuitively, this solves the same optimum discretization problem, but the cluster exclusivity constraint has been modified to represent the condition that some samples may belong to more than one cluster. Similar to how we used non-maximal suppression in (15) to solve the problem previously, the solution to the modified problem is again given by non-maximal suppression, with the exception that for samples belonging to more than one group, we set the largest two entries to 1, while zeroing out the others. Mathematically,

Overlap Detection

Our proposed overlap-aware diarization method relies heavily on the performance of f(U)f(U), the overlap detector (OD). In this section, we detail an HMM-DNN based overlap detector. Our model is similar to the speech activity detector previously used in the CHiME-6 baseline system .

We first trained a neural network classifier to assign each frame in an utterance a label from C\cal C = {silence, single, overlap}, denoting silence, single speaker, or overlapping regions, respectively. We used the architecture shown in Figure 1, consisting of time-delay neural network (TDNN) layers to capture long temporal contexts , interleaved with bidirectional long short term memory (BLSTM) layers with projection, to incorporate utterance-level statistics.

The posteriors obtained from the classifier were scaled with an external bias parameter tuned on the development data to reduce the false alarm rate. We then post-processed the per-frame classifier outputs to enforce minimum and maximum silence/single/overlap durations, by constructing a simple HMM whose state transition diagram encodes these constraints. Treating the per-frame posteriors like emission probabilities, we performed Viterbi decoding to obtain the most likely label-sequence. Furthermore, state transitions between the silence and overlap states were prohibited, mimicking real-world observations where it is highly unlikely for two speakers to start or stop speaking simultaneously.

Experimental Setup

We performed experiments on two datasets – the AMI meeting corpus , and the LibriCSS data . AMI consists of 100 hours of recorded meetings containing 4 speakers per session, with speech from close-talk, single distant microphone (SDM), and array microphones. For our experiments, we used the mix-headset recordings, which are obtained by summing the individual headset signals from the participants in the meeting. The dataset contains approximately 20% overlap ratio, i.e., 20% of the total speech contains overlaps. LibriCSS is a recently released corpus consisting of multi-channel audio recordings of “simulated conversations.” It comprises 10 sessions, where each session is approximately one hour long. Each session is made up of six 10-minute-long “mini sessions” that have different overlap ratios, ranging from 0 to 40%, and contain 8 speakers. The recordings were made in a regular meeting room by using a seven-channel circular microphone array. For our experiments, we selected the recordings from the first channel of the array. We used this dataset to conduct a performance analysis of our proposed method on different overlap conditions.

2 Baselines

We first have single-speaker baselines: (i) agglomerative hierarchical clustering (AHC) of x-vectors with probabilistic linear discriminant analysis (PLDA) scoring , (ii) spectral clustering of x-vectors with cosine scoring (using the Ng-Jordan-Weiss method) , and (iii) Bayesian HMM based x-vector clustering (VBx) . We used the same x-vector extractor for all the baselines (described in Section 4.3), such that the difference in their performance was only due to the clustering process. Furthermore, we used the same PLDA model (trained on a subset of the AMI training data) for the AHC and VBx baselinesThe VBx diarization system has been shown to obtain significant gains with a PLDA interpolated between general data and in-domain data, but we did not use this method in this paper.. For both these baselines, hyperparameters were tuned on the development set. No hyperparameter selection is required for spectral clustering since it is auto-tuned. We used a ground-truth VAD for these baselines as well as for our proposed method.

We also compare our approach with diarization methods that are not overlap-agnostic. These include: (i) overlap-aware VB resegmentation and (ii) region proposal networks (RPNs) . For the former, we present the official results from the paper, which uses a neural VAD. For RPNs, we filtered out non-speech regions using the ground truth VAD.

3 Implementation details

X-vector extractor. We used a similar x-vector extractor as described in earlier studies . The model consists of TDNN layers with statistics pooling, and we extracted 128-dim embeddings from the pre-final layer. It was trained on VoxCeleb data with simulated room impulse responses using the Kaldi toolkit , and released as part of the CHiME-6 baseline .

Overlap detector. We trained an HMM-DNN overlap detector (described in Section 3) using Kaldi. We used 40-dim MFCCs features as input, and trained the classifier on in-domain training data. For AMI, we used targets obtained from annotations of the official training set. Since LibriCSS does not have a corresponding training data, we generated simulated mixtures with reverberation using Librispeech training utterances and used force-aligned targets for training our overlap detector. The decoding graph was created with additional constraints on the minimum (maximum) durations for single speaker and overlapping regions as 0.03s (10.0s) and 0.1s (5.0s), respectively. Note that our clustering method itself is independent of the overlap detector used.

Overlap-aware spectral clustering. We extended the spectral clustering algorithm in scikit-learn for our implementation. Since the overlap detector provides frame-level classification decisions whereas x-vectors were extracted for 1.5s segments, we assumed that a segment is “overlapping” if at least half of it lies in overlapping regions. For estimating the number of speakers K^\widehat{K}, we swept the binarization factor pp in the range from 2 to 20, similar to what was done in .

Results and Discussion

We present the results obtained by our overlap detector using 40-dim MFCC features on AMI mix-headset data in Table 1. We can see that the performance is comparable to previous studies on this dataset, without using waveform-level learned features. We used the output from this overlap detector for further experiments.

2 Diarization results for AMI

Table 2 shows our proposed overlap-aware spectral clustering method compared with baselines, evaluated on the AMI mix-headset eval data. Using our overlap detector trained on the AMI train set, we were able to improve the DER from 28.3% for the AHC/PLDA baseline, to 24.0%, which is a relative improvement of 15.2%. This compares favorably with the performance of other diarization methods like overlap-aware VB resegmentation and RPNs. Furthermore, it is possible to reduce the DER to 21.5% using an oracle overlap detector.

A detailed analysis of the results reveals that although the missed speech reduces substantially (from 19.9% to 11.3%) as a result of overlap detection, there is also a significant increase in speaker confusion errors (from 8.4% to 10.5%). We conjecture that since the x-vector extractor was trained only on single-speaker utterances, a mismatch in the overlap regions of the recording results in noisy samples. The speaker confusion improved by 0.4% when we used an x-vector extractor trained with noise augmentation, using noises from the MUSAN corpus . To verify our hypothesis further, we show the T-SNE plots for the non-overlapping and overlapping segments in Fig. 2(a) and 2(b), respectively. We can see that while the embeddings for the non-overlapping segments are well separated, those for overlapping segments may often be noisy, leading to clustering errors.

3 Analysis on LibriCSS

Table 3 shows a breakdown of diarization errors obtained by our system, compared with some of the baselines. It is evident that as the overlap ratio increases from 0 to 40%, the difference in performance becomes more significant. On average, our method provided a 42.9% relative DER improvement compared to a baseline AHC system, and this increased to 46.0% relative on using an oracle overlap detector. We note here that since LibriCSS does not have a corresponding training set, the PLDA was trained on Librispeech utterances. The mismatch between clean training data versus overlapping mixtures at test time may be particularly detrimental to the performance of the AHC system. As the overlap ratio increases, RPN performs better than our method. We again attribute this to the fact that our RPN model is trained on closely matched overlapping speech, whereas the x-vector extractor was trained on single-speaker utterances, which results in a higher speaker confusion (cf. Section 5.2).

Related Work

Spectral clustering was first applied to speaker diarization in using the Ng-Jordan-Weiss (NJW) algorithm . extended this to the case of an unknown number of speakers by using the eigengap criterion. Agglomerative and spectral clustering methods for meeting diarization were compared in . After i-vectors were proposed for speaker recognition , they were combined with cosine scoring and spectral clustering to perform diarization in . It was further observed that spectral clustering was more robust to non-stationary environmental noise compared to other clustering methods . More recently, with the ubiquitousness of deep neural networks, several researchers have proposed methods to incorporate DNNs with spectral clustering. proposed a supervised method to measure the similarity matrix between all segments of an audio recording with BLSTMs, and applied spectral clustering on top of the similarity matrix. Other approaches use DNN-based speaker embeddings, such as x-vectors , to compute the similarty matrix between segment pairs . Additionally, introduced pp-binarization and normalized maximum eigengap (NME) techniques to automatically estimate the number of speakers in the recording.

Speaker diarization in overlapping settings has been studied extensively (cf. Section 1). However, to the best of our knowledge, there is no prior work on incorporating overlap awareness into spectral clustering based diarization. The recently proposed target speaker voice activity detection (TS-VAD) method uses x-vector based spectral clustering for initial estimate of speaker i-vectors, and thereafter performs frame-level multi-label classification to predict speaker activities in a speech frame.

Conclusion

We proposed a new method for overlap-aware speaker diarization using spectral clustering. We leveraged an external overlap detector to identify the overlapping subsegments, and then assigned these segments to multiple speakers during clustering. The clustering approach itself was reformulated by first relaxing the discrete constraints, and then solving an optimal discretization problem with the additional overlap constraints. Our method provided significant improvements over conventional single-speaker clustering models on the AMI meeting corpus, and was competitive with other overlap-aware diarization methods. Analysis on LibriCSS showed that overlapping regions benefit strongly from this approach, although speaker confusion may increase due to an inadequate speaker embedding extractor. We conjecture that this problem may be alleviated by training the x-vector extractor additionally on overlapping segments. This investigation is left as future work.

Acknowledgment

The authors thank Takuya Yoshioka for providing simulation scripts for the LibriCSS training data, and Leibny Paola García-Perera for helpful discussions and insights. This work was partially supported by research grants from Johns Hopkins Applied Physics Laboratory, Government of Israel, Hitachi Ltd., Japan, and Nanyang Technological University, Singapore.

References