Integration of speech separation, diarization, and recognition for multi-speaker meetings: System description, comparison, and analysis
Desh Raj, Pavel Denisov, Zhuo Chen, Hakan Erdogan, Zili Huang, Maokui He, Shinji Watanabe, Jun Du, Takuya Yoshioka, Yi Luo, Naoyuki Kanda, Jinyu Li, Scott Wisdom, John R. Hershey
Introduction
Multi-speaker speech recognition is defined as a task wherein, given a long unsegmented recording consisting of a conversation between an unknown number of participants, the expected output contains “who said what and when.” Although the last two decades have seen incredible leaps in speech technology — deep learning systems have matched human parity in the single-speaker conversational speech settings — multi-speaker speech processing remains a challenge as such recordings may contain almost 20% overlapped speech . This scenario has implications for both diarization and automatic speech recognition (ASR) — single-speaker diarization systems miss the interfering speaker completely, and ASR systems trained on clean utterances are more error-prone on overlapped regions. These factors, along with the effect of far-field acoustic conditions, may increase the speaker-attributed word error rates by up to 86% in meetings. To tackle such problems, recent editions of DIHARD and CHiME have focused on these tasks in very challenging settings.
A straightforward approach for solving the overlap problem in multi-speaker conversations is through speech separation. Deep learning based separation methods have been progressing consistently over the last few years — a signal-to-distortion ratio of 19.0 dB has been obtained on the popular WSJ0-2mix dataset , making the enhanced signals almost indistinguishable from clean utterances . However, these techniques are often evaluated on short, fully overlapping (and often simulated) mixtures, and do not measure distortions introduced in overlap-free regions. They may also make unrealistic assumptions, such as prior knowledge of the number of speakers in the mixture. Recently, there have been efforts towards “continuous speech separation,” which situates separation techniques in more realistic settings of long-form conversations containing partially overlapped speech , such as those found in multi-speaker conversations.
There is increasing interest towards combining the advances made in speech separation, diarization, and ASR for tackling multi-speaker speech recognition. The CHiME-6 challenge included a track aimed at recognizing unsegmented dinner-party conversations recorded on multiple microphone arrays, and baseline evaluations indicated that replacing oracle segments with diarization outputs could degrade downstream WERs by almost 52%. Chen et al. proposed the new LibriCSS meeting dataset, and combined speaker-independent continuous speech separation with a strong hybrid ASR to recognize long unsegmented meeting recordings. Similar datasets such as LibriMix have been created to satisfy the need for multi-speaker conversational data containing partial overlaps. Joint modeling of speaker counting, speaker identification, and ASR using end-to-end models has also been investigated .
Nevertheless, the number of available methods to perform separation, diarization, or recognition are plenty — each with their own sets of requirements and assumptions. Together with the several possibilities of combining them, this creates a daunting task to analyze systems that recognize unsegmented multi-speaker recordings. In this paper, we make a first attempt at tackling this problem, using the LibriCSS dataset. We propose a novel end-to-end modular pipeline that performs separation, diarization, and recognition, in that order. For each of these modules, we compare several existing methods, describing the advantages and disadvantages of each alternative. We also demonstrate the impact of separation on downstream tasks by presenting corresponding results on the original mixed recording. Since the modules only interact at the I/O level, they can be implemented using different tools and optimized independently. Our best system with and without separation achieves a concatenated minimum-permutation WER (cpWER) of 12.7% and 23.9%, respectively, on the LibriCSS evaluation set. We will release code to reproduce (and extend) our pipeline here: https://desh2608.github.io/pages/jsalt/.
System Overview
Our pipeline consists of three components: Separation Diarization ASR. An unsegmented multichannel audio recording is provided as input to a continuous speech separation system, such as . The separation module works on small windowed segments containing at most 2 or 3 speakers, and produces the corresponding number of separated audio streams, which are passed to the diarization module. Since a speaker may have been split into different streams across different windows, the diarization is performed by considering all the audio streams simultaneously. We will discuss the implications of this requirement on different diarization methods in Section 4. After diarization, the single-speaker homogenenous segments are fed into an ASR decoder.
Fig. 1 shows our proposed approach, and situates it in comparison with alternative approaches proposed in and , which we refer to as the CHiME-6 pipeline and the CSS pipeline, respectively. All 3 methods perform explicit separation, as opposed to an implicit “separation+recognition” usually performed by target-speaker ASR methods . Stacking the components in different orders elicits unique advantages and limitations for each approach. Since diarization is performed first in the CHiME-6 pipeline, it makes it possible to use the source activity pattern derived from the diarization output to guide the estimation of the mixture model parameters in guided source separation (GSS) , or use the speaker information to perform target speaker extraction, as in SpeakerBeam . However, poor diarization can have significant impacts on separation results in this pipeline. The CSS pipeline performs separation in the first stage, followed by ASR and diarization, allowing the system to be deployed in a streaming setting. The diarization performance may itself also benefit from reduced false alarms, since segments are obtained from the ASR output. However, it makes speaker-biasing infeasible at the ASR stage. Furthermore, the ASR output may contain spurious insertions resulting from cross-talk in non-overlap regions of the separated audio streams.
In contrast, our pipeline makes it feasible to employ a speaker-biased ASR, and also to remove extra insertions through simple post-processing of the diarization output (cf. Section 3.3). Unlike the CHiME-6 pipeline, our separation is not dependent on the performance of the diarizer, and speaker-independent separation techniques can be effectively used in the pipeline. Here we demonstrate two of these — mask-based minimum variance distortionless response (MVDR), and sequential multi-frame separation. Furthermore, since the diarization stage sees only single-speaker recordings, it circumvents the need for an overlap-aware diarization component.
We will now describe the performance of different methods in each module – namely separation, diarization, and ASR — in Sections 3, 4, and 5, respectively. All our experiments are conducted on LibriCSS , which consists of multi-channel audio recordings of “simulated conversations,” generated by mixing Librispeech utterances . It comprises 10, approximately one-hour long, sessions. Each session is made up of six 10-minute-long “mini sessions” that have different overlap ratios, ranging from 0% to 40%. The recordings were made in a regular meeting room by using a seven-channel circular microphone array. We used session 0 as the development set, and the remaining 9 sessions for evaluation (which we use to report results). Diarization results are reported using diarization error rate (DER), and ASR performance is evaluated in terms of cpWER. It is computed by concatenating all the utterances of a speaker in the reference and hypothesis, scoring all speaker pairs, and then finding the speaker permutation that minimizes the total WER.
Speech Separation
The speech separation module follows the continuous speech separation scheme , consisting of three steps. First, the recording is uniformly segmented into smaller chunks with overlap. For each chunk, a multi-channel separation network , trained with permutation invariant training (PIT) criterion, is applied to estimate three time-frequency (TF) masks: two for speech sources and one for noise. In this scheme, we used a chunk size of 2.4s with 0.8s hop.
Thereafter, a stitching algorithm is used to track local permutations and glue the masks from each chunk into the meeting-wise mask. This stitching is performed using mask similarity between adjacent chunks on the overlapped region (which is 1.6s in our case), by finding the permutation that has minimum distance between chunks. Based on the resulting permutation, the chunk-wise masks are connected to obtain a mask stream for the long recording. We averaged the overlap region between chunks for this process.
Finally, given the stitched masks for the entire meeting, a mask-based adaptive MVDR beamforming is performed to get the final separation result. Since the noise is mostly stationary, we used the noise mask for the entire meeting to estimate the noise spatial covariance.
Note that the number of output channels in each local separation is highly correlated with the chunk size. As suggests, most 2.4s chunks contains at most 2 speakers, therefore two output channels suffice for local processing. However, when the chunk size is larger (as in the next separation model), it may contain more than 2 speakers, even though LibriCSS contains at most 2-speaker overlaps.
2 Sequential multi-frame separation
We also used a novel sequential separation system which has multiple mask-based beamforming steps . A mask-based multi-frame multi-channel Wiener filter (MCWF) beamformer is used. We used three untied sequential steps inside the model by feeding previously beamformed signals into the next mask-prediction step. The output of this model can be the last mask-network output or the one previous beamformed output. This model was trained using 10 second long mixtures of three speaker utterances from Libri-Light database . During training, the utterances were mixed on-the-fly with room impulse responses obtained using an image-method based room simulator that simulates data acquisition in shoe-box shaped rooms with an 8-microphone array . This pre-trained model was used to separate sources in the LibriCSS dev and eval datasets using 8 second long blocks with 4 seconds overlap. Since LibriCSS has 7 microphones, we added another microphone signal by shifting the first microphone signal with one sample and adding white Gaussian noise with variance 1e-6. The source estimate outputs were stitched using a stitching loss of magnitude STFT domain mean-squared error between time-domain signals to resolve the permutation across blocks. This is similar to the stitching method described in Section 3.1.
3 Experimental results
Table 1 shows the performance of the separation methods on the LibriCSS eval set. Separation performance is calculated on a simulated eval set different from LibriCSS since reference signals are required. Since the models work with varying window sizes, we first mapped the output tracks to the source tracks for every 0.8s block0.8s is the highest common factor between the chunk size of the models., where is the number of participating speakers, using magnitude-domain distance and assuming at most two active speakers per 0.8s block, and then generated speaker tracks from the separated outputs. We call this process “oracle track mapping”. We report separation performance in terms of average meeting-level SDR . Sequential multi-frame model achieved a higher SDR of 14.1 dB as compared to the mask-based MVDR’s 5.8 dB, since implemented MVDR beamformer does not attempt to reconstruct target signals exactly.
We also evaluated the separated audio with a spectral clustering based diarizer (Section 4.1) and a hybrid HMM-DNN ASR (Section 5.1), and the results are shown in the table in terms of DER and cpWER, respectively. From the table, we note that the diarization performance of the two methods were comparable, but the 3-stream sequential multi-frame separation outperformed the 2-stream mask-based MVDR in terms of cpWER results. Although the sequential model was trained with an 8-microphone cubic geometry, it generalized well to LibriCSS, which has a completely different geometry, since the model is inherently geometry-independent.
Furthermore, a simple post-processing trick applied on the diarization output was found to be useful for the sequential model – using this trick improved the cpWER by 15.4% relative. For this post-processing, we removed a segment from a stream if it was completely enclosed within a same-speaker segment in a different stream. This is akin to filtering out cross-talk prior to ASR decoding. This trick was not useful for the mask-based MVDR, providing only a marginal improvement of 0.2% relative in terms of cpWER. Note that although sequential model achieved significantly better SDR, their cpWER are similar. This indicates the potential objective mismatch between speech separation and recognition.
Speaker Diarization
We can categorize diarization methods based on whether or not they can assign overlapping speaker segments. Since our pipeline separates the recording prior to diarization, overlap-awareness is not a strict requirement. At the same time, since the separation is done on windowed segments, the diarization needs to be performed across audio streams (since a speaker can be present in different streams at different times). Based on these conditions, we selected the following diarization methods for our study.
This method consists of a speech activity detection (SAD) component followed by clustering of small subsegment embeddings. We used a similar SAD as that described in , consisting of a TDNN-Stats based classifier with Viterbi decoding for inference. The speech segments were divided into subsegments with a window size of 1.5s and a stride of 0.75s, and 128-dimensional embeddings were extracted using an x-vector extractor trained on VoxCeleb data with simulated room impulse response . We conducted experiments with 3 variants of clustering: (i) agglomerative hierarchical clustering (AHC) on PLDA scores , (ii) spectral clustering (SC) on cosine similarity , and (iii) VBx clustering initialized from the AHC system . For these methods, we did not make any assumptions about the number of speakers in the recording. We used the same PLDA trained on Librispeech for both AHC and VBx, and did not use the PLDA interpolation technique from .
2 Region proposal networks (RPN)
This is a supervised method, which combines the segmentation and embedding extraction steps into a single neural network and jointly optimizes them . The region embeddings are then clustered (using K-means clustering) and a non-maximal suppression is applied. We trained the RPN on simulated meeting-style recordings with partial overlaps generated using utterances from the Librispeech training set. Since we used K-means clustering, we assumed that the oracle number of speakers for each recording is known.
3 Target-speaker voice activity detection (TS-VAD)
The TS-VAD model takes conventional speech features (e.g., MFCC) along with i-vectors for each speaker as inputs and produces frame-level activities for each speaker using a neural network with a set of binary classification output layers . Since the number of these binary output nodes are fixed for training, we assumed that the maximum possible number of speakers in any session is at most 8. The initial estimates for the speaker i-vectors were obtained using the SC system. For training, we created simulated meeting-style data similar to that used for training the RPN model.
4 Experimental results
We conducted experiments with “mixed” as well as “separated” audio to analyze the impact of separation on diarization performance. For the mixed recording, we selected the first channel as our input. For evaluation on separated streams, we fixed the separation component as mask-based MVDR, and the ASR was chosen to be a speaker-biased hybrid HMM-DNN (described in Section 5.1). Table 2 shows the diarization performance on mixed LibriCSS, with a breakdown by overlap condition. The SAD error for the clustering-based systems was 4.8%. It is immediately evident that assigning overlapping speech is important to perform well on this task – both RPN and TS-VAD outperformed clustering-based methods. Even on low overlap regions, there is a significant difference, and we conjecture that this arises from a mismatch between the training data for the PLDA used for scoring the AHC and VBx models, and the evaluation set. Since SC uses cosine scoring, it performed better than the other clustering methods. For the RPN and TS-VAD systems, creating a simulated mixture which closely resembled the test set was found to be important, and using cepstral mean normalization (CMN) was critical for this performance.
Next, we evaluated the systems on separated audio streams obtained using the mask-based MVDR method. Table 3 shows these results, along with the downstream cpWER obtained using a hybrid HMM-DNN ASR model. Since AHC and SC perform subsegment-level clustering, they can be naturally extended to diarization across streams. VBx, on the other hand, estimates speaker changes through HMM state transitions, so it is not directly applicable to this scenario. RPN and TS-VAD can also be extended to this new setting, since RPN uses clustering of region embeddings, and TS-VAD predicts frame-level speaker activity based on the corresponding i-vectors. Our first observation is that among clustering-based methods, SC performed significantly better than AHC, likely because the PLDA used for AHC was trained on clean Librispeech utterances, which are acoustically very different from the separated audio. To verify this, we performed the AHC also on a cosine similarity matrix, and it resulted in an absolute DER improvement of 13.8%. For RPN and TS-VAD, performance on mixed recording was not indicative of results on separated audio. Without any post-processing, we obtained a DER of 26.9% using RPN. This improved to 22.4% on filtering out non-speech segments using our SAD from Section 4.1. Similarly, TS-VAD performance degraded severely on going from mixed to separated audio. We found this degradation to be consistent for missed speech, false alarms, and speaker confusions, indicating a likely mismatch in train vs. test conditions. Consequently, the cpWER for both RPN and TS-VAD was found to be significantly higher than that for SC.
Speech Recognition
We conducted our ASR experiments on a hybrid TDNN-F based, and a Transformer-based end-to-end ASR model. Pretrained models (trained on Librispeech) for both of these are publicly available, enabling reproducibility of our results.
Following the Kaldi Librispeech recipe, we trained a 17-layer deep neural network consisting of factored TDNN layers using the lattice-free MMI objective . We used 40-dim MFCC features, and additionally appended 100-dim i-vectors estimated online. The model was trained on the 960h Librispeech data with 3x speed perturbation. We call this our base model. On the Librispeech test-clean and test-other evaluation sets, this model obtains WERs of 3.8% and 8.8%, respectively. We additionally fine-tuned it for 1 epoch on Librispeech train set augmented with simulated room impulse responses to match the acoustic conditions of mixed LibriCSS recordings; this is referred to as the fine-tuned model. For decoding on LibriCSS, we used the 3-gram language model provided with the Librispeech data, and the lattices were rescored using a pruned TDNN-LSTM based RNNLM trained on the training transcripts . We used a 2-pass decoding strategy, where the i-vectors were re-estimated from the non-silence regions for the second pass . The fine-tuned model was used to evaluate the downstream ASR performance of the separation and diarization methods.
2 Transformer-based end-to-end ASR
Our end-to-end (E2E) ASR is a state-of-the-art ESPNet-based Transformer encoder-decoder model . It was trained on the 960h Librispeech corpus with SpecAugment , using 83-dim log-mel filterbank features with pitch. The encoder consists of 2 convolutional layers and 12 self-attention blocks, and the decoder contains 6 self-attention blocks. The training loss jointly minimizes sequence-to-sequence (S2S) and connectionist temporal classification (CTC) objectives . The decoder predicts subword units generated using SentencePiece . For decoding, we used beam search which combines scores from the S2S, CTC, and an external Transformer-based language model. This ASR model obtains WERs of 2.2% and 5.5% on the Librispeech test-clean and test-other evaluation sets, respectively, using a beam size of 60.
3 Experimental results
Similar to our diarization experiments, we evaluated the ASR models on mixed and separated audio. Table 4 shows the results obtained on the mixed LibriCSS data. For these experiments, we used oracle segments and speaker information. From the table, we see that a strong Librispeech model also performed well on mixed LibriCSS utterances. Among the TDNNF-based hybrid models, fine tuning on reverberated data provided a 31.5% relative WER improvement. For transformer-based E2E models, decoding with larger beams (of size 30) improved WER by almost 10% relative, compared to decoding with smaller beam sizes. We did not get any significant gains by increasing beam sizes further, so we used this setting for further experiments.
Next, we present ASR results in the context of our pipeline, i.e., using separated audio streams from the mask-based MVDR model and segments from spectral clustering based diarization. We used the fine-tuned variant of the hybrid HMM-DNN ASR model. Additionally, to emphasize the importance of the separation module, we also show cpWERs obtained on mixed recordings. The results are shown in Table 5. On replacing oracle segments with those obtained from diarization, the downstream cpWERs for hybrid and E2E ASR systems degraded by 23.4% and 37.6%, respectively. Interestingly, this increase was more prominent in the low and medium overlap conditions, which suggests that errors in these cases occured primarily from incorrect speaker assignment. On applying separation, the cpWERs improved significantly (although this was partly due to better diarization). In particular, we found that the E2E model benefited more from separated audio streams, with its cpWER improving from 27.1% to 13.4%. For both the models, the cpWER with separation was found to outperform the corresponding results on mixed recording even using oracle segments.
Finally, with all our experimental evaluations in place, we combined the best performing models at each stage of the pipeline. We prepared two variants: (a) without separation, and (b) with separation. For (a), we selected TS-VAD based diarization and the Transformer-based ASR model. Pipeline (b) consists of sequential multi-frame separation, followed by SC-based diarization and the same ASR. We present the final results for both variants in Table 6. It is evident that in the absence of an explicit separation module, even a well performing diarization and ASR combo is hamstrung. We discuss some more implications of these results in the next section.
Discussion
Our extensive experiments using a variety of separation, diarization, and ASR methods elicit important lessons for solving the multi-speaker speech recognition problem. First, the importance of explicit separation cannot be overstated — the best cpWER for a pipeline with separation is 46.9% relative better than one without it (Table 6). For separation, we found training and inference with larger chunks to perform better; there is a caveat, however — the improved performance is obtained through diarization tricks that can filter out the increased cross-talk with these models. Second, we observed that although new (supervised) diarization methods like RPN and TS-VAD provide substantial gains on mixed recordings, the performance does not carry over well to separated audio streams, and traditional clustering approaches can still outperform them in these settings (Table 3). In particular, the best diarization without separation was found to be 45.8% relative better than one with separation, which is highly counter-intuitive. This is particularly relevant for pipelines such as ours and the CSS pipeline, where separation is performed before diarization, and we regard this as an important direction for further investigation. There have also been concurrent efforts to apply diarization on separated audio streams . Finally, unlike diarization, ASR performance on mixed recordings was strongly indicative of the performance on separated audio streams. We found that a state-of-the-art ASR model trained on clean, single speaker utterances integrates well in the pipeline and results in the best available cpWER on this dataset. This is encouraging, especially in view of recent advances in end-to-end speech recognition.
Acknowledgment. The work reported here was started at JSALT 2020 at JHU, with support from Microsoft, Amazon, and Google.