Investigation of End-To-End Speaker-Attributed ASR for Continuous Multi-Talker Recordings

Naoyuki Kanda, Xuankai Chang, Yashesh Gaur, Xiaofei Wang, Zhong Meng, Zhuo Chen, Takuya Yoshioka

Introduction

Speaker-attributed automatic speech recognition (SA-ASR), which recognizes ”who spoke what”, is essential to meeting transcription. SA-ASR requires to count the number of speakers, transcribe the utterances, and identify or diarize the speaker of each utterance from conversational recordings where some utterances are usually overlapped. It has a long research history, from the projects in the early 2000’s to the recent international efforts such as the CHiME and DIHARD challenges. While significant progress has been made especially in multi-microphone settings (e.g., ), SA-ASR for monaural audio remains challenging due to the difficulty in handling overlapped speech for both ASR and speaker diarization/identification.

One dominant approach to SA-ASR is applying speech separation (e.g., ) before ASR and speaker diarization/iden-tification. However, a speech separation module is often designed and trained with a signal-level criterion and therefore suboptimal for the downstream modules. To overcome this problem, joint modeling of multiple modules has been investigated from a variety of view points. For example, a number of studies have investigated joint modeling of speech separation and ASR (e.g., ). Several methods were also proposed for integrating speaker identification and speech separation . A few studies attempted to improve the speaker diarization by leveraging ASR results .

However, only a limited number of research works investigated the joint modeling of all necessary modules of SA-ASR. proposed to generate transcriptions for different speakers interleaved by speaker role tags to recognize doctor-patient conversations based on a recurrent neural network transducer (RNN-T). Although promising results were shown, the method cannot deal with speech overlaps due to the monotonicity constraint of RNN-T. Furthermore, their method is difficult to extend to an arbitrary number of speakers because the target speaker roles need to be uniquely defined. In , the authors applied a similar technique to by interleaving multiple utterances with speaker identity tags instead of speaker role tags. To handle speakers who were unseen in the training data, the authors used speaker identity tags from the training data even for the unseen test speakers, or they simply applied a separated speaker diarization module. However, their method showed severe degradation of ASR and speaker diarization accuracy when the oracle utterance boundaries were not used. proposed a joint decoding framework for overlapped speech recognition and speaker diarization, where speaker embedding estimation and target-speaker ASR were performed alternately. While their formulation is applicable to any number of speakers, the method was actually implemented and evaluated in a way that could be used only for the two-speaker case, as target-speaker ASR was performed with an auxiliary output branch representing a single interference speaker .

Recently, an end-to-end (E2E) SA-ASR model has been proposed as a joint model of speaker counting, speech recognition, and speaker identification for monaural (possibly) overlapped speech . It was trained to maximize the joint probability for multi-talker speech recognition and speaker identification, and achieved a significantly lower speaker-attributed word error rate (SA-WER) than a system that separately performs overlapped speech recognition and speaker identification. However, the model only works with a speaker inventory that includes the profiles (i.e., embeddings) of all speakers involved in the input speech. This requirement strongly limited its application to real scenarios.

In this paper, we extend the previous E2E SA-ASR work to address the case where no speaker profile is available. Specifically, we propose to cluster the internal speaker representations of the E2E SA-ASR model to diarize the utterances of the speakers whose speaker profiles are not included in the speaker inventory. Combined with a silence-region detector, this also allows a very long-form signal spanning an entire meeting to be handled. We also propose a simple modification to the reference label construction for the E2E SA-ASR training to handle continuous multi-talker recordings more effectively. Comprehensive experimental results using the monaural LibriCSS dataset , consisting of eight-speaker sessions, show the effectiveness of the proposed method.

Review: E2E SA-ASR

In this section, we review the E2E SA-ASR method proposed in . The goal of this method is to estimate a multi-speaker transcription Y={y1,...,yN}Y=\{y_{1},...,y_{N}\} and the speaker identity of each token S={s1,...,sN}S=\{s_{1},...,s_{N}\} given acoustic input X={x1,...,xT}X=\{x_{1},...,x_{T}\} and a speaker inventory D={d1,...,dK}\mathcal{D}=\{d_{1},...,d_{K}\}. Here, NN is the number of the output tokens, TT is the number of the input frames, and KK is the number of the speaker profiles (e.g., d-vector ) in the inventory D\mathcal{D}. Following the idea of serialized output training (SOT) , the multi-speaker transcription YY is represented by concatenating individual speakers’ transcriptions interleaved by a special symbol ⟨sc⟩\langle sc\rangle representing the speaker change.

In the E2E SA-ASR modeling, it is assumed that the profiles of all the speakers involved in the input speech are included in D\mathcal{D}. Note that, as long as this assumption holds, the speaker inventory may include irrelevant speakers’ profiles.

2 Model architecture

Figure 1 shows the architecture of the E2E SA-ASR model. It consists of ASR-related blocks (shown in green), and speaker identification-related blocks (shown in yellow). The computation consists of the following five steps.

Given the acoustic input XX, an ASR encoder firstly converts XX into a sequence, HencH^{enc}, of embeddings for ASR, i.e.,

At the same time, a speaker encoder converts XX into a sequence, HspkH^{spk}, of embeddings representing the speaker features of the input XX as follows:

2.2 Step2: attention weight estimation

Secondly, at each decoder step nn, an attention module generates attention weight αn={αn,1,...,αn,T}\alpha_{n}=\{\alpha_{n,1},...,\alpha_{n,T}\} as

where unu_{n} is a decoder state vector at the nn-th step, and cn−1c_{n-1} is a context vector at the previous time step.

2.3 Step3: calculating context vector for ASR

Then, context vector cnc_{n} for the current decoder step nn is generated as a weighted sum of the encoder embeddings as follows:

2.4 Step4: speaker identification

At every decoder step nn, the attention weight αn\alpha_{n} is also applied to HspkH^{spk} to extract an attention-weighted average, pnp_{n}, of the speaker embeddings as

Note that pnp_{n} could be contaminated by interfering speech because some time frames include two or more speakers.

The speaker query RNN in Fig. 1 then generates a speaker query qnq_{n} given the speaker embedding pnp_{n}, the previous output yn−1y_{n-1}, and the previous speaker query qn−1q_{n-1}, i.e.,

With the speaker query qnq_{n}, an attention module for speaker inventory (shown as InventoryAttention in the diagram) estimates attention weight βn,k\beta_{n,k} for each profile in D\mathcal{D}:

The attention weight βn,k\beta_{n,k} can be seen as a posterior probability of person kk speaking the nn-th token given all the previous tokens and speakers as well as XX and D\mathcal{D}, i.e.,

Attention-weighted speaker profile dˉn\bar{d}_{n} is also calculated based on the attention weight βn,k\beta_{n,k} and input profile dkd_{k} as

2.5 Step5: ASR using context and speaker vectors

Finally, the output distribution for yny_{n} is estimated given the context vector cnc_{n}, the decoder state vector unu_{n}, and the weighted speaker vector dˉn\bar{d}_{n} as follows:

Here, it is assumed that cnc_{n} and unu_{n} have the same dimensionality, and WdW_{d} is a matrix to change the dimension of dˉn\bar{d}_{n} to that of cnc_{n}. Variable WoutW_{out} is the affine transformation matrix of the final layer. Typically, DecoderOut{\rm DecoderOut} consists of a single affine transform with a softmax output layer. However, in this work, we insert one LSTM just before the affine transform as it improves the efficacy of the SOT model as shown in .

3 Training

All network parameters are optimized by maximizing the speaker-attributed maximum mutual information criterion as follows:

Here, γ\gamma is a scaling parameter for the speaker estimation probability and is set to 0.1 per .

4 Decoding

An extended beam search algorithm is used for decoding for the E2E SA-ASR. With the conventional beam search, each hypothesis contains estimated tokens accompanied by the posterior probability of the hypothesis. In addition to these, a hypothesis for the E2E SA-ASR method contains speaker estimation βn,k\beta_{n,k}. Each hypothesis expands until ⟨eos⟩\langle eos\rangle is detected, and the estimated tokens in each hypothesis are grouped by ⟨sc⟩\langle sc\rangle to form multiple utterances. For each utterance, the speaker with the highest βn,k\beta_{n,k} value at the point of ⟨sc⟩\langle sc\rangle or ⟨eos⟩\langle eos\rangle token is selected as the predicted speaker of that utterance We observed slight performance improvement by using the speaker estimation at the end of an utterance (i.e., the ⟨sc⟩\langle sc\rangle or ⟨eos⟩\langle eos\rangle position) instead of the original scheme proposed in which uses the average βn,k\beta_{n,k} values calculated over all tokens of the utterance. Finally, when the same speaker is predicted for multiple utterances, those utterances are concatenated to form a single utterance.

Extensions of E2E SA-ASR

This section describes our proposed extensions of the E2E SA-ASR for recognizing continuous multi-talker recordings without prior speaker knowledge.

The E2E SA-ASR requires the speaker inventory to include the profiles of all speakers involved in the input speech. However, it is often difficult to prepare such a speaker inventory for various reasons, including the participation of guest speakers who are not originally invited to a meeting and the privacy concern about voice enrollment.

To cope with the case where no prior speaker knowledge is available, we combine the E2E SA-ASR and speaker clustering. Here, we assume we have a well-trained E2E SA-ASR model. Then, our proposed procedure to recognize long audio recordings is as follows.

Firstly, we apply a silence-region detector to divide an input long audio recording into multiple shorter segments at every silence regions. Each segment may include multiple utterances of different speakers with overlaps.

Then, we apply the E2E SA-ASR for each segment with a set of example speaker profiles who do not appear in the input audio.

Finally, we cluster the speaker query vectors qnq_{n} of the recognized hypotheses (i.e., the query vectors obtained at the last token of each utterance) to count and diarize the speakers. Specifically, we first determine the number of clusters based on normalized maximum eigengap (NME) , and then perform spectral clustering with a normalized graph Laplacian matrix . In , spectral clustering was applied to a binarized and unnormalized graph Laplacian matrix after speaker counting. However, we applied the conventional spectral clustering with a normalized graph Laplacian as this yielded slightly better results in our preliminary experiments.

One may have multiple questions about this procedure. For example, how many example (irrelevant) speaker profiles are necessary in step 2? How does the silence-region detector in step 1 affect the final result? How about using the weighted profile dˉn\bar{d}_{n} for speaker counting and clustering in step 3 instead of the speaker query qnq_{n}? We will experimentally examine these questions in Section 4.

2 Modified FIFO training

We also introduce a simple yet effective modification of the reference transcription construction for the E2E SA-ASR training. In the previous work , the authors trained the E2E SA-ASR model with overlapped speech of up to three utterances. However, in real conversation, there are many cases where the same speaker utters multiple times in one continuous audio segment as illustrated in the upper part of Fig. 2. In this example, three people are speaking in one audio segment, and rjir^{i}_{j} represents the jj-th reference token of speaker ii. The term Ni,uN^{i,u} represents the end position of uu-th utterance of speaker ii.

The previous work employed the first-in first-out (FIFO) training scheme , where the reference labels of different speakers are sorted by their start times and concatenated by ⟨sc⟩\langle sc\rangle token. Since the ⟨sc⟩\langle sc\rangle token represents the speaker change, the transcriptions of individual speakers are sorted by the times they start speaking. We call this original version speaker-based FIFO training, and shows an example in Fig. 2.

Alternatively, we may sort the reference labels according to the start time of each utterance and join the utterances with the ⟨sc⟩\langle sc\rangle token. Note that this scheme implicitly assumes that we can define what an end of an utterance is in continuous speech. We call this modified version utterance-based FIFO training, as illustrated in Fig. 2. In the next section, we experimentally investigate which FIFO training scheme results in better performance.

Experiments

We evaluated the effectiveness of the proposed method by using the LibriCSS dataset , which comprises conversation-like recordings created based on the LibriSpeech corpus . The dataset consists of 10 hours of recordings of concatenated LibriSpeech utterances that were played back by multiple loudspeakers in a meeting room and captured by a seven-channel microphone array. While the recordings have seven channels, we used only the first channel data (i.e. monaural audio) for all our experiments.

The LibriCSS dataset consists of 10 sessions, each being one hour long and comprising eight speakers. Per , each session is decomposed to six 10-minute-long “mini-sessions” that have different overlap ratios ranging from 0% to 40%. The recordings of the first session (Session 0) was used to tune the decoding parameters, and those in the rest of 9 sessions (Session 1–9) were used for the evaluation. Note that there are two types of mini-sessions for the 0% overlap case: one has only 0.1-0.5 sec of silence between adjacent utterances (called “0S”); one has 2.9-3.0 sec of silence between the adjacent utterances (called “0L”).

1.2 Training data

For the E2E SA-ASR training, we used multi-speaker signals that were generated by room simulation from the 960 hours of LibriSpeech training data (“train_960”) . We generated 500,000 training samples, each of which was a mixture of multiple utterances randomly selected from train_960. When the utterances were mixed, each utterance was shifted by a random delay to simulate partially overlapped conversational recordings. Each training sample was generated under the following conditions.

The number of speakers was randomly chosen from 1 to 5.

The number of utterances was randomly chosen from 1 to 5.

The start times of different utterances were apart by 0.5 sec or longer.

Every utterance in each mixed audio sample had at least one speaker-overlapped region with other utterances.

Utterances of the same speakers do not overlap.

Before mixing the source utterances, a room impulse response generated by the image method was applied to each utterance . In addition, random noise was generated by following , and added at a random SNR from 10 to 40 dB after mixing the utterances. Finally, the volume of the mixed audio was changed by a random scale between 0.125 and 2.0.

In addition to the multi-speaker signals, speaker profiles were generated for each training sample as follows. For a training sample consisting of SS speakers, the number of the profiles was randomly selected from SS to 8. Among those profiles, SS profiles were for the speakers involved in the overlapped speech. The utterances for creating the profiles of these speakers were different from those constituting the input overlapped speech. The rest of the profiles were randomly extracted from different speakers in train_960. Each profile was extracted by using 10 utterances.

1.3 Evaluation metric

The main evaluation metric used in this paper is the concatenated minimum-permutation word error rate (cpWER) . The cpWER is computed as follows: (i) concatenate all reference transcriptions for each speaker; (ii) concatenate all hypothesis transcriptions for each detected speaker; (iii) compute the WER between the reference and hypothesis and repeat this for all possible speaker permutations; and (iv) pick the lowest WER among them. The cpWER is affected by both the speech recognition and speaker diarization results.

Besides cpWER, we evaluated the mean speaker counting error, which is the absolute difference between the estimated number of speakers and the actual number of speakers (= 8 in LibriCSS) averaged over all mini-sessions. We also analyzed the source-target attention of our system in terms of the diarization error rate (DER). It should be noted that the mean speaker counting error and the DER are not the performance metrics we care, and they were evaluated only for analysis purposes. The hyper-parameters of our systems were tuned on the development set to improve only the cpWER.

1.4 Model settings

In our experiments, an 80-dim log mel filterbank extracted every 10 msec was used for the input feature. 3 frames of features were stacked, and the model was applied on top of the stacked features. For the speaker profile, we used a 128-dim d-vector , whose extractor was separately trained on VoxCeleb Corpus . The d-vector extractor consisted of 17 convolution layers followed by an average pooling layer, which was a modified version of the one presented in .

The AsrEncoder consisted of 5 layers of 1024-dim bidirectional long short-term memory (BLSTM), interleaved with layer normalization . The DecoderRNN consisted of 2 layers of 1024-dim unidirectional LSTM, and the DecoderOut consisted of 1 layer of 1024-dim unidirectional LSTM. We used a conventional location-aware content-based attention with a single attention head. The SpeakerEncoder had the same architecture as the d-vector extractor except for not having the final average pooling layer. Our SpeakerQueryRNN consisted of 1 layer of 512-dim unidirectional LSTM. We used 16k subwords based on a unigram language model as a recognition unit.

In addition to the E2E SA-ASR model described above, we trained an external language model (LM) that consisted of 4 layers of 2,048-dim LSTM. As training data, we generated a text corpus by (1) shuffling the official training text corpus for LibriSpeech and the transcription of train_960, and (2) concatenating every consecutive rand(1,5)rand(1,5) utterances interleaved by ⟨sc⟩\langle sc\rangle token. We used the shallow fusion (i.e. simple weighted sum) to combine the E2E SA-ASR and the LM scores with an LM weight calibrated by using the development set.

2 Evaluation with oracle silence boundary

We firstly evaluated the proposed method with an oracle silence-region detector. Namely, we divided each recording at every silence position obtained from the oracle utterance boundary information. Note that each segmented audio still consisted of multiple overlapped utterances of different speakers. The minimum and maximum numbers of utterances were found to be 1 and 24, respectively. In this subsection, we used the oracle silence detection. The performance using an automatic silence detector is reported in the next subsection.

As a baseline, we evaluated the E2E SA-ASR with a speaker inventory consisting only of the eight relevant speakers. Each speaker’s profile was extracted by using 5 utterances that were not included in the recording used for the evaluation. We firstly compared the speaker-based and utterance-based FIFO training schemes that we described in Section 3.2. The result is shown in Table 1. We can see that the utterance-based FIFO training significantly outperformed the speaker-based FIFO training. Therefore, we always used the E2E SA-ASR model based on the utterance-based FIFO training in the remaining experiments.

Next, we evaluated the accuracy of the E2E SA-ASR when the speaker inventory included irrelevant speaker profiles. In this experiment, irrelevant speakers were randomly chosen from train_960 of LibriSpeech, and a randomly selected one utterance was used to extract the speaker profile of each irrelevant speaker. The result is shown in the first five rows of Table 2. When no irrelevant profiles were included in the speaker inventory, the E2E SA-ASR achieved the best cpWER of 16.9%. The cpWER gradually deteriorated as the addition of irrelevant profiles, but the system still achieved 22.1% of cpWER even with 100 irrelevant profiles.

Finally, to analyze the impact of the speaker profiles, we also evaluated the E2E SA-ASR with no relevant speaker profiles. The results of this experiment are shown from the 6th to 8th rows of Table 2, where we provided 10, 20, or 100 irrelevant profiles as an input while not using any profiles for the relevant speakers. Speaker diarization was conducted purely based on the speaker identification result for each utterance. As expected, we observed a very high cpWER of 72.1–81.3%. Note that, the mean speaker counting error for the 10 irrelevant profile case was relatively small (= 1.19) just because the given (10) and correct (8) numbers of speakers were close.

2.2 Results of the proposed method

We then evaluated the proposed procedure of the combination of the E2E SA-ASR and speaker clustering. The results are shown in the last two rows of Table 2. In this experiment, we used 100 irrelevant speaker profiles as a set of example profiles. When we applied the speaker clustering with the oracle number of speakers, the proposed method achieved 16.7% of cpWER, which was even better than the best number obtained by the E2E SA-ASR with the relevant speaker inventory. This is because spectral clustering can access to the speaker embeddings of all utterances while the speaker identification inside the E2E SA-ASR was done by accessing only the information of the single segment. When we estimated the number of speakers by using NME with a maximum possible number of speakers of {8, 12, 16}, the cpWER was slightly degraded to 17.9–20.0%. Nonetheless, it was still as good as the E2E SA-ASR with 10-20 irrelevant profiles.

We also evaluated the effect of the number of irrelevant profiles (= example profiles) for the combination of the E2E SA-ASR and speaker clustering. The result of this study is shown in Table 3. It can be seen that using too few irrelevant profiles resulted in the degradation of cpWER. It is because we cannot calculate an appropriate weighted profile dˉn\bar{d}_{n} when we have too few profiles, which ends up with degrading the overall accuracy. Note that the computational cost of the inventory attention (Eq. (8)–(11)) was negligible even with 100 profiles. Thus, we used 100 irrelevant speaker profiles in the following experiments unless otherwise stated.

We also compared clustering using the weighted profile dˉn\bar{d}_{n} and that using speaker query qnq_{n}. The results are shown in Table 4. In this experiment, we applied the E2E SA-ASR with 100 irrelevant speaker profiles, and then applied speaker clustering given the oracle number of speakers. As seen in the table, the use of the speaker query qnq_{n} resulted in significantly better speaker clustering performance.

3 Evaluation with automatic silence-region detector

We finally evaluated the proposed method with an automatic silence-region detector. In this experiment, we applied the WebRTC Voice Activity Detectorhttps://github.com/wiseman/py-webrtcvad for each recording, and segmented the audio whenever silence regions were detected.

The result with the automatic silence-region detector is shown in Table 5. The original E2E SA-ASR with the relevant speaker inventory achieved 18.6% to 26.0% of cpWER depending on the number of the additional irrelevant profiles. On the other hand, the proposed combination of the E2E SA-ASR and speaker clustering achieved 19.2% of cpWER with oracle speaker counting, and 21.8% of cpWER with NME-based speaker counting, respectively.

Compared with the case using the oracle silence-region information, the cpWER was degraded by 3.1%. Especially, we noticed that “0S” setting showed a severe cpWER degradation even though the overlap ratio was 0%. With “0S”, there was very short silence (0.1-0.5 sec) between adjacent utterances of different speakers. As a result, segments in “0S” often consisted of consecutive speech of multiple speakers. We observed the E2E SA-ASR sometimes misrecognized the speaker change point for such speech, which resulted in the degradation of cpWER. Note that the speaker change detection for non-overlapped speech could be more difficult than that for overlapped speech because speech overlaps could be used as a clue of speaker change besides the difference of voice characteristics.

3.2 Analysis of the source-target attention with DER

We analyzed the source-target attention αn\alpha_{n} of the E2E SA-ASR. We estimated the start and end times of each utterance based on αn\alpha_{n} as follows and calculated the DER accordingly.

For each utterance hypothesis, the attention (αn\alpha_{n})-weighted average of the frame indices was calculated for each token other than ⟨sc⟩\langle sc\rangle or ⟨eos⟩\langle eos\rangle.

The minimum frame index fminf_{min} and the maximum frame index fmaxf_{max} were calculated.

Here, TfT_{f} is the frame shift in second, and it was 0.03 sec according to our model settings. The term TmT_{m} is a heuristic margin tuned by the development set, and it was determined as 0.5 sec in our experiment.

The DER result is shown in Table 6. In this evaluation, we calculated the DER without a collar margin, and the overlapping regions were included in the DER calculation. As shown in the table, the E2E SA-ASR systems showed 15.23–16.75% of DER on average. In the high overlap test sets (with the overlap ratios of 20%–40%), the DERs were significantly better than the overlap ratios of the input audio, which indicates that the source-target attention scanned the encoder embeddings back and forth to recognize overlapped utterances one by one as originally designed by SOT . On the other hand, the DER was as high as 11.15% even for the non overlapped speech (0L). This could be because the our model is optimized to achieve good SA-ASR accuracy, unlike other diarization methods, such as the end-to-end neural diarization or target-spekaer voice activity detection , that are optimized for DER. That being said, the result shows that the source-target attention in the E2E SA-ASR model provides information about the start and end times of the hypotheses and thus can be used for applications requiring both the time boundary and the recognition result.

Conclusion

In this paper, we proposed to apply speaker counting and clustering to the speaker query of an E2E SA-ASR model to diarize utterances of speakers whose speaker profiles are not included in the speaker inventory. We also proposed a simple yet effective modification to the reference label construction for E2E SA-ASR training, which helps cope with the continuous multi-talker recordings. In the evaluation, compared with the original E2E SA-ASR with a speaker inventory consisting only of relevant speaker profiles, the proposed method achieved a close cpWER even without any prior speaker knowledge.

References