Streaming end-to-end multi-talker speech recognition

Liang Lu, Naoyuki Kanda, Jinyu Li, Yifan Gong

I Introduction

Overlapped speech is ubiquitous among natural conversations and meetings. For automatic speech recognition (ASR), recognizing overlapped speech has been a long-standing problem. A common practice is to follow the divide-and-conquer strategy, e.g., applying speech separation cascaded with a single-speaker speech recognition model . While this approach has enjoyed significant progress thanks to the achievement in deep learning based speech separation , there are two key drawbacks with this paradigm. Firstly, the overall system is cumbersome, especially given the increasing complexity of both speech separation and speech recognition modules. Consequently, maintaining and developing the cascaded system requires significant engineering effort. Secondly, each module in the cascaded system is optimized independently, which does not guarantee the overall performance improvement.

Recently, there have been considerable amount of work on the end-to-end approach for overlapped speech recognition. End-to-end speech recognition models, such as Connectionist Temporal Classification (CTC) , attention-based sequence-to-sequence model (S2S) , and Recurrent Neural Network Transducer (RNN-T) have been explored to address this challenge. In particular, Settle et al. proposed a model with joint speech separation and recognition training. Chang et al. applied multi-task learning with CTC and S2S to train an end-to-end model for overlapped speech recognition. Kanda et al. proposed Serialized Output Training (SOT) for S2S-based end-to-end multi-talker speech recognition. RNN-T has also been investigated for overlapped speech recognition in in an offline setting with bidirectional long short-term memory (LSTM) networks and auxiliary masking loss functions. Compared with the joint speech separation and recognition approach using an hybrid model, the end-to-end approach enjoys lower system complexity and high flexibility . While the progress in end-to-end overlapped speech recognition is promising, to the best of our knowledge, all previous studies only consider the offline condition, which assumes that the overlapped audio has been segmented. Thus, these systems cannot be deployed for streaming first-pass speech recognition scenarios that require low recognition latency such as online speech transcription for meetings and conversations.

In this paper, we propose the Streaming Unmixing and Recognition Transducer (SURT) for multi-talker speech recognition. Our model relies on RNN-T as the backbone, and it can transcribe the overlapped speech into multiple streams of transcriptions simultaneously with very low latency. In this work, we investigate two different network architectures. The first architecture employs a mask encoder to separate the feature representations, while the second model uses a speaker-differentiator encoder for this purpose. To train SURT, we study an approach similar to the one applied in , which we refer to as Heuristic Error Assignment Training (HEAT) for the clarity of presentation. This approach can be viewed as a simplified version of the widely used Permutation Invariant Training (PIT) by picking only one label assignment based on heuristic information. Compared with PIT, HEAT consumes much less memory, and is more computationally efficient. To evaluate the proposed SURT model, we performed experiments using the LibriSpeechMix dataset , which simulate the overlapped speech data from the LibriSpeech corpus . We show that SURT can achieve strong recognition accuracy with 150 milliseconds algorithmic latency compared with an offline S2S model trained with PIT.

The contributions of the paper are summarized as follows.

We perform the first study on streaming end-to-end multi-talker speech recognition, and propose an RNN-T based model for this problem. We also demonstrate a strong recognition accuracy compared with an offline system.

PIT is the mostly commonly used loss function for multi-talker speech processing, while an approach similar to HEAT has only been used in . We present a rigid comparison of the two approaches in dealing with the label ambiguity problem, and show the superiority of HEAT over PIT in terms of computational efficiency and model accuracy for our problem.

We propose two Unmixing architectures that are inspired from related works.

II Related Work

There have been a few studies on S2S and joint CTC/attention models for end-to-end overlapped speech recognition , however, these works are all in the category of offline condition. To the best of our knowledge, our work is the first study on streaming end-to-end overlapped ASR. The work that is most closely related to our work is RNN-T based approach for end-to-end overlapped ASR done by Tripathi et al. . However, the authors also focus on the offline scenario in their work. In addition, the authors in applied carefully designed auxiliary loss functions for signal reconstruction to train the RNN-T model, while in our work, we apply a single ASR loss function for model training, which simplifies the system development. Besides, the model architectures and loss functions are also different in this work.

III RNN-T

RNN-T is a time-synchronous model for sequence transduction, which works naturally for end-to-end streaming speech recognition. Given an acoustic feature sequence X={x1,⋯ ,xT}X=\{x_{1},\cdots,x_{T}\} and its corresponding label sequence Y={y1,⋯ ,yU}Y=\{y_{1},\cdots,y_{U}\}, where TT is the length of the acoustic sequence, and UU is the length of the label sequence, RNN-T is trained to directly maximizing the conditional probability

where ft{\bm{f}}_{t} and gu{\bm{g}}_{u} are the output vectors from the audio encoder network and the label encoder network followed by an affine transform at the time step tt and uu respectively, and J(⋅)J(\cdot) denotes a nonlinear activation function followed by an affine transform. Vˉ\bar{\mathcal{V}} denotes the set of the vocabulary V\mathcal{V} with an additional blank token, i.e., Vˉ=V∪\O\bar{\mathcal{V}}=\mathcal{V}\cup\O. Given the distribution of each timestep (t,u)(t,u), the sequence-level conditional probability Eq. (1) can be obtained by the forward-backward algorithm, where the forward variable is defined as

with the initial condition α(1,0)=1\alpha(1,0)=1, while the backward variable can be defined similarly. The probability P(Y∣X)P(Y\mid X) can be computed as

RNN-T is trained by minimizing the negative log-likelihood as:

IV Streaming Unmixing and Recognition Transducer

In this work, we focus on the 2-speaker case for overlapped speech recognition, and the proposed SURT model is shown in Fig.1-a). We denote the overlapped acoustic sequence as XX, and the label sequences are Y1Y^{1} and Y2Y^{2}. Given the overlapped speech XX, the Unmixing module extracts the speaker-dependent features representations, H1H_{1} and H2H_{2}, which are then fed into the RNN-T module. The role of the Unmixing module is similar to speech separation, however, we do not apply any speech separation loss in training. Instead, the whole SURT model is trained end-to-end using a speech recognition loss as defined in section IV-C. In this section, we firstly discuss two network structures for the Unmixing module, before explaining the loss functions to train the models.

Inspired by , we use two speaker-differentiator (SD) encoders to construct the Unmixing module as shown in Fig.1-b). The speaker-dependent feature representations H1H_{1} and H2H_{2} are obtained as

where MixEnc is an encoder used to pre-process the overlapped speech signals; SD1 and SD2 are two difference encoders to generate the two feature sequences. While many different neural network encoders are applicable, we focus on convolution neural networks (CNNs) in this work as detailed in the experimental section.

IV-B Mask-based Unmixing Model

Inspired by works in speech separation , we define a mask-based Unmixing module as Fig.1-c), in which, H1H_{1} and H2H_{2} are obtained as

where σ\sigma denotes the Sigmoid function, and MaskEnc is the encoder to estimate the mask MM; MixEnc is the pre-processing encoder as discussed before, and \mathds1\mathds{1} is a tensor of the same shape as MM, and each of its elements is 1; ∗* denotes element-wise multiplication.

IV-C Loss Functions

For model training, we study two loss functions, i.e., Permutation Invariant Training (PIT) and Heuristic Error Assignment Training (HEAT).

PIT has been widely used for speech separation and multi-talker speech recognition due to its simplicity and superior performance. The key problem in overlapped speech separation and recognition, as argued in , is the label ambiguity issue, i.e., it is unclear if the feature representation H1H_{1} corresponds to Y1Y^{1} or Y2Y^{2}. To address this problem, PIT considers all the possible error assignments when computing the loss, and hence, it is invariant to the label permutations. For the 2-speaker case studied in this work, the PIT loss can be expressed as:

While being simple and effective, PIT also has drawbacks. In particular, it is not very scalable to the number of speakers in the mixed signal. For the S−S-speaker case, the total number of permutations is S!S!, which will require to compute the RNN-T loss S!S! times in the framework of SURT. We could use Hungarian algorithm [kuhn1955hungarian] to reduce the computation from O(S!)O(S!) to O(S3)O(S^{3}), but it is still clearly not affordable due to the high computational and memory cost of the RNN-T loss.

IV-C2 Heuristic Error Assignment Training

Different from PIT, HEAT only picks one possible error assignment based on some heuristic information that can disambiguate the labels. In this work we particularly use the heuristic to disambiguate the labels based on the start times that they were spoken, e.g.,

where Y1Y^{1} always refers to the utterance that was spoken first in our setting. Similar approach has been used in , and the authors also tried other heuristic information such as the time boundaries which were used to mask the encoder embedding vectors and define the mapping between (H1,H2)(H_{1},H_{2}) and (Y1,Y2)(Y^{1},Y^{2}). They also introduced auxiliary loss functions, while in our work, we prefer Eq. (6) for simplicity. With HEAT, the model will be trained to produce the hidden representations H1H_{1} that match the label sequence Y1Y^{1}. Note that, it does not make any difference if we swap H1H_{1} and H2H_{2}, as before model training, the model parameters do not have any label correspondence yet. However, once the mapping function is chosen, we have to fix it during model training. Compared with PIT, HEAT is more scalable and memory efficient, as it only evaluates the RNN-T loss SS times for the SS-speaker case.

V Experiments and Results

Our experiments were performed on the simulated LibriSpeechMix dataset , which is derived from the 1,000 hour LibriSpeech corpus by simulating the overlapped audio segments. We used the same protocol to simulate the training and evaluation data as in . The source code to reproduce our evaluation data is publicly availablehttps://github.com/NaoyukiKanda/LibriSpeechMix. To generate the simulated training data, for each utterance in the original LibriSpeech train_960 set, we randomly pick another utterance from a different speaker, and mix the latter with the previous one with a random delay sampled from [τ,ν][\tau,\nu], in which τ\tau and ν\nu are the minimum and maximum delay in seconds respectively, as shown in Figure 2. ν\nu is always the same as the length of the first utterance, and we evaluate two different values of τ\tau in our experiments, i.e., τ=0\tau=0 and τ=0.5\tau=0.5. We used the same approach to generate the dev-clean and test-clean datasets. The number of mixed audio is the same as the number of utterances in the original LibriSpeech dataset. For both training and evaluation data, each utterance only has 2 speakers after simulation.

V-B Experimental Setup

In our experiments, we used the magnitude of the 257-dimensional short-time Fourier transform (STFT) as raw input features, which are sampled as the 10 milliseconds frame rate. The features were then spliced by a context window of 3 and downsampled by a factor of 3, results in 771-dimensional features at the frame rate of 30 milliseconds. We then reshaped the feature sequences to have 3 input channels. We used 4,000 word-pieces as the output tokens for RNN-T, which are generated by byte-pair encoding (BPE) . We set the dropout ratio as 0.2 for LSTM layers, and applied one layer of time-reduction to further reduce the input sequence length by the factor of 2 . We also applied speed perturbation for data augmentation with perturbation ratios as 0.9 and 1.1 [ko2015audio].

We used 4-layer 2D CNN encoder for MixEnc in the Unmxing module of the SD-based SURT model. The detailed configuration is shown in Table I. We used a 2-layer unidirectional LSTM with 1024 hidden units for SD1 encoder, SD2 encoder, the audio encoder and the label encoder of RNN-T. For the mask-based SURT model, we used the same CNN encoder as in Table I for MixEnc and MaskEnc for the Umixing module. The label encoder of RNN-T is a 2-layer unidirectional LSTM with 1024 hidden units as in the SD-based model, and the audio encoder is a 6-layer unidirectional LSTM with 1024 hiddent unit. The total number of model parameters is around 80 million (M) for both model architectures, and the algorithmic latency for both types of model is 5 frames, corresponding to 150 milliseconds, which is incurred by the convolution module. In our experiments, the models were trained using Adam optimizer with the intial learning rate as 4×10−44\times 10^{-4}, and halved the learning rate every 40,000 updates. We used data parallelism across 16 GPUs, and the mini-batch size for each GPU is 5,000 frames for both model architectures. During evaluation, the model produces two transcriptions in the 2-speaker case. For scoring, we follow the same protocol as in by choosing the label permutation yielding the lowest word error rate (WER).

V-C Results

Table II shows the WER results of the SD-based model. In particular, we evaluated two conditions when generating the mixed speech signals, i.e., τ=[0,0.5]\tau=[0,0.5], for both training and evaluation data. From the results in Table II, we observe that using the training data with the minimum delay τ=0.5\tau=0.5, the model achieved consistent lower WERs in both evaluation conditions compared with the model trained with data of τ=0\tau=0. Our interpretation is that the starting region of the speech signal that has no overlap can provide a strong cue for the model to track the first speaker and disentangle the overlapped signals. This information also makes the recognition task easier, as we observe that the model can achieve consistent lower WER for the evaluation condition τ=0.5\tau=0.5 compared with the evaluation condition of τ=0\tau=0.

The results also shown that HEAT can achieve lower WERs compared with PIT. To further understand the behaviors of the two loss functions, we plot the convergence curves of the models trained with PIT and HEAT in Figure 3. The horizontal axis indicates the validation loss values, while the vertical axis represents the number of model updates. In this comparison, we used exactly the same experimental setting for model training. The figure shows that the two approaches can result in very similar convergence speed, and HEAT can reach to a lower validation loss. As discussed before, HEAT is also faster than PIT, and we can use a larger mini-batch size as HEAT requires less memory.

Table III compares the SD-based model with the Mask-based model with (w/) or without (w/o) the MixEnc encoder, and the results show that the Mask-based model achieved much lower WERs. Finally, Table IV compares the proposed Mask-based SURT model (w/ MixEnc) with an offline LSTM-based S2S model trained with PIT . SURT achieved comparable results with half of the number of model parameters and with a very low latency constraint.

VI Conclusions

Overlapped speech recognition remains a challenging problem in the speech research community. While all the existing end-to-end approaches tackling this problem work in the offline condition, we proposed Streaming Unmixing and Recognition Transducer (SURT) for end-to-end multi-talker speech recognition, which can meet various latency constraints. In this work, SURT relies on RNN-T as the backbone, while other types of streaming transducers such as Transformer Transducers are also applicable. We investigated two different model architectures, and two different loss functions for the proposed SURT model. Based on experiments using the LibrispeechMix dataset, we achieved strong recognition accuracy with very low latency and a much smaller model compared with an offline PIT-S2S model.

References