Segmental Audio Word2Vec: Representing Utterances as Sequences of Vectors with Applications in Spoken Term Detection

Yu-Hsuan Wang, Hung-yi Lee, Lin-shan Lee

Introduction

In natural language processing, it is well known that Word2Vec transforming words (in text) into vectors of fixed dimensionality is very useful in various applications, because those vectors carry semantic information. In speech signal processing, it has been shown that audio Word2Vec transforming spoken words into vectors of fixed dimensionality is also useful for example in spoken term detection or data augmentation, because those vectors carry phonetic structure for the spoken words. It has been shown that this audio Word2Vec can be trained in a completely unsupervised way from an unlabeled dataset, except the spoken word boundaries are needed. The need for spoken word boundaries is a major limitation for audio Word2Vec, because word boundaries are usually not available for given speech utterances or corpora.

Although it is possible to use some automatic processes to estimate word boundaries followed by the audio Word2Vec, it is highly desired that the signal segmentation and audio Word2Vec may be integrated and jointly learned, because in that way they may enhance each other. This means the machine learns to segment the utterances into a sequence of spoken words, and transform these spoken words into a sequence of vectors at the same time. This is the segmental audio Word2Vec proposed here: representing each utterance as a sequence of fixed-dimensional vectors, each of which hopefully carries the phonetic structure information for a spoken word. This actually extends the audio Word2Vec from word-level up to utterance-level. Such segmental audio Word2Vec can have plenty of potential applications in the future, for example, speech information summarization, speech-to-speech translation or voice conversion. Here we show the very attractive first application in spoken term detection.

The segmental audio Word2Vec proposed in this paper is based on a segmental sequence-to-sequence autoencoder (SSAE) for learning a segmentation gate and a sequence-to-sequence autoencoder jointly. The former determines the word boundaries in the utterance, and the latter represents each audio segment with an embedding vector. These two processes can be jointly learned from an unlabeled corpus in a completely unsupervised way. During training, the model learns to convert the utterances into sequences of embeddings, and then reconstructs the utterances with these sequences of embeddings. A guideline for the proper number of vectors (or words) within an utterance of a given length is needed, in order to prevent the machine from segmenting the utterances into more segments (or words) than needed. Since the number of embeddings is a discrete variable and not differentiable, the standard back-propagation is not applicable. The policy gradient for reinforcement learning is therefore used. How these generated word vector sequences carry the phonetic structure information of the original utterances was evaluated with the real application task of query-by-example spoken term detection on four languages: English (on TIMIT), Czech, French, German (on GlobalPhone corpora).

Proposed Approach

The proposed structure for SSAE is depicted in Fig. 1, in which the segmentation gate is inserted into the recurrent autoencoder. For an input utterance X == {x1\mathbf{x}_{1}, x2\mathbf{x}_{2}, …, xT\mathbf{x}_{T}}, where xt\mathbf{x}_{t} represents the t-th acoustic feature like MFCC and TT is the length of the utterance, the model learns to determine the word boundaries and produce the embeddings for the NN generated audio segments, Y = { e1,e2,...,eN\mathbf{e}_{1},\mathbf{e}_{2},...,\mathbf{e}_{N}}, where en\mathbf{e}_{n} is the n-th embedding and N≤TN\leq T.

The proposed SSAE consists of an encoder RNN (ER) and a decoder RNN (DR) just like the conventional autoencoder. But the encoder includes an extra segmentation gate, controlled by another RNN (shown as a sequence of blocks S in Fig. 1). The segmentation problem is formulated as a reinforcement learning problem. At each time tt, the segmentation gate agent performs an action ata_{t}, ”segment” or ”pass”, according to a given state st\textbf{s}_{t}. xt\textbf{x}_{t} is taken as a word boundary if ata_{t} is ”segment”.

For the segmentation gate, the state at time tt, st\textbf{s}_{t}, is defined as the concatenation of the input xt\mathbf{x}_{t}, the gate activation signal (GAS) gt\textbf{g}_{t} extracted from the gates of the GRU in another pre-trained RNN autoencoder , and the previous action at−1a_{t-1} taken ,

The output ht\mathbf{h}_{t} of layers of the segmentation gate RNN (blocks S in Fig. 1) followed by a linear transform (WπW^{\pi},bπ\mathbf{b}^{\pi}) and a softmax nonlinearity models the policy πt\pi_{t} at time t,

This πt\pi_{t} gives two probabilities respectively for ”segment” and ”pass”. An action ata_{t} is then sampled from this distribution during training to encourage exploration. During testing ata_{t} is ”segment” whenever its probability is higher.

When ata_{t} is ”segment”, the time tt is viewed as a word boundary, and the segmentation gate passes the output of encoder RNN as an embedding. The state of the encoder RNN is also reset to its initial value. So the embedding en\mathbf{e}_{n} is generated based on the acoustic features of the audio segment only, independent of the previous input in spite of the recurrent structure,

where t1t_{1}, t2t_{2} refers to the beginning and ending time for the n-th audio segment.

The input utterance X should be reconstructed with the embedding sequence Y = { e1,e2,...,eN\mathbf{e}_{1},\mathbf{e}_{2},...,\mathbf{e}_{N}}. Because the decoder RNN (DR) is backward in order as shown in Fig. 1 , for the embedding en\mathbf{e}_{n} for the input segment from t1t_{1} to t2t_{2} in Eq.(4) above, the reconstructed feature vector is,

The decoder RNN is also reset when beginning decoding each audio segment to remove the information flow from the following segment.

2 Encoder and Decoder Training

3 Segmentation Gate Training

The segmentation gate is trained with the reinforcement learning. After the segmentation gate performs the segmentation for each utterance, it receives a reward rr and a reward baseline rbr_{b} for the utterance for updating the parameters. rr and rbr_{b} will be defined in the next subsection. We can write the expected reward for the gate under policy π\pi as J(θ)=Eπ[r]J(\theta)=\mathbf{E}_{\pi}[r], where θ\theta is the parameter set. The updates of the segmentation gate are simply given by:

where πt(θ)(at)\pi^{(\theta)}_{t}(a_{t}) is the probability for the action ata_{t} taken as in Eq.(3).

3.2 Rewards

The reconstruction error is certainly a good indicator to see whether the segmentation boundaries are good, since the embeddings are generated based on the segmentation. We hypothesize that good boundaries, for example those close to word boundaries, would result in smaller reconstruction errors, because the audio segments for words would appear more frequently in the corpus and thus the embeddings would be trained better giving lower reconstruction errors. So the smaller the reconstruction errors the higher the reward:

This is very similar to Eq.(6) except for a specific utterance here.

On the other hand, a guideline for the proper number of segments (words) NN in an utterance of a given length TT is important, otherwise for minimizing the reconstruction error as many segments as possible will be generated. So the smaller number of segments NN normalized by the utterance length TT, the higher the reward:

where NN and TT are respectively the numbers of segments and frames for the utterance as in Fig. 1.

The total rewards rr is obtained by choosing the minimum between rMSEr_{\textit{MSE}} and rN/Tr_{\textit{N/T}}:

where λ\lambda is a hyperparameter to be tuned for a reasonable guideline for estimating the proper number of segments for an utterance of length TT. In our experiments, this minimum function gave better results than linear interpolation.

We further use utterance-wise reward baseline to remove the bias between utterances. For each utterance, MM different sets of segment boundaries are sampled by the segmentation gate, each used to evaluate a reward rmr_{m} with Eq.(10). The reward baseline rbr_{b} for the utterance is then the average of them:

4 Iterative Training Process

Although all the models described in sections 2.2 and 2.3 can be trained simultaneously, we actually trained our model with an iterative process consisting of two phases. The first phase is to train the encoder and decoder with Eq.(6) while fixing the parameters of the segmentation gate. The second phase is to update the parameters of the segmentation gate with rewards provided by the encoder and decoder while fixing their parameters. The two phases are performed iteratively. In phase one, the encoder and decoder should be learned from random initialized parameters each time, instead of taking the parameters learned in the previous iteration as the initialization, which was found to offer better training stability.

Example Application: Unsupervised Query-by-example Spoken Term Detection

This approach can be used in many potential applications. Here we consider the unsupervised query-by-example spoken term detection (QbE STD) as the first example application. The task of unsupervised QbE STD is to locate the occurrence regions of the input spoken query in a large spoken archive without performing speech recognition. With the SSAE proposed here, this can be achieved as illustrated in Fig. 2. Given frame sequences of a spoken query and a spoken document, SSAE can represent these sequences as embeddings, q = { q1\mathbf{q}_{1}, q2\mathbf{q}_{2}, …, qNq\mathbf{q}_{N_{q}}} for the query and d = { d1\mathbf{d}_{1}, d2\mathbf{d}_{2}, …, dNd\mathbf{d}_{N_{d}} } for the document. With the embeddings, simply subsequence matching can be used to evaluate the relevance score S(q,d)S(\textbf{q},\textbf{d}) between q and d:

Cosine similarity can be used in the similarity measure in Eq.(13). As is clear in the right part of Fig. 2, S1S_{1} = sim(q1,d1)⋅sim(q2,d2)sim(\mathbf{q}_{1},\mathbf{d}_{1})\cdot sim(\mathbf{q}_{2},\mathbf{d}_{2}), S2S_{2} = sim(q1,d2)⋅sim(q2,d3)sim(\mathbf{q}_{1},\mathbf{d}_{2})\cdot sim(\mathbf{q}_{2},\mathbf{d}_{3}) and so on. The relevance score S(q,d)S(\mathbf{q},\mathbf{d}) in Eq.(12) between the query and document is then the maximum out of all SnS_{n}’s obtained in Eq.(13). In this way, the frame-based template matching such as DTW can be replaced by segment-based subsequence matching with much less on-line computation requirements.

Experiments

We performed the experiments on four different languages: English, Czech, French, German. The English corpus was TIMIT and the corpus for the other languages was the GlobalPhone. The ground truth word boundaries for English were provided by TIMIT, while for the other three languages we used the forced aligned word boundaries. Both the encoder and decoder RNNs of the SSAE consisted of one hidden layer of 100 LSTM units. The segmentation gate consisted of 2 layers of LSTM of size 256. All parameters were trained with Adam. M=5M=5 in Eq.(11) in estimating the reward baseline for each utterance. The proximal policy optimization algorithm was used to train the reinforcement learning model. The tolerance window for word segmentation evaluation was taken as 40 ms. The acoustic features used were 39-dim MFCCs with utterance-wise cepstral mean and variance normalization (CMVN) applied. In our experiments λ=5\lambda=5 in Eq.(10), which was obtained empirically and obviously had to do with the average duration of the segmented spoken words. In spoken term detection, 5 words for each language containing a variety of phonemes were randomly selected to be the query words as listed in Table 1, and several occurrences for each of them in training set were used as the spoken queries. The testing set utterances were used as spoken documents . The numbers of spoken queries used for evaluation on English, Czech, French and German were 29, 21, 25 and 23 respectively.

2 Spoken Word Segmentation Evaluation

Fig. 3 shows the learning curves for SSAE on the Czech validation set. From the figure, we can see that SSAE gradually learned to segment utterances into spoken words because both the precision and recall (blue curves in Fig. 3(a)(b) respectively) got higher when the reward rN/Tr_{N/T} in Eq.(9) (red curves) converged to a reasonable number. Similar trends were found in the other three languages.

We evaluated the spoken word segmentation performance of the proposed SSAE by comparison with the random segmentation baseline and two segmentation methods, one using Gate Activation Signals (GAS) and the other using the hierarchical agglomerative clustering (HAC), and the results are shown in Table 2 in terms of F1 score. Precision (P) and recall (R) were also provided for English. We see that the proposed SSAE performed significantly better than the other two methods on all languages except comparable to GAS for German. Also, for English recall about 50% was achieved while the precision was significantly lower, which implies many of the word boundaries were actually identified, but many spoken words were in fact segmented into subword units. Similar trends were found for other languages.

3 Spoken Term Detection (STD) Evaluation

We evaluated the quality of embeddings generated by SSAE with the real application of spoken term detection using the method presented in section 3, compared with other kinds of audio Word2Vec embeddings trained with signal segments generated from different segmentation methods. Mean Average Precision (MAP) was used as the performance measure.

The results are listed in Table 3. The performance for embeddings trained with ground truth word boundaries (oracle) in the last column serves as the upper bound. The random baseline in the first column simply assigned a random score to each pair of query and document. We also list the performance of standard frame-based dynamic time warping (DTW) as a primary baseline in the second column. From the table, it is clear that the oracle achieved the best and significantly better performance than all other methods on all languages. SSAE outperformed the DTW baseline by a wide gap. This is probably because DTW may not be able to identify the spoken words if the speaker or gender characteristics are very different, but such different signal characteristics may be better absorbed in the audio Word2Vec training. These experimental results verified that the embeddings obtained with SSAE did carry the sequential phonetic structure information in the utterances, leading to the better performance in STD here. The performance of embeddings trained with GAS and HAC are not too far from random in most cases. It seems the performance of the spoken word segmentation has to be above some minimum level, otherwise the audio Word2Vec couldn’t be reasonably trained, or spoken word segmentation boundaries had the major impact on the STD performance.

However, interestingly, although the segmentation performance of GAS was slightly better than SSAE for German, SSAE outperformed GAS a lot on spoken term detection for German. The reason is not clear yet, probably due to some special characteristics of the German language.

Conclusion

We propose in this paper the segmental sequence-to-sequence autoencoder (SSAE), which jointly learns and performs the spoken word segmentation and audio word embedding together. This actually extends the audio Word2Vec from word-level to utterance-level. This is achieved by reinforcement learning considering both the reconstruction errors obtained with the embeddings and the reasonable number of words within the utterances. Due to the reset mechanism in SSAE, an embedding is generated only based on an audio segment, therefore can be regarded as the audio word vector representing the segment. This is verified by the improved performance in experiments on unsupervised word segmentation and spoken term detection on four languages.

References