SpecAugment on Large Scale Datasets

Daniel S. Park, Yu Zhang, Chung-Cheng Chiu, Youzheng Chen, Bo Li, William Chan, Quoc V. Le, Yonghui Wu

Introduction

Data augmentation has been a successful method for improving generalization performance in Automatic Speech Recognition (ASR). Recently, SpecAugment , an augmentation scheme that directly augments the spectrogram of the input utterance, has shown surprising effectiveness in improving the performance of ASR networks on the 960h Librispeech and 300h Switchboard datasets. One natural question that arises is whether the effectiveness of SpecAugment persists for large scale tasks.

In this paper, we address this question by applying SpecAugment to the Google Multidomain Dataset introduced in . The Google Multidomain Dataset is a large scale multi-domain dataset, with multiple test sets from disparate domains. All data in the data set is anonymized. We compare the performance of the trained network with respect to the various forms of augmentation applied to the data, the results of which are summarized in table 1. In , Multistyle TRaining (MTR) , where a mixed room simulator is used to combine clean audio with a large library of noise audio, is employed to augment the input data. We take this as the baseline when studying the performance of SpecAugment.

As summarized in table 1, we compare the performance of the network when trained on clean data, data with MTR applied, data with SpecAugment applied, data with both SpecAugment and MTR applied, and data obtained by mixing SpecAugmented and MTR data. We find that SpecAugment, when applied to clean data, performs better than the baseline on all natural test sets, while it performs worse only on a synthetic test set obtained by applying MTR to test utterances. To our surprise, applying SpecAugment on top of MTR degrades performance across most domains. Meanwhile, we are able to achieve improvement across all domains by mixing SpecAugmented data with MTR data.

SpecAugment requires a negligible amount of additional computational resources, does not require additional audio data, can be applied online and is thus highly scalable as the training set becomes large. Our results therefore suggest that SpecAugment can be considered as a serious alternative to more sophisticated resource-heavy augmentation methods.

SpecAugment policies consist of frequency masking, time masking and time warping. The augmentation policies considered in have a fixed number of time masks regardless the length of the utterance. On large scale tasks spanning multiple domains, we expect the length of the utterances to have a large variance. We thus introduce adaptive time masking, where the number of time masks and/or the size of the time mask vary depending on the length of the input. We experiment with several adaptive policies on the Google Multidomain Dataset and LibriSpeech 960h . So far, we have not found adaptive time policies that perform better than vanilla SpecAugment on the Google Multidomain Dataset. Meanwhile, we find adaptive policies that yield performance gains on LibriSpeech relative to , as we are able to train a Listen, Attend and Spell network to have 2.2%2.2\% WER on test-clean and 5.2%5.2\% WER on test-other.

There is a vast literature on augmentation in ASR, only a part of which we survey here. Artificial data augmentation for low resource speech recognition tasks has been studied in . Vocal Tract Length Perturbation has been introduced in the context of data augmentation for ASR in , and explored further in . Noisy audio signals have been used for augmentation in . Speed perturbation has been an integral part of augmentation of speech data. The work studies the effect of using acoustic room simulators. The works examine application of data augmentation for keyword spotting. Drop-out for features have been used for training multi-stream ASR systems in . Systematic omission of frequency channels of the input spectrogram has been studied in the context of CNN ASR networks in . We have commented on SpecAugment in the introduction.

Data augmentation has also been successfully applied to large scale industrial datasets. As noted earlier, Multistyle TRaining (MTR) is a popular technique where clean audio is combined with background noise using a room simulator . MTR has been successfully applied to HMM-based systems and end-to-end LAS models . A natural question is how SpecAugment compares to or can complement existing data augmentation techniques like MTR, especially on large scale datasets.

Our contribution in this paper is three-fold:

We scale up SpecAugment to large scale industrial datasets. We compare to existing MTR data augmentation, and present how we can improve upon it.

We demonstrate that SpecAugment improves the performance of streaming models.

We present an adaptive version of SpecAugment, where the degree of time masking is adaptive to the input sequence length.

SpecAugment and Adaptive Masking

We briefly review SpecAugment in this section, and introduce its adaptive variants. A SpecAugment policy is obtained by composing three basic augmentations—time warping, frequency masking and time masking. We denote the time and frequency dimensions of the spectrogram as τ\tau and ν\nu.

Time warping with parameter WW: A displacement ww is chosen from a uniform distribution from −W-W to WW. A start point w0w_{0} is chosen from the time interval [W,τ−W)[W,\tau-W). A linear warping function W(t)\mathcal{W}(t) is defined so that the start point w0w_{0} is mapped to the point w0+ww_{0}+w and that the boundary points t=0t=0 and t=τ−1t=\tau-1 are fixed:

Warping is defined so that the warped features xwarp(t)\mathbf{x}_{\text{warp}}(t) (in our case, log-mel frequency coefficients) at time tt are related to the original features xorig(t)\mathbf{x}_{\text{orig}}(t) by

We note that the original implementation of time warping presented in , for all practical purposes, is equivalent to this alternative definition.

Frequency masking with parameter FF: A mask size ff is chosen from a uniform distribution from 0 to FF. The consecutive log-mel frequency channels [f0,f0+f)[f_{0},f_{0}+f) are then masked, where f0f_{0} is chosen from [0,ν−f)[0,\nu-f).

Time masking with parameter TT: A mask size tt is chosen from a uniform distribution from 0 to TT. The consecutive time steps [t0,t0+t)[t_{0},t_{0}+t) are masked, where t0t_{0} is chosen from [0,τ−t)[0,\tau-t).

The SpecAugment policies in consist of applying these three augmentations a fixed number of times.

In large scale datasets that contain disparate domains of inputs, we expect there to be a large variance in the length of the input audio. Thus, a fixed number of time masks may not be adequate for such tasks, as the time masking may be too weak for longer utterances, or too severe for shorter ones. We thus introduce two different ways time masking can be made adaptive with respect to length of the spectrogram τ\tau:

Adaptive multiplicity: The number, or multiplicity, of time masks Mt-maskM_{\text{t-mask}} is set to be Mt-mask=⌊pM⋅τ⌋M_{\text{t-mask}}=\lfloor p_{M}\cdot\tau\rfloor for the multiplicity ratio pMp_{M}.

Adaptive size: The time mask parameter is set to be T=⌊pS⋅τ⌋T=\lfloor p_{S}\cdot\tau\rfloor for the size ratio pSp_{S}.

In this paper, we cap the number of time masks at 20 when using adaptive time masking, so that Mt-maskM_{\text{t-mask}} is given by

Experiments

Our set-up for LibriSpeech 960h is based on that of . We use the model LAS-6-1280 of that work and train with training schedule “L”(ong). We use shallow fusion with an LSTM language model (LM) with two fusion parameters—the LM weight and coverage penalty . In this work, we use a 3-layer LSTM with width 4096, with a resulting word-level perplexity of 63.6 on the dev-set transcripts. We tune the fusion parameters on the dev-set using grid-search and apply them to the test set to report the final results.

1.2 Adaptive SpecAugment Policies

We compare three augmentation policies. The baseline policy is the policy coined “LibriSpeech Double” in . This policy has two frequency masks with F=27F=27, two time masks with T=100T=100 which are applied after time warping with W=80W=80.

Let us introduce a hand-crafted adaptive policy, which we denote LibriFullAdapt. This policy has two frequency mask applications with F=27F=27 and time masks with both adaptive multiplicity and size with pM=0.04p_{M}=0.04 and pS=0.04p_{S}=0.04 applied on top of time warping applied with W=80W=80.

1.3 Results

We list the results of our training in table 2. We find that the adaptive policy performs better than the fixed policy, and observe gain in performance both before and after shallow fusion with the language model.

2 Google Multidomain Dataset

We study the effect of SpecAugment when training on the Google Multidomain Dataset . We consider five test sets—Search, Search-Noisy, TTS-Audiobook, Telephony and YouTube—to measure the performance of the network. All training and testing data is anonymized.

As a baseline for our experiments, we augment the input data by using a room simulator described in . For training, various factors of the room simulator, including room-size, reverberation time, microphone positions, speech and noise sources, signal to noise ratio are randomly selected and applied to all input utterances. The injected noise is sampled from either anonymized YouTube audio or a collection of real-life noises. The test set Search-Noisy is constructed by applying these perturbations to the Search test set.

The network input is a log-mel frequency spectrogram obtained from the audio using 32 msec frame windows with 10 msec shift. The log-mel frequency coefficients have 128 dimensions, and are stacked with height 512 with stride 3. The text is tokenized using a Word Piece Model (WPM) of vocabulary size 4k.

We consider five different input configurations: MTR data, clean data, MTR data with SpecAugment applied, clean data with SpecAugment applied and finally data obtained by mixing clean data with SpecAugment applied and MTR data with an 8:2 ratio. Augmentation is applied to the spectrogram after unstacking the features to obtain an array of 128 dimensional features. The augmented spectrogram is then restacked to the original form and fed into the acoustic model.

We present the result of training with a vanilla SpecAugment policy, which we denote SpecAugBasic. This policy has two frequency masks and two time masks with T=50T=50. Time warping has not been used. As a control experiment, we also train the network on data augmented only using frequency masking with two masks of F=27F=27.

2.2 SpecAugmemt on RNN Transducer (RNN-T)

We train an RNN-T model described in . The encoder is an 8-layer uni-directional LSTM with cell size 2048, while the decoder is a 2-layer LSTM with the same cell size. No language model is used.

We note that this model produces weaker context information due to its streaming nature. We nevertheless get gains from time masking, as we demonstrate shortly.

As explained in , our RNN-T model heavily relies on layer normalization . Note that the application of time masks make the variance of hidden activations vanish, which destabilizes training in the presence of layer normalization. Even when using an aggressive variance floor, this still leads to huge gradients when the network becomes deeper. To alleviate this instability, we add Gaussian noise to the time masked regions, which stabilizes training.

2.3 Results

The results of training the acoustic model using the different augmentation methods are presented in table 3. Note that when SpecAugment is applied on top of MTR, the performance degrades below the baseline across all test sets.

Meanwhile, we find that when SpecAugBasic is applied to the clean utterances, it out-performs the baseline across all “natural test sets,” while it performs worse on the synthetic test set obtained by applying MTR to Search-domain utterances. This degradation, however, can be addressed by ensembling SpecAugmented data with MTR data, as shown in the last row of the table.

We note that while we have experimented with adaptive time masking policies, we have not discovered one that out-performs fixed policy SpecAugBasic. The benefit of adaptive time masking on this dataset has yet to be seen.

We emphasize that the trained model is a streaming model, whose performance SpecAugment is still able to noticeably improve. Furthermore, we see that time masking plays an important role in improving the performance of this network, which is evident from the evaluation results on the YouTube dataset.

Summary and Discussion

We find that SpecAugment, despite its simplicity, yields better gains on large scale datasets compared to time-tested and more sophisticated augmentation methods. Given the computational advantage that SpecAugment has, we find it has rich potential for being incorporated into the data pipeline of industrial-scale tasks.

We have introduced adaptive time-masking for SpecAugment. While we have not been able to find an adaptive policy that out-performs a non-adaptive policy on the Google Multidomain Dataset, we have demonstrated the effectiveness of adaptive masking on LibriSpeech 960h. We expect further exploration of adaptive masking to bring improvements when SpecAugment is applied to large scale tasks.

Acknowledgement

We thank Yuan Cao, Ekin Dogus Cubuk, Yanping Huang, Luke Metz, Arun Narayanan, Ruoming Pang, Tara Sainath, Qizhe Xie and Barret Zoph for useful discussions and helping with our experiments.

References