Towards Learning a Universal Non-Semantic Representation of Speech
Joel Shor, Aren Jansen, Ronnie Maor, Oran Lang, Omry Tuval, Felix de Chaumont Quitry, Marco Tagliasacchi, Ira Shavitt, Dotan Emanuel, Yinnon Haviv
Introduction
One of the most powerful uses of deep learning is finding a good representation for a given domain. Despite progress on representations in the visual domain and the language domain , no such universal representation exists for the speech domain. One reason is a lack of standard benchmark tasks to compare different methods; for example, existing speech representations tend to focus on one problem set such as speaker recognition or speech emotion recognition . In this paper, we propose a set of benchmark speech tasks that are diverse, to require that “good” representations contain general speech information, and targeted, to allow good performance when compared with task-specific representations.
We propose a specific set of publicly available tasks, called the “NOn-Semantic Speech benchmark” (NOSS), to assess the general usefulness of speech representations on “non-semantic” tasks. What we call “non-semantic” tasks do not include tasks like automatic speech recognition and phone classification, which require sub-second granularity. They do include paralinguistic tasks such as speech emotion recognition, as well as tasks such as speaker identification, language identification, and medical diagnosis. In addition, we introduce a new set of intra-speaker sub-tasks from existing tasks where a model is trained and evaluated on speech from a single speaker. These intra-speaker tasks help measure which representations are most useful for personalization, which is an increasingly-relevant use-case as more computation is performed on-device.
A good speech representation should be high-performing on a diverse set of downstream tasks using simple models. In addition, it should be useful in transfer learning with small amounts of data for a new task. This use-case is relevant for model personalization, such as user-specific emotion recognition or speaker identification.
We introduce a representation, TRILL (TRIpLet Loss network), which is learned in a self-supervised manner on speech containing clips from AudioSet . Using the techniques of , the network represents audio such that segments which are closer in time are also closer in the embedding space. We demonstrate that this simple proxy objective is highly effective in learning a strong representation for multiple non-semantic speech tasks.
We evaluate TRILL and other representations on our benchmark by training small models built on top of the representations and comparing their performances. In addition, we explore transfer learning by fine-tuning TRILL using data from the downstream tasks. This is an advantage of learned representations over non-learned ones. Pre-training via transfer learning can sometimes outperform models trained on a single dataset , and this is also the case in our benchmark. Using transfer learning, we are able to achieve a new state-of-the-art in many of the tasks, surpassing previously published results which sometimes were hand-crafted for those specific datasets.
We define a new benchmark for comparing representations on non-semantic speech tasks using previously published data. In addition, we add a sub-category of personalization tasks.
We demonstrate that a single representation learned in an unsupervised manners performs best on this benchmark. We compare it to existing representations, feature-based and learned.
We fine-tune our best performing representation, further boosting results. This method sets a new state-of-the-art on many previously published tasks.
We distill our learned representation to a model that can run inference and training on-device, and we open-source the original and distilled models.
Background
Transfer learning and domain adaptation have been extensively studied in machine learning . Recent research has mostly focused on deep representation learning methods, either supervised, semi-supervised, or unsupervised. Successful representations improve the sample efficiency of ML algorithms by extracting most information out of the raw signal from the new data before any task-specific learning takes place. This strategy has been used successfully in many application domains .
An important step in learning a good representation is having a standard benchmark to evaluate it. Such benchmarks should contain a variety of downstream tasks, representing different tasks in the domain. Such benchmarks have been developed in vision and NLP .
There are three standard approaches to adapting a representation to multiple, potentially heterogeneous, downstream tasks. One approach is to train a task-specific linear classifier on the embeddings produced by a pre-trained network, whose parameters are kept frozen . A second approach is to fully fine-tune. Generally, fine-tuning matches or outperforms the performance of fully-supervised models trained on the downstream tasks , especially when the amount of labeled data is small. A third approach is multi-task learning. This has been applied in the speech domain , although not on a wide range of tasks. It is usually favored when the downstream tasks are all applied on the same input set.
There are many methods for learning audio representations. The work in trained an embedding for audio classification on AudioSet. Other work has demonstrated the value of supervised , semi-supervised , or unsupervised representations. The unsupervised audio representation literature is especially diverse (L3 , AuDeep , Autoregressive Predictive Coding , Contrastive Predictive Coding , metric learning , autoencoding ). However, these methods were evaluated on just one or a limited set of downstream tasks.
In other domains, training a strong representation requires a very large, general dataset. AudioSet is the largest dataset for general purpose audio machine learning, serving as an audio equivalent of ImageNet. Even when restricted to only the samples with speech tags, it surpasses all datasets in size and variability. It has been used to learn a general purpose audio embedding in , and can be used for multiple speech tasks.
Non-Semantic Speech Benchmark (NOSS)
To standardize the assessment of non-semantic speech representations, we introduce NOSS. This section describes the benchmark in detail (summarized in Table 1). These tasks reflect different properties of the speech signal, and they vary in size and difficulty. Personalization and on-device training is increasingly important, so we include “Intra Speaker” tasks for the datasets when applicable. Intra-speaker tasks are an important addition because they also test task adaptation for small amounts of data, and that representations do not just rely on speaker identity.
Inter-speaker tasks are described in Table 1. They were chosen so that they span a wide variety of non-semantic tasks, including “paralinguistic” tasks such as speech emotion recognition, as well as other kinds of tasks such as speaker identification, language identification, and one task of voice-based medical diagnosis. Each dataset has a different number of target classes, and different number of examples.
An important use-case of task adaptation is personalization—training and evaluating a model on data only of a specific person. We call these tasks intra-speaker tasks. The accuracy is averaged over all speakers. Note that most inter-speaker tasks divide the speakers into disjoint groups in the train/test split, and intra-speaker tasks use all speakers for both training and testing. Not all tasks have a meaningful intra-speaker versions; it is meaningless to train and test on the same speaker in tasks with labels that depend only on the speaker’s identity, such as language identification, medical diagnosis, and speaker identification. The intra-speaker tasks are: CREMA-D, SAVEE, and Speech Commands.
Experiments
Non-semantic aspects of the speech signal (e.g., speaker identity, language, and emotional state) generally change more slowly than the phonetic and lexical aspects used to explicitly convey meaning. Therefore, we expect a good representation for non-semantic downstream tasks to be considerably more stable in time than what is required for ASR applications. To take advantage of this intuition, we follow the work of (Section 3.2.5) and use temporal proximity as a self-supervision signal.
where is the norm, is standard hinge loss, and is a nonnegative margin hyperparameter. We use the now-standard within-batch semi-hard negative mining technique .
The TRILL model is trained on the subset of AudioSet training set clips possessing the speech label. We set to 10 seconds, the maximum duration of each AudioSet clip. This makes the training task is primarily same clip / different clip discrimination. Following , we (i) take as input log mel spectrogram context windows with mel bands and frames representing 0.96 s of input audio (STFT computed with 25 ms windows with step 10 ms); and (ii) employ the Resnetish variant of the standard ResNet-50 architecture followed by a dimensional embedding layer. Since the ResNet’s final average pooling operation destroys the sub-second temporal structure, we also consider representations defined by earlier convolutional blocks.
2 Other representations
We compared TRILL to several representations used frequently on non-semantic tasks. These include classical methods like Mel spectrograms and OpenSmile , which is the de-facto standard in emotion recognition. For Mel spectrograms, we tested many different configurations and selected the best one on each task, and for OpenSmile we used the ComParE16 acoustic parameter set .
We also compared to existing learned representations. YAMNet is a state-of-the-art network for audio classification , trained on AudioSet. VGGish is an audio embedding produced by training a modified VGGNet model to predict video-level tags from the Youtube-8M dataset. We also compared to our own network with random initialization, a technique which has been shown to produce good embeddings .
3 Experimental Method
In order to evaluate the usefulness of the representations described in Section 4.2, we train small models to solve the downstream NOSS tasks (Section 3). For each representation / task pair, we explore different downstream models, representation aggregation techniques, and normalization methods.
We train shallow models using Scikit-Learn . We experimented with logistic regression, random forest classifiers, and linear discriminator analysis (LDA) as the downstream models. Despite the shallowness of the above models, we achieve competitive results with the best previously-reported results on some of the benchmark tasks.
The embeddings were aggregated over time using average pooling, except for VoxCeleb which used NetVLAD aggregation . Datasets which have multiple utterances per speaker were normalized for each speaker using normalization.
Some of the downstream tasks have fixed canonical splits. For the inter-speaker tasks on those datasets, we report numbers on the canonical splits (Speech Commands and VoxCeleb). For the other datasets, and for all the intra-speaker tasks, we perform five random train / test splits and report the average. For intra-speaker tasks (Section 3), we train and test on one speaker at a time, then average results across splits and across speakers.
3.2 Network layer and model distillation
For the representations generated from pre-trained neural networks (TRILL, VGGish, YAMNet), we experimented with both the final output and two intermediate representations. For TRILL, we tried the final 512-dimensional embedding layer and the pre-ReLU output of the first 512-depth convolutional layer (subsequently referred to as layer 19 due to TensorFlow convention). We found that layer 19 performed best on our tasks. For VGGish, we tried the final layer, and the first fully-connected layer. For YAMNet, we test the final pre-logit layer, and the fifth depth-separable convolutional layer outputs.
To make the network more useful, we used distilliation to train a smaller model of similar quality. We used a truncated MobileNet architecture to predict the original TRILL network’s layer 19 embeddings. This distillation reduces model parameters and multiplies by factors of 5.6X (9M 1.6M) and 25X (1.5B 59M), respectively.
3.3 Fine-tuning
Shallow models trained on top of frozen embeddings sometimes do not have enough degrees of freedom to adapt to mismatches between the train and target distributions, so we also experimented with fine-tuning the entire model end-to-end. For benchmarks tasks with relatively small amounts of data, we applied early stopping.
We also experimented with intra-speaker fine-tuning. For this we used the CREMA-D dataset, with fixed train/dev/test splits of 80%/10%/10%. We first fine-tuned a common model on the training partitions of all the speakers, then further fine-tuned on each speaker separately.
Results
TRILL outperforms previously reported results on three of the six benchmark tasks, and is competitive with previous best on two of the three remaining tasks. On these two tasks, the best previously reported numbers use other modalities in addition to audio (visual or textual features) (Table 2). Of the representations we compared against, TRILL performed the best on five-of-six tasks (Table 2) and two-of-three intra-speaker tasks (Table 3). We successfully distilled TRILL to a much smaller model that can be used on a mobile device. The distilled model has no performance degradation on five-of-nine tasks, statistically insignificant degradation on one, and minor degradation on the remaining tasks. We compare across all tasks by fitting a linear regression on the observed accuracies, with both the model and task as the explanatory variables (see Figure 1).
The distilled model results are presented in the one-before-last line of Table 2. With the exception of VoxForge language identification and SAVEE speech emotion recognition, the reduction in model capacity and dimensionality has performance within the standard deviation of the larger embedding (variance calculated over 5-fold data splits). In the personalized tasks the quality of the distilled model is the same as the larger model.
Analysis
As can be seen in Table 2, fine-tuning the final embedding gives a clear boost on most tasks and sets a new state-of-the-art in 3 out of the 6 datasets. This approach of learning a strong representation on a large benchmark and then fine-tuning it for the downstream task has been proven very effective in other domains , and in this paper we demonstrate the same is true in the speech domain.
An important observation of our research is that the effective representation learned by both YAMNet and TRILL is not at their final layer, but in fact in one of their intermediate layers. This intermediate layer must capture information which is later discarded in the final layers. The reasoning might be that when learning the triplet loss, the network might learn to discard properties which vary temporally such as tone of voice or semantic information, but this embedding is still learned in the intermediate layer. Another evidence for that can be seen in the results on the Speech Commands dataset, where the performance of the intermediate layers of all of the learned representations is much better than the performance of their respective top layer.
Table 4 shows that fine-tuning per speaker allows to further personalize to each speaker, and generally improves accuracy. When breaking down the performance impact on each speaker, we can see per-speaker fine-tuning improves accuracy for 31 speakers, is mostly unchanged for 49 speakers, and decreases for 12 speakers.
Conclusions
In this work, we explore the importance of clearly defining benchmarks when comparing representations of speech. We propose NOSS (Section 3) to help fairly compare non-semantic speech representations, and we introduce a sub-category of personalization tasks to help measure progress in the age of on-device computation. We also demonstrate that TRILL (Section 4.1), based on a self-supervised training criteria, simultaneously performs well on all benchmark tasks. We show that finetuning TRILL on a small amount of data outperforms or is competitive with almost all previously reported numbers for the NOSS tasks, and that TRILL is significantly better than other representations. Finally, we distill TRILL to be on-device with very little or no performance loss. The NOSS benchmark, the evaluation code, TRILL, and TRILL-distilled are all publicly available.