Contrastive Learning of General-Purpose Audio Representations
Aaqib Saeed, David Grangier, Neil Zeghidour
Introduction
Self-supervised pre-training has recently emerged as a successful technique to leverage unlabeled data to learn representations beneficial to supervised problems. This success spans a wide range of tasks and modalities . Among these methods, Discriminative Pre-Training (DPT) is particularly effective. This approach learns a representation from pairs of similar inputs from unlabeled data, exploiting e.g. temporal consistency or data augmentation and trains a model to recognize similar elements among negative distractors. In contrast with generative encoder-decoder approaches , DPT is computationally efficient as it avoids input reconstruction entirely.
Amidst DPT models for audio, used a metric learning approach with a triplet loss to minimize the distance between embeddings of anchor and positive pairs and maximize it among the negatives. The instance generation is achieved through noise injection, shifting along time-frequency dimensions, and extracting samples in temporally close neighborhoods. Along similar lines, proposed a benchmark for comparing speech representations on non-semantic tasks. Through utilizing a triplet loss as an unsupervised objective with a subset of AudioSet for model training, they showed improved performance on several downstream speech classification tasks. Inspired from seminal work in NLP , the work in adopted a similar approach to learn audio representations (i.e. Audio2Vec) along with another “pretext” task of estimating temporal distance between audio segments. The pre-trained models are tested on several downstream tasks, from speaker identification to music recognition. Despite recent progress, most work on learning representations of audio focuses on speech tasks (with the exception of ) and ignores other audio tasks such as acoustic scene detection or animal vocalizations. Moreover, triplet-based objectives heavily rely on the mining of negative samples, and the quality of learned features can vary significantly with the sample generation scheme.
In this work, we propose COLA (COntrastive Learning for Audio), a simple contrastive learning framework to learn general-purpose representations of sounds beyond speech. We build upon recent advances in contrastive learning for computer vision (SimCLR , MoCo ) and reinforcement learning (CURL ). We generate similar pairs by simply sampling segments from the same audio clip, which avoids exploring augmentation strategies entirely unlike SimCLR, MoCo, CURL and others . Our dissimilar pairs simply associate segments from different clips in the same batch, which does not require maintaining a memory bank of distractors as in MoCo. Our approach allows us to consider a large number of negatives for each positive pair in the loss function and bypass the need for a careful choice of negative examples, unlike triplet-based approaches . COLA is also different from CPC as it does not predict future latent representations from past ones.
We demonstrate the effectiveness of COLA over challenging and diverse downstream tasks, including speech, music, acoustic scenes, and animal sounds. After pre-training on the large-scale AudioSet database , we show that a linear classifier trained over a COLA embedding gets close to the performance of a fully-supervised in-domain convolutional network and exceeds it when using fine-tuning. Moreover, our system outperforms previous unsupervised approaches on most downstream tasks. These experiments demonstrate that COLA offers a simple, easy-to-implement method to learn general-purpose audio representations without supervision.
Method
We learn general-purpose audio representations from unlabeled data by pre-training a neural network with a contrastive loss function. Our objective function maximizes an agreement between the latent embedding of segments extracted from the same audio clip while using different audio clips as negative classes, as shown in Figure 1. This objective pre-trains a convolutional feature extractor on unlabeled audio data. After pre-training, we combine our feature extractor with an additional classification layer for solving various audio understanding tasks across several datasets.
Contrastive learning extracts a latent space in which the similarity between an anchor example and a related example should be greater than the similarity between the same anchor and unrelated examples. In our case, an anchor and its corresponding positive are audio segments from the same clip. This contrasts with approaches that generate positives as perturbations of the anchor . For negative examples, we take segments from different audio clips in the current training batch. This strategy allows to consider a large number of negatives and is efficient since batch examples are used both as positives and negatives without additional computation.
Bilinear similarity has been used in the past but is less common than cosine similarity, e.g. SimCLR and MoCo. In Section 3, we perform an ablation study on the choice of similarity measure. Table 3 shows that a bilinear similarity outperforms a simple cosine similarity () on all downstream tasks. In the rest of this paper, we use this method when not stated otherwise.
As an objective function, we rely on multi-class cross entropy applied to similarities, i.e.
where is the positive associated to anchor , while refers to the set of negative distractors. This loss, unlike the triplet loss , leverages multiple distractors at a time.
As mentioned earlier, we train our model with positive segment pairs sampled from the same audio clip. For each pair, we use one segment as the anchor and the other element as the positive. Positive segments are used as negatives for all other anchors in the batch. This strategy is more efficient than keeping a memory bank of negatives since the representation of an example is paired with every anchor in the batch either as a positive or as a negative segment. In particular, we experiment with batch sizes varying from to , as shown in Table 4. A large batch size allows the model to see many negative samples per anchor and helps accuracy on end tasks. It is important to note that we sample segment pairs on-the-fly and reshuffle the data at each training epoch to maximize the diversity of positive and negative pairs seen during training. The sample generation procedure is illustrated in Figure 1.
Experiments
We evaluate our method by pre-training COLA embeddings on a large-scale audio dataset and then transferring it to downstream tasks in the following ways: 1) training a linear classifier on top of a frozen embedding, used as a feature extractor and 2) fine-tuning the entire network on the end-task. Importantly, we assess the performance on several diverse datasets to determine the transferability of learned representations across audio domains and recording conditions.
We pre-train COLA embeddings on the diverse, large-scale Audioset . It contains millions excerpts of seconds audio from YouTube videos that are annotated in a multi-label fashion with over classes. This dataset has been used by for self-supervised pre-training. Since our method is self-supervised, we never use Audioset labels. As described earlier, we randomly sample audio clips to generate examples. Likewise, for the extraction of anchors and positives, segments of audio are selected uniformly at random inside a sequence.
We perform downstream evaluation on a variety of tasks, including both speech and non-speech. To allow for comparison with previous methods, we rely on datasets that have been previously used by . For speaker identification, we use a -hours subset of LibriSpeech (LBS) that contains audio of books read by speakers, as well as the Voxceleb subset used in , with speakers. For keyword spotting, we use Speech Commands (SPC) V and V to recognize and spoken commands (classes) from one second of audio, respectively. For acoustic scene classification, we use TUT Urban Acoustic Scenes (TUT) , consisting of labeled audio segments from different acoustic scenes. For animal vocalizations, we use the Bird Song Detection (BSD) dataset from DCASE Challenge to solve a binary classification problem. For music recognition, we use MUSAN that differentiates audio samples across classes (speech, music and noise), as well as the NSynth dataset of musical notes, labeled with the family of the instrument (11 classes). For language identification, we use the Voxforge dataset to categorize audio clips into six classes based on the spoken language.
2 Model Architecture and Implementation Details
Given an audio input sequence, we extract log-compressed mel-filterbanks with a window size of ms, a hop size of ms, and mel-spaced frequency bins in the range – Hz for frames, corresponding to ms. These features are passed through an encoder based on EfficientNet-B0 , a lightweight and highly scalable convolutional neural network. Even though EfficientNet-B0 has been originally proposed for computer vision, the 2D structure of mel-filterbanks allows using this architecture without any adjustment. We apply a global max-pooling to the last layer of the encoder to get an embedding of size . During pre-training, we pass through the projection head , which contains a fully-connected layer with units followed by a Layer Normalization and a activation. We discard the projection head for the downstream tasks and train a linear classifier on top of the encoder directly. We pre-train all our models with ADAM and a learning rate of , for epochs. We explore the impact of the batch size and report the results in Table 4. We train the downstream classifiers with a batch size of and a learning rate of , on randomly selected ms segments, as for pre-training. However, we evaluate downstream classifiers on entire sequences using the following procedure: we split the sequence into non-overlapping ms segments, pass them through the encoder and linear classifier, and average the predictions.
3 Results
Table 1 reports the accuracy on the downstream datasets. We compare our approach against multiple baselines: a linear classifier trained on a randomly initialized fixed encoder and a fully-supervised model trained directly on downstream datasets which indicates the performance achievable with EfficientNet-B0 on these datasets. First, we evaluate pre-trained COLA embeddings with a linear classifier on top of frozen representations, following the same procedure as . This outperforms drastically the performance of a linear classifier trained on a random embedding ( against on average), showing that the encoder has learned useful representations. This is remarkable as we pre-train a single COLA embedding, which performs well across many tasks. Next, we also use a pre-trained COLA as initialization and fine-tune one model per downstream task. Table 1 shows that on all tasks but language identification, initializing a supervised model with COLA improves the performance over training from scratch ( against on average), which demonstrates the benefits of transferring COLA representations even in a fully supervised setting.
To investigate the role of the similarity measure in the quality of learned representations, we perform an ablation study to compare model pre-training with cosine and bilinear similarity. With the cosine similarity, we use a temperature to normalize the scores before computing the loss. Table 3 reports the results obtained on downstream classifiers using encoders pre-trained with each of the similarity estimation techniques. We observe that the best results are obtained using bilinear similarity in all cases. We also conduct an experiment to measure the impact of pre-training batch size, as larger batch sizes result in more negative samples and facilitate convergence . Table 4 shows that, on average, a batch size as large as provides better representations compared to smaller ones. However, increasing the batch size up to worsens the performance in most cases.
Conclusions
We introduce COLA , a simple, easy-to-implement, self-supervised contrastive algorithm for general-purpose audio representation learning. Our approach achieves remarkable performance improvements over earlier unsupervised methods on a wide variety of challenging downstream tasks in a linear evaluation protocol as well as significantly improves results over supervised baselines through fine-tuning. We believe that the simplicity of our system, combined with its strong transferability across audio tasks, will pose it as a go-to baseline for future work in self-supervised learning for audio.