Deep Convolutional Neural Networks and Data Augmentation for Environmental Sound Classification

Justin Salamon, Juan Pablo Bello

I Introduction

The problem of automatic environmental sound classification has received increasing attention from the research community in recent years. Its applications range from context aware computing and surveillance to noise mitigation enabled by smart acoustic sensor networks .

To date, a variety of signal processing and machine learning techniques have been applied to the problem, including matrix factorization , dictionary learning , wavelet filterbanks and most recently deep neural networks . See for further reviews of existing approaches. In particular, deep convolutional neural networks (CNN) are, in principle, very well suited to the problem of environmental sound classification: first, they are capable of capturing energy modulation patterns across time and frequency when applied to spectrogram-like inputs, which has been shown to be an important trait for distinguishing between different, often noise-like, sounds such as engines and jackhammers . Second, by using convolutional kernels (filters) with a small receptive field, the network should, in principle, be able to successfully learn and later identify spectro-temporal patterns that are representative of different sound classes even if part of the sound is masked (in time/frequency) by other sources (noise), which is where traditional audio features such as Mel-Frequency Cepstral Coefficients (MFCC) fail . Yet the application of CNNs to environmental sound classification has been limited to date. For instance, the CNN proposed in obtained comparable results to those yielded by a dictionary learning approach (which can be considered an instance of “shallow” feature learning), but did not improve upon it.

Deep neural networks, which have a high model capacity, are particularly dependent on the availability of large quantities of training data in order to learn a non-linear function from input to output that generalizes well and yields high classification accuracy on unseen data. A possible explanation for the limited exploration of CNNs and the difficulty to improve on simpler models is the relative scarcity of labeled data for environmental sound classification. While several new datasets have been released in recent years (e.g., ), they are still considerably smaller than the datasets available for research on, for example, image classification .

An elegant solution to this problem is data augmentation, that is, the application of one or more deformations to a collection of annotated training samples which result in new, additional training data . A key concept of data augmentation is that the deformations applied to the labeled data do not change the semantic meaning of the labels. Taking an example from computer vision, a rotated, translated, mirrored or scaled image of a car would still be a coherent image of a car, and thus it is possible to apply these deformations to produce additional training data while maintaining the semantic validity of the label. By training the network on the additional deformed data, the hope is that the network becomes invariant to these deformations and generalizes better to unseen data. Semantics-preserving deformations have also been proposed for the audio domain, and have been shown to increase model accuracy for music classification tasks . However, in the case of environmental sound classification the application of data augmentation has been relatively limited (e.g., ), with the author of (which used random combinations oftime shifting, pitch shifting and time stretching for data augmentation) reporting that “simple augmentation techniques proved to be unsatisfactory for the UrbanSound8K dataset given the considerable increase in training time they generated and negligible impact on model accuracy”.

In this paper we present a deep convolutional neural network architecture with localized (small) kernels for environmental sound classification. Furthermore, we propose the use of data augmentation to overcome the problem of data scarcity and explore different types of audio deformations and their influence on the model’s performance. We show that the proposed CNN architecture, in combination with audio data augmentation, yields state-of-the-art performance for environmental sound classification.

II Method

Given our input XX, the network is trained to learn the parameters Θ\Theta of a composite nonlinear function F(⋅∣Θ)\mathcal{F}(\cdot|\Theta) which maps XX to the output (prediction) ZZ:

The proposed CNN architecture is parameterized as follows:

II-B Data Augmentation

We experiment with 4 different audio data augmentations (deformations), resulting in 5 augmentation sets, as detailed below. Each deformation is applied directly to the audio signal prior to converting it into the input representation used to train the network (log-mel-spectrogram). Note that for each augmentation it is important that we choose the deformation parameters such that the semantic validity of the label is maintained. The deformations and resulting augmentation sets are described below:

Time Stretching (TS): slow down or speed up the audio sample (while keeping the pitch unchanged). Each sample was time stretched by 4 factors: {0.81,0.93,1.07,1.23}\{0.81,0.93,1.07,1.23\}.

Pitch Shifting (PS1): raise or lower the pitch of the audio sample (while keeping the duration unchanged). Each sample was pitch shifted by 4 values (in semitones): {−2,−1,1,2}\{-2,-1,1,2\}.

Pitch Shifting (PS2): since our initial experiments indicated that pitch shifting was a particularly beneficial augmentation, we decided to create a second augmentation set. This time each sample was pitch shifted by 4 larger values (in semitones): {−3.5,−2.5,2.5,3.5}\{-3.5,-2.5,2.5,3.5\}.

Dynamic Range Compression (DRC): compress the dynamic range of the sample using 4 parameterizations, 3 taken from the Dolby E standard and 1 (radio) from the icecast online radio streaming server : {music standard, film standard, speech, radio}.

Background Noise (BG): mix the sample with another recording containing background sounds from different types of acoustic scenes. Each sample was mixed with 4 acoustic scenes: {street-workers, street-traffic, street-people, park}We ensured these scenes did not contain any of the target sound classes.. Each mix zz was generated using z=(1−w)⋅x+w⋅yz=(1-w)\cdot{}x+w\cdot{}y where xx is the audio signal of the original sample, yy is the signal of the background scene, and ww is a weighting parameter that was chosen randomly for each mix from a uniform distribution in the range [0.1,0.5][0.1,0.5].

The augmentations were applied using the MUDA library , to which the reader is referred for further details about the implementation of each deformation. MUDA takes an audio file and corresponding annotation file in JAMS format , and outputs the deformed audio together with an enhanced JAMS file containing all the parameters used for the deformation. We have ported the original annotations provided with the dataset used for evaluation in this study (see below) into JAMS files and made them available online along with the post-deformation JAMS files.https://github.com/justinsalamon/UrbanSound8K-JAMS

II-C Evaluation

To evaluate the proposed CNN architecture and the influence of the different augmentation sets we use the UrbanSound8K dataset . The dataset is comprised of 8732 sound clips of up to 4 s in duration taken from field recordings. The clips span 10 environmental sound classes: air conditioner, car horn, children playing, dog bark, drilling, engine idling, gun shot, jackhammer, siren and street music. By using this dataset we can compare the results of this study to previously published approaches that were evaluated on the same data, including the dictionary learning approach proposed in (spherical k-means, henceforth SKM) and the CNN proposed in (PiczakCNN) which has a different architecture to ours and did not employ augmentation during training. PiczakCNN has 2 convolutional layers followed by 3 dense layers, the filters of the first layer are “tall” and span almost the entire frequency dimension of the input, and the network operates on 2 input channels: log mel-spectra and their deltas.

The proposed approach and those used for comparison in this study are evaluated in terms of classification accuracy. The dataset comes sorted into 10 stratified folds, and all models were evaluated using 10-fold cross validation, where we report the results as a box plot generated from the accuracy scores of the 10 folds. For training the proposed CNN architecture we use 1 of the 9 training folds in each split as a validation set for identifying the training epoch that yields the best model parameters when training with the remaining 8 folds.

III Results

The classification accuracy of the proposed CNN model (SB-CNN) is presented in Figure 1. To the left of the dashed line we present the performance of the proposed model on the original datast without augmentation. For comparison, we also provide the accuracy obtained on the same dataset by the dictionary learning approach proposed in (SKM, using the best parameterization identified by the authors in that study) and the CNN proposed by Piczak (PiczakCNN, using the best performing model variant (LP) proposed by the author). To the right of the dashed line we provide the performance of the SKM model and the proposed SB-CNN once again, this time when using the augmented dataset (all augmentations described in Section II-B combined) for training.

We see that the proposed SB-CNN performs comparably to SKM and PiczakCNN when training on the original dataset without augmentation (mean accuracy of 0.74, 0.73 and 0.73 for SKM, PiczakCNN and SB-CNN, respectively). The original dataset is not large/varied enough for the convolutional model to outperform the “shallow” SKM approach. However, once we increase the size/variance in the dataset by means of the proposed augmentations, the performance of the proposed model increases significantly, yielding a mean accuracy of 0.79. The corresponding per-class accuracies (with respect to the list of classes provided in Section II-C) are 0.49, 0.90, 0.83, 0.90, 0.80, 0.80, 0.94, 0.68, 0.85, 0.84. Importantly, we note that while the proposed approach performs comparably to the “shallow” SKM learning approach on the original dataset, it significantly outperforms it (p=0.0003p=0.0003 according to a paired two-sided t-test) using the augmented training set. Furthermore, increasing the capacity of the SKM model (by increasing the dictionary size from k=2000k=2000 to k=4000k=4000) did not yield any further improvement in classification accuracy. This indicates that the superior performance of the proposed SB-CNN is not only due to the augmented training set, but rather thanks to the combination of an augmented training set with the increased capacity and representational power of the deep learning model.

In Figure 2(a) we provide the confusion matrix yielded by the proposed SB-CNN model using the augmented training set, and in Figure 2(b) we provide the difference between the confusion matrices yielded by the proposed model with and without augmentation. From the latter we see that overall the classification accuracy is improved for all classes with augmentation. However, we observe that augmentation can also have a detrimental effect on the confusion between specific pairs of classes. For instance, we note that while the confusion between the air conditioner and drilling classes is reduced with augmentation, the confusion between the air conditioner and the engine idling classes is increased.

To gain further insight into the influence of each augmentation set on the performance of the proposed model for each sound class, in Figure 3 we present the difference in classification accuracy (the delta) when adding each augmentation set compared to using only the original training set, broken down by sound class. At the bottom of the plot we provide the delta scores for all classes combined. We see that most classes are affected positively by most augmentation types, but there are some clear exceptions. In particular, the air conditioner class is negatively affected by the DRC and BG augmentations. Given that this sound class is characterized by a continuous “hum” sound, often in the background, it makes sense that the addition of background noise that can mask the presence of this class will deteriorate the performance of the model. In general, the pitch augmentations have the greatest positive impact on performance, and are the only augmentation sets that do not have a negative impact on any of the classes. Only half of the classes benefit from applying all augmentations combined more than they would from the application of a subset of augmentations. This suggests that the performance of the model could be improved further by the application of class-conditional augmentation during training – one could use the validation set to identify which augmentations improve the model’s classification accuracy for each class, and then selectively augment the training data accordingly. We intend to explore this idea further in future work.

IV Conclusion

In this article we proposed a deep convolutional neural network architecture which, in combination with a set of audio data augmentations, produces state-of-the-art results for environmental sound classification. We showed that the improved performance stems from the combination of a deep, high-capacity model and an augmented training set: this combination outperformed both the proposed CNN without augmentation and a “shallow” dictionary learning model with augmentation. Finally, we examined the influence of each augmentation on the model’s classification accuracy. We observed that the performance of the model for each sound class is influenced differently by each augmentation set, suggesting that the performance of the model could be improved further by applying class-conditional data augmentation.

Acknowledgment

The authors would like to thank Brian McFee and Eric Humphrey for their valuable feedback, and Karol Piczak for providing details on the results reported in . This work was partially supported by NSF award 1544753.

References