ESResNet: Environmental Sound Classification Based on Visual Domain Models
Andrey Guzhov, Federico Raue, Jörn Hees, Andreas Dengel
I Introduction
With the increasing popularity of voice assistants, many of which use Deep Learning techniques, the currently most apparent task from the audio domain is probably automatic speech recognition. However, apart from this very prominent example, many other challenges still exist in the audio domain. One of these challenges is Environmental Sound Classification (ESC), which is concerned with correctly differentiating between sound classes that we experience in our everyday environment (e.g., “baby crying”, “car honking”, “children playing”, “dog barking”, “siren”, “snoring”, “street music”). While ESC has many potential application areas, one of the most obvious ones is multimedia retrieval, in which ESC could be used to improve the performance of video retrieval systems by making better use of the audio modality . Another application area is the automated analysis of urban sounds, for example to offer more detailed insights for high noise levels .
While ESC is a comparably young field, a lot of progresses were made after great datasets such as the ESC-50 and UrbanSound8K (US8K) found wide acceptance in the community. However, we observed that the general trend in the ESC community is to design audio-domain-specific architectures and combine them with specially engineered features. On one side, this approach makes more difficult to benefit from advances made in other fields, such as the computer vision community. On the other, this scenario sparked our interest to investigate how well a current state-of-the-art approach from the image domain would perform on ESC. During our investigations, we found, that while our approach immediately out-performed all previous ones on the ESC-50 dataset, it initially seemed to perform quite poorly on US8K, despite the fact that it can actually make use of stereo inputs. However, during our follow-up, we noticed that there are reproducibility problems wrt. prior publications reporting on the US8K dataset: many existing approaches lack necessary details for reproduction and about the used dataset splits. Given a fair comparison, our approach in fact out-performs all prior ESC models also on the US8K.
The remainder of this paper is organized as follows. In Section II we discuss prior models that were used for environmental sound classification. We then describe our proposed approach based on log-power STFT spectrograms and a well-known CNN model in Section III, how it was trained and evaluate in Section IV, before presenting our results in Section V and concluding with a summary and future work in Section VI.
II Related Work
Unlike image-related tasks (image classification, segmentation, object detection, etc.), the environmental sound classification task implies the usage of locally correlated one-dimensional signals, so the input is stretched along a single axis. The most widely known datasets in the field of environmental sound classification are the ESC-50 / -10 and the UrbanSound8K (US8K) , further detailed in Section IV.
The representation of audio is quite different from visual signals (e.g., photo) that have local correlations in both spatial dimensions. Thus, many methods were proposed specifically tailored towards the audio domain. We can divide them into the following major groups. A comprehensive overview and comparison of all methods can be found in the Table I.
The use of a raw signal as an input provides a straightforward solution to build a model that handles any sort of time-frequency transformation internally. The most important property of this class of models is that data pre-processing is not needed. First, and proposed a one-dimensional architecture called EnvNet v1 / 2 that was able to achieve state-of-the-art results at that time. Later, in the concept of 1D-CNNs was extended into a model that operated on an input signal at different time scales. Another way to improve performance of this type of models was proposed in where the use of gammatone filterbanks for the initialization of model allowed to improve results in comparison to the otherwise random weight initialization.
In contrast, for simplicity and to reduce the amount of trainable parameters (considering limited training data), we decided to rely on a fixed time-frequency transform with a wide spectrum range in our model. However, we mention potential improvements in this direction in our future work.
II-B Learnable Filterbanks and 2D-CNN
While one-dimensional CNNs handle all transformations of input signal internally, which makes it possible to apply them in an end-to-end fashion, this approach involves a lack of control over the transformed representation. The use of gammatone initialization in helped to overcome this issue partially. The uniqueness of is that the model was split into two parts, namely Convolutional Restricted Boltzman Machines (ConvRBM) for feature extraction (instead of the fixed procedure that is used in our work) and the CNN proposed in for the actual classification task.
II-C Pre-computed Time-Frequency Representation and 2D-CNN
The first model that set the baseline for the environmental sound classification was the CNN proposed in (Piczak-CNN) that operated on Mel-scaled spectrograms. The use of a fixed feature extraction procedure made it possible to obtain a model’s input that possessed the required characteristics. Further development of single-feature input and research on data augmentation techniques were done in . Follow-up studies involved the extension of input features to others that based mostly on the Short-Time Fourier Transform (STFT) to Mel-Frequency Cepstral Coefficients (MFCC) , ; Cross Recurrence Plot (CRP) , ; Teager’s Energy Operator (TEO) , ; (Phase-Encoded) FilterBank Energies ((PE)FBE) ; gammatone-spectrogram , , ; chromagram , ; spectral contrast , ; and Tonnetz , .
However, all of the aforementioned features were developed with a reduction of computational complexity or compression in mind. With the growth of computational capacity, it seems that we can now make use of a single-feature that covers the full range without any reduction.
The model we propose belongs to this major group, it is also handling single-feature input (log-power spectrograms).
We provide detailed description of features used in the aforementioned studies in the Table I.
II-D Data Augmentation
Data augmentation is a powerful technique that allows to increase variability in the training data and thus acts as a regularizer preventing overfitting. According to and there are the following transformations that augment audio training data:
This method changes the duration of the audio, while keeping its spectral characteristic untouched.
II-D2 Pitch Shift
In opposite to time stretching, this method allows to manipulate spectral characteristics and preserve duration of the track.
II-D3 Time Inversion
Time inversion that was applied in is an effective data augmentation technique that is related to random flip of images during the training on the visual classification datasets.
II-E Comparability of Results on the UrbanSound8K (US8K) Dataset
According to our findings, there are at least five papers (three in the 2019), whose reported results are not directly comparable with others. In particular, as reported by , the authors of used an unofficial split of the US8K dataset. Also, the authors of stated that the results were obtained on a non-standard split, whereas the authors of provide the description of a custom snippets generation strategy. Finally, we determined that results published by and are incomparable with those acquired on the official split of the US8K dataset as well. We provide further details on this in the Table I and Section V-D.
III Model
In this paper, we propose a visual domain convolutional neural network in conjunction with log-power spectrograms to solve the environmental sound classification task. This section describes the architecture of the model and how it is extended by the attention mechanism. We also describe its application to stereo audio, the initialization of the network’s weights and the process of log-power spectrogram computations.
Residual neural networks are characterized by the additional skip connections that bypass some of the layers and merge their input and output. The motivation for this is to prevent gradient vanishing that made it very difficult to design deep neural networks before . In our work, we propose the ESResNet model based on the vanilla ResNet-50 architecture in order to demonstrate its ability to achieve state-of-the-art results on a domain the model was not designed for. The overall structure of the model is presented in Figure 1.
III-B Attention
The attention mechanism was presented initially for the use in conjunction with recurrent models, in particular, in sequence modelling tasks . The main goal of it was to highlight relevant parts of a long sequence and to get rid of irrelevant ones. In the visual domain, one uses attention blocks in order to produce weighting for the input signal. Usually, there are several attention sub-branches consisting of one or many convolutional layers that process feature maps in parallel with the main branch.
For the environmental sound classification task, the main purpose of the attention blocks is to focus the model on the most important information in both the time and frequency domain. To implement the attention mechanism, we extended our ESResNet model (inspired by ), by adding a stack of attention blocks in parallel (ESResNet-Attention, Figure 1). Each block among the first 4 handles either frequency- or time-related information. For instance, the first attention block receives the same input as the first layer , then it processes the signal using frequency-dedicated convolutional filters and provides an output of the same shape as the one provided by the . Finally, the input of the second layer is constructed by the element-wise multiplication of outputs of and blocks (Equation 1).
The last attention block handles a joint time-frequency representation. The core of the attention block is a depth-wise separable convolution . Output of each attention block is given by the logistic function.
III-C Spectrogram
A spectrogram is an image-like representation of the spectrum of frequencies varying with time. In relation to digital signal processing, there are several ways to obtain a spectrogram. It can be generated using filterbanks, Fourier (or more generally wavelet) transform, etc. In our work, we compute log-power spectrogram from the STFT of an audio signal (Equation 2).
STFT belongs to the family of Fourier-related transforms and is used to determine magnitude and phase of basis sinusoidal frequencies at different time points in a time-domain signal .
In practice, to compute Equation 3, one splits input signal into overlapping frames multiplied by window function , then the Fast Fourier Transform (FFT) is being applied to each frame separately.
III-C2 Window Function
In order to reduce spectrum perturbances caused by the framing, a window function is applied. The use of windowing reduces the amount of noise in the spectrum and therefore improves the signal-to-noise ratio. The drawback of the usage of a window function is so-called spectral leakage. Spectral leakage is a common name for the non-zero values produced by Fourier transform at frequencies other than fundamental. The choice of window function is a trade-off between many characteristics. In our work, we decided to choose the minimum 4-term Blackman-Harris window which is given by Equation 4 as it provides reasonable bandwidth and very low spectral leakage making it a good choice as a general-purpose window :
III-D Input Channel Transformation
For image classification models, such as ours, the usual way to represent input data is an RGB model with 3 input channels (red, green and blue). However, our spectrograms only provide input in form of a single-channel (grayscale values). One way to tackle this issue is to replicate the spectrogram to other channels or to pass zeros instead. The major drawback of this solution is either unnecessary redundancy or the loss of information, and increased computational cost.
III-E Handling Stereo Audio Using Siamese-like Architecture
The way how humans perceive audible information is inherently stereo. In this work, we exploit advantages brought by additional audio channels and show on the US8K dataset that a minor architectural tweak helps us to out-perform state-of-the-art results.
Siamese neural networks were developed to produce a similarity measure for two input samples . In this work, we however use the common broader notation to call any network Siamese that applies the same set of weights to two different inputs and thus produces two comparable vectors (or embeddings). As Figure 3 illustrates, we take a two-channel audio input followed by log-power spectrogram computation (via STFT) and pass each channel separately through the layers. After we obtain the network’s outputs, we fuse them by an element-wise addition and pass the resulting embedding through the last fully-connected layer that performs the final classification.
III-F ImageNet Training as Weight Initializer
The ESC datasets contain a limited number of samples. This setup becomes especially important in case of the ESC-50 dataset as it provides the challenging task to distinguish between 50 classes using only 1600 training samples .
To leverage the full power of deep neural networks, the amount of data should grow exponentially with the amount of parameters. If the number of training samples is restricted, one way is to perform fine-tuning. In this work, we decided to employ a model that was trained from scratch on the ImageNet dataset . The ImageNet dataset provides more than 1 million training samples divided into 1000 classes. As we will see in Section V, the initialization of weights based on a pre-training on the ImageNet image classification task is beneficial for the environmental sound classification.
IV Experimental Setup
In this section we describe the setup of our experiments, starting with the datasets, their pre-processing and how our model was trained. We also describe our reproduction of previous results / re-implementation of their approaches for comparison.
IV-A2 UrbanSound8K
We would like to explicitly highlight the importance of using the officially provided folds by describing the way training samples were acquired by authors of . As at the time of collection the number of qualitatively labeled recordings provided by the Freesound project was restricted , each track was split into snippets that had an overlap of 50 % . Let us consider two tracks A and B belonging to the same class (Figure 4). Applying a sliding window and moving it with the overlap of 50 %, we obtain snippets called A1–4 and B1–4, respectively. Two subsequent snippets share a part of the original track, so one has to make sure that they are presented either in training or evaluation set (the official split) and not in both (as happens by random shuffling, which is underlying many unofficial splits).
IV-A3 Data Pre-processing
IV-B Model Training
During the training phase, the following augmentations were applied (see Section II-D): random time inversion and time scaling . The later can be considered as a combination of time stretching and pitch shift. The main advantage of such combined transformation is its computational cheapness in comparison to aforementioned ones. For instance, pitch shift implies forward and inverse STFT which makes it inefficient to apply this transformation on-the-fly during the training. The probability of the time inversion was set to . The scaling factor was sampled uniformly from the continuous range .
IV-C Re-implementation
As the results reported by and especially were very high, we focused on them to find the key to a such high performance. As the description of models and / or setups did not allow us to determine the crucial component, we decided to reproduce their results. Sadly, the authors of did not publish their code, nor did any of them respond to our email within a month. Hence, we re-implemented the part of the TSCNN-DS model called LMCNet and evaluated it on the official and unofficial random split of the US8K dataset using all available implementation details provided by authors.
The authors of had published parts of their source code including hyper-parameters for their TFNet model (only ESC-50), which allowed us to reproduce their results with minor re-implementations (by extending to US8K). Sadly, after contacting them by email, the original repository disappeared.
For our evaluations on an unofficial random split of the US8K dataset we used a StratifiedKFold as provided by scikit-learn . The number of splits was set to 10, all experiments were conducted with the same random seed.
We report the reproduced results in Table I (emphasized by italic font) and discuss them in Section V-D.
V Results
As can be observed in Table I, our presented approaches out-perform all previous approaches in a fair comparison.
As we discussed in the Section IV-A, the amount of available training samples plays a crucial role for deep learning models. In this work, we compared performance differences between a model that was trained from scratch and one that was pre-trained on the ImageNet dataset and then fine-tuned. The largest relative change can be observed on the ESC-50 dataset (from 81.15 % to 90.80 %, ESResNet) as it presents a challenging problem in conjunction with a restricted number of training samples. We still find strong improvements on the ESC-10 (from 92.50 % to 96.75 %, ESResNet) and US8K dataset (from 79.91 % to 83.59 %, ESResNet).
V-B Stereo vs. Mono
Further, despite the availability of stereo recordings in the US8K, we identified, that the competing previous models only consider single-channel audio. As described, we use a Siamese-like extension to the vanilla input processing of the ResNet-50 network in order to enable our ESResNet architecture to process multi-channel inputs where possible (US8K in Table I). Further, Table II presents a comparison of results of our model achieved on mono and stereo inputs. The results show that between-channel difference provides useful information that allows to out-perform previous state-of-the-art results on the US8K dataset even without the use of additional attention blocks. For instance, the ESResNet model trained from scratch is able to achieve accuracy of 79.91 % on the US8K dataset using mono audio as an input, however the use of stereo input allows to classify 81.31 % of the test samples correctly whereas the extension by attention blocks (ESResNet-Attention) provides a smaller performance gain when operating only on mono input (81.00 %). A similar situation can be observed in the case of our ESResNet model that was pre-trained on the ImageNet dataset . The use of stereo input for the ESResNet model out-performs (84.90 %) the vanilla model on mono input (83.59 %) as well as the attention-boosted model on mono input (84.21 %).
V-C Attention-boosted vs. Vanilla
Combining a powerful visual model and descriptive time-frequency representation (ESResNet) already allows us to out-perform previous results. However, further improvement is possible by including attention (ESResNet-Attention, Figure 1). The use of the attention blocks allows us to out-perform previous state-of-the-art results on all three datasets (ESC-50 / -10 and US8K) achieving 91.50 %, 97.00 % and 84.21 %, respectively. Additionally, the combination of stereo input and attention blocks provides further improvement of the achieved accuracy on the US8K, allowing the ESResNet-Attention model to achieve a new highest state-of-the-art accuracy of 85.42 %.
V-D Official and Unofficial Splits and Reproducibility Problems
As stated in the Section IV-C, we reproduced the approaches presented in and .
The performance achieved by the re-implemented LMCNet model on the US8K (Table I) allows us to attribute the results of to those that did not perform evaluation on the official split.
For TFNet, we re-ran the temporarily available code for the ESC-50 dataset (without data-augmentation). We then slightly adapted the code to also run it on the US8K. In both cases, we surprisingly reached significantly lower accuracies than stated by the authors . However, when running their code on a completely random unofficial US8K split, we achieved significantly higher results than previously reported. We conclude from this, that either, the shared code lacks crucial steps for reproducibility of the reported results, or that the authors neither ran their experiments on the official nor a completely random unofficial US8K split.
In order to roughly quantify the influence of an unofficial (random) splitting strategy on our results, we also report them in Table I. To point out, that these very high numbers do not constitute a basis for fair comparison, we put them in parenthesis.
VI Conclusion
In this work we demonstrated how a well-known visual domain model could successfully be applied to Enviromental Sound Classification. Being applied in conjunction with regular log-power spectrograms, our ESResNet model is able to perform competitive to humans (ESC-50), whereas pre-training on the ImageNet dataset already allows us to out-perform all current state-of-the-art methods. We also showed that the presence of multiple channels in the input signal gives an additional performance gain on the UrbanSound8K dataset with only minor architectural changes (Siamese-like processing). Further improvement is possible with the help of attention blocks supporting the network in focusing on the relevant parts of its input in time and frequency domain (ESResNet-Attention). Such a configuration reached the highest accuracy and out-performed all previous state-of-the-art models significantly in a fair comparison on the ESC-50 and UrbanSound8K datasets.
Finally, we highlighted the importance of the strict adherence to the evaluation procedure, demonstrated the influence of a random splitting strategy on evaluation results on the UrbanSound8K dataset and differentiated previously reported results into official and unofficial splits. For reproducibility, we provide all code, also including our re-implementations of models that were sadly published without code before.
In the future, we would like to investigate learning time-frequency representations instead of using the current fixed feature extraction. Also as we have seen that ImageNet helps in the initialization of our model, we would like to investigate which classes benefit more and which less from domain transfer.
Acknowledgments
This work was supported by the TU Kaiserslautern CS PhD scholarship program, the BMBF project DeFuseNN (Grant 01IW17002) and the NVIDIA AI Lab (NVAIL) program. Further, we thank all members of the Deep Learning Competence Center at the DFKI for their comments and support.