Deep Convolutional Neural Networks and Data Augmentation for Acoustic Event Detection
Naoya Takahashi, Michael Gygli, Beat Pfister, Luc Van Gool
Introduction
Scenes typically contain many sound sources. While speech is arguably one of the most important types, non-speech sounds such as music or laughter provide important information as well. In most conversations no mention is made of the environment, like its location or people and objects present. Automatic speech recognition (ASR) could benefit from having such contextual knowledge though . Knowing the type of non-speech sounds improves the performance of source separation and speech enhancement . Furthermore, multi-media tasks such as video classification and video summarization have been shown to improve when including audio information. Acoustic Event Detection (AED) is attracting more and more attention also due to new applications incl. surveillance , multimedia content retrieval and audio segmentation .
Traditional methods for AED apply techniques from ASR directly. For instance, Mel Frequency Cepstral Coefficients (MFCC) were modeled with Gaussian Mixture Models (GMM) or Support Vector Machines (SVM) . Yet, applying standard ASR approaches leads to inferior performance due to differences between speech and non-speech signals. Thus, more discriminative features were developed. Most were hand-crafted and derived from low-level descriptors such as MFCC , filter banks or time-frequency descriptors . These descriptors are frame-by-frame representations (typically frame length is in the order of ) and are usually modeled by GMMs to deal with the sounds of entire acoustic events that normally last seconds at least. Another common method to aggregate frame level descriptors is the Bag of Audio Words (BoAW) approach, followed by an SVM . These models however discard the temporal order of the frame level features, causing considerable information loss. Moreover, methods based on hand-crafted features optimize the feature extraction process and the classification process separately, rather than learning end-to-end.
Recently Deep Neural Networks (DNNs) have been very successful at many tasks, including ASR , image classification , and visual object detection . One advantage of DNNs is their capability to jointly learn feature representations and appropriate classifiers. Supported by large amounts of training data, more recently, deeper architectures further pushed the state-of-the art for several competitions in computer vision . In comparison, few AED methods rely on DNNs. One reason is the lack of large, publicly available datasets. In , DNNs were built on top of MFCCs. Miquel et al. utilize a Convolutional Neural Network (CNN) to extract features from spectrograms. These networks are still relatively shallow (e.g. 3 layers). Furthermore, the networks take only a few frames as input and the complete acoustic events are modeled by Hidden Markov Models (HMM) or simply by calculating the mean of the network outputs, which is too simple to model complicated acoustic event structures.
In this work, we introduce novel network architectures with up to 9 layers and a large input field. The large input field allows the networks to directly model entire audio events and be trained end-to-end, as depicted in Fig. 1. Our network architecture is inspired by “VGG Net” which obtained second place in the ImageNet 2014 competition and was successfully applied for ASRs . The main idea of VGG Net is to replace large (typically 99) convolutional kernels by a stack of 33 kernels without pooling between these layers. Advantages of this architecture are (1) additional non-linearity hence more expressive power, and (2) a reduced number of parameters (i.e. one 99 convolution layer with channel has weights while three-layer 33 convolution stack has weights). Our first goal is to adapt the VGG Net architecture to AED. In order to train our network we further propose a novel data augmentation method, especially effective for AED. For our experiments, we created a new dataset harvested from the Freesound repository and conducted acoustic event classification. Experimental results show that our deeper CNN significantly outperforms several baseline techniques, including state-of-the-art methods based on BoAW and classical DNNs. We further show that the proposed data augmentation method improves the performance by more than 12%.
Architectural and Training Novelties
We propose two CNN architectures, adapted to AED, as outlined in Table 1. Architecture has 4 convolutional and 3 fully connected layers, while Architecture has 9 weight layers: 6 convolutional and 3 fully connected. In this table, the convolutional layers are described as conv(input feature maps, output feature maps). All convolutional layers have 33 kernels, thus henceforth kernel size is omitted. The convolution stride is fixed to 1. The max-pooling layers are indicated as in Table 1. They have a stride equal to the pool size. All hidden layers except the last fully-connected layer are equipped with the Rectified Linear Unit (ReLU) non-linearity. In contrast to , we do not apply zero padding before convolution since the output size of the last pooling layer is still large enough in our case. The networks were trained by minimizing the cross entropy loss with regularization using back propagation:
where are the th input, label and network parameters, respectively. is a constant which is set to in this work.
2 Large input field
In ASR, few-frames descriptors are typically concatenated and modeled by GMM or DNN . This is reasonable since they aim to model sub-word units like phonemes which typically last less than a few hundreds of . The sequence of sub-word units is typically modeled by a HMM. Most works in AED also follow similar frameworks, where signals lasting from tens to hundreds of are modeled first. These small input field representations are then aggregated to model longer signals by HMM, GMM or a combination of BoAW and SVM . Yet, unlike speech signals, non-speech signals are much more diverse even within a category and it is not clear that a sub-word approach is suitable for AED. Hence, we design a network architecture that directly models the entire acoustic event, based on a single input of multiple seconds. This also enables the networks to optimize the parameter end-to-end.
3 Data Augmentation
Since the proposed CNN architectures have many hidden layers and a large input field, the number of parameters is high, as shown in the last row of Table 1. A large number of training data is vital to train such networks. Jaitly et al. showed that the data augmentation based on Vocal Tract Length Perturbation (VTLP) is effective to improve ASR performance. VTLP attempts to alter the vocal tract length during extraction of descriptors, such as a log filter bank, and perturbs the data in a certain non-linear way. In order to introduce more data variation, we propose a different augmentation technique. For most sounds coming with an acoustic event, mixed sounds from the same class also belong to that class, except when the class is differentiated by the number of sound sources. For example, when mixing two different ocean surf sounds, or of breaking glass, or birds tweeting, the result still belongs to the same class. Considering this property we produce augmented sounds by randomly mixing two sounds of a class, with randomly selected timings. In addition to mixing sounds, we further perturb the sound by moderately modifying frequency characteristics of each source sound by boosting/attenuating a particular frequency band. An augmented data sample is generated from source signals for the same class as the one both and belong to, as follows:
where are uniformly distributed random values, is the maximum delay and is an equalizing function parametrized by . In this work, we used a second order parametric equalizer parametrized by where is the center frequency, is a gain and is a Q-factor. An arbitrary number of such synthetic samples can be obtained by randomly selecting parameters for each data augmentation. We refer to this approach as Equalized Mixture Data Augmentation (EMDA).
4 Multiple Instance Learning
Since we used web data to build our dataset (see Sec. 3.1), the training data is expected to be noisy and to contain outliers. In order to alleviate the negative effects of outliers, we also employed multiple instance learning (MIL) . In MIL, data is organized as bags and within each bag there are a number of instances . Labels are provided only at the bag level, while labels of instances are unknown. A positive bag means that at least one instance in the bag is positive, while a negative bag means that all instances in the bag are negative. We adapted our CNN architecture for MIL as shown in Fig. 2. instances in a bag are fed to a replicated CNN which shares parameters. The last softmax layer is replaced with an aggregation layer where the outputs from each network are aggregated. Here, is the number of classes. The distribution of class of bag is calculated as where is an aggregation function. In this work, we investigate 2 aggregation functions: max aggregation
Since it is unknown which sample is an outlier, we can not be sure that a bag has at least one positive instance. However, the probability that all instances in a bag are negative exponentially decreases with , thus the assumption becomes very realistic.
Experiments
The proposed methods are evaluated on a novel acoustic event classification database The dataset is available at https://data.vision.ee.ethz.ch/cvl/ae_dataset harvested from Freesound , which is a repository of audio samples uploaded by users. The database consists of 28 events as described in Table 2. Note that since the sounds in the repository are tagged in free-form style and the words used vary a lot, the harvested sounds contain irrelevant sounds. For instance, a sound tagged ’cat’ sometime does not contain a real cat meow, but instead a musical sound produced by a synthesizer. Furthermore sounds were recorded with various devices under various conditions (e.g. some sounds are very noisy and in others the acoustic event occurs during a short time interval among longer silences). This makes our database more challenging than previous datasets such as .
In order to reduce the noisiness of the data, we first normalized the harvested sounds and eliminated silent parts. If a sound was longer than 12 sec, we split the sound in pieces so that split sounds were less than 12 sec. All audio samples were converted to 16 kHz sampling rate, 16 bits/sample, mono channel. Similar to , the data was randomly split into training set (75%) and test set (25%). Only the test set was manually checked and irrelevant sounds not containing the target acoustic event, were omitted. The data augmentation was applied only to the training set.
2 Implementation details
Through all experiments, 49 band log-filter banks, log-energy and their delta and delta-delta were used as a low-level descriptor using 25 ms frames with 10 ms shift, except for the BoAW baseline described in Sec. 3.3.1. Input patch length was set to 400 frames (i.e. 4 sec). The effects of this length were further investigated in Sec. 3.3.2. During training, we randomly crop 4 sec for each sample. The networks were trained using mini-batch gradient descent based on back propagation with momentum. We applied dropout to each fully-connected layer with keeping probability . The batch size was set to 128, the momentum to 0.9. For data augmentation we used VTLP and the proposed EMDA. The number of augmented samples is balanced for each class. During testing, 4 sec patches with 50% shift were extracted and used as input to the Neural Networks. The class with the highest probability was considered the detected class. The models were implemented using the Lasagne library .
3 Experimental Results and Discussions
In our first set of experiments we compared our proposed deeper CNN architectures to three different state-of-the-art baselines, namely, BoAW , HMM+DNN/CNN as in , and classical DNN/CNN with large input field. BoAW We used MFCC with delta and delta-delta as a low-level descriptor. K-means clustering was applied to generate an audio word code book with 1000 centers. We evaluated both SVM with a kernel and a 4 layer DNN as a classifier. The layer sizes of the DNN classifier were (1024, 256, 128, 28). DNN/CNN+HMM We evaluate the DNN-HMM system. The neural network architectures are described in the left 2 columns in Table 1. Both DNN and CNN models are trained to estimate HMM state posteriors. The HMM topology consists of one state per acoustic event, and an ergodic architecture in which all states have a self-transition and equal transitions to all other states, as in . The input patch length for CNN, DNN is 30 frames with 50% shift. DNN/CNN+Large input field In order to evaluate the effect of using the proposed CNN architectures, we also evaluated the baseline DNN/CNN architectures with the same large input field, namely, 400 frame patches. The classification accuracies of these systems trained with and without data augmentation are shown in Table 3. Even without data augmentation, the proposed CNN architectures outperform all previous methods. Furthermore, the performance is significantly improved by applying data augmentation, achieving 12.5% improvement for the architecture. The best result was obtained by the architecture with data augmentation. It is important to note that the architecture outperforms classical DNN/CNN even though it has less parameters as shown in Table 1. This result supports the efficiency of deeper CNNs with small kernels for modelling large input fields.
3.2 Effectiveness of large input field
Our second set of experiments focuses on input field size. We tested our CNN with different patch size 50, 100, 200, 300, 400 frames (i.e. from 0.5 to 4 sec). The architecture was used for this experiment. As a baseline we evaluated the CNN+HNN system described in Sec. 3.3.1 but using our architecture , rather than a classical CNN. The performance improvement over the baseline is shown in Fig. 3. The result shows that larger input fields improve the performance. Especially the performance with patch length less than 1 sec sharply drops. This proves that modeling long signals directly with deeper CNN is superior to handling long sequences with HMMs.
3.3 Effectiveness of data augmentation
We verified the effectiveness of our EMDA data augmentation method in more detail. We evaluated 3 types of data augmentation: EMDA only, VTLP only, and a mixture of EMDA and VTLP (50%, 50%) with different numbers of augmented sample 10k, 20k, 30k, 40k. Fig. 4 shows that using both EDMA and VTLP always outperforms EDMA or VTLP only. This shows that EDMA and VTLP perturbs original data and create new samples in a different way, providing more effective variation of data and helping to train the network to learn a more robust and general model from limited amount of data.
3.4 Effects of Multiple Instance Learning
Finally, the and architectures with a large input field were adapted to MIL to handle the noise in the database. The number of parameters were identical since both max and Noisy OR aggregation methods are parameter free. The number of instances in a bag was set to 2. We randomly picked 2 instances from the same class during each epoch of the training. Table 4 shows that MIL didn’t improve performance. However, MIL with a medium size input field (i.e. 2 sec) performs as good as or even slightly better than single instance learning with a large input field. This is perhaps due to the fact that the MIL took the same size input length (2 sec 2 instances 4 sec), while it had less parameter. Thus it managed to learn a more robust model.
Conclusions
We proposed new CNN architectures and showed that they allow to learn a model for AED end-to-end, by directly modeling a several seconds long signal. We further proposed a method for data augmentation that prevents over-fitting and leads to superior performance even when training data is fairly limited. Experimental results shows that proposed methods significantly outperforms state of the arts. We further validated the effectiveness of deeper architectures, large input fields and data augmentation one by one. Future work will be directed towards applying the proposed AED to different applications such as video segmentation and summarization.