Slow-Fast Auditory Streams For Audio Recognition
Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, Dima Damen
Introduction
Recognising objects, interactions and activities from audio is distinct from prior efforts for scene audio recognition, due to the need for recognising sound-emitting objects (e.g. alarm clock, coffee-machine), sounds generated from interactions with objects (e.g. put down a glass, close drawer), and activities (e.g. wash, fry). This introduces challenges related to variable-length audio associated with these activities. Some can be momentary (e.g. close) while others are repetitive over a longer period (e.g. fry), and many exhibit intra-class variations (e.g. cut onion vs cut cheese). Background or irrelevant sounds are often captured with these activities. We focus on two activity-based datasets, VGG-Sound and EPIC-KITCHENS , captured from YouTube and egocentric videos respectively, and target activity recognition solely from the audio signal associated with these videos.
There is strong evidence in neuroscience for the existence of two streams in the human auditory system, the ventral stream for identifying sound-emitting objects and the dorsal streams for locating these objects. Studies suggest the ventral stream accordingly exhibits high spectral resolution for object identification, while the dorsal stream has a high temporal resolution and operates at a higher sampling rate.
Using this evidence as the driving force for designing our architecture, and inspired by a similar vision-based architecture , we propose two streams for auditory recognition: a Slow and a Fast stream, that realise some of the properties of the ventral and dorsal auditory pathways respectively. Our streams are variants of residual networks and use 2D separable convolutions that operate on frequency and time independently. The streams are fused in multiple representation levels with lateral connections from the Fast to the Slow stream, and the final representation is obtained by concatenating the global average pooled representations for action recognition.
The contributions of this paper are the following: i) we propose a novel two-stream architecture for auditory recognition that respects evidence in neuroscience; ii) we achieve state-of-the-art results on both EPIC-KITCHENS and VGG-Sound; and finally iii) we showcase the importance of fusing our specialised streams through an ablation analysis. Our pretrained models and code is available at https://github.com/ekazakos/auditory-slow-fast.
Related work
Single-stream architectures. A common approach in audio recognition for both scene and activity recognition, is to use a single-stream convolutional architecture . SoundNet uses 1D ConvNet trained in a teacher-student manner, and fine-tuned for acoustic scene classification. Single-stream 2D ConvNets have been extensively used by high-ranked entries of DCASE challenges , for acoustic scene classification. These consider spectograms as input and utilise 2D convolutions with square filters, processing frequency and time together , similarly to image ConvNets. However, symmetric filtering in frequency and time might not be optimal as the statistics of spectrograms are not homogeneous. One alternative is to utilise rectangular filters as in . Another is separable convolutions with and filters, which have recently been used in audio .
Multi-stream architectures. Late fusion of multiple streams for audio recognition was used in . Most approaches utilise modality-specific streams . In addition to late fusion, integrate multi-level fusion in their architecture in the form of attention.
In , all streams digest the same input. In , one stream takes as input low frequencies and the second inputs high frequencies. applies median filtering with different kernels at the input of each stream to model long duration sound events, medium, and short duration impulses separately. In , 1D convolutions are used with different dilation rates at each stream to model convolutional streams that operate on different temporal resolutions. The architectures of these multiple streams remain identical.
Similar to these works, we propose to utilise two-streams that consider the same input. Different from these, we design each stream with varying number of channels and temporal resolution, in addition to convolutional separation. Furthermore, we integrate the streams through multi-level fusion.
Network Architecture
Next, we describe in detail the design principles of our architecture, depicted in Figure 1. The Slow stream operates on a low sampling rate with high channel capacity to capture frequency semantics, while the Fast stream operates on a high sampling rate with more temporal convolutions and less channels to capture temporal patterns.
Input. Both streams operate on the same audio length, from which a log-mel-spectrogram is extracted. The Fast stream takes as input the whole log-mel-spectrogram without any striding, while the Slow stream uses a temporal stride of on the input log-mel-spectrogram, where .
Slow and Fast streams. The two streams are variants of ResNet50 . Each stream is comprised of an initial convolutional block with a pooling layer followed by 4 residual stages, where each stage contains multiple residual blocks. The two streams differ in their ability to capture frequency semantics and temporal patterns. The details of each stream including the number of blocks per stage and numbers of channels can be seen in Table 1.
The Slow stream has a high channel capacity, with times more channels than the Fast stream, while operating on a low sampling rate. As the input spectrogram is strided temporally by , the intermediate feature maps have a lower temporal resolution. Moreover, the Slow stream has temporal convolutions only in and (see the brown and green blocks in Fig. 1 right). By restricting the temporal resolution and the temporal kernels of the Slow stream while keeping a high channel capacity, this stream can focus on learning frequency semantics.
The Fast stream on the other hand uses no temporal striding in the input. Therefore, the intermediate feature maps have a higher temporal resolution, with temporal convolutions throughout the stream. With a high temporal resolution and more temporal kernels while having less channels, it is easier for the Fast stream to focus on learning temporal patterns.
Separable convolutions. We use separable convolutions in frequency and time as can be seen in the green block in Fig. 1 right. We break a kernel in two kernels, followed by . Separable convolutions have proven useful for video recognition . We utilise them with the motivation to separately attend to time and frequency of the input signal. We contrast separable convolutions to two-dimensional filters that convolve across both frequency and time.
Multi-level fusion. Following the approach in , we fuse the information from the Fast to the Slow stream with lateral connections, at multiple levels. We first apply a 2D temporal convolution with a kernel and a stride of to the output of the Fast stream to match the Slow stream sampling rate, and then we concatenate the downsampled feature map with the Slow stream feature map. Fusion is applied after and each residual stage.
The final representation fed to the classifier is obtained by applying time-frequency global average pooling after the last convolutional layer of both Slow and Fast streams and concantenating the pooled representations. We set and in all our experiments.
Differences compared to visual Slow-Fast . Our two-stream architecture is inspired by its visual counterpart which produces state of the art results for visual action recognition. However, key differences are introduced: Our input is 2D rather than 3D, as we operate on time-frequency while the visual Slow-Fast operates on time-space. Hence, we use 2D separable convolutions decomposed as and filters, whereas uses 3D separable convolutions decomposed as and filters. Additionally, the sampling rate for audio is naturally significantly higher than that of video, e.g. 24kHz vs 50fps in EPIC-KITCHENS-100, and the dimensionality in video is significantly higher. Accordingly, the approach in only considers a few temporal samples (8 and 32 frames in the Slow and Fast streams respectively). In contrast, our audio spectogram (see Sec 4.2) contains 100 and 400 temporal dimensions in the Slow and Fast streams respectively. To compensate for the high sampling rate of audio, we temporally downsample the representations of both streams by a factor of 4, using a temporal stride=2 in and of both streams. The remaining stages do not perform any temporal downsamplingIn preliminary experiments, we tried different downsampling schemes, such as strided convolutions throughout the whole network but they resulted in inferior performance..
Experiments
VGG-Sound. VGG-Sound is a large-scale audio dataset obtained from YouTube. It contains over 200k clips of 10s for 309 classes capturing human actions, sound-emitting objects as well as interactions. These are visually-grounded where sound emitting objects are visible in the corresponding video clip, utilising image classifiers to find correspondence between sound and image labels. Audio is sampled at 16kHz.
EPIC-KITCHENS-100. EPIC-KITCHENS-100 is the largest egocentric audio-visual dataset, containing unscripted daily activities in kitchen environments. The data are recorded in 45 different kitchens. It contains 100 hours of data, split across 700 untrimmed videos, and 90K trimmed action clips. These capture hand-object interactions as well as activities, formed as the combination of a verb and a noun (e.g. “cut onion” and “wash plate”), where there are 97 verb classes, 300 noun classes, and 4025 action classes (many verbs and nouns do not co-occur). The classes are highly unbalanced. Actions are mainly short-term (average action length is 2.6s with minimum length 0.25s). Audio is sampled at 24kHz.
2 Experimental protocol
Feature extraction. We extract log-mel-spectrograms with 128 Mel bands using the Librosa library. For VGG-Sound, we use 5.12s of audio with a window of 20ms and a hop of 10ms, resulting in spectrograms of size . For EPIC-KITCHENS-100, we use 2s of audio with a 10ms window and a 5ms hop, resulting in spectrograms of size . For clips 2s in EPIC-KITCHENS-100, we duplicate the last time-frame of the log-mel-spectrogram.
Train / Val details. All models are trained using SGD with momentum set to 0.9 and cross-entropy loss. We train on EPIC-KITCHENS-100 as a multitask learning problem, as in , using two prediction heads, one for verbs and one for nouns. We train on VGG-Sound from random initialisation for 50 epochs and fine-tune on EPIC-KITCHENS-100 using the VGG-Sound pretrained models for 30 epochs. We drop the learning rate by 0.1 at epochs 30 and 40 for VGG-Sound, and at epochs 20 and 25 for EPIC-KITCHENS-100. For fine-tuning, we freeze Batch-Normalisation layers except the first one, as done in . For regularisation, we use dropout on the concatenation of Slow and Fast streams with probability 0.5, plus weight decay in all trainable layers using the value of . For data augmentation during training, we use the implementation of SpecAugment from and set its parameters as follows: 2 frequency masks with F=27, 2 time masks with T=25, and time warp with W=5. During training we randomly extract one audio segment from each clip. During testing we average the predictions of 2 equally distanced segments for VGG-Sound, and 10 for EPIC-KITCHENS-100.
Evaluation metrics. For VGG-Sound, we follow the evaluation protocol of and report mAP, AUC, and d-prime, as defined in . Additionally we report top-1/5% accuracy. For EPIC-KITCHENS-100, we follow the evaluation protocol of and report top-1 and top-5 % accuracy for the validation and test sets separately, as well as for the subset of unseen participants within val/test.
Baselines and ablation study. We compare to published state-of-the-art results in each dataset. For VGG-Sound, we also compare against using their publicly available code, which is the closest work to ours in motivation, as it uses two audio streams separating input into low/high frequencies.
We also perform an ablation study investigating the importance of the two streams as follows:
Slow, Fast: We compare to each single stream individually.
Enriched Slow stream: We combine two Slow streams with late fusion of predictions, as well as a deeper Slow stream (ResNet101 instead of ResNet50).
Slow-Fast without multi-level fusion: Streams are fused by averaging their predictions, without lateral connections.
3 Results
EPIC-KITCHENS-100 Our proposed network achieves state-of-the-art results as can be seen in Table 2 for both Val and Test. Our previous results use a TSN with BN-Inception architecture , initialised from ImageNet, while here we utilise pre-training from VGG-Sound. Our proposed architecture outperforms by a good margin. We report the ablation comparison using the published Val split. The significant improvement in our proposed Slow-Fast architecture when compared to Slow and Fast streams independently shows that there is complementary information in the two streams that benefit audio recognition. The Slow stream performs better than Fast, due to the increased number of channels. When comparing to the enriched Slow architectures (see the last column of Table 2 for number of parameters), our proposed model still significantly outperforms these baselines, showcasing the need for the two different pathways. We conclude that the synergy of Slow and Fast streams is more important than simply increasing the number of parameters of the stronger Slow stream. Finally, our proposed architecture consistently outperforms late fusion, indicating the importance of multi-level fusion with lateral connections.
VGG-Sound. We report results in Table 3 comparing to state-of-the-art from , which uses a single-stream ResNet50 architecture, which uses a ResNet variant with 19 layers as backbone for their two-stream architecture with significantly less parameters than our model at 3.2M parameters, as well as ablations of our model. We report the best performing model on the test set in each case. Our proposed Slow-Fast architecture outperforms and . The rest of our observations on the ablations from EPIC-KITCHENS-100 hold for VGG-Sound as well, with a key difference: the gap in performance between single streams and our proposed two-stream architecture is even bigger for VGG-Sound, indicating more complementary information in the two streams. The fact that Slow-Fast outperforms Slow by such a large accuracy gap with an insignificant increase in parameters indicates the efficient interaction between Slow and Fast streams.
Conclusion
We propose a two-stream architecture for audio recognition, inspired by the two pathways in the human auditory system, fusing Slow and Fast streams with multi-level lateral connections. We showcase the importance of our fusion architecture through ablations on two activity-based datasets, EPIC-KITCHENS-100 and VGG-Sound, achieving state-of-the-art performance. For future work, we will explore learning the stride parameter and assessing the impact of the number of channels. We hope that this work will pave the path for efficient multi-stream training in audio.
Acknowledgements. Publicly-available datasets were used for this work. Kazakos is supported by EPSRC DTP, Damen by EPSRC Fellowship UMPIRE (EP/T004991/1) and Nagrani by Google PhD fellowship. Research is also supported by Seebibyte (EP/M013774/1).
References
Appendix
In this additional material, we provide further insight into what each of the Slow and Fast streams learn, through class analysis and visualising feature maps from each stream. We also offer an ablation on separable convolutions. Finally, we detail the hyperparameters used to train on VGG-Sound.
Appendix A Class performance of two streams
In Figure 2, we distinguish between VGG-Sound classes that are better predicted from the Slow stream to the left, and classes that are better predicted from the Fast stream to the right. To obtain these, we calculated per-class accuracy and retrieved classes for which the accuracy difference is above a threshold. Particularly, we used to retrieve classes best predicted from Slow and to retrieve classes best predicted from Fast. We used a higher threshold for the Slow stream as it more frequently outperforms the Fast stream, as shown in our earlier results.
As can be seen in Figure 2, Slow predicts better animals and scenes. This matches our intuition that Slow focuses on learning frequency patterns as different animals make distinct sounds at different frequencies, e.g. mosquito buzzing vs whale calling, requiring a network with fine spectral resolution to distinguish between those. In Scenes, there are classes such as sea waves, airplane and wind chime, that contain slow evolving sounds.
The Fast stream, in contrast, can better predict classes with percussive sounds like playing drum kit, tap dancing, woodpecker pecking tree, and popping popcorn. This also matches our design motivation that the Fast stream learns better temporal patterns as these classes contain temporally localised sounds that require a model with fine temporal resolution. Interestingly, Fast is better at human speech, laughter, singing, and other human voices, where we speculate that it can better capture articulation.
Appendix B Visualising Feature Maps
We show examples of feature maps from Slow and Fast streams, when trained independently (Fig 3). In each case, we show two samples from classes that are better predicted from the corresponding stream. For Slow, these are sea waves and mosquito buzzing, compared to woodpecker pecking tree and playing vibraphone for Fast. In each case, we show the input spectogram as well as feature maps from residual stages 3 and 5. In each plot, the horizontal axis represents time while the vertical axis corresponds to frequency. We visualise a single channel from each feature map, manually chosen.
In Figure 3, we demonstrate that Fast is capable of detecting the hits of the woodpecker on the tree as well as the hits on the vibraphone, while Slow extracts frequency patterns that do not seem to be useful for discriminating these classes that contain temporally localised sounds. For sea waves and mosquito buzzing, Slow extracts frequency patterns over time, while Fast aims to temporally localise events, which does not assist the discrimination of these classes.
Appendix C Ablation of separable convolutions
We provide an ablation of separable convolutions in Table 4. We trained the ResNet50 architecture as proposed in without separable convolutions, as well as a variant with separable convolutions. We compare this to the published results by Chen et al. that also uses a ResNet50 architecture. Our reproduced results already outperform . ResNet50-separable has separable convolutions as used in our Slow-Fast network (see Figure 1 and Table 1).
Results show that ResNet50-separable achieves slightly better results than ResNet50 in all metrics except Top-5. Although accuracy is not significantly increased in this ablation, we employ separable convolutions in our proposed architecture, following our motivation to attend differently to frequency and time. These results also show that a single stream ResNet50 has comparable performance to our two stream proposal, however ours performs better in accuracy and the two streams accommodate different characteristics of audio classes as shown previously.
Appendix D Hyperparameter Details
Training the publicly available code of McDonnell & Gao with the default hyperparameters on VGG-Sound provided poor results. We tuned the hyperparameters as follows: We set the maximum learning rate to , train the network for epochs, with for mixup. Lastly, we adjusted the number of FFT points to for log-mel-spectrogram extraction, to apply a window and hop length similar to the ones in (their datasets are sampled at 48kHz and 44.1kHz, while VGG-Sound is sampled at 16kHz).