Naver at ActivityNet Challenge 2019 -- Task B Active Speaker Detection (AVA)
Joon Son Chung
Introduction
Multimodal speech perception has received increasing attention in recent years, with significant breakthroughs in audio-visual methods using deep learning . For many applications, it is important to be able to determine which of the visible people are speaking at any given time. To this end, the new AVA-ActiveSpeaker dataset has been an important contribution to the field since it enables deep-learning based active speaker detection (ASD) models to be trained with full supervision. This report gives a brief overview of the dataset, and describes the method that underlies our submission to the challenge.
The method is trained on the AVA-ActiveSpeaker dataset. The dataset consists of training, validation and test sets, and the splits are given in Table 1. Ground truth labels are given for the training and validation sets.
The dataset is challenging for a few reasons. The durations of speaking segments are extremely short, the average Speaking & Audible segment being 1.11 seconds. As a result, the system must be able to make accurate detection using only a small number of frames. Existing methods that require the output to be smoothed over several-second time window cannot be effective for this condition.
The dataset also contains a significant number of older videos in which the audio and the video appear to have been recorded separately or significantly out-of-sync. This means that the temporal correspondence between audio and video speech representations is not a reliable indicator of whether a person is speaking.
Model
The active speaker detection model consists of front-end feature extractors and a back-end classifier, each of which will be described in the following sections.
We use pre-trained networks to extract audio and video representations. The encoder networks have been trained on the audio-visual correspondence task using self-supervision on unlabelled videos Model and code available from the author’s website..
The video encoder is a convolutional neural network (CNN) that ingests 5 RGB image frames and outputs a 512-D representation. The architecture is based on the compact but efficient VGG-M network , but with a 3D convolution in the first layer in place of the standard 2D convolution.
The input to the audio encoder is 20 frames in the time direction and 13 cepstral coefficients in the other direction. The encoder generates a 512-D representation in the same embedding space as the video representation. The detailed architecture is given in Figure 2.
2 Back-end architecture
Both audio and video encoders ingest 5 video frame (0.2-second) input, moving 1 video frame (0.04-second) at a time. Hence for a frame input, the output dimensions are . Here, we compare two simple back-end classifiers. In our experiments, is used, but we observed no significant effect on performance for changes to the value of in the range .
LSTM classifier. The audio and video representations are fed into two separate bi-directional LSTM networks, each of 2 layers and hidden size of 128. The outputs from both networks are concatenated, then passed through a linear classification layer that predicts whether the person is speaking or not. The classifier is trained with the softmax cross-entropy loss.
TC classifier. Instead of the LSTM layers, the encoder outputs are passed to two temporal convolution layers each with 128 filters. The outputs are also concatenated and passed to the classifier, as in the case of the LSTM classifier.
Ensemble. In many applications of machine learning, it has been shown that ensemble methods can often perform better than any single classifier . Here, the predictions from both LSTM and TC classifiers are averaged with equal weighting in order to generate the final prediction.
Smoothing. In order to remove noise in the predictions, the output of the classifiers are smoothed in the temporal domain using median or Wiener filter, both over 0.5-second windows.
Experiments
Implementation details. Our implementation is based on the PyTorch library and trained on a single Tesla M40 card with 24GB memory. The network is trained using the ADAM optimiser with the default parameters and a fixed learning rate of . To eliminate the effect of bias in the training data, the number of samples for positive and negative classes are balanced in each mini-batch during training.
Evaluation metrics. The metric used in this task is the mean Average Precision (mAP). The evaluation code is provided by the organisers of the challenge.
Results. The results on the validation set using the different back-end classifiers are given in Table 2.
Our best performing model also achieved mAP of 0.878 on the held-out test set for the challenge. In comparison, the GRU-based baseline model produces mAP of 0.821.
Qualitative results of the proposed method far exceed existing correspondence-based method on this dataset, since it does not rely on precise audio-to-video synchronisation.