Multi-target DoA Estimation with an Audio-visual Fusion Mechanism
Xinyuan Qian, Maulik Madhavi, Zexu Pan, Jiadong Wang, Haizhou Li
Introduction
In human-robot interaction, a robot relies on its Sound Source Localization (SSL) mechanism to direct its attention. Traditionally, SSL approaches only use audio signals and attempt as a signal processing problem . However, those approaches are adversely affected by acoustically challenged conditions, such as noise and reverberation scenarios . To address that, several Neural Networks (NN)-based approaches were explored assuming a sufficient amount of data are available. Specifically, location-related Short-Time-Fourier-Transform (STFT) cues are mapped to sound DoA information in while the Generalized Cross Correlation with Phase Transform (GCC-PHAT) cues are used in . Despite the progress, many research problems remain. One of them is multi-speaker localization in real multi-party human-robot interaction scenarios under acoustic challenging conditions .
Considering seeing and hearing are the two most essential human cognitive abilities, studies observed that audio and video convey complementary information and may help to overcome uni-modal limitations of a degradation condition for scene analysis . There is a very broad literature of audio-visual approaches for speaker localization over the past decades . However, it was not until recently that the deep learning-based approaches have attracted more attention, thanks to the increasing computational power and rapid development in NN techniques. Nevertheless, most of these methods aim at locating sound sources in visual scenes . Specifically, an attention mechanism is incorporated into the individual sound and vision network to model the audio-visual image correspondence . A visual saliency network is employed in , together with an audio representation network, to feature a SSL module for producing an audio-visual saliency map. An attention network is proposed in to learn the visual regions of a sounding event. By fusing audio and visual features using LSTM and bilinear pooling, the audio assisted visual feature extraction is described in . All the research studies use audio as a supplementary modality for visual localization and require the sound sources to be both audible and visible.
Unlike the prior studies, we aim to perform audio-visual speaker localization in the spatial DoA domain where targets can appear either inside (visible) or outside (invisible) the camera’s Field-of-View (FoV). We propose two neural network architectures and make the following contributions in this paper: (1) we propose a novel video simulation method to deal with the lack of video data; (2) for the first time, we design a deep learning network for audio-visual multi-speaker DoA estimation, and (3) we adopt an adaptive weighting mechanism in a simple feedforward network to estimate the multi-modal reliability under different conditions.
Proposed Method
Given a sequence of frame-synchronized audio and video signals captured by a microphone array and a calibrated camera, we aim to estimate the DoA information for each sound source at each frame. Next, we describe the way we characterize audio and video signals, the video simulation method, and the proposed neural networks.
The GCC-PHAT is widely used to calculate the time different of arrival (TDOA) between any two microphones in a microphone array . We adopt it as the audio feature due to its robustness in the noisy and reverberant environment and the fewer tunable parameters than the other counterparts e.g. STFT . Let and be the Fourier transforms of audio sequence at and channels of the microphone array, respectively. We compute the GCC-PHAT features with different delay lags as:
where denotes the complex conjugate operation, denotes the real part of complex number and denotes the FFT length. Here, the delay lag between two signals arrived is reflected in the steering vector in Eq. 1.
2 Visual features and simulation
With the advent of deep learning, accurate face detection at low computational cost becomes widely available . Let us define as the face detection bounding box , where ⊺ denotes transpose, are the horizontal and vertical positions of the top-left point, are the width and height, and is the number of detected faces. The central point of detection is thus computed as:
The visual feature is encoded as the exponential part of the multi-variant Gaussian distribution (in and direction) with the standard deviations specified by the detection width and height and achieves the maximum at the central point:
where indicates the potential image positions, is a diagonal covariance matrix, and indicates uniform distribution. The components in are re-sampled to the same length of GCC-PHAT.
Audio-visual parallel data are not abundantly available. However, it is possible to obtain the camera’s extrinsic and intrinsic calibration parameters and , the 3D location of a sound source. We propose a novel method to synthesize visual features in synchrony with the audio features by Eq. 3. The overall pipeline of visual feature generation is illustrated in Fig. 2 and the process is formulated next.
We first add three-variant Gaussian distributed spatial noise to the target 3D location to account for possible face detection error, and transfer the resulting point to the camera coordinates given the extrinsic parameters:
with noise covariance matrix assuming that the additive noises to are independent, and is the transformation using the pin-hole camera model .
Then, we geometrically create the 3D face bounding box whose plane is perpendicular to the camera’s optical axis ( in Fig. 2), and project to the image plane:
where is the 3D-to-image projection, is the translation vector which equals to for the top-left point and for the bottom-right point , respectively. and are the width and height assumptions of a real human face.
Finally, the simulated face detection bounding box is computed as , where denotes a concatenation operation to form a column vector.
3 Neural network architecture
We propose two NN architectures for audio-visual speaker DoA estimation based on Multilayer Perceptron (MLP), namely MLP Audio-Visual Concatenation (MLP-AVC) and MLP Audio-Visual Adaptive Weighting (MLP-AVAW), which specify different ways of audio-visual feature fusion and classifier design as illustrated in Fig. 3.
MLP-AVC consists of three hidden layers, denoted as MLP3 in Fig. 3(a) by a dotted blue box, each one is a fully-connected layer with ReLU activation and batch normalization . It takes the flattened and concatenated GCC-PHAT and visual features as an input vector. The network is trained to predict the probability of DoA labels, as in , using a sigmoid output layer. MLP-AVC adopts an early fusion strategy by concatenating audio and visual features. We hypothesize that such early fusion doesn’t learn to pay selective attention to uni-modal features, that are crucial in face of missing data or noisy data.
MLP-AVAW introduces an adaptive weighting mechanism, which uses a tiny NN with two fully-connected layers, colored in purple in Fig. 3(b)), to learn three adaptive weights for the audio GCC-PHAT feature, video image horizontal and vertical features, respectively. A softmax activation function is applied for weights normalization. We call this as ‘adaptive weighting’ mechanism as the weights are adapted according to the live input during inference. Finally, the weighted multi-modal features are concatenated for MLP3 to compute DoA.
Experiments
The existing audio-visual datasets, such as AV16.3 , CAV3D , and AVASM , are either of limited size, or don’t provide the spatial ground truth. We, therefore, simulate the synchronized visual features for a SSL dataset of the loudspeaker cases. We choose the recently released SSLR datasetSSLR dataset: https://www.idiap.ch/dataset/sslr/ , that is recorded in a physical setup from one or two concurrent speakers, and with adequate target 3D annotations. It consists of 4-channel audio recordings at sampling rate, that is organized into three subsets, namely train (loudspeaker), test-human, and test-loudspeaker.
We evaluate the performance of DoA estimates using the same metrics of i.e. Mean Absolute Error (MAE) and Accuracy (ACC), where MAE is defined as the mean absolute error between the actual and the estimated DoA, while the accuracy allowance of ACC is in the classification prediction.
For the test-human subset, we apply the RetinaFace detector [Deng_2020_CVPR] to achieve the face bounding boxes. For the train and test-loudspeaker subsets, the visual features are simulated with the method proposed in Sec. 2.2 with a noise covariance matrix . Fig. 4(a) illustrates the ground truth camera (magenta) and target 3D locations for the train (blue), test-loudspeaker (green) and test-human (red) subsets for all frames. Targets in the gray region are inside the camera’s FoV, therefore, visible to the camera. We only generate face bounding boxes of visible targets, as visualized in Fig. 4(b-c) and formulated in Eq. 4-5 with the simulated bounding box . Fig. 4 shows that the face bounding boxes spread well across the FoV with a balanced distribution. We don’t generate bounding boxes for speakers that are outside the FoV. As a result, the visual features for the invisible speakers become missing data (the normal distribution in Eq. 3 for visual feature representation) in the audio-visual dataset.
The statistics of simulated visual features are summarized in Tab. 1 where DR represents the percentage of video frames having targets inside the FoV. Low DR means a high percentage of missing visual features. We also report in Tab. 1 the DoA MAE and ACC of the simulated visual features, indicating that the simulated data is of enough difficulty to represent real scenarios.
2 Parameter settings
The GCC-PHAT is computed for every segments with delay lags , resulting in 51 coefficients for each microphone pair as in . With 6 microphone pairs, each pair contributing 51 GCC-PHAT coefficients, we obtain 306 GCC-PHAT coefficients. For visual features, the human face width and height are assumed to have , respectively as such in . We adjust the size of the horizontal and vertical visual feature encoding to 51 to match that of GCC-PHAT coefficients.
We use the Adam optimizer . All models are trained for 10 epochs with a batch size of 256 samples and a learning rate of 0.001. Since multi-speaker localization is not a single-label classification problem, we use Mean Square Error (MSE) instead of cross-entropy as the loss function.
3 Results
Tab. 2 provides the experimental results on the SSLR test set. Results are separately reported for different subsets and the speaker number (assumed to be known). The best result for each column is in the bold font. We compare the results of MLP-AVC and MLP-AVAW with two audio baseline methods: the traditional Steered Response Power PHAse Transform (SRP-PHAT) method and the state-of-the-art MLP-GCC method . As speakers are not always visible, we don’t provide the video-only baseline to avoid unfair comparison. Furthermore, Tab. 1 suggests that it is challenging to expect visual features alone to outperform the audio DoA estimation.
Tab. 2 shows that, by both early fusion of audio-visual features. In particular, MLP-AVC reduces MAE from (MLP-GCC) to , which confirms the audio-visual fusion benefits. For the test-human subset, speakers are mostly inside the camera’s FoV (the red points locate in the gray region in Fig. 4(a)) and DR of the RetinaFace detector [Deng_2020_CVPR] achieves 100 %, which is much higher than DR in test-loudspeaker (9.2 %). Thus, the MAE degradation in test-human (from to and from to ) is more significant than in test-loudspeaker (from to and from to ). Besides, further improvements are introduced by the adaptive weighting mechanism in MLP-AVAW, which achieves the best results in most cases with the overall MAE at and ACC at %.
Next, we further evaluate the noise robustness of the proposed networks. For audio, we apply additive white Gaussian noise of SNRs varying from to on the original SSLR audio signals. For video, we randomly swap up to face detections to the other frames to generate false positives and false negatives. Tab. 3 lists the overall MAE and ACC of MLP-AVAW in comparison with those under clean audio condition. We also provide the MLP-GCC results in the first two columns indicating the audio-only performance without swapping the face detection. From the results, we can see that fusing visual features always brings benefits. Additionally, audio is of more importance than video since with the degradation of SNR, both MAE and ACC are getting worse as Face Detection Swap Percentage (FDSP) increases, the performance degradation is also obvious but not so significant. Even at FDSP=, the proposed network still outperforms the MLP-GCC. The performance gains by MLP-AVAW suggest that visual features provide additional information in degraded acoustic conditions.
Conclusions
This paper presented two neural network architectures for multi-speaker DoA estimation using audio-visual signals. The comprehensive evaluation results confirm the benefits of audio-visual fusion and the adaptive weighting mechanism. Besides, we proposed a technique to synthesize visual features from geometric information about the sound sources to deal with lack of annotated audio-visual data. Future work will include exploring network models that can generalize with limited training data.