Deep Multimodal Learning for Audio-Visual Speech Recognition

Youssef Mroueh, Etienne Marcheret, Vaibhava Goel

Introduction

Human speech perception is not only about hearing but also about seeing: our brain integrates the waveforms representing the speech information as well as the lips poses and motions, often called visemes, which carry important visual information about what is being said. This has been demonstrated by the so called McGurk effect [MM76], which shows that a voicing of ba and a mouthing of ga is perceived as being da. In the presence of noise and multiple speakers (cocktail party effect), humans rely on lip reading in order to enhance speech recognition [CHLN08]. The visual information is also important in a clean speech scenario as it helps in disambiguating voices with similar acoustics [Sum92]. In Audio-Visual Automatic Speech Recognition (AV-ASR), both audio recordings and videos of the person talking are available at training time. It is challenging to build models that integrates both visual and audio information, and that enhance the recognition performance of the overall system. While most previous works in AV-ASR focused on enhancing the performance in the noisy case [PNLM04, NKK+11], where the visual information can be crucial, we focus in this paper on showing that the visual information is indeed helpful even in the clean speech scenario. Multimodal learning consists of fusing and relating information coming from different sources, hence AV-ASR is an important multimodal problem. Finding correlations between different modalities, and modeling their interactions, has been addressed in various learning frameworks and has been applied to AV-ASR [Gea09, LS06, MHD96, PKPM07, PKPM09, PKPM06]. Deep Neural Networks (DNN) have shown impressive performance in both audio and visual classification tasks, which is why we restrict ourselves to the deep multimodal learning framework [YGS89, NKK+11, SS, SCMN, AALB13]. In this paper, we propose methods in deep learning to fuse modalities, and validate them on the IBM AV-ASR Large Vocabulary Studio Dataset (Section 2). First we consider the training of two networks on the audio and the visual modality separately. Then, considering the last layer of each network as a better feature space, and concatenating them, we train a classifier on that joint representation, and obtain gains in Phone Error Rates (PER), with respect to an audio-only trained network. We then propose a new bilinear network that accounts for correlations between modalities and allows for joint training of the two networks, we show that a committee of such bilinear networks, fused at the level of posteriors, achieves a better PER in a clean speech scenario. The paper is organized as follows. In Section 2 we present the IBM AV-ASR large vocabulary studio dataset, our feature extraction pipeline for the audio and the visual channels. Next, in Section 3, we present results for the fusion of networks separately trained on each modality. In Section 4 we introduce the bilinear DNN that allows for a joint training and captures correlations between the two modalities, and derive its back-propagation algorithm in Section 5. Finally we present posterior combination of bimodal and bilinear bimodal DNNs in Section 6.

Audio-Visual Data Set & Feature Extraction

In this Section we present the IBM AV-ASR Large Vocabulary Studio dataset, and our feature extraction pipeline.

The IBM AV-ASR Large Vocabulary Studio Dataset consists of 4040 hours of audio-visual recordings from 262262 speakers. These were carried out in clean, studio conditions. The audio is sampled at 1616 KHz along with the video frame rate of 3030 frames per second at 704×480704\times 480 resolution. The vocabulary size in these recordings is 10,40010,400 words. This data set was divided into a test set of 22 hours of audio+video from 2222 speakers, with the rest used for training.

2 Feature Extraction

For the audio channel we extract 24 MFCC coefficients at 100 frames per second. Nine consecutive frames of MFCC coefficients are stacked and projected to 40 dimensions using an LDA matrix. Input to the audio neural network is formed by concatenating ±4\pm 4 LDA frames to the central frame of interest, resulting in an audio feature vector of dimension 360360.

For the visual channel we start by detecting the face in the image using the openCV implementation of the Viola-Jones algorithm. We then do a mouth carving by an openCV mouth detection model. Both these utilize the ENCARA2 model as described in [MDHL11]. In order to get an invariant representation to small distortions and scales we then extract level 1 and level 2 scattering coefficients [BM13] on the 64×6464\times 64 mouth region of interest and then reduce their dimension to 60 using LDA (Linear discriminant Analysis). In order to match the audio frame rate we replicate video frames according to audio and video time stamps. We also add ±4\pm 4 context frames to the central frame of interest, and obtain finally a visual feature vector of dimension 540540.

3 Context-dependent Phoneme Targets

Each audio+video frame is labeled with one of 13281328 targets that represent context dependent phonemes. 4242 phones in phonetic context of ±2\pm 2 are clustered using decision trees down to 13281328 classes. We measure classification error rate at the level of these 13281328 classes, this is referred to as phone error rate (PER).

Uni-modal DNNs & Feature Fusion

In the supervised multimodal scenario, we are given a training set SS of NN labeled examples, and CC classes:

The first multimodal modeling approach we study is to train two separate networks DNNaDNN_{a} and DNNvDNN_{v} on the audio and the visual features, respectively. The networks are optimized under the cross-entropy objective (1) using the stochastic gradient descent. We formed a joint Audio-Visual feature representation by concatenating the outputs of final hidden layers of these two networks, as shown in Figure 1. This feature space is then kept fixed while a deep or a shallow (softmax only) network is trained in this fused space up to the targets. To keep the feature space dimension manageable, we configure the individual audio and video networks to have a low dimensional final hidden layer.

We consider for DNNaDNN_{a} and DNNvDNN_{v} the following architecture dim/1024/1024/1024/1024/1024/200/1328dim/1024/1024/1024/1024/1024/200/1328, where dim=360dim=360 for DNNaDNN_{a} and dim=540dim=540 for DNNvDNN_{v}. The fused feature space dimension is 400400.

While DNNaDNN_{a} achieves a PER of 41.25%41.25\%, DNNvDNN_{v} alone achieves a PER of 69.36%69.36\%, showing that the visual information alone carries some information but that is not enough in itself to get a low error rate. A deep network built in the fused feature space results in a PER of 35.77%35.77\% while a softmax layer only in this feature space yields PER of 35.83%35.83\%. This substantial PER gain from joint audio-visual representation, even in clean audio conditions, demonstrates the value of visual information for the phoneme classification task. Interestingly, the deep and the shallow fusion are roughly on par in terms of PER. Results are summarized in the following table:

Bilinear Deep Neural Network

In the previous section the training was done separately on the two modalities, in this section we address the joint training problem, and introduce the bilinear bimodal DNN.

The fusion happens at the last hidden layer, where the posteriors capture the correlation between the intermediate non-linear features of the two modalities produced by the DNN layers, through a bilinear term. Let vL1=hL−11, vL2=hL−12v^{1}_{L}=h^{1}_{L-1},~{}v^{2}_{L}=h^{2}_{L-1}, the posteriors have the following form:

As the number of classes increases, the bilinear model becomes cumbersome computationally, and we need large training sets to get better estimates of the parameters. In order to decrease the computational complexity of the model, we propose the use of a factorization of the bilinear term, that is similar to the one in [MZHP], but is motivated in our case by Canonical Correlation Analysis (CCA) [HSST04]:

For fixed weights wyw_{y}, learning (U1,U2)(U^{1},U^{2}) corresponds to a class specific weighted CCA-like learning where we are looking for projections that maximize alignment between the intermediate features of the two modalities, in a discriminative way. Deep CCA of [AALB13] shares similarities with this model. On the other hand, for fixed (U1,U2)(U^{1},U^{2}), we can rewrite the log-posteriors in the following way:

2 Factored Bilinear Softmax With Sharing

When the classes we would like to predict are organized as the leaves of a tree structure of depth two, we can further reduce the computational complexity by sharing weights between leaves having the same parent node. This is the case in AV-ASR as the 13281328 contextual phoneme states are organized as leaves of a tree, where the parent nodes correspond to 4242 different phoneme categories. In that case we share the bilinear term across leaves having the same parents. By doing so in the case of AV-ASR, we are only taking into account the correlations between the audio and the visual channel at the phoneme level, rather than on a fine grained grid of contextual states. We can think of this sharing as a pooling operation at the phoneme level. More formally, assume that the label set Y\mathcal{Y} is partitioned into GG non overlapping groups {Yg}g=1…G\{\mathcal{Y}_{g}\}_{g=1\dots G}, we assume that:

Hence we reduce the number of weights to learn for the joint representation from C×FC\times F to G×FG\times F.

Back-propagation with the Factored Bilinear DNN with Sharing

In this section we give the back-propagation algorithm and the update rules for the bilinear DNN with sharing (bi2-DDN-wS). Recall that our classes have a tree structure with leaves yy, and parent nodes gg; a training example is therefore labeled by its leave label yy (States) as well its parent node gg (Phonemes), (x1,x2,y,g)(x^{1},x^{2},y,g), y∈{1…C}y\in\{1\dots C\}, and g∈{1…G}g\in\{1\dots G\}. We use the notation g(y)g(y) to note the group to which yy belongs, and we set Rootg(y)=1,Rootg=0,g=1…G,g≠g(y)Root_{g(y)}=1,Root_{g}=0,g=1\dots G,g\neq g(y). For the bilinear softmax with sharing, we keep track of the errors at the level of the labels (States), as well as the groups level (Phonemes):

For the layer right before the Bilinear softmax, we have a double projection to the first modality network (audio stream) and to the second modality network (visual stream). We need to compute:

Let Vj=[V1j…VCj], j∈{1,2}V^{j}=[V^{j}_{1}\dots V^{j}_{C}],~{}j\in\{1,2\}, hence the errors we propagate to each network have the following form:

For the bilinear softmax without sharing the update rules are similar (δG\delta_{G} is replaced by δL\delta_{L}).

Combining Posteriors from Bimodal and Bilinear Bimodal Networks

We experiment with various factored bi2-DNN-wS architectures, initialized at random on the IBM AV-ASR Large Vocabulary Studio Dataset. We use the following notation for the architecture of the bilinear network: [archa∣archv∣F][arch_{a}|arch_{v}|F], where archaarch_{a} and archvarch_{v} are the architectures of the audio and the visual network respectively, and FF is the dimension of the fused space. We consider architectures by increasing complexity Arch=[360,500,500,200,1328∣540,500,500,200,1328∣F=200]Arch=[360,500,500,200,1328|540,500,500,200,1328|F=200], Arch1=[360,600,600,400,100,1328∣540,600,600,400,100,1328∣F=100]Arch_{1}=[360,600,600,400,100,1328|540,600,600,400,100,1328|F=100], and Arch2=[360,500,500,500,500,500,200,1328∣540,500,500,500,500,500,200,1328∣F=200]Arch_{2}=[360,500,500,500,500,500,200,1328|540,500,500,500,500,500,200,1328|F=200]. In all our experiments we set λ=2\lambda=2. Recall that the bimodal DNN using the separate training paradigm introduced in Section 3 achieves 35.83%35.83\% PER. As shown in Table 2, each architecture alone does not improve on the bimodal DNN, but averaging the posteriors of the three architectures we obtain a small gain. A gain of 1.8%1.8\% absolute is obtained by averaging the posteriors of the bimodal and the bilinear bimodal networks, showing that the bilinear networks have uncorrelated errors with the bimodal network.

Conclusion

In this paper we have studied deep multimodal learning for the task of phonetic classification from audio and visual modalities. We demonstrate that even in clean acoustic conditions using visual channel in addition to speech results in signifiantly improved classification performance. A bilinear bimodal DNN is introduced which leverages correlation between the audio and visual modalities, and leads to further error rate reduction.

References