Learning Factorized Multimodal Representations

Yao-Hung Hubert Tsai, Paul Pu Liang, Amir Zadeh, Louis-Philippe Morency, Ruslan Salakhutdinov

Introduction

Multimodal machine learning involves learning from data across multiple modalities (Baltrušaitis et al., 2017). It is a challenging yet crucial research area with real-world applications in robotics (Liu et al., 2017), dialogue systems (Pittermann et al., 2010), intelligent tutoring systems (Petrovica et al., 2017), and healthcare diagnosis (Frantzidis et al., 2010). At the heart of many multimodal modeling tasks lies the challenge of learning rich representations from multiple modalities. For example, analyzing multimedia content requires learning multimodal representations across the language, visual, and acoustic modalities (Cho et al., 2015). Although the presence of multiple modalities provides additional valuable information, there are two key challenges to address when learning from multimodal data: 1) models must learn the complex intra-modal and cross-modal interactions for prediction (Zadeh et al., 2017), and 2) trained models must be robust to unexpected missing or noisy modalities during testing (Ngiam et al., 2011).

In this paper, we propose to optimize for a joint generative-discriminative objective across multimodal data and labels. The discriminative objective ensures that the representations learned are rich in intra-modal and cross-modal features useful towards predicting the label, while the generative objective allows the model to infer missing modalities at test time and deal with the presence of noisy modalities. To this end, we introduce the Multimodal Factorization Model (MFM in Figure 1) that factorizes multimodal representations into multimodal discriminative factors and modality-specific generative factors. Multimodal discriminative factors are shared across all modalities and contain joint multimodal features required for discriminative tasks. Modality-specific generative factors are unique for each modality and contain the information required for generating each modality. We believe that factorizing multimodal representations into different explanatory factors can help each factor focus on learning from a subset of the joint information across multimodal data and labels. This method is in contrast to jointly learning a single factor that summarizes all generative and discriminative information (Srivastava & Salakhutdinov, 2012). To sum up, MFM defines a joint distribution over multimodal data, and by the conditional independence assumptions in the assumed graphical model, both generative and discriminative aspects are taken into account. Our model design further provides interpretability of the factorized representations.

Through an extensive set of experiments, we show that MFM learns improved multimodal representations with these characteristics: 1) The multimodal discriminative factors achieve state-of-the-art or competitive performance on six multimodal time series datasets. We also demonstrate that MFM can generalize by integrating it with other existing multimodal discriminative models. 2) MFM allows flexible generation concerning multimodal discriminative factors (labels) and modality-specific generative factors (styles). We further show that we can perform reconstruction of missing modalities from observed modalities without significantly impacting discriminative performance. Finally, we interpret our learned representations using information-based and gradient-based methods, allowing us to understand the contributions of individual factors towards multimodal prediction and generation.

Multimodal Factorization Model

Multimodal Factorization Model (MFM) is a latent variable model (Figure 1(a)) with conditional independence assumptions over multimodal discriminative factors and modality-specific generative factors. According to these assumptions, we propose a factorization over the joint distribution of multimodal data (Section 2.1). Since exact posterior inference on this factorized distribution can be intractable, we propose an approximate inference algorithm based on minimizing a joint-distribution Wasserstein distance over multimodal data (Section 2.2). Finally, we derive the MFM objective by approximating the joint-distribution Wasserstein distance via a generalized mean-field assumption.

Notation: We define X1:M\mathbf{X}_{1:M} as the multimodal data from MM modalities and Y\mathbf{Y} as the labels, with joint distribution PX1:M,Y=P(X1:M,Y){P}_{\mathbf{X}_{1:M},\mathbf{Y}}={P}(\mathbf{X}_{1:M},\mathbf{Y}). Let X^1:M\mathbf{\hat{X}}_{1:M} denote the generated multimodal data and Y^\mathbf{\hat{Y}} denote the generated labels, with joint distribution PX^1:M,Y^=P(X^1:M,Y^){P}_{\mathbf{\hat{X}}_{1:M},\mathbf{\hat{Y}}}={P}(\mathbf{\hat{X}}_{1:M},\mathbf{\hat{Y}}).

To factorize multimodal representations into multimodal discriminative factors and modality-specific generative factors, MFM assumes a Bayesian network structure as shown in Figure 1(a). In this graphical model, factors Fy\mathbf{F_{y}} and Fa{1:M}\mathbf{F_{a}}_{\{1:M\}} are generated from mutually independent latent variables Z=[Zy,Za{1:M}]\mathbf{Z}=[\mathbf{Z_{y}},\mathbf{Z_{a}}_{\{1:M\}}] with prior PZP_{\mathbf{Z}}. In particular, Zy\mathbf{Z_{y}} generates the multimodal discriminative factor Fy\mathbf{F_{y}} and Za{1:M}\mathbf{Z_{a}}_{\{1:M\}} generate modality-specific generative factors Fa{1:M}\mathbf{F_{a}}_{\{1:M\}}. By construction, Fy\mathbf{F_{y}} contributes to the generation of Y^\mathbf{\hat{Y}} while {Fy,Fai}\{\mathbf{F_{y}},\mathbf{F_{a}}_{i}\} both contribute to the generation of X^i\mathbf{\hat{X}}_{i}. As a result, the joint distribution P(X^1:M,Y^){P}(\mathbf{\hat{X}}_{1:M},\mathbf{\hat{Y}}) can be factorized as follows:

with dF=dFy∏i=1MdFaid\mathbf{F}=d\mathbf{F_{y}}\prod_{i=1}^{M}d\mathbf{F_{a}}_{i} and dZ=dZy∏i=1MdZaid\mathbf{Z}=d\mathbf{Z_{y}}\prod_{i=1}^{M}d\mathbf{Z_{a}}_{i}.

Exact posterior inference in Equation 1 may be analytically intractable due to the integration over Z\mathbf{Z}. We therefore resort to using an approximate inference distribution Q(Z∣X1:M,Y)Q(\mathbf{Z}|\mathbf{X}_{1:M},\mathbf{Y}) as detailed in the following subsection. As a result, MFM can be viewed as an autoencoding structure that consists of encoder (inference) and decoder (generative) modules (Figure 1(c)). The encoder module for Q(⋅∣⋅)Q(\cdot|\cdot) allows us to easily sample Z\mathbf{Z} from an approximate posterior. The decoder modules are parametrized according to the factorization of P(X^1:M,Y^∣Z){P}(\mathbf{\hat{X}}_{1:M},\mathbf{\hat{Y}}|\mathbf{Z}) as given by Equation 1 and Figure 1(a).

2 Minimizing Joint-Distribution Wasserstein Distance over Multimodal Data

Two common choices for approximate inference in autoencoding structures are Variational Autoencoders (VAEs) (Kingma & Welling, 2013) and Wasserstein Autoencoders (WAEs) (Zhao et al., 2017; Tolstikhin et al., 2017). The former optimizes the evidence lower bound objective (ELBO), and the latter derives an approximation for the primal form of the Wasserstein distance. We consider the latter since it simultaneously results in better latent factor disentanglement (Zhao et al., 2017; Rubenstein et al., 2018) and better sample generation quality than its counterparts (Chen et al., 2016; Higgins et al., 2016; Kingma & Welling, 2013). However, WAEs are designed for unimodal data and do not consider factorized distributions over latent variables that generate multimodal data. Therefore, we propose a variant for handling factorized joint distributions over multimodal data.

As suggested by Kingma & Welling (2013), we adopt the design of nonlinear mappings (i.e. neural network architectures) in the encoder and decoder (Figure 1 (c)). For the encoder Q(Z∣X1:M,Y)Q(\mathbf{Z}|\mathbf{X}_{1:M},\mathbf{Y}), we learn a deterministic mapping Qenc:X1:M,Y→ZQ_{enc}:\mathbf{X}_{1:M},\mathbf{Y}\rightarrow\mathbf{Z} (Rubenstein et al., 2018; Tolstikhin et al., 2017). For the decoder, we define the generation process from latent variables as Gy:Zy→FyG_{y}:\mathbf{Z_{y}}\rightarrow\mathbf{F_{y}}, Ga{1:M}:Za{1:M}→Fa{1:M}G_{a\{1:M\}}:\mathbf{Z}_{\mathbf{a}\{1:M\}}\rightarrow\mathbf{F}_{\mathbf{a}\{1:M\}}, D:Fy→Y^D:\mathbf{F_{y}}\rightarrow\mathbf{\hat{Y}}, and F1:M:Fy,Fa{1:M}→X^1:MF_{1:M}:\mathbf{F_{y}},\mathbf{F}_{\mathbf{a}\{1:M\}}\rightarrow\mathbf{\hat{X}}_{1:M}, where Gy,Ga{1:M},DG_{y},G_{a\{1:M\}},D and F1:MF_{1:M} are deterministic functions parametrized by neural networks.

Let Wc(PX1:M,Y,PX^1:M,Y^)W_{c}({P}_{\mathbf{X}_{1:M},\mathbf{Y}},{P}_{\mathbf{\hat{X}}_{1:M},\mathbf{\hat{Y}}}) denote the joint-distribution Wasserstein distance over multimodal data under cost function cXi{c_{X}}_{i} and cYc_{Y}. We choose the squared cost c(a,b)=∥a−b∥22c(a,b)=\|a-b\|_{2}^{2}, allowing us to minimize the 2-Wasserstein distance. The cost function can be defined not only on static data but also on time series data such as text, audio and videos. For example, given time series data X=[X1,X2,⋯ ,XT]\mathbf{X}=[X^{1},X^{2},\cdots,X^{T}] and X^=[X^1,X^2,⋯ ,X^T]\mathbf{\hat{X}}=[\hat{X}^{1},\hat{X}^{2},\cdots,\hat{X}^{T}], we define c(X,X^)=∑t=1T∥Xt−X^t∥22c(\mathbf{X},\mathbf{\hat{X}})=\sum_{t=1}^{T}\|X^{t}-\hat{X}^{t}\|_{2}^{2}.

With conditional independence assumptions in Equation 1, we express Wc(PX1:M,Y,PX^1:M,Y^)W_{c}({P}_{\mathbf{X}_{1:M},\mathbf{Y}},{P}_{\mathbf{\hat{X}}_{1:M},\mathbf{\hat{Y}}}) as:

For any functions Gy:Zy→FyG_{y}:\mathbf{Z_{y}}\rightarrow\mathbf{F_{y}}, Ga{1:M}:Za{1:M}→Fa{1:M}G_{a\{1:M\}}:\mathbf{Z}_{\mathbf{a}\{1:M\}}\rightarrow\mathbf{F}_{\mathbf{a}\{1:M\}}, D:Fy→Y^D:\mathbf{F_{y}}\rightarrow\mathbf{\hat{Y}}, and F1:M:Fa{1:M},Fy→X^1:MF_{1:M}:\mathbf{F}_{\mathbf{a}\{1:M\}},\mathbf{F_{y}}\rightarrow\mathbf{\hat{X}}_{1:M}, we have Wc(PX1:M,Y,PX^1:M,Y^)=W_{c}({P}_{\mathbf{X}_{1:M},\mathbf{Y}},{P}_{\mathbf{\hat{X}}_{1:M},\mathbf{\hat{Y}}})=

where PZ{P_{\mathbf{Z}}} is the prior over Z=[Zy,Za{1,M}]\mathbf{Z}=[\mathbf{Z_{y}},\mathbf{Z_{a}}_{\{1,M\}}] and QZ{Q_{\mathbf{Z}}} is the aggregated posterior of the proposed approximate inference distribution Q(Z∣X1:M,Y){Q}({\mathbf{Z}|\mathbf{X}_{1:M},\mathbf{Y}}).

Proof: The proof is adapted from Tolstikhin et al. (Tolstikhin et al., 2017). The two differences are: (1) we show that P(X^1:M,Y^∣Z=z)P(\mathbf{\hat{X}}_{1:M},\mathbf{\hat{Y}}|\mathbf{Z}=z) are Dirac for all z∈Zz\in\mathcal{Z}, and (2) we use the fact that c((X1:M,Y),(X^1:M,Y^))=∑i=1McXi(Xi,X^i)+cY(Y,Y^)c((\mathbf{X}_{1:M},\mathbf{Y}),(\mathbf{\hat{X}}_{1:M},\mathbf{\hat{Y}}))=\sum_{i=1}^{M}{c_{X}}_{i}(\mathbf{X}_{i},\mathbf{\hat{X}}_{i})+c_{Y}(\mathbf{Y},\mathbf{\hat{Y}}). Please refer to the supplementary material for proof details. ■\blacksquare

The constraint on QZ=PZQ_{\mathbf{Z}}=P_{\mathbf{Z}} in Proposition 1 is hard to satisfy. To obtain a numerical solution, we first relax the constraint by performing a generalized mean field assumption on QQ according to the conditional independence as shown in the inference network of Figure 1 (b):

The intuition here is based on our design that Zy\mathbf{Z_{y}} generates the multimodal discriminative factor Fy\mathbf{F_{y}} and Za{1:M}\mathbf{Z_{a}}_{\{1:M\}} generate modality-specific generative factors Fa{1:M}\mathbf{F_{a}}_{\{1:M\}}. Therefore, the inference for Zy\mathbf{Z_{y}} should depend on all modalities X1:M\mathbf{X}_{1:M} and the inference for Zai\mathbf{Z_{a}}_{i} should depend only on the specific modality Xi\mathbf{X}_{i}. Following this assumption, we define Q\mathcal{Q} as a nonparametric set of all encoders that fulfill the factorization in Equation 3. A penalty term is added into our objective to find the Q(Z∣⋅)∈QQ(\mathbf{Z}|\cdot)\in\mathcal{Q} that is the closest to prior PZP_{\mathbf{Z}}, thereby approximately enforcing the constraint QZ=PZQ_{\mathbf{Z}}=P_{\mathbf{Z}}:

where λ\lambda is a hyper-parameter and MMD\mathcal{MMD} is the Maximum Mean Discrepancy (Gretton et al., 2012) as a divergence measure between QZ{Q}_{\mathbf{Z}} and PZ{P}_{\mathbf{Z}}. The prior PZ{P}_{\mathbf{Z}} is chosen as a centered isotropic Gaussian N(0,I)\mathcal{N}(\mathbf{0},\mathbf{I}), so that it implicitly enforces independence between the latent variables Z=[Zy,Za{1,M}]\mathbf{Z}=[\mathbf{Z_{y}},\mathbf{Z_{a}}_{\{1,M\}}] (Higgins et al., 2016; Kingma & Welling, 2013; Rubenstein et al., 2018).

Equation 4 represents our hybrid generative-discriminative optimization objective over multimodal data: the first loss term ∑i=1McXi(Xi,F(Gai(Zai),Gy(Zy)))\sum_{i=1}^{M}c_{X_{i}}(\mathbf{X}_{i},F(G_{ai}(\mathbf{Z}_{\mathbf{a}i}),G_{y}(\mathbf{Z_{y}}))) is the generative objective based on reconstruction of multimodal data and the second term cY(Y,D(Gy(Zy)))c_{Y}(\mathbf{Y},D(G_{y}(\mathbf{Z_{y}}))) is the discriminative objective. In practice we compute the expectations in Equation 4 using empirical estimates over the training data. The neural architecture of MFM is illustrated in Figure 1(c).

3 Surrogate Inference for Missing Modalities

A key challenge in multimodal learning involves dealing with missing modalities. A good multimodal model should be able to infer the missing modality conditioned on the observed modalities and perform predictions based only on the observed modalities. To achieve this objective, the inference process of MFM can be easily adapted using a surrogate inference network to reconstruct the missing modality given the observed modalities. Formally, let Φ\Phi denote the surrogate inference network. The generation of missing modality X^1\hat{\mathbf{X}}_{1} given the observed modalities X2:M\mathbf{X}_{2:M} can be formulated as

4 Encoder and Decoder Design

We now discuss the implementation choices for the MFM neural architecture in Figure 1(c). The encoder Q(Zy∣X1:M)Q(\mathbf{Z_{y}}|\mathbf{X}_{1:M}) can be parametrized by any model that performs multimodal fusion (Morency et al., 2011; Zadeh et al., 2017). For multimodal image datasets, we adopt Convolutional Neural Networks (CNNs) and Fully-Connected Neural Networks (FCNNs) with late fusion (Nojavanasghari et al., 2016) as our encoder Q(Zy∣X1:M)Q(\mathbf{Z_{y}}|\mathbf{X}_{1:M}). The remaining functions in MFM are also parametrized by CNNs and FCNNs. For multimodal time series datasets, we choose the Memory Fusion Network (MFN) (Zadeh et al., 2018a) as our multimodal encoder Q(Zy∣X1:M)Q(\mathbf{Z_{y}}|\mathbf{X}_{1:M}). We use Long Short-term Memory (LSTM) networks (Hochreiter & Schmidhuber, 1997) for functions Q(Za{1:M}∣X1:M)Q(\mathbf{Z_{a}}_{\{1:M\}}|\mathbf{X}_{1:M}), decoder LSTM networks (Cho et al., 2014) for functions F1:MF_{1:M}, and FCNNs for functions GyG_{y}, Ga{1:M}{G_{a}}_{\{1:M\}} and DD. Details are provided in the appendix and the code is available at https://github.com/pliang279/factorized/.

Experiments

In order to show that MFM learns multimodal representations that are discriminative, generative and interpretable, we design the following experiments. We begin with a multimodal synthetic image dataset that allows us to examine whether MFM displays discriminative and generative capabilities from factorized latent variables. Utilizing image datasets allows us to clearly visualize the generative capabilities of MFM. We then transition to six more challenging real-world multimodal video datasets to 1) rigorously evaluate the discriminative capabilities of MFM in comparison with existing baselines, 2) analyze the importance of each design component through ablation studies, 3) assess the robustness of MFM’s modality reconstruction and prediction capabilities to missing modalities, and 4) interpret the learned representations using information-based and gradient-based methods to understand the contributions of individual factors towards multimodal prediction and generation.

In this section, we study MFM on a synthetic image dataset that considers SVHN (Netzer et al., 2011) and MNIST (Lecun et al., 1998) as the two modalities. SVHN and MNIST are images with different styles but the same labels (digits 0∼90\sim 9). We randomly pair 100,000100,000 SVHN and MNIST images that have the same label, creating a multimodal dataset which we call SVHN+MNIST. 80,00080,000 pairs are used for training and the rest for testing. To justify that MFM is able to learn improved multimodal representations, we show both classification and generation results on SVHN+MNIST in Figure 2.

Prediction: We perform experiments on both unimodal and multimodal classification tasks. UM denotes a unimodal baseline that performs prediction given only one modality as input and MM denotes a multimodal discriminative baseline that performs prediction given both images (Nojavanasghari et al., 2016). We compare the results for UM(SVHN), UM(MNIST), MM and MFM on SVHN+MNIST in Figure 2(b). We achieve better classification performance from unimodal to multimodal which is not surprising since more information is given. More importantly, MFM outperforms MM, which suggests that MFM learns improved factorized representations for discriminative tasks.

Generation: We generate images using the MFM generative network (Figure 2(a)). We fix one variable out of Z=[Za1\mathbf{Z}=[\mathbf{Z_{a}}_{1}, Za2\mathbf{Z_{a}}_{2}, and Zy]\mathbf{Z_{y}}] and randomly sample the other two variables from prior PZP_{\mathbf{Z}}. From Figure 2(c), we observe that MFM shows flexible generation of SVHN and MNIST images based on labels and styles. This suggests that MFM is able to factorize multimodal representations into multimodal discriminative factors (labels) and modality-specific generative factors (styles).

2 Multimodal Time Series Datasets

In this section, we transition to more challenging multimodal time series datasets. All the datasets consist of monologue videos. Features are extracted from the language (GloVe word embeddings (Pennington et al., 2014)), visual (Facet (iMotions, 2017)), and acoustic (COVAREP (Degottex et al., 2014)) modalities. For a detailed description of feature extraction, please refer to the appendix.

We consider the following six datasets across three domains: 1) Multimodal Personality Trait Recognition: POM (Park et al., 2014) contains 903 movie review videos annotated for the following personality traits: confident (con), passionate (pas), voice pleasant (voi), dominant (dom), credible (cre), vivid (viv), expertise (exp), entertaining (ent), reserved (res), trusting (tru), relaxed (rel), outgoing (out), thorough (tho), nervous (ner), persuasive (per) and humorous (hum). The short form is indicated in parenthesis. 2) Multimodal Sentiment Analysis: CMU-MOSI (Zadeh et al., 2016) is a collection of 2199 monologue opinion video clips annotated with sentiment. ICT-MMMO (Wöllmer et al., 2013) consists of 340 online social review videos annotated for sentiment. YouTube (Morency et al., 2011) contains 269 product review and opinion video segments from YouTube each annotated for sentiment. MOUD (Perez-Rosas et al., 2013) consists of 79 product review videos in Spanish. Each video consists of multiple segments labeled as either positive, negative or neutral sentiment. 3) Multimodal Emotion Recognition: IEMOCAP (Busso et al., 2008) consists of 302 videos of recorded dyadic dialogues. The videos are divided into multiple segments each annotated for the presence of 6 discrete emotions (happy, sad, angry, frustrated, excited and neutral), resulting in a total of 7318 segments in the dataset. We report results using the following metrics: Acc_CC = multiclass accuracy across CC classes, F1 = F1 score, MAE = Mean Absolute Error, rr = Pearson’s correlation.

Prediction: We first compare the performance of MFM with existing multimodal prediction methods. For a detailed description of the baselines, please refer to the appendix. From Table 1, we first observe that the best performing baseline results are achieved by different models across different datasets (most notably MFN, MARN, and TFN). On the other hand, MFM consistently achieves state-of-the-art or competitive results for all six multimodal datasets. We believe that the multimodal discriminative factor Fy\mathbf{F_{y}} in MFM has successfully learned more meaningful representations by distilling discriminative features. This highlights the benefit of learning factorized multimodal representations towards discriminative tasks. Furthermore, MFM is model-agnostic and can be applied to other multimodal encoders Q(Zy∣X1:M)Q(\mathbf{Z_{y}}|\mathbf{X}_{1:M}). We perform experiments to show consistent improvements in discriminative performance for several choices of the encoder: EF-LSTM (Morency et al., 2011) and TFN (Zadeh et al., 2017). For Acc_22 on CMU-MOSI, our factorization framework improves the performance of EF-LSTM from 74.374.3 to 75.2 and TFN from 74.674.6 to 75.5.

Ablation Study: In Figure 3, we present the models M{A,B,C,D,E}\mathbf{M_{\{A,B,C,D,E\}}} used for ablation studies. These models are designed to analyze the effects of using a multimodal discriminative factor, a hybrid generative-discriminative objective, factorized generative-discriminative factors and modality-specific generative factors towards both modality reconstruction and label prediction. The simplest variant is MA\mathbf{M_{A}} which represents a purely discriminative model without a joint multimodal discriminative factor (i.e. early fusion (Morency et al., 2011)). MB\mathbf{M_{B}} models a joint multimodal discriminative factor which incorporates more general multimodal fusion encoders (Zadeh et al., 2018a). MC\mathbf{M_{C}} extends MA\mathbf{M_{A}} by optimizing a hybrid generative-discriminative objective over modality-specific factors. MD\mathbf{M_{D}} extends MB\mathbf{M_{B}} by optimizing a hybrid generative-discriminative objective over a joint multimodal factor (resembling Srivastava & Salakhutdinov (2012)). ME\mathbf{M_{E}} factorizes the representation into separate generative and discriminative factors. Finally, MFM is obtained from ME\mathbf{M_{E}} by using modality-specific generative factors instead of a joint multimodal generative factor.

From the table in Figure 3, we observe the following general trends. For sentiment prediction, using 1) a multimodal discriminative factor outperforms modality-specific discriminative factors (MD>MC\mathbf{M_{D}}>\mathbf{M_{C}}, MB>MA\mathbf{M_{B}}>\mathbf{M_{A}}), and 2) adding generative capabilities to the model improves performance (MC>MA\mathbf{M_{C}}>\mathbf{M_{A}}, ME>MB\mathbf{M_{E}}>\mathbf{M_{B}}). For both sentiment prediction and modality reconstruction, 3) factorizing into separate generative and discriminative factors improves performance (ME>MD\mathbf{M_{E}}>\mathbf{M_{D}}), and 4) using modality-specific generative factors outperforms multimodal generative factors (MFM >ME>\mathbf{M_{E}}). These observations support our design decisions of factorizing multimodal representations into multimodal discriminative factors and modality-specific generative factors.

Table 2 shows that MFM with missing modalities outperforms the generative (ΦG\Phi_{G}) or discriminative baselines (ΦD\Phi_{D}) in terms of modality reconstruction and sentiment prediction. Additionally, MFM with missing modalities performs close to MFM with all modalities observed. This fact indicates that MFM can learn representations that are relatively robust to missing modalities. In addition, discriminative performance is most affected when the language modality is missing, which is consistent with prior work which indicates that language is most informative in human multimodal language (Zadeh et al., 2017). On the other hand, sentiment prediction is more robust to missing acoustic and visual features. Finally, we observe that reconstructing the low-level acoustic and visual features is easier as compared to the high-dimensional language features that contain high-level semantic meaning.

Interpretation of Multimodal Representations: We devise two methods to study how individual factors in MFM influence the dynamics of multimodal prediction and generation. These interpretation methods represent both overall trends and fine-grained analysis that could be useful towards deeper understandings of multimodal representation learning. For more details, please refer to the appendix.

Firstly, an information-based interpretation method is chosen to summarize the contribution of each modality towards the multimodal representations. Since Fy\mathbf{F_{y}} is a common cause of X^1:M\hat{\mathbf{X}}_{1:M}, we can compare MI(Fy,X^1),\textnormal{MI}(\mathbf{F_{y}},\hat{\mathbf{X}}_{1}), ⋯ ,MI(Fy,X^M)\cdots,\textnormal{MI}(\mathbf{F_{y}},\hat{\mathbf{X}}_{M}), where MI(⋅,⋅)\textnormal{MI}(\cdot,\cdot) denotes the mutual information measure between Fy\mathbf{F_{y}} and generated modality X^i\hat{\mathbf{X}}_{i}. Higher MI(Fy,X^i)\textnormal{MI}(\mathbf{F_{y}},\hat{\mathbf{X}}_{i}) indicates greater contribution from Fy\mathbf{F_{y}} to X^i\hat{\mathbf{X}}_{i}. Figure 4 reports the ratios ri=MI(Fy,X^i)/MI(Fai,X^i)r_{i}=\textrm{MI}(\mathbf{F_{y}},\hat{\mathbf{X}}_{i})/\textrm{MI}(\mathbf{F_{a}}_{i},\hat{\mathbf{X}}_{i}) which measure a normalized version of the mutual information between Fai\mathbf{F_{a}}_{i} and X^i\hat{\mathbf{X}}_{i}. We observe that on CMU-MOSI, the language modality is most informative towards sentiment prediction, followed by the acoustic modality. We believe that this result represents a prior over the expression of sentiment in human multimodal language and is closely related to the connections between language and speech (Kuhl, 2000).

Secondly, a gradient-based interpretation method to used analyze the contribution of each modality for every time step in multimodal time series data. We measure the gradient of the generated modality with respect to the target factors (e.g., Fy\mathbf{F_{y}}). Let {x1,x2,⋯ ,xM}\{x_{1},x_{2},\cdots,x_{M}\} denote multimodal time series data where xix_{i} represents modality ii, and x^i=[x^i1,⋯ ,x^it,⋯ ,x^iT]\hat{x}_{i}=[\hat{x}_{i}^{1},\cdots,\hat{x}_{i}^{t},\cdots,\hat{x}_{i}^{T}] denote generated modality ii across time steps t∈[1,T]t\in[1,T]. The gradient ∇fy(x^i)\nabla_{f_{y}}(\hat{x}_{i}) measures the extent to which changes in factor fy∼P(Fy∣X1:M=x1:M)f_{y}\sim P(\mathbf{F_{y}}|\mathbf{X}_{1:M}=x_{1:M}) influences the generation of sequence x^i\hat{x}_{i}. Figure 4 plots ∇fy(x^i)\nabla_{f_{y}}(\hat{x}_{i}) for a video in CMU-MOSI. We observe that multimodal communicative behaviors that are indicative of speaker sentiment such as positive words (e.g. “very profound and deep”) and informative acoustic features (e.g. hesitant and emphasized tone of voice) indeed correspond to increases in ∇fy(x^i)\nabla_{f_{y}}(\hat{x}_{i}).

Related Work

The two main pillars of research in multimodal representation learning have considered the discriminative and generative objectives individually. Discriminative representation learning (Liang et al., 2018; Chen et al., 2017; Chaplot et al., 2017; Frome et al., 2013; Socher et al., 2013; Tsai et al., 2017) models the conditional distribution P(Y∣X1:M)P(\mathbf{Y}|\mathbf{X}_{1:M}). Since these approaches are not concerned with modeling P(X1:M)P(\mathbf{X}_{1:M}) explicitly, they use parameters more efficiently to model P(Y∣X1:M)P(\mathbf{Y}|\mathbf{X}_{1:M}). For instance, recent works learn visual representations that are maximally dependent with linguistic attributes for improving one-shot image recognition (Tsai & Salakhutdinov, 2017) or introduce tensor product mechanisms to model interactions between the language, visual and acoustic modalities (Liu et al., 2018; Zadeh et al., 2017). On the other hand, generative representation learning captures the interactions between modalities by modeling the joint distribution P(X1,⋯ ,XM)P(\mathbf{X}_{1},\cdots,\mathbf{X}_{M}) using either undirected graphical models (Srivastava & Salakhutdinov, 2012), directed graphical models (Suzuki et al., 2016b), or neural networks (Sohn et al., 2014). Some generative approaches compress multimodal data into lower-dimensional feature vectors which can be used for discriminative tasks (Pham et al., 2018; Ngiam et al., 2011). To unify the advantages of both approaches, MFM factorizes multimodal representations into generative and discriminative components and optimizes for a joint objective.

Factorized representation learning resembles learning disentangled data representations which have been shown to improve the performance on many tasks (Kulkarni et al., 2015; Lake et al., 2017; Higgins et al., 2016; Bengio et al., 2013). Several methods involve specifying a fixed set of latent attributes that individually control particular variations of data and performing supervised training (Cheung et al., 2014; Karaletsos et al., 2015; Yang et al., 2015; Reed et al., 2014; Zhu et al., 2014), assuming an isotropic Gaussian prior over latent variables to learn disentangled generative representations (Kingma & Welling, 2013; Rubenstein et al., 2018) and learning latent variables in charge of specific variations in the data by maximizing the mutual information between a subset of latent variables and the data (Chen et al., 2016). However, these methods study factorization of a single modality. MFM factorizes multimodal representations and demonstrates the importance of modality-specific and multimodal factors towards generation and prediction. A concurrent and parallel work that factorizes latent factors in multimodal data was proposed by Hsu & Glass (2018). They differ from us in the graphical model design, discriminative objective, prior matching criterion, and scale of experiments. We provide a detailed comparison with their model in the appendix.

Conclusion

In this paper, we proposed the Multimodal Factorization Model (MFM) for multimodal representation learning. MFM factorizes the multimodal representations into two sets of independent factors: multimodal discriminative factors and modality-specific generative factors. The multimodal discriminative factor achieves state-of-the-art or competitive results on six multimodal datasets. The modality-specific generative factors allow us to generate data based on factorized variables, account for missing modalities, and have a deeper understanding of the interactions involved in multimodal learning. Our future work will explore extensions of MFM for video generation, semi-supervised learning, and unsupervised learning. We believe that MFM sheds light on the advantages of learning factorizing multimodal representations and potentially opens up new horizons for multimodal machine learning.

Acknowledgements

This work was supported in part by the DARPA grants D17AP00001 and FA875018C0150, Office of Naval Research, Apple, and Google focused award. We would also like to acknowledge NVIDIA’s GPU support. This material is also based upon work partially supported by the National Science Foundation (Award #1750439). Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of National Science Foundation, and no official endorsement should be inferred.

References

A Proof of Proposition 1

To simplify the proof, we first prove it for the unimodal case by considering the Wasserstein distance between PX,Y{P}_{\mathbf{X},\mathbf{Y}} and PX^,Y^{P}_{\mathbf{\hat{X}},\mathbf{\hat{Y}}}.

For any functions Gy:Zy→FyG_{y}:\mathbf{Z_{y}}\rightarrow\mathbf{F_{y}}, Ga:Za→FaG_{a}:\mathbf{Z_{a}}\rightarrow\mathbf{F_{a}}, D:Fy→Y^D:\mathbf{F_{y}}\rightarrow\mathbf{\hat{Y}}, and F:Fa,Fy→X^F:\mathbf{F_{a},F_{y}}\rightarrow\mathbf{\hat{X}}, we have

where WcW_{c} is the Wasserstein distance under cost function cXc_{X} and cYc_{Y}, PZ{P_{\mathbf{Z}}} is the prior over Z=[Za,Zy]\mathbf{Z}=[\mathbf{Z_{a}},\mathbf{Z_{y}}] and QZ{Q_{\mathbf{Z}}} is the aggregated posterior of the proposed inference distribution Q(Z∣X){Q}({\mathbf{Z}|\mathbf{X}}).

To begin the proof, we abuse some notations as follows.

By definition, the Wasserstein distance under cost function cc between PX,Y{P}_{\mathbf{X,Y}} and PX^,Y^{P}_{\mathbf{\hat{X},\hat{Y}}} is

where c\Big{(}(\mathbf{X},\mathbf{Y}),(\mathbf{\hat{X}},\mathbf{\hat{Y}})\Big{)}:(\mathcal{X},\mathcal{Y})\times(\mathcal{X},\mathcal{Y})\rightarrow\mathcal{R}_{+} is any measurable cost function. \mathcal{P}\Big{(}(\mathbf{X},\mathbf{Y})\sim{P}_{\mathbf{X},\mathbf{Y}},(\mathbf{\hat{X}},\mathbf{\hat{Y}})\sim{P}_{\mathbf{\hat{X}},\mathbf{\hat{Y}}}\Big{)} is the set of all joint distributions of \Big{(}(\mathbf{X},\mathbf{Y}),(\mathbf{\hat{X}},\mathbf{\hat{Y}})\Big{)} with marginals PX,Y{P}_{\mathbf{X,Y}} and PX^,Y^{P}_{\mathbf{\hat{X},\hat{Y}}}, respectively. Note that c\Big{(}(\mathbf{X},\mathbf{Y}),(\mathbf{\hat{X}},\mathbf{\hat{Y}})\Big{)}=c_{X}\Big{(}\mathbf{X},\mathbf{\hat{X}}\Big{)}+c_{Y}\Big{(}\mathbf{Y},\mathbf{\hat{Y}}\Big{)}.

Next, we denote the set of all joint distributions of (X,Y,X^,Y^,Z\mathbf{X},\mathbf{Y},\mathbf{\hat{X}},\mathbf{\hat{Y}},\mathbf{Z}) such that (X,Y)∼PX,Y(\mathbf{X},\mathbf{Y})\sim{P}_{\mathbf{X},\mathbf{Y}}, (X^,Y^,Z)∼PX^,Y^,Z(\mathbf{\hat{X}},\mathbf{\hat{Y}},\mathbf{Z})\sim{P}_{\mathbf{\hat{X}},\mathbf{\hat{Y}},\mathbf{Z}}, and \Big{(}(\mathbf{X},\mathbf{Y})\mathchoice{\mathrel{\hbox to0.0pt{\displaystyle\perp\hss}\mkern 2.0mu{\displaystyle\perp}}}{\mathrel{\hbox to0.0pt{\textstyle\perp\hss}\mkern 2.0mu{\textstyle\perp}}}{\mathrel{\hbox to0.0pt{\scriptstyle\perp\hss}\mkern 2.0mu{\scriptstyle\perp}}}{\mathrel{\hbox to0.0pt{\scriptscriptstyle\perp\hss}\mkern 2.0mu{\scriptscriptstyle\perp}}}(\mathbf{\hat{X}},\mathbf{\hat{Y}})|\mathbf{Z}\Big{)} as PX,Y,X^,Y^,Z\mathcal{P}_{\mathbf{X},\mathbf{Y},\mathbf{\hat{X}},\mathbf{\hat{Y}},\mathbf{Z}}. PX,Y,X^,Y^\mathcal{P}_{\mathbf{X},\mathbf{Y},\mathbf{\hat{X}},\mathbf{\hat{Y}}} and PX,Y,Z\mathcal{P}_{\mathbf{X},\mathbf{Y},\mathbf{Z}} are the sets of the marginals (X,Y,X^,Y^)(\mathbf{X,Y,\hat{X},\hat{Y}}) and (X,Y,Z)(\mathbf{X,Y,Z}) induced by PX,Y,X^,Y^,Z\mathcal{P}_{\mathbf{X},\mathbf{Y},\mathbf{\hat{X}},\mathbf{\hat{Y}},\mathbf{Z}}.

We now introduce two Lemmas to help the proof.

P(X^,Y^∣Z=z)P(\mathbf{\hat{X}},\mathbf{\hat{Y}}|\mathbf{Z}=z) are Dirac for all z∈Zz\in\mathcal{Z}.

Proof: First, we have X^=F(Ga(Za),Gy(Zy))\mathbf{\hat{X}}=F(G_{a}(\mathbf{Z_{a}}),G_{y}(\mathbf{Z_{y}})) and Y^=D(Gy(Zy))\mathbf{\hat{Y}}=D(G_{y}(\mathbf{Z_{y}})) with Z={Za,Zy}\mathbf{Z}=\{\mathbf{Z_{a}},\mathbf{Z_{y}}\}. Since the functions F,Ga,Gy,DF,G_{a},G_{y},D are all deterministic, then P(X^,Y^∣Z)P(\mathbf{\hat{X}},\mathbf{\hat{Y}}|\mathbf{Z}) are Dirac measures. □\square

\mathcal{P}\Big{(}{P}_{\mathbf{X},\mathbf{Y}},{P}_{\mathbf{\hat{X}},\mathbf{\hat{Y}}}\Big{)} = PX,Y,X^,Y^\mathcal{P}_{\mathbf{X},\mathbf{Y},\mathbf{\hat{X}},\mathbf{\hat{Y}}} when P(X^,Y^∣Z=z)P(\mathbf{\hat{X}},\mathbf{\hat{Y}}|\mathbf{Z}=z) are Dirac for all z∈Zz\in\mathcal{Z}.

Proof: When X^,Y^\mathbf{\hat{X}},\mathbf{\hat{Y}} are deterministic functions of Z\mathbf{Z}, for any AA in the sigma-algebra induced by X^,Y^\mathbf{\hat{X}},\mathbf{\hat{Y}}, we have

Therefore, this implies that (\mathbf{X},\mathbf{Y})\mathchoice{\mathrel{\hbox to0.0pt{\displaystyle\perp\hss}\mkern 2.0mu{\displaystyle\perp}}}{\mathrel{\hbox to0.0pt{\textstyle\perp\hss}\mkern 2.0mu{\textstyle\perp}}}{\mathrel{\hbox to0.0pt{\scriptstyle\perp\hss}\mkern 2.0mu{\scriptstyle\perp}}}{\mathrel{\hbox to0.0pt{\scriptscriptstyle\perp\hss}\mkern 2.0mu{\scriptscriptstyle\perp}}}(\mathbf{\hat{X}},\mathbf{\hat{Y}})|\mathbf{Z} which concludes the proof. A similar argument is made in Lemma 1 of (Tolstikhin et al., 2017).

Now, we use the fact that \mathcal{P}\Big{(}{P}_{\mathbf{X},\mathbf{Y}},{P}_{\mathbf{\hat{X}},\mathbf{\hat{Y}}}\Big{)} = PX,Y,X^,Y^\mathcal{P}_{\mathbf{X},\mathbf{Y},\mathbf{\hat{X}},\mathbf{\hat{Y}}} (Lemma 1 + Lemma 2), c\Big{(}(\mathbf{X},\mathbf{Y}),(\mathbf{\hat{X}},\mathbf{\hat{Y}})\Big{)}=c_{X}\Big{(}\mathbf{X},\mathbf{\hat{X}}\Big{)}+c_{Y}\Big{(}\mathbf{Y},\mathbf{\hat{Y}}\Big{)}, \mathbf{\hat{X}}=F\big{(}G_{a}(\mathbf{Z_{a}}),G_{y}(\mathbf{Z_{y}})\big{)}, and \mathbf{\hat{Y}}=D\big{(}G_{y}(\mathbf{Z_{y}})\big{)}, Eq. equation 7 becomes

Note that in Eq. equation 8, \mathcal{P}_{\mathbf{X},\mathbf{Y},\mathbf{Z}}=\mathcal{P}\Big{(}(\mathbf{X},\mathbf{Y})\sim{P}_{\mathbf{X},\mathbf{Y}},\mathbf{Z}\sim{P}_{\mathbf{Z}}\Big{)} and with a proposed Q(Z∣X){Q}(\mathbf{Z|X}), we can rewrite Eq. equation 8 as

A.2 From Unimodal to Multimodal

The proof is similar to Proposition 2, and we present a sketch to it. We can first show P(X^1:M,Y^∣Z=z)P(\mathbf{\hat{X}}_{1:M},\mathbf{\hat{Y}}|\mathbf{Z}=z) are Dirac for all z∈Zz\in\mathcal{Z}. Then we use the fact that c\Big{(}(\mathbf{X}_{1:M},\mathbf{Y}),(\mathbf{\hat{X}}_{1:M},\mathbf{\hat{Y}})\Big{)}=\sum_{i=1}^{M}{c_{X}}_{i}\Big{(}\mathbf{X}_{i},\mathbf{\hat{X}}_{i}\Big{)}+c_{Y}\Big{(}\mathbf{Y},\mathbf{\hat{Y}}\Big{)}. Finally, we follow the tower rule of expectation and the conditional independence property similar to the proof in Proposition 2 and this concludes the proof.

B Full Baseline Models & Results

For a detailed description of the baselines, we point the reader to MFN (Zadeh et al., 2018a), MARN (Zadeh et al., 2018b), TFN (Zadeh et al., 2017), BC-LSTM (Poria et al., 2017), MV-LSTM (Rajagopalan et al., 2016), EF-LSTM (Hochreiter & Schmidhuber, 1997; Graves et al., 2013; Schuster & Paliwal, 1997), DF (Nojavanasghari et al., 2016), MV-HCRF (Song et al., 2012; 2013), EF-HCRF (Quattoni et al., 2007; Morency et al., 2007), THMM (Morency et al., 2011), SVM-MD (Zadeh et al., 2016) and RF (Breiman, 2001).

We use the following extra notations for full descriptions of the baseline models described in Section 3.2, paragraph 3:

Variants of EF-LSTM: EF-LSTM (Early Fusion LSTM) uses a single LSTM (Hochreiter & Schmidhuber, 1997) on concatenated multimodal inputs. We also implement the EF-SLSTM (stacked) (Graves et al., 2013), EF-BLSTM (bidirectional) (Schuster & Paliwal, 1997) and EF-SBLSTM (stacked bidirectional) versions.

Variants of EF-HCRF: EF-HCRF: (Hidden Conditional Random Field) (Quattoni et al., 2007) uses a HCRF to learn a set of latent variables conditioned on the concatenated input at each time step. EF-LDHCRF (Latent Discriminative HCRFs) (Morency et al., 2007) are a class of models that learn hidden states in a CRF using a latent code between observed concatenated input and hidden output. EF-HSSHCRF: (Hierarchical Sequence Summarization HCRF) (Song et al., 2013) is a layered model that uses HCRFs with latent variables to learn hidden spatio-temporal dynamics.

Variants of MV-HCRF: MV-HCRF: Multi-view HCRF (Song et al., 2012) is an extension of the HCRF for Multi-view data, explicitly capturing view-shared and view specific sub-structures. MV-LDHCRF: (Morency et al., 2007) is a variation of the MV-HCRF model that uses LDHCRF instead of HCRF. MV-HSSHCRF: (Song et al., 2013) further extends EF-HSSHCRF by performing Multi-view hierarchical sequence summary representation.

In the following, we provide the full results for all baselines models described in Section 3.2, paragraph 3. Table 3 contains results for multimodal speaker traits recognition on the POM dataset. Table 4 contains results for the multimodal sentiment analysis on the CMU-MOSI, ICT-MMMO, YouTube, and MOUD datasets. Table 5 contains results for multimodal emotion recognition on the IEMOCAP dataset. MFM consistently achieves state-of-the-art or competitive results for all six multimodal datasets. We believe that by our MFM design, the multimodal discriminative factor Fy\mathbf{F_{y}} has successfully learned more meaningful representations by distilling discriminative features. This highlights the benefit of learning factorized multimodal representations towards discriminative tasks.

C Multimodal Features

For each of the multimodal time series datasets as mentioned in Section 3.2, paragraph 3, we extracted the following multimodal features: Language: We use pre-trained word embeddings (glove.840B.300d) (Pennington et al., 2014) to convert the video transcripts into a sequence of 300 dimensional word vectors. Visual: We use Facet (iMotions, 2017) to extract a set of features including per-frame basic and advanced emotions and facial action units as indicators of facial muscle movement (Ekman et al., 1980; Ekman, 1992). Acoustic: We use COVAREP (Degottex et al., 2014) to extract low level acoustic features including 12 Mel-frequency cepstral coefficients (MFCCs), pitch tracking and voiced/unvoiced segmenting features, glottal source parameters, peak slope parameters and maxima dispersion quotients. To reach the same time alignment between different modalities we choose the granularity of the input to be at the level of words. The words are aligned with audio using P2FA (Yuan & Liberman, 2008) to get their exact utterance times. We use expected feature values across the entire word for visual and acoustic features since they are extracted at a higher frequencies.

We make a note that the features for some of these datasets are constantly being updated. The authors of Zadeh et al. (2018a) notified us of a discrepancy in the sampling rate for acoustic feature extraction in the ICT-MMMO, YouTube and MOUD datasets which led to inaccurate word-level alignment between the three modalities. They publicly released the updated multimodal features. We performed all experiments on the latest versions of these datasets which can be accessed from https://github.com/A2Zadeh/CMU-MultimodalSDK. All baseline models were retrained with extensive hyperparameter search for fair comparison.

D Information and Gradient-Based Interpretation

Information-Based Interpretation: We choose the normalized Hilbert-Schmidt Independence Criterion (Gretton et al., 2005; Wu et al., 2018) as the approximation (see Sugiyama & Yamada (2012); Wu et al. (2018)) of our MI measure:

The most common choice for the kernel is the RBF kernel. However, if we consider time series data with various time steps, we need to either perform data augmentation or choose another kernel choice. For example, we can adopt the Global Alignment Kernel (Cuturi et al., 2007) which considers the alignment between two varying-length time series when computing the kernel score between them. To simplify our analysis, we choose to augment data before we calculate the kernel score with the RBF kernel. More specifically, we perform averaging over time series data:

The bandwidth of the RBF kernel is set as 1.0 throughout the experiments.

Here, we provide an additional interpretation result for the POM dataset in Table 6. We observe that the language modality is also the most informative while the visual and acoustic modalities are almost equally informative. This result is in agreement with behavioral studies which have observed that non-verbal behaviors are particularly informative of personality traits (Guimond & Massrieh, 2012; Levine et al., 2009; Mohammadi et al., 2010). For example, the same sentence “this movie was great” can convey significantly different messages on speaker confidence depending on whether it was said in a loud and exciting voice, with eye contact, or powerful gesticulation.

Gradient-Based Interpretation: MFM reconstructs xix_{i} as follows:

Equation equation 12 also explains how we obtain fy∼P(Fy∣X1:M=x1:M)f_{y}\sim P(\mathbf{F_{y}}|\mathbf{X}_{1:M}=x_{1:M}). The gradient flow through time is defined as:

E Encoder and Decoder Design for Multimodal Synthetic Image Dataset

For experiments on the multimodal synthetic image dataset, we use convolutional+fully-connected layers for the encoder and deconvolutional+fully-connected layers for the decoder (Zeiler et al., 2010). Different convolutional layers are each applied on the input SVHN and MNIST images to learn modality-specific generative factors. Next, we concatenate the features from two more convolutional layers on SVHN and MNIST to learn the multimodal-discriminative factor. The multimodal discriminative factor is passed through fully-connected layers to predict the label. For generation, we concatenate the multimodal discriminative factors and the modality-specific generative factor together and use a deconvolutional layer to generate digits.

F Encoder and Decoder Design for Multimodal Time Series Datasets

Figure 5 illustrates how MFM operates on multimodal time series data. The encoder Q(Zy∣X1:M)Q(\mathbf{Z_{y}}|\mathbf{X}_{1:M}) can be parametrized by any model that performs multimodal fusion (Nojavanasghari et al., 2016; Zadeh et al., 2018a). We choose the Memory Fusion Network (MFN) (Zadeh et al., 2018a) as our encoder Q(Zy∣X1:M)Q(\mathbf{Z_{y}}|\mathbf{X}_{1:M}). We use encoder LSTM networks and decoder LSTM networks (Cho et al., 2014) to parametrize functions Q(Za1:M∣X1:M)Q(\mathbf{Z_{a}}_{1:M}|\mathbf{X}_{1:M}) and F1:MF_{1:M} respectively, and FCNNs to parametrize functions GyG_{y}, Ga{1:M}{G_{a}}_{\{1:M\}} and DD.

G Surrogate Inference Graphical Model

We illustrate the surrogate inference for addressing the missing modalities issue in Figure 6. The surrogate inference model infers the latent codes given the present modalities. These inferred latent codes can then be used for reconstructing the missing modalities or label prediction in the presence of missing modalities.

H Comparison with Hsu & Glass (2018)

A similar approach for factorizing the latent factors was recently proposed by Hsu & Glass (2018) in work that was performed independently and in parallel. In comparison with MFM, there are several major differences that can be categorized into the (1) prior matching discrepancy, (2) inference network, (3) discriminative objective, (4) multimodal fusion, (5) scale of experiments.

MFM uses MMD(QZ,PZ)\mathcal{MMD}(Q_{\mathbf{Z}},P_{\mathbf{Z}}) (see Equation 4) as the prior matching discrepancy while Hsu & Glass (2018) use KL(QZ∣X,PZ)\mathcal{KL}(Q_{\mathbf{Z|X}},P_{\mathbf{Z}})) (see Section 2.2.1 in Hsu & Glass (2018)).

MFM considers multimodal and unimodal inference in a single network (see Figure 1(b)), while Hsu & Glass (2018) considers separate networks (see Figure 1 in Hsu & Glass (2018)). They further propose to match the coherence between these two networks using an additional loss term (see Equation 7 in Hsu & Glass (2018)).

MFM learns to predict the labels using a generative framework (see Figure 1(a)), while Hsu & Glass (2018) use an additional hinge loss to separate the latent factors from different labels (see Equation 9 in Hsu & Glass (2018)).

MFM is a flexible framework that can be combined with any multimodal fusion encoder (see Section 2.4), while Hsu & Glass (2018) considers a fixed multimodal encoder (similar to early fusion) (see Section 4.1 in Hsu & Glass (2018)).

We evaluate the performance of MFM over a much larger scale of datasets. We perform experiments on six multimodal time-series datasets that take on the form of videos with the language, visual, and acoustic modalities. These datasets span three core research areas of multimodal personality traits recognition, multimodal sentiment analysis, and multimodal emotion recognition. On the other hand, Hsu & Glass (2018) evaluates their model on a spoken digit dataset which randomly combines a digit image with a spoken digit (see Section 4 in Hsu & Glass (2018)). MFM further considers experiments to evaluate reconstruction and prediction in the presence of missing modalities (see Section 3.2) which Hsu & Glass (2018) do not. Lastly, we compares to over 20 baseline models in our experiments (see Section 3.2) and explore the choice of various multimodal encoders in MFM. Hsu & Glass (2018) only compares to the JMVAE baseline model (Suzuki et al., 2016a) which resembles the MD\mathbf{M_{D}} model in our ablation study (see Section 4.4 in Hsu & Glass (2018)).

In Table 7, we provide a comparison on the CMU-MOSI, ICT-MMMO, YouTube and MOUD datasets to test the disentanglement and prediction performance for the model described in Hsu & Glass (2018). These experimental results show that across these datasets and metrics, MFM performs better than the model proposed in Hsu & Glass (2018). We would like to highlight that at the time of submission, the code for (Hsu & Glass, 2018) had not been made public and we reimplemented their model to experiment on our datasets.

I Comparison with β𝛽\beta-VAE

Although β\beta-VAE (Higgins et al., 2017) was designed to handle unimodal data, we provide an extension to multimodal data. To achieve this, we set the choice of prior matching discrepancy as the KL-divergence KL(QZ∣X,PZ)\mathcal{KL}(Q_{\mathbf{Z|X}},P_{\mathbf{Z}}) and set β\beta large (i.e. β∈{10,50,100,200}\beta\in\{10,50,100,200\}) to encourage disentanglement of latent variables. We train a β\beta-VAE to model multimodal data using the same factorization as proposed in our model (i.e. modality-specific generative factors Za{1:M}\mathbf{Z_{a}}_{\{1:M\}} and a multimodal discriminative factor Zy\mathbf{Z_{y}}). To provide a fair comparison to our discriminative model, we fine tune by training a classifier on top of the multimodal discriminative factor Zy\mathbf{Z_{y}} to the label Y\mathbf{Y}. We provide experimental results in Table 8 on the CMU-MOSI, ICT-MMMO, YouTube and MOUD datasets. MFM outperforms β\beta-VAE across these datasets and metrics.