Memory Fusion Network for Multi-view Sequential Learning
Amir Zadeh, Paul Pu Liang, Navonil Mazumder, Soujanya Poria, Erik Cambria, Louis-Philippe Morency
Introduction
In many natural scenarios, data is collected from diverse perspectives and exhibits heterogeneous properties: each of these domains present a different view of the same data, where each view can have its own individual representation space and dynamics. Such forms of data are known as multi-view data. In a multi-view setting, each view of the data may contain some knowledge that other views do not have access to. Therefore, multiple views must be employed together in order to describe the data comprehensively and accurately. Multi-view learning has been an active area of machine learning research [\citeauthoryearXu, Tao, and Xu2013]. By exploring the consistency and complementary properties of different views, multi-view learning can be more effective, more promising, and has better generalization ability than single-view learning.
Multi-view sequential learning extends the definition of multi-view learning to manage with different views all in the form of sequential data, i.e. data that comes in the form of sequences. For example, a video clip of an orator can be partitioned into three sequential views – text representing the spoken words, video of the speaker, and vocal prosodic cues from the audio. In multi-view sequential learning, two primary forms of interactions exist. The first form is called view-specific interactions; interactions that involve only one view. For example, learning the sentiment of a speaker based only on the sequence of spoken works. More importantly, the second form of interactions are defined across different views. These are known as cross-view interactions. Cross-view interactions span across both the different views and time – for example a listener’s backchannel response or the delayed rumble of distant lightning in the video and audio views. Modeling both the view-specific and cross-view interactions lies at the core of multi-view sequential learning.
This paper introduces a novel neural model for multi-view sequential learning called the Memory Fusion Network (MFN). At a first layer, the MFN encodes each view independently using a component called the System of Long Short Term Memories (LSTMs). In this System of LSTMs, each view is assigned one LSTM function to model the dynamics in that particular view. The second component of MFN is called the Delta-memory Attention Network (DMAN) which finds cross-view interactions across memories of the System of LSTMs. Specifically, the DMAN identifies the cross-view interactions by associating a relevance score to the memory dimensions of each LSTM. The third component of the MFN stores the cross-view information over time in the Multi-view Gated Memory. This memory updates its contents based on the outputs of the DMAN and its previously stored contents, acting as a dynamic memory module for learning crucial cross-view interactions throughout the sequential data. Prediction is performed by integrating both view-specific and cross-view and information.
We perform extensive experimentation to benchmark the performance of MFN on 6 publicly available multi-view sequential datasets. Throughout, we compare to the state-of-the-art approaches in multi-view sequential learning. In all the benchmarks, MFN is able to outperform the baselines, setting new state-of-the-art results across all the datasets.
Related Work
Researchers dealing with multi-view sequential data have largely focused on three major types of models.
The first category of models have relied on concatenation of all multiple views into a single view to simplify the learning setting. These approaches then use this concatenated view as input to a learning model. Hidden Markov Models (HMMs) [\citeauthoryearBaum and Petrie1966, \citeauthoryearMorency, Mihalcea, and Doshi2011], Support Vector Machines (SVMs) [\citeauthoryearCortes and Vapnik1995], Hidden Conditional Random Fields (HCRFs) [\citeauthoryearQuattoni et al.2007] and their variants [\citeauthoryearMorency, Quattoni, and Darrell2007] have been successfully used for structured prediction. More recently, with the advent of deep learning, Recurrent Neural Networks, specially Long-short Term Memory (LSTM) networks [\citeauthoryearHochreiter and Schmidhuber1997], have been extensively used for sequence modeling. Some degree of success for modeling multi-view problems is achieved using this concatenation. However, this concatenation causes over-fitting in the case of a small size training sample and is not intuitively meaningful because each view has a specific statistical property [\citeauthoryearXu, Tao, and Xu2013] which is ignored in these simplified approaches.
The second category of models introduce multi-view variants to the structured learning approaches of the first category. Multi-view variations of these models have been proposed including Multi-view HCRFs where the potentials of the HCRF are changed to facilitate multiple views [\citeauthoryearSong, Morency, and Davis2012, \citeauthoryearSong, Morency, and Davis2013]. Recently, multi-view LSTM models have been proposed for multimodal setups where the LSTM memory is partitioned into different components for different views [\citeauthoryearRajagopalan et al.2016].
The third category of models rely on collapsing the time dimension from sequences by learning a temporal representation for each of the different views. Such methods have used average feature values over time [\citeauthoryearPoria, Cambria, and Gelbukh2015]. Essentially these models apply conventional multi-view learning approaches, such as Multiple Kernel Learning [\citeauthoryearPoria, Cambria, and Gelbukh2015], subspace learning or co-training [\citeauthoryearXu, Tao, and Xu2013] to the multi-view representations. Other approaches have trained different models for each view and combined the models using decision voting [\citeauthoryearNojavanasghari et al.2016], tensor products [\citeauthoryearZadeh et al.2017] or deep neural networks [\citeauthoryearPoria et al.2017]. While these approaches are able to learn the relations between the views to some extent, the lack of the temporal dimension limits these learned representations, eventually affect their performance. Such is the case for long sequences where the learned representations do not sufficiently reflect all the temporal information in each view.
The proposed model in this paper is different from the first category models since it assigns one LSTM to each view instead of concatenating the information from different views. MFN is also different from the second category models since it considers each view in isolation to learn view-specific interactions. It then uses an explicitly designed attention mechanism and memory to find and store cross-view interactions over time. MFN is different from the third category models since view-specific and cross-view interactions are modeled over time.
Memory Fusion Network (MFN)
The Memory Fusion Network (MFN) is a recurrent model for multi-view sequential learning that consists of three main components: 1) System of LSTMs consists of multiple Long-short Term Memory (LSTM) networks, one for each of the views. Each LSTM encodes the view-specific dynamics and interactions. 2) Delta-memory Attention Network is a special attention mechanism designed to discover both cross-view and temporal interactions across different dimensions of memories in the System of LSTMs. 3) Multi-view Gated Memory is a unifying memory that stores the cross-view interactions over time. Figure 1 shows the overview of MFN pipeline and its components.
Delta-memory Attention Network
The goal of the Delta-memory Attention Network (DMAN) is to outline the cross-view interactions at timestep between different view memories in the System of LSTMs. To this end, we use a coefficient assignment technique on the concatenation of LSTM memories at time . High coefficients are assigned to the dimensions jointly form a cross-view interaction and low coefficients to the other dimensions. However, coefficient assignment using only memories at time is not ideal since the same cross-view interactions can happen over multiple time instances if the LSTM memories in those dimensions remain unchanged. This is especially troublesome if the recurring dimensions are assigned high coefficients, in which case they will dominate the coefficient assignment system. To deal with this problem we add the memories of time so DMAN can have the freedom of leaving unchanged dimensions in the System of LSTMs memories and only assign high coefficient to them if they are about to change. Ideally each cross-view interaction is only assigned high coefficients once before the state of memories in System of LSTMs changes. This can be done by comparing the memories at the two time-steps (hence the name Delta-memory).
are softmax activated scores for each LSTM memory at time and . Applying softmax at the output layer of allows for regularizing high-value coefficients over the . The output of the DMAN is defined as:
is the attended memories of the LSTMs. Applying this element-wise product amplifies the relevant dimensions of the while marginalizing the effect of remaining dimensions. DMAN is also able to find cross-view interactions that do not happen simultaneously since it attends to the memories in the System of LSTMs. These memories can carry information about the observed inputs across different timestamps.
Multi-view Gated Memory
This update proposes changes to Multi-view Gated Memory based on observations about cross-view interactions at time .
At each time-step of MFN recursion, is updated using retain and update gates, and , as well as the current cross-view update proposal with the following formulation:
is activated using squashing function to improve model stability by avoiding drastic changes to the Multi-view Gated Memory. The Multi-view Gated Memory is different from LSTM memory in two ways. Firstly, the Multi-view Gated Memory has a more complex gating mechanism: both gates are controlled by neural networks while LSTM gates are controlled by a non-linear affine transformation. As a result, the Multi-view Gated Memory has superior representation capabilities as compared to the LSTM memory. Secondly, the value of the Multi-view Gated Memory does not go through a sigmoid activation in each iteration. We found that this helps in faster convergence.
Output of MFN
The outputs of the MFN are the final state of the Multi-view Gated Memory and the outputs of each of the LSTMs:
representing individual sequence information. denotes vector concatenation.
Experimental Setup
In this section we design extensive experiments to evaluate the performance of MFN. We choose three multi-view domains: multimodal sentiment analysis, emotion recognition and speaker traits analysis. All benchmarks involve three views with completely different natures: language (text), vision (video), and acoustic (audio). The multi-view input signal is the video of a person speaking about a certain topic. Since humans communicate their intentions in a structured manner, there are synchronizations between intentions in text, gestures and tone of speech. These synchronizations constitute the relations between the three views.
In all the videos in the datasets described below, only one speaker is present in front of the camera.
Sentiment Analysis The first domain in our experiments is multimodal sentiment analysis, where the goal is to identify a speaker’s sentiment based on online video content. Multimodal sentiment analysis extends the conventional text-based definition of sentiment analysis to a multimodal setup where different views contribute to modeling the sentiment of the speaker. We use four different datasets for English and Spanish sentiment analysis in our experiments. The CMU-MOSI dataset [\citeauthoryearZadeh et al.2016] is a collection of 93 opinion videos from online sharing websites. Each video consists of multiple opinion segments and each segment is annotated with sentiment in the range . The MOUD dataset [\citeauthoryearPerez-Rosas, Mihalcea, and Morency2013] consists of product review videos in Spanish. Each video consists of multiple segments labeled to display positive, negative or neutral sentiment. To maintain consistency with previous works [\citeauthoryearPoria et al.2017, \citeauthoryearPerez-Rosas, Mihalcea, and Morency2013] we remove segments with the neutral label. The YouTube dataset [\citeauthoryearMorency, Mihalcea, and Doshi2011] introduced tri-modal sentiment analysis to the research community. Multi-dimensional data from the audio, visual and textual modalities are collected in the form of 47 videos from the social media web site YouTube. The collected videos span a wide range of product reviews and opinion videos. These are annotated at the segment level for sentiment. The ICT-MMMO dataset [\citeauthoryearWöllmer et al.2013] consists of online social review videos that encompass a strong diversity in how people express opinions, annotated at the video level for sentiment.
Emotion Recognition The second domain in our experiments is multimodal emotion recognition, where the goal is to identify a speakers emotions based on the speakers verbal and nonverbal behaviors. These emotions are categorized as basic emotions [\citeauthoryearEkman1992] and continuous emotions [\citeauthoryearGunes2010]. We perform experiments on IEMOCAP dataset [\citeauthoryearBusso et al.2008]. IEMOCAP consists of 151 sessions of recorded dialogues, of which there are 2 speakers per session for a total of 302 videos across the dataset. Each segment is annotated for the presence of emotions (angry, excited, fear, sad, surprised, frustrated, happy, disappointed and neutral) as well as valence, arousal and dominance.
Speaker Traits Analysis The third domain in our experiments is speaker trait recognition based on communicative behavior of the speaker. The goal is to identify 16 different speaker traits. The POM dataset [\citeauthoryearPark et al.2014] contains 1,000 movie review videos. Each video is annotated for various personality and speaker traits, specifically: confident (con), passionate (pas), voice pleasant (voi), dominant (dom), credible (cre), vivid (viv), expertise (exp), entertaining (ent), reserved (res), trusting (tru), relaxed (rel), outgoing (out), thorough (tho), nervous (ner), persuasive (per) and humorous (hum). The short form of these speaker traits is indicated inside parentheses and used for the rest of this paper.
Sequence Features
The chosen system of sequences are the three modalities: language, visual and acoustic. To get the exact utterance time-stamp of each word we perform forced alignment using P2FA [\citeauthoryearYuan and Liberman2008] which allows us to align the three modalities together. Since words are considered the basic units of language we use the interval duration of each word utterance as a time-step. We calculate the expected video and audio features by taking the expectation of their view feature values over the word utterance time interval [\citeauthoryearZadeh et al.2017]. For each of the three modalities, we process the information from videos as follows.
Language View For the language view, Glove word embeddings [\citeauthoryearPennington, Socher, and Manning2014] were used to embed a sequence of individual words from video segment transcripts into a sequence of word vectors that represent spoken text. The Glove embeddings used are 300 dimensional word embeddings trained on 840 billion tokens from the common crawl dataset, resulting in a sequence of dimension after alignment. The timing of word utterances is extracted using P2FA forced aligner. This extraction enables alignment between text, audio and video.
Visual View For the visual view, the library Facet [\citeauthoryeariMotions2017] is used to extract a set of visual features including facial action units, facial landmarks, head pose, gaze tracking and HOG features [\citeauthoryearZhu et al.2006]. These visual features are extracted from the full video segment at 30Hz to form a sequence of facial gesture measures throughout time, resulting in a sequence of dimension .
Acoustic View For the audio view, the software COVAREP [\citeauthoryearDegottex et al.2014] is used to extract acoustic features including 12 Mel-frequency cepstral coefficients, pitch tracking and voiced/unvoiced segmenting features [\citeauthoryearDrugman and Alwan2011], glottal source parameters [\citeauthoryearChilders and Lee1991, \citeauthoryearDrugman et al.2012, \citeauthoryearAlku1992, \citeauthoryearAlku, Strik, and Vilkman1997, \citeauthoryearAlku, Bäckström, and Vilkman2002], peak slope parameters and maxima dispersion quotients [\citeauthoryearKane and Gobl2013]. These visual features are extracted from the full audio clip of each segment at 100Hz to form a sequence that represent variations in tone of voice over an audio segment, resulting in a sequence of dimension after alignment.
Experimental Details
The time steps in the sequences are chosen based on word utterances. The expected (average) visual and acoustic sequences features are calculated for each word utterance to ensure time alignment between all LSTMs. In all the aforementioned datasets, it is important that the same speaker does not appear in both train and test sets in order to evaluate the generalization of our approach. The training, validation and testing splits are performed so that the splits are speaker independent. The full set of videos (and segments for datasets where the annotations are at the resolution of segments) in each split is detailed in Table 1. All baselines were re-trained using these video-level train-test splits of each dataset and with the same set of extracted sequence features. Training is performed on the labeled segments for datasets annotated at the segment level and on the labeled videos otherwise. All the code and data required to recreate the reported results are available at https://github.com/A2Zadeh/MFN.
Baseline Models
We compare the performance of the MFN with current state-of-the-art models for multi-view sequential learning. To perform a more extensive comparison we train all the following baselines across all the datasets. Due to space constraints, each baseline name is denoted by a symbol (in parenthesis) which is used in Table 2 to refer to specific baseline results.
Song2013 (): This is a layered model that uses CRFs with latent variables to learn hidden spatio-temporal dynamics. For each layer an abstract feature representation is learned through non-linear gate functions. This procedure is repeated to obtain a hierarchical sequence summary (HSS) representation [\citeauthoryearSong, Morency, and Davis2013].
Morency2011 (): Hidden Markov Model is a statistical Markov model in which the system being modeled is assumed to be a Markov process with unobserved (i.e. hidden) states [\citeauthoryearBaum and Petrie1966]. We follow the implementation in [\citeauthoryearMorency, Mihalcea, and Doshi2011] for tri-modal data.
Quattoni2007 (): Concatenated features are used as input to a Hidden Conditional Random Field (HCRF) [\citeauthoryearQuattoni et al.2007]. HCRF learns a set of latent variables conditioned on the concatenated input at each time step.
Morency2007 (): Latent Discriminative Hidden Conditional Random Fields (LDHCRFs) are a class of models that learn hidden states in a Conditional Random Field using a latent code between observed input and hidden output [\citeauthoryearMorency, Quattoni, and Darrell2007].
Hochreiter1997 (): A LSTM with concatenation of data from different views as input [\citeauthoryearHochreiter and Schmidhuber1997]. Stacked, bidirectional and stacked bidirectional LSTMs are also trained in a similar fashion for stronger baselines.
Multi-view Sequential Learning Models
Rajagopalan2016 (): Multi-view (MV) LSTM [\citeauthoryearRajagopalan et al.2016] aims to extract information from multiple sequences by modeling sequence-specific and cross-sequence interactions over time and output. It is a strong tool for synchronizing a system of multi-dimensional data sequences.
Song2012 (): MV-HCRF [\citeauthoryearSong, Morency, and Davis2012] is an extension of the HCRF for Multi-view data. Instead of view concatenation, view-shared and view specific sub-structures are explicitly learned to capture the interaction between views. We also implement the topological variations - linked, coupled and linked-couple that differ in the types of interactions between the modeled views. Song2012LD (): is a variation of this model that uses LDHCRF instead of HCRF.
Song2013MV (): MV-HSSHCRF is an extension of Song2013 that performs Multi-view hierarchical sequence summary representation.
Dataset Specific Baselines
Poria2015 (): Multiple Kernel Learning [\citeauthoryearBach, Lanckriet, and Jordan2004] classifiers have been widely applied to problems involving multi-view data. Our implementation follows a previously proposed model for multimodal sentiment analysis [\citeauthoryearPoria, Cambria, and Gelbukh2015].
Nojavanasghari2016 (): Deep Fusion Approach [\citeauthoryearNojavanasghari et al.2016] trains single neural networks for each view’s input and combine the views with a joint neural network. This baseline is current state of the art in POM dataset.
Zadeh2016 (): Support Vector Machine [\citeauthoryearCortes and Vapnik1995] is a widely used classifier. This baseline is closely implemented similar to a previous work in multimodal sentiment analysis [\citeauthoryearZadeh et al.2016].
Ho1998 (): We also compare to a Random Forest [\citeauthoryearHo1998] baseline as another strong non-neural classifier.
Dataset Specific State-of-the-art Baselines
Poria2017 (): Bidirectional Contextual LSTM [\citeauthoryearPoria et al.2017] performs context-dependent fusion of multi-sequence data that holds the state of the art for emotion recognition on IEMOCAP dataset and sentiment analysis on MOUD dataset.
Zadeh2017 (): Tensor Fusion Network [\citeauthoryearZadeh et al.2017] learns explicit uni-view, bi-view and tri-view concepts in multi-view data. It is the current state of the art for sentiment analysis on CMU-MOSI dataset.
Wang2016 (): Selective Additive Learning Convolutional Neural Network [\citeauthoryearWang et al.2016] is a multimodal sentiment analysis model that attempts to prevent identity-dependent information from being learned so as to improve generalization based only on accurate indicators of sentiment.
MFN Ablation Study Baselines
MFN : These baselines use only individual views – for language, for visual, and for acoustic. The DMAN and Multi-view Gated Memory are also removed since only one view is present. This effectively reduces the MFN to one single LSTM which uses input from one view.
MFN (no ): This variation of our model shrinks the context to only the current timestamp in the DMAN. We compare to this model to show the importance of having the memory temporal information – memories at both time and .
MFN (no mem): This variation of our model removes the Delta-memory Attention Network and Multi-view Gated Memory from the MFN. Essentially this is equivalent to three disjoint LSTMs. The output of the MFN in this case would only be the outputs of LSTM at the final timestamp . This baseline is designed to evaluate the importance of spatio-temporal relations between views through time.
MFN Results and Discussion
The comparison between MFN and MFN (no ) indicates the crucial role of the memories of time . The comparison between MFN and MFN (no mem) shows the essential role of the Multi-view Gated Memory. The final observation comes from comparing all multi-view variations of MFN with single view MFN . This indicates that using multiple views results in better performance even if various crucial components are removed from MFN. Increasing The DMAN Input Region Size: In our set of experiments increasing the to cover instead of did not significantly improve the performance of the model. We argue that this is because additional memory steps do not add any information to the DMAN internal mechanism.
Conclusion
This paper introduced a novel approach for multi-view sequential learning called Memory Fusion Network (MFN). The first component of MFN is called System of LSTMs. In System of LSTMs, each view is assigned one LSTM function to model the interactions within the view. The second component of MFN is called Delta-memory Attention Network (DMAN). DMAN outlines the relations between views through time by associating a cross-view relevance score to the memory dimensions of each LSTM. The third component of the MFN unifies the sequences and is called Multi-view Gated Memory. This memory updates its content based on the outputs of DMAN calculated over memories in System of LSTMs. Through extensive experimentation on multiple publicly available datasets, the performance of MFN is compared with various baselines. MFN shows state-of-the-art performance in multi-view sequential learning on all the datasets.
Acknowledgements
This project was partially supported by Oculus research grant. We thank the reviewers for their valuable feedback.