On Multiplicative Integration with Recurrent Neural Networks
Yuhuai Wu, Saizheng Zhang, Ying Zhang, Yoshua Bengio, Ruslan Salakhutdinov
Introduction
Recently there has been a resurgence of new structural designs for recurrent neural networks (RNNs) . Most of these designs are derived from popular structures including vanilla RNNs, Long Short Term Memory networks (LSTMs) and Gated Recurrent Units (GRUs) . Despite of their varying characteristics, most of them share a common computational building block, described by the following equation:
In this work, we propose an alternative design for constructing the computational building block by changing the procedure of information integration. Specifically, instead of utilizing sum operation “+", we propose to use the Hadamard product “” to fuse and :
The result of this modification changes the RNN from first order to second order , while introducing no extra parameters. We call this kind of information integration design a form of Multiplicative Integration. The effect of multiplication naturally results in a gating type structure, in which and are the gates of each other. More specifically, one can think of the state-to-state computation (where for example represents the previous state) as dynamically rescaled by (where for example represents the input). Such rescaling does not exist in the additive building block, in which is independent of . This relatively simple modification brings about advantages over the additive building block as it alters RNN’s gradient properties, which we discuss in detail in the next section, as well as verify through extensive experiments.
In the following sections, we first introduce a general formulation of Multiplicative Integration. We then compare it to the additive building block on several sequence learning tasks, including character level language modelling, speech recognition, large scale sentence representation learning using a Skip-Thought model, and teaching a machine to read and comprehend for a question answering task. The experimental results (together with several existing state-of-the-art models) show that various RNN structures (including vanilla RNNs, LSTMs, and GRUs) equipped with Multiplicative Integration provide better generalization and easier optimization. Its main advantages include: (1) it enjoys better gradient properties due to the gating effect. Most of the hidden units are non-saturated; (2) the general formulation of Multiplicative Integration naturally includes the regular additive building block as a special case, and introduces almost no extra parameters compared to the additive building block; and (3) it is a drop-in replacement for the additive building block in most of the popular RNN models, including LSTMs and GRUs. It can also be combined with other RNN training techniques such as Recurrent Batch Normalization . We further discuss its relationship to existing models, including Hidden Markov Models (HMMs) , second order RNNs and Multiplicative RNNs .
Structure Description and Analysis
The key idea behind Multiplicative Integration is to integrate different information flows and , by the Hadamard product “”. A more general formulation of Multiplicative Integration includes two more bias vectors and added to and :
Note that the number of parameters of the Multiplicative Integration is about the same as that of the additive building block, since the number of new parameters (, and ) are negligible compared to total number of parameters. Also, Multiplicative Integration can be easily extended to LSTMs and GRUsSee exact formulations in the Appendix., that adopt vanilla building blocks for computing gates and output states, where one can directly replace them with the Multiplicative Integration. More generally, in any kind of structure where information flows () are involved (e.g. residual networks ), one can implement pairwise Multiplicative Integration for integrating all information sources.
2 Gradient Properties
The Multiplicative Integration has different gradient properties compared to the additive building block. For clarity of presentation, we first look at vanilla-RNN and RNN with Multiplicative Integration embedded, referred to as MI-RNN. That is, versus . In a vanilla-RNN, the gradient can be computed as follows:
where . By looking at the gradient, we see that the matrix and the current input is directly involved in the gradient computation by gating the matrix , hence more capable of altering the updates of the learning system. As we show in our experiments, with directly gating the gradient, the vanishing/exploding problem is alleviated: dynamically reconciles , making the gradient propagation easier compared to the regular RNNs. For LSTMs and GRUs with Multiplicative Integration, the gradient propagation properties are more complicated. But in principle, the benefits of the gating effect also persists in these models.
Experiments
In all of our experiments, we use the general form of Multiplicative Integration (Eq. 4) for any hidden state/gate computations, unless otherwise specified.
To further understand the functionality of Multiplicative Integration, we take a simple RNN for illustration, and perform several exploratory experiments on the character level language modeling task using Penn-Treebank dataset , following the data partition in . The length of the training sequence is 50. All models have a single hidden layer of size 2048, and we use Adam optimization algorithm with learning rate . Weights are initialized to samples drawn from uniform. Performance is evaluated by the bits-per-character (BPC) metric, which is of perplexity.
1.2 Scaling Problem
When adding two numbers at different order of magnitude, the smaller one might be negligible for the sum. However, when multiplying two numbers, the value of the product depends on both regardless of the scales. This principle also applies when comparing Multiplicative Integration to the additive building blocks. In this experiment, we test whether Multiplicative Integration is more robust to the scales of weight values. Following the same models as in Section 3.1.1, we first calculated the norms of and for both vanilla-RNN and MI-RNN for different after training. We found that in both structures, is a lot smaller than in magnitude. This might be due to the fact that is a one-hot vector, making the number of updates for (columns of) be smaller than . As a result, in vanilla-RNN, the pre-activation term is largely controlled by the value of , while becomes rather small. In MI-RNN, on the other hand, the pre-activation term still depends on the values of both and , due to multiplication.
We next tried different initialization of and to test their sensitivities to the scaling. For each model, we fix the initialization of to uniform and initialize to uniform where varies in . Table 1, top left panel, shows results. As we increase the scale of , performance of the vanilla-RNN improves, suggesting that the model is able to better utilize the input information. On the other hand, MI-RNN is much more robust to different initializations, where the scaling has almost no effect on the final performance.
1.3 On different choices of the formulation
In our third experiment, we evaluated the performance of different computational building blocks, which are Eq. 1 (vanilla-RNN), Eq. 2 (MI-RNN-simple) and Eq. 4 (MI-RNN-general)We perform hyper-parameter search for the initialization of in MI-RNN-general.. From the validation curves in Figure 1 (b), we see that both MI-RNN, simple and MI-RNN-general yield much better performance compared to vanilla-RNN, and MI-RNN-general has a faster convergence speed compared to MI-RNN-simple. We also compared our results to the previously published models in Table 1, bottom left panel, where MI-RNN-general achieves a test BPC of 1.39, which is to our knowledge the best result for RNNs on this task without complex gating/cell mechanisms.
2 Character Level Language Modeling
In addition to the Penn-Treebank dataset, we also perform character level language modeling on two larger datasets: text8http://mattmahoney.net/dc/textdata and Hutter Challenge Wikipediahttp://prize.hutter1.net/. Both of them contain 100M characters from Wikipedia while text8 has an alphabet size of 27 and Hutter Challenge Wikipedia has an alphabet size of 205. For both datasets, we follow the training protocols in and respectively. We use Adam for optimization with the starting learning rate grid-searched in . If the validation BPC (bits-per-character) does not decrease for 2 epochs, we half the learning rate.
We implemented Multiplicative Integration on both vanilla-RNN and LSTM, referred to as MI-RNN and MI-LSTM. The results for the dataset are shown in Table 1, bottom middle panel. All five models, including some of the previously published models, have the same number of parameters (4M). For RNNs without complex gating/cell mechanisms (the first three results), our MI-RNN (with initialized as ) performs the best, our MI-LSTM (with initialized as ) outperforms all other models by a large margin reports better results but they use much larger models (16M) which is not directly comparable..
On Hutter Challenge Wikipedia dataset, we compare our MI-LSTM (single layer with 2048 unit, 17M, with initialized as ) to the previous stacked LSTM (7 layers, 27M) , GF-LSTM (5 layers, 20M) , and grid-LSTM (6 layers, 17M) . Table 1, bottom right panel, shows results. Despite the simple structure compared to the sophisticated connection designs in GF-LSTM and grid-LSTM, our MI-LSTM outperforms all other models and achieves the new state-of-the-art on this task.
3 Speech Recognition
We next evaluate our models on Wall Street Journal (WSJ) corpus (available as LDC corpus LDC93S6B and LDC94S13B), where we use the full 81 hour set “si284” for training, set “dev93” for validation and set “eval92” for test. We follow the same data preparation process and model setting as in , and we use 59 characters as the targets for the acoustic modelling. Decoding is done with the CTC based weighted finite-state transducers (WFSTs) as proposed by .
Our model (referred to as MI-LSTM+CTC+WFST) consists of 4 bidirectional MI-LSTM layers, each with 320 units for each direction. CTC is performed on top to resolve the alignment issue in speech transcription. For comparison, we also train a baseline model (referred to as LSTM+CTC+WFST) with the same size but using vanilla LSTM. Adam with learning rate is used for optimization and Gaussian weight noise with zero mean and 0.05 standard deviation is injected for regularization. We evaluate our models on the character error rate (CER) without language model and the word error rate (WER) with extended trigram language model.
Table 1, top right panel, shows that MI-LSTM+CTC+WFST achieves quite good results on both CER and WER compared to recent works, and it has a clear improvement over the baseline model. Note that we did not conduct a careful hyper-parameter search on this task, hence one could potentially obtain better results with better decoding schemes and regularization techniques.
4 Learning Skip-Thought Vectors
Next, we evaluate our Multiplicative Integration on the Skip-Thought model of . Skip-Thought is an encoder-decoder model that attempts to learn generic, distributed sentence representations. The model produces sentence representation that are robust and perform well in practice, as it achieves excellent results across many different NLP tasks. The model was trained on the BookCorpus dataset that consists of 11,038 books with 74,004,228 sentences. Not surprisingly, a single pass through the training data can take up to a week on a high-end GPU (as reported in ). Such training speed largely limits one to perform careful hyper-parameter search. However, with Multiplicative Integration, not only the training time is shortened by a factor of two, but the final performance is also significantly improved.
We exactly follow the authors’ Theano implementation of the skip-thought modelhttps://github.com/ryankiros/skip-thoughts: Encoder and decoder are single-layer GRUs with hidden-layer size of 2400; all recurrent matrices adopt orthogonal initialization while non-recurrent weights are initialized from uniform distribution. Adam is used for optimization. We implemented Multiplicative Integration only for the encoder GRU (embedding MI into decoder did not provide any substantial gains). We refer our model as MI-uni-skip, with initialized as . We also train a baseline model with the same size, referred to as uni-skip(ours), which essentially reproduces the original model of .
During the course of training, we evaluated the skip-thought vectors on the semantic relatedness task, using SICK dataset, every 2500 updates for both MI-uni-skip and the baseline model (each iteration processes a mini-batch of size 64). The results are shown in Figure 2a. Note that MI-uni-skip significantly outperforms the baseline, not only in terms of speed of convergence, but also in terms of final performance. At around 125k updates, MI-uni-skip already exceeds the best performance achieved by the baseline, which takes about twice the number of updates.
We also evaluated both models after one week of training, with the best results being reported on six out of eight tasks reported in : semantic relatedness task on SICK dataset, paraphrase detection task on Microsoft Research Paraphrase Corpus, and four classification benchmarks: movie review sentiment (MR), customer product reviews (CR), subjectivity/objectivity classification (SUBJ), and opinion polarity (MPQA). We also compared our results with the results reported on three models in the original skip-thought paper: uni-skip, bi-skip, combine-skip. Uni-skip is the same model as our baseline, bi-skip is a bidirectional model of the same size, and combine-skip takes the concatenation of the vectors from uni-skip and bi-skip to form a 4800 dimension vector for task evaluation. Table 2 shows that MI-uni-skip dominates across all the tasks. Not only it achieves higher performance than the baseline model, but in many cases, it also outperforms the combine-skip model, which has twice the number of dimensions. Clearly, Multiplicative Integration provides a faster and better way to train a large-scale Skip-Thought model.
5 Teaching Machines to Read and Comprehend
In our last experiment, we show that the use of Multiplicative Integration can be combined with other techniques for training RNNs, and the advantages of using MI still persist. Recently, introduced Recurrent Batch-Normalization. They evaluated their proposed technique on a uni-directional Attentive Reader Model for the question answering task using the CNN corpusNote that used a truncated version of the original dataset in order to save computation.. To test our approach, we evaluated the following four models: 1. A vanilla LSTM attentive reader model with a single hidden layer size 240 (same as ) as our baseline, referred to as LSTM (ours), 2. A multiplicative integration LSTM with a single hidden size 240, referred to as MI-LSTM, 3. MI-LSTM with Batch-Norm, referred to as MI-LSTM+BN, 4. MI-LSTM with Batch-Norm everywhere (as detailed in ), referred to as MI-LSTM+BN-everywhere. We compared our models to results reported in (referred to as LSTM, BN-LSTM and BN-LSTM everywhere) Learning curves and the final result number are obtained by emails correspondence with authors of ..
For all MI models, were initialized to . We follow the experimental protocol of https://github.com/cooijmanstim/recurrent-batch-normalization.git. and use exactly the same settings as theirs, except we remove the gradient clipping for MI-LSTMs. Figure. 2b shows validation curves of the baseline (LSTM), MI-LSTM, BN-LSTM, and MI-LSTM+BN, and the final validation errors of all models are reported in Table 2, bottom right panel. Clearly, using Multiplicative Integration results in improved model performance regardless of whether Batch-Norm is used. However, the combination of MI and Batch-Norm provides the best performance and the fastest speed of convergence. This shows the general applicability of Multiplication Integration when combining it with other optimization techniques.
Relationship to Previous Models
2 Relations to Second Order RNNs and Multiplicative RNNs
MI-RNN is related to the second order RNN and the multiplicative RNN (MRNN) . We first describe the similarities with these two models:
There are however several differences that make MI a favourable model: (1) Simpler Parametrization: MI uses a rank-1 approximation compared to the second order RNNs, and a diagonal approximation compared to Multiplicative RNN. Moreover, MI-RNN shares parameters across the first and second order terms, whereas the other two models do not. As a result, the number of parameters are largely reduced, which makes our model more practical for large scale problems, while avoiding overfitting. (2) Easier Optimization: In tensor decomposition methods, the products of three different (low-rank) matrices generally makes it hard to optimize . However, the optimization problem becomes easier in MI, as discussed in section 2 and 3. (3) General structural design vs. vanilla-RNN design: Multiplicative Integration can be easily embedded in many other RNN structures, e.g. LSTMs and GRUs, whereas the second order RNN and MRNN present a very specific design for modifying vanilla-RNNs.
Moreover, we also compared MI-RNN’s performance to the previous HF-MRNN’s results (Multiplicative RNN trained by Hessian-free method) in Table 1, bottom left and bottom middle panels, on Penn-Treebank and text8 datasets. One can see that MI-RNN outperforms HF-MRNN on both tasks.
3 General Multiplicative Integration
Multiplicative Integration can be viewed as a general way of combining information flows from two different sources. In particular, proposed the ladder network that achieves promising results on semi-supervised learning. In their model, they combine the lateral connections and the backward connections via the “combinator” function by a Hadamard product. The performance would severely degrade without this product as empirically shown by . explored neural embedding approaches in knowledge bases by formulating relations as bilinear and/or linear mapping functions, and compared a variety of embedding models on the link prediction task. Surprisingly, the best results among all bilinear functions is the simple weighted Hadamard product. They further carefully compare the multiplicative and additive interactions and show that the multiplicative interaction dominates the additive one.
Conclusion
In this paper we proposed to use Multiplicative Integration (MI), a simple Hadamard product to combine information flow in recurrent neural networks. MI can be easily integrated into many popular RNN models, including LSTMs and GRUs, while introducing almost no extra parameters. Indeed, the implementation of MI requires almost no extra work beyond implementing RNN models. We also show that MI achieves state-of-the-art performance on four different tasks or 11 datasets of varying sizes and scales. We believe that the Multiplicative Integration can become a default building block for training various types of RNN models.
Acknowledgments
The authors acknowledge the following agencies for funding and support: NSERC, Canada Research Chairs, CIFAR, Calcul Quebec, Compute Canada, Disney research and ONR Grant N00014-14-1-0232. The authors thank the developers of Theano and Keras , and also thank Jimmy Ba for many thought-provoking discussions.
References
Appendix A Implementation Details
Our MI-LSTM (without peephole connection) in experiments has the follow formulation:
where are bias vectors, denotes the sigmoid function.
A.2 MI-GRU
Our MI-GRU in experiments has the follow formulation:
where are bias vectors, denotes the sigmoid function.