Distance-based Self-Attention Network for Natural Language Inference

Jinbae Im, Sungzoon Cho

Introduction

Sequence modeling has been employing Recurrent Neural Networks (RNN) or Convolutional Neural Networks (CNN) mostly. More recently, models incorporating attention mechanisms have shown good performance in machine translation (Bahdanau et al., 2014; Sutskever et al., 2014), Natural Language Inference (NLI) (Liu et al., 2016), and Question Answering (QA) (Hermann et al., 2015; Sukhbaatar et al., 2015) etc. Attention mechanisms used to be exploited in conjunction with RNN or CNN as an ancillary means to help improve performance. Lately, Vaswani et al. (2017) presented the first fully attention-based model, which recorded the state-of-the-art result in machine translation. As a fully attention-based model can consider all words in a sentence at once, parallelization leads to great reduction in training time.

Motivated by Vaswani et al. (2017), Shen et al. (2017) proposed the first fully attention-based sentence encoder. Shen et al. (2017) recorded good performance in a variety of tasks. In particular, they recorded the state-of-the-art result with Stanford Natural Language Inference (SNLI) dataset (Bowman et al., 2015) which is a representative dataset of NLI. The NLI task aims to classify the relationship between two sentences as entailment, contradiction, or neutral. One of the approaches to solving the NLI task is to use sentence-encoding based models.The NLI task can be solved through two different approaches: sentence encoding-based models and joint models. The former separately encode each sentence, whereas the latter take into account the direct relationship between two sentences. Between them, sentence-encoding based models focus on training sentence encoder that can represent sentences in vector form well. We focus on the former approach, since the objective of our work is to develop an advanced sentence-encoding model. Shen et al. (2017) presented a sentence-encoding based model reflecting directional information in a sentence. However, the distance between words was not considered at all in their model, and the directional information simply involved words before and after the reference word. Altogether, positional information of words was not fully taken into account. As a result, the difference of importance between the distant words and the nearby words was not appropriately reflected. Hence local dependency was not properly modeled, which in turn failed to capture the context information in long sentences.

To tackle this limitation, we propose Distance-based Self-Attention Network which introduces a distance mask which models the relative distance between words. In conjunction with a directional mask, the distance mask allows us to incorporate complete positional information of words in our model. Our Distance-based Self-Attention Network achieved good performance with NLI data, and recorded the state-of-the-art result with SNLI. Our model worked exceptionally well with long sentences, in particular. We also visualized the effect of the distance mask to show that our model can grasp both local dependency and global dependency.

Related Works

NLI tasks have been studied through models of various structures. Most of all, models combining attention with Long Short-Term Memory (LSTM) have performed well. Liu et al. (2016) improved the performance by adding the mean pooling vector to the conventional attention model in which attention is applied to hidden states of LSTM. Chen et al. (2017) used the input gates of the LSTM as attention weights to simplify the model structure. In Chen et al. (2017) and Ni and Bansal (2017), short-cut connections in stacked LSTM, in combination with max-pooling originally suggested by Conneau et al. (2017), were proven effective in improving performance, recording the state-of-the-art performance in MultiNLI. And Munkhdalai and Yu (2016a) used the memory for sentence encoding motivated by Neural Turing Machine (Graves et al., 2014).

Vaswani et al. (2017) was the first study to construct an end-to-end model with attention alone, and recorded the state-of-the-art performance in machine translation tasks. Vaswani et al. (2017)’s encoder-decoder framework consists of a multi-head attention and a position-wise feed forward network as a basic building block which is deeply stacked combined with residual connection. The multi-head attention projects the input sentences to multiple subspaces and then computes the scaled dot-product attention in each subspace. The results in each subspace are then concatenated and projected again. Position-wise feed forward network adds non-linearity to vector representations of each position. In this way, the fully attention-based model was constructed without using RNN or CNN, and the training cost was greatly reduced.

Shen et al. (2017), a very recent work, constructed a fully attention-based sentence encoder motivated by Vaswani et al. (2017). They proposed a multi-dimensional attention mechanism that computes the attention by each dimension through modification of additive attention. In addition, their model exploits directional attention as well as fusion gate motivated by bi-directional LSTM. Directional information was reflected by introducing a simple directional mask. By adding a directional mask to the logit of attention, words in a specific direction in the sentence were masked to avoid attention. The extent to which attention results are ultimately reflected was determined through fusion gate. In our study, we construct our model based on Vaswani et al. (2017)’s basic building block, as well as Shen et al. (2017)’s key model structures. In order to model the distance between words, which was not considered in their works, we transform the multi-head attention in Vaswani et al. (2017), in particular, to fit our objective. Details can be found in section 4.

Background

In Vaswani et al. (2017), the attention function is defined as follows by introducing the concept of query, key, and value. “An attention function can be described as mapping a query and a set of key-value pairs to an output, where the query, keys, values, and output are all vectors. The output is computed as a weighted sum of the values, where the weight assigned to each value is computed by a compatibility function of the query with the corresponding key (Vaswani et al., 2017).” The two most commonly used attentions are additive attention (Bahdanau et al., 2014; Shang et al., 2015) and dot-product attention (Kim et al., 2016; Sukhbaatar et al., 2015; Vaswani et al., 2017).

Let query, iith key, and iith value be qq, kik_{i}, and viv_{i} respectively. (q∈Rdkq\in R^{d_{k}}, ki∈Rdkk_{i}\in R^{d_{k}}, and vi∈Rdvv_{i}\in R^{d_{v}})

Compatibility function of the query with the iith key is represented by the following equation 1.

where u∈Rdku\in R^{d_{k}}, and σ(⋅)\sigma(\cdot) is an activation function usually chosen as tanh.

And attention weight assigned to each iith value is computed by applying the softmax function to lil_{i} and final output is weighted sum of value as following equations.

2 Dot-product Attention

Dot-product attention is the same as additive attention except for compatibility function. In dot-product attention, compatibility function is computed by the following equation 4 in place of the equation 1.

On implementation, dot-product attention is much faster and more space-efficient than additive attention due to optimized matrix multiplication.

In practice, however, additive attention outperforms dot product attention for large values of dkd_{k}. So Vaswani et al. (2017) used scaled dot-product attention instead of normal dot-product attention to prevent performance loss in large dimension as following equation 5.

Proposed Model

Our model’s overall architecture is shown in Figure 1. We follow the conventional architecture for training NLI data. First, the two input sentences, premise and hypothesis, are encoded as vectors, uu and vv respectively, through identical sentence encoders. For the encoded vectors uu and vv, the representation of relation between the two vectors is generated by the concatenation of uu, vv, ∣u−v∣|u-v|, and u∗vu*v. Thereafter, a probability for each of the 3-class is generated through the 300D ReLU layer and the 3-way softmax output layer. We configured the model with the setting of 1layer 300D as in Shen et al. (2017) to focus on the performance evaluation of the sentence encoder itself. Layer normalization (Ba et al., 2016) and dropout are applied to 300D ReLU layer.

2 Sentence Encoder

The sentence encoder structure proposed in this paper is shown in Figure 2. The term “Norm” in Figure 2 stands for layer normalization. The sentence encoder of Figure 2 encodes the premise and hypothesis in a vector form. We describe each component of our sentence encoder in detail in the following subsections.

Let an input sentence be a sequence of discrete words x=[x1,x2,⋅⋅⋅,xn]\mathbf{x}=[x_{1},x_{2},\cdot\cdot\cdot,x_{n}], where xi∈RNx_{i}\in R^{N} is a one-hot representation of the word ii, and NN is the vocabulary size. These one-hot representations are transformed into dense representations by using the pre-trained word embedding.

Let We∈Rde×NW_{e}\in R^{d_{e}\times N} be a pre-trained word embedding matrix. Then a sequence of dense word representations can be written as w\mathbf{w} = Wex=[w1,w2,⋅⋅⋅,wn]{W_{e}}\mathbf{x}=[w_{1},w_{2},\cdot\cdot\cdot,w_{n}], where wi∈Rdew_{i}\in R^{d_{e}} is dense representation of the word ii.

2.2 Masked Multi-Head Attention

The masked multi-head attention is a variation of the multi-head attention employed by Vaswani et al. (2017). The scaled dot-product attention of Vaswani et al. (2017) is expressed as following:

where Q,K,VQ,K,V are matrices composed of a set of queries, keys, and values, respectively.

We transform equation 6 and express the masked attention as following:

Here, Mdir∈Rn×nM_{dir}\in R^{n\times n} is the directional mask as proposed in Shen et al. (2017), while Mdis∈Rn×nM_{dis}\in R^{n\times n} is the distance mask proposed in this model. Hyper parameter α\alpha is the distance-alpha tuned through validation data.

MdirM_{dir} consists of the forward mask and backward mask as explained in Figure 3. In the Forward Masked Multi-Head Attention phase, the forward mask is selected, and in the Backward Masked Multi-Head Attention phase, the backward mask. The forward masks prevent words that appear after a given word from being considered in the attention process, while backward masks prevent words that appear before from consideration by adding −∞-\infty to the logits before taking the softmax at the attention phase. The diagonal component of MdirM_{dir} is also set to −∞-\infty so that each token does not consider itself to attention, and the information of each token is later transmitted through the fusion gate of section 4.2.3

MdisM_{dis} is shown in the Figure 4. The (i,j)(i,j) component of the distance mask is −∣i−j∣-|i-j|, representing the distance between (i+1)(i+1)th word and (j+1)(j+1)th word multiplied by −1-1. By multiplying this value by α\alpha and adding it to logit, the attention weight becomes smaller as distance increases. That is, the distance mask serves to concentrate on the local words around the reference word. Such a structure may appear similar to a CNN filter extracting a local feature. Yet, the big difference is that CNN only uses information in the window size, whereas our model considers all words in a sentence at once, concentrating on the local words by taking account of the relative distance between words.

By using the distance mask, the distance between words, not considered through the directional mask of Shen et al. (2017), was considered additionally, so the complete positional information of words was taken into consideration.In Vaswani et al. (2017), the positional information of the word was used through positional encoding. By adding the positional encoding vector to the word embedding vector, the embedding changed according to the absolute position of the word in the sentence. However, in sentence modeling, the relative position with respect to the other words is important, not the absolute position of the word. In other words, what words are placed in order before and after the word is important, not the absolute position of the word in a sentence. Therefore, we take the relative position directly into account in our model through the distance mask instead of the positional encoding which considers the relative position indirectly.

The masked multi-head attention can be expressed as following:

where headi=Masked(QWiQ,KWiK,VWiV)head_{i}=\text{Masked}(QW_{i}^{Q},KW_{i}^{K},VW_{i}^{V}), with hh as the number of heads, WiQW_{i}^{Q}, WiKW_{i}^{K}, WiVW_{i}^{V} ∈Rde×de/h\in R^{d_{e}\times d_{e}/h}, and WOW^{O} ∈Rde×de\in R^{d_{e}\times d_{e}}. Q,K,VQ,K,V ∈Rn×de\in R^{n\times d_{e}} are matrices created from nn word embedding vectors of sentences and expressed as equation 9.

The masked multi-head attention first projects Q,K,VQ,K,V into hh subspaces, respectively, and performs masked attention of equation 7 for each Q,K,VQ,K,V projection combination. The hh attention result is concatenated before projection.Multi-head attention (Vaswani et al., 2017) is fast and efficient because it is based on dot-product attention. However, multi-dimensional attention (Shen et al., 2017) has a disadvantage in that it consumes a lot of gpu memory because it requires several 4-dimensional tensors on implementation. So, in our model, the multi-head attention was used as a base structure instead of the multi-dimensional attention. In addition, the performance of the actual implementation was also better with multi-head attention.

2.3 Fusion Gate

At the fusion gate, raw word embedding S∈Rn×deS\in R^{n\times d_{e}} and the result of masked multi-head attention H∈Rn×deH\in R^{n\times d_{e}} in equation 10 are used as input.

First, we generate SF,HFS^{F},H^{F} by projecting S,HS,H using WS,WH∈Rde×deW^{S},W^{H}\in R^{d_{e}\times d_{e}}. Mathematically:

Then create gate FF as shown in equation 12 where bF∈Rdeb^{F}\in R^{d_{e}}.

Finally, we obtain the gated sum by using FF. It is common in many papers including Shen et al. (2017) to use raw SS and HH in gated sum. We, however, use the gated sum of SFS^{F} and HFH^{F} which resulted in a significant increase in accuracy.

2.4 Position-wise Feed Forward Networks

We used position-wise feed forward network structure of Vaswani et al. (2017) as it is. The position-wise feed forward network employs the same fully connected network to each position of sentence, in which the fully connected layer consists of two linear transformations, with the ReLU activation in between. Mathematically:

where x∈R1×dex\in R^{1\times d_{e}}, W1P∈Rde×dffW_{1}^{P}\in R^{d_{e}\times d_{ff}}, W2P∈Rdff×deW_{2}^{P}\in R^{d_{ff}\times d_{e}}, b1P∈Rdffb_{1}^{P}\in R^{d_{ff}}, and b2P∈Rdeb_{2}^{P}\in R^{d_{e}}.

The FFN function of the above equation 13 is applied to each position of the result of the fusion gate. Note that position-wise feed forward network is combined with the residual connection as shown in Figure 2. That is, FFN learns the residuals. In our model, dffd_{ff} was set to 4de4d_{e}.

2.5 Pooling Layer

The vector representation of input sentence is generated through the pooling layer after the concatenation of the results of forward directional self attention and backward directional self attention. That is, the input of pooling layer is U=[Ufw;Ubw]∈Rn×2deU=[U^{fw};U^{bw}]\in R^{n\times 2d_{e}} where each directional self attention output is Ufw∈Rn×deU^{fw}\in R^{n\times d_{e}}, Ubw∈Rn×deU^{bw}\in R^{n\times d_{e}}.

We use the multi-dimensional source2token self-attention of Shen et al. (2017) for our multi-dimensional self-attention.

For iith row vector of UU, uiu_{i}, logit l(ui)l(u_{i}) is computed as following:

where ui=Ui∗∈R1×2deu_{i}=U_{i*}\in R^{1\times 2d_{e}}, W1M,W2M∈R2de×2deW_{1}^{M},W_{2}^{M}\in R^{2d_{e}\times 2d_{e}}, and b1M,b2M∈R2deb_{1}^{M},b_{2}^{M}\in R^{2d_{e}}.

The calculations of logit consist of two linear transformations, with the Exponential Linear Units (ELU) activation function (Clevert et al., 2015) in between. Multi-dimensional attention differs from general attention in that the logit for an input vector is not a scalar but a vector with dimensions equal to the dimensions of the input vector. This allows each dimension of the input vector to have a scalar logit, and we can perform attention to nn word tokens in each dimension, as illustrated below by equation 15, 16. Note that softmax is performed on the row dimension of LL, not the column dimension.

The 2de2d_{e}-dimensional output vector of multi-dimensional attention and the 2de2d_{e}-dimensional vector obtained by applying max pooling to UU are concatenated to encode the input sentence as a 4de4d_{e}-dimensional vector.

Experiments and Results

The dataset used in the experiments are SNLI (Bowman et al., 2015) and MultiNLI (Williams et al., 2017) datasets. The SNLI dataset consists of 549,367 / 9,842 / 9,824 (train / valid / test) premise and hypothesis pairs; and the MultiNLI dataset, 392,702 / 9,815 / 9,832 / 9,796 / 9,847 (train / valid_matched / valid_mismatched / test_matched / test_mismatched) sentence pairs. The two datasets have the same format, but sentences in the MultiNLI dataset are much longer than those in SNLI dataset. In addition, MultiNLI dataset consists of various genre information. If genres included in the train data are also found in valid (test) data, then the dataset is called “matched”; if valid (test) data includes genres that are not in the train data, then the dataset is called “mismatched”.

2 Training Details

We used the Glove 840B 300Dhttps://nlp.stanford.edu/projects/glove/ (de=300d_{e}=300) for the pre-trained word embedding without any fine-tuning. This is to train the more universally usable sentence encoder.

Layer normalization (Ba et al., 2016) was applied to all linear projections of masked multi-head attention, fusion gate, and multi-dimensional attention. We applied residual dropout as used in Vaswani et al. (2017), with dropout to the output of masked multi-head attention and SF+HF+bFS^{F}+H^{F}+b^{F} of fusion gate.

3 SNLI Results

Experimental results of SNLI data compared with the existing models on the SNLI leader-boardhttps://nlp.stanford.edu/projects/snli/ are shown in Table 1. Compared with the existing state-of-the-art model (Shen et al., 2017), the number of parameters and the training time increased, but our results show the new state-of-the-art record. We also looked at the model with distance mask removed to verify the effect of the distance mask proposed in this paper. Results show that the addition of the distance mask improved the performance without significantly affecting the training time or increasing the number of parameters.

The improvement of the test accuracy by introducing the distance mask is only by 0.3% point, potentially because SNLI data mostly consist of short sentences. Hence, we additionally examined how the effect of the distance mask changes as the average length of the two sentences of premise and hypothesis pair changes. The distribution of the average length of the two sentences of the SNLI test data is shown in Figure 5, and the effect of the distance mask according to the average length change can be seen from Figure 6. Figure 6 shows that the accuracy is similar until the average length is less than 25, yet the test accuracy of the model without the distance mask deteriorates drastically for data of an average length exceeding 25. This demonstrates that the distance mask has an advantage with long sentences or documents.

4 MultiNLI Results

The results of applying SNLI best model to MultiNLI dataset without additional parameter tuning are presented in Table 2. Note that matched-test accuracy and mismatched-test accuracy were obtained by submitting our test results to Kaggle open evaluation platforms: MultiNLI Matched Open Evaluationhttps://www.kaggle.com/c/ multinli-matched-open-evaluation and MultiNLI Mismatched Open Evaluationhttps://www.kaggle.com/c/ multinli-mismatched-open-evaluation. First, the average test accuracy difference is greater than 2% when compared to the Directional Self-Attention Network (Shen et al., 2017). This once again confirms our model’s advantage in long sentences, given that the sentence is much longer in MultiNLI.

Compared with the result of RepEVAL 2017 (Nangia et al., 2017), we can see that the Distance-based Self-Attention Network performs well. When compared with the model of Chen et al. (2017), our model showed similar average test accuracy with much lower number of parameters. Also, considering that the model of Chen et al. (2017) is a complex LSTM model, our model has an advantage in training time as a fully attention-based model.

Ni and Bansal (2017) showed the best performance with 74.5% accuracy in Matched Test. However, it is a very deep structured LSTM model with 140.2m parameters. In our model, the inference layer is simply composed of 1 layer of 300D in order to focus on the training of sentence encoder. Both in Chen et al. (2017) and Ni and Bansal (2017) models, the inference layer was set very complex in order to improve the MultiNLI accuracy. Taking this into consideration, it can be seen that our Distance-based Self-Attention Network performs competitively given its simpler structure.

5 Case Study

A case study was conducted to investigate the role of each structure of the Distance-based Self-Attention Network. For this, a sentence “A lady stands outside of a Mexican market.” is picked among the premise sentences of SNLI test data. We focused on training encoders that can represent each sentence in a vector form well. Therefore, a case study was conducted on a single sentence, not a sentence pair.

Masked Multi-Head Attention We first look at the attention weights in masked multi-head attention. Attention weights represent a nn by nn matrix corresponding to softmax(QKTdk+Mdir+αMdis)(\frac{QK^{T}}{\sqrt{d_{k}}}+M_{dir}+\alpha M_{dis}) of equation 7, which is different for each head. Here we look at the average attention weights obtained by averaging the attention weights of each head. The attention weights for each head can be found in Appendix.

The row of the matrix of Figure 7 represents each word of the sentence, and the column represents the attention weights for each word at each row. It can be seen that the attention weights are heavier to the nearby words as compared to those distant from the reference word. At the same time, ‘outside’ in the forward mask and ‘Mexican’ in the backward mask have high attention weights for several words. From this, it can be seen that important word is considered in the attention process.

Distance Mask We compared the masked multi-head average attention weights for the longest sentence example in the SNLI test data, with length of 57 words to further verify the effect of the distance mask. Panels (a) and (b) of Figure 8 show results without considering distance, while (c) and (d) show the results with the distance mask. In panels (a) and (b), very distant words are considered in the attention and the overall attention weights were reduced. This implies that each word does not focus on the important words in the attention process, but rather takes into account almost every word, resulting in noisier figures.

However, in panels (c) and (d), the neighboring words are seen more intensively, which implies that the local dependency has been well captured by our model. In addition, as shown in panel (c), even if the word is far apart, it is still considered in the attention process if it is important. This demonstrates the effectiveness of the distance mask to identify local dependencies without losing the ability to grasp the global dependency.

Fusion Gate We visualize the role of the fusion gate F∈Rn×deF\in R^{n\times d_{e}} at forward directional self attention. Figure 9 represents the average gate value that averages ded_{e}-dimensional gate value for each word. If look at the results of both extremes, keyword ‘Mexican’ has a low gate value, resulting in an output that greatly reflects the multi-head attention result. In contrast, ‘of’, ‘.’, the words of little importance, have large gate values, which indicates that the original word embedding is greatly reflected, not the multi-head attention result.

Position-wise FFN For the FFN function of equation 13, Figure 10(a) represents the deactivation ratio in the first hidden layer of position-wise ffn.

As shown in Figure 2, position-wise ffn is used in conjunction with a residual connection. That is, the final output of position-wise ffn for input xx is the ded_{e}-dimensional vector of LayerNorm(x+FFN(x))(x+\text{FFN}(x)). Figure 10(b) visualizes the maximum value of this final output vector.

In Figure 10, keywords with a high deactivation ratio is shown in panel (a) and a high final max value in panel (b). In case of a word corresponding to a keyword, deactivation occurs frequently in (a), and residual learning is hardly achieved in the position-wise ffn, so that the output of the fusion gate is almost maintained. On the other hand, in case of non-important words, residual learning is performed in position-wise ffn because there is less deactivation in (a), so that the max value of final output becomes smaller in (b). This results in preventing non-important words from consideration in the subsequent pooling layer. In summary, position-wise ffn plays a key role in ensuring that non-critical words are paid less attention to in pooling layers.

Pooling Layer For the multi-dimensional attention corresponding to Figure 11(a), we visualized the attention weights averaged for each word, where attention weights correspond to softmax(L)∈Rn×2de(L)\in R^{n\times 2d_{e}} in equation 15.

In max pooling, the max value is selected for each column of U∈Rn×2deU\in R^{n\times 2d_{e}}. Thus, in Figure 11(b), we visualize the percentage at which each word is selected in the max pooling operation for the 2de2d_{e} dimension.

It can be seen that panels (a) and (b) of Figure 11 are similar on the whole. In other words, both multi-dimensional attention and max pooling utilize information about key words intensively. A similar result can be expected by using only one of the pooling layers. However, experiment results show that using both multi-dimensional attention and max pooling layer gives better performance.

Conclusion

In this paper, we propose the Distance-based Self-Attention Network reflecting the distance between words. By reflecting the word distance information, our model learns the local dependency without losing the ability to capture the global dependency. This was achieved through a simple distance mask, so that the performance of the NLI task could be improved while maintaining the number of parameters and training time. In particular, we recorded the new state-of-the-art performance for SNLI data. The introduction of the distance mask improves the performance with longer sentences.

As the research on universal sentence encoders using NLI data was proposed by Conneau et al. (2017), we plan to carry out research on fully attention-based networks for universal sentence embedding as future work. We will also study the fully attention-based network in image data and speech data. Especially, regarding image data, capsule network (Sabour et al., 2017) recently proposed, and as research on new structure to replace CNN is going on, our future work will move in similar directions.

Acknowledgments

We would like to thank Hyejin Lee, Hyunjoong Kim, Taewook Kim, Jinwon An, Inbeom Park, Minki Chung, and many others in SNUDM center, for critical feedback and discussions.

This work was supported by the BK21 Plus Program(Center for Sustainable and Innovative Indus- trial Systems, Department of Industrial Engineering & Institute for Industrial Systems Innovation, Seoul National University) funded by the Ministry of Education, Korea (No. 21A20130012638), the National Research Foundation (NRF) grant funded by the Korea government (MSIP) (No. 2011- 0030814), and the Institute for Industrial Systems Innovation of SNU.

References

Appendix

Masked Multi-Head Attention The attention weights for each head in the masked multi-head attention are shown in Figures 12 and 13. Figure 12 shows the result of using a forward directional mask, and Figure 13 is the result of using a backward directional mask. It can be seen that the attention weights are different for each head. This allows our model to capture various dependencies between words in a sentence.