A Study on ReLU and Softmax in Transformer

Kai Shen, Junliang Guo, Xu Tan, Siliang Tang, Rui Wang, Jiang Bian

Introduction

Transformer (Vaswani et al., 2017) models have achieved great success in various natural language processing tasks including large-scale language model pretraining (Devlin et al., 2018; Brown et al., 2020) and machine translation (Vaswani et al., 2017; Hassan et al., 2018). A lot of works have been conducted to explore, analyze and explain the architecture. A series of works (Dai et al., 2021; Sukhbaatar et al., 2019; Yao et al., 2022) investigate how Transformers understand and store a huge amount of knowledge from data, mainly by building the connections between the network modules (i.e., self-attention network (SAN) and feed-forward network (FFN)) and key-value memory according to the similarity in their formats. For example, they take the two layers in FFN as key and value parameters that store different patterns of knowledge (Geva et al., 2020).

However, previous works have not considered the difference in the activation function, which plays an important role in neural networks. Specifically, traditional key-value memory (and also SAN) usually utilizes Softmax to normalize the output of query and key in order to highlight important values, whereas FFN utilizes ReLU as the activation function which does not introduce normalization. In this paper, we conduct extensive studies and revisit the connections between SAN, FFN, and key-value memory while taking activation functions into account.

Specifically, for FFN, we first replace ReLU with Softmax to keep consistent with the format of memory, but observe performance degradation because the activation results become too small to carry enough information after being normalized by Softmax. We introduce an additional layer normalization module on Softmax results to adjust the variance and show the identity of FFN and key-value memory. In addition, we explore their scalability as the hidden dimension (i.e., the dimension of the inner hidden vector between two layers) of FFN is usually very large on large-scale models (Shazeer et al., 2017). By changing the total number of key-value slots, we find that ReLU performs better than Softmax when the number of slots is larger. We explore the reason by calculating the ratio of top scores among all activations and find that the activation weights are highly centralized in a small number of slots, thus insufficient to utilize the context information of other slots, while ReLU is able to alleviate this problem.

Given the superior performance of ReLU when scaling to a large number of value slots, we then explore how ReLU performs on SAN where Softmax may have a trouble modeling long-sequences (Sun et al., 2022). Unfortunately, directly alternating Softmax to ReLU does not converge. With theoretical and experimental analysis, we find that the variance of SAN results with ReLU activation grows with the length of the input sequence, and the dynamic variance will lead to an unstable training process. Therefore, a variance reduction factor and regularization loss functions are introduced to solve this problem. As a result, we make it possible to utilize ReLU on self-attention, which performs better than Softmax when dealing with long input sequences.

In summary, this paper provides insights into the difference and relationship between ReLU and Softmax activation functions in the Transformer architecture.

Softmax provides exponential normalization over all value slots and therefore highlights a small number of them while neglecting others, which may cause performance degradation when the number of slots is large, e.g., FFN with large hidden dimensions or SAN with long input lengths. ReLU bypasses this problem but faces variance exploding which varies over different training samples.

We revisit the relations between the FFN and key-value memory and find that they are equivalent when additional layer normalization is introduced.

With the ReLU activation function in both SAN and FFN components, we propose a fully-ReLU Transformer architecture (ReLUFormer). We claim that our ReLUFormer can be viewed as an integration of key-value memory, with FFNs as global key-value memories and SANs as local memories. We evaluate the model on the task of long-document translation and verify the superiority of the full ReLU architecture over the Transformer baseline.

The rest of the paper is organized as follows. We introduce the backgrounds of FFN, SAN, and key-value memory in Section 2. And then we revisit the connections between FFN and key-value memory in Section 3. We then explore the connections between the ReLU and Softmax in SAN, and compare the ReLU and Softmax in the task of long document translation with the proposed fully-ReLU architecture in Section 4. We summarize and provide insights on our findings in Section 5.

Background

Transformer (Vaswani et al., 2017) has achieved great success on natural language processing tasks, such as machine translation and language modeling. Recently, some works have been proposed to analyze the architecture and investigate the secrete of the success of Transformers. Some works have revealed the relation between feed-forward network (FFN) and key-value memory (Geva et al., 2020; Sukhbaatar et al., 2019). They intuitively regard the FFN as key-value memory by unifying them in the formulation.

where dhd_{h} is the number of memory slots.

Previous works (Sukhbaatar et al., 2019; Geva et al., 2020) reveal that the FFN and key-value memory are similar in the formulation, i.e., by regarding W1W_{1} as keys and W2W_{2} as values the FFN can be viewed as a kind of key-value memory. In addition, Dai et al. (2021) conducts experiments on how knowledge is stored in a pre-trained model and finds that there are some knowledge neurons in the FFN layer related to the expression of factual knowledge. Lample et al. (2019) introduces a large-scale external memory based on product keys and successfully integrates it into transformer architecture by replacing the FFN layer. However, there still exists a difference in the choice of activation functions, where the FFN usually adopts ReLU and the key-value memory uses Softmax, which may lead to different model performance. In this paper, we will explore the connections between FFN and key-value memory by studying the ReLU and Softmax.

2 Self-Attention Network and Key-Value Memory

In conclusion, although FFN, SAN, and key-value memory are similar in formulation, previous works have not carefully discussed the differences in activation functions. In practice, it is a convention to use ReLU in FFN and Softmax in SAN and key-value memory. In this paper, we provide in-depth analyses of ReLU and Softmax as well as their performance on FFN and SAN. We start by revisiting the connections between FFN and key-value memory.

Connections Between FFN and Key-Value Memory

We first investigate whether the difference in ReLU and Softmax activation functions will influence the performance of FFN and key-value memory. We conduct a preliminary experiment that simply replaces FFN layers with key-value memories (i.e., replacing Equation (1) with Equation (2)) in a vanilla Transformer and keeps other components the same. We test on the machine translation task and the IWSLT14 De-En benchmark dataset (refer to Appendix A for more details). We report the BLEU score (Papineni et al., 2002) as the evaluation metric.

As shown in Table 1, we find the model performance drops from 34.2234.22 to 33.0833.08 in the BLEU score. According to the discussion in the previous section, the only difference lies in the activation function, we can conclude that the drop comes from changing ReLU to Softmax, which has a significant influence and can not be ignored.

2 Bridge the Gap between FFN and Key-value Memory

Then, we analyze the reason for the performance drop. When computing results, Softmax normalizes the score of all slots while ReLU does not. Intuitively, the results of Softmax will have a much smaller variance than the results of ReLU. In consequence, the residual from previous layers will dominate the output of the current FFN layer with the Softmax activation function, resulting in inefficient utilization of parameters and degeneration of model capacity. To demonstrate this phenomenon, we compute the ratio between the variance of FFN output and residual for both ReLU and Softmax. The results in Table 1 verify our claim, where the ratio of output with Softmax is much smaller than that of ReLU.

To alleviate this problem, we add a layer normalization (Ba et al., 2016) module after the FFN layer to learn and adjust the variance ratio of the output, i.e., Equation (2) becomes to H=LN(Softmax(X⋅KT)⋅V)H=\textrm{LN}(\textrm{Softmax}(X\cdot K^{T})\cdot V). The experimental results are shown in Table 1. With layer normalization, both the variance ratio and BLEU scores of Softmax are promoted, and the FFN with Softmax performs similarly to ReLU. Therefore, we amend the previous findings (Geva et al., 2020) that the FFN and key-value memory are equivalent when an additional layer normalization module is introduced.

3 Scaling to Large Number of Values

With the progress of large-scale Transformer networks, the hidden dimension (or the number of slots from the view of memories) dhd_{h} of FFN layers becomes larger and brings better results as shown in previous works (Shazeer et al., 2017; Fedus et al., 2021). To verify the performance and show the generalization ability of different activation functions when the number of values is large, we vary dhd_{h} from 3232 to 40964096 and train three models including FFN with ReLU (ReLU), FFN with Softmax (Softmax), and FFN with Softmax and layer normalization (Softmax with LN). Results are shown in Figure 1. The observations are twofold. 1) The ReLU activation function is superior to Softmax consistently on all sizes. 2) When equipped with LN, Softmax performs comparably to ReLU. However, when the memory size is large (i.e., 30723072, 40964096), the ReLU performs better than Softmax with LN. In conclusion, ReLU shows stronger capacity when dealing with a large number of values than Softmax.

4 Quantitative Analysis between ReLU and Softmax

We conjecture the reason is the exponential normalization in Softmax. Concretely, since Softmax provides the exponential normalization on the elements while ReLU does not, Softmax provides over-centralized distribution over elements, which means only a few elements are highlighted while occupying most weights. Then when the memory size is large, Softmax will overlook most value slots and only utilize a few of them, which does not benefit from the large size of memory. In contrast, there is no competition among elements in ReLU, which is able to aggregate more knowledge. A straightforward method to alleviate this problem is to increase the temperature in Softmax to flatten the output distribution. However, we empirically find it has little effect in experiments.

To verify the conjecture, we qualitatively analyze the competition by visualizing the score distribution, i.e., Softmax(X⋅KT)\textrm{Softmax}(X\cdot K^{T}) in Equation (2). For each distribution, we sort scores in descending order and calculate the summation of the top-p%p\% elements. We normalize ReLU scores to make sure the summation of all scores is 11 as ReLU(xikjT)/∑p=1nReLU(xikpT)\text{ReLU}(x_{i}k_{j}^{T})/\sum_{p=1}^{n}\text{ReLU}(x_{i}k_{p}^{T}), where xi,kjx_{i},k_{j} is the ii-th, jj-th elements of query XX, key KK mentioned in Section 2.1. Therefore, a higher top-p%p\% score sum indicates a more centralized distribution, i.e., all scores are concentrated on a small number of elements. From Figure 2, we can find that Softmax provides a highly centralized distribution as top 0.2%0.2\% elements occupy more than 85%85\% scores, and becomes more severe when the memory size grows. In contrast, ReLU can alleviate this problem and therefore utilize the information of more memory slots.

As a result of the over-centralized distribution, the optimization of memory slots with Softmax will be sub-optimal as most of them cannot receive enough gradient due to small scores. We then quantitatively evaluate the quality of the learned values by measuring their anisotropy (Ethayarajh, 2019), which is defined as the average of pair-wise similarities and therefore the lower the better, i.e., value slots are different from each other and able to contain discriminative and diverse information. The anisotropy score (ANI) is defined as follows:

where ViV_{i} indicates the ii-th value slot. Visualization results are illustrated in Figure 3, from which we can find that the values of Softmax collapse and fail to store diverse knowledge, and adding LN alleviates this problem while utilizing ReLU achieves the best performance.

5 Summary

With these explorations, we have the following insights.

The results of Softmax and ReLU have different properties, including variance and normalization. For variance, the ReLU has a larger variance compared with Softmax, which is more expressive. The hidden output layer with a small variance may be dominated by the residual during end-to-end training, which can lead to the waste of the parameters and sub-optimal results. For normalization, since Softmax provides exponential normalization on the elements while ReLU does not, the distribution of Softmax is more centralized. When the memory size is large, Softmax will overlook most of the elements, thus resulting in a less diverse and discriminative memory value space.

When the Softmax is equipped with layer normalization in key-value memory, the FFN and key-value memory can be equivalent. The layer normalization can largely alleviate the small variance and over-centralized distribution brought by Softmax in the key-value memory, thus the Softmax with layer normalization can achieve comparable performance with FFN.

We find that the ReLU is more capable to deal with a large number of memory slots. When the number of memory slots is larger, the distribution of Softmax is more centralized, which results in the inefficient utilization of the memory slots.

The last observation also inspires us that when handling long sequences in self-attention, it is beneficial to pay attention to more elements instead of centralizing on a small portal. However, since self-attention is similar to key-value memory and it is conventional to use Softmax as the activation function, will ReLU perform better than Softmax when handling long sequences? We will explore the differences of ReLU and Softmax in self-attention in the next section.

ReLU vs Softmax in Self-Attention

Given the findings that ReLU outperforms Softmax in FFN, when dealing with a large number of value slots, a natural question is how will ReLU perform on the self-attention network (SAN). As discussed in Section 2.2, SAN can be straightforwardly formatted as key-value memory, with queries, keys, and values as different representations of the input. Then the memory size is denoted by the length nn of the input instead of the hidden dimension dhd_{h} in FFN. Therefore, we expect to observe the superiority of ReLU over Softmax when dealing with long sequences, following the conclusions of the previous section.

Similarly, we conduct preliminary experiments by directly replacing Softmax with ReLU, but we find the model fails to converge. Specifically, we find the variance of SAN results exploding. To solve the variance exploding problem, we add a layer normalization layer succeeding to the SAN results to adjust the variance of SAN. Unfortunately, there still occurs performance degradation. Such phenomenon is also reported in previous studies (Zhang et al., 2021). Therefore, in this section, we first analyze the reasons that ReLU fails and then propose our solutions. We then compare the performance of ReLU and Softmax on different sequence lengths.

Recall the formulation of SAN with ReLU activation:

where qi,kj,vjq_{i},k_{j},v_{j} is the i,j,ji,j,j-th element of query X^\hat{X}, key K^\hat{K}, and value V^\hat{V} mentioned in Section 2.2 respectively, hih_{i} is the output representation of the ii-th token, and nn is the sequence length. We find that the variance of hih_{i} is dependent on the sequence length nn by the following theory introduced by He et al., 2015:

Given nn random variables xi∼N(0,1),i∈[1,n]x_{i}\sim\mathcal{N}(0,1),i\in[1,n] and vj∼N(0,1),j∈[1,n]v_{j}\sim\mathcal{N}(0,1),j\in[1,n], yiy_{i} defined as:

Then yiy_{i} follows Gaussian distribution N(0,n2)\mathcal{N}(0,\frac{n}{2}).

Based on Theorem 4.1, in Equation (5), the output hih_{i} will follow the distribution N(0,n2)N(0,\frac{n}{2}). Therefore, the variance of results grows with the sequence length, and directly replacing Softmax with ReLU will lead to instability of the training process. Empirically, although adding layer normalization is supposed to learn and re-scale the variance, we find it still leads to sub-optimal BLEU results. We conjecture the different performance of LN on FFN and SAN is due to the dynamic memory size in SAN (i.e., the sequence length nn) which is static in FFN (i.e., the hidden dimension dhd_{h}), and thus the LN module is not able to learn appropriate variance for sentences with different lengths.

Motivated by these analyses, to stabilize the variance of ReLU output, we propose the variance reduction factor defined as γn/2\gamma\sqrt{n/2}, where γ\gamma is a hyper-parameter. This is similar to the implementation of Kaiming Normalization (He et al., 2015) but we adapt it during end-to-end training instead of initialization since the sequence length nn is dynamic. Formally, the Equation 5 goes to:

Then we apply it to the SAN and obtain 33.1933.19 BLEU scores on IWSLT14 De-En machine translation task as shown in Table 2. Although this is a big step towards a successful ReLU-based self-attention model, it still has a gap of 1.031.03 BLEU scores to the vanilla Softmax-based self-attention. We then go a step further and close the gap between ReLU and Softmax for SAN in the following section.

2 Closing the Gap Between ReLU and Softmax

To better analyze the performance gap between ReLU and Softmax-based SAN, we denote the weight distribution over values as s=(s1,...,sn)s=(s_{1},...,s_{n}) where sj=ReLU(qiTkj)/γn/2s_{j}=\textrm{ReLU}(q_{i}^{T}k_{j})/\gamma\sqrt{n/2} for ReLU w/ variance reduction, and sj=Softmax(qiTkj)s_{j}=\textrm{Softmax}(q_{i}^{T}k_{j}) for Softmax, and then compute the entropy of ss, i.e., H(s)=−∑i=1nsilog⁡(si)H(s)=-\sum_{i=1}^{n}{s_{i}\log(s_{i})}. We normalize the output of ReLU to ensure the summation is 11 with the same method used in Section 3.4 as ReLU(si)/∑j=1nReLU(sj)\text{ReLU}(s_{i})/\sum_{j=1}^{n}\text{ReLU}(s_{j}). Theoretically, a larger entropy indicates the distribution is more uniform, while a smaller one indicates the distribution is more centralized. And in our case, we want to find a balance where the entropy is neither too small nor too large, i.e., the weight distribution is not too uniform or over-centralized. In this way, the context information can be well utilized.

The computed entropy results are listed in Table 3. The distribution learned with ReLU has a very small entropy, and 94%94\% weights are all zeros. Therefore, the ReLU on self-attention leads to weight distributions that are too sparse to cover enough context. To alleviate this problem, we propose a regularization loss including two parts. Firstly, we introduce a normalization regularization to enlarge the weights. Secondly, we constrain the entropy of the learned distribution by an entropy-margin regularization to make it more informative. The loss functions are shown as follows:

where ∣⋅∣|\cdot| indicates the absolute value, CC is a constant that represents the upper bound of H(s)H(s). The normalization regularization encourages the summation of weights ∑i=1nsi\sum_{i=1}^{n}{s_{i}} to be 11, and therefore results in more non-zero elements. Note that different from the normalization in Softmax, it is not a hard constraint and therefore we do not observe the over-centralized distribution as that in Softmax. In contrast, the distribution becomes flat as we find the entropy grows drastically after the loss function is added. The entropy-margin regularization will encourage the model to try to keep the entropy of weight distribution sis_{i} smaller than the upper bound CC, where CC is a hyper-parameter which we describe in the Appendix A.

It is worth noting that the Transformer contains casual self-attention and cross-attention in the decoder, which are slightly different from the self-attention we have discussed aforementioned. To generalize the ReLU-based SAN to the decoder, we propose two solutions. 1) For the causal self-attention, we assign different lengths to each token as tokens after the current one are masked off. 2) The cross-attention does not require additional adaptations thus we treat it as the self-attention in the encoder. By packing all the components, we propose a fully ReLU Transformer named ReLUFormer.

3 ReLUFormer and Its Performance on Translation

In this section, we will first demonstrate the effectiveness of the proposed ReLUFormer on traditional sentence-level machine translation benchmarks. Then, we compare ReLU with Softmax when dealing with long sequences on document-level benchmarks.

We evaluate ReLUFormer on the sentence-level translation task, a seminal task in NLP. We consider two benchmark machine translation datasets, i.e., IWSLT14 German-English and WMT14 English-German. Details of datasets can be found in Appendix A.

In addition to the vanilla Transformer (Vaswani et al., 2017), we also consider other sparse activation baselines which alternate Softmax with other functions: 1) Sparsemax (Martins & Astudillo, 2016), 2) 1.5Entmax (Peters et al., 2019), and 3) Rectified Linear Attention (ReLA) (Zhang et al., 2021). The Sparsemax and 1.5Entmax are specially designed sparse attention activation functions similar to ReLU. And ReLA is another baseline simply replacing Softmax with ReLU, in which they propose the RMS normalization mechanism to address the variance problem caused by ReLU. We leave the detailed introduction of baselines to Appendix C.

The results of ReLUFormer and baselines are listed in Table 4, from which we have the following observations. 1) Our proposed ReLUFormer outperforms the vanilla Transformer baseline, showing the effectiveness of the proposed techniques. 2) When comparing with sparse attention-based baselines, our model also achieves consistent improvements over the Sparsemax, 1.5Entmax, and ReLA baselines by a large margin. 3) By comparing the latency during inference, we first find that ReLUFormer is comparable with the vanilla Transformer, while slightly faster than the ReLA and at 1.7 times faster than the Sparsemax and 1.5Entmax methods.

In this section, we conduct ablation studies on the proposed ReLUFormer to further discuss the effectiveness of the proposed three techniques: 1) attention scale factor (scale factor), and 2) regularization loss (reg loss). We demonstrate their effectiveness by removing each part individually. Besides the BLEU scores, we also use the entropy mentioned in Section 4.2 to quantitatively analyze the quality of the weight distribution.

The results are shown in Table 5, and the observations are as follows. 1) The model cannot converge by removing the scale factor because of the exploding of the variance, and adding the variance reduction factor stabilizes the training of the model. 2) By removing the normalization loss, the performance drops by 1.37 BLEU and the entropy decreases by 1.52. It illustrates that the regularization loss can provide a more informative attention distribution.

3.2 Experiments on Document-Level Translation

In this section, to verify our findings in Section 3 that ReLU performs better than Softmax when the number of value slots is large, we conduct experiments on the document-level translation task with long input and output sequences, because the length of the sequence represents the number of local memory slots in self-attention networks. We construct the long documents from a widely used document translation dataset Europarl7 En-De (Maruf et al., 2019; Zheng et al., 2020). To compare our method under different lengths of input sequences, we reconstruct the documents by concatenating sentences to different limitations of lengths. The details can be found in Appendix B. We conduct our experiment on 5 generated datasets with length limit {128,256,512,1024,2048}\{128,256,512,1024,2048\}.

We compare the proposed ReLUFormer with the vanilla Transformer (Vaswani et al., 2017) and the Sparsemax (Martins & Astudillo, 2016) baselines in document neural translation with different lengths. The results are shown in Table 6.

We have the following observations. 1) We find that when the sequence length is small (i.e., 128, 256), the ReLUFormer is comparable with the vanilla transformer and the Sparsemax baseline, showing that both methods are able to deal with relatively smaller lengths. 2) When the sequence length is large (i.e., 512, 1024, 2048), we find that our ReLUFormer consistently outperforms the vanilla transformer and the Sparsemax baselines. For example, when the sequence length is 1024 in Europarl7, ReLUFormer achieves 1.15 BLEU gains in translation quality. And the Sparsemax fails to converge when facing long sequences. It confirms that when the sequence is long, the ReLU is more effective.

In this section, we intuitively explain why ReLU outperforms Softmax in long document translation. Similar to the study in FFN and key-value memory, we also visualize the top-pp% elements mentioned in Section 3.4 of the activated scores for both Softmax and ReLU activation functions from the Europarl7 with 1024 length. A higher top-p% score sum indicates a more centralized distribution. In practice, we observe that the specific token rating in top-1.5% only occupies 2.1% and 4.4% scores for Softmax and ReLU, respectively. It means the tokens rating before top-1.5% dominate the quality of the performance. Thus we only visualize the elements rating before top-1.5%. From Figure 4, we can find that the Softmax provides a more centralized distribution compared with ReLU, which is similar to the observation in Section 3.3. When modeling long sequence inputs in self-attention, the over-centralized distribution will pay attention to fewer contexts, which results in sub-optimal performance.

We also provide self-attention visualization to further demonstrate the superiority of ReLU. We randomly select a case in the Europarl7 test set with 1024 length and visualize the attention map in Figure 5 for both ReLU and Softmax activation functions. We have the following observations. 1) We observe that the ReLU can capture more distant correlations. For example, the word “Madam” and “President” have relatively large attention weights in ReLU while small weights in Softmax. It shows the ReLU can capture more distant correlations compared to Softmax, which is beneficial to long sequence modeling. 2) We observe that the ReLU has less noise than Softmax. The ReLU assigns smaller attention values to some stop words such as “,” and “this”. Since the stop words contain little contextual information, it shows the ReLU has less noise.

4 Summary

In this section, we are motivated to explore whether ReLU outperforms Softmax in self-attention when handling long sequences.

We find that the variance of SAN results produced by ReLU is dependent on sequence length and therefore dynamic. Thus when ReLU is directly applied to replace Softmax in SAN, it can cause variance exploding and training instability.

With the extensive analysis of the performance degradation of ReLU, we propose solutions correspondingly to make the ReLU performs competitively to Softmax in SAN on the sentence-level translation task.

Similar to the observation in Section 3, we also find that the Softmax tends to generate a more centralized distribution which restricts the utilization of more context information especially when the sequence is long, while ReLU does not have the restriction.

By applying the ReLU-based Transformer to the long document translation task, we verify that the ReLU outperforms Softmax when the input sequence is long.

Insights and Findings

We summarize the insights and findings of the paper in this section.

Softmax and ReLU are different in Transformer from the perspective of variance and normalization. 1) Regarding the variance, the result of ReLU has a larger variance compared with that of Softmax. The hidden representation with a small variance will be dominated by the residual, leading to the inefficient usage of parameters and sub-optimal performance. In addition, the variance of SAN results produced by ReLU is related to the sequence length which will lead to the variance exploding problem. 2) Regarding the normalization, since Softmax provides exponential normalization on the elements while ReLU does not, the distribution of Softmax is more centralized compared with ReLU. When dealing with a large number of value slots, Softmax restricts the utilization of more context information and leads to sub-optimal performance, while ReLU does not.

When Softmax is equipped with layer normalization in key-value memory, the FFN and key-value memory are equivalent. The layer normalization can largely alleviate the scale and over-centralized problem caused by Softmax, which can boost performance.

The ReLU is good at handling a large number of key-value slots (in FFN and key-value memory) and long sequences (in SAN). With quantitive analyses, we find that ReLU is less centralized, thus it can integrate the context information of more tokens when the sequence is long.

As a whole, the Transformer can be viewed as the memory network, where FFN and SAN are global and local memory respectively. For FFN, the keys and values are parameters of two linear projection weights, which are globally shared by all input queries. For the SAN, the keys and values are constructed from the input sequences locally.

Conclusion

In this work, we revisit the relations between the Transformer components: the self-attention and feed-forward network, and key-value memory. Then we propose the full ReLU architecture which can achieve competitive performance with the vanilla Transformer. We have the following findings: 1) The FFN and key-value memory are equivalent when layer normalization is introduced; 2) Compared with Softmax, ReLU performs better when the number of memory value slots is large; 3) With specific designs, the proposed the full ReLU architecture can work effectively in sentence-level translation and significantly outperform vanilla Transformer in long document translation.

Limitation and Future Works

This work has the following limitations and will be extended from several aspects. First, we only conduct experiments on the machine translation task currently. Secondly, although we replace Softmax with a more efficient function ReLU, we obtain slight latency gains since it is still O(N2)O(N^{2}) complexity in self-attention. In the future, we will conduct experiments on more tasks, including language modeling, text summarization, etc. Then, since our work is parallel to the work designing the efficient Transformer, we will investigate better methods in self-attention to improve latency.

References

Appendix A Datasets and Implementation Details of Sentence-Level Translation

We evaluate our model on two widely used public sentence-level machine translation datasets: IWSLT14 De-En and WMT14 En-De, which have 153K/4.5M bilingual sentence pairs in corresponding training sets. For IWSLT14 De-En, following prior works (Ranzato et al., 2015; Bahdanau et al., 2016; Guo et al., 2019), we use 7K data split from the training set as the validation set and use the concatenation of dev2010, tst2010, tst2011 and tst2012 as the test set. For WMT14 En-De task, we use newstest2013 and newstest2014 as the validation and test set respectively. We use byte-pair encoding (BPE) (Sennrich et al., 2015) to tokenize and segment all the data into subword tokens. We share the source and target vocabulary and the embedding in each language pair. And the vocabulary is extracted by Moses (Koehn et al., 2007) for each training set with default hyperparameters.

A.2 Implementation Details

We follow the same encoder and decoder architecture as in Transformer (Vaswani et al., 2017). For the IWSLT14 De-En task, we use 6 encoder and decoder layers, 4 multi-heads, 512 and 1024 as the hidden and FFN inner dimension. For the WMT14 En-De task, we use 8 multi-heads and 2048 as the FFN inner dimension. During training, we apply dropout to the residual connections and attention weights with a rate of 0.1. We tune model parameters using Adam (Kingma & Ba, 2014) (β1=0.9,β2=0.98\beta_{1}=0.9,\beta_{2}=0.98) with label smoothing of 0.1. We schedule the learning rate following (Vaswani et al., 2017) with a warmup step of 4K. Each training batch contains around 8192 tokens, and the model is trained with 10w steps. While inference, we set the length penalty to 1.1 for IWSLT14 De-En task and 0.6 for WMT14 De-En task. The beam size is set to 5 for all tasks. We report tokenized BLEU scores as the evaluation metric. Importantly, the choice of entropy upper bound CC in Equation (7) should depend on the sequence length nn as the entropy varies on different sizes. Therefore we set C=0.7log(n)C=0.7log(n) in all experiments.

Appendix B Datasets of Document Translation

We evaluate our model on a widely used document translation dataset Europarl7 En-De (Maruf et al., 2019; Zheng et al., 2020). To compare the models under different lengths of input sequences, we reconstruct the documents by concatenating sentences to different limitations of lengths. Given a length limitation LL and KK sentences with {l1,l2,...,lK}\{l_{1},l_{2},...,l_{K}\} tokens in a document, we select mm consecutive sentences that satisfy ∑1mli<=L\sum_{1}^{m}l_{i}<=L and ∑1m+1li>L\sum_{1}^{m+1}l_{i}>L to form a new document. The next document will start with the sentence lm+1l_{m+1}. We vary the document length limitation in L∈{128,256,512,1024,2048}L\in\{128,256,512,1024,2048\}. We construct the validation and test sets with the same procedure as the training set.

Appendix C Baseline Details for Sentence-Level Translation

The Sparsemax (Martins & Astudillo, 2016) is a wildly used sparse activation function. Compared with Softmax, Sparsemax sets the low score elements to zero according to a pre-defined threshold to make the results sparse.

The 1.5Entmax is a wild variant of the α\alpha-Entmax sparse activation function family (Peters et al., 2019), with setting α\alpha to 1.5. It is a more general form that includes Sparsemax (α=1\alpha=1) and Softmax (α=2)(\alpha=2) as particular cases.

Rectified Linear Attention (ReLA) (Zhang et al., 2021) is an efficient sparse attention framework. Despite the Softmax activation function, they also use the ReLU function. They propose the RMS normalization with the gating mechanism to stable the training process and boost performance.