Rational Recurrences
Hao Peng, Roy Schwartz, Sam Thomson, Noah A. Smith
Introduction
Neural models, and in particular gated variants of recurrent neural networks (RNNs, e.g., Hochreiter and Schmidhuber, 1997; Cho et al., 2014), have become a core building block for state-of-the-art approaches in NLP Goldberg (2016). While these models empirically outperform classical NLP methods on many tasks (Zaremba et al., 2014; Bahdanau et al., 2015; Dyer et al., 2016; Peng et al., 2017, inter alia), they typically lack the intuition offered by classical models, making it hard to understand the roles played by each of their components. In this work we show that many neural models are more interpretable than previously thought, by drawing connections to weighted finite state automata (WFSAs). We study several recently proposed RNN architectures and show that one can use WFSAs to characterize their recurrent updates. We call such models rational recurrences (§3). Where the term regular is used with unweighted FSAs (e.g., regular languages, regular expressions), rational is the weighted analog (e.g., rational series, Sakarovitch, 2009; rational kernels, Cortes et al., 2004). Analyzing recurrences in terms of WFSAs provides a new view of existing models and facilitates the development of new ones.
In recent work, Schwartz et al. (2018) introduced SoPa, an RNN constructed from WFSAs, and thus rational by our definition. They also showed that a single-layer max-pooled CNN LeCun (1998) can be simulated by a set of simple WFSAs (one per output dimension), and accordingly are also rational. In this paper we broaden such efforts, and show that rational recurrences are in frequent use (Mikolov et al., 2014; Balduzzi and Ghifary, 2016; Lei et al., 2016, 2017a, 2017b; Bradbury et al., 2017; Foerster et al., 2017). For instance, we will show in §4 that the WFSA diagrammed in Figure 1 has strong connections to several of the models mentioned above.
Based on these observations, we then discuss potential approaches to deriving novel neural architectures from WFSAs (§5). As a case study, we present a new model motivated by the interpolation of a two-state WFSA and a three-state one, capturing (soft) unigram and bigram features, respectively. Our experiments show that in two tasks—language modeling and text classification—the proposed model outperforms recently proposed rational models (§6). Further extensions might lead to larger gains, and the rational recurrence view could facilitate easier exploration of such extensions. To promote such exploration, we publicly release our implementation at https://github.com/Noahs-ARK/rational-recurrences.
Background: Weighted Finite State Automata (WFSAs)
This section reviews weighted finite-state automata and semirings, which underly our analyses in §3. WFSAs extend nondeterministic unweighted finite-state automata by assigning weights to transitions, start states, and final states. Instead of simply accepting or rejecting a string, a WFSA returns a score for the string, and this score summarizes the weights along all paths through the WFSA that consume the string. In order for this summary score to be efficiently computable, weights are taken from a semiring.
marks special /̄transitions that may be taken without consuming any input. assigns a score to a string by summing over the scores of all possible paths deriving . The score of each individual path is the product of the weights of the transitions it consists of. Formally:
Let be a sequence of adjacent transitions in , with each transition . The path derives string , which is the substring of that excludes symbols (for example, if , then ). ’s score in is given by
Let denote the set of all paths in that derive . Then the score assigned by to is defined to be
gives the total score of all paths that derive and end in state .
Figure 1 diagrams a WFSA , consisting of two states. A path starts from the initial state (with ); it then takes any number of “self-loop” transitions, each consuming an input without changing the path score (since it’s weighted by ); it then consumes an input symbol and takes a transition weighted by , and reaches the final state (with ); it may further consume more input by taking self-loops at , updating the path score by multiplying it by for each symbol . Then from Definition 4, we can calculate that gives the empty string score , and gives any nonempty string score
can be seen as capturing soft unigram patterns (Davidov et al., 2010), in the sense that it consumes one input symbol to reach the final state from the initial state. It is straightforward to design WFSAs capturing longer patterns by including more states (Schwartz et al., 2018), as we will discuss later in §4 and §5.
Rational Recurrences
Before formally defining rational recurrences in §3.2, we highlight the connection between WFSAs and RNNs using a motivating example (§3.1).
We describe a simplified RNN which strips away details of some recent RNNs, in order to highlight the behaviors of the forget gate and the input.
For an input sequence , let the word embedding vector for be . As in many gated RNN variants (Hochreiter and Schmidhuber, 1997; Cho et al., 2014), we use a forget gate , which is computed with an affine transformation followed by an elementwise sigmoid function . The current input representation is similarly computed, but with an optional nonlinearity (e.g., ) . The hidden state can be seen as a weighted sum of the previous state and the new input, controlled by the forget gate.
The hidden state can then be used in downstream computation, e.g., to calculate output state , which is then fed to an MLP classifier. We focus only on the recurrent computation.
In Example 6, both and depend only on the current input token (through ), and not the previous state. Importantly, the interaction with the previous state is not via affine transformations followed by nonlinearities, as in, e.g., an Elman network (Elman, 1990), where . As we will discuss later, this is important in relating this recurrent update function to WFSAs.
Since the recurrent update in Equation 5c is elementwise, for simplicity we focus on just the th dimension. Unrolling it in time steps, we get
where denotes the th dimension of a vector. As noted by Lee et al. (2017), the hidden state at time step can be seen as a sum of previous input representations, weighted by the forget gate; longer histories typically get a smaller weight, since the forget gate values are between 0 and 1 due to the sigmoid function.
Denote the resulting WFSA by , and we have:
Running a single layer RNN in Example 6 over any nonempty input string , the th dimension of its hidden state at time step equals the score assigned by to :
In other words, the th dimension of the RNN in Example 6 can be seen as a WFSA structurally equivalent to . Its weight functions are implemented as the th dimension of Equations 5, and the learned parameters are the th row of and . Then it is straightforward to recover the full -dimensional RNN, by collecting such WFSAs, each of which is parametrized by a row in the s and s. Based on this observation, we are now ready to formally define rational recurrences.
2 Recurrences and Rationality
It directly follows from Proposition 7 that
Relationship to Existing Neural Models
This section studies several recently proposed neural architectures, and relates them to rational recurrences. §4.1 begins by relating some of them to the RNN defined in Example 6, and then to the WFSA (Example 5). We then describe a WFSA similar to , but with one additional state, and discuss how it provides a new view of RNN models motivated by -gram features (§4.2). In §4.3 we study rational recurrences that are not elementwise, using an existing model.
In the following discussion, we shall assume the real seimiring, unless otherwise noted.
Despite its simplicity, Example 6 corresponds to several existing neural architectures. For instance, quasi-RNN (QRNN; Bradbury et al., 2017) and simple recurrent unit (SRU; Lei et al., 2017b) aim to speed up the recurrent computation. To do so, they drop the matrix multiplication dependence on the previous hidden state, resulting in similar recurrences to that in Example 6.The SRU architecture discussed through this work is based on Lei et al. (2017b). In a later updated version, Lei et al. (2018) introduce diagonal matrix multiplication interaction in the hidden state updates, inspired by (Li et al., 2018), which yields a recurrence not obviously rational. Other works start from different motivations, but land on similar recurrences, e.g., strongly-typed RNNs (T-RNN; Balduzzi and Ghifary, 2016) and its gated variants (T-LSTM and T-GRU), and structurally constrained RNNs (SCRN; Mikolov et al., 2014).
The analysis in §3.1 directly applies to SRU, T-RNN, and SCRN. In fact, Example 6 presents a slightly more complicated version of them. In these models, input representations are computed without the bias term or any nonlinearity: . By Proposition 7 and Corollary 9:
The recurrences of single-layer SRU, T-RNN, and SCRN architectures are rational.
It is slightly more complicated to analyze the recurrences of the QRNN, T-LSTM, and T-GRU. Although their hidden states are updated in the same way as Equation 5c, the input representations and gates may depend on previous inputs. For example, in T-LSTM and T-GRU, the forget gate is a function of two consecutive inputs:
QRNNs are similar, but may depend on up to tokens, due to the -window convolutions. Eisner (2002) discuss finite state machines for second (or higher) order probabilistic sequence models. Following the same intuition, we sketch the construction of WFSAs corresponding to QRNNs with 2-window convolutions in Appendix A, and summarize the key results here:
The recurrences of single-layer T-GRU, T-LSTM, and QRNN are rational. In particular, a single-layer -dimensional QRNN using -window convolutions can be recovered by a set of WFSAs, each with states.
The size of WFSAs needed to recover QRNN grows exponentially in the window size. Therefore, at least for QRNNs, Proposition 11 has more conceptual value than practical.
2 More than Two States
So far our discussion has centered on , a two-state WFSA capturing unigram patterns (Example 5). In the same spirit as going from unigram to -gram features, one can use WFSAs with more states to capture longer patterns (Schwartz et al., 2018). In this section we augment by introducing more states, and explore its relationship to some neural architectures motivated by -gram features. We start with a three-state WFSA as an example, and then discuss more general cases.
Figure 2 diagrams a WFSA , augmenting with another state. To reach the final state , at least two transitions must be taken, in contrast to one in . History information is decayed by the self-loop at the final state , assuming is between 0 and 1. has another self-loop over , weighted by . The motivation is to allow (but down-weight) nonconsecutive bigrams, as we will soon show.
The scores assigned by can be inductively computed by applying the Forward algorithm (§2). Given input sequence longer than one, let , then
and . Unrolling in time, we get
Due to the self-loop over state , can be seen as a weighted sum of the terms up to (Equaltion 14). The second product term in Equation 12 then provides multiplicative interactions between , and the weighted sum of s. In this sense, it captures nonconsecutive bigram features.
At a first glance, Equations 12 and 13 resemble recurrent convolutional neural networks (RCNN; Lei et al., 2016). RCNN is inspired by nonconsecutive -gram features and low rank tensor factorization. It is later studied from a string kernel perspective (Lei et al., 2017a). Here we review its nonlinear bigram version:
where the s are computed similarly to Equation 5b, and is used as output for onward computation. Different strategies to computing were explored (Lei et al., 2015, 2016). When is a constant, or depends only on , e.g., , the th dimension of Equations 15 can be recovered from Equation 12, by letting
It is straightforward to generalize the above discussion to higher order cases: -gram RCNN corresponds to WFSAs with states, constructed similarly to how we build from (Appendix B).
For a single-layer RCNN with being a constant or depending only on , the recurrence is rational.
As noted later in §4.3, its recurrence may not be rational when .
3 Beyond Elementwise Operations
So far we have discussed rational recurrences for models using elementwise recurrent updates (e.g., Equation 5c). This section uses an existing model as an example, to study a rational recurrence that is not elementwise. We focus on the input switched affine network (ISAN; Foerster et al., 2017). Aiming for efficiency and interpretability, it does not use any explicit nonlinearity; its affine transformation parameters depend only on the input:
Due to the matrix multiplication, the recurrence of a single-layer ISAN is not elementwise. Yet, we argue that it is rational. We will sketch the proof for a 2-dimensional case, and it is straightforward to generalize to higher dimensions (Appendix C).
We define two WFSAs, each recovering one dimension of ISAN’s recurrent updates. Figure 3 diagrams one of them, . The other one, , is identical (including shared weights), except using instead of as the final state. For any nonempty input sequence , the scores assigned by and can be inductively computed by applying the Forward algorithm. Letting , for
Then Equation 17, in the case of hidden size 2, is recovered by letting and .
The recurrence of a single-layer ISAN is rational.
For a single-layer Elman network, in the absence of any nonlinearity, the recurrence is rational.
It is known that an Elman network can approximate any recursively computable partial function (Siegelmann and Sontag, 1995). On the other hand, in their single-layer cases, WFSAs (and thus models with rational recurrences) are restricted to rational series (Schützenberger, 1961). Therefore, we hypothesize that models like Elman networks, LSTMs, and GRUs, where the recurrences depend on previous states through affine transformations followed by nonlinearities, are not rational.
This work does not intend to propose rational recurrences as a concept general enough to include most existing RNNs. Rather, we wish to study a more constrained class of methods to better understand the connections between WFSAs and RNNs. Therefore in Definition 8, we restrict the semirings to be “simple,” in the sense that both operations take constant time and space. Such a restriction aims to exclude the possibility of hiding arbitrarily complex computations inside the semiring, which might allow RNNs to satisfy the definition in a trivial and unilluminating way.
Such theoretical limitations might be less severe than they appear, since it is not yet entirely clear what they correspond to in practice, especially when multiple vertical layers of these models are used (Leshno and Schocken, 1993). We defer to future work the further study of the connections between WFSAs and Elman-style RNNs.
Closing this section, Table 1 summarizes the discussed recurrent neural architectures and their corresponding WFSAs.
Deriving Neural Models from WFSAs
Rational recurrences provide a new view of several recently proposed neural models. Based on such observations, this section aims to explore potential approaches to designing neural architectures in a more interpretable and intuitive way: by deriving them from WFSAs. §5.1 studies an interpolation of unigram and bigram features by combining 2-state and 3-state WFSAs (Figures 1 and 2). We then explore alternative semirings (§5.2), an approach orthogonal to what we’ve discussed so far.
We note that our goal is not to devise new state-of-the-art architectures. Rather, we illustrate a new design process for neural architectures that draws inspiration from WFSAs. That said, in our experiments (§6), one of our new architectures performs as well as or better than strong baselines.
We start by presenting a straightforward extension to 2-state and 3-state rational models: one combining both. It is inspired by many classical NLP models, where unigram features and higher-order ones are interpolated.
Figure 4 diagrams a 4-state WFSA . Compared to (Figure 2), uses as a second final state, aiming to capture both unigram and bigram patterns, since a path is allowed to stop at after consuming one input. The final states are weighted by and respectively. Another notable modification is the additional state , which is used to create a “shortcut” to reach , together with an -transition. Specifically, starting from , a path can now take the -transition and reach , and then take a transition with weight to reach . Recall from §2, that -transitions do not consume any input, yet they can still be weighted by a (parameterized) function not depending on the inputs. The -transition allows for skipping the first word in a bigram. It can be discouraged by using , just as we do in our experiments.
As in §3, we relate hidden states of an RNN to the scores assigned by WFSAs to input strings. We then derive the neural architecture with a dynamic program. Here we keep the discussion self-contained by explicitly overviewing the procedure. It is a direct application of the Forward algorithm (§2), though now in a form that deals with the -transition. Such an approach applies, of course, to more general cases, as noted by Schwartz et al. (2018).
Given an input string , let denote the total score of all paths landing in state just after consuming . Let , then for ,
We now collect of these WFSAs to construct an RNN, and we parameterize their weight functions with the technique we’ve been using:
The vectors correspond to the final state weights and . Despite the similarities, are different from output gates (Bradbury et al., 2017), since the former do not depend on the input, and are parameterized (through a sigmoid) by two leanred vectors . The same applies to and , which correspond to the weights for -transitions .
2 Alternative Semirings
Example 15 does not use the forget gate when computing (Equation 22b), which is different from its plus-times counterpart, where \mathbf{u}_{t}=(\mathbf{1}-\mathbf{f}_{t})\odot\bm{g}\bigl{(}\mathbf{W}_{u}\mathbf{v}_{t}+\mathbf{b}_{u}\bigr{)}. The reason is that, unlike the real semiring, the max-plus semiring lacks a well-defined negation. Possible alternatives include taking the of a separate input gate, or using , which we leave for future work.
Example 15 can be seen as replacing sum-pooling with max-pooling. Both max and sum-pooling have been used successfully in vision and NLP models. Intuitively, max-pooling “detects” the occurrence of a pattern while sum-pooling “counts” the occurrence of a pattern. One advantage of max operator is that the model’s decisions can be back-traced and interpreted, as argued by Schwartz et al. (2018). Such a technique is applicable to all the models with rational recurrences.
Experiments
This section evaluates four rational RNNs on language modeling (§6.2) and text categorization (§6.3). Our goal is to compare the behaviors of models derived from different WFSAs, showing that our understanding of WFSAs allows us to improve existing rational models.
Our comparisons focus on the recurrences of the models, i.e., how the hidden states are computed (e.g., Equations 5c and 20c). Therefore we follow Lei et al. (2017b) and use across all compared models, listed below and as well as in Table 2:
rrnn(), with real semiring (§4.1);
rrnn(), with max-plus semiring (§5.2);
rrnn(), with real semiring (§4.2);
rrnn(), with real semiring (§5.1).
We also compare to an LSTM baseline. Aiming to control for comfounding factors, we do not use highway connections in any of the models.Thus rrnn() is essentially an SRU without highway connections. We denote it differently, to note its differences from the original implementation Lei et al. (2017b). Similarly, we do not denote rrnn() as RCNN (Lei et al., 2016). In the interest of space, the full architectures and hyperparameters are detailed in Appendices D and E.
2 Language Modeling
We experiment with the Penn Treebank corpus (PTB; Marcus et al., 1993). We use the preprocessing and splits from Mikolov et al. (2010), resulting in a vocabulary size of 10K and 1M tokens.
Following standard practice, we treat the training data as one long sequence, split into mini batches, and train using BPTT truncated to 35 time steps (Williams and Peng, 1990). The input embeddings and output softmax weights are tied (Press and Wolf, 2017).
Results.
Following Collins et al. (2017) and Melis et al. (2018), we compare models controlling for parameter budget. Table 3 summarizes language modeling perplexities on PTB test set. The middle block compares all models with two layers and 10M trainable parameters. rrnn() and rrnn() achieve roughly the same performance; interpolating both unigram and bigram features, rrnn() outperforms others by more than 2.9 test perplexity. For the three-layer and 24M setting (the bottom block), we observe similar trends, except that rrnn() slightly underperforms rrnn(). Here rrnn() outperforms others by more than 2.1 perplexity.
Using a max-plus semiring, rrnn() underperforms rrnn() under both settings. Possible reasons could be the suboptimal design choice for computing input representations in the former (§5.2). Finally, most compared models outperform the LSTM baselines, whose numbers are taken from Lei et al. (2017b).Melis et al. (2018) point out that carefully tuning LSTMs can achieve much stronger performance, at the cost of exceptionally large amounts of computational resources for tuning.
3 Text Classification
We use unidirectional 2-layer architectures for all compared models. To build the classifiers, we feed the final RNN hidden states into a 2-layer -MLP. Further implementation details are described in Appendix E.
Datasets.
We experiment with four binary text classification datasets, described below.
Amazon (electronic product review corpus; McAuley and Leskovec, 2013).http://riejohnson.com/cnn_data.html We focus on the positive and negative reviews.
SST (Stanford sentiment treebank; Socher et al., 2013).nlp.stanford.edu/sentiment/index.html We focus on the binary classification task. SST provides labels for syntactic phrases; we experiment with a more realistic setup, and consider only complete sentences at either training or evaluating time.
subj (Subjectivity dataset; Pang and Lee, 2004). As subj doesn’t come with official splits, we randomly split it to train (80%), development (10%), and test (10%) sets.
CR (customer reviews dataset; Hu and Liu, 2004).http://www.cs.uic.edu/?liub/FBS/sentiment-analysis.html As with subj, we randomly split this dataset using the same ratio.
Table 4 summarizes the sizes of the datasets.
Results.
Table 5 summarizes text classification test accuracy. We report the average performance of 5 trials different only in random seeds. rrnn() outperforms all other models on 3 out of the 4 datasets. For Amazon, the largest one, we do not observe significant differences between rrnn() and rrnn(), while both outperform others. This may suggest that the interpolation of unigram and bigram features by rrnn() is especially useful in small data setups. As in the language modeling experiments, rrnn() underperforms all other models in most cases, and in particular rrnn(). These results provide evidence that replacing the real semiring in rational models might be challenging. We leave further exploration to future work.
Related Work
WFSAs were once popular among many sequential tasks (Mohri et al., 2002; Kumar and Byrne, 2003; Cortes et al., 2004; Pardo and Birmingham, 2005; Moore et al., 2006, inter alia), and are still successful in morphology (Dreyer, 2011; Cotterell et al., 2015; Rastogi et al., 2016, inter alia). Compared to neural networks, WFSAs are better understood theoretically and arguably more interpretable. They were recently revisited in combination with the former in, e.g., text generation (Ghazvininejad et al., 2016, 2017; Lin et al., 2017) and automatic music accompaniment (Forsyth, 2016).
Recurrent neural networks.
RNNs (Elman, 1990; Jordan, 1989) prove to be strong models for sequential data (Siegelmann and Sontag, 1995). Besides the perhaps most notable gated variants (Hochreiter and Schmidhuber, 1997; Cho et al., 2014), extensive efforts have been devoted to developing alternatives (Balduzzi and Ghifary, 2016; Miao et al., 2016; Zoph and Le, 2017; Lee et al., 2017; Lei et al., 2017a; Vaswani et al., 2017; Gehring et al., 2017, inter alia). Departing from the above approaches, this work derives RNN architectures drawing inspiration from WFSAs.
Another line of work studied the connections between WFSAs and RNNs in terms of modeling capacity, both empirically (Kolen, 1993; Giles et al., 1992; Weiss et al., 2018, inter alia) and theoretically (Cleeremans et al., 1989; Visser et al., 2001; Chen et al., 2018, inter alia).
Conclusion
We presented rational recurrences, a new construction to study the recurrent updates in RNNs, drawing inspiration from WFSAs. We showed that rational recurrences are in frequent use by several recently proposed recurrent neural architectures, providing new understanding of them. Based on such connections, we discussed approaches to deriving novel neural architectures from WFSAs. Our empirical results demonstrate the potential of doing so. We publicly release our implementation at https://github.com/Noahs-ARK/rational-recurrences.
Acknowledgments
We thank Jason Eisner, Luheng He, Tao Lei, Omer Levy, members of the ARK lab at the University of Washington, and researchers at the Allen Institute for Artificial Intelligence for their helpful comments on an early version of this work, and the anonymous reviewers for their valuable feedback. We also thank members of the Aristo team at the Allen Institute for Artificial Intelligence for their support with the Beaker experimentation system. This work was supported in part by NSF grant IIS-1562364 and by the NVIDIA Corporation through the donation of a Tesla GPU.
References
Appendix A Proof of Proposition 11
Let’s consider a single-layer QRNN with 2-window convolutions:
A similar analysis applies to T-GRUs and T-LSTMs directly, and it should be straightforward to generalize the discussion to QRNNs with larger convolution windows.
Let denote the alphabet, and let be a nonempty input string. consider a WFSA over the real semiring with states, where is the initial state with ; of them are final states , with , and the remaining states are denoted by .
The transition weights are constructed by
otherwise. Then one dimension of the reccurent updates of a 2-window QRNN is recovered by parameterizing the weight functions as
The recurrent computation of a 2-window QRNN of hidden size can then be recovered by collecting such WFSAs. ∎
Appendix B Proof of Proposition 12
We present the construction of WFSAs for a single layer -gram RCNNs of hidden size .
Let’s assume a given input sequence , with , since otherwise one only needs include paddings, just as in a RCNN. Consider a WFSA with states over the real semiring. Use as the initial state with , and as the final state with . The transition weight function is defined by
Let denote the total score of all paths landing in state just after consuming . Let . By the forward algorithm
Applying similar parametrization to that in §4.2, recovers one dimension of the recurrence. Collecting such WFSAs we recover the recurrence of a single layer -gram RCNNs, with being a constant, or depending only on . ∎
Appendix C Proof of Proposition 13
Closely following the 2-dimensional case in §4.3, let’s discuss a single layer ISAN of hidden size .
Consider a WFSA over the real semiring with states. Let of them, denoted by be the initial states, with . Denote the other half by . Define transition weight by:
Using as the final state with , and denote the resulting WFSA by . By Forward algorithm, recovers the th dimension of the single layer ISAN by letting , and ; the -dimensional recurrent computation is recovered by a set of WFSAs constructed similarly. ∎
Appendix D Compared Models
This section formally describes the models compared in the experiments (§6.1).
rrnn() is derived from (§4.1).
rrnn(ℬℬ\mathscr{B})m+m+{}_{\text{m+}}.
Also derived from , but uses the max-plus semiring (§5.2).
rrnn(𝒞𝒞\mathscr{C}).
rrnn() is derived from (§4.2):
rrnn(ℱℱ\mathscr{F}).
The output gates (Equations 25d, 26d, 27e, and 28h) are optional. They are only used in language modeling experiments, where we empirically find that they improve performance.
Appendix E Experimental Setup
Our implementation is based on Lei et al. (2017b)https://github.com/taolei87/sru and Peng et al. (2018),https://github.com/Noahs-ARK/SPIGOT using PyTorch.https://pytorch.org/
E.2 Language Modeling
For hyperparameters, we do not deviate much from the language modeling experiments in Lei et al. (2017b). We change the hidden sizes for all compared models based on the trainable parameter budget, and adjust the dropout probabilities accordingly to keep the number of remaining hidden units is roughly the same in expectation. Besides, we observe that rrnn() and rrnn() fail to converge when optimized with the SGD algorithm using 1.0 initial learning rate. And thus we use 0.5 for both models. Other hyperparameters are kept the same as Lei et al. (2017b).
E.3 Text classification
We tune the hyperparameters of our model on the development set by running 20 epochs of random search. We then take the best development configuration, and train five models with it using different random seeds. We report the average test results. The hyperparameters values explored are summarized in Table 6. We train all models for 500 epochs, stopping early if development accuracy does not improve for 30 epochs. During training, we halve the learning rate if development accuracy does not improve for 10 epochs.