Cross Temporal Recurrent Networks for Ranking Question Answer Pairs
Yi Tay, Luu Anh Tuan, Siu Cheung Hui
Introduction
Learning-to-rank for QA (question answering) is a long standing problem in NLP and IR research which benefits a wide assortment of subtasks such as community-based question answering (CQA) and factoid based question answering. The problem is mainly concerned with computing relevance scores between questions and prospective answers and subsequently ranking them. Across the rich history of answer or document retrieval, statistical approaches based on feature engineering are commonly adopted. These models are largely based on complex lexical and syntactic features (?; ?; ?) and a learning-to-rank classifier such as Support Vector Machine (SVM) (?; ?).
Today, we see a shift into neural question answering. Specifically, end-to-end deep neural networks are used for both automatically learning features and scoring of QA pairs. Popular neural encoders for neural question answering include long short-term memory (LSTM) networks (?) and convolutional neural networks (CNN). The key idea behind neural encoders is to learn to compose (?), i.e., compressing an entire sentence into a single feature vector.
While it is possible to encode questions and answers independently, and later merge them with multi-layer perceptrons (MLP) (?), tensor layers (?) or holographic layers (?), it would be desirable for question and answer pairs to benefit from information available from their partner. There have been many models proposed for doing so which adopt techniques for jointly learning question and answer representations. Many of these recent techniques adopt soft-attention matching (?; ?; ?) to learn attention weights that are jointly influenced by both question and answer. Subsequently, the joint attention weights are applied accordingly to learn a final representation of question and answer. Performance results have shown that incorporating the interactions between QA pairs can indeed improve the performance of QA systems.
Temporal gates form the cornerstone of modern recurrent neural encoders such as long short-term memory (LSTM) or gated recurrent units (GRU), serving as one of the key mitigation strategies against vanishing gradients. In these models, temporal gates control the inner recursive loop along with the amount of information being discarded and retained at each time step, allowing fine-grained control over the semantic compositionality of learned representations. Our work explores the idea of jointly learning temporal gates for sequence pairs, aiming to learn fine-grained representations of QA pairs which benefit from information pertaining to what each other is remembering or forgetting.
The key idea here is as follows: By exploiting information about the question, can we learn an optimal way to semantically compose the answer? (and vice versa). First, consider the following example in Table 1 which highlights the importance of semantic compositionality.
First, it would be easy for many soft-attention and matching-based models to classify this question and answer pair with a high relevance score due to the underlined words (‘deep learning’ and ‘tensorflow’). However, the reason why this is a negative example is in the intricate details which can be effectively learned only via semantic compositionality (e.g., ‘without programming experience’, ‘get into’). On the other hand, by exploiting joint temporal gates, our method learns to compose the sentence given the information about its partner. For instance, without the knowledge of the answer which contains the phrase ‘advanced research’, the question encoder will not know if it should retain the word ‘beginner’. Hence, joint learning of temporal gates can help our model learn to compose, by influencing what it remembers and forgets. As a result, this additional knowledge can allow the words in boldface (‘beginner’,‘advanced research’) to be strongly retained in the final representation. This is in similar spirit to neural attention. However, our approach jointly learns to compose instead of learning to attend. The difference is at the level which representations are influenced at.
We introduce a new method for using temporal gates to synchronously and jointly learn the interactions between text pairs. In the context of question answering, we learn which information to remember or discard in the answer while being aware of the context of the question. To the best of our knowledge, this is the first work that performs question answer matching at the temporal gate level.
We propose a novel neural architecture for QA ranking. Our proposed Cross Temporal Recurrent Network (CTRN) model is largely inspired by the recently incepted Quasi Recurrent Neural Network (QRNN) (?) and can be considered as a natural extension of QRNN to sequence pairs. Our model takes after QRNN in the sense that gates are first learned (via 1D convolutional layers) and then subsequently applied to temporally adjust the representations. Hence, the facilitation of information flow can be interpreted as joint pairwise gating.
Our proposed CTRN model achieves state-of-the-art performance on two community-based QA (CQA) datasets, namely the Yahoo Answers dataset and the QatarLiving dataset from SemEval 2016. Moreover, our model also achieves highly competitive performance on the TrecQA dataset for factoid based QA. Experimental results show that CTRN outperforms models that utilize attention-based matching while being significantly more efficient. Experimental results also confirm that CTRN improves the underlying QRNN model.
Related Work
This section introduces prior work in the field of neural QA ranking. We also introduce the Quasi Recurrent Neural Network (QRNN) model, which lives at the heart of our proposed approach.
Convolutional neural network (CNN) (?; ?) and recurrent models like the long short-term memory (LSTM) (?) network are popular neural encoders for the QA ranking problem. (?) proposed to use CNN for learning features and subsequently apply logistic regression for generating QA relevance scores. Subsequently, an end-to-end neural architecture based on CNN and bilinear matching was proposed in (?).
In many recent works, the key innovation in most models is the technique used to model interaction between question and answer pairs. The CNN model introduced in (?) uses a MLP to compose vectors of questions and answers. (?) adopted tensor layers for richer modeling capabilities. (?) proposed holographic memory layers. (?) proposed Multi-Perspective CNN which matches ‘multiple perspectives’ based on variations of pooling and convolution schemes. Recent work has showed the effectiveness of learning QA embeddings in non-Euclidean spaces such as Hyperbolic space (?). Models based on soft-attention such as Attentive Pooling networks (AP-BiLSTM and AP-CNN) (?), AI-CNN (?) and aNMM (attention-based neural matching) (?) have also been proposed. These models learn weighted representations of QA pairs using similarity matrix based attentions.
Quasi Recurrent Neural Network (QRNN)
where is the cell state and is the hidden state. are the forget and output gates respectively at time step . is the sigmoid activation which nonlinearly projects each element of its input to $$. Z can be regarded as the convolved base representation similar to what a traditional CNN model learns. F and O are then applied recursively to temporally adjust and influence the semantic compositionality of Z. As such, this makes it ‘quasi-recurrent’. The key difference between QRNN and recurrent models like LSTM is that gates are prelearned via convolution while RNN models like LSTM learn their gates sequentially during the recursive forward operation. In short, the forward operation in QRNN is still sequentially applied but is comparatively much cheaper than traditional LSTM cells since gates are merely applied in the case of QRNN. As such, the parallelization of gate learning improves the speed of QRNN as compared to LSTM. In the original paper, QRNN achieved around 4 times less computational time as compared to LSTM models while achieving similar or better performance. For the sake of brevity, we refer interested readers to (?) for more details.
Inspired by the computational benefits of QRNN, we adopt it as our base model. Next, we also notice an attractive property of QRNN. In QRNN models, because gates are prelearned, it enables us to align temporal gates between two QRNNs easily. Conversely, considering the fact that questions and answers might not have similar sequence length, trying to sequentially align temporal gates in LSTM models can be extremely cumbersome and inefficient. More importantly, temporal gates of LSTM cells do not have global information, i.e., each step is only aware of all steps that precede it. On the other hand, temporal gates from QRNNs have global information about the entire sequence.
Our Proposed Approach
Our model accepts two sequences of indices (question and answer inputs) which are passed through an embedding layer (shared between and inputs) and returns a sequence of dimensional vectors. In practice, we initialize and fix this layer with pretrained embeddings while connecting to a projection layer. As such, the output of the embedding + projection layer is a dimensional vector. Note that this layer is shared between question and answer inputs.
Quasi-Recurrent Layer
Lightweight Temporal Crossing (LTC)
In this section, we use CTRN-Q as an example but note that CTRN-A and CTRN-Q are functionally symmetrical. Figure 2 illustrates a single CTRN-Q cell. Each CTRN-Q cell contains two cell states denoted as and and two hidden states denoted as and . As such, there are two learned representations in the CTRN-Q cell, denoted by and respectively. The first representation is learned as per normal, i.e., applying on . The second representation is learned by applying partner gates on the question representation . The following equations depict the forward operation of the CTRN-Q cell.
where denote the forget and output gates for text at time step . is an aligned time step between question and answer sequences as the sequence length of question and answer might be different. For simplicity, we consider . Similarly, the forward operation for CTRN-A is as follows:
Finally, to obtain a single representation for each question and answer. We simply apply the Hadamard product between hidden states of each time step, i.e., and . This enables joint representations of temporal gates which form the crux of our LTC mechanism.
Notably, since gates are learned via parameterized convolutional layers, our learned gates (F and O) not only contain ‘local’ index-specific information but also ‘global’ information of the entire text sequence. This is modeled by the parameters of the convolutional layers which produce F and O. As such, it would suffice to compose them index-wise since the goal is to enable information flow between the temporal gates of question and answer. Our intuition here is to cross apply question and answer gates to both question and answer representations so as to enable gradients flow across question and answer during back-propagation. Since the goal is to fuse and not to ‘match’, we empirically found that soft-attention alignment of gates to yield no performance benefits over a simple index-wise alignment.
Temporal Mean Pooling Layer
The output of each CTRN cell is an array of hidden states . In this layer, we apply temporal mean pooling for both CTRN-Q and CTRN-A. The operation of this layer is a simple element-wise average of all output hidden vectors.
Dense Layers (MLP)
The inputs of this layer are two vectors which are the final representations of question and answer respectively. In this layer, we concatenate the two vectors and pass them through a series of fully-connected dense layers (or MLP). Likewise, the number of layers is also a hyperparameter to be tuned.
Softmax Layer and Optimization
The final output of the hidden layer is then passed through a 2-class softmax layer. The final score of each QA pair is described as follows:
where is the output of the softmax layer. contains all the parameters of the network and is the L2 regularization. The parameters of the network are updated using the Adam Optimizer (?).
Complexity Analysis
In this section, we study the memory complexity of our model to further justify the lightweight aspect in our LTC mechanism. First, our CTRN does not incur any parameter cost over the vanilla QRNN model ( for a single QRNN model). As such, the memory complexity and parameter size remain equal to QRNN. This is easy to see as there is no additional parameters added since our model, from a computational graph perspective, is simply adding connections between nodes. Next, we consider the runtime complexity of our model (forward pass). Let be the number of filters of the convolution layer and be the maximum sequence length. The computational complexity of a single QRNN cell is excluding convolution operations used to generate . Though the number of operations is approximately doubled (due to cross applying gates), the complexity of a CTRN cell is still , i.e., our model still runs in linear time as compared to LSTM models with quadratic time complexity. Overall, our model, though seemingly more complicated, does not increase the parameter size and only incurs a slight increase in computational cost as compared to the already efficient QRNN model. We are also able to leverage the computational benefits of QRNN over the vanilla LSTM model. Table 2 shows a simple comparison of our proposed CTRN model against the standard LSTM and AP-BiLSTM models. We observe that QRNN and CTRN are much more parameter efficient as compared to recurrent models, only taking up the parameter size of the vanilla LSTM model and being smaller than AP-BiLSTM.
Experiments
To ascertain the effectiveness of our proposed approach, we conduct experiments on three popular benchmark datasets.
This section describes the datasets used, baselines compared and evaluation metrics.
We select three popular benchmark datasets which are described as follows:
YahooQA - Yahoo Answers is a CQA platform. This is a moderately large dataset containing QA pairs which are obtained from the CQA platform. More specifically, preprocessing and testing splitsSplits be obtained at https://github.com/vanzytay/YahooQA_Splits. are obtained from (?). In their setting, questions and answers that are not in the range of tokens are filtered. Additionally, negative samples are generated for each question by sampling from the top hits using Lucene search.
QatarLiving - This is another CQA dataset which was obtained from the popular SemEval-2016 Task 3 Subtask A (CQA). This is a real world dataset obtained from Qatar Living Forums. In this dataset, there are ten answers per thread (question) which are marked as ‘Good’, ‘Potentially Useful’ or ‘Bad’. Following (?), we treat ‘Good’ as positive and anything else as negative labels.
TrecQA - This is a popular QA ranking benchmark obtained from the TREC QA Tracks 8-13. QA pairs are generally short and factoid-based consisting trivia like questions. In this dataset, there are two training sets, namely TRAIN and TRAIN-ALL. TRAIN consists of QA pairs that have been manually judged and annotated. TRAIN-ALL is an automatically judged dataset of QA pairs and contains a larger number of QA pairs. TRAIN-ALL, being a larger dataset, also contains more noise. Nevertheless, both datasets enable the comparison of all models with respect to the availability and volume of training samples.
The statistics of all datasets, i.e., training sets, development sets and testing sets, are given in Table 3.
Evaluation Metrics
For each dataset, we adopt the evaluation metrics used in prior work. For YahooQA, we follow (?) that uses P@1 (Precision@1) and MRR (Mean Reciprocal Rank). For QatarLiving, we follow (?) and evaluate on P@1 and MAP (Mean Average Precision). For TrecQA, we follow the experiment procedure in (?) using the official evaluation metrics of MAP and MRR. Since the evaluation metrics are commonplace in ranking tasks, we omit any further details for the sake of brevity.
Implementation Details and Baselines
For our CTRN model, we tune the output dimension (number of filters) within $128\{10^{-3},10^{-4},10^{-5}\}\{64,128,256,512\}0.54\times 10^{-6}$. Word embedding matrices are all non-trainable and are learned by the projection layer instead. For the three datasets, we adopt dataset-specific baselines largely based on prior published works.
QatarLiving - The key competitors of this dataset are the CNN-based ARC-I/II architecture by Hu et al. (?), the Attentive Pooling CNN (?), Kelp (?) a feature engineering based SVM method, ConvKN (?) a combination of convolutional tree kernels with CNN and finally AI-CNN (Attentive Interactive CNN) (?), a tensor-based attentive pooling neural model. We initialize with pretrained GloVE embeddings of trained using the domain-specific unannotated corpus provided by the task.
TrecQA - We compare against published works which include both traditional models and neural models. Moreover, we compare with models reported in (?) on TRAIN and TRAIN-ALL datasets to observe the effect of different dataset sizes. The evaluation procedure follows (?) closely. We initialize the embedding layers with the same pretrained word embeddings of as (?) for fair comparisons against competitor approaches. These embeddings are trained with the Skip-gram model using the Wikipedia and AQUAINT corpus. Four word overlap features are also concatenated before the dense layers following (?). We train our model for epochs for TRAIN and epochs for TRAIN-ALL and report the test score from the best performing model on the development set. Hyperparameters are also tuned on the development set. Early stopping is adopted and training is terminated if the validation performance doesn’t improve after epochs.
Experimental Results
In this section, we report some observations pertaining to our empirical results.
Table 4 reports the experimental results on the YahooQA dataset. Firstly, we observe that our proposed CTRN achieves state-of-the-art performance on this dataset. Notably, we outperform HD-LSTM (?) by in terms of P@1 and in terms of MRR. CTRN also outperforms attention based models such as AP-BiLSTM and AP-CNN (?) by a considerable margin, i.e., of about . At this junction, we make several observations about our proposed CTRN model. Firstly, this shows that our LTC mechanism is more effective than soft-attention matching on this dataset. Secondly, the merits of this mechanism can be further observed by the performance difference in the QRNN and CTRN. Our proposed CTRN comfortably outperforms QRNN by in terms of P@1 and MRR. Surprisingly, we see that a simple baseline QRNN performs quite well on this dataset which outperforms other complex models such as NTN-LSTM and HD-LSTM (?).
Experimental Results on QatarLiving
Table 5 reports our experimental results on the QatarLiving dataset. Our CTRN model outperforms AI-CNNFor fair comparison, we compare against the reported results of AI-CNN that does not use handcrafted features. by in terms of P@1 while maintaining similar performance on MRR. The performance of the CTRN model also outperforms the baseline QRNN by on P@1. Similar to the Yahoo QA dataset, we also found that the baseline QRNN performed surprisingly well, i.e., outperforming ConvKN and other CNN based models such as ARC-I and ARC-II. Overall, our proposed approach achieves very competitive results on this dataset.
Results on TrecQA
Table 6 reports the results on TRAIN and TRAIN-ALL settings of the TrecQA task. CTRN achieves the top results comparing to the multitude of neural baselines. Notably, the performance of CTRN is about better than QRNN. QRNN performs very competitively on the TRAIN-ALL setting but fails in comparison for the TRAIN setting. This might be because QRNN, with three 1D convolutional layers, might overfit on the smaller dataset. However, CTRN performs well on smaller TRAIN as well which hints at possibly some regularizing effect of the LTC mechanism. The performance of the vanilla QRNN model on the TRAIN-ALL setting is also surprisingly competitive, outperforming more complex models such as HD-LSTM and NTN-LSTM.
Table 7 reports the results of the CTRN model against other published competitors. We can see that CTRN outperforms many complex neural architectures such as the aNMM model (?), HD-LSTM (?) and MP-CNN model (?; ?).
Runtime Comparison
Figure 3 shows the runtime comparison for recurrent models on the TRAIN-ALL dataset. We observe that CTRN is a very scalable and efficient model. Notably, our CTRN model benefits from the training speed brought from the QRNN model, which is clearly significantly faster than LSTM models. Moreover, we also show that CTRN does not significantly increase the runtime of the base QRNN, only incurring an additional per epoch. Moreover, we achieve times faster runtime compared to vanilla LSTM models and times faster than AP-BiLSTM.
Conclusion
We introduced a novel method for jointly learning to compose QA pairs. This is achieved by aligning temporal gates. We show that our lightweight temporal crossing (LTC) mechanism is an effective method of modeling interactions between QA pairs without incurring any parameter cost. Our CTRN model performs competitively on two CQA benchmarks and one factoid QA benchmark while being much faster than LSTM and AP-BiLSTM models.
Acknowledgements
The authors thank anonymous reviewers for their hardwork and feedback.