Multi-task Learning of Pairwise Sequence Classification Tasks Over Disparate Label Spaces
Isabelle Augenstein, Sebastian Ruder, Anders Søgaard
Introduction
Multi-task learning (MTL) and semi-supervised learning are both successful paradigms for learning in scenarios with limited labelled data and have in recent years been applied to almost all areas of NLP. Applications of MTL in NLP, for example, include partial parsing Søgaard and Goldberg (2016), text normalisation Bollman et al. (2017), neural machine translation Luong et al. (2016), and keyphrase boundary classification (Augenstein and Søgaard, 2017).
Contemporary work in MTL for NLP typically focuses on learning representations that are useful across tasks, often through hard parameter sharing of hidden layers of neural networks Collobert et al. (2011); Søgaard and Goldberg (2016). If tasks share optimal hypothesis classes at the level of these representations, MTL leads to improvements Baxter (2000). However, while sharing hidden layers of neural networks is an effective regulariser Søgaard and Goldberg (2016), we potentially loose synergies between the classification functions trained to associate these representations with class labels. This paper sets out to build an architecture in which such synergies are exploited, with an application to pairwise sequence classification tasks. Doing so, we achieve a new state of the art on topic-based sentiment analysis.
For many NLP tasks, disparate label sets are weakly correlated, e.g. part-of-speech tags correlate with dependencies Hashimoto et al. (2017), sentiment correlates with emotion Felbo et al. (2017); Eisner et al. (2016), etc. We thus propose to induce a joint label embedding space (visualised in Figure 2) using a Label Embedding Layer that allows us to model these relationships, which we show helps with learning.
In addition, for tasks where labels are closely related, we should be able to not only model their relationship, but also to directly estimate the corresponding label of the target task based on auxiliary predictions. To this end, we propose to train a Label Transfer Network (LTN) jointly with the model to produce pseudo-labels across tasks.
The LTN can be used to label unlabelled and auxiliary task data by utilising the ‘dark knowledge’ Hinton et al. (2015) contained in auxiliary model predictions. This pseudo-labelled data is then incorporated into the model via semi-supervised learning, leading to a natural combination of multi-task learning and semi-supervised learning. We additionally augment the LTN with data-specific diversity features Ruder and Plank (2017) that aid in learning.
Our contributions are: a) We model the relationships between labels by inducing a joint label space for multi-task learning. b) We propose a Label Transfer Network that learns to transfer labels between tasks and propose to use semi-supervised learning to leverage them for training. c) We evaluate MTL approaches on a variety of classification tasks and shed new light on settings where multi-task learning works. d) We perform an extensive ablation study of our model. e) We report state-of-the-art performance on topic-based sentiment analysis.
Related work
Existing approaches for learning similarities between tasks enforce a clustering of tasks Evgeniou et al. (2005); Jacob et al. (2009), induce a shared prior Yu et al. (2005); Xue et al. (2007); Daumé III (2009), or learn a grouping Kang et al. (2011); Kumar and Daumé III (2012). These approaches focus on homogeneous tasks and employ linear or Bayesian models. They can thus not be directly applied to our setting with tasks using disparate label sets.
Multi-task learning with neural networks
Recent work in multi-task learning goes beyond hard parameter sharing (Caruana, 1993) and considers different sharing structures, e.g. only sharing at lower layers (Søgaard and Goldberg, 2016) and induces private and shared subspaces (Liu et al., 2017; Ruder et al., 2017). These approaches, however, are not able to take into account relationships between labels that may aid in learning. Another related direction is to train on disparate annotations of the same task Chen et al. (2016); Peng et al. (2017). In contrast, the different nature of our tasks requires a modelling of their label spaces.
Semi-supervised learning
There exists a wide range of semi-supervised learning algorithms, e.g., self-training, co-training, tri-training, EM, and combinations thereof, several of which have also been used in NLP. Our approach is probably most closely related to an algorithm called co-forest Li and Zhou (2007). In co-forest, like here, each learner is improved with unlabeled instances labeled by the ensemble consisting of all the other learners. Note also that several researchers have proposed using auxiliary tasks that are unsupervised Plank et al. (2016); Rei (2017), which also leads to a form of semi-supervised models.
Label transformations
The idea of manually mapping between label sets or learning such a mapping to facilitate transfer is not new. Zhang et al. (2012) use distributional information to map from a language-specific tagset to a tagset used for other languages, in order to facilitate cross-lingual transfer. More related to this work, Kim et al. (2015) use canonical correlation analysis to transfer between tasks with disparate label spaces. There has also been work on label transformations in the context of multi-label classification problems Yeh et al. (2017).
Multi-task learning with disparate label spaces
In our multi-task learning scenario, we have access to labelled datasets for tasks at training time with a target task that we particularly care about. The training dataset for task consists of examples and their labels . Our base model is a deep neural network that performs classic hard parameter sharing Caruana (1993): It shares its parameters across tasks and has task-specific softmax output layers, which output a probability distribution for task according to the following equation:
The MTL model is then trained to minimise the sum of the individual task losses:
where is the negative log-likelihood objective and is a parameter that determines the weight of task . In practice, we apply the same weight to all tasks. We show the full set-up in Figure 1(a).
2 Label Embedding Layer
In order to learn the relationships between labels, we propose a Label Embedding Layer (LEL) that embeds the labels of all tasks in a joint space. Instead of training separate softmax output layers as above, we introduce a label compatibility function that measures how similar a label with embedding is to the hidden representation :
where is the dot product. This is similar to the Universal Schema Latent Feature Model introduced by Riedel et al. (2013). In contrast to other models that use the dot product in the objective function, we do not have to rely on negative sampling and a hinge loss Collobert and Weston (2008) as negative instances (labels) are known. For efficiency purposes, we use matrix multiplication instead of a single dot product and softmax instead of sigmoid activations:
3 Label Transfer Network
The LEL allows us to learn the relationships between labels. In order to make use of these relationships, we would like to leverage the predictions of our auxiliary tasks to estimate a label for the target task. To this end, we introduce the Label Transfer Network (LTN). This network takes the auxiliary task outputs as input. In particular, we define the output label embedding of task as the sum of the task’s label embeddings weighted with their probability :
where designates concatenation. The mapping of the tasks in the LTN yields another signal that can be useful for optimisation and act as a regulariser. The LTN can also be seen as a mixture-of-experts layer Jacobs et al. (1991) where the experts are the auxiliary task models. As the label embeddings are learned jointly with the main model, the LTN is more sensitive to the relationships between labels than a separately learned mixture-of-experts model that only relies on the experts’ output distributions. As such, the LTN can be directly used to produce predictions on unseen data.
4 Semi-supervised MTL
The downside of the LTN is that it requires additional parameters and relies on the predictions of the auxiliary models, which impacts the runtime during testing. Instead, of using the LTN for prediction directly, we can use it to provide pseudo-labels for unlabelled or auxiliary task data by utilising auxiliary predictions for semi-supervised learning.
We train the target task model on the pseudo-labelled data to minimise the squared error between the model predictions and the pseudo labels produced by the LTN:
We add this loss term to the MTL loss in Equation 2. As the LTN is learned together with the MTL model, pseudo-labels produced early during training will likely not be helpful as they are based on unreliable auxiliary predictions. For this reason, we first train the base MTL model until convergence and then augment it with the LTN. We show the full semi-supervised learning procedure in Figure 1(c).
5 Data-specific features
When there is a domain shift between the datasets of different tasks as is common for instance when learning NER models with different label sets, the output label embeddings might not contain sufficient information to bridge the domain gap.
To mitigate this discrepancy, we augment the LTN’s input with features that have been found useful for transfer learning Ruder and Plank (2017). In particular, we use the number of word types, type-token ratio, entropy, Simpson’s index, and Rényi entropy as diversity features. We calculate each feature for each example.For more information regarding the feature calculation, refer to Ruder and Plank (2017). The features are then concatenated with the input of the LTN.
6 Other multi-task improvements
Hard parameter sharing can be overly restrictive and provide a regularisation that is too heavy when jointly learning many tasks. For this reason, we propose several additional improvements that seek to alleviate this burden: We use skip-connections, which have been shown to be useful for multi-task learning in recent work Ruder et al. (2017). Furthermore, we add a task-specific layer before the output layer, which is useful for learning task-specific transformations of the shared representations Søgaard and Goldberg (2016); Ruder et al. (2017).
Experiments
For our experiments, we evaluate on a wide range of text classification tasks. In particular, we choose pairwise classification tasks—i.e. those that condition the reading of one sequence on another sequence—as we are interested in understanding if knowledge can be transferred even for these more complex interactions. To the best of our knowledge, this is the first work on transfer learning between such pairwise sequence classification tasks. We implement all our models in Tensorflow Abadi et al. (2016) and release the code at https://github.com/coastalcph/mtl-disparate.
We use the following tasks and datasets for our experiments, show task statistics in Table 1, and summarise examples in Table 2:
Topic-based sentiment analysis aims to estimate the sentiment of a tweet known to be about a given topic. We use the data from SemEval-2016 Task 4 Subtask B and C Nakov et al. (2016) for predicting on a two-point scale of positive and negative (Topic-2) and five-point scale ranging from highly negative to highly positive (Topic-5) respectively. An example from this dataset would be to classify the tweet “No power at home, sat in the dark listening to AC/DC in the hope it’ll make the electricity come back again” known to be about the topic “AC/DC”, which is labelled as a positive sentiment. The evaluation metrics for Topic-2 and Topic-5 are macro-averaged recall () and macro-averaged mean absolute error () respectively, which are both averaged across topics.
Target-dependent sentiment analysis
Target-dependent sentiment analysis (Target) seeks to classify the sentiment of a text’s author towards an entity that occurs in the text as positive, negative, or neutral. We use the data from Dong et al. Dong et al. (2014). An example instance is the expression “how do you like settlers of catan for the wii?” which is labelled as neutral towards the target “wii’.’ The evaluation metric is macro-averaged ().
Aspect-based sentiment analysis
Aspect-based sentiment analysis is the task of identifying whether an aspect, i.e. a particular property of an item is associated with a positive, negative, or neutral sentiment Ruder et al. (2016). We use the data of SemEval-2016 Task 5 Subtask 1 Slot 3 Pontiki et al. (2016) for the laptops (ABSA-L) and restaurants (ABSA-R) domains. An example is the sentence “For the price, you cannot eat this well in Manhattan”, labelled as positive towards both the aspects “restaurant prices” and “food quality”. The evaluation metric for both domains is accuracy ().
Stance detection
Stance detection (Stance) requires a model, given a text and a target entity, which might not appear in the text, to predict whether the author of the text is in favour or against the target or whether neither inference is likely Augenstein et al. (2016). We use the data of SemEval-2016 Task 6 Subtask B Mohammad et al. (2016). An example from this dataset would be to predict the stance of the tweet “Be prepared - if we continue the policies of the liberal left, we will be #Greece” towards the topic “Donald Trump”, labelled as “favor”. The evaluation metric is the macro-averaged score of the “favour” and “against” classes ().
Fake news detection
The goal of fake news detection in the context of the Fake News Challengehttp://www.fakenewschallenge.org/ is to estimate whether the body of a news article agrees, disagrees, discusses, or is unrelated towards a headline. We use the data from the first stage of the Fake News Challenge (FNC-1). An example for this dataset is the document “Dino Ferrari hooked the whopper wels catfish, (…), which could be the biggest in the world.” with the headline “Fisherman lands 19 STONE catfish which could be the biggest in the world to be hooked” labelled as “agree”. The evaluation metric is accuracy ()We use the same metric as Riedel et al. (2017)..
Natural language inference
Natural language inference is the task of predicting whether one sentences entails, contradicts, or is neutral towards another one. We use the Multi-Genre NLI corpus (MultiNLI) from the RepEval 2017 shared task Nangia et al. (2017). An example for an instance would be the sentence pair “Fun for only children”, “Fun for adults and children”, which are in a “contradiction” relationship. The evaluation metric is accuracy ().
2 Base model
Our base model is the Bidirectional Encoding model Augenstein et al. (2016), a state-of-the-art model for stance detection that conditions a bidirectional LSTM (BiLSTM) encoding of a text on the BiLSTM encoding of the target. Unlike Augenstein et al. (2016), we do not pre-train word embeddings on a larger set of unlabelled in-domain text for each task as we are mainly interested in exploring the benefit of multi-task learning for generalisation.
3 Training settings
We use BiLSTMs with one hidden layer of dimensions, -dimensional randomly initialised word embeddings, a label embedding size of . We train our models with RMSProp, a learning rate of , a batch size of , and early stopping on the validation set of the main task with a patience of .
Results
Our main results are shown in Table 3, with a comparison against the state of the art. We present the results of our multi-task learning network with label embeddings (MTL + LEL), multi-task learning with label transfer (MTL + LEL + LTN), and the semi-supervised extension of this model. On 7/8 tasks, at least one of our architectures is better than single-task learning; and in 4/8, all our architectures are much better than single-task learning.
The state-of-the-art systems we compare against are often highly specialised, task-dependent architectures. Our architectures, in contrast, have not been optimised to compare favourably against the state of the art, as our main objective is to develop a novel approach to multi-task learning leveraging synergies between label sets and knowledge of marginal distributions from unlabeled data. For example, we do not use pre-trained word embeddings Augenstein et al. (2016); Palogiannidi et al. (2016); Vo and Zhang (2015), class weighting to deal with label imbalance Balikas and Amini (2016), or domain-specific sentiment lexicons Brun et al. (2016); Kumar et al. (2016). Nevertheless, our approach outperforms the state-of-the-art on two-way topic-based sentiment analysis (Topic-2).
The poor performance compared to the state-of-the-art on FNC and MultiNLI is expected; as we alternate among the tasks during training, our model only sees a comparatively small number of examples of both corpora, which are one and two orders of magnitude larger than the other datasets. For this reason, we do not achieve good performance on these tasks as main tasks, but they are still useful as auxiliary tasks as seen in Table 4.
Analysis
Our results above show that, indeed, modelling the similarity between tasks using label embeddings sometimes leads to much better performance. Figure 2 shows why. In Figure 2, we visualise the label embeddings of an MTL+LEL model trained on all tasks, using PCA. As we can see, similar labels are clustered together across tasks, e.g. there are two positive clusters (middle-right and top-right), two negative clusters (middle-left and bottom-left), and two neutral clusters (middle-top and middle-bottom).
Our visualisation also provides us with a picture of what auxilary tasks are beneficial, and to what extent we can expect synergies from multi-task learning. For instance, the notion of positive sentiment appears to be very similar across the topic-based and aspect-based tasks, while the conceptions of negative and neutral sentiment differ. In addition, we can see that the model has failed to learn a relationship between MultiNLI labels and those of other tasks, possibly accounting for its poor performance on the inference task. We did not evaluate the correlation between label embeddings and task performance, but Bjerva (2017) recently suggested that mutual information of target and auxiliary task label sets is a good predictor of gains from multi-task learning.
2 Auxilary Tasks
For each task, we show the auxiliary tasks that achieved the best performance on the development data in Table 4. In contrast to most existing work, we did not restrict ourselves to performing multi-task learning with only one auxiliary task Søgaard and Goldberg (2016); Bingel and Søgaard (2017). Indeed we find that most often a combination of auxiliary tasks achieves the best performance. In-domain tasks are less used than we assumed; only Target is consistently used by all Twitter main tasks. In addition, tasks with a higher number of labels, e.g. Topic-5 are used more often. Such tasks provide a more fine-grained reward signal, which may help in learning representations that generalise better. Finally, tasks with large amounts of training data such as FNC-1 and MultiNLI are also used more often. Even if not directly related, the larger amount of training data that can be indirectly leveraged via multi-task learning may help the model focus on relevant parts of the representation space Caruana (1993). These observations shed additional light on when multi-task learning may be useful that go beyond existing studies Bingel and Søgaard (2017).
3 Ablation analysis
We now perform a detailed ablation analysis of our model, the results of which are shown in Table 5. We ablate whether to use the LEL (+ LEL), whether to use the LTN (+ LTN), whether to use the LEL output or the main model output for prediction (main model output is indicated by , main model), and whether to use the LTN as a regulariser or for semi-supervised learning (semi-supervised learning is indicated by + semi). We further test whether to use diversity features (– diversity feats) and whether to use main model predictions for the LTN (+ main model feats).
Overall, the addition of the Label Embedding Layer improves the performance over regular MTL in almost all cases.
4 Label transfer network
To understand the performance of the LTN, we analyse learning curves of the relabelling function vs. the main model. Examples for all tasks without semi-supervised learning are shown in Figure 3. One can observe that the relabelling model does not take long to converge as it has fewer parameters than the main model. Once the relabelling model is learned alongside the main model, the main model performance first stagnates, then starts to increase again. For some of the tasks, the main model ends up with a higher task score than the relabelling model. We hypothesise that the softmax predictions of other, even highly related tasks are less helpful for predicting main labels than the output layer of the main task model. At best, learning the relabelling model alongside the main model might act as a regulariser to the main model and thus improve the main model’s performance over a baseline MTL model, as it is the case for TOPIC-5 (see Table 5).
To further analyse the performance of the LTN, we look into to what degree predictions of the main model and the relabelling model for individual instances are complementary to one another. Or, said differently, we measure the percentage of correct predictions made only by the relabelling model or made only by the main model, relative to the number of correct predictions overall. Results of this for each task are shown in Table 6 for the LTN with and without semi-supervised learning. One can observe that, even though the relabelling function overall contributes to the score to a lesser degree than the main model, a substantial number of correct predictions are made by the relabelling function that are missed by the main model. This is most prominently pronounced for ABSA-R, where the proportion is 14.6.
Conclusion
We have presented a multi-task learning architecture that (i) leverages potential synergies between classifier functions relating shared representations with disparate label spaces and (ii) enables learning from mixtures of labeled and unlabeled data. We have presented experiments with combinations of eight pairwise sequence classification tasks. Our results show that leveraging synergies between label spaces sometimes leads to big improvements, and we have presented a new state of the art for topic-based sentiment analysis. Our analysis further showed that (a) the learned label embeddings were indicative of gains from multi-task learning, (b) auxiliary tasks were often beneficial across domains, and (c) label embeddings almost always led to better performance. We also investigated the dynamics of the label transfer network we use for exploiting the synergies between disparate label spaces.
Acknowledgments
Sebastian Ruder is supported by the Irish Research Council Grant Number EBPPG/2014/30 and Science Foundation Ireland Grant Number SFI/12/RC/2289. Anders Søgaard is supported by the ERC Starting Grant Number 313695. Isabelle Augenstein is supported by Eurostars grant Number E10138. We further gratefully acknowledge the support of NVIDIA Corporation with the donation of the Titan Xp GPU used for this research.