Hierarchical Pre-training for Sequence Labelling in Spoken Dialog
Emile Chapuis, Pierre Colombo, Matteo Manica, Matthieu Labeau, Chloe Clavel
Introduction
The identification of both Dialog Acts (DA) and Emotion/Sentiment (E/S) in spoken language is an important step toward improving model performances on spontaneous dialogue task. Especially, it is essential to avoid the generic response problem, i.e., having an automatic dialog system generate an unspecific response — that can be an answer to a very large number of user utterances Yi et al. (2019); Colombo et al. (2019). DA and emotion identification Witon et al. (2018); Jalalzai et al. (2020) are done through sequence labelling systems that are usually trained on large corpora (with over labelled utterances) such as Switchboard Godfrey et al. (1992), MRDA Shriberg et al. (2004) or Daily Dialog Act Li et al. (2017). Even though large corpora enable learning complex models from scratch (e.g., seq2seq Colombo et al. (2020)), those models are very specific to the labelling scheme employed. Adapting them to different sets of emotions or dialog acts would require more annotated data. Generic representations Mikolov et al. (2013); Pennington et al. (2014); Peters et al. (2018); Devlin et al. (2018); Yang et al. (2019); Liu et al. (2019) have been shown to be an effective way to adapt models across different sets of labels. Those representations are usually trained on large written corpora such as OSCAR Suárez et al. (2019), Book Corpus Zhu et al. (2015) or Wikipedia Denoyer and Gallinari (2006). Although achieving state-of-the-art (SOTA) results on written benchmarks Wang et al. (2018), they are not tailored to spoken dialog (SD). Indeed, Tran et al. (2019) have suggested that training a parser on conversational speech data can improve results, due to the discrepancy between spoken and written language (e.g., disfluencies Stolcke and Shriberg (1996), fillers Shriberg (1999); Dinkar et al. (2020), different data distribution). Furthermore, capturing discourse-level features, which distinguish dialog from other types of text Thornbury and Slade (2006), e.g., capturing multi-utterance dependencies, is key to embed dialog that is not explicitly present in pre-training objectives Devlin et al. (2018); Yang et al. (2019); Liu et al. (2019), as they often treat sentences as a simple stream of tokens. The goal of this work is to train on SD data a generic dialog encoder capturing discourse-level features that produce representations adapted to spoken dialog. We evaluate these representations on both DA and E/S labelling through a new benchmark SILICONE (Sequence labellIng evaLuatIon benChmark fOr spoken laNguagE) composed of datasets of varying sizes using different sets of labels. We place ourselves in the general trend of using smaller models to obtain lightweight representations Jiao et al. (2019); Lan et al. (2019) that can be trained without a costly computation infrastructure while achieving good performance on several downstream tasks Henderson et al. (2020). Concretely, since hierarchy is an inherent characteristic of dialog Thornbury and Slade (2006), we propose the first hierarchical generic multi-utterance encoder based on a hierarchy of transformers. This allows us to factorise the model parameters, getting rid of long term dependencies and enabling training on a reduced number of GPUs. Based on this hierarchical structure, we generalise two existing pre-training objectives. As embeddings highly depend on data quality Le et al. (2019) and volume Liu et al. (2019), we preprocess OpenSubtitles Lison et al. (2019): a large corpus of spoken dialog from movies. This corpora is an order of magnitude bigger than corpora Budzianowski et al. (2018b); Lowe et al. (2015); Danescu-Niculescu-Mizil and Lee (2011) used in previous works Mehri et al. (2019); Hazarika et al. (2019). Lastly, we evaluate our encoder along with other baselines on SILICONE, which lets us draw finer conclusions of the generalisation capability of our modelsUpon publication, we will release the code, models and especially the preprocessing scripts to replicate our results..
Method
We start by formally defining the Sequence Labelling Problem. At the highest level, we have a set of conversations composed of utterances, i.e., with being the corresponding set of labels (e.g., DA, E/S). At a lower level each conversation is composed of utterances , i.e with being the corresponding sequence of labels: each is associated with a unique label . At the lowest level, each utterance can be seen as a sequence of words, i.e . Concrete examples with dialog act can be found in Table 1.
Our work builds upon existing objectives designed to pre-train encoders: the Masked Language Model (MLM) from Devlin et al. (2018); Liu et al. (2019); Lan et al. (2019); Zhang et al. (2019a) and the Generalized Autoregressive Pre-training (GAP) from Yang et al. (2019).
MLM Loss: The MLM loss corrupts sequences (or in our case, utterances) by masking a proportion of tokens. The model learns bidirectional representations by predicting the original identities of the masked-out tokens. Formally, for an utterance , a random set of indexed positions is selected and the associated tokens are replaced by a masked token [MASK] to obtain a corrupted utterance . The set of parameters is learnt by maximizing :
GAP Loss: the GAP loss consists in computing a classic language modelling loss across different factorisation orders of the tokens. In this way, the model will learn to gather information across all possible positions from both directions. The set of parameters is learnt by maximising:
2 Hierarchical Encoding
Capturing dependencies at different granularity levels is key for dialog embedding. Thus, we choose a hierarchical encoder Chen et al. (2018b); Li et al. (2018a). It is composed of two functions and , satisfying:
3 Hierarchical Pre-training
Current self-supervised pre-training objectives such as MLM and GAP are trained at the sequence level, which for us translates to only learning . In this section, we extend both the MLM and GAP losses at the dialog level in order to pre-train . Following previous work on both multi-task learning Argyriou et al. (2007); Ruder (2017) and hierarchical supervision Garcia et al. (2019); Sanh et al. (2019), we argue that optimising simultaneously at both levels rather than separately improves the quality of the resulting embeddings. Thus, we write our global hierarchical loss as:
where is either the MLM or GAP loss at the utterance level and is its generalisation at the dialog level.
3.2 MLM Loss
The MLM loss at the utterance level is defined in Equation 1. Our generalisation at the dialog level masks a proportion of utterances and generates the sequences of masked tokens (a concrete example can be found in Appendix B). Thus, at the dialog level the MLM loss is defined as:
3.3 GAP Loss
The GAP loss at the utterance level is defined in Equation 2. A possible generalisation of the GAP at the dialog level is to compute the loss of the generated utterance across all factorization orders of the context utterances. Formally, the GAP loss is defined at the dialog level as:
4 Architecture
Commonly, The functions and are either modelled with recurrent cells Serban et al. (2015) or Transformer blocks Vaswani et al. (2017). Transformer blocks are more parallelizable, offering shorter paths for the forward and backward signals and requiring significantly less time to train compared to recurrent layers. To the best of our knowledge this is the first attempt to pre-train a hierarchical encoder based only on transformersAlthough it is possible to relax the fixed size imposed by transformers Dai et al. (2019) in this paper we follow Colombo et al. (2020) and fix the context size to and the max utterance length to — these choices are made to work with OpenSubtitles, since the number of available dialogs drops when considering a number of utterances greater than .. The structure of the model can be found in Figure 1. In order to optimize dialog level losses as described in Equation 5, we generate (through ) the sequence with a Transformer Decoder (). For downstream tasks, the context embedding is fed to a simple MLP (simple classification), or to a CRF/GRU/LSTM (sequential prediction) — see Appendix B for more details. In the rest of the paper, we will name our hierarchical transformer-based encoder and the hierarchical RNN-based encoder . We use to refer to the set of model parameters learnt using the pre-training objective (either MLM or GAP) at the level if solely utterance level training is used, if solely dialog level is used and if multi level supervision is used ( according to the case.).
5 Pre-training Datasets
Datasets used to pre-train dialog encoders Hazarika et al. (2019); Mehri et al. (2019) are often medium-sized (e.g. Cornell Movie Corpus Danescu-Niculescu-Mizil and Lee (2011), Ubuntu Lowe et al. (2015), MultiWOz Budzianowski et al. (2018a)). In our work, we focus on OpenSubtitles Lison and Tiedemann (2016)http://opus.nlpl.eu/OpenSubtitles-alt-v2018.php because (1) it contains spoken language, contrarily to the Ubuntu corpus Lowe et al. (2015) based on logs; (2) as Wizard of Oz Budzianowski et al. (2018a) and Cornell Movie Dialog Corpus Danescu-Niculescu-Mizil and Lee (2011), it is a multi-party dataset; and (3) OpenSubtitles is an order of magnitude larger than any other spoken language dataset used in previous work. We segment OpenSubtitles by considering the duration of the silence between two consecutive utterances. Two consecutive utterances belong to the same conversation if the silence is shorter than We choose . Conversations shorter than the context size are droppedUsing pre-training method based on the next utterance proposed by Mehri et al. (2019) requires dropping conversation shorter than leading to a non-negligible loss in the preprocessing stage.. After preprocessing, Opensubtitles contains subtitles from movies or series which represent conversations and over billion of words.
6 Baseline Encoder
We compare the different methods we presented with two different types of baseline encoders: pre-trained encoders, and hierarchical encoders based on recurrent cells. The latter, achieve current SOTA performance in many sequence labelling tasks Li et al. (2018a); Colombo et al. (2020); Lin et al. (2017).
Pre-trained Encoder Models. We use BERT Devlin et al. (2018) through the pytorch implementation provided by the Hugging Face transformers library Wolf et al. (2019). The pre-trained model is fed with a concatenation of the utterances. Formally given an input context the concatenation is fed to BERT.
Hierarchical Recurrent Encoders. In this work we rely on our own implementation of the model based on . Hyperparameters are described in Appendix B.
Evaluation of Sequence Labelling
Sequence labelling tasks for spoken dialog mainly involve two different types of labels: DA and E/S. Early work has tackled the sequence labelling problem as an independent classification of each utterance. Deep neural network models that currently achieve the best results Keizer et al. (2002); Surendran and Levow (2006); Stolcke et al. (2000) model both contextual dependencies between utterances Colombo et al. (2020); Li et al. (2018b) and labels Chen et al. (2018b); Kumar et al. (2018); Li et al. (2018c).
The aforementioned methods require large corpora to train models from scratch, such as: Switchboard Dialog Act (SwDA) Godfrey et al. (1992), Meeting Recorder Dialog Act (MRDA) Shriberg et al. (2004), Daily Dialog Act Li et al. (2017), HCRC Map Task Corpus (MT) Thompson et al. (1993). This makes harder their adoption to smaller datasets, such as: Loqui human-human dialogue corpus (Loqui) Passonneau and Sachar. (2014), BT Oasis Corpus (Oasis) Leech and Weisser (2003), Multimodal Multi-Party Dataset (MELD) Poria et al. (2018a), Interactive emotional dyadic motion capture database (IEMO), SEMAINE database (SEM) Mckeown et al. (2013).
2 Presentation of SILICONE
Despite the similarity between methods usually employed to tackle DA and E/S sequential classification, studies usually rely on a single type of label. Moreover, despite the variety of small or medium-sized labelled datasets, evaluation is usually done on the largest available corpora (e.g., SwDA, MRDA). We introduce SILICONE, a collection of sequence labelling tasks, gathering both DA and E/S annotated datasets. SILICONE is built upon preexisting datasets which have been considered by the community as challenging and interesting. Any model that is able to process multiple sequences as inputs and predict the corresponding labels can be evaluated on SILICONE. We especially include small-sized datasets, as we believe it will ensure that well-performing models are able to both distil substantial knowledge and adapt to different sets of labels without relying on a large number of examples. The description of the datasets composing the benchmark can be found in the following sections, while corpora statistics are gathered in Table 2.
Switchboard Dialog Act Corpus (SwDA) is a telephone speech corpus consisting of two-sided telephone conversations with provided topics. This dataset includes additional features such as speaker id and topic information. The SOTA model, based on a seq2seq architecture with guided attention, reports an accuracy of Colombo et al. (2020) on the official split.
ICSI MRDA Corpus (MRDA) has been introduced by Shriberg et al. (2004). It contains transcripts of multi-party meetings hand-annotated with DA. It is the second biggest dataset with around utterances. The SOTA model reaches an accuracy of Li et al. (2018a) and uses Bi-LSTMs with attention as encoder as well as additional features, such as the topic of the transcript.
DailyDialog Act Corpus () has been produced by Li et al. (2017). It contains multi-turn dialogues, supposed to reflect daily communication by covering topics about daily life. The dataset is manually labelled with dialog act and emotions. It is the third biggest corpus of SILICONE with utterances. The SOTA model reports an accuracy of Li et al. (2018a), using Bi-LSTMs with attention as well as additional features. We follow the official split introduced by the authors.
HCRC MapTask Corpus (MT) has been introduced by Thompson et al. (1993). To build this corpus, participants were asked to collaborate verbally by describing a route from a first participant’s map by using the map of another participant. This corpus is small ( utterances). As there is no standard train/dev/test splitWe split according to the code in https://github.com/NathanDuran/Maptask-Corpus. performances depends on the split. Tran et al. (2017) make use of a Hierarchical LSTM encoder with a GRU decoder layer and achieves an accuracy of .
Bt Oasis Corpus (Oasis) contains the transcripts of live calls made to the BT and operator services. This corpus has been introduced by Leech and Weisser (2003) and is rather small ( utterances). There is no standard train/dev/test split We use a random split from https://github.com/NathanDuran/BT-Oasis-Corpus. and few studies use this dataset.
2.2 S/E Datasets
In S/E recognition for spoken language, there is no consensus on the choice the evaluation metric (e.g., Ghosal et al. (2019); Poria et al. (2018b) use a weighted F-score while Zhang et al. (2019b) report accuracy). For SILICONE, we choose to stay consistent with the DA research and thus follow Zhang et al. (2019b) by reporting the accuracy. Additionally, emotion/sentiment labels are neither merged nor prepossessedComparison with concurrent work is more difficult as system performance heavily depends on the number of classes and label processing varies across studies Clavel and Callejas (2015)..
DailyDialog Emotion Corpus () has been previously introduced and contains eleven emotional labels. The SOTA model De Bruyne et al. (2019) is based on BERT with additional Valence Arousal and Dominance features and reaches an accuracy of 85% on the official split.
Multimodal EmotionLines Dataset (MELD) has been created by enhancing and extending EmotionLines dataset Chen et al. (2018a) where multiple speakers participated in the dialogues. There are two types of annotations and : three sentiments (positive, negative and neutral) and seven emotions (anger, disgust, fear, joy,neutral, sadness and surprise). The SOTA model with text only is proposed by Zhang et al. (2019b) and is inspired by quantum physics. On the official split, it is compared with a hierarchical bi-LSTM, which it beats with an accuracy of % () and % () against and .
IEMOCAP database (IEMO) is a multimodal database of ten speakers. It consists of dyadic sessions where actors perform improvisations or scripted scenarios. Emotion categories are: anger, happiness, sadness, neutral, excitement, frustration, fear, surprise, and other. There is no official split on this dataset. One proposed model is built with bi-LSTMs and achieves , with text only Zhang et al. (2019b).
SEMAINE database (SEM) comes from the Sustained Emotionally coloured Machine human Interaction using Nonverbal Expression project Mckeown et al. (2013). This dataset has been annotated on three sentiments labels: positive, negative and neutral by Barriere et al. (2018). It is built on Multimodal Wizard of Oz experiment where participants held conversations with an operator who adopted various roles designed to evoke emotional reactions. There is no official split on this dataset.
Results on SILICONE
This section gathers experiments performed on the SILICONE benchmark. We first analyse an appropriate choice for the decoder, which is selected over a set of experiments on our baseline encoders: a pre-trained BERT model and a hierarchical RNN-based encoder (). Since we focus on small-sized pre-trained representations, we limit the sizes of our pre-trained models to TINY and SMALL (see Table 7). We then study the results of the baselines and our hierarchical transformer encoders () on SILICONE along three axes: the accuracy of the models, the difference in performance between the E/S and the DA corpora, and the importance of pre-training. As we aim to obtain robust representations, we do not perform an exhaustive grid search on the downstream tasks.
Current research efforts focus on single label prediction, as it seems to be a natural choice for sequence labelling problems (subsection 2.1). Sequence labelling is usually performed with CRFs Chen et al. (2018b); Kumar et al. (2018) and GRU decoding Colombo et al. (2020), however, it is not clear to what extent inter-label dependencies are already captured by the contextualised encoders, and whether a plain MLP decoder could achieve competitive results. As can be seen in Table 3, we found that in the case of E/S prediction there is no clear difference between CRFs and MLPs, while GRU decoders exhibit poor performance, probably due to a lack of training data. It is also important to notice, that training a sequential decoder usually requires thorough hyper-parameter fine-tuning. As our goal is to learn and evaluate general representations that are decoder agnostic, in the following, we will use a plain MLP decoder for all the models compared.
2 General Performance Analysis
Table 4 provides an exhaustive comparison of the different encoders over the SILICONE benchmark. As previously discussed, we adopt a plain MLP as a decoder to compare the different encoders. We show that SILICONE covers a set of challenging tasks as the best performing model achieves an average accuracy of . Moreover, we observe that despite having half the parameters of a BERT model, our proposed model achieves an average result that is higher on the benchmark. SILICONE covers two different sequence labelling tasks: DA and E/S. In Table 4 and Table 3, we can see that all models exhibit a consistently higher average accuracy (up to ) on DA tagging compared to E/S prediction. This performance drop could be explained by the different sizes of the corpora (see Table 2). Despite having a larger number of utterances per label (), E/S tasks seem generally harder to tackle for the models. For example, on Oasis, where the is inferior than those of most E/S datasets (, , IEMO and SEM), models consistently achieve better results.
3 Importance of Pre-training for SILICONE
Results reported in Table 4 and Table 3 show that pre-trained transformer-based encoders achieve consistently higher accuracy on SILICONE, even when they are not explicitly considering the hierarchical structure. This difference can be observed both in small-sized datasets (e.g. MELD and SEM) and in medium/large size datasets (e.g SwDA and MRDA). To validate the importance of pre-training in a regime of low data, we train different (with random initialisation) on different portions of SEM and . Results shown in Figure 2 illustrate the importance of pre-trained representations.
Model Analysis
In this section, we dissect our hierarchical pre-trained models in order to better understand the relative importance of each component. We show how a hierarchical encoder allows us to obtain a light and efficient model. Additional experiments can be found in Appendix C.
First, we explore the differences in training representations on spoken and written corpora. Experimentally, we compare the predictions on SILICONE made by and the one made by (). The latter is a hierarchical encoder where utterance embeddings are obtained with the hidden vector representing the first token [CLS] (see Devlin et al. (2018)) of the second layer of BERT. In both cases, predictions are performed using an MLPWe consider the two first layer for a fair comparison based on the number of model parameters.. Results in Table 5 show higher accuracy when the pre-training is performed on spoken data. Since SILICONE is a spoken language benchmark, this result might be due to the specific features of colloquial speech (e.g. disfluencies, sentence length, vocabulary, word frequencies).
2 Hierarchy and Multi-Level Supervision
We study the relative importance of three aspects of our hierarchical pre-training with multi-level supervision. We first show that accounting for the hierarchy increases the performance of fine-tuned encoders, even without our specific pre-training procedure. We then compare our two proposed hierarchical pre-training procedures based on the GAP or MLM loss. Lastly, we look at the contribution of the possible levels of supervision on reduced training data from SEM.
We compare the performance of BERT-4layers with the () previously described. Results reported in Table 5 demonstrate that fine-tuning on downstream tasks with a hierarchical encoder yields to higher accuracy, with fewer parameters, even when using already pre-trained representations.
2.2 MLM vs GAP
In this experiment, we compare the different pre-training objectives at utterance and dialog level. As a reminder () and () are respectively trained using the standard MLM loss Devlin et al. (2018) and the standard GAP loss Yang et al. (2019). In Table 6 we report the different pre-training objective results. We observe that pre-training at the dialog level achieves comparable results to the utterance level pre-training for MLM and slightly worse for GAP. Interestingly, we observe that () compared to () achieves worse results, which is not consistent with the performance observed on other benchmarks, such as GLUE Wang et al. (2018). The lower accuracy of the models trained using a GAP-based loss could be due to several factors (e.g., model size, pre-training using the GAP loss could require a finer choice of hyper-parameters). Finally, we see that supervising at both dialog and utterance level helps for MLMWe investigate a similar setting for GAP which lead to poor results, the loss hit a plateau suggesting that objectives are competing against each other. More advanced optimisations techniques Sener and Koltun (2018) are left for future work..
2.3 Multi level Supervision for pre-training
In this section, we illustrate the advantages of learning using several levels of supervision on small datasets. We fine-tune different model on SEM using different size of the training set. Results are shown in Figure 2. Overall we see that introducing sequence level supervision induces a consistent improvement on SEM. Results on are provided in Appendix C.
3 Other advantages of hierarchy
Introducing a hierarchical design in the encoder allows to break dialog into utterances and to consider inputs of size instead of size . First, it allows parameters sharing, reducing the number of model parameters. The different model sizes are reported in Table 7. Our TINY model contains half the parameters of BERT (4-layers). Furthermore, modelling long-range dependencies hierarchically makes learning faster and allows to get rid of learning tricks (e.g., partial order prediction Yang et al. (2019), two-stage pre-training based on sequence length Devlin et al. (2018)) required for non-hierarchical encoders. Lastly, original BERT and XLNET are pre-trained using respectively 16 and 512 TPUs. Pre-training lasts several days with over iterations. Our TINY hierarchical models are pre-trained during iterations (1.5 days) on 4 NVIDIA V100.
Conclusions
In this paper, we propose a hierarchical transformer-based encoder tailored for spoken dialog. We extend two well-known pre-training objectives to adapt them to a hierarchical setting and use OpenSubtitles, the largest spoken language dataset available, for encoder pre-training. Additionally, we provide an evaluation benchmark dedicated to comparing sequence labelling systems for the NLP community, SILICONE, on which we compare our models and pre-training procedures with previous approaches. By conducting ablation studies, we demonstrate the importance of using a hierarchical structure for the encoder, both for pre-training and fine-tuning. Finally, we find that our approach is a powerful method to learn generic representations on spoken dialog, with less parameters than state-of-the-art transformer models.
These results open new future research directions: (1) to investigate new pre-training objectives leveraging the hierarchical framework in order to achieve better results on SILICONE while keeping light models (2) to provide multilingual models using the whole pre-training corpus (OpenSubtitles) available in 62 languages, (3) investigate robust methods Staerman et al. (2020a) and the application of our embedding to different anomaly detection settings Staerman et al. (2019, 2020b). We hope that the SILICONE benchmark, experimental results, and publicly available code encourage further research to build stronger sequence labelling systems for NLP.
Acknowledgement
This work was supported by a grant overseen from the French National Research Agency (ANR-17-MAOI).
References
Appendix A Additional Details on data composing SILICONE
In this section, we illustrate the diversity of the dataset composing SILICONE. In Figure 3, we plot two histograms representing the different utterance lengths for DA and E/S. As expected, for spoken dialog, lengths are shorter than for written benchmarks (e.g., GLUE).
Appendix B Additional Details for Models
In this section we report model hyper-parameters and as well as additional descriptions of our baselines. For all models we use a tokenizer based on WordPiece Wu et al. (2016). We also provide a concrete example of corrupted context for the MLM Loss.
We report in Table 8 the main hyper-parameters used fo our model pre-training. We used GELU Hendrycks and Gimpel (2016) activations and the dropout rate Srivastava et al. (2014) is set to .
B.2 MLM Loss example
In this section we propose a visual illustration of the corrupted context Figure 4 by the MLM Loss.
B.3 Experimental Hyper-parameters for SILICONE
For all models, we use a batch size of and automatically select the best model on the validation set according to its loss. We do not perform exhaustive grid search either on the learning rate (that is set to ), nor on other hyper-parameters to perform a fair comparison between all the models. We use ADAMW Kingma and Ba (2014); Loshchilov and Hutter (2017) with a linear scheduler on the learning rate and the number of warm-up steps is set to .
B.4 Additional Details on Baselines
A representation for all the baselines can be found in Figure 5. For all models, both hidden dimension and embedding dimension is set to to ensure fair comparison with the proposed model. The MLP used for decoding contains 3 layers of sizes . We use RELU Agarap (2018) to introduce non linearity inside our architecture.
Appendix C Additional Experimental Results
In this section we report the detailed results on SILICONE, including the ones presented in Table 4. We report results on two new experiments: importance of pre-training time for both a TINY and SMALL model, we report the convergence time of a TINY model and finally we extend subsubsection 5.2.3 by reporting results on IEMO.
We show in Table 9 the results on the SILICONE benchmark for all the models mentioned in the paper.
C.2 Improvement over pre-training
In this experiment we illustrate how pre-training improves performance on SEM (see Figure 6). As expected accuracy improves when pre-training.
C.3 Multi level Supervision for pre-training MELD
In this experiment we report results of the experiment mentioned in subsubsection 5.2.3. In this experiment we see that the training process seems to be noisier for fractions lower than 40%. For larger percentages, we observe that including higher supervision (at the dialog level) during pre-training leads to a consistent improvement.
Appendix D Negative Results on GAP
We briefly describe few ideas we tried to make GAP works at both the utterance and dialog level. We hypothesise that:
giving the same weight to the utterance level and the dialog level (see Equation 3) was responsible of the observed plateau. Different combinations lead to fairly poor improvements.
the limited model capacity was part of the issue. Larger models does not give the expected results.