Supervised Adversarial Contrastive Learning for Emotion Recognition in Conversations
Dou Hu, Yinan Bao, Lingwei Wei, Wei Zhou, Songlin Hu
Introduction
Emotion recognition in conversations (ERC) aims to detect emotions expressed by speakers during a conversation. The task is a crucial topic for developing empathetic machines Ma et al. (2020). Existing works mainly focus on context modeling Majumder et al. (2019); Ghosal et al. (2019); Hu et al. (2021a) and emotion representation learning Zhu et al. (2021); Yang et al. (2022); Li et al. (2022a) to recognize emotions. However, these methods have limitations in discovering the intrinsic structure of data relevant to emotion labels, and struggle to extract generalized and robust representations, resulting in mediocre recognition performance.
In the field of representation learning, label-based contrastive learning Khosla et al. (2020); Martín et al. (2022) techniques are used to learn a generalized representation by capturing similarities between examples within a class and contrasting them with examples from other classes. Since similar emotions often have similar context and overlapping feature spaces, these techniques that directly compress the feature space of each class are likely to hurt the fine-grained features of each emotion, thus limiting the ability of generalization.
To address these, we propose a supervised adversarial contrastive learning (SACL) framework to learn class-spread structured representations in a supervised manner. SACL applies contrast-aware adversarial training to generate worst-case samples and uses a joint class-spread contrastive learning objective on both original and adversarial samples. It can effectively utilize label-level feature consistency and retain fine-grained intra-class features.
Specifically, we adopt softThe soft version means a cross-entropy term is added to alleviate the class collapse issue Graf et al. (2021), wherein each point in the same class has the same representation. SCL (Gunel et al., 2021) on original samples to obtain contrast-aware adversarial perturbations. Then, we put perturbations on the hidden layers to generate hard positive examples with a min-max training recipe. These generated samples can spread out the representation space for each class and confuse robust-less networks. After that, we utilize a new soft SCL on obtained adversarial samples to maximize the consistency of class-spread representations with the same label. Under the joint objective on both original and adversarial samples, the network can effectively learn label-consistent features and achieve better generalization.
In context-dependent dialogue scenarios, directly generating adversarial samples interferes with the correlation between utterances, which is detrimental to context understanding. To avoid this, we design a contextual adversarial training (CAT) strategy to adaptively generate context-level worst-case samples and extract more diverse features from context. This strategy applies adversarial perturbations to the context-aware network structure in a multi-channel way, instead of directly putting perturbations on context-free layers in a single-channel way Goodfellow et al. (2015); Miyato et al. (2017). After introducing CAT, SACL can further learn more diverse features and smooth representation spaces from context-dependent inputs, as well as enhance the model’s context robustness.
Under SACL framework, we design a sequence-based method SACL-LSTM to recognize emotion in the conversation. It consists of a dual long short-term memory (Dual-LSTM) module and an emotion classifier. Dual-LSTM is a modified version of the contextual perception module Hu et al. (2021a), which can effectively capture contextual features from a dialogue. With the guidance of SACL, the model can learn label-consistent and context-robust emotional features for the ERC task.
We conduct experiments on three public benchmark datasets. Results consistently demonstrate that our SACL-LSTM significantly outperforms other state-of-the-art methods on the ERC task, showing the effectiveness and superiority of our method. Moreover, extensive experiments prove that our SACL framework can capture better structured and robust representations for classification.
The main contributions are as follows: 1) We propose a supervised adversarial contrastive learning (SACL) framework to extract class-spread structured representations for classification. It can effectively utilize label-level feature consistency and retain fine-grained intra-class features. 2) We design a contextual adversarial training (CAT) strategy to learn more diverse features from context-dependent inputs and enhancing the model’s context robustness. 3) We develop a sequence-based method SACL-LSTM under the framework to learn label-consistent and context-robust emotional features for ERCTo the best of our knowledge, this is the first attempt to introduce the idea of adversarial training into the ERC task.. 4) Experiments on three benchmark datasets show that SACL-LSTM significantly outperforms other state-of-the-art methods, and prove the effectiveness of the SACL frameworkThe source code is available at https://github.com/zerohd4869/SACL.
Methodology
In this section, we first present the methodology of SACL framework. Besides, for better adaptation to context-independent scenarios, we introduce a CAT strategy to SACL framework. Finally, we apply the proposed SACL framework for emotion recognition in conversations and provide a sequence-based method SACL-LSTM.
In the field of representation learning, label-based contrastive learning Khosla et al. (2020); Martín et al. (2022) techniques are used to learn a generalized representation by capturing similarities between examples within a class and contrasting them with examples from other classes. However, directly compressing the feature space of each class is prone to harming fine-grained intra-class features, which limits the model’s ability to generalize.
To address this, we design a supervised adversarial contrastive learning (SACL) framework for learning class-spread structured representations. The framework applies contrast-aware adversarial training to generate worst-case samples and uses a joint class-spread contrastive learning objective on both original and adversarial samples. It can effectively utilize label-level feature consistency and retain fine-grained intra-class features. Figure 1 visualizes the difference between SACL and two representative optimization objectives (i.e., CE and soft SCL (Gunel et al., 2021)) on a toy example.
Formally, let us denote as the set of samples in a mini-batch. Define is the set of indices of all positives in the mini-batch distinct from , and is its cardinality. The loss function of soft SCL is a weighted average of CE loss and SCL loss with a trade-off scalar parameter , i.e.,
and denote the value of one-hot vector and probability vector at class index k, respectively. . refers to the hidden representation of the network’s output for the -th sample. is a pairwise similarity function, i.e., dot product. is a scalar temperature parameter that controls the separation of classes.
At each step of training, we apply an adversarial training strategy with the soft SCL objective on original samples to produce anti-contrast worst-case samples. The training strategy can be implemented using a context-free approach such as FGM (Miyato et al., 2017) or our context-aware CAT. These samples can be seen as hard positive examples, which spread out the representation space for each class and confuse the robust-less model. After that, we utilize a new soft SCL on obtained adversarial samples to maximize the consistency of class-spread representations with the same label. Following the above calculation process of on original samples, the optimization objective on corresponding adversarial samples can be easily obtained in a similar way, i.e., .
The overall loss of SACL is defined as a sum of two soft SCL losses on both original and adversarial samples, i.e.,
2 Contextual Adversarial Training
Adversarial training (AT) Goodfellow et al. (2015); Miyato et al. (2017) is a widely used regularization method for models to improve robustness to small, approximately worst-case perturbations. In context-dependent scenarios, directly generating adversarial samples interferes with the correlation between samples, which is detrimental to context understanding.
To avoid this, we design a contextual adversarial training (CAT) strategy for a context-aware network, to obtain diverse context features and a robust model. Different from the standard AT that put perturbations on context-free layers (e.g., word/sentence embeddings), we add adversarial perturbations to the context-aware network structure in a multi-channel way. Under a supervised training objective, it can obtain diverse features from context and enhance model robustness to contextual perturbations.
Here, we take the LSTM network (Hochreiter and Schmidhuber, 1997) with a sequence input as an example, and the corresponding representations of the output are . Adversarial perturbations are put on context-aware hidden layers of the LSTM in a multi-channel way, including three gated layers and a memory cell layer in the LSTM structure, as shown in Figure 2.
With contextual perturbations on the network, there is a reasonable interpretation of the formulation in Eq. (5). The inner maximization problem is finding the context-level worst-case samples for the network, and the outer minimization problem is to train a robust network to the worst-case samples. After introducing CAT, our SACL can further learn more diverse features and smooth representation spaces from context-dependent inputs, as well as enhance the model’s context robustness.
3 Application for Emotion Recognition in Conversations
In this subsection, we apply SACL framework to the task of emotion recognition in conversations (ERC), and present a sequence-based method SACL-LSTM. The overall architecture is illustrated in Figure 3. With the guidance of SACL with CAT, the method can learn label-consistent and context-robust emotional features for better emotion recognition.
The ERC task aims to recognize emotions expressed by speakers in a conversation. Formally, let be a conversation with utterances and speakers/parties. Each utterance is spoken by the party , where maps the utterance index into the corresponding speaker index. For each , represents the set of utterances spoken by the party , i.e., . The goal is to identify the emotion label for each utterance from the pre-defined emotions .
3.2 Textual Feature Extraction
3.3 Model Structure
The network structure of SACL-LSTM consists of a dual long short-term memory (Dual-LSTM) module and an emotion classifier.
After extracting textual features, we design a Dual-LSTM module to capture situation- and speaker-aware contextual features in a conversation. It is a modified version of the contextual perception module in Hu et al. (2021a).
Specifically, to alleviate the speaker cold-start issueIn multi-party interactions, some speakers have limited interaction with others, making it difficult to capture context-aware speaker characteristics directly with sequence-based networks, especially with the short speaker sequence., we modify the speaker perception module. If the number of utterances of the speaker is less than a predefined integer threshold , the common characteristics of these cold-start speakers are directly represented by a shared general speaker vector . The speaker-aware features are computed as:
where indicates a BiLSTM to obtain speaker embeddings. is the -th hidden state of the party with a dimension of . . refers to all utterances of in a conversation. The situation-aware features are defined as,
where is a BiLSTM to obtain situation-aware embeddings and is the hidden vector with a dimension of .
We concatenate the situation-aware and speaker-aware features to form the context representation of each utterance, i.e.,
Finally, according to the context representation, an emotion classifier is applied to predict the emotion label of each utterance.
3.4 Optimization Process
Under SACL framework, we apply contrast-aware CAT to generate worst-case samples and utilize a joint class-spread contrastive learning objective on both original and adversarial samples. At each step of training, we apply the CAT strategy with the soft SCL objective on original samples to produce context-level adversarial perturbations. The perturbations are put on context-aware hidden layers of Dual-LSTM in a multi-channel way, and then obtain adversarial samples. After that, we leverage a new soft SCL on these worst-case samples to maximize the consistency of emotion-spread representations with the same label. Under the joint objective on both original and adversarial samples, SACL-LSTM can learn label-consistent and context-robust emotional features for ERC.
Experimental Setups
We evaluate our model on three benchmark datasets. IEMOCAP Busso et al. (2008) contains dyadic conversation videos between pairs of ten unique speakers, where the first eight speakers belong to train sets and the last two belong to test sets. The utterances are annotated with one of six emotions, namely happy, sad, neutral, angry, excited, and frustrated. MELD Poria et al. (2019a) contains multi-party conversation videos collected from Friends TV series. Each utterance is annotated with one of seven emotions, i.e., joy, anger, fear, disgust, sadness, surprise, and neutral. EmoryNLP Zahiri and Choi (2018) is a textual corpus that comprises multi-party dialogue transcripts of the Friends TV show. Each utterance is annotated with one of seven emotions, i.e., sad, mad, scared, powerful, peaceful, joyful, and neutral.
The statistics are reported in Table 1. In this paper, we focus on ERC in a textual setting. Other multimodal knowledge (i.e., acoustic and visual modalities) is not used. We use the pre-defined train/val/test splits in MELD and EmoryNLP. Following previous studies Hazarika et al. (2018b); Ghosal et al. (2019), we randomly extract 10% of the training dialogues in IEMOCAP as validation sets since there is no predefined train/val split.
2 Comparison Methods
The fourteen baselines compared are as follows. 1) Sequence-based methods: bc-LSTM Poria et al. (2017) employs an utterance-level LSTM to capture contextual features. DialogueRNN Majumder et al. (2019) is a recurrent network to track speaker states and context. COSMIC Ghosal et al. (2020) uses GRUs to incorporate commonsense knowledge and capture complex interactions. DialogueCRN Hu et al. (2021a) is a cognitive-inspired network with multi-turn reasoning modules that captures implicit emotional clues in a dialogue. CauAIN Zhao et al. (2022) uses causal clues in commonsense knowledge to enrich the modeling of speaker dependencies.
2) Graph-based methods: DialogueGCN Ghosal et al. (2019) uses GRUs and GCNs with relational edges to capture context and speaker dependency. RGAT Ishiwatari et al. (2020) applies position encodings to RGAT to consider speaker and sequential dependency. DAG-ERC Shen et al. (2021b) adopts a directed GNN to model the conversation structure. SGED+DAG Bao et al. (2022) is a speaker-guided framework with a one-layer DAG that can explore complex speaker interactions.
3) Transformer-based methods: KET Zhong et al. (2019) incorporates commonsense knowledge and context into a Transformer. DialogXL Shen et al. (2021a) adopts a modified XLNet to deal with longer context and multi-party structures. TODKAT Zhu et al. (2021) enhances the ability of Transformer by incorporating commonsense knowledge and a topic detection task. CoG-BART Li et al. (2022a) uses a SupCon loss Khosla et al. (2020) and a response generation task to enhance BART’s ability. SPCL+CL Song et al. (2022) applies a prompt-based BERT with supervised prototypical contrastive learning Wang et al. (2021); Martín et al. (2022) and curriculum learning Bengio et al. (2009).
3 Evaluation Metrics
Following previous works Hu et al. (2021a); Li et al. (2022a), we report the accuracy and weighted-F1 score to measure the overall performance. Also, the F1 score per class and macro-F1 score are reported to evaluate the fine-grained performance. For the structured representation evaluation, we choose three supervised clustering metrics (i.e., ARI, NMI, and FMI) and three unsupervised clustering metrics (i.e., SC, CHI, and DBI) to measure the clustering performance of learned representations. For the empirical robust evaluation Carlini and Wagner (2017), we use the robust weighted-F1 score on adversarial samples generated from original test sets. Besides, the paired t-test Kim (2015) is used to verify the statistical significance of the differences between the two approaches.
4 Implementation Details
All experiments are conducted on a single NVIDIA Tesla V100 32GB card. The validation sets are used to tune hyperparameters and choose the optimal model. For each method, we run five random seeds and report the average result of the test sets. The network parameters of our model are optimized by using Adam optimizer (Kingma and Ba, 2015). More experimental details are listed in Appendix B.
Results and Analysis
The overall resultsWe noticed that DialogueRNN and CauAIN present a poor weighted-F1 but a fine accuracy score on EmoryNLP, which is most likely due to the highly class imbalance issue. are reported in Table 2. SACL-LSTM consistently obtains the best weighted-F1 score over comparison methods on three datasets. Specifically, SACL-LSTM obtains +1.1% absolute improvements over other state-of-the-art methods in terms of the average weighted-F1 score on three datasets. Besides, SACL-LSTM obtains +1.2% absolute improvements in terms of the average accuracy score. The results indicates the good generalization ability of our method to unseen test sets.
We also report fine-grained results on three datasets in Table 3. SACL-LSTM achieves better results for most emotion categories (17 out of 20 classes), except three classes (i.e., disgust and anger in MELD, and scared in EmoryNLP). It is worth noting that SACL-LSTM obtains +2.0%, +1.6% and +0.8% absolute improvements in terms of the macro-F1 (average score of F1 for all classes) on IEMOCAP, MELD and EmoryNLP, respectively.
2 Ablation Study
We conduct ablation studies to evaluate key components in SACL-LSTM. The results are shown in Table 4. When removing the proposed SACL framework (i.e., - w/o SACL) and replacing it with a simple cross-entropy objective, we obtain inferior performance in terms of all metrics. When further removing the context-aware Dual-LSTM module (i.e., - w/o SACL - w/o Dual-LSTM) and replacing it with a context-free MLP (i.e., a fully-connected neural network with a single hidden layer), the results decline significantly on three datasets. It shows the effectiveness of both components.
3 Comparison with Different Optimization Objectives
To demonstrate the superiority of SACL, we include control experiments that replace it with the following optimization objectives, i.e., CE+SCL (soft SCL) (Gunel et al., 2021), CE+SupConThe idea of SupCon is very similar to SCL. Their implementations are slightly different. Combined with CE, they achieved very close performance, as shown in Table 5. Khosla et al. (2020), and cross-entropy (CE).
Table 5 shows results against various optimization objectives. SACL significantly outperforms the comparison objectives on three datasets. CE+SCL and CE+SupCon objectives apply label-based contrastive learning to extract a generalized representation, leading to better performance than CE. However, they compress the feature space of each class and harm fine-grained intra-class features, yielding inferior results than our SACL. SACL uses a joint class-spread contrastive learning objective on both original and adversarial samples. It can effectively utilize label-level feature consistency and retain fine-grained intra-class features.
4 Comparison with Different Training Strategies
To evaluate the effectiveness of contextual adversarial training (CAT), we compare with different training strategies, i.e., adversarial training (AT) Miyato et al. (2017), contextual random training (CRT), and vanilla training (VT). CRT is the strategy in which we replace in CAT with random perturbations from a multivariate Gaussian with the scaled norm on context-aware hidden layers.
The results are reported in Table 6. Compared with other strategies, our CAT obtains better performance consistently on three datasets. It shows that CAT can enhance the diversity of emotional features by adding adversarial perturbations to the context-aware structure with a min-max training recipe. We notice that AT strategy achieves the worst performance on MELD and EmoryNLP with the extremely short length of conversations. It indicates that AT is difficult to improve the diversity of context-dependent features with a limited context.
5 Structured Representation Evaluation
To evaluate the quality of structured representations, we measure the clustering performance based on the representations learned with different optimization objectives on the test set of IEMOCAP and MELD. Table 7 reports the clustering results of the Dual-LSTM network under three optimization objectives, including CE, CE+SCL, and our SACL.
According to supervised clustering metrics, the proposed SACL outperforms other optimization objectives by +1.3% and +1.4% in ARI, +0.9% and +1.1% in NMI, +1.1% and +0.9% in FMI for IEMOCAP and MELD, respectively. The more accurate clustering results show that our SACL can distinguish different data categories and assign similar data points to the same categories. It indicates that SACL can discover the intrinsic structure of data relevant to labels and extract generalized representations for emotion recognition.
According to unsupervised clustering metrics, SACL achieves better results than other optimization objectives by +0.03 and +0.07 in SC, +464.86 and +586.55 in CHI, and +0.07 and +0.25 in DBI for IEMOCAP and MELD, respectively. Better performance on these metrics suggests that SACL can learn more clear, separated, and compact clusters. This indicates that SACL can better capture the underlying structure of the data, which can be beneficial for subsequent emotion recognition.
Overall, the results demonstrate the effectiveness of the SACL framework in learning structured representations for improving clustering performance and quality, as evidenced by the significant improvements in various clustering metrics.
6 Context Robustness Evaluation
We further validate context robustness against different optimization objectives. We adjust different attack strengths of CE-based contextual adversarial perturbations on the test set and report the robust weighted-F1 scores. The context robustness results of SACL, CE with AT, and CE objectives on IEMOCAP and MELD are shown in Figure 4. CE with AT means using a cross-entropy objective with traditional adversarial training, i.e., FGM.
Our SACL consistently gains better robust weighted-F1 scores over other optimization objectives on both datasets. Under different attack strengths (), SACL-LSTM achieves up to 2.2% (average 1.3%) and 17.2% (average 13.4%) absolute improvements on IEMOCAP and MELD, respectively. CE with AT obtains sub-optimal performance since generating context-free adversarial samples interferes with the correlation between utterances, which is detrimental to context understanding. Our SACL using CAT can generate context-level worst-case samples for better training and enhance the model’s context robustness.
Moreover, we observe that SACL achieves a significant improvement on MELD with limited context. The average number of dialogue turns in MELD is relatively small, making it more likely for any two utterances to be strongly correlated. By introducing CAT, SACL learns more diverse features from the limited context, obtaining better context robustness results on MELD than others.
7 Representation Visualization
We qualitatively visualize the learned representations on the test set of MELD with t-SNE Van der Maaten and Hinton (2008). Figure 5 shows the visualization of the three speakers. Compared with using CE objective, the distribution of each emotion class learned by our SACL is more tight and united. It indicates that SACL can learn cluster-level structured representations and have a better ability to generalization. Besides, under SACL, the representations of surprise are away from neutral, and close to both joy and anger, which is consistent with the nature of surpriseSurprise is a non-neutral complex emotion that can be expressed with positive or negative valence Poria et al. (2019a).. It reveals that SACL can partly learn inter-class intrinsic structure in addition to intra-class feature consistency.
8 Error Analysis
Figure 6 shows an error analysis of SACL-LSTM and its ablated variant on the test set of IEMOCAP and MELD. The normalized confusion matrices are used to evaluate the quality of each model’s predicted outputs. From the diagonal elements of the matrices, SACL-LSTM reports better true positives against others on most fine-grained emotion categories. It suggests that SACL-LSTM is unbiased towards the under-represented emotion labels and learns better fine-grained features. Compared with the ablated variant w/o SACL, SACL-LSTM obtains better performances at similar categories, e.g., excited to happy, angry to frustrated, and frustrated to angry on IEMOCAP. It indicates that the SACL framework can effectively mitigate the misclassification problem of similar emotions. The poor effect of happy to excited may be due to the small proportion of happy samples used for training. For MELD, some categories (i.e., fear, sadness, and disgust) that account for a small proportion are easily misclassified as neutral accounting for nearly half, which is caused by the class imbalance issue.
Conclusion
We propose a supervised adversarial contrastive learning framework to learn class-spread structured representations for classification. It applies a contrast-aware adversarial training strategy and a joint class-spread contrastive learning objective. Besides, we design a contextual adversarial training strategy to learn more diverse features from context-dependent inputs and enhance the model’s context robustness. Under the SACL framework with CAT, we develop a sequence-based method SACL-LSTM to learn label-consistent and context-robust features on context-dependent data for better emotion recognition. Experiments verified the effectiveness of SACL-LSTM for ERC and SACL for learning generalized and robust representations.
Limitations
In this paper, we present a supervised adversarial contrastive learning (SACL) framework with contextual adversarial training to learn class-spread structured representations for context-dependent emotion classification. However, the framework is somewhat limited by the class imbalance issue, as illustrated in Section 4. To more comprehensively evaluate the generalization of SACL, it is necessary to test its transferability in low-resource and out-of-distribution scenarios, and evaluate its performance across a wider range of tasks. Additionally, it would be beneficial to explore the theoretical underpinnings and potential applications of the framework in greater depth. The aforementioned limitations will be left for future research.
Acknowledgements
This work was supported by the National Key Research and Development Program of China (No. 2022YFC3302102) and the National Natural Science Foundation of China (No. 62102412). The authors thank the anonymous reviewers and the meta-reviewer for their helpful comments on the paper.
References
Appendix Overview
In this supplementary material, we provide: (i) the related work, (ii) a detailed description of experimental setups, and (iii) detailed results.
Appendix A Related Work
Unlike traditional sentiment analysis (Zhou et al., 2019; Wei et al., 2020; Hu et al., 2022c; Li et al., 2022b), context information plays a significant role in identifying the emotion in conversations Poria et al. (2019b). Existing works usually utilize deep learning techniques to identify the emotion by context modeling and emotion representation learning. These works can be roughly divided into sequence-, graph- and Transformer-based methods.
Sequence-based methods Poria et al. (2017); Hazarika et al. (2018b, a); Majumder et al. (2019); Ghosal et al. (2020); Jiao et al. (2020a, b); Hu et al. (2021a); Zhao et al. (2022) generally utilize sequential information in a dialogue to capture different levels of contextual features, i.e., situation, speakers and emotions. For example, Poria et al. (2017) employ an LSTM to capture context-level features from surrounding utterances. Hazarika et al. (2018b, a); Jiao et al. (2020b) use memory networks to capture contextual features. Majumder et al. (2019) use GRUs to capture speaker, context and emotion features. Jiao et al. (2020a) introduce a conversation completion task based on unsupervised data to benefit the ERC task. Ghosal et al. (2020); Zhao et al. (2022) utilize GRUs to fuse commonsense knowledge and capture complex interactions in the dialogue. Hu et al. (2021a) propose a cognitive-inspired network that uses multi-turn reasoning modules to capture implicit emotional clues in conversations. In this paper, we propose a supervised adversarial contrastive learning framework with contextual adversarial training to learn class-spread structured representations for better emotion recognition.
A.1.2 Graph-based Methods
Graph-based methods Ghosal et al. (2019); Zhang et al. (2019b); Ishiwatari et al. (2020); Shen et al. (2021b); Hu et al. (2021b, 2022b); Bao et al. (2022) usually design a specific graph structure to capture complex dependencies in the conversation. For example, Ghosal et al. (2019); Zhang et al. (2019b); Shen et al. (2021b) leverage GNNs to capture complex interactions in a conversation. In order to simultaneously consider speaker interactions and sequence information, Ishiwatari et al. (2020) introduce a positional encoding module into RGAT. Hu et al. (2021b, 2022b) respectively design a graph-based fusion method that can simultaneously fuse multimodal knowledge and contextual features.
A.1.3 Transformer-based Methods
Transformer-based methods (Zhong et al., 2019; Wang et al., 2020; Shen et al., 2021a; Zhu et al., 2021; Li et al., 2021a; Lee and Choi, 2021; Lee and Lee, 2022; Li et al., 2022a; Song et al., 2022) usually exploit general knowledge in pre-trained language models (Devlin et al., 2019; Liu et al., 2019; Hu et al., 2022a), and model the conversation by a Transformer-based architecture. For example, Zhong et al. (2019) design a Transformer with graph attention to incorporate commonsense knowledge and contextual features. Wang et al. (2020) use a Transformer with an LSTM-CRF module to learn emotion consistency. Shen et al. (2021a) adopt a modified XLNet to deal with longer context and multi-party structures. Lee and Choi (2021) leverage LSTM and GCN to enhance BERT’s ability of context modeling. Yang et al. (2022) apply curriculum learning to deal with the learning problem of difficult samples. Li et al. (2022a) utilize a supervised contrastive term and a response generation task to enhance BART’s ability for ERC.
A.2 Contrastive Learning and Adversarial Training
Contrastive learning is a representation learning technique to learn generalized embeddings such that similar data sample pairs are close while dissimilar sample pairs stay far apart Chopra et al. (2005). Sohn (2016); van den Oord et al. (2018); Bachman et al. (2019); Tian et al. (2020); Hénaff (2020); Chen et al. (2020) utilize self-supervised contrastive learning to learn powerful representations. But these self-supervised techniques are generally limited by the risk of sampling bias and non-trivial data augmentation. Li et al. (2021b) propose prototypical contrastive learning to encode the semantic structure of data into the embedding space. Kim et al. (2020); Jiang et al. (2020b); Fan et al. (2021) add instance-wise adversarial examples during self-supervised contrastive learning to improve model robustness. Recently, Khosla et al. (2020); Gunel et al. (2021) use supervised contrastive learning to avoid the above risks and boost performance on downstream tasks by introducing label-level supervised signals. Wang et al. (2021); Martín et al. (2022) use supervised contrastive learning over prototype-label embeddings to learn representations for classification. Lin et al. (2022) employ supervised contrastive learning and CE-based adversarial training to learn domain-adaptive features for low-resource rumor detection. In this paper, we propose a supervised adversarial contrastive learning framework with contextual adversarial training to learn class-spread structured representations for classification on context-dependent data.
A.2.2 Adversarial Training
Adversarial training is a widely used regularization method to improve model robustness by generating adversarial examples with a min-max training recipe Szegedy et al. (2014). For example, Szegedy et al. (2014) train neural networks on a mixture of adversarial examples and clean examples. Goodfellow et al. (2015) further propose a fast gradient sign method to produce adversarial examples during training. Miyato et al. (2017) extend adversarial and virtual adversarial training to the text domain by applying perturbations to the word embeddings. After that, there are many variants established for supervised/semi-supervised learning Shafahi et al. (2019); Zhang et al. (2019a); Qin et al. (2019); Jiang et al. (2020a); Zhu et al. (2020).
Appendix B Experimental Setups
We report the detailed hyperparameter settings of SACL-LSTM on three datasets in Table 9. The class weights in the CE loss are applied to alleviate the class imbalance issue and are set by their relative ratios in the train and validation sets, except for MELD, which presents a poor effect. For MELD and EmoryNLP, we use focal loss (Lin et al., 2017), a modified version of the CE loss, to balance the weights of easy and hard samples during training.
Appendix C Experimental Results
The detailed results of context robustness evaluation on IEMOCAP and MELD are listed in Table 8.
C.2 Parameter Analysis
Figure 7 illustrates the effect of the temperature parameter in SACL framework on the ERC task.