Heavy-tailed Representations, Text Polarity Classification & Data Augmentation

Hamid Jalalzai, Pierre Colombo, Chloé Clavel, Eric Gaussier, Giovanna Varni, Emmanuel Vignon, Anne Sabourin

Introduction

EVT in a machine learning framework has received increasing attention in the past few years. Learning tasks considered so far include anomaly detection , anomaly clustering , unsupervised learning , online learning , dimension reduction and support identification . The present paper builds upon the methodological framework proposed by Jalalzai et al. for classification in extreme regions. The goal of Jalalzai et al. is to improve the performance of classifiers g^(x)\widehat{g}(x) issued from Empirical Risk Minimization (ERM) on the tail regions {∥x∥>t}\{\|x\|>t\} Indeed, they argue that for very large tt, there is no guarantee that g^\widehat{g} would perform well conditionally to {∥X∥>t}\{\|X\|>t\}, precisely because of the scarcity of such examples in the training set. They thus propose to train a specific classifier dedicated to extremes leveraging the probabilistic structure of the tails. Jalalzai et al. demonstrate the usefulness of their framework with simulated and some real world datasets. However, there is no reason to assume that the previously mentioned text embeddings satisfy the required regularity assumptions. The aim of the present work is to extend ’s methodology to datasets which do not satisfy their assumptions, in particular to text datasets embedded by state of the art techniques. This is achieved by the algorithm Learning a Heavy Tailed Representation (in short LHTR) which learns a transformation mapping the input data XX onto a random vector ZZ which does satisfy the aforementioned assumptions. The transformation is learnt by an adversarial strategy .

In Appendix C we propose an interpretation of the extreme nature of an input in both LHTR and BERT representations. In a word, these sequences are longer and are more difficult to handle (for next token prediction and classification tasks) than non extreme ones.

Our second contribution is a novel data augmentation mechanism GENELIEX which takes advantage of the scale invariance properties of ZZ to generate synthetic sequences that keep invariant the attribute of the original sequence. Label preserving data augmentation is an effective solution to the data scarcity problem and is an efficient pre-processing step for moderate dimensional datasets . Adapting these methods to NLP problems remains a challenging issue. The problem consists in constructing a transformation hh such that for any sample xx with label y(x)y(x), the generated sample h(x)h(x) would remain label consistent: ~{}y\big{(}h(x)\big{)}=y(x) . The dominant approaches for text data augmentation rely on word level transformations such as synonym replacement, slot filling, swap deletion using external resources such as wordnet . Linguistic based approaches can also be combined with vectorial representations provided by language models . However, to the best of our knowledge, building a vectorial transformation without using any external linguistic resources remains an open problem. In this work, as the label y\big{(}h(x)\big{)} is unknown as soon as h(x)h(x) does not belong to the training set, we address this issue by learning both an embedding φ\varphi and a classifier gg satisfying a relaxed version of the problem above mentioned, namely ∀λ≥1\forall\lambda\geq 1

For mathematical reasons which will appear clearly in Section 2.2, hλh_{\lambda} is chosen as the homothety with scale factor λ\lambda, hλ(x)=λxh_{\lambda}(x)=\lambda x. In this paper, we work with output vectors issued by BERT . BERT and its variants are currently the most widely used language model but we emphasize that the proposed methodology could equally be applied using any other representation as input. BERT embedding does not satisfy the regularity properties required by EVT (see the results from statistical tests performed in Appendix B.5) Besides, there is no reason why a classifier gg trained on such embedding would be scale invariant, i.e. would satisfy for a given sequence uu, embedded as xx, g(hλ(x))=g(x)g(h_{\lambda}(x))=g(x) ∀λ≥1\forall\lambda\geq 1. On the classification task, we demonstrate on two datasets of sentiment analysis that the embedding learnt by LHTR on top of BERT is indeed following a heavy-tailed distribution. Besides, a classifier trained on the embedding learnt by LHTR outperforms the same classifier trained on BERT. On the dataset augmentation task, quantitative and qualitative experiments demonstrate the ability of GENELIEX to generate new sequences while preserving labels.

The rest of this paper is organized as follows. Section 2 introduces the necessary background in multivariate extremes. The methodology we propose is detailed at length in Section 3. Illustrative numerical experiments on both synthetic and real data are gathered in sections 4 and 5. Further comments and experimental results are provided in the supplementary material.

Background

for any Borelian set A⊂[0,∞]dA\subset[0,\infty]^{d} which is bounded away from and such that the limit measure μ\mu of the boundary ∂A\partial A is zero. For a complete introduction to the theory of Regular Variation, the reader may refer to . The measure μ\mu may be understood as the limit distribution of tail events. In (2), μ\mu is homogeneous of order −1-1, that is μ(tA)=t−1μ(A)\mu(tA)=t^{-1}\mu(A), t>0,A⊂[0,∞]d∖{0}t>0,A\subset[0,\infty]^{d}\setminus\{0\}. This scale invariance is key for our purposes, as detailed in Section 2.2. The main idea behind extreme value analysis is to learn relevant features of μ\mu using the largest available data.

2 Classification in extreme regions

In the present work we do not assume that the baseline representation XX for text data satisfies the assumptions of Theorem 1. Instead, our goal is is to render the latter theoretical framework applicable by learning a representation which satisfies the regular variation condition given in (2), hereafter referred as Condition (2) which is the main assumption for Theorem 1 to hold. Our experiments demonstrate empirically that enforcing Condition (2) is enough for our purposes, namely improved classification and label preserving data augmentation, see Appendix B.3 for further discussion.

Heavy-tailed Text Embeddings

A major difference distinguishing LHTR from existing auto-encoding schemes is that the target distribution on the latent space is not chosen as a Gaussian distribution but as a heavy-tailed, regularly varying one. A workable example of such a target is provided in our experiments (Section 4). As the Bayes classifier (i.e. the optimal one among all possible classifiers) in the extreme region has a potentially different structure from the Bayes classifier on the bulk (recall from Section 2 that the optimal classifier at infinity depends on the angle Θ(x)\Theta(x) only), LHTR trains two different classifiers, gextg^{\text{ext}} on the extreme region of the latent space on the one hand, and gbulkg^{\text{bulk}} on its complementary set on the other hand. Given a high threshold tt, the extreme region of the latent space is defined as the set {z:∥z∥>t}\{z:\|z\|>t\}. In practice, the threshold tt is chosen as an empirical quantile of order (1−κ1-\kappa) (for some small, fixed κ\kappa) of the norm of encoded data ∥Zi∥=∥φ(Xi)∥\|Z_{i}\|=\|\varphi(X_{i})\|. The classifier trained by LHTR is thus of the kind g(z)=gext(z)\mathds1{∥z∥>t}+gbulk(z)\mathds1{∥z∥≤t}.g(z)=g^{\text{ext}}(z)\mathds{1}\{\|z\|>t\}+g^{\text{bulk}}(z)\mathds{1}\{\|z\|\leq t\}. If the downstream task is classification on the whole input space, in the end the bulk classifier gbulkg^{\text{bulk}} may be replaced with any other classifier g′g^{\prime} trained on the original input data XX restricted to the non-extreme samples (i.e. {Xi,∥φ(Xi)∥≤t}\{X_{i},\|\varphi(X_{i})\|\leq t\}). Indeed training gbulkg^{\text{bulk}} only serves as an intermediate step to learn an adequate representation φ\varphi.

Recall from Section 2.2 that the optimal classifier in the extreme region as t→∞t\to\infty depends on the angular component θ(x)\theta(x) only, or in other words, is scale invariant. One can thus reasonably expect the trained classifier gext(z)g^{\text{ext}}(z) to enjoy the same property. This scale invariance is indeed verified in our experiments (see Sections 4 and 5) and is the starting point for our data augmentation algorithm in Section 3.2. An alternative strategy would be to train an angular classifier, i.e. to impose scale invariance. However in preliminary experiments (not shown here), the resulting classifier was less efficient and we decided against this option in view of the scale invariance and better performance of the unconstrained classifier.

The goal of LHTR is to minimize the weighted risk

2 A heavy-tailed representation for dataset augmentation

We now introduce GENELIEX (Generating Label Invariant sequences from Extremes), a data augmentation algorithm, which relies on the label invariance property under rescaling of the classifier for the extremes learnt by LHTR. GENELIEX considers input sentences as sequences and follows the seq2seq approach . It trains a Transformer Decoder GextG^{\text{ext}} on the extreme regions.

Note that the proposed method only augments data on the extreme regions. A general data augmentation algorithm can be obtained by combining this approach with any other algorithm on the original input data XX whose latent code Z=φ(XU)Z=\varphi(X_{U}) does not lie in the extreme regions.

Experiments : Classification

We start with a simple bivariate illustration of the heavy tailed representation learnt by LHTR. Our goal is to provide insight on how the learnt mapping φ\varphi acts on the input space and how the transformation affects the definition of extremes (recall that extreme samples are defined as those samples which norm exceeds an empirical quantile).

Labeled samples are simulated from a Gaussian mixture distribution with two components of identical weight. The label indicates the component from which the point is generated. LHTR is trained on 22502250 examples and a testing set of size 750750 is shown in Figure 2. The testing samples in the input space (Figure 2(a)) are mapped onto the latent space via φ\varphi (Figure 2(c)) In Figure 2(b), the extreme raw observations are selected according to their norm after a component-wise standardisation of XiX_{i}, refer to Appendix B for details. The extreme threshold tt is chosen as the 75%75\% empirical quantile of the norm on the training set in the input space. Notice in the latter figure the class imbalance among extremes. In Figure 2(c), extremes are selected as the 25%25\% samples with the largest norm in the latent space. Figure 2(d) is similar to Figure 2(b) except for the selection of extremes which is performed in the latent space as in Figure 2(c). On this toy example, the adversarial strategy appears to succeed in learning a code which distribution is close to the logistic target, as illustrated by the similarity between Figure 2(c) and Figure 5(a) in the supplementary. In addition, the heavy tailed representation allows a more balanced selection of extremes than the input representation.

2 Application to positive vs. negative classification of sequences

In this section, we dissect LHTR to better understand the relative importance of: (i) working with a heavy-tailed representation, (ii) training two independent classifiers: one dedicated to the bulk and the second one dedicated to the extremes. In addition, we verify experimentally that the latter classifier is scale invariant, which is neither the case for the former, nor for a classifier trained on BERT input. Experimental settings. We compare the performance of three models. The baseline NN model is a MLP trained on BERT. The second model LHTR1 is a variant of LHTR where a single MLP (CC) is trained on the output of the encoder φ\varphi, using all the available data, both extreme and non extreme ones. The third model (LHTR) trains two separate MLP classifiers CextC^{\text{ext}} and CbulkC^{\text{bulk}} respectively dedicated to the extreme and bulk regions of the learnt representation φ\varphi. All models take the same training inputs, use BERT embedding and their classifiers have identical structure, see Appendix A.2 and B.6 for a summary of model workflows and additional details concerning the network architectures. Comparing LHTR1 with NN model assesses the relevance of working with heavy-tailed embeddings. Since LHTR1 is obtained by using LHTR with Cext=CbulkC^{\text{ext}}=C^{\text{bulk}}, comparing LHTR1 with LHTR validates the use of two separate classifiers so that extremes are handled in a specific manner. As we make no claim concerning the usefulness of LHTR in the bulk, at the prediction step we suggest working with a combination of two models: LHTR with CextC^{ext} for extreme samples and any other off-the-shelf ML tool for the remaining samples (e.g. NN model). Datasets. In our experiments we rely on two large datasets from Amazon (231k reviews) and from Yelp (1,450k reviews) . Reviews, (made of multiple sentences) with a rating greater than or equal to \nicefrac4  5\nicefrac{{4\ }}{{\ 5}} are labeled as +1+1, while those with a rating smaller or equal to \nicefrac2  5\nicefrac{{2\ }}{{\ 5}} are labeled as −1-1. The gap in reviews’ ratings is designed to avoid any overlap between labels of different contents. Results. Figure 3 gathers the results obtained by the three considered classifiers on the tail regions of the two datasets mentioned above. To illustrate the generalization ability of the proposed classifier in the extreme regions we consider nested subsets of the extreme test set Ttest\mathcal{T}_{\text{test}}, Tλ={z∈Ttest,∥z∥≥λt}\mathcal{T^{\lambda}}=\{z\in\mathcal{T}_{\text{test}},\|z\|\geq\lambda t\}, λ≥1\lambda\geq 1. For all factor λ≥1\lambda\geq 1, Tλ⊆Ttest\mathcal{T^{\lambda}}\subseteq\mathcal{T}_{\text{test}}. The greater λ\lambda, the fewer the samples retained for evaluation and the greater their norms. On both datasets, LHTR1 outperforms the baseline NN model. This shows the improvement offered by the heavy-tailed embedding on the extreme region. In addition, LHTR1 is in turn largely outperformed by the classifier LHTR, which proves the importance of working with two separate classifiers.

The performance of the proposed model respectively on the bulk region, tail region and overall, is reported in Table 1, which shows that using a specific classifier dedicated to extremes improves the overall performance.

Scale invariance. On all datasets, the extreme classifier gextg^{\text{ext}} verifies Equation (1) for each sample of the test set, gext(λZ)=gext(Z)g^{\text{ext}}(\lambda Z)=g^{\text{ext}}(Z) with λ\lambda ranging from 11 to 2020, demonstrating scale invariance of gextg^{\text{ext}} on the extreme region. The same experiments conducted both with NN model and a MLP classifier trained on BERT and LHTR1 show label changes for varying values of λ\lambda: none of them are scale invariant. Appendix B.5 gathers additional experimental details. The scale invariance property will be exploited in the next section to perform label invariant generation.

Experiments : Label Invariant Generation

Comparison with existing work. We compare GENELIEX with two state of the art methods for dataset augmentation, Wei and Zou and Kobayashi . Contrarily to these works which use heuristics and a synonym dictionary, GENELIEX does not require any linguistic resource. To ensure that the improvement brought by GENELIEX is not only due to BERT, we have updated the method in with a BERT language model (see Appendix B.7 for details and Table 7 for hyperparameters). Evaluation Metrics. Automatic evaluation of generative models for text is still an open research problem. We rely both on perceptive evaluation and automatic measures to evaluate our model through four criteria (C1, C2, C3,C4). C1 measures Cohesion (Are the generated sequences grammatically and semantically consistent?). C2 (named Sent. in Table 3) evaluates label conservation (Does the expressed sentiment in the generated sequence match the sentiment of the input sequence?). C3 measures the diversity (corresponding to dist1 or dist2 in Table 3dist nn is obtained by calculating the number of distinct nn-grams divided by the total number of generated tokens to avoid favoring long sequences.) of the sequences (Does the augmented dataset contain diverse sequences?). Augmenting the training set with very diverse sequences can lead to better classification performance. C4 measures the improvement in terms of F1 score when training a classifier (fastText ) on the augmented training set (Does the augmented dataset improve classification performance?). Datasets. GENELIEX is evaluated on two datasets, a medium and a large one (see ) which respectively contains 1k and 10k labeled samples. In both cases, we have access to Dgn\mathcal{D}_{g_{n}} a dataset of 80k unlabeled samples. Datasets are randomly sampled from Amazon and Yelp. Experiment description. We augment extreme regions of each dataset according to three algorithms: GENELIEX (with scaling factor λ\lambda ranging from 1 to 1.5), Kobayashi , and Wei and Zou . For each train set’s sequence considered as extreme, 1010 new sequences are generated using each algorithm. Appendix B.7 gathers further details. For experiment C4 the test set contains 10410^{4} sequences.

2 Results

Automatic measures. The results of C3 and C4 evaluation are reported in Table 2. Augmented data with GENELIEX are more diverse than the one augmented with Kobayashi and Wei and Zou . The F1-score with dataset augmentation performed by GENELIEX outperforms the aforementioned methods on Amazon in medium and large dataset and on Yelp for the medium dataset. It equals state of the art performances on Yelp for the large dataset. As expected, for all three algorithms, the benefits of data augmentation decrease as the original training dataset size increases. Interestingly, we observe a strong correlation between more diverse sequences in the extreme regions and higher F1 score: the more diverse the augmented dataset, the higher the F1 score. More diverse sequences are thus more likely to lead to better improvement on downstream tasks (e.g. classification).

Perceptive Measures. To evaluate C1, C2, three turkers were asked to annotate the cohesion and the sentiment of 100100 generated sequences for each algorithm and for the raw data. F1 scores of this evaluation are reported in Table 3. Grammar evaluation confirms the findings of showing that random swaps and deletions do not always maintain the cohesion of the sequence. In contrast, GENELIEX and Kobayashi , using vectorial representations, produce more coherent sequences. Concerning sentiment label preservation, on Yelp, GENELIEX achieves the highest score which confirms the observed improvement reported in Table 2. On Amazon, turker annotations with data from GENELIEX obtain a lower F1-score than from Kobayashi . This does not correlate with results in Table 2 and may be explained by a lower Krippendorff Alphameasure of inter-rater reliability in $:isperfectdisagreementand: is perfect disagreement and1isperfectagreement.onAmazon(is perfect agreement. on Amazon (\alpha=0.20)thanonYelp() than on Yelp (\alpha=0.57$).

Broader Impact

In this work, we propose a method resulting in heavy-tailed text embeddings. As we make no assumption on the nature of the input data, the suggested method is not limited to textual data and can be extended to any type of modality (e.g. audio, video, images). A classifier, trained on aforementioned embedding is dilation invariant (see Equation 1) on the extreme region. A dilation invariant classifier enables better generalization for new samples falling out of the training envelop. For critical application ranging from web content filtering (e.g. spam , hate speech detection , fake news or multi-modal classification ) to medical case reports to court decisions it is crucial to build classifiers with lower generalization error. The scale invariance property can also be exploited to automatically augment a small dataset on its extreme region. For application where data collection requires a huge effort both in time and cost (e.g. industrial factory design, classification for rare language ), beyond industrial aspect, active learning problems involving heavy-tailed data may highly benefit from our data augmentation approach.

Acknowledgement

Anne Sabourin was partly supported by the Chaire Stress testing from Ecole Polytechnique and BNP Paribas. Concerning Eric Gaussier, this project partly fits within the MIAI project (ANR-19-P3IA-0003).

References

Appendix A Models

Adversarial networks, introduced in , form a system where two neural networks are competing. A first model GG, called the generator, generates samples as close as possible to the input dataset. A second model DD, called the discriminator, aims at distinguishing samples produced by the generator from the input dataset. The goal of the generator is to maximize the probability of the discriminator making a mistake. Hence, if PinputP_{\text{input}} is the distribution of the input dataset then the adversarial network intends to minimize the distance (as measured by the Jensen-Shannon divergence) between the distribution of the generated data PGP_{G} and PinputP_{\text{input}}. In short, the problem is a minmax game with value function V(D,G)V(D,G)

Auto-encoders and derivatives form a subclass of neural networks whose purpose is to build a suitable representation by learning encoding and decoding functions which capture the core properties of the input data. An adversarial auto-encoder (see ) is a specific kind of auto-encoders where the encoder plays the role of the generator of an adversarial network. Thus the latent code is forced to follow a given distribution while containing information relevant to reconstructing the input. In the remaining of this paper, a similar adversarial encoder constrains the encoded representation to be heavy-tailed.

A.2 Models Overview

Figure 4 provides an overview of the different algorithms proposed in the paper. Figure 4(a) describes the pipeline for LHTR detailed in Algorithm 1. Figure 4(b) describes the pipeline for the comparative baseline LHTR1 where Cext=CbulkC^{\text{ext}}=C^{\text{bulk}}. Figure 4(c) illustrates the pipeline for the baseline classifier trained on BERT. Figure 4(d) describes GENELIEX described in Algorithm 2, note that the hatched components are inherited from LHTR and are not used in the workflow.

A.3 LHTR and GENELIEX algorithm

This subsection provides detailed algorithm for both models LHTR and GENELIEX.

Appendix B Extreme Value Analysis: additional material

To the best of our knowledge, selection of kk in extreme value analysis (in particular in Algorithm 1 and Algorithm 2) is still a vivid problem in EVT for which no absolute answer exists. As kk gets large the number of extreme points increases including samples which are not large enough and deviates from the asymptotic distribution of extremes. Smaller values of kk increase the variance of the classifier/generator. This bias-variance trade-off is beyond the scope of this paper.

B.2 Preliminary standardization for selecting extreme samples

In Figure 2(b) selecting the extreme samples on the input space is not a straightforward step as the two components of the vector are not on the same scale, componentwise standardisation is a natural and necessary preliminary step. Following common practice in multivariate extreme value analysis it was decided to standardise the input data (Xi)i∈{1,…,n}(X_{i})_{i\in\{1,\ldots,n\}} by applying the rank-transformation:

B.3 Enforcing regularity assumptions in Theorem 1

Uniform convergence (4) is not enforced in LHTR and the question of how to enforce it together with regular variation of each class separately remains open. However, our experiments in sections 4 and 5 demonstrate that enforcing Condition (2) is enough for our purposes, namely improved classification and label preserving data augmentation.

B.4 Logistic distribution

B.5 Scale invariance comparison of BERT and LHTR

In this section, we compare LHTR and BERT and show that the latter is not scale invariant. For this preliminary experiment we rely on labeled fractions of both Amazon and Yelp datasets respectively denoted as Amazon small dataset and Yelp small dataset detailed in , each of them containing 10001000 sequences from the large dataset. Both datasets are divided at random in a train set Ttrain\mathcal{T}_{\text{train}} and Ttest\mathcal{T}_{\text{test}}. The train set represents \nicefrac34\nicefrac{{3}}{{4}} of the whole dataset while the remaining samples represent the test set. We use the hyperparameters reported in Table 4.

BERT is not regularly varying. In order to show that XX is not regularly varying, independence between ∥X∥\|X\| and a margin of Θ(X)\Theta(X) can be tested , which is easily done via correlation tests. Pearson correlation tests were run on the extreme samples of BERT and LHTR embeddings of Amazon small dataset and Yelp small dataset. The statistical tests were performed between all margins of \big{(}\Theta(X_{i})\big{)}_{1\geq i\geq n} and \big{(}\|X_{i}\|\big{)}_{1\geq i\geq n}.

Each histogram in Figure 6 displays the distribution of the pp-values of the correlation tests between the margins XjX_{j} and the angle Θ(X)\Theta(X) for j∈{1,…d}j\in\{1,\ldots d\}, in a given representation (BERT or LHTR) for a given dataset. For both Amazon small dataset and Yelp small dataset the distribution of the pp-values is shifted towards larger values in the representation of LHTR than in BERT, which means that the correlations are weaker in the former representation than in the latter. This phenomenon is more pronounced with Yelp small dataset than with Amazon small dataset. Thus, in BERT representation, even the largest data points exhibit a non negligible correlation between the radius and the angle and the regular variation condition does not seem to be satisfied. As a consequence, in a classification setup such as binary sentiment analysis detailed in Section 4.2), classifiers trained on BERT embedding are not guaranteed to be scale invariant. In other words for a representation XX of a sequence UU with a given label YY, the predicted label g(λX)g(\lambda X) is not necessarily constant for varying values of λ≥1\lambda\geq 1. Figure 7 illustrates this fact on a particular example taken from Yelp small dataset. The color (white or black respectively) indicates the predicted class (respectively −1-1 and +1+1). For values of λ\lambda close to 11, the predicted class is −1-1 but the prediction shifts to class +1+1 for larger values of λ\lambda.

Scale invariance of LHTR. We provide here experimental evidence that LHTR’s classifier gextg^{\text{ext}} is scale invariant (as defined in Equation (1)). Figure 8 displays the predictions gext(λZi)g^{\text{ext}}(\lambda Z_{i}) for increasing values of the scale factor λ≥1\lambda\geq 1 and ZiZ_{i} belonging to Ttest\mathcal{T}_{\text{test}}, the set of samples considered as extreme in the learnt representation. For any such sample ZZ, the predicted label remains constant as λ\lambda varies, i.e. it is scale invariant, gext(λZ)=gext(Z)g^{\text{ext}}(\lambda Z)=g^{\text{ext}}(Z), for all λ≥1\lambda\geq 1.

B.6 Experimental settings (Classification): additional details

Toy example. For the toy example, we generate 30003000 points distributed as a mixture of two normal distributions in dimension two. For training LHTR, the number of epochs is set to 100100 with a dropout rate equal to 0.40.4, a batch size of 6464 and a learning rate of 5∗10−45*10^{-4}. The weight parameter ρ3\rho_{3} in the loss function (Jensen-Shannon divergence from the target) is set to 10−310^{-3}. Each component φ\varphi, CbulkC^{\text{bulk}} and CextC^{\text{ext}} is made of 33 fully connected layers, the sizes of which are reported in Table 5. Datasets. For Amazon, we work with the video games subdataset from http://jmcauley.ucsd.edu/data/amazon/. For Yelp , we work with 1,450,000 reviews after that can be found at https://www.yelp.com/dataset.

B.7 Experiments for data generation

As mentioned in Section 5.1, hyperparameters for dataset augmentation are detailed in Table 7.

B.7.2 Influence of the scaling factor on the linguistic content

Table 8 gathers some extreme sequences generated by GENELIEX for λ\lambda ranging from 11 to 1.51.5. No major linguistic change appears when λ\lambda varies. The generated sequences are grammatically correct and share the same polarity (positive or negative sentiment) as the input sequence. Note that for greater values of λ\lambda, a repetition phenomenon appears. The resulting sequences keep the label and polarity of the input sequence but repeat some words .

Appendix C Extremes in Text

The aim of this section is double: first, to provide some intuition on what characterizes sequences falling in the extreme region of LHTR. Second, to investigate the hypothesis that extremes from LHTR are input sequences which tend to be harder to model than non extreme ones

Regarding the first aim ( (i) Are there interpretable text features correlated with the extreme nature of a text sample?, since we characterize extremes by their norm in LHTR representation, in practice the question boils down to finding text features which are positively correlated with the norm of the text samples in LHTR, which we denote by ∥φ(X)∥\|\varphi(X)\| and referred to as the ‘LHTR norm’ in the sequel. Preliminary investigations did not reveal semantic features (related to the meaning or the sentiment expressed in the sequence ) displaying such correlation. However we have identified two features which are positively correlated both together and with the norm in LHTR, namely the sequence length ∣U∣|U| as measured by the number of tokens of the input (recall that in our case an input sequence UU is a review composed of multiple sequences ), and the norm of the input in BERT representation (‘BERT norm’, denoted by ∥X∥\|X\|).

As for the second question ( (ii) Are LHTR’s extremes harder to model? ) we consider the next token prediction loss (‘LM loss’ in the sequel) obtained by training a language model on top of BERT. The next token prediction loss can be seen as a measure of hardness to model the input sequence. The question is thus to determine whether this prediction loss is correlated with the norm in LHTR (or in BERT, or with the sequence length).

Figure 9 displays pairwise scatterplots for the four considered variables on Yelp dataset (left) and Amazon dataset (right). These scatterplot suggest strong dependence for all pairs of variables. For a more quantitative assessment, Figure 10 displays the correlation matrices between the four quantities ∥φ(X)∥\|\varphi(X)\|, ∥X∥\|X\|, ∣U∣|U| and ‘LM Loss’ described above on Amazon and Yelp datasets. Pearson and Spearman two-sided correlation tests are performed on all pairs of variables, both tests having as null hypothesis that the correlation between two variables is zero. For all tests, pp-values are smaller than 10−1610^{-16}, therefore null hypotheses are rejected for all pairs.

These results prove that the four considered variables are indeed significantly positively correlated, which answers questions (i)(i) and (ii)(ii) above.

Figure 11 provides additional insight about the magnitude of the shift in sequence length between extremes in the LHTR representation and non extreme samples. Even though the histograms overlap (so that two different sequences of same length may be regarded as extreme or not depending on other factors that are not understood yet), there is a visible shift in distribution for both Yelp and Amazon datasets, both for the positive and negative class in the classification framework for sentiment analysis. Kolmogorov-Smirnoff tests between the length distributions of the two considered classes for each label were performed, which allows us to reject the null hypothesis of equality between distributions, as the maximum pp-values is less than 0.050.05.

We summarize the empirical findings of this section:

An ‘extreme’ text sequence in LHTR representation is more likely to have a greater length (number of tokens) than a non extreme one.

Positive correlation between the BERT norm and the LHTR norm indicates that a large sample in the BERT representation is likely to have a large norm in the LHTR representation as well: the learnt representation LHTR taking BERT as input keeps invariant (in probability) the ordering implied by the norm.

A consequence of the two above points is that long sequences tend to have a large norm in BERT.

Extreme text samples (regarding the BERT norm or the LHTR norm) tend to be harder to model than non-extreme ones.

Since extreme texts are harder to model and also somewhat harder to classify in view of the BERT classification scores reported in Table 1, there is room for improvement in their analysis and it is no wonder that a method dedicated to extremes i.e. relying on EVT such as LHTR outperforms the baseline.