XFORMAL: A Benchmark for Multilingual Formality Style Transfer
Eleftheria Briakou, Di Lu, Ke Zhang, Joel Tetreault
Introduction
Style Transfer (st) is the task of automatically transforming text in one style into another (for example, making an impolite request more polite). Most work in this growing field has focused primarily on style transfer within English, while covering different languages has received disproportional interest. Concretely, out of st papers we reviewed, all of them report results for st within English text, while there is just a single work covering each of the following languages: Chinese, Russian, Latvian, Estonian, and French Shang et al. (2019); Tikhonov et al. (2019); Korotkova et al. (2019); Niu et al. (2018). Notably, even though some efforts have been made towards multilingual st, researchers are limited to providing system outputs as a means of evaluation, and progress is hampered by the scarcity of resources for most languages.
At the same time, st lies at the core of human communication: when humans produce language, they condition their choice of grammatical and lexical transformations to a target audience and a specific situation. Among the many possible stylistic variations, Heylighen et al. (1999) argue that “a dimension similar to formality appears as the most important and universal feature distinguishing styles, registers or genres in different languages”. Consider the informal excerpts and their formal reformulations in French (fr) and Brazilian Portuguese (br-pt) in Table 1. Both informal-formal pairs share the same content. However, the informal language conveys more information than is contained in the literal meaning of the words Hovy (1987). These examples relate to the notion of deep formality Heylighen et al. (1999), where the ultimate goal is that of adding the context needed to disambiguate an expression. On the other hand, variations in formality might just reflect different situational and personal factors, as shown in the Italian (it) example.
This work takes the first step towards a more language-inclusive direction for the field of st by building the first corpus of style transfer for non-English languages. In particular, we make the following contributions: 1. Building upon prior work on Formality Style Transfer (fost) Rao and Tetreault (2018), we contribute an evaluation dataset, xformal that consists of multiple formal rewrites of informal sentences in three Romance languages: Brazilian Portuguese (br-pt), French (fr), and Italian (it); 2. Without assuming access to any gold-standard training data for the languages at hand, we benchmark a myriad of leading st baselines through automatic and human evaluation methods. Our results show that fost in non-English languages is particularly challenging as complex neural models perform on par with a simple rule-based system consisting of hand-crafted transformations. We make xformal, our annotations protocols, and analysis code publicly available and hope that this study facilitates and encourages more research towards Multilingual st.
Related Work
Controlling style aspects in generation tasks is studied in monolingual settings with an English-centric focus (intra-language) and cross-lingual settings together with Machine Translation (mt) (inter-language). Our work rests in intra-language st with a multilingual focus, in contrast to prior work.
that consist of parallel pairs in different styles include: gyafc for formality Rao and Tetreault (2018), Yelp Shen et al. (2017) and Amazon Product Reviews for sentiment He and McAuley (2016), political slant and gender controlled datasets Prabhumoye et al. (2018), Expert Style Transfer Cao et al. (2020), pastel for imitating personal Kang et al. (2019), simile for simile generation Chakrabarty et al. (2020), and others.
Intra-language st
was first cast as generation task by Xu et al. (2012) and is addressed through methods that use either parallel data or unpaired corpora of different styles. Parallel corpora designed for the task at hand are used to train traditional encoder-decoder architectures Rao and Tetreault (2018), learn mappings between latent representation of different styles Shang et al. (2019), or fine-tune pre-trained models Wang et al. (2019). Other approaches use parallel data from similar tasks to facilitate transfer in the target style via domain adaptation Li et al. (2019), multi-task learning Niu et al. (2018); Niu and Carpuat (2020), and zero-shot transfer Korotkova et al. (2019) or create pseudo-parallel data via data augmentation techniques Zhang et al. (2020); Krishna et al. (2020). Approaches that rely on non-parallel data include disentanglement methods based on the idea of learning style-agnostic latent representations (e.g., Shen et al. (2017); Hu et al. (2017)). However, they are recently criticized for resulting in poor content preservation Xu et al. (2018); Jin et al. (2019); Luo et al. (2019); Subramanian et al. (2018) and alternatively, translation-based models are proposed that use reconstruction and back-translation losses (e.g., Logeswaran et al. (2018); Prabhumoye et al. (2018)). Another line of work, focuses on manipulation methods that remove the style-specific attribute of text (e.g., Li et al. (2018); Xu et al. (2018)), while recent approaches use reinforcement learning (e.g., Wu et al. (2019); Gong et al. (2019), probabilistic formulations He et al. (2020), and masked language models Malmi et al. (2020).
Inter-language st
is introduced by Mirkin and Meunier (2015) who proposed personalized mt for en-French and en-German. Subsequent mt works control for politeness Sennrich et al. (2016a), voice Yamagishi et al. (2016), personality traits Rabinovich et al. (2017), user-provided terminology Hasler et al. (2018), gender Vanmassenhove et al. (2018), formality Niu et al. (2017); Feely et al. (2019), morphological variations Moryossef et al. (2019), complexity Agrawal and Carpuat (2019) and reading level Marchisio et al. (2019).
xformal Collection
We describe the process of collecting formal rewrites using data statements protocols Bender and Friedman (2018); Gebru et al. (2018).
To collect xformal, we firstly curate informal excerpts in multiple languages. To this end, we follow the procedures described in Rao and Tetreault (2018) (henceforth rt) who create a corpus of informal-formal sentence-pairs in English (en) entitled Grammarly’s Yahoo Answers Formality Corpus (gyafc).
Concretely, we use the L Yahoo! Answers corpus that consists of questions and answers posted to the Yahoo! Answers platform.https://webscope.sandbox.yahoo.com/catalog.php?datatype=l&did=11 The corpus contains a large number of informal text and allows control for different languages and different domains. More details are included under A.F. Similar to the collection of gyafc, we extract all answers from the Family & Relationships (F&R) topic that correspond to the three languages of interest: Família e Relacionamentos (br-pt), Relazioni e famiglia (it ), and Amour et relations (fr) (Step 1). We follow the same pre-processing steps as described in rt for consistency (Step 2). We filter out answers that: a) consist of questions; b) include urls; c) have fewer than five or more than tokens; or d) constitute duplicates.We tokenize with nltk: https://www.nltk.org/api/nltk.tokenize.html We automatically extract informal candidate sentences, as described in §5.3 (Step 3). Finally, we randomly sample sentences from the pool of informal candidates for each language. Table 2 presents statistics of the curation steps.
Procedures
We use the Amazon Mechanical Turk (mturk) platform to collect formal rewrites for our informal sentences. For each language, we split the annotation into batches of Human Intelligence Tasks (hits). In each hit, Turkers are given an informal excerpt and asked to generate its formal rewrite in the same language without changing its meaning. We collect rewrites per excerpt and release detailed instructions under A.G.
Annotation Workflow & Quality Control
Our annotation protocol consists of multiple Quality Control (qc) steps to ensure the recruitment of high-quality annotators. As a first step, we use location restrictions (qc1) to limit the pool of workers to countries where native speakers are most likely to be found. Next, we run several small pilot studies (of hits) to recruit potential workers. To participate in the pilot study, Turkers have to pass a qualification test (qc2) consisting of multiple-choice questions that test workers’ understanding of formality (see A.L). The pilot study results are reviewed by a native speaker (qc3) of each language to exclude workers who performed consistently poorly. We find that the two main reasons for poor quality are: a) rewrites of minimum-level edits, or b) rewrites that change the input’s meaning. Table 3 presents the number of workers at each qc step. Only workers passing all quality control steps (last row of Table 3) contribute to the final task. Finally, we post-process the collected rewrites by a) removing instances consisting of normalization-based edits only and b) correcting minor spelling errors using an off-the-shelf tool.https://languagetool.org/
Turkers’ demographics
We recruit Turkers from Brazil, France/Canada, and Italy for br-pt, fr, and it, respectively. Beyond their country of residence, no further information is available.
Compensation
We compensate at a rate of $ per hit with additional one-time bonuses that bumps them up to a target rate of over $/hour.
After this entire process, we have constructed a high-quality corpus of formality rewrites of sentences for three languages. In the next section, we provide statistics and an analysis of xformal.
xformal Statistics & Analysis
Following Pavlick and Tetreault (2016), we analyze the most frequent edit operations Turkers perform when formalizing the informal sentences. We conduct both an automatic analysis (details in A.I) of the whole set of rewrites, and a human analysis (details in A.H) of a random sample of rewrites per language (we recruited a native speaker for each language). Table 4 presents both analyses’ results, where we also include the corresponding statistics for the English language (gyafc). In general, we observe similar trends across languages: humans make edits covering both the "noisy-text" sense of formality (e.g., fixing punctuation, spelling errors, capitalization) and the more situational sense (paraphrase-based edits). Although cross-language trends are similar, we also observe differences: deleting fillers and word completion seems to be more prominent in the English rewrites than in other languages; normalizing abbreviations is a considerably frequent edit type for Brazilian Portuguese; paraphrasing is more frequent in the three non-English languages.
Surface differences of informal-formal pairs
We quantify surface-level differences between the informal sentences and formal rewrites via computing their character-level Levenshtein distance (Figure 1) and their pairwise Lexical Difference (LeD) based on the percentages of tokens that are not found in both sentences (Table 5). Both analyses show that Italian rewrites have the most edits compared to their corresponding informal sentences. French and Brazilian Portuguese follow, with English rewrites being closer to the informal inputs.
Diversity of formal rewrites
Are Turkers making similar choices when formalizing text? Since a large number of reformulations consist of paraphrase-based edits (more than %), we want to quantify the extent to which the formal rewrites of each sentence are diverse, in terms of their lexical choices. To that end, we quantify diversity via measuring self-bleu Zhu et al. (2018): considering one set of formal sentences as the hypothesis set and the others as references, we compute bleu for each formal set and define the average bleu score as a measure of the dataset’s diversity. Higher scores imply less diversity of the set. Results (last row of Table 4) show that xformal consists of more diverse rewrites compared to gyafc.
Formality shift of rewrites
We analyze the formality distribution of the original informal sentences with their formal rewrites in gyafc and xformal, as predicted by formality mbert models (§5.3). The distributions of formal rewrites are skewed towards positive values (Figure 2).
Multilingual fost Experiments
We benchmark eight st models on xformal to serve as baseline scores for future research. We describe the models (§5.1), the experimental setting (§5.2), the human and automatic evaluation methods (§5.3 and §5.4), and results (§5.5).
We define three baselines: 1. copyMotivated by Pang and Gimpel (2019) who notice that untransferred sentences with no alterations have the highest bleu score by a large margin for st tasks, we use this simple baseline as a lower bound; 2. rule-basedBased on the quantitative analysis of §4 and similarly to rt, we develop a rule-based approach that performs a set of predefined edits-based operations defined by hand-crafted rules. Example transformations include fix casing, remove repeated punctuation, handcraft a list of contraction expansions—a detailed description is found at A.C; 3. round-trip mtInspired by Zhang et al. (2020) who identify useful training pairs from the paraphrases generated by round-trip translations of millions of sentences, we devise a simpler baseline that starts from a text in language , pivots to en and then backtranslates to , using the aws translation service.https://aws.amazon.com/translate/
NMT-based models with synthetic parallel data
We follow the translate train Conneau et al. (2018); Artetxe et al. (2020) approach to collect data in multilingual settings: we obtain pseudo-parallel corpora in each language via machine translating an en resource of informal-formal pairs (§5.2).Details on the aws performance are found in A.B. Then, starting with translate train we benchmark the following nmt-based models: 1. translate train tagextends a leading en fost approach Niu et al. (2018) and trains a unified model that handles either formality direction via attaching a source tag that denotes the desired target formality; 2. multi-task tag-styleNiu et al. (2018)augments the previous approach with bilingual data that is automatically identified as formal (§5.3). The models are then trained in a multi-task fashion; 3. backtranslateaugments the translate train data with back-translated sentences of automatically detected informal text Sennrich et al. (2016b), using as the base model. We exclude backtranslated pairs consisting of copies. The output of the rule-based system is given as input to each model at inference time. For all three models, we run each system with random seeds, and combine them in a linear ensemble for decoding.
Unsupervised approaches
We benchmark two unsupervised methods that are used for en st: 1. unpsupervised neural machine translation (unmt)Subramanian et al. (2018)defines a pseudo-supervised setting and combines denoising auto-encoding and back-translation losses; 2. deep latent sequence model (dlsm)He et al. (2020)defines a probabilistic generative story that treats two unpaired corpora of separate styles as a partially observed parallel corpus and learns a mapping between them, using variational inference.
2 Experimental setting
For translate train tag we use gyafc, a large set of K en informal-formal parallel sentence-pairs obtained through crowdsourcing. Additionally, we augment the translated resource with OpenSubtitles Lison and Tiedemann (2016) bilingual data used for training mt models.Data are available at: http://opus.nlpl.eu/. Given that bilingual sentence-pairs can be noisy, we perform a filtering step to extract noisy bitexts using the Bicleaner toolkit Sánchez-Cartagena et al. .We use the publicly available pretrained Bicleaner models: https://github.com/bitextor/bicleaner, and discard sentences with a score lower than . Furthermore, we apply the same filtering steps as in §3 (Curation rational). Finally, each of the remaining sentences is assigned a formality score (§5.3), resulting in two pools of informal and formal text. Training instances are then randomly sampled from those pools: formal parallel pairs are used for multi-task tag-style; informal target side sentences are backtranslated for backtranslate; both informal and formal target-side texts are independently sampled from the two pools for training unsupervised models. Finally, for unsupervised fost in fr, we additionally experiment with in-domain data from the L French Yahoo! Answer Corpus that consists of M fr questions. https://webscope.sandbox.yahoo.com/catalog.php?datatype=l&did=74, Split into M/M formal/informal sentences. Table 6 includes statistics on training sizes.Bilingual data statistics are in A.K.
Preprocessing
We preprocess data consistently across languages using moses Koehn et al. (2007). Our pipeline consists of three steps: a) normalization; b) tokenization; c) true-casing. For nmt-based approaches, we also learn joint source-target bpe with K operations Sennrich et al. (2016b).
Model Implementations
For NMT-based and unsupervised models we use the open-sourced impementations of Niu et al. (2018) and He et al. (2020), respectively.https://github.com/xingniu/multitask-ft-fsmt, https://github.com/cindyxinyiwang/deep-latent-sequence-model We include more details on model architectures in A.D.
3 Automatic Evaluation
Recent work on st evaluation highlights the lack of standard evaluation practices Yamshchikov et al. (2020); Pang (2019); Pang and Gimpel (2019); Mir et al. (2019). We follow the most frequent evaluation metrics used in en tasks and measure the quality of the system’s outputs with respect to four dimensions, while we leave an extensive evaluation of automatic metrics for future work.
We compute self-bleu Papineni et al. (2002) which compares system outputs with the informal sentences.
Formality
We average the style transfer score of transferred sentences computed by a formality regression model. We fine-tune mbert Devlin et al. (2019) pre-trained language models on the machine-translated answers genre from Pavlick and Tetreault (2016) that consists of about K human-annotated sentences rated on a -point formality scale. To acquire an annotated corpus in the languages of interest, we follow the translate train transfer approach: we propagate the original en training data’s human ratings to their corresponding translations, assuming that translation preserves formality.See A.A for discussion on this assumption. To evaluate the multilingual formality regression models’ performance, we crowdsourced human judgments of Turkers for sentences per language. We report Spearman correlations of (br-pt), (it), (fr), and (en).
Fluency
We compute the logarithm of each sentence’s probability—computed by a -gram Kneser-Ney language model Kneser and Ney (1995)—and normalize it by the sequence length. We train each language model on M random sample of the non-English side of OpenSubtitles formal data.
Overall
We compute multi-bleu Post (2018) via comparing with multiple formal rewrites on xformal. Freitag et al. (2020) shows that correlation with human judgments improves when considering multiple references for mt evaluation.
4 Human evaluation
Given that automatic evaluation of st lacks standard evaluation practices—even in cases when en is considered—we turn to human evaluation to reliably assess our baselines following the protocols of rt. We sample a subset of sentences from xformal per language, evaluate outputs of systems, and collect judgments per instance.We open the task to all workers passing qc2 in Table 3. We include inter-annotator agreement results in A.E.
We collect formality ratings for the original informal reference, the formal human rewrite, and the formal system outputs on a -point discrete scale of to , following Lahiri (2015) (Very informal Informal Somewhat Informal Neutral Somewhat Formal Formal Very Formal).
Fluency
We collect fluency ratings for the original informal reference, the formal human rewrite, and the formal system outputs on a discrete scale of to , following Heilman et al. (2014) (Other Incomprehensible Somewhat Comprehensible Comprehensible Perfect).
Meaning Preservation
We adopt the annotation scheme of Semantic Textual Similarity Agirre et al. (2016): given the informal reference and formal human rewrite or the formal system outputs, Turkers rate the two sentences’ similarity on a to scale (Completely dissimilar Not equivalent but on same topic Not equivalent but share some details Roughly equivalent Mostly equivalent Completely equivalent).
Overall
We collect overall judgments of the system outputs using relative ranking: given the informal reference and a formal human rewrite, workers are asked to rank system outputs in the order of their overall formality, taking into account both fluency and meaning preservation. An overall score is then computed for each model via averaging results across annotating instances.
5 Results
Table 7 shows automatic results for all models across the four dimensions as well as human ratings for selected top models.
Concretely, the rule-based baselines are significantly () the best performing models in terms of meaning preservation across languages. This result is intuitive as the rule-based models act at the surface level and are unlikely to change the informal sentence’s meaning. The backtranslate ensemble systems are the second-best performing models in terms of meaning preservation, while the round-trip mt outputs diverge semantically from the informal sentences the most. Those results are consistent across languages and human/automatic evaluations. On the other hand, when we compare systems in terms of their formality, we observe the opposite pattern: the rule-based and backtranslate outputs are the most informal compared to the other ensemble nmt-based approaches across languages. Interestingly, the round-trip mt outputs exhibit the largest formality shift for br-pt and fr as measured by human evaluation. The trade-off between meaning preservation and formality among models was also observed in en (rt). Moreover, when we move to fluency, we notice similar results across systems. Specifically, human evaluation assigns almost all models an average score of , denoting that system outputs are comprehensible on average, with small differences between systems not being statistically significant. Notably, perplexity tells a different story: all system outputs are significantly better compared to the rule-based systems across configurations and languages. This result denotes that perplexity might not be a reliable metric to measure fluency in this setting, as noticed in Mir et al. (2019) and Krishna et al. (2020). When it comes to the overall ranking of systems, we observe that the nmt-based ensembles are better than the rule-based baselines for br-pt and fr, yet by a small margin as denoted by both multi-bleu and human evaluation. However, the corresponding results for it denote that there is no clear win, and the nmt-based ensembles still fail to surpass the naive rule-based models, yet by a small margin. Finally, all ensembles outperform the trivial copy baseline. Table 8 presents examples of system outputs. As a side note, we followed the recommendation of Tikhonov and Yamshchikov (2018) to show the performance of st models of individual runs and visualize trade-offs between metrics better. Unlike their work which found that reruns of the same model showed wide performance discrepancies, we found that most of our nmt-based models did not vary in performance on xformal. The results can be visualized in A.M.
Unsupervised model evaluation
We also benchmark the unsupervised models but focus solely on automatic metrics since they lag behind their supervised counterparts. As shown in Table 7, when using out-of-domain data (e.g., OpenSubtitles) for training, the models perform worse than their nmt counterparts across all three languages. The difference is most stark when considering self-bleu and multi-bleu scores. However, given access to large in-domain corpora (e.g., L Yahoo! French Answers) the gap between the two model classes closes with dlsm achieving a multi-bleu score of 42.1 compared to 48.3 for the best performing nmt model backtranslate. This shows the promise of unsupervised methods, assuming a large amount of in-domain data, on multilingual st tasks.
Lexical differences of system outputs
Finally, in Figure 3 we analyze the diversity of outputs by leveraging LeD scores resulting from pair-wise comparisons of different nmt systems. A larger LeD score denotes a larger difference between the lexical choices of the two systems under comparison. First, we observe that the round-trip mt outputs have the smallest lexical overlap with the informal input sentences. However, when this observation is examined together with human evaluation results, we conclude that the large number of lexical edits happens at the cost of diverging semantically from the input sentences. Moreover, we observe that the average lexical differences within nmt-based systems are small. This indicates that different systems perform similar edit operations that do not deviate a lot from the input sentence in terms of their lexical choices. This is unfortunate given that multilingual fost requires systems to perform more phrase-based operations, as shown in the analysis in §4.
Evaluation Metric
While evaluating evaluation metrics is not a goal of this work (though the data can be used for that purpose), we observe that the top models identified by the automatic metrics generally align with the top models identified by humans. While promising, further work is required to confirm if the automatic measures really do correlate with human judgments.
Conclusions & Future Directions
This work extends the task of formality style transfer to a multilingual setting. Specifically, we contribute xformal, an evaluation testbed consisting of informal sentences and multiple formal rewrites spanning three languages: br-pt, fr, and it. As in Rao and Tetreault (2018) Turkers can be effective in creating high quality st corpora. In contrast to the aforementioned en corpus, we find that the rewrites in xformal tend to be more diverse, making it a more challenging task.
Additionally, inspired by work on cross-lingual transfer and en fost, we benchmark several methods and perform automatic and human evaluations on their outputs. We found that nmt-based ensembles are the best performing models for fr and br-pt—a result consistent with en—however, they perform comparably to a naive rule-based baseline for it. To further facilitate reproducibility of our evaluations and corpus creation processes, as well as drive future work, we will release our scripts, rule-based baselines, source data, and annotation templates, on top of the release of xformal.
Our results open several avenues for future work in terms of benchmarking and evaluating fost in a more language inclusive direction. Notably, current supervised and unsupervised approaches for en fost rely on parallel in-domain data—with the latter treating the parallel set as two unpaired corpora—that are not available in most languages. We suggest that benchmarking fost models in multilingual settings will help understand their ability to generalize and lead to safer conclusions when comparing approaches. At the same time, multilingual fost calls for more language-inclusive consideration for automatic evaluation metrics. Model-based approaches have been recently proposed for evaluating different aspects of st. However, most of them rely heavily on English resources or pre-trained models. How those methods can be extended to multilingual settings and how we evaluate their performance remain open questions.
Acknowledgements
We thank Sudha Rao for providing references and materials of the gyafc dataset, Chris Callison-Burch and Courtney Napoles for discussions on MTurk annotations, Svebor Karaman for helping with data collection, our colleagues at Dataminr, and the naacl reviewers for their helpful and constructive comments.
Ethical Considerations
Finally, we address ethical considerations for dataset papers given that our work proposes a new corpus xformal. We reply to the relevant questions posed in the naacl Ethics faq.https://2021.naacl.org/ethics/faq/
The underlying data for our dataset as well as training our fost models and formality classifiers are from Yahoo! Answers L6 dataset. We were granted written permission by Yahoo (now Verizon) to make the resulting dataset public for academic use.
2 Dataset Collection Process
Turkers are paid over usd an hour. We targeted a rate higher than the us national minimum wage of usd given discussions with other researchers who use crowdsourcing. We include more information on collection procedures in §3.
3 IRB Approval
This question is not applicable for our work.
4 Dataset Characteristics
We follow Bender and Friedman (2018) and Gebru et al. (2018) and report characteristics in §3 and §4.
References
Appendix A Does translation preserve formality?
We examine the extend to which machine translation—through the aws service—affects the formality level of an input sentence: starting from a set of English sentences we have formality judgments for (ORIGINAL-EN), we perform a round-trip translation via pivoting through an auxiliary language (PIVOT-X). We then compare the formality prediction scores of the English Formality regression model for the two versions of the English input. In terms of Spearman correlation, the model’s performance drops by points on average when tested on round-trip translations. To better understand what causes this drop in performance, we present a per formality bin analysis in Table 9. On average we observe that translation preserves the formality level of formal sentences considerably well, while at the same time it tends to shift the formality level of informal sentences towards formal values—by a margin smaller than point—most of the times. To account for the formalization effect of translation, we draw the line between formal and informal sentences at the value of for scores predicted by multilingual regression models. This decision is based on the following intuition: if the formality shift of machine translated informal sentences is around value, the propagation of English formality labels imposes a negative shift of formal sentences in the model’s predictions.
Appendix B Amazon Web Service details
We compute the performance of the aws system on K randomly sentences from OpenSubtitles, as a sanity check of translation performance. We report bleu of (br-pt), (fr), and (it).
Appendix C Rule-based baselines
We develop a set of rules to automatically make an informal sentence more formal via performing surface-level edits similar to the en rule-based system of Rao and Tetreault (2018). The set of extracted rules are shared across languages with the only difference being the list of abbreviations:
We remove punctuation symbols that are repeated times.
Character repetition
We trailed characters repeated times (e.g., ciaooo ciao (it)).
Normalize casing
Several sentences might consist of words that are written in upper case. We lower case all characters apart from the first characters of the first word.
Normalize abbreviations
Informal text might contain slang words that are be abrreviated. We hand-craft a list of expansions for each language. For example, we replace kra cara (br-pt), ta ti amo (it), and bjr bonjour (fr), using publicly available resources. https://braziliangringo.com/brazilianinternetslang/, https://www.dummies.com/languages/italian/texting-and-chatting-in-italian/, https://frenchtogether.com/french-texting-slang/ The resulting lists sizes are: (br-pt), (it), and (fr).
Appendix D NMT architecture details
We used the nmt implementations of Niu et al. (2018) that are publicly available: https://github.com/xingniu/multitask-ft-fsmt. The nmt models are implemented as bi-directional lstms on Sockeye Hieber et al. (2018), using the same configurations across languages to establish fair comparisons. We use single lstms consisting of a single of size , multilayer perceptron attention with a layer size of , and word representations of size . We apply layer normalization and tie the source and target embeddings as well as the output layer’s weight matrix. We add dropout with probability (for the embeddings and lsmt cells in both the encoder and the decoder). For training, we use the Adam optimizer with a batch size of sentences and checkpoint the model every updates. Training stops after checkpoints without improvement of validation perplexity. For decoding, we use a beam size of .
Appendix E Inter-annotator agreement on human evaluation
To quantify inter-annotator agreement for the tasks of formality, meaning preservation, and fluency we measure the correlation of their ordinal ratings using inter-class correlation (icc) and their categorical agreement using a variations of Cohen’s coefficient. For the latter, given that we collect human evaluation judgments through crowd-sourcing, we follow the simulation framework of Pavlick and Tetreault (2016) to quantify agreement. Concretely, we simulate two annotators (Annotator , Annotator ) via randomly choosing one annotator’s judgment for a given instance as the rating of Annotator and taking the mean rating of the rest judgments as the rating of Annotator . We then compute Cohen’s for these two simulated annotators. We repeat this process times, and report the median and standard deviation of results. For measuring agreement of the overall ranking evaluation task we use the same simulated framework and report results of Kendall’s . Table 10 presents Inter-annotator agreement results on human evaluation across evaluation aspects and languages.
Appendix F Yahoo! L6 language statistics
Table 11 presents the total number of questions included in L Yahoo! Corpus, for each language. Although almost of questions are in English, the corpus contains a non-neglibible number of informal sentences in Spanish, French, Portuguese, Italian, and German, with a long tail of few questions for other languages.
Appendix G Instructions for xformal
Given an informal sentence in French (or Portuguese/Italian), generate its formal rewrite, without changing its meaning.
Detailed Instructions
Given an informal sentence, provide us with its informal rewrite. The informal rewrite should only change the formality attribute of the original sentence and preserve its meaning. Each sentence should be treated independently while rewrites should only rely on the information available in the sentences. There is no need to guess what additional information might be available in the documents the sentences come from.
Examples
Following we include examples of good and bad rewrites given to Turkers: informal Wow, I am very dumb in my observation skills…… good formal I do not have good observation skills. reasoning Formality is properly transferred and meaning is preserved. bad formal Wow, I am very foolish in my observation skills.. reasoning Formality is not properly transferred. bad formal I am very unintelligent and I don’t have good observation skills. reasoning Meaning has changed.
Appendix H Qualitative analysis of xformal
Following, we include the instructions given to native speakers for the qualitative analysis of xformal, as described in §4.
In this task you will be asked to judge the quality of formal rewrites of a set of informal sentences. To give more background context, your work would serve as a quality check over annotations obtained from the Amazon Mechanical Turk (AMT) platform. AMT workers were given an informal sentence, e.g., “I’d say it is punk though”, and then asked to provide us with its formal rewrite while maintaining the original content and being grammatical/fluent, e.g., “However, I do believe it to be punk”. In our work, workers were presented with sentences in French, Italian and Brazilian Portuguese. For our analysis we want to know a) whether the quality of the collected annotations is good (Task 1), b) what are the types of edits workers performed when formalizing the input sentence (Task 2). Both tasks consist of the same informal-formal sentences-pairs and could be performed either in parallel (e.g., judging a single informal-formal sentence-pair both in terms of quality and types of edit at the same time; Task 1 and 2 in Google sheet), or individually (Task 1, Task 2 in Google sheet). More information and examples for the two tasks are included below.
Task 1
Given an informal sentence you are asked to assess the quality of its formal rewrite. For each sentence pair type excellent, acceptable, or poor under the rate column. Read the instructions below before performing the task! What constitutes a good formal rewrite?
The style of the rewrite should be formal
The overall meaning of the informal sentence should be preserved
How should I interpret the provided options? Below we include detailed instructions on how to interpret the provided options.
Excellent the rewrite is formal, fluent and the original meaning is maintained. There is very little that could be done to make a better rewrite.
Acceptable the rewrite is generally an improvement upon the original but there are some minor issues (such as the rewrite contains a typo or missed transforming some informal parts of the sentence, for example).
Poor the rewrite offers a marginal or no improvement over the original. There are many aspects that could be improved.
Task 2
In this task the goal is to characterize the types of edits workers made while formalizing the informal input sentence. For each informal-formal sentence pair you should check each of the provided boxes. Note that: a) Multiple types of edits might hold at the same type (e.g., in the informal—formal “Lexus cars are awesome!”—“Lexus cars are very nice.” you should check both the paraphrase and the punctuation boxes.); b) At the end of the Google sheet there is a column named ‘Other’ you are welcome to write down in plain text any additional type that you observed and it is not covered by the existing classes. The provided classes are: capitalization, punctuation, paraphrase, delete fillers, completion, add context, contractions, spelling, normalization, slang/idioms, politeness, split sentences, and relativizers.
Appendix I Quantitative analysis of xformal
Following, we include more details on the qualitative analysis procedures.
A rewrite performs a capitalization edit if it contains tokens that appear in capital letters in the informal text but in lowercase letters in the formal text.
Punctuation
A rewrite contains a punctuation edits if any of the punctuation of the informal-formal texts differs.
Spelling
We identify spelling errors based on the character-based Levenshtein distance between informal-formal tokens.
Normalization
We identify normalization edits based on a hand-crafted list of abbreviations for each language.
Split sentences
We split sentences using the nltk toolkit.
Paraphrase
A formal rewrite is considered to contain a paraphrase edit if the token level Levenshtein distance between the informal-formal text is greater than 3 tokens.
Lowercase
A rewrite performs a capitalization edit if it contains tokens that appear in lower case in the informal text but in capital letters in the formal text.
Repetition
We identify repetition tokens using regular expressions (a token that appears more than 3 times in a row is considered a repetition.)
Appendix J xformal: Data Quality
We ask a native speaker to judge the quality of a sample of rewrites for each language via choosing one of the following three options: “excellent”, “acceptable”, and “poor”. The details of this analysis are included in A.H (Task 1).Results indicate that on average less than of the rewrites were identified as of poor quality across languages while more than as of excellent quality. The small number of rewrites identified as “poor” consists mostly of edits where humans add context not appearing in the original sentence. We choose not to exclude any of the rewrites from the dataset as we provide multiple reformulations of each informal instance.
Appendix K OpenSubtitles data
Appendix L Qualification tests
Table 13 presents questions and answers used in qc2 of our annotation protocol. Turkers have to score and above to participate in the task. To compute an average score for each test, we assume that an answer is incorrect if it deviates more than point from the gold-standard scores given in the second column of Table 13.
Appendix M Trade-off plots
Figure 4 presents trade-off plots between multi-bleu vs. formality score and fluency, across different reruns as proposed by Tikhonov and Yamshchikov (2018). First, we observe that models exhibit small variations in terms of formality and fluency scores across different reruns, and larger variations across bleu for most cases. Notably, the single seed backtranslate systems are the most consistent across reruns for all metrics and languages. Furthermore, in almost all cases, models trained on M data, perform better that the naive copy baseline, across metrics. However, single seed models fail to consistently outperform the rule-based baselines in almost all cases, with the exception of backtranslate for br-pt and fr which report an improvement of about bleu score.
Appendix N Compute time &\& Infrastracture
All experiments for benchmarking both nmt-based and unsupervised approaches use Amazon ec P instances: https://aws.amazon.com/ec2/instance-types/p3/, on Tesla V gpus. Concretely, nmt-based experiments are run on gpus ( - hours), while unsupervised models are run on single gpus, with average run time spanning from hours (for out-of-domain data—K), to hours (for in domain data—M).