Evaluating the Evaluation Metrics for Style Transfer: A Case Study in Multilingual Formality Transfer
Eleftheria Briakou, Sweta Agrawal, Joel Tetreault, Marine Carpuat
Introduction
Textual style transfer (st) is defined as a generation task where a text sequence is paraphrased while controlling one aspect of its style Jin et al. (2020). For instance, the informal sentence in Italian “in bocca al lupo!” (i.e., “good luck”) is rewritten to the formal version “Ti rivolgo un sincero augurio!” (i.e., “I send you a sincere wish!”). Despite the growing attention on st in the nlp literature (Jin et al., 2020), progress is hampered by a lack of standardized and reliable automatic evaluation metrics. Standardizing the latter would allow for quicker development of new methods and comparison to prior art without relying on time and cost-intensive human evaluation that is currently employed by more than of st papers Briakou et al. (2021a).
st is usually evaluated across three dimensions: style transfer (i.e., has the style of the generated output changed as intended?), meaning preservation (i.e., are the semantics of the input preserved?), and fluency (i.e., is the output well-formed?). As we will see, a wide range of automatic evaluation metrics and models has been used to quantify each of these dimensions. For example, prior work has employed as many as nine different automatic systems to rate formality alone (see Table 1). However, it is not clear how different automatic metrics compare to each other and how well they agree with human judgments. Furthermore, previous studies of automatic evaluation have exclusively focused on the English language Yamshchikov et al. (2021); Pang (2019); Pang and Gimpel (2019); Tikhonov et al. (2019); Mir et al. (2019); yet, st requires evaluation methods that generalize reliably beyond English.
We address these limitations by conducting a controlled empirical comparison of commonly used automatic evaluation metrics. Concretely, for all three evaluation dimensions, we compile a list of different automatic evaluation approaches used in prior st work and study how well they correlate with human judgments. We choose to build on available resources as collecting human judgments across the evaluation dimensions is a costly process that requires recruiting fluent speakers in each language addressed in evaluation. While there are many stylistic transformations in st, we conduct our study through the lens of formality style transfer (fost), which is one of the most popular style dimensions considered by past st work (Jin et al., 2020; Briakou et al., 2021a) and for which reference outputs and human judgments are available for four languages: English, Brazilian-Portuguese, French, and Italian.
We contribute a meta-evaluation study that is not only the first large-scale comparison of automatic metrics for st but is also the first work to investigate the robustness of these metrics in multilingual settings.
We show that automatic evaluation approaches based on a formality regression model fine-tuned on xlm-r and the chrf metric correlate well with human judgments for style transfer and meaning preservation, respectively, and propose that the field adopts their usage. These metrics are shown to work well across languages, and not just in English.
We show that framing style transfer evaluation as a binary classification task is problematic and propose that the field treats it as a regression task to better mirror human evaluation.
Our analysis code and meta-evaluation files with system outputs are made public to facilitate further work in developing better automatic metrics for st: https://github.com/Elbria/xformal-FoST-meta.
Background
Recent work highlights the need for research to improve evaluation practices for st along multiple directions. Not only does st lack standardized evaluation practices Yamshchikov et al. (2021), but commonly used methods have major drawbacks which hamper progress in this field. Pang (2019) and Pang and Gimpel (2019) show that the most widely adopted automatic metric, bleu, can be gamed. They observe that untransferred text achieves the highest bleu score for the task of sentiment transfer, questioning complex models’ ability to surpass this trivial baseline. Mir et al. (2019) discuss the inherent trade-off between st evaluation aspects and propose that models are evaluated at specific points of their trade-off plots. Tikhonov et al. (2019) argue that, despite their cost, human-written references are needed for future experiments with style transfer. They also show that comparing models without reporting error margins can lead to incorrect conclusions as state-of-the-art models sometimes end up within error margins from one another.
2 Structured Review of st Evaluation
We systematically review automatic evaluation practices in st with formality as a case study. We select fost for this work since it is one of the most frequently studied styles Jin et al. (2020) and there is human annotated data including human references available for these evaluations Rao and Tetreault (2018); Briakou et al. (2021b). Tables 1 and 2 summarize evaluation details for all fost methods in papers from the st survey by Jin et al. (2020).The complete list is hosted at: https://github.com/fuzhenxin/Style-Transfer-in-Text Most works employ automatic evaluation for style () and meaning preservation (). Fluency is the least frequently evaluated dimension (), while of papers employ automatic metrics to assess the overall quality of system outputs that captures all desirable aspects.
Across dimensions, papers also frequently rely on human evaluation (, , , and for style, meaning, fluency, and overall). However, human judgments and automatic metrics do not always agree on the best-performing system. In of evaluations, the top-ranked system is the same according to human and automatic evaluation (marked as ✓ in Table 1), and their ranking disagrees in of evaluations (marked as ✗ in Table 1). When there is a disagreement, human evaluation is trusted more and viewed as the standard. This highlights the need for a systematic evaluation of automatic evaluation metrics.
Finally, almost all papers () consider fost for English (en), as summarized in Table 2. There are only two exceptions: Korotkova et al. (2019) study fost for Latvian (lv) and Estonian (et) in addition to en, while Briakou et al. (2021b) study fost for Romance languages: Brazilian Portuguese (br-pt), French (fr), and Italian (it). The former provides system output samples as a means of evaluation, and the latter employs human evaluations, highlighting the challenges of automatic evaluation in multilingual settings.
Next, we review the automatic metrics used for each dimension of evaluation in fost papers. As we will see, a wide range of approaches is used. Yet, it remains unclear how they compare to each other, what their respective strengths and weaknesses are, and how they might generalize to languages other than English.
3 Automatic Metrics for fost
Style transfer is often evaluated using model-based approaches. The most frequent method consists of training a binary classifier on human written formal vs. informal pairs. The classifier is later used to predict the percentage of generated outputs that match the desired attribute per evaluated system—the system with the highest percentage is considered the best performing with respect to style. Across methods, the corpus used to train the classifier is the gyafc parallel-corpus Rao and Tetreault (2018) consisting of K parallel informal-formal human-generated excerpts. This corpus is curated for fost in en, while similar resources are not available for other languages. Different model architectures have been used by prior work (e.g., cnn, lstm, gru, fine-tuning on pre-trained language models such as Roberta and bert; Table 1). In most papers, the resulting classifier is evaluated on the test side of the gyafc corpus, reporting accuracy scores in the range of %. Despite the high accuracy scores, the best ranking system under the classifier is very often in disagreement with human evaluations (marked as ✗ under the third subcolumn of style of Table 1). A few works train regression-based models instead, using the training data of Pavlick and Tetreault (2016) that are human-annotated for formality on a -point scale—while, again, this resource is only available for en.
Evaluation of this dimension is performed using a wider spectrum of approaches, as presented in the third column of Table 1. The most frequently used metric is reference-bleu (r-bleu), which is based on the -gram precision of the system output compared to human rewrites of the desired formality. Other approaches include self-bleu (s-bleu), where the system output is compared to its input, measuring the semantic similarity between the system input and its output, or regression models (e.g., cnn, bert) trained on data annotated for similarity-based tasks, such as the Semantic Textual Similarity task (sts) Agirre et al. (2016).
Fluency is typically evaluated with model-based approaches (see fourth column of Table 1). Among those, the most frequent method is that of computing perplexity (ppl) under a language model. The latter is either trained from scratch on the same corpus used to train the fost models (i.e., gyafc) using different underlying architectures (e.g., Kenlm, lstm), or employ large pre-trained language models (e.g., gpt). A few other works train models on en data annotated for grammaticality (Heilman et al., 2014) or linguistic acceptability (Warstadt et al., 2019) instead.
Systems’ overall quality (see fifth column of Table 2) is mostly evaluated using r-bleu or by combining independently computed metrics into a single score (e.g., geometric mean - gm(.), harmonic mean - hm(.), f(.)). Moreover, out of approaches that rely on combined scores do not include fluency scores in their overall evaluation.
Since most of the current work on fost and st is in en, prior work relies heavily on en resources for designing automatic evaluation methods. For instance, resources for training stylistic classifiers or regression models are not available for other languages. For the same reason, it is unclear whether model-based approaches for measuring meaning preservation and fluency can be ported to multilingual settings. Furthermore, reference-based evaluations (e.g., r-bleu) require human rewrites that are only available for en, br-pt, it, and fr. Finally, even though perplexity does not rely on annotated data, without standardizing the data language models are trained on, we cannot make meaningful cross-system comparisons.
4 Summary
Reviewing the literature shows the lack of standardized metrics for st evaluation, which hampers comparisons across papers, the lack of agreement between human judgments and automatic metrics, which hampers system development, and the lack of portability to languages other than English which severely limits the impact of the work. These issues motivate the controlled multilingual evaluation of evaluation metrics in our paper.
Evaluating Evaluation Metrics
We evaluate evaluation metrics (described in §3.2) for multilingual fost, in four languages for which human evaluation judgments (described in §3.1) on fost system outputs are available.
We use human judgments collected by prior work of Rao and Tetreault (2018) for en and Briakou et al. (2021b) for br-pt, fr, and it. We include details on their annotation frameworks, the quality of human judges, and the evaluated systems below.
We briefly describe the annotation frameworks employed by Rao and Tetreault (2018) and Briakou et al. (2021b) to collect human judgments for each evaluation aspect: 1. formalityratings are collected—for each system output—on a -point discrete scale, ranging from to , as per Lahiri (2015) (Very informal, Informal, Somewhat Informal, Neutral, Somewhat Formal, Formal. Very Formal); 2. meaning preservationjudgments adopt the Semantic Textual Similarity annotation scheme of Agirre et al. (2016), where an informal input and its corresponding formal system output are rated on a scale from to based on their similarity (Completely dissimilar, Not equivalent but on same topic, Not equivalent but share some details, Roughly equivalent, Mostly equivalent, Completely equivalent); 3. fluencyjudgments are collected for each system output on a discrete scale of to , as per Heilman et al. (2014) (Other, Incomprehensible, Somewhat Comprehensible, Comprehensible, Perfect); 4. overalljudgments are collected following a relative ranking approach: all system outputs are ranked in order of their formality, taking into account both meaning preservation and fluency.
Both studies recruited workers from the Amazon Mechanical Turk platform after employing quality control methods to exclude poor quality workers (i.e., manual checks for en, and qualification tests for br-pt, fr, and it). For all human evaluations and languages Briakou et al. (2021b) report at least moderate inter-annotator agreement.
The evaluated system outputs were sampled from fost models for each language, spanning a range of simple baselines to neural architectures Rao and Tetreault (2018); Briakou et al. (2021b). We also include detailed descriptions of them in Appendix B. For each evaluation dimension outputs are evaluated for en and outputs per system for br-pt, fr, and it.
2 Evaluation Metrics
For the fost evaluation aspects described below, we cover a broad spectrum of approaches that range from dedicated models for the tasks at hand to more lightweight methods relying on unsupervised approaches and automated metrics.
We benchmark model-based approaches that fine-tune multilingual pre-trained language models (i.e., xlm-r, mbert), where the task of formality detection is modeled either as a binary classification task (i.e., formal vs. informal), or as a regression task that predicts different formality levels on an ordinal scale.
We evaluate the bleu score Papineni et al. (2002) of the system output compared to the reference rewrite (r-bleu) since it is the dominant metric in prior work. Prior reviews of meaning preservation metrics for paraphrase and sentiment st tasks in en (Yamshchikov et al., 2021) cover -gram metrics and embedding-based approaches. We consider three additional metric classes to compare system outputs with inputs, as human annotators do:
n-gram based metrics include: s-bleu (self-bleu that compares system outputs with their inputs as opposed to references, i.e., r-bleu), meteor Banerjee and Lavie (2005) based on the harmonic mean of unigram precision and recall while accounting for synonym matches, and chrF Popović (2015) based on the character -gram F-score;
embedding-based methods fall under the category of unsupervised evaluation approaches that rely on either contextual word representations extracted from pre-trained language models or non-contextual pre-trained word embeddings (e.g., word2vec Mikolov et al. (2013); Glove Pennington et al. (2014)). For the former, we use bert-score Zhang et al. (2020a) which computes the similarity between each output token and each reference token based on bert contextual embeddings. For the latter, we experiment with two similarity metrics: the first is the cosine distance between the sentence-level feature representations of the compared texts extracted via averaging their word embeddings; the second is the Word Mover’s Distance (wmd) metric of Kusner et al. (2015) that measures the dissimilarity between two texts as the minimum amount of distance that the embedded words of one text need to “travel" to reach the word embeddings of the other;
semantic textual similarity (sts) models constitute supervised methods that we model via fine-tuning multilingual pre-trained language models (i.e., xlm-r, mbert) to predict a semantic similarity score for a pair of texts on an ordinal scale.
We experiment with perplexity (ppl) and likelihood (ll) scores based on probability scores of language models trained from scratch (e.g., Kenlm Heafield (2011)), as well as pseudo-likelihood scores (pseudo-ll) extracted from pre-trained masked language models similarly to Salazar et al. (2020), by masking sentence tokens one by one.
Experiment Settings
For supervised model-based methods that rely on the availability of human-annotated instances to train dedicated models for specific tasks, we experiment with three standard cross-lingual transfer approaches (e.g., Hu et al. (2020)): 1. zero-shottrains a single model on the en training data and evaluates it on the original test data for each language; 2. translate-trainuses machine translation (mt) to obtain training data in each language through translating the en training set—and trains independent systems for each language; 3. translate-testtrains a single model on the en training data and evaluates it on the test data that are translated into en using mt.
For meaning preservation metrics, we use the open-sourced implementations of: Post (2018) for bleu Papineni et al. (2002); Banerjee and Lavie (2005) for meteor; Popović (2015) for chrf.https://github.com/mjpost/sacrebleu,https://www.cs.cmu.edu/~alavie/METEOR/,https://github.com/m-popovic/chrF For bert-score we use the implementation of Zhang et al. (2020a);https://github.com/Tiiiger/bert_score non-contextualized embeddings-based approaches are based on fastText pre-trained embeddings.https://fasttext.cc For fluency metrics, we use the implementation of Salazar et al. (2020) for computing pseudo-likelihood.https://github.com/awslabs/mlm-scoring ppl and ll scores are extracted from a -gram Kenlm model (Heafield, 2011).https://github.com/kpu/kenlm
Table 3 presents statistics on the training data used for supervised and unsupervised models across the st evaluation aspects. For datasets that are only available for en, we use the already available machine translated resources for sts https://github.com/PhilipMay/stsb-multi-mt and formality datasets Briakou et al. (2021b). The former employs the DeepL service (no information of translation quality is available) while the latter uses the aws translation servicehttps://aws.amazon.com/translate (with reported bleu scores of (br-pt), (fr), and (it)).bleu scores were computed on randomly sampled data from OpenSubtitles. The Kenlm models for all the languages are trained on M randomly sampled sentences from the OpenSubtitles dataset Lison and Tiedemann (2016).
Experimental Results
We analyze the results of comparing the outputs from the several automatic metrics to their human-generated counterparts for formality style transfer (§5.1), meaning preservation (§5.2), fluency (§5.3) via conducting segment-level analysis—and then, turn into analyzing system-level rankings to evaluation overall task success (§5.4).
The field is divided on the best way to evaluate the style dimension – formality in our case. Practitioners use either a binary approach (is the new sentence formal or informal?) or a regression approach (how formal is the new sentence?). We discuss the first approach and its limitations in § 5.1.1, before moving to regression in § 5.1.2.
As discussed in §2, the vast majority of fost works evaluate style transfer based on the accuracy of a binary classifier trained to predict whether human-written segments are formal or informal. Yet, as Table 1 indicates, this approach fails to identify the best system in this dimension 59% of the time. To better understand this issue, we evaluate these classifiers on human-written texts versus st system outputs.
Table 4 presents F scores when testing the binary formality classifiers on the task they are trained on: predicting whether human-written sentences from gyafc and xformal are formal or informal. First, the last column (i.e., (xlm-r, mbert)) shows that xlm-r is a better model than mbert for this task, across languages, with the largest improvements in the zero-shot setting where xlm-r beats mbert by , , for br-pt, fr, and it respectively.
Second, zero-shot is surprisingly the best strategy to port en models to other languages. translate-train and translate-test hurt F1 by and points on average compared to zero-shot, despite exploiting more resources in the form of machine translation systems and their training data. However, transfer accuracy is likely affected by regular translation errors (as suggested by larger F1 drops for languages with lower mt bleu scores) and by formality-specific errors. Machine translation has been found to produce outputs that are more formal than its inputs (Briakou et al., 2021b), which yields noisy training signals for translate-train and alters the formality of test samples for translate-test.
We now evaluate the best performing binary classifier (i.e., xlm-r in zero-shot setting) on real system outputs—a setup in line with automatic evaluation frameworks. Figure 1 presents a breakdown of the number of formal vs. informal predictions of the classifiers binned by human-rated formality levels. Across languages, the performance of the classifier deteriorates as we move away from extreme formality ratings (i.e., very informal () and very formal ()). This lack of sensitivity to different formality levels is problematic since system outputs across languages are concentrated around neutral formality values. In addition, when testing on br-pt, fr, and it (zero-shot settings), the classifier is more biased towards the formal class, which leads one to question its ability to correctly evaluate more formal outputs in multilingual settings. Taken together, these results suggest that validating the classifiers against human rewrites rather than system outputs is unrealistic and potentially misleading.
1.2 Regression Models
Table 5 presents Spearman’s correlation of regression models’ predictions with human judgments. Again, xlm-r with zero-shot transfer yields the highest correlation across languages. More specifically, the trends across different transfer approaches and different pre-trained language models are similar to the ones observed on evaluation of binary classifiers: xlm-r outperforms mbert for almost all settings, while zero-shot is the most successful transfer approach, followed by translate-train, with translate-test yielding the lowest correlations across languages. Interestingly, regression models highlight the differences between the generalization abilities of xlm-r and mbert more clearly than the previous analysis on binary predictions: zero-shot transfer on xlm-r yields , , and higher correlations than mbert for br-pt, fr, and it—while both models yield similar correlations for en.
2 Meaning Preservation Metrics
Table 6 presents Spearman’s correlation of meaning preservation metrics with human judgments. chrF consistently yields the highest correlations across languages—this result is in line with prior observations on evaluating meaning preservation metrics for en st tasks Yamshchikov et al. (2021) and is now confirmed in a multilingual setting. This trend might be explained by chrf’s ability to match spelling errors within words via character -grams. xlm-r trained on sts with zero-shot transfer is a close second to chrf, consistent with this model’s top-ranking behavior as a formality transfer metric. However, chrf outperforms the remaining more complex and expensive metrics, including bert-score and mbert models. In contrast to Yamshchikov et al. (2021), embedding-based methods (i.e., cosine, wmd) show no advantage over -gram metrics, perhaps due to differences in word embedding quality across languages. Finally, it should be noted that r-bleu is the worst performing metric across languages, and its correlation with human scores is particularly poor for languages other than English. This is remarkable because it has been used in of automatic evaluations for fost meaning preservation evaluation (as seen in Table 1). We, therefore, recommend discontinuing its use.
3 Fluency Metrics
Table 7 presents Spearman’s correlation of various fluency metrics with human judgments. Pseudo-likelihood (pseudo-ll) scores obtained from xlm-r correlate with human fluency ratings best across languages. Their correlations are strong across languages, while other methods only yield weak (i.e., Kenlm, mbert) to moderate correlations (i.e, Kenlm-ppl) for it. We, therefore, recommend evaluating fluency using Pseudo-likelihood scores derived from xlm-r to help standardize fluency evaluation across languages.
4 System-level Rankings
Finally, we turn to predict the overall ranking of systems by focusing on how many correct pairwise system comparisons each metric gets correct. For each language, there are systems, which means there are pairwise comparisons, for a total of given the languages. We analyze corpus-level r-bleu, commonly used for this dimension, along with leading metrics from the other dimensions: xlm-r formality regression models, chrf and xlm-r pseudo-likelihood. r-bleu gets out of comparisons correct while the other metrics get , , and respectively. This indicates that r-bleu correlates with human judgments better at the corpus-level than at the sentence-level, as in machine translation evaluation (Mathur et al., 2020). We caution that these results are not definitive but rather suggestive of the best performing metric, given the ideal evaluation would be a larger number of systems with which to perform a rank correlation. The complete analysis for each language is in Appendix A.
Conclusions
Automatic (and human) evaluation processes are well-known problems for the field of Natural Language Generation Howcroft et al. (2020); Clinciu et al. (2021) and the burgeoning subfield of st is not immune. st, in particular, has suffered from a lack of standardization of automatic metrics, a lack of agreement between human judgments and automatics metrics, as well as a blindspot to developing metrics for languages other than English. We address these issues by conducting the first controlled multilingual evaluation for leading st metrics with a focus on formality, covering metrics for evaluation dimensions and overall ranking for languages. Given our findings, we recommend the formality style transfer community adopt the following best practices:
xlm-r formality regression models in the zero-shot cross-lingual transfer setting yields the clear best metrics across all four languages as it correlates very well with human judgments. However, the commonly used binary classifiers do not generalize across languages (due to misleadingly over-predicting formal labels). We propose that the field use regression models instead since they are designed to capture a wide spectrum of formality rates.
We recommend using chrf as it exhibits strong correlations with human judgments for all four languages. We caution against using bleu for this dimension, despite its overwhelming use in prior work as both its reference and self variants do not correlate as strongly as other more recent metrics.
xlm-r is again the best metric (in particular for French). However, it does not correlate well with human judgments as compared to the other two dimensions.
chrf and xlm-r are the best metrics using a pairwise comparison evaluation. However, an ideal evaluation would be to have a large number of systems with which to draw reliable correlations.
Our results support using zero-shot transfer instead of machine translation to port metrics from English to other languages for formality transfer tasks.
We view this work as a strong point of departure for future investigations of st evaluation. Our work first calls for exploring how these evaluation metrics generalize to other styles and languages. Across the different ways of defining style evaluation (either automatic or human), prior work has mostly focused on the three main dimensions covered in our study. As a result, although our meta-evaluation on st metrics focuses on formality as a case study, it can inform the evaluation of other style definitions (e.g., politeness, sentiment, gender, etc.). However, more empirical evidence is needed to test the applicability of the best performing metrics for evaluating style transfer beyond formality. Our work suggests that the top metrics based on xlm-r and chrf are robust across Romance languages; yet, our conclusions and recommendations are currently limited to this set of languages. We hope that future work in multilingual style transfer will allow for testing their generalization to a broader spectrum of languages and style definitions. Furthermore, our study highlights that more research is needed on automatically ranking systems. For example, one could build a metric that combines metrics’ outputs for the three dimensions, or one could develop a singular metric. In line with Briakou et al. (2021a), our study also calls for releasing more human evaluations and more system outputs to enable robust evaluation. Finally, there is still room for improvement in assessing how fluent a rewrite is. Our study provides a framework to address these questions systematically and calls for st papers to standardize and release data to support larger-scale evaluations.
Acknowledgements
We thank Sudha Rao for providing references and materials of the gyafc dataset, Jordan Boyd-Graber, Pedro Rodriguez, the clip lab at umd, and the emnlp reviewers for their helpful and constructive comments.
References
Appendix A System-level Analysis
Table 8 presents the number of correct system-level pair-wise comparisons of automatic metrics based on human judgments. For sts, chrf, f.reg*, f.class*, and pseudo-lkl*, system-level scores are extracted via averaging sentence-level scores. For s-bleu and r-bleu the system scores are extracted at the corpus-level. The total number of pairwise comparisons for each language is (given access to systems). Among the meaning preservation metrics (i.e., sts, s-bleu, and chrf), chrf yields the highest number of correct comparisons (i.e., out of for all languages). The formality regression models (i.e., f.reg*) result in correct rankings more frequently than the formality classifiers (i.e., f.class*) yielding out of correct comparisons. Reference-bleu (i.e., r-bleu) is compared with overall ranking judemnts. It ranks out of systems correctly for en, fr, and br-pt and only for it. Finally, perplexity (i.e., ppl) results in the fewest correct rankings at system-level (i.e., out of ), despite correlating well with human judgments at the segment-level.
Additionally, in Figure 2 we visualize the differences between relative rankings induced by human judgments and the best segment-level correlated metrics for each dimension, averaged per system.
Appendix B Evaluated Systems Details
For each of br-pt, it, and fr, outputs are sampled from:
Rule-based systems consisting of hand-crafted transformations (e.g., fixing casing, normalizing punctuation, expanding contractions, etc.);
Round-trip translation models that pivot to en and backtranslate to the original language;
Bi-directional neural machine translation (mt) models that employ side constraints to perform style transfer for both directions of formality (i.e., informalformal)—trained on (machine) translated informal-formal pairs of an English parallel corpus (i.e., gyafc);
Bi-directional nmt models that augment the training data of 3. via backtranslation of informal sentences;
A multi-task variant of 3. that augments the training data with parallel-sentences from bilingual resources (i.e., OpenSubtitles) and learns to translate jointly between and across languages.
A rule-based system of similar transformations to ones for br-pt, fr, and it;
A phrase-based machine translation model trained on informal-formal pairs of gyafc;
An nmt model trained on gyafc to perform style transfer uni-directionaly;
A variant of 3. that incorporates a copy-enriched mechanism that enables direct copying of words from input;
A variant of 4. trained on additional back-translated data of target style sentences using 2.
In general, neural models performed best for all languages according to overall human judgments, while the simpler baselines perform closer to the more advanced neural models for br-pt, fr, and it. For each evaluation dimension outputs are evaluated for en and outputs per system for br-pt, fr, and it.
Appendix C Meaning Preservation Metrics (reference-based)
Table 9 presents supplemental results on meaning preservation metrics for reference-based settings.