X-Eval: Generalizable Multi-aspect Text Evaluation via Augmented Instruction Tuning with Auxiliary Evaluation Aspects
Minqian Liu, Ying Shen, Zhiyang Xu, Yixin Cao, Eunah Cho, Vaibhav Kumar, Reza Ghanadan, Lifu Huang
Introduction
Recent advancements of pre-training Chung et al. (2022); Touvron et al. (2023a, b), prompting Brown et al. (2020); Wei et al. (2022b); Wang et al. (2023); Yao et al. (2023); Qi et al. (2023), and instruction tuning Wei et al. (2022a) have improved the quality of machine generated texts by a significant degree. Nevertheless, the evaluation of various Natural Language Generation (NLG) tasks still lags far behind compared with the rapid progress of large language models (LLMs). Previous similarity-based metrics such as ROUGE Lin (2004), BLUE Papineni et al. (2002), METEOR Banerjee and Lavie (2005), and BERTScore Zhang* et al. (2020) predominantly measures the similarity between the generated and reference text, failing to accurately reflect the quality of generated text Gehrmann et al. (2023), especially for open-ended generation tasks.
To obtain a more comprehensive assessment of text quality, multi-aspect evaluation Fabbri et al. (2021) has been proposed to evaluate the generated text from multiple fine-grained evaluation aspects, such as fluency and consistency. While most existing studies Mehri and Eskenazi (2020b); Yuan et al. (2021); Zhong et al. (2022) consider a closed set of aspects, in many realistic scenarios, the users may need to evaluate the text with their customized aspects and specifications, calling for building an evaluator that can be flexibly extended to any unseen aspects without the need of training data. Recent studies Fu et al. (2023); Liu et al. (2023) propose to leverage LLMs (e.g., GPT-4 OpenAI (2023)) as NLG evaluators, yielding promising zero-shot performance on unseen aspects. However, such evaluations, especially with proprietary LLMs, are cost-intensive, time-consuming, and pose concerns about data privacy and reproducibility.
In this work, we propose X-Eval, an automatic evaluation framework that can conduct fine-grained evaluation on both seen and unseen aspects across various NLG tasks with a single model. X-Eval follows a two-stage training paradigm: we first instruction-finetune an LLM to equip the model with the capability of following human-written instructions for evaluation. Then, motivated by the observation that the evaluation aspects usually exhibit inter-connections Fu et al. (2023) and thus their evaluations can benefit each other, we introduce an additional training stage to fine-tune the LLM on the instruction-tuning tasks enriched with the evaluations of a set of auxiliary aspects, which are expected to provide clues for evaluating the target aspect and encourage consistent evaluations across multiple aspects. During training, for each target aspect, we take all the remaining aspects defined in the corresponding original dataset as auxiliary aspects, convert their human-annotated evaluations into natural languages based on templates, and incorporate them into the instruction. During inference, we first obtain the evaluation predicted from the model for each aspect. Then, we take each aspect as the target, select its auxiliary aspects based on the similarity of the aspect definitions, and re-perform evaluation by considering the evaluations of auxiliary aspects.
To support our proposed two-stage training of X-Eval, we construct AspectInstruct, the first multi-aspect evaluation instruction tuning dataset spanning 27 diverse evaluation aspects over 65 tasks. This dataset is anchored around three core categories of NLG tasks: dialogue, summarization, and data-to-text. In light of insights from previous studies in instruction tuning Wei et al. (2022a); Xu et al. (2023b), which emphasize the advantage of task diversity in enhancing zero-short generalization, we further augment the dataset by converting the original human rating task into diverse forms of NLG evaluation tasks, including scoring, comparison, ranking and Boolean question answering. In addition, in order to incorporate auxiliary aspects, we manually create templates that convert the numerical evaluation scores of each aspect into natural language based descriptions.
The main advantages of our approach are highlighted as follows: (1) Generalization ability: we introduce X-Eval that can be flexibly generalized to evaluate the unseen NLG tasks or the aspects customized by user instructions in a zero-shot manner with a single model; (2) Strong performance with high efficiency: with significantly less amount of model parameters (780M), X-Eval achieves strong performance compared to the state-of-the-art LLM-based evaluators (including GPT-4) demonstrated through comprehensive experiments; (3) Reference-free and open-source: our evaluator does not require gold reference to perform evaluation and it is more reliable and transparent thanks to its open-source nature.
Related Work
The previously dominant text evaluation paradigm is to predict a one-size-fits-all evaluation score, where most of them are similarity-based metrics: One line of research, such as ROUGE Lin (2004), BLUE Papineni et al. (2002), and METEOR Banerjee and Lavie (2005) measures the surface overlap between the generated and reference text. Another line of work, such as BERTScore Zhang* et al. (2020) and MoverScore Zhao et al. (2019), measures the distance between the contextualized embeddings of the generated text and the reference as the similarity score. Although these metrics are widely adopted, they often overlook fine-grained aspects and later study Gehrmann et al. (2023) has proven that they fail to truly capture the quality of text with the coarse-grained score.
Multi-Aspect Metrics
To conduct a more holistic evaluation, recent studies Wang et al. (2020a); Huang et al. (2020) propose to evaluate the NLG systems via multiple fine-grained aspects. UniEval Zhong et al. (2022) proposes to re-frame NLG evaluation into a question answering format and performs multi-aspect evaluation with a single model via continual learning Madotto et al. (2021); Liu et al. (2022); Liu and Huang (2023). However, UniEval cannot maintain robust performance when generalizing to novel aspects. To obtain an evaluator that can be generalized to customized aspects, some recent studies Fu et al. (2023); Liu et al. (2023) harness proprietary large language models (LLMs), such as GPT-4 OpenAI (2023) to perform fine-grained evaluation in a zero-shot manner. However, due to the closed-source nature, these evaluation metrics suffer from issues of reproducibility and are prohibitively expensive. More recently, some works that are concurrent to ours Xu et al. (2023a); Jiang et al. (2023) proposed to extract instruction-following data from proprietary LLMs for fine-tuning a more lightweight language model as the evaluator. Nevertheless, they still require high costs to call the APIs to obtain a large amount of training data and it is non-trivial to ensure the data are of high quality. In addition, to the best of our knowledge, we are the first to meticulously curate the instruction-tuning dataset and train an instruction-based evaluator for dialogue evaluation.
AspectInstruct
Multi-aspect automatic text evaluation aims to evaluate the quality of NLG system’s output given a set of evaluation aspects (e.g., coherence, naturalness and so on), and optionally an additional set of texts (e.g., the source documents for text summarization, or context for dialogue evaluation). The evaluation task can be formulated as:
where is the fine-grained aspect to be evaluated, and is the scoring function that provides an assessment w.r.t. the aspect .
2 Data Collection
We aim to build a unified automatic evaluation framework that can assess the text quality for both seen and unseen NLG tasks and evaluation aspects via instruction tuning. To this end, we build an instruction-tuning dataset tailored for multi-aspect evaluation, namely AspectInstruct, with the following steps:
We first collect 10 existing evaluation datasets with human annotations for 3 representative categories of NLG tasks, including dialogue generation Sai et al. (2020); Gunasekara et al. (2020); Pang et al. (2020); Gopalakrishnan et al. (2019); Mehri and Eskenazi (2020a), text summarization Völske et al. (2017); Fabbri et al. (2021); Wang et al. (2020b); Zhong et al. (2022), and data-to-text Wen et al. (2015).
Task Augmentation
The original datasets we collect only contain numerical scores annotated by human experts, which severely limits the diversity of instruction-tuning tasks. Thus, we further derive diverse forms of evaluation tasks from the original annotations to enhance the diversity of task formats. Denote the ground truth score for text as . We derive four types of tasks based on this annotation: (1) Scoring: we ask the model to directly predict a discrete score (e.g., in the Likert scale) by mapping the continuous ground truth into a discrete scale; (2) Comparison: we sample two texts and for an identical context, e.g., two versions of summaries for the same source document, and ask the model to select the text with the higher ground truth score; (3) Ranking: we further extend the comparison task into ranking by sampling three (or more) candidates under the same context and ask the model to predict the correct ranking of the candidates based on the text quality; (4) Boolean Question Answering: we also formulate evaluation as a Boolean QA task following Zhong et al. (2022) by asking the model a question such as "Is this response fluent?" and let the model predict "Yes" or "No".
Instruction Creation
Finally, we define a unified format for the instructions of tasks included in AspectInstruct. Each instruction consists of three parts: (1) task description that briefly introduces the evaluation task, (2) aspect definition, and (3) evaluation protocol that details what the model should output to perform the evaluation. To curate the definition for each aspect, we first refer to the definition of the aspect in the original annotation guideline. When a definition is absent from the guideline, three human experts construct and revise the definition until they reach an agreement. The task descriptions and evaluation protocols are also written by three human experts in similar ways. We provide an illustrative example of the original annotation, the derived four types of evaluation tasks along with the curated instructions in Figure 2. The full list of evaluation aspects and the collected instructions can be found in Appendix A.1.
Statistics
In total, we constructed 65 tasks in AspectInstruct, where we split 32 tasks with 14 seen aspects for instruction tuning and 33 tasks with 13 unseen aspects for meta-evaluation. We collected 72,637 instances in total with 55,602 instances for training and 17,035 instances for inference. Note that there is no overlap among the datasets used for training and inference. We consider two aspects that have identical aspect names but are in different NLG tasks as distinct aspects since their definitions are drastically different. We include more details on the source datasets, constructed instruction-tuning tasks, and the number of instances of each task in Appendix A.2.
X-Eval
Figure 3 presents an overview of X-Eval, which consists of two stages of instruction tuning:
The first training stage aims to equip the model with the ability to follow instructions to perform diverse evaluation tasks. In this stage, we adopt FLAN-T5 Chung et al. (2022), an open-source large language model to act as the base model for our evaluator. Here, we perform standard instruction tuning on the mixture of four types of instruction tuning tasks: scoring, comparison, ranking, and Boolean QA, as elaborated in Section 3.2.
Instruction Tuning with Auxiliary Aspects
Through our study, we discern that certain evaluation aspects could be interrelated. As evidence, within the Topical-Chat dataset Gopalakrishnan et al. (2019), the aspect naturalness usually shows a notable correlation with engagingness. When a dialogue response is not natural, it is very likely that human considers the response to be not engaging. While these two aspects are not interchangeable given their different definitions, the evaluation of one aspect can offer useful clues for the evaluation of another aspect that has certain underlying connections. Motivated by this, we enrich our training regimen with an additional instruction tuning stage with auxiliary aspects to leverage their potential connections to the target evaluation aspect. As such, the model can provide a more accurate assessment with the additional information from the auxiliary aspects.
More precisely, for each instruction-tuning task detailed in Section 3.2, we augment it based on the ground truth evaluation results of a predefined set of auxiliary aspects, where we select all other auxiliary aspects collected in the corresponding source dataset. To convert the evaluation results of auxiliary aspects into natural language that can be fed into the input, we employ a template-based evaluation verbalizer, denoted as , which takes in an aspect and its corresponding evaluation score for a particular instance, mapping it into a verbalized auxiliary evaluation . For example, with the aspect Consistency on the Data2Text task and the evaluation score 0.9 out of 1.0, the verbalized result is phrased as "This sentence is consistent with the source.", as depicted in Figure 3. We construct the set of verbalized results using the evaluation verbalizer for each auxiliary aspect (except for the target aspect). Once crafted, this set is then concatenated into the additional set of texts in the evaluator’s input, formalized as . Within the second training stage, the evaluator undergoes further finetuning on the instruction tasks enriched with these evaluation results.
2 Inference with Auxiliary Aspects
After X-Eval is fine-tuned with two-stage instruction tuning, the model has been trained to follow both basic evaluation instructions and leverage the instructions enhanced by auxiliary aspects. At the inference stage, we perform the following steps to obtain the evaluation of the target aspect: First, based on the definitions of the target aspect and a pool of candidate aspects, we employ Sentence-T5 Ni et al. (2022) to encode the definitions and then measure the similarity between the sentence embeddings of target aspect definition and each candidate aspect definition. Since the candidate pool can be large, we select the aspects with the top- similarity scores as the auxiliary aspects to limit the inference cost, where is the hyperparameter. Second, we run an inference process on each aspect (including both seen and unseen aspects) and convert the prediction into natural language results. These verbalized results, denoted as , are subsequently integrated into the additional set of texts for evaluating the target aspect. Finally, we utilize the Boolean question-answering task format where the model predicts either "Yes" or "No", as outlined in Section 3.2, to compute the evaluation score of the target aspect:
where denotes the probability of the model generating a specific word. Note that we adopt the same format during the evaluation of the auxiliary aspects. Algorithm 1 provides the pseudo-code detailing our proposed inference pipeline.
Experiment Setup
We meta-evaluate our X-Eval on the test split of AspectInstruct, where the details of the test set are introduced as follows. For text summarization, we adopt SummEval Fabbri et al. (2021) and QAGS Wang et al. (2020b). For dialogue generation, we employ Topical-Chat Gopalakrishnan et al. (2019) and FED Mehri and Eskenazi (2020a). For data-to-text generation, we utilize SFHOT & SFRES Wen et al. (2015). AspectInstruct contains the following unseen aspects: topic depth (DEP), likeability (LIK), understandability (UND), flexibility (FLE), informativeness (INF), inquisitiveness (INQ), interestingness (INT), specificity (SPE), correctness (COR), and semantic appropriateness (SEM). More detailed descriptions of the benchmarks, as well as seen and unseen evaluation aspects, can be found in Appendix A.3.
2 Baselines and Variants of X-Eval
We compare our X-Eval with the following state-of-the-art NLG evaluation metrics: (1) BERTScore Zhang* et al. (2020) is a similarity-based evaluator. It uses the contextualized representation from BERT Devlin et al. (2019) to compute the similarity between the generated text and the reference text; (2) MoverScore Zhao et al. (2019) goes beyond BERTScore by utilizing soft alignments (many-to-one) and new aggregation methods on the layer-wise information, resulting in a more powerful similarity-based evaluator; (3) USR Mehri and Eskenazi (2020b) is an unsupervised and reference-free evaluation metric for dialog generation. It employs various variants to predict multiple scores, reflecting the diverse qualities of dialogue; (4) BARTScore Yuan et al. (2021) is a unified evaluator based on BART Lewis et al. (2019), which uses the average likelihood of the model output as the metric; (5) UniEval Zhong et al. (2022) is a unified multi-dimensional evaluator that can evaluate different aspects of text generation by re-framing the evaluation process as a Boolean Question Answering (QA) task. (6) GPTScore Fu et al. (2023) is a customized, multi-faceted, and training-free evaluation framework that utilizes the emergent abilities of generative pre-trained models to score generated texts; (7) G-Eval Liu et al. (2023) proposes to leverage large language models such as GPT-3.5 or GPT-4 to assess the text quality with chain-of-thoughts and form-filling paradigm.
Variants of X-Eval
We also design several variants of X-Eval for ablation studies: (1) X-Eval w/o Training denotes the vanilla FLAN-T5-large model without any further finetuning on our proposed AspectInstruct; (2) X-Eval w/o Instructions: based on Flan-T5, we only conduct multi-task training and inference in the similar way as Zhong et al. (2022) and we do not provide any instructions to the model; (3) X-Eval w/o Stage-Two Tuning: for this variant, we only conduct vanilla instruction tuning in Stage 1 based on Flan-T5 and we do not apply the instruction tuning with auxiliary aspects stage. During inference, we directly perform evaluation based on instructions without using auxiliary aspects.
3 Implementation Details
We adopt FLAN-T5-large as our base LM for subsequent finetuning. In the first instruction tuning stage, we set the number of epochs to 2, the learning rate to 5e-05, and the maximum source length to 1024. The second training stage shares the same setup except the number of epochs set to 1. We set the maximum source length during inference to 2048 and pick the top-1 aspect during inference, i.e., . We use sentence-T5-large to compute the embeddings for aspect definition for auxiliary aspect selection.
Main Results
To assess X-Eval’s ability to generalize to unseen aspects, we present the dialogue-level Spearman correlation results on the FED benchmark in Table 1. The table’s upper section delineates the performance of traditional metrics and evaluators based on lightweight open-source models. X-Eval significantly surpasses baseline metrics in the top section, with over 45% improvement in Spearman correlation, demonstrating greater adaptability to new aspects than conventional evaluators. In addition, despite its modest 780M parameter size, X-Eval matches the performance of proprietary LLM-based baselines shown in the table’s middle section. It is also worth noting that UniEval achieves notably poor performance on the dialogue-level evaluation on FED. One plausible reason is that UniEval has been overfitted to the turn-level evaluation and failed to generalize to dialogue-level evaluation.
The bottom section of the table provides a comparison among different variants of X-Eval. “X-Eval w/o Training” exhibits a sub-optimal average correlation, even when inference with identical instructions, highlighting the importance of finetuning on AspectInstruct. By comparing the correlations between X-Eval and “X-Eval w/o Instructions”, we observe a clear performance drop after removing detailed evaluation instructions during training, indicating the effectiveness of instruction tuning on comprehensive evaluation instructions. Furthermore, the average correlation of X-Eval improves from to with the inclusion of the second training stage, emphasizing the effectiveness of the auxiliary aspects in evaluation. We observe that X-Eval achieves relatively poorer performance on likeability and flexibility. The potential reason is that it’s more challenging for the model to evaluate the aspects that are more subjective since it requires capturing more nuances and subtleties of language.
Seen Aspect Evaluation on Topical-Chat
We further evaluate how various evaluators align with human ratings for the seen aspects of dialogue response on the Topical-chat dataset. Table 2 reports both the Pearson and Spearman correlation for each aspect. In alignment with previous findings, X-Eval outperforms traditional metrics and evaluators based on lightweight open-source models in averaged Pearson and Spearman correlations. Notably, X-Eval also surpasses all the proprietary LLM-based baselines in averaged Spearman correlations. The bottom section of the table highlights the performance enhancements achieved from training on AspectInstruct, incorporating comprehensive evaluation instructions, and integrating auxiliary aspects. We notice that X-Eval achieves relatively lower correlation on naturalness and engagingness. The plausible reason is that it’s challenging for X-Eval to capture subjective characteristics, such as human likeness and whether the response fulfills human’s implicit intends.
2 Results on Summarization Evaluation
Following Zhong et al. (2022), we use summary-level Spearman and Kendall-Tau correlation to assess the performance of various evaluators on the seen aspects of SummEval. As shown in Table 3, X-Eval surpasses traditional metrics and evaluators based on lightweight models in averaged Spearman correlation. Against proprietary LLM-based metrics, X-Eval outperforms both GPTScore and G-Eval (GPT-3.5) in both averaged Spearman and Kendal-Tau correlations. However, G-Eval (GPT-4) consistently excels across all aspects, while X-Eval slightly lags in the summarization task. We speculate this might stem from GPT’s strong ability to handle long input contexts. In addition, we report the results on QAGS in Table 4.
3 Results on Unseen NLG Task Evaluation
In this experimental scenario, we also evaluate X-Eval on the unseen data-to-text generation task. Table 5 reveals that while X-Eval experiences a slight performance loss in the naturalness aspect compared to G-Eval (GPT-4), it consistently excels over all other baselines across all aspects, achieving the highest averaged Spearman correlation. This underscores the generalization capability of X-Eval on unseen NLG tasks.
Discussions
We conduct ablation studies to investigate the contribution of incorporating diverse forms of evaluation tasks during instruction tuning. Table 6 shows the averaged Spearman correlation across all meta-evaluation datasets. In general, X-Eval trained on the combination of all forms of evaluation tasks, including scoring, comparison, ranking, achieves the highest averaged correlation for nearly all tasks.
Qualitative Correlation Analysis on Instruction Tuning
To further investigate the effect of instruction tuning, in Figure 4, we visualize the correlation of our X-Eval and Flan-T5 (i.e., “X-Eval w/o Training”) based on naturalness on Topical-Chat and consistency on SummEval. The red lines are linear regression fits to show how well the predicted scores correlate to human judgments linearly. We observe that before instruction tuning, most output scores of Flan-T5 are in the range of 0.5-0.7. Also, the predicted scores are more uniformly distributed regardless of ground truth scores, which results in poor correlation. On the contrary, our X-Eval can predict scores that not only achieve better correlation but also are more distinctive (either close to 1 or 0), demonstrating the effectiveness of our instruction tuning.
Error Propagation from Auxiliary Aspects during Inference
During inference, the evaluator may predict inaccurate evaluations for auxiliary aspects. To investigate their impact on the evaluation of target aspects, we tailor several baselines including (1) directly applying the model after two-stage tuning to perform inference without auxiliary aspects; (2) using the ground truth (“GT”) evaluation results instead of predicted results for auxiliary aspects, and; (3) letting the model have a random guess on the quality of auxiliary aspect and using random evaluation results for evaluating the target aspect. We find that removing auxiliary aspects makes the overall performance drop, which demonstrates its effectiveness. The variant with GT results (upperbound) gains improvement on all aspects by a large margin, indicating our framework can effectively leverage gold evaluation to improve the evaluation. Using random results, on the other hand, deteriorates the performance significantly.
Impact of Auxiliary Aspect Selection
We investigate the impact of different aspect selection strategies by examining the choice of in selecting the top- relevant aspects during inference. Table 8 shows that inference with the top- relevant aspect generally achieves superior correlation across all meta-evaluation tasks in comparison with selections like top- or top- aspects. We speculate that this may stem from the error propagation during inference on auxiliary aspects, where using more auxiliary aspects potentially introduces more inaccuracies, offsetting their potential performance benefits. In Figure 5, we also report the cosine similarity between the sentence embeddings of the aspect definitions used in turn-level dialogue evaluation as the qualitative analysis of our aspect selection strategy. In general, our strategy is able to select semantically related aspects for target-aspect evaluation. In addition, we conducted an experiment to compare the performance of selecting the auxiliary aspect based on seen, unseen, or all aspects, as well as randomly selecting the auxiliary aspect regardless of the definition. We set the number of auxiliary aspects to 1 in these variants. From Table 9, selecting the auxiliary aspect based on all the aspects achieves the best overall performance considering both Topical-Chat and FED (Turn-level) datasets. Also, we observe a substantial performance degradation when the auxiliary aspect is randomly selected, which demonstrates the effectiveness of our aspect selection strategy based on sentence embeddings.
Conclusion
In this work, we present X-Eval, a unified automatic evaluation framework capable of evaluating unseen customized aspects in NLG, guided by human instructions. To facilitate training, we collect AspectInstruct, the first multi-aspect evaluation instruction tuning dataset consisting of 27 diverse evaluation aspects and is augmented with diverse forms of NLG evaluation tasks. In addition, we propose a novel instruction-tuning algorithm that exploits the interrelationship between evaluation aspects both during training and inference. Extensive experiments on various multi-aspect evaluation benchmark datasets demonstrate strong zero-shot performance of X-Eval compared to the state-of-the-art NLG evaluators including GPT-4.
Limitations
In this work, we mainly target evaluation tasks in English. Future work can explore evaluation tasks in a more diverse language setting and augment our AspectInstruct dataset. In addition, our dataset focuses on a limited subset of NLG tasks including dialogue, summarization, and data2text. More NLG tasks can be considered in the future.
Inference Efficiency
Our algorithm requires multiple rounds of predictions in order to generate evaluation results from auxiliary aspects in the inference time. This hint generation process imposes additional computational cost and hence decreases the inference efficiency.
Error Propagation
During inference, the hints of related aspects are generated by the evaluator and may contain some errors. The errors in the hint can accumulate and affect the final prediction of the evaluator. In Table 8, we can observe that in most cases, the best performance is achieved when the hint of the top-1 related aspect is used and appending more hints may hurt the performance. Later work can develop more robust inference algorithms to address the error propagation problem.
References
Appendix A More Details on AspectInstruct
We present the annotated definitions in AspectInstruct in the following. We show the definitions of seen aspects on dialogue evaluation on Table 10, unseen aspects on dialogue evaluation on Table 11, and the aspects on summarization on Table 12.
A.2 Augmenting Instruction-tuning Tasks
We show the seen aspects, their corresponding source datasets where we collect the training data, constructed tasks, and the number of training instances for each task in Table 13 and Table 14.
A.3 Source Datasets for Meta Evaluation
is an evaluation benchmark for summarization which contains human ratings of 100 summaries along four evaluation dimensions: fluency, coherence, consistency, and relevance.
QAGS Wang et al. (2020b)
is a benchmark for identifying and evaluating hallucinations in the summarization task. It aims to measure the factual inconsistencies of generated summaries.
Topical-Chat Gopalakrishnan et al. (2019)
is a knowledge-grounded human-human conversation dataset. Following Zhong et al. (2022), we utilize human ratings collected by Mehri and Eskenazi (2020b) for Topical-Chat as the benchmark for evaluating dialog response generation. The assessment consider five aspects: naturalness, coherence, engagingness, groundedness, and understandability.
FED Mehri and Eskenazi (2020a)
is an evaluation benchmark for fine-grained dialog evaluation. It comprises human annotations evaluated across eighteen dialog aspects at both the turn-level and the dialog-level.
SFHOT & SFRES Wen et al. (2015)
are evaluation benchmarks for data-to-text task. They provide information about restaurants and hotels in San Francisco. The generated text is evaluated based on two aspects: informativeness and naturalness.