Few-Shot Table-to-Text Generation with Prototype Memory
Yixuan Su, Zaiqiao Meng, Simon Baker, Nigel Collier
Introduction
Generating natural language from structured table Gatt and Krahmer (2018), i.e. table-to-text generation, is an important research problem for various NLP applications, such as biographical descriptions Lebret et al. (2016), restaurant information Novikova et al. (2017), basketball game summaries Wiseman et al. (2017), and open-domain question answering Chen et al. (2021).
The main challenge of table-to-text generation stems from the structural difference between the table and the natural language text. With recent advances in neural networks, many sophisticated neural models Liu et al. (2018); Gehrmann et al. (2018); Puduppully et al. (2019a, b); Su et al. (2021b) have been proposed to address this problem. While achieving impressive results, such neural models are data-hungry, i.e. large amounts of training data are required for them to learn the mapping between tables and texts. This can prohibit these models from being applied to real-world applications due to the huge data curation overhead Chen et al. (2020b). This motivates us to investigate few-shot table-to-text generation Ma et al. (2019); Chen et al. (2020b), that allows the model to learn a satisfactory table-to-text mapping with limited labelled training data. In this work, we propose to address this problem by augmenting data-to-text generation models with prototype memory acquired from a large unlabelled corpus. Our motivation is two-fold: (1) Relevant human-authored texts, termed “prototypes”, are informative and can teach the model how to better describe the table when limited training data is available. (2) However, traditional lexical-based IR systems, e.g. BM25, are inaccurate and the quality of their results are not guaranteed. Therefore, a BERT-based prototype selector is required to further select the prototypes, from the results retrieved by the IR system, that are closely related to the table for better guiding the neural generation model.
Figure 1 illustrates the proposed Prototype-to-Generate (P2G) framework. Given the table, an IR system is first applied to retrieve candidates that are potentially related to the table from a large unlabelled corpus. Based on the retrieved candidates, a prototype selector then selects the top prototypes based on the table-text pairwise similarity. Lastly, a sequence generator takes the table and the selected prototypes as input to produce the output. To prevent the model from uncritically copying the information contained in the prototypes that is irrelevant to the table, we introduce a content-aware learning objective when training the generator. In recent years, retrieval-based (i.e. template-based) text generation has been studied in different NLP areas, including machine translation Gu et al. (2017), unconditional text generation Guu et al. (2018), dialogue systems Wu et al. (2019); Su et al. (2021c), paraphrase generation Kazemnejad et al. (2020); Su et al. (2021a), and question answering Lewis et al. (2020b). Despite their differences, we identify two major limitations in previous studies compared to our approach. Firstly, most previous research Gu et al. (2017); Wu et al. (2019); Kazemnejad et al. (2020) build their retrieval corpus based on data consisting of aligned source-target pairs, which precludes the use of abundant unlabelled data. Secondly, current retrieval mechanisms are either based on lexical similarity (e.g. BM25) where its accuracy cannot be guaranteed, or large neural networks Karpukhin et al. (2020) which require a large amount of data to train. Notably, our framework is independent of the choice of generation model. For a comprehensive evaluation, we test our approach on three representative models, including the current state of the art. The experimental results on three datasets show that our framework leads to remarkable performance improvements across all evaluation metrics.
Methodology
Figure 1 depicts an overview of our framework. Given a linearized table , where is an attribute-value pair, an IR system first retrieves a set of candidates from the large unlabelled corpus. Then, a prototype selector (§2.1) selects the top prototypes from that are most related to . Lastly, a sequence generator (§2.2) takes and to produce the output .
As illustrated in Figure 1, given the table , the IR system relies on lexical features (e.g., word overlaps between the table and texts as colored in blue) to retrieve candidates . However, such lexical features are inaccurate and the semantic relevance between and cannot be guaranteed. To remedy this problem, we utilize a prototype selector to select the top prototypes from based on the table-text pairwise similarity. Formally, given the table and a text , their pairwise similarity score is defined as and is then defined as:
Figure 1 shows examples of the selected prototypes, . We see that are better related to the table and being closer to the reference text, i.e., the reference and could share similar contexts like the words in red. Thus, can be deemed as an guiding signal which teaches the model how to describe the table.
In this work, we use BERT Devlin et al. (2019) to build the prototype selector. The score is computed by a linear projection over the average embeddings of , where denotes concatenation operation. During training, given the table , the reference text , and the retrieved candidate set provided by the IR system, the learning objective of the prototype selector is defined as:
where and is the number of negatives sampled from . After training , we can obtain the prototype-augmented dataset for the learning of the generator.
2 Sequence Generator
Experiment
We conduct experiments on three benchmark few-shot table-to-text datasets Chen et al. (2020b) from different domains: Humans, Books, and Songs. Following previous studies Chen et al. (2020b); Gong et al. (2020), we train our model on different settings by varying the training size from , and evaluate our model using BLEU Papineni et al. (2002) and ROUGE Lin (2004) metrics. Test sets of Humans, Books, and Songs contain 13587, 5252 and 11879 instances. To build the IR system, we use Lucenehttps://lucene.apache.org/core/ to pre-index all sentences contained in the English Wikipedia (Dec. 2018 dump). For each table, the IR system retrieves 100 sentences as the candidates . The prototype selector then select the top 3 results from as the prototypes To avoid the data leakage problem, when building the dataset, we make sure the prototypes do not contain the reference.. When training the prototype selector, we set in Eq. (2) as 5. We compare our approach with both existing table-to-text methods that are not retrieval-based and also with the existing retrieval-based methods which we adapt for our concerned task. The existing table-to-text methods include Struct-Aware Liu et al. (2018), Pivot Ma et al. (2019), Switch-GPT Chen et al. (2020b), KGPT Chen et al. (2020a), Table-GPT Gong et al. (2020), and T5-Prefix Ribeiro et al. (2020). The latter four are based on pre-trained language models (PLMs). The retrieval-based approaches include Retri-Gen Wu et al. (2019) and RA-Gen Lewis et al. (2020b), where RA-Gen is based on PLMs. We select three representative models (Switch-GPT, Table-GPT, and T5-Prefix) to test the proposed framework.
2 Main Results
Table 1 lists the experiment results, where P2G+X indicates using model X under our framework. We can see that the proposed framework consistently and significantly improves the performance of all three models on all metrics, showing the robustness and universality of our approach. The notable performance gains suggest that the incorporation of retrieved prototypes greatly benefit the model’s ability in bridging the gap between tables and texts. It is worth noting that the RA-Gen model applies a strong BART Lewis et al. (2020a) as the generator. However, their retrieval module is purely based on a large neural models Karpukhin et al. (2020) that requires a large amount of data to train, and its accuracy degenerates when training data is limited, leading to the reduced generation performance.
3 Further Analysis
In this section, we present further discussions and empirical analysis of the proposed model.
First, we perform ablation analysis on the T5-Prefix model by progressively incorporating each proposed technique. The +Ret model directly utilizes the top retrieved results from the IR system as input. The +Ret&PS model utilizes the prototypes selected by the prototype selector as input. Finally, we include the proposed content-aware objective (+Ret&PS&CA) which results in the same model as P2G+T5-Prefix. The experiments are conducted on the Humans dataset with different training size. Table 2 lists the results which show that each component positively contributes to the overall performance. By comparing T5-Prefix with +Ret, we only observe a marginal improvement, suggesting that the retrieved results from the IR system are inaccurate (i.e., unrelated to the table) which brings little help to the generator. Next, from the results of +Ret&PS model we see that the incorporation of prototype selector significantly boosts the performance. This is inline with our hypothesis that the prototype selector can select more accurate (i.e., related to the table and similar to the reference) prototypes that can effectively teach the generator about how to describe the table. Lastly, the results of +Ret&PS&CA show that the proposed content-aware learning objective also benefits the model performance.
Effect of the Number of Prototypes.
Next, we examine how the number of prototypes ( in Eq. (1)) affects the model performance. To this end, we train P2G+T5-Prefix with 100 instances on the Humans dataset by varying the size of . Figure 2 depicts the results of BLEU and ROUGE. We observe that, when is small (i.e., ), the model performances are relatively the same. However, as approaching , the results drop notably. The reason is that, as increases, the top prototypes are likely to contain more information that is irrelevant to the table (i.e. noisy information), which leads to the degeneration of model performances.
4 Human Evaluation
We also conduct a human evaluation to assess the P2G+T5-Prefix model against several strong baselines, using graders proficient in English from an internal grading platform. Experiments are conducted on Humans dataset using 100 training instances and we randomly select 300 test cases for evaluation. All generated results, plus the reference, are evaluated by three graders on two aspects: (1) factual correctness; and (2) language fluency. Firstly, the graders are asked to count how many facts contained in the output are consistent with the table (#Support), and are contradicted to the table (#Contradict). Secondly, the graders are asked to assess the output in terms of language fluency on a 3-point Likert scale (0, 1, or 2).
Table 3 lists the evaluation results, with the first row showing strong inter-annotator agreements as measured by Fleiss kappa coefficient Fleiss et al. (1971). The results show that our model (P2G+T5-Prefix) significantly outperforms other baseline models on all metrics (Sign Test with p-value < 0.05). The performance gains of P2G+T5-Prefix over T5-Prefix further suggest that the prototypes help the model to produce not only more syntactically fluent but also more factually correct outputs.
Case Study
In Table 4, we present two generated examples from our model. For comparison, we also show the results generated by the strongest baseline (T5-Prefix) along with the reference sentence. As for our model, we show the selected prototypes along with the generated output. Both our model and the baseline model are trained with 100 instances.
As seen in the first case, the T5-Prefix fails to produce a correct output which describes the band. Instead, it just elaborates the name of the band members based on the table. In contrast, by relying on the prototypes that are related to the table, our model (P2G+T5-Prefix) produces an output that properly describes the band. Similarly, in the second case, the result of our model is more diverse and contains more facts that are supported by the table. These results further demonstrate that the prototypes can be deemed as effective guiding signals which teach the model how to describe the table. For better illustration, we highlight the parts, with red color, of prototypes on which the model relies when producing the output.
Conclusion
In this study, we introduced a new retrieval-based framework, Prototype-to-Generate (P2G), which augments table-to-text models with prototype memory from unlabelled data. Extensive experiments and analysis on three benchmark datasets show that our approach can significantly improve the performance of various strong generation models on all evaluation metrics. Our code, models and other related resources can be found in https://github.com/yxuansu/Few-Shot-Table-to-Text-Generation
Acknowledgments
The authors wish to thank our anonymous reviewers for their insightful suggestions and comments.
References
Appendix A Related Work
Table-to-text generation is a long-standing problem Reiter and Dale (1997) that aims at producing natural language descriptions of structured table. Traditional systems are primarily built on template-based algorithms Oh and Rudnicky (2000); Stent et al. (2004); Kondadadi et al. (2013). With recent advances in neural networks, researchers have built different neural models based on various strategies, e.g. latent-variables Wiseman et al. (2018); Ye et al. (2020), structure awareness Liu et al. (2018); Colin and Gardent (2019), copy mechanism Gehrmann et al. (2018); Puduppully et al. (2019a, b), and pre-trained language models (PLMs) Chen et al. (2020a); Kale (2020); Ribeiro et al. (2020). More recently, to alleviate the data-hungry nature of neural models, Ma et al. (2019) applied a pipeline model which first selects key facts from the table before producing the output. Chen et al. (2020a) designed a knowledge-grounded strategy for language model pre-training. Chen et al. (2020b) and Gong et al. (2020) adapted the pre-trained GPT-2 model with different architectural designs, e.g. switch policy Chen et al. (2020b) and content matching Gong et al. (2020), to address the few-shot table-to-text generation problem.
Retrieval-Based Text Generation.
In the last few years, retrieval-based text generation has attracted much attention. Gu et al. (2017) utilized a search engineer to assist the neural machine translation model. Guu et al. (2018) addressed unconditional text generation with a neural editor model that edits the retrieved prototypes. Wu et al. (2019) and Su et al. (2021c) incorporated retrieval frameworks into Seq2seq models to enrich the information contained in the dialogue responses. Kazemnejad et al. (2020) applied a retrieval model to assist the generation of paraphrased sentence. Lewis et al. (2020b) incorporated external knowledge using a retrieval model for knowledge-intensive question answering. To the best of our knowledge, our work is the first one which explores how retrieval-based approach could benefit neural models for table-to-text generation task.
Appendix B Human Evaluation Guidelines
In the human evaluation, the graders are asked to assess the results from two aspects. Following previous research Chen et al. (2020b); Gong et al. (2020), in the first study, the graders evaluate the factual correctness of the generated results by counting how many facts contained in the output are consistent with the table (#Support), and are contradicted to the table (#Contradict). In the second study, the graders assess the language fluency of the generated results following a 3-point Likert scale (0, 1, or 2). The definitions of different scores are provided as following:
: The result is grammatically fluent and is easy to understand.
: The result contains small errors but the errors does not affect your understanding.
: The result does not make sense and it is unreadable.