Plan-then-Generate: Controlled Data-to-Text Generation via Planning

Yixuan Su, David Vandyke, Sihui Wang, Yimai Fang, Nigel Collier

Introduction

Generating natural language from structured data Gatt and Krahmer (2018), i.e. data-to-text generation, is a research problem that is crucial to many downstream NLP applications. Some examples are dialogue systems Wen et al. (2016), restaurant assistant Novikova et al. (2017), and open domain question answering Chen et al. (2021).

To address this task, many researchers have designed sophisticated neural models based on various methods, such as soft-template Wiseman et al. (2018), copy mechanism Gehrmann et al. (2018), and pre-trained language models Kale and Rastogi (2020); Ribeiro et al. (2020). While achieving impressive results, most existing studies only focused on producing results that are close to the references. On the other hand, the controllability of such models is still under-explored, i.e. what to generate and in what order (the output structure) in their outputs cannot be explicitly controlled by the users. We argue that the model’s ability to control the structure of its output is highly

desirable for at least two reasons. (1) Arranging the structure of the output in a certain form enables it to have greater naturalness, as the structure of the sentence often reflects the salience of the entities it contains Poesio et al. (2004). Suppose we have a digital assistant which replies to user queries based on knowledge tables like Table 1. Then, for a user query “Who played Evelyn in Kids in Love?”, a natural response is “Evelyn in Kids in Love was played by Alma Jodorowsky.”. In contrast, to a different query “What role did Alma Jodorowsky play in Kids in Love?”, a natural response would be “Alma Jodorowsky played Evelyn in Kids in Love.”. While both answers are semantically equivalent, producing the answer with the most appropriate structure allows the system to sound less robotic and be easily understood. (2) It allows the model to generate outputs with diverse structures by simply changing the input planning information (i.e. a content plan), which could potentially benefit other applications such as paraphrasing and data augmentation. To control the output structure, we need an intermediate “planning”

signal (i.e. a content plan) which informs the model what to generate and in what order. To this end, we propose a Plan-then-Generate (PlanGen) framework which consists of two components: a content planner and a sequence generator. Given the input data, the content planner first predicts the most plausible content plan that the output should follow. Then, the sequence generator takes the data and the content plan as input to generate the result. To further ensure the controllability of our model, we propose a structure-aware reinforcement learning (RL) objective that encourages the generated output to adhere to the given content plan. In this work, we formulate the intermediate content plan as an ordered list of tokens for its simplicity and wide applicability to data with different structures. For tabular data, each token in the content plan is a slot key from the table. As for graphical data with RDF structure, each token represents the predicate from an RDF triple. In Figure 1, we provide examples for both cases.

To fully evaluate our approach, we test the proposed model on two benchmarks with different data structures: (i) ToTTo dataset Parikh et al. (2020) with tabular data, and (ii) WebNLG dataset Colin et al. (2016); Gardent et al. (2017) with graphical data. Compared with previous state-of-the-art approaches, our model achieves better performance in terms of generation quality as judged by both human and automatic evaluations. In particular, the results also show that the outputs of our model are highly controllable and contain diverse structures.

In summary, our contributions are: (1) A novel Plan-then-Generate (PlanGen) framework that consists of a content planner and a sequence generator for data-to-text generation. (2) Extensive automatic and human evaluations reporting state-of-the-art results on two benchmark datasets. (3) In-depth analysis revealing the merits of the proposed approach in terms of controllability and diversity.

Related Work

Data-to-text generation is a long-standing problem Reiter and Dale (1997) that aims at producing natural language descriptions of structured data. Traditional systems are primarily built on template-based algorithms Oh and Rudnicky (2000); Stent et al. (2004); Kondadadi et al. (2013). With recent advances in deep learning, researchers have shifted their attention to neural generation models that can be summarized into two categories.

Many existing studies are dedicated to building end-to-end neural models with different strategies like soft-templates Wiseman et al. (2018); Ye et al. (2020), attention awareness Liu et al. (2018); Colin and Gardent (2019), and retrieved prototypes Li et al. (2020); Su et al. (2021b). Gehrmann et al. (2018), Puduppully et al. (2019a, b), and Chen et al. (2020b) adopted copy mechanism for content selection to improve the information coverage of the outputs. With recent advance in pre-trained language models (PLMs) Devlin et al. (2019); Liu et al. (2019); Raffel et al. (2020); Lewis et al. (2020), several researchers Chen et al. (2020a, b); Kale and Rastogi (2020); Ribeiro et al. (2020) have studied the ways to adapt PLMs into the data-to-text generation task.

Pipeline Models.

Another line of research investigates ways to tackle the generation problem in a pipeline framework. Ma et al. (2019) proposed to first use a classifier to select the key contents. The planning and surface realisation of the selected contents are then addressed by a subsequent Seq2seq model. More related to our work, some researchers studied how neural models can benefit from traditional NLG steps Kukich (1983); McKeown (1992), that is, (i) content planning and (ii) surface realisation. To simultaneously select the key contents and arrange their orderings (i.e. content planning), different strategies are proposed such as the most probable traversal of graph trees Moryossef et al. (2019), the ordering of graph nodes Zhao et al. (2020), and the multi-step pipeline that includes discourse ordering, lexicalization, and regular expression generation Ferreira et al. (2019). While achieving satisfactory results, these approaches can only be applied to data with graphical structure. Compared with previous studies, we show that our content planning approach is more accurate and less dependent on the data structure. In addition, by providing the desired content plan, our model can control the output structure on both the intra-sentence and inter-sentence levels (section 7.3).

Preliminaries

In this study, our training dataset is defined as D={(T,C,S)i}i=1∣D∣\mathcal{D}=\{(T,C,S)_{i}\}_{i=1}^{|D|}. (1) TT is the linearized structured data and it is defined as T={t1,...,t∣T∣}T=\{t_{1},...,t_{|T|}\}. For data with tabular structure, each item ti={ki,vi}t_{i}=\{k_{i},v_{i}\} is a pair of slot key kik_{i} and slot value viv_{i} (e.g., (Date, 1956) in Figure 1(a)). As for graphical data with RDF structure, each item ti={si,pi,oi}t_{i}=\{s_{i},p_{i},o_{i}\} represents a RDF triple, where sis_{i}, pip_{i}, and oio_{i} are subject, predicate, and object, respectively. For instance, in Figure 1(b), (“Alan Bean”, “status”, “Retired”) is a RDF triple. (2) The reference content plan CC is defined as C={c1,..,c∣C∣}C=\{c_{1},..,c_{|C|}\}, where each token cic_{i} either denotes a slot key (for tabular data) or a predicate (for graphical data). The content plan is thus a selection of the content from the structured data that should appear in the output, in a particular order. (3) The S={s1,..,s∣S∣}S=\{s_{1},..,s_{|S|}\} denotes the reference text.

Content Plan Construction.

Note that the original ToTTo and WebNLG datasets only consist of pairs of structured data and reference text. Thus, we use a heuristic delexicalizer F\mathcal{F} to construct the reference content plan. For a tabular data TT, given the reference text SS, the content plan C=F(T,S)C=\mathcal{F}(T,S) is built by replacing the parts of the reference text that comes from the table slot values with the corresponding slot keys. For instance, suppose we have a text “Alma Jodorowsky played Evelyn in Kids in Love.” and Table 1, then the resulting content plan is {“Name”→\rightarrow“Role”→\rightarrow“Title”}. For graphical data with RDF structure, we apply a similar procedure to build the reference content plan by replacing the parts of the reference text that comes from the objects of the RDF triples with the corresponding predicates. In Figure 1, we show examples of reference content plan for both cases.

Methodology

Figure 2 depicts the proposed Plan-then-Generate (P2G) framework. Given the input data, the content planner (section 4.1) first predicts the most probable content plan. The sequence generator (section 4.2) then takes the structured data and the predicted content plan to generate the output. In the following, we elaborate the details of the proposed framework.

During training, the likelihood of the ordering sequence YY defined by the content plan is

Here, Φyi(hic)\Phi_{y_{i}}(h^{c}_{i}) is the label score of yiy_{i} at step ii, where label yiy_{i} indicates the position of the token in the final content plan. Taking Figure 2 as an example, the position of the “Name” key is 11, meaning that “Name” should appear in the front of the content plan. By predicting the positions instead of the actual slot keys, at test time, our model can handle tables with out-of-vocabulary slot keys that did not appear in the training set. In practice, Φ\Phi is parameterized by a feed-forward layer. The Myi−1,yiM_{y_{i-1},y_{i}} denotes the transition score from position yi−1y_{i-1} to position yiy_{i}, and MM is a learnable transition matrix.

2 Sequence Generator

Our sequence generator is built on a BART-base model Lewis et al. (2020) which consists of a transformer based encoder-decoder architecture.

Given the structured data TT, the reference content plan CC, and the reference text SS, the learning objective of the sequence generator is defined as

where EE, GG are the encoder and decoder, and [⋅:⋅][\cdot:\cdot] denotes the concatenation operation.

3 Structure-Aware RL Training

We note that the structure of the generated sequence can only be accurately measured on the sequence-level, which is not directly optimized by the token-level objective (Eq. (2)). Therefore, to encourage the generator to follow the sequence-level structure defined by the content plan, we incorporate reinforcement learning into our training process.

Formally, in training, given the structured data TT and the reference content plan CC, the generator first samples an output sequence S′=(S1′,...,S∣S′∣′)S^{\prime}=(S^{\prime}_{1},...,S^{\prime}_{|S^{\prime}|}), where St′S^{\prime}_{t} is the token sampled at time step tt. The generator parameters θ\theta are then updated using the REINFORCE algorithm Williams (1992) as

The reward function R(S,S′,T,C)R(S,S^{\prime},T,C) measures the structure of the sampled sequence S′S^{\prime} against the input content plan CC, and its surface form against the reference text SS as

where B(⋅,⋅)B(\cdot,\cdot) is the BLEU score Papineni et al. (2002). C′=F(T,S′)C^{\prime}=\mathcal{F}(T,S^{\prime}), and F\mathcal{F} is described in section 3. By optimizing Eq. (3), the structure of the output is encouraged to follow the content plan.

4 Learning

The learning objective of the content planner is LCRF=−log⁡PCRF\mathcal{L}_{\textup{CRF}}=-\log P_{\textup{CRF}} and PCRFP_{\textup{CRF}} is defined in Eq. (LABEL:score_function). For the sequence generator, at the first 10k steps, we train it with LLM\mathcal{L}_{\textup{LM}} as described in Eq. (2). Then, we incorporate the structure-aware RL objective (Eq. (3)) and further train the sequence generator with LLM+LRL\mathcal{L}_{\textup{LM}}+\mathcal{L}_{\textup{RL}} for 5k more steps.

Experiment Setup

Parikh et al. (2020) consists of Wikipedia tables paired with human-written descriptions. Each input is a full table with highlighted cells and the model is required to generate the text that describes the highlighted cells. Similar to previous studies Parikh et al. (2020); Kale and Rastogi (2020), we only use the highlighted cells as the model input. We report the automatic result of BLEU-4, PARENTPARENT is a word-overlap based metric that reflects the factual accuracy of the generated text in relation to both the input table and the reference sentence. Dhingra et al. (2019), and a learnt metric BLEURT Sellam et al. (2020). Note that ToTTo features a hidden test set with two splits: Overlap and Non-Overlap. The Non-Overlap set contains out-of-domain examples. To get the test set result, a submission must be made to the leaderboard.

WebNLG Dataset

is used in the WebNLG challenge Gardent et al. (2017). For each data instance, the input is a set of RDF triples from DBPedia and the output is their textual description. The test set of WebNLG features a Seen and Unseen subset. The Unseen subset contains out-of-domain instances. Following previous studies, we report the the automatic result of BLEU and METEOR Banerjee and Lavie (2005).

2 Implementation Details

Our implementation is based on the Huggingface Library Wolf et al. (2019). We optimize the model using Adam Kingma and Ba (2015) with a learning rate of 22e−5-5 and a batch size of 6464.

Results

In this section, we report the experimental results.

We compare our model with the latest models on ToTTo dataset, including NCP Puduppully et al. (2019a), Pointer-Generator See et al. (2017), BERT-to-BERT Rothe et al. (2020) and T5-3B Kale and Rastogi (2020). Similar to our model, the later two are also based on pre-trained language models.

Table 3 lists the results on ToTTo test set. For most of the metrics, our model with 140M parameters outperforms the current state-of-the-art T5-3B model which has over 2.8B parameters. The results on the PARENT metric suggest that our model can generate more factually accurate text. Moreover, in the Non-Overlap subset, our model achieves the best result on all metrics, showing its robustness to out-of-domain examples.

2 WebNLG Results

We compare our approach with two types of models on WebNLG dataset. The first type of models does not use pre-trained language models (PLMs), including GTR-LSTM Trisedya et al. (2018), Transformer Ferreira et al. (2019), Step-by-Step Moryossef et al. (2019), and PLANENC Zhao et al. (2020). Similar to ours, the latter three are pipeline models that utilize different methods to decide the output planning before generating the result. The second line of research utilizes PLMs, including Switch-GPT Chen et al. (2020b), T5 Kale and Rastogi (2020), and T5+Prefix Ribeiro et al. (2020). The Switch-GPT model applies a copy mechanism to copy content from the source to the output. We also include the top systems of the WebNLG challenge, including ADAPT, TILB-SMT, and MELBOURNE.

Table 3 lists the results of different methods in terms of text generation. We see that our approach outperforms all prior works. Compared with previous models that utilize PLMs, our performance improvements suggest that the incorporation of an explicit content plan can provide effective guiding signal for the model to achieve better generation results.

Evaluation on Content Planning.

Next, we compare our content planner with other pipeline models in terms of content planning performance. Following Zhao et al. (2020), we report the results on planning accuracy (P-A) and planning BLEU-2 score (B-2) against the human-generated plansThe human-generated plans are provided in the enriched WebNLG dataset Ferreira et al. (2018).. In addition, we examine two ablated variants of our content planner by either removing the CRF layer (w/o CRF) or using randomly initialized parameters instead of the pre-trained BERT (w/o PLMs). Table 4 lists the results. We see that our content planner outperforms all the baselines on both measures. Moreover, the results show that both the CRF layer and the pre-trained parameters positively contribute to the overall performance which further justifies our design of the content planner.

3 Human Evaluation

We also conduct a human evaluation to assess our model, using graders proficient in English from an internal grading platform. We randomly selected 200 samples from the ToTTo validation set. For each sample, we first use our sequence generator to produce the result with the content plan (CP) predicted by the content planner. Next, we randomly shuffle the predicted content plan and generate five different results (Shuffled CP). For comparison, we also include results of BERT-to-BERT and T5-3B using greedy decoding. All generated results, plus the reference sentence, are evaluated by three graders on a 3-point Likert scale (0, 1, or 2) for each of the following featuresMore evaluation details are provided in the Appendix A.:

Faithfulness: Whether the sentence is factually consistent with the input data.

Fluency: Whether the sentence is fluent and easy to understand.

Accuracy: How accurately the sentence follows the input content plansAs BERT-to-BERT and T5-3B do not take the content plan as input, thus we do not report their accuracy score..

Table 5 lists the results, with the first row showing strong inter-annotator agreements as measured by Fleiss\textprime\textprime kappa coefficient Fleiss et al. (1971). Comparing with BERT-to-BERT and T5-3B, our model achieves best results on both measures. Furthermore, on the faithfulness and fluency metrics, our model with both CP and Shuffled CP performs comparably with the reference sentence (Sign Test with p-value > 0.4). On the accuracy metric, our CP model also performs comparably with the reference as judged by the Sign Test. However, with randomly shuffled content plan, our model (Shuffled CP) fails to match the accuracy of the reference (p-value < 0.05). Our analysis is that the random content plans could contain patterns that are rare or unseen during training. In such cases, our model might fail to produce results that precisely follow the content plan, resulting in a lower accuracy score. Nonetheless, the human results suggest that, while being able to produce fluent and correct sentences, our model is also highly controllable. Finally, we note that on the accuracy metric, even the reference sentence does not score a perfect 2.02.0. This suggests that our simple heuristic delexicalizer F\mathcal{F} introduced in section 3 still lapses behind human performance. We leave to future work of designing better F\mathcal{F}.

Further Analysis

In this section, we present and discuss more empirical analyses of the proposed model.

We first evaluate the ability of different models in generating diverse results on the overall ToTTo validation set. We compare our model with two strong baselines, BERT-to-BERT (B2B) and T5-3B. Given the input data, the baseline models generate the results with different decoding strategiesFor each decoding strategy, five results are generated. , including greedy search, beam search (beam size of 1010), top-kk sampling (k=50k=50) Fan et al. (2018), and Nucleus sampling (p=0.9p=0.9) Holtzman et al. (2020). For our model, to generate diverse results, we simply vary the input content plan and use greedy decoding. We use two variants of the input content plan: (1) the content plan predicted by the content planner (Predict), or (2) the reference content plan (Oracle). For each variant, five results are generated by either using the input content plan (CP), or using five randomly shuffled forms of the content plan (Shuffled CP). The outputs are expected to vary in the latter case only.

Metric.

To measure the output quality, BLEU and PARENT scores are reported. To evaluate the generation diversity, we use Self-BLEU Zhu et al. (2018) and iBLEU Sun and Zhou (2012) metricsFor all evaluation metrics, we use the same hyper-parameters as in the original works that proposed the metric..

Results.

Table 6 lists the results in which our model ranks best on all metrics. On the quality metrics, we observe notable performance improvements from our model by using the reference content plan (Oracle), suggesting that the choice of content plan has a significant impact on the outputs. By shuffling the content plan, our model shows the largest decrease in BLEU and PARENT, showing that the variation of content plan encourages our model to produce diverse results that have different structures than the reference.

Furthermore, we see that, even with different decoding strategies, the baseline models still generate results that are very similar to the ones acquired from greedy search, with their BLEU and PARENT scores relatively unchanged. The results on the diversity metrics also verify the superiority of our model which outperforms the strong T5-3B model by over 5757 and 88 points on Self-BLEU and iBLEUBy definition, models using greedy search get 100 Self-BLEU as the generated results are always the same.. The performance gains suggest that the controllable property of our model is beneficial in producing high-quality as well as diverse results.

2 Ablation Study

In this part, we evaluate the importance of each component of our model on the overall ToTTo validation set. Specifically, we study the effect of content plan (CP) and the RL training by removing them iteratively. In addition to BLEU and PARENT, we measure the structure of the model output against the reference content plan with a S-BLEU metric. Given the data TT, the reference content plan CC, and the model output S′S^{\prime}, S-BLEU is defined as B(C,C′)B(C,C^{\prime}), where B(⋅,⋅)B(\cdot,\cdot) measures the BLEU score, C′=F(T,S′)C^{\prime}=\mathcal{F}(T,S^{\prime}), and F\mathcal{F} is the heuristic delexicalizer described in section 3. The results are listed in Table 7 with the first row showing the baseline results of BART model.

By comparing models with and without the content plan (model 1 vs. 3 and model 2 vs. ours), we observe that the content plan is an effective guiding signal that leads to better results. Moreover, we see that the Oracle results outperform the Predict results by a large margin, showing that the quality of the content plan is an important factor of the model performance and future research can focus more on this aspect.

Effect of RL.

By comparing the models trained with and without RL (model 1 vs. 2 and model 3 vs. ours), we see that training with our proposed RL objective consistently improves the model performance. The most notable improvement is observed in S-BLEU which means that the generated outputs better follow the input content plan. This is in line with our hypothesis that our reward function in Eq. (4) helps to improve the model’s adherence to the output structure defined by the content plan.

3 Case Study

To gain more insights into our model, we present generated examples from ToTTo and WebNLG datasetsMore examples are shown in the Appendix B. in Table 8 and Table 9, respectively.

In Table 8, we compare our model with predicted content plan against T5-3B. We see that T5-3B fails to produce the key game result (i.e. 13-0) in its outputs. In contrast, by following the content plan, our model is able to maintain all key information in its generated results.

Diversity and Controllability.

Next, we examine the output diversity and controllability. For the T5-3B model, when using beam search, only the position of the term “1956” varies, showing its reduced ability to generate diverse outputs. For our model, the variation of content plan leads to outputs with diverse structures. Furthermore, the results show that our model is not only able to control the intra-sentence output structure as shown in Table 8 but also to control the inter-sentence output structure as shown in Table 9.

Error Analysis.

We show one failure case in the bottom right cell of Table 8, in which it repeats the Title key twice in the output. Our analysis for such error is that the randomly shuffled content plan might contain patterns that are rarely seen in training. One possible solution is filtering out rare content plan patterns via statistical approaches such as bigram statistics.

Conclusion

In this study, we propose a new Plan-then-Generate (PlanGen) framework for data-to-text generation which can be easily applied to data with different structures. Extensive experiments and analyses are conducted on two benchmark datasets. Both automatic and human evaluation results demonstrate that our model is highly controllable. Furthermore, compared with previous studies, our model achieves better results both in terms of the generation quality as well as the output diversity. Our code, models and other related resources can be found in https://github.com/yxuansu/PlanGen/

Acknowledgments

The authors wish to thank Ehsan Shareghi, Zaiqiao Meng, Piji Li, and Benjamin Muller for their insightful discussions and support. Many thanks to our anonymous reviewers for their suggestions and comments.

References

Appendix A Details of Human Evaluation Setup

To perform human evaluation, we randomly sample 200 samples from the ToTTo validation set. For each sampled data, we use each baseline model (BERT-to-BERT and T5-3B) to produce one result. As for our model, we produce 6 different results (one with the predicted content plan, the other five with five randomly shuffled versions of the predicted content plan). Therefore, for each case, we have 9 different results (1 from BERT-to-BERT, 1 from T5-3B, 6 from our model, and 1 reference). To reduce human bias, we randomly shuffle these 1800 data points before presenting them to three annotators. Each annotator is asked to assess all these 1800 data points. Because BERT-to-BERT and T5-3B do not take the content plan as input, thus we only measure the accuracy score for the results generated by our model and the reference sentence. Note that the accuracy score of the reference sentence is measured against the reference content plan. In Figure 3, we show an example of the human evaluation interface.

Appendix B More Examples of Generated Result

In this part, we provide more generated examples of our model. The generated results on samples from WebNLG and ToTTo datasets are shown in Table 10 and 11, respectively. From the results, we can see that our model is able to generate fluent and diverse sentence while maintaining the structure defined by the desired content plan. In particular, our model is able to control the output structure both on the inter-sentence level (i.e. the structure across multiple sentences) as shown in Table 9 and on the intra-sentence level (i.e. the structure within a single sentence) as shown in Table 11. These results further demonstrate the applicability and generalization ability of our model.