A Survey of Knowledge-Enhanced Text Generation
Wenhao Yu, Chenguang Zhu, Zaitang Li, Zhiting Hu, Qingyun Wang, Heng Ji, Meng Jiang
Introduction
Text generation, which is often formally referred as natural language generation (NLG), is one of the most important yet challenging tasks in natural language processing (NLP) (Garbacea and Mei, 2020). NLG aims at producing understandable text in human language from linguistic or non-linguistic data in a variety of forms such as textual data, numerical data, image data, structured knowledge bases, and knowledge graphs. Among these, text-to-text generation is one of the most important applications and thus often shortly referred as “text generation”. Researchers have developed numerous technologies for this task in a wide range of applications (Gatt and Krahmer, 2018; Wang et al., 2020a; Iqbal and Qureshi, 2020). Text generation takes text (e.g., a sequence, keywords) as input, processes the input text into semantic representations, and generates desired output text. For example, machine translation generates text in a different language based on the source text; summarization generates an abridged version of the source text to include salient information; question answering (QA) generates textual answers to given questions; dialogue system supports chatbots to communicate with humans with generated responses.
With the recent resurgence of deep learning technologies (LeCun et al., 2015), deep neural NLG models have achieved remarkable performance in enabling machines to understand and generate natural language. A basic definition of the text generation task is to generate an expected output sequence from a given input sequence, called sequence-to-sequence (Seq2Seq). The Seq2Seq task and model were first introduced in 2014 (Sutskever et al., 2014). It maps an input text to an output text under encoder-decoder schemes. The encoder maps the input sequence to a fixed-sized vector, and the decoder maps the vector to the target sequence. Since then, developing NLG systems has rapidly become a hot topic. Various text generation models have been proposed under deep neural encoder-decoder architectures. Popular architectures include recurrent neural network (RNN) encoder-decoder (Sutskever et al., 2014), convolutional neural network (CNN) encoder-decoder (Gehring et al., 2017), and Transformer encoder-decoder (Vaswani et al., 2017).
Nevertheless, the input text alone contains limited knowledge to support neural generation models to produce the desired output. Meanwhile, the aforementioned methods generally suffer from an inability to well comprehend language, employ memory to retain and recall knowledge, and reason over complex concepts and relational paths; as indicated by their name, they involve encoding an input sequence, providing limited reasoning by transforming their hidden state given the input, and then decoding to an output. Therefore, the performance of generation is still far from satisfaction in many real-world scenarios. For example, in dialogue systems, conditioning on only the input text, a text generation system often produces trivial or non-committal responses of frequent words or phrases in the corpus (Xing et al., 2017; Zhou et al., 2018b), such as “Me too.” or “Oh my god!” given the input text “My skin is so dry.” These mundane responses lack meaningful content, in contrast to human responses rich in knowledge. In comparison, humans are constantly acquiring, understanding, and storing knowledge from broader sources so that they can be employed to understand the current situation in communicating, reading, and writing. For example, in conversations, people often first select concepts from related topics (e.g., sports, food), then organize those topics into understandable content to respond; for summarization, people tend to write summaries containing keywords used in the input document and perform necessary modifications to ensure grammatical correctness and fluency; in question answering (QA), people use commonsense or professional knowledge pertained to the question to infer the answer. Therefore, it is often the case that knowledge beyond the input sequence is required to produce informative output text.
In general, knowledge is the familiarity, awareness, or understanding that coalesces around a particular subject. In NLG systems, knowledge is an awareness and understanding of the input text and its surrounding context. These knowledge sources can be categorized into internal knowledge and external knowledge (see Figure 1). Internal knowledge creation takes place within the input text(s), including but not limited to keyword, topic, linguistic features, and internal graph structure. External knowledge acquisition occurs when knowledge is provided from outside sources, including but not limited to knowledge base, external knowledge graph, and grounded text. These sources provide information (e.g., commonsense triples, topic words, reviews, background documents) that can be used as knowledge through various neural representation learning methods, and then applied to enhance the process of text generation. In addition, knowledge introduces interpretability for models with explicit semantics. This research direction of incorporating knowledge into text generation is named as knowledge-enhanced text generation.
Given a text generation problem where the system is given an input sequence , and aims to generate an output sequence . Assume we also have access to additional knowledge denoted as . Knowledge-enhanced text generation aims to incorporate the knowledge to enhance the generation of given , through leveraging the dependencies among the input text, knowledge, and output text.
Many existing knowledge-enhanced text generation systems have demonstrated promising performance on generating informative, logical, and coherent texts. In dialogue systems, a topic-aware Seq2Seq model helped understand the semantic meaning of an input sequence and generate a more informative response such as “Then hydrate and moisturize your skin.” to the aforementioned example input “My skin is so dry.” In summarization, knowledge graph produced a structured summary and highlight the proximity of relevant concepts, when complex events related with the same entity may span multiple sentences. A knowledge graph enhanced Seq2Seq model generated summaries that were able to correctly answer 10% more topically related questions (Huang et al., 2020). In question answering (QA) systems, facts stored in knowledge bases completed missing information in the question and elaborate details to facilitate answer generation (He et al., 2017; Fan et al., 2019a). In story generation, using commonsense knowledge acquired from knowledge graph facilitated understanding of the storyline and better narrate following plots step by step, so each step could be reflected as a link on the knowledge graph and the whole story would be a path (Guan et al., 2019).
2. Why a Survey of Knowledge-enhanced Text Generation?
Recent years have witnessed a surge of interests in developing methods for incorporating knowledge in NLG beyond input text. However, there is a lack of comprehensive survey of this research topic. Related surveys have laid the foundation of discussing this topic. For example, Garbacea et al. (Garbacea and Mei, 2020) and Gatt et al. (Gatt and Krahmer, 2018) reviewed model architectures for core NLG tasks but did not discuss knowledge-enhanced methods. Ji et al. (Ji et al., 2020c) presented a review on knowledge graph techniques which could be used for enhancing NLG. Wang et al. (Wang et al., 2020a) summarized how to represent structural knowledge such as knowledge base and knowledge graph for reading comprehension and retrieval.
To the best of our knowledge, this is the first survey that presents a comprehensive review of knowledge-enhanced text generation. It aims to provide NLG researchers a synthesis and pointer to related research. Our survey includes a detailed discussion about how NLG can benefit from recent progress in deep learning and artificial intelligence, including technologies such as graph neural network, reinforcement learning, and neural topic modeling.
3. What are the Challenges in Knowledge-enhanced Text Generation?
To start with, we note that the first challenge in knowledge-enhanced NLG is to obtain useful related knowledge from diverse sources. There has been a rising line of work that discovers knowledge from topic, keyword, knowledge base, knowledge graph and knowledge grounded text. The second challenge is how to effectively understand and leverage the acquired knowledge to facilitate text generation. Multiple methods have been explored to improve the encoder-decoder architecture (e.g., attention mechanism, copy and pointing mechanism).
Based on the first challenge, the main content of our survey is divided into two parts: (1) general methods of integrating knowledge into text generation (Section 2); (2) specific methods and applications according to different sources of knowledge enhancement (Sections 3–4). More concretely, since knowledge can be obtained from different sources, we first divide existing knowledge enhanced text generation work into two categories: internal knowledge enhanced and external knowledge enhanced text generation. The division of internal and external knowledge is widely adopted by management science (Menon and Pfeffer, 2003), which can be analogous with knowledge enhanced text generation. Based on the second challenge, we categorize recent knowledge-enhanced text generation methods evolved from how knowledge is extracted and incorporated into the process of text generation in each section (named as M1, M2, and etc). Furthermore, we review methods for a variety of natural language generation applications in each section to help practitioners choose, learn, and use the methods. In total, we discuss seven mainstream applications presented in more than 80 papers that were published or released in or after the year of 2016.
As shown in Figure 2, the remainder of this survey is organized as follows. Section 2 presents basic NLG models and general methods of integrating knowledge into text generation. Sections 3 reviews internal knowledge-enhanced NLG methods and applications. The internal knowledge is obtained from topic, keyword, linguistic features and internal graph structures. Sections 4 reviews external knowledge-enhanced NLG methods and applications. The external knowledge sources include knowledge bases, knowledge graphs, and grounded text. Section 5 presents knowledge-enhanced NLG benchmarks. Section 6 discusses future work and concludes the survey.
General Methods of Integrating Knowledge into NLG
Early encoder-decoder frameworks are often based on recurrent neural network (RNN) such as RNN-Seq2Seq (Sutskever et al., 2014). Convolutional neural network (CNN) based encoder-decoder (Gehring et al., 2017) and Transformer encoder-decoder (Vaswani et al., 2017) have been increasingly widely used. From a probabilistic perspective, the encoder-decoder frameworks learn the conditional distribution over a variable length sequence conditioned on yet another variable length sequence:
The encoder learns to encode a variable length sequence into a fixed length vector representation. RNN encoder reads the input sentence sequentially. CNN encoder performs convolutional operations on a word and its surrounding word(s) in a sequential window. Transformer encoder eschews recurrence and instead relying entirely on the self-attention mechanism to draw global dependencies between different tokens in the input . We denote them uniformly as:
where is the word embedding of word , is the contextualized hidden representation of .
The decoder is to decode a given fixed length vector representation into a variable length sequence (Sutskever et al., 2014). Specially, the decoder generates an output sequence one token at each time step. At each step the model is auto-regressive, consuming the previously generated tokens as additional input when generating the next token. Formally, the decoding function is represented as:
where is a nonlinear multi-layered function that outputs the probability of .
A generation process is regarded as a sequential multi-label classification problem. It can be directly optimized by the negative log likelihood (NLL) loss. Therefore, the objective of a text generation model via maximum likelihood estimation (MLE) is formulated as:
2. Knowledge-enhanced Model Architectures
The most popular idea of incorporating knowledge is designing specialized architectures of text generation models that can reflect the particular type of knowledge. In the context of neural networks, several general neural architectures are widely used and customized to bake the knowledge about the problems being tackled into the models.
It is useful to capture the weight of each time step in both encoder and decoder (Bahdanau et al., 2015). During the decoding phase, the context vector is added, so the hidden state is:
Unlike Eq.(3), here the probability is conditioned on the distinct context vector for target word , and depends on a sequence of hidden states that were mapped from input sequence.
In RNN-Seq2Seq decoder, the is computed as a weighted sum of :
where is parametrized as a multi-layer perception to compute a soft alignment. enables the gradient of loss function to be backpropagated. There are six alternatives for the function (see Table 2 in (Garbacea and Mei, 2020)). The probability reflects the importance of the hidden state of input sequence in presence of the previous hidden state for deciding the next hidden state.
In Transformer decoder, on top of the two sub-layers in the encoder, the decoder inserts a third sub-layer, which performs multi-head attention over the output of the encoder stack H. Efficient implementations of the transformer use the cached history matrix to generate next token. To compare with RNN-Seq2Seq, we summarize the Transformer decoder using recurrent notation:
where = , where corresponds to the key-value pairs from the -th layer generated at all time-steps from to . Instead of noting a specific name, we will use Encoder() and Decoder() to represent encoder and decoder in the following sections.
Attention mechanism has been widely used to incorporate knowledge representation in recent knowledge-enhanced NLG work. The general idea is to learn a knowledge-aware context vector (denoted as ) by combining both hidden context vector () and knowledge context vector (denoted as ) into decoder update, such as . The knowledge context vector () calculates attentions over knowledge representations (e.g., topic vectors, node vectors in knowledge graph). Table 1 summarizes a variety of knowledge attentions, including keyword attention (Li et al., 2018; Li et al., 2020, 2019), topic attention (Xing et al., 2017; Liu et al., 2019a; Wei et al., 2019; Zhang et al., 2016), knowledge base attention (He et al., 2017; Fu and Feng, 2018), knowledge graph attention (Zhang et al., 2020; Huang et al., 2020; Koncel-Kedziorski et al., 2019), and grounded text attention (Meng et al., 2020; Bi et al., 2020).
2.2. Copy and Pointing Mechanisms
CopyNet and Pointer-generator (PG) are used to choose subsequences in the input sequence and put them at proper places in the output sequence.
CopyNet and PG have a differentiable network architecture (Gu et al., 2016). They can be easily trained in an end-to-end manner. In CopyNet and PG, the probability of generating a target token is a combination of the probabilities of two modes, generate-mode and copy-mode. First, they represent unique tokens in the global vocabulary and the vocabulary of source sequence . They build an extended vocabulary . The difference between CopyNet and PG is the way to calculate distribution over the extended vocabulary. CopyNet calculates the distribution by
where and stand for the probability of generate-mode and copy-mode. Differently, PG explicitly calculates a switch probability between generate-mode and copy-mode. It recycles the attention distribution to serve as the copy distribution. The distribution over is calculated by
A knowledge-related mode chooses subsequences in the obtained knowledge and puts them at proper places in the output sequence. It helps NLG models to generate words that are not included in the global vocabulary () and input sequence (). For example, by adding the model of knowledge base, the extended vocabulary () adds entities and relations from the knowledge base, i.e., . The probability of generating a target token is a combination of the probabilities of three modes: generate-mode, copy-mode and knowledge base-mode. Therefore, knowledge-related mode is not only capable of regular generation of words but also operation of producing appropriate subsequences in knowledge sources. Table 1 summarizes different kinds of knowledge-related modes such as topic mode (Xing et al., 2017), keyword mode (Li et al., 2020), knowledge base mode (He et al., 2017), knowledge graph mode (Zhou et al., 2018b; Zhang et al., 2020), and background mode (Meng et al., 2020; Ren et al., 2020).
2.3. Memory Network
Memory networks (MemNNs) are recurrent attention models over a possibly large external memory (Sukhbaatar et al., 2015). They write external memories into several embedding matrices, and use query (generally speaking, the input sequence ) vectors to read memories repeatedly. This approach encodes long dialog history and memorize external information.
Given an input set to be stored in memory. The memories of MemNN are represented by a set of trainable embedding matrices , where each maps tokens to vectors, and a query (i.e., input sequence) vector is used as a reading head. The model loops over hops and it computes the attention weights at hop for each memory using:
where is the memory content in -th position, i.e., mapping into a memory vector. Here, is a soft memory selector that decides the memory relevance with respect to the query vector . Then, the model reads out the memory by the weighted sum over ,
Then, the query vector is updated for the next hop by using . The result from the encoding step is the memory vector and becomes the input for the decoding step.
Memory augmented encoder-decoder framework has achieved promising progress for many NLG tasks. For example, MemNNs are widely used for encoding dialogue history in task-oriented dialogue systems (Wu et al., 2019; Reddy et al., 2019). Such frameworks enable a decoder to retrieve information from a memory during generation. Recent work explored to model external knowledge with memory network such as knowledge base (Madotto et al., 2018; Yang et al., 2019) and topic (Fu and Feng, 2018; Zhou et al., 2018a).
2.4. Graph Network
Graph network captures the dependence of graphs via message passing between the nodes of graphs. Graph neural networks (GNNs) (Wu et al., 2020c) and graph-to-sequence (Graph2Seq) (Beck et al., 2018) potentiate to bridge up the gap between graph representation learning and text generation. Knowledge graph, dependency graph, and other graph structures can be integrated into text generation through various GNN algorithms. Here we denote a graph as , where is the set of entity nodes and is the set of (typed) edges. Modern GNNs typically follow a neighborhood aggregation approach, which iteratively updates the representation of a node by aggregating information from its neighboring nodes and edges. After iterations of aggregation, a node representation captures the structural information within its -hop neighborhood. Formally, the -th layer of a node is:
where denotes the set of edges containing node , and are feature vectors of a node and the edge between and at the -th iteration/layer. The choice of and in GNNs is crucial. A number of architectures for have been proposed in different GNN works such as GAT (Veličković et al., 2018). Meanwhile, the function used in labeled graphs (e.g., a knowledge graph) is often taken as those GNNs for modeling relational graphs (Schlichtkrull et al., 2018). To obtain the representation of graph (denoted as ), the function (either a simple permutation invariant function or sophisticated graph-level pooling function) pools node features from the final iteration ,
Graph network has been commonly used in integrating knowledge in graph structure such as knowledge graph and dependency graph. Graph attention network (Veličković et al., 2018) can be combined with sequence attention and jointly optimized (Zhou et al., 2018b; Zhang et al., 2020). We will introduce different graph structure knowledge in subsequent sections such as knowledge graph (Section 4.2), dependency graph (Section 3.3.2-3.3.3), and open knowledge graph (OpenKG) (Section 3.4).
2.5. Pre-trained Language Models
Pre-trained language models (PLMs) aims to learn universal language representation by conducting self-supervised training on large-scale unlabeled corpora. Recently, substantial PLMs such as BERT (Devlin et al., 2019) and T5 (Raffel et al., 2020) have achieved remarkable performance in various NLP downstream tasks. However, these PLMs suffer from two issues when performing on knowledge-intensive tasks. First, these models struggle to grasp structured world knowledge, such as concepts and relations, which are very important in language understanding. For example, BERT cannot deliver great performance on many commonsense reasoning and QA tasks, in which many of the concepts are directly linked on commonsense knowledge graphs (Yu et al., 2022b). Second, due to the domain discrepancy between pre-training and fine-tuning, these models do not perform well on domain-specific tasks. For example, BERT can not give full play to its value when dealing with electronic medical record analysis task in the medical field (Liu et al., 2020).
Recently, a lot of efforts have been made on investigating how to integrate knowledge into PLMs (Yu et al., 2022b; Liu et al., 2021a, 2020; Xiong et al., 2020; Guan et al., 2020; Zhou et al., 2021). Specifically, we will introduce some PLMs designed for NLG tasks. Overall, these approaches can be grouped into two categories: The first one is to explicitly inject entity representation into PLMs, where the representations is pre-computed from external sources (Zhang et al., 2019; Liu et al., 2021a). For example, KG-BART encoded the graph structure of KGs with knowledge embedding algorithms like TransE (Bordes et al., 2013), and then took the informative entity embeddings as auxiliary input (Liu et al., 2021a). However, the method of explicitly injecting entity representation into PLMs has been argued that the embedding vectors of words in text and entities in KG are obtained in separate ways, making their vector-space inconsistent (Liu et al., 2020). The second one is to implicitly modeling knowledge information into PLMs by performing knowledge-related tasks, such as concept order recovering (Zhou et al., 2021), entity category prediction (Yu et al., 2022b). For example, CALM proposed a novel contrastive objective for packing more commonsense knowledge into the parameters, and jointly pre-trained both generative and contrastive objectives for enhancing commonsense NLG tasks (Zhou et al., 2021).
3. Knowledge-enhanced Learning and Inference
Besides specialized model architectures, one common way of injecting knowledge to generation models is through the supervised knowledge learning. For example, one can encode knowledge into the objective function that guides the model training to acquire desired model behaviors (Dinan et al., 2019; Kim et al., 2020a). Such approaches enjoy the flexibility of integrating diverse types of knowledge by expressing them as certain forms of objectives. In general, knowledge-enhanced learning is agnostic to the model architecture, and can be combined with the aforementioned architectures.
One could devise learning tasks informed by the knowledge so that the model is trained to acquire the knowledge information.
The methods can be mainly divided into two categories as shown in Figure 3. The first category of knowledge-related tasks creates learning targets based on the knowledge, and the model is trained to recover the targets. These tasks can be combined as auxiliary tasks with the text generation task, resulting in a multi-task learning setting. For example, knowledge loss is defined as the cross entropy between the predicted and true knowledge sentences, and it is combined with the standard conversation generation loss to enhance grounded conversation (Dinan et al., 2019; Kim et al., 2020a). Similar tasks include keyword extraction loss (Li et al., 2020), template re-ranking loss (Cao et al., 2018; Wang et al., 2019c), link prediction loss on knowledge graph (Ji et al., 2020b), path reasoning loss (Liu et al., 2019b), mode loss (Zhou et al., 2018b; Wu et al., 2020b), bag-of-word (BOW) loss (Xu et al., 2020a; Lian et al., 2019), etc. The second category of methods directly derive the text generation targets from the knowledge, and use those (typically noisy) targets as supervisions in the standard text generation task. The approach is called weakly-supervised learning. Weakly-supervised learning enforces the relevancy between the knowledge and the target sequence. For example, in the problem of aspect based summarization, the work (Tan et al., 2020) automatically creates target summaries based on external knowledge bases, which are used to train the summarization model in a supervised manner.
The second way of devising knowledge-related tasks is to augment the text generation task by conditioning the generation on the knowledge. That is, the goal is to learn a function , where is the input sequence, is the target text and is the knowledge. Generally, the knowledge is first given externally (e.g., style, emotion) or retrieved from external resources (e.g., facts from knowledge base, a document from Wikipedia) or extracted from the given input text (e.g., keywords, topic words). Second, a conditional text generation model is used to incorporate knowledge and generate target output sequence. In practice, knowledge is often remedied by soft enforcing algorithms such as attention mechanism (Bahdanau et al., 2015) and copy/pointing mechanism (Gu et al., 2016; See et al., 2017). Regarding knowledge as condition is widely used in knowledge-enhanced text generation. For examples, work has been done in making personalized dialogue response by taking account of persona (Zhang et al., 2018) and emotion (Zhou et al., 2018a), controlling various aspects of the response such as politeness (Niu and Bansal, 2018), grounding the responses in external source of knowledge (Zhou et al., 2018b; Dinan et al., 2019; Ghazvininejad et al., 2018) and generating topic-coherent sequence (Tang et al., 2019; Xu et al., 2020a). Besides, using variational autoencoder (VAE) to enforce the generation process conditioned on knowledge is one popular approach to unsupervised NLG. By manipulating latent space for certain attributes, such as topic (Wang et al., 2019a) and style (Hu et al., 2017), the output sequence can be generated with desired attributes without supervising with parallel data.
3.2. Learning with knowledge constraints
Instead of creating training objectives in standalone tasks that encapsulate knowledge, another paradigm of knowledge-enhanced learning is to treat the knowledge as the constraints to regularize the text generation training objective.
where is the slack variable. The PR framework is also related to other constraint-driven learning methods (Chang et al., 2007; Mann and McCallum, 2007). We refer readers to (Ganchev et al., 2010) for more discussions.
3.3. Inference with knowledge constraints
Pre-trained language models leverage large amounts of unannotated data with a simple log-likelihood training objective. Controlling language generation by particular knowledge in a pre-trained model is difficult if we do not modify the model architecture to allow for external input knowledge or fine-tuning with specific data (Dathathri et al., 2020). Plug and play language model (PPLM) opened up a new way to control language generation with particular knowledge during inference. At every generation step during inference, the PPLM shifts the history matrix in the direction of the sum of two gradients: one toward higher log-likelihood of the attribute under the conditional attribute model and the other toward higher log-likelihood of the unmodified pre-trained generation model (e.g., GPT). Specifically, the attribute model makes gradient based updates to as follows:
where is the scaling coefficient for the normalization term; is update of history matrix (see Eq.(8)) and initialized as zero. The update step is repeated multiple times. Subsequently, a forward pass through the generation model is performed to obtain the updated as . The perturbed is then used to generate a new logit vector. PPLMs is efficient and flexible to combine differentiable attribute models to steer text generation (Qin et al., 2020).
NLG enhanced by Internal Knowledge
Topic, which can be considered as a representative or compressed form of text, has been often used to maintain the semantic coherence and guide the NLG process. Topic modeling is a powerful tool for finding the high-level content of a document collection in the form of latent topics (Blei et al., 2003). A classical topic model, Latent Dirichlet allocation (LDA), has been widely used for inferring a low dimensional representation that captures latent semantics of words and documents (Blei et al., 2003). In LDA, each topic is defined as a distribution over words and each document as a mixture distribution over topics. LDA generates words in the documents from topic distribution of document and word distribution of topic. Recent advances of neural techniques open a new way of learning low dimensional representations of words from the tasks of word prediction and context prediction, making neural topic models become a popular choice of finding latent topics from text (Cao et al., 2015; Guo et al., 2020).
Next, we introduce popular NLG applications enhanced by topics:
Dialogue system. A vanilla Seq2Seq often generates trivial or non-committal sentences of frequent words or phrases in the corpus (Xing et al., 2017). For example, a chatbot may say “I do not know”, “I see” too often. Though these off-topic responses are safe to reply to many queries, they are boring with very little information. Such responses may quickly lead the conversation to an end, severely hurting user experience. Thus, on-topic response generation is highly needed.
Machine translation. Though the input and output languages are different (e.g., translating English to Chinese), the contents are the same, and globally, under the same topic. Therefore, topic can serve as an auxiliary guidance to preserve the semantics information of input text in one language into the output text in the other language.
Paraphrase. Topic information helps understand the potential meaning and determine the semantic range to a certain extent. Naturally, paraphrases concern the same topic, which can serve as an auxiliary guidance to promote the preservation of source semantic.
As shown in Figure 4, we summarize topic-enhanced NLG methods into three methodologies: (M1) leverage topic words from generative topic models; (M2) jointly optimize generation model and CNN topic model; (M3) enhance NLG by neural topic models with variational inference.
Topics help understand the semantic meaning of sentences and determine the semantic spectrum to a certain extent. To enhanced text generation, an effective solution is to first discover topics using generative topic models (e.g., LDA), and then incorporate the topics representations into neural generation models, as illustrated in Figure 4(a). In existing work, there are two mainstream methods to represent topics obtained from generative topic models. The first way is to use the generated topic distributions for each word (i.e., word distributions over topics) in the input sequence (Zhang et al., 2016; Narayan et al., 2018). The second way is to assign a specific topic to the input sequence, then picks the top- words with the highest probabilities under the topic, and use word embeddings (e.g., GloVe) to represent topic words (Xing et al., 2017; Liu et al., 2019a). Explicitly making use of topic words can bring stronger guidance than topic distributions, but the guidance may deviate from the target output sequence when some generated topic words are irrelevant. Zhang et al. proposed the first work of using a topic-informed Seq2Seq model by concatenating the topic distributions with encoder and decoder hidden states (Zhang et al., 2016). Xing et al. designed a topic-aware Seq2Seq model in order to use topic words as prior knowledge to help dialogue generation (Xing et al., 2017).
1.2. M2: Jointly Optimize Generation Model and CNN Topic Model
The LDA models were separated from the training process of neural generation model and were not able to adapt to the diversity of dependencies between input and output sequences. Therefore, the idea of addressing this issue is to use neural topic models. Convolutional neural networks (CNN) were used to learn latent topic representations through iterative convolution and pooling operations. There are growing interests of using the CNNs to map latent topics implicitly into topic vectors that can be used to enhance text generation tasks (Wei et al., 2019; Gao and Ren, 2019). Empirical analyses showed that convolution-based topic extractors could outperform LDA-based topic models for multiple applications (e.g., dialogue system, text summarization, machine translation). However, theoretical analysis was missing to ensure the quality of the topics captured by the convolutions. And their interpretability is not as satisfactory as the LDA-based topic models.
1.3. M3: Enhance NLG by Neural Topic Models with Variational Inference
where is variational distributions for . The combined object function is given by:
1.4. Discussion and Analysis of Different Methods
For M1, topic models (e.g., LDA) has a strict probabilistic explanation since the semantic representations of both words and documents are combined into a unified framework. Besides, topic models can be easily used and integrated into generation frameworks. For example, topic words can be represented as word embeddings; topic embeddings can be integrated into the decoding phase through topic attention. However, LDA models are separated from the training process of generation, so they cannot adapt to the diversity of dependencies between input and output sequences.
For M2, it is an end-to-end neural framework that simultaneously learns latent topic representations and generates output sequences. Convolutional neural networks (CNN) are often used to generate the latent topics through iterative convolution and pooling operations. However, theoretical analysis is missing to ensure the quality of the topics captured by the convolutions. And their interpretability is not as good as the LDA-based topic models.
For M3, neural topic models combine the advantages of neural networks and probabilistic topic models. They enable back propagation for joint optimization, contributing to more coherent topics, and can be scaled to large data sets. Generally, neural topic models can provide better topic coherence than LDAs (Cao et al., 2015; Wang et al., 2019a; Xu et al., 2020a). However, neural variational approaches share a same drawback that topic distribution is assumed to be an isotropic Gaussian, which makes them incapable of modeling topic correlations. Existing neural topic models assume that the documents should be i.i.d. to adopt VAE, while they are commonly correlated. The correlations are critical for topic modeling.
2. NLG Enhanced by Keywords
Keyword (aka., key phrase, key term) is often referred as a sequence of one or more words, providing a compact representation of the content of a document. The mainstream methods of keyword acquisition for documents can be divided into two categories (Siddiqi and Sharan, 2015): keyword assignment and keyword extraction. Keyword assignment means that keywords are chosen from a controlled vocabulary of terms or predefined taxonomy. Keyword extraction selects the most representative words explicitly presented in the document, which is independent from any vocabulary. Keyword extraction techniques (e.g., TF-IDF, TextRank, PMI) have been widely used over decades. Many NLG tasks can benefit from incorporating such a condensed form of essential content in a document to maintain the semantic coherence and guide the generation process.
Next, we introduce popular NLG applications enhanced by keywords:
Dialogue system. Keywords help enlighten and drive the generated responses to be informative and avoid generating universally relevant replies which carry little semantics. Besides, recent work introduced personalized information into the generation of dialogue to help deliver better dialogue response such as emotion (Li and Sun, 2018; Zhou et al., 2018a; Song et al., 2019b), and persona (Zhang et al., 2018; Zheng et al., 2020).
Summarization. Vanilla Seq2Seq models often suffer when the generation process is hard to control and often misses salient information (Li et al., 2018). Making use of keywords as explicit guidance can provide significant clues of the main points about the document (Li et al., 2018; Li et al., 2020). It is closer to the way that humans write summaries: make sentences to contain the keywords, and then perform necessary modifications to ensure the fluency and grammatically correctness.
Question generation. It aims to generate questions from a given answer and its relevant context. Given an answer and its associated context, it is possible to raise multiple questions with different focuses on the context and various means of expression.
Researchers have developed a great line of keyword-enhanced NLG methods. These methods can be categorized into two methodologies: (M1) Incorporate keyword assignment into text generation; (M2) Incorporate keyword extraction into text generation.
When assigning a keyword to an input document, the set of possible keywords is bounded by a pre-defined vocabulary (Siddiqi and Sharan, 2015). The keyword assignment is typically implemented by a classifier that maps the input document to a word in the pre-defined vocabulary (Choudhary et al., 2017; Li and Sun, 2018; Zhou et al., 2018a; Song et al., 2019b). Unfortunately, some NLG scenarios do not hold an appropriate pre-defined vocabulary, so keyword assignment cannot be widely used to enhance NLG tasks. One applicable scenario is to use a pre-determined domain specific vocabulary to maintain relevance between the input and the output sequence (Choudhary et al., 2017). Another scenario is to generate dialogue with specific attributes such as persona (Song et al., 2019a; Xu et al., 2020a), emotion (Li and Sun, 2018; Zhou et al., 2018a; Song et al., 2019b).
A straightforward method of keyword assignment is to assign the words from pre-defined vocabulary and use them as the keywords (Song et al., 2019a; Xu et al., 2020a). Sometimes, the input sequence does not have an explicit keyword, but we can find one from the pre-defined vocabulary. For example, a dialogue utterance “If you had stopped him that day, things would have been different.” expresses sadness but it does not have the word “sad.” To address this issue, Li et al. propose a method to predict an emotion category by fitting the sum of hidden states from encoder into a classifier (Li and Sun, 2018). Then, the response will be generated with the guidance of the emotion category. In order to dynamically track how much the emotion is expressed in the generated sequence, Zhou et al. propose a memory module to capture the emotion dynamics during decoding (Zhou et al., 2018a). Each category is initialized with an emotion state vector before the decoding phase starts. At each step, the emotion state decays by a certain amount. Once the decoding process is completed, the emotion state decays to zero, indicating that the emotion is completely expressed.
As mentioned in (Song et al., 2019b), explicitly incorporating emotional keywords suffers from expressing a certain emotion overwhelmingly. Instead, Song et al. propose to increase the intensity of the emotional experiences not by using emotional words explicitly, but by implicitly combining neutral words in distinct ways on emotion (Song et al., 2019b). Specifically, they use an emotion classifier to build a sentence-level emotion discriminator, which helps to recognize the responses that express a certain emotion but not explicitly contain too many literal emotional words. The discriminator is connected to the end of the decoder.
2.2. M2: Incorporate Keyword Extraction into Text Generation
Keyword extraction selects salient words from input documents (Siddiqi and Sharan, 2015). Recent work has used statistical keyword extraction techniques (e.g., PMI (Li et al., 2019), TextRank (Li et al., 2018)), and neural-based keyword extraction techniques (e.g., BiLSTM (Li et al., 2020)). The process of incorporating extracted keywords into generation is much like the process discussed in Section 3.2.1. It takes keywords as an additional input into decoder. Recent work improves encoding phase by adding another sequence encoder to represent keywords (Li et al., 2018; Li et al., 2020). Then, the contextualized keywords representation is fed into the decoder together with input sequence representation. To advance the keyword extraction, Li et al. propose to use multi-task learning for training a keyword extractor network and generating summaries (Cho et al., 2019; Li et al., 2020). Because both summarization and keyword extraction aim to select important information from input document, these two tasks can benefit from sharing parameters to improve the capacity of capturing the gist of the input text. In practice, they take overlapping words between the input document and the ground-truth summary as keywords, and adopt a BiLSTM-Softmax as keyword extractor. Similar idea has also been used in question generation tasks (Cho et al., 2019; Wang et al., 2020d). They use overlapping words between the input answer context and the ground-truth question as keywords.
2.3. Discussion and Analysis of Different Methods
For M1, the primary advantage of keyword assignment is that the quality of keywords is guaranteed, because irrelevant keywords are not included in the pre-defined vocabulary. Another advantage is that even if two semantically similar documents do not have common words, they can still be assigned with the same keyword. However, there are mainly two drawbacks. On one hand, it is expensive to create and maintain dictionaries in new domains. So, the dictionaries might not be available. On the other hand, potential keywords occurring in the document would be unfortunately ignored if they were not in the vocabulary. Therefore, keyword assignment is suitable for the task that requires specific categories of keywords to guide the generated sentences with these key information. For example, dialogue systems generate responses with specific attitudes.
For M2, keyword extraction selects the most representative words explicitly presented in the document, which is independent from any vocabulary. It is easy to use but has two drawbacks. First, it cannot guarantee consistency because similar documents may still be represented by different keywords if they do not share the same set of words. Second, when an input document does not have a proper representative word, and unfortunately, the keyword extractor selects an irrelevant word from the document as a keyword, this wrong guidance will mislead the generation. Therefore, keyword extraction is suitable for the task that the output sequence needs to keep important information in the input sequence such as document summarization and paraphrase.
Table 3 summarizes tasks and datasets used in keyword-enhanced NLG work. Comparing with keyword-enhanced methods (E-SCBA (Li and Sun, 2018)) and the basic Seq2Seq attention model, keyword-enhanced methods can greatly improve both generation quality (evaluated by BLEU) and emotional expression (evaluated by emotion-w and emotion-s) on the NLPCC dataset. Besides, as shown in Table 3(a), EmoDS (Song et al., 2019b) achieved the best performance among three M1 methods, which indicates taking keyword assignment as a discriminant task can make better improvement than assigning keyword before the sentence decoding. For M2 methods, since most methods were evaluated on different tasks, we can only compare the performance between “without using keyword” and “using keyword”. As shown in Table 3(b), leveraging extracted keywords from input sequence into Seq2Seq model can improve the generation quality on summarization and question generation tasks. Comparing with KGAS (Li et al., 2020) and KIGN (Li et al., 2018), we can observe using BiLSTM-Softmax to extract keyword (a supervised manner by using overlapping words between and as labels) can make better performance than using TextRank (an unsupervised manner).
3. NLG Enhanced by Linguistic Features
Feature enriched encoder means that the encoder not only reads the input sequence, but also incorporates auxiliary hand-crafted features (Zhou et al., 2017; Sennrich and Haddow, 2016; Yu et al., 2021b). Linguistic features are the most common hand-crafted features, such as part-of-speech (POS) tags, dependency parsing, and semantic parsing.
Part-of-speech tagging (POS) assigns token tags to indicate the token’s grammatical categories and part of speech such as noun (N), verb (V), adjective (A). Named-entity recognition (NER) classifies named entities mentioned in unstructured text into pre-defined categories such as person (P), location (L), organization (O). CoreNLP is the most common used tool (Manning et al., 2014). In spite of homonymy and word formation processes, the same surface word form may be shared between several word types. Incorporating NER tags and POS tags can detect named entities and understand input sequence better, hence, further improve NLG (Zhou et al., 2017; Nallapati et al., 2016; Dong et al., 2021).
3.2. Syntactic dependency graph
Syntactic dependency graph is a directed acyclic graph representing syntactic relations between words (Bastings et al., 2017). For example, in the sentence “The monkey eats a banana”, “monkey” is the subject of the predicate “eats”, and “banana” is the object. Enhancing sequence representations by utilizing dependency information captures source long-distance dependency constraints and parent-child relation for different words (Chen et al., 2018; Bastings et al., 2017; Aharoni and Goldberg, 2017). In NLG tasks, dependency information is often modeled in three different ways as follows: (i) linearized representation: linearize dependency graph and then use sequence model to obtain syntax-aware representation (Aharoni and Goldberg, 2017); (ii) path-based representation: calculate attention weights based on the linear distance between a word and the aligned center position, i.e., the greater distance a word to the center position on the dependency graph is, the smaller contribution of the word to the context vector is (Chen et al., 2018); and (iii) graph-based representation: use GNNs to aggregate information from dependency relations (Bastings et al., 2017).
3.3. Semantic dependency graph
Semantic dependency graph represents predicate-argument relations between content words in a sentence and have various semantic representation schemes (e.g., DM) based on different annotation systems. Nodes in a semantic dependency graph are extracted by semantic role labeling (SRL) or dependency parsing, and connected by different intra-semantic and inter-semantic relations (Pan et al., 2020). Since semantic dependency graph introduces a higher level of information abstraction that captures commonalities between different realizations of the same underlying predicate-argument structures, it has been widely used to improve text generation (Pan et al., 2020; Jin et al., 2020; Liao et al., 2018). Jin et al. propose a semantic dependency guided summarization model (Jin et al., 2020). They incorporate the semantic dependency graph and the input text by stacking encoders to guide summary generation process. The stacked encoders consist of a sequence encoder and a graph encoder, in which the sentence encoder first reads the input text through stacked multi-head self-attention, and then the graph encoder captures semantic relationships and incorporates the semantic graph structure into the contextual-level representation.
4. NLG Enhanced by Open Knowledge Graphs
For those KGs (e.g., ConceptNet) constructed based on data beyond the input text, we refer them as external KGs. On the contrary, an internal KG is defined as a KG constructed solely based on the input text. In this section, we will discuss incorporating internal KG to help NLG (Huang et al., 2020; Fan et al., 2019a).
Internal KG plays an important role in understanding the input sequence especially when it is of great length. By constructing an internal KG intermediary, redundant information can be merged or discarded, producing a substantially compressed form to represent the input document (Fan et al., 2019a). Besides, representations on KGs can produce a structured summary and highlight the proximity of relevant concepts, when complex events related with the same entity may span multiple sentences (Huang et al., 2020). One of the mainstream methods of constructing an internal KG is using open information extraction (OpenIE). Unlike traditional information extraction (IE) methods, OpenIE is not limited to a small set of target entities and relations known in advance, but rather extracts all types of entities and relations found in input text (Niklaus et al., 2018). In this way, OpenIE facilitates the domain independent discovery of relations extracted from text and scales to large heterogeneous corpora.
After obtaining an internal KG, the next step is to learn the representation of the internal KG and integrate it into the generation model. For example, Zhu et al. use a graph attention network (GAT) to obtain the representation of each node, and fuse that into a transformer-based encoder-decoder architecture via attention (Zhu et al., 2021). Their method generates abstractive summaries with higher factual correctness. Huang et al. extend by first encoding each paragraph as a sub-KG using GAT, and then connecting all sub-KGs with a Bi-LSTM (Huang et al., 2020). This process models topic transitions and recurrences, which enables the identification of notable content, thus benefiting summarization.
NLG enhanced by External Knowledge
One of the biggest challenges in NLG is to discover the dependencies of elements within a sequence and/or across input and output sequences. The dependencies are actually various types of knowledge such as commonsense, factual events, and semantic relationship. Knowledge base (KB) is a popular technology that collects, stores, and manages large-scale information for knowledge-based systems like search engines. It has a great number of triples composed of subjects, predicates, and objects. People also call them “facts” or “factual triplets”. Recently, researchers have been designing methods to use KB as external knowledge for learning the dependencies easier, faster, and better.
Next, we introduce popular NLG applications enhanced by knowledge base:
Question answering. It is often difficult to generate proper answers only based on a given question. This is because, depending on what the question is looking for, a good answer may have different forms. It may completes the question precisely with the missing information. It may elaborate details of some part of the question. It may need reasoning and inference based on some facts and/or commonsense. So, only incorporating input question into neural generation models often fails the task due to the lack of commonsense/factual knowledge (Bi et al., 2019). Related structured information of commonsense and facts can be retrieved from KBs.
Dialogue system. The needs of KB in generating conversations or dialogues are relevant with QA but differ from two aspects. First, a conversation or dialogue could be open discussions when started by an open topic like “Do you have any recommendations?” Second, responding an utterance in a certain step needs to recall previous contexts to determine involved entities. KB will play an important role to recognize dependencies in the long-range contexts.
To handle different kinds of relationships between KB and input/output sequences, these methods can be categorized into two methodologies which is shown in Figure 5: (M1) design supervised tasks around KB for joint optimization; (M2) enhance incorporation by selecting KB or facts.
Knowledge bases (KBs) that acquire, store, and represent factual knowledge can be used to enhance text generation. However, designing effective incorporation to achieve a desired enhancement is challenging because a vanilla Seq2Seq often fails to represent discrete isolated concepts though they perform well to learn smooth shared patterns (e.g., language diversity). To fully utilize the knowledge bases, the idea is to jointly train neural models on multiple tasks. For example, the target task is answer sequence generation, and additional tasks include question understanding and fact retrieval in the KB. Knowledge can be shared across a unified encoder-decoder framework design. Typically, question understanding and fact retrieval are relevant and useful tasks, because a question could be parsed to match (e.g., string matching, entity linking, named entity recognition) its subject and predicate with the components of a fact triple in KB, and the answer is the object of the triple. KBCopy was the first work to generate responses using factual knowledge bases (Eric and Manning, 2017). During the generation, KBCopy is able to copy words from the KBs. However, the directly copying relevant words from KBs is extremely challenging. CoreQA used both copying and retrieving mechanisms to generate answer sequences with an end-to-end fashion (He et al., 2017). Specifically, it had a retrieval module to understand the question and find related facts from the KB. Then, the question and all retrieved facts are transformed into latent representations by two separate encoders. During the decoding phase, the integrated representations are fed into the decoder by performing a joint attention on both input sequence and retrieved facts. Figure 5(a) demonstrates a general pipeline that first retrieves relevant triples from KBs, then leverages the top-ranked triples into the generation process.
1.2. M2: Enhance Incorporation by Selecting KB or Facts in KB
Ideally, the relevance of the facts is satisfactory with the input and output sequence dependencies, however, it is not always true in real cases. Lian et al. addressed the issue of selecting relevant facts from KBs based on retrieval models (e.g. semantic similarity) might not effectively achieve appropriate knowledge selection (Lian et al., 2019). The reason is that different kinds of selected knowledge facts can be used to generate diverse responses for the same input utterance. Given a specific utterance and response pair, the posterior distribution over knowledge base from both the utterance and the response may provide extra guidance on knowledge selection. The challenge lies in the discrepancy between the prior and posterior distributions. Specifically, the model learns to select effective knowledge only based on the prior distribution, so it is hard to obtain the correct posterior distribution during inference.
To tackle this issue, the work of Lian et al. (Lian et al., 2019) and Wu et al. (Wu et al., 2020b) (shown in Figure 5(b)) approximated the posterior distribution using the prior distribution in order to select appropriate knowledge even without posterior information. They introduced an auxiliary loss, called Kullback-Leibler divergence loss (KLDivLoss), to measure the proximity between the prior distribution and the posterior distribution. The KLDivLoss is defined as follows:
where is the number of retrieved facts. When minimizing KLDivLoss, the posterior distribution can be regarded as labels to apply the prior distribution for approximating . Finally, the total loss is written as the sum of the KLDivLoss and NLL (generation) loss.
1.3. Discussion and Analysis of Different Methods
The relevance between triples in KBs and input sequences plays a central role in discovering knowledge for sequence generation. Methods in M1 typically follows the process that parses input sequence, retrieves relevant facts, and subsequently, a knowledge-aware output can be generated based on the input sequence and previously retrieved facts. Even though the improvement by modeling KB with memory network (Madotto et al., 2018), existing KG-enhanced methods still suffer from effectively selecting precise triples.
Methods of M2 improve the selection of facts, in which the ground-truth responses used as the posterior context knowledge to supervise the training of the prior fact probability distribution. Wu et al. used exact match and recall to measure whether the retrieved triples is used to generate the target outputs (Wu et al., 2020a). Table 4 shows the entity recall scores of M1-based methods and M2-based methods reported in (Wu et al., 2020a, b). We observe that compared to M1-based methods, M2-based methods can greatly improve the accuracy of triple retrieval, as well as the generation quality.
There are still remaining challenges in KB-enhanced methods. One is that retrieved facts may contain noisy information, making the generation unstable (Kim et al., 2020a). This problem is extremely harmful in NLG tasks, e.g., KB-based question answering and task-oriented dialogue system, since the information in KB is usually the expected entities in the response.
2. NLG Enhanced by Knowledge Graph
Knowledge graph (KG), as a type of structured human knowledge, has attracted great attention from both academia and industry. A KG is a structured representation of facts (a.k.a. knowledge triplets) consisting of entitiesFor brevity, we use “entities” to denote both entities (e.g., prince) and concepts (e.g., musician) throughout the paper., relations, and semantic descriptions (Ji et al., 2020c). The terms of “knowledge base” and “knowledge graph” can be interchangeably used, but they do not have to be synonymous. The knowledge graph is organized as a graph, so the connections between entities are first-class citizens in it. In the KG, people can easily traverse links to discover how entities are interconnected to express certain knowledge. Recent advances in artificial intelligence research have demonstrated the effectiveness of using KGs in various applications like recommendation systems (Wang et al., 2019d).
Next, we introduce popular NLG applications that have been enhanced by knowledge graph:
Commonsense reasoning. It aims to empower machines to capture the human commonsense from KG during generation. The methods exploit both structural and semantic information of the commonsense KG and perform reasoning over multi-hop relational paths, in order to augment the limited information with chains of evidence for commonsense reasoning. Popular tasks in commonsense reasoning generation include abductive reasoning (e.g., the NLG task) (Bhagavatula et al., 2020; Ji et al., 2020b), counterfactual reasoning (Ji et al., 2020b, a), and entity description generation (Cheng et al., 2020).
Dialogue system. It frequently makes use of KG for the semantics in linked entities and relations (Zhou et al., 2018b; Niu et al., 2019; Tuan et al., 2019; Zhang et al., 2020). A dialogue may shift focus from one entity to another, breaking one discourse into several segments, which can be represented as a linked path connecting the entities and their relations.
Creative writing. This task can be found in both scientific and story-telling domains. Scientific writing aims to explain natural processes and phenomena step by step, so each step can be reflected as a link on KG and the whole explanation is a path (Koncel-Kedziorski et al., 2019; Wang et al., 2019b). In story generation, the implicit knowledge in KG can facilitate the understanding of storyline and better predict what will happen in the next plot (Guan et al., 2019, 2020; Liu et al., 2021a).
Compared with separate, independent knowledge triplets, knowledge graph provides comprehensive and rich entity features and relations for models to overcome the influence of the data distribution and enhance its robustness. Therefore, node embedding and relational path have played important roles in various text generation tasks. The corresponding techniques are knowledge graph embedding (KGE) (Wang et al., 2017) and path-based knowledge graph reasoning (Chen et al., 2020a). Furthermore, it has been possible to encode multi-hop and high-order relations in KGs using the emerging graph neural network (GNN) (Wu et al., 2020c) and graph-to-sequence (Graph2Seq) frameworks (Beck et al., 2018).
A knowledge graph (KG) is a directed and multi-relational graph composed of entities and relations which are regarded as nodes and different types of edges. Formally, a KG is defined as , where is the set of entity nodes and is the set of typed edges between nodes in with a certain relation in the relation schema .
Then given the input/output sequences in the text generation task, a subgraph of the KG which is associated with the sequences can be defined as below.
A sequence-associated K-hop subgraph is defined as , where is the union of the set of entity nodes mapped through an entity linking function and their neighbors within K-hops. Similarly, is the set of typed edges between nodes in .
Sequence-associated subgraph provides a graphical form of the task data (i.e., sequences) and thus enables the integration of KGs and the sequences into graph algorithms.
Many methods have been proposed to learn the relationship between KG semantics and input/output sequences. They can be categorized into four methodologies as shown in Figure 6: (M1) incorporate knowledge graph embeddings into language generation; (M2) transfer knowledge into language model with triplet information; (M3) perform reasoning over knowledge graph via path finding strategies; and (M4) improve the graph embeddings with graph neural networks.
2.2. M2: Transfer Knowledge into Language Model with Knowledge Triplet Information
The vector spaces of entity embeddings (from KGE) and word embeddings (from pre-trained language models) are usually inconsistent (Liu et al., 2021a). Beyond a simple concatenation, recent methods have explored to fine-tune the language models directly on knowledge graph triplets. Guan et al. transformed the commonsense triplets (in ConceptNet and ATOMIC) into readable sentences using templates, as illustrated in Figure 6(b). And then the language model (e.g., GPT-2) is fine-tuned on the transformed sentences to learn the commonsense knowledge to improve text generation.
2.3. M3: Perform Reasoning over Knowledge Graph via Path Finding Strategies
KGE learns node representations from one-hop relations through a certain semantic relatedness (e.g. TransE). However, Xiong et al. argued that an intelligent machine is supposed to be able to conduct explicit reasoning over relational paths to make multiple inter-related decisions rather than merely embedding entities in the KGs (Xiong et al., 2017). Take the QA task an example. The machine performs reasoning over KGs to handle complex queries that do not have an obvious answer, infer potential answer-related entities, and generate the corresponding answer. So, the challenge lies in identifying a subset of desired entities and mentioning them properly in a response (Moon et al., 2019). Because the connected entities usually follow natural conceptual threads, they help generate reasonable and logical answers to keep conversations engaging and meaningful. As shown in Figure 6(c), path-based methods explore various patterns of connections among entity nodes such as meta-paths and meta-graphs. Then they learn from walkable paths on KGs to provide auxiliary guidance for the generation process. The path finding based methods can be mainly divided into two categories: (1) path ranking based methods and (2) reinforcement learning (RL) based path finding methods.
Path ranking algorithm (PRA) emerges as a promising method for learning and inferring paths on large KGs (Lao et al., 2011). PRA uses random walks to perform multiple bounded depth-first search processes to find relational paths. Coupled with elastic-net based learning (Zou and Hastie, 2005), PRA picks plausible paths and prunes non-ideal, albeit factually correct KG paths. For example, Tuan et al. proposed a neural conversation model with PRA on dynamic knowledge graphs (Tuan et al., 2019). In the decoding phase, it selected an output from two networks, a general GRU decoder network and a PRA based multi-hop reasoning network, at each time step. Bauer et al. ranked and filtered paths to ensure both the information quality and variety via a 3-step scoring strategy: initial node scoring, cumulative node scoring, and path selection (Bauer et al., 2018). Ji et al. heuristically pruned the noisy edges between entity nodes and proposed a path routing algorithm to propagate the edge probability along multi-hop paths to the entity nodes (Ji et al., 2020a).
Reinforcement learning (RL) based methods make an agent to perform reasoning to find a path in a continuous space. These methods incorporate various criteria in their reward functions of path finding, making the path finding process flexible. Xiong et al. proposed DeepPath, the first work that employed Markov decision process (MDP) and used RL based approaches to find paths in KGs (Xiong et al., 2017). Leveraging RL based path finding for NLG tasks typically consists of two stages (Niu et al., 2019; Liu et al., 2019b). First, they take a sequence as input, retrieve a starting node on , then perform multi-hop graph reasoning, and finally arrive at a target node that incorporates the knowledge for output sequence generation. Second, they represent the sequence and selected path through two separate encoders. They decode a sequence with multi-source attentions on the input sequence and selected path. Path-based knowledge graph reasoning converts the graph structure of a KG into a linear path structure that can be easily represented by sequence encoders (e.g, RNN) (Niu et al., 2019; Tuan et al., 2019; Fan et al., 2019a). For example, Niu et al. encoded selected path and input sequence with two separate RNNs and generated sequence with a general attention-based RNN decoder (Niu et al., 2019). To enhance the RL process, Xu et al. proposed six reward functions for training an agent in the reinforcement learning process. For example, the functions looked for accurate arrival at the target node as well as the shortest path between the start and target node, i.e., minimize the length of the selected path (Xu et al., 2020b).
2.4. M4: Improve the Graph Embeddings with Graph Neural Networks
The contexts surrounding relevant entities on KGs play an important role in understanding the entities and generating proper text about their interactions (Guan et al., 2019; Koncel-Kedziorski et al., 2019). For example, in scientific writing, it is important to consider the neighboring nodes of relevant concepts on a taxonomy and/or the global context of a scientific knowledge graph (Koncel-Kedziorski et al., 2019). However, neither KGE nor relational path could fully represent such information. Graph-based representations aim at aggregating the context/neighboring information on graph data; and recent advances of GNN models demonstrate a promising advancement in graph-based representation learning (Wu et al., 2020c). In order to improve text generation, graph-to-sequence (Graph2Seq) models encode the structural information of the KG in a neural encoder-decoder architecture (Beck et al., 2018). Since then, GNNs have been playing an important role in improving the NLG models. They have been applied to both encoding and decoding phases.
For encoding phase, a general process of leveraging GNNs for incorporating KG is to augment semantics of a word in the input text by combining with the vector of the corresponding entity node vector to the word on the KG (Zhou et al., 2018b; Guan et al., 2019; Zhang et al., 2020; Huang et al., 2020; Zeng et al., 2021). A pre-defined entity linking function maps words in the input sequence to entity nodes on the KG. Given an input sequence, all the linked entities and their neighbors within -hops compose a sequence-associated K-hop subgraph (formally defined in Definition 4.2). For each entity node in , it uses the KG structure as well as entity and edge features (e.g., semantic description if available) to learn a representation vector u. Specifically, a GNN model follows a neighborhood aggregation approach that iteratively updates the representation of a node by aggregating information from its neighboring nodes and edges. After iterations of aggregation, the node representation captures the structural information within its -hop neighborhood. Formally, the -th layer of a node is:
The sub-graph representation is learned thorough a function from all entity node representations (i.e., \textbf{h}_{subG}=\textsc{Readout}(\big{\{}\textbf{u}^{(k)},u\in\mathcal{U}_{sub}\big{\}}). Zhou et al. was the first to design such a knowledge graph interpreter to enrich the context representations with neighbouring concepts on ConceptNet using graph attention network (GAT) (Zhou et al., 2018b).
The sequence decoder uses attention mechanism to find useful semantics from the representation of KG as well as the hidden state of the input text, where the KG’s representation is usually generated by GNNs. Specially, the hidden state is augmented by subgraph representation , i.e., (Beck et al., 2018). Then, the decoder attentively reads the retrieved subgraph to obtain a graph-aware context vector. Then it uses the vector to update the decoding state (Zhou et al., 2018b; Guan et al., 2019; Zhang et al., 2020; Ji et al., 2020b; Liu et al., 2021a). It adaptively chooses a generic word or an entity from the retrieved subgraph to generate output words. Because graph-level attention alone might overlook fine-grained knowledge edge information, some recent methods adopted the hierarchical graph attention mechanism (Zhou et al., 2018b; Guan et al., 2019; Liu et al., 2021a). It attentively read the retrieved subgraph and then attentively read all knowledge edges involved in . Ji et al. added a relevance score that reflected the relevancy of the knowledge edge according to the decoding state (Ji et al., 2020b).
2.5. Discussion and Analysis of the Methodologies and Methods
Knowledge graph embedding (M1) was the earliest attempt to embed components of a KG including entities and relations into continuous vector spaces and use them to improve text generation. Those entity and relation embeddings can simply be used to enrich input text representations (e.g., concatenating embeddings), bridging connections between entity words linked from input text in latent space. Because the graph projection and text generation are performed as two separate steps, the embedding vectors from knowledge graph and the hidden states from input text were in two different vector spaces. The model would have to learn to bridge the gap, which might make a negative impact on the performance of text generation.
Fine tuning pre-trained language models on the KG triplets (M2) can eliminate the gap between the two vector spaces. Nevertheless, M1 and M2 share two drawbacks. First, they only preserve information of direct (one-hop) relations in a KG, such as pair-wise proximity in M1 and KG triplet in M2, but ignore the indirect (multi-hop) relations of concepts. The indirect relations may provide plausible evidence of complex reasoning for some text generation tasks. Second, from the time KGs were encoded in M1 or M2 methods, the generation models would no longer be able to access the KGs but their continuous representations. Then the models could not support reasoning like commonsense KG reasoning for downstream tasks. Due to these two reasons, M1 and M2 were often used to create basic KG representations upon which the KG path reasoning (M3) and GNNs (M4) could further enrich the hidden states (Zhou et al., 2018b; Zhang et al., 2020).
The path finding methods of KG reasoning (M3) perform multi-hop walks on the KGs beyond one-hop relations. It enables reasoning that is needed in many text generation scenarios such as commonsense reasoning and conversational question answering. At the same time, it provides better interpretability for the entire generation process, because the path selected by the KG reasoning algorithm will be explicitly used for generation. However, the selected paths might not be able to capture the full contexts of the reasoning process due to the limit of number. Besides, reinforcement-learning based path finding uses heuristic rewards to drive the policy search, making the model sensitive to noises and adversarial examples.
The algorithms of GNN and Graph2Seq (M4) can effectively aggregate semantic and structural information from multi-hop neighborhoods on KGs, compared to M3 that considers multi-hop paths. Therefore, the wide range of relevant information can be directly embedded into the encoder/decoder hidden states. Meanwhile, M4 enables back propagation for jointly optimizing text encoder and graph encoder. Furthermore, the attention mechanism that has been applied in GNN and Graph2Seq (e.g., graph attention) can explain the model’s output at some extent, though the multi-hop paths from M3 has better interpretability.
M3 and M4 are able to use multi-hop relational information, compared to M1 and M2. However, they have two weak points. First, they have higher complexity than M1 and M2. In M3, the action space of path finding algorithms can be very large due to the large size and sparsity of the knowledge graph. In M4, the decoder has to attentively read both input sequence and knowledge graph. Second, the subgraphs retrieved by M3 and M4 might provide low coverage of useful concepts for generating the output. For example, people use ConceptNet, a widely used commonsense KG, to retrieve the subgraph on three generative commonsense reasoning tasks. The task datasets are ComVE (Ji et al., 2020b), -NLG (Bhagavatula et al., 2020), and ROCSories (Guan et al., 2019). We found 25.1% / 24.2% / 21.1% of concepts in the output could be found on ConceptNet, but only 11.4% / 8.1% / 5.7% of concepts in the output can be found on the retrieved 2-hop sequence-associated subgraph, respectively. It means that a large portion of relevant concepts on the KG are not utilized in the generation process.
Table 5 summarizes tasks, datasets, and KG sources used in existing KG-enhanced works. Three important things should be mentioned. First, all the datasets in the table are public, and we include their links in Table 12. CommonGen (Lin et al., 2020), ComVE (Wang et al., 2020b) and -NLG (Bhagavatula et al., 2020) have a public leaderboard for competition. Second, for KG sources, we observe that eight (57.1%) papers use ConceptNet as external resource, while six (42.9%) papers constructed their own KGs from domain-specific corpus. For example, Koncel et al. created a scientific knowledge graph by applying the SciIE tool (science domain information extraction) (Koncel-Kedziorski et al., 2019). Besides, Zhao et al. compared the performance of models between using ConceptNet and using a self-built KG, and found the model with self-built KG could work better on story generation and review generation tasks (Zhao et al., 2020). Third, we observed that KG-enhanced NLG methods made the largest improvement on generative commonsense reasoning tasks, in which the average improvement is +2.55% in terms of BLEU, while the average improvement on all different tasks is +1.32%.
Table 6 compares different KG-enhanced methods from three dimensions: multi-hop information aggregation, multi-hop path reasoning, and auxiliary knowledge graph related tasks. M3 is commonly used for multi-hop path reasoning and M4 is used for multi-hop information aggregation, except that CCM (Zhou et al., 2018b) only aggregates one-hop neighbors. Besides, the auxiliary KG-related tasks are often used to further help the model learn knowledge from the KG. For example, ablation studies in (Liu et al., 2019b; Ji et al., 2020a, b) show that the tasks of path selection, concept selection and link prediction can further boost the generation performance. GRF (Ji et al., 2020b) learns these three abilities at the same time. It achieves the state-of-art performance on three generation tasks.
3. NLG enhanced by Grounded Text
Knowledge grounded text refers to textual information that can provide additional knowledge relevant to the input sequence. The textual information may not be found in training corpora or structured databases, but can be obtained from massive textual data from online resources. These online resources include encyclopedia (e.g., Wikipedia), social media (e.g., Twitter), shopping websites (e.g., Amazon reviews). Knowledge grounded text plays an important role in understanding the input sequence and its surrounding contexts. For example, Wikipedia articles may offer textual explanations or background information for the input text. Amazon reviews may contain necessary descriptions and reviews needed to answer a product-related question. Tweets may contain people’s comments and summaries towards an event. Therefore, knowledge grounded text is often taken as an important external knowledge source to help with a variety of NLG applications.
Next, we introduce popular NLG applications enhanced by knowledge grounded text:
Dialogue system. Building a fully data-driven dialogue system is difficult since most of the universal knowledge is not presented in the training corpora (Ghazvininejad et al., 2018). The lack of universal knowledge considerably limits the appeal of fully data-driven generation methods, as they are bounded to respond evasively or defectively and seldom include meaningfully factual contents. To infuse the response with factual information, an intelligent machine is expected to obtain necessary background information to produce appropriate response.
Summarization. Seq2Seq models that purely depend on the input text tend to “lose control” sometimes. For example, 3% of summaries contain less than three words, and 4% of summaries repeat a word for more than 99 times as mentioned in (Cao et al., 2018). Furthermore, Seq2Seq models usually focus on copying source words in their exact order, which is often sub-optimal in abstractive summarization. Therefore, leveraging summaries of documents similar as the input document as templates can provide reference for the summarization process (Cao et al., 2018; Wang et al., 2019c).
Question answering (QA). It is often difficult to generate proper answers only based on the given question. For example, without knowing any information of an Amazon product, it is hard to deliver satisfactory answer to the user questions such as “Does the laptop have a long battery life?” or “Is this refrigerator frost-free?” So, the product description and customer reviews can be used as a reference for answering product-related questions (Chen et al., 2019; Bi et al., 2020).
To handle different kinds of relationships between grounded text and input/output sequences, these methods can be categorized into two methodologies as shown in Figure 7: (M1) guiding generation with retrieved information; (M2) modeling background knowledge into response generation.
Because knowledge grounded text is not presented in the training corpora, an idea is to retrieve relevant textual information (e.g., a review, a relevant document, a summary template) from external sources based on the input text and to incorporate the retrieved grounded text into the generation process. This process is similar to designing knowledge acquisition and incorporation of KBs and KGs in text generation tasks. The difference is that ground text is unstructured and noisy. So, researchers design knowledge selection and incorporation methods to address the challenges. Based on the number of stages, we further divide related methods into two categories: retrieve-then-generate (also known as retrieval-augmented generation, short as RAG, in many existing papers (Lewis et al., 2020b; Petroni et al., 2021; Krishna et al., 2021)) methods (2-stage methods) and retrieve, rerank and rewrite methods (3-stage methods).
RAG follows a two-stage process: retrieval and generation. Specially, as shown in Figure 7(a), a retriever first returns (usually top-K truncated) distributions over text passages given a query , and then a generator generates a current token based on a context of the previous tokens , the original input and a retrieved passage . Methods for retrieving fact or review snippets are various, including matching from a collection of raw text entries indexed by named entities (Ghazvininejad et al., 2018); scoring relevant documents within a large collection by statistical approaches such as BM25 (Dinan et al., 2019), or neural-based retrieval approaches such as dense paragraph retrieval (DPR) (Lewis et al., 2020b). For training the retriever and generator, most of existing work has jointly optimized these two components, without any direct supervision on what document should be retrieve (Lewis et al., 2020b; Krishna et al., 2021). However, by asking human experts to label what document should be retrieved and adding the retrieval loss (resulting in a multi-task learning setting), the generation performance can be greatly improved (Dinan et al., 2019; Kim et al., 2020a), though the labelling process is an extremely time-consuming and labor-intensive task.
Ghazvininejad et al. proposed a knowledge grounded neural conversation model (KGNCM), which is the first work to retrieve review snippets from Foursquare and Twitter. Then it incorporates the snippets into dialogue response generation (Ghazvininejad et al., 2018). It uses an end-to-end memory network (Sukhbaatar et al., 2015) to generate responses based on the selected review snippets. Lewis et al. introduced a general retrieval-augmented generation (RAG) framework by leveraging a pre-trained neural retriever and generator. It can be easily fine-tuned on downstream tasks, and it has demonstrated state-of-the-art performance on various knowledge intensive NLG tasks (Lewis et al., 2020b). Recently, the fusion-in-decoder methods (i.e., the decoder performs attention over the concatenation of the resulting representations of all retrieved passages (Li et al., 2021; Yu et al., 2022a)) could even outperform RAG as reported in KILT benchmark (Petroni et al., 2021).
Different from RAG, a -based method is expected to retrieve a most precise reference document that can be directly used for rewriting/editing. -based method has proved successful in a number of NLG tasks such as machine translation (Gu et al., 2018), and summarization (Cao et al., 2018; Wang et al., 2019c). In summarization, Seq2Seq models that purely depend on the input document to generate summaries tend to deteriorate with the accumulation of word generation, e.g., they generate irrelevant and repeated words frequently (Cao et al., 2018; Wang et al., 2019c). Template-based summarization assume the golden summaries of the similar sentences (i.e., templates) can provide a reference point to guide the input sentence summarization process (Cao et al., 2018; Wang et al., 2019c). These templates are often called soft templates in order to distinguish from the traditional rule-based templates. Soft template-based summarization typically follows a three-step design: retrieve, rerank, and rewrite. The step of retrieval aims to return a few candidate templates from a summary collection. The reranking identifies the best template from the retrieved candidates. And the rewriting leverages both the source document and template to generate more faithful and informative summaries.
Compared with -based methods, RAG-based have several differences, including less of emphasis on lightly editing a retrieved item, but on aggregating content from several pieces of retrieved content, as well as learning latent retrieval, and retrieving evidence documents rather than related training pairs.
4. M2: Modeling Background Knowledge into Response Generation
Background document, with more global and comprehensive knowledge, has been often used for generating informative responses and ensuring a conversation to not deviate from its topic. Keeping a conversation grounded on a background document is referred as background based conversation (BBC) (Moghe et al., 2018; Bi et al., 2020). Background knowledge plays an important role in human-human conversations. For example, when talking about a movie, people often recall important points (e.g., a scene or review about the movie) and appropriately mention them in the conversation context. Therefore, an intelligent NLG model is expected to find an appropriate background snippet and generate response based on the snippet. As shown in Figure 7(b), the task of BBC is often compared with machine reading comprehension (MRC), in which a span is extracted from the background document as a response to a question (Rajpurkar et al., 2016). However, since BBC needs to generate natural and fluent responses, the challenge lies in not only locating the right semantic units in the background, but also referring to the right background information at the right time in the right place during the decoding phase.
As MRC models tie together multiple text segments to provide a unified and factual answer, many BBC models use the same idea to connect different pieces of information and find the appropriate background knowledge based on which the next response is to be generated (Qin et al., 2019; Meng et al., 2020). For instance, Qin et al. proposed an end-to-end conversation model that jointly learned response generation together with on-demand machine reading (Qin et al., 2019). The MRC models can effectively encode the input utterance by treating it as a question in a typical QA task (e.g., SQuAD (Rajpurkar et al., 2016)) and encode the background document as the context. Then, they took the utterance-aware background representation as input into decoding phase.
For M1, guiding generation with retrieved information explicitly exposes the role of world knowledge by asking the model to decide what knowledge to retrieve and use during language generation. Since retrieval-augmented generation (RAG) captures knowledge in a interpretable and modular way, it is often used for knowledge-intensive tasks such as long-form QA and argument generation. However, a knowledge retriever is expected to retrieve documents from a large-scale corpus, e.g., the entire Wikipedia, which causes significant computational challenge. Besides, one input often requires retrieved text whose amount is much larger than the input itself (as indicated in Table 7), leading to serious information overwhelming for the generation model.
For M2, background based conversations (BBCs) avoid generating generic responses in a dialogue system and are able to generate more informative responses by exploring related background information. However, existing methods still cannot solve inherent problems effectively, such as tending to break a complete semantic unit and generate shorter responses (Meng et al., 2020).
Table 7 summarizes tasks, datasets and evidence sources used in existing grounded text enhanced work. Three important things should be mentioned. First, all the datasets in the table are public, and we include their links in Table 12. Second, Wikipedia is the most commonly used evidence source since it is the largest free online encyclopedia. Besides, some online platforms contain plenty of product-related textural information, e.g., product reviews on Amazon, which are often used to build up task/goal oriented dialogue systems for business purpose. Third, the retrieval space of candidate documents are usually larger than 1 million and only 7-10 documents are selected. So, the process of retrieving relevant documents is challenging.
Table 8 compares different grounded text enhanced methods from three dimensions: retrieval supervision, pre-training of the retriever, and number of stages. First, as mentioned above, retrieving relevant documents from a large candidate set is a challenging task. To improve the retrieval accuracy, four (57.1%) papers added the retrieval supervision either by human annotated labels or pseudo labels, resulting in a multi-task learning setting. Besides, three (42.9%) papers used pre-trained language models to produce document representation for better retrieval. Though existing work has greatly improved the retrieval accuracy, the performance is still far from satisfactory in many text generation tasks (Lewis et al., 2020b; Krishna et al., 2021). How to learn mutually enhancement between retrieval and generation is still a promising direction in the grounded text enhanced text generation systems.
Benchmark, Toolkit and Leaderboard Performance
The development of general evaluation benchmarks for text generation helps to promote the development of research in related fields. Existing text generation benchmarks did not specially focus on choosing the tasks and datasets that have been widely used for knowledge-enhanced text generation. Therefore, we re-screened from the existing four text generation benchmarks, i.e., GLGE (Liu et al., 2021b), GEM (Gehrmann et al., 2021), KilT (Petroni et al., 2021), GENIE (Khashabi et al., 2021), and determined 9 benchmark datasets for evaluating knowledge-enhanced NLG methods. Here is our criteria for selection:
We only consider benchmark datasets that have open-access downloading link.
We focus on diverse text generation tasks, involving various applications.
We select at most three benchmark datasets for each text generation task.
We include a mix of internal and external knowledge focused datasets.
We prefer multi-reference datasets for robust automatic evaluation.
Based on the benchmark selection criteria, we finalize 9 knowledge-centric tasks that covers various NLG tasks, including commonsense reasoning, text summarization, question generation, generative question answering, and dialogue. The data statistics is shown in Table 9. Descriptions and dataset links are listed as follows:
Wizard of Wikipedia (WOW): It is an open-domain dialogue dataset, where two speakers conduct an open-ended conversion that is directly grounded with knowledge retrieved from Wikipedia. (Data link: https://parl.ai/projects/wizard_of_wikipedia/)
CommonGen: It is a generative commonsense reasoning dataset. Given a set of common concepts, the task is to generate a coherent sentence describing an everyday scenario using these concepts. (Data link: https://inklab.usc.edu/CommonGen/)
NLG-ART: It is a generative commonsense reasoning dataset. Given the incomplete observations about the world, the task it to generate a valid hypothesis about the likely explanations to partially observable past and future. (Data link: http://abductivecommonsense.xyz/)
ComVE: It is a generative commonsense reasoning dataset. The task is to generate an explanation given a counterfactual statement for sense-making. (Data link: https://github.com/wangcunxiang/SemEval2020-Task4-Commonsense-Validation-and-Explanation
ELI5: It is a dataset for long-form question answering. The task is to produce explanatory multi-sentence answers for diverse questions. Web search results are used as evidence documents to answer questions. (Data link: https://facebookresearch.github.io/ELI5/)
SQuAD: It is a dataset for answer-aware question generation. The task is to generate a question asks towards the given answer span based on a given text passage or document. (Data link: https://github.com/magic282/NQG)
CNN/DailyMail (CNN/DM): It is a dataset for summarization. Given a news aticles, the goal is to produce a summary that represents the most important or relevant information within the original content. (Data link: https://www.tensorflow.org/datasets/catalog/cnn_dailymail)
Gigaword: It is a dataset for summarization. Similar with CNN/DM, the goal is to generate a headline for a news article. (Data link: https://www.tensorflow.org/datasets/catalog/gigaword)
PersonaChat: It is an open-domain dialogue dataset. It presents the task of making chit-chat more engaging by conditioning on profile information. (Data link: https://github.com/facebookresearch/ParlAI/tree/master/projects/personachat)
Discussion on Future Directions
Many efforts have been conducted to tackle the problem of knowledge-enhanced text generation and its related applications. To advance the field, there remains several open problems and future directions. Designing more effective ways to represent knowledge and integrate them into the generation process is still the most important trend in knowledge-enhanced NLG systems. From a broader perspective, we provide three directions that make focusing such efforts worthwhile now: (i) incorporating knowledge into visual-language generation tasks, (ii) learning knowledge from broader sources, especially pre-trained language models, (iii) learning knowledge from limited resources, (iv) learning knowledge in a continuous way.
Beyond text-to-text generation tasks, recent years have witnessed a growing interest in visual-language (VL) generation tasks, such as describing visual scenes (Hossain et al., 2019), and answering visual-related questions (Mao et al., 2019). Although success has been achieved in recent years on VL generation tasks, there is still room for improvement due to the fact that image-based factual descriptions are often not enough to generate high-quality captions or answers (Zhou et al., 2019). External knowledge can be added in order to generate attractive image/video captions. We observed some pioneer work has attempted to utilize external knowledge to enhance the image/video captioning tasks. For example, Tran et al. proposed to detect a diverse set of visual concepts and generate captions by using an external knowledge base (i.e., Freebase), in recognizing a broad range of entities such as celebrities and landmarks (Tran et al., 2016). Zhou et al. used a commonsense knowledge graph (i.e., ConceptNet), to infer a set of terms directly or indirectly related to the words that describe the objects found in the scene by the object recognition module (Zhou et al., 2019). In addition, Mao et al. (2019) proposed a neuro-symbolic learner for improving visual-language generation tasks (e.g., visual question answering).
However, existing approaches for knowledge-enhanced visual-language generation tasks still have a lot of space for exploration. Some promising directions for future work include using other knowledge sources, such as retrieving image/text to help solve open-domain visual question answering and image/video captioning tasks; bringing structured knowledge for providing justifications for the captions that they produce, tailoring captions to different audiences and contexts, etc.
2. Learning Knowledge from Broader Sources
More research efforts should be spent on learning to discover knowledge more broadly and combine multiple forms of knowledge from different sources to improve the generation process. More knowledge sources can be but not limited to network structure, dictionary and table. For examples, Yu et al. (Yu et al., 2020) and An et al. (An et al., 2021) augmented the task of scientific papers intention detection and summarization by introducing the citation graph; Yu et al. augmented the rare word representations by retrieving their descriptions from Wiktionary and feed them as additional input to a pre-trained language model (Yu et al., 2021a). Besides, structured knowledge and unstructured knowledge can play a complementary role in enhancing text generation. To improve knowledge richness, Fu et al. combined both structured (knowledge base) and unstructured knowledge (grounded text) (Fu and Feng, 2018).
Pre-trained language models can learn a substantial amount of in-depth knowledge from data without any access to an external memory, as a parameterized implicit knowledge base (Lewis et al., 2020b; Raffel et al., 2020). However, as mentioned in (Guan et al., 2020), directly fine-tuning pre-trained language generation models on the story generation task still suffers from insufficient knowledge by representing the input text thorough a pre-trained encoder, leading to repetition, logic conflicts, and lack of long-range coherence in the generated output sequence. Therefore, discovering knowledge from pre-trained language models can be more flexible, such as knowledge distillation, data augmentation, and using pre-trained models as external knowledge (Petroni et al., 2019). More efficient methods of obtaining knowledge from pre-trained language models are expected.
3. Learning Knowledge from Limited Resources
Most of current NLG research conduct on extensively labelled data to favor model training. However, this is in contrast to many real-world application scenarios, where only a few shots of examples are available for new domains. Limited data resources lead to limited knowledge that can be learnt in new domains. For examples, learning topical information of a dialogue occurring under a new domain is difficult since the topic may be rarely discussed before; constructing a syntactic dependency graph of a sequence in a low-resource language is hard since many linguistic features are of great uniqueness. Besides, external knowledge bases are often incomplete and insufficient to cover full entities and relationships due to the human costs of collecting domain-specific knowledge triples. Therefore, quick domain adaptation is an essential task in text generation tasks. One potential route towards addressing these issues is meta-learning, which in the context of NLG means a generation model develops a broad set of skills and pattern recognition abilities at training time, and quickly adapt to a new task given very few examples without retraining the model from scratch. Recently, there has been raising interests in both academia and industry to investigate meta-learning in different NLG tasks. Thus, it is a promising research direction to build efficient meta-learning algorithms that only need a few task-specific fine-tuning to learn the new task quickly. And for knowledge-enhanced text generation, it is of crucial importance to adapt the model quickly on new domains with limited new knowledge (e.g., only a few knowledge triples).
4. Learning Knowledge in a Continuous Way
A machine learning is expected to learn continuously, accumulate the knowledge learned in previous tasks, and use it to assist future learning. This research direction is referred as lifelong learning (Chen and Liu, 2018). In the process, the intelligent machine becomes more and more knowledgeable and effective at learning new knowledge. To make an analogy, humans continuously acquire new knowledge and constantly update the knowledge system in the brain. However, existing knowledge-enhanced text generation systems usually do not keep updating knowledge in real time (e.g., knowledge graph expansion). A meaningful exploration of was discussed in (Mazumder et al., 2018). They built a general knowledge learning engine for chatbots to enable them to continuously and interactively learn new knowledge during conversations. Therefore, it is a promising research direction to continuously update knowledge obtained from various information sources, empowering intelligent machines with incoming knowledge and improving the performance on new text generation tasks.
Conclusions
In this survey, we present a comprehensive review of current representative research efforts and trends on knowledge-enhanced text generation, and expect it can facilitate future research. To summarize, this survey aims to answer two questions that commonly appears in knowledge-enhanced text generation: how to acquire knowledge and how to incorporate knowledge to facilitate text generation. Base on knowledge acquisition, the main content of our survey is divided into three sections according to different sources of knowledge enhancement. Based on knowledge incorporation, we first present general methods of incorporating knowledge into text generation and further discuss a number of specific ideas and technical solutions that incorporate the knowledge to enhance the text generation systems in each section. Besides, we review a variety of text generation applications in each section to help practitioners learn to choose and employ the methods.
Acknowledgements
We thank all anonymous reviewers for valuable comments. We also appreciate the suggestions from readers of the pre-print version. We thank Dr. Michael Zeng (Microsoft) and Dr. Nazneen Rajani (Saleforce) for their constructive comments and suggestions. Wenhao Yu and Dr. Meng Jiang’s research is supported by National Science Foundation grants IIS-1849816, CCF-1901059, and IIS-2119531. Qingyun Wang and Dr. Heng Ji’s research is based upon work supported by Agriculture and Food Research Initiative (AFRI) grant no. 2020-67021-32799/project accession no.1024178 from the USDA National Institute of Food and Agriculture, U.S. DARPA SemaFor Program No. HR001120C0123, DARPA AIDA Program No. FA8750-18-2-0014, and DARPA KAIROS Program No. FA8750-19-2-1004. The views and conclusions contained herein are those of the authors and should not be interpreted as necessarily representing the official policies, either expressed or implied, of DARPA, or the U.S. Government. The U.S. Government is authorized to reproduce and distribute reprints for governmental purposes notwithstanding any copy-right annotation therein.
References
Appendix A Appendix
Figure 8 demonstrates the statistics of selected publications in this survey. The left figure shows the paper publishing venues. Most papers were published in top machine learning, artificial intelligence, and natural language processing conferences, such as ACL, EMNLP, AAAI, ICLR, NeurIPS. Besides, many selected papers were published in high-impact journals, such as TNNLS, JMLR, TACL. The right figure shows the paper categories. Among 160 selected papers, 87 papers (“general methods (General)”, “topic”, “keyword”, “knowledge base (KG)”, “knowledge graph (KG)”, “grounded text (Text)”) are directly relevant to the different kinds of knowledge-enhanced text generation methods; 10 papers are relevant to benchmark datasets; 10 papers are related survey papers. Besides, other 43 papers are about basic (pre-trained) generation methods (e.g., Seq2Seq, CopyNet, BART, T5), or necessary background (e.g., TransE, OpenIE, GNN, LDA), or future direction.
Figure 9 summarized different papers according to years, knowledge sources, and methods.
Table 10(e) lists the leaderboard performance on ten knowledge-enhanced generation benchmarks.
Table 12 lists code links and programming language of representative open-source knowledge-enhanced text generation systems that have been introduced in this survey.
BLEU- (short as B-): BLEU is a weighted geometric mean of -gram precision scores.
ROUGE- (short as R-): ROUGE measures the overlap of n-grams between the reference and hypothesis; ROUGE-L measures the longest matched words using longest common sub-sequence.
Distinct- (short as D-k): Distinct measures the total number of unique -grams normalized by the total number of generated -gram tokens to avoid favoring long sentences.