Learning to generate one-sentence biographies from Wikidata

Andrew Chisholm, Will Radford, Ben Hachey

Introduction

Despite massive effort, Wikipedia and other collaborative knowledge bases (kbs) have coverage and quality problems. Popular topics are covered in great detail, but there is a long tail of specialist topics with little or no text. Other text can be incorrect, whether by accident or vandalism. We report on the task of generating textual summaries for people, mapping slot-value facts to one-sentence encyclopaedic biographies. In addition to initialising stub articles with only structured data, the resulting model could be used to improve consistency and accuracy of existing articles. Figure 1 shows a Wikidata entry for Mathias Tuomi, with fact keys and values flattened into a sequence, and the first sentence from his Wikipedia article. Some values are in the text, others are missing (e.g. male) or expressed differently (e.g. dates).

We treat this knowlege-to-text task like translation, using a recurrent neural network (rnn) sequence-to-sequence model [Sutskever et al., 2014] that learns to select and realise the most salient facts as text. This includes an attention mechanism to focus generation on specific facts, a shared vocabulary over input and output, and a multi-task autoencoding objective for the complementary extraction task. We create a reference dataset comprising more than 400,000 knowledge-text pairs, handling the 15 most frequent slots. We also describe a simple template baseline for comparison on bleu and crowd-sourced human preference judgements over a heldout test set.

Our model obtains a bleu score of 41.0, compared to 33.1 without the autoencoder and 21.1 for the template baseline. In a crowdsourced preference evaluation, the model outperforms the baseline and is preferred 40% of the time to the Wikipedia reference. Manual analysis of content selection suggests that the model can infer knowledge but also makes mistakes, and that the autoencoding objective encourages the model to select more facts without increasing sentence length. The task formulation and models are a foundation for text completion and consistency in kbs.

Background

rnn sequence-to-sequence models [Sutskever et al., 2014] have driven various recent advances in natural language understanding. While initial work focused on problems that were sequences of the same units, such as translating a sequence of words from one language to another, other work been able to use these models by coercing different structures into sequences, e.g., flattening trees for parsing [Vinyals et al., 2015], predicting span types and lengths over byte input [Gillick et al., 2016] or flattening logical forms for semantic parsing [Xiao et al., 2016].

rnns have also been used successfully in knowledge-to-text tasks for human-facing systems, e.g., generating conversational responses [Vinyals and Le, 2015], abstractive summarisation [Rush et al., 2015]. Recurrent lstm models have been used with some success to generate text that completely expresses a set of facts: restaurant recommendation text from dialogue acts [Wen et al., 2015], weather reports from sensor data and sports commentary from on-field events [Mei et al., 2015]. Similarly, we learn an end-to-end model trained over key-value facts by flattening them into a sequence.

Choosing the salient and consistent set of facts to include in generated output is also difficult. Recent work explores unsupervised autoencoding objectives in sequence-to-sequence models, improving both text classification as a pretraining step [Dai and Le, 2015] and translation as a multi-task objective [Luong et al., 2016]. Our work explores an autoencoding objective which selects content as it generates by constraining the text output sequence to be predictive of the input.

Biographic summarisation has been extensively researched and is often approached as a sequence of subtasks [Schiffman et al., 2001]. A version of the task was featured in the Document Understanding Conference in 2004 [Blair-Goldensohn et al., 2004] and other work learns policies for content selection without generating text [Duboue and McKeown, 2003, Zhang et al., 2012, Cheng et al., 2015]. While pipeline components can be individually useful, integrating selection and generation allows the model to exploit the interaction between them.

kbs have been used to investigate the interaction between structured facts and unstructured text. Generating textual templates that are filled by structured data is a common approach and has been used for conversational text [Han et al., 2015] and biographical text generation [Duma and Klein, 2013]. Wikipedia has also been a popular resource for studying biography, including sentence harvesting and ordering [Biadsy et al., 2008], unsupervised discovery of distinct sequences of life events [Bamman and Smith, 2014] and fact extraction from text [Garera and Yarowsky, 2009]. There has also been substantial work in generating from other structured kbs using template induction [Kondadadi et al., 2013], semantic web techniques [Power and Third, 2010], tree adjoining grammars [Gyawali and Gardent, 2014], probabilistic context free grammars [Konstas and Lapata, 2012] and probabilistic models that jointly select and realise content [Angeli et al., 2010].

?) present the closest work to ours with a similar task using Wikipedia infoboxes in place of Wikidata. They condition an attentional neural language model (nlm) on local and global properties of infobox tables, including copy actions that allow wholesale insertion of values into generated text. They use 723k sentences from Wikipedia articles with 403k lower-cased words mapping to 1,740 distinct facts. They compare to a 5-gram language-model with copy actions, and find that the nlm has higher bleu and lower perplexity than their baseline. In contrast, we utilise a deep recurrent model for input encoding, minimal slot value templating and greedy output decoding. We also explore a novel autoencoding objective that measures whether input facts can be re-created from the generated sentence.

Evaluating generated text is challenging and no one metric seems appropriate to measure overall performance. ?) report bleu scores [Papineni et al., 2002] which calculate the n-gram overlap between text produced by the system with respect to a human-written reference. Summarisation evaluations have concentrated on the content that is included in the summary, with semantic content typically extracted manually for comparison [Lin and Hovy, 2003, Nenkova and Passonneau, 2004]. We draw from summarisation and generation to formulate a comprehensive evaluation based on automated metrics and human validation. Our final system comparison follows ?) in running a crowd task to collect pairwise preferences for evaluating and comparing both systems and references.

Task and Data

We formulate the one-sentence biography generation task as shown in Figure 1. Input is a flat string representation of the structured data from the kb, comprising slot-value pairs (the subject being the topic of the kb record, e.g., Mathias Tuomi), ordered by slot frequency from most to least common. Output is a biography string describing the salient information in one sentence.

We validate the task and evaluation using a closely-aligned set of resources: Wikipedia and Wikidata. In addition to the kb maintenance issues discussed in the introduction, Wikipedia first sentences are of particular interest because they are clear and concise biographical summaries. These could be applied to entities outside Wikipedia for which one can obtain comparable parallel structured/textual data, e.g., movie summaries from IMDb, resume overviews from LinkedIn, product descriptions from Amazon.

We use snapshots of Wikidata (2015/07/13) and Wikipedia (2015/10/02) and batch process them to extract instances for learning. We select all entities that are INSTANCE_OF human in Wikidata. We then use sitelinks to identify each entity’s Wikipedia article text and nltk [Bird et al., 2009] to tokenize and extract the lower-cased first sentence. This results in 1,268,515 raw knowledge-text pairs. The summary sentences can be long and the most frequent length is 21 tokens. We filter to only include those between the 10th and 90th percentiles: 10 and 37 tokens. We split this collection into train, dev and test collections with 80%, 10% and 10% of instances allocated respectively. Given the large variety of slots which may exist for an entity, we restrict the set of slots used to the top-15 by occurrence frequency. This criteria covers 72.8% of all facts. Table 1 shows the distribution of fact slots in the structured data and the percentage of time tokens from a fact value occur in the corresponding Wikipedia summary.

Additionally, some Wikidata entities remain underpopulated and do not contain sufficient facts to reconstruct a text summary. We control for this information mismatch by limiting our dataset to include only instances with at least 6 facts present. The final dataset includes 401,742 train, 50,017 dev and 50,030 test instances. Of these instances, 95% contain 6 to 8 slot values while 0.1% contain the maximum of 10 slots. 51% of unique slot-value pairs expressed in test and dev are not observed in train so generalisation of slot usage is required for the task. The kb facts give us an opportunity to measure the correctness of the generated text in a more precise way than text-to-text tasks. We use this for analysis in Section 7.3, driving insight into system characteristics and implications for use.

Wikipedia first sentences exhibit a relatively narrow domain of language in comparison to other generation tasks such as translation. As such, it is not clear how complex the generation task is, and we first try to use perplexity to describe this.

We train both rnn models until dev perplexity stops improving. Our basic sequence-to-sequence model (s2s) reaches perplexity of 2.82 on train and 2.92 on dev after 15,000 batches of stochastic gradient descent. The autoencoding sequence-to-sequence model (s2s+ae) takes longer to fit, but reaches a lower minimum perplexity of 2.39 on train and 2.51 on dev after 25,000 batches.

To help ground perplexity numbers and understand the complexity of sentence biographies we train a benchmark language model and evaluate perplexity on dev. Following ?), we build Kneser-Ney smoothed 5-gram language models using the KenLM toolkit [Heafield, 2011].

Table 2 lists perplexity numbers for the benchmark LM models with different templating schemes on dev. We observe decreasing perplexity for data with greater fact value templating. title indicates templating of entity names only, while full indicates templating of all fact values by token index as described in ?). This shows that templating is an effective way to reduce the sparsity of a task, and that titles account for a large component of this.

Although ?) evaluate on a different dataset, we are able to draw some comparisons given the similarity of our task. On their data, the benchmark LM baseline achieves a similar perplexity of 10.5 to ours when following their templating scheme on our dataset - suggesting both samples are of comparable complexity.

Model

We model the task as a sequence-to-sequence learning problem. In this setting, a variable length input sequence of entity facts is encoded by a multi-layer rnn into a fixed-length distributed representation. This input representation is then fed into a separate decoder network which estimates a distribution over tokens as output. During training, parameters for both the encoder and decoder networks are optimized to maximize the likelihood of a summary sequence given an observed fact sequence.

Our setting differs from the translation task in that the input is a sequence representation of structured data rather than natural human language. As described above in Section 3, we map Wikidata facts to a sequence of tokens that serves as input to the model as illustrated at the top of Figure 2. Experiments below demonstrate that this is sufficient for end-to-end learning in the generation task addressed here. To generate summaries, our model must both select relevant content and transform it into a well formed sentence. The decoder network includes an attention mechanism [Vinyals et al., 2015] to help facilitate accurate content selection. This allows the network to focus on different parts of the input sequence during inference.

To generate language, we seed the decoder network with the output of the encoder and a designated GO token. We then generate symbols greedily, taking the most likely output token from the decoder at each step given the preceding sequence until an EOS token is produced. This approach follows [Sutskever et al., 2014] who demonstrate a larger model with greedy sequence inference performs comparably to beam search. In contrast to translation, we might expect good performance on the summarization task where output summary sequences tend to be well structured and often formulaic. Additionally, we expect a partially-shared language across input and output. To exploit this, we use a tied embedding space, which allows both the encoder and decoder networks to share information about word meaning between fact values and output tokens.

Our model uses a 3-layer stacked Gated Recurrent Unit rnn for both encoding and decoding, implemented using TensorFlow.https://www.tensorflow.org, v0.8. We limit the shared vocabulary to 100,000 tokens with 256 dimensions for each token embedding and hidden layer. Less common tokens are marked as UNK, or unknown. To account for the long tail of entity names, we replace matches of title tokens with templated copy actions (e.g. TITLE0 TITLE1…). These template are then filled after generation, as well as any initial unknown tokens in the output, which we fill with the first title token. We learn using minibatch Stochastic Gradient Descent with a batch size of 64 and a fixed learning rate of 0.5.

2 s2s with autoencoding (s2s+ae)

One challenge for vanilla sequence-to-sequence models in this setting is the lack of a mechanism for constraining output sequences to only express those facts present in the data. Given a fact extraction oracle, we might compare facts expressed in the output sequence with those of the input and appropriately adjust the loss for each instance. While a forward-only model is only constrained to generate text sequences predicted by the facts, an autoencoding model is additionally constrained to generate text predictive of the input facts. In place of this ideal setting, we introduce a second sequence-to-sequence model which runs in reverse - re-encoding the text output sequence of the forward model into facts.

This closed-loop model is detailed in Figure 3. The resulting network is trained end-to-end to minimize both the input-to-output sequence loss L(x,y)L(x,y) and output-to-input reconstruction loss L(x,x′)L(x,x^{\prime}). While gradients cannot propagate through the greedy forward decode step, shared parameters between the forward and backward network are fit to both tasks. To generate language at test time, the backward network does not need to be evaluated.

Experimental methodology

The evaluation suite here includes standard baselines for comparison, automated metrics for learning, human judgement for evaluation and detailed analysis for diagnostics. While each are individually useful, their combination gives a comprehensive analysis of a complex problem space.

We use the first sentence from Wikipedia both as a gold standard reference for evaluating generated sentences, and as an upper bound in human preference evaluation.

base

Template-based systems are strong baselines, especially in human evaluation. While output may be stilted, the corresponding consistency can be an asset when consistency is important. We induce common patterns from the train set, replacing full matches of values with their slot and choosing randomly on ties. Multiple non-fact tokens are collapsed to a single symbol. A small sample of the most frequent patterns were manually examined to produce templates, roughly expressed as: TITLE, known as GIVEN_NAME, (born DATE_OF_BIRTH in PLACE_OF_BIRTH; died DATE_OF_DEATH in PLACE_OF_DEATH) is an POSITION_HELD and OCCUPATION from CITIZENSHIP, with some sensible back-offs where slots are not present, and rules for determiner agreement and is versus was where a death date is present. For example, ollie freckingham (born 12 november 1988) is a cricketer from the united kingdom. In total, there are 48 possible template variations.

2 Metrics

We also report bleu n-gram overlap with respect to the reference Wikipedia summary. With a large dev/test sets (10,000 sentences here), bleu is a reasonable evaluation of generated content. However, it does not give an indication of well-formedness or readability. Thus we complement bleu with a human preference evaluation.

Human preference

We use crowd-sourced judgements to evaluate the relative quality of generated sentences and the reference Wikipedia first sentence. We obtain pairwise judgements, showing output from two different systems to crowd workers and asking each to give their binary preference. The system name mappings are anonymized and ordered pseudo-randomly. We request 3 judgements and dynamically increase this until we reach at least 70% agreement or a maximum of 5 judgements. We use CrowdFlowerhttp://www.crowdflower.com to collect judgements at the cost of 31 USD for all 6 pairwise combinations over 82 randomly selected entities. 67 workers contributed judgements to the test data task, each providing no more than 50 responses. We use the majority preference for each comparison. The CrowdFlower agreement is 80.7%, indicating that roughly 4 of 5 votes agree on average.

3 Analysis of content selection

Finally, no system is perfect, and it can be challenging to understand the inherent difficulty of the problem space and the limitations of a system. Due to the limitations of the evaluation metrics mentioned above, we propose that manual annotation is important and still required for qualitative analysis to guide system improvement. The structured data in knowledge-to-text tasks allows us, if we can identify expressions of facts in text, cases where facts have been omitted, incorrectly mentioned, or expressed differently.

Results

Table 3 shows bleu scores calculated over 10,000 entities sampled from dev and test using the Wikipedia sentence as a single reference, using uniform weights for 1- to 4-grams, and padding sentences with fewer than 4 tokens. Scores are similar across dev and test, indicating that the samples are of comparable difficulty. We evaluate significance using bootstrapped resampling with 1,000 samples. Each system result lies outside the 95% confidence intervals of other systems. base has reasonable scores at 21, with s2s higher at around 32, indicating that the model is at least able to generate closer text than the baseline. s2s+ae scores higher still at around 41, roughly double the baseline scores, indicating that the autoencoder is indeed able to constrain the model to generate better text.

2 Human preference evaluation

Table 4 shows the results of our human evaluation over 82 entities sampled from test. For each pair of systems, we show the percentage of entities where the crowd preferred A over B. Significant differences are annotated with ∗\ast and ∗∗\ast\ast for pp values << 0.05 and 0.01 using a one-way χ2\chi^{2} test. wiki is uniformly preferred to any system, as is appropriate for an upper bound. The s2s model is the least-preferred with respect to wiki. The s2s+ae model is more-preferred than the base and s2s models, by a larger margin for the latter. These results show that without autoencoding, the sequence-to-sequence model is less effective than a template-based system. Finally, although wiki is more preferred than s2s+ae, the distributions are not significantly different, which we interpret as evidence that the model is able to generate good text from the human point-of-view, but autoencoding is required to do so.

Analysis

While results presented above are encouraging and suggest that the model is performing well, they are not diagnostic in the sense that they can drive deeper insights into model strengths and weaknesses. While inspection and manual analysis is still required, we also leverage the structured factual data inherent to our task to perform quantitative as well as qualitative analysis.

Figure 4 shows the effects of input fact count on generation performance. While more input facts give more information for the model to work with, longer inputs are also both rarer and more complex to encode. Interestingly, we observe the s2s+ae model maintains performance for more complex inputs while s2s performance declines.

Table 7.1 shows some dev entities and their summaries. The model learns interesting mappings: between numeric and string dates, and country demonyms. The model also demonstrates the ability to work around edge cases where templates fail, i.e. stripping parenthetical disambiguations (e.g. (actor)) and emitting the name Robert when the input is Bob. Output also suggests the model may perform inference across multiple facts to improve generation precision, e.g. describing an entity as english rather than british given information about both citizenship and place of birth. Unfortunately, the model can also infer unsubstantiated facts into the text (i.e. jazz drummer).

We randomly sample 50 entities from dev and manually annotate the Wikipedia and system text. We note which fact slots are expressed as well as whether the expressed values are correct with respect to Wikidata. Given two sets of correctly extracted facts, we can consider one gold, one system and calculate set-based precision, recall and F1.

Firstly, to understand how Wikipedia editors select content for the first sentence of articles, we measure recall with the real facts as gold, and Wikipedia as system. Overall, the recall is 0.61 indicating that 61% of input facts are expressed in the reference summary from Wikipedia. The entity name (TITLE) is always expressed. Four slots are nearly always expressed when available: OCCUPATION (90%), DATE_OF_BIRTH (84%), CITIZENSHIP (81%), DATE_OF_DEATH (80%). Six slots are infrequently expressed in the analysis sample: PLACE_OF_BIRTH (33%), POSITION_HELD (25%), PARTICIPANT_OF (20%), POLITICAL_PARTY (20%), EDUCATED_AT (14%), SPORTS_TEAM (9%). Two are never expressed explicitly: PLACE_OF_DEATH (0%), SEX_OR_GENDER (0%). AWARD_RECEIVED and SPORT are not in the analysis sample.

Do systems select the same facts found in the reference summaries?

Table 6 shows content selection scores for systems with respect to the Wikipedia text as reference. This suggests that the autoencoding in s2s+ae helps increase fact recall without sacrificing precision. The template baseline also attains this higher recall, but at the cost of precision. For commonly expressed facts found in most person biographies, recall is over 0.95 (e.g., CITIZENSHIP, BIRTH_DATE, DEATH_DATE and OCCUPATION). Facts that are infrequently expressed are more difficult to select, with system F1 ranging from 0.00 to 0.50. Interestingly, macro-averaged F1 across infrequently expressed facts mirror human preference rather than bleu results, with s2s+ae (0.26) >> base (0.17) >> s2s (0.07). However, all systems perform poorly on these facts and no reliable differences are observed.

How does autoencoding effect fact density?

Interestingly, we observe that the autoencoding objective encourages the model to select more facts (5.2 for s2s+ae vs. 4.5 for s2s), without increasing sentence length (19.1 vs. 19.7 tokens). base is similarly productive (5.1 facts) but wordier (21.2 tokens), while the wiki reference produces both more facts (6.1) and longer sentences (23.7).

Do systems hallucinate facts?

To quantify the effect of hallucinated facts, we asses content selection scores of systems with respect to the input Wikidata relations (Table 7). Our best model achieves a precision of 0.93 with respect to Wikidata input. Notably, the template-driven baseline maintains a precision of 1.0 as it is constrained to emit Wikidata facts verbatim.

Our experiments show that rnns can generate biographic summaries from structured data, and that a secondary autoencoding objective is able to account for some of the information mismatch between input facts and target output sentences. In the future, we will explore whether results improve with explicit modelling of facts and conditioning of generation and autoencoding losses on slots. We expect this could benefit generation for diverse and noisy slot schemas like Wikipedia Infoboxes.

Another natural extension is to investigate the performance of the network running in reverse, from summary text back to facts. We plan to isolate the performance of the s2s+ae backward model when inferring facts and compare it to standard relation extraction systems. Finally, similar rnn models have been applied extensively to language translation tasks. We plan to explore whether a joint model of machine translation and fact-driven generation can help populate kb entries for low-coverage languages by leveraging a shared set of facts.

We present a neural model for mapping between structured and unstructured data, focusing on creating Wikipedia biographic summary sentences from Wikidata slot-value pairs. We introduce a sequence-to-sequence autoencoding rnn which improves upon base models by jointly learning to generate text and reconstruct facts. Our analysis of the task suggests evaluation in this domain is challenging. In place of a single score, we analyse statistical measures, human preference judgements and manual annotation to help characterise the task and understand system performance. In the human preference evaluation, our best model outperforms template baselines and is preferred 40% of the time to the gold standard Wikipedia reference.

Code and data is available at https://github.com/andychisholm/mimo.

This work was supported by a Google Faculty Research Award (Chisholm) and an Australian Research Council Discovery Early Career Researcher Award (DE120102900, Hachey). Many thanks to reviewers for insightful comments and suggestions, and to Glen Pink, Kellie Webster, Art Harol and Bo Han for feedback at various stages.