PlotMachines: Outline-Conditioned Generation with Dynamic Plot State Tracking
Hannah Rashkin, Asli Celikyilmaz, Yejin Choi, Jianfeng Gao
Introduction
Composing a story requires a complex planning process. First, the writer starts with a rough sketch of what key characters and events the story will contain. Then, as they unfold the story, the writer must keep track of the elaborate plot that weaves together the characters and events in a coherent and consistent narrative.
We study this complex storytelling process by formulating it as the task of outline-conditioned story generation, illustrated in Figure 1. Given an outline, a set of phrases describing key characters and events to appear in a story, the task is to generate a coherent narrative that is consistent with the provided outline. This task is challenging as the input provides only the rough elements of the plotHere, we define plot as the main sequence of events in the story.. Thus, the model needs to flesh out how these plot elements will intertwine with each other across different parts of the story. The flowchart in Figure 1 demonstrates an example of a latent plot structure: different key phrases from the outline appear and re-appear jointly throughout different sentences and paragraphs. Notably, the way that outline points are interwoven needs to be determined dynamically based on what’s already been composed while also staying true to the original outline and overall narrative structure.
We present PlotMachines, a novel narrative transformer that simulates the outline-conditioned generation process described above.code available at https://github.com/hrashkin/plotmachines Our model learns to transform an outline into a multi-paragraph story using dynamic memory blocks that keep track of the implicit plot states computed using the outline and the story generated thus far. We draw inspiration from prior work in dialogue state tracking (Thomson and Young, 2010; Lee, 2013; Chao and Lane, 2019), entity tracking (Henaff et al., 2017; Bosselut et al., 2018), and memory networks (Sukhbaatar et al., 2015) for keeping track of plot states. We also inform our model with high-level narrative structure using discourse labels so that it can learn different styles of writing corresponding to different parts of the narrative (i.e. beginning, middle, and end). PlotMachines is, to the best of our knowledge, the first model designed to generate multi-paragraph stories conditioned on outlines and can be trained end-to-end to learn the latent plot patterns without explicit plot annotations for supervision.
To support research on outline-conditioned generation, we present three datasets, including both fiction and non-fiction domains, where multi-paragraph narratives from existing datasets are paired with automatically constructed outlines using state-of-the-art key phrase extraction. Importantly, our task formulation of outline-conditioned generation is general and can be applied to various forms of grounded language generation. Comprehensive experiments on these datasets demonstrate that recently introduced state-of-the-art large-scale language models such as GPT-2 and Grover (Radford et al., 2019; Zellers et al., 2019), despite their impressive generation performance, still struggle to generate coherent narratives that are consistent with input outlines. Our experiments indicate that dynamic plot state tracking is important for constructing narratives with tighter and more consistent plots compared to competitive baselines.
Our main contributions are: (1) a new task formulation of outline-conditioned story generation, (2) the presentation of three new datasets for this task, (3) PlotMachines, a novel narrative transformer that learns to transform outlines to full stories with dynamic plot state tracking, and (4) empirical results demonstrating the limitations of state-of-the-art large-scale language models and the advantage of PlotMachines compared to competitive baselines.
Outline-Conditioned Generation
Our primary goal is to design a task for investigating how story generation models can plan long narrative according to controllable story elements. To that end, we introduce the outline-conditioned story generation task, which takes a plot outline as input and produces a long, multi-paragraph story.
In order to be flexible to multiple forms of control that might be required for different downstream tasks, we envision plot outlines to be defined loosely as lists of an arbitrary number of un-ordered plot points that should guide a story being generated. Plot points could consist of high-level concepts, low-level events, or even detailed sentences. For practical reasons, in this work, we limit the scope of plot points to events and phrases since these can be automatically extracted. Future work could explore alternate methods of defining plot outlines, perhaps using an event-based planning systems Porteous and Cavazza (2009); Riedl (2009); Riedl and Young (2010); Fan et al. (2019) for generating key points.
More concretely, in this paper, we formulate the outline as a list of un-ordered bullet points which reflect key phrases to be loosely integrated in the output narrative. These plot outlines are inspired, in part, by previous work in short-form story generation tasks that conditioned on storylines (Peng et al., 2018; Yao et al., 2019), which were defined as an ordered list of exactly five single-word points. We extend this concept to long-form story generation by defining a plot outline more flexibly as: an un-ordered list of an arbitrary number of multi-word plot elements. An outline also differs from a writing prompt, such as those found in other controllable writing tasks Fan et al. (2018), which are more abstract and often just a starting point for a story. Unlike a prompt, an outline is a list of concrete points that must appear somewhere in the narrative.
One challenge of this task is to create stories that have appropriate discourse and narrative flow. A second challenge is for stories to include the outline in a natural way. For example, it may be appropriate for certain outline points to be used only later on in the story (e.g. the protagonist dying may be more typically used at the end).
We construct three datasets for outline-conditioned generationCode for replicating data creation available at www.github.com/hrashkin/plotmachines by creating novel plot outlines to be used as inputs to generating stories from three existing story datasets. Table 1 shows statistics and examples from each dataset. We focus on fictitious generation, but also include the news domain for generalization. We build on existing story datasets for the target narratives, which we pair with automatically constructed input outlines as described below:
Wikiplots corpus www.github.com/markriedl/WikiPlots consists of plots of TV shows, movies, and books scraped from Wikipedia.
WritingPrompts Fan et al. (2018) is a story generation dataset, collected from the /r/WritingPrompts subreddit a forum where Reddit users compose short stories inspired by other users’ prompts. We use the same train/dev/test split from the original dataset paper.
NYTimes Sandhaus (2008) contains news articles rather than fictional stories, unlike the other two datasets.Due to concerns over fake news creation, we will not release the model trained on this data.
We extract a list of plot outlines from each dataset to use as input using the RAKE (Rapid Automatic Keyword Extraction) algorithm Rose et al. (2010)https://pypi.org/project/rake-nltk/. RAKE is a domain independent keyword extraction algorithm, which determines key phrases in a document based on the word frequency and co-occurrence statistics. We filtered key-points with overlapping n-grams. This is inspired by similar RAKE-based methods for creating storylines (Peng et al., 2018), but differs in that we extract longer outline points (3-8 words each) with no particular order.
PlotMachines
Our approach to this task is to design a model that combines recent success in text generation with transformer-based architectures (Vaswani et al., 2017) with memory mechanisms that keep track of the plot elements from the outline as they are used in the story. We also incorporate special discourse features into the modelling to learn a structure over the long multi-paragraph story format.
We introduce PlotMachines (PM), an end-to-end trainable transformer built on top of the GPT modelWe build on top of GPT, though our approach could be used with most transformer-based LMs. In experiments, we also look at a version of PlotMachines using GPT-2 (Radford et al., 2019) as a base. Radford et al. (2018), as shown in Figure 2. Given an outline as input, the model generates paragraphs, recurrently, while updating a memory matrix that keeps track of plot elements from the outline. This generation framework is motivated by human writing styles, in which each paragraph is a distinct section of related sentences.
where is the outline representation (Sec. 3.1), is the discourse representation associated with paragraph (Sec. 3.2), is a vector representation of the preceding story context (Sec. 3.3), and is the previous memory (Sec. 3.4).
The plot outline (i.e. the input to the model) is treated as a sequence of tokens, , and used as input for the transformer for each paragraph that is generated. We use special _kw_ tokens to delimit each plot point in the outline and end the sequence with a special _endkw_ token. We truncate the entire outline to maximum of tokens. For example, an outline containing two plot points ({‘strange phone call’, ‘detective’}) is turned into the input sequence: strange phone call _kw_ detective _endkw_
2 Discourse Representation
We posit that there are stylistic differences in how the beginning, middle and end of a story are written. To learn these differences, we introduce , discourse information about whether the -th paragraph is an introduction, body, or conclusion paragraph. We append a special token to the outline representation as part of the input sequence: _i_, _b_, _c_ for the introduction, body, and conclusion paragraphs respectivelyWe make the simplifying assumption that the first paragraph is an introduction, the last paragraph is the conclusion paragraph, and the other paragraphs are all body paragraphs..
3 Preceding Context Representation
With the goal of incorporating previous story context in generating each paragraph, we use , an embedded representation of the previous paragraph, which is added to the model input. More concretely, is computed as the average embedding of GPT output representations of words from the previous paragraph (using a GPT model that is static, i.e. not finetuned). The vector is used as an initial input to the transformer architecture, as shown in Figure 2.
4 Memory Representation
We implement memory to address two key task challenges. First, we want to keep track of the portions of the outline that have been mentioned. Second, we want to maintain semantic coherence across the entire story. To address these two challenges, we implement the memory as consisting of two parts: , a set of vectors keeping track of outline points, and , a matrix that stores a latent topic distribution of what’s been written so far.
Updating memory: The memory is updated (top left corner of Fig. 2) using , the average GPT output representation of the previous paragraph. We use update equations based on those in entity-based models such as Henaff et al. (2017). We use a gating mechanism, , to allow the model to learn to flexibly control how much each cell in memory is updated, as below:
Transformer Blocks with Memory: Lastly, we must alter the GPT transformer blocks to include the memory in the language modeling. We alter the attention used within the transformer blocks to contain two parallel attention modules, as shown in Figure 2. One attention module (on the left in the figure) performs the standard GPT self-attention using transformer inputs to create queries, keys, and values. The other attention module uses transformer input to attend over the memory vectors (i.e., using the memory for creating key and value vectors). The outputs of both attention modules are averagedWe experimented with a few other variants of implementing multiple attention mechanisms within the transformer blocks, but found this to be empirically effective. before performing the remaining transformer block operations.
5 Training and Decoding
At training time, the model is trained end-to-end on the cross-entropy loss of predicting each paragraph. Gold representations of previous paragraphs in the story are used to update the memory and compute . At decoding time, the model must decode a document starting with the first paragraph and use its own predictions to compute and update the memory. Additionally, at decoding time, we assume a five paragraph structure (introduction, three body paragraphs, and conclusion) as a pre-set discourse structure to decode from.
Experiments
We present experiments comparing PlotMachines with competitive baselines and ablations using automatic metrics and human judgements targeting multiple aspects of performance. In Sec. A.3, we also include example generations.
We compare with two models that have been used in related conditional story generation tasks. First, we train a Fusion model, from the original WritingPrompts dataset paper (Fan et al., 2018), using delimited outlines as a single input in place of a prompt. We also compare with the static storyline-to-story variant of Plan-and-Write (P&W-Static) from Yao et al. (2019), which is an LSTM-based model that we train by using the plot outline as delimited input.
Additionally, given the recent successes in text generation using large pre-trained LM’s, we compare with these models, as well. We finetune the large-scale Grover (Zellers et al., 2019) (equivalent to GPT-2 medium, 345M param) , which is a transformer-based language model that has been pre-trained for controllable text generation. To finetune Grover, we give the outline as a delimited form of metadata. Grover (345M param) has significantly larger capacity than PlotMachines (160M param). Therefore, for more direct comparison, we also investigate a 460M parameter version of PlotMachines that is built on top of GPT-2 medium (Radford et al., 2019) instead of GPT.
Unlike our models, the baselines are trained with the traditional generation framework, to generate an entire document conditioned on outlines without generating each paragraph recurrently.
We also show results in Table 2 on ablated versions of our model. First, we use the base GPT and GPT2 models, that are fine-tuned similarly to our model but using only outline inputs (without memory, preceding context, or discourse representations). Second, we investigate the effects of using the preceding context representation but still excluding memory and discourse tokens (PM-NoMem-NoDisc). Lastly, we use PM-NoMem, a model variant that excludes the memory but uses outline, discourse, and preceding context representations as input.
We use the HuggingFace implementations of GPT and GPT-2, and we fine-tune using ADAM. For generating with our models, we use nucleus sampling with repetition penalties (Holtzman et al., 2019; Keskar et al., 2019) using and for GPT and and for GPT-2 (based on a hyperparameter sweep using grid search with cross-entropy on the dev. data). We use a minimum sequence length of 100 bpe tokens per paragraph and a maximum sequence length of 400, 922 bpe per paragraph for GPT and GPT-2, respectively. We set , the maximum number of outline tokens and memory dimensions to 100. We used the settings for the baselines from their respective papers and codebases.
2 Automatic Metrics
In this section, we evaluate performance using different automatic metrics. We compute ROUGE scores (Lin, 2004) and self-BLEU (Zhu et al., 2018) following from previous work (Shen et al., 2019; Zhu et al., 2018) showing that a large ROUGE score together with a low self-BLEU score can demonstrate a model’s ability to generate realistic-looking as well as diverse generations.
We compute ROUGE scores (Lin, 2004) with respect to the gold stories (Table 2). Results show that the full PlotMachines achieves comparable or higher ROUGE on all three datasets. Both PlotMachines variants (using GPT or GPT-2 as a base) achieve improvements over Grover, even though Grover includes significantly more parameters than the model using GPT.
In the bottom block of Table 2, we compare performance of ablated versions of PlotMachines. First, we compare GPT-2 with PM-NoMem-NoDisc, which differs by including preceding context representations. We observe that PM-NoMem-NoDisc performs slightly better than GPT-2, emphasizing the importance of including context from the previous paragraph. Second, we investigate the impact of discourse structure representations. We compare PM-NoMem-NoDisc, which omits the discourse token, with PM-NoMem, which uses the discourse token. As shown in Table 2, PM-NoMem generally has higher ROUGE scores than PM-NoMem-NoDisc, indicating that the discourse representation is beneficial to the model. Lastly, we compare PM-NoMem with the full PlotMachines to determine the effects of having a memory component. Our full model with memory has large ROUGE score improvements over PM-NoMem, underscoring the importance of the plot state tracking.
We evaluate the diversity of generated paragraphs from our models using self-BLEU scores (Zhu et al., 2018). In Table 3, we report the self-BLEU scores along with the average length of each generated story. Using all the generated documents from a model, we take one generated document as hypothesis and the others as reference, and calculate BLEU score for every generated document, and define the average BLEU score to be the self-BLEU of the model. While the Fusion model achieved relatively high ROUGE scores, it has generally worse diversity scores (much higher self-BLEU in Table 3). It may be that this model’s high ROUGE scores were obtained by producing text that is more repetitive and generic.We show an example output from the Fusion model in Figure 13. In contrast, PlotMachines generally achieves good performance on both ROUGE and diversity scores, with self-BLEU scores that are lower than most other models. Notably, they generally have more similar self-BLEU scores to the actual gold stories, indicating that the language diversity is more similar to what humans write.
3 Human Evaluations
Due to the limitations of automatic metrics, we also perform extensive human evaluations. We conduct human studies to explore how generated stories compare along three dimensions: outline utilization, narrative flow, and ordering. We ask human ratersusing Amazon Mechanical Turk. We include screenshots from the tasks Appendix A.2. In total, over 700 humans participated in all of our studies. to evaluate each single or pair of generations from the Wikiplots test set, collecting 3 responses per story.For sampling from Fusion, we restrict to stories or excerpts with two or fewer unk tokens to ensure legibility for workers.
To scale up the crowdsourcing work pragmatically, we split evaluations into two studies: one small-scale study evaluating full-length stories, and one large-scale study evaluating single-paragraph excerpts. In the first study, humans perform head-to-head ratings of 20 randomly sampled stories per pair of models. In the second study, humans rate story excerpts from 100 randomly sampled outputs per model.
We give human raters a pair of stories generated from the same outlines and ask them to choose which one is better in different aspects related to outline utilization, narrative flow, and ordering. In Figure 3, we show how often PlotMachines (PM) was selected over the other models (values above 50% indicate that PM was preferred) using the majority vote for each example. PM was selected over base GPT and Fusion in all of the categories, demonstrating that the memory and discourse features are vitally important to improving the base model. While humans rated Grover as using the outline more, PM is ranked higher in all of the questions about narrative flow and ordering.
3.2 Excerpt Ratings
We give raters two paragraphs each generated by different models and ask them to select which is utilizing the outline better. We perform two trials, one with random paragraphs from each story and one with the paragraph from each story that has the most n-gram overlap with the outline (i.e. the closest). In both cases, we compute the majority vote over the three responses and report the percentage of examples where our model is preferred. Results in Table 4 show that, when looking at single paragraphs, humans tend to choose our PM as using the outlines in a more natural way, particularly when looking at the “closest” paragraph from both models. Fusion and GPT, in particular, are judged to be utilizing the outline much less than PlotMachines.
In this task, we give raters a generated paragraph (with the previous paragraph as context). They are asked to rate on a scale from 1 to 5 how much the paragraph: (a) repeats content from the previous paragraph, (b) transitions naturally from the previous paragraph, and (c) stays relevant and on-topic throughout the paragraph.
In the left side of Table 5, we show the average ratings of each model. GPT is the least repetitive between paragraphs but has very low subscores for transitions and relevance. We posit that this behavior is likely due to GPT often generating unrelated content from one paragraph to the next. PM tends to have the highest rated transitions and achieve highest relevancy within paragraphs while being much less repetitive between paragraphs than Grover or Fusion.
It’s challenging for humans to directly rate the ordering of a story based on a short excerpt. We instead set up a related proxy task: we give raters a pair of consecutive generated paragraphs, presented in a random order, and ask them to attempt to decipher the order. The intuition is that if the model output is very well-structured then it should be easier for humans to decipher the order. We compute the accuracy of the majority vote compared to the actual order in the right side of Table 5. Accuracy for PM approaches 60% accuracy and is much better than the base GPT. Grover and Fusion are easiest for humans to re-order (62%, 73% respectively). This result differs slightly from the full story analysis where the humans preferred PM over Grover and Fusion in the ordering-based questions. One possible explanation is that these two models, which decode word-by-word, without an explicit notion of paragraph, may be better at resolving coreference problems between paragraphs. This may make it easier for humans to re-order short excerpts even though they generally prefer the overall narrative order of PM due to it having better beginnings, endings, etc. (as indicated in our full story human study).
4 N-gram Based Outline Usage Analysis
We perform an additional quantitative study to further investigate how outline points are used in generated stories. For fifty stories in the Wikiplots dev. set, we compute how many paragraphs mention each outline point using exact matching or partial matching ( of the n-grams in the outline point also appear in the paragraph). We report the results in Figure 4.
We observe that Grover tends to over-repeat outline points (about twice as much as the gold story). This mirrors our human evaluations that Grover is more repetitive. This may also explain why human raters in the full story ratings in Sec. 4.3.1 judged Grover as using the outline more but having worse narrative flow and order. Similar observations have been made about pre-trained language models in See et al. (2019) that the models followed story prompts very closely but often copied too much compared to human writing.
In contrast, the Fusion model tends to leave out portions of the outline. This may reflect the way Fusion was originally designed – for use with a task using more abstract prompts as input. The GPT and PM-NoMem models, while more inclusive than Fusion, are also likely to exclude outline points. The full PM model is generally more inclusive and more similar to the gold reference than the other models. The gold story mentions each outline point in around one paragraph on average, indicating that there is an ideal balance between the more conservative coverage achieved by our model and the over-repetitive coverage of Grover.
5 Qualitative Examples
In the Appendix, (Sec. A.3), we include examples of model outputs on the validation set with annotations for incorporated outline points. Examples indicate that Grover often finishes the story and then starts a new story partway through the document. This may help explain why Grover over-repeats outline points and why humans judge it to be more repetitive and less consistently relevant. In contrast, our model adheres more to a beginning-middle-ending structure.
We also look at examples of introduction and conclusion paragraphs generated by PlotMachines, investigating the discourse the model has learned (Table 11). The model often starts stories by setting the scene (e.g. “In the early 1950s, a nuclear weapons testing continues ….”) and tends to write conclusions with a definitive closing action (e.g. “… the film ends with humperdinck and buttercup riding off into the sunset.”)
Related Work
There is a plethora of work in state tracking for dialogue where memory states are updated after each utterance (Thomson and Young, 2010; Young et al., 2010; Lee, 2013; Chao and Lane, 2019). Similarly, SC-LSTMs (Wen et al., 2015) dynamically updated dialogue act representations as a form of sentence planning in spoken dialogue generation. Memory and entity networks Henaff et al. (2017); Sukhbaatar et al. (2015) and neural checklists Kiddon et al. (2016) also used similar methods for tracking entities for other tasks. We adapt these techniques for generating stories while tracking plot state that is updated after each paragraph. Our method of decoding paragraphs recurrently also draws on existing work in hierarchical decoding (Li et al., 2015; Shen et al., 2019), which similarly decodes in multiple levels of abstraction over paragraphs, sentences, and words.
There has been a variety of work focusing on generating stories in plot-controllable, plan-driven, or constrained ways (e.g. (Riedl and Young, 2010; Fan et al., 2018; Peng et al., 2018; Jain et al., 2017; Lebowitz, 1987; Ippolito et al., 2019; Pérez y Pérez and Sharples, 2001)). Similar work in creative generation has conditioned on keywords for poetry generation Yan (2016); Ghazvininejad et al. (2016); Wang et al. (2016). Outline-conditioned generation is complementary to these tasks in that outlines provide more flexibility than very fine-grained srl-based, event-based, or graph-based plans (Fan et al., 2019; Martin et al., 2017; Harrison et al., 2017; Li et al., 2013) and more structured grounding than coarse-grained prompts (Fan et al., 2018; Xu et al., 2018) or ending goals (Tambwekar et al., 2019). Another similar task generates five line stories from five keywords (Peng et al., 2018; Yao et al., 2019). We generalize to a similar set-up for long-form narratives. Similar to many recent works in this area, we use seq2seq-based approaches, implemented using transformers. We further expand upon the modeling for the challenges specific to our task by using state tracking and applying discourse structure.
Conclusion
We present outline-conditioned story generation, a new task for generating stories from outlines representing key plot elements. We facilitate training by altering three datasets to include plot outlines as input for long story generation. In order to keep track of plot elements, we create PlotMachines which generates paragraphs using a high-level discourse structure and a dynamic plot memory keeping track of both the outline and story. Quantitative analysis shows that PlotMachines is effective in composing tighter narratives based on outlines compared to competitive baselines.
Acknowledgements
We would like to thank anonymous reveiwers for their insightful feedback. We also thank Rowan Zellers and Ari Holtzman for their input on finetuning Grover and other language models, Maarten Sap for his feedback on human evaluations, and Elizabeth Clark for consulting on baselines and related work. We would also like to thank various members of the MSR AI and UW NLP communities who provided feedback on various other aspects of this work. This research was supported in part by DARPA under the CwC program through the ARO (W911NF-15-1-0543), DARPA under the MCS program through NIWC Pacific (N66001-19-2-4031), and the National Science Foundation Graduate Research Fellowship Program under Grant No. DGE-1256082.
References
Appendix A Supplementary Materials
We show full stories in Tables 6-8 corresponding to the excerpts shown in the Dataset sub-section of Outline-Conditioned Generation in the main text.
A.2 Human Evaluation Details
In Figures 5-8, we show the questionaires we asked the human raters. In question 2 of the full story task, we asked about which story was more repetitive, but we flip their answers in Figure 3 to show the model that was less repetitive in the Figure (i.e. for ease of reading, we made higher better as with the other metrics).
A.3 Qualitative Examples
In this section, we include examples of model outputs on the validation set with annotations for incorporated outline points.
We show example full stories from the Wikiplots validation set comparing outputs from:
Grover (Figure 9) and PlotMachines (Figure 10)
Grover (Figure 11) and PlotMachines (Figure 12)
Fusion Fan et al. (2018) (Figure 13) and PlotMachines (Figure 14)
In the examples, we highlight outline points that are mentioned in red. We also bold a few sections in the Grover output where the model notably ends the story and starts a new one. Examples indicate that Grover often finishes the story and then starts a new story partway through the document. This shortcoming may help explain why Grover over-repeats outline points and why humans judge it to be more repetitive and less consistently relevant. In contrast, our models adhere more to a beginning-middle-ending structure.
We also show additional examples of introduction and conclusion paragraphs generated by PlotMachines (Table 11), demonstrating the discourse the model has learned. For example, the model often starts stories by setting the scene (e.g. “In the early 1950s, a nuclear weapons testing continues ….”) and often ends with a definitive closing action (e.g. “… the film ends with humperdinck and buttercup riding off into the sunset.”)