Make-A-Story: Visual Memory Conditioned Consistent Story Generation

Tanzila Rahman, Hsin-Ying Lee, Jian Ren, Sergey Tulyakov, Shweta Mahajan, Leonid Sigal

Introduction

Multimodal deep learning approaches have pushed the quality and the breadth of conditional generation tasks such as image captioning and text-to-image synthesis . Owing to the technical leaps made in generative models, such as generative adversarial networks (GANs) , variational autoencoders (VAEs) and the more recent diffusion models , approaches for text-to-image synthesis can now generate images with high visual fidelity representative of the textual descriptions. The captions, however, in such cases, are generally short self-contained sentences representing the high-level semantics of a scene. This is rather restrictive in the real-world applications where fine-grained understanding of object interactions, motion and background information described by multiple sentences becomes necessary. One such task is that of story generation or visualization – the goal of which is to generate a sequence of illustrative image frames with coherent semantics given a sequence of sentences .

Characteristic features of a good visual story is high visual quality over multiple frames; this includes rendering of discernible objects, actors, poses and realistic interactions of those actors with objects and within the scene. Moreover, for text-based story generation it is crucial to maintain consistency between the generated frames and the multi-sentence descriptions. Not only the actor context, but also the background of the generated story should be in-line with the description demonstrating effortless transition and adaptation to the changing environments within the story .

Recent advances on the task of story generation have made significant advances along these lines, showing high visual fidelity and character consistency for story sentences that are self-contained and unambiguous (explicitly mentioning characters and the setting each time). While impressive, this setup is fundamentally unrealistic. Realistic story text is considerably more complex and referential in nature; requiring ability to resolve ambiguity and references (or co-references) through reasoning. As shown in Fig. 1, while the description corresponding to the first frame has an explicit reference to the character names, the (typical) subsequent frame descriptions, provided by human, contain references such as “she, he, they”. Moreover, while maintaining character consistency, current approaches are limited in preserving, or transitioning through, the background information in agreement with the text (cf. Fig. 3) .

In natural language processing (NLP) co-reference resolution in text is an important and core task . While it maybe possible to apply such methods to story text to first resolve ambiguous references and then generate corresponding images using existing story generation approaches, this is sub-optimal. The reason, is that co-reference resolution in the text domain, at best, would only allow to resolve references and maintain consistency across identity of the character. Appearance across frames would still lack consistency and require some form of visual reasoning. As also noted in , reference resolution in the visual domain, or visio-lingual domain, is more powerful.

In this work, for the first time (to our knowledge), we study co-reference resolution in story generation. Prior work offers limited performance when faced with text containing references (see Sec. 5). We address this by proposing a new autoregressive diffusion-based framework with a visual memory module that implicitly captures the actor and background context across the generated frames. Sentence-conditioned soft attention over the memories enables effective visio-lingual co-reference resolution and learns to maintain scene and actor consistency when needed. Further, given the lack of datasets that contain references and more complex sentence structure, we extend the MUGEN dataset and introduce additional characters, backgrounds and referencing in multi-sentence storylines.

Contributions. Our contributions are three-fold: (i) First, we introduce a novel autoregressive deep generative framework, Story-LDM, that adopts and extends latent diffusion models for the task of story generation. As part of Story-LDM, we propose a meticulously designed memory-attention mechanism capable of encoding and leveraging contextual relevance between the part of the story-line that has already been generated, and the current frame being generated based on learned semantic similarity of corresponding sentences. Equipped with this, our sequential diffusion model can generate consistent stories by resolving and then capturing temporal character and background context. (ii) Second, to validate our approach for co-reference resolution, and character and background consistency in the visual domain, we extend existing datasets to include more complex scenarios and, importantly, referential text. Specifically, we extend the MUGEN dataset to include multiple characters and diverse backgrounds. We also modify FlintstonesSV and PororoSV dataset to include character references. These enhancements allow us to increase the complexity of the aforementioned datasets by introducing co-references in the sentences of a story. (iii) Finally, to evaluate different approaches for foreground (character) as well as background consistency we propose novel evaluation metrics. Our results on the MUGEN , the PororoSV and the FlintstonesSV datasets show that we outperform the prior state-of-the-art on consistency metrics by a large margin.

Related work

Text-to-image synthesis. Deep generative models, particularly, generative adversarial networks (GANs) , variational autoencoders (VAEs) and normalizing flows have been applied to multimodal tasks at the intersection of vision and language. Typical such tasks include image captioning and text-to-image synthesis . Early work on text conditioned image synthesis built upon the success of GANs . More recent approaches have utilized multi-stage generators and normalizing flow-based priors in the latent space to model the distribution of images given text. Various approaches have found cross-domain contrastive loss to improve text-to-image generation models . DALL-E and Cogview harness the power of transformers and discrete variational autoencoders (VQ-VAE) yielding very high quality image samples.

More recent are the advances in diffusion models which have revolutionized the domain of image generation . Diffusion models progressively add noise to the data and learn a reverse diffusion process to reconstruct it. Nichol et al. adapted diffusion models for text-to-image generation and explore CLIP guided generation as well as classifier-free modeling. Standard diffusion models are employed directly in the high-dimensional pixel space and therefore, cannot directly be used for the more complex task of story generation. Recent work instead use encodings from pre-trained models as input to the diffusion models, thereby reducing the complexity of the task by working in a lower-dimensional space. In this work, we build upon this idea and extend it for sequential story generation.

Text-to-video synthesis. One of the challenges of text-to-video synthesis is the smoothness of motion in a video . Early work focused on generating short clips . To effectively learn the motion, various approaches disentangle the motion features from the background information . Wu et al. propose a novel two-dimensional VQ-VAE and sparse attention module for real-world text-to-video generation. Singer et al. decompose the temporal U-Net and the attention modules to approximate them in space and time to extend the text-to-image diffusion models to model text-to-video generation. Ligong et al. propose a transformer framework to jointly model various modalities.

Story Generation. Li et al. proposed the initial idea and task of story generation. A two-level StoryGAN framework is applied to ensure image-level consistency between each sentence and image pair, and a global discriminator enforces global consistency between the entire image sequence and the story. Various approaches have proposed improvements to the StoryGAN architecture. Zeng et al. introduce sentence-level alignment and word-based attention to improve relevance. Li et al. further improve the performance with enhanced discriminators and dilated convolutions. In foreground-background information is provided as additional supervision and use video captioning for semantic alignment between text and frames. Recently, Chen et al. adopted visual planning and character token alignment to improve character consistency.

Story Completion. Recently, another task for text-to-story synthesis referred to as story completion has been proposed . In this task, in addition to sentences, the first frame of the story is provided as input. In effect, story completion is a simplified variant of story generation. StoryDALL-E leverages models pre-trained for text-to-image synthesis to perform story completion. Datasets for this task include CLEVR-SV and Pororo-SV which are derived from the CLEVR dataset , and the Flintstones dataset for text-to-video synthesis has also been modified for the task of story visualization . Additionally, to evaluate the generalization performance, popular DiDeMo dataset for video captioning is adapted for the task in .

Reference Resolution. Co-reference resolution is an important and well-researched topic in NLP and focuses on resolving the pronouns and their associated entities. Classic methods in NLP to co-reference resolution employ decision trees , maximum-entropy modeling , cluster-ranking and classification algorithms . More recent approaches leverage neural network architectures to obtain improved performance with Transformers . Seo et al. proposed visual co-reference resolution for the task of Visual Question-Answering (VQA) dialogs. We take inspiration from , but propose a much more sophisticated memory-attention module that allows us to perform visio-lingual co-reference resolution (and visual consistency modeling) for visual story generation.

Approach

To generate temporally consistent stories based solely on the linguistic story-line, we develop a deep generative approach with autoregressive structure. We build upon the success of diffusion models in modeling the underlying data distribution of images to produce high quality samples, and learn the generative conditional distribution of the visual story based on the textual descriptions. Given that the multi-frame stories involve high-dimensional data input, we employ Latent Diffusion Models , such that diffusion models can be applied in a computationally efficient manner. Besides, to ensure temporal consistency and smooth story progression, we propose a novel memory attention mechanism which not only attends to the multimodal representations of the current frame but also takes into account the already generated semantics of the previous frames. This module also allows us to resolve ambiguous references (e.g., he/she, they, etc.) using visual memory and is the core of our technical contribution. We first provide an overview of the Diffusion Models and the Latent Diffusion Models, following which we present our autoregressive latent diffusion model for stories called Story-LDMhttps://github.com/ubc-vision/Make-A-Story.

Diffusion Models. Diffusion models are a class of generative models that approximate the underlying data distribution p(x)p(\mathbf{x}) by denoising a base (Gaussian) distribution in multiple steps using a reverse process of a fixed Markov Chain of length TT. To estimate p(x)p(\mathbf{x}), the forward diffusion process starts from the input data x0=x\mathbf{x}_{0}=\mathbf{x} and gradually adds noise to obtain a set of noisy samples x1,…,xT{\mathbf{x}_{1},\dots,\mathbf{x}_{T}} such that xT∼N(0,1)\mathbf{x}_{T}\sim\mathcal{N}(0,1) represents a sample from a Gaussian distribution. Under the Markov assumption, the probability of the forward process modeling the distribution q(x0:T∣x0)q(\mathbf{x}_{0:T}\mid\mathbf{x}_{0}) and the reverse diffusion process estimating probability at an earlier time-step are formulated as:

Here, {βi}i=1T\{\beta_{i}\}_{i=1}^{T} is the variance schedule for each time-step such that xT\mathbf{x}_{T} is nearly a Gaussian. The model parameters θ\theta are learnt with the following objective,

where ϵ∼N(0,1)\epsilon\sim\mathcal{N}(0,1) and ϵθ(xt,t)\epsilon_{\theta}(\mathbf{x}_{t},t), t=1,…,Tt=1,\ldots,T is a sequence of denoising autoencoders with noisy input xt\mathbf{x}_{t} predicting the noise that was added to the original input x\mathbf{x}.

Despite yielding state-of-the-art results in various image generation tasks, diffusion models directly operating in the high-dimensional pixel are computationally expensive and resource exhaustive. This limits their application to an even higher-dimensional data such as multi-frame stories or video datasets, which is the focus of this work.

During training, a forward diffusion process is applied to generate Z\mathbf{Z}, which are mapped to the original image space using a decoder D(⋅)D(\cdot).

2 Story-Latent Diffusion Models

Given a textural story, characterized by sequence of MM sentences Stxt={S0,…,SM}\mathbf{S}_{txt}=\{\mathbf{S}^{0},\ldots,\mathbf{S}^{M}\}, the goal of story generation is to produce a sequence of corresponding frames Simg={I0,…,IM}\mathbf{S}_{img}=\{\mathbf{I}^{0},\ldots,\mathbf{I}^{M}\} that visualize the story. We note that this is a more difficult problem than one of story continuation , where in addition to the textual story Stxt\mathbf{S}_{txt} approaches have access to a source frame I0\mathbf{I}^{0} for additional context at inference time. During training it is assumed that we have access to paired dataset of NN samples D={Stxt(i),Simg(i)}i=1N\mathcal{D}=\{\mathbf{S}^{(i)}_{txt},\mathbf{S}^{(i)}_{img}\}_{i=1}^{N}. We extend latent diffusion models to this task, by allowing them to generate multi-frame stories autoregressively, and by introducing rich conditional structure that takes into account current sentence as well as context from earlier generated frames through visual memory module. This visual memory allows the model to incorporate character/background consistency and resolve text references when needed, resulting in improved performance.

Given a condition y\mathbf{y}, LDM utilizes a cross-attention layer with key (K)(\mathbf{K}), query (Q)(\mathbf{Q}) and value (V)(\mathbf{V}) where,

Note that the denoising autoencoders ϵθ\epsilon_{\theta} now additionally depend on the condition encoding f(y)f(\mathbf{y}).

For sequential generation, the model in addition to the current state, requires information from all the previous states. To enable this, the diffusion process for any frame representation Zm\mathbf{Z}^{m} is conditioned on the visual representations of the previous frames Z0,…,Zm−1\mathbf{Z}^{0},\ldots,\mathbf{Z}^{m-1} as well as the sentence descriptions S0,…,Sm\mathbf{S}^{0},\ldots,\mathbf{S}^{m}. This conditioning is realized through a novel Memory-attention module which forms the basis of our autoregressive approach.

To capture the spatio-temporal interactions across multiple frames and sentences for a story, in our conditional diffusion model, we condition the frame Zm\mathbf{Z}^{m} not only on the corresponding text Sm\mathbf{S}^{m} but also on the previous texts Si\mathbf{S}^{i}, for i∈{0,…m−1}i\in\{0,\ldots m-1\}. This conditioning is applied throughout the TT time-steps of the diffusion process for Zm\mathbf{Z}^{m}. The conditional denoising autoencoder thus models the conditional distribution p(Zm∣Z<m,S≤m)p(\mathbf{Z}^{m}\mid\mathbf{Z}^{<m},\mathbf{S}^{\leq m}). The (conditional) generative process of our Story-LDM approach over the TT steps of the diffusion process for a single frame is thus given by

The key motivation for this approach is to propagate the semantic (visual and textual) features from the already processed story-line based on the relevance of the current description to the previous frames as well as the previous descriptions. We achieve this by implementing a special attention layer, called memory-attention module. Similar to the cross-attention layer, we utilize the attention mechanism based on the key, query and value formulation. In this case Eq. 4 becomes,

where f^(.)\hat{f}(.) is applied to align the dimensions of the values, V\mathbf{V} with the keys, K\mathbf{K}. In the memory attention module, the relevance the query Q\mathbf{Q} which depends on the current sentence Sm\mathbf{S}^{m} and the keys K\mathbf{K} which represent the previous sentences S<m\mathbf{S}^{<m} is used to weight the feature representations Z<m\mathbf{Z}^{<m}. The aggregated representation now contains the information relevant for the current frame Zm\mathbf{Z}^{m} from the already generated story-line (see Fig. 2(b)). That is, our mechanism based on the similarity of the current sentence to the previous sentences in the story, identifies the features in the previous frames which are of importance to the context of the current frame. This may include recurrence of certain semantics with-in the story such as characters or backgrounds. In all, this formulation of the diffusion process allows us to maintain temporal consistency as we amplify the visual feature information from the sequence of story already generated. This allows the model to implicitly capture temporal dependencies in storylines for resolving ambiguities in character and background information.

Given the above conditioning, the objective for the story latent diffusion model for a single frame is formalized as

where ϵθm\epsilon_{\theta^{m}} are the denoising autoencoders for the frame mm.

Having formalized the diffusion process for single frame generation, the generative process for the entire story-line using Eq. 6 for autoregressive conditional frame generation, is given by

Notably, the conditioning is applied to the all states within the diffusion process i.e., for all Ztm\mathbf{Z}^{m}_{t}, t∈{1,…T}t\in\{1,\ldots T\} at each diffusion step, we apply the cross attention as well as the memory attention module allowing us effectively capture the temporal context.

Network Architecture.

To generate visual storylines, in our Story-LDM, we first introduce an autoregressive structure and modify the two-dimensional U-Net in to so as to process the temporal information in the storyline. As shown in Fig. 2(a), the frame encoder, EE equipped with the positional information of the frame in the sequence, is applied to get the low-dimensional representation Zm\mathbf{Z}^{m} for all the frames in a datapoint from D\mathcal{D}. A text-based transformer is applied to get a suitable representation for the sentence Sm\mathbf{S}^{m}. The U-Net is then applied to model the diffusion process over TT time-steps. The layers within the U-Net are augmented with the cross-attention layer and our memory attention layer. After each downsampling or upsampling operation we apply the attention mechanism to reinforce the conditioning on the already encoded (learned) story-line up to the previous time-step. For any frame mm, the cross-attention CattnC_{attn} is given by

where f^(Zm)\hat{f}(\mathbf{Z}^{m}) and f(Sm)f(\mathbf{S}^{m}) are the representations of the frame encoding Zm\mathbf{Z}^{m} within the neural network and sentence Sm\mathbf{S}^{m} respectively such that they have same dimensions. Similarly, the memory attention MattnM_{attn} is computed as,

The output of the attention-module is then computed as the aggregation, Cattn+MattnC_{attn}+M_{attn}.

Starting from the noise sample ZTm\mathbf{Z}^{m}_{T} the output of the reverse diffusion process Z0m\mathbf{Z}^{m}_{0} is reconstructed using the frame decoder DD to get the final image. Having outlined the details of our Story-LDM framework, we show through extensive experiments on the task on story generation, the effectiveness and the benefits of our powerful conditioning based on memory-attention.

Datasets and Evaluation Metrics

In this paper, we formulate story generation with co-references to actors and backgrounds across frames.

Datasets. Since reference resolution has not been studied in story generation, to validate our approach on this much harder task, we construct the following datasets: (i) We take an existing story-generation dataset – FlintstonesSV , and modify the sentences by replacing the named entities (characters) with references where possible; including pronouns such as he, she, or they (cf. Fig. 1). This dataset contains 2013220132-training, 20712071-validation and 23092309-test stories with 77 main characters and 323323 backgrounds. (ii) MUGEN is a video dataset collected from the open-sourced platform game CoinRun . The dataset is divided into 104,796104,796-train and 11,80211,802 test stories with 96 to 602 frames. We extend the MUGEN dataset by introducing two additional characters Lisa and Jhon (we rename Mugen to Tony). We construct stories of four frames and corresponding text, ensuring consistent co-referencing in the story; each story has 3 such references. Moreover, we augment the existing two backgrounds (Planet and Snow) with four additional backgrounds: Sand, Dirt, Grass and Stone. (iii) We also modify existing PororoSV dataset which contains 10191/2334/2208 train/val/test set. Similarly, we reference characters by pro-nouns to generate more natural story. We show in Fig. 1, example stories from the two modified datasets and enlist the complete statistics in Tab. 1.

Evaluation Metrics. To measure the consistency of the characters as well as the backgrounds in the generated stories, we consider following evaluation metrics: (i) Character Classification: Following , we consider fine-tuned Inception-v3 to measure the classification accuracy and F1-score. Frame accuracy evaluates the character match to the ground-truth and F1-score measures the quality of generated characters in the predicted images. (ii) Background Classification: Similar to character classification, we use fine-tuned Inception-v3 to measure the correspondence of the background to the ground-truth and consider F1-score as a measure of quality. (iii) Frechet Inception Distance (FID): To assess the quality of images,we consider FID score which is the distance between feature vectors from real and generated images.

Experiments

In this section, we evaluate our Story-LDM approach for consistent story generation with reference resolution.

Baselines. We construct a strong baseline with the LDMhttps://github.com/CompVis/latent-diffusion which contains a cross-attention layer to generate text-to-image based story, without using our proposed autoregressive memory modules as our baselines for MUGEN, PororoSV and FlintstonesSV datasets. The parameters of the diffusion model within the Story-LDM are initialized with the pre-trained LDM . Similarly, for the textual embedding, we use BERT-tokenizer and use the pre-trained text-transformer from LDM.

Quantitative Results. Table 2 shows quantitative results for consistent story generation on the FlintstoneSV dataset. We compare the performance of our approach (row 4) to the LDM which we train/test with both original (row 2) as well as the co-referenced (row 3) descriptions. Furthermore, we include the results of the state-of-art VLCStoryGAN (row 1) with the original text of the datasetResults for were obtained using pretrained model provided by original authors in private communication. (i.e. without co-references). We note that VLCStoryGAN was shown to be better than Duco-StoryGAN , CP-CSV and original StoryGAN (see ).

Based on Table 2 we make three observations: (1) Our LDM baseline is better than VLCStoryGAN on the original reference-free text (cf. Tab. 2, rows 1 & 2). (2) Reference resolution makes the task considerably harder. With the reference text in our modified dataset, we observe a drop in performance in terms of character and background classification scores (cf. Tab. 2, rows 2 & 3). (3) Our model, with memory-attention module, significantly outperforms the baseline (cf. Tab. 2, rows 3 & 4) both in terms of generative image quality and character consistency; and outperforms SoTA of VLCStoryGAN by ∼41%\sim 41\% percentage points on character accurary (while performing a more difficult version of the task). Further, our model, that is required to conduct reference resolution, comes close to the LDM trained with original, reference-free, text (cf. Tab. 2, rows 3 & 4), which can be viewed as a sort of an upper bound.

On the MUGEN dataset, our method outperforms the strong LDM baseline with gains of ∼62%\sim 62\% on character accuracy and ∼76%\sim 76\% on the background accuracy, thereby showing the advantages of the memory-attention mechanism for consistent story generation. We note that MUGEN dataset has more references across story scenes. Flintstones while contains more references per story overall, many of those references are within scenes as opposed to across scenes. Meaning that in terms of reference impact on consistency, MUGEN dataset is actually harder. Experimental results on the PororoSV dataset are provided in the Supplemental.

Qualitative Results. Figure 3 illustrates qualitative results on the MUGEN dataset. Rows 1, 2 & 3 show ground truth, LDM and our Story-LDM approach, respectively. Here, we see that our method is able to maintain consistency in terms of both character and background. Similarly, in Figure 4 we can show the results on FlintstoneSV dataset which further validates the strong performance of our method when generating high-quality, consistent story. Compared to the LDM, our approach is able to adapt to the diverse backgrounds in the story descriptions.

Additional Results. We compare the qualitative results of our method to both story generation and story continuation in Figs. 5 and 6 respectively. The comparative images are taken directly from respective papers. We note that story continuation Fig. 6 is solving a different (easier) problem and with text that contains no-references. This makes the comparison to our method, which receives fewer inputs, not very meaningful. Nether-the-less, our approach, that can resolve references and is solving a harder story generation task, obtains highly competitive results. Furthermore, to show that our autoregressive visual memory module can generate diverse stories conditioned on the current and previous condition, we create different story-lines starting for a single sentence. In Fig. 8, we can see for reference ‘they’, the model can generate both the characters according to the storyline already parsed. Moreover, in Fig. 7 we show that our approach can not only generate consistent visual stories, but also diverse frames for the same text (cf. Fig. 7). Additional results are provided in the Supplemental.

Conclusion

In this paper, we formulate consistent story generation in a more realistic way by co-referencing actors/backgrounds in the story descriptions. We develop an autoregressive Story-LDM approach with memory attention capable of maintaining consistency across the frames based on the previously generated frames and their corresponding descriptions. We introduced modified datasets to evaluate the performance for reference resolution. We expect our proposed formulation and models to be conductive to the real-world use cases and further the research.

References