Paint by Word

Alex Andonian, Sabrina Osmany, Audrey Cui, YeonHwan Park, Ali Jahanian, Antonio Torralba, David Bau

Introduction

A writer can create vivid verbal imagery using just a few carefully-chosen words, but written scenes are abstract, without any specific physical form. In this paper, we ask how to enable an artist to use words to build an image concretely, painting textually described visual concepts at specific locations in a scene. Our work is inspired and enabled by recent dramatic progress in text-to-image generation . However, rather than having the AI compose the whole image, we ask how large models can be used collaboratively to let a person apply words as a paintbrush. We wish to enable a user to paint brushstrokes on specific objects and regions within a generated image, then restyle, alter, or insert new objects by describing the desired visual concept with words.

The current work introduces the problem of zero-shot semantic image manipulation, in which the words and painting gestures available to the user are unconstrained and unknown ahead of time during model training. Our goal is to enable a user to point at an image and apply an arbitrary new concept that may include visual concepts as wide-ranging as “rustic” or “opulent” or “happy dog.”

To paint with words, we pair large-scale generative adversarial networks (GANs ) that are trained unconditionally or with simple class conditioning , together with a full-text image retrieval network trained to measure the semantic similarity between text and an image . We ask whether such a semantic similarity model can provide enough information to drive the image synthesis process to match a specific description within or beyond the domain on which the GAN is trained. Then we ask how the structure of a GAN can be exploited to generate realistic images where a specific part of the image is modified to match a textual description. We propose a simple architecture to enable painting with words, and we apply it using CLIP paired with StyleGAN2 and BigGAN .

We conduct user studies to quantify the realism and accuracy of the changes made when using our method to edit a part of a generated image; we compare our method to several baselines, and we ask which types of visual concepts are easier or harder to paint by word. Finally we investigate several new semantic image manipulations that are enabled using our method. Code and data will be available upon publication.

Related Work

Our application is inspired by recent progress in the text-to-image synthesis problem as well as paint-based semantic image manipulation methods .

Text-to-image synthesis. Synthesizing an image based on a text description is an ambitious problem that has attracted a variety of proposed solutions: initial RNN-based generators have been followed a series of by GAN-based methods that have produced increasingly plausible images. The GAN methods have adopted two distinct approaches to generating realistic images from text. One is to generate images corresponding to text by sampling GAN latents that match semantics according to a separately trained image-text matching model ; this idea can be refined by viewing the sampling as an energy-based modeling procedure . The second approach is to train a conditional generator that explicitly takes a language embedding as input ; the image quality of this approach can be improved using multi-stage generators , and semantic consistency can be improved through careful design of architectures and training losses . As an alternative approach, generating images from scene graphs instead of sentences can allow a generator to exploit logical structure explicitly. Recent work stands apart by eschewing the use of GANs: the DALL-E model generates images from text autoregressively using a transformer based on GPT-3 to jointly generate natural text and image tokens using a VQ-VAE encoding ; outputs are sampled to maximize semantic consistency via a state-of-the-art image-text matching model CLIP . Artist Murdock has observed that CLIP can be used as a source of gradients to guide a generator . The current paper is different from prior text-to-image work, because our goal is not to generate an image from text, but to define a paintbrush using text that enables a user to manipulate a semantics of a generated image at a specific painted location.

Semantic image painting. Our work is also inspired by a series of methods that have enabled users to construct realistic images by painting a finite selection of chosen concepts at user-specified locations; these can be thought of as “paint by number.” In this setting, there have again been two approaches enabled by GANs. One approach is to find GAN latents to generate images that match a user’s painted intention: this can be done to match color or a finite vocabulary of visual concepts at a given location. The second approach is to train an image-conditional Pix2Pix model on the task of generating a realistic image with a semantic segmentation that matches a given painted input; the SPADE method refines this approach to improve image quality given the limited information available in flat segmentation inputs. While these methods all enable a user to paint images using a finite vocabulary of semantic concepts, the goal of the current work is to enable a user to paint an unlimited vocabulary of visual concepts, specified by free text.

GAN latent space methods. Semantic manipulations in GAN latent spaces have been explored from several other perspectives. Early work on GANs observed the presence of some latent vector directions corresponding to semantic concepts , and further explorations have developed methods to identify interesting vector directions that steer latent space by using reconstruction losses or by learning from classification losses . One disadvantage of this approach is that a semantic direction can only be found if we train a model to look for it in advance. On the other hand, it has been observed that interior latents in GANs are partially disentangled . So recent work has developed unsupervised methods for identifying disentangled interpretable directions in these interior latents . Our approach differs from the above because our method for identifying latent manipulations is supervised neither by a set of finite classes nor does it rely on disentanglement; rather it identifies semantic latent changes in a zero-shot manner using full-text image semantic similarity.

GAN inversion methods. We note that the leading state-of-the-art generative models are not trained with an encoder. Therefore, in order to apply GAN manipulation methods on real user-provided images, one must solve GAN inversion. That is, we must be able to encode a given image in the latent space of the GAN. Because several powerful editing methods are enabled by GAN latent manipulation, the problem of inverting a GAN is of widespread interest and is the subject of an active line of ongoing research . The GAN inversion problem is complementary to our work. In this paper, we shall assume that the image to be edited is represented in the latent space of the generator.

Method

There are two challenging problems that must be simultaneously solved in order to implement paint-by-word: first, we must achieve semantic consistency, altering the painted part of the image to match the text description given by the user; and second, we must achieve realism: the altered part of the image should have a plausible appearance. In particular, the newly painted content should be consistent with the context of the rest of the image in terms of pose, color, lighting, and style.

We adopt a framework that addresses these two concerns using two separately trained networks: the first is a semantic similarity network C(x,t)C(x,t) that scores the semantic consistency between an image xx and a text description tt. This network need not be concerned with the overall realism of the image. The second is a convolutional generative network G(z)G(z) that is trained to synthesize realistic images given a random zz; this network enforces realism. Given GG and CC, we can formulate the following optimization to generate a realistic image G(z∗)G(z^{*}) that matches descriptive text tt:

This simple approach factors the problem into two models that can be trained at large scale without any awareness of each other, and we shall use it as a starting point. It allows us to take advantage of recent progress in state-of-the-art models that can be used for GG and CC (Such as StyleGAN , BigGAN , CLIP , and ALIGN ). Because our method allows the direct use of such large-scale models, this simple architecture is a promising approach for generating images from a text description even without any spatial conditions. We study this approach empirically in Section 4.1.

When providing semantic paint, the method of Equation 1 is not sufficient, because it does not direct changes specifically towards a user’s chosen painted area. To focus effects on to one area, we direct the matching network CC to attend only to the region of the user’s brushstroke instead of the whole image. This can be done in a straightforward way: given a user-supplied mask mm, we define a masked semantic similarity model

Here x⊙mx\odot m is a projection that zeroes the components of xx outside the region of the mask mm.

By hiding the regions outside the mask from the semantic similarity model, the masked model Ct,m(x)C_{t,m}(x) focuses the optimization on the selected region. However, in practice we find that this is not enough to direct changes to only a single object. Figure 2(c) shows the effect of optimizing arg max⁡zCt,m(G(z))\operatorname*{arg\,max}_{z}C_{t,m}(G(z)) to match a text description in a StyleGAN2 model. Although no gradients pass from CC to GG outside the masked region, GG has a computational structure that links the appearance of objects outside the mask to objects within the mask.

In order to allow a user to edit a single object, we must obtain a generative model that allows the region inside of the mask to be unlinked. Interestingly, this can be done without training a new large-scale model from scratch.

2 Generative modeling in a region

To create a generative model that decouples the region inside and outside the mask, we exploit the convolutional structure of GG. We work with an intermediate latent representation of the image ww that has a spatial structure over the field of the image. To access this structure, we express GG as two steps:

Given a mask mm provided by the user’s brushstrokes, we then decompose ww into a portion outside the mask w0w_{0} and a portion inside the mask w1w_{1}, as follows:

Here w⊙mw\odot m is a projection that zeroes the components of ww outside the region of the mask mm. In practice, we downsample mm to the relevant feature map resolution(s) of ww, and use the Hadamard product to zero the outside components.

For a given original image G(z)G(z) we can fix the original w0=w−w⊙mw_{0}=w-w\odot m while allowing w1w_{1} to vary. This split gives us a new generative model, where the representation outside the mask mm is fixed as w0w_{0}, while the representation inside the mask is parameterized by ww:

Splitting w0w_{0} and ww in this way relaxes the problem, and allows Gz,m(w)G_{z,m}(w) to generate images that the original model G(z)G(z) could not generate.

In BigGAN, we use as ww the featuremap output of a layer of the network. This featuremap has a direct natural spatial structure, and w⊙mw\odot m simply zeros the featuremap locations outside the mask. In our BigGAN experiments we split the featuremap output of the first convolutional block of the generative network.

In StyleGAN, the natural interior latent, the ww vector, modulates featuremaps by changing channel normalizations uniformly across the spatial extent of the featuremap at all layers. In our formulation, we split the ww style latent vector spatially by applying a one style modulation outside the user-specified mask, fixed as w0w_{0}, and another style modulation ww inside the mask. In StyleGAN, the style modulation is applied across all layers, so we apply this split at the appropriate resolution at every layer of the model: the effect is to have a model that has two style vectors w0w_{0} and ww, instead of one.

This additional flexibility is instrumental for allowing the generator to create user-specified attributes in the region of interest without changing attributes of unrelated parts of the image. Figure 2(d) shows the effect of splitting the model: the appearance of one object can be changed without changing the appearance of other objects in the scene. Combining (2) with (7) allows us to define the following masked semantic consistency loss in the region

In order to explicitly limit changes in the synthesized image outside the painited region, we apply the following image consistency loss:

Here dd denotes an image similarity metric: in our experiments, for dd we use a sum of an L2 pixel difference and the LPIPS perceptual similarity. In Limg\mathcal{L}_{\text{img}}, the inverse of the user’s masked region is applied, so this loss term does not limit changes inside the painted region.

3 Avoiding adversarial attack using CMA

When using gradients to optimize a single image to a deep network objective, it is well-understood that it is easy to obtain an adversarial example that fools the network, achieving a strong score without being a typical representative of the distribution modeled by the network. For example, it is easy to obtain an image that is classified as one class while having the appearance of an unrelated class.

The problem of adversarial attack also arises in our application. Figure 3(a) plots the loss when optimizing Equation 1 using the Adam optimizer, where GG is a StyleGAN2 trained on LSUN birds , and where CC is the pretrained CLIP ViT-B/32 network . The convergence plot shows a rapid and stable-looking improvement as the image is modified to better match the phrase “A photo of a yellow bird flying through the air,” but the longer the optimization runs, the less like a flying bird the image becomes in practice. Figure 3(b) shows the image that results after several minutes of optimization: it achieves a strong score and yet has many artifacts that look obviously synthetic. It does not resemble a photo of a bird.

One way to think about the problem is that our networks CC and GG were trained for optimal expected behavior over a distribution, not worst-case behavior for every individual. So when optimizing a single instance it is not hard to find an individual point at which both models misbehave.

We solve this problem by switching optimization strategies. Instead of seeking out a single optimal image, we use the Covariance Matrix Adaptation evolution strategy (CMA-ES) , which is a non-gradient method that aims to optimize a Gaussian distribution to have minimal loss when random samples are drawn from the distribution. Although CMA-ES may not achieve individual point losses that are as low as a gradient method like Adam, we find that in practice, the images that it obtains match modeled semantics very well, while remaining more realistic than the images produced by gradient descent: Figure 3(c) shows a typical image from the distribution after running CMA optimization for an equivalent amount of time as Adam. In Section 4.1 we compare CMA-ES to Adam in a human evaluation.

Results

To understand the strengths and weaknesses of our method, we conduct user studies in two problem settings. Then we explore the types of new image manipulations that are enabled by paint-by-word.

Without a user-provided mask, our method is a simple factorization of the text-conditional image generation problem (Equation 1). We wish to understand if this straightforward split of the problem has disadvantages compared to other approaches that explicitly condition image generation, so we begin by evaluating this architecture in comparison to a state-of-the-art model.

For GG we train a 256-pixel StyleGAN2 with Adaptive Data Augmentation on the CUB birds data set . Standard training settings are used: the generator is unconditional, so no bird class labels nor text descriptions are revealed to the network during training. For CC we use an off-the-shelf CLIP model, using the pretrained ViT-B/32 weights; this model was trained on an extensive proprietary data set of 400 million image/caption pairs. We generate images using two different techniques to maximize C(G(z),t)C(G(z),t). As a simple baseline we optimize zz using first-order gradient descent (Adam ). Then we also optimize zz using CMA followed by Adam.

To evaluate our method, we generate images from 500 descriptions of birds from a test dataset also used for evaluation of DALL-E in . To compare the images generated by our method to images generated by DALL-E, we conduct a user study on Amazon Mechanical Turk in which workers are asked to choose which of a pair of images are more realistic, and in a separately asked question, which of a pair more closely resemble a given text description.

The results are shown in Figure 5. We find that, for this test setting, our method gives competitive results. The CMA optimization process is able to steer GG to generate an image that is found to be a better fit for the description than the DALL-E method 65.4% of the time, and more realistic 89% of the time. The CMA method is also better than Adam alone, more accurate in 66.2%; more realistic in 75.2%.

This experiment is not intended to show that our simple method is superior to DALL-E: it is not. Our network is trained only on birds; it utterly fails to draw any other type of subject. Because of this narrow focus, it is unsurprising that it might be better at drawing realistic bird images than the DALL-E model, which is trained on a far broader variety of unconstrained images. Nevertheless, this experiment demonstrates that it is possible to obtain state-of-the-art semantic consistency, at least within a narrow image domain, without explicitly training the generator to take information about the textual concept as input.

2 A large scale user study of bedroom image edits

Next, we conduct a large-scale user study on localized editing of objects in bedroom scenes. For this study, we sample 300 images generated by StyleGAN2 trained on LSUN bedrooms and manually paint a single bed in each image. Then for each image we perform 10 paint-by-word tests by applying one of a set of 50 text descriptions of colors, textures, styles, states, and shapes. Humans compare each of the 3000 edited images to the corresponding original image: they evaluate whether one image is better described by the text description than the other, and separately whether one image is more realistic than the other. Each comparison is done without revealing which image is the original. Each comparison is evaluated by two people.

Results are tabulated by category in Table 1, and charted by word in Figure 7. Our method can decisively edit a range of colors and textures, but it is also effective for a range of other visual classes to a lesser degree. In all tested categories, most visual words can have some positive effect discerned by raters, however our method is weakest at being able to alter the shapes of objects. Examples of success and failure cases from the test are shown in Figure 6.

In Table 1, we note that ratings of realism of edited images are worst in the same categories where the strength of the semantic consistency is best. Some of the loss of realism is due to visible artifacts such as the blurred colors seen in Figure 6(e). However, other cases without noticeable visual artifacts are rated as less realistic by raters as in Figure 6(c,d), including many of the most interesting edits in the study. In these cases, the new painted styles are noticeable and stand out as less realistic than the original because the painted object now has an atypical style for the context - such as a gold bed in a cream-colored room, or an ugly bed in a hotel room. The portion of effective edits that are rated as unrealistic is shown as orange bars in Figure 6.

3 Demonstrating paint-by-word on a broader diversity of image domains

To demonstrate the usefulness of our method beyond the narrow domains of birds and bedrooms, we apply our method on BigGAN models . We apply two BigGAN models. First, we apply a BigGAN model trained on MIT Places ; and then we apply a BigGAN model trained on Imagenet .

Because the BigGAN models are trained on broad distributions of images, these generative models allow us to experiment with a wide variety of different editing operations. In Figure 8, We demonstrate our method’s ability to alter the style of buildings outdoors (a); we also demonstrate the ability to add new objects that did not previously exist in the scene (b). Both (a) and (b) demonstrate the potential of of BigGAN trained on Places to be used to manipulate general scene images. On BigGAN Imagenet, we discover the ability to change whether a dog is happy or sad (c-d). We can also add new objects to an appropriate context (e). A challenging task is to force the model to synthesize objects outside its training domain. Although the Imagenet training data contains no blue firetrucks, we can use language to specify that we wish for the model to render a firetruck as blue (f). Using our method, it can create blue parts of a vehicle, but the out-of-domain task of creating a whole blue firetruck remains difficult and introduces distortions.

Discussion

We have introduced the problem of paint-by-word, a human-AI collaborative task which enables a user to edit a generated image by painting a semantic modification specified by any text description, to any location of the image. Our work is enabled by the recent development of large-scale high-performance networks for realistic image generation and accurate text-image similarity matching. We have shown that a natural combination of these powerful components can work well enough to enable paint-by-word. Our user studies and demonstrations of effective edits have taken a first step in characterizing the opportunities, challenges and potential utility in this new application.

Acknowledgements. We thank Aditya Ramesh at OpenAI for assistance with DALL-E evaluation data. We thank OpenAI, Google, and Nvidia for publishing weights for pretrained large-scale CLIP, BigGAN and StyleGAN models that make this work possible. We thank Hendrik Strobelt and Daksha Yadav for their insights, encouragement, and valuable discussions, and we are grateful for the support of DARPA XAI (FA8750-18-C-0004), and Signify Lighting Research.

Alex Andonian contributed to developing the method, creating data sets, implementing and running experiments, analyzing results, and writing. Sabrina Osmany developed the concept and prototype for abstract words like minimal, rustic, etc, including data sets, experiments, analysis. Audrey Cui contributed to developing the method, implemented experiments. YeonHwan Park contributed to implementing experiments, graphing and comparing results with other techniques. Ali Jahanian contributed to developing the method, implementing experiments, and writing. David Bau contributed to developing the method, created data sets, implemented experiments, analyzed results, and writing.

References