GANalyze: Toward Visual Definitions of Cognitive Image Properties

Lore Goetschalckx, Alex Andonian, Aude Oliva, Phillip Isola

Introduction

Why do we remember the things we do? Decades of work have provided numerous explanations: we remember things that are out of context , that are emotionally salient , that involve people , etc. But a picture is, as they say, worth a thousand words. What does it look like to make an image more or less memorable? The same questions can be asked for many cognitive visual properties: what visual changes can take a bland foggy seascape and add just the right colors and tones to make it serenely beautiful.

Attributes like memorability, aesthetics, and emotional valence are of special interest because we do not have concrete definitions of what they entail. This contrasts with attributes like “object size" and "smile". We know exactly what it means to zoom in on a photo, and it’s easy to imagine what a face looks like as it forms a smile. It’s an open question, on the other hand, what exactly do changes in “memorability" look like? Previous work has built powerful predictive models of image memorability but these have fallen short of providing a fine-grained visual explanation of what underlies the predictions.

In this paper, we propose a new framework, GANalyze, based on Generative Adversarial Networks (GAN) , to study the visual features and properties that underlie high-level cognitive attributes. We focus on image memorability as a case study, but also show that the same methods can be applied to study image aesthetics and emotional valence.

Our approach leverages the ability of GANs to generate a continuum of images with fine-grained differences in their visual attributes. We can learn how to navigate the GAN’s latent space to produce images that have increasing or decreasing memorability, according to an off-the-shelf memorability predictor . Starting with a seed image, this produces a sequence of images of increasing and decreasing predicted memorability (see Figure 1). By showing this visualization for a diverse range of seed images, we come up with a catalog of different image sequences showcasing a variety of visual effects related to memorability. We call this catalog a visual definition of image memorability. GANalyze thereby offers an alternative to the non-parametric approach in which real images are simply sorted on their memorability score to visualize what makes them memorable (example shown in Figure S1). The parametric, fine-grained visualizations generated by GANalyze provide much clearer visual definitions.

These visualizations surface several correlates of memorability that have been overlooked by prior work, including “object size", “circularity", and “colorfulness". Most past work on modeling image memorability focused on semantic attributes, such as object category (e.g., “people" are more memorable than “trees") . By applying our approach to a class-conditional GAN, BigGAN , we can restrict it to only make changes that are orthogonal to object class. This reveals more fine-grained changes that nonetheless have large effects on predicted memorability. For example, consider the cheeseburgers in Figure 4. Our model visualizes more memorable cheeseburgers as we move to the right. The apparent changes go well beyond semantic category – the right-most burger is brighter, rounder, more canonical, and, we think, looks tastier.

Since our visualizations are learned based on a model of memorability, a critical step is to verify that what we are seeing really has a causal effect on human behavior. We test this by running a behavioral experiment that measures the memorability of images generated by our GAN, and indeed we find that our manipulations have a causal effect: navigating the GAN manifold toward images that are predicted to be more memorable actually results in generating images that are measurably more memorable in the behavioral experiment.

Introducing GANalyze, a framework that uses GANs to provide a visual definition of image properties, like memorability and aesthetics, that we can measure but are not easy, in words, to define.

Showing that this framework surfaces previously overlooked attributes that correlate with memorability.

Demonstrating that the discovered transformations have a causal effect on memorability.

Showing that GANalyze can be applied to provide visual definitions for aesthetics and emotional valence.

Generative Adversarial Networks or GANs. GANs introduced a revolutionary framework to synthesize natural-looking images . Among the many applications for GANs are style transfer , visual prediction , and “sim2real" domain adaptation . Here, we show how they can also be applied to the problem of understanding high-level, cognitive image properties, such as memorability.

Understanding CNN representations The internal representations of a CNN can be unveiled using methods like network dissection including for a CNN trained on memorability . For instance, Khosla et al. showed that units with strong positive correlations with memorable images specialized for people, faces, body parts, etc., while those with strong negative correlations where more sensitive to large regions in landscapes scenes. Here, our framework introduces a new way of defining what memorability, and aesthetic, variability look like.

Modifying Memorability. The memorability of an image, like faces, can be manipulated using warping techniques . Concurrent work has also explored using a GAN for this purpose . Another approach is a deep style transfer which taps into more artistic qualities. Now that GANs have reached a quality that is often almost indistinguishable from real images, they offer a powerful tool to synthesize images with different cognitive qualities. As shown here, our GANalyze framework successfully modified GAN-generated images across a wide range of image categories to produce a second generation of GAN realistic photos with different mnemonic qualities.

Model

Note that this is simply the MSE loss between the target memorability score, i.e. the seed image’s score A(G(z,y))A(G(\mathbf{z},\mathbf{y})) increased by α\alpha, and the memorability score of the transformed clone image A(G(Tθ(z,α),y))A(G(T_{\theta}(\mathbf{z},\alpha),\mathbf{y})). The scalar α\alpha acts as a metaphorical knob with which one can use to turn up or turn down memorability. The optimizing problem is θ∗=arg⁡ ⁣min⁡θL(θ)\theta^{*}=\arg\!\min_{\theta}\mathcal{L}(\theta). The Transformer TT is defined as:

Figure 2 presents a schematic of the model. Finally, note that when α=0\alpha=0, TT becomes a null operation and G(Tθ(z,α),y)G(T_{\theta}(\mathbf{z},\alpha),\mathbf{y}) then equals G(z,y)G(\mathbf{z},\mathbf{y}).

2 Implementation

For the results presented here, we used the Generator of BigGAN , which generates state-of-the art GAN images and is pretrained on ImageNet . The Assessor was implemented as MemNet , a CNN predicting image memorability. Note, however, that training our model with different Generators or different Assessors can easily be achieved by substituting the respective modules. We discuss an Assessor for image aesthetics in Section 4. Furthermore, we present additional results for implementations with a StyleGAN Generator in the supplementary materials.

To train our model and find θ∗\theta^{*}, we built a training set by randomly sampling 400K z\mathbf{z} vectors from a standard normal distribution truncated to the range $.Each. Each\mathbf{z}wasaccompaniedbyanwas accompanied by an\alphavalue,randomlydrawnfromauniformdistributionbetween−0.5and0.5,andarandomlychosenvalue, randomly drawn from a uniform distribution between -0.5 and 0.5, and a randomly chosen\mathbf{y}$. We used a batch size of 4 and an Adam optimization procedure.

In view of the behavioral experiments (see Section 3), we restricted the test set to 750 randomly chosen ImageNet classes and two z\mathbf{z} vectors per class. Each z\mathbf{z} vector was then paired with five different α\alpha values: [−0.2,−0.1,0,0.1,0.2]{[-0.2,-0.1,0,0.1,0.2]}. Note that this includes an α\alpha of 0, representing the original image G(z,y)G(\mathbf{z},\mathbf{y}). Finally, the test set consisted of 1.5K sets of five images, or 7.5K test images in total.

Experiments

Did our model learn to navigate the latent space such that it can increase (or decrease) the Assessor score of the generated image with positive (or negative) α\alpha values?

Figure 3.A suggests the model learned. The mean MemNet score of test set images increases with every increment of α\alpha. To test this formally, we fitted a linear mixed-effects regression model to the data and found a (unstandardized) slope (β\beta) of 0.68 (95%CI=[0.66,0.70],p<0.001)95\%CI=[0.66,0.70],p<0.001), confirming that the Memnet score increases significantly with α\alpha.

2 Emerging factors

We observe that the model can successfully change the memorability of an image, given its z\mathbf{z} vector. Next, we ask which image factors it altered to achieve this. The answer to this question can provide further insight into what the Assessor has learned about the to-be-assessed image property, in this case what MemNet has learned about memorability. From a qualitative analysis of the test set (examples shown in Figures 4, S2, and S3), a number of candidate factors stand out.

First, MemNet assigns higher memorability scores when the size of the object (or animal) in the image is larger, as our model is in many cases zooming in further on the object with every increase of α\alpha.

Second it is centering the subject in the image frame.

Third, it seems to strive for square or circular shapes in classes where it is realistic to do so (e.g., snake, cheeseburger, necklace, and espresso in Figure 4).

Fourth, it is often simplifying the image from low to high α\alpha, by reducing the clutter and/or number of objects, such as in the cheeseburger or flamingo, or by making the background more homogeneous, as in the snake example (see Figure 4).

A fifth observation is that the subject’s eyes sometimes become more pronounced and expressive, in particular in the dog classes (see Figure 1).

Sixth, one can also detect color changes between the different α\alpha conditions. Positive α′s\alpha^{\prime}s often produce brighter and more colorful images, and negative α′s\alpha^{\prime}s often produce darker images with dull colors. Finally, for those classes where multiple object hues can be considered realistic (e.g., the the bell pepper and the necklace in Figure 1 and Figure 4), the model seems to prefer a red hue.

To verify our observations, we quantified the factors listed above for the images in the test set (except for "expressive eyes", which is more subjective and harder to quantify). Brightness was measured as the average pixel value after transforming the image to grayscale. For colorfulness, we used the metric proposed by , and for redness we computed the normalized number of red pixels. Finally, the entropy of the pixel intensity histogram was taken as proxy for simplicity. For the remaining three factors, a pretrained Mask R-CNN was used to generate an instance-level segmentation mask of the subject. To capture object size, we calculated the difference in the mask’s area (normalized number of pixels) as the step size α\alpha varied. To measure centeredness, we computed the deviation of the mask’s centroid from the center of the frame. Finally, we calculated the length of minor and major axes of an ellipse that has the same normalized second central moments as the mask, and used their ratio as a metric of squareness. Figure 3.B shows that the emerging factor scores increase with α\alpha.

3 Realness

While BigGAN achieves state-of-the-art to generate highly realistic images, there remains a certain variability in the “realness" of the generated images. How best to evaluate the realness of a set of GAN-images is still an open question. Below, we discuss two automatically computed realness measures and a human measure in relation to our data. We discuss an additional human measure, based on a different task, in the supplementary materials.

In Figure 5.A, we plot two popular automatic measures in function of α\alpha: the Frechet Inception Distance (FID) and the Inception Score (IS) . A first observation is that the FID is below 40 in all α\alpha conditions. An FID as low as 40 already corresponds to reasonably realistic images. Thus the effects of our model’s modifications on memorability are not explained by making the images unrealistic. But we do observe interesting differences in FID- and IS-differences related to α\alpha, suggesting that more memorable images have more interpretable semantics.

3.2 Human measure

In addition to the two automatic measures, we conducted an experiment to collect human realness scores. The experiment consisted of a two-alternative forced choice (2AFC) task, hosted on Amazon Mechanical Turk (AMT), in which workers had to discriminate GAN-images from real ones. Workers were shown a series of pairs, consisting of one GAN-image and one real image. They were presented side by side for a duration of 1.6 s. Once a pair had disappeared off the screen, workers pressed the j-key when they thought the GAN-image was shown on the right, or the f-key when they thought it was shown on the left. The position of the GAN-image was randomized across trials. The set of real images used in this experiment was constructed by randomly sampling 10 real ImageNet exemplars per GAN-image class. The set of GAN-images was the same as the one quantified on memorability in Section 3.4. A GAN-image was randomly paired with one of the 10 real images belonging to the same class. Each series consisted of 100 trials, of which 20 were vigilance trials. For the vigilance trials, we generated GAN-images from z\mathbf{z} vectors that were sampled from the tails of a normal distribution (to make them look less real). For a worker’s first series, we prepended 20 trials with feedback as practice (not included in the analyses). Workers could complete up to 17 series, but were blocked if they scored less than 65% correct on the vigilance trials. Series that failed this criterion were also excluded from the analyses. The pay rate equaled 0.50percompletedseries.Onaverage,eachofourtestimageswasseenby2.76workers,meaning4137datapointsper0.50 per completed series. On average, each of our test images was seen by 2.76 workers, meaning 4137 data points per\alpha$ condition.

We did not observe differences in task performance between different α\alpha (see Figure 5.B). Indeed, a logistic mixed-effects regression fitted to the raw, binary data (correct/incorrect) did not reveal a statistically significant regression weight for α\alpha (β=−0.08,95%CI=[−0.33,0.18],p=0.55\beta=-0.08,95\%CI=[-0.33,0.18],p=0.55). In other words, the model’s image modifications did not affect workers’ ability to correctly identify the fake image, indicating that perceptually, the image clones of a seed image did not differ in realness.

4 Do our changes causally affect memory?

In addition to the MemNet scores, is our model also successful at changing the probability of an image being recognized by participants in an actual memory experiment?

We tested people’s memory for the images of a test set (see Section 2.2) using a repeat-detection visual memory game, which was hosted on AMT (see Figure 6). . AMT workers watched a series of one image at the time and had to press a key whenever they saw a repeat of a previously shown image. Each series consisted of 215 images, shown each for 600 ms with a blank interstimulus interval of 800 ms. Sixty images were targets, sampled from our test set, and repeated after 34 to 139 intervening images. The remaining images were either filler or vigilance images and were sampled from a separate set. This set was created with 10 z\mathbf{z} vectors per class and the same five α\alpha values as the test set: [−0.2,−0.1,0,0.1,0.2][-0.2,-0.1,0,0.1,0.2], making a total of 37.5K images. Filler images were only presented once and ensured spacing between a target and its repeat. Vigilance images were presented twice, with 0 to 3 intervening images in-between the two presentations. The vigilance repeats constituted easy trials to keep workers attentive. Care was taken to ensure that a worker never saw more than one G(Tθ(z,α),y)G(T_{\theta}(\mathbf{z},\alpha),\mathbf{y}) for a given z\mathbf{z}. Workers could complete up to 25 series, but were blocked if they missed more than 55% of the vigilance repeats in a series or made more than 30% false positives. Series that failed this were excluded from the analyses. The pay rate was 0.50percompletedseries.Onaverage,atestimagewasseenby3.16workers,with4740datapointsper0.50 per completed series. On average, a test image was seen by 3.16 workers, with 4740 data points per\alpha$ condition.

Workers could either recognize a repeated test image (hit, 1), or miss it (miss, 0). Figure 7.A shows the hit rate across all images and workers. The hit rate increases with every step of α\alpha. Fitting a logistic mixed-effects regression model to the raw, binary data (hit/miss), we found that the predicted log odds of image being recognized increase with 0.19 for an increase in α\alpha of 0.01 (β=1.92,95%CI=[1.71−2.12],p<0.001\beta=1.92,95\%CI=[1.71-2.12],p<0.001). This shows that our model can successfully navigate the BigGAN latent space in order to make an image more (or less) memorable to humans.

Given human memory data for images modified for memorability, we evaluate how the images’ emerging factor scores relate to their likelihood of being recognized. We fitted mixed-effects logistic regression models, each with a different emerging factor as the predictor, see Table 1. Except for entropy, all the emerging factors show a significant, positive relation to the likelihood of a hit in the memory game, but none fit the data as well as the model’s α\alpha. This indicates that a single emerging factor is not enough to fully explain the effect observed in Figure 7.A. Note that the emerging factor results are correlational and the factors are intercorrelated. This makes it hard to draw conclusions about which individual factors truly causally affect human memory performance. As an example of how this can be addressed within the GANalyze framework, we conducted an experiment focusing on the effect of one salient emerging factor: object size. As seen in Figure 4, more memorable images tend to center and enlarge the object class.

We trained a version of our model with an Object size Assessor, instead of the MemNet Assessor. This is the same Object size Assessor used to quantify the object size in the images modified according to MemNet (e.g., for the results in Figure 3.B), now teaching the Transformer to perform “enlarging" modifications. After training with 161750 z\mathbf{z} vectors, we generated a test set as described in Section 2.2, except that we used a different set of α\alpha’s: [−0.8,−0.4,0,0.4,0.8][-0.8,-0.4,0,0.4,0.8]. We chose these values to qualitatively match the degree of object size changes achieved by the MemNet version of the model. Figure 8.A visualizes the results achieved on the test set. The model successfully enlarges the object with increasing alpha’s, as confirmed by a linear mixed-effects regression analysis (β=0.07,95%CI=[0.06,0.07],p<0.001\beta=0.07,95\%CI=[0.06,0.07],p<0.001). Figure 10 shows example images generated by that model, after having been trained with 161750 z\mathbf{z} vectors. A comparison with images modified according to MemNet suggests that the latter model was doing more than just enlarging the object.

To study how the new size modifications affect memorability, we generated a new set of images (7.5K targets, 37.5K fillers) with α\alpha’s [−0.8,−0.4,0,0.4,0.8][-0.8,-0.4,0,0.4,0.8]. The new images were then quantified using the visual memory game (on average 2.36 data points per image and 3540 per α\alpha condition). Figure 7.B shows the results. Memory performance increases with α\alpha, as confirmed by a logistic mixed-effects analysis (β=0.11,95%CI=[0.06,0.18],p<0.001\beta=0.11,95\%CI=[0.06,0.18],p<0.001, although mostly for positive α\alpha values.

Other properties

As mentioned in Section 2.2, the proposed method can be applied to other image properties, simply by substituting the Assessor module. To show our framework can generalize, we trained a model for aesthetics, using Kong et al’s CNN (hereinafter referred to as AestheticsNet) as the Assessor. In addition, we also trained a model for emotional valence. Emotional valence refers to the extent to which the emotions evoked by an image are experienced as positive (or negative). For this property, we trained our own Assessor by fine-tuning a ResNet50 model , pretrained on the Moments database , to the Cornell Emotion6 Image Database . We refer to this Assessor as EmoNet. Finally, we generated a test set for each of the two new models, like we did for the memorability model.

Figure 8.B shows the average AestheticsNet scores per α\alpha condition. The scores significantly increase with α\alpha, as evidenced by the results of a linear mixed-effects regression (β=0.72,95%CI=[0.70,0.74],p<0.001\beta=0.72,95\%CI=[0.70,0.74],p<0.001). We can successfully train the model to increase (or decrease) an image’s aesthetic score as shown in Figure 9 (left) and Figure S4. Similarly, Figure 8.C shows the average EmoNet scores per α\alpha condition. Here too, the scores significantly increase with α\alpha (β=0.44,95%CI=[0.43,0.45],p<0.001\beta=0.44,95\%CI=[0.43,0.45],p<0.001). Example visualizations generated by this model are presented in Figure 9 (right) and Figure S5.

Based on a qualitative inspection of such visualizations, we observed that the aesthetics model is modifying factors like depth of field, color palette, and lighting, suggesting that the AestheticsNet is sensitive to those factors. Indeed, the architecture of the AestheticsNet includes attribute-adaptive layers to predict these factors, now highlighted by our visualizations. The emotional valence model often averts the subject’s gaze away from the "camera" when decreasing valence. To increase valence, it often makes images more colorful, introduces bokeh, and makes the skies more blue in landscape images. Finally, the teddy bear in Figure 1 (right) seems to smile more. Interestingly, the model makes different modifications for every property (see Figure 10), suggesting that what makes an image memorable is different from what makes it aesthetically pleasing or more positive in its emotional valence.

A final question we asked is whether an image modified to become more (less) aesthetic also becomes more (less) memorable? To test this, we quantified the images of the aesthetic test set on memorability by presenting them to workers in the visual memory game (we collected 1.54 data points per image and 2306 data points per α\alpha condition). Figure 7.C shows the human memory performance in function of an α\alpha that is tuning aesthetics. A logistic mixed-effects regression revealed that with an 0.1 increase in the aesthetics α\alpha, the predicted log odds of an image being recognized increase with 0.07 (β=0.72,95%CI=[0.44,1.00],p<0.001\beta=0.72,95\%CI=[0.44,1.00],p<0.001). While modifying an image to make it more aesthetic does increase its memorability, the effect is rather small, suggesting that memorability is more than only aesthetics and that our model was right to modify memorability and aesthetics in different ways.

Conclusion

We introduce GANalyze, a framework that shows how a GAN-based model can be used to visualize what another model (i.e. CNN as an Assessor) has learned about its target image property. Here we applied it to memorability, yielding a kind of “visual definition" of this high-level cognitive property, where we visualize what it looks like for an image to become more or less memorable. These visualizations surface multiple candidate features that may help explain why we remember what we do. Importantly, our framework can also be generalized to other image properties, such as aesthetics or emotional valence: by replacing the Assessor module, the framework allows us to explore the visual definition for any property we can model as a differentiable function of the image. We validated that our model successfully modified GAN images to become more (or less) memorable via a behavioral human memory experiment on manipulated images.

GANalyze’s intended use is to contribute to the scientific understanding of otherwise hard to define cognitive properties. Note that this was achieved by modifying images for which the encoding into the latent space of the GAN was given. In other words, it is currently only possible to modify seed images that are GAN-images themselves, not user-supplied, real images. However, should advances in the field lead to an encoder network, this would become possible and it would open applications in graphics and education, for example, where selected images can be made more memorable. One should also be wary, though, of potential misuse, especially when applied to images of people or faces. Note that the BigGAN generator used here was trained on ImageNet categories which only occasionally include people, and that it does not allow to render realistically looking people. Nevertheless, with generative models yielding ever more realistic output, an increasingly important challenge in the field is to develop powerful detection methods to allow us to reliably distinguish generated, fake images from real ones .

Acknowledgments

This work was partly funded by NSF award 1532591 in Neural and Cognitive Systems (to A.O), by a fellowship (Grant 1108116N) and a travel grant (Grant V4.085.18N) awarded to Lore Goetschalckx by the Research Foundation - Flanders (FWO).

References