Image Captioning with Semantic Attention

Quanzeng You, Hailin Jin, Zhaowen Wang, Chen Fang, Jiebo Luo

Introduction

Automatically generating a natural language description of an image, a problem known as image captioning, has recently received a lot of attention in Computer Vision. The problem is interesting not only because it has important practical applications, such as helping visually impaired people see, but also because it is regarded as a grand challenge for image understanding which is a core problem in Computer Vision. Generating a meaningful natural language description of an image requires a level of image understanding that goes well beyond image classification and object detection. The problem is also interesting in that it connects Computer Vision with Natural Language Processing which are two major fields in Artificial Intelligence.

There are two general paradigms in existing image captioning approaches: top-down and bottom-up. The top-down paradigm starts from a ‘‘gist’’ of an image and converts it into words, while the bottom-up one first comes up with words describing various aspects of an image and then combines them. Language models are employed in both paradigms to form coherent sentences. The state-of-the-art is the top-down paradigm where there is an end-to-end formulation from an image to a sentence based on recurrent neural networks and all the parameters of the recurrent network can be learned from training data. One of the limitations of the top-down paradigm is that it is hard to attend to fine details which may be important in terms of describing the image. Bottom-up approaches do not suffer from this problem as they are free to operate on any image resolution. However, they suffer from other problems such as there lacks an end-to-end formulation for the process going from individual aspects to sentences. There leaves an interesting question: Is it possible to combine the advantages of these two paradigms? This naturally leads to feedback which is the key to combine top-down and bottom-up information.

Visual attention is an important mechanism in the visual system of primates and humans. It is a feedback process that selectively maps a representation from the early stages in the visual cortex into a more central non-topographic representation that contains the properties of only particular regions or objects in the scene. This selective mapping allows the brain to focus computational resources on an object at a time, guided by low-level image properties. The visual attention mechanism also plays an important role in natural language descriptions of images biased towards semantics. In particular, people do not describe everything in an image. Instead, they tend to talk more about semantically more important regions and objects in an image.

In this paper, we propose a new image captioning approach that combines the top-down and bottom-up approaches through a semantic attention model. Please refer to Figure 1 for an overview of our algorithm. Our definition for semantic attention in image captioning is the ability to provide a detailed, coherent description of semantically important objects that are needed exactly when they are needed. In particular, our semantic attention model has the following properties: 1) able to attend to a semantically important concept or region of interest in an image, 2) able to weight the relative strength of attention paid on multiple concepts, and 3) able to switch attention among concepts dynamically according to task status. Specifically, we detect semantic concepts or attributes as candidates for attention using a bottom-up approach, and employ a top-down visual feature to guide where and when attention should be activated. Our model is built on top of a Recurrent Neural Network (RNN), whose initial state captures global information from the top-down feature. As the RNN state transits, it gets feedback and interaction from the bottom-up attributes via an attention mechanism enforced on both network state and output nodes. This feedback allows the algorithm to not only predict more accurately new words, but also lead to more robust inference of the semantic gap between existing predictions and image content.

The main contribution of this paper is a new image captioning algorithm that is based on a novel semantic attention model. Our attention model naturally combines the visual information in both top-down and bottom-up approaches in the framework of recurrent neural networks. Our algorithm yields significantly better performance compared to the state-of-the-art approaches. For instance, on Microsoft COCO and Flickr 30K, our algorithm outperforms competing methods consistently across different evaluation metrics (Bleu-1,2,3,4, Meteor, and Cider). We also conduct an extensive study to compare different attribute detectors and attention schemes.

It is worth pointing out that also considered using attention for image captioning. There are several important differences between our work and . First, in attention is modeled spatially at a fixed resolution. At every recurrent iteration, the algorithm computes a set of attention weights corresponding to pre-defined spatial locations. Instead, we can use concepts from anywhere at any resolution in the image. Indeed, we can even use concepts that do not have direct visual presence in the image. Second, in our work there is a feedback process that combines top-down information (the global visual feature) with bottom-up concepts which does not exist in . Third, in uses pretrained feature at a particular spatial location. Instead, we use word features that correspond to detected visual concepts. This way, we can leverage external image data for training visual concepts and external text data for learning semantics between words.

Related work

There is a growing body of literature on image captioning which can be generally divided into two categories: top-down and bottom-up. Bottom-up approaches are the ‘‘classical’’ ones, which start with visual concepts, objects, attributes, words and phrases, and combine them into sentences using language models. and detect concepts and use templates to obtain sentences, while pieces together detected concepts. and use more powerful language models. and are the latest attempts along this direction and they achieve close to the state-of-the-art performance on various image captioning benchmarks.

Top-down approaches are the ‘‘modern’’ ones, which formulate image captioning as a machine translation problem . Instead of translating between different languages, these approaches translate from a visual representation to a language counterpart. The visual representation comes from a convolutional neural network which is often pretrained for image classification on large-scale datasets . Translation is accomplished through recurrent neural networks based language models. The main advantage of this approach is that the entire system can be trained from end to end, i.e., all the parameters can be learned from data. Representative works include . The differences of the various approaches often lie in what kind of recurrent neural networks are used. Top-down approaches represent the state-of-the-art in this problem.

Visual attention is known in Psychology and Neuroscience for long but is only recently studied in Computer Vision and related areas. In terms of models, approach it with Boltzmann machines while does with recurrent neural networks. In terms of applications, studies it for image tracking, studies it for image recognition of multiple objects, and uses for image generation. Finally, as we discuss in Section 1, we are not the first to consider it for image captioning. In , Xu et al., propose a spatial attention model for image captioning.

Semantic attention for image captioning

We extract both top-down and bottom-up features from an input image. First, we use the intermediate filer responses from a classification Convolutional Neural Network (CNN) to build a global visual description denoted by v\bm{v}. In addition, we run a set of attribute detectors to get a list of visual attributes or concepts {Ai}\{A_{i}\} that are most likely to appear in the image. Each attribute AiA_{i} corresponds to an entry in our vocabulary set or dictionary Y\mathcal{Y}. The design of attribute detectors will be discussed in Section 4.

Different from previous image captioning methods, our model has a unique way to utilize and combine different sources of visual information. The CNN image feature v\bm{v} is only used in the initial input node x0\bm{x}_{0}, which is expected to give RNN a quick overview of the image content. Once the RNN state is initialized to encompass the overall visual context, it is able to select specific items from {Ai}\{A_{i}\} for task-related processing in the subsequent time steps. Specifically, the main working flow of our system is governed by the following equations:

where a linear embedding model is used in Eq. (1) with weight Wx,v\bm{W}^{x,v}. For conciseness, we omit all the bias terms of linear transformations in the paper. The input and output attention models in Eq. (3) and (4) are designed to adaptively attend to certain cognitive cues in {Ai}\{A_{i}\} based on the current model status, so that the extracted visual information will be most relevant to the parsing of existing words and the prediction of future word. Eq. (2) to (4) are recursively applied, through which the attended attributes are fed back to state ht\bm{h}_{t} and integrated with the global information from v\bm{v}. The design of Eq. (3) and (4) is discussed below.

2 Input attention model

Once calculated, the attention scores are used to modulate the strength of attention on different attributes. The weighted sum of all attributes is mapped from word embedding space to the input space of xt\bm{x}_{t} together with the previous word:

3 Output attention model

The output attention model φ\varphi is designed similarly as the input attention model. However, a different set of attention scores are calculated since visual concepts may be attended in different orders during the analysis and synthesis processes of a single sentence. With all the information useful for predicting YtY_{t} captured by the current state ht\bm{h}_{t}, the score βti\beta_{t}^{i} for each attribute AiA_{i} is measured with respect to ht\bm{h}_{t}:

Again, {βti}\{\beta^{i}_{t}\} are used to modulate the attention on all the attributes, and the weighted sum of their activations is used as a compliment to ht\bm{h}_{t} in determining the distribution pt\bm{p}_{t}. Specifically, the distribution is generated by a linear transform followed by a softmax normalization:

4 Model learning

The training data for each image consist of input image features v\bm{v}, {Ai}\{A_{i}\} and output caption words sequence {Yt}\{Y_{t}\}. Our goal is to learn all the attention model parameters ΘA={U,V,W∗,∗,w∗,∗}\mathbf{\Theta}_{A}=\{\bm{U},\bm{V},\bm{W}^{*,*},\bm{w}^{*,*}\} jointly with all RNN parameters ΘR\mathbf{\Theta}_{R} by minimizing a loss function over training set. The loss of one training example is defined as the total negative log-likelihood of all the words combined with regularization terms on attention scores {αti}\{\alpha^{i}_{t}\} and {βti}\{\beta^{i}_{t}\}:

where α\bm{\alpha} and β\bm{\beta} are attention score matrices with their (t,i)(t,i)-th entries being αti\alpha^{i}_{t} and βti\beta^{i}_{t}. The regularization function gg is used to enforce the completeness of attention paid to every attribute in {Ai}\{A_{i}\} as well as the sparsity of attention at any particular time step. This is done by minimizing the following matrix norms of α\bm{\alpha} (same for β\bm{\beta}):

where the first term with p>1p{>}1 penalizes excessive attention paid to any single attribute AiA_{i} accumulated over the entire sentence, and the second term with 0<q<10{<}q{<}1 penalizes diverted attention to multiple attributes at any particular time. We use a stochastic gradient descent algorithm with an adaptive learning rate to optimize Eq. (10).

Visual attribute prediction

The prediction of visual attributes {Ai}\{A_{i}\} is a key component of our model in both training and testing. We propose two approaches for predicting attributes from an input image. First, we explore a non-parametric method based on nearest neighbor image retrieval from a large collection of images with rich and unstructured textual metadata such as tags and captions. The attributes for a query image can be obtained by transferring the text information from the retrieved images with similar visual appearances. The second approach is to directly predict visual attributes from the input image using a parametric model. This is motivated by the recent success of deep learning models on visual recognition tasks . The unique challenge for attribute detection is that usually there are more than one visual concepts presented in an image, and therefore we are faced with a multi-label problem instead of a multi-class problem. Note that the two approaches to obtain attributes are complementary to each other and can be used jointly. Figure 3 shows an example of visual attributes predicted for an image using different methods.

Thanks to the popularity of social media, there is a growing number of images with weak labels, tags, titles and descriptions available on Internet. It has been shown that these weakly annotated images can be exploited to learn visual concepts , text-image embedding and image captions . One of the fundamental assumptions is that similar images are likely to share similar and correlated annotations. Therefore, it is possible to discover useful annotations and descriptions from visual neighbors in a large-scale image dataset.

We extract key words as the visual attributes for our model from a large image dataset. For fair comparison with other existing work, we only do nearest neighbor search on our training set to retrieve similar ones to test images. It is expected that the attribute prediction accuracy can be further improved by using a larger web-scale database. We use the GoogleNet feature to evaluate image distances, and employ simple Term-Frequency (TF) to select the most frequent words in the ground-truth captions of the retrieved training images. In this way, we are able to build a list of words for each image as the detected visual attributes.

2 Parametric attribute prediction

In addition to retrieved attributes, we also train parametric models to extract visual attributes. We first build a set of fixed visual attributes by selecting the most common words from the captions in the training data. The resulting attributes are treated as a set of predefined categories and can be learned as in a conventional classification problem.

The advance of deep learning has enabled image analysis to go beyond the category level. In this paper we mainly investigate two state-of-the-art deep learning models for attribute prediction: using a ranking loss as objective function to learn a multi-label classifier as in , and using a Fully Convolutional Network (FCN) to learn attributes from local patches as in . Both two methods produce a relevance score between an image and a visual attribute, which can be used to select the top ranked attributes as input to our captioning model. Alternatives may exist which can potentially yield better results than the above two models, which is not in the scope of this work.

Experiments

We perform extensive experiments to evaluate the proposed models. We report all the results using Microsoft COCO caption evaluation toolhttps://github.com/tylin/coco-caption, including BLEU, Meteor, Rouge-L and CIDEr . We will first briefly discuss the datasets and settings used in the experiments. Next, we compare and analyze the results of the proposed model with other state-of-the-art models on image captioning.

We choose the popular Flickr30k and MS-COCO to evaluate the performance of our models. Flickr30k has a total of 31,78331,783 images. MS-COCO is more challenging, which has 123,287123,287 images. Each image is given at least five captions by different AMT workers. To make the results comparable to others, we use the publicly available splitshttps://github.com/karpathy/neuraltalk of training, testing and validating sets for both Flickr30k and MS-COCO. We also follow the publicly available code to preprocess the captions (i.e. building dictionaries, tokenizing the captions).

Our captioning system is implemented based on a Long Short-Term Memory (LSTM) network . We set n=m=512n=m=512 for the input and hidden layers, and use tanh⁡\tanh as nonlinear activation function σ\sigma. We use Glove feature representation with d=300d=300 dimensions as our word embedding E\bm{E}.

The image feature v\bm{v} is extracted from the last 10241024-dimensional convolutional layer of the GoogleNet CNN model. Our attribute detectors are trained for the same set of visual concepts as in for Microsoft COCO data set. We build and train another independent set of attribute detectors for Flickr30k following the steps in using the training split of Flickr30k. The top 1010 attributes with highest detection scores are selected to form the set {Ai}\{A_{i}\} in our best attention model setting. An attribute set of such size can maintain a good tradeoff between precision and recall.

In training, we use RMSProp algorithm to do model updating with a mini-batch size of 256256. The regularization parameters are set as p=2,q=0.5p=2,q=0.5 in (3.4).

In testing, a caption is formed by drawing words from RNN until a special end word is reached. All our results are obtained with the ensemble of 5 identical models trained with different initializations, which is a common strategy adopted in other work .

In the following experiments, we evaluate different ways to obtain visual attributes as described in Section 4, including one non-parametric method (kk-NN) and two parametric models trained with ranking-loss (RK) and fully-connected network (FCN). Besides the attention model (ATT) described in Section 3, two fusion-based methods to utilize the detected attributes {Ai}\{A_{i}\} are tested by simply taking the element-wise max (MAX) or concatenation (CON) of the embedded attribute vectors {Eyi}\{\bm{E}\bm{y}^{i}\}. The combined attribute vector is used in the same framework and applied at each time step.

2 Performance on MS-COCO

Note that the overall captioning performance will be affected by the employed visual attributes generation method. Therefore, we first assume ground truth visual attributes are given and evaluate different ways (CON, MAX, ATT) to select these attributes. This will indicate the performance limit of exploiting visual attributes for captioning. To be more specific, we select the most common words as visual attributes from their ground-truth captions to help the generation of captions. Table 1 shows the performance of the three models using the ground-truth visual attributes. These results can be considered as the upper bound of the proposed models, which suggest that all of the proposed models (ATT, MAX and CON) can significantly improve the performance of image captioning system, if given visual attributes of high quality.

Now we evaluate the complete pipeline with both attribute detection and selection. The right half of Table 2 shows the performance of the proposed model on the validation set of MS-COCO. In particular, our proposed attention model outperforms all the other state-of-the-art methods in most of the metrics, which are commonly used together for fair and overall performance measurement. Note that B-1 is related to single word accuracy, the performance gap of B-1 between our model and may be due to different preprocessing for word vocabularies.

In Table 2, the entries with prefix ‘‘Ours" show the performance of our method configured with different combinations of attribute detection and selection methods. In general, attention model ATT with attributes predicted by FCN model yields better performance than other combinations over all benchmarks.

For attribute fusion methods MAX and CON, we find using the top 33 attributes gives the best performance. Due to the lack of attention scheme, too many keywords may increase the parameters for CON and may reduce the distinction among different groups of keywords for MAX. Both models have comparable performance. The results also suggest that FCN gives more robust visual attributes. MAX and CON can also outperform the state-of-the-art models in most evaluation metrics using visual attributes predicted by FCN. Attention models (ATT) on FCN visual attributes show the best performance among all the proposed models. On the other hand, visual attributes predicted by ranking loss (RK) based model seem to have even worse performance than kk-NN. This is possible due to the lack of local features in training the ranking loss based attribute detectors.

Performance on MS-COCO 2014 test server We also evaluate our best model, Ours-ATT-FCN, on the MS COCO Image Captioning Challenge sets c5 and c40 by uploading results to the official test server. In this way, we could compare our method to all the latest state-of-the-art methods. Despite the popularity of this contest, our method has held the top 1 position by many metrics at the time of submission. Table 3 lists the performance of our model and other leading methods. Besides the absolute scores, we provide the rank of our model among all competing methods for each metric. By comparing with two other leading methods, we can see that our method achieves better ranking across different metrics. All the results are up-to-date at time of submission.

3 Performance on Flickr30k

We now report the performance on Flickr30k dataset. Similarly, we first train and test our models by using the ground-truth visual attributes to get an upper-bound performance. The obtained results are listed in Table 1. Clearly, with correct visual attributes, our model is able to improve caption results by a large margin comparing to other methods. We then conduct the full evaluation. As shown in the left half of Table 2, the performance of our models are consistent with that on MS-COCO, and Ours-ATT-FCN achieves significantly better results over all competing methods in all metrics, except B-1 score, for which we have discussed potential causes in previous section.

4 Visualization of attended attributes

We now provide some representative captioning examples in Figure 4 for better understanding of our model. For each example, Figure 4 contains the generated captions for several images with the input attention weights αti\alpha_{t}^{i} and the output attention weights βti\beta_{t}^{i} plotted at each time step. The generated caption sentences are shown under the horizontal time axis of the curve plots, and each word is positioned at the time step it is generated. For visual simplicity, we only show the attention weights of top attributes from the generated sentence. As captions are being generated, the attention weights at both input and output layers vary properly as sentence context changes, while the distinction between their weights shows the underlying attention mechanisms are different. In general, the activations of both α\alpha and β\beta have strong correlation with the words generated. For example, in the Figure 4, the attention on ‘‘swimming’’ peaks after ‘‘ducks’’ is generated for both α\alpha and β\beta. In Figure 4, the concept of ‘‘motorcycle’’ attracts strong attention for both α\alpha and β\beta. The β\beta peaks twice during the captioning process, one after ‘‘photo of’’ and the other after ‘‘riding a’’, and both peaks reasonably align with current contexts. It is also observed that, as the output attention weight, β\beta correlates with output words more closely; while the input weights α\alpha are allocated more on background context such as the ‘‘plate’’ in Figure 4 and the ‘‘group’’ in Figure 4. This temporal analysis offers an intuitive perspective on our visual attributes attention model.

5 Analysis of attention model

As described in Section 3.2 and Section 3.3, our framework employs attention at both input and output layers to the RNN module. We evaluate the effect of each of the individual attention modules on the final performance by turning off one of the attention modules while keeping the other one in our ATT-FCN model. The two model variants are trained on MS-COCO dataset using the ground-truth visual attributes, and compared in Table 4. The performance of using output attention is slightly better than only using input attention on some metrics. However, the combination of this two attentions improves the performance by several percents on almost every metric. This can be attributed to that fact that attention mechanisms at input and output layers are not the same, and each of them attend to different aspects of visual attributes. Therefore, combining them may help provide a richer interpretation of the context and thus lead to improved performance.

6 The role of visual attributes

We also conduct a qualitative analysis on the role of visual attributes in caption generation. We compare our attention model (ATT) with Google NIC, which corresponds to the LSTM model used in our framework. Figure 5 shows several examples. We can find that visual attributes can help our model to generate better captions, as shown by the examples in the green box. However, irrelevant visual attributes may disrupt the model to attend on incorrect concepts. For example, in the left example in the red dashed box, ‘‘clock’’ distracts our model to the clock tower in background from the main objects in foreground. In the rightmost example, and ‘‘tower’’ may be the culprit of the word ‘‘building’’ in the predicted caption.

Conclusion

In this work, we proposed a novel method for the task of image captioning, which achieves state-of-the-art performance across popular standard benchmarks. Different from previous work, our method combines top-down and bottom-up strategies to extract richer information from an image, and couples them with a RNN that can selectively attend on rich semantic attributes detected from the image. Our method, therefore, exploits not only an overview understanding of input image, but also abundant fine-grain visual semantic aspects. The real power of our model lies in its ability to attend on these aspects and seamlessly fuse global and local information for better caption. For next steps, we plan to experiment with phrase-based visual attribute with its distributed representations, as well as exploring new models for our proposed semantic attention mechanism.

Acknowledgment

This work was generously supported in part by Adobe Research and New York State through the Goergen Institute for Data Science at the University of Rochester.

References