GRIT: Faster and Better Image captioning Transformer Using Dual Visual Features
Van-Quang Nguyen, Masanori Suganuma, Takayuki Okatani
Introduction
Image captioning is the task of generating a semantic description of a scene in natural language, given its image. It requires a comprehensive understanding of the scene and its description reflecting the understanding. Therefore, most existing methods solve the task in two corresponding steps; they first extract visual features from the input image and then use them to generate a scene’s description. The key to success lies in the problem of how we can extract good features.
Researchers have considered several approaches to the problem. There are two primary methods, referred to as grid features and region features . Grid features are local image features extracted at the regular grid points, often obtained directly from a higher layer feature map(s) of CNNs/ViTs. Region features are a set of local image features of the regions (i.e., bounding boxes) detected by an object detector.
The current state-of-the-art methods employ the region features since they encode detected object regions directly. Identifying objects and their relations in an image will be useful to correctly describing the image. However, the region features have several issues. First, they do not convey contextual information such as objects’ relation since the regions do not cover the areas between objects. Second, there is a risk of erroneous detection of objects; important objects could be overlooked, etc. Third, computing the region feature is computationally costly, which is especially true when using a high-performance CNN-based detector, such as Faster R-CNN .
The grid features are extracted from the entire image, typically a high-layer feature map of a backbone network. While they do not convey object-level information, they are free from the first two issues with the region features. They may represent contextual information such as objects’ relations in images, and they are free from the risk of erroneous object detection.
In this study, we consider using such region and grid features in an integrated manner, aiming to build a better model for image captioning. The underlying idea is that properly integrating the two types of features will provide a better representation of input images since they are complementary, as explained above. While a few recent studies consider their integration , it is still unclear what the best way is. In this study, we reconsider how to extract each from input images and then consider how to integrate them.
There is yet another issue with the region features, usually obtained by a CNN-based detector. At the last stage of its computation, CNN-based detectors employ non-maximum suppression (NMS) to eliminate redundant bounding boxes. This makes the end-to-end training of the entire model hard, i.e., jointly training the decoder part of the image captioning model and the detector by minimizing a single loss. Recent studies detach the two parts in training; they first train a detector on the object detection task and then train only the decoder part on image captioning. This could be a drag on achieving optimal performance of image captioning.
To overcome this limitation of CNN-based detectors and also cope with their high-computational cost, we employ the framework of DETR , which does not need NMS. We choose Deformable DETR , an improved variant, for its high performance, and also replace a CNN backbone used in the original design with Swin Transformer to extract initial features from the input image. We also obtain the grid features from the same Swin Transformer. We input its last layer features into a simple self-attention Transformer and update them to obtain our grid features. This aims to model spatial interaction between the grid features, retrieving contextual information absent in our region features.
The extracted two types of features are fed into the second half of the model, the caption generator. We design it as a lightweight Transformer generating a caption sentence in an autoregressive manner. It is equipped with a unique cross-attention mechanism that computes and applies attention from the two types of visual features to caption sentence words.
These components form a Transformer-only neural architecture, dubbed GRIT (Grid- and Region-based Image captioning Transformer). Our experimental results show that GRIT has established a new state-of-the-art on the standard image captioning benchmark of COCO . Specifically, in the offline evaluation using the Karpathy test split, GRIT outperforms all the existing methods without vision and language (V&L) pretraining. It also performs at least on a par with SimVLMhuge leveraging V&L pretraining on 1.8B image-text pairs.
Related Work
Recent image captioning methods typically employ an encoder-decoder architecture. Specifically, given an image, the encoder extracts visual features; the decoder receives the visual features as inputs and generates a sequence of words. Early methods use a CNN to extract a global feature as a holistic representation of the input image . Although it is simple and compact, this holistic representation suffers from information loss and insufficient granularity. To cope with this, several studies employed more fine-grained grid-based features to represent input images and also used attention mechanisms to utilize the granularity for better caption generation. Later, Anderson et al. introduced the method of using an object detector, such as Faster R-CNN, to extract object-oriented features, called region features, showing that this leads to performance improvement in many V&L tasks, including image captioning and visual question answering. Since then, region features have become the de facto choice of visual representation for image captioning. Pointing out the high computational cost of the region features, Jiang et al. showed that the grid features extracted by an object detector perform well on the VQA task. RSTNet has recently applied these grid features to image captioning.
2 Application of Transformer in Vision/Language Tasks
Transformer has long been a standard neural architecture in natural language processing , and started to be extended to computer vision tasks. Besides ViT for image classification, it was also applied to object detection, leading to DETR , followed by several variants . A recent study applied the framework of DETR to pretraining for various V&L tasks, where they did not use it to obtain the region features.
Transformer has been applied to image captioning, where it is used as an encoder for extracting and encoding visual features and a decoder for generating captions. Specifically, Yang et al. proposed to use the self-attention mechanism to encode visual features. Li et al. used Transformer for obtaining the region features in combination with a semantic encoder that exploits knowledge from an external tagger. Several following studies proposed several variants of Transformer tailored to image captioning, such as Attention on Attention , X-Linear Attention , Memory-augmented Attention , etc. Transformer is naturally employed also as a caption decoder .
Grid- and Region-based Image captioning Transformer
This section describes the architecture of GRIT (Grid- and Region-based Image captioning Transformer). It consists of two parts, one for extracting the dual visual features from an input image (Sec. 3.1) and the other for generating a caption sentence from the extracted features (Sec. 3.2).
A lot of efforts have been made to apply the Transformer architecture to various computer vision tasks since ViT applied it to image classification. ViT divides an input image into small patches and computes global attention over them. This is not suitable for tasks requiring spatially dense prediction, e.g., object detection since the computational complexity increases quadratically with the image resolution.
Swin Transformer mitigates this issue to a great extent by incorporating operations such as patch reduction and shifted windows that support local attention. It is currently a de facto standard as a backbone network for various computer vision tasks. We employ it to extract initial visual features from the input image in our model.
We briefly summarize its structure, explaining how we extract features from the input image and send them to the components following the backbone. Given an input image of resolution , Swin Transformer computes and updates feature maps through multiple stages; it uses the patch merging layer after every stage (but the last stage) to downsample feature maps in their spatial dimension by the factor of 2. We apply another patch merging layer to downsample the last layer’s feature map. We then collect the feature maps from all the stages, obtaining four multi-scale feature maps, i.e., where , which have the resolution from to . These are inputted to the subsequent modules, i.e., the object detector and the network for generating grid features.
1.2 Generating Region Features
As in previous image captioning methods, ours also rely on an object detector to create region features. However, we employ a Transformer-based decoder framework, i.e., DETR instead of CNN-based detectors, such as Faster R-CNN, which is widely employed by the SOTA image captioning models . DETR formulates object detection as a direct set prediction problem, which makes the model free of the unideal computation for us, i.e., NMS and RoI alignment. This enables the end-to-end training of the entire model from the input image to the final output, i.e., a generated caption, and also leads to a significant reduction in computational time while maintaining the model’s performance on image captioning compared with the SOTA models.
Specifically, we employ Deformable DETR , a variant of DETR. Deformable DETR extracts multi-scale features from an input image with its encoder part, which are fed to the decoder part. We use only the decoder part, to which we input the multi-scale features from the Swin Transformer backbone. This leads to further reduction in computational time. We will refer this decoder part as “object detector” in what follows; see Fig. 2.
Although we train it as a part of our entire model, we pretrain our “object detector” including the vision backbone on object detection before the training of image captioning. For the pretraining, we follow the procedure of Deformable DETR; placing a three-layer MLP and a linear layer on its top to predict box coordinates and class category, respectively. We then minimize a set-based global loss that forces unique predictions via bipartite matching.
Following , we pretrain the model (i.e., our object detector including the vision backbone) in two steps. We first train it on object detection following the training method of Deformable DETR. We then fine-tune it on a joint task of object detection and object attribute prediction, aiming to make it learn fine-grained visual semantics with the following loss:
where and are the attribute and class probabilities, , is the loss for normalized bounding box regression for object .
1.3 Grid Feature Network
2 Caption Generation Using Dual Visual Features
The caption generator consists of a stack of identical layers. The initial layer receives the sequence of predicted words and the output from the last layer is input to a linear layer whose output dimension equals the vocabulary size to predict the next word.
Each transformer layer has a sub-layer of masked self-attention over the sentence words and a sub-layer(s) of cross-attention between them and the visual features in this order, followed by a feedforward network (FFN) sub-layer. The masked self-attention sub-layer at the -th layer receives an input sequence at time step , and computes and applies self-attention over the sequence to update the tokens with the attention mask to prevent the interaction from the future words during training.
The cross-attention sub-layer in the layer , located after the self-attention sub-layer, fuses its output with the dual visual features by cross-attention between them, yielding . We consider the three design choices shown in Fig. 3 and described below. We examine their performance through experiments.
2.2 Cross-attention between Caption Word and Dual Visual Features
We show three designs of cross-attention between the word features and the dual visual features (i.e., the region features and the grid features ) as below.
The simplest approach is to concatenate the two visual features and use the resultant features as keys and values in the standard multi-head attention sub-layer, where the words serve as queries; see Fig. 3(a).
Another approach is to perform cross-attention computation separately for the two visual features. The corresponding design is to place two independent multi-head attention sub-layers in a sequential fashion, and uses one for the grid features and the other for the region features (or the opposite combination); see Fig. 3(b). Note that their order could affect the performance.
The third approach is to perform multi-head attention computation on the two visual features in parallel. To do so, we use two multi-head attention mechanisms with independent learnable parameters. The detailed design is as follows. Let be the word features inputted to the meta-layer containing this cross attention sub-layer. As shown in Fig. 2, they are first input to the self-attention sub-layer, converted into (layer index omitted for brevity) and then input to this cross attention sub-layer. In this sub-layer, multi-head attention (MHA) is computed with as queries and the region features as keys and values, yielding attended features . The same computation is performed in parallel with the grid features as keys and values, yielding . Next, we concatenate them with as and , projecting them back to -dimensional vector using learnable affine projections. Normalizing them with sigmoid into probabilities and , respectively, we have
We then multiply them with and , add the resultant vectors to , and finally feed to layer normalization, obtaining as follows:
2.3 Caption Generator Losses
Following a standard practice of image captioning studies, we pre-train our model with a cross-entropy loss (XE) and finetune it using the CIDEr-D optimization with self-critical sequence training strategy . Specifically, the model is first trained to predict the next word at , given the ground-truth sentence . This is equal to minimize the following XE loss with respect to the model’s parameter :
We then finetune the model with the CIDEr-D optimization, where we use the CIDEr score as the reward and the mean of the rewards as the reward baseline, following . The loss for self-critical sequence training is given by
where is the -th sentence in the beam; is the reward function; and is the reward baseline; and is the number of samples in the batch.
Experiments
As mentioned earlier, we train our object detector (including the backbone) in two steps. In the first step, we train it on object detection using either Visual Genome or a combination of four datasets: COCO , Visual Genome, Open Images , and Object365 , depending on what previous methods we experimentally compare. In the second step, we train the model on object detection plus attribute prediction using Visual Genome. Note that following the standard practice, we exclude the duplicated samples appearing in the testing and validation splits of the COCO and nocaps datasets to remove data contamination. See the supplementary material for more details.
1.2 Image Captioning
We conduct our experiments on the COCO dataset, the standard for the research of image captioning . The dataset contains 123,287 images, each annotated with five different captions. For offline evaluation, we follow the widely adopted Karpathy split , where 113,287, 5,000, and 5,000 images are used for training, validation, and testing respectively.
To test our method’s effectiveness on other image captioning datasets, we also report the performances on the nocaps dataset and the Artemis dataset . See the supplementary material for more details.
2 Implementation Details
We employ the standard evaluation protocol for the evaluation of methods. Specifically, we use the full set of captioning metrics: BLEU@N , METEOR , ROUGE-L , CIDEr , and SPICE . We will use the abbreviations, B@N, M, R, C, and S, to denote BLEU@N, METEOR, ROUGE-L, CIDEr, and SPICE, respectively.
2.2 Hyperparameters Settings
In our model, we set the dimension of each layer to , the number of heads to eight. We employ dropout with the dropout rate of on the output of each MHA and FFN sub-layer following . We set the number of layers as for the object detector, as for the grid feature network, and as for the caption generator. Following previous studies, we convert all the captions to lower-case, remove punctuation characters, and perform tokenization with the SpaCy toolkit . We build the vocabularies, excluding the words which appear less than five times in the training and validation splits.
3 Training Details
In the first stage, we pretrain the object detector with the backbone. We consider several existing region-based methods for comparison, which employ similar pretraining of an object detector but use different datasets. For a fair comparison, we consider two settings. One uses Visual Genome for training, following most previous methods. We train our detector for 150,000 iterations with a batch size of 32. The other (results indicated with in what follows) uses the four datasets mentioned above, following . We train the detector for 125,000 iterations with a batch size of 256. In both settings, the input image is resized so that the maximum for the shorter side is 800 and for the longer side is 1333. We use Adam optimizer with a learning rate of , decreased by 10 at iteration 120,000 and 100,000 in the first and second settings, respectively. We follow for other training procedures. After this, we finetune the models on object detection plus attribute prediction using Visual Genome for additional five epochs with a learning rate of , following . The supplementary material presents the details of implementation and experimental results on object detection.
3.2 Second Stage
We train the entire model for the image captioning task in the second stage. We employ the standard method for word representation, i.e., linear projections of one-hot vectors to vectors of dimension = 512. In this stage, we resize all the input images so that the maximum dimensions for the shorter side and longer side are 384 and 640, respectively. We train models, as explained earlier. Specifically, we train models with the cross-entropy loss for ten epochs, in which we warmp up the learning rates for the grid feature network and the caption generator from to in the first epoch, while we fix those for the backbone network and the object detector at . Then, we finetune the model based on the CIDEr-D optimization for ten epochs, where we set the fixed learning rate to for the entire model. We use the Adam optimizer with a batch size of . For the CIDEr-D optimization, we use beam search with a beam size of 5 and a maximum length of 20.
4 Performance of Different Configurations
Our method has several design choices. We conduct experiments to examine which configuration is the best. The results are shown in Table 1. We used an identical configuration unless otherwise noted. Specifically, we use the feature extractor pretrained on the four datasets and parallel cross-attention for fusing the region and grid features.
The first block of Table 1(a) shows the effects of different (pre)training strategies of the visual backbone on image captioning performance. The ‘ImageNet’ column shows the result of the model using a Swin Transformer backbone pretrained on ImageNet21K and the grid features alone; ‘VG’ and ‘4DS’ indicate the models with a detector pretrained on Visual Genome and the four datasets, respectively. They show that using more datasets leads to better performance.
The second block of Table 1(a) shows the effects of the number of object queries, or equivalently region features. The performance increases as they vary as 50, 100, and 150. We also confirmed that the performance is saturated for more region features, while the computational cost and false detection increase.
The third block shows the effect of the end-to-end training of the entire model. ‘Yes’ indicates the end-to-end training of the entire model and ‘No’ indicates training the model but the vision backbone. The results show that the end-to-end training considerably improves CIDEr score (from 139.6 to 144.3) with little sacrifice of B@4. This validates our expectation about the effectiveness of the end-to-end training; it arguably helps reduce the domain gap between object detection and image captioning.
The first block of Table 1(b) shows the performances of the model employing the concatenated cross-attention and its two variants using the grid features alone or the region features alone. They show that the region features alone work better than the grid features alone, and their fusion achieves the highest performance.
The three blocks of Table 1(b) show the performances of the three cross-attention architectures explained in Sec. 3.2. The second block shows the two variants of the sequential cross-attention, and the third block shows the two variants of the parallel cross-attention with different gated activation functions, i.e., sigmoid and identity. By identity activation, we mean setting all the values of and in Eq.(4) to one. These results show that the parallel cross-attention with sigmoid activation function performs the best; the sequential cross-attention in the order attains the second best result.
5 Results on the COCO Dataset
We next show complete results on the COCO dataset by the offline and online evaluations. We present example results in the supplementary material.
Table 2 shows the performances of our method and the current state-of-the-art methods on the offline Karpathy test split. The compared methods are as follows: grid-based methods , region-based methods , the methods employing both grid and region features , and also the methods relying on large-scale pretraining on vision and language (V&L) tasks using a large image-text corpus , including SimVLMhuge, a model pretrained on an extremely large dataset (i.e., 1.8 billion image-caption pairs) .
For fair comparison with the region-based methods, we report the results of two variants of our model, one with the object detector pretrained on Visual Genome alone and the other (marked with †) with the object detector pretrained on the four datasets, as explained earlier. It is seen from Table 2 that our models, regardless of the datasets used for the detector’s pretraining, outperform all the methods that do not use large-scale pretraining of vision and language tasks (i.e., the methods in the second block entitled ‘w/o VL pretraining’). Moreover, our model with the detector pretrained solely on Visual Genome (i.e., ‘GRIT’) performs better than those relying on large-scale V&L pretraining but SimVLMhuge. Finally, our model with the pretrained detector on multiple datasets (i.e., ‘GRIT†’) outperforms SimVLMhuge leveraging large-scale V&L pretraining in CIDEr score (i.e., 144.2 vs 143.3).
5.2 Online Evaluation
We also evaluate our models (i.e., a single model and an ensemble of six models) on the 40K testing images by submitting their results on the official evaluation server. Table LABEL:tab:online_test shows the results and those of all the published methods on the leaderboard. Table LABEL:tab:online_test presents the metric scores based on five (c5) and 40 reference captions (c40) per image. We can see that our method achieves the best scores for all the metrics. Note that even our single model outperforms all the published methods that use ensembles.
6 Results on the ArtEmis and nocaps Datasets
As explained above, we evaluate our method on the ArtEmis and nocaps datasets. For nocaps, we evaluate zero-shot inference performance, i.e., the performance of the model trained on COCO. For ArtEmis, we train the model in the same way as COCO except for the number of training epochs, precisely, five epochs each for the training with the XE loss and that with the CIDEr-D optimization.
Table 4(a) shows the results of our method on the test split of ArtEmis . It also show the results of existing methods reported in , which are grid-based , region-based , and a nearest neighbor method using a holistic vector to encode images (denoted as ). Our method outperforms all these methods by a large margin.
Table 4 shows the results on the nocaps dataset, including the baseline methods reported in . All the models are trained on the training split of the COCO datasets and tested on the validation split of nocaps, which consists of images with novel objects and captions with unseen vocabularies. Our method surpasses all the other methods including region-based methods in both in-domain and out-of-domain images. See the supplementary material for the full results.
7 Computational Efficiency
We measured the inference time of GRIT and two representative region-based methods, VinVL and Transformer . It is the computational time per image from image input to caption generation. Specifically, we measured the time to generate a caption of length 20 with a beam size of five on a V100 GPU. The input image resolution was set to for VinVL and Transformer as reported in . We set it to for GRIT since it already achieves higher accuracy. Figure 1 shows the breakdown of the inference time for the three methods. GRIT reduces the time for feature extraction by a factor of 10 compared with the others. Similar to Transformer, GRIT has a lightweight caption generator and thus spends much less time than VinVL for generating a caption after receiving the visual features. GRIT can run with minibatch size up to on a single V100 GPU, while others cannot afford large minibatch. With minibatch size , the per-image inference time decreases to about 32ms. More details are given in the supplementary material.
Summary and Conclusion
In this paper, we have proposed a Transformer-based architecture for image captioning named GRIT. It integrates the region features and the grid features extracted from an input image to extract richer visual information from input images. Previous SOTA methods employ a CNN-based detector to extract region features, which prevents the end-to-end training of the entire model and makes to high computational costs. Using the Swin Transformer for a backbone extracting the initial visual feature, GRIT resolves these two issues by employing a DETR-based detector. Furthermore, GRIT obtains grid features by updating the feature from the same backbone using a self-attention Transformer, aiming to extract richer context information complementing the region feature. These two features are fed to the caption generator equipped with a unique cross-attention mechanism, which computes and applies attention from the dual features on the generated caption sentence. The integration of all these components led to significant performance improvement. The experimental results validated our approach, showing that GRIT outperforms all published methods by a large margin in inference accuracy and speed.
This work was supported by JST [Moonshot Research and Development], Grant Number [JPMJMS2032] and by JSPS KAKENHI Grant Number 20H05952 and 19H01110.
Appendix 0.A Additional Details for Object Detection
When pretraining our model on the four datasets (i.e., Visual Genome (VG), COCO, OpenImages, and Objects365), we follow to build a unified training corpus with the statistics shown in Table 5 except that we do not use the annotations from COCO stuff . The resultant corpus has images with 1848 categories.
A.2 Implementation Details
For the object detector, we set the number of queries , the number of sampling points equal to 4, and the hidden dimension . The backbone network weights are intialized by the weights of Swin-Base () pretrained on ImageNet21K . Following , the loss for normalized bounding box regression for object , , is computed as the weighted summation of a box distance and a GIoU loss :
where , , and outputs the largest box covering and . We also employ two training strategies, i.e., iterative bounding box refinement and auxiliary losses; see and our configuration files for details.
A.3 Object Detection Results
Table 6 shows the performance on the COCO validation split and the Visual Genome test split of our object detector compared with VinVL and BUTD . It is seen that the object detector of GRIT attains comparable or higher performance on the two datasets as compared with BUTD and VinVL when pretrained on the similar datasets.
Appendix 0.B Additional Details for Image Captioning
B.0.2 Boundary Tokens
Following previous studies, we prepend a special token to the beginning of captions, and append another special token to the end of captions during training. During inference, we start the generation by setting the first token to .
B.1 Image Captioning on the COCO dataset
Table 7 reports a breakdown of SPICE F-scores over various sub-categories on the “Karpathy” test split, in comparison with the region-based methods: Up-Down , vanilla Transformer , and Transformer . These scores give a quantitative assessment of performance on different aspects when describing the content of images. As seen in Table 7, our method attains better scores over all sub-categories, showing significant improvement on identifying and counting objects, attributes, and relationships between objects. The table also reports the CLIP scores of the two methods, showing consistent improvement of our method over the compared method.
B.2 Image Captioning on the ArtEmis dataset
This dataset consists of 80,031 unique images divided into the training, validation, and test splits with the ratios of 85%, 5%, and 10%, respectively. Each caption of a given image is annotated with an emotion label. In total, there are 454,684 captions along with 8 unique emotion categories; see for details.
B.2.2 Emotion Grounded Model
Following , we also trained an emotion grounded model, which predicts the emotion associated with the caption. Specifically, we mapped the updated class embedding into an -dimensional vector using a linear projection. During training, we minimized the summation of the two losses, i.e., emotion prediction and caption generation.
B.2.3 Full Results
Table 8 shows the full results of different models on the test split of the Artemis dataset including the emotion grounded models. It is noted that the ground truth emotion labels are not provided during inference.
B.3 Image Captioning on the nocaps Dataset
We report the full results on the validation split of the nocaps dataset for different domains, i.e., in-domain, near-domain, and out-of-domain, in Table 9.
B.4 Computational Efficiency
We measured the inference time of GRIT and two representative region-based methods, VinVL and Transformer , on the same machine having a Tesla V100-SXM2 of 16GB memory with CUDA version 10.0 and Driver version 410.104. It has Intel(R) Xeon(R) Gold 6148 CPU. The comparison was conducted following . Specifically, we excluded the time of preprocessing the image and loading it to the GPU device. Also, the images are rescaled to the resolutions such that all the compared methods achieve its highest performance for image captioning. For the compared methods, we used the official implementations of Transformerhttps://github.com/aimagelab/meshed-memory-transformer and VinVLhttps://github.com/pzzhang/VinVL.
Regarding feature extraction, we extracted the region features from Faster R-CNN using the original implementationhttps://github.com/peteanderson80/bottom-up-attention used by Transformer and another implementationhttps://github.com/microsoft/scene_graph_benchmark used by VinVL. It is seen that VinVL and Transformer spend considerable time on feature extraction due to the forward pass through the CNN backbone with high resolution inputs and the computationally expensive regional operations. It is also noted that VinVL introduced class-agnostic NMS operations, which reduce a great amount of time consumed by class-aware NMS operations in the standard Faster R-CNN. On the other hand, we employ a Deformable DETR-based detector to extract region features without using all such operations. Table 10 shows the comparison on feature extraction.
Regarding caption generation, all the methods use beam search as the decoding strategy, with beam size of 5 and the maximum caption length of 20. Both Transformer and GRIT employ a lightweight caption generator (caption decoder) having only 3 transformer layers with hidden dimension of 512 while VinVLlarge has 24 transformer layers with hidden dimension of 1024; see Table 11. Thus, with the visual features as inputs, Transformer and GRIT spend less inference time generating words than VinVLlarge in the autoregressive manner.
B.5 Qualitative Examples
Figure 4, 5, 6, and 7 show some examples of the captions generated by our proposed method (GRIT) and another region-based method ( Transformer) given the same input images from the COCO test split. It is observed that the generated captions from GRIT are qualitatively better than those generated by the baseline method in terms of detecting and counting objects as well as describing their relationships in the given images. The inaccuracy of the captions generated by the baseline method might be due to the drawbacks of the region features extracted by a frozen pretrained object detector which produces wrong detection and lacks of contextual information.