Show, Control and Tell: A Framework for Generating Controllable and Grounded Captions

Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara

Introduction

Image captioning brings vision and language together in a generative way. As a fundamental step towards machine intelligence, this task has been recently gaining much attention thanks to the spread of Deep Learning architectures which can effectively describe images in natural language . Image captioning approaches are usually capable of learning a correspondence between an input image and a probability distribution over time, from which captions can be sampled either using a greedy decoding strategy , or more sophisticated techniques like beam search and its variants .

As the two main components of captioning architectures are the image encoding stage and the language model, researchers have focused on improving both phases, which resulted in the emergence of attentive models on one side, and of more sophisticated interactions with the language model on the other . Recently, attentive models have been improved by replacing the attention over a grid of features with attention over image regions . In these models, the generative process attends a set of regions which are softly selected while generating the caption.

Despite these advancements, captioning models still lack controllability and explainability – i.e., their behavior can hardly be influenced and explained. As an example, in the case of attention-driven models, the architecture implicitly selects which regions to focus on at each timestep, but it cannot be supervised from the exterior. While an image can be described in multiple ways, such an architecture provides no way of controlling which regions are described and what importance is given to each region. This lack of controllability creates a distance between human and machine intelligence, as humans can manage the variety of ways in which an image can be described, and select the most appropriate one depending on the task and the context at hand. Most importantly, this also limits the applicability of captioning algorithms to complex scenarios in which some control over the generation process is needed. As an example, a captioning-based driver assistance system would need to focus on dangerous objects on the road to alert the driver, rather than describing the presence of trees and cars when a risky situation is detected. Eventually, such systems would also need to be explainable, so that their behavior could be easily interpreted in case of failures.

In this paper, we introduce Show, Control and Tell, that explicitly addresses these shortcomings (Fig. 1). It can generate diverse natural language captions depending on a control signal which can be given either as a sequence or as a set of image regions which need to be described. As such, our method is capable of describing the same image by focusing on different regions and in a different order, following the given conditioning. Our model is built on a recurrent architecture which considers the decomposition of a sentence into noun chunks and models the relationship between image regions and textual chunks, so that the generation process can be explicitly grounded on image regions. To the best of our knowledge, this is the first captioning framework controllable from image regions.

Contributions. Our contributions are as follows:

We propose a novel framework for image captioning which is controllable from the exterior, and which can produce natural language captions explicitly grounded on a sequence or a set of image regions.

The model explicitly considers the hierarchical structure of a sentence by predicting a sequence of noun chunks. Also, it takes into account the distinction between visual and textual words, thus providing an additional grounding at the word level.

We evaluate the model with respect to a set of carefully designed baselines, on Flickr30k Entities and on COCO, which we semi-automatically augment with grounding image regions for training and evaluation purposes.

Our proposed method achieves state of the art results for controllable image captioning on Flick30k and COCO both in terms of diversity and caption quality, even when compared with methods which focus on diversity.

Related work

A large number of models has been proposed for image captioning . Generally, all integrate recurrent neural networks as language models, and a representation of the image which might be given by the output of one or more layer of a CNN , or by a time-varying vector extracted with an attention mechanism selected either from a grid over CNN features, or integrating image regions eventually extracted from a detector . Attentive models provided a first way of grounding words to parts of the image, although with a blurry indication which was rarely semantically significant. Regarding the training strategies, notable advances have been made by using Reinforcement Learning to train non-differentiable captioning metrics . In this work, we propose an extended version of this approach which deals with multiple output distributions and rewards the alignment of the caption to the control signal.

Recently, more principled approaches have been proposed for grounding a caption on the image : DenseCap generates descriptions for specific image regions. Further, the Neural Baby Talk approach extends the attentive model in a two-step design in which a word-level sentence template is firstly generated and then filled by object detectors with concepts found in the image. We instead decompose the caption at the level of noun chunks, and explicitly ground each of them to a region. This approach has the additional benefit of providing an explicability method at the chunk level.

Another related line of work is that of generating diverse descriptions. Some works have extended the beam-search algorithm to sample multiple captions from the same distribution , while different GAN-based approaches have also appeared . Most of these improve on diversity, but suffer on accuracy and do not provide controllability over the generation process. Others have conditioned the generation with a specific style or sentiment . Our work is mostly related to , which uses a control input as a sequence of part-of-speech tags. This approach, while generating diversity, is hardly employable to effectively control the generation of the sentence; in contrast, we use image regions as a controllability method.

Method

Sentences are natural language structures which are hierarchical by nature . At the lowest level, a sentence might be thought as a sequence of words: in the case of a sentence describing an image, we can further distinguish between visual words, which describe something visually present in the image, and textual words, that refer to entities which are not present in the image . Analyzing further the syntactic dependencies between words, we can recover a higher abstraction level in which words can be organized into a tree-like structure: in a dependency tree , each word is linked together with its modifiers (Fig. 3).

Given a dependency tree, nouns can be grouped with their modifiers, thus building noun chunks. For instance, the caption depicted in Fig. 3 can be decomposed into a sequence of different noun chunks: “a young boy”, “a cap”, “his head”, “striped shirt”, and “gray and sweat jacket”. As noun chunks, just like words, can be visually grounded into image regions, a caption can also be mapped to a sequence of regions, each corresponding to a noun chunk. A chunk might also be associated with multiple image regions of the same class if more than one possible mapping exists.

The number of ways in which an image can be described results in different sequences of chunks, linked together to form a fluent sentence. Therefore, captions also differ in terms of the set of considered regions, the order in which they are described, and their mapping to chunks given by the linguistic abilities of the annotator.

Following these premises, we define a model which can recover the variety of ways in which an image can be described, given a control input expressed as a sequence or set of image regions. We begin by presenting the former case, and then show how our model deals with the latter scenario.

Given an image I\bm{I} and an ordered sequence of set of regions R=(r0,r1,...,rN)\bm{R}=(\bm{r}_{0},\bm{r}_{1},...,\bm{r}_{N})For generality, we will always consider sequences of sets of regions, to deal with the case in which a chunk in the target sentence can be associated to multiple regions in training and evaluation data., the goal of our captioning model is to generate a sentence y=(y0,y1,...,yT)\bm{y}=\left(y_{0},y_{1},...,y_{T}\right) which in turns describes all the regions in R\bm{R} while maintaining the fluency of language.

Our model is conditioned on both the input image I\bm{I} and the sequence of region sets R\bm{R}, which acts as a control signal, and jointly predicts two output distributions which correspond to the word-level and chunk-level representation of the sentence: the probability of generating a word at a given time, i.e. p(yt∣R,I;θ)p(y_{t}|\bm{R},\bm{I};\bm{\theta}), and that of switching from one chunk to another, i.e. p(gt∣R,I;θ)p(g_{t}|\bm{R},\bm{I};\bm{\theta}), where gtg_{t} is a boolean chunk-shifting gate. During the generation, the model maintains a pointer to the current region set ri\bm{r}_{i} and can shift to the next element in R\bm{R} by means of the gate gtg_{t}.

To generate the output caption, we employ a recurrent neural network with adaptive attention. At each timestep, we compute the hidden state ht\bm{h}_{t} according to the previous hidden state ht−1\bm{h}_{t-1}, the current image region set rt\bm{r}_{t} and the current word wtw_{t}, such that ht=RNN(wt,rt,ht−1)\bm{h}_{t}=\text{RNN}(w_{t},\bm{r}_{t},\bm{h}_{t-1}). At training time, rt\bm{r}_{t} and wtw_{t} are the ground-truth region set and word corresponding to timestep tt; at test time, wtw_{t} is sampled from the first distribution predicted by the model, while the choice of the next image region is driven by the values of the chunk-shifting gate sampled from the second distribution:

where {gk}k\{g_{k}\}_{k} is the sequence of sampled gate values, and NN is the number of region sets in R\bm{R}.

Chunk-shifting gate. We compute p(gt∣R)p(g_{t}|\bm{R}) via an adaptive mechanism in which the LSTM computes a compatibility function between its internal state and a latent representation which models the state of the memory at the end of a chunk. The compatibility score is compared to that of attending one of the regions in rt\bm{r}_{t}, and the result is used as an indicator to switch to the next region set in R\bm{R}.

The LSTM is firstly extended to obtain a chunk sentinel stc\bm{s}^{c}_{t}, which models a component extracted from the memory encoding the state of the LSTM at the end of a chunk. The sentinel is computed as:

We then compute a compatibility score between the internal state ht\bm{h}_{t} and the sentinel vector through a single-layer neural network; analogously, we compute a compatibility function between ht\bm{h}_{t} and the regions in rt\bm{r}_{t}.

The probability of shifting from one chunk to the next one is defined as the probability of attending the sentinel vector stc\bm{s}^{c}_{t} in a distribution over stc\bm{s}^{c}_{t} and the regions in rt\bm{r}_{t}:

where ztir\bm{z}^{r}_{ti} indicates the ii-th element in ztr\bm{z}^{r}_{t}, and we dropped the dependency between nn and tt for clarity. At test time, the value of gate gt∈{0,1}g_{t}\in\{0,1\} is then sampled from p(gt∣R)p(g_{t}|\bm{R}) and drives the shifting to the next region set in R\bm{R}.

Adaptive attention with visual sentinel. While the chunk-shifting gate predicts the end of a chunk, thus linking the generation process with the control signal given by R\bm{R}, once rt\bm{r}_{t} has been selected a second mechanism is needed to attend its regions and distinguish between visual and textual words. To this end, we build an adaptive attention mechanism with a visual sentinel .

The visual sentinel vector models a component of the memory to which the model can fall back when it chooses to not attend a region in rt\bm{r}_{t}. Analogously to Eq. 2, it is defined as:

where [⋅][\cdot] indicates concatenation. Based on the attention distribution, we obtain a context vector which can be fed to the LSTM as a representation of what the network is attending:

Notice that the context vector will be, mostly, an approximation of one of the regions in rt\bm{r}_{t} or the visual sentinel. However, rt\bm{r}_{t} will vary at different timestep according to the chunk-shifting mechanism, thus following the control input. The model can alternate the generation of visual and textual words by means of the visual sentinel.

The captioning model is trained using a loss function which considers the two output distributions of the model. Given the target ground-truth caption y1:T∗\bm{y}^{*}_{1:T}, the ground-truth region sets r1:T∗\bm{r}^{*}_{1:T} and chunk-shifting gate values corresponding to each timestep g1:T∗g^{*}_{1:T}, we train both distributions by means of a cross-entropy loss. The relationship between target region sets and gate values will be further expanded in the implementation details. The loss function for a sample is defined as:

Following previous works , after a pre-training step using cross-entropy, we further optimize the sequence generation using Reinforcement Learning. Specifically, we use the self-critical sequence training approach , which baselines the REINFORCE algorithm with the reward obtained under the inference model at test time.

Given the nature of our model, we extend the approach to work on multiple output distributions. At each timestep, we sample from both p(yt∣R)p(y_{t}|\bm{R}) and p(gt∣R)p(g_{t}|\bm{R}) to obtain the next word wt+1w_{t+1} and region set rt+1\bm{r}_{t+1}. Once a EOS tag is reached, we compute the reward of the sampled sentence ws\bm{w}^{s} and backpropagate with respect to both the sampled word sequence ws\bm{w}^{s} and the sequence of chunk-shifting gates gs\bm{g}^{s}. The final gradient expression is thus:

where b=r(w^)b=r(\hat{\bm{w}}) is the reward of the sentence obtained using the inference procedure (i.e. by sampling the word and gate value with maximum probability). We then build a reward function which jointly considers the quality of the caption and its alignment with the control signal R\bm{R}.

Rewarding caption quality. To reward the overall quality of the generated caption, we use image captioning metrics as a reward. Following previous works , we employ the CIDEr metric (specifically, the CIDEr-D score) which has been shown to correlate better with human judgment .

Rewarding the alignment. While captioning metrics can reward the semantic quality of the sentence, none of them can evaluate the alignment with respect to the control inputAlthough METEOR creates an alignment with respect to the reference caption, this is done for each unigram, thus mixing semantic and alignment errors.. Therefore, we introduce an alignment score based on the Needleman-Wunsch algorithm .

Given a predicted caption y\bm{y} and its target counterpart y∗\bm{y}^{*}, we extract all nouns from both sentences, and evaluate the alignment between them, recalling the relationships between noun chunks and region sets. We use the following scoring system: the reward for matching two nouns is equal to the cosine similarity between their word embeddings; a gap gets a negative reward equal to the minimum similarity value, i.e. −1-1. Once the optimal alignment is computed, we normalize its score, al(y,y∗)al(\bm{y},\bm{y}^{*}) with respect to the length of the sequences. The alignment score is thus defined as:

where #y\#\bm{y} and #y∗\#\bm{y}^{*} represent the number of nouns contained in y\bm{y} and y∗\bm{y}^{*}, respectively. Notice that NW(⋅,⋅)∈\text{NW}(\cdot,\cdot)\in. The final reward that we employ is a weighted version of CIDEr-D and the alignment score.

The proposed architecture, so far, can generate a caption controlled by a sequence of region sets R\bm{R}. To deal with the case in which the control signal is unsorted, i.e. a set of regions sets, we build a sorting network which can arrange the control signal in a candidate order, learning from data. The resulting sequence can then be given to the captioning network to produce the output caption (Fig. 3).

To this aim, we train a network which can learn a permutation, taking inspiration from Sinkhorn networks . As shown in , the non-differentiable parameterization of a permutation can be approximated in terms of a differentiable relaxation, the so-called Sinkhorn operator. While a permutation matrix has exactly one entry of 1 in each row and each column, the Sinkhorn operator iteratively normalizes rows and columns of any matrix to obtain a “soft” permutation matrix, i.e. a real-valued matrix close to a permutation one.

Given a set of region sets R={r1,r2,...,rN}\mathcal{R}=\{\bm{r}_{1},\bm{r}_{2},...,\bm{r}_{N}\}, we learn a mapping from R\mathcal{R} to its sorted version R∗\bm{R}^{*}. Firstly, we pass each element in R\mathcal{R} through a fully-connected network which processes every item of a region set independently and produces a single output feature vector with length NN. By concatenating together the feature vectors obtained for all region sets, we thus get a N×NN\times N matrix, which is then passed to the Sinkhorn operator to obtain the soft permutation matrix P\bm{P}. The network is then trained by minimizing the mean square error between the scrambled input and its reconstructed version obtained by applying the soft permutation matrix to the sorted ground-truth, i.e. PTR∗\bm{P}^{T}\bm{R}^{*}.

At test time, we take the soft permutation matrix and apply the Hungarian algorithm to obtain the final permutation matrix, which is then used to get the sorted version of R\mathcal{R} for the captioning network.

Language model and image features. We use a language model with two LSTM layers (Fig. 3): the input of the bottom layer is the concatenation of the embedding of the current word, the image descriptor, as well as the hidden state of the second layer. This layer predicts the context vector via the visual sentinel as well as the chunk-gate. The second layer, instead, takes as input the context vector and the hidden state of the bottom layer and predicts the next word.

To represent image regions, we use Faster R-CNN with ResNet-101 . In particular, we employ the model finetuned on the Visual Genome dataset provided by . As image descriptor, following the same work , we average the feature vectors of all the detections.

The hidden size of the LSTM layers is set to 10001000, and that of attention layers to 512512, while the input word embedding size is set to 10001000.

Ground-truth chunk-shifting gate sequences. Given a sentence where each word of a noun chunk is associated to a region set, we build the chunk-shifting gate sequence {gt∗}t\{g^{*}_{t}\}_{t} by setting gt∗g^{*}_{t} to 11 on the last word of every noun chunk, and otherwise. The region set sequence {rt∗}t\{\bm{r}_{t}^{*}\}_{t} is built accordingly, by replicating the same region set until the end of a noun chunk, and then using the region set of the next chunk. To compute the alignment score and for extracting dependencies, we use the spaCy NLP toolkithttps://spacy.io/. We use GloVe as word vectors.

Sorting network. To represent regions, we use Faster R-CNN vectors, the normalized position and size and the GloVe embedding of the region class. Additional details on architectures and training can be found in the Supplementary material.

We experiment with two datasets: Flickr30k Entities, which already contains the associations between chunks and image regions, and COCO, which we annotate semi-automatically. Table 1 summarizes the datasets we use.

Flickr30k Entities . Based on Flickr30k , it contains 31,00031,000 images annotated with five sentences each. Entity mentions in the caption are linked with one or more corresponding bounding boxes in the image. Overall, 276,000276,000 manually annotated bounding boxes are available. In our experiments, we automatically associate each bounding box with the image region with maximum IoU among those detected by the object detector. We use the splits provided by Karpathy et al. .

COCO Entities. Microsoft COCO contains more than 120,000120,000 images, each of them annotated with around five crowd-sourced captions. Here, we again follow the splits defined by and automatically associate noun chunks with image regions extracted from the detector .

We firstly build an index associating each noun of the dataset with the five most similar class names, using word vectors. Then, each noun chunk in a caption is associated by using either its name or the base form of its name, with the first class found in the index which is available in the image. This association process, as confirmed by an extensive manual verification step, is generally reliable and produces few false positive associations. Naturally, it can result in region sets with more than one element (as in Flickr30k), and noun chunks with an empty region set. In this case, we fill empty training region sets with the most probable detections of the image and let the adaptive attention mechanism learn the corresponding association; in validation and testing, we drop those captions. Some examples of the additional annotations extracted from COCO are shown in Fig. 4.

2 Experimental setting

The experimental settings we employ is different from that of standard image captioning. In our scenario, indeed, the sequence of set of regions is a second input to the model which shall be consider when selecting the ground-truth sentences to compare against. Also, we employ additional metrics beyond the standard ones like BLEU-4 , METEOR , ROUGE , CIDEr and SPICE .

When evaluating the controllability with respect to a sequence, for each ground-truth regions-image input (R,I)(\bm{R},\bm{I}), we evaluate against all captions in the dataset which share the same pair. Also, we employ the alignment score (NW) to evaluate how the model follows the control input.

Similarly, when evaluating the controllability with respect to a set of regions, given a set-image pair (R,I)(\mathcal{R},\bm{I}), we evaluate against all ground-truth captions which have the same input. To assess how the predicted caption covers the control signal, we also define a soft intersection-over-union (IoU) measure between the ground-truth set of nouns and its predicted counterpart, recalling the relationships between region sets and noun chunks. Firstly, we compute the optimal assignment between the two set of nouns, using distances between word vectors and the Hungarian algorithm , and define an intersection score between the two sets as the sum of assignment profits. Then, recalling that set union can be expressed in function of an intersection, we define the IoU measure as follows:

where I(⋅,⋅)\text{I}(\cdot,\cdot) is the intersection score, and the #\# operator represents the cardinality of the two sets of nouns.

3 Baselines

Controllable LSTM. We start from a model without attention: an LSTM language model with a single visual feature vector. Then, we generate a sequential control input by feeding a flattened version of R\bm{R} to a second LSTM and taking the last hidden state, which is concatenated to the visual feature vector. The structure of the language model resembles that of , without attention.

Controllable Up-Down. In this case, we employ the full Up-Down model from , which creates an attentive distribution over image regions and make it controllable by feeding only the regions selected in R\bm{R} and ignoring the rest. This baseline is not sequentially controllable.

Ours without visual sentinel. To investigate the role of the visual sentinel and its interaction with the gate sentinel, in this baseline we ablate our model by removing the visual sentinel. The resulting baseline, therefore, lacks a mechanism to distinguish between visual and textual words.

Ours with single sentinel. Again, we ablate our model by merging the visual and chunk sentinel: a single sentinel is used for both roles, in place of stc\bm{s}_{t}^{c} and stv\bm{s}_{t}^{v}.

As further baselines, we also compare against non-controllable captioning approaches, like FC-2K , Up-Down , and Neural Baby Talk .

4 Quantitative results

Controllability through a sequence of detections. Firstly, we show the performance of our model when providing the full control signal as a sequence of region sets. Table 2 shows results on COCO Entities, in comparison with the aforementioned approaches. We can see that our method achieves state of the art results on all automatic evaluation metrics, outperforming all baselines both in terms of overall caption quality and in terms of alignment with the control signal. Using the cross-entropy pre-training, we outperform the Controllable LSTM and Controllable Up-Down by 32.0 on CIDEr and 0.112 on NW. Optimizing the model with CIDEr and NW further increases the alignment quality while maintaining outperforming results on all metrics, leading to a final 0.649 on NW, which outperforms the Controllable Up-Down baseline by a 0.25. Recalling that NW ranges from −1-1 to 11, this improvement amounts to a 12.5%12.5\% of the full metric range.

In Table 3, we instead show the results of the same experiments on Flickr30k Entities, using CIDEr+NW optimization for all controllable methods. Also on this manually annotated dataset, our method outperforms all the compared approaches by a significant margin, both in terms of caption quality and alignment with the control signal.

Controllability through a set of detections. We then assess the performance of our model when controlled with a set of detections. Tables 4 and 5 show the performance of our method in this setting, respectively on COCO Entities and Flickr30k Entities. We notice that the proposed approach outperforms all baselines and compared approaches in terms of IoU, thus testifying that we are capable of respecting the control signal more effectively. This is also combined with better captioning metrics, which indicate higher semantic quality.

Diversity evaluation. Finally, we also assess the diversity of the generated captions, comparing with the most recent approaches that focus on diversity. In particular, the variational autoencoder proposed in and the approach of , which allows diversity and controllability by feeding PoS sequences. To test our method on a significant number of diverse captions, given an image we take all regions which are found in control region sets, and take the permutations which result in captions with higher log-probability. This approach is fairly similar to the sampling strategy used in , even if ours considers region sets. Then, we follow the experimental approach defined in : each ground-truth sentence is evaluated against the generated caption with the maximum score for each metric. Higher scores, thus, indicate that the method is capable of sampling high accuracy captions. Results are reported in Table 6, where to guarantee the fairness of the comparison, we run this experiments on the full COCO test split. As it can be seen, our method can generate significantly diverse captions.

We presented Show, Control and Tell, a framework for generating controllable and grounded captions through regions. Our work is motivated by the need of bringing captioning systems to more complex scenarios. The approach considers the decomposition of a sentence into noun chunks, and grounds chunks to image regions following a control signal. Experimental results, conducted on Flickr30k and on COCO Entities, validate the effectiveness of our approach in terms of controllability and diversity.

Appendix A Sorting network

We provide additional details on the architecture and training strategy of the sorting network. For the ease of the reader, a schema is reported in Fig. 7. Given a scrambled sequence of NN region sets, each region is encoded through a fully connected network which returns a NN-dimensional descriptor. The fully connected network employs visual, textual and geometric features: the Faster R-CNN vector of the detection (20482048-d), the GloVe embedding of the region class (300300-d) and the normalized position and size of the bounding-box (44-d). The visual vector is processed by two layers (512512-d, 128128-d), while the textual feature is processed by a single layer (128128-d). The outputs of the visual and textual branches are then concatenated with the geometric features and fed through another fully connected layer (256256-d). A final layer produces the resulting NN-dimensional descriptors. All layers have ReLU activations, except for the last fully-connected which has a tanh⁡\tanh activation. In case the region set contains more than one detection, we average-pool the resulting NN-dimensional descriptors to obtain a single feature vector for a region set.

Once the feature vectors of the scrambled sequence are concatenated, we get a N×NN\times N matrix, which is then converted into a “soft” permutation matrix P\bm{P} through the Sinkhorn operator. The operator processes a NN-dimensional square matrix X\bm{X} by applying LL consecutive row-wise and column-wise normalization, as follows:

where Tr(X)=X⊘(X1N1NT)\mathcal{T}_{r}(\bm{X})=\bm{X}\oslash(\bm{X}\bm{1}_{N}\bm{1}_{N}^{T}), and Tc(X)=X⊘(1N1NTX)\mathcal{T}_{c}(\bm{X})=\bm{X}\oslash(\bm{1}_{N}\bm{1}_{N}^{T}\bm{X}) are the row-wise and column-wise normalization operators, with ⊘\oslash denoting element-wise division, 1N\bm{1}_{N} a column vector of NN ones. At test time, once LL normalizations (L=20L=20 in our experiments) have been performed, the resulting “soft” permutation matrix can be converted into a permutation matrix via the Hungarian algorithm .

At training time, instead, we measure the mean square error between the scrambled sequence and its reconstructed version obtained by applying the soft permutation matrix to the sorted ground-truth sequence R∗\bm{R}^{*}, i.e. PTR∗\bm{P}^{T}\bm{R}^{*}. On the implementation side, all tensors are appropriately masked to deal with variable-length sequences and sets. We set the maximum length of input scrambled sequences to 1010.

In Table 7 we evaluate the quality of the rankings in terms of accuracy (proportion of completely correct rankings) and Kendall’s Tau (correlation between GT and predicted ranking, between −1-1 and 11). We compare with a predefined local ranking (sorting detections with their probability), a predefined global ranking based on detection classes, and compare the Sinkorn network with a SVM Rank model trained on the same features. As it can be seen, the Sinkhorn network performs better than other baselines and can generate accurate rankings.

Appendix B Training details

We used a weight of 0.2 for the word loss and 0.8 for the two chunk-level terms in Eq. 11. To train both the captioning model and the sorting network, we use the Adam optimizer with an initial learning rate of 5×10−45\times 10^{-4} decreased by a factor of 0.80.8 every epoch. For the captioning model, we run the reinforcement learning training with a fixed learning rate of 5×10−55\times 10^{-5}. We use a batch size of 100100 for all our experiments. During caption decoding, we employ for all experiments the beam search strategy with a beam size of 55: similarly to what has been done when training with Reinforcement Learning, we sample from both output distribution to select the most probable sequence of actions. We use early stopping on validation CIDEr for the captioning network, and validation accuracy of the predicted permutations for the sorting network.

Appendix C The COCO Entities dataset

In Fig. 8, we report additional examples of the semi-automatic annotation procedure used to collect COCO Entities. As in the main paper, we use different colors to visualize the correspondences between noun chunks and image regions. For the ease of visualization, we display a single region for chunk, even though multiple associations are possible. In this case, the region set would contain more than one element.

In the last two rows, we also report samples in which at least one noun chunk could not be assigned to any detection. Recall that in this case, at training time, we use the most probable detections of the image and let the adaptive attention mechanism learn the corresponding association: we found that this procedure, overall, increases the final accuracy of the network rather than feeding empty region sets. Captions with missing associations are dropped in validation and testing.

Appendix D Additional experimental results

Tables 8, 9 and 10 report additional experimental results which have not been reported in the main paper for space constraints. In particular, Table 8 integrates Table 3 of the main paper by evaluating the controllability via a sequence of region sets on Flickr30K, when training with cross-entropy only, and when optimizing with CIDEr and CIDEr+NW. Analogously, Tables 9 and 10 analyze the controllabilty via a set of regions, on both Flickr30K and COCO Entities and with all training strategies.

We observe that the CIDEr+NW fine-tuning approach is effective on all settings, and that our model outperforms by a clear margin the baselines both when controlled via a sequence and when controlled by a scrambled set of regions, regardless of the careful choice of the baselines. The performance of the Controllable LSTM baseline is constantly significantly lower than that of the Controllable Up-Down, thus indicating both the importance of an attention mechanism and that of having a good representation of the control signal. The Controllable Up-Down baseline, however, shows lower performance when compared to our approach, in both sequence- and set-controlled scenarios.

Appendix E Additional qualitative results

Finally, Fig. 9 and 10 report other qualitative results on COCO Entities. As in the main paper, the same image is reported multiple times with different control inputs: our method generates multiple captions for the same image, and can accurately follow the control input.