Grounding Referring Expressions in Images by Variational Context

Hanwang Zhang, Yulei Niu, Shih-Fu Chang

Introduction

Grounding natural language in visual data is a hallmark of AI, since it establishes a communication channel between humans, machines, and the physical world, underpinning a variety of multimodal AI tasks such as robotic navigation , visual Q&A , and visual chatbot . Thanks to the rapid development in deep learning-based CV and NLP, we have witnessed promising results not only in grounding nouns (e.g., object detection ), but also short phrases (e.g., noun phrases and relations ). However, the more general task: grounding referring expressions , is still far from resolved due to the challenges in understanding of both language and scene compositions . As illustrated in Figure 1, given an input referring expression “largest elephant standing behind baby elephant” and an image with region proposals, a model that can only localize “elephant” is not satisfactory as there are multiple elephants. Therefore, the key for referring expression grounding is to comprehend and model the context. Here, we refer to context as the visual objects (e.g., “elephant”), attributes (e.g., “largest” and “baby”), and relationships (e.g., “behind”) mentioned in the expression that help to distinguish the referent from other objects.

One straightforward way of modeling the relations between the referent and context is to: 1) use external syntactic parsers to parse the expression into entities, modifiers, and relations , and then 2) apply visual relation detectors to localize them . However, this two-stage approach is not practical due to the limited generalization ability of the detectors applied in the highly unrestricted language and scene compositions. To this end, recent approaches use multimodal embedding networks that jointly comprehend language and model the visual relations . Due to the prohibitively high cost of annotating both referent and context of referring expressions in images, multiple instance learning (MIL) is usually adopted in them to handle the weak supervision of the unannotated context objects, by maximizing the joint likelihood of every region pair. However, for a referent, the MIL framework essentially oversimplifies the number of context configurations of NN regions from O(2N)\mathcal{O}(2^{N}) to O(N)\mathcal{O}(N). For example, to localize the “elephant” in Figure 1, we may need to consider the other three elephants all together as a multinomial subset for modeling the context such as “largest”, “behind” and “baby elephant”.

In this paper, we propose a novel model called Variational Context for grounding referring expressions in images. Compared to the previous MIL-based approaches , our model approximates the combinatorial context configurations with weak supervision using a variational Bayesian framework . Intuitively, it exploits the reciprocity between referent and context, given either of which can help to localize the other. As shown in Figure 1, for each region xx, we first estimate a coarse context zz, which will help to refine the true localizations of the referent. This reciprocity is formulated into the variational lower-bound of the grounding likelihood p(x∣L)p(x|L), where LL is the text expression and the context is considered as a hidden variable zz (cf. Section 3). Specifically, the model consists of three multimodal modules: context posterior q(z∣x,L)q(z|x,L), referent posterior p(x∣z,L)p(x|z,L), and context prior pz(z∣L)p_{z}(z|L), each of which performs a grounding task (cf. Section 4.3) that aligns image regions with a cue-specific language feature; each cue dynamically encodes different subsets of words in the expression LL that help the corresponding localization (cf. Section 4.2).

Thanks to the reciprocity between referent and context, our model can not only be used in the conventional supervised setting, where there is annotation for referent , but also in the challenging unsupervised setting, where there is no instance-level annotation (e.g., bounding boxes) of both referent and context. We perform extensive experiments on four benchmark referring expression datasets: RefCLEF , RefCOCO , RefCOCO+ , and RefCOCOg . Our model consistently outperforms previous methods in both supervised and unsupervised settings. We also qualitatively show that our model can ground the context in the expression to the corresponding image regions (cf. Section 5).

Related Work

Grounding Referring Expressions. Grounding referring expression is also known as referring expression comprehension, whose inverse task is called referring expression generation . Different from grounding phrases and descriptive sentences , the key for grounding referring expression is to use the context (or pragmatics in linguistics ) to distinguish the referent from other objects, usually of the same category . However, most previous works resort to use holistic context such as the entire image or visual feature difference between regions . Our model is similar to the works on explicitly modeling the referent and context region pairs , however, due to the lack of context annotation, they reduce the grounding task into a multiple instance learning framework . As we will discuss later, this framework is not a proper approximation to the original task. There are also studies on visual relation detection that detect objects and their relationships . However, they are limited to a fixed-vocabulary set of relation triplets and hence are difficult to be applied in natural language grounding. Our cue-specific language feature is similar to the language modular network that learns to decompose a sentence into referent/context-related words, which are different from other approaches that treat the expression as a whole .

Variational Bayesian Model vs. Multiple Instance Learning. Our proposed variational context model is in a similar vein of the deep neural network based variational autoencoder (VAE) , which uses neural networks to approximate the posterior distribution of the hidden value q(z∣x)q(z|x), i.e., encoder, and the conditional distribution of the observation p(x∣z)p(x|z), i.e., decoder. VAE shows efficient and effective end-to-end optimization for the intractable log-sum likelihood log⁡∑zp(x,z)\log\sum_{z}p(x,z) that is widely used in generative processes such as image synthesis and video frame prediction . Considering the unannotated context as the hidden variable zz, the referring expression grounding task can also be formulated into the above log-sum marginalization (cf. Eq. (2)). The MIL framework is essentially a sum-log approximation of the log-sum, i.e., ∑zlog⁡p(x,z)\sum_{z}\log p(x,z). To see this, the max-pooling function log⁡max⁡zp(x,z)\log\max_{z}p(x,z) used in can be viewed as the sum-log ∑zlog⁡p(x∣z)p(z)\sum_{z}\log p(x|z)p(z), where p(z)=1p(z)=1 if zz is the correct context and 0 otherwise, indicating there is only one positive instance; maximizing the noisy-or function log⁡(1−∏z(1−p(x,z)))\log(1-\prod_{z}(1-p(x,z))) used in is equivalent to maximize ∑zlog⁡p(x,z)\sum_{z}\log p(x,z), assuming there is at least one positive instance. However, due to the numerical property of the log function, this sum-log approximation will unnecessarily force every (x,z)(x,z) pair to explain the data . Instead, we use the variational Bayesian upper-bound to obtain a better sum-log approximation. Note that visual attention models simplify the variational lower bound by assuming p(z)=q(z∣x)p(z)=q(z|x); however, we explicitly use the KL divergence KL(q(z∣x)∣∣p(z))KL(q(z|x)||p(z)) in the lower bound to regularize the approximate posterior q(z∣x)q(z|x) not being too far from the prior p(z)p(z).

Variational Context

In this section, we derive the variational Bayesian formulation of the proposed variational context model and the objective function for training and test.

The task of grounding a referring expression LL in an image II, represented by a set of regions x∈Xx\in\mathcal{X}, can be viewed as a region retrieval task with the natural language query LL. Formally, we maximize the log-likelihood of the conditional distribution to localize the referent region x∗∈Xx^{*}\in\mathcal{X}:

where we omit the image II in p(x∣I,L)p(x|I,L).

As there is usually no annotation for the context, we consider it as a hidden variable zz. Therefore, Eq. (1) can be rewritten as the following maximization of the log-likelihood of the conditional marginal distribution:

Note that zz is NOT necessary to be one region as assumed in recent MIL approaches , i.e., z∈Xz\in\mathcal{X}. For example, the contextual objects “surrounding elephants” in “a bigger elephant than the surrounding elephants” should be composed by a multinomial subset of X\mathcal{X}, resulting in an extremely large sample space that requires O(2∣X∣)\mathcal{O}(2^{|\mathcal{X}|}) search complexity. Therefore, the marginalization in Eq (2) is intractable in general.

To this end, we use the variational lower-bound to approximate the marginal distribution in Eq. (2) as:

where KL(⋅)KL(\cdot) is the Kullback-Leibler divergence, ϕ\phi, θ\theta, and ω\omega are independent parameter sets for the respective distributions. As shown in Figure 1, the lower bound Q(x,L)\mathcal{Q}(x,L) offers a new perspective for exploiting the reciprocal nature of referent and context in referring expression grounding: Localization. This term calculates the localization score for xx given an estimated context zz, using the referent-cue of LL parameterized by θ\theta. In particular, we design a new posterior qϕ(z∣x,L)q_{\phi}(z|x,L) that approximates the true context prior p(z∣x,L)p(z|x,L), which models the context zz using the context-cue of LL parameterized by ϕ\phi. In the view of variational auto-encoder , this term works in an encoding-decoding fashion: qϕq_{\phi} is the encoder from xx to zz, and pθp_{\theta} is the decoder from zz to xx. Regularization. As KLKL is non-negative, maximizing Q(x,L)\mathcal{Q}(x,L) would encourage that the posterior qϕq_{\phi} is similar to the prior pωp_{\omega}, i.e., the estimated context zz sampled from qϕ(z∣x,L)q_{\phi}(z|x,L) should not be too far from the referring expression, which is modeled by pω(z∣L)p_{\omega}(z|L) with the generic-cue of LL parameterized by ω\omega. This term is necessary as the estimated zz could be overfitted to region features that are inconsistent with the visual context described in the expression.

2 Training and Test

Deterministic Context. The lower-bound Q(x,L)\mathcal{Q}(x,L) transforms the intractable log-sum in Eq. (2) into the efficient sum-log in Eq. (3), which can be optimized by using Monte Carlo unbiased gradient estimator such as REINFORCE . However, due to that ϕ\phi is dependent on the sampling of zz over O(2∣X∣)\mathcal{O}(2^{|\mathcal{X}|}) configurations, its gradient variance is large. To this end, we implement qϕ(z∣x,L)q_{\phi}(z|x,L) as a differentiable but biased encoder:

where we slightly abuse qϕq_{\phi} as a score function such that ∑x′qϕ(x′∣x,L)=1\sum_{x^{\prime}}q_{\phi}(x^{\prime}|x,L)=1. Note that this deterministic context can be viewed as applying the “re-parameterization” trick as in Variational Auto-Encoder : rewriting z∼qϕ(z∣x,L)z\sim q_{\phi}(z|x,L) to z=f(x,L;ϵ),ϵ∼p(ϵ)z=f(x,L;\epsilon),\epsilon\sim p(\epsilon), where the stochasticity of the auxiliary random variable ϵ\epsilon comes from training samples x∈X(ϵ)x\in\mathcal{X}(\epsilon). A clear example is Adversarial Autoencoder which shows that such stochasticity achieves similar test-likelihood compared to other distributions such as Gaussian.

Objective Function. Applying Eq. (4) to Eq. (3), we can rewrite Q(x,L)Q(x,L) into a function of only one sample estimation, which is a common practice in SGD:

In supervised setting where the ground truth of the referent is known, to distinguish the referent from other objects, we need to train a model that outputs a high p(x∣L)p(x|L) (i.e., Q(x,L)\mathcal{Q}(x,L)), while maintaining a low p(x′∣L)p(x^{\prime}|L) (i.e., Q(x′,L)\mathcal{Q}(x^{\prime},L)), whenever x′≠xx^{\prime}\neq x. Therefore, we use the so-called Maximum Mutual Information loss as in −log⁡{Q(x,L)/∑x′Q(x′,L)}-\log\{\mathcal{Q}(x,L)/\sum_{x^{\prime}}\mathcal{Q}(x^{\prime},L)\}, where we do not need to explicitly model the distributions with normalizations; we use the following score function:

where zz is omitted as it is a function of xx in Eq. (4). sθs_{\theta}, sϕs_{\phi}, and sωs_{\omega} are the score functions (e.g., pθ∝sθp_{\theta}\propto s_{\theta}) for pθp_{\theta}, qϕq_{\phi}, and pωp_{\omega}, respectively. These functions will be detailed in Section 4.3. In this way, maximizing Eq. (5) is equivalent to minimizing the following softmax loss:

where the softmax is over x∈Xx\in\mathcal{X} and xgtx_{gt} is the ground truth referent region.

Note that the reciprocity between referent and context can be extended to unsupervised learning, where neither of the referent and context has annotation. In this setting, we adopt the image-level max-pooled MIL loss functions for unsupervised referring expression grounding:

where the softmax is over x∈Xx\in\mathcal{X}. Note that the max-pooled MIL function is reasonable since there is only one ground truth referent given an expression and image training pair.

At test stage, in both supervised and unsupervised settings, we predict the referent region x∗x^{*} by selecting the region x∈Xx\in\mathcal{X} with the highest score:

Model Architecture

The overall architecture of the proposed variational context model is illustrated in Figure 2. Thanks to the deterministic context in Eq. (4), the five modules in our model can be integrated into an end-to-end differentiable fashion. Next, we will detail the implementation of each module.

Given an image with a set of Region of Interests (RoIs) X\mathcal{X}, obtained by any off-the-shelf proposal generator or object detectors , this module extracts the feature vector xi\mathbf{x}_{i} for every RoI. In particular, xi\mathbf{x}_{i} is the concatenation of visual feature vi\mathbf{v}_{i} and spatial feature pi\mathbf{p}_{i}. For vi\mathbf{v}_{i}, we can use the output of a pre-trained convolutional network (cf. Section 5). If the object category of each RoI is available, we can further utilize the comparison between the referent and other objects to capture the visual difference such as “the largest/baby elephant”. Specifically, we append the visual difference feature δvi=1n∑j≠ivi−vj∣∣vi−vj∣∣\delta\mathbf{v}_{i}=\frac{1}{n}\sum_{j\neq i}\frac{\mathbf{v}_{i}-\mathbf{v}_{j}}{||\mathbf{v}_{i}-\mathbf{v}_{j}||} to the original vi\mathbf{v}_{i} visual feature, where nn is the number of objects chosen for comparison (e.g., the number of RoI in the same object category). For spatial feature, we use the 5-d spatial attributes pi=[xtlW,ytlH,xbrW,ybrH,w⋅hW⋅H]\mathbf{p}_{i}=[\frac{x_{tl}}{W},\frac{y_{tl}}{H},\frac{x_{br}}{W},\frac{y_{br}}{H},\frac{w\cdot h}{W\cdot H}], where xx and yy are the coordinates the top left (tl) and bottom right (br) RoI of the size w×hw\times h, and the image is of the size W×HW\times H.

2 Cue-Specific Language Features

The cue-specific language feature representation for a referring expression is inspired by the attention weighted sum of word vectors , where the weights are parameterized by context-cue ϕ\phi, referent-cue θ\theta, and generic-cue ω\omega. The context-cue language feature yc=[yc1,yc2]\mathbf{y}^{c}=[\mathbf{y}^{c1},\mathbf{y}^{c2}] is a concatenation of yc1\mathbf{y}^{c1} for language-vision association between single RoI and the expression, and yc2\mathbf{y}^{c2} for the association between pairwise RoIs; the referent-cue language feature yr\mathbf{y}^{r} can be represented in a similar way to yc\mathbf{y}^{c}; the generic-cue language feature yg\mathbf{y}^{g} is only for single RoI association as it is an independent prior. The weights of each cue are calculated from the hidden state vectors of a 2-layer bi-directional LSTM (BLSTM) , scanning through the expression. The hidden states encode forward and backward compositional semantic meanings of the sentences, beneficial for selecting words that are useful for single and pairwise associations. Specifically, suppose hj\mathbf{h}_{j} as the 4,000-d concatenation of forward and backward hidden vectors of the jj-th word, without loss of generality, the word attention weight αj\alpha_{j} and the language feature y\mathbf{y} for single/pairwise association of any cue can be calculated as:

where wj\mathbf{w}_{j} is a 300-d vector. Note that the BLSTM module can be jointly trained with the entire model.

Figure 3 shows that the cue-specific language features dynamically weight words in different expressions. We can have two interesting observations. First, c1 is almost uniform while c2 is highly skewed; although r2 is more skewed than c1, it is still less skewed than r1. This is reasonable since: 1) without ground-truth, individual score (c1) does not help much for context estimation from scratch; context is more easily found by the pairwise score (c2) induced by relationships or other objects (e.g., “left” or “frisbee”); 2) in referent grounding with ground truth, individual score (r1) is sufficient (e.g., “dog lying” and “black white dog”) and pairwise score (r2) is helpful; 3) g is adaptive to the number of object categories in the expression, i.e., if the context object is of the same category as the referent, g weighs descriptive or relationship words higher (e.g., “lying, standing, left”), and nouns higher (e.g., “frisbee”), otherwise; moreover, it demonstrates that the deterministic guess of zz in Eq. (4) is meaningful.

3 Score Functions

For any image and expression pair, given the RoI feature xi\mathbf{x}_{i}, and the cue-specific language feature yc\mathbf{y}^{c}, yr\mathbf{y}^{r}, and yg\mathbf{y}^{g}, we implement the final grounding score in Eq. (6) as:

where the right-hand side functions are defined as below.

Context Estimation Score: sϕ(xi,xj,yc)s_{\phi}(\mathbf{x}_{i},\mathbf{x}_{j},\mathbf{y}^{c}). It is a score function for modeling the context posterior qϕ(z∣x,L)q_{\phi}(z|x,L), i.e., given an RoI xi\mathbf{x}_{i} as the candidate referent, we calculate the likelihood of any RoI xj\mathbf{x}_{j} to be the context. We can also use this function to estimate the final context posterior score sϕ(xi,zi,yc)s_{\phi}(\mathbf{x}_{i},\mathbf{z}_{i},\mathbf{y}^{c}). Specifically, the context estimation score is a sum of the single and pairwise vision-language association scores: xj\mathbf{x}_{j} and yc1\mathbf{y}^{c1}, [xi,xj][\mathbf{x}_{i},\mathbf{x}_{j}] and yc2\mathbf{y}^{c2}. Each associate score is an fc output from the input of a normalized feature:

where the element-wise multiplication ⊙\odot is an effective way for multimodal features . According to Eq. (4), we can obtain the estimated context zz as zi=∑jβjxj\mathbf{z}_{i}=\sum\nolimits_{j}\beta_{j}\mathbf{x}_{j}, where βj=softmaxj(sϕ(xi,xj,yc))\beta_{j}=\textrm{softmax}_{j}(s_{\phi}(\mathbf{x}_{i},\mathbf{x}_{j},\mathbf{y}^{c})).

Referent Grounding Score: sθ(xi,zi,yr)s_{\theta}(\mathbf{x}_{i},\mathbf{z}_{i},\mathbf{y}^{r}). After obtaining the context feature zi\mathbf{z}_{i}, we can use this score function to calculate how likely a candidate RoI xi\mathbf{x}_{i} is the referent given the context zi\mathbf{z}_{i}. This function is similar to Eq. (12).

Context Regularization Score: sω(zi,yg)−sϕ(xi,zi,yc)s_{\omega}(\mathbf{z}_{i},\mathbf{y}^{g})-s_{\phi}(\mathbf{x}_{i},\mathbf{z}_{i},\mathbf{y}^{c}). As discussed in Eq. (6), this function scores how likely the estimated context feature zi\mathbf{z}_{i} is consistent with the content mentioned in the expression. In particular, sω(zi,yg)s_{\omega}(\mathbf{z}_{i},\mathbf{y}^{g}) is only dependent on single RoI:

Experiment

We used four popular benchmarks for the referring expression grounding task.

RefCOCO . It has 142,210 referring expressions for 50,000 referents (e.g., object instances) in 19,994 images from MSCOCO . The expressions are collected in an interactive way . The dataset is split into train, validation, Test A, and Test B, which has 120,624, 10,834, 5,657 and 5,095 expression-referent pairs, respectively. An image contains multiple people in Test A and multiple objects in Test B.

RefCOCO+ . It has 141,564 expressions for 49,856 referents in 19,992 images from MSCOCO. The difference from RefCOCO is that it only allows appearances but no locations to describe the referents. The split is 120,191, 10,758, 5,726 and 4,889 expression-referent pairs for train, validation, Test A, and Test B respectively.

RefCOCOg . It has 95,010 referring expressions for 49,822 objects in 25,799 images from MSCOCO. Different from RefCOCO and RefCOCO+, this dataset not collected in an interactive way and contains longer sentences containing both appearance and location expressions. The split is 85,474 and 9,536 expression-referent pairs for training and validation. Note that there is no open test split for RefCOCOg, so we used the hyper-parameters cross-validated on RefCOCO and RefCOCO+.

RefCLEF . It contains 20,000 images with annotated image regions. It has some ambiguous (e.g. “anywhere”) phrases and mistakenly annotated image regions that are not described in the expressions. For fair comparison, we used the split released by , i.e., 58,838, 6,333 and 65,193 expression-referent pairs for training, validation and test, respectively.

2 Settings and Metrics

We used an English vocabulary of 72,704 words contained in the GloVe pre-trained word vectors , which was also used for the initialization of our word vectors. We used a “unk” symbol for the input word of the BLSTM if the word is out of the vocabulary; we set the sentence length to 20 and used “pad” symbol to pad expression sentence <20<20. For RoI visual features on RefCOCO, RefCOCO+, and RefCOCOg which have MSCOCO annotated regions with object categories, we used the concatenation of the 4,096-d fc7 output of a VGG-16 based Faster-RCNN network trained on MSCOCO and its corresponding 4,096-d visdiff feature ; although RefCLEF regions also have object categories, for fair comparison with , we did not use the visdiff feature.

The model training is single-image based, with all referring expressions annotated. We applied SGD of 0.95-momentum with initial learning rate of 0.01, multiplied by 0.1 after every 120,000 iterations, up to 160,000 iterations. Parameters in BILSTM and fc-layers were initialized by Xavier with 0.0005 weight decay. Other settings were default in TensorFlow. Note that our model is trained without bells and whistles, therefore, other optimization tricks such as batch normalization and GRU are expected to further improve the results reported here. Besides the ground truth annotations, grounding to automatically detected objects is a more practical setting. Therefore, we also evaluated with the SSD-detected bounding boxes on the four datasets provided by . A grounding is considered as correct if the intersection-over-union (IoU) of the top-1 scored region and the ground-truth object is larger than 0.50.5. The grounding accuracy (a.k.a, P@1) is the fraction of correctly grounded test expressions.

3 Evaluations of Supervised Grounding

We compared our variational context model (VC) with state-of-the-art referring expression methods published in recent years, which can be categorized into: 1) generation-comprehension based such as MMI , Attr , Speaker , Listener , and SCRC ; 2) localization based such as GroundR , NegBag , CMN . Note that NegBag and CMN are MIL-based models. In particular, we used the author-released code to obtain the results of CMN on RefCLEF, RefCOCO, and RefCOCO+.

From the results on RefCOCO, RefCOCO+, and RefCOCOg in Table 1 and that on RefCLEF in Table 2, we can see that VC achieves the state-of-the-art performance. We believe that the improvement is attributed to the variational Bayesian modeling of context. First, on all datasets, except for the most recent reinforcement learning based , VC outperforms all the other sentence generation-comprehension methods that do not model context. Second, compared to VC without the regularization term in Eq. (3) (VC w/o reg), VC can boost the performance by around 2% on all datasets. This demonstrates the effectiveness of the KL divergence for the prevention of the overfitted context estimation.

In particular, we further demonstrate the superiority of VC over the most recent MIL-based method CMN. As illustrated in Figure 4, VC has better context comprehension in both of the language and image regions than CMN. For example, in the top two rows where VC is correct and CMN is wrong, for the grounding in the second column, CMN unnecessarily considers the “girl” as context but the expression only describes using “elephant”; in the last column, CMN misses the key context “frisbee”. Even in the failure cases where VC is wrong and CMN is correct, VC still localizes reasonable context. For example, in the fourth column, although CMN grounds the correct TV, but it is based on incorrect context of other TVs; while VC can predict the correct context “children”. In addition, we observed that most of the cases that CMN is better than VC involves multiple humans. This demonstrates that VC is better at grounding objects of different categories.

VC is also effective in images with more objects. Figure 5 shows the performances of VC and CMN with various number of bounding boxes. We can observe that VC considerably outperforms CMN over all bounding boxes numbers. Recall that context is the key to distinguish objects of the same category. In particular, on the Test A sets of RefCOCO and RefCOCO+ where the grounding is only about people, i.e., the same object category, the gap between VC and CMN is becoming larger as the box number increases. This demonstrates that MIL is ineffective in modeling context, especially when the number of image regions is large.

4 Evaluations of Unsupervised Grounding

We follow the unsupervised setting in GroundR . To our best knowledge, it is the only work on unsupervised referring expression grounding. Note that it is also known as “weakly supervised” detection as there is still image-level ground truth (i.e., the referring expression). Table 2 reports the unsupervised results on the RefCLEF. We can see that VC outperforms the state-of-the-art GroundR, which is a generation-comprehension based method. This demonstrates that using context also helps unsupervised grounding. As there is no published unsupervised results on RefCOCO, RefCOCO+, and RefCOCOg, we only compared our baselines on them in Table 3. We can have the following three key observations which highlight the challenges of unsupervised grounding:

Context Prior. VC w/o reg is the baseline without the KL divergence as a context regularization in Eq. (3). We can see that in most of the cases, VC considerably outperforms VC w/o reg by over 2%, even over 5% on RefCOCO+ (det) and RefCOCOg (det). Note that this improvement is significantly higher than that in supervised setting (e.g., <3%<3\% as reported in Table 1). The reason is that the context estimation in Eq. (4) would be easier to be stuck in image regions that are irrelevant to the expression in unsupervised setting, therefore, context prior is necessary.

Language Feature. Except on RefCOCOg, we consistently observed the ineffectiveness of the cue-specific language feature in unsupervised setting, i.e., VC w/o α\alpha outperforms VC in Table 2 and 3. Here α\alpha represents the cue-specific word attention. This is contrary to the observation in the supervised setting as listed in Table 1, where VC w/o α\alpha is consistently lower than VC. Note that without the cue-specific word attention α\alpha in Eq. (10), the language feature is merely the average value of the word embedding vectors in the expression. In this way, VC w/o α\alpha does not encode any structural language composition as illustrated in Figure 3, thus, it is better for short expressions. However, when the expression is long in RefCOCOg, discarding the language structure still degrades the performance on RefCOCOg.

Unsupervised Relation Discovery. Although we demonstrated that VC improves the unsupervised grounding by modeling context, we believe that there is still a large space for improving the quality of modeling the context. As the failure examples shown in Figure (6), 1) many context estimations are still out of the scope of the expression, e.g., we may localize the “cup” and “table” as context even though the expression is “woman with green t-shirt”; 2) we may mistake due to the wrong comprehension of the relations, e.g., “right” as “left”, even if the objects belong to the same category, e.g., “elephant”. For further investigation, Figure 7 visualizes the cue-specific word attentions in supervised and unsupervised settings. The almost identical word attentions in unsupervised setting reflect the fact that the relation modeling between referent and context is not as successful as in supervised setting. This inspires us to exploit stronger prior knowledge such as language structure and spatial configurations .

Conclusions

We focused on the task of grounding referring expressions in images and discussed that the key problem is how to model the complex context, which is not effectively resolved by the multiple instance learning framework used in prior works. Towards this challenge, we introduced the Variational Context model, where the variational lower-bound can be interpreted by the reciprocity between the referent and context: given any of which can help to localize the other, and hence is expected to significantly reduce the context complexity in a principled way. We implemented the model using cue-specific language-vision embedding network that can be efficiently trained end-to-end. We validated the effectiveness of this reciprocity by promising supervised and unsupervised experiments on four benchmarks. Moving forward, we are going to 1) incorporate expression language generation in the variational framework, 2) use more structural features of language rather than word attentions, and 3) further investigate the potential of our model in the unsupervised referring expression grounding.

Supplementary Material

Recall that we are going to transform the log-sum objective function in Eq. (2) to sum-log for tractable training. Without loss of generality, we omit the conditional LL in the derivation. By using the concavity of the log function:

2 More Examples on Language Features

Figure 9 shows more cue-specific language features on RefCOCOg.

3 More Results on Unsupervised Grounding

4 External Parsers

As discussed in Section 5.4 that the language feature in unsupervised VC is not as good as that in the supervised setting. An alternative is to use external NLP parsers to obtain the compositions. However, conventional parsers (e.g., Standford Dependency) are observed to be suboptimal to the visual grounding task . Therefore, we adopt the parser jointly trained on the referring expression grounding task . As illustrated in Figure 8, this parser assigns word-level attention weights of subject, relation, and object. In particular, we consider the language features of c1, r2 as the average word embeddings, c2 as the relation weights, r1 as the subject weights, g as the object weights. Table 4 shows the performances on unsupervised grounding. We can see that there is no significant improvement of VC w/ parser over VC w/o α\alpha.

5 More Qualitative Results

Figure 10 shows more qualitative results on supervised and unsupervised grounding results on RefCOCO, RefCOCO+, and RefCOCOg.

References