Referring Expression Object Segmentation with Caption-Aware Consistency

Yi-Wen Chen, Yi-Hsuan Tsai, Tiantian Wang, Yen-Yu Lin, Ming-Hsuan Yang

Introduction

Object segmentation aims to separate foreground objects from the background. In this work, we focus on object segmentation from referring expressions, in which the segmentation is guided by a natural language description that identifies a particular object instance in a scene, e.g\bmvaOneDot, the man in a blue jacket or the laptop on the left.

Transferring knowledge between language and visual domains is an important but challenging task. Two relevant tasks are: 1) referring expression comprehension for localizing or segmenting an object according to a natural language description, and 2) referring expression generation for producing a sentence that identifies a particular object in an image. Existing methods [Mao et al.(2016)Mao, Huang, Toshev, Camburu, Yuille, and Murphy, Yu et al.(2016)Yu, Poirson, Yang, Berg, and Berg] address both tasks by constructing a generation model and inferring the region which maximizes the expression posterior in the comprehension task. However, such joint information is usually exploited to only enhance the generation performance.

In this paper, we focus on referring expression object segmentation. Unlike existing methods, our model jointly considers both tasks to benefit the comprehension task. Intuitively, when one signal, e.g\bmvaOneDot, a sentence, is transferred from the language domain to the visual domain, and then transferred back to the language domain, the transferred-back signal is supposed to be similar to the original one. By exploiting this property, we develop a network that jointly considers referring expression comprehension and generation, and enforces a caption-aware consistency between the visual and language domains.

To this end, we first design a comprehension network that contains the language and visual encoders to extract the feature representations of respective domains. To connect these two domains, we further propose to use spatial-aware dynamic filters to bridge the language and visual encoders. Meanwhile, these filters provide visual representations with the localization ability from the input referring expression. Based on the proposed baseline model, we then employ a caption generation model that takes feature representations from the comprehension network as inputs. The generated referring expression should be similar to the original sentence, and we leverage this property as an additional consistency cue to enhance the language and visual representations. The main steps of the proposed model are illustrated in Figure 1.

To evaluate the proposed method, we conduct extensive experiments on the RefCOCO [Yu et al.(2016)Yu, Poirson, Yang, Berg, and Berg] and RefCOCOg [Mao et al.(2016)Mao, Huang, Toshev, Camburu, Yuille, and Murphy, Nagaraja et al.(2016)Nagaraja, Morariu, and Davis] datasets. Experimental results show that our model performs favorably against the state-of-the-art methods. In addition, we provide the ablation study to demonstrate the effectiveness of each component in the proposed framework, including the spatial-aware dynamic filters and caption-aware consistency. The main contributions of this work are summarized as follows: 1) We integrate referring expression generation into referring expression comprehension so that the two complementary tasks can benefit each other via enforcing the caption-aware consistency. 2) We develop the spatial-aware dynamic filters that bridge the visual and language domains and facilitate the feature learning process. 3) We design an end-to-end trainable network for referring expression comprehension, achieving the state-of-the-art performance.

Related Work

The task of referring expression comprehension aims to localize or segment an object given a natural language description. Existing methods [Hu et al.(2016b)Hu, Xu, Rohrbach, Feng, Saenko, and Darrell, Luo and Shakhnarovich(2017), Mao et al.(2016)Mao, Huang, Toshev, Camburu, Yuille, and Murphy] mainly rely on recurrent caption generation models, and select the object with the maximum posterior probability of the expression among all object proposals. By exploring the relationship between the object and its context [Nagaraja et al.(2016)Nagaraja, Morariu, and Davis, Yu et al.(2016)Yu, Poirson, Yang, Berg, and Berg, Zhang et al.(2018)Zhang, Niu, and Chang], the target object can be better localized. Recent approaches adopt various learning strategies, such as embedding images and sentences into a common feature space [Wang et al.(2016)Wang, Li, and Lazebnik, Rohrbach et al.(2016)Rohrbach, Rohrbach, Hu, Darrell, and Schiele], or learning attributes [Liu et al.(2017b)Liu, Wang, and Yang] to help differentiate objects of the same category. In addition, Hu et al\bmvaOneDot [Hu et al.(2017)Hu, Rohrbach, Andreas, Darrell, and Saenko] analyze the inter-object relationships by parsing the sentence into subject, relationship and object parts. To jointly consider the associated factors such as attributes and relationships between objects, Yu et al\bmvaOneDot [Yu et al.(2018)Yu, Lin, Shen, Yang, Lu, Bansal, and Berg] propose a modular attention network to decompose the expression into subject appearances, locations, and relationships to other objects.

While the aforementioned methods mainly localize an object by a bounding box, algorithms that focus on segmentation [Hu et al.(2016a)Hu, Rohrbach, and Darrell, Li et al.(2018)Li, Li, Kuo, Shu, Qi, Shen, and Jia, Liu et al.(2017a)Liu, Lin, Shen, Yang, Lu, and Yuille, Shi et al.(2018)Shi, Li, Meng, and Wu, Margffoy-Tuay et al.(2018)Margffoy-Tuay, Pérez, Botero, and Arbeláez] usually encode the referring expression through the LSTM network and use a fully convolutional network for foreground/background segmentation by using both the language and visual features. Different from these approaches, our proposal-based model first localizes objects and performs segmentation via learning better feature representations through a referring generation network that considers the caption consistency. We note that the approach in [Rohrbach et al.(2016)Rohrbach, Rohrbach, Hu, Darrell, and Schiele] also considers the consistency between the generated sentence and input sentence but does not target at segmenting objects. Furthermore, this approach uses pre-defined and fixed region proposals, in which the visual representations are not updated through the proposals. In contrast, our unified framework is end-to-end trainable while bridging features across visual and language domains.

The generation task of referring expressions is a special case of image captioning. Rather than describing the whole image, the generated sentence uniquely identifies an object within the image. A referring expression is considered good if one can localize the corresponding object by comprehending this referring expression. Therefore, referring expression comprehension is often employed in the generation task [Mao et al.(2016)Mao, Huang, Toshev, Camburu, Yuille, and Murphy, Liu et al.(2017b)Liu, Wang, and Yang, Luo and Shakhnarovich(2017)] to improve the performance.

CNN-LSTM based models are widely used for image captioning [Karpathy and Fei-Fei(2015), Rennie et al.(2017)Rennie, Marcheret, Mroueh, Ross, and Goel, Vinyals et al.(2015)Vinyals, Toshev, Bengio, and Erhan, Xu et al.(2015)Xu, Ba, Kiros, Cho, Courville, Salakhutdinov, Zemel, and Bengio]. While a CNN model extracts visual features, an LSTM module produces captions. To address referring expression generation, Mao et al\bmvaOneDot [Mao et al.(2016)Mao, Huang, Toshev, Camburu, Yuille, and Murphy] combine the extracted visual features with the location and size of the target object. Furthermore, this method uses a CNN-LSTM model for the comprehension task and jointly trains the generation and comprehension modules. Yu et al\bmvaOneDot [Yu et al.(2017)Yu, Tan, Bansal, and Berg] further propose a joint speaker-listener-reinforcer model where a reward function is introduced to guide the expression sampling. Their approach jointly trains the generation and comprehension networks, but does not specifically consider the caption consistency as our framework. While the aforementioned methods mainly utilize the comprehension model to generate high-quality sentences, in this work, we focus on the comprehension task and demonstrate that the generation model also facilitates the comprehension performance by enforcing the proposed caption-aware consistency between the visual and language domains.

Proposed Framework

In this work, we focus on referring expression object segmentation. The overview of the proposed framework is illustrated in Figure 2. Given an image II and a natural language description rr, we aim to segment the object in II specified by rr. To this end, we propose an end-to-end trainable network that contains a language encoder EE, a visual encoder VV, a Mask R-CNN head DD, and a caption generator CC. The encoders EE and VV extract language and visual features, respectively. Motivated by the dynamic filter network [Brabandere et al.(2016)Brabandere, Jia, Tuytelaars, and Gool], we enhance the ability of specific object localization via introducing the spatial-aware dynamic filters to transfer knowledge from text to image. The yielded cross-modal information allows the Mask R-CNN head DD to produce more accurate segmentation results. To further improve our model, we employ the caption generation network CC and a consistency loss LcapL_{cap} to jointly train the comprehension and generation networks. We describe each component of the proposed network below.

In this subsection, we introduce how the proposed network generates the object segment given the query referring expression. To this end, the language encoder EE, visual encoder VV, and spatial-aware dynamic filters are elaborated.

Similar to [Yu et al.(2018)Yu, Lin, Shen, Yang, Lu, Bansal, and Berg], we use a bi-directional LSTM model to extract features of a referring expression. Given a referring expression r={wt}t=1Tr=\{w_{t}\}_{t=1}^{T} of TT words with each word wtw_{t} represented by a one-hot vector ete_{t}, the bi-directional LSTM SS is applied to encode the whole sentence in both forward and backward directions:

where h→t\overrightarrow{h}_{t} and h←t\overleftarrow{h}_{t} are the forward and backward hidden states at time step tt, respectively. We concatenate the final hidden states in both directions to yield the feature representation FrefF_{ref} of the referring expression.

Given an input image II, we aim at pixel-wise segmentation. Different from the approaches based on the fully convolutional network (FCN) that does not generate instance-aware results, we adopt the proposal-based Mask R-CNN [He et al.(2017)He, Gkioxari, Dollar, and Girshick] framework to generate an object mask based on each detected object bounding box. We use the ResNet-101 [He et al.(2016)He, Zhang, Ren, and Sun] model as the backbone network and extract features over the entire image. The feature from the final convolutional layer of the fourth block, denoted by Fvis=V(I)F_{vis}=V(I), serves as the representation of image II.

Motivated by the recent work [Li et al.(2017)Li, Tao, Gavves, Snoek, and Smeulders] on tracking with natural language, we utilize dynamic convolutional filters as a bridge to connect the language and visual domains. Unlike conventional convolutional filters that apply the same weights to all input images, dynamic convolutional filters are generated depending on the input sentence. Given the feature representation FrefF_{ref} of a sentence rr, a single fully connected layer parameterized by the weights Wd1W_{d}^{1} and the bias bd1b_{d}^{1} is adopted to generate a set of dynamic filters:

where tanh⁡\tanh is the hyperbolic tangent function, and fd1f_{d}^{1} is a set of 1×11\times 1 convolutional filters with the same number of channels as the visual representation FvisF_{vis}. We then convolve the visual representation FvisF_{vis} with the generated dynamic filters fd1f_{d}^{1} to obtain a response map Rref1R_{ref}^{1}:

With this formulation, knowledge is transferred from the language domain through learning the dynamic filters, with which the response map reflects the information inferred from the referring expression.

However, such filters consider the entire image and thus may only be able to catch the global structure but ignore spatially distributed objects. As such, we propose to utilize spatial-aware dynamic convolutional filters that consider local regions of the image, including up, down, left, right, horizontal and vertical middle regions, and each region covers a half area of the entire image. We thereby apply six additional fully connected layers to generate spatial-aware dynamic filters {fdi}i=27\{f_{d}^{i}\}_{i=2}^{7} corresponding to each region ii via (2). The six dynamic filters are then convolved with the visual feature FvisF_{vis}, where the values outside the defined regions are set to 0. Then we obtain six spatial-aware response maps similar to (3), denoted by {Rrefi}i=27\{R^{i}_{ref}\}_{i=2}^{7}, in which each map focuses on its defined region.

To combine these spatial response maps and the one from (3), we adopt another set of dynamic filters fwf_{w} with 7 channels, which are also generated from the sentence representation FrefF_{ref}, to account for the importance of each region depending on the input sentence. We convolve fwf_{w} with the concatenation of the 77 response maps Rcon=concat(Rrefi)R_{con}=\text{concat}(R^{i}_{ref}) and obtain a final response map RR with one channel, i.e\bmvaOneDot,

where σ\sigma is the sigmoid function with output range $.Ideally,. Ideally,Rrepresentsamapoftheobjectspecifiedbytheinputreferringexpression.Thus,weapplyabinarycross−entropylossrepresents a map of the object specified by the input referring expression. Thus, we apply a binary cross-entropy lossL_{res}tosupervisetheresponsemapto supervise the response mapR$ with respect to the ground-truth object mask.

Based on the response map in (4), we take the element-wise multiplication of RR and FvisF_{vis} to be the caption-aware feature representation F^vis\hat{F}_{vis}, which carries the information from both the language and visual domains. To obtain the final segmentation result, we then feed F^vis\hat{F}_{vis} into the Mask R-CNN [He et al.(2017)He, Gkioxari, Dollar, and Girshick] RoI head DD, which includes the bounding box and the binary segmentation branches. The overall objective can be written as:

where LroiL_{roi} includes the classification loss, bounding box loss and mask loss, the same as those defined in Mask R-CNN. With this formulation, we construct an end-to-end trainable network that produces the referring expression object segmentation. Unlike the state-of-the-art methods, such as MAttNet [Yu et al.(2018)Yu, Lin, Shen, Yang, Lu, Bansal, and Berg], that require multiple training stages and pre-processing steps, our model can be efficiently learned, through the help of spatial-aware dynamic filters which provide the spatial information from the input sentence.

2 A Joint Framework

In light of the cycle consistency work [Zhu et al.(2017)Zhu, Park, Isola, and Efros] that solves the domain transfer problem in cross-directions, we integrate both the referring expression comprehension and generation tasks into a joint framework, where their feature representations are shared and can be jointly optimized through back-propagation.

To generate a sentence describing a particular object within an image, we adopt the attention-based image captioning model [Xu et al.(2015)Xu, Ba, Kiros, Cho, Courville, Salakhutdinov, Zemel, and Bengio]. To train the caption generation model CC, we input the feature representation FvisF_{vis} extracted from Mask R-CNN and concatenate it with F^vis\hat{F}_{vis} which contains the spatial information about the object. As a result, during training the caption generation model, gradients can be back-propagated through FvisF_{vis} to update the Mask R-CNN feature extractor, as well as through F^vis\hat{F}_{vis} to optimize dynamic filters and the language encoder.

Given the ground-truth sentence r={wt}t=1Tr=\{w_{t}\}_{t=1}^{T}, which is the input to the language encoder, the objective for caption generation is to minimize the cross-entropy loss LcapL_{cap}:

where pθc(w^t∣w^1,...,w^t−1)p_{\theta_{c}}(\hat{w}_{t}|\hat{w}_{1},...,\hat{w}_{t-1}) is the probability of predicting a particular word from the caption generation network parameterized by θc\theta_{c}. Here, this loss function in our framework enforces that the predicted sentence r^\hat{r} generated by the feature F^vis\hat{F}_{vis}, i.e\bmvaOneDot, r^=C(F^vis,⋅)\hat{r}=C(\hat{F}_{vis},\cdot), should be consistent with the input query rr that generates the same feature, i.e\bmvaOneDot, F^vis=F(E(r))\hat{F}_{vis}=\mathcal{F}(E(r)), where F\mathcal{F} is a mixed operation involving the visual encoder VV and dynamic filters in the proposed method. Hence, our caption-aware consistency actually enforces r≈C(F(E(r)),⋅)r\approx C(\mathcal{F}(E(r)),\cdot).

To exploit the caption-aware consistency, we jointly train the comprehension model, including the language encoder EE, visual encoder VV, Mask R-CNN head DD in Section 3.1 and the caption generation model CC in Section 3.2. The total loss function is extended from (5) to:

where α\alpha is the coefficient of the consistency loss. We note that adding LcapL_{cap} enables the joint optimization between the language and the visual domains. That is, the intermediate feature F^vis\hat{F}_{vis} would be updated by the guidance from the first two loss functions in (7), which are supervised by the comprehension task, and in the meanwhile LcapL_{cap} updates the feature based on the caption generation task.

3 Model Training and Implementation Details

To train the joint network model, we adopt a sequential training strategy to optimize the objective in (7). First, we only update the comprehension network by optimizing (5). Then, we pre-train the caption generation network by optimizing (6) as a warm-up. Finally, we update the entire framework with the objective in (7). With the trained model, we choose the detected object with the largest score during testing.

We implement our model with PyTorch using the SGD optimizer. For the language encoder, the dimension of the LSTM hidden states is set to 512512. By concatenating the forward and backward hidden states, the feature FrefF_{ref} is a 10241024-dimensional vector. In the visual encoder, the visual feature FvisF_{vis} is of dimension 10241024. Thus, we also generate the dynamic filters of dimension 10241024. For the caption generation model, the input spatial features are resized to 14×1414\times 14 and have the same number of channels as that of the concatenation of FvisF_{vis} and F^vis\hat{F}_{vis}. When training the full model, the loss weight α\alpha in (7) is set to 0.10.1 for all experiments. The codes and models are available at: https://github.com/wenz116/lang2seg.

Experimental Results

We evaluate the proposed framework on two referring expression datasets: RefCOCO [Yu et al.(2016)Yu, Poirson, Yang, Berg, and Berg] and RefCOCOg (with two splitsThe first split [Mao et al.(2016)Mao, Huang, Toshev, Camburu, Yuille, and Murphy] randomly partitions objects into training and validation sets. We denote the validation set as “val*” in this paper. The second split [Nagaraja et al.(2016)Nagaraja, Morariu, and Davis] randomly partitions images into training, validation and testing sets, where we denote the validation and testing ones as “val” and “test”, respectively.) [Mao et al.(2016)Mao, Huang, Toshev, Camburu, Yuille, and Murphy, Nagaraja et al.(2016)Nagaraja, Morariu, and Davis]. The two datasets are collected from the Microsoft COCO images [Lin et al.(2014)Lin, Maire, Belongie, Hays, Perona, Ramanan, Dollár, and Zitnick], with different properties of expressions. We show both detection and segmentation results with comparisons against the state-of-the-art algorithms. In addition, we present an ablation study to demonstrate the importance of each component in the proposed framework. More results are provided in the supplementary material.

For evaluating the detection performance, the predicted bounding box is considered correct if the intersection-over-union (IoU) of the prediction and the ground truth is above 0.5. As for the segmentation quality, we use Intersection-over-Union (IoU) as metric.

In Table 1, we show comparisons with existing state-of-the-art algorithms [Liu et al.(2017b)Liu, Wang, and Yang, Luo and Shakhnarovich(2017), Nagaraja et al.(2016)Nagaraja, Morariu, and Davis, Yu et al.(2018)Yu, Lin, Shen, Yang, Lu, Bansal, and Berg, Yu et al.(2017)Yu, Tan, Bansal, and Berg, Zhang et al.(2018)Zhang, Niu, and Chang]. Since each method adopts diverse information to help the comprehension task, we further summarize the major cues that each approach relies on, such as context information [Nagaraja et al.(2016)Nagaraja, Morariu, and Davis, Zhang et al.(2018)Zhang, Niu, and Chang], attribute prediction [Liu et al.(2017b)Liu, Wang, and Yang, Yu et al.(2018)Yu, Lin, Shen, Yang, Lu, Bansal, and Berg], and joint training with referring expression generation [Liu et al.(2017b)Liu, Wang, and Yang, Luo and Shakhnarovich(2017), Yu et al.(2017)Yu, Tan, Bansal, and Berg].

Table 1 shows that the proposed method performs favorably against most methods by significant margins, and competitively with MAttNet [Yu et al.(2018)Yu, Lin, Shen, Yang, Lu, Bansal, and Berg]. We note that, the MAttNet method utilizes various cues, including attention module, attribute prediction, location information, and relations between objects to achieve good performance, while our model only focuses on the location cue and joint training with referring expression generation. It is also worth mentioning that our model is a unified framework that is end-to-end trainable, while MAttNet requires multiple separate training stages to obtain the final model. The runtime speed of our method is 0.17 seconds per image, which is much faster than MAttNet with 0.67 seconds per image on an Intel Xeon 2.5 GHz machine and an NVIDIA GTX 1080 Ti GPU with 11 GB memory.

2 Segmentation Results

We present the experimental results with comparisons to the state-of-the-art algorithms including D+RMI+DCRF [Liu et al.(2017a)Liu, Lin, Shen, Yang, Lu, and Yuille], RRN+LSTM+DCRF [Li et al.(2018)Li, Li, Kuo, Shu, Qi, Shen, and Jia], MAttNet [Yu et al.(2018)Yu, Lin, Shen, Yang, Lu, Bansal, and Berg], KWAN [Shi et al.(2018)Shi, Li, Meng, and Wu] and DMN [Margffoy-Tuay et al.(2018)Margffoy-Tuay, Pérez, Botero, and Arbeláez] on the two datasets in Table 2. Overall, our method consistently and significantly outperforms other segmentation-based approaches that use a similar backbone network (i.e\bmvaOneDot, Deeplab [Chen et al.(2016)Chen, Papandreou, Kokkinos, Murphy, and Yuille] with ResNet-101) as ours. Different from the DMN [Margffoy-Tuay et al.(2018)Margffoy-Tuay, Pérez, Botero, and Arbeláez] scheme that utilizes dynamic filters in a sequential manner for capturing the information of each word in a sentence, our model generates the dynamic filters in a spatial-aware manner, where each set of filters produces a response map to certain region of the image. We note that the proposed method achieves better performance. Similar to the localization results, MAttNet [Yu et al.(2018)Yu, Lin, Shen, Yang, Lu, Bansal, and Berg] that fuses multiple cues performs competitively with our model. We present qualitative examples of referring expression object segmentation in Figure 3. The proposed model can segment different objects according to various query expressions, such as the location, color, or action information, and further demonstrates the effectiveness of the proposed caption-aware consistency framework.

3 Ablation Study

We present the results of an ablation study in Table 1. We first show that using the proposed spatial-aware dynamic filters improves the baseline with only a single dynamic filter or the spatial-aware mechanism [Hu et al.(2016a)Hu, Rohrbach, and Darrell] that concatenates spatial coordinates and feature maps. Second, the referring expression generation network with caption-aware consistency performs favorably against the baseline model. In the full model with both spatial-aware filters and caption-aware consistency, higher performance gains are achieved over other baselines.

We present sample segmentation results predicted by different variants of our model in Figure 4. Compared with the baseline and the model with spatial-aware filters, the proposed full model can localize objects accurately while the baseline model predicts the wrong object. In addition to improving the localizing ability, our full model enhances feature representations around the object. For instance, the elephant in back is well segmented by our model even if it is surrounded by complex background and similar instances.

Concluding Remarks

In this paper, we propose an end-to-end trainable framework for referring expression segmentation. We design a comprehension model that consists of language and visual encoders to extract feature representations in the respective domains. By introducing the spatial-aware dynamic filters, knowledge can be transferred from language domain to visual domain, while capturing the useful location cue. In addition to the proposed baseline model, we employ a caption generation network to connect referring expression comprehension and generation. Considering the consistency that the generated sentence is supposed to be similar to the given referring expression, we enforce a caption-aware consistency loss and further enhance the language and visual representations. Extensive experiments and an ablation study on two referring expression datasets show that the proposed algorithm achieves favorable performance against the state-of-the-art methods.

This work was supported in part by Ministry of Science and Technology (MOST) under grants 107-2628-E-001-005-MY3 and 108-2634-F-007-009.

References