Cross-Modal Progressive Comprehension for Referring Segmentation
Si Liu, Tianrui Hui, Shaofei Huang, Yunchao Wei, Bo Li, Guanbin Li
Introduction
In this paper, we target at an emerging task called referring segmentation . Given a natural language expression and an image/video as inputs, the goal of referring segmentation is to segment the entities referred by the subject of the input expression. Traditional semantic segmentation methods aim to classify each pixel as one of a fixed set of categories denoted by short words (e.g., “person”, “cell phone”). Referring segmentation can be regarded as a generalized semantic segmentation task where the categories belong to an open set denoted by expressions with various grammars and diverse contents, such as entities, attributes, relationships and actions, etc. Combining rich visual and linguistic information, referring segmentation has a wide range of potential applications such as language-based robot controlling , interactive image editing , etc.
Previous works tackle the referring image or video segmentation task using a straightforward concatenation-and-convolution scheme involving dynamic filters , convolutional RNNs or cross-modal attention mechanism to fuse visual and linguistic features. Instead of the above one-stage implicit approaches, human tends to comprehend the referring expression in a progressive way . As human reads the expression, different nouns and adjectives will first be located in the image/video to find all the candidate entities. Then, relational and action words like prepositions and verbs are extracted to reason the relationships among different entities, where the disturbing entities are excluded and the target one is found out.
To mimic the more natural processing way of human beings, we introduce a Cross-Modal Progressive Comprehension (CMPC) scheme to solve the referring segmentation task in multiple stages. We define the entity referred by the expression as the referent. For referring image segmentation, we illustrate the CMPC scheme in the left part of Fig. 1. If the referent is described by “The man holding a white frisbee”, the referring process is divided into two progressive stages. First, the model can utilize entity words and attribute words, e.g., “man” and “white frisbee”, to perceive all the possible entities mentioned in the expression. Second, as one image may contain multiple entities of the same category, for example, the three men in Fig. 1 (b), the model needs to further highlight the referent matched with the relationship cue in the expression while suppressing other mismatched ones by reasoning relationships among entities. In Fig. 1 (c), under the guidance of the relational word “holding” which associates “man” with “white frisbee”, the model can focus on the referent who holds a white frisbee rather than the other two men. After comprehending multimodal information progressively, the model can make correct prediction as shown in Fig. 1 (d).
Different from image, expressions containing action descriptions are often used to refer entities in a video. In the right part of Fig. 1, we illustrate the video version of our CMPC scheme. If the entity in a video is described by “The girl in white shirt bouncing on the ball”, the referring process is divided into three stages. The first entity perception stage and the second relation-aware reasoning stage are conducted on the center frame of a video snippet, which is almost same as those of the image model. Then in the third stage, action words (e.g., “bouncing”) are exploited by the model to capture temporal cues among frames in the video to further highlight the referent conducting the described action, as shown in Fig. 1 (g). Finally, multimodal spatio-temporal information is comprehended by the model to make correct predictions along the video as shown in Fig. 1 (h).
To tackle image and video inputs respectively, we develop two versions of CMPC scheme. For image data, we propose a CMPC-I (Image) module which progressively exploits different types of words in the expression to segment the referent in an image. Concretely, our CMPC-I module consists of two stages. First, we extract linguistic features of entity words and attribute words (e.g., “man” and “white frisbee”) from the expression and then fuse them with visual features extracted from the image to build multimodal features. During this process, all the entities that may be referred by the expression are perceived. Second, we construct a fully-connected spatial graph where each image region is regarded as a vertex and multimodal information of the entity is contained in each vertex. Appropriate edges are required for vertexes to communicate with each other. Naive edges which connect all the vertexes equally will introduce abundant information and hinder the identification of the referent. Thus, our CMPC-I module employs relational words (e.g., “holding”) of the expression as a group of routers to build adaptive edges to connect spatial vertexes, i.e., entities, which are involved with the relationship described in the expression. Particularly, spatial vertexes (e.g., “man”) yielding strong responses to the relational words (e.g., “holding”) will exchange information with other vertexes (e.g., “frisbee”) that also highly correlate with the relational words. Meanwhile, less interaction will occur among spatial vertexes yielding weak responses to the relational words. After relation-aware reasoning on the multimodal graph, our CMPC-I module can highlight feature of the referent while suppressing those of the irrelevant entities, which assists in generating accurate segmentation. For video data, we further extend our CMPC-I module to CMPC-V (Video) module with an additional action-aware reasoning stage based on action words to exploit temporal information. Concretely, our CMPC-V module extracts global multimodal features of all the frames in the video snippet based on action words (e.g., “bouncing”) after the same entity perception and relation-aware reasoning stages. A fully-connected temporal graph is constructed using each frame as a vertex where frame feature is served as vertex feature. Information propagation among temporal vertexes is performed to extract temporal multimodal context among frames. Finally, features of the temporal graph are aggregated with feature of annotated frame to supplement temporal multimodal context for better segmentation.
As prior works show multiple levels of visual features can complement each other, we also propose a Text-Guided Feature Exchange (TGFE) module to exploit information of multimodal features refined by our CMPC modules from different levels. Combining CMPC-I or CMPC-V module with TGFE module forms our image or video version referring segmentation framework. For each level of multimodal features, our TGFE module utilizes linguistic features as guidance to select useful feature channels from other levels to enable information communication. After multiple rounds of communication, Our TGFE further fuses multi-level features by ConvLSTM to comprehensively integrate low-level visual details and high-level semantics for precise segmentation results.
The main contributions of our paper are summarized as follows:
We introduce a Cross-Modal Progressive Comprehension (CMPC) scheme to align multimodal features in multiple stages based on informative words in the expression, which provides a general solution to referring segmentation task and is robust to different visual modalities.
We instantiate the CMPC scheme by proposing a CMPC-I module containing entity perception and relation-aware reasoning stages for referring image segmentation. We further propose a CMPC-V module by extending CMPC-I with an action-aware reasoning stage for referring video segmentation.
We also propose a TGFE module to aggregate multi-level multimodal features to enhance feature of the referent for better segmentation.
Combining CMPC-I or CMPC-V with TGFE, our image and video version frameworks achieves current state-of-the-art performances on four referring image segmentation benchmarks and three referring video segmentation benchmarks, respectively.
This paper is an extension of our previous conference version . The current work adds to the initial version with significant aspects. First, we extend our CMPC scheme from referring image segmentation to referring video segmentation by introducing an additional action-aware reasoning stage, which effectively extracts temporal multimodal context to enhance feature of the referent. Second, we improve our initial CMPC-I module by changing the way of multimodal feature concatenation to better highlight feature of the referent. Third, we add considerable new experimental results including ablation study, model setting and visualization analysis. Our image model in this paper also obtains better performance than our conference version.
Related Work
Based on Fully Convolutional Networks (FCN) , semantic segmentation has made a huge progress in recent years. FCN uses convolution layers to replace all the fully-connected layers in original classification networks and becomes the most popular architecture in the semantic segmentation community. DeepLab series incorporates FCN with dilated convolutions with different dilation rates, which enlarges the receptive field of filters to aggregate multi-scale visual context. PSPNet proposes similar pyramid pooling operations to extract multi-scale context as well. Later works such as DANet and CFNet exploit self-attention mechanism to capture long-range dependencies among image positions and achieve notable performance. In this paper, we target at the more challenging semantic segmentation problem whose semantic categories are specified by diverse natural language sentences.
2 Referring Expression Grounding
Given a natural language expression, referring expression grounding aims to localize the entities matched with the expression in the given image or video. Many works conduct localization in bounding box level. Liao et al. performs cross-modality correlation filtering to match features from different modalities in real time. Graph models involving attention mechanism are explored in to find the most related objects for the expression. Yu et al. propose modular networks to decompose the referring expression into subject, location and relationship in order to finely compute the matching score. Most box-based methods are two-stage where a pretrained detector is utilized to first generate RoI proposals for later grounding. This design paradigm achieves high localization performance but lacks global context information and heavily relies on the quality of proposal candidates. In addition, it can only ground a single object and cannot locate the stuff or multiple objects.
Beyond bounding box, the referred object can also be localized more precisely with segmentation mask. Hu et al. first proposes the referring image segmentation problem and directly concatenates and fuses multimodal features from CNN and LSTM to generate the segmentation mask. In and , multimodal LSTM is employed to sequentially fuse visual and linguistic features in multiple time steps. Multi-level feature fusion is explored in to recurrently refine the local details of segmentation mask. As context information is critical to segmentation task, recent works employ cross-modal attention and self-attention to extract multimodal context between image regions and referring words. Cycle-consistency learning and adversarial training are also investigated to boost the segmentation performance. Gavrilyuk et al. further introduce referring segmentation task into video data in which sentences contain action descriptions of the actors in the videos. They utilize dynamic filters and multi-resolution decoder to generate mask of the referent. Based on , Wang et al. propose asymmetric cross-guided attention between visual and linguistic modalities to segment the referent more precisely. Different from box-based methods, most mask-based methods are one-stage where FCN is utilized to directly generate the mask of the referent. This design paradigm can extract rich global context information and does not rely on pretrained detectors. However, generating masks based on monolithic representations of referring expressions and visual contents may be difficult to distinguish between different instances, thus producing false positive masks and harming localization performance. In this paper, we tackle the above issue of one-stage methods by proposing to progressively highlight the referent via entity perception, relation-aware reasoning and action-aware reasoning, which effectively distinguishes different instances and reasons the referent on both image and video data.
3 Graph-Based Reasoning
Recently, graph-based models have shown their effectiveness in context reasoning for many tasks. Graph Convolution Networks (GCN) becomes popular for its superiority on semi-supervised classification. Wang et al. uses RoI proposals as vertexes to construct a spatial-temporal graph and conduct context reasoning with GCN, which improve performance on video recognition task. Chen et al. propose a global reasoning module which projects visual feature into an interactive space and performs graph convolution for global context reasoning. Then they reproject the reasoned global context back to the coordinate space to enhance original visual feature. There are several concurrent works sharing the same spirit with while using different implementations. Graph-based reasoning has also been widely applied in vision and language tasks such as Hu et al. and Yu et al. . Hu et al. build a fully-connected visual graph where each node corresponds to an object proposal generated by a pre-trained detector and formulates the message passing among graph nodes as a recurrent process. Yu et al. build two heterogeneous graphs where the primal vision-to-answer graph utilizes object proposals and answer words as graph nodes, while the dual question-to-answer graph utilizes query words and answer words as graph nodes. They conduct message passing between visual nodes and answer nodes in the primal graph, while between query nodes and answer nodes in the dual graph. Different from above two methods, we propose to regard image regions and video frames as vertexes to build spatial and temporal graphs for effectively reasoning multimodal context based on informative words in the expression which is more sutible for the segmentation task. Besides, we exploit the relational words as routers to connect each pair of visual nodes on the feature map and message passing among all visual nodes is guided by relational words in a more effective way.
Method
In this section, we elaborate the instantiations of our introduced CMPC scheme to tackle the referring segmentation task on image and video data respectively. The proposed modules are denoted as CMPC-I (Image) and CMPC-V (Video) in the rest of our paper.
1.2 Entity Perception
Since the input image may contain many entities, it is natural to progressively narrow down the candidate set from all the entities to the actual referent. The first stage of our CMPC-I (Image) module is entity perception. We associate linguistic features of entity words and attribute words with the correlated visual features of spatial regions using bilinear fusion to perceive all the candidate entities.
1.3 Relation-Aware Reasoning
After perceiving all the possible entities in the image, the second stage of our CMPC-I module is relation-aware reasoning. We construct a fully-connected multimodal graph based on using relational words as a group of routers to connect vertexes. Each vertex of the graph represents a spatial region on . By reasoning among vertexes of the multimodal graph, our model can highlight the responses of the referent which are involved with the relationship cue while suppressing those of the non-referred ones.
Each element of represents the normalized magnitude of information flow from the spatial region to the region , which depends on their affinities with relational words in the expression. If a vertex has high attention weight with a certain word, and this word also has high attention weight with another vertex, then these two vertexes will have high attention weight with each other in a relational context. In this way, adaptive edges connecting spatial vertexes can be built by leveraging relational words of the expression as a group of routers.
After building the multimodal graph , we conduct graph convolution over it as follow:
2 CMPC on Video
2.2 Action-Aware Reasoning
As video data usually contains temporal information, reasoning only static relationship between entities are not enough to identify the referent in a video. Therefore, we further introduce an action-aware reasoning stage to highlight entities matched with the temporal action cues and fuse them with entities matched with the static relationship cues to distinguish the referent.
As shown in Fig. 4, in order to highlight features of entities which are matched with the temporal action cues in the expression, we conduct cross-modal attention between and with necessary reshaping operations:
We further conduct graph convolution among temporal vertexes as follows:
3 Text-Guided Feature Exchange
where denotes sigmoid function. At each round, feature of each level will select its relevant features from the other two levels under the guidance of textual information. After rounds of exchange, the output features , and are further fused with ConvLSTM to produce the final mask prediction.
Experiments
We conduct extensive experiments on four referring image segmentation benchmarks including UNC , UNC+ , G-Ref and ReferIt , and also on three referring video segmentation benchmarks including A2D Sentences , J-HMDB Sentences and Refer-Youtube-VOS .
UNC, UNC+ and G-Ref are all collected on MS-COCO . They contain , and images with , and referring expressions for over objects, respectively. Expressions in UNC+ contain no location words while those in G-Ref have much longer length than others. ReferIt is collected on IAPR TC-12 and contains images with expressions for objects (including stuff). A2D Sentences is extended from the Actor-Action Dataset by providing textual descriptions for each video. It contains videos annotated with action classes performed by actor classes. J-HMDB sentences is extended from the J-HMDB dataset which contains different actions, videos and corresponding sentences. All the actors in JHMDB dataset are humans and one natural language query is annotated to describe the action performed by each actor. Refer-Youtube-VOS is a large-scale referring video segmentation dataset extended from Youtube-VOS dataset which contains videos, objects and expressions with both first-frame expression and full-video expression annotated.
1.2 Evaluation Metrics
We adopt Prec@X and overall Intersection-over-Union (overall IoU) as metrics to evaluate our image model. Prec@X measures the percentage of test samples whose IoU with ground-truth masks are higher than the threshold . Overall IoU accumulates the total intersection regions over total union regions of all the test samples. For video model, we additionally use mean Average Precision (mAP) and mean IoU as metrics in addition to Prec@X and overall IoU.
1.3 Implementation Details
We adopt DeepLab-ResNet101 which is pretrained on the PASCAL-VOC dataset as the CNN backbone to extract visual features for the input image and video. The output of Res, Res and Res are used for multi-level feature fusion. Input images and video frames are resized to . The input video clip contains frames for CMPC-v model. Besides, for video segmentation, we further adopt I3D pretrained on Kinetics dataset as backbone of CMPC-v model to compare fairly with previous methods . Due to the temporal downsamping operation in I3D, the input video clip of our model contains frames following PRPE to retain temporal information throughout the network. The output of last three stages of I3D are used for multi-level feature fusion. Channel dimensions of features are set as and the cell size of ConvLSTM is set to . When comparing with other methods, the hyper-parameter of bilinear fusion is set to and the number of feature exchange rounds is set to . GloVe word embeddings pretrained on Common Crawl 840B tokens are adopted following . Number of graph convolution layers is set as on G-Ref dataset and on others. We train the network using Adam optimizer with the initial learning rate of and weight decay of . Parameters of CNN backbone are fixed during training. Binary cross-entropy loss averaged over all pixels is used for training. DenseCRF is adopted to refine the segmentation masks for fair comparison with prior works.
2 Comparison with State-of-the-arts
To demonstrate the superiority of our method for referring image segmentation, we evaluate it on four benchmark datasets. As illustrated in Table I, our method outperforms previous state-of-the-arts on all the datasets with large margins. Comparing with STEP which densely fuses levels of features for times, our method utilizes fewer levels of features and fusion times while consistently obtaining - performance gains on all the four datasets, demonstrating the effectiveness of our modules. Particularly, our method yields IoU improvement over STEP on G-Ref val set, indicating our method can better handle long sentences with progressive comprehension. Besides, ReferIt is a challenging dataset and previous methods only obtain marginal improvements on it. For instance, STEP and CMSA achieve only and improvements on ReferIt test set respectively, while our method enlarges the performance gain to , which shows that our model can well segment both objects and stuff. In addition, our method also outperforms MAttNet by a large margin in overall IoU. MAttNet depends on Mask R-CNN pretrained on much more COCO images () than ours pretrained on PASCAL-VOC images () to generate RoI proposals. Therefore, it may not be completely fair to directly compare performances of MAttNet with ours.
2.2 Referring Video Segmentation
Comparisons with state-of-the-art methods on A2D Sentences dataset are summarized in Table II. Since prior works adopt I3D as visual backbone to encode video features, we also build a 3D-version of our CMPC-V network for fair comparison. We present results of both 2D backbone and 3D backbone in Table II, which are denoted as ‘Ours-R2D’ and ‘Ours-I3D’ respectively. Our I3D-based model achieves notable improvements comparing with the R2D-based model, indicating that 3D backbone extracts more temporal information. Our method also outperforms previous state-of-the-arts, PRPE , on most evaluation metrics except Overall IoU, where our model achieves comparable result with PRPE.
To further demonstrate the generalization ability of our video model, we conduct experiments on the J-HMDB Sentences dataset and the Refer-Youtube-VOS dataset . We follow prior works to use the best model pretrained on A2D Sentences dataset to directly evaluate all the test samples of J-HMDB Sentences dataset without finetuning. As shown in Table III, our video model achieves notable performance gain over previous methods for most evaluation metrics ( Overall IoU, Prec@, etc.), indicating that our method obtains stronger generalization ability. Please note that all the methods including ours produce or on Prec@, which is probably because without training on J-HMDB Sentences, models cannot generate particularly fine masks on unseen samples.
For Refer-Youtube-VOS dataset, we train our video model for iterations with as initial learning rate (divided by at th iteration). As shown in Table IV, Our video model outperforms URVOS on most metrics except Prec@ without further refining segmentation masks using memory attention between frames. The comparison shows that our model is able to recognize the referred objects without too much interactions among frames.
3 Ablation Studies
We perform ablation studies on UNC val set, G-Ref val set and A2D Sentence test set to testify the effectiveness of each proposed module for referring image and video segmentation.
We first explore the effectiveness of each component of our proposed CMPC-I module. Experimental results are summarized in Table V. EP and RAR denotes the entity perception and relation-aware reasoning stages in CMPC-I module respectively. GloVe means using GloVe word embeddings to initialize the parameters of embedding layer. CMF means concatenating multimodal feature from EP instead of pure visual feature to produce the output of CMPC-I, as mentioned in the last paragraph of Section 3.1.3. Results in rows to are all based on single-level visual feature, i.e. Res. Our baseline (row ) simply concatenates the visual feature from DeepLab- and linguistic feature from LSTM and makes predictions on the fusion of them.
As shown in row of Table V, including EP brings IoU improvement over the baseline, indicating the perception of candidate entities can help model to eliminate noisy backgrounds. In row , RAR alone brings IoU improvement over baseline, demonstrating that the referent can be effectively highlighted by leveraging relational words as routers to reason among spatial regions, thus boosting the performance notably. Combining EP with RAR, our CMPC-I module can achieve IoU with single level visual feature, enlarging the performance margin to IoU. This indicates that our model can accurately identify the referent by progressively comprehending the input image and expression. Integrated with GloVe word embeddings, the IoU gain achieves with the aid of large-scale corpus. As shown in row ad , our proposed TGFE has significant influence on the performance while ConvLSTM only yields marginal improvements, demonstrating the effectiveness of TGFE. Particularly, CMF further boosts the performance gain to , which shows concatenating multimodal feature can provide richer context than pure visual feature.
We further conduct ablation studies based on multi-level visual features as shown in rows to of Table V. Row is the multi-level version of row using ConvLSTM to fuse multi-level features. The TGFE module in rows to adopts single round of feature exchange. As shown in Table V, our model yields consistent performance gains with the single-level version, which demonstrates the effectiveness of our CMPC-I module under multi-level situation.
3.2 Components of CMPC-V Module
We build the CMPC-V module based on CMPC-I module by introducing an additional action-aware reasoning (AAR) stage to exploit temporal information for identifying the referent. Our video-version baseline contains GloVe and CMF by default and we evaluate TGFE, EP, RAR and AAR respectively for clarity. The experimental results are summarized in Table VI. We can observe that TGFE can bring notable gains on most of metrics benefited from multi-level visual features. Combining EP and RAR stages can improve the performance slightly by utilizing only spatial information. It should be noticed that incorporating our proposed AAR stage is able to further obtain large performance gain over our strong baseline using TGFE, EP and RAR stages, demonstrating that temporal context information is critical to the referring video segmentation task.
We also tried to use action words as routers to obtain the adjacency matrix of AAR as in RAR, and the results are shown in Table VII. AR and DR represent “Adaptive Relevance” in RAR and “Direct Relevance” in original AAR respectively. Adaptive relevance yields inferior performance than direct relevance, indicating that direct relevance is more suitable to propagate information among different frames.
3.3 TGFE module
Table VIII presents the ablation results of TGFE module for referring image segmentation. is the number of feature exchange rounds. Our experiments are conducted upon CMPC-I module without CMF. Results show that only one round of feature exchange in TGFE could improve the IoU from to . The IoU performance increases as the number of feature exchange rounds increases, which well proves the effectiveness of our TGFE module. We further directly incorporate TGFE module with the baseline model and results are shown in row and row of Table V. TGFE with single round of feature exchange boosts the IoU from to , indicating our TGFE module can effectively utilize rich contexts in multi-level features.
We also tried to remove the additional necessary words extracting process in TGFE stage and the results are shown in Table IX. Results shows that extracting necessary words features can yield slight improvements. It indicates the words extraction is not redundant.
3.4 Number of Graph Convolution Layer
In Table X, we explore the number of graph convolution layers in CMPC-I module based on single-level feature without CMF. denotes the number of graph convolution layers in CMPC-I. Results on UNC val set show that more graph convolution layers lead to performance degradation. However, on G-Ref val set, layers of graph convolution in CMPC-I obtains better performance than layer while layers decreasing the performance. Since G-Ref has much longer expressions (average length of words) than UNC (average length words), we suppose that stacking more graph convolution layers in CMPC-I can appropriately improve the reasoning ability for longer referring expressions. However, too many graph convolution layers may introduce noises and harm the performance.
4 Visualization Analyses
As mentioned in Section 3.1.2, we supervise the learning of word classification by the final segmentation loss due to the lack of annotations for word types. The types of words, i.e., entity, attribute, relation and unnecessary, are implicitly defined according to the role each word plays in the whole expression and it is hard to quantitatively evaluate words classification accuracy. Thus we show the visualization of the words classification probabilities which are predicted by our model in Fig. 5 (d). Among the three blue blocks, from left to right are the probabilities of the word being entity, attribute and relation types where darker color denotes larger probability. In the first example of expression “back right top donut”, words “back”, “right” and “top” have largest probabilities of being relation type, while word “donut” has largest probability of being entity type. This indicates our model can well recognize the type of each word in a soft manner and further utilize these words to highlight features of the referent by our CMPC module.
4.2 Correlations of Word Classification and Segmentation
To demonstrate our model can learn meaningful word classification results, we randomly set the classification probabilities for words in the expression during testing. Experimental results are summarized in Table XI. Assigning random classification probabilities to each word leads to notable performance degradation, which shows the implicit learning of word classification is able to guide the progressive comprehension of referring expressions.
We also visualized the segmentation results with random word classification probabilities in the expressions during testing in Fig. 6. The segmentation results show that with randomly assigned word categories, the model cannot identify the correct object described in the expression. Taking the first row as an example, after modifying the word categories, the model mis-recognizes the man in the middle of the image as referred object, while model with original word categories could make correct prediction. These results show that our model can learn meaningful word classification results without direct supervision.
4.3 Qualitative Results
In Fig. 7, we presents qualitative comparison between the multi-level baseline model (row in Table V) and our full model (row in Table V) for referring image segmentation. From the top-left example we can observe that the baseline model fails to make clear judgement between the two girls, while our model is able to distinguish the correct girl involving the relationship with the phone. Similar result is shown in the top-right example of Fig. 7. As illustrated in the bottom row of Fig. 7, attributes and location relationship can also be well handled by our model, demonstrating its effectiveness.
We also provide qualitative results of our full video model (row in Table VI) and baseline model (row in Table VI) in Fig. 8. Colors of expressions correspond to masks of instances. As shown in Fig. 8 (c) and (d), our full video model well segments the baby and dog with coherent masks while the baseline model fails to distinguish different actors, indicating the effectiveness of our proposed modules.
4.4 Visualization of Affinity Maps
In Fig. 9, we visualize the affinity maps between multimodal feature and the first word of the expression in our CMPC-I module. As shown in (b) and (c), our model can progressively produce more concentrated responses on the referent as the expression becomes more informative from only entity words to the full sentence. It should be noticed that when we manually modify the expression to refer to other entities in the image, our model can still comprehend the new expression and correctly identify the referent. For example, in the third row of Fig. 9(e), when the expression changes from “Donut at the bottom” to “Donut at the left”, high response area shifts from bottom donut to the left donut accordingly. It indicates that our model can adapt to new expressions flexibly.
4.5 Effectiveness of Relation-Aware Reasoning and Action-Aware Reasoning
To verify the effectiveness of relation-aware reasoning and action-aware reasoning, we present the comparison with and without them in Fig. 10.
First, we show the comparison of with relation-aware reasoning (EP + RAR) and without relation-aware reasoning (EP) in the top rows of Fig. 10. It is shown that comparing with EP, EP + RAR is able to discriminate the right referent from others. Taking the nd row as an example, there are several boys in the images and EP only fails to recognize the right most boy. While EP + RAR makes the right prediction.
In addition, we also show the comparison of with action-aware reasoning (EP + RAR + AAR) and without action-aware reasoning (EP + RAR) in the bottom of Fig. 10. In the first row, EP + RAR fails to discriminate the man who is moving his head up and down from others who also stand behind camera. With AAR to recognize action in the video, EP + RAR + AAR is able to locate the correct man. The visualization results demonstrate the necessity of action reasoning.
Conclusion and Future Work
To address the referring segmentation problem for image and video, we propose a CMPC scheme which first perceives candidate entities which might be considered by the expression using entity and attribute words, then conduct graph-based reasoning with the aid of relational words and action words to further highlight the referent while suppressing others. We implement CMPC scheme as two modules, namely CMPC-I and CMPC-V for image and video inputs. We also propose a TGFE module which exploits textual information to selectively integrate multi-level features to refine the mask prediction. Our model consistently outperforms previous state-of-the-art methods on four referring image segmentation benchmarks and three referring video segmentation benchmarks, demonstrating its effectiveness. In the future, we plan to analyze the linguistic information more structurally and explore more compact graph formulation. Our code is available at https://github.com/spyflying/CMPC-Refseg.