Zero-Shot Video Object Segmentation via Attentive Graph Neural Networks

Wenguan Wang, Xiankai Lu, Jianbing Shen, David Crandall, Ling Shao

Introduction

Automatically identifying the primary objects in videos is an important problem that could benefit a wide variety of applications, by reducing or eliminating manual effort needed to process and understand video. However, discovering the most prominent and distinct objects across video frames without having prior knowledge of what those foreground objects are is a challenging task.

Traditional methods tend to tackle this issue by using handcrafted or learnable features in a local or sequential manner. For instance, handcrafted feature based methods use objectness , motion boundary , and saliency cues over a few successive video frames, or explore trajectories , i.e., link optical flow over multiple frames to capture long-term motion information. These are typically non-learning methods working in a purely unsupervised manner. Recent deep learning based methods learn more powerful video object features from large-scale training data, yielding a zero-shot solution (still no annotation used for any testing frame). Many of these employ two-stream networks to combine local motion and appearance information, and apply recurrent neural networks to model the dynamics in a frame-by-frame manner.

Though these methods greatly promoted the development of this field and gained promising results, they generally suffer from two limitations. First, they focus primarily on the local pair-wise or sequential relations between successive frames, while ignoring the ubiquitous, high-order relationships among the frames (since frames from the same video are usually correlated). Second, since they do not fully leverage the rich relationships, they fail to completely capture the video content and hence may easily get inferior foreground estimates. From another perspective, as video objects usually suffer from underlying object occlusions, huge scale variations and appearance changes (Fig. 1 (a)), it is difficult to correctly infer the foreground when only considering successive or local pair-wise relations in videos.

To alleviate these issues, we need to explore an effective framework that can comprehensively model the high-order relationships among video frames into modern neural networks. In this work, an attentive graph neural network (AGNN) is proposed to addresses zero-shot video object segmentation (ZVOS), which recasts ZVOS as an end-to-end, message passing based graph information fusion procedure (Fig. 1 (b)). Specifically, we construct a fully connected graph where video frames are represented as nodes and the pair-wise relations between two frames are described as the edge between their corresponding nodes. The correlation between two frames is efficiently captured by an attention mechanism, which avoids time-consuming optical flow estimation . By using recursive message passing to iteratively propagate information over the graph, i.e., each node receives the information from other nodes, AGNN can capture higher-order relationships among video frames and obtain more optimal results from a global view. In addition, as video object segmentation is a per-pixel prediction task, AGNN has a desirable, spatial information preserving property, which significantly distinguishes it from previous fully connected graph neural networks (GNNs).

AGNN operates on multiple frames, bringing the added advantage of natural training data augmentation, as the combination candidates are numerous. In addition, since AGNN offers a powerful tool for representing and mining much richer and higher-order relationships among video frames, it brings a more complete understanding of video content. More significantly, due to its recursive property, AGNN is flexible enough to process variable numbers of nodes during inference, enabling it to consider more input information and gain better performance (Fig. 1 (c)).

We extensively evaluate AGNN on three widely-used video object segmentation datasets, namely DAVIS16 , Youtube-Objects and DAVIS17 , showing its superior performance over current state-of-the-art methods.

AGNN is a fully differential, end-to-end trainable framework that allows rich and high-order relations among frames (images) to be captured and is highly applicable to spatial prediction problems. To further demonstrate its advantages and generalizability, we apply AGNN to an additional task: image object co-segmentation (IOCS), which aims to extract the common objects from a group of semantically related images. It also gains promising results on two popular IOCS benchmarks, PASCAL VOC and Internet , compared to existing IOCS methods.

Experiments on the ZVOS and additional IOCS tasks clearly demonstrate that AGNN is able to not only capture the relationships among correlated video frame images, but also mine the semantics among semantically related static images. Notably, this work can be viewed as a very early attempt to apply and extend GNNs for pixel-wise prediction tasks, which provides an effective video object segmentation solution and new insight into this task.

Related Work

GNN was first proposed in and further developed in to handle the underlying relationships among structured data. In , recurrent neural networks were used to model the state of each node, and the underlying correlation between nodes are learned via parameterized message passing over neighbors. Li et al. further adapted GNN to sequential outputs. Gilmer et al. Later formulated the message passing module in GNNs as a learnable neural network. Recently, GNNs have been successfully applied in many fields, including molecular biology , computer vision , machine learning and natural language processing . Another popular trend in GNNs is to generalize the convolutional architecture over arbitrary graph-structured data , which is called graph convolution neural network (GCNN).

The proposed AGNN falls into the former category; it is a message passing based GNN, where all the nodes, edges, and message passing functions are parameterized by neural networks. It shares the general idea of mining relationships over graphs but has significant differences. First, our AGNN is unique in its spatial information preserving nature, which is opposed to conventional fully connected GNNs and crucial for per-pixel prediction task. Second, to efficiently capture the relationship between two image frames, we introduce a differentiable attention mechanism which addresses the correlated information and produces further discriminative edge features. Third, as far as we know, there is no prior attempt to explore GNNs in ZVOS.

2 Automatic Video Object Segmentation

To automatically separate primary objects from the background, conventional methods typically use handcrafted features (e.g., color, optical flow) and certain heuristic assumptions related to the foreground (i.e., local motion differences , background priors ). Some others explore more efficient object representations, such as dense point trajectories or object proposals . Most of these methods work in a purely unsupervised manner without using any training data.

Recently, with the renaissance of deep learning, more research efforts have been devoted to tackling this in deep learning frameworks, leading to a zero-shot solution . For instance, a multi-layer perception based detector was designed in to detect moving objectness. Li et al. integrated deep learning based instance embedding and motion saliency to boost performance. Some others turned to fully convolutional networks (FCNs) . They introduced two-stream networks to fuse appearance and motion information , or explored more efficient feature extraction models and LSTM variants , to better locate the foreground objects.

The differences from previous methods are multifold: our AGNN 1) provides a unified, end-to-end trainable, graph model based ZVOS solution; 2) efficiently mines diverse and high-order relations within videos, through iteratively propagating and fusing messages over the graph; and 3) utilizes a differentiable attention mechanism to capture the correlated information between frame pairs.

3 Image Object Co-Segmentation

IOCS aims to jointly segment common objects belonging to the same semantic class in a given set of related images. Early methods usually formulate IOCS as an energy function defined over the whole or a part of the image set and consider intra- and inter-image cues . To capture the relationships between images, some methods applied scene matching techniques , global appearance models , discriminative clustering methodologies , manifold ranking or saliency heuristics . There are only a very few deep IOCS models , mainly due to the lack of a proper, end-to-end modeling strategy for this problem. tackled IOCS through a pair-wise comparison protocol and employed a Siamese network to capture the similarity between two related images. Our AGNN based ICOS solution is significantly different from . First, consider IOCS as a pair-wise image matching problem, while we formulate IOCS as an information propagation and fusion process among multiple images. That means our model can capture richer relations from a global view. Second, the Siamese network based systems only handle pair-wise relations, while our message passing based iterative inference can learn higher-order relations among multiple images. Third, our method is based on the graph model, yielding a more general and elegant framework for modeling IOCS.

Our Algorithm

Before elaborating on our proposed AGNN (§3.2), we first give a brief introduction to generic formulations of GNN models (§3.1). Finally, in §3.3, we provide detailed information on our network architecture.

Based on deep neural networks and graph theory, GNNs are powerful for collectively aggregating information from data represented in graph domains . Specifically, a GNN model is defined according to a graph G ⁣= ⁣(V,E)\mathcal{G}\!=\!(\mathcal{V},\mathcal{E}). Each node vi ⁣∈ ⁣Vv_{i}\!\in\!\mathcal{V} takes a unique value from {1,…,∣V∣}\{1,\dots,|\mathcal{V}|\}, is associated with an initial node representation (or node state or node embedding) vi\mathbf{v}_{i}. Each edge ei,j ⁣∈ ⁣Ee_{i,j}\!\in\!\mathcal{E} is a pair ei,j ⁣= ⁣(vi,vj) ⁣∈ ⁣∣V∣ ⁣× ⁣∣V∣e_{i,j}\!=\!(v_{i},v_{j})\!\in\!|\mathcal{V}|\!\times\!|\mathcal{V}|, with an edge representation ei,j\mathbf{e}_{i,j}. For each node viv_{i}, we learn an updated node representation hi\mathbf{h}_{i} through aggregating representations of its neighbors. Here hi\mathbf{h}_{i} is used to produce an output oi\mathbf{o}_{i}, i.e., a node label. More specifically, GNNs map graph G\mathcal{G} to the node outputs {oi}i=1∣V∣\{\mathbf{o}_{i}\}_{i=1}^{|\mathcal{V}|} through two phases. First, a parametric message passing phase runs for KK steps, which recursively propagates messages and updates node representations. At the kk-th iteration, for each node viv_{i}, we update its state according to its received message mik\mathbf{m}^{k}_{i} (i.e., summarized information from its neighbors Ni\mathcal{N}_{i}) and its previous state hik−1\mathbf{h}^{k-1}_{i}:

where hi0 ⁣= ⁣vi\mathbf{h}_{i}^{0}\!=\!\mathbf{v}_{i}, M(⋅)M(\cdot) and U(⋅)U(\cdot) are the message function and state update function, respectively. After kk iterations of aggregation, hik\mathbf{h}^{k}_{i} captures the relations within the kk-hop neighborhood of node viv_{i}.

Second, a readout phase maps the node representation hiK\mathbf{h}_{i}^{K} of the final KK-iteration to a node output, through a readout function R(⋅)R(\cdot):

The message function MM, update function UU, and readout function RR are all learned differentiable functions.

Next, we present our AGNN based ZVOS solution, which essentially extends traditional fully connected GNNs to (1) preserve spatial features; and (2) capture pair-wise relations (edges) via a differentiable attention mechanism.

2 Attentive Graph Neural Network

The core idea of our AGNN is to perform KK message propagation iterations over G\mathcal{G} to efficiently mine rich and high-order relations within I\mathcal{I}. This helps to better capture the video content from a global view and obtain more accurate foreground estimates. We then readout the segmentation predictions S^\hat{\mathcal{S}} from the final node states {hiK}i=1N\{\mathbf{h}^{K}_{i}\}_{i=1}^{N}. Next, we describe each component of our model in detail.

FCN-Based Node Embedding. We leverage DeepLabV3 , a classical FCN based semantic segmentation architecture, to extract effective frame features, as node representations (see Fig. 2 (b) and Fig. 3 (a)). For node viv_{i}, its initial embedding hi0\mathbf{h}^{0}_{i} can be computed as:

where hi0\mathbf{h}^{0}_{i} is a 3D tensor feature with W ⁣× ⁣HW\!\times\!H spatial resolution and CC channels, which preserves spatial information as well as high-level semantic information.

Intra-Attention Based Loop-Edge Embedding. A loop-edge ei,i ⁣∈ ⁣Ee_{i,i}\!\in\!\mathcal{E} is a special edge that connects a node to itself. The loop-edge embedding ei,ik\mathbf{e}_{i,i}^{k} is used to capture the intra relations within node representation hik\mathbf{h}_{i}^{k} (i.e., internal frame representation). We formulate ei,ik\mathbf{e}_{i,i}^{k} as an intra-attention mechanism , which has been proven complementary to convolutions and helpful for modeling long-range, multi-level dependencies across image regions . In particular, the intra-attention calculates the response at a position by attending to all the positions within the same node embedding (see Fig. 2 (c) and Fig. 3 (b)):

where ​‘∗*’​ represents the convolution operation, W\mathbf{W}s indicate learnable convolution kernels, and α\alpha is a learnable scale parameter. Eq. 4 makes the output element of each position in hik\mathbf{h}_{i}^{k} encode contextual information as well as its original information, thus enhancing the representability.

Inter-Attention Based Line-Edge Embedding. A line-edge eij ⁣∈ ⁣Ee_{ij}\!\in\!\mathcal{E} connects two different nodes viv_{i} and vjv_{j}. The line-edge embedding ei,jk\mathbf{e}_{i,j}^{k} is used to mine the relation from node viv_{i} to vjv_{j}, in the node embedding space (see Fig. 2 (b)). Here we compute an inter-attention mechanism to capture the bi-directional relations between two nodes viv_{i} and vjv_{j} (see Fig. 2 (c) and Fig. 3 (c)):

Gated Message Aggregation. In our AGNN, for the message passed in the self-loop, we view the loop-edge embedding ei,jk−1\mathbf{e}_{i,j}^{k-1} itself as a message (see Fig. 3 (b)), since it already contains the contextual and original node information (see Eq. 4):

For the message mj,i\mathbf{m}_{j,i} passed from node vjv_{j} to viv_{i} (see Fig. 3 (c)), we have:

where softmax(⋅\cdot) normalizes each row of the input. Thus, each row (position) of mj,ik\mathbf{m}_{j,i}^{k} is a weighted combination of each row (position) of hjk−1\mathbf{h}_{j}^{k-1}, where the weights come from the corresponding column of ei,jk−1\mathbf{e}^{k-1}_{i,j}. In this way, the message function M(⋅)M(\cdot) assigns its edge-weighted feature (i.e., message) to the neighbor nodes . Then, mj,ik\mathbf{m}^{k}_{j,i} is reshaped back to a 3D tensor with a size of W ⁣× ⁣H ⁣× ⁣CW\!\times\!H\!\times\!C.

In addition, because some nodes are noisy due to camera shift or out-of-view, their messages may be useless or even harmful. We apply a learnable gate G(⋅)G(\cdot) to measure the confidence of a message mj,i\mathbf{m}_{j,i}:

where FGAP(⋅)F_{\text{GAP}}(\cdot) indicates the use of global average pooling to generate channel-wise responses, σ\sigma is the logistic sigmoid function σ(x) ⁣= ⁣1/(1 ⁣+ ⁣exp⁡(−x))\sigma(x)\!=\!1/(1\!+\!\exp(-x)), and Wg\mathbf{W}_{g} and bgb_{g} are the trainable convolution kernel and bias.

Following Eq. 1, we collect the messages from the neighbors and self-loop via gated summarization (see Fig. 2 (d)):

where ‘⋆\star’ denotes the channel-wise Hadamard product. Here, the gate mechanism is used to filter out irrelevant information from noisy frames. See §4.3 for a quantitative study of this design.

ConvGRU based Node-State Update. In step kk, after aggregating all the information from the neighbor nodes and itself (Eq. 9), viv_{i} gets a new state hik\mathbf{h}_{i}^{k} by taking into account its prior state hik−1\mathbf{h}_{i}^{k-1} and its received message mik\mathbf{m}_{i}^{k}. To preserve the spatial information conveyed in hik−1\mathbf{h}_{i}^{k-1} and mik\mathbf{m}_{i}^{k}, we leverage ConvGRU to update the node state (Fig. 2 (e)):

ConvGRU is proposed as a convolutional counterpart to previous fully connected GRU , and introduces convolution operation into input-to-state and state-to-state transitions.

Readout Function. After KK message passing iterations, we obtain the final state hiK\mathbf{h}_{i}^{K} for each node viv_{i}. Finally, in the readout phase, we get a segmentation prediction map S^ ⁣∈ ⁣W×H\hat{S}\!\in\!^{W\times H} from hiK\mathbf{h}_{i}^{K} through a readout function R(⋅)R(\cdot) (see Fig. 2 (f)). Slightly different from Eq. 2, we concatenate the final node state hiK\mathbf{h}_{i}^{K} and the original node feature vi\mathbf{v}_{i} (i.e., hi0\mathbf{h}_{i}^{0}) together and feed the combined feature into R(⋅)R(\cdot):

Again, to preserve spatial information, the readout function is implemented as a small FCN network, which has three convolution layers with a sigmoid function to normalize the prediction to $$.

The convolution operations in the intra-attention (Eq. 4) and update function (Eq. 10) are realized with 1 ⁣× ⁣11\!\times\!1 convolutional layers. The readout function (Eq. 11) consists of two 3 ⁣× ⁣33\!\times\!3 convolutional layers cascaded by a 1 ⁣× ⁣11\!\times\!1 convolutional layer. As a message passing based GNN model, these functions share weights among all the nodes. Moreover, all the above functions are carefully designed to avoid disturbing spatial information, which is essential for ZVOS since it is a pixel-wise prediction task.

3 Detailed Network Architecture

Training Phase. As we operate on batches of a certain size (which is allowed to vary, depending on the GPU memory size), we leverage a random sampling strategy to train AGNN. Specifically, we split each training video I\mathcal{I} with a total of NN frames into N′N^{\prime} segments (N′ ⁣≤ ⁣NN^{\prime}\!\leq\!N) and randomly select one frame from each segment. Then we feed the N′N^{\prime} sampled frames into a batch and train AGNN. Thus the relationships among all the N′N^{\prime} sampling frames in each batch are represented using an N′N^{\prime}-node graph. Such a sampling strategy provides robustness to variations and enables the network to fully exploit all frames. The diversity among the samples enables our model to better capture the underlying relationships and improve its generalizability. Let us denote the ground-truth segmentation mask and predicted foreground map for a training frame IiI_{i} as S ⁣∈ ⁣{0,1}60×60S\!\in\!\{0,1\}^{60\times 60} and S^ ⁣∈ ⁣60×60\hat{S}\!\in\!^{60\times 60}. Our model is trained through the weighted binary cross entropy loss (see Fig. 2):

where η\eta indicates the foreground-background pixel number ratio in SS. It is worth mentioning that, as AGNN handles multiple video frames at the same time, it leads to a remarkably efficient training data augmentation strategy, as the combination candidates are numerous. In our experiments, during training, we randomly select 2 videos from the training video set and sample 3 frames (N′ ⁣= ⁣3N^{\prime}\!=\!3) per video, due to the computation limitation. In addition, we set the total number of iterations as K ⁣= ⁣3K\!=\!3. Quantitative experimental settings can be found in §4.3.

Testing Phase. After training, we can apply the learned AGNN model to perform per-pixel object prediction over unseen videos. For an input test video I\mathcal{I} with NN frames (with 473 ⁣× ⁣473473\!\times\!473 resolution), we split I\mathcal{I} into TT subsets: {I1,I2,…,IT}\{\mathcal{I}_{1},\mathcal{I}_{2},\dots,\mathcal{I}_{T}\}, where T ⁣= ⁣N/N′T\!=\!N/N^{\prime}. Each subset contains N′N^{\prime} frames with an interval of TT frames: It ⁣= ⁣{It,It+T,…,IN−T+t}\mathcal{I}_{t}\!=\!\{I_{t},I_{t+T},\dots,I_{N-T+t}\}. Then we feed each subset into AGNN to obtain the segmentation maps of all the frames in the subset. In practice, we set N′ ⁣= ⁣5N^{\prime}\!=\!5 during testing. We quantitatively study this setting in §4.3. As our AGNN does not require time-consuming optical flow computation and processes N′N^{\prime} frames in one feed-forward propagation, it achieves a fast speed of 0.28s0.28s per frame. Following the widely used protocol , we apply CRF as a post-processing step, which takes about 0.50s0.50s per frame. More implementation details can be found in §4.1.1.

Experiments

We first report performance on the main task: unsupervised video object segmentation (§4.1). Then, in §4.2, to further demonstrate the advantages of our AGNN model, we test it on an additional task: image object co-segmentation. Finally, we conduct an ablation study in §4.3.

Datasets and Metrics: We use two well-known datasets:

DAVIS16 is a challenging video object segmentation dataset which consists of 50 videos in total (30 for training and 20 for val) with pixel-wise annotations for every frame. Three evaluation criteria are used in this dataset, i.e., region similarity (Intersection-over-Union) J\mathcal{J}, boundary accuracy F\mathcal{F}, and time stability T\mathcal{T}.

Youtube-Objects comprises 126 video sequences which belong to 10 object categories and contain more than 20,000 frames in total. Following its protocol, we use J\mathcal{J} to measure the segmentation performance.

DAVIS17 consists of 60 videos in the training set, 30 videos in the validation set and 30 videos in the test-dev set. Different from DAVIS2016 and Youtube-Objects, which only focus on object-level video object segmentation, DAVIS17 provides instance-level annotations.

Implementation Details: Following , both static data from image salient object segmentation datasets, MSRA10K , DUT , and video data from the training set of DAVIS16 are iteratively used to train our model. In a ‘static-image’ iteration, we randomly sample 6 images from the static training data to train our backbone network (DeepLabV3) to extract more discriminative foreground features. To train the backbone network, a 1 ⁣× ⁣11\!\times\!1 convolution layer with sigmoid function is appended as an intermediate output layer, which can access the static image supervision signal. This is followed by a ‘dynamic-video’ iteration, in which we use the sampling strategy described in §3.3 to sample 6 video frames to train our whole AGNN model. The ‘static-image’ and ‘dynamic-video’ iterations are executed alternately. To apply the trained AGNN model on DAVIS17, we first use category agnostic mask-RCNN to generate instance-level object proposals for each frame. Then, we run AGNN on the whole video and generate a coarse mask for the primary objects in each frame. Then the object-level masks are used to filter out the proposals from the background and highlight the foreground proposals. Through combining an instance bounding proposals and coarse masks, we obtain the instance-level mask for each primary object. Finally, to connect multiple instances across different frames, we use overlap ratio and optical flow as an association metric to match different instance-level masks.

1.2 Quantitative Performance

Val-set of DAVIS16. We compare the proposed AGNN with the top ZVOS methods from the DAVIS16 benchmarkhttps://davischallenge.org/davis2016/soa_compare.html, deadline: Mar. 2019 . Table 1 shows the detailed results. We can see that our AGNN outperforms the best reported results (i.e., AGS ) on DAVIS16 benchmark by a significant margin in terms of mean J\mathcal{J} (80.7 vs 79.7) and F\mathcal{F} (79.1 vs 77.4). Compared to PDB , which uses the same training protocol and training datasets, our AGNN yields significant performance gains of 3.5%\% and 4.6%\% in terms of mean J\mathcal{J} and mean F\mathcal{F}, respectively.

Youtube-Objects. Table 2 gives the detailed per-category performance and average results on Youtube-Objects. As can be seen, our AGNN performs favorably according to mean J\mathcal{J} criterion. Furthermore, unlike other methods whose performance fluctuates across categories, AGNN mains a stable performance.This further proves its robustness and generalizability.

Test-dev set of DAVIS17. In Table 3 we report the performance comparison with the recent instance-level ZVOS method, RVOS , on the DAVIS17 test-dev set. We can find that AGNN significantly outperforms RVOS over most evaluation criteria.

1.3 Qualitative Performance

Fig. 4 depicts visual results for the proposed AGNN on two challenging video sequences soapbox and judo of DAVIS16 and DAVIS17, respectively. For soapbox, the primary objects undergo huge scale variation, deformation and view changes, but our AGNN still generates accurate foreground segments. Our AGNN also handles judo well, although the different foreground instances suffer from similar appearance and rapid motions.

2 Additional Task: IOCS

Our AGNN model can be viewed as a framework for capturing high-order relations among images (or frames). To demonstrate its generalizability, we extend AGNN for IOCS task. Rather than extracting the foreground objects across multiple relatively similar video frames in videos, IOCS needs to infer the common objects from a group of semantically related images.

Datasets and Metrics: We perform experiments on two well-known IOCS datasets:

PASCAL VOC has 1,464 training images and 1,449 validation images. Following , we split the validation set into 724 validation and 725 test images, and use mean J\mathcal{J} as the performance measure.

Internet contains 1,306 car, 879 horse, and 561 airplane images. Following , we measure the performance on a subset of Internet (100 images per class are sampled) with mean J\mathcal{J}.

Implementation Details: Following , we employ PASCAL VOC to train our model. In each iteration, we randomly sample a group of N′ ⁣= ⁣3N^{\prime}\!=\!3 images that belong to the same semantic class, and feed two groups with randomly selected classes (6 images in total) to the network. All other experimental settings are the same as ZVOS.

After training, we evaluate the performance of our method on the test sets of PASCAL VOC and Internet dataset. When processing an image, IOCS must leverage information from the whole image group (as the images are typically different and some are irrelevant) . To this end, for each image IiI_{i} to be segmented, we uniformly split the other N ⁣− ⁣1N\!-\!1 images into TT groups, where T ⁣= ⁣(N−1)/(N′−1)T\!=\!(N-1)/(N^{\prime}-1). Then we feed the first image group and IiI_{i} to a batch of size N′N^{\prime}, and store the node state for IiI_{i}. After that, we feed the next group and the store node state of IiI_{i} to get a new state of IiI_{i}. After TT steps, the final state of IiI_{i} contains its relationships to all other images and is used to produce its final co-segmentation result.

2.2 Quantitative Performance

PASCAL VOC. It is very challenging to segment the common objects in this dataset, since the objects undergo drastic variation in scale, position and appearance. In addition, some images have multiple objects belonging to different categories. On this dataset, we compare AGNN with six representative methods, including Siamese-based co-segmentation methods , as well as deep semantic segmentation models (e.g., FCNs ).

Table 4 shows detailed results in terms of mean J\mathcal{J}. FCNs segment each image individually (without considering other related images), and thus give poor performance. Both and consider pairs of images and gain better results. Our AGNN achieves the best performance because it considers high-order information from multiple images during inference, enabling it to capture richer semantic relations within the image groups.

Internet. We evaluate our model (pre-trained on PASCAL VOC) on Internet . Quantitative results in Table 5 again demonstrate the superiority of AGNN (4.5% performance gain compared with the second best method). The result of AGNN is higher than compared methods for three classes: Car (84.0%), Horse (72.6%), Airplane (76.1%).

2.3 Qualitative Results

Fig. 5 shows some sample results. Specifically, the first four images in the top row belong to the Cat category (red circle), while the last four images contain the Person category (yellow circle) with significant intra-class variation. For both cases, our AGNN successfully detects the common object instances amongst background clutter. For the second row, AGNN also performs well in the cases with remarkable intra-class appearance change.

3 Ablation Study

We perform an ablation study on DAVIS16 to investigate the effect of each essential component of AGNN.

Effectiveness of Our AGNN. To quantify the contribution of our AGNN, we derive a baseline w/o. AGNN, which indicates the results from our backbone model, DeepLabV3. As shown in Table 6, AGNN indeed brings significant performance improvements (72.2→\rightarrow80.6 in term of mean J\mathcal{J}).

Gated Message Aggregation Strategy. In Eq. 9, we equip the message passing with a channel-wise gated mechanism to decrease the negative influence of irrelevant frames. To evaluate this design, we offer a baseline w/o. Gated Message, which aggregates messages directly. A performance degradation is observed after excluding the gates.

Message Passing Iterations KK. To investigate the message passing iterations KK, we report the performance as a function of KKs. We find that, with more iterations (1 ⁣→ ⁣31\!\rightarrow\!3), better results can be obtained. The performance of the message passing converges at K ⁣= ⁣3K\!=\!3.

Node Numbers N′N^{\prime} During Inference. To evaluate the impact of the number of nodes N′N^{\prime} during inference, we report performance with different values of N′N^{\prime}. We observe that, with more input frames (3 ⁣→ ⁣53\!\rightarrow\!5), the performance raises accordingly. When even more frames are considered (5 ⁣→ ⁣75\!\rightarrow\!7), the final performance does not change obviously. This may be due to the redundant content in video sequences.

Conclusion

This paper proposes a novel AGNN based ZVOS framework for capturing the relations among videos frames and inferring the common foreground objects. It leverages an attention mechanism to capture the similarity between nodes and performs recursive message passing to mine the underlying high-order correlations. Meanwhile, we demonstrate the generalizability of AGNN by extending it to IOCS task. Extensive experiments on three ZVOS and two IOCS datasets indicate that our AGNN performs favorably against current state-of-the-art methods. This further illustrates the importance of AGNN which can capture diverse relations among similar video frames or semantically related images.

Acknowledgements This work was supported in part by ARO grant W911NF-18-1-0296, Beijing Natural Science Foundation under Grant 4182056, CCF-Tencent Open Fund, Zhijiang Lab’s International Talent Fund for Young Professionals, and the National Science Foundation (CAREER IIS-1253549).

References