Local-Global Context Aware Transformer for Language-Guided Video Segmentation

Chen Liang, Wenguan Wang, Tianfei Zhou, Jiaxu Miao, Yawei Luo, Yi Yang

Introduction

Language-guided video segmentation (LVS) , also known as language-queried video actor segmentation , aims to segment a specific object/actor in a video referred by a linguistic phrase. LVS is a challenging task as it delivers high demands for understanding whole video content and diverse language concepts, and, most essentially, comprehending alignments between linguistic and spatio-temporal visual clues at pixel level. Till now, popular LVS solutions are mainly upon FCN-style architectures with different visual-linguistic information fusion modules, such as dynamic convolution , cross-modal attention .

Though impressive, most existing LVS solutions fail to fully exploit the global video context and the connection between global video context and language descriptions. This is because they usually adopt 3D CNNs for video representation learning. Due to the heavy computational cost and locality nature of 3D kernels, many algorithms can only gather limited information from short video clips (with 8-/16-frame length ) but discard long-term temporal context, which, however, is crucial for LVS task. First, due to the underlying low-level video processing challenges such as object occlusion, appearance change and fast motion, it is hard to segment the target referent safely by only considering short-term temporal context . Second, from a perspective of high-level visual-linguistic semantics understanding, short-term temporal context is insufficient for reaching a complete understanding of video content and struggles for tackling complex activity related phrases. As shown in Fig. 1, given the description of “a boy is skiing on the snow and suddenly falling ⋯\cdots”, the target skier cannot be identified unless with a holistic understanding of the video content, as the ‘fall’ event happens in the second half of the video.

To cope with these temporal and cross-modal challenges posed in LVS task, we propose a new model, named as Locater (local-global context aware Transformer). Building upon Transformer encoder-decoder architecture , Locater first utilizes self-attention to enhance per-frame representations with intra-frame visual context and linguistic features. To further incorporate temporal cues into per-frame representation, Locater builds an external, finite memory, which encodes multi-temporal-scale context and also bases content retrieve on the attention operation. This makes Locater a fully attentional model, yet with greatly reduced space and computation complexities.

In particular, the external memory has two components: i) the former is to persistently memorize global temporal context, i.e., highly compact descriptors summarized from frames sampled over the entire span of a video; ii) the latter is to online gather local temporal context and segmentation history, from past segmented frames. The global memory is maintained unchanged during the whole segmentation procedure while the local memory is dynamically updated with the segmentation processes. Hence, Locater gains a holistic understanding of video content and captures temporal coherence, leading to contextualized visual representation learning. Conditioned on the stored context and particular content of one frame, Locater vividly interprets the expression by adaptively attending to informative words, and forms an expressive query vector that specifically suits that frame. The specific query vector is then used to query corresponding contextualized visual feature for mask decoding.

With such a memory design, Locater is capable of comprehensively modeling temporal dependencies and cross-modal interactions in LVS, and processing arbitrary length videos with low time complexity O(N)\mathcal{O}(N) and constant space cost O(1)\mathcal{O}(1). In contrast, Transformer-style self-attention requires to maintain all O(N2)\mathcal{O}(N^{2}) cross-frame dependencies. In addition, we introduce a deeply supervised learning strategy that feeds supervision signal into intermediate layers of Locater, for easing training and boosting performance.

We further notice that in A2D-S , the most popular LVS dataset, a large portion of videos only contain very few but obvious objects/actors, making such task trivial. To address this limitation, we synthesize A2D-S+, a harder dataset, from A2D-S. Each video in A2D-S+ is either selected or created to contain several semantically similar objects through a semi-automatic contrasting sampling process. It doubles the dataset difficulty in terms of the number of grounding-required examples per video, while with a negligible cost of human labour.

The contributions can be summarized into three folds:

We propose the pilot work that tackles LVS task with a memory augmented, fully attentional Transformer framework. Several essential designs (i.e., finite memory, progressive cross-modal fusion, contextualized query embedding, deeply supervision) significantly facilitate the network learning and finally lead to impressive performance.

The finite memory enables elegant long-term memorization and extraction of cross-modal context, while in the meantime, getting rid of the unaffordable space and computation cost brought by the quadratic complexity of the conventional attention in Transformers.

We alleviate the excess of trivial cases in the current most popular LVS benchmark, i.e., A2D-S, through introducing a harder synthesized dataset with little human effort.

We empirically demonstrate that Locater consistently surpasses existing state-of-the-arts across three popular ben- chmark datasets (i.e., 3.6%/5.9%/1.1% on A2D-S ⁣{}_{\!}  ⁣{}_{\!}(§5.1)/ J-HMDB-S  ⁣{}_{\!}(§5.2)/R-YTVOS  ⁣{}_{\!}(§5.3) in mIoU/mIoU/ J&F\mathcal{J}\&\mathcal{F}) and also performs robust on our challenging A2D-S+ dataset (§5.5). Moreover, based on Locater, we ranked 1st place in the Referring Video Object Segmentation (RVOS) track in the 3rd Large-scale Video Object Segmentation Challenge ⁣{}_{\!} (YTB-VOS21{}_{\text{21}}), outperforming the runner-up by a large margin (e.g., 11.3% in J&F\mathcal{J}\&\mathcal{F}; see §5.6). Further, we conduct a series of diagnostic experiments on several variants of Locater, verifying both the efficacy and efficiency of our core model designs (§5.7).

Related Work

To offer necessary background, we review literature in LVS (§2.1), and discuss relevant work in other areas (§2.2-2.7).

Studies of LVS are initiated by . With the theme of efficiently capturing the multi-modal nature of the task, existing efforts investigate different visual-linguistic embedding schemes, such as capsule routing , dynamic convolution , and cross-modal attention . Owning to the difficulty in handling variable duration of videos, existing algorithms are mainly built on 3D CNNs, suffering from an intrinsic limitation in modeling dependencies among distant frames . To remedy this issue, we exploit self-attention to gather cross-modal and temporal cues in a holistic and efficient manner. This is achieved by augmenting Transformer architecture with an explicit memory, which gathers and stores both global and local temporal context with learnable operations. Although also enjoys the advantage of the outside memory, it does not consider global video context and requires multi-round complicated inference with a heuristic memory update rule. In essence, we formulate the task in an encoder-decoder attention framework, instead of following the widely-used CNN-style architecture. Two concurrent works , inspired by the query-based paradigm in object detection and instance segmentation , formulate LVS as a sequence prediction problem and introduce Transformers for querying object sequences. Though effective, they have to translate the whole video as input tokens. In contrast, with the aid of the external memory, Locater yields comparable performance with improved efficiency.

2 Referring Image Segmentation (RIS)

As the counterpart of LVS in image domain, RIS has longer research history, dating back to the work of Hu and others in 2016. Primitive solutions directly fuse concatenated language and image features with a segmentation network to infer the referent mask . More recent approaches focus on designing modules to promote visual-linguistic interactions, e.g., cross-modal attention , progressive dual-modal encoding , linguistic structure .

3 Referring Expression Comprehension (REC)

REC is to localize linguistic phrases in images (phase grounding) or videos (language-guided object tracking, LOT), in a form of bounding box. The majority of studies, to date, follow a two-stage procedure to select the best-matching region from a set of bounding box proposals . Despite their promising results, the performance of two-stage methods is capped by the speed and accuracy of the proposal generator. Alternatively, a few recent works resort to a single-stage paradigm, which embeds linguistic features into one-stage detectors, and directly predicts the bounding box . Compared with REC, LVS is more challenging since it requires pixel-level joint video-language understanding. As for LOT, it is an emerging research domain . Existing solutions typically equip famous trackers with a language grounding module . Note that there are many differences between LOT and LVS: i) they explain phrases with different visual constituents (pixel vs tight bounding box); ii) LOT adopts first-frame oriented phrases due to its tracking nature, while LVS is more aware of video understanding, i.e., considering arbitrary objects and unconstrained referring expressions .

4 Semi-automatic absent{}_{\!}Video absent{}_{\!}Object absent{}_{\!}Segmentation absent{}_{\!}(SVOS)

SVOS aims to selectively segment video objects based on first-frame masks. LVS also has a close connection with SVOS, as it replaces the expensive pixel-level intervention with easily acquired linguistic guidance. Depending on the utilization of test-time supervision, current SVOS models can be categorized into three classes : i) online fine-tuning based methods first train a generic segmentation network and then fine-tune it with the given masks ; ii) propagation based methods use the previous frame mask to infer the current one ; and iii) matching based methods classify each pixel’s label according to its similarity to the annotated target . Although some recent matching based SVOS approaches also leverage memory to reuse past segmentation information, we focus on a multi-modal task, and enhance Transformer self-attention with a fixed size memory to address the efficiency issue in modeling intra- and inter-modal long-term dependencies.

5 Transformer in Vision-Language Tasks

The remarkable successes of Transformer in NLP spur increasing efforts applying Transformer for vision-language tasks, e.g., video captioning , text-to-visual retrieval , visual question answering , and temporal language localization . For generating coherent captions, also utilizes memory to better summarize history information. However, our target task, LVS, requires fine-grained grounding on the spatio-temporal visual space, and our Locater exploits local and global video context as well as segmentation history in a comprehensive and efficient manner. Transformer-based pretrained models also showed great potential in joint image-language embedding . Some other efforts were made towards unsupervised video-language representation learning , while they still often utilize 3D CNNs for video content encoding.

6 Transformer in Visual Referring Tasks

Concurrent to our work, only a handful of attempts apply Transformer-style models for LVS , RIS , and REC . Compared with these efforts, our approach is unique in several aspects, including memory design, language-guided progressive encoding, contextualized query embedding, and deep supervision strategy.

7 Neural Networks with External Memory

Modeling structures within sequential data has long been considered a foundational problem in machine learning. Recurrent networks (RNNs) , e.g., LSTM , GRU , as a traditional class of methods, have been shown Turing-Complete and successfully applied into a wide range of relevant fields, e.g., machine translation , and video recognition . However, due to the limited capacity of the latent network states, they struggle for long-term context modeling. To address this limitation, researchers enrich RNNs’ dynamic storage updates with external memory design, whose long-term memorization/reasoning with explicit storage manipulations earns remarkable success in robot control , language modeling , etc.

Drawing inspiration from these novel designs, we equip Locater with an external memory. This allows for persistent and precise information storage, as well as flexible manipulations using learnable read/write operations. Hence such memory-augmented architecture well supports long-term modeling, which is critical for LVS.

Methodology

Combining several paralleled single-head attention, one can derive multi-head attention. In Transformer, each encoder block has two layers, i.e., multi-head self-attention (MSA) and a multi-layer perceptron (MLP). MSA is a variant of multi-head attention, where the query, key and value are from the same source. Each encoder block is formulated as:

where residual connection and layer-normalization are applied; X\bm{X} and Y\bm{Y} are the input and output, respectively.

The decoder block has a similar structure. A multi-head attention layer is first adopted, where a specific query embedding is generated to gather information from the encoder side, with the output of corresponding encoder block as key and value embeddings.

Despite the strong expressivity, applying Transformer to LVS is not trivial. As the self-attention function (cf. Eq. 1) comes with quadratic time and space complexity, letting Transformer capture all the intra- and inter-modal relations is not feasible. To address the efficiency issue, some studies exploit sparse attention , linear attention , and recurrence . With a similar spirit, we develop Locater that specifically focuses on LVS and models temporal and cross-modal dependencies with constant size memory and time complexity linear in sequence length.

2 Local-Global Context Aware Transformer

Given the input video {It}t ⁣\{I_{t}\}_{t\!} and language expression EE, our Locater mainly consists of three parts (see Fig. 2): i) a visual-linguistic encoder (§3.2.1) that gradually fuses the linguistic embedding E\bm{E} into visual embedding It\bm{I}_{t} and genera- tes ⁣{}_{\!} a ⁣{}_{\!} language-enhanced ⁣{}_{\!} visual ⁣{}_{\!} feature ⁣{}_{\!} Vt\bm{V}_{t} for ⁣{}_{\!} each ⁣{}_{\!} frame ⁣{}_{\!} ItI_{t}; ii) a local-global memory (§3.2.2) that gathers diverse temporal context ⁣{}_{\!} from ⁣{}_{\!} {It}t\{I_{t}\}_{t} so ⁣{}_{\!} as ⁣{}_{\!} to ⁣{}_{\!} render ⁣{}_{\!} Vt\bm{V}_{t} as ⁣{}_{\!} a ⁣{}_{\!} contextualized ⁣{}_{\!} fea- ture ⁣{}_{\!} Gt\bm{G}_{t} and ⁣{}_{\!} comprehend ⁣{}_{\!} the ⁣{}_{\!} expression ⁣{}_{\!} EE into ⁣{}_{\!} an ⁣{}_{\!} expressive yet frame-specific query vector qt{\bm{q}}_{t}; and iii) a referring decoder (§3.2.3) that queries Gt ⁣\bm{G}_{t\!} with qt ⁣{\bm{q}}_{t\!} for mask prediction S^t\hat{S}_{t}.

As illustrated in Fig. 2 (a), the visual-linguistic encoder first extracts visual and linguistic features from each frame and the referring expression respectively, and then aggregates them into a compact, language-enhanced visual feature, through self-attention.

Cross-Modality Encoder. To fuse heterogeneous features of each frame image and language expression early, we devise a cross-modality encoder Evw\mathcal{E}_{vw}. It adopts a Transformer encoder architecture, that captures all the correlations between image patches and language words, with several cascaded blocks. Specifically, the kk-th module Fvwk\mathcal{F}_{vw}^{k} is formulated as:

2.2 Local-Global Memory

Although Vt\bm{V}_{t} yields a powerful multi-modal representation, it is less suitable for LVS, as it is computed for each frame ItI_{t} individually without accounting for temporal information. Directly using all the video patches {It,p ⁣}t,p\{\bm{I}_{t,p\!}\}_{t,p} as the visual tokens of the cross-modality attention encoder Evw\mathcal{E}_{vw} will cause unaffordable computational and memory costs, since the complexity of self-attention computation scales quadratically with the sequence length (cf. §3.1). We thus frame a local-global memory module M\mathcal{M}, which helps Locater to capture the temporal context in an efficient manner.

After processing all the Nt′N_{t^{\prime}} sampled frames, the global memory Mg\mathcal{M}^{g} is constructed and maintained unchanged during the whole segmentation process of I\mathcal{I}. Thus Locater can pay continuous attention to global context, so as to better interpret long-term or complex activity related phrases (e.g., “a boy is skiing and suddenly falling” in Fig. ⁣{}_{\!} 1) and handle video ⁣{}_{\!} processing ⁣{}_{\!} challenges ⁣{}_{\!} (e.g., ⁣{}_{\!} occlusion, ⁣{}_{\!} fast ⁣{}_{\!} motion, ⁣{}_{\!} etc.).

With Ml\mathcal{M}^{l}, Locater can access and leverage local temporal context to comprehend simple action-related descriptions (e.g., “a woman is running”), and generate temporally coherent results with the aid of segmentation history.

Contextualized Visual Embedding. Locater next enriches the linguistic-enhanced visual feature Vt\bm{V}_{t} of ItI_{t} with local-global context, by looking up its current memory Mt ⁣= ⁣{Mg,Mtl} ⁣= ⁣{mt,n}n=1Ng+Nl\mathcal{M}_{t}\!=\!\{\mathcal{M}^{g},\mathcal{M}^{l}_{t}\}\!=\!\{{\bm{m}}_{t,n}\}_{n=1}^{N_{g}+N_{l}}. An attention-based read operator is used to retrieve context from the memory Mt\mathcal{M}_{t}:

where ctg{\bm{c}}^{g}_{t} and ctl{\bm{c}}^{l}_{t} refer to the gathered global and local temporal context, respectively. Then we generate a contextualized visual embedding Gt\bm{G}_{t} for frame ItI_{t}:

With this local-global memory-induced attention operation, Gt ⁣\bm{G}_{t\!} encodes multi-modal information and rich temporal context, which is used in the decoder for querying the referent.

2.3 Referring Decoder

In this way, we can obtain a mask sequence S^t ⁣ ⁣= ⁣ ⁣{s^t,p}p=1 ⁣ ⁣Nv ⁣\hat{S}_{t\!}\!=_{\!}\!\{\hat{s}_{t,p}\}_{p=1\!\!}^{N_{v}\!} ∈ ⁣ ⁣Nv ⁣\in_{\!}\!^{N_{v}\!} that collects all the patch-level responses. Recall that each patch is of O ⁣× ⁣OO\!\times\!O size and Nv ⁣ ⁣= ⁣ ⁣WH/O2 ⁣N_{v\!}\!=_{\!}\!WH/O^{2\!}. S^t\hat{S}_{t} is then reshaped into a 2D mask with W/O ⁣× ⁣H/OW/O\!\times\!H/O size and bilinearly upsampled to the original image resolution, i.e., W ⁣× ⁣HW\!\times\!H, as the final segmentation result for the frame ItI_{t}.

Deeply-Supervised Transformer Learning. In practice, we find our Transformer based model is hard to train. Deeply-supervised learning , i.e., introducing supervision signals into intermediate network layers, has been shown effective in CNN-based network training. Given the fact that Transformer also has a layer-by-layer architecture, we suppose that deeply-supervised learning can also help ease the learning of our model. To explore this idea, we separately forward the output features {F2k}k=1K\{\bm{F}_{2}^{k}\}_{k=1}^{K} of all the KK modules in the cross-modality encoder Evw\mathcal{E}_{vw} (Eq. 4), into an auxiliary readout layer. The readout layer is a small MLP followed by sigmoid activation, and learns to predict a segmentation mask. For each training frame II, given the auxiliary outputs {S^k ⁣∈ ⁣W×H ⁣}k=1K\{\hat{S}_{k}\!\in\!^{W\times H\!}\}_{k=1}^{K} from Evw\mathcal{E}_{vw} and final segmentation prediction S^ ⁣∈ ⁣W×H ⁣\hat{S}\!\in\!^{W\times H\!}, the training loss is:

where L\mathcal{L}CE{}_{\text{CE}} refers to pixel-wise, binary cross-entropy loss, S ⁣ ⁣∈ ⁣ ⁣{0,1}H ⁣× ⁣W ⁣ ⁣{S}_{\!}\!\in_{\!}\!\{0,1\}^{H_{\!}\times_{\!}W\!\!} is the ground-truth mask, and λ\lambda balances the two terms. Such a learning strategy can better supervise our early-stage cross-modal information fusion (cf. §3.2.1).

3 Implementation Details

Detailed Architecture. Each video frame ItI_{t} is resized to 320 ⁣× ⁣320320\!\times\!320, and then patch-wise separated with a patch size of 1616, i.e., O ⁣= ⁣16O\!=\!16 and Nv ⁣= ⁣400N_{v}\!=\!400, as in . Each language description is fixed to Nw ⁣ ⁣= ⁣20N_{w\!}\!=\!20 word length, with padding and truncation for the mismatched ones. The visual encoder Ev\mathcal{E}_{v} is implemented as six Transformer encoder blocks, while the linguistic encoder Ew\mathcal{E}_{w} is implemented as a bi-LSTM as in . For the cross-modality encoder Evw\mathcal{E}_{vw}, it has K ⁣ ⁣= ⁣ ⁣3K_{\!}\!=_{\!}\!3 modules. The hidden sizes of all the modules are set to D ⁣= ⁣768D\!=\!768. For the global memory Mg\mathcal{M}^{g} and local memory Ml ⁣\mathcal{M}^{l\!}, we set the capacity as Ng ⁣ ⁣= ⁣ ⁣1.5Nv{N_{g\!}\!=_{\!}\!1.5N_{v}} and Nl ⁣ ⁣= ⁣ ⁣2Nv{N_{l\!}\!=_{\!}\!2N_{v}} respectively. Unless otherwise specified, we sample representative frames with an interval of 1010 frames to construct Mg\mathcal{M}^{g} , i.e., Nt′≈Nt/10N_{t^{\prime}}\approx N_{t}/10. Related experiments can be found in §5.7.

Training Details. Our model is trained for 3030 epochs using Adam optimizer with initial learning rate 4\times10−54\text{\times}{10}^{-5}, batch size 3232 and weight decay 1\times10−41\text{\times}{10}^{-4}. We adopt polynomial annealing policy to schedule the learning rate. The λ\lambda in Eq. 12 is set to 0.40.4 by default. Random horizontal flipping is employed as data augmentation where video frames are flipped with the corresponding positional description being modified from “right” to “left” and vice versa.

Inference. During inference, for each video, Locater first builds the global memory Mg\mathcal{M}^{g} and maintains it unchanged. Then, Locater conducts segmentation in a sequential manner. The local memory Ml\mathcal{M}^{l} is gradually updated during segmentation. The continuous segmentation prediction mask S^\hat{S} for each frame II is binarized with a threshold of 0.50.5.

Our A2D-S+ Dataset

Till now, A2D-S is the most popular LVS dataset. However, in practice, we find that the majority of its test videos (459459 of 757757) only contain one single actor. With these trivial cases, such cross-modal task tends to degrade to a single-modal problem of novel object segmentation. Moreover, many of the remaining videos only contain very few objects yet with distinctive semantics. To better examine the visual grounding capability of LVS models, we construct a harder dataset – A2D-S+. It consists of three subsets, i.e., A2D-SM+{}^{+}_{\text{M}}, A2D-SS+{}^{+}_{\text{S}}, and A2D-S  ⁣T+{}^{+}_{~{}\!\text{T}}, which are all built upon A2D-S but fully aware of the limitation of A2D-S. Specifically, each video in A2D-S+ ⁣{}^{+\!} is selected/created to contain multiple instances of the same object or action category. Thus our A2D-S+ ⁣{}^{+\!} places a higher demand for the grounding ability of LVS models as the distinction among the similar instances is necessary. Next, we will detail the construction process in §4.1 and discuss dataset features and statistics in §4.2.

We first ⁣{}_{\!} curate ⁣{}_{\!} videos ⁣{}_{\!} with ⁣{}_{\!} multiple ⁣{}_{\!} actors ⁣{}_{\!} from ⁣{}_{\!} A2D-S ⁣{}_{\!} test ⁣ ⁣{}_{\!\!} as the subset A2D-SM ⁣ ⁣ ⁣+ ⁣{}^{+\!}_{\text{M}\!\!\!} (A2D-SMultiple+{}^{+}_{\text{Multiple}}). Although each video in A2D-SM ⁣ ⁣ ⁣+ ⁣{}^{+\!}_{\text{M}\!\!\!} has multiple instances, most of these instances are with different semantic categories, making the distinction among them relatively easy. To mitigate this, we further create two challenging subsets: A2D-SS+{}^{+}_{\text{S}} and A2D-S  ⁣T+{}^{+}_{~{}\!\text{T}}, by emp- loying a contrasting sampling strategy , which synthesizes new test samples from videos that contain similar but not exactly the same instances as described by the language expression. Specifically, the contrasting sampling based construction process of A2D-SS+{}^{+}_{\text{S}} and A2D-S  ⁣T+{}^{+}_{~{}\!\text{T}} includes three steps: i) parse free-form language descriptions into semantic-roles (e.g., who did what to whom); ii) for each query video, sample a video with the description of the same semantic-roles structure as the queried description, but the role is realized by a different noun or a verb; and iii) generate a new sample by concatenating the query and sampled videos along the width axis (for A2D-SS+{}^{+}_{\text{S}}) or time axis (for A2D-S  ⁣T+{}^{+}_{~{}\!\text{T}}).

Semantic-Role Labeling (SRL). We first parse each free- form ⁣{}_{\!} language ⁣{}_{\!} description ⁣{}_{\!} in ⁣{}_{\!} A2D-S ⁣{}_{\!} test ⁣{}_{\!} into ⁣{}_{\!} semantic-roles ⁣{}_{\!} , i.e., given the description, answer the high-level question of “who (Arg0) did what (Verb) to whom (Arg1) at where (ARGM-LOC) and when (ARGM-TMP)There are other types of semantic-roles that are not considered in our case, e.g., How (ARGM-MNR), Why (ARGM-CAU), which only applicable for ∼\scriptstyle\sim1%1\% of the language queries in A2D-S test.”. In particu- lar, we first use a BERT-based semantic-role labeling model to predict the semantic-roles for each language description. The model is trained on OntoNote5 with PropBank annotations . Then the labeled sentences are further cleaned by: i) removing words without any roles, e.g., “is”, “are”, “a”, “the”; ii) abandoning non-semantic-role labeled sentences; and iii) lemmatization . An example of our semantic-role labeling process is given in Table I.

Contrasting Sampling. After assigning semantic-roles to language descriptions of A2D-S ⁣{}_{\!} test ⁣{}_{\!} videos, we conduct contrasting sampling to find similar but not exactly the same examples, according to their parsed semantic-role structures. Specifically, for each description, we sample one other description from the dataset that contains at least  ⁣{}_{\!}one, but not all, of semantic roles, which are realized with the same phrase. For example, with a description {(boy, Arg0), (run, Verb), (on grass, ARGM-LOC)}, an eligible sampled description can be {(boy, Arg0), (run, Verb), (on snow, ARGM-LOC)}. Finally, all sampled example pairs are checked by human for diminishing visual and linguistic confusions.

Concatenation Strategies. For each valid contrasting example pair, two concatenation strategies are adopted to merge them into one video, i.e., concatenating along the space axis as an instance of A2D-SS+{}^{+}_{\text{S}} (A2D-SSpatial+{}^{+}_{\text{Spatial}}), and along the time axis as an instance of A2D-S ⁣ T+{}^{+}_{\!~{}\text{T}} (A2D-S ⁣ Temporal+{}^{+}_{\!~{}\text{Temporal}}). Specifically:

For A2D-SS+{}^{+}_{\text{S}}, videos are concatenated along the width axis exclusively, since height-axis concatenation might violate the natural up-down order in real world, e.g., sky should be on top of grass. The sampled videos are resized to have the same height as the query videos, and truncated or padded to fit the time durations of the query videos.

For A2D-S ⁣ T+{}^{+}_{\!~{}\text{T}}, videos are concatenated along the time axis. For a query video with NtN_{t} frames, the sampled contrasting video is resized into the same resolution as the query video. There will be 2Nt2N_{t} frames in each newly constructed example, in which NtN_{t} frames are from the query video, and the other NtN_{t} frames are sampled from the contrasting video. For the frames from the contrasting videos, LVS models are expected to predict all-zero masks to indicate that there is no target referent.

Notably, there is no constraint on the concatenation order of videos. By default, we put the query video at the left and ahead, for spatial and temporal concatenation, respectively. We provide representative examples for A2D-SM,S,T+{}^{+}_{\text{M,S,T}} in Fig. 4.

2 Dataset Features and Statistics

Dataset Features. Through the above contrasting sampling process, A2D-S+ ⁣{}^{+\!} has two distinctive characteristics:

Dense Grounding-ability Required: A2D-S+ ⁣{{}^{+}\!} is manufactured to guarantee that each video does contain multiple semantically similar objects where grounding among them is necessarily required.

Low Human-labour Cost: The entire dataset creation process is semi-automatic so that the human labour cost is greatly reduced. Only a few inevitable efforts have been made to diminish the visual and linguistic ambiguity.

Dataset Statistics. The detailed statistics of A2D-SM,S,T+\text{A2D-S}^{+}_{\text{M,S,T}} are shown in Table II. As our A2D-S+ ⁣{{}^{+}\!} is only used to evaluate LVS models, the statistic of A2D-S test is also provided for clear comparison. As seen, compared to A2D-S test, the newly constructed A2D-SM,S,T+\text{A2D-S}^{+}_{\text{M,S,T}} significantly increase the number of existing semantically similar objects in each video, i.e., 1.761.76 vs 3.333.33/5.025.02/5.135.13, which together provide a stronger testing bed for evaluating LVS methods.

Experiment

Overview. To thoroughly examine the efficacy of Locater, we first report quantitative results on three standard LVS datasets, i.e., A2D Sentences (A2D-S) (§5.1), J-HMDB Sentences (J-HMDB-S) (§5.2), and Refer-Youtube-VOS (R-YTVOS) (§5.3), followed by qualitative results (§5.4). And we further perform experiments on our proposed A2D-S+ dataset in §5.5. Then in §5.6, we report the model performance on the RVOS Track in YTB-VOS21{}_{\text{21}} Challenge ⁣{}_{\!} , where our Locater based solution achieved the 1st place. Later, in §5.7, we conduct a set of ablative studies to examine the core ideas and essential components of Locater. We finally analyse several typical failure modes in §5.8.

Evaluation Criteria. For A2D-S and J-HMDB-S, we follow conventions to use intersection-over-union (IoU) and precision for evaluation. We report overall IoU as the ratio of the total intersection area divided by the total union area over testing samples, as well as mean IoU as the average IoU of all samples. We also measure precision@KK as the percentage of testing samples whose IoU scores are higher than an overlap threshold KK. We report precision at five thresholds ranging from 0.50.5 to 0.90.9 and mean average precision (mAP) over 0.500.50: 0.050.05: 0.950.95. For R-YTVOS and YTB-VOS21{}_{\text{21}}, we follow their standard evaluation protocols to report the region similarity (J\mathcal{J}), contour accuracy (F\mathcal{F}) and their average score J&F\mathcal{J}\&\mathcal{F} over all video sequences .

We first conduct experiments on A2D-S , which is the most popular dataset in the field of LVS.

Dataset. A2D-S contains 3,7823,782 videos with 88 action classes performed by 77 actors, and 6,6556,655 actor and action related descriptions. In each video, 55 to 77 frames are provided with segmentation annotations. As in , we use 3,0173,017/737737 split for train/test, and ignore the 2828 unlabeled videos.

Quantitative Performance. Table III reports the comparison results of Locater against 3D CNN based methods and two latest fully attentional works on the A2D-S test. For fair comparisons with the concurrent competitors , we report the model performance with the strong Video-Swin-T backbone. As seen, Locater yields state-of-the-art performance for all metrics when compared with 3D CNN based methods. Concretely, it advances the SOTA in mAP by 2.9%, Mean IoU by 3.6%, Overall IoU by 2.8%, and also produces great improvements in terms of precision scores under all overlap thresholds. Equipped with the strong Video-Swin-T backbone, Locater achieves comparable or even better performance than the contemporary Transformer based methods . For completeness, we report here the standard deviations of the best performed variants (Row 13): ±\pm  ⁣{}_{\!}0.354 and ±\pm  ⁣{}_{\!}0.518 in terms of mIoU and oIoU. These experimental results well demonstrate the superiority of Locater on local-global semantics understanding brought by the memory design.

Runtime Analysis. Table III reports the efficiency comparison with several famous methods, including four FCN-style models , and two latest fully attentional models . For fairness, we conduct runtime analysis on a video clip of 3636 frames with resolution 512 ⁣× ⁣512512\!\times\!512 and a text sequence of 2020 words length. The window size of 3D CNN/Transformer backbones, i.e., I3D and Video-Swin-T , are all set to 1616. The inference speed is measured on a single NVIDIA GeForce RTX 2080 Ti GPU. As seen, Locater is much faster than existing LVS methods, owing to its memory-augmented fully attentional architecture design.

2 Results on J-HMDB-S Dataset

To investigate the generalization ability of Locater, we also report the performance of our A2D-S trained model on J-HMDB-S , following .

Dataset. J-HMDB-S contains 928928 short videos with 2121 different action categories and 928928 language descriptions.

Quantitative Performance. Table IV presents performance comparison on J-HMDB-S. All models are trained under the same setting on A2D-S train exclusively without fine-tuning. As seen, our model surpasses other competitors across most metrics. Notably, Locater yields Mean IoU 66.3%, Overall IoU 67.3% and mAP 36.3%, while the corresponding scores for the previous SOTA method are 62.7%62.7\%, 65.2%65.2\% and 33.5%33.5\%, respectively. Compared with the recent works , competitive performance is also achieved. These results confirm again the effectiveness of our Locater.

3 Results on R-YTVOS Dataset

We further train and test our model on a recently new proposed large-scale dataset, R-YTVOS .

Dataset. R-YTVOS has 3,9783,978 videos, with 131K segmentation masks and 1515K expressions. Only 3,4713,471/202202 train/ val videos are public available for training and evaluation.

Quantitative Performance. Following the official leaderboard, we report experimental results of Locater along with the SOTA LVS models and two concurrent works in Table V. We observe that Locater still performs well on this large-scale dataset. Locater consistently outperforms these methods on all metrics with the new SOTA scores, i.e., 56.5%/ 54.8%/58.1% J&F\mathcal{J}\&\mathcal{F}/J\mathcal{J}/F\mathcal{F}.

4 Qualitativeabsent{}_{\!} Analysisabsent{}_{\!} onabsent{}_{\!} A2D-Sabsent{}_{\!} andabsent{}_{\!} R-YTVOS

Fig. 5 depicts visual comparison results on A2D-S test (left) and Refer-Youtube-VOS val (right). Locater produces more precise segmentation results against ACGA and CSTM . It shows strong robustness in handling occlusions and complex textual descriptions, especially when facing ambiguity caused by scene dynamics, where the referent is hard to locate relying on narrowly local perspective. Particularly as shown in Fig. 5 (left), beyond the jump action, ACGA (1st row, 3rd column) and CSTM (3rd row, 4th column) both fail to distinguish the two actors.

5 Results on Our A2D-S+ Dataset

Dataset. Basically, A2D-SM+{}^{+}_{\text{M}}, A2D-SS+{}^{+}_{\text{S}}, and A2D-S  ⁣T+{}^{+}_{~{}\!\text{T}} yield increased challenges by collecting or synthesizing videos with multiple objects/actors. For A2D-SM+{}^{+}_{\text{M}}, it is a subset of A2D-S test, which contains 257257 videos with multiple actors. For A2D-SS+{}^{+}_{\text{S}} and A2D-S  ⁣T+{}^{+}_{~{}\!\text{T}}, they are built upon the concept of contrasting examples (cf. §4.1), i.e., videos that are similar to but not exactly the same as described by the language query. They contain 1,2901,290/1,2931,293 distinct videos with 1,2951,295/1,2951,295 linguistic queries respectively.

Quantitative Performance. Our A2D-S+ dataset is only used for evaluation. All the benchmarked methods are the ones in Table III trained on A2D-S train set. As shown in Table VI, although obtain compelling performance on A2D-S and J-HMDB-S (cf. Table III-IV), they cannot handle our challenging cases well. In contrast, our Locater yields better overall performance, especially on mIoU (3.2% averaged improvement over all subsets), verifying its strong ability in fine-grained visual-linguistic understanding.

Performance over Varied Number of Objects. Furthermore, for a thorough evaluation, we study the model performance with varied numbers of objects. For either A2S-SS+{}^{+}_{\text{S}} or A2S-ST+{}^{+}_{\text{T}}, we group the generated examples into three clusters according to the number of existing objects, and report the benchmarked results in Table VII. The performance on the single actor subset of A2D-S test is also provided for reference. As seen, as the number of semantically similar objects increases, the dataset difficulty also gradually elevated. Yet, the performance gap between Locater and other competitors becomes even larger, which well verifies the strong grounding ability of Locater.

6 Results on RVOS Track in YTB-VOS2121{}_{\text{21}} Challenge

Experimental Setup. We first detail the experimental setup for the challenge dataset in YTB-VOS21{}_{\text{21}}.

Dataset: R-YTVOS (cf. §5.3) is the standard benchmark. Challenge solutions are first developed on the test-dev set, which contains the same video sequences as R-YTVOS val, and finally evaluated on the private test-challenge set, which contains 305305 videos.

Evaluation Metric: Following the official evaluation metrics, we use J&F\mathcal{J}\&\mathcal{F}, J\mathcal{J} and F\mathcal{F} to evaluate our model.

Implementation Details: We resize video frames to 384 ⁣× ⁣384384\!\times\!384 and separate them with a patch size O ⁣= ⁣8O\!=\!8, which results in Nv ⁣= ⁣2,304N_{v}\!=\!2,304. The linguistic encoder Ew\mathcal{E}_{w} is implemented as BERTBASE{}_{\text{BASE}} . Our model is trained with an initial learning rate 2\times10−42\text{\times}{10}^{-4} which decays polynomially , batch size 4848, weight decay 1\times10−41\text{\times}{10}^{-4} and max epoch 5050. During testing, we adopt multi-scale inference with horizonal flip and scales of [0.5,0.75,1.0,1.25,1.5,1.75][0.5,0.75,1.0,1.25,1.5,1.75]. Other settings are kept unchanged (cf. §3.3).

Modifications. In addition to the common challenges in LVS benchmarks, we observe a unique issue in YTB-VOS21{}_{\text{21}}: A part of objects in the test-challenge set belongs to novel categories that are unseen in the train set. Thus, along with the Locater trained on YTB-VOS21{}_{\text{21}} challenge dataset, we further ensemble other RES models, i.e., MCN , which are trained on open-set data sources, including image-level datasets like RefCOCO , RefCOCOg and RefCOCO+ . To effectively combine these predictions, we propose a standalone grounding module. Specifically, the grounding module takes several sequence-level mask predictions as inputs and predicts the similarity scores based on the referring expression. It is implemented as four stacked Transformer blocks followed by one MLP layer for score prediction. This module complements the precise segmentation masks from Locater with novel yet coarse object masks from other models.

Quantitative Results. In Table VIII and 3, we respectively report the final results of our final solution and other top-leading teams on the test-dev and test-challenge sets of the RVOS track in YTB-VOS21{}_{\text{21}}. For completeness, in Table VIII, we also report the performance of Locater trained under the same setup to our full solution. Other competitors mainly adopt an image-level referring object grounding strategy and simply generate video-level predictions with a fixed tracking module. They not only neglect the indispensable long-term cues within linguistic expressions but also overlook the intrinsic low-level challenges posed within video sequences; In contrast, our model well-addresses these issues. As seen, our final solution significantly surpasses the second-place solution with a large gap of 11.3%/11.0%/11.7% in terms of J&F\mathcal{J}\&\mathcal{F}/J\mathcal{J}/F\mathcal{F} on test-challenge set. Notably, Locater has already surpassed all other competitors significantly.

7 Diagnostic Experiments

In this section, we conduct a series of ablative studies on both A2D-S test and our newly proposed A2D-S+ to fully examine the efficacy of our algorithm design. For A2D-S+, the score is reported as the average over all the three subsets, i.e., A2D-SM+{}^{+}_{\text{M}}, A2D-SS+{}^{+}_{\text{S}} and A2D-ST+{}^{+}_{\text{T}}.

Key Component Analysis. To study the effect of essential components of Locater, we first establish a baseline, that only remains single-modal encoders and directly concatenates visual and linguistic features for mask prediction (i.e., the first row in Table IXa). Then we gradually add different modules, i.e., cross-modality encoder (Evw\mathcal{E}_{vw}), local-global memory (M\mathcal{M}), and deeply-supervised learning (DSL) strategy, into the baseline. As reported in Table IXa, all these components indeed boost segmentation and combining them together yields the best performance.

Visual Encoder Ev\mathcal{E}_{v}. Next, to verify the advantage of our fully attentional model design, we replace Transformer based visual encoder Ev\mathcal{E}_{v} with traditional I3D , the conventional backbone in previous works , and observe performance degradation in Table IXd. We further notice that, even with I3D as visual backbone, Locater still outperforms existing algorithms, if we compare the results with Table III, i.e., 58.3%58.3\% vs 56.1%56.1\% (+2.2%2.2\%) in terms of mIoU.

Cross-Modality Encoder Evw\mathcal{E}_{vw}. Then we study the efficacy of our cross-modality encoder Evw ⁣\mathcal{E}_{vw\!} design, which has K ⁣ ⁣= ⁣ ⁣3K_{\!}\!=_{\!}\!3 attention-based modules. As shown in Table IXb, the performance increases when stacking more modules (K ⁣: ⁣1 ⁣→ ⁣3K{\!}:_{\!}1\!\rightarrow\!3), and then the gain becomes marginal (K ⁣: ⁣3 ⁣→ ⁣5K{\!}:_{\!}3\!\rightarrow\!5).

Local-Global Memory M\mathcal{M}. Table ⁣{}_{\!} IXc summarizes the impact of the capacity of the local-global memory M\mathcal{M} in segmentation. As seen, without using either local (Nl ⁣= ⁣0N_{l}\!=\!0) or global memory (Ng ⁣= ⁣0N_{g}\!=\!0), or even both (Nl ⁣= ⁣0N_{l}\!=\!0 and Ng ⁣= ⁣0N_{g}\!=\!0), Locater suffers from significant performance drop. With the increase of the local or global memory capacity, the performance is improved but the gain becomes less stark, confirming the value of the learnable memory operations.

Contextualized Query Embedding. In §3.2.3, we generate a compact query embedding q{\bm{q}} based on visual context (Eq. ⁣{}_{\!} 10). To investigate such a design, we consider an alternative strategy, i.e., applying self-attention to generate q{\bm{q}}. As in ⁣{}_{\!} Table ⁣{}_{\!} IXe, ⁣{}_{\!} visual-guided ⁣{}_{\!} query ⁣{}_{\!} embedding ⁣{}_{\!} is ⁣{}_{\!} more ⁣{}_{\!} favored.

Efficiency of Memory Design. We further demonstrate the efficiency of the finite memory design in Table IXg. To clearly reveal the gap, we directly feed image features to both one Transformer block with 12 heads and our memory module (cf. §3.2.2). Every frame is patch-wise separated into 256256 tokens with 768768 feature channels for each. Compared with using conventional Transformer-style network design, i.e., flattening all tokens as input, the advantage of our method in efficiency is more and more significant with the increase of video length.

Frame Sampling Interval. Moreover, we evaluate the impact of frame sampling interval used during global memory construction (§3.2.2). As shown in Table IXh, sampling more frames with smaller intervals (≤\leq10) cannot provide performance gain, showing the high redundancy among video frames and explaining the high efficiency of our approach. We, therefore, set the default sampling interval as 1010 to construct the global memory Mg\mathcal{M}^{g}.

8 Failure Case Analysis

Despite the stronger cross-modal video understanding ability of Locater against previous methods, it still suffers from difficulties in some challenging scenarios. We present three typical failure modes on A2D-S test in Fig. 6. The first type of mistakes is caused by the ambiguity of natural language. For example, in 1st1\textit{st} column, Locater with an LSTM-based language model, which is trained from scratch, is hard to distinguish the two reference, i.e., white man and man in white, with limited language training data. This issue can be alleviated by introducing a large-scale pretrained language model, e.g., BERT , or harvesting finer language modeling. The second type of challenges comes from weak or ambiguous descriptions. For example, in 2nd2\textit{nd} column, the referent cannot be well differentiated from other objects, as both of the two black dogs are matched well with the query “Black dog running on the grass”. The third type of challenges is brought by highly similar referents. For example, in 3rd3\textit{rd} column, our model faces difficulties in distinguishing the parrot and its reflection, as their appearance and motion information are almost exactly the same.

Conclusion

This work presents a memory augmented, fully attentional model, Locater, for LVS. It effectively aligns cross-modal representations, and efficiently models long-term temporal context as well as short-term segmentation history through an external memory. With visual context guided expression attention, Locater produces frame-specific query vectors for mask generation. We further observe and mitigate the critical absence issue of grounding-required objects in the current most popular A2D-S benchmark  ⁣{}_{\!}with  ⁣{}_{\!}a  ⁣{}_{\!}newly  ⁣{}_{\!}created A2D-S+ dataset.  ⁣{}_{\!}It,  ⁣{}_{\!}as  ⁣{}_{\!}a  ⁣{}_{\!}sufficient  ⁣{}_{\!}complement, significantly increases the number of semantically similar objects in testing examples. Experiments demonstrate that Locater dramatically advances state-of-the-arts with high efficiency on both standard benchmarks and our proposed challenging A2D-S+. Furthermore, our Locater based solution achieved the 1st place in RVOS Track of YTB-VOS21{}_{\text{21}} Challenge, surpassing other competitors by large margins.

References