End-to-End Referring Video Object Segmentation with Multimodal Transformers
Adam Botach, Evgenii Zheltonozhskii, Chaim Baskin
Introduction
Attention-based deep neural networks exhibit impressive performance on various tasks across different fields, from computer vision to natural language processing . These advancements make networks of this sort, such as the Transformer , particularly interesting candidates for solving multimodal problems. By relying on the self-attention mechanism, which allows each token in a sequence to globally aggregate information from every other token, Transformers excel at modeling global dependencies and have become the cornerstone in most NLP tasks . Transformers have also started showing promise in solving computer vision tasks, from recognition to object detection and even outperforming the long-used CNNs as general-purpose vision backbones .
The referring video object segmentation task (RVOS) involves the segmentation of a text-referred object instance in the frames of a given video. Compared with the referring image segmentation task (RIS) , in which objects are mainly referred to by their appearance, in RVOS objects can also be referred to by the actions they are performing or in which they are involved. This renders RVOS significantly harder than RIS, as text expressions that refer to actions often cannot be properly deduced from a single static frame. Furthermore, unlike their image-based counterparts, RVOS methods may be required to establish data association of the referred object across multiple frames (tracking) in order to deal with disturbances such as occlusions or motion blur.
To solve these challenges and effectively align video with text, existing RVOS approaches typically rely on complicated pipelines. In contrast, here we propose a simple, end-to-end Transformer-based approach to RVOS. Using recent advancements in Transformers for textual feature extraction , visual feature extraction and object detection , we develop a framework that significantly outperforms existing approaches. To accomplish this, we employ a single multimodal Transformer and model the task as a sequence prediction problem. Given a video and a text query, our model generates prediction sequences for all objects in the video before determining the one the text refers to. Additionally, our method is free of text-related inductive bias modules and utilizes a simple cross-entropy loss to align the video and the text. As such, it is much less complicated than previous approaches to the task.
The proposed pipeline is schematically depicted in Fig. 1. First, we extract linguistic features from the text query using a standard Transformer-based text encoder, and visual features from the video frames using a spatio-temporal encoder. The features are then passed into a multimodal Transformer, which outputs several sequences of object predictions . Next, to determine which of the predicted sequences best corresponds to the referred object, we compute a text-reference score for each sequence. For this we propose a temporal segment voting scheme that allows our model to focus on more relevant parts of the video when making the decision.
We present a Transformer-based RVOS framework, dubbed Multimodal Tracking Transformer (MTTR), which models the task as a parallel sequence prediction problem and outputs predictions for all objects in the video prior to selecting the one referred to by the text.
Our sequence selection strategy is based on a temporal segment voting scheme, a novel reasoning scheme that allows our model to focus on more relevant parts of the video with regards to the text.
The proposed method is end-to-end trainable, free of text-related inductive bias modules, and requires no additional mask refinement. As such, it greatly simplifies the RVOS pipeline compared to existing approaches.
We thoroughly evaluate our method. On the A2D-Sentences and JHMDB-Sentences , MTTR significantly outperforms all existing methods across all metrics. We also show strong results on the public validation set of Refer-YouTube-VOS , a challenging dataset that has yet to receive attention in the literature.
Related Work
The RVOS task was introduced by Gavrilyuk et al. , whose goal was to attain pixel-level segmentation of actors and their actions in video content. To effectively aggregate and align visual, temporal and lingual information from video and text, state-of-the-art RVOS approaches typically rely on complicated pipelines . Gavrilyuk et al. proposed an I3D-based encoder-decoder architecture that generated dynamic filters from text features and convolved them with visual features to obtain the masks. Following them, Wang et al. added spatial context to the kernels with deformable convolutions . For a more effective representation, VT-Capsule encoded each modality in capsules , while ACGA utilized a co-attention mechanism to enhance the multimodal features. To improve positional relation representations in the text, PRPE explored a positional encoding mechanism based on polar coordinates. URVOS improved tracking capabilities by performing language-based object segmentation on a key frame and then propagating its mask throughout the video. AAMN utilized a top-down approach where an off-the-shelf object detector is used to localize objects in the video prior to parsing relations between visual and textual features. CMPC-V achieved state-of-the-art results by constructing a temporal graph from video and text features, and applying graph convolution to detect the referred entity.
The Transformer was introduced as an attention-based building block for sequence-to-sequence machine translation, and since then has become the cornerstone for most NLP tasks . Unlike previous architectures, the Transformer relies entirely on the attention mechanism to draw dependencies between input and output.
Recently, the introduction of Transformers to computer vision tasks has demonstrated spectacular performance. DETR , which utilizes a non-auto-regressive Transformer, simplifies the traditional object detection pipeline while achieving performance comparable to that of CNN-based detectors . Given a fixed set of learned object queries, DETR reasons about the global context of an image and the relations between its objects and then outputs a final set of detection predictions in parallel. VisTR extends the idea behind DETR to video instance segmentation. It views the task as a direct end-to-end parallel sequence prediction problem. By supervising video instances at the sequence level as a whole, VisTR is able to output an ordered sequence of masks for each instance in a video directly (i.e., natural tracking).
ViT introduced the Transformer to image recognition by using linearly projected patches as tokens for a Transformer encoder. Swin Transformer proposed a general-purpose backbone for computer vision based on a hierarchical Transformer whose representations are computed inside shifted windows. This architecture was also extended to the video domain , which we adapt as our temporal encoder.
Another recent relevant work is MDETR , a DETR-based end-to-end multimodal detector that detects objects in an image conditioned on a text query. Different from our method, their approach is designed to work on static images, and its performance largely depends on well-annotated datasets that contain aligned text and box annotations, the types of which are not available in the RVOS task.
Method
We begin by extracting features from each frame in the sequence using a deep spatio-temporal encoder. Simultaneously, linguistic features are extracted from the text query using a Transformer-based text encoder. Then, the spatio-temporal and linguistic features are linearly projected to a shared dimension .
In the next step, the features of each frame of interest are flattened and separately concatenated with the text embeddings, producing a set of multimodal sequences. These sequences are fed in parallel into a Transformer . In the Transformer’s encoder layers, the textual embeddings and the visual features of each frame exchange information. Then, the decoder layers, which are fed with object queries per input frame, query the multimodal sequences for entity-related information and store it in the object queries. Corresponding queries of different frames share the same trainable weights and are trained to attend to the same instance in the video (each one in its designated frame). We refer to these queries (represented by the same unique color and shape in Figs. 1 and 2) as queries belonging to the same instance sequence. This design allows for natural tracking of each object instance in the video .
For each output instance sequence, we generate a a corresponding mask sequence using an FPN-like spatial decoder and dynamically generated conditional convolution kernels . Finally, we use a novel text-reference score function that, based on text associations, determines which of the object query sequences has the strongest association with the object described in , and returns its segmentation sequence as the model’s prediction.
2 Temporal Encoder
A suitable temporal encoder for the RVOS task should be able to extract both visual characteristics (e.g., shape, size, location) and action semantics for each instance in the video. Several previous works utilized the Kinetics-400 pre-trained I3D network as their temporal encoder. However, since I3D was originally designed for action classification, using its outputs as-is for tasks that require fine details (e.g., instance segmentation) is not ideal as the features it outputs tend to suffer from spatial misalignment caused by temporal downsampling. To compensate for this side effect, past state-of-the-art approaches came up with different solutions, from auxiliary mask refinement algorithms to utilizing additional backbones that operate alongside the temporal encoder . In contrast, our end-to-end approach does not require any additional mask refinement steps and utilizes a single backbone.
Recently, the Video Swin Transformer was proposed as a generalization of the Swin Tranformer to the video domain. While the original Swin was designed with dense predictions (such as segmentation) in mind, Video Swin was tested mainly on action recognition benchmarks. To the best of our knowledge, we are the first to utilize it (with a slight modification) for video segmentation. As opposed to I3D, Video Swin contains just a single temporal downsampling layer and can be easily modified to output per-frame feature maps (we refer to Sec. C.1 for more details). As such, it is a much better choice for processing a full sequence of consecutive video frames for segmentation purposes.
3 Multimodal Transformer
4 The Instance Segmentation Process
Our segmentation process, as shown in Fig. 2, consists of several steps. First, given , the updated multimodal sequences output by the last Transformer encoder layer, we extract and reshape the video-related part of each sequence (i.e., the first tokens) into the set . Then, we take , the outputs of the first blocks of our temporal encoder, and hierarchically fuse them with using an FPN-like spatial decoder . This process results in semantically-rich, high resolution feature maps of the video frames, denoted as .
Finally, a sequence of segmentation masks is generated for by convolving each segmentation kernel with its corresponding frame features, followed by a bilinear upsampling operation to resize the masks into ground-truth resolution,
5 Instance Sequence Matching
During the training process we need to determine which of the predicted instance sequences best fits the referred object. However, if the video sequence contains additional annotated instances, we found that supervising their detection (as negative examples) alongside that of the referred instance helps stabilize the training process.
Let us denote by the set of ground-truth sequences that are available for , and by the set of the predicted instance sequences. We assume that the number of predicted sequences () is chosen to be strictly greater than the number of annotated instances (denoted ) and that the ground-truth sequences set is padded with (no object) to fill any missing slots. Then, we want to find a matching between the two sets . Accordingly, we search for a permutation with the lowest total cost:
where is a pair-wise matching cost. The optimal permutation can be computed efficiently using the Hungarian algorithm . Each ground-truth sequence is of the form
where is a ground-truth mask, and is a one-hot referring vector, i.e., the positive class means that corresponds to the text-referred object and that this object is visible in the corresponding video frame . Note that if is a padding sequence then .
Thus, each prediction of our model is a pair of sequences:
We define the pair-wise matching cost function as the sum
6 Loss Functions
Let us denote (with a slight abuse of notation) by the set of predicted instance sequences permuted according to the optimal permutation . Then, we can define our loss function as follows:
Following VisTR , the first term, dubbed , ensures mask alignment between the predicted and ground-truth sequences. As such, this term is defined as a combination of the Dice and the per-pixel Focal loss functions:
The second loss term, denoted , is a cross-entropy term that supervises the sequence reference predictions:
7 Inference
For a given sample of video and text, let us denote by the set of reference prediction sequences output by our model. Additionally, we denote by the probability of the positive (“referred”) class for a given reference prediction . During inference we return the segmentation mask sequence that corresponds to , the predicted reference sequence with the highest positive score:
This sequence selection scheme, which we term the “temporal segment voting scheme” (TSVS), grades each prediction sequence based on the total association of its terms with the text referred object. Thus, it allows our model to focus on more relevant parts of the video (in which the referred object is visible), and disregard less relevant parts (which may depict irrelevant objects or in which the referred object is occluded) when making the decision. We refer to Sec. D.2 for further analysis of the effect of TSVS.
Experiments
To evaluate our approach, we conduct experiments on three referring video object segmentation datasets. The first two, A2D-Sentences and JHMDB-Sentences , were created by adding textual annotations to the original A2D and JHMDB datasets. Each video in A2D has 3–5 frames annotated with pixel-level segmentation masks, while in JHMDB, 2D articulated human puppet masks are available for all frames. We refer to Sec. B.1 for more details. We adopt Overall IoU, Mean IoU, and precision@K to evaluate our method on these datasets. Overall IoU computes the ratio between the total intersection and the total union area over all the test samples. Mean IoU is the averaged IoU over all the test samples. Precision@K considers the percentage of test samples whose IoU scores are above a threshold K, where . We also compute mean average precision (mAP) over 0.50:0.05:0.95 .
We want to note that we found inconsistencies in the mAP metric calculation in previous studies. For example, examination of published code revealed incorrect calculation of the metric as the average of the precision@K metric over several K values. To avoid further confusion and ensure a fair comparison, we suggest adopting the COCO APIhttps://github.com/cocodataset/cocoapi for mAP calculation. For reference, a full implementation of the evaluation that utilizes the API is released with our code.
We further evaluate MTTR on the more challenging Refer-YouTube-VOS dataset, introduced by Seo et al. , who provided textual annotations for the original YouTube-VOS dataset . Each video has pixel-level instance segmentation annotations for every fifth frame. The original release of Refer-YouTube-VOS contains two subsets. One subset contains first-frame expressions that describe only the first frame. The other contains full-video expressions that are based on the whole video and are, therefore, more challenging. Following the introduction of the RVOS competitionhttps://youtube-vos.org/dataset/rvos/, only the more challenging subset of the dataset is publicly available now. Since ground-truth annotations are available only for the training samples and the test server is currently inaccessible, we report results on the validation samples by uploading our predictions to the competition’s serverhttps://competitions.codalab.org/competitions/29139. We refer to Sec. B.2 for more details. The primary evaluation metrics for this dataset are the average of the region similarity () and the contour accuracy () .
As our temporal encoder we use the smallest (“tiny”) Video Swin Transformer pretrained on Kinetics-400 . The original Video Swin consists of four blocks with decreasing spatial resolution. We found the output of the fourth block to be too small for small object detection and hence we only utilize the first three blocks. We use the output of the third block as the input of the multimodal Transformer, while the outputs of the earlier blocks are fed into the spatial decoder. We also modify the encoder’s single temporal downsampling layer to output per-frame feature maps as required by our model. As our text encoder we use the Hugging Face implementation of RoBERTa-base . For A2D-Sentences we feed the model windows of frames with the annotated target frame in the middle. Each frame is resized such that the shorter side is at least 320 pixels and the longer side is at most 576 pixels. For Refer-YouTube-VOS , we use windows of consecutive annotated frames during training, and full-length videos (up to 36 annotated frames) during evaluation. Each frame is resized such that the shorter side is at least 360 pixels and the longer side is at most 640 pixels. We do not use any segmentation-related pretraining, e.g., on COCO , which is known to boost segmentation performance . We refer the reader to Appendix C for more implementation details.
2 Comparison with State-of-the-Art Methods
We compare our method with existing approaches on the A2D-Sentences dataset. For fair comparison with existing works , our model is trained and evaluated for this purpose with windows of size 8. As shown in Tab. 1, our method significantly outperforms existing approaches across all metrics. For example, our model shows a 4.3 mAP gain over current state of the art, and an absolute improvement of 6.6% on the most stringent metric P@0.9, which demonstrates its ability to generate high-quality masks. We also note that our top configuration () achieves a massive 5.7 mAP gain and 6.7% absolute improvement on both Mean and Overall IoU compared to the current state of the art. Impressively, this configuration is able to do so while processing 76 frames per second on a single RTX 3090 GPU.
Following previous works , we evaluate the generalization ability of our model by evaluating it on JHMDB-Sentences without fine-tuning. We uniformly sample three frames from each video and evaluate our best model on these frames. As shown in Tab. 2, our method generalizes well and outperforms all existing approaches. Note that all methods (including ours) produce low results on P@0.9. This can be attributed to JHMDB’s imprecise mask annotations which were generated by a coarse human puppet model.
Finally, we report our results on the public validation set of Refer-YouTube-VOS in Tab. 3. As mentioned earlier, this subset contains only the more challenging full-video expressions from the original release of Refer-YouTube-VOS. Compared with existing methods which trained and evaluated on the full version of the dataset, our model demonstrates superior performance across all metrics despite being trained on less data and evaluated exclusively on a more challenging subset. Additionally, our method shows competitive performance compared with the methods that led in the 2021 RVOS competition . We note, however, that these methods use ensembles and are trained on additional segmentation and referring datasets .
3 Ablation Studies
We conduct ablation studies on A2D-Sentences to evaluate our model’s design and robustness. Unless stated otherwise, we use window size . An ablation study on the number of object queries can be found in Sec. D.1.
To evaluate MTTR’s performance independently of the temporal encoder, we compare it with CMPC-I, the image-targeted version of CMPC-V . Following CMPC-I, we use DeepLab-ResNet101 pretrained on PASCAL-VOC as a visual feature extractor. We train our model using only the target frames (i.e., without additional frames for temporal context). As shown in Fig. 4(a), our method significantly surpasses CMPC-I across all metrics, with a 6.1 gain in mAP and 8.7% absolute improvement in Mean IoU. In fact, this configuration of our model surpasses all existing methods regardless of the temporal context.
In Fig. 4(b) we study the effect of the temporal context size on MTTR’s performance. A larger temporal context enables better extraction of action-related information. For this purpose, we train and evaluate our model using different window sizes. As expected, widening the temporal context leads to large performance gains, with an mAP gain of 4.3 and an absolute Mean IoU improvement of 3.7% when gradually changing the window size from 1 to 10. Intriguingly, however, peak performance on A2D-Sentences is obtained using , as widening the window even further (e.g., ) results in a performance drop.
To study the effect of the selected word embeddings on our model’s performance, we train our model using two additional widely-used Transformer-based text encoders, namely BERT-base and Distill-RoBERTa-base , a distilled version of RoBERTa . Additionally, we experiment with GloVe and fastText , two simpler word embedding methods. As shown in Fig. 4(c), our model achieves comparable performance when relying on the different Transformer-based encoders, which demonstrates its robustness to this change. Unsurprisingly, however, performance is slightly worse when relying on the simpler methods. This may be explained by the fact that while Transformer-based encoders are able to dynamically encode sentence context within their output embeddings, simpler methods disregard this context and merely rely on fixed pretrained embeddings.
To study the effect of supervising the detections of un-referred instances alongside that of the referred instance in each sample, we train different configurations of our model without supervision of un-referred instances. Intriguingly, in all such experiments our model immediately converges to a local minimum of the text loss (), where the same object query is repeatedly matched with all ground-truth instances, thus leaving the rest of the object queries untrained. In some experiments our model manages to escape this local minimum after a few epochs and then achieves comparable performance with our original configuration. Nevertheless, in other experiments this phenomenon significantly hinders its final mAP score.
4 Qualitative Analysis
As illustrated in Fig. 3, MTTR can successfully track and segment the referred objects even in challenging situations where they are surrounded by similar instances, occluded, or completely outside of the frame in large parts of the video.
Conclusion
We introduced MTTR, a simple Transformer-based approach to RVOS that models the task as a sequence prediction problem. Our end-to-end method considerably simplifies existing RVOS pipelines by simultaneously processing both text and video frames in a single multimodal Transformer. Extensive evaluation of our approach on standard benchmarks reveals that our method outperforms existing state-of-the-art methods by a large margin (e.g., a 5.7 mAP improvement on A2D-Sentences). We hope our work will inspire others to see the potential of Transformers for solving complex multimodal tasks.
References
Appendix A Additional Method Details
In this section we present the full definitions of the Dice and Focal cost and loss functions used in Sec. 3.5 and Sec. 3.6.
Given two segmentation masks , the Dice coefficient between the two masks is defined as follows:
where for a mask , and . Note that in practice we also add a smoothing constant to both the numerator and denominator of the above expression to avoid possible division by .
Given the above, the Dice cost between the mask sequences and is defined as
Similarly, the Dice loss between the two sequences is defined as
A.1.2 Focal Loss
The Focal loss between two corresponding segmentation masks and for time step is defined as
where is the probability predicted for the ground-truth class of the pixel:
and is a class balancing factor defined as
Following we use . We refer to for more information about these hyperparameters.
Given the above, the Focal loss between the ground-truth and predicted mask sequences and is defined as
Appendix B Additional Dataset Details
A2D-Sentences contains 3,754 videos (3,017 train, 737 test) with 7 actors classes performing 8 action classes. Additionally, the dataset contains 6,655 sentences describing the actors in the videos and their actions. JHMDB-Sentences contains 928 videos along with 928 corresponding sentences describing 21 different action classes.
B.2 Refer-YouTube-VOS
The original release of Refer-YouTube-VOS contains 27,899 text expressions for 7,451 objects in 3,975 videos. The objects belong to 94 common categories. The subset with the first-frame expressions contains 10,897 expressions for 3,412 videos in the train split and 1,993 expressions for 507 videos in the validation split. The subset with the full-video expressions contains 12,913 expressions for 3,471 videos in the train split and 2,096 expressions for 507 videos in the validation split. Following the introduction of the RVOS competitionhttps://youtube-vos.org/dataset/rvos/, only the more challenging full-video expressions subset is publicly available now, so we use this subset exclusively in our experiments. Additionally, this subset’s original validation set was split into two separate competition validation and test sets of 202 and 305 videos respectively. Since ground-truth annotations are available only for the training set and the test server is currently closed, we report results exclusively on the competition validation set by uploading our predictions to the competition’s serverhttps://competitions.codalab.org/competitions/29139.
Appendix C Additional Implementation Details
The original architecture of Video Swin Transformer contains a single temporal down-sampling layer, realized as a 3D convolution with kernel and stride of size (the first dimension is temporal). However, since our multimodal Transformer expects per-frame embeddings, we removed this temporal down-sampling step by modifying the kernel and stride of the above convolution to size . In order to achieve this while maintaining support for the Kinetics-400 pretrained weights of the original Swin configuration, we summed the pretrained kernel weights of the aforementioned convolution on its temporal dim, resulting in a new kernel. This solution is equivalent to (but more efficient than) duplicating each frame in the input sequence before inserting it into the temporal encoder.
C.2 Multimodal Transformer
We employ the same Transformer architecture proposed by Carion et al. . The decoder layers are fed with a set of object queries per input frame. For efficiency reasons we only utilize 3 layers in both the encoder and decoder, but note that more layers may lead to additional performance gains, as demonstrated by Carion et al. . Also, similarly to Carion et al. , fixed sine spatial positional encodings are added to the features of each frame before inserting them into the Transformer. No positional encodings are used for the text embeddings, as in our experiments using sine embeddings have led to reduced performance and learnable encodings had no effect compared to using no encodings at all.
C.3 Instance Segmentation
The spatial decoder is an FPN-like module consisting of several 2D convolution, GroupNorm and ReLU layers. Nearest neighbor interpolation is used for the upsampling steps. The segmentation kernels and the feature maps of are of dimension following .
C.4 Additional Training Details
We use as the feature dimension of the multimodal Transformer’s inputs and outputs. The hyperparameters for the loss and matching cost functions are .
Following Carion et al. we utilize AdamW as the optimizer with weight decay set to during training. We also apply gradient clipping with a maximal gradient norm of . A learning rate of is used for the Transformer and for the temporal encoder. The text encoder is kept frozen.
Similarly to Carion et al. we found that utilizing auxiliary decoding losses on the outputs of all layers in the Transformer decoder expedites training and improves the overall performance of the model.
During training, to enhance model’s position awareness, we randomly flip the input frames horizontally and swap direction-related words in the corresponding text expressions accordingly (e.g., the word ’left’ is replaced with ’right’).
We train the model for 70 epochs on A2D-Sentences . The learning rate is decreased by a factor of 2.5 after the first 50 epochs. In the default configuration we use window size and batch size of 6 on 3 RTX 3090 24GB GPUs. Training takes about 31 hours in this configuration. On Refer-YouTube-VOS the model is trained for 30 epochs, and the learning rate is decreased by a factor of 2.5 after the first 20 epochs. In the default configuration we use window size and batch size of 4 on 4 A6000 48GB GPUs. Training takes about 45 hours in this configuration.
Appendix D Additional Experiments
To study the effect of the number of object queries on MTTR’s performance, we train and evaluate our model on A2D-Sentences using window size and different values of . As shown in Tab. D.1, the best performance is achieved for . Our hypothesis is that when using lower values of the resulting set of object queries may not be diverse enough to cover a large set of possible object detections. On the other hand, using higher values of may require a longer training schedule to obtain good results, as the probability of each query being matched with a ground-truth instance (and thus updated) at each training iteration is lower.
D.2 Analysis of the Effect of TSVS
To further illustrate and analyze the effect of the temporal segment voting scheme (TSVS), we refer to the zebras example in the third row of Fig. 3 and the orange text query. Without TSVS, a prediction would have to be made for each frame in the video separately. Hence, as the correct zebra (marked in orange) is not yet visible in the first two frames, one of the other visible zebras in each of these frames may be wrongly selected. With TSVS, however, the predictions of the correct zebra in the final three frames vote together as part of a sequence, and due to the high reference scores of these predictions, this sequence is then selected over all other instance sequences. This results in only the correct zebra being segmented throughout the video (i.e., no zebra is segmented in the first two frames), as expected.