A Survey on Deep Learning Technique for Video Segmentation

Tianfei Zhou, Fatih Porikli, David Crandall, Luc Van Gool, Wenguan Wang

I Introduction

Video segmentation — identifying the key objects with some specific properties or semantics in a video scene — is a fundamental and challenging problem in computer vision, with numerous potential applications including autonomous driving, robotics, automated surveillance, social media, augmented reality, movie production, and video conferencing.

The problem has been addressed using various traditional computer vision and machine learning techniques, including hand-crafted features (e.g., histogram statistics, optical flow, etc.), heuristic prior knowledge (e.g., visual attention mechanism ⁣{}_{\!} , motion boundaries ⁣{}_{\!} , etc.), low/mid-level visual representations (e.g., super-voxel ⁣{}_{\!} , trajectory ⁣{}_{\!} , object proposal ⁣{}_{\!} , etc.), and classical machine learning models (e.g., clustering ⁣{}_{\!} , graph models ⁣{}_{\!} , random walks ⁣{}_{\!} , support vector machines ⁣{}_{\!} , random decision forests ⁣{}_{\!} , markov random fields ⁣{}_{\!} , conditional random fields ⁣{}_{\!} , etc.). Recently, deep neural networks, and Fully Convolutional Networks (FCNs) ⁣{}_{\!} in particular, have led to remarkable advances in video segmentation. These deep learning-based video segmentation algorithms are significantly more accurate (and sometimes even more efficient) than traditional approaches.

With the rapid advance of this field, there is a huge body of new literature being produced. However, most existing surveys predate the modern deep learning era ⁣{}_{\!} , and often take a narrow view, such as focusing only on video foreground/background segmentation ⁣{}_{\!} . In this paper, we offer a state-of-the-art review that addresses the wide area ⁣{}_{\!} of ⁣{}_{\!} video ⁣{}_{\!} segmentation, ⁣{}_{\!} especially ⁣{}_{\!} to ⁣{}_{\!} help ⁣{}_{\!} new ⁣{}_{\!} researchers ⁣{}_{\!} enter this rapidly-developing field. We systematically introduce recent advances in video segmentation, spanning from task formulation to taxonomy, from algorithms to datasets, and from unsolved issues to future research directions. We cover crucial aspects including task categories (i.e., foreground/background separation vs semantic segmentation), inference modes (i.e., automatic, semi-automatic, and interactive), and learning paradigms (i.e., supervised, unsupervised, and weakly supervised), and we try to clarify terminology (e.g., background subtraction, motion segmentation, etc.). We hope that this survey helps accelerate progress in this field.

This survey mainly focuses on recent progress in two major branches of video segmentation, namely video object segmentation (Fig. 1(a-e)) and video semantic segmentation (Fig. ⁣{}_{\!} 1(f-h)), ⁣{}_{\!} which ⁣{}_{\!} are ⁣{}_{\!} further ⁣{}_{\!} divided ⁣{}_{\!} into ⁣{}_{\!} eight ⁣{}_{\!} sub-fields. Even after restricting our focus to deep learning-based video segmentation, there are still hundreds of papers in this fast-growing field. We select influential work published in prestigious journals and conferences. We also include some non-deep learning video segmentation models and relevant literature in other areas, e.g., visual tracking, to give necessary background. Moreover, in order to promote the development of this field, we provide an accompanying webpage which catalogs algorithms and datasets addressing video segmentation: https://github.com/tfzhou/VS-Survey.

Fig. 2 shows the structure of this survey. Section §II gives some brief background on taxonomy, terminology, study history, and related research areas. We review representative papers on deep learning algorithms and video segmentation datasets in §III and §IV, respectively. Section §V conducts performance evaluation and analysis, while §VI raises open questions and directions. Finally, we make concluding remarks in §VII.

II Background

In this section, we first formalize the task, categorize research directions, and discuss key challenges and driving factors in §II-A. Then, §II-B offers a brief historical background covering early work and foundations, and §II-C establishes linkages with relevant fields.

Formally, let X\mathcal{X} and Y\mathcal{Y} denote the input space and output segmentation space, respectively. Deep learning-based video segmentation solutions generally seek to learn an ideal video-to-segment mapping f∗ ⁣:X↦Yf^{*\!}:\bm{\mathcal{X}}\mapsto\bm{\mathcal{Y}}.

According to how the output space Y\mathcal{Y} is defined, video segmentation can be broadly categorized into two classes: video object (foreground/background) segmentation, and video semantic segmentation.

∙\bullet Video Foreground/Background Segmentation (Video Object Segmentation, VOS). VOS is the classic video segmentation setting and refers to segmenting dominant objects (of unknown categories). In this case, Y\bm{\mathcal{Y}} is a binary, foreground/background segmentation space. VOS is typically used in video analysis and editing application scenarios, such as object removal in movie editing, content-based video coding, and virtual background creation in video conferencing. It typically is not concerned with the exact semantic categories of the segmented objects.

∙\bullet Video Semantic Segmentation (VSS). As a direct extension of image semantic segmentation to the spatio-temporal domain, VSS aims to extract objects within predefined semantic categories (e.g., car, building, pedestrian, road) from videos. Thus, Y\bm{\mathcal{Y}} corresponds to a multi-class, semantic parsing space. VSS serves as a perception foundation for many application fields, such as robot sensing, human-machine interaction, and autonomous driving, which require high-level understanding of the physical environment.

Remark. VOS and VSS share some common challenges, such as fast motion and object occlusion. However, due to differences in application scenarios, many challenges are different. For instance, VOS often focuses on human created media, which often have large camera motion, deformation, and appearance changes. VSS instead often focuses on applications like autonomous driving, which requires a good trade off between accuracy and latency, accurate detection of small objects, model parallelization, and cross-domain generalization ability.

II-A2 Inference Modes for Video Segmentation

VOS methods can be further classified into three types: automatic, semi-automatic, and interactive, according to how much human intervention is involved during inference.

∙\bullet Automatic Video Object Segmentation (AVOS). AVOS, or unsupervised video segmentation or zero-shot video segmentation, performs VOS in an automatic manner, without any manual initialization (Fig. 1(a-b)). The input space X\bm{\mathcal{X}} refers to the video domain V\bm{\mathcal{V}} only. AVOS is suitable for video analysis but not for video editing that requires segmenting arbitrary objects or their parts flexibly; a typical application is virtual background creation in video conferencing.

∙\bullet Semi-automatic Video Object Segmentation (SVOS). SVOS, also known as semi-supervised video segmentation or one-shot video segmentation , involves limited human inspection (typically provided in the first frame) to specify the desired objects (Fig. 1(c)). For SVOS, X ⁣ ⁣= ⁣ ⁣V ⁣× ⁣M\bm{\mathcal{X}}_{\!}\!=_{\!}\!\bm{\mathcal{V}}\!\times\!\bm{\mathcal{M}}, where V ⁣\bm{\mathcal{V}}_{\!} indicates the video space and M\bm{\mathcal{M}} refers to human input. Typically the human input is an object mask in the first video frame, in which case SVOS is also called pixel-wise tracking or mask propagation. Other forms of human input include bounding boxes and scribbles . From this perspective, language-guided video object segmentation (LVOS) is a sub-branch of SVOS, in which the human input is given as linguistic descriptions about the desired objects (Fig. 1(e)). Compared to AVOS, SVOS is more flexible in defining target objects, but requires human input. SVOS is typically applied in a user-friendly setting (without specialized equipment), such as video content creation in mobile phones. One of the core challenges in SVOS is how to fully utilize target information from limited human intervention.

∙\bullet Interactive Video Object Segmentation (IVOS). SVOS models are designed to operate automatically once the target has been identified, while systems for IVOS incorporate user guidance throughout the analysis process (Fig. 1(d)). IVOS can obtain high-quality segments and works well for computer-generated imagery and video post-production, where tedious human supervision is possible. IVOS is also studied in the graphics community as video cutout. The input space X\bm{\mathcal{X}} for IVOS is V ⁣× ⁣S\bm{\mathcal{V}}\!\times\!\bm{\mathcal{S}}, where S\bm{\mathcal{S}} typically refers to human scribbling. Key challenges include: 1) allowing users to easily specify segmentation constraints; 2) incorporating human specified constraints into the segmentation algorithm; and 3) giving quick response to the constraints.

In contrast to VOS, VSS methods typically work in an automatic mode (Fig. 1(f-h)), i.e., X ⁣≡ ⁣V\bm{\mathcal{X}}\!\equiv\!\bm{\mathcal{V}}. Only a few early methods address the semi-automatic setting, called label propagation .

Remark. The terms “unsupervised” and “semi-supervised” are conventionally used in VOS to specify the amount of human interaction involved during inference. But they are easily confused with “unsupervised ⁣{}_{\!} learning” ⁣{}_{\!} and ⁣{}_{\!} “semi-supervised ⁣{}_{\!} learning.” ⁣{}_{\!} We ⁣{}_{\!} urge ⁣{}_{\!} the ⁣{}_{\!} community ⁣{}_{\!} to ⁣{}_{\!} replace ⁣{}_{\!} these ⁣{}_{\!} ambiguous ⁣{}_{\!} terms with “automatic” and “semi-automatic.”

II-A3 Learning Paradigms for Video Segmentation

Deep learning-based video segmentation models can be grouped into three categories according to the learning strategy they use to approximate f∗f^{*}: supervised, unsupervised, and weakly supervised.

∙\bullet Supervised Learning Methods. Modern video segmentation models are typically learned in a fully supervised manner, requiring NN input training samples and their desired outputs yn ⁣ ⁣:= ⁣f∗ ⁣(xn)y_{n}\!\!:=\!f^{*\!}(x_{n}), where {(xn,yn)}n ⁣ ⁣⊂<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mimathvariant="script">X</mi></mrow><annotationencoding="application/x−tex">X</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6833em;"></span><spanclass="mordmathcal"style="margin−right:0.1464em;">X</span></span></span></span></span>×\{(x_{n},y_{n})\}_{n\!}\!\subset<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mi mathvariant="script">X</mi></mrow><annotation encoding="application/x-tex">\mathcal{X}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathcal" style="margin-right:0.1464em;">X</span></span></span></span></span>\timesY\mathcal{Y}. The standard method for evaluating learning outcomes follows an empirical risk/loss minimization formulation: We omit the regularization term for brevity.

∙\bullet Unsupervised (Self-supervised) Learning Methods. When only data samples {xn}n ⁣ ⁣ ⁣⊂\{x_{n}\}_{n\!\!}\!\subsetX\mathcal{X} are given, the problem of approximating f∗ ⁣f^{*\!} is known as unsupervised learning. Unsupervised learning includes fully unsupervised learning methods in which the methods do not need any labels at all, as well as self-supervised learning methods in which networks are explicitly trained with automatically-generated pseudo labels without any human annotations . Almost all existing unsupervised learning-based video segmentation models are self-supervised learning methods, where the prior knowledge Z\mathcal{Z} refers to pseudo labels derived from intrinsic properties of video data (e.g., cross-frame consistency). We thus use “unsupervised learning” and “self-supervised learning” interchangeably.

∙\bullet Weakly-Supervised Learning Methods. In this case, Z\mathcal{Z} is typically a more easily-annotated domain, such as tags, bounding boxes, or scribbles, and f∗f^{*} is approximated using a finite number of samples from X<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>Z\mathcal{X}<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>\mathcal{Z}.

Remark. So far, deep supervised learning-based methods are dominant in the field of video segmentation. However, exploring the task in an unsupervised or weakly supervised setting is more appealing, not only because it alleviates the annotation burden of Y\mathcal{Y}, but because it inspires an in-depth ⁣{}_{\!} understanding ⁣{}_{\!} of ⁣{}_{\!} the ⁣{}_{\!} nature ⁣{}_{\!} of ⁣{}_{\!} the ⁣{}_{\!} task ⁣{}_{\!} by ⁣{}_{\!} exploring ⁣{}_{\!} Z\mathcal{Z}.

II-B History and Terminology

Digital image segmentation has been studied for at least 50 years, starting with the Roberts operator for identifying object boundaries. Since then, numerous algorithms for image segmentation have been proposed, and many are extended to the video domain. The field of video segmentation has evolved quickly and undergone great change.

Earlier attempts focus on video over-segmentation, i.e., partitioning a video into space-time homogeneous, perceptually distinct-regions. Typical approaches include hierarchical video segmentation , temporal superpixel , and super-voxels , based on the discontinuity and similarity of pixel intensities in a particular location, i.e., separating pixels according to abrupt changes in intensity or grouping pixels with similar intensity together. These methods are instructive for early stage video preprocessing, but cannot solve the problem of object-level pattern modeling, as they do not provide any principled approach to flatten the hierarchical video decomposition into a binary segmentation .

To extract foreground objects from video sequences, background subtraction techniques emerged beginning in the late 70s , and became popular following the work of . They assume that the background is known a priori, and that the camera is stationary​ or undergoes a predictable, parametric 2D​ or 3D motion with 3D parallax​ . These geometry-based methods fit well for specific application scenarios such as surveillance systems , but they are sensitive to model selection (2D or 3D), and cannot handle non-rigid camera movements.

Another group of video segmentation solutions tackled the task of motion segmentation, i.e., finding objects in motion. Background subtraction can also be viewed as a specific case of motion segmentation. However, most motion segmentation models are built upon motion analysis , factorization​ , and/or statistical​ techniques that comprehensively model the characteristics of moving scenes without prior knowledge of camera motion. Among the big family of motion segmentation algorithms, trajectory segmentation attained particular attention . Trajectories are generated through tracking points over multiple frames and can represent long-term motion patterns, serving as an informative cue for segmentation. Though impressive, motion-based methods heavily rely on the accuracy of optical flow estimation and can fail when different parts of an object exhibit heterogeneous motions.

To overcome these limitations, the task of extracting generic objects from unconstrained video sequences, i.e., AVOS, has drawn increasing research interest . Several methods explored object hypotheses or proposals as middle-level object representations. They generate a large number of object candidates in every frame and cast the task of segmenting video objects as an object region selection problem. The main drawbacks of the proposal-based algorithms are the high computational cost and complicated object inference schemes. Some others explored heuristic hypotheses such as visual attention and motion boundary , but easily fail in scenarios where the heuristic assumptions do not hold.

As argued earlier, an alternative to the above unattended solutions is to incorporate human-marked initialization, i.e., SVOS. Older SVOS methods often rely on optical flow and share a similar spirit with object tracking . In addition, some pioneering IVOS methods were proposed to address high-quality video segmentation under extensive human guidance, including rotoscoping , scribble , contour , and points . Significant engineering is typically needed to allow IVOS systems to operate at interactive speeds. In short, SVOS and IVOS pay for the improved flexibility and accuracy: they are infeasible at large scale due to their human-in-the-loop nature.

In the pre-deep learning era, relatively few papers ⁣{}_{\!} considered VSS due to the complexity of the task. The approaches typically relied on supervised classifiers such as SVMs and video over-segmentation techniques.

Overall, traditional approaches for video segmentation, though giving interesting results, are constrained by hand-crafted features and heavy engineering. But deep learning brought the performance of video segmentation to a new level, as we will review in §III.

II-C Related Research Areas

There are several research fields closely related to video segmentation, which we now briefly describe.

∙\bullet Visual Tracking. To infer the location of a target object over time, current tracking methods usually assume that the target is determined by a bounding box in the first frame . However, in more general tracking scenarios, and in particular the cases studied in early tracking methods, diverse object representations are explored , including centroids, skeletons, and contours. Some video segmentation techniques, such as background subtraction, are also merged into older trackers . Hence, visual tracking and video segmentation encounter some common challenges (e.g., object/camera motion, appearance change, occlusion, etc.), fostering their mutual collaboration.

∙\bullet Image Semantic Segmentation. The success of end-to-end image semantic segmentation has sparked the rapid development of VSS. Rather than directly applying image semantic segmentation techniques frame by frame, recent VSS systems explore temporal continuity to increase both accuracy and efficiency. Nevertheless, image semantic segmentation techniques continue to serve as a foundation for advancing segmentation in video.

∙\bullet Video Object Detection. To generalize object detection in the video domain , video object detectors incorporate temporal cues over the box- or feature- level. There are many key technical steps and challenges, such as object proposal generation, temporal information aggregation, and cross-frame object association, that are shared between video object detection and (instance-level) video segmentation.

III Deep Learning-based Video Segmentation

VOS extracts generic foreground objects from video sequences with no concern for semantic category recognition. Based on how much human intervention is involved in inference, VOS models can be divided into three classes (§II-A2): automatic (AVOS, §III-A1), semi-automatic (SVOS, §III-A2), and interactive (IVOS, §III-A3). Moreover, although language-guided video object segmentation (LVOS) falls in the broader category of SVOS, LVOS methods are reviewed alone (§III-A4), due to the specific multi-modal task setup.

Instead of using heuristic priors and hand-crafted features to automatically execute VOS, modern AVOS methods learn generic video object patterns in a data-driven fashion. We group landmark efforts based on their key techniques.

∙\bullet Deep Learning Module based Methods. In 2015, Fragkiadaki et al. made an early effort that learns a multi-layer perceptron to rank proposal segments and infer foreground objects. In 2016, Tsai et al. proposed a joint optimization framework for AVOS and optical flow estimation with a naïve use of deep features from a pre-trained classification network. Later methods learn FCNs to predict initial, pixel-level foreground estimates from frame images or optical flow fields , while several post-processing steps are still needed. Basically, these primitive solutions largely rely on traditional AVOS techniques; the learning ability of neural networks is under-explored.

∙\bullet Pixel Instance Embedding based Methods. A group of AVOS models has been developed to make use of stronger deep learning descriptors – instance embeddings – learned from image instance segmentation data . They first generate pixel-wise instance embeddings, and select representative embeddings which are clustered into foreground and background. Finally, the labels of the sampled embeddings are propagated to the other ones. The clustering and propagation can be achieved without video specific supervision. Though using fewer annotations, these methods suffer from a fragmented and complicated pipeline.

∙\bullet End-to-end Methods with Short-term Information Encoding. End-to-end model designs became the mainstream in this field. For example, convolutional recurrent neural networks (RNNs) were used to learn spatial and temporal visual patterns jointly . Another big family is built upon two-stream networks , wherein two parallel streams are built to extract features from raw image and optical flow, which are further fused for segmentation prediction. Two-stream methods make explicit use of appearance and motion cues, at the cost of optical flow computation and vast learnable parameters. These end-to-end methods improve accuracy and show the advantages of applying neural networks to this task. However, they only consider local content within very limited time span; they stack appearance and/or motion information from a few successive frames as input, ignoring relations among distant frames. Although RNNs are usually adopted, their internal hidden memory creates the inherent limits in modeling longer-term dependencies .

∙\bullet End-to-end Methods with Long-term Context Encoding. Current leading AVOS models use global context over long time spans. In a seminal work , Lu et al. proposed a Siamese architecture-based model that extracts features for arbitrary frame pairs and captures cross-frame context by calculating pixel-wise feature correlations. During inference, for each test frame, context from several other frames (within the same video) is aggregated to locate objects. A contemporary work exploited a similar idea but only used the first frame as reference. Several papers extended by making better use of information from multiple frames , encoding spatial context , and incorporating temporal consistency to improve representation power and computation efficiency .

∙\bullet  ⁣{}_{\!}Un-/Weakly-supervised ⁣{}_{\!} based ⁣{}_{\!} Methods. Only a handful of methods learn to perform AVOS from unlabeled or weakly labeled data. In , static image salient object segmentation and dynamic eye fixation data, which are more easily acquired compared with video segmentation data, are used to learn video generic object patterns. In , visual patterns are learned through exploring several intrinsic properties of video data at multiple granularities, i.e., intra-frame saliency, short-term visual coherence, long-range semantic correspondence, and video-level discriminativeness. In , an adversarial contextual model is developed to segment moving objects without any manual annotation, achieved by minimizing the mutual information between the motions of an object and its context. This method is further enhanced in by adopting a bootstrapping strategy and enforcing temporal consistency. In , motion is exclusively exploited to discover moving objects, and a Transformer-based model is designed and trained by self-supervised flow reconstruction using unlabeled video data.

∙\bullet Instance-level AVOS Methods. Instance-level AVOS, also referred as multi-object unsupervised video segmentation, was introduced with the launch of the DAVIS19 challenge . This task setting is more challenging as it requires not only separating the foreground objects from the background, but also discriminating different object instances. To tackle this task, current solutions typically work in a top-down fashion, i.e., generating object candidates for each frames, and associating instances over different frames. In an early attempt , Ventura et al. delivered a recurrent network-based model that consists of a spatial LSTM for per-frame instance discovery and a temporal LSTM for cross-frame instance association. This method features an elegant model design, while its representation ability is too weak to enumerate all the object instances and to capture complex interactions between instances over the temporal domain. Thus later methods strengthen the two-step pipeline through: i) employing image instance segmentation models (e.g., Mask R-CNN ) to detect object candidates, and ii) leveraging tracking/re-identification techniques and manually designed rules for instance association. Foreground/background AVOS techniques are also used to filter out nonsalient candidates . More recent methods, e.g., , generate object candidates first and obtain corresponding tracklets via advanced SVOS techniques. Overall, current instance-level AVOS models follow the classic tracking-by-detection paradigm, involving several ad-hoc designs. There is still considerable room for further improvement in accuracy and efficiency.

III-A2 Semi-automatic Video Object Segmentation (SVOS)

Deep learning-based SVOS methods mainly focus on the first-frame mask propagation setting. They are categorized by their utilization of the test-time provided object masks.

∙\bullet Online Fine-tuning based Methods. Following the one-

shot principle, this family of methods trains a segmentation model separately on each given object mask in an online fashion. Fine-tuning methods essentially exploit the transfer learning capabilities of neural networks and often follow a two-step training procedure: i) offline pre-training: learn general segmentation features from images and video sequences, and ii) online fine-tuning: learn target-specific representations from test-time supervision. The idea of fine-tuning was first introduced in , where only the initial image-mask pair is used for training an online, one-shot, but merely appearance-based FCN model. Then, in , more pixel samples in the unlabeled frames are mined as online training samples to better adapt to further changes over time. As have no notion of individual objects, further incorporates instance segmentation models (e.g., Mask R-CNN ) during inference. While elegant through their simplicity, fine-tuning methods have several weaknesses: i) pre-training is fixed and not optimized for subsequent fine-tuning, ii) hyperparameters of online fine-tuning are often excessively hand-crafted and fail to generalize between test cases, iii) the common existing fine-tuning setups suffer from high test runtimes (up to 1,000 training iterations per segmented object online ). The root cause is that these approaches choose to encode all the target-related cues (i.e., appearance, mask) into network parameters implicitly. Towards efficient and automated fine-tuning, some recent methods turn to meta learning techniques, i.e., optimize the fine-tuning policies (e.g., generic model initialization, learning rates, etc.) or even directly modify network weights .

∙\bullet Propagation-based Methods. Two recent lines of research – built upon mask propagation and template matching techniques respectively – try to refrain from the online optimization to deliver compact, end-to-end SVOS solutions. In particular, propagation-based methods use the previous frame mask to infer the current mask ⁣{}_{\!} . For example, Jampani et al. propose a bilateral network for long-range video-adaptive mask propagation. Perazzi et al. approach SVOS by employing a modified FCN, where the previous frame mask is considered as an extra input channel. Follow-up work adopts optical flow guided mask alignment , heavy first-frame data augmentation ⁣{}_{\!} , and multi-step segmentation refinement ⁣{}_{\!} . Others apply re-identification to retrieve missing objects after prolonged occlusions ⁣{}_{\!} , design a reinforcement learning agent that tackles ⁣{}_{\!} SVOS as ⁣{}_{\!} a ⁣{}_{\!} conditional ⁣{}_{\!} decision-making ⁣{}_{\!} process ⁣{}_{\!} , or propagate masks in a spatiotemporal MRF model to improve temporal coherency . Some researchers leverage location-aware embeddings to sharpen the feature ⁣{}_{\!} , or directly learn sequence-to-sequence mask propagation . Advanced tracking techniques are also exploited in . Propagation-based methods are found to easily suffer from error accumulation due to occlusions and drifts during mask propagation. Conditioning propagation on the initial frame-mask pair ⁣{}_{\!} seems a feasible ⁣{}_{\!} solution ⁣{}_{\!} to ⁣{}_{\!} this. ⁣{}_{\!} Although ⁣{}_{\!} target-specific ⁣{}_{\!} mask ⁣{}_{\!} is ⁣{}_{\!} exp- licitly encoded into the segmentation network, making up for ⁣{}_{\!} the ⁣{}_{\!} deficiencies ⁣{}_{\!} of ⁣{}_{\!} fine-tuning ⁣{}_{\!} methods ⁣{}_{\!} to ⁣{}_{\!} a ⁣{}_{\!} certain ⁣{}_{\!} extent, propagation-based ⁣{}_{\!} methods still embed object appearance into hidden network weights. Clearly, such implicit target-appearance modeling strategy hurts flexibility and adaptivity (while ⁣{}_{\!} is an exception – a generative model of target and background is explicitly built to aid mask propagation).

∙\bullet ⁣{}_{\!} Matching-based ⁣{}_{\!} Methods. ⁣{}_{\!} This type of methods, might the most promising SVOS solution so far, constructs an embedding space to memorize the initial object embeddings, and classifies each pixel’s label according to their similarities to the target object in the embedding space. Thus the initial object appearance is explicitly modeled, and test-time fine-tuning is not needed. The earliest effort in this direction can be tracked back to ⁣{}_{\!} . Inspired by the advance in visual tracking , Yoon et al. proposed a Siamese network to perform pixel-level matching between the first frame and upcoming frames. Later, proposed to learn an embedding space from the first-frame supervision and pose VOS as a task of pixel retrieval: pixels are simply their respective nearest neighbors in the learned embedding space. The idea of is also explored in , while it computes two ma- tching maps for each upcoming frame, with respect to the foreground and background annotated in the first frame. In , pixel-level similarities computed from the first frame and from the previous frame are used as a guide to segment succeeding frames. Later, many matching-based solutions were proposed , perhaps most notably Oh et al., who propose a space-time memory (STM) model to explicitly store previously computed segmentation information in an external memory . The memory facilitates learning the evolution of objects over time and allows for comprehensive use of past segmentation cues even over long period of time. Almost all current top-leading SVOS solutions are built upon STM; they improve the target adaption ability , incorporate local temporal continuity , explore instance-aware cues , and develop more efficient memory designs . Recently, introduced a Transformer based model, which performs matching-like computation through attending over a history of multiple frames. In general, matching-based solutions enjoy the advantage of flexible and differentiable model design as well as long-term correspondence modeling. On the other hand, feature matching relies on a powerful and generic feature embedding, which may limit its performance in challenging scenarios.

It is also worth mentioning that, as an effective technique for target-specific model learning, online learning is applied by many propagation and matching methods to boost performance.

∙\bullet Box-initialization based Methods. As pixel-wise annotations are time-consuming or even impractical to acquire in realistic scenes, some work has considered the situation where the first-frame annotation is provided in the form of a bounding box. Specifically, in , Siamese trackers are augmented with a mask prediction branch. In , reinforcement learning is introduced to make decisions for target updating and matching. Later, in , an outside memory is utilized to build a stronger Siamese track-segmenter.

∙\bullet Un-/Weakly-supervised based Methods. To alleviate the demand for large-scale, pixel-wise annotated training samples, several un-/weakly-supervised learning-based SVOS solutions were recently developed. They are typically built as ⁣{}_{\!} a ⁣{}_{\!} reconstruction ⁣{}_{\!} scheme ⁣{}_{\!} (i.e., ⁣{}_{\!} each ⁣{}_{\!} pixel ⁣{}_{\!} from ⁣{}_{\!} a ⁣{}_{\!} ‘query’ frame is reconstructed by finding and assembling related pixels from adjacent frame(s)) , and/or adopt a cycle-consistent tracking paradigm (i.e., pixels/patches are encouraged to fall into the same location after one cycle of forward and backward tracking) .

∙\bullet Other Specific Methods. Other papers make specific contributions that deserve a separate look. In , Zeng et al. extract mask proposals per frame and formulate the matching between object templates and proposals in a differentiable manner. Instead of using only the first frame annotation, learns to select the best frame from the whole video for user interaction, so as to boost mask propagation. In , Li et al. introduce a forward-backward data flow based cycle consistency mechanism to improve both traditional SVOS training and offline inference protocols, through mitigating the error propagation problem. To accelerate processing speed, a dynamic network is proposed to selectively allocate computation source for each frame according to the similarity to the previous frame.

III-A3 Interactive Video Object Segmentation (IVOS)

AVOS, without any human involvement, loses flexibility in segmenting arbitrary objects of user interest. SVOS addi- setting has gained increasing attention. Unlike classic models requiring extensive and professional user intervention, recent deep learning-based IVOS solutions usually work with multiple rounds of scribble supervision, to minimize the user’s effort. In this scenario , the user draws scribbles on a selected frame and an algorithm computes the segmentation maps for all video frames in a batch process. For refinement, user intervention and segmentation are repeated. This round-based interaction is useful for consumer-level applications and rapid prototyping for professional usage, where efficiency is the main concern. One can control the segmentation quality at the expense of time, as more rounds of interaction will provide better results.

∙\bullet Interaction-propagation based Methods. The majority of current studies follow an interaction-propagation scheme. In the preliminary attempt , IVOS is achieved by a simple combination of two separate modules: an interactive image segmentation model for producing segmentation based on user annotations; and a SVOS model for propagating masks from the user-annotated frames to the others. Later, devised a more compact solution, with also two modules for interaction and propagation, respectively. However, the two modules are internally connected through intermediate feature exchanging, and also externally connected, i.e., each of them is conditioned on the other’s output. In , a similar model design is also adopted, however, the propagation part is specifically designed to address both local mask tracking (over adjacent frames) and global propagation (among distant frames), respectively. However, these techniques have to start a new feed-forward computation in each interaction round, making them inefficient as the number of rounds grows. A more efficient solution was developed in . The critical idea is to build a common encoder for discriminative pixel embedding learning, upon which two small network branches are added for interactive segmentation and mask propagation, respectively. Thus the model extracts pixel embeddings for all frames only once (in the first round). In the following rounds, the feed-forward computation is only made within the two shallow branches.

∙\bullet Other Methods. Chen et al. propose a pixel embedding learning-based model, applicable to both SVOS and IVOS. With a similar idea of , IVOS is formulated as a pixel-wise retrieval problem, i.e., transferring labels to each pixel according to its nearest reference pixel. This model supports different kinds of user input, such as masks, clicks and scribbles, and can provide immediate feedback after user interaction. In , an interactive annotation tool is proposed for VOS. The annotation has two phases: annotating objects with tracked boxes, and labeling masks inside these tracks. Box tracks are annotated efficiently by approximating the trajectory using a parametric curve with a small number of control points which the annotator can interactively correct. Segmentation masks are corrected via scribbles which are propagated through time. In , a reinforcement learning framework is exploited to automatically determine the most valuable frame for interaction.

III-A4 Language-guided{}_{\!} Video{}_{\!} Object{}_{\!} Segmentation{}_{\!} (LVOS){}_{\!\!\!}

LVOS is an emerging area, dating back to 2018 . Although there have already existed some efforts in the intersection of language and video understanding, none of them addresses pixel-level video-language reasoning. Most efforts in LVOS are made around the theme of efficient alignment between visual and linguistic modalities. According to the multi-modal information fusion strategy, existing models can be divided into three groups.

∙\bullet Dynamic Convolution-based Methods. The first initiate was proposed in​ that applies dynamic networks​ for visual-language relation modeling. Specifically, convolution filters, dynamically generated from linguistic query, are used to adaptively transform visual features into desired segments. In the same line of work, incorporate spatial context into filter generation. However, as indicated by , linguistic variation of input description may greatly impact sentence representation and subsequently make dynamic filters unstable, causing inaccurate segmentation. For example, “car in blue is parked on the grass” and “blue car standing on the grass” have the same meaning but different generated filters, leading to poor performance.

∙\bullet Capsule Routing-based Methods. In , both video and textual inputs are encoded through capsules , which are considered effective in modeling visual/textual entities. Then, dynamic routing is applied over the video and text capsules for visual-textual information integration.

∙\bullet Attention-based Methods. Neural attention technique is also widely adopted in the filed of LVOS ⁣{}_{\!} , for fully capturing global visual/textual context. In , vision-guided language attention and language-guided vision attention were developed to capture visual-textual correlations. In , two different attentions are learned to ground spatial and temporal relevant linguistic cues to static and dynamic visual embeddings, respectively.

III-B Deep Learning-based VSS Models

Video semantic segmentation aims to group pixels with different semantics (e.g., category or instance membership), where different semantics result in different types of segmentation tasks, such as (instance-agnostic) video semantic segmentation (VSS, §III-B1), video instance segmentation (VIS, §III-B2) and video panoptic segmentation (VPS, §III-B3).

Extending the success of deep learning-based image semantic segmentation techniques to the video domain has become a research focus in computer vision recently. To achieve this, the most straightforward strategy is the naïve application of an image semantic segmentation model in a frame-by-frame manner. But this strategy completely ignores temporal continuity and coherence cues provided in videos. To make better use of temporal information, research efforts in this field are mainly made along two lines.

∙\bullet Efforts towards More Accurate Segmentation. A major stream of methods exploits cross-frame relations to boost the prediction accuracy. They typically first apply the very same segmentation algorithms to each frame independently. Then they add extra modules on top, e.g., optical flow-guided feature aggregation , and sequential network based temporal information propagation , to gather multi-frame context and get better results. For example, in some pioneer work , after performing static semantic segmentation for each frame individually, optical ⁣{}_{\!} flow ⁣{}_{\!}  ⁣{}_{\!} or ⁣{}_{\!} 3D ⁣{}_{\!} CRF ⁣{}_{\!}  ⁣{}_{\!} based ⁣{}_{\!} post ⁣{}_{\!} processing is applied for gaining temporally consistent segments. Later, jointly learns CNN-based per-frame segmentation and CRF-based spatio-temporal reasoning. In , features warped from previous frames with optical flow are combined with the current frame features for prediction. These methods require additional feature aggregation modules, which increase the computational costs during the inference. Recently, proposes to only incorporate flow-guided temporal consistency into the training phase, without bringing any extra inference cost. But its processing speed is still bounded to the adopted per-frame segmentation algorithms, as all features must be recomputed at each frame. For these methods, the utility in time-sensitive application areas, such as mobile and autonomous driving, is limited.

∙\bullet Efforts towards Faster Segmentation. Yet another complementary line of work tries to leverage temporal information to accelerate computation. They approximate the expensive per-frame forward pass with cheaper alternatives, i.e., reusing the features in neighbouring frames. In , parts of segmentation networks are adaptively executed across frames, thus reducing the computation cost. Later methods use keyframes to avoid processing of each frame, and then propagate the outputs or the feature maps to other frames. For instance, employs optical flow to warp the features between the keyframe and non-key frames. Adaptive keyframe selection is later exploited in , further enhanced by adaptive feature propagation . In , Jain et al. use a large, strong model to predict the keyframe and use a compact one in non-key frames. Keyframe-based methods have different computational loads between keyframes and non-key frames, causing high maximum latency and unbalanced occupation of computation resources that may decrease system efficiency . Additionally, the spatial misalignment of other frames with respect to the keyframes is challenging to compensate for and often leads to different quantity results between keyframes and non-key frames. In , a temporal consistency guided knowledge distillation technique is proposed to train a compact network, which is applied to all frames. In , several weight-sharing sub-networks are distributed over sequential frames, whose extracted shallow features are composed for final segmentation. This trend of methods indeed speeds up inference, but still with the cost of reduced accuracy.

∙\bullet Semi-/Weakly-supervised based Methods. Away from these main battlefields, some researchers made efforts to learn VSS under annotation efficient settings. In , classifier heatmaps are used to learn VSS from image tags only. use both labeled and unlabeled video frames to learn VSS. They propagate annotations from labeled frames to other unlabeled, neighboring frames , or alternatively train teacher and student networks with groundtruth annotations and iteratively generated pseudo labels .

III-B2 Video Instance Segmentation (VIS)

In 2019, Yang et al. extended image instance segmentation to the video domain , which requires simultaneous detection, segmentation and tracking of instances in videos. This task is also known as multi-object tracking and segmentation (MOTS) . Based on the patterns of generating instance sequences, existing frameworks can be roughly categorized into four paradigms: i) track-detect, ii) clip-match, iii) propose-reduce, iv) segment-as-a-whole. Track-detect methods detect and segment instances for each individual frame, followed by frame-by-frame instance tracking . For example, in , Mask R-CNN is adapted for VIS/MOTS by adding a tracking branch for cross-frame instance association. Alternatively, models spatial attention to describe instances, tackling the task from a novel single-stage yet elegant perspective. Clip-match methods divide an entire video into multiple overlapped clips, and perform VIS independently for each clip through mask propagation or spatial-temporal embedding . Final instance sequences are generated by merging neighboring clips. Both of the two paradigms need two independent steps to generate a complete sequence. They both generate multiple incomplete sequences (i.e., frames or clips) from a video, and merge (or complete) them by tracking/matching at the second stage. Intuitively, these paradigms are vulnerable to error accumulation in the process of merging sequences, especially when occlusion or fast motion exists. To address these limitations, a propose-reduce paradigm is proposed in . It first samples several key frames and obtains instance sequences by propagating the instance segmentation results from each key frame to the entire video. Then, the redundant sequence proposals of the same instances are removed. This paradigm not only discards the step of merging incomplete sequences, but also achieves robust results considering multiple key frames. However, these three types of methods still need complex heuristic rules to associate instances and/or multiple steps to generate instance sequences. The segment-as-a-whole paradigm is more elegant; it poses the task as a direct sequence prediction problem using Transformer .

Almost all VIS models are built upon fully supervised learning, while are the exceptions. Specifically, in , motion and temporal consistency cues are leveraged to generate pseudo-labels from tag labeled videos for weakly supervised VIS learning. In , a semi-supervised embedding learning approach is proposed to learn VIS from pixel-wise annotated images and unlabeled videos.

III-B3 Video Panoptic Segmentation (VPS)

Very recently, Kim et al. extended image panoptic segmentation to the video domain , which aims at a holistic segmentation of all foreground instance tracklets and background regions, and assigning a semantic label to each video pixel. They adapt an image panoptic segmentation model for VPS, by adding two modules for temporal feature fusion and cross-frame instance association, respectively. Later, temporal correspondence was explored in through learning coarse segment-level and fine pixel-level matching. Qiao et al. propose to learn monocular depth estimation and video panoptic segmentation jointly.

IV Video Segmentation Datasets

Several datasets have been proposed for video segmentation over the past decades. Fig. 3 shows example frames from twenty commonly used datasets. We summarize their essential features in Table VI and give detailed review below.

∙\bullet Youtube-Objects is a large dataset of 1, ⁣4071,\!407 videos collected from 155 web videos belonging to 10 object categories (e.g., dog, cat, plane, etc.). VOS models typically test the generalization ability on a subset having totally 126 shots with 20, ⁣64720,\!647 frames that provides coarse pixel-level fore-/background annotations on every 10th frames.

∙\bullet FBMS59​ consists of 59 video sequences with 13, ⁣86013,\!860 frames in total. However, only 720 frames are annotated for fore-/background separation. The dataset is split into 29 and 30 sequences for training and evaluation, respectively.

∙\bullet DAVIS16​ has 50 videos (30 for train set and 20 for val set) with 3, ⁣4553,\!455 frames in total. For each frame, in addition to high-quality fore-/background segmentation annotation, a set of attributes (e.g., deformation, occlusion, motion blur, etc.) are also provided to highlight the main challenges.

∙\bullet DAVIS17​ contains 150 videos, i.e., 60/30/30/30 videos for train/val/test-dev/test-challenge sets. Its train and val sets are extended from the respective sets in DAVIS16. There are 10,459 frames in total. DAVIS17 provides instance-level annotations to support SVOS. Then, DAVIS18 challenge​ provides scribble annotations to support IVOS. Moreover, as the original annotations of DAVIS17 are biased towards the SVOS scenario, DAVIS19 challenge​ re-annotates val and test-dev sets of DAVIS17 to support AVOS.

∙\bullet YouTube-VOS​ is a large-scale dataset, which is split into a train (3, ⁣4713,\!471 videos), val (507 videos), and test (541 videos) set, in its newest 2019 version. Instance-level precise annotations are provided every five frames in a 30FPS frame rate. There are 94 object categories (e.g., person, snake, etc.) in total, of which 26 are unseen in train set.

Remark. Youtube-Objects, FBMS59 and DAVIS16 are used

for instance-agnostic AVOS and SVOS evaluation. DAVIS17 is unique in comprehensive annotations for instance-level AVOS, SVOS as well as IVOS, but its scale is relatively small. YouTube-VOS is the largest one but only supports SVOS benchmarking. There also exist some other VOS datasets, such as SegTrackV1 and SegTrackV2 , but they were less used recently, due to the limited scale and difficulty.

IV-A2 LVOS Datasets

∙\bullet A2D Sentence​ augments A2D with phrases. It contains 3, ⁣7823,\!782 videos, with 88 action classes performed by 77 actors. In each video, 33 to 55 frames are provided with segmentation masks. It contains 6, ⁣6556,\!655 sentences describing actors and their actions. The dataset is split into 3, ⁣0173,\!017/737737 for train/test, and 2828 unlabeled videos are ignored .

∙\bullet J-HMDB Sentence​ is built upon J-HMDB . It is comprised of 928928 short videos with 928928 corresponding sentences describing 2121 different action categories.

∙\bullet DAVIS17-RVOS​ extends DAVIS17 by collecting referring expressions for the annotated objects. 90 videos from train and val sets are annotated with more than 1,500 referring expressions. They provide two types of annotations, which describe the highlighted object: 1) based on the entire video (i.e., full-video expression) and 2) using only the first frame of the video (i.e., first-frame expression).

∙\bullet Refer-Youtube-VOS​ includes 3,975 videos from YouTube-VOS​ , with 27,899 language descriptions of target objects. Similar to DAVIS17-RVOS , both full-video and first-frame expression annotations are provided.

Remark. To date, A2D Sentence and J-HMDB Sentence are the main test-beds. However, the phrases are not produced with the aim of reference, but description, and limited to only a few object categories corresponding to the dominant ‘actors’ performing a salient ‘action’ . But newly introduced DAVIS17-RVOS and Refer-Youtube-VOS show improved difficulties in both visual and linguistic modalities.

IV-B VSS Datasets

∙\bullet CamVid​ is composed of 4 urban scene videos with 11-class pixelwise annotations. Each video is annotated every 30 frames. The annotated frames are usually grouped into 467/100/233 for train/val/test .

∙\bullet CityScapes ⁣ {}_{\!~}  ⁣ {}_{\!~} is ⁣ {}_{\!~} a ⁣ {}_{\!~} large-scale ⁣ {}_{\!~} VSS ⁣ {}_{\!~} dataset ⁣ {}_{\!~} for ⁣ {}_{\!~} street

views. It has 2,975/500/1,525 snippets for train/val/ test, captured at 17FPS. Each snippet contains 30 frames, and only the 20th frame is densely labelled with 19 semantic classes. 20,000 coarsely annotated frames are also provided.

∙\bullet NYUDv2​ contains 518 indoor RGB-D videos with high-quality ground-truths (every 10th video frame is labeled). There are 795 training frames and 654 testing frames being rectified and annotated with 40-class semantic labels.

∙\bullet VSPW​ is a recently proposed large-scale VSS dataset. It addresses video scene parsing in the wild by considering diverse scenarios. It consists of 3,536 videos, and provides pixel-level annotations for 124 categories at 15FPS. The train/val/test sets contain 2,806/343/387 videos with 198,244/24,502/28,887 frames, respectively.

∙\bullet YouTube-VIS​ is built upon YouTube-VOS with instance-level annotations. Its newest 2021 version has 3,859 videos (2,985/421/453 for train/val/test) with 40 semantic categories. It provides 232K high-quality annotations for 8,171 unique video instances.

∙\bullet KITTI MOTS​ extends the 21 training sequences of KITTI tracking dataset with VIS annotations – 12 for training and 9 for validation, respectively. The dataset contains 8,008 frames with a resolution of 375×1242375\times 1242, 26,899 annotated cars and 11,420 annotated pedestrians.

∙\bullet MOTSChallenge​ annotates 4 of 7 training sequences of MOTChallenge2017 . It has 2,862 frames with 26,894 annotated pedestrians and presents many occlusion cases.

∙\bullet BDD100K​ is a large-scale dataset with 100K driving videos (40 seconds and 30FPS each) and supports various tasks, including VSS and VIS. For VSS, 7,000/1,000/2,000 frames are densely labelled with 40 semantic classes for train/val/test. For VIS, 90 videos with 8 semantic categories are annotated by 129K instance masks – 60 training videos, 10 validation videos, and 20 testing videos.

∙\bullet OVIS​ is a new challenging VIS dataset, where object occlusions usually occur. It has 901 videos and 296K high-quality instance masks for 25 semantic categories. It is split into 607 training, 140 validation and 154 test videos.

∙\bullet VIPER-VPS​ re-organizes VIPER into the video panoptic format. VIPER, extracted from the GTA-V game engine, has annotations of semantic and instance segmentations for 10 thing and 13 stuff classes on 254K frames of ego-centric driving scenes at 1080 ⁣× ⁣19201080\!\times\!1920 resolution.

∙\bullet Cityscapes-VPS​ is built upon CityScapes​ . Dense panoptic annotations for 8 thing and 11 stuff classes for 500 snippets in Cityscapes val set are provided every five frames and temporally consistent instance ids to the thing objects are also given, leading to 3000 annotated frames in total. These videos are split into 400/100 for train/val.

Remark. CamVid, CityScapes, NYUDv2, and VSPW are built for VSS benchmarking. YouTube-VIS, OVIS, KITTI MOTS, and MOTSChallenge are VIS datasets, but the diversity of the last two are limited. BDD100K has both VSS and VIS annotations. VIPER-VPS and Cityscapes-VPS are aware of VPS evaluation, but VIPER-VPS is a synthesized dataset.

V Performance Comparison

Next we tabulate the performance of previously discussed algorithms. For each of the reviewed fields, the most widely used dataset is selected for performance benchmarking. The performance scores are gathered from the original articles, unless specified. For the running speed, we obtain the FPS for most methods by running their codes on a RTX 2080Ti GPU. For a small set of methods whose implementations are not well organized or publicly available, we directly borrow the values from the corresponding papers. Despite this, it is essential to remark the difficulty when comparing runtime. As different methods are with different code bases and levels of optimization, it is hard to make completely fair runtime comparison ; the values are only provided for reference.

Presently, three metrics are frequently used to measure how object-level AVOS methods perform on this task:

∙\bullet Region Jaccard J\mathcal{J} is calculated by the intersection-over-union (IoU) between the segmentation results Y^ ⁣∈ ⁣{0,1}w×h{\hat{Y}}\!\in\!\{0,1\}^{w\times h} and the ground-truth Y ⁣∈ ⁣{0,1}w×h{Y}\!\in\!\{0,1\}^{w\times h}: J=∣Y^∩Y∣/∣Y^∪Y∣\mathcal{J}={|\hat{Y}\cap Y}|/|{\hat{Y}\cup Y|}, which computes the number of pixels of the intersection between Y^\hat{Y} and Y{Y}, and divides it by the size of the union.

∙\bullet Boundary Accuracy F\mathcal{F} is the harmonic mean of the boundary precision Pc\text{P}_{c} and recall Rc\text{R}_{c}. The value of F\mathcal{F} reflects how well the segment contours c(Y^)c(\hat{Y}) match the ground-truth contours c(Y)c(Y). Usually, the value of Pc\text{P}_{c} and Rc\text{R}_{c} can be computed via bipartite graph matching , then the boundary accuracy F\mathcal{F} can be computed as: F=2PcRc/(Pc+Rc)\mathcal{F}={2\text{P}_{c}\text{R}_{c}}/({\text{P}_{c}+\text{R}_{c}}).

∙\bullet Temporal Stability T\mathcal{T} is informative of the stability of segments. It is computed as the pixel-level cost of matching two successive segmentation boundaries. The match is achieved by minimizing the shape context descriptor distances between matched points while preserving the order in which the points are present in the boundary polygon. Note that T\mathcal{T} will compensate motion and small deformations, but not penalize inaccuracies of the contours .

V-A2 Results

We select DAVIS16​ , the most widely used dataset in AVOS, for performance benchmarking. Table VII presents the results of those reviewed AVOS methods DAVIS16 val set. The current best solution, RTNet , reaches 85.6 region similarity J\mathcal{J}, significantly outperforming the earlier deep learning-based methods, such as SFL , proposed in 2017.

V-B Instance-level AVOS Performance Benchmarking

In instance-level AVOS setting, region Jaccard J\mathcal{J}, boundary accuracy F\mathcal{F}, and J&F\mathcal{J}\&\mathcal{F} – the mean of J\mathcal{J} and F\mathcal{F} – are used for evaluation . Each of the annotated object tracklets will be matched with one of predicted tracklets according to J&F\mathcal{J}\&\mathcal{F}, using bipartite graph matching. For a certain criterion, the final score will be computed between each ground-truth object and its optimal assignment.

V-B2 Results

Regarding instance-level AVOS, we take into account DAVIS17 in which the vast majority of methods are evaluated. From Table 8 we can find that UnOVOST is the top scorer, with 67.9 J\mathcal{J} at the time of this writing.

V-C SVOS Performance Benchmarking

Region Jaccard J\mathcal{J}, boundary accuracy F\mathcal{F}, and J&F\mathcal{J}\&\mathcal{F} are also widely adopted for SVOS performance evaluation .

V-C2 Results

DAVIS17 is also one of the most important SVOS dataset. Table 9 shows the results of recent SVOS methods on DAVIS17 val set. In this case, all the top-leading solutions, such as EGMN , LCM , and RMNet , are built upon the memory augmented architecture – STM .

V-D IVOS Performance Benchmarking

Area under the curve (AUC) and Jaccard at 60 seconds (J\mathcal{J}@60s) are two widely used IVOS evaluation criteria .

∙\bullet AUC is designed to measure the overall accuracy of the evaluation. It is computed over the plot Time vs Jaccard. Each sample in the plot is computed considering the average time and the average Jaccard for a certain interaction.

∙\bullet J\mathcal{J}@60 measures the accuracy with a limited time budget (60 seconds). It is achieved by interpolating the Time vs Jaccard plot at 60 seconds. This evaluates which quality an IVOS method can obtain in 60 seconds.

V-D2 Results

DAVIS17 is also widely used for IVOS performance benchmarking. Results summarized in Table V-A1 show that the method proposed by Cheng et al. is the top one.

V-E LVOS Performance Benchmarking

As , overall IoU, mean IoU and precision are adopted.

∙\bullet IoU: overall IoU is computed as total intersection area of all test data over the total union area, while mean IoU refers to average over IoU of each test sample.

∙\bullet Precision: Precision@KK is computed as the percentage of test samples whose IoU scores are higher than a threshold KK. Precision at five thresholds ranging from 0.5 to 0.9 and mean ⁣{}_{\!} average ⁣{}_{\!} precision ⁣{}_{\!} (mAP) ⁣{}_{\!} over ⁣{}_{\!} 0.5:0.05:0.95 ⁣{}_{\!} are ⁣{}_{\!} reported.

V-E2 Results

A2D Sentence is arguably the most popular dataset in

LVOS. Table V-A1 gives the results of six recent methods on A2D Sentence test set. It shows clear improvement trend from the first LVOS model proposed in 2018, to recent complicated solution . For runtime comparison, all the methods are tested on a video clip of 16 frames with resolution 512 ⁣× ⁣512512\!\times\!512 and a textual sequence of 20 words.

V-F VSS Performance Benchmarking

IoU metric is the most widely used metric in VSS. Moreover, in Cityscapes – the gold-standard benchmark dataset in this field, two IoU scores, IoUcategory{}_{\text{category}} and IoUclass{}_{\text{class}}, defined over two semantic granularities, are reported. Here, ‘category’ refers to high-level semantic categories (e.g., vehicle, human), while ‘class’ indicates more fine-grained semantic classes (e.g., car, bicycle, person, rider). In total, considers 1919 classes, which are further grouped into 88 categories.

V-F2 Results

Table 12 summarizes the results of eleven VSS approaches on Cityscapes val set. As seen, EFC performs the best currently, with 83.5%83.5\% in terms of IoUclass{}_{\text{class}}.

V-G VIS Performance Benchmarking

As in , precision and recall metrics are used for VIS performance evaluation. Precision at IoU thresholds 0.5 and 0.75, as well as mean average precision (mAP) over 0.50:0.05:0.95 are reported. Recall@NN is defined as the maximum recall given NN segmented instances per video. These two metrics are first evaluated per category and then averaged over the category set. The IoU metric is similar to region Jaccard J\mathcal{J} used in instance-level AVOS (§V-B1).

V-G2 Results

Table 13 gathers VIS results for on YouTube-VIS val set, showing that Transformer-based architecture, i.e., VisTR , and redundant sequence proposal based solution Propose-Reduce , greatly improve the state-of-the-art.

V-H VPS Performance Benchmarking

In , the panoptic quality (PQ) metric used in image panoptic segmentation is modified as video panoptic quality (VPQ) to adapt to video panoptic segmentation.

∙\bullet ⁣{}_{\!} VPQ: ⁣{}_{\!} Given ⁣{}_{\!} a ⁣{}_{\!} snippet ⁣{}_{\!} Vt:t+k ⁣V^{t:t+k\!} with ⁣{}_{\!} time ⁣{}_{\!} window ⁣{}_{\!} kk, ⁣{}_{\!} true ⁣{}_{\!} po- ⁣{}_{\!} sitive (TP) is defined by TP ⁣= ⁣{(u,u^) ⁣ ⁣∈ ⁣U ⁣ ⁣× ⁣U^ ⁣ ⁣: ⁣IoU(u,u^) ⁣> ⁣0.5}\text{TP}\!=\!\{(u,\hat{u})_{\!}\!\in\!U_{\!}\!\times\!\hat{U}_{\!}\!:_{\!}\text{IoU}(u,\hat{u})\!>\!0.5\} where UU and U^\hat{U} are the set of the ground-truth and predicted tubes, respectively. False Positives (FP) and False Negatives (FN) are defined accordingly. After accumulating TPc, FPc, and FNc on all the clips with window size kk and class cc, we define: ⁣{}_{\!} VPQ ⁣k ⁣= ⁣ ⁣1Nclass ⁣∑c ⁣∑(u,u^)∈TPc ⁣IoU(u,u^)∣TPc∣+12∣FPc∣+12∣FNc∣\text{VPQ}^{k}_{\!}\!=_{\!}\!\frac{1}{N_{\text{class}}}\!\sum_{c}\!\frac{\sum_{(u,\hat{u})\in\text{TP}_{c}}\!\text{IoU}(u,\hat{u})}{|\text{TP}_{c}|+\frac{1}{2}|\text{FP}_{c}|+\frac{1}{2}|\text{FN}_{c}|}. ⁣{}_{\!} When ⁣{}_{\!} k ⁣= ⁣1k\!=\!1, ⁣{}_{\!} VPQ1

is equivalent to PQ. For evaluation, VPQk\text{VPQ}^{k} is reported over k ⁣∈ ⁣{0,5,10,15}k\!\in\!\{0,5,10,15\} and finally, VPQ ⁣= ⁣14∑k∈{0,5,10,15} ⁣VPQk\text{VPQ}\!=\!\frac{1}{4}\sum_{k\in\{0,5,10,15\}\!}\text{VPQ}^{k}.

V-H2 Results

​Cityscapes-VPS​  ⁣{}_{\!} is ⁣{}_{\!} chosen ⁣{}_{\!} for ⁣{}_{\!} testing ⁣{}_{\!} VPS ⁣{}_{\!} methods. ⁣{}_{\!} As ⁣{}_{\!} shown ⁣{}_{\!} in ⁣{}_{\!} Table​ 14, ViP-DeepLab​  ⁣{}_{\!} is ⁣{}_{\!} the ⁣{}_{\!} top ⁣{}_{\!} one.

V-I Summary

From the results, we can draw several conclusions. The most important of them is related to reproducibility. Across different video segmentation areas, many methods do not describe the setup for the experimentation or do not provide the source code for implementation. Some of them even do not release segmentation masks. Moreover, different methods use various datasets and backbone models. These make fair comparison impossible and hurt reproducibility.

Another important fact discovered thanks to this study is the lack of information about execution time and memory use. Many methods particularly in the fields of AVOS, LVOS, and VPS, do not report execution time and almost no paper reports memory use. This void is due to the fact that many methods focus only on accuracy without any concern about running time efficiency or memory requirements. However, in many application scenarios, such as mobile devices and self-driving cars, computational power and memory are typically limited. As benchmark datasets and challenges serve as a main driven factor behind the fast evolution of segmentation techniques, we encourage organizers of future video segmentation datasets to give this kind of metrics its deserved importance in benchmarking.

Finally, performance on some extensively studied video segmentation datasets, such as DAVIS16 in AVOS, DAVIS17 in SVOS, A2D Sentence in LVOS, have nearly reached saturation. Though some new datasets are proposed recently and claim huge space for performance improvement, the dataset collectors just gather more challenging samples, without necessarily figuring out which exact challenges have and have not been solved.

VI Future Research Directions

Based on the reviewed research, we list several future research directions that we believe should be pursued.

∙\bullet Long-Term Video Segmentation: Long-term video segmentation is much closer to practical applications, such as video editing. However, as the sequences in existing datasets often span several seconds, the performance of VOS models over long video sequences (e.g., at the minute level) are still unexamined. Bringing VOS into the long-term setting will unlock new research lines, and put forward higher demand of the re-detection capability of VOS models.

∙\bullet Open World Video Segmentation: Despite the obvious dynamic and open nature of the world, current VSS algorithms are typically developed in a closed-world paradigm, where all the object categories are known as a prior. These algorithms are often brittle once exposed to the realistic complexity of the open world, where they are unable to efficiently adapt and robustly generalize to unseen categories. For example, practical deployments of VSS systems in robotics, self-driving cars, and surveillance cannot afford to have complete knowledge on what classes to expect at inference time, while being trained in-house. This calls for smarter VSS systems, with a strong capability to identify unknown categories in their environments .

∙\bullet Cooperation across Different Video Segmentation Sub-fields: VOS and VSS face many common challenges, e.g., object occlusion, deformation, and fast motion. Moreover, there are no precedents for modeling these tasks in a unified framework. Thus we call for closer collaboration across different video segmentation sub-fields.

∙\bullet Annotation-Efficient Video Segmentation Solutions: Though great advances have been achieved in various videos segmentation tasks, current top-leading algorithms are built on fully-supervised deep learning techniques, requiring a huge amount of annotated data. Though semi-supervised, weakly supervised and unsupervised alternatives were explored in some literature, annotation-efficient solutions receive far less attention and typically show weak performance, compared with the fully supervised ones. As the high temporal correlations in video data can provide additional cues for supervision, exploring existing annotation-efficient techniques in static semantic segmentation in the area of video segmentation is an appealing direction.

∙\bullet Adaptive Computation: It is widely recognized that there exist high correlations among video frames. Though such data redundancy and continuity are exploited to reduce the computation cost in VSS, almost all current video segmentation models are fixed feed-forward structures or work alternatively between heavy and light-weight modes. We expect more flexible segmentation model designs towards more efficient and adaptive computation , which allows network architecture change on-the-fly – selectively activating part of the network in an input-dependent fashion.

∙\bullet Neural Architecture Search: Video segmentation models are typically built upon hand-designed architectures, which may be suboptimal for capturing the nature of video data and limit the best possible performance. Using neural architecture search techniques to automate the design of video segmentation networks is a promising direction.

VII CONCLUSION

To ⁣{}_{\!} our ⁣{}_{\!} knowledge, ⁣{}_{\!} this ⁣{}_{\!} is ⁣{}_{\!} the ⁣{}_{\!} first ⁣{}_{\!} survey ⁣{}_{\!} to ⁣{}_{\!} comprehensively review recent progress in video segmentation. We provided the reader with the necessary background knowledge and summarized more than 150 deep learning models according to various criteria, including task settings, technique contributions, and learning strategies. We also presented a structured survey of 20 widely used video segmentation datasets and benchmarking results on 7 most widely-used ones. We discussed the results and provided insight into the shape of future research directions and open problems in the field. In conclusion, video segmentation has achieved notable progress thanks to the striking development of deep learning techniques, but several challenges still lie ahead.

References