Efficient Video Instance Segmentation via Tracklet Query and Proposal
Jialian Wu, Sudhir Yarram, Hui Liang, Tian Lan, Junsong Yuan, Jayan Eledath, Gerard Medioni
Introduction
Video Instance Segmentation (VIS) is a challenging video task recently introduced in . It aims to predict a tracklet segmentation mask with a class label for every appeared object instance in a video as illustrated in Fig. 1. Existing methods typically solve the VIS problem at either frame-level or clip-level. The frame-level methods follow a tracking-by-segmentation paradigm, which first performs image instance segmentation and then links the current masks with history tracklets via data association as shown in Fig. 1 (a). This paradigm typically requires complex data association algorithms and exploits limited temporal context, making it susceptible to object occlusions. By contrast, the clip-level methods jointly perform segmentation and tracking clip-by-clip. Within each clip, object information is propagated back and forth. Such a paradigm usually performs stronger than the frame-level methods thanks to larger temporal receptive field. However, most clip-level methods are not end-to-end and require elaborated inference that causes slow speed. These issues are addressed by the recent VIS transformer (VisTR) which extends image object detector DETR to the VIS task. VisTR generates VIS predictions within each video clip in one end-to-end pass, which greatly simplifies the clip-level paradigm and makes VIS within a clip end-to-end trainable.
However, the VIS transformer is confronted with two issues: (i) The convergence speed of VisTR is slow since the frame-wise dense attention weights in the transformer need long training epochs to search for properly attended regions among all video frame pixels. It is difficult for researchers to experiment with this algorithm as it requires long development cycles. (ii) When a video is too long to fit into GPU memory in one forward-pass, it has to be divided into multiple successive clips for sequential processing. In this case, VisTR becomes partially end-to-end. As shown in Fig. 1 (b), VisTR requires a hand-crafted data association to link tracklets between clips. Such a hand-crafted scheme is not only more complicated but also less effective than an end-to-end learnable association as evidenced in Table 1(d).
In this paper, we propose a new clip-level VIS framework, coined as EfficientVIS. EfficientVIS delivers a fully end-to-end framework that can achieve fast convergence and inference with strong performance. Our work is inspired by the recent success of the query-based R-CNN architecture in image object detection. EfficientVIS extends the spirit to the video domain to model the complex space-time interactions. Specifically, EfficientVIS uses tracklet query paired with tracklet proposal to represent an individual object instance in video. Tracklet query is latent embeddings that encode appearance information for a target instance, while tracklet proposal is a space-time tube that locates the target in the video. Our tracklet query collects the target information by interacting with video in a clip-by-clip fashion through a designed temporal dynamic convolution. This way enriches temporal object context that is an important cue for handling object occlusion and motion blur in videos. We also design a factorised temporo-spatial self-attention allowing tracklet queries to exchange information over not only space but also time. It enables one query to correlate a target across multiple frames so as to end-to-end generate target tracklet mask as a whole in a video clip. Compared to prior works, EfficientVIS enjoys three remarkable properties:
(i) Fast convergence: In EfficientVIS, tracklet queries interact with video CNN features only in the region of the space-time RoIs defined by the tracklet proposal. This is different from VisTR that is interacting with all video pixels on the transformer encoded features using dense attention. Our RoI-wise design drastically reduces video redundancies and therefore allows EfficientVIS for faster convergence than transformer as shown in Table 1(g).
(ii) Fully end-to-end learnable framework: EfficientVIS goes beyond a short clip and is fully end-to-end learnable over the whole video. When there are multiple successive clips, one needs to link instance tracklets between clips. In contrast to prior clip-level works that manually stitch the tracklets, we design a correspondence learning that enables tracklet query to be shared among clips for seamlessly associating a same instance. In other words, the tracklet query output from one clip is enabled to be fed into the next clip to associate and segment the same instance. Meanwhile, the query is dynamically updated in terms of the next clip content so as to achieve a continuous tracking for the future. Such a scheme makes EfficientVIS fully end-to-end, without any explicit data association for either inner-clip or inter-clip tracking as shown in Fig. 1 (c).
(iii) Tracking in low frame rate videos: Tracking object instances that have dramatic movements is a great challenge for motion-based trackers , as they suppose instances move smoothly over time. In contrast, our method retrieves a target instance in a frame conditioned on its query representation regardless of where it is in nearby frames. Thus, EfficientVIS is robust to dramatic object movements and can track instances in low frame rate videos as shown in Fig. 5 and Table 1(h).
We summarize our major contributions as follows:
EfficientVIS is the first RoI-wise clip-level VIS framework that runs in real-time. The RoI-wise design enables a fast convergence by drastically reducing video redundancies. Fully end-to-end learnable tracking and rich temporal context of the clip-by-clip workflow together bring a strong performance. EfficientVIS ResNet-50 achieves 37.9 AP on Youtube-VIS in 36 FPS by training 33 epochs, which is 15 training epochs fewer and 2.3 AP higher than VIS transformer.
EfficientVIS is the first fully end-to-end neural net for VIS. Given a video as input despite its length, EfficientVIS directly produces VIS predictions without any data association or post-processing. We will demonstrate by diagnostic experiments that this fully end-to-end paradigm is not only simpler but also more effective than the previous partially/non end-to-end frameworks.
Related Works
Frame-level VIS: Most video instance segmentation methods work at the frame-level fashion, a.k.a. tracking-by-segmentation . This paradigm produces instance segmentation frame-by-frame and achieves tracking by linking the current instance mask to the history tracklet. So it requires an explicit data association algorithm in either a hand-crafted or trainable way. Some attempts strive to exploit temporal context to improve VIS, yet their temporal context is limited to one or only a few frames ignoring fertile resources of videos. In contrast to the above methods, our method does not require data association algorithm at all and we take advantage of richer temporal context by performing VIS clip-by-clip.
Clip-level VIS: Recent clip-level works demonstrate promising VIS accuracy by exploiting rich temporal knowledge from multiple frames of a video clip. This paradigm propagates object information back and forth in a clip, which can well handle object occlusion and motion blur. However, most clip-level methods have a complex inference process (e.g. extra mask refinement, box ensemble, etc) that makes them neither end-to-end learnable nor real-time in inference. To address these limitations, the VIS transformer (VisTR) extends DETR from images to videos, achieving efficient and end-to-end VIS within each clip. Nevertheless, VisTR suffers from slow convergence due to the frame-wise dense attention. Besides, if there are two or more video clips, VisTR is not fully end-to-end and requires a hand-crafted data association to link tracklets between clips. In contrast to the above methods, EfficientVIS is fully end-to-end learnable without the need of data association for either inner- or inter-clips. Moreover, EfficientVIS performs VIS in an RoI-wise manner through an efficient query-video interaction, enabling fast convergence and inference.
EfficientVIS
EfficientVIS aims to deliver four key features: (i) Fast training and inference; (ii) Fully end-to-end learnable framework without any data association/post-processing; (iii) Tracking in low frame rate videos; (iv) Strong performance. EfficientVIS performs VIS clip-by-clip, where we present the main VIS architecture for inner-clip in Sec. 3.1 and the inter-clip tracklets correspondence learning in Sec. 3.3.
The temporal self-attention allows the embeddings of one instance query from different frames to exchange the target instance information, such that the embeddings can jointly associate the same instance over time. Therefore, instance association/tracking in each clip is implicitly achieved and end-to-end learned by our temporal self-attention. The spatial self-attention is then computed at the same frame, across the embedding pairs from different tracklet queries as:
The spatial self-attention allows each query to acquire object context from other queries in each frame, reasoning about the relations among different object instances. We observe in experiments that the order of temporal and spatial self-attention does not cause a noticeable performance difference.
Joint temporo-spatial self-attention that computes MSA across all embedding pairs for once is another alternative for query communication. Compared to this scheme, our FTSA saves more computation and is more effective as we will demonstrate in Table 1(a).
Temporal Dynamic Convolution (TDC): The aim of TDC is to collect the target instance information from the video clip. Specifically, we generate a dynamic convolutional filter conditioned on tracklet query embedding. We then use it to perform convolution on an RoI region of the base feature specified by the corresponding tracklet proposal. Different from the still-image dynamic convolution , we perform a 3D dynamic convolution in order to collect temporal instance context from nearby frames as well. Let denote the dynamic convolutional filter generated from tracklet query embedding . The TDC is computed as:
Since the dynamic filter is generated from , has distinct cues and appearance semantics of the -th instance. The TDC exploits to exclusively collect the -th instance information from the video base feature yet filter out irrelevant information. Therefore, tracklet feature shall be highly activated by in the region where the -th instance appears but will be suppressed in the region of uncorrelated instances or background clutter as evidenced in Fig. 4.
Head Networks: For each instance , we employ light-weight head networks on its tracklet feature to output its VIS predictions and renewed tracklet query and proposal. Since tracklet feature exclusively carries the -th instance information across the clip, these outputs can be readily derived by applying regular convolutions or fully connected layers (FCs). Specifically, we apply convolutions to segment tracklet mask, and FCs to classify tracklet, regress tracklet proposal, and update tracklet query. The updated query and proposal are fed into the next iteration. The head network architectures are similar to . In a nutshell, our tracklet (proposal) for a video clip is directly generated as a whole determined by tracklet query. So we do not need explicit data association to link instances/boxes across frames in a clip as frame-level VIS methods .
2 Training
For each video clip, EfficientVIS is trained by the one-to-one matching loss , which first obtains a one-to-one assignment between predictions and ground truths by bipartite matching and then computes the training loss in terms of the assignment. Let denote the tracklet predictions and denote the tracklet ground truths, where () is the total number of ground truths. We denote as a -elements permutation of . The bipartite matching between and is found by searching for a with the lowest cost as:
where is the binary cross-entropy loss and is the dice loss . All the losses and balancing weights are similar to , except that we add a to enforce predicted mask to be zero if the instance disappears. Eq. 5 is averaged by the total number of entries. For those unmatched tracklet predictions, we only impose classification loss to supervise them to be background class.
3 Correspondence Learning
If a long video cannot fit into GPU memory in one forward-pass, it needs to be split into multiple successive clips for sequential processing. To stitch tracklets between two clips, prior works turn to hand-crafted data association. However, human-designed rules are cumbersome and not end-to-end learnable. Therefore, we design a correspondence learning that makes the tracklet linking between clips end-to-end learnable without data association, which is not only simpler but also more effective.
Correspondence learning: As shown in Fig. 3, every two clips from the input video are pairedThe order of the two clips in a pair is randomized. in training. For the first clip of a pair, the initial tracklet queries are the model parameters as described in Sec. 3.1, and the training procedure is the same as Sec. 3.2, i.e. finding the optimal assignment and backpropagating . For the initial queries of the second clip, we use the output tracklet queries from the first clip averaged along the time dimension. The average operation can extract a comprehensive abstract for the instance representation. As for the assignment of the second clip, we do not perform bipartite matching. Instead, we use the same assignment that is already obtained by the first clip to train the second clip. Such a training manner enforces the tracklet queries output from one clip to segment and correspond to the same object instances in the next clip. In this way, tracklet query is able to serve as a generic video-level representation to seamlessly associate and correlate an instance across a whole video instead of only a clip.
Inference: EfficientVIS simply takes output queries from the last clip as the initial queries to correlate and segment the same instances in the current clip without data association. Meanwhile, the queries are continuously updated by our query-video interaction to collect the current clip information so as to achieve a continuous tracking for the future. If a tracklet query has been classified as background for many clips, one may re-initialize it with in order to allocate slots. Since instance positions vary in different clips, we use the same initial proposals for every clip, i.e. the trained model parameters rather than the last clip output proposals.
Experiments
Datasets: Our experiments are conducted on the YouTube-VIS benchmark in two versions. YouTube-VIS 2019 is the first dataset for video instance segmentation, which contains 2,883 videos labeled at 6 FPS, 131K instance masks, and 40 object classes. YouTube-VIS 2021 is an increased version that contains 3859 videos.
Evaluation Metric: The video-level average precision (AP) and average recall (AR) are the metrics. Different from the image domain, the video intersection over union (IoU) is computed between the predicted tracklet mask and ground truth tracklet mask. So, in order to achieve high performance, a model needs to not only correctly classify and segment the target instance but also accurately track the target over time.
Architecture Settings: Following , the default video clip length and the number of tracklet queries are set to and , respectively. Following , the number of iterations is and the default CNN backbone is ResNet-50. We use the above default settings for all experiments unless otherwise specified. Image frame size and pretrained model are the same as .
Training: EfficientVIS is trained by AdamW with an initial learning rate of . We train the model for epochs where the learning rate is dropped by a factor of at the -th epoch. For example, the model can be trained in 12 hours with 4 RTX 3090 GPUs on YouTube-VIS 2019. The correspondence learning is enabled if the maximum frame number of videos in a dataset is larger than . No training data augmentation is used unless otherwise specified.
Fully end-to-end Inference: EfficientVIS does not include any data association or post-processing. Reason: Target tracklet within a clip is generated as a whole in a single forward-pass thanks to the implicit association of our FTSA (Sec. 3.1). Target tracklets between two clips are automatically correlated owing to our correspondence learning (Sec. 3.3). The one-to-one matching loss (Sec. 3.2) further enables our method to get rid of post-processing like NMS.
2 Ablation Studies
Time Disentangled Query: As described in Sec. 3.1, our tracklet query is designed to be disentangled in time, i.e. each query contains embeddings instead of one embedding. As shown in Table 1(f), we compare it with the time shared query scheme that only uses one single embedding for each tracklet query. Our time disentangled query achieves 1.5 AP higher than the time shared one. We argue that this is because objects in different time may have dramatic appearance changes like heavy occlusion or motion blur. It is more reasonable to use multiple embeddings to separately encode different appearance patterns of an instance.
Factorised Temporo-Spatial Self-Attention: To evaluate our factorised temporo-spatial self-attention (FTSA), we experiment with different self-attention schemes as shown in Table 1(a). “S” indicates the spatial self-attention that only performs multi-head self-attention (MSA) within the same frame across the embeddings of different tracklet queries. “T” denotes the temporal self-attention that only performs MSA within the same tracklet query across the embeddings of different time. “Joint T-S” is the joint temporo-spatial self-attention that performs MSA across all embeddings as described in Sec. 3.1. We see from the table that the FTSA and “Joint T-S” achieve better performance than using either “S” or “T” standalone. This is because the temporal self-attention can let a query in different time exchange information so as to associate the same instance, while the spatial self-attention models different instances relations in space. Such a query communication in both space and time is necessary for the VIS task. Compared to “Joint T-S”, our FTSA yields better VIS results. We think the reason is that all attendees in FTSA are strongly related to each other. For example, the embeddings in temporal self-attention are all from the same tracklet query, while those in spatial self-attention are all from the same frame. However, by mixing up all embeddings in a single self-attention pass, much less relevant information from other tracklets at far temporal positions is directly involved during “Joint T-S”.
Temporal Dynamic Convolution: To assess our temporal dynamic convolution (TDC), we compare it with the still-image dynamic convolution that performs 2D convolution on the current frame only. As shown in Table 1(b), our temporal dynamic convolution improves 1 AP over the still-image version. It suggests that the temporal object context from nearby frames is also an informative cue. Moreover, we visualize the tracklet feature. As shown in Fig. 4, tracklet feature is highly activated by our TDC in the region where the target instance appears. We also see that tracklet feature is suppressed when the target instance disappears, even if its tracklet proposal has drifted to background clutter or uncorrelated instance. This demonstrates the dynamic filter conditioned on tracklet query is target-specific, which exclusively collects the query’s target information from the video base feature yet discards irrelevant information.
Video Clip Length: As shown in Table 1(c), we experiment with different video clip lengths. EfficientVIS improves 1.7 AP by increasing the number of frames from 9 to 36. It is reasonable because more frames can provide richer temporal context which is important to video tasks. Moreover, since our method runs clip-by-clip in 36 FPS as shown in Table 2, it can be regarded as a near-online fashion with a delay of . For example, there is a 0.25s delay when .
Correspondence Learning: To validate the effectiveness of our correspondence learning, we train an EfficientVIS without the correspondence learning. Concretely, we treat every clip as an individual video and do not use tracklet queries from other clips as input during training. In inference, we keep our fully end-to-end paradigm, i.e. the output tracklet queries from one clip are used as the initial queries of the next clip. As shown in Table 1(e), the performance significantly drops if EfficientVIS is trained without the correspondence learning. It suggests that our correspondence learning is the key to enabling tracklet query to be a video-level instance representation that can be shared among different video clips.
Fully End-to-end Learnable Framework: Thanks to the correspondence learning, EfficientVIS fully end-to-end performs VIS by taking tracklet queries from one clip as the initial queries of the next clip. To evaluate this fully end-to-end framework, we compare it with two partially end-to-end schemes, “hand-craft” and VisTR in Table 1(d). These two methods perform VIS end-to-end within each clip but require a hand-crafted data association to link tracklets between clips. For the “hand-craft” scheme, all the model architectures keep the same as EfficientVIS except that tracklets between two clips are linked by a human-designed rule: 1) We first construct an affinity matrix by taking into account box IoU, box L1 distance, mask IoU, and class label matching like prior tracking works ; 2) We then solve the tracklet association by the Hungarian algorithm. As shown in Table 1(d), our fully end-to-end scheme performs better than the partially end-to-end framework in different settings while greatly simplifying the previous VIS paradigms. The reason for this performance gap is that the tracklet linking between clips in hand-crafted scheme tends to be sub-optimal, while that in our fully end-to-end framework is learned and optimized by ground truths during training.
Tracking in Low Frame Rate Videos: Tracking object instances in low frame rate videos is a hard problem, where motion-based trackers usually fail. Motion-based trackers suppose instances move smoothly over time and impose spatial movement constraints to prevent trackers from associating very distant instances. In contrast, our tracking is determined by instance appearance rather than motion, as we retrieve a target instance solely conditioned on its query representation regardless of the target spatial distance among different frames. Therefore, our tracking is not limited by dramatic instance movements. To demonstrate this property, we downsample the frame rate of the YouTube-VIS dataset to 1.5 FPS and test our method. As shown in Table 1(h), EfficientVIS successfully maintains its VIS performance in the low frame rate video setting. As shown in Fig. 5, we also visualize the results of EfficientVIS on a very low frame rate video whose frame is sampled every 5 seconds, i.e. 0.2 FPS.
Fast Convergence: In our method, we crop video using tracklet proposal and collect target information from the proposal. Compared to VIS transformer (VisTR) that uses frame-wise dense attention, this RoI-wise design eliminates much background clutter and redundancies in videos and enforces EfficientVIS to focus more on informative regions, making our convergence 15 faster as shown in Table 1(g). This RoI-wise pipeline also leads to better accuracy than VisTR. This is because tracklet proposal can avoid regions outside the target being segmented (second figure in appendix), which results in more precise masks and significantly higher AP75 as evidenced in Table 1(g).
3 Comparison to State of the Art
YouTube-VIS 2019: In Table 2, we compare EfficientVIS with the real-time state-of-the-art VIS models. EfficientVIS is the only fully end-to-end framework while achieving superior accuracy. We attribute the strong performance to two main aspects: 1) Our object tracking is learned via the end-to-end framework; 2) Our clip-by-clip processing and temporal dynamic convolution enrich temporal object context that is helpful for handling occlusion or motion blur. As shown in Table 2, we achieve notably higher recall (AR) than other competitors. We think the reason for the high recall is that we recognize more heavily occluded or blurred objects, which are usually missed by common methods.
For a fair comparison, Table 2 only lists real-time methods using single model without extra data. There are two works with 40+ AP, MaskProp (5.6 FPS) and Propose-Reduce (1.8 FPS), which however are not real-time and adopt multiple models or extra training data. MaskProp uses DCN backbone , HTC detector , extra High-Resolution Mask Refinement network, etc. Propose-Reduce is trained with additional DAVIS-UVOS and COCO pseudo videos datasets. These elaborated implementations are beyond the scope of our work, as our main goal is to present a simple end-to-end model with real-time inference and fast training.
YouTube-VIS 2021: We experiment with EfficientVIS on YouTube-VIS 2021 in Table 3. EfficientVIS achieves state-of-the-art AP without bells and whistles. Similar to the findings on YouTube-VIS 2019, EfficientVIS significantly improves AR over other state of the arts.
Conclusion
This work presents a new VIS model, EfficientVIS, which simultaneously classifies, segments, and tracks multiple object instances in a single end-to-end pass and clip-by-clip fashion. EfficientVIS adopts tracklet query and proposal to respectively represent instance appearance and position in videos. An efficient query-video interaction is proposed for associating and segmenting instances in each clip. A correspondence learning is designed to correlate instance tracklets between clips without data association. The above designs enable a fully end-to-end framework achieving state-of-the-art VIS accuracy with fast training and inference.
Acknowledgement. This paper is supported in part by a gift grant from Amazon Go and National Science Foundation Grant CNS1951952.
References
Appendix A Qualitative Results on YouTube-VIS
As shown in Fig. 6, we visualize our video instance segmentation results on the YouTube-VIS 2019 dataset. We see from the figure that EfficientVIS can well recognize heavily occluded object instances. The reason is that the target instance information is propagated back and forth in a clip thanks to our clip-by-clip processing and temporal dynamic convolution. In this way, the non-occluded instance appearances from nearby frames are propagated to provide strong cues to recognize those heavily occluded instances. We also see in the figure that our tracklet proposal can successfully track target instance even its shape dramatically changes over time. This is because tracklet proposal is regressed conditioned on the target query representation, and we do not impose space-time constraints or smoothness. Thus, tracklet proposal is not limited by the target positions or shapes in nearby frames.
Appendix B Qualitative Comparison
As shown in Fig. 7, we compare EfficientVIS with the VIS transformer (VisTR) . Since VisTR produces an instance mask by segmenting the whole frame, one instance mask may easily contaminate other instances or background regions as shown in the figure. This suggests that it is hard to enforce the query representation in VisTR to be very discriminative to distinguish target object instance from the whole scene. In contrast to this frame-wise scheme, EfficientVIS filters out many irrelevant instances and regions by our tracklet proposals. Our method only needs to enforce tracklet query to distinguish target instance from the proposal region, which is easier for the model to achieve.