Locality-Aware Inter-and Intra-Video Reconstruction for Self-Supervised Correspondence Learning
Liulei Li, Tianfei Zhou, Wenguan Wang, Lu Yang, Jianwu Li, Yi Yang
Introduction
As a fundamental problem in computer vision, correspondence matching facilitates many applications, such as scene understanding , object dynamics modeling , and 3D reconstruction . However, supervising representation for visual correspondence is not trivial, as obtaining pixel-level manual annotations is costly, and sometimes even prohibi- tive (due to occlusions and free-form object deformations). Although synthetic data would serve an alternative in some low-level visual correspondence tasks (e.g., optical flow estimation ), they limit the generalization to real scenes.
Using natural videos as a source of free supervision, i.e., self-supervised temporal correspondence learning, is consi- dered as appealing . This is because videos contain rich realistic appearance and shape variations with almost infinite supply, and deliver valuable supervisory signals from the intrinsic coherence, i.e., correlations among frames.
Along this direction, existing solutions are typically built upon a reconstruction scheme (i.e., each pixel from a ‘query’ frame is reconstructed by finding and assembling relevant pixels from adjacent frame(s)) , and/or adopt a cycle-consistent tracking paradigm (i.e., pixels/patches are encouraged to fall into the same location after one cycle of forward and backward tracking) .
Unfortunately, these successful approaches largely neglect three crucial abilities for robust temporal correspondence learning, namely instance discrimination, location awareness, and spatial compactness. First, many of them share a narrow view that only considers intra-video context for correspondence learning. As it is hard to derive a free signal from a single video for identifying different object instances, the learned features are inevitably less instance-discriminative. Second, existing methods are typically built without explicit position representation. Such design seems counter-intuitive, given the extensive evidence that spatial position is encoded in human visual system and plays a vital role when human track objects . Third, as the visual world is continuous and smoothly-varying, both spatial and temporal coherence naturally exist in videos. While numerous strategies are raised to address smoothness on the time axis, far less attention has been paid to the spatial case.
To fill in these three missing pieces to the puzzle of self- supervised correspondence learning, we present a locality- aware inter-and intra-video reconstruction framework – Liir. First, we augment existing intra-video analysis based correspondence learning strategy with inter-video context, which is informative for instance-level separation. This leads to an inter-and intra-video reconstruction based training objective, that inspires intra-video positive correspondence matching, but penalizes unreliable pixel associations within and cross videos. We empirically verify that our inter-and intra-video reconstruction strategy can yield more discriminative features, that encode higher-level semantics beyond low-level intra-instance invariance modeled by previous algorithms. Second, to make our Liir more location-sensitive, we learn to encode position information into the representation. Al- though position bias is favored for intra-video correspon- dence matching, it is undesired in the inter-video case. We thus devise a position shifting strategy to foster the strength and circumvent the weaknesses of position encoding. We ex- perimentally show that, explicit position embedding benefits correspondence matching. Third, we involve a spatial compactness prior in intra-video pixel-wise affinity estimation, resulting in sparse yet compact associations. For each query pixel, the distribution of related pixels is fit by a mixture of Gaussians. This enforces each query pixel to match only a handful of spatially close pixels in an adjacent frame. Our experiments show that such compactness prior not only regularizes training, but also removes outliers during inference.
These three contributions together make Liir a powerful framework for self-supervised correspondence learning. Without any adaptation, the learned representation is effec- tive for various correspondence-related tasks, i.e., video object segmentation, semantic part propagation, pose tracking. On these tasks, Liir consistently outperforms unsupervised state-of-the-arts and is comparable to, or even better than, some task-specific fully-supervised methods (e.g., Fig. 1).
Related Work
Self-Supervised Temporal Correspondence Learning. In the video domain, correspondence matching plays a central role in many tasks (e.g., video segmentation , flow esti- mation and object tracking ). An emerging line of work tackles this problem in a self-supervised learning paradigm, by exploiting the temporal coherence in videos. One may group these work into two major classes. The first class of methods poses a colorization proxy task (Fig. 2(a)), i.e., reconstruct a query frame from an adjacent frame, according to their correspondence. The latter type of methods performs forward and backward tracking and penalizes the inconsistency between the start and end positions of the tracked pixels or regions (Fig. 2(b)). The basic idea – cycle-consistency – is also adopted in un- supervised tracking , optical flow and depth estimation. Though impressive, these methods miss three key elements for robust correspondence matching: instance discrimination, location awareness, and spatial compactness. In response, Liir is equipped with three specific modules (Fig. 2(c)). First, for instance discriminative repre- sentation learning, it adopts an inter-and intra-video recons- truction scheme that formulates contrast over inter-and intra- video affinities. Second, it involves position encoding into representation learning. Third, it imposes a spatial compact- ness prior to both correspondence learning and inference.
We are not the first to explore inter-video context. In , Lu et al. raise an unsupervised learning objective, i.e., discriminate between a set of surrogate video classes, but they compute this over video-level embeddings, not sufficient for pixel-wise correspondence learning. In , Wang et al. also simultaneously consider inter-and intra-video representation associations, but requiring pre-aligned patch pairs. In addition, they use three loss terms to address the desired intra-inter constraints, which are, however, formulated as a single unified training objective in our case. Moreover, we take a further step by addressing location awareness and spatial compactness, instead inter-video context only. In , Xu et al. revisit the idea of image-level similarity learning and consider frame pairs from the same videos are positive samples and pairs from different videos are negative, however, they find negative samples (i.e., inter-video context) hurt the performance of their model. In contrast, we formulate inter-and intra-video context through a unified, pixel-wise affinity framework that boosts instance-level discrimination without sacrificing intra-instance invariance. Our results suggest that how to make a good use of negative samples for temporal correspondence learning is still an intriguing question.
Self-Supervised Video Representation Learning. Corres- pondence learning approaches that use unlabeled video data fall in a broad field of self-supervised video representation learning. Towards learning transferable video representation, diverse pretext tasks are proposed to explore different intrinsic properties of videos as free supervisory signals, including temporal sequence ordering , predicting motion patterns , solving space-time cubic puzzles , anticipating future representations , and temporally aligning videos . The learned representations are compact video descriptors that can be generalized to va- rious downstream tasks (e.g., action recognition , video captioning , video retrieval ), while in this work we specifically focus on learning fine-grained visual features for pixel-level correspondence matching.
Self-Supervised Image Representation Learning. Basically, self-supervised approaches for image representation learning share a similar key idea with the counterparts for videos – design pretext tasks so as to mine supervision signals from the inherent information inside images . Examples of such tasks include estimating spatial context , predicting image rotations , solving jigsaw puzzles , among many others . Recent efforts were mainly devoted to improving deep metric learning techniques in large scale, i.e., adopt an instance discrimination pretext task where the cross-entropy objective is used to discriminate each image from other different images (i.e., negative sam- ples) . In this work, we assimilate the idea of contrasting positive samples (intra-video affinities) over numerous negative ones (inter-video affinities) for robust temporal correspondence matching, and formulate this within a unified inter-and intra-video reconstruction framework.
Video Mask Propagation. This task, a.k.a. semi-automatic video segmentation, aims at propagating first-frame object masks over the whole video sequence . It addresses classic automatic video segmentation techniques’ lack of flexibility in defining target objects . Depending on how they make use of the first-frame supervision, existing mask propagation models can be classified into three groups: i) online fine-tuning based methods that fine-tune a generic segmentation network with masks during inference ; and ii) matching based methods that directly condition the segmentation network on the first-frame masks and/or previous segments . As shown in Fig. 1, extensive annotations are typically required to train such systems, i.e., first pre-train on ImageNet and then fine-tune with COCO , DAVIS , Youtube-VOS , etc. In contrast, we pursue a more annotation-efficient solution; similar to previous self-supervised correspondence learning methods , Liir is trained using unlabeled videos only. Once trained, it can be directly applied for mask propagation, without adaptation.
Our Approach
We present Liir, a self-supervised framework that learns dense correspondence from raw videos. Before elaborating on our model design (cf. §3.2), we first review the classic reconstruction based temporal correspondence learning stra- tegy (cf. §3.1), which serves as the basis of our Liir.
Due to the appearance continuity in video, one can con- sider pixels in a ‘query’ frame as being copied from some locations of other ‘reference’ frames. In light of this, a few studies raise a reconstruction-based correspondence learning scheme: each query pixel struggles to find pixels in a reference frame that can best reconstruct itself.
where refers to -th element in , signifying the similarity between pixel in and pixel in , and ‘’ stands for the dot product. In this way, gives the strength of all the pixel pair-wise correspondence between and , according to which pixel in can be reconstructed by a weighted sum of pixels in :
The training objective of is hence defined as a reconstruction loss:
In practice, to avoid trivial solutions caused by information leakage, an information bottleneck is adopted over training samples, e.g., RGB2gray operation , channel-wise drop- out in RGB or Lab colorspace. After training, the representation encoder is used for correspondence matching: similar to Eq. 2, the affinity is estimated and used to propagate desired pixel-level entities (e.g., instance masks, key-point maps), from a reference frame to a query frame.
2 Liir: Locality-Awareabsent{}_{\!} Inter-andabsent{}_{\!} Intra-Videoabsent{}_{\!} Reconstruction Framework
Building upon the reconstruction-by-copy scheme, Liir is empowered with three crucial yet long overlooked abilities for robust correspondence learning: instance discrimination, location awareness, and spatial compactness.
Inter-and Intra-Video Reconstruction. With the computa- tion of the intra-video affinity (Eq. 1), each query pixel is forced to distinguish its counterpart (positive) reference pixels from unrelated (negative) ones within a same video, with the indictor of the reconstruction quality (Eqs. 2-3). As both the positive and negative samples are sourced from the same video, there is less evidence for distinction among similar object instances with intra-video appearance only (Fig. 3(a)). As one single video only contains limited content, conducting correspondence matching within videos is less challenging, and inevitably hinders the discrimination potential of the learned representation . These insights motivate us to improve the intra-video affinity based reconstruction scheme by further accounting for negative across-video correspondence. Concretely, given the query () and reference () frames from the same video, an intra-inter video affinity is computed:
where refer to a collection of frames, which are sampled from the whole training dataset, except the source video of (). By additionally considering other irrelevant videos during affinity computation, both the quantity and diversity of negative samples are greatly improved, allowing us to derive a more challenging inter-and intra-video reconstruction scheme (Fig. 3(b)):
With Eqs. 4-5, each pixel in the query frame is required to distinguish its counterpart pixels from massive unrelated ones, which are from not only the reference frame in current video, but a huge amount of irrelevant frames in other videos. This powerful idea, yet, is elegantly achieved by the same training objective as in Eq. 3. Note that Eq. 4 normalizes intra-video correspondence over both inter-and intra-video pixel-to-pixel relevance, while Eq. 5 only uses the pixels in the reference frame for reconstruction. The rationale here is that, even if the encoder wrongly matches a query pixel with a negative but similar-looking pixel in , i.e., will be large and , the synthesized color will be still very different to and will receive a large gradient from Eq. 3. Thus is driven to mine more high-level semantics and context-related clues, hence reinforcing the instance-level discrimination ability (Fig. 3(c)). Fig. 3(d) shows that the representation learned with our inter-and intra-video reconstruction strategy can distinguish similar-looking dogs nearby.
Although also addresses inter-video analysis based reconstruction, it conducts embedding association on patch level, relying on a pre-trained tracker for patch alignment. Besides, the method consumes three loss terms for supervision, which is much more complicated than ours. Further, it separately conducts inter-and intra-video affinity based reconstruction. This is problematic; when both the reference frame in current video and an irrelevant frame from other videos contain query-like pixels, there is no explicit supervision signal to determine which one should be matched.
Position Encoding and Position Shifting. Plenty of literature in neuroscience has revealed that human visual system encodes both appearance and position information when we perceive and track objects . Yet existing unsuper- vised correspondence methods put all focus on improving appearance based representation by ConvNets, ignoring the value of position information. Although suggest that ConvNets can implicitly capture position information by uti- lizing image boundary effects, explicit position encoding has already been a core of full attention networks (e.g., Trans- former ), and facilitated a variety of tasks (e.g., instance segmentation , tracking , video segmentation ). All these indicate that position encoding deserves more attention in the field of temporal correspondence learning.
where is added with the output feature of the first conv layer of and has the same size and dimension as the conv feature. We explore three position encoding strategies:
2D Sinusoidal Position Embedding (2DSPE): is given with a family of pre-defined sinusoidal functions, without introducing new trainable parameters:
where , specify the horizontal and vertical positions, specify the dimension, and . The horizontal (vertical) positions are encoded in the first (second) half of the dimensions. 2DSPE naturally handles resolutions that are unseen during training.
1D Absolute Position Embedding (1DAPE): 1DAPE is the most heavy-weight strategy: the whole is a learnable parameter matrix without any constraint.
With our intra-inter video affinity (Eq. 4), exploiting position information in intra-video correspondence matching, i.e., , addresses local continuity resides in videos. However, for inter-video pixel relevance computation, i.e., , such position prior is undesirable, as it inspires the query pixel in to prefer matching these pixels with similar positions in other irrelevant videos . To eliminate such position encoding induced bias from inter-video correspondence matching, we design a position shifting strategy (Fig. 4(a)). During training, for from other videos, we circularly shift the position encoding vectors in by a random step in horizontal and vertical axes, respectively. The reason why we adopt random shifting with circular boundary conditions, instead of random shuffling, is to preserve the spatial layout in the modulated position encoding map . Then and are fed into for inter-video correspondence matching, and related gradients are abandoned if learnable 1DAPE or 2DAPE is used. Note that the standard position encoding is applied for the query () and reference () frames and updated normally. Fig. 4(b) intuitively shows that merging position information into visual representation can enable robust correspondence matching even with confusing background and fast motion. In §4.4, we will quantitatively verify that 1DAPE is more favored and indeed boosts the performance.
Spatial Compactness Prior. As the visual world is conti- nuous and smoothly-varying, it is reasonable to assume appearances in video data change smoothly both in spatial and temporal dimensions. For correspondence learning, the tem- poral coherence has been extensively studied, while the spatial continuity received far less attention. To reduce search region, some existing methods restrict corres- pondence matching within a local window, considering spatial regularities in a simple way. To make a better use of the spatial continuity, we augment the original reconstruction objective with an additional prior, termed as spatial compa- ctness. Such a prior poses constraints on the spatial distribution of associated pixels, leading to sparse and coherent solutions. Specifically, given the query () and reference () frames, we expect i) each query pixel to be only matched with a small number of reference pixels, and ii) the matched reference pixels to be clustered. For a query pixel and its matching ‘heatmap’: , w.r.t. , we assume follows a mixture of 2D Gaussian distributions:
3 Implementation Details
Network Configuration. For fair comparison, our feature encoder is implemented as ResNet-18 as in . Following , 2 downsampling is only made in the third residual block. Thus finally outputs 256 feature maps of 1/4 size of the input, i.e., . The position embedding is added to feature after the first 77 Conv-BN-ReLU layer, i.e., .
Training: Liir is trained from scratch on two NVIDIA RTX- GPUs and only uses the raw videos from Youtube-VOS . Each training image is resized into 256256 and channel-wise dropout in Lab colorspace is adopted as the information bottleneck. Adam optimizer is used. At the initial epochs, only intra-video reconstruction learning is adopted for warm-up, with a learning rate of and batch size of . Then we conduct inter-video reconstruction learning with spatial compactness based regularization at the next epochs, with a learning rate of and batch size of . We online maintain a memory bank of 1,440 frames from different videos. For each query pixel, we sample 4 feature points from each stored frame, i.e., a total of 1,4404 negative samples are used for the inter-video correspondence computation, and we employ the moving average strategy for parameter updating.
Experiment
We evaluate the learned representation on diverse video label propagation tasks, i.e., video object segmentation (§4.1), body part propagation (§4.2), and pose keypoint tracking (§4.3). As in conventions , all these tasks are to propagate the first frame annotation to the whole video sequence, and we use our model to compute inter-frame dense correspondences. In §4.4, we conduct a set of ablative studies to examine the efficacy of our essential model designs.
Dataset. We first test our method on val sets of two popular video object segmentation datasets, i.e., DAVIS17 and YouTube-VOS . There are and videos in DAVIS17 and YouTube-VOS val sets, respectively.
Evaluation Metric. Following the official protocol , we use the region similarity () and contour accuracy () as the evaluation metrics. Note that the scores on YouTube-VOS are respectively reported for seen and unseen categories, obtained from the official evaluation server.
Performance on DAVIS17: As illustrated in Table 1, our Liir consistently outperforms all existing self-supervised methods across all the evaluation metrics. For example, it surpasses current best-performing self-supervised method, i.e., CLTC , in terms of mean (72.1 vs. 70.3). In addition, even without using any manual annotations for training, Liir achieves very competitive segmentation performance in comparison with some famous supervised models trained with massive pixel-wise annotations.
Performance on YouTube-VOS val. Table 2 reports performance comparison of Liir against four self-supervised competitors on YouTube-VOS val. It can be observed that Liir sets new state-of-the-art. In particular, Liir yields an overall score of 69.3%, surpassing the second-best (i.e., CLTC ) and third-best (i.e., MAST ) approaches by 2.0% and 5.1%, respectively. Further, Liir even outperforms some famous supervised methods (i.e., OSVOS and PreMVOS ), especially for the unseen categories, clearly demonstrating its remarkable generalization ability.
Qualitative Results. Fig. 6 depicts visual results on representative videos in the datasets. As seen, Liir is able to establish accurate correspondences under various challenging scenarios, e.g., scale changes, small objects and occlusions.
2 Results for Body Part Propagation
Dataset. We next evaluate our model performance for body part propagation. Experiments are conducted on VIP val , which contains 50 videos with annotations of human semantic part categories (e.g., hair, face, dress).
Evaluation Metric. As suggested by VIP , we adopt mean intersection-over-union (mIoU) and mean Average Precision (mAP) metrics for evaluation of semantic-level and instance-level parsing, respectively.
Performance. As shown in Table 3, Liir achieves state-of-the-art performance on both semantic-level and instance-level parsing. This indicates that Liir can generate strong representation which models both cross-instance discrimination and intra-instance invariance well. Fig. 7 depicts some visualization results on two representative videos. Liir achieves temporally stable results and shows robustness to typical challenges (e.g., pose variations, occlusions).
3 Results for Pose Keypoint Tracking
Dataset. We then examine the model performance on human keypoint tracking, using JHMDB val. JHMDB val has videos. For each person, a total of body joints, e.g., torso, head, shoulder, elbow, are annotated.
Evaluation Metric. We use probability of correct keypoint (PCK) to measure the accuracy between each tracking result and corresponding ground-truth with a threshold .
Performance. Table 3 shows that Liir exhibits compelling overall performance. Note that CLTC uses different checkpoints and model architectures for different tasks and datasets, while we only use a single model for evaluation. The visual results in Fig. 7 also demonstrate the strong capability of Liir in establishing precise correspondence.
4 Diagnostic Experiment
For further detailed analysis, we conduct a series of ablative studies on DAVIS17 val and VIP val sets.
Key Component Analysis. We first examine the efficacy of essential components of Liir, i.e., inter-and intra-video reconstruction, position encoding, and spatial compactness. The results are summarized in Table A.1, where position encoding is implemented as 1DAPE, and compactness prior is used during both training and inference stages. When separately comparing row #2 - #4 with the baseline (MAST ) in row #1, we can observe that each individual module in- deed boosts the performance. For example, on DAVIS17 val, inter-and intra-video reconstruction, position encoding, and spatial compactness prior respectively bring 3.4%, 1.6%, and 3.1% gains. This verifies our core insight that these three elements are crucial for correspondence learning. Finally, in row #5, we combine all the three components together – Liir, and obtain the best performance. This suggests that these modules are complementary to each other, and confirms the effectiveness of our whole design.
Inter-and Intra-Video Reconstruction. We next study the impact of increasing the number of negative samples, i.e., frames from other irrelevant videos used for inter-video correspondence computation (Eq. 4). In Table 5a, row #1 gives scores of learning without considering inter-video correspondence. In this case, the results are unsatisfactory. When more negative samples are involved (i.e., 01,440), better performance can be achieved (i.e., 69.272.1 on DAVIS17 val, 38.441.2 on VIP val). Finally we use 1,440 negative samples for inter-video reconstruction based learning, which is the maximum number allowed by our GPU.
Position Encoding. To determine the effect of our position encoding module, we then report the performance with different encoding strategies in Table 5b. As seen, the non-learnable strategy, 2DSPE, hinders the performance, while the learnable alternatives, i.e., 1DAPE and 2DAPE, lead to better results. Compared with 2DAPE, 1DAPE is more favored, probably due to its high flexibility and capacity.
Position Shifting. We further study the influence of our position shifting strategy on performance. As shown in Table 5c, we consider two alternatives, i.e., ‘NAN’ and ‘position shuffling’. ‘NAN’ refers to using the normal position encoding map without any modulation during inter-video correspondence matching. Compared with ‘position shifting’, ‘NAN’ suffers from performance degradation (i.e., 72.171.3 on DAVIS17 val, 41.240.6 on VIP val), showing the negative effect of the position-induced bias. The other baseline, ‘position shuffling’, i.e., randomly shuffling the position encoding map for inter-video affinity computation, though better than ‘NAN’, is still worse than ‘position shifting’. This is because it destroys the spatial layouts.
Spatial Compactness Prior. The spatial compactness prior (Eq. 7) is used to regularize intra-video correspondence matching during both training and inference stages. In Table 5d, we quantitatively identify the performance contribution of our spatial compactness prior in different stages.
Conclusion
We presented a self-supervised temporal correspondence learning approach, Liir, that makes contributions in three aspects. First, going beyond the popular intra-video analysis based learning scheme, we further enforce separation between intra- and inter-video pixel associations, enhancing instance-level feature discrimination. Second, with a clever position shifting strategy, we bring the advantages of position encoding into full play, while avoiding its undesirable impact at the same time. Third, a spatial compactness prior is introduced to regularize representation learning and improve correspondence inference. The effectiveness was thoroughly validated over various label propagation tasks.
References
Appendix A Analysis of Spatial Compactness Prior
We use 2D Gaussian distributions to approximate the affinity matrix between two frames. Table A.1 provides a detailed analysis of the hyper-parameter . The st row corresponds to a baseline model that disregards the spatial compactness prior in both training and inference phases. We see from the table that 1) when in training, the model performs worse than the baseline model, as such a rigorous matching constraint easily leads to overconfident pre-dictions; 2) when (in training) becomes larger, the performance greatly improves; 3) the models trained with shows consistently better performance than those trained with ; and 4) in the inference stage, always leads to the best performance. Accordingly, we set to in both training and inference stages.
Appendix B Analysis of Long-Term Dependence
Liir leverages multiple reference frames during testing as in . We analysis the impact of long-term dependence in Table B.1. As seen, our method is more robust and shows smaller drop wrt : 4.7% vs 6.8%.
Appendix C Analysis of Feature Point Sampling
We give the ablative study on the number of feature points sampled during inter-video reconstruction in Table C.1. It can be seen that, with 1440 frames, better results can be achieved if we sample more points per frame (70.4 71.472.1), supporting our claim about instance separation. With the same number of total sampled points, we gain better performance if we consider more frames (70.9 71.772.1). It is reasonable as more frames can provide much rich/challenging context. With our limited GPU capacity, we choose to sample 1440 frames and four feature points per frame, so as to maximize the performance. But we can speculate that, if with enough GPU capacity, sampling more features points from more videos will further improve the performance.
Appendix D Visualization of Ablation Study
Fig. D.1 depicts visual effect of each essential component in Liir. Starting from the baseline model (b), we progressively add position encoding (c), inter-video reconstruction (d) and spatial compactness (e). As seen, with explicit positional encoding, our model is able to heavily suppress background regions (e.g., shelves in the first row). The inter-video reconstruction enables more accurate discrimination between targets and semantically similar distractors (e.g., “motorcycle” in the third row). Last, incorporating the spatial compactness prior facilitates more precise correspondences, leading to high-quality segmentation results.
Appendix E Additional Qualitative Results
We provide additional video propagation results on four datasets, including DAVIS17 val in Fig. D.2, Youtube-VOS val in Fig. D.3, VIP val in Fig. D.4 and JHMDB val in Fig. D.5. We observe that even training with no annotations, Liir is able to produce highly exquisite results.
Appendix F Limitation
Although Liir demonstrates remarkable performance and high generalizability in correspondence matching, we still see a large performance gap between Liir and current top-leading supervised models (e.g., STM in VOS). However, as a self-supervised method, Liir can be easily scaled to leverage any available collection of video data for training. This could lead to more accurate correspondence learning from massive unlabeled data instead of using small-scale datasets only (e.g., DAVIS, YouTube-VOS). Apart from that, inter-video reconstruction spares massive space to bank the negative samples, this coerces the memory size of GPUs and extends the training time. Fortunately, the performance of Liir has improved by leaps and bounds to compensate for it, for instance, up to 2.9% on DAVIS17 . Further more, note that the inter-video reconstruction is only applied during training, thus it does not introduce extra computational cost at inference. In our future work, we will explore towards the above direction to narrow the performance gap between supervised methods, and find the way to reduce the memory space taken up by negative videos and speed up training simultaneously.