Locality-Aware Inter-and Intra-Video Reconstruction for Self-Supervised Correspondence Learning

Liulei Li, Tianfei Zhou, Wenguan Wang, Lu Yang, Jianwu Li, Yi Yang

Introduction

 ⁣{}_{\!}As ⁣{}_{\!} a ⁣{}_{\!} fundamental ⁣{}_{\!} problem ⁣{}_{\!} in ⁣{}_{\!} computer ⁣{}_{\!} vision, ⁣{}_{\!} correspondence matching facilitates many applications, such as ⁣{}_{\!} scene ⁣{}_{\!} understanding ⁣ ⁣{}_{\!\!} , ⁣ ⁣{}_{\!\!} object ⁣ ⁣{}_{\!\!} dynamics ⁣{}_{\!} modeling ⁣ ⁣{}_{\!\!} , ⁣ ⁣{}_{\!\!} and ⁣{}_{\!} 3D reconstruction ⁣{}_{\!} . ⁣{}_{\!} However, ⁣{}_{\!} supervising ⁣{}_{\!} representation ⁣{}_{\!} for visual ⁣{}_{\!} correspondence ⁣{}_{\!} is ⁣{}_{\!} not ⁣{}_{\!} trivial, ⁣{}_{\!} as obtaining pixel-level manual annotations is costly, and sometimes even prohibi- tive ⁣{}_{\!} (due ⁣{}_{\!} to occlusions and free-form object deformations). Although synthetic data would serve an alternative in some low-level visual correspondence tasks (e.g., optical flow ⁣{}_{\!} estimation ), they limit the generalization to real scenes.

Using ⁣{}_{\!} natural ⁣{}_{\!} videos ⁣{}_{\!} as ⁣{}_{\!} a ⁣{}_{\!} source ⁣{}_{\!} of ⁣{}_{\!} free ⁣{}_{\!} supervision, ⁣{}_{\!} i.e., self-supervised ⁣{}_{\!} temporal ⁣{}_{\!} correspondence ⁣{}_{\!} learning, ⁣{}_{\!} is ⁣{}_{\!} consi- dered ⁣{}_{\!} as ⁣{}_{\!} appealing ⁣{}_{\!} . This is because videos ⁣{}_{\!} contain ⁣{}_{\!} rich realistic appearance and shape variations with almost infinite supply, and deliver valuable supervisory signals from the intrinsic coherence, i.e., correlations among frames.

Along this direction, existing solutions are typically built upon a reconstruction scheme (i.e., ⁣{}_{\!} each ⁣{}_{\!} pixel ⁣{}_{\!} from ⁣{}_{\!} a ⁣{}_{\!} ‘query’ frame is reconstructed by finding and assembling relevant pixels from adjacent frame(s)) , and/or adopt a cycle-consistent tracking paradigm (i.e., pixels/patches are encouraged to fall into the same location after one cycle of forward and backward tracking) .

Unfortunately, these successful approaches largely neglect three crucial abilities for robust temporal correspondence learning, namely instance discrimination, location awareness, and spatial compactness. First, many of them share a narrow view that only considers intra-video context for correspondence learning. As it is hard to derive a free signal from a single video for identifying different object instances, the learned features are inevitably less instance-discriminative. Second, existing methods are typically built without explicit position representation. Such design seems counter-intuitive, given the extensive evidence that spatial position is encoded in human visual system and plays a vital role when human track objects . Third, as the visual world is continuous and smoothly-varying, both spatial and temporal coherence naturally exist in videos. While numerous strategies are raised to address smoothness on the time axis, far less attention has been paid to the spatial case.

To fill in these three missing pieces to the puzzle of self- supervised correspondence learning, we present a locality- aware ⁣{}_{\!} inter-and ⁣{}_{\!} intra-video ⁣{}_{\!} reconstruction ⁣{}_{\!} framework ⁣{}_{\!} – ⁣{}_{\!} Liir. ⁣ ⁣{}_{\!\!} First, we augment existing intra-video analysis based correspondence learning strategy with inter-video context, which is informative for instance-level separation. This leads to an inter-and ⁣{}_{\!} intra-video ⁣{}_{\!} reconstruction ⁣{}_{\!} based ⁣{}_{\!} training ⁣{}_{\!} objective, that inspires intra-video positive correspondence matching, but ⁣{}_{\!} penalizes ⁣{}_{\!} unreliable ⁣{}_{\!} pixel ⁣{}_{\!} associations ⁣{}_{\!} within ⁣{}_{\!} and ⁣{}_{\!} cross videos. We empirically verify that our inter-and intra-video reconstruction strategy can yield more discriminative features, that encode higher-level semantics beyond low-level intra-instance invariance modeled by previous algorithms. Second, to make our Liir more location-sensitive, we learn to encode position information into the representation. Al- though position bias is favored for intra-video correspon- dence matching, it is undesired in the inter-video case. We thus devise a position shifting strategy to foster the strength and ⁣{}_{\!} circumvent ⁣{}_{\!} the ⁣{}_{\!} weaknesses ⁣{}_{\!} of ⁣{}_{\!} position ⁣{}_{\!} encoding. ⁣{}_{\!} We ⁣{}_{\!} ex- perimentally ⁣{}_{\!} show ⁣{}_{\!} that, ⁣{}_{\!} explicit ⁣{}_{\!} position ⁣{}_{\!} embedding ⁣{}_{\!} benefits correspondence matching. Third, we involve a spatial compactness prior in intra-video pixel-wise affinity estimation, resulting in sparse yet compact associations. ⁣{}_{\!} For ⁣{}_{\!} each query pixel, the distribution of related pixels is fit by a mixture ⁣{}_{\!} of ⁣{}_{\!} Gaussians. ⁣{}_{\!} This ⁣{}_{\!} enforces ⁣{}_{\!} each ⁣{}_{\!} query ⁣{}_{\!} pixel ⁣{}_{\!} to ⁣{}_{\!} match ⁣{}_{\!} only ⁣{}_{\!} a handful of spatially close pixels in an adjacent frame. Our experiments show that such compactness prior not only regularizes training, but also removes outliers during inference.

These three contributions together make Liir a powerful framework for self-supervised correspondence learning. Without any adaptation, the learned representation is effec- tive ⁣{}_{\!} for ⁣{}_{\!} various ⁣{}_{\!} correspondence-related ⁣{}_{\!} tasks, ⁣{}_{\!} i.e., ⁣{}_{\!} video ⁣{}_{\!} object segmentation, semantic part propagation, pose tracking. On these tasks, Liir consistently outperforms unsupervised state-of-the-arts and is comparable to, or even better than, some task-specific fully-supervised methods (e.g., Fig. ⁣{}_{\!} 1).

Related Work

Self-Supervised Temporal Correspondence Learning. In the ⁣{}_{\!} video ⁣{}_{\!} domain, ⁣{}_{\!} correspondence ⁣{}_{\!} matching ⁣{}_{\!} plays ⁣{}_{\!} a ⁣{}_{\!} central role in many tasks (e.g., video segmentation ⁣{}_{\!} , flow esti- mation ⁣{}_{\!}  ⁣{}_{\!} and ⁣{}_{\!} object ⁣{}_{\!} tracking ⁣{}_{\!} ). ⁣{}_{\!} An ⁣{}_{\!} emerging ⁣{}_{\!} line of work tackles this problem in a self-supervised learning paradigm, by exploiting the temporal coherence in videos. One may group these work into two major classes. The ⁣{}_{\!} first ⁣{}_{\!} class ⁣{}_{\!} of ⁣{}_{\!} methods ⁣{}_{\!}  ⁣{}_{\!} poses ⁣{}_{\!} a ⁣{}_{\!} colorization proxy task (Fig. ⁣{}_{\!} 2(a)), i.e., reconstruct a query frame from an adjacent ⁣{}_{\!} frame, ⁣{}_{\!} according ⁣{}_{\!} to ⁣{}_{\!} their ⁣{}_{\!} correspondence. ⁣{}_{\!} The ⁣{}_{\!} latter type of methods ⁣{}_{\!}  ⁣{}_{\!} performs ⁣{}_{\!} forward ⁣{}_{\!} and ⁣{}_{\!} backward ⁣{}_{\!} tracking and penalizes the inconsistency between the start and end positions of the tracked ⁣{}_{\!} pixels ⁣{}_{\!} or ⁣{}_{\!} regions ⁣{}_{\!} (Fig. ⁣{}_{\!} 2(b)). The ⁣{}_{\!} basic ⁣{}_{\!} idea ⁣{}_{\!} – ⁣{}_{\!} cycle-consistency ⁣{}_{\!} – ⁣{}_{\!} is ⁣{}_{\!} also ⁣{}_{\!} adopted ⁣{}_{\!} in ⁣{}_{\!} un- supervised ⁣{}_{\!} tracking ⁣{}_{\!} , ⁣{}_{\!} optical ⁣{}_{\!} flow ⁣{}_{\!}  ⁣{}_{\!} and ⁣{}_{\!} depth ⁣{}_{\!} estimation ⁣  ⁣ ⁣{}_{\!~{}\!\!}. ⁣{}_{\!} Though impressive, ⁣{}_{\!} these methods miss three key elements for robust correspondence matching: instance discrimination, location awareness, and spatial ⁣{}_{\!} compactness. In response, Liir is equipped with three specific modules (Fig. ⁣{}_{\!} 2(c)). First, for instance discriminative repre- sentation learning, it adopts an inter-and intra-video recons- truction ⁣{}_{\!} scheme ⁣{}_{\!} that ⁣{}_{\!} formulates ⁣{}_{\!} contrast ⁣{}_{\!} over ⁣{}_{\!} inter-and ⁣{}_{\!} intra- video affinities. Second, it involves position encoding into representation ⁣{}_{\!} learning. ⁣{}_{\!} Third, ⁣{}_{\!} it ⁣{}_{\!} imposes ⁣{}_{\!} a ⁣{}_{\!} spatial ⁣{}_{\!} compact- ness prior to both correspondence learning and inference.

 ⁣{}_{\!}We ⁣{}_{\!} are ⁣{}_{\!} not ⁣{}_{\!} the ⁣{}_{\!} first ⁣{}_{\!} to ⁣{}_{\!} explore ⁣{}_{\!} inter-video ⁣{}_{\!} context. ⁣{}_{\!} In ⁣{}_{\!} , ⁣{}_{\!} Lu et al. raise an unsupervised learning objective, i.e., discriminate between a set of surrogate video classes, but they compute this over video-level embeddings, ⁣{}_{\!} not ⁣{}_{\!} sufficient ⁣{}_{\!} for pixel-wise correspondence ⁣{}_{\!} learning. ⁣{}_{\!} In ⁣{}_{\!} , Wang ⁣{}_{\!} et al.  ⁣{}_{\!} also simultaneously consider inter-and intra-video representation associations, but requiring pre-aligned patch pairs. In addition, ⁣{}_{\!} they ⁣{}_{\!} use ⁣{}_{\!} three ⁣{}_{\!} loss ⁣{}_{\!} terms ⁣{}_{\!} to ⁣{}_{\!} address ⁣{}_{\!} the ⁣{}_{\!} desired ⁣{}_{\!} intra-inter constraints, which are, however, formulated as a single unified ⁣{}_{\!} training ⁣{}_{\!} objective ⁣{}_{\!} in ⁣{}_{\!} our ⁣{}_{\!} case. ⁣{}_{\!} Moreover, ⁣{}_{\!} we ⁣{}_{\!} take a further step by addressing location awareness and spatial compactness, instead inter-video context only. In ⁣{}_{\!} , Xu et al. revisit the idea of image-level similarity learning and consider frame pairs from the same videos are positive samples and pairs from different videos are negative, however, they find negative samples (i.e., inter-video context) hurt the performance of their model. In contrast, we formulate inter-and intra-video context through a unified, pixel-wise affinity framework that boosts instance-level discrimination without sacrificing intra-instance invariance. Our results suggest that how to make a good use of negative samples for temporal correspondence learning is still an intriguing question.

 ⁣{}_{\!}Self-Supervised ⁣{}_{\!} Video ⁣{}_{\!} Representation ⁣{}_{\!} Learning. ⁣{}_{\!} Corres- pondence ⁣{}_{\!} learning ⁣{}_{\!} approaches ⁣{}_{\!} that ⁣{}_{\!} use ⁣{}_{\!} unlabeled video data fall ⁣{}_{\!} in ⁣{}_{\!} a ⁣{}_{\!} broad ⁣{}_{\!} field ⁣{}_{\!} of ⁣{}_{\!} self-supervised ⁣{}_{\!} video representation learning. Towards learning transferable video representation, diverse pretext tasks are proposed to explore different intrinsic properties of videos as free supervisory signals, including temporal sequence ordering ⁣{}_{\!} , predicting motion patterns ⁣{}_{\!} , solving space-time cubic puzzles ⁣{}_{\!} , anticipating future representations ⁣{}_{\!} , and temporally ⁣{}_{\!} aligning ⁣{}_{\!} videos ⁣{}_{\!} . ⁣{}_{\!} The ⁣{}_{\!} learned ⁣{}_{\!} representations are ⁣{}_{\!} compact ⁣{}_{\!} video ⁣{}_{\!} descriptors ⁣{}_{\!} that ⁣{}_{\!} can ⁣{}_{\!} be ⁣{}_{\!} generalized ⁣{}_{\!} to ⁣{}_{\!} va- rious ⁣{}_{\!} downstream ⁣{}_{\!} tasks ⁣{}_{\!} (e.g., ⁣{}_{\!} action ⁣{}_{\!} recognition ⁣{}_{\!} , ⁣{}_{\!} video ⁣{}_{\!} captioning ⁣{}_{\!} , ⁣{}_{\!} video ⁣{}_{\!} retrieval ⁣{}_{\!} ), while ⁣{}_{\!} in ⁣{}_{\!} this work we specifically focus on learning fine-grained visual features for pixel-level correspondence matching.

Self-Supervised Image Representation Learning. Basically, self-supervised approaches for image representation learning share a similar key idea with the counterparts for videos – design pretext tasks so as to mine supervision signals from the inherent information inside images ⁣{}_{\!} . Examples of such tasks include estimating spatial context ⁣{}_{\!} , predicting image rotations ⁣{}_{\!} , solving jigsaw puzzles ⁣{}_{\!} , among many others ⁣{}_{\!} . Recent efforts were mainly devoted to improving deep metric learning techniques in large ⁣{}_{\!} scale, ⁣{}_{\!} i.e., ⁣{}_{\!} adopt ⁣{}_{\!} an ⁣{}_{\!} instance ⁣{}_{\!} discrimination ⁣{}_{\!} pretext ⁣{}_{\!} task where the cross-entropy objective is used to discriminate each image from other different images (i.e., negative sam- ples) ⁣{}_{\!} . In this work, we assimilate the idea of contrasting positive samples (intra-video affinities) over numerous negative ones (inter-video affinities) for robust temporal correspondence matching, and formulate this within a unified inter-and intra-video reconstruction framework.

Video Mask Propagation. This task, a.k.a. semi-automatic video segmentation, aims at propagating first-frame object masks over the whole video sequence ⁣{}_{\!} . It addresses classic ⁣{}_{\!} automatic ⁣{}_{\!} video ⁣{}_{\!} segmentation ⁣{}_{\!} techniques’ ⁣{}_{\!}  ⁣{}_{\!} lack ⁣{}_{\!} of ⁣{}_{\!} flexibility ⁣{}_{\!} in ⁣{}_{\!} defining ⁣{}_{\!} target ⁣{}_{\!} objects ⁣{}_{\!} . ⁣{}_{\!} Depending on how they make use of the first-frame supervision, existing mask propagation models can be classified into three groups: i) online fine-tuning based methods that fine-tune a generic segmentation network with masks during inference ⁣{}_{\!} ; and ii) matching based methods that directly condition the segmentation network on the first-frame masks and/or previous segments ⁣{}_{\!} . As shown in Fig. ⁣{}_{\!} 1, extensive annotations are typically required to train such systems, i.e., first pre-train on ImageNet ⁣{}_{\!} and then fine-tune with COCO ⁣{}_{\!} , DAVIS ⁣{}_{\!} , Youtube-VOS ⁣{}_{\!} , etc. In contrast, we pursue a more annotation-efficient solution; similar to previous self-supervised correspondence learning methods ⁣{}_{\!} , Liir is trained using unlabeled videos only. Once trained, it can be directly applied for mask propagation, without adaptation.

Our Approach

We present Liir, a self-supervised framework that learns dense ⁣{}_{\!} correspondence ⁣{}_{\!} from ⁣{}_{\!} raw ⁣{}_{\!} videos. ⁣{}_{\!} Before ⁣{}_{\!} elaborating on ⁣{}_{\!} our ⁣{}_{\!} model ⁣{}_{\!} design ⁣{}_{\!} (cf. §3.2), ⁣{}_{\!} we ⁣{}_{\!} first ⁣{}_{\!} review ⁣{}_{\!} the ⁣{}_{\!} classic ⁣{}_{\!} reconstruction based ⁣{}_{\!} temporal ⁣{}_{\!} correspondence ⁣{}_{\!} learning ⁣{}_{\!} stra- tegy (cf. §3.1), which serves as the basis of our Liir.

Due to the appearance continuity in video, one can con- sider pixels in a ‘query’ frame as being copied from some locations ⁣{}_{\!} of ⁣{}_{\!} other ⁣{}_{\!} ‘reference’ ⁣{}_{\!} frames. ⁣{}_{\!} In light of this, a few studies ⁣{}_{\!} raise ⁣{}_{\!} a ⁣{}_{\!} reconstruction-based ⁣{}_{\!} correspondence learning scheme: each query pixel struggles to find pixels in a reference frame that can best reconstruct itself.

where A(i,j) ⁣∈ ⁣ ⁣A(i,j)\!\in_{\!}\! refers to (i,j)(i,j)-th element in AA, signifying the similarity between pixel ii in Iq{I}_{q} and pixel jj in Ir{I}_{r}, and ‘⋅\cdot’ stands for the dot product. In this way, AA gives the strength of all ⁣{}_{\!} the ⁣{}_{\!} pixel ⁣{}_{\!} pair-wise ⁣{}_{\!} correspondence ⁣{}_{\!} between ⁣{}_{\!} Iq ⁣\bm{I}_{q\!} and ⁣{}_{\!} Ir\bm{I}_{r}, ⁣{}_{\!} according ⁣{}_{\!} to ⁣{}_{\!} which pixel ii in Iq{I}_{q} can be reconstructed by a weighted sum of pixels in Ir{I}_{r}:

The training objective of ϕ\phi is hence defined as a reconstruction loss:

In practice, to avoid trivial solutions caused by information leakage, an information bottleneck is adopted over training samples, ⁣{}_{\!} e.g., ⁣{}_{\!} RGB2gray ⁣{}_{\!} operation ⁣{}_{\!} , ⁣{}_{\!} channel-wise ⁣{}_{\!} drop- out in RGB ⁣{}_{\!} or Lab ⁣{}_{\!} colorspace. After training, the representation encoder ϕ\phi is used for correspondence matching: similar to Eq. ⁣{}_{\!} 2, the affinity AA is estimated and used to propagate desired pixel-level entities (e.g., instance masks, key-point maps), from a reference frame to a query frame.

2 Liir: Locality-Awareabsent{}_{\!} Inter-andabsent{}_{\!} Intra-Videoabsent{}_{\!} Reconstruction Framework

Building upon the reconstruction-by-copy scheme, Liir is empowered with three crucial yet long overlooked abilities for robust correspondence learning: instance discrimination, location awareness, and spatial compactness.

Inter-and ⁣{}_{\!} Intra-Video ⁣{}_{\!} Reconstruction. ⁣{}_{\!} With ⁣{}_{\!} the ⁣{}_{\!} computa- tion of the intra-video affinity AA (Eq. ⁣{}_{\!} 1), each query pixel is forced to distinguish its counterpart (positive) reference pixels from unrelated (negative) ones within a same video, with ⁣{}_{\!} the ⁣{}_{\!} indictor ⁣{}_{\!} of ⁣{}_{\!} the ⁣{}_{\!} reconstruction ⁣{}_{\!} quality ⁣{}_{\!} Lres ⁣\mathcal{L}_{\text{res}\!} (Eqs. ⁣{}_{\!} 2-3). ⁣ ⁣{}_{\!\!} As both the positive and negative samples are sourced from the same video, there is less evidence for distinction among similar ⁣{}_{\!} object ⁣{}_{\!} instances ⁣{}_{\!} with ⁣{}_{\!} intra-video ⁣{}_{\!} appearance ⁣{}_{\!} only ⁣{}_{\!} (Fig. ⁣{}_{\!} 3(a)). ⁣{}_{\!} As ⁣{}_{\!} one ⁣{}_{\!} single ⁣{}_{\!} video ⁣{}_{\!} only ⁣{}_{\!} contains ⁣{}_{\!} limited content, ⁣{}_{\!} conducting ⁣{}_{\!} correspondence ⁣{}_{\!} matching ⁣{}_{\!} within ⁣{}_{\!} videos ⁣{}_{\!} is less challenging, and inevitably hinders the discrimination potential of the learned representation ⁣{}_{\!} . These insights motivate us to improve the intra-video affinity based reconstruction scheme by further accounting for negative across-video correspondence. Concretely, given the query (IqI_{q}) and reference (IrI_{r}) frames from the same video, an intra-inter video affinity A ⁣′ ⁣∈ ⁣hw ⁣× ⁣hw ⁣A^{\prime}_{\!}\!\in\!^{hw_{\!}\times_{\!}hw\!} is computed:

where {In}n\{{I}_{n}\}_{n} refer to a collection of frames, which are sampled from the whole training dataset, except the source video of IqI_{q} (IrI_{r}). By additionally considering other irrelevant videos during affinity computation, both the quantity and diversity of negative samples are greatly improved, allowing us to derive a more challenging inter-and intra-video reconstruction scheme (Fig. ⁣{}_{\!} 3(b)):

With Eqs. ⁣{}_{\!} 4-5, each pixel ii in the query frame Iq ⁣{I}_{q\!} is required to distinguish its counterpart pixels from massive unrelated ones, which are from not only the reference frame Ir{I}_{r} in current video, but a huge amount of irrelevant frames {In}n\{{I}_{n}\}_{n} in other videos. This powerful idea, yet, is elegantly achieved by the same training objective as in Eq. ⁣{}_{\!} 3. Note that Eq. ⁣{}_{\!} 4 normalizes intra-video correspondence over both inter-and intra-video pixel-to-pixel relevance, while Eq. ⁣{}_{\!} 5 only uses the pixels in the reference frame Ir{I}_{r} for reconstruction. The rationale here is that, even if the encoder ϕ\phi wrongly matches ⁣{}_{\!} a ⁣{}_{\!} query ⁣{}_{\!} pixel ⁣{}_{\!} ii with ⁣{}_{\!} a ⁣{}_{\!} negative ⁣{}_{\!} but ⁣{}_{\!} similar-looking ⁣{}_{\!} pixel ⁣{}_{\!} kk in In{I}_{n}, i.e., exp⁡(Iq(i) ⁣⋅ ⁣In ⁣(k))\exp(\bm{I}_{q}(i)_{\!}\cdot_{\!}\bm{I}_{n\!}(k)) will be large and ∑jA′(i,j) ⁣≪ ⁣1\sum_{j}A^{\prime}(i,j)\!\ll\!1, the synthesized color I^q(i)\hat{I}_{q}(i) will be still very different to Iq(i){I}_{q}(i) and ϕ\phi will receive a large gradient from Eq. ⁣{}_{\!} 3. Thus ϕ\phi is driven to mine more high-level semantics and context-related clues, hence reinforcing the instance-level discrimination ability (Fig. ⁣{}_{\!} 3(c)). Fig. ⁣{}_{\!} 3(d) shows that the representation learned with our inter-and intra-video reconstruction strategy can distinguish similar-looking dogs nearby.

Although also addresses inter-video analysis based reconstruction, it conducts embedding association on patch level, relying on a pre-trained tracker for patch alignment. Besides, the method consumes three loss terms for supervision, which is much more complicated than ours. Further, it separately conducts inter-and intra-video affinity based reconstruction. This is problematic; when both the reference frame in current video and an irrelevant frame from other videos contain query-like pixels, there is no explicit supervision signal to determine which one should be matched.

Position Encoding and Position Shifting. Plenty of literature in neuroscience has revealed that human visual system encodes both appearance and position information when we perceive ⁣{}_{\!} and ⁣{}_{\!} track ⁣{}_{\!} objects ⁣{}_{\!} . ⁣{}_{\!} Yet ⁣{}_{\!} existing ⁣{}_{\!} unsuper- vised correspondence methods put all focus on improving appearance based representation by ConvNets, ignoring the value ⁣{}_{\!} of ⁣{}_{\!} position ⁣{}_{\!} information. ⁣{}_{\!} Although ⁣{}_{\!}  ⁣{}_{\!} suggest that ConvNets ⁣{}_{\!} can ⁣{}_{\!} implicitly ⁣{}_{\!} capture ⁣{}_{\!} position ⁣{}_{\!} information ⁣{}_{\!} by ⁣{}_{\!} uti- lizing ⁣{}_{\!} image ⁣{}_{\!} boundary ⁣{}_{\!} effects, ⁣{}_{\!} explicit ⁣{}_{\!} position ⁣{}_{\!} encoding ⁣{}_{\!} has already been a core of full ⁣{}_{\!} attention ⁣{}_{\!} networks ⁣{}_{\!} (e.g., Trans- former ⁣{}_{\!} ), and ⁣{}_{\!} facilitated ⁣{}_{\!} a ⁣{}_{\!} variety ⁣{}_{\!} of ⁣{}_{\!} tasks ⁣{}_{\!} (e.g., instance segmentation ⁣{}_{\!} , tracking ⁣{}_{\!} , video segmentation ⁣{}_{\!} ). All these indicate that position encoding deserves more attention in the field of temporal correspondence learning.

where P\bm{P} is added with the output feature of the first conv layer of ϕ\phi and has the same size and dimension as the conv feature. We explore three position encoding strategies:

2D Sinusoidal Position Embedding (2DSPE): P\bm{P} is given with a family of pre-defined sinusoidal functions, without introducing new trainable parameters:

where ⁣{}_{\!} x ⁣ ⁣∈ ⁣ ⁣[0,w′)x_{\!}\!\in_{\!}\![0,w^{\prime}), ⁣{}_{\!} y ⁣ ⁣∈ ⁣ ⁣[0,h′)y_{\!}\!\in_{\!}\![0,h^{\prime}) ⁣{}_{\!} specify ⁣{}_{\!} the ⁣{}_{\!} horizontal ⁣{}_{\!} and vertical positions, u,v ⁣ ⁣∈ ⁣ ⁣[0,c′/4)u,v_{\!}\!\in_{\!}\![0,c^{\prime}/4) specify the dimension, and ε ⁣ ⁣= ⁣ ⁣10−4\varepsilon_{\!}\!=_{\!}\!10^{-4}. The horizontal (vertical) positions are encoded in the first (second) half of the dimensions. 2DSPE naturally handles resolutions that are unseen during training.

1D Absolute Position Embedding (1DAPE): 1DAPE is the most heavy-weight strategy: the whole P\bm{P} is a learnable parameter matrix without any constraint.

With ⁣{}_{\!} our ⁣{}_{\!} intra-inter ⁣{}_{\!} video ⁣{}_{\!} affinity ⁣{}_{\!} (Eq. ⁣{}_{\!} 4), ⁣{}_{\!} exploiting ⁣{}_{\!} position information ⁣{}_{\!} in ⁣{}_{\!} intra-video ⁣{}_{\!} correspondence ⁣{}_{\!} matching, ⁣{}_{\!} i.e., {exp⁡(Iq(i) ⁣⋅ ⁣Ir(j))}j\{\exp(\bm{I}_{q}(i)\!\cdot\!\bm{I}_{r}(j))\}_{j}, addresses local continuity resides in videos. However, for inter-video pixel relevance computation, i.e., {exp⁡(Iq(i) ⁣⋅ ⁣In ⁣(k))}n,k\{\exp(\bm{I}_{q}(i)\!\cdot\!\bm{I}_{n\!}(k))\}_{n,k}, such position prior is undesirable, as it inspires the query pixel ii in Iq ⁣I_{q\!} to prefer matching these pixels with similar positions in other irrelevant videos {In}n\{{I}_{n}\}_{n}. To eliminate such position encoding induced bias from inter-video correspondence matching, we design a position shifting strategy (Fig. ⁣{}_{\!} 4(a)). During training, for In ⁣{I}_{n\!} from other videos, we circularly shift the position encoding vectors in P\bm{P} by a random step in horizontal and vertical axes, respectively. The reason why we adopt random shifting with circular boundary conditions, instead of random shuffling, is to preserve the spatial layout in the modulated position encoding map Pˉ\bar{\bm{P}}. Then Pˉ\bar{\bm{P}} and In ⁣{I}_{n\!} are fed into ϕ\phi for inter-video correspondence matching, and Pˉ\bar{\bm{P}} related gradients are abandoned if learnable 1DAPE or 2DAPE is used. Note ⁣{}_{\!} that ⁣{}_{\!} the ⁣{}_{\!} standard ⁣{}_{\!} position ⁣{}_{\!} encoding ⁣{}_{\!} P\bm{P} ⁣{}_{\!} is ⁣{}_{\!} applied ⁣{}_{\!} for ⁣{}_{\!} the query (Iq{I}_{q}) and reference (Ir{I}_{r}) frames and updated normally. Fig. ⁣{}_{\!} 4(b) intuitively shows that merging position information into visual representation can enable robust correspondence matching even with confusing background and fast motion. In §4.4, we will quantitatively verify that 1DAPE is more favored and indeed boosts the performance.

Spatial Compactness Prior. As the visual world is conti- nuous and smoothly-varying, it is reasonable to assume appearances in video data change smoothly both in spatial and temporal ⁣{}_{\!} dimensions. ⁣{}_{\!} For ⁣{}_{\!} correspondence ⁣{}_{\!} learning, ⁣{}_{\!} the ⁣{}_{\!} tem- poral coherence has been extensively studied, while the spatial continuity received far less attention. To reduce search region, some existing methods ⁣{}_{\!} restrict corres- pondence matching within a local window, considering spatial regularities in a simple way. To make a better use of the spatial continuity, we augment the original reconstruction objective with an additional prior, termed as spatial compa- ctness. Such a prior poses constraints on the spatial distribution of associated pixels, leading to sparse and coherent solutions. Specifically, given the query (Iq{I}_{q}) and reference (Ir{I}_{r}) frames, we expect i) each query pixel ii to be only matched with a small number of reference pixels, and ii) the matched reference pixels to be clustered. For a query pixel ii and its matching ‘heatmap’: ⁣{}_{\!} Ai ⁣ ⁣= ⁣ ⁣[A(i,j)]j ⁣ ⁣∈ ⁣ ⁣h ⁣× ⁣wA_{i\!\!}=_{\!\!}[A(i,j)]_{j\!}\!\in_{\!}\!^{h_{\!}\times_{\!}w}, w.r.t. Ir{I}_{r}, we assume AiA_{i} follows a mixture of MM 2D Gaussian distributions:

3 Implementation Details

Network ⁣{}_{\!} Configuration. ⁣{}_{\!} For ⁣{}_{\!} fair ⁣{}_{\!} comparison, ⁣{}_{\!} our ⁣{}_{\!} feature ⁣{}_{\!} encoder ⁣{}_{\!} ϕ\phi ⁣{}_{\!} is ⁣{}_{\!} implemented ⁣{}_{\!} as ⁣{}_{\!} ResNet-18​ as in . Following , ×\times2 downsampling is only made in the third residual block. Thus ϕ\phi finally outputs 256 feature maps of 1/4 size of the input, i.e., h ⁣ ⁣= ⁣ ⁣H4,w ⁣ ⁣= ⁣ ⁣W4,c ⁣ ⁣= ⁣ ⁣256h_{\!}\!=_{\!}\!\frac{H}{4},w_{\!}\!=_{\!}\!\frac{W}{4},c_{\!}\!=_{\!}\!256. The position embedding is added to feature after the first 7×\times7 Conv-BN-ReLU layer, i.e., h ⁣′ ⁣= ⁣ ⁣H2,w ⁣′ ⁣= ⁣ ⁣H2,c ⁣′ ⁣= ⁣ ⁣64h^{\prime}_{\!}\!=_{\!}\!\frac{H}{2},w^{\prime}_{\!}\!=_{\!}\!\frac{H}{2},c^{\prime}_{\!}\!=_{\!}\!64.

Training: Liir is trained from scratch on two NVIDIA RTX-30903090 GPUs and only uses the raw videos from Youtube-VOS​ . Each training image is resized into 256×\times256 and channel-wise dropout in Lab colorspace ⁣{}_{\!} is adopted as the information bottleneck. Adam optimizer is used. At the initial 3030 epochs, only intra-video reconstruction learning is adopted for warm-up, with a learning rate of 10−3 ⁣10^{-3\!} and batch size of 3232. Then we conduct inter-video reconstruction learning with spatial compactness based regularization at the next 55 epochs, with a learning rate of 10−4 ⁣10^{-4\!} and batch size of 1212. We online maintain a memory bank of 1,440 frames from different videos. For each query pixel, we sample 4 feature points from each stored frame, i.e., a total of 1,440×\times4 negative samples are used for the inter-video correspondence computation, and we employ the moving average strategy for parameter updating.

Experiment

We evaluate the learned representation on diverse video label ⁣{}_{\!} propagation ⁣{}_{\!} tasks, ⁣{}_{\!} i.e., ⁣ ⁣{}_{\!\!} video ⁣{}_{\!} object ⁣{}_{\!} segmentation ⁣{}_{\!} (§4.1), ⁣ ⁣{}_{\!\!} body part propagation (§4.2), and pose keypoint tracking (§4.3). As in conventions , all these tasks are to propagate the first frame annotation to the whole video sequence, and we use our model to compute inter-frame dense correspondences. In §4.4, we conduct a set of ablative studies to examine the efficacy of our essential model designs.

Dataset. We first test our method on val sets of two popular video object segmentation datasets, i.e., DAVIS17 and YouTube-VOS . There are 3030 and 474474 videos in DAVIS17 and YouTube-VOS val sets, respectively.

Evaluation Metric. Following the official protocol , we use the region similarity (J\mathcal{J}) and contour accuracy (F\mathcal{F}) as the evaluation metrics. Note that the scores on YouTube-VOS are respectively reported for seen and unseen categories, obtained from the official evaluation server.

Performance on DAVIS17: As illustrated in Table ⁣{}_{\!} 1, our Liir consistently outperforms all existing self-supervised methods across all the evaluation metrics. For example, it surpasses current best-performing self-supervised method, i.e., CLTC ⁣{}_{\!} , in terms of mean J&F\mathcal{J}\&\mathcal{F} (72.1  ⁣{}_{\!}vs. ⁣{}_{\!} 70.3). In addition, even without using any manual annotations for training, Liir achieves very competitive segmentation performance in comparison with some famous supervised models trained with massive pixel-wise annotations.

Performance on YouTube-VOS val. Table 2 reports performance comparison of Liir against four self-supervised competitors on YouTube-VOS val. It can be observed that Liir sets new state-of-the-art. In particular, Liir yields an overall score of 69.3%, surpassing the second-best (i.e., CLTC ⁣{}_{\!} ) and third-best (i.e., MAST ⁣{}_{\!} ) approaches by 2.0% and 5.1%, respectively. Further, Liir even outperforms some famous supervised methods (i.e., OSVOS ⁣{}_{\!} and PreMVOS ⁣{}_{\!} ), especially for the unseen categories, clearly demonstrating its remarkable generalization ability.

Qualitative Results. Fig. ⁣{}_{\!} 6 depicts visual results on representative videos in the datasets. As seen, Liir is able to establish accurate correspondences under various challenging scenarios, e.g., scale changes, small objects and occlusions.

2 Results for Body Part Propagation

Dataset. We next evaluate our model performance for body part propagation. Experiments are conducted on VIP val ⁣{}_{\!} , which contains 50 videos with annotations of 1919 human semantic part categories (e.g., hair, face, dress).

Evaluation Metric. As suggested by VIP , we adopt mean intersection-over-union (mIoU) and mean Average Precision (mAP) metrics for evaluation of semantic-level and instance-level parsing, respectively.

Performance. As shown in Table 3, Liir achieves state-of-the-art performance on both semantic-level and instance-level parsing. This indicates that Liir can generate strong representation which models both cross-instance discrimination and intra-instance invariance well. Fig. ⁣{}_{\!} 7 depicts some visualization results on two representative videos. Liir achieves temporally stable results and shows robustness to typical challenges (e.g., pose variations, occlusions).

3 Results for Pose Keypoint Tracking

Dataset. We then examine the model performance on human keypoint tracking, using JHMDB​ val. JHMDB val has 268268 videos. For each person, a total of 1515 body joints, e.g., torso, head, shoulder, elbow, are annotated.

Evaluation Metric. We use probability of correct keypoint (PCK) to measure the accuracy between each tracking result and corresponding ground-truth with a threshold τ\tau.

Performance. Table 3 shows that Liir exhibits compelling overall performance. Note that CLTC​ uses different checkpoints and model architectures for different tasks and datasets, while we only use a single model for evaluation. The visual results in Fig. ⁣{}_{\!} 7 also demonstrate the strong capability of Liir in establishing precise correspondence.

4 Diagnostic Experiment

For further detailed analysis, we conduct a series of ablative studies on DAVIS17 val and VIP ⁣{}_{\!} val sets.

Key Component Analysis. We first examine the efficacy of essential components of Liir, i.e., inter-and intra-video reconstruction, position encoding, and spatial compactness. The results are summarized in Table ⁣{}_{\!} A.1, where position encoding is implemented as 1DAPE, and compactness prior is used during both training and inference stages. When separately comparing row #2 - #4 with the baseline (MAST ⁣{}_{\!} ) in row #1, we can observe that each individual module in- deed ⁣{}_{\!} boosts ⁣{}_{\!} the ⁣{}_{\!} performance. ⁣{}_{\!} For ⁣{}_{\!} example, ⁣{}_{\!} on ⁣{}_{\!} DAVIS17 ⁣{}_{\!} val, ⁣{}_{\!} inter-and intra-video reconstruction, position encoding, and spatial compactness prior respectively bring 3.4%, 1.6%, and 3.1% J&F\mathcal{J}\&\mathcal{F} gains. This verifies our core insight that these three elements are crucial for correspondence learning. Finally, in row #5, we combine all the three components together – Liir, and obtain the best performance. This suggests that these modules are complementary to each other, and confirms the effectiveness of our whole design.

Inter-and Intra-Video Reconstruction. We next study the impact of increasing the number of negative samples, i.e., frames from other irrelevant videos used for inter-video correspondence computation (Eq. ⁣{}_{\!} 4). In Table ⁣{}_{\!} 5a, row #1 gives scores of learning without considering inter-video correspondence. In this case, the results are unsatisfactory. When more negative samples are involved (i.e., 0→\rightarrow1,440), better performance can be achieved (i.e., 69.2→\rightarrow72.1 on DAVIS17 val, 38.4→\rightarrow41.2 on VIP val). Finally we use 1,440 negative samples for inter-video reconstruction based learning, which is the maximum number allowed by our GPU.

Position Encoding. To determine the effect of our position encoding module, we then report the performance with different encoding strategies in Table ⁣{}_{\!} 5b. As seen, the non-learnable strategy, 2DSPE, hinders the performance, while the learnable alternatives, i.e., 1DAPE and 2DAPE, lead to better results. Compared with 2DAPE, 1DAPE is more favored, probably due to its high flexibility and capacity.

Position Shifting. We further study the influence of our position shifting strategy on performance. As shown in Table ⁣{}_{\!} 5c, we consider two alternatives, i.e., ‘NAN’ and ‘position shuffling’. ‘NAN’ refers to using the normal position encoding map P\bm{P} without any modulation during inter-video correspondence matching. Compared with ‘position shifting’, ‘NAN’ suffers from performance degradation (i.e., 72.1→\rightarrow71.3 on DAVIS17 val, 41.2→\rightarrow40.6 on VIP val), showing the negative effect of the position-induced bias. The other baseline, ‘position shuffling’, i.e., randomly shuffling the position encoding map for inter-video affinity computation, though better than ‘NAN’, is still worse than ‘position shifting’. This is because it destroys the spatial layouts.

Spatial Compactness Prior. The spatial compactness prior (Eq. ⁣{}_{\!} 7) is used to regularize intra-video correspondence matching during both training and inference stages. In Table ⁣{}_{\!} 5d, we quantitatively identify the performance contribution of our spatial compactness prior in different stages.

Conclusion

We presented a self-supervised temporal correspondence learning approach, Liir, that makes contributions in three aspects. First, going beyond the popular intra-video analysis based learning scheme, we further enforce separation between intra- and inter-video pixel associations, enhancing instance-level feature discrimination. Second, with a clever position shifting strategy, we bring the advantages of position encoding into full play, while avoiding its undesirable impact at the same time. Third, a spatial compactness prior is introduced to regularize representation learning and improve correspondence inference. The effectiveness was thoroughly validated over various label propagation tasks.

References

Appendix A Analysis of Spatial Compactness Prior

We use MM 2D Gaussian distributions to approximate the affinity matrix AA between two frames. Table A.1 provides a detailed analysis of the hyper-parameter MM. The 11st row corresponds to a baseline model that disregards the spatial compactness prior in both training and inference phases. We see from the table that 1) when M ⁣= ⁣1M\!=\!1 in training, the model performs worse than the baseline model, as such a rigorous matching constraint easily leads to overconfident pre-dictions; 2) when MM (in training) becomes larger, the performance greatly improves; 3) the models trained with M ⁣= ⁣2M\!=\!2 shows consistently better performance than those trained with M ⁣= ⁣3M\!=\!3; and 4) in the inference stage, M ⁣= ⁣2M\!=\!2 always leads to the best performance. Accordingly, we set MM to 22 in both training and inference stages.

Appendix B Analysis of Long-Term Dependence

Liir leverages multiple reference frames during testing as in . We analysis the impact of long-term dependence in Table B.1. As seen, our method is more robust and shows smaller drop wrt : 4.7% vs 6.8%.

Appendix C Analysis of Feature Point Sampling

We give the ablative study on the number of feature points sampled during inter-video reconstruction in Table C.1. It can be seen that, with 1440 frames, better results can be achieved if we sample more points per frame (70.4→\rightarrow 71.4→\rightarrow72.1), supporting our claim about instance separation. With the same number of total sampled points, we gain better performance if we consider more frames (70.9→\rightarrow 71.7→\rightarrow72.1). It is reasonable as more frames can provide much rich/challenging context. With our limited GPU capacity, we choose to sample 1440 frames and four feature points per frame, so as to maximize the performance. But we can speculate that, if with enough GPU capacity, sampling more features points from more videos will further improve the performance.

Appendix D Visualization of Ablation Study

Fig. ⁣{}_{\!} D.1 depicts visual effect of each essential component in Liir. Starting from the baseline model (b), we progressively add position encoding (c), inter-video reconstruction (d) and spatial compactness (e). As seen, with explicit positional encoding, our model is able to heavily suppress background regions (e.g., shelves in the first row). The inter-video reconstruction enables more accurate discrimination between targets and semantically similar distractors (e.g., “motorcycle” in the third row). Last, incorporating the spatial compactness prior facilitates more precise correspondences, leading to high-quality segmentation results.

Appendix E Additional Qualitative Results

We provide additional video propagation results on four datasets, including DAVIS17​ val in Fig. ⁣{}_{\!} D.2, Youtube-VOS​ val in Fig. ⁣{}_{\!} D.3, VIP​ val in Fig. ⁣{}_{\!} D.4 and JHMDB​ val in Fig. ⁣{}_{\!} D.5. We observe that even training with no annotations, Liir is able to produce highly exquisite results.

Appendix F Limitation

Although Liir demonstrates remarkable performance and high generalizability in correspondence matching, we still see a large performance gap between Liir and current top-leading supervised models (e.g., STM in VOS). However, as a self-supervised method, Liir can be easily scaled to leverage any available collection of video data for training. This could lead to more accurate correspondence learning from massive unlabeled data instead of using small-scale datasets only (e.g., DAVIS, YouTube-VOS). Apart from that, inter-video reconstruction spares massive space to bank the negative samples, this coerces the memory size of GPUs and extends the training time. Fortunately, the performance of Liir has improved by leaps and bounds to compensate for it, for instance, up to 2.9% on DAVIS17​ . Further more, note that the inter-video reconstruction is only applied during training, thus it does not introduce extra computational cost at inference. In our future work, we will explore towards the above direction to narrow the performance gap between supervised methods, and find the way to reduce the memory space taken up by negative videos and speed up training simultaneously.