Do Different Tracking Tasks Require Different Appearance Models?

Zhongdao Wang, Hengshuang Zhao, Ya-Li Li, Shengjin Wang, Philip H. S. Torr, Luca Bertinetto

Introduction

Unlike popular image-based computer vision tasks such as classification and object detection, which are (for the most part) unambiguous and clearly defined, the problem of object tracking has been considered under different setups and scenarios, each motivating the design of a separate set of benchmarks and methods. For instance, for the Single Object Tracking (SOT) and Video Object Segmentation (VOS) communities , tracking means estimating the location of an arbitrary user-annotated target object throughout a video, where the location of the object is represented by a bounding box in SOT and by a pixel-wise mask in VOS. Instead, in multiple object tracking settings (MOT , MOTS and PoseTrack ), tracking means connecting sets of (often given) detections across video frames to address the problem of identity association and forming trajectories. Despite these tasks only differing in the number of objects per frame to consider and observation format (bounding boxes, keypoints or masks), the best practices developed by the methods tackling them vary significantly.

Though the proliferation of setups, benchmarks and methods is positive in that it allows specific use cases to be thoroughly studied, we argue it makes increasingly harder to effectively study one of the fundamental problems that all these tasks have in common, i.e. what constitutes a good representation to track objects throughout a video? Recent advancements in large-scale models for language and vision have suggested that a strong representation can help addressing multiple down-stream tasks. Similarly, we speculate that a good representation is likely to benefit many different tracking tasks, regardless of their specific setup. In order to validate our speculation, in this paper we present a framework that allows to adopt the same appearance model to address five different tracking tasks (Figure 3). In our taxonomy (Figure 4), we consider existing tracking tasks as problems that have either propagation or association at their core. When the core problem is propagation (as in SOT and VOS), one has to localise a target object in the current frame given its location in the previous one. Instead, in association problems (MOT, MOTS, and PoseTrack), target states in both previous and current frames are given, and the goal is to determine the correspondence between the two sets of observations. We show how most tracking tasks currently considered by the community can be simply expressed starting from the primitives of propagation or association. For propagation tasks, we employ existing box and mask propagation algorithms . For association tasks, we propose a novel reconstruction-based metric that leverages fine-grained correspondence to measure similarities between observations. In the proposed framework, each individual task is assigned to a dedicated “head” that allows to represent the object(s) in the appropriate format to compare against prior arts on the relevant benchmarks.

Note that, in our framework, only the appearance model contains parameters that can be learned via back-propagation, and that we do not experiment with appearance models that have been trained on specific tracking tasks. Instead, we adopt models trained via recent self-supervised learning (SSL) techniques and that have already demonstrated their effectiveness on a variety of image-based tasks. Our motivation is twofold. First, SSL models are particularly interesting for our use-case, as they are explicitly conceived to be of general purpose. As a byproduct, our work also serves the purpose of evaluating and comparing appearance models obtained from self-supervised learning approaches (see Figure 1). Second, we hope to facilitate the tracking community in directly benefiting from the rapid advancements of the self-supervised learning literature.

To summarise, the contributions of our work are as follows:

We propose UniTrack, a framework that supports five tracking tasks: SOT , VOS , MOT , MOTS , and PoseTrack ; and that can be easily extended to new ones.

We show how UniTrack can leverage many existing general-purpose appearance models to achieve a performance that is competitive with the state-of-the-art on several tracking tasks.

We propose a novel reconstruction-based similarity metric for association that preserves fine-grained visual features and supports multiple observation formats (box, mask and pose).

We perform an extensive evaluation of self-supervised models, significantly extending the empirical analysis of prior literature to video-based tasks.

The UniTrack Framework

Inspecting existing tracking tasks and benchmarks, we noticed that their differences can be roughly categorised across four axes, illustrated in Figure 3 and detailed below.

Whether the requirement is to track a single object (SOT , VOS ), or multiple objects (MOT , MOTS , PoseTrack ).

Whether the targets are specified by a user in the first frame only (SOT, VOS), or instead are given in every frame, e.g. by a pre-trained detector (MOT, MOTS, PoseTrack).

Whether the target objects are represented by bounding-boxes (SOT, MOT), pixel-wise masks (VOS, MOTS) or pose annotations (PoseTrack).

Whether the task is class-agnostic, i.e. the target objects can be of any class (SOT, VOS); or if instead they are from a predefined set of classes (MOT, MOTS, PoseTrack).

Typically, in single-object tasks the target is specified by the user in the first frame, and it can be of any class. Instead, for multi-object tasks detections are generally considered as given for every frame, and the main challenge is to solve identity association for the several objects. Moreover, in multi-object tasks the set of classes to address is generally known (e.g. pedestrians or cars).

Figure 3 depicts a schematic overview of the proposed UniTrack framework, which can be understood as conceptually divided in three “levels”. The first level is represented by the appearance model, responsible for extracting high-resolution feature maps from the input frame (Section 2.2). The second level consists of the algorithmic primitives addressing propagation (Section 2.3) and association (Section 2.4). Finally, the last level comprises multiple task-specific algorithms that make direct use of the primitives of the second level. In this work, we illustrate how UniTrack can be used to obtain competitive performance on all of the five tracking tasks of level-3 from Figure 3. Moreover, new tracking tasks can be easily integrated.

Importantly, note that the appearance model is the only component containing trainable parameters. The reason we opted for a shared and non task-specific representation is twofold. Firstly, the large amount of different setups motivated us to investigate whether having separately-trained models for each setup is necessary. Since training on specific datasets can bias the representation towards a limited set of visual concepts (e.g. animals or vehicles) and limit its applicability to “open-world” settings, we wanted to understand how far can a shared representation go. Second, we wanted to provide the community with multiple baselines that can be used to better assess newly proposed contributions, and that can be immediately used on new datasets and tasks without the need of retraining.

2 Base appearance model

In order to learn fine-grained correspondences, fully-supervised methods are only amenable for synthetic datasets (e.g. Flying Chairs for optical flow ). With real-world data, it is intractable to label pixel-level correspondences and train models in a fully-supervised fashion. To overcome this obstacle, in this paper we adopt representations obtained with self-supervision. We experiment both with models trained with approaches that leverage pixel-wise pretext tasks and, inspired by prior works that have pointed out how fine-grained correspondences emerge in middle-level features , with models obtained from image-level tasks (e.g. MoCo , SimCLR ).

3 Propagation

Problem definition. Figure 4a schematically illustrates the problem of propagation, which we use as a primitive to address SOT and VOS tasks. Considering the single-object case, given video frames {It}t=1T\{I_{t}\}_{t=1}^{T} and an initial ground truth observation z1{z}_{1} as input, the goal is to predict object states {z^t}t=2T\{\hat{{z}}_{t}\}_{t=2}^{T} for each time-step tt. In this work we consider three formats to represent objects: bounding boxes, segmentation masks and pose skeletons.

where ⟨⋅,⋅⟩\braket{\cdot,\cdot} indicates inner product, and τ\tau is a temperature hyperparameter. As in , we only keep the top KK values for each row and set other values to zero. Then, the mask for the current frame at time tt is predicted by propagating the previous prediction: zt=Kt−1tzt−1z_{t}=K_{t-1}^{t}z_{t-1}. Mask propagation proceeds in a recurrent fashion: the output mask of the current frame is used as input for the next one.

Pose propagation. In order to represent pose keypoints, we use the widely adopted Gaussian belief maps . For a keypoint pp, we obtain a belief map zp∈sz^{p}\in^{s} by using a Gaussian with mean equal to the keypoint’s location and variance proportional to the subject’s body size. In order to propagate a pose, we can then individually propagate each belief map in the same manner as mask propagation, again as ztp=Kt−1tzt−1pz_{t}^{p}=K^{t}_{t-1}z_{t-1}^{p}.

Box propagation. The position of an object can also be more simply expressed with a four-dimensional vector z=(u,v,w,h){z=(u,v,w,h)}, where (u,v)(u,v) are the coordinates of the bounding-box center, and (w,h)(w,h) are its width and height. While one could reuse the strategy adopted above by simply converting the bounding-box to a pixel-wise mask, we observed that using this strategy leads to inaccurate predictions. Instead, we use the approach of SiamFC , which consists in performing cross-correlation (XCorr) between the target template zt−1z_{t-1} and the frame XtX_{t} to find the new location of the target in frame tt. Cross-correlation is performed at different scales, so that the bounding-box representation can be resized accordingly. We also provide a Correlation Filter-based alternative (DCF) (see Appendix B.1).

4 Association

Problem definition. Figure 4b schematically illustrates the association problem, which we use as primitive to address the tasks of MOT, MOTS and PoseTrack. In this case, observations for object states {Z^t}t=1T\{\hat{\mathcal{Z}}_{t}\}_{t=1}^{T} are given for all the frames {It}t=1T\{I_{t}\}_{t=1}^{T}, typically via the output of a pre-trained detector. The goal here is to form trajectories by connecting observations across adjacent frames according to their identity.

Association algorithm. We adopt the association algorithm proposed in JDE for MOT, MOTS and PoseTrack tasks, of which detailed description can be found in Appendix C.1. In summary, we compute an N×MN\times M distance matrix between NN already-existing tracklets and MM “new” detections from the last processed frame. We then use the Hungarian algorithm to determine pairs of matches between tracklets and detections, using the distance matrix as input. To obtain the matrix of distances used by the algorithm, we compute the linear combination of two terms accounting for motion and appearance cues. For the former, we compute a matrix indicating how likely a detection corresponds to the object state predicted by a Kalman Filter .

Instead, the appearance component is directly computed by using feature-map representations obtained by processing individual frames with the appearance model (Section 2.2). While object-level features for box and mask observations can be directly obtained by cropping frame-level feature maps, when an object is represented via a pose it first needs to be converted to a mask (via a procedure described in Appendix C.2).

A key issue of this scenario is how to measure similarities between object-level features. We find existing methods limited. First, objects are often compared by computing the cosine similarity of average-pooled object-level feature maps . However, the operation of average inherently discards local information, which is important for fine-grained recognition. Approaches that instead to some extent do preserve fine-grained information, such as those computing the cosine similarity of (flattened) feature maps, do not support objects with differently-sized representation (situation that occurs for instance with pixel-level masks). To cope with the above limitations, we propose a reconstruction-based similarity metric that is able to deal with different observation formats, while still preserving fine-grained information.

where t^i←j\hat{t}_{i\leftarrow j} represents tit_{i} reconstructed from djd_{j} and d^j←i\hat{d}_{j\leftarrow i} represents djd_{j} reconstructed from tit_{i}. In multi-object tracking scenarios, observations are often incomplete due to frequent occlusions. As such, directly comparing features between incomplete and complete observations often fails because of misalignment between local features. Suppose djd_{j} is a detection feature representing a severely occluded pedestrian, while tit_{i} a tracklet feature representing the same person, but unoccluded. Likely, directly computing the cosine similarity between the two will not be very telling. RSM addresses this issue by introducing a step of reconstruction after which the co-occurring parts of point features will be better aligned, thus making the final similarity more likely to be meaningful.

RSM can be interpreted from an attention perspective. The feature map of a tracklet tit_{i} being reconstructed can be seen as a set of queries, and the “source” detection feature djd_{j} can be interpreted both as keys and values. The goal is to reconstruct the queries by linear combination of the values. The linear combination (attention) weights are computed using the affinity between queries and keys. Specifically, we first compute a global affinity matrix between tit_{i} and all the dj′d_{j^{\prime}} for j′=1,...,Mj^{\prime}=1,...,M, and then extract the corresponding sub-matrix for tit_{i} and dj′d_{j^{\prime}} as the attention weights. Our formulation leads to a desired property: if the attention weights approach zero, the corresponding reconstructed point vectors will approach zero and so the RSM between tit_{i} and djd_{j}.

Measuring similarity by reconstruction is popular in problems such as few-shot learning , self-supervised learning , and person re-identification . However, reconstruction is typically framed as a ridge regression or optimal transport problem. With O(n2)O(n^{2}) complexity, RSM is more efficient than ridge regression and it has a similar computation cost to calculating the Earth Moving Distance for the optimal transport problem. Appendix D shows a series of ablation studies illustrating the importance of the proposed RSM for the effectiveness of UniTrack on association-type tasks.

Experiments

Since UniTrack does not require task-specific training, we were able to experiment with many alternative appearance models (see Figure 3) with little computational cost. In Section 3.1 we perform an extensive evaluation to benchmark a wide variety of off-the-shelf, modern self-supervised models, showing their strengths and weaknesses on all five tasks considered. In this section we also conduct a correlation study with the so-called “linear probe” strategy , which became a popular way to evaluate representations obtained with self-supervised learning. Then, in Section 3.2 we compare UniTrack (equipped with supervised or unsupervised appearance models) against recent and task-specific tracking methods.

Implementation details. We use ResNet-18 or ResNet-50 as the default architecture. With ImageNet-supervised appearance model, we refer to the ImageNet pre-trained weights made available in PyTorch’s “Model Zoo”. To prevent excessive downsampling, we modify the spatial stride of layer3 and layer4 to 11, achieving a total stride of r=8r=8. We extract features from both layer3 and layer4. We report results with layer3 features when comparing against task-specific methods (Section 3.2), and with both layer3 and layer4 when evaluating multiple different representations (Section 3.1). Further implementation details are deferred to Appendix B and C.

Datasets and evaluation metrics. For fair comparison with existing methods, we report results on standard benchmarks with conventional metrics for each task. Please refer to Appendix A for details.

The process of evaluating representations obtained via self-supervised learning (SSL) often involves additional training , for instance via the use of linear probes , which require to fix the pre-trained model and train an additional linear classifier on top of it. In contrast, using UniTrack as evaluation platform (1) does not require any additional training and (2) enables the evaluation on a battery of important video tasks, which have generally been neglected in self-supervised-learning papers in favour of more established image-level tasks such as classification.

In this section, we evaluate three types of SSL representations: (a) Image-level representations learned from images, e.g. MoCo and BYOL ; (b) Pixel-level representations learned from images (such as DetCo and PixPro ) and (c) videos (such as UVC and CRW ). For all methods considered, we use the pre-trained weights provided by the authors.

Results are shown in Table 7 and 7, where we report the results obtained by using features from either layer3 or layer4 of the pre-trained ResNet backbone. We report both results and separate them by a ‘/’ in the table. Note that, for this analysis only, for association-type tasks motion cues are discarded to better highlight distinctions between different representations and avoid potential confounding factors. Figure 1 and 6 provides a high-level summary of the results by focusing on the ranking obtained by different SSL methods on the five tasks considered (each represented by a vertex in the radar-style plot). Several observations can be made:

(1) There is no significant correlation between “linear probe accuracy” on ImageNet and overall tracking performance. The linear probe approach has become a standard way to compare SSL representations. In Figure 7, we plot tracking performance on five tasks (y-axes) against ImageNet top-1 accuracy of 16 different models (x-axes), and report Pearson and Spearman (rank) correlation coefficients. We observe that the correlation between ImageNet accuracy and tracking performance is small, i.e. the Pearson’s rr ranges from −0.38-0.38 to +0.20+0.20, and Spearman’s ρ\rho ranges from −0.36-0.36 to +0.26+0.26. For most tasks, there is almost no correlation, while for VOS the two measures are mildly inversely correlated. The result suggests that evaluating SSL models on five extra tasks with UniTrack could constitute a useful complement to ImageNet linear probe evaluation, and encourage the SSL community to pursue the design of even more general purpose representations.

(2) A vanilla ImageNet-trained supervised representation is surprisingly effective across the board. On most tasks, it reports a performance competitive with the best representation for that task. This is particularly evident from Figure 1, where its performance is outlined as a gray dashed line. This result suggests that results obtained with vanilla ImageNet features should be reported when investigating new tracking methods.

(3) The best self-supervised representation ranks first on most tasks. Recently, it has been shown how SSL-trained representations can match or surpass their supervised counterparts on ImageNet classification (e.g. ) and many downstream tasks . Within UniTrack, although no individual SSL representation is able to beat the vanilla ImageNet-trained representation on every single task, we observe that the recently proposed VFS ranks first on every task, except for single-object tracking. This suggests that advancements of the self-supervised learning literature can directly benefit the tracking community: it is reasonable to expect that newly-proposed representations will further improve performance across the board.

(4) Pixel-level SSL representations do not seem to have a consistent advantage in pixel-level tasks. In Table 7 and at the bottom of Table 7 we compare recent SSL representations trained with pixel-level proxy tasks: PixPro , DetCo , TimeCycle , Colorization , UVC and Contrastive Random Walk (CRW) . Considering that pixel-level models leverage more fine-grained information during training, one may expect them to outperform image-based models in the tracking tasks where this is important. It is not straightforward to compare pixel-level SSL models with image-level ones, as the two types employ different default backbone networks. However, note how good image-based models (MoCo-v1, SimCLR-v2) are on par with their supervised counterpart in all tasks, while good pixel-level models (DetCo, CRW) still have gaps with respect to their supervised counterparts in tasks like SOT and MOT. Moreover, from Table 7, one can notice how the last three rows, despite representing methods leveraging pixel-level information during training, are actually outperformed by image-level representations on the pixel-level tasks of VOS, MOTS and PoseTrack.

(5) Video data can benefit representation learning for video tasks. The top-ranking VFS is similar to MoCo, SimCLR and BYOL in terms of learning scheme: they all perform contrastive learning on image level features. The most important distinction is the training data. Previous SSL methods mostly train on still-image based datasets (typically ImageNet), while VFS employs a large-scale video dataset Kinetics . Clearly, this is not very surprising, as training on video data can help closing the domain gap with the (video-based) downstream tasks considered in this paper.

2 Comparison with task-specific tracking methods

Unsupervised methods. We observe that UniTrack performs competitively against unsupervised state-of-the-art methods in both the propagation-type tasks we considered (Table 5d and 5e). For SOT, UniTrack with a DCF head outperforms UDT (a strong recent method) by 2.42.4 AUC points, while it is surpassed by LUDT+ by 2.12.1 points. Considering that LUDT+ adopts an additional online template update mechanism while ours does not, we believe the gap could be closed. In VOS, existing unsupervised methods are usually trained on video datasets , and some of the most recent outperform UniTrack (with an ImageNet-trained representation). Nonetheless, when we use a VFS-trained representation, this performance difference is reduced to 2%. Finally, note that for association-type tasks we are not aware of any existing unsupervised learning method, and thus in this case we limit the comparison to supervised methods.

Comparison with supervised methods. In general, UniTrack with a ResNet-18 appearance model already performs on par with several existing task-specific supervised methods, and in several tasks it even shows superior accuracy, especially for identity-related metrics. (1) For SOT, UniTrack with a DCF head outperforms SiamFC by 3.63.6 AUC points. This is a significant margin considering that SiamFC is trained with a large amount of crops from video datasets with annotated bounding boxes. (2) For VOS, UniTrack surpasses SiamMask by 4.14.1 J\mathcal{J}-mean points, despite this being trained on the joint set of three large-scale video datasets . (3) For MOT, we employ the same detections used by the state-of-the-art tracker FairMOT . The appearance embedding in FairMOT is trained with 270K bounding boxes of 8.7K labeled identities, from a MOT-specific dataset. In contrast, despite our appearance model not being trained with any MOT-specific data, our IDF1 score is quite competitive (71.871.8 v.s. 72.872.8 of FairMOT), and the ID switches are considerably reduced by 36.4%36.4\%, from 1074 to 683. (4) For MOTS, we start from the same segmentation masks used by the COSTA tracker, and observe a degradation in terms of ID switches (622 vs the 421 of the state of the art), and also a gap in IDF1 and sMOTA. (5) Finally, for pose tracking, we employ the same pose estimator used by LightTrack . Compared with LightTrack, the MOTA of UniTrack degrades of 1.31.3 points because of an increased amount of ID switches. However, the IDF-1 score is improved by a significant margin (+21.021.0 points). This shows UniTrack preserves identity more accurately for long tracklets: even if ID switches occur more frequently, after a short period UniTrack is able to correct the wrong association, leading to a higher IDF-1.

Notice how, overall, UniTrack obtains more competitive performance on tasks that have association at their core, i.e. MOT, MOTS and PoseTrack. Upon inspection, we observed that most failure cases in propagation-type tasks regard the “drift” occurring when the scale of the object is improperly estimated. In future work, this could be addressed for instance by a bounding-box regression module to refine predictions, or by carefully designing a motion model. For association-type tasks, the consequences of any type of inaccuracy are isolated to individual pairs of frames, and thus much less catastrophic by nature.

Related Work

To the best of our knowledge, sharing the appearance model across multiple tracking tasks has not been extensively studied in the computer vision literature, and especially not in the context of SSL representations. Some existing methods do share a common backbone architecture across tasks. For instance, STEm-Seg addresses VIS and MOTS; while TraDeS addresses MOT, MOTS and VIS. However, both methods need to be trained separately and on different datasets for every task. Conversely, we reuse the same representation across five tasks. A promising direction for future work would be to use UniTrack to train a shared representation in a multi-task fashion. Only a few relevant works do adopt a multi-task approach , and they usually consider SOT and VOS tasks only. In general, despite the multi-task direction being surely interesting, it requires the availability of large-scale datasets with annotations in multiple formats, and costly training. These are two of the main reasons for which we believe that having a framework that allows to achieve competitive performance on multiple tasks with previously-trained models is a worthwhile endeavour.

Self-supervised model evaluation. Given the difference between the pretext tasks used to train self-supervised models and the downstream tasks used to evaluate them, the comparison between self-supervised approaches has always been a delicate matter. Existing evaluation strategies typically require additional training once a general-purpose representation has been obtained. One strategy keeps the representation fixed, and then trains additional task-specific heads with very limited capacity (e.g. a linear classifier or a regression head for object detection ). A second strategy, instead, leverages SSL to obtain particularly effective initializations, and then proceeds to fine-tune such initialized models on the downstream task of interest. A wider range of tasks can be tested using this setup, such as semantic segmentation and surface normal estimation . In contrast, UniTrack provides a simpler way to evaluate SSL models, one that does not require additional training or fine-tuning. Also, this work is the first to extend SSL evaluation to a set of diverse video tasks. We believe this contribution will allow the study of self-supervised learning methods with a broader scope of applicability. Our work is also related to a line of self-supervised learning methods that learn their representations in a task-agnostic fashion, and then test it on propagation tasks (SOT and VOS). The design of UniTrack is inspired by their task-agnostic philosophy, while significantly extending their scope to a new set of tasks.

Conclusion

Do different tracking tasks require different appearance models? In order to address this question, the proposed UniTrack framework has been instrumental, as it has allowed to easily experiment with alternative representations on a wide variety of downstream problems. Although the answer is not a resounding “no”, as only sometimes a single shared appearance model can outperform dedicated methods, we argue that a unified framework is an appealing alternative to task-specific methods. The main reason is that it allows us to make the most of the progress made in the representation learning literature at no extra cost. With the rapid development of self-supervised learning, and the large amount of computational resources dedicated to it, we believe it is reasonable to expect that, in the future, a general-purpose representation will be able to outperform task-specific methods across the board. Until then, UniTrack could still serve as a useful evaluation tool for novel representations, especially considering the lack of correlation with the standard linear-probe approach. We believe this will encourage the community to develop self-supervised representations that are of “general purpose” in a broader sense. Broader impact. Upon reflection, we believe that progress in tracking applications and self-supervised learning is beneficial for society, as it can significantly impact (for instance) the development of autonomous vehicles, which we consider a net positive for society. We also recognise that the same technologies could constitute a threat if deployed for surveillance by entities hostile to civil liberties.

Funding Transparency Statement

This work was supported by the National Natural Science Foundation of China under Grant No. 61771288, Cross-Media Intelligent Technology Project of Beijing National Research Center for Information Science and Technology (BNRist) under Grant No. BNR2019TD01022 and the research fund under Grant No. 2019GQG0001 from the Institute for Guo Qiang, Tsinghua University.

This work was also supported by the EPSRC grant: Turing AI Fellowship: EP/W002981/1, EPSRC/MURI grant EP/N019474/1. We would also like to thank the Royal Academy of Engineering and FiveAI.

References

Appendix A Datasets and Evaluation Metrics

The table below summarizes the datasets (all publicly available) and evaluation metrics used in this work. In general, to compare with existing task-specific methods, we use the most popular benchmark for each task and report the standard metrics.

For association-type tasks (MOT, MOTS and PoseTrack), we first report the MOTA metric since it highly-correlates with human’s perception in measuring tracking accuracy . However, the MOTA metric disproportionately overweights good detection accuracy . Since most multi-object trackers (included UniTrack) adopt off-the-shelf detectors, it is desirable to also adopt detection-independent measures of performance. For this reason, we also report identity based metrics such as IDF-1 and ID-switch. We also adopt the recently-introduced higher-order HOTA , to replace MOTA and to represent the overall tracking accuracy when comparing self-supervised methods.

For pose tracking, results are averaged for IDF-1 and MOTA, and summed for ID-switch, over 15 key points. In the main text, we only report results for the first five tasks from the table below. For the rest tasks (PoseProp and VIS) we provide additional results in Appendix E. We also provide SOT results on many more recent large-scale datasets in Appendix F.

A single run of the evaluation on five tasks takes about 2 hours in a Titan Xp GPU.

Appendix B Propagation

In order to propagate bounding boxes, we adopt two methods relying on fully-convolutional Siamese networks. Given a target image patch IxI_{{x}} that contains the object of interest, and a search image patch IzI_{{z}} (typically a larger search area in the next frame), the appearance model ϕ\phi processes both patches and outputs their feature maps x=ϕ(Ix){x}=\phi(I_{{x}}) and z=ϕ(Iz){z}=\phi(I_{{z}}).

Cross-correlation (XCorr) head. As in SiamFC , we simply cross-correlate the two feature maps, yielding the response map

Eq. 3 is equivalent to performing an exhaustive search of the pattern x{x} over the search region z{z}. The location of the target object can be determined by finding the maximum value of response map.

Discriminative Correlation Filter (DCF) head. The DCF head is similar to the XCorr head, with two major differences. The first one is that it involves solving a ridge-regression problem to find the template w=ω(x){w}={\omega}({x}) rather than using the original template x{x}, so that the response map is given by

More specifically, the DCF template w=ω(x){w}={\omega}({x}) is a more discriminative template compared with the original template, and is obtained by solving

where yy is an ideal response (here represented as a Gaussian function peaked at the center) and λ≥0\lambda\geq 0 is the regularization coefficient typical of ridge regression. The solution to Eq. 5 can be computed efficiently in the Fourier domain as

where the hat notation x^=F(x)\hat{x}=\mathcal{F}(x) indicates the discrete Fourier Transform of xx, y∗y^{*} represents the complex conjugate of yy and ⊙\odot denotes the Hadamard (element-wise) product. The response map can be computed via inverse Fourier Transform F−1\mathcal{F}^{-1},

Another difference w.r.t the XCorr head is that it is effective to update the template online by simple moving average , i.e. , w^t=αx^t⊙y^∗+(1−α)x^t−1⊙y^∗α(x^t⊙x^t∗+λ)+(1−α)(x^t−1⊙x^t−1∗+λ)\hat{w}_{t}=\frac{\alpha\hat{x}_{t}\odot\hat{y}^{*}+(1-\alpha)\hat{x}_{t-1}\odot\hat{y}^{*}}{\alpha(\hat{x}_{t}\odot\hat{x}_{t}^{*}+\lambda)+(1-\alpha)(\hat{x}_{t-1}\odot\hat{x}_{t-1}^{*}+\lambda)}. In contrast, with the XCorr head every frame is compared against the first one.

As shown in Table 2 and Table 3 from the main paper, for the tested architectures and appearance models we can see a clear advantage of DCF of XCorr (note that the difference was less significant in the original paper, though the experiments were done with a shallower architecture).

Hyper-parameters. Following common practice , we provide the Correlation Filter with a larger region of context in the template patch. To be specific, the template patch IxI_{x} is determined by expanding the height and width of the target bounding box by k=4.5k=4.5 times. The search patch is also determined by expanding the bounding box by same amount, and its center corresponds the latest estimated location of the target. To handle scale variation of the object, we consider s=3s=3 different search patches at different scales 0.985{1,0,1}0.985^{\{1,0,1\}}. Template and search patches are cropped and resized to 520×520520\times 520. This means that with a total stride of r=8r=8, we have feature maps of size 65×6565\times 65. In the DCF head, we set the regularization coefficient to λ=1e−4\lambda=1e^{-4}, and the moving average momentum to α=1e−2\alpha=1e^{-2}.

B.2 Mask and Pose Propagation

In Section 2.3 we introduced the recursive mask propagation as zt=Kt−1tzt−1z_{t}=K_{t-1}^{t}z_{t-1}. In practice, to provide more temporal context, we use a memory bank consisting of multiple former label maps as the source label zmz_{m} instead of a single label map zt−1z_{t-1}, i.e. zt=Kmtzmz_{t}=K_{m}^{t}z_{m}. More specifically, the resulting source label map is obtained by concatenating all the label maps inside the memory bank, zm∈Msz_{m}\in^{Ms}, where ss is the spatial size of a single label map and MM is the size of the memory bank. The softmax computed for KmtK_{m}^{t} is applied over all MsMs points in the memory bank. The memory bank includes the first frame of the video, together with the latest M−1M-1 frames, and we choose M=6M=6. As suggested by MAST and CRW , we also introduce the local attention technique, which restricts the source points considered for each target point to a local circle with radius r=12r=12. The hyper-parameter kk for the kk-NN used when computing the transition matrix KmtK_{m}^{t} is set to k=10k=10.

Propagating pose key points is cast as propagating the mask of each individual key point, represented with the widely adopted Gaussian belief maps . Each Gaussian has mean equal to the corresponding keypoint’s location, and variance proportional to the subject’s body size σ=max(ηsbody,0.5)\sigma=max(\eta s_{body},0.5). The body size is determined by,

where (xp,yp)(x_{p},y_{p}) are the coordinates of the pp-th key point.

Appendix C Association

Motion cues: object states and Kalman Filtering. We employ a Kalman filter with constant velocity and linear motion model to handle motion cues in algorithms of the association type. We assume a generic setting where the camera is not calibrated and the ego-motion is not known. The object states are defined in an eight-dimensional space (u,v,γ,h,u˙,v˙,γ˙,h˙)(u,v,\gamma,h,\dot{u},\dot{v},\dot{\gamma},\dot{h}), where (u,v)(u,v) indicate the position bounding box center, hh the bounding-box height and γ=hw\gamma=\frac{h}{w} the aspect ratio. The latter four dimensions represent the respective velocities of the first four terms.

For the sake of simplicity we convert mask representations to bounding boxes. Let the coordinates of “in-mask” pixels form a set {(xj,yj)∣j=1,...N}\{(x_{j},y_{j})|j=1,...N\}, where NN is the number of mask pixels. Then, the center of the corresponding bounding box is obtained by averaging these coordinates, as (u,v)=1N∑j=1N(xj,yj)(u,v)=\frac{1}{N}\sum_{j=1}^{N}(x_{j},y_{j}). We estimate the height of the bounding box as h=2N∑j=1N∥yj−h∥1h=\frac{2}{N}\sum_{j=1}^{N}\|y_{j}-h\|_{1}. This estimation is analogous to the one suggested in the continuous case . Consider a rectangle with scale (2w,2h)(2w,2h) whose center locates at the origin of a 2D coordinate plane; by integrating over the points inside of the rectangle, we have 1h∫−hh∥y∥1dy=2h∫0hydy=h\frac{1}{h}\int_{-h}^{h}\|y\|_{1}dy=\frac{2}{h}\int_{0}^{h}ydy=h. For objects represented as a pose, we first convert pose keypoints to masks following Appendix C.2, and then convert masks to boxes.

For each timestep, the Kalman Filter predicts current states of existing tracklets. If a new detection is associated to a tracklet, then the state of the detection is used to update the tracklet state. If a tracklet is not associated with any detection, its state is simply predicted without correction.

We use the (squared) Mahalanobis distance to measure the “motion distance” between a newly arrived detection and an existing tracklet. Let us project the state distribution of the ii-th tracklet into the measurement space and denote mean and covariance as μi\bm{\mu}_{i} and Σi\bm{\Sigma}_{i}, respectively. Then, the motion distance is given by

where oj\bm{o}_{j} indicates the observed (4D) state of the jj-th detection. We observe that the Mahalanobis distance consistently outperforms Euclidean distance and IOU distance, likely thanks to the consideration of state estimation uncertainty. Using this metric also allows us to filter out unlikely matches by simply thresholding at 95%95\% confidence interval . We denote the filtering with an indicator function

The threshold η\eta can be computed from the inverse X2\mathcal{X}^{2} distribution. In our case the degrees of freedom of the X2\mathcal{X}^{2} distribution is 4, so the threshold η=9.4877\eta=9.4877.

Association algorithm. Algorithm 1 outlines the association procedure for a single timestamp. The algorithm takes as input a set of tracklets T={1,...,N}\mathcal{T}=\{1,...,N\} and detections D={1,...,M}\mathcal{D}=\{1,...,M\}. First, we predict the current states of the all tracklets using the Kalman Filter. Then we perform the main matching stage. In this stage, we compute a motion cost matrix Cm\bm{C}^{m} using Eq 9, and compute an appearance cost matrix Ca\bm{C}^{a} using the RSM metric described in Section 2.4,

The final cost matrix is the linear combination of the two cost matrices C=λCa+(1−λ)Cm\bm{C}=\lambda\bm{C}^{a}+(1-\lambda)\bm{C}^{m}. We set λ=0.99\lambda=0.99. A Hungarian solver takes the cost matrix C\bm{C} as input and outputs matches [xi,j][x_{i,j}]. We then filter out unrealistic matches using Eq 10. For the remaining tracklets and detections which failed matching, we perform a second matching stage using IOU distance as the cost matrix. Remaining tracklets and detections are output by the association algorithm, further steps (described below) determine if a remaining tracklet should be terminated or if a new identity should be initialized from a remaining detection.

Tracklet termination and initialization. If a tracklet fails to be matched with a newly arrived detection with Algorithm 1, we mark it as inactive. To account for short occlusions, inactive tracklets can still be restored if they are found to be matching with a new detection. We record a “lost age” for each inactive tracklet. If the lost age is greater than a pre-given time, the tracklet would be removed from the current tracklet pool. The lost age is set to 11 second in our experiments.

If a detection fails to match existing tracklets with Algorithm 1, it could correspond to a new tracklet. However, this would result in the creation of frequent brief “spurious” tracklets, containing one detection only. To cope with this issue, similarly to we only initialize a new tracklet if a new detection appears in two consecutive frames (and the IOU between consecutive boxes is at least 0.80.8).

C.2 Pose-to-Mask Conversion

Given the key points’ location of a target person, we convert the pose into a binary mask in two steps. First, the key points are connected to form a skeleton, where the width of each segment forming this skeleton is proportional to the body size with a linear coefficient ηp=0.05\eta_{p}=0.05, and the body size is computed with Eq. 8. Second, we fill closed polygons inside the pose skeleton, since the parts inside the polygon usually belong to the target object.

Appendix D Ablations for the Reconstruction Similarity Metric (RSM)

In Section 2.4 we claimed that the good tracking performance of UniTrack on association-type tasks is largely attributed to the proposed Reconstruction Similarity Metric (RSM). In this section, we provide results of several baseline methods in order to validate the effectiveness of RSM. These baseline are described below.

Global-pooled feature (GPF). Similar to the global feature, but averaging is performed along the sdjs_{d_{j}} dimension to obtain a single feature vector with length CC. Cosine similarity then is computed to measure how likely it is that the two observations belong to the same identity. A large body of re-identification (ReID) approaches employ global-pooled feature (on fully supervised learned feature maps). The benefit and drawback are similar to center feature.

Supervised ReID feature (ReID). For a given image cropped from a bounding box, we employ an strong, off-the-shelf person ReID model to extract a single feature vector with length CC, and compute cosine similarity between observations. The model uses a ResNet-50 architecture and is trained with the joint set of three widely-used datasets: Market-1501 , CUHK-03 , and DukeMTMC-ReID . Using supervised ReID models to extract appearance features is widely used in existing multi-object tracking approaches . Considering large amount of identity labels are leveraged in training, supervised ReID models usually show good association accuracy.

Note that for CF, GF, GPF, and the proposed RSM, we employ the same appearance model (ImageNet pre-trained ResNet-18) for fair comparison. For a broad comparison, we provide results obtained with different detectors and on different datasets. We adopt the following detectors and test on MOT-16 train split (listed with detection accuracy from low to high): DPM , Faster R-CNN (FRCNN), SDP , and FairMOT .

Results are shown in Table 6. We first apply the full association algorithm, i.e. using both appearance and motion cues. In this case (first half of the table), RSM consistently outperforms CF, GF, GPF baselines, and even surpasses the supervised ReID features in several cases, e.g. with FRCNN and FairMOT detectors. In the second half of the table, we show results in which only appearance cues are used, so that the difference between metrics (which are based on appearance) can be better emphasized. In this case, the gaps between different methods are more significant than in the previous case, and RSM still consistently outperforms CF, GF, and GPF. Furthermore, RSM also surpasses the strong supervised ReID feature with all detectors, except for DPM. This suggests that RSM can be an effective similarity metric for tasks that have association at their core.

To show the generality of the results, we also experiment on different datasets and different tasks. Table D shows comparisons on the MOT-20 train split for the MOT task (box observations). The MOT-20 dataset is specialized for the extreme crowded person tracking scenario. Table D presents results on MOTS train split for the MOTS task (mask observations). Note for the MOTS task, since the observations (masks) vary in size, it is not feasible to apply the GF strategy. Results show that the proposed RSM yields significantly higher IDF1 scores on both datasets.

Appendix E More Tracking Tasks

In this section we present two more tasks that UniTrack can address.

The first task is human Pose Propagation on the JHMDB dataset: each video contains a single person of interest, and the pose keypoints are provided in the first frame of the video only. The goal here is to predict the pose of the person throughout the video. Note that this is different from the previously mentioned PoseTrack task: PoseTrack mainly focuses on association between different identities, while in Pose Propagation we aim at propagating the pose of a single identity.

Results are shown in Table E. We report a higher result with ImageNet pre-trained ResNet-18 compared with in previous work (58.3 v.s. 53.8 PCK@0.1). With this result, we observe the best self-supervised method CRW does not beat the ImageNet pre-trained representation by a significant margin (only +0.7+0.7 PCK@1). This again validates our second finding in Section 3.2: a vanilla ImageNet-trained representation is surprisingly effective.

The second task is Video Instance Segmentation (VIS). The problem of VIS is similar to Multiple Object Tracking and Segmentation (MOTS), but its setup differs in the following aspects: first, the object categories are fairly diverse (40 different categories), while in MOTS objects are mostly persons and vehicles. This also requires the trackers tackling the VIS task to handle objects from different classes within the same scene. Second, the evaluation metrics are different. In MOTS, the MOT-like metrics (CLEAR , IDF-1/IDs, and HOTA ) are used, which implicitly encourages methods to focus on outputting temporally consistent trajectories. Instead, for VIS the evaluation metric is spatial-temporal mAP, a temporal extension of the vanilla mAP which is usually used in detection and segmentation tasks. The mAP metric significantly biases towards segmentation and classification accuracy in single frames, thus being less informative for evaluating “tracking” accuracy.

Results on VIS task are shown in Table E. We adopt an identical segmentation model to the one of MaskTrackRCNN , and observe only a 0.20.2 difference in mAP. For further comparison, we also provide results of two other association methods, OSMN and DeepSORT , providing them with the same observations as used by UniTrack. Note how UniTrack boasts better accuracy than both methods (30.030.0 v.s. 27.527.5 and 26.126.1 mAP). Comparing with an state-of-the-art model, SipMask , our result is also comparable with −2.4-2.4 point mAP. We believe if equipped with more advanced single frame segmentation model, the mAP would be further improved.

Appendix F SOT results on more datasets

To further validate the general validity of our experiments, we provide more results for the SOT task by testing on more recent datasets that contain large-scale and long-term videos.

The results in Table 11 show a very similar trend to the one already observed for OTB (Table 3e in the main text): For the SOT task, UniTrack with ImageNet features has comparable performance to the one of the recent LUDT+, which like UniTrack does not require task-specific supervision, but can only be used for SOT. Again, similarly to what was reported for OTB, UniTrack is outperformed by recent methods such as SiamRPN++. This is to be expected, as SiamRPN++ is specifically designed for SOT and trained in a supervised fashion on several large-scale video datasets.

Appendix G Additional Correlation Studies

In Section 3.3 (main paper) we investigated the correlation between tracking performance and ImageNet “linear probe” accuracy for different SSL models. In this section, we provide more results and discussions by studying the correlation between tracking performance and several other downstream tasks when using the appearance model from the many SSL methods under consideration. For non-tracking tasks, we report numbers from and plot them against tracking performance in Figure 8.

We report three tasks: surface normal estimation on the NYUv2 dataset, where the mean angular error is used as the evaluation metric (the lower the better); Object detection on Pascal VOC , with performance measured in mAP (the higher the better); Semantic segmentation on ADE20k dataset, with performance measured in mean IOU (the higher the better). In each subfigure, we plot the performance of five tracking tasks along the y-axes, and performance of the other task along the x-axes. Note that we actually use negative mean error for surface normal estimation, to represent accuracy. As in the main paper, we compute two types of correlation coefficient: Spearman’ rr and Pearson’s ρ\rho, and report them in the left bottom corner of each plot. Several interesting findings can be observed:

(a) Correlation between tracking and surface normal prediction performance is fairly strong. Results are shown in Figure 8(a). For instance, r=0.70r=0.70 for surface normal error v.s. MOT accuracy, and 0.560.56 for surface normal error v.s. PoseTrack accuracy. Interestingly, the behavior of SOT is in contrast with MOT and PoseTrack: SOT accuracy is moderately negative correlated (r=−0.50r=-0.50) with surface normal estimation accuracy. VOS presents a similar trend to the one of SOT, but with a lower correlation coefficient.

(b) Object detection is moderately correlated with association-type tracking tasks. For object detection, we consider two setups: one is to freeze the representation and only train the additional classification/regression head; the other is to finetune the whole network in an end-to-end manner. Results are shown in Figure 8(b) and 8(c) respectively. In general, MOT and PoseTrack are moderately correlated with object detection under the frozen setting (r=0.48r=0.48 for MOT and and r=0.42r=0.42 for PoseTrack), and MOTS is moderately correlated with object detection under the finetune setting (r=0.51r=0.51). Propagation-type tasks are poorly correlated with object detection results under both settings (∣ρ∣<0.10|\rho|<0.10). We speculate that, in this case, positive correlation might be due to the fact that both object detection and association-type tracking require discriminative features at the level of the object.

(c) Semantic segmentation is slightly negative correlated with tracking tasks. As can be observed in Figure 8(d), correlation coefficients between segmentation accuracy and tracking performance are mildly negative. Among these results, VOS is the task that is most (negatively) correlated with segmentation, with r=−0.50r=-0.50. MOTS and PoseTrack are also mildly correlated, with r=−0.41r=-0.41 and r=−0.25r=-0.25 respectively. We speculate that negative correlation might be cause to the fact that tracking and segmentation require features with contradictory properties. Consider two different instances that belongs to the same category, i.e. two different pedestrian. For segmentation, the task requires pixel-wise classification, meaning that pixels inside the two instances should be equally classified into the same “pedestrian” class, thus their features should be similar (close to the class center). In contrast, for tracking tasks, it is required to distinguish different instances from the same class, otherwise a tracker would easily fail when objects overlap with each other. Therefore, point features inside the two different pedestrian are expected to be dissimilar.