Joint Feature Learning and Relation Modeling for Tracking: A One-Stream Framework
Botao Ye, Hong Chang, Bingpeng Ma, Shiguang Shan, Xilin Chen
Introduction
Visual object tracking (VOT) aims at localizing an arbitrary target in each video frame, given only its initial appearance. The continuously changing and arbitrary nature of the target poses a challenge to learn a target appearance model that can effectively discriminate the specified target from the background. Current mainstream trackers typically address this problem with a common two-stream and two-stage pipeline, which means that the features of the template and the search region are separately extracted (two-stream), and the whole process is divided into two sequential steps: feature extraction and relation modeling (two-stage). Such a natural pipeline employs the strategy of “divide-and-conquer” and achieves remarkable success in terms of tracking performance.
However, the separation of feature extraction and relation modeling suffers from the following limitations. Firstly, the feature extracted by the vanilla two-stream two-stage framework is unaware of the target. In other words, the extracted feature for each image is determined after off-line training, since there is no interaction between the template and the search region. This is against with the continuously changing and arbitrary nature of the target, leading to limited target-background discriminative power. On some occasions when the category of the target object is not involved in the training dataset (\ie, one-shot tracking), the above problems are particularly serious. Secondly, the two-stream, two-stage framework is vulnerable to the performance-speed dilemma. According to the computation burden of the feature fusion module, two different strategies are commonly utilized. The first type, shown in Fig 2(a), simply adopts one single operator like cross-correlation or discriminative correlation filter , which is efficient but less effective since the simple linear operation leads to discriminative information loss . The second type, shown in Fig 2(b), addresses the information loss by complicated non-linear interaction (Transformer ), but is less efficient due to a large number of parameters and the use of iterative refinement (\eg, for each search image, STARK-S50 takes 7.5 ms for the feature extraction and 14.1 ms for relation modeling on an RTX2080Ti GPU).
In this work, we set out to address the aforementioned problems via a unified one-stream one-stage tracking framework. The core insight of the one-stream framework is to bridge a free information flow between the template and search region at the early stage (\ie, the raw image pair), thus extracting target-oriented features and avoiding the loss of discriminative information. Specifically, we concatenate the flattened template and search region and feed them into staked self-attention layers (widely used Vision Transformer (ViT) is chosen in our implementation), and the produced search region features can be directly used for target classification and regression without further matching. The staked self-attention operations enable iteratively feature matching between the template and the search region, thus allowing mutual guidance for target-oriented feature extraction. Therefore, both template and search region features can be extracted dynamically with strong discriminative power. Additionally, the proposed framework achieves a good balance between performance and speed because the concatenation of the template and the search region makes the one-stream framework highly parallelizable and does not require additional heavy relational modeling networks.
Moreover, the proposed one-stream framework provides a strong prior about the similarity of the target and each part of the search region (\iecandidates) as shown in Fig. 4, which means that the model can identify background regions even at the early stage. This phenomenon verifies the effectiveness of the one-stream framework and motivates us to propose an in-network early candidate elimination module for progressively identifying and discarding the candidates belonging to the background in a timely manner. The proposed candidate elimination module not only significantly boosts the inference speed, but also avoids the negative impact of uninformative background regions on feature matching.
Despite its simple structure, the proposed trackers achieve impressive performance and set a new state-of-the-art (SOTA) on multiple benchmarks. Moreover, it maintains adorable inference efficiency and shows faster convergence compared to SOTA Transformer based trackers. As shown in Fig. 1, our method achieves a good balance between the accuracy and inference speed.
The main contributions of this work are three-fold: (1) We propose a simple, neat, and effective one-stream, one-stage tracking framework by combining the feature extraction and relation modeling. (2) Motivated by the prior of the early acquired similarity score between the target and each part of the search region, an in-network early candidate elimination module is proposed for decreasing the inference time. (3) We perform comprehensive experiments to verify that the one-stream framework outperforms the previous SOTA two-stream trackers in terms of performance, inference speed, and convergence speed. The resulting tracker OSTrack sets a new state-of-the-art on multiple tracking benchmarks.
Related Work
In this section, we briefly review different tracking pipelines, as well as the adaptive inference methods related to our early candidate elimination module.
Tracking Pipelines. Based on the different computational burdens of feature extraction and relation modeling networks, we compare our method with two different two-stream two-stage archetypes in Fig. 2. Earlier Siamese trackers and discriminative trackers belong to Fig. 2(a). They first extract the features of the template and the search region separately by a CNN backbone , which shares the same structure and parameters. Then, a lightweight relation modeling network (\eg, the cross-correlation layer in Siamese trackers and correlation filter in discriminative trackers) takes responsibility to fuse these features for the subsequent state estimation task. However, the template feature cannot be adjusted according to the search region feature in these methods. Such a shallow and unidirectional relation modeling strategy may be insufficient for information interaction. ARecently, stacked Transformer layers are introduced for better relation modeling. These methods belong to Fig. 2(b) where the relation modeling module is relatively heavy and enables bi-directional information interaction. TransT proposes to stack a series of self-attention and cross-attention layers for iterative feature fusion. STARK concatenates the pre-extracted template and search region features and feeds them into multiple self-attention layers. The bi-directional heavy structure brings performance gain but inevitably slows down the inference speed. Differently, our one-stream one-stage design belongs to Fig. 2(c). For the first time, we seamlessly combine feature extraction and relation modeling into a unified pipeline. The proposed method provides free information flow between the template and search region with minor computation costs. It not only generates target-oriented features by mutual guidance but also is efficient in terms of both training and testing time.
Adaptive Inference. Our early candidate elimination module can be seen as a progressive process of adaptively discarding potential background regions based on the similarity between the target and the search region. One related topic is the adaptive inference in vision transformers, which is proposed to accelerate the computation of ViT. DynamicViT trains extra control gates with the Gumbel-softmax trick to discard tokens during inference. Instead of directly discarding non-informative tokens, EViT fuses them to avoid potential information loss. These works are tightly coupled with the classification task and are therefore not suitable for tracking. Instead, we treat each token as a target candidate and then discard the candidates that are least similar to the target by means of a free similarity score calculated by the self-attention operation. To the best of our knowledge, this is the first work that attempts to eliminate potential background candidates within the tracking network.
Method
This section describes the proposed one-stream tracker (OSTrack). The input image pairs are fed into a ViT backbone for simultaneous feature extraction and relation modeling, and the resulting search region features are directly adopted for subsequent target classification and regression. An overview of the model is shown in Fig. 3(a).
To verify whether adding addition identity embeddings (to indicate a token belonging to the template or search region as in BERT ) or adopting relative positional embeddings are beneficial to the performance, we also conduct ablation studies and observe no significant improvement, thus they are omitted for simplicity (details can be found in the supplementary material).
Token sequences and are then concatenated as , and the resulting vector is then fed into several Transformer encoder layers . Unlike the vanilla ViT , we insert the proposed early candidate eliminating module into some of encoder layers as shown in Fig. 3(b) for inference efficiency, and the technical details are presented in Sec. 3.2. Notably, adopting the self-attention of concatenated features makes the whole framework highly parallelized compared to the cross-attention . Although template images are also fed into the ViT for each search frame, the impact on the inference speed is minor due to the highly parallel structure and the fact that the number of template tokens is small compared to the number of search region tokens.
Analysis. From the perspective of the self-attention mechanism , we further analyze the intrinsic reasons why the proposed framework is able to realize simultaneous feature extraction and relation modeling. The output of self-attention operation in our approach can be written as:
where is a measure of similarity between the template and the search region, and the rest are similar. The output can be further written as:
In the right part of Eq. 5, is responsible for aggregating the iter-image feature (relation modeling) and aggregating the intra-image feature (feature extraction) based on the similarity of different image parts. Therefore, the feature extraction and relation modeling can be done with a self-attention operation. Moreover, Eq. 5 also constructs a bi-direction information flow that allows mutual guidance of target-oriented feature extraction through the similarity learning.
Comparisons with Two-Stream Transformer Fusion Trackers. 1) Previous two-stream Transformer fusion trackers all adopt a Siamese framework, where the features of the template and search region are separately extracted first, and the Transformer layer is only adopted to fuse the extracted features. Therefore, the extracted features of these methods are not adaptive and may lose some discriminative information, which is irreparable. In contrast, OSTrack directly concatenates linearly projected template and search region images at the first stage, so feature extraction and relation modeling are seamlessly integrated and target-oriented features can be extracted through the mutual guidance of the template and the search region. 2) Previous Transformer fusion trackers only employ ImageNet pre-trained backbone networks and leave Transformer layers randomly initialized, which degrades the convergence speed, while OSTrack benefits from pre-trained ViT models for faster convergence. 3) The one-stream framework provides the possibility of identifying and discarding useless background regions for further improving the model performance and inference speed as presented in Sec. 3.2.
2 Early Candidate Elimination
Each token of the search region can be regarded as a target candidate and each template token can be considered as a part of the target object. Previous trackers keep all candidates during feature extraction and relation modeling, while background regions are not identified until the final output of the network (\ie, classification score map). However, our one-stream framework provides a strong prior on the similarity between the target and each candidate. As shown in Fig. 4, the attention weights of the search region highlight the foreground objects in the early stage of ViT (\eg, layer 4), and then progressively focus on the target. This property makes it possible to progressively identify and eliminate candidates belonging to the background regions inside the network. Therefore, we propose an early candidate elimination module that progressively eliminates candidates belonging to the background in the early stages of ViT to lighten the computational burden and avoid the negative impact of noisy background regions on feature learning.
Candidate Elimination. Recall that the self-attention operation in ViT can be seen as a spatial aggregation of tokens with normalized importances , which is measured by the dot product similarity between each token pair. Specifically, each template token is calculated as:
where , , and denote the query vector of token , the key matrix corresponding to the template, the key matrix corresponding to the search region and the value matrix. The attention weight determines the similarity between the template part and all search region tokens (candidates). The item (, is the number of input search region tokens) of determines the similarity between and the candidate. However, the input templates usually include background regions that introduce noise when calculating the similarity between the target and each candidate. Therefore, instead of summing up the similarity of each candidate to all template parts , , we take ( token corresponding to the center part of the original template image) as the representative similarity. This is fairly reasonable as the center template part has aggregated enough information through self-attention to represent the target. In the supplementary, we compare the effect of different template token choices. Considering that multi-head self-attention is used in ViT, there are multiple similarity scores , where and is the total number of attention heads . We average the similarity scores of all heads by , which serves as the final similarity score of the target and each candidate. One candidate is more likely to be a background region if its similarity score with the target is relatively small. Therefore, we only keep the candidates corresponding to the largest (top-) elements in ( is a hyperparameter, and we define the token keeping ratio as ), while the remaining candidates are eliminated. The proposed candidate elimination module is inserted after the multi-head attention operation in the encoder layer, which is illustrated in Fig. 3(b). In addition, the original order of all remaining candidates is recorded so that it can be recovered in the final stage.
Candidate Restoration. The aforementioned candidate elimination module disrupts the original order of the candidates, making it impossible to reshape the candidate sequence back into the feature map as described in Sec. 3.3, so we restore the original order of the remaining candidates and then pad the missing positions. Since the discarded candidates belong to the irrelevant background regions, they will not affect the classification and regression tasks. In other words, they just act as placeholders for the reshaping operation. Therefore, we first restore the order of the remaining candidates and then zero-pad in between them.
Visualization. To further investigate the behavior of the early candidate elimination module, we visualize the progressive process in Fig. 5. By iteratively discarding the irrelevant tokens in the search region, OSTrack not only largely lightens the computation burden but also avoids the negative impact of noisy background regions on feature learning.
3 Head and Loss
We first re-interpret the padded sequence of search region tokens to a 2D spatial feature map and then feed it into a fully convolutional network (FCN), which consists of stacked Conv-BN-ReLU layers for each output. Outputs of the FCN contain the target classification score map , the local offset to compensate the discretization error caused by reduced resolution and the normalized bounding box size (\iewidth and height) . The position with highest classification score is considered to be target position, \ie, and the finial target bounding box is obtained as:
where and are the regularization parameters in our experiments as in .
Experiments
After introducing the implementation details, this section first presents a comparison of OSTrack with other state-of-the-art methods on seven different benchmarks. Then, ablation studies are provided to analyze the impact of each component and different design choices.
Our trackers are implemented in Python using PyTorch. The models are trained on 4 NVIDIA A100 GPUs and the inference speed is tested on a single NVIDIA RTX2080Ti GPU.
Model. The vanilla ViT-Base model pre-trained with MAE is adopted as the backbone for joint feature extraction and relation modeling. The head is a lightweight FCN, consisting of 4 stacked Conv-BN-ReLU layers for each of three outputs. The keeping ratio of each candidate elimination module is set as 0.7, and a total of three candidate elimination modules are inserted at layers 4, 7, and 10 of ViT respectively, following . We present two variants with different input image pair resolution for showing the scalability of OSTrack:
OSTrack-256. Template: 128128 pixels; Search region: 256256 pixels.
OSTrack-384. Template: 192192 pixels; Search region: 384384 pixels.
Training. The training splits of COCO , LaSOT , GOT-10k (1k forbidden sequences from GOT-10k training set are removed following the convention ) and TrackingNet are used for training. Common data augmentations including horizontal flip and brightness jittering are used in training. Each GPU holds 32 image pairs, resulting in a total batch size of 128. We train the model with AdamW optimizer , set the weight decay to , the initial learning rate for the backbone to and other parameters to , respectively. The total training epochs are set to 300 with 60k image pairs per epoch and we decrease the learning rate by a factor of 10 after 240 epochs.
Inference. During inference, Hanning window penalty is adopted to utilize positional prior in tracking following the common practice . Specifically, we simply multiply the classification map by the Hanning window with the same size, and the box with the highest score after multiplication will be selected as the tracking result.
2 Comparison with State-of-the-arts
To demonstrate the effectiveness of the proposed models, we compare them with state-of-the-art (SOTA) trackers on seven different benchmarks.
GOT-10k. GOT-10k test set employs the one-shot tracking rule, \ie, it requires the trackers to be trained only on the GOT-10k training split, and the object classes between train and test splits are not overlapped. We follow this protocol to train our model and evaluate the results by submitting them to the official evaluation server. As reported in Tab. 1, OSTrack-384 and OSTrack-256 outperform SwinTrack-B by 1.6% and 4.3% in AO. The SR0.75 score of OSTrack-384 reaches 70.8%, outperforming SwinTrack-B by 6.5%, which verifies the capability of our trackers in both accurate target-background discrimination and bounding box regression. Moreover, the high performance on this one-shot tracking benchmark demonstrates that our one-stream tracking framework can extract more discriminative features for unseen classes by mutual guidance.
LaSOT. LaSOT is a challenging large-scale long-term tracking benchmark, which contains 280 videos for testing. We compare the result of the OSTrack with previous SOTA trackers in Tab. 1. The results show that the proposed tracker with smaller input resolution, \ie, OSTrack-256, already obtains comparable performance with SwinTrack-B . Besides, OSTrack-256 runs at a fast inference speed of 105.4 FPS, being 2x faster than SwinTrack-B (52 FPS), which indicates that OSTrack achieves an excellent balance between accuracy and inference speed. By increasing the input resolution, OSTrack-384 further improves the AUC on LaSOT to 71.1% and sets a new state-of-the-art.
TrackingNet. The TrackingNet benchmark contains 511 sequences for testing, which covers diverse target classes. Tab. 1 shows that OSTrack-256 and OSTrack-384 surpass SwinTrack-B by 0.6% and 1.4% in AUC separately. Moreover, both models are faster than SwinTrack-B.
LaSOT. LaSOT is a recently released extension of LaSOT, which consists of 150 extra videos from 15 object classes. Tab 1 presents the results. Previous SOTA tracker KeepTrack designs a complex association network and runs at 18.3 FPS. In contrast, our simple one-stream tracker OSTrack-256 shows slightly lower performance but runs at 105.4 FPS. OSTrack-384 sets a new state-of-the-art AUC score of 50.5% while runs in 58.1 FPS, which is 2.3% higher in AUC score and 3x faster in speed.
NFS, UAV123 and TNL2K. We also evaluate our tracker on three additional benchmarks: NFS , UAV123 and TNL2K includes 100, 123, and 700 video sequences, separately. The results in Tab. 2 show that OSTrack-384 achieves the best performance on all three benchmarks, demonstrating the strong generalizability of OSTrack.
3 Ablation Study and Analysis
The Effect of Early Candidate Elimination Module Tab. 1 shows that increasing the input resolution of the input image pairs can bring significant performance gain. However, the quadratic complexity with respect to the input resolution makes simply increasing the input resolution unaffordable in inference time. The proposed early candidate elimination module addresses the above problem well. We present the effect of the early candidate elimination module from the aspects of inference speed (FPS), multiply-accumulate computations (MACs), and tracking performance on multiple benchmarks in Tab. 3. The effect on different input search region resolutions is also presented. Tab. 3 shows that the early candidate elimination module can significantly decrease the calculation and increase the inference speed, while slightly boosting the performance in most cases. This demonstrates that the proposed module alleviates the negative impact brought by the noisy background regions on feature learning. For example, adding the early candidate elimination module in OSTrack-256 decreases the MACs by 25.9% and increases the tracking speed by 13.2%, and the LaSOT AUC is increased by 0.4%. Furthermore, larger input resolution benefits more from this module, \eg, OSTrack-384 shows a 40.3% increase in speed.
Different Pre-training Methods. While previous Transformer fusion trackers random initialize the weights of Transformer layers, our joint feature learning and relation modeling module can directly benefit from the pre-trained weights. We further investigate the effect of different pre-training methods on the tracking performance by comparing four different pre-training strategies: no pre-training; ImageNet-1k pre-trained model provided by ; ImageNet-21k pre-trained model provided by ; unsupervised pre-training model MAE . As the results in Tab. 4 show, pre-training is necessary for the model weights initialization. Interestingly, we also observe that the unsupervised pre-training method MAE brings better tracking performance than the supervised pre-training ones using ImageNet. We hope this can inspire the community for designing better pre-training strategies tailored for the tracking task.
Aligned Comparison with SOTA Two-stream Trackers. One may wonder whether the performance gain is brought by the proposed one-stream structure or purely by the superiority of ViT. We thus compare our method with two SOTA two-stream Transformer fusion trackers by eliminating the influencing factors of backbone and head structure. To be specific, we align two previous SOTA two-stream trackers (STRAK-S and SwinTrack ) with ours for fair comparison as follows: replacing their backbones with the same pre-trained ViT and setting the same input resolution, head structure, and training objective as OSTrack-256. The remaining experimental settings are kept the same as in the original paper. As shown in Tab. 5, our re-implemented two-stream trackers show comparable or stronger performance compared to the initially published performance, but still lag behind OSTrack, which demonstrates the effectiveness of our one-stream structure. We also observe that OSTrack significantly outperforms the previous two-stream trackers on the one-shot benchmark GOT-10k, which further proves the advantage of our one-stream framework in the challenging scenario. Actually, the discriminative power of features extracted by the two-stream framework is limited since the object classes in the testing set are completely different from the training set. Whereas, by iterative interaction between the features of the template and search region, OSTrack can extract more discriminative features through mutual guidance. Different from the two-stream SOTA trackers, OSTrack neglects the extra heavy relation modeling module while still keeping the high parallelism of joint feature extraction and relation modeling module. Therefore, when the same backbone network is adopted, the proposed one-stream framework is much faster than STARK (40.2 FPS faster) and SwinTrack (25.6 FPS faster). Besides, OSTrack requires fewer training image pairs to converge.
Discriminative Region Visualization. To better illustrate the effectiveness of the proposed one-stream tracker, we visualize the discriminative regions of the backbone features extracted by OSTrack and a SOTA two-stream tracker (SwinTrack-aligned) in Fig. 6. As can be observed, due to the lack of target awareness, features extracted by the backbone of SwinTrack-aligned show limited target-background discriminative power and may lose some important target information (\eg, head and helmet in Fig. 6), which is irreparable. In contrast, OSTrack can extract discriminative target-oriented features, since the proposed early fusion mechanism enables relation modeling between the template and search region at the first stage.
Conclusion
This work proposes a simple, neat, and high-performance one-stream tracking framework based on Vision Transformer, which breaks out of the Siamese-like pipeline. The proposed tracker combines the feature extraction and relation modeling tasks, and shows a good balance between performance and inference speed. In addition, we further propose an early candidate elimination module that progressively discards search region tokens belonging to background regions, which significantly boosts the tracking inference speed. Extensive experiments show that the proposed one-stream trackers perform much better than previous methods on multiple benchmarks, especially under the one-shot protocol. We expect this work can attract more attention to the one-stream tracking framework.
Acknowledgments. This work is partially supported by Natural Science Foundation of China (NSFC): 61976203 and 61876171. Thanks Zhipeng Zhang for his helpful suggestions.
References
Appendix
A More Implementation Details
Training Details. In OSTrack-256, the input sizes of templates and search regions are pixels and pixels respectively, corresponding to and times of the target bounding box area. In OSTrack-384, the input sizes of templates and search regions are pixels and pixels, corresponding to and times of the target bounding box area. For the GOT-10k test benchmark , which requires training the models with only the training split of GOT-10k (one-shot setting), we set the total training epoch to 100 with 60k image pairs per epoch, and we decrease the learning rate by a factor of 10 after 80 epochs. The other settings are kept consistent with the models trained with all datasets.
where and are hyper-parameters and we set and as in .
Position Embeddings. The length of the position embeddings in the pre-trained ViT is different from the length of the input template and search region embeddings. Therefore, the pre-trained positional embeddings are interpolated (2D bicubic interpolation is adopted) to the sizes of the template and search region embeddings separately, which are further added to the patch embeddings.
Model Details. In Sec. 4.3, we compare our OSTrack (without the early candidate elimination module) with aligned two-stream trackers (\ie, STARK-aligned and SwinTrack-aligned), and we further present the detailed structures in this section. The proposed one-stream framework, as shown in Fig. A1(a), combines feature extraction and relation modeling modules into a single ViT backbone. The aligned two-stream framework, as shown in Fig. A1(b), first extracts features of the template and the search region separately with the same ViT backbone and then models the feature relation with several extra Transformer encoder layers. As presented in Sec. 4.3, this relation modeling module is instantiated with the encoder structure proposed in STARK (STARK-aligned) and SwinTrack (SwinTrack-aligned) separately.
Discriminative Regions Visualization. We show the method used to obtain the visualization of the discriminative regions in Fig. 6. Zagoruyko \etal show that the importance of a hidden neuron activation can be indicated by its absolute value. In this work, we adopt a similar approach to obtain the discriminative regions of each feature map . We first calculate the absolute mean values of each pixel along the channel dimension:
where is the number of channels. Then, the relative importance of each pixel is then calculated by:
where and return the maximum and minimum values of all pixels, respectively.
B More Ablation Studies
As pointed out in Sec. 3.2, the goal of the early candidate elimination module is to identify and discard candidates belonging to background regions based on the ranking of similarity between the target and each candidate. However, the input template also contains background regions, which introduces noisy information when calculating the similarity score. Therefore, different choices of template parts (tokens) used for the similarity calculation may influence the candidate elimination results and consequently affect the tracking performance. We compare four different template token choices (the similarity scores of all chosen template tokens are summed up for the final ranking): 1) all template tokens; 2) all template tokens within the ground truth target bounding box; 3) template tokens within a 4x4 region around the center of the template image; 4) the template token corresponding to the center of the template image. The result comparison of these template choices is shown in Tab. A1. The results demonstrate that different template token choices do affect the quality of identifying background candidates. Since the input template contains background regions, directly using “All Template Tokens” clearly degrades the tracking performance compared with the baseline (“No Early Candidate Elimination”), \ie, 0.6% lower in LaSOT AUC. Compared to other choices, using the central template token shows better performance, probably because the central token does not contain any background region and has aggregated the entire target information through self-attention.
B.2 Identity Embeddings and Relative Positional Embeddings
We additionally verify the effect of adding identity embeddings and relative positional embeddings. Specifically, for the identity embeddings, we add learnable identity embeddings (to indicate a token belonging to the template or search region as in BERT ) to template tokens and search region tokens separately. For the relative positional embeddings, the same method as in SwinTrack is adopted. The results are presented in Tab. A2, these two components do not bring performance gain compared to the original design, thus not adopted in our model.
B.3 Additional Relation Modeling Module
To investigate whether our one-stream framework does not require an extra feature relation module, we add an additional transformer-based feature fusion module proposed in , which consists of 4 self-attention layers and 1 cross-attention layer, to further fusion the extracted template and search region features. As the results in Tab. A3 show, adding such a relation modeling module instead degrades the tracking performance, indicating that the output search region features of the ViT backbone have been sufficiently fused with the template features.
B.4 Fewer Relation Modeling Layers
In the implementation of vanilla OSTrack, all encoder layers in ViT-Base (12 layers in total) are used for simultaneous feature extraction and relation modeling. In this subsection, we try to decrease the number of layers used for relation modeling. Specifically, only the last encoder layers are used for simultaneous feature extraction and relation modeling, and the first layers are only used for the template and search region feature extraction. is set to be 6 and 3 separately and the results are presented in Tab. A4. The results show that using fewer encoder layers for simultaneous feature extraction and relation modeling will degrade the tracking performance, showing the necessity of sufficient feature fusion.
B.5 Different Token Drop Rate
We also try to apply a different keeping ratio for the early candidate elimination module. As the results in Tab. A5 show, using leads to performance drop on the LaSOT tracking benchmark since small may cause a significant information loss. However, the reduction in computational cost that comes with large is limited. Setting shows a decent decrease in computational cost with a slight improvement in tracking performance. Therefore, we use in our experiments.
C Results on VOT2020
VOT2020 is a challenging short-term tracking benchmark that is evaluated by target segmentation results. To evaluate OSTrack on VOT2020, we use AlphaRefine to generate segmentation masks, and the results are shown in Tab. A6. Since the wide existence of distractors in VOT2020, updating the template during the tracking process has become a common practice to avoid tracking drift, which can bring significant performance gain (\eg, STARK-ST50 citestark raises the EAO from 0.462 to 0.505 by simply adding a dynamic template). OSTrack-256 obtains an EAO of 0.518, which already outperforms the STARK-ST50 with an online template updating mechanism. This demonstrates the great potential of OSTrack which serves as a neat and strong baseline model.
D Results on ITB
ITB benchmark is a newly collected benchmark with 9 representative scenarios and 180 diverse videos, which contains more informative tracking sequences. Tab. A7 shows the results of OSTrack compared with other SOTA tackers. Our OSTrack-384 achieves 64.8% in mIoU, surpassing the previous best tracker STARK by a large margin (7.2%).
E More Visualization
We first provide more visualization results for attention weights of the search region corresponding to the center part of the template (which can be seen as the target) in Fig. A2. The results show that the model attends to the foreground objects at an early stage (see “Layer 4” in Fig. A2) and finally shows great discriminative power between the target and distractors (see “Layer 12” in Fig. A2). These phenomenons demonstrate that the proposed OSTrack can extract target-oriented features with strong target-distractor discriminability.
In Fig. A3, more visualization results of the early candidate elimination module are presented. The results validate that the proposed method can effectively identify and discard background regions under various target categories and challenge scenarios (\eg, target deformation, occlusion, motion blur, \etc).