Transformer Tracking with Cyclic Shifting Window Attention

Zikai Song, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang

Introduction

Visual object tracking (VOT) is one of the fundamental problems in computer vision research with a wide range of applications in video surveillance, autonomous vehicles, human-machine interaction, and others. It aims to estimate the position of a target object in each video frame, commonly represented as a bounding box encapsulating the target. The target object is given as a template in the initial frame, and the tracker is required to extract proper features about the target and localize the target in the following frames. Most of the popular trackers adopt the Siamese network structure, which conducts tracking by calculating the similarity between the template and search region in the current frame. The similarity metric of cross-correlation used in Siamese trackers is prone to lose much semantic information for it is a single-level linear computational process. This deficiency can be well tackled by using the attention mechanism to learn the global context. Recently, transformer-based approaches have reported new state-of-the-art performance on image recognition, object detection, and semantic segmentation benchmarks. This is no wonder as transformer has a powerful cross-attention mechanism to reasoning between patches. Particularly, transformer trackers have shown their great strength by introducing the attention mechanism to enhance and fuse the features of the target and the tracked object. However, we observe that these transformer trackers simply put the flattened features of the template and search region into pixel-level attention, each pixel of a flattened feature (Query) matches all pixels of another flattened feature (Key) in a complete and disordered manner, as shown in Figure 1(b). This pixel-level attention destroys the integrity of the target object and leads to information loss of relative positions between pixels.

In this paper, we propose a novel multi-scale cyclic shifting window transformer for visual object tracking to further lift pixel-level attention to window-level attention, calculating attention between indivisible windows by treating each window as a whole keeps the location information within the window. The proposed method is inspired by the seminal work of the Swin Transformer, which adopts a hierarchical transformer structure by starting from small-sized patches and gradually increasing the size through merging to achieve a broader receptive field. Different from the Swin Transformer, we calculate the cross-window attention between the template and search region directly, which helps to discriminate the target from the background by ensuring the integrity of the object. Further, we propose a multi-head multi-scale attention where each head of the transformer measures the relevance among partitioned windows at a specific scale. The key idea here is to apply a cyclic shifting strategy on each window, as shown in Figure 1(a), for generating more accurate attention results. To address performance drop around boundaries caused by the cyclic shifting operation, we design a spatially regularized attention mask which turns out to be very effective in alleviating the boundary artifacts. Finally, we present some efficient computation strategies to avoid redundant computation introduced by multi-scale cyclic shifting windows, which greatly reduce the time and computational cost. Extensive experiments demonstrate that our tracker performs remarkably better than other state-of-the-art algorithms.

To summarize, our main contributions include:

We propose a novel transformer architecture with multi-scale cyclic shifting window attention for visual object tracking, uplifting the original pixel-level attention to the new deliberately designed window-level attention. The cross-window attention ensures the integrity of the tracking object, and the cyclic shifts bring greater accuracy by expanding window samples.

We design a spatially regularized attention mask and some computational optimization strategies to improve the accuracy and speed of the window attention. Specifically, a spatially regularized attention mask is used to address performance drop around boundaries caused by the cyclic shifts, and we propose three computational optimization strategies to remove redundant computations.

Related Work

Visual object tracking. Existing visual object tracking approaches can be roughly divided into two categories, the Correlation Filter (CF) based trackers and Deep Neural Network (DNN) based trackers. CF based approaches exploit the convolution theorem and train a filter in the Fourier domain that maps known target images to the desired output. The filter is learned through circular shifting patches around the target object to discriminate background against the target. DNN based trackers refer to the methods adopting deep neural networks in the tracking process. Many methods treat the tracking task as a basic recognition task, i.e., using a convolutional backbone network to extract features and locate the target by classification heads in the form of fully connected layers.

Recent years, tracking algorithms that adopt a Siamese network structure have shown great success. A Siamese network usually consists of two branches, one for template and the other for search regions, and similarities between them are reported through cross-correlations. However, such a strategy is unable to effectively explore the semantic correlation between template and search regions. This issue leads to the further exploration of using the powerful cross-attention mechanism of transformer structure for object tracking.

Vision transformers. Vaswani et alet~{}al. propose the very first Transformer structure for handling long-range dependencies in Natural Language Processing (NLP). The basic block in a transformer is the attention module, which takes a sequence as input and measures the relevance of different parts of the sequence, aggregating the global information from the input sequence. Transformer not only conducts the self-attention within a single input but also calculates the cross-attention between different inputs. ViT first introduces transformer to image recognition tasks. Ever since, transformer has been widely applied in image classification, object detection, semantic segmentation, visual object tracking and etc.

The seminal work of Swin Transformer proposes an effective hierarchical architecture with shifted windows and achieves the state-of-the-art performance on COCO object detection, and ADE20K semantic segmentation. Though our approach is inspired by the Swin Transformer, we have three fundamental differences: (1) where attention is applied are different. Swin Transformer partitions the image into windows and then conducts pixel-level attention inside each window, while we do window partitioning in feature maps, and calculate attention between windows by treating each window as a whole. (2) multi-scaling strategy is different. Swin transformer uses the same window size in one layer and merges windows to form a larger window in deeper layers. In contrast, we use windows with different sizes as heads for multi-scale matching. (3) window shifting is applied differently. Swin Transformer shifts the whole feature map, in order to exchange information and provide connectivity between different windows. We apply independent cyclic shifts in each window in a non-exchangeable way. Additionally, in contrast to Swin Transformer, where each window is shifted only once, in our algorithm each window is shifted multiple times depending on its size.

Recently, transformer-based visual object tracking methods have become more and more popular. TrDiMP separates the encoder-decoder transformer into two Siamese-like branches, the encoder reinforces the template features and the decoder propagates the tracking cues from previous templates to the current frame. TransT proposes a feature fusion network and employs an attention mechanism to combine the features of template and search region. This feature fusion network consists of an ego-context augment module based on self-attention and a cross-feature augment module based on cross-attention. STARK develops a spatial-temporal architecture based on the encoder-decoder transformer, the encoder learns the relationship between template and search region and the decoder learns a query embedding to predict the target positions. Moreover, STARK introduces a corner-based prediction head used for estimating the bounding box and a score head for controlling the updates of the template image. Most of the previous tracking algorithms such as use encoder-decoder structure to enhance or fuse the features, while we consider the transformer as a feature matching module to calculate the similarity between template and search region. Moreover, previous approaches use the transformer naively and do touch the attention mechanism within. On the contrary, we carefully design a multi-head multi-scale window-level attention transformer with the cyclic shifting strategy, to fully exploit the transformer structure for object tracking.

Method

In this section, we present our multi-scale cyclic shifting window transformer tracker, namely CSWinTT. We use the transformer as a matching module for measuring the relevance between template and search region to fully exploit the powerful cross-attention capabilities of the transformer.

The tracking architecture is visualized in Figure 2, which consists of three major components: a feature extraction backbone, a transformer matching module, and a bounding box estimation head. We choose the ResNet-50 as our backbone for feature extraction, which takes a pair of image inputs, i.e., the template image and the search region image. The pair of output features are then partitioned into window sequences and fed into the transformer matching module. The matching module concatenates the two window sequences and sends them to the multi-head 6-layer transformer. The multi-head transformer uses a specific window size for each head for scale adaption. Finally, the outputs of each transformer head are concatenated together, passed through the corner-based box estimation head as in to get the result bounding box.

Multi-head attention. Multi-head attention is the fundamental component in our architecture. As described in , given queries Q\mathbf{Q}, keys K\mathbf{K}, and values V\mathbf{V}. The multi-head attention is computed as:

where nhn_{h} is the number of heads, and dkd_{k} is the dimension of key. For a clearer description of the post-order steps, we define the attention score as:

To address the aforementioned problems, we propose a cyclic shifting strategy on the proposed window-level attention. It enhances the effectiveness of cross-window attention while preserving positional information and the integrity of objects, as shown in Figure 3. Within a specific head, consider a window with size r×rr\times r referred to as the base sample. We define the shift operator shift(x,y)\text{shift}(x,y) of the base sample as translating the sample by xx pixels in the horizontal direction, and yy pixels in the vertical direction. Our cyclic shifts of a base sample then are performed at a single-pixel distance and move the sample into bottom-right directions with boundaries being warped back to the top-left. The operation shift(x,y),x,y∈[−r+1,r−1]\text{shift}(x,y),x,y\in[-r+1,r-1] generates r×rr\times r into (2r−1)2(2r-1)^{2} samples for the base sample with window size r×rr\times r. Obviously, these cyclic shifts generate a lot of duplicates, we will discuss how to effectively remove duplicate computations in the section 3.2.

2 Efficient Computation

Spatially regularized attention mask. In practice, we find the shifted samples near the center contribute more to the final attention. This is reasonable as samples close to the boundaries are more likely to break the integrity and position information of the tracking object in the window. Hence we design a weighting scheme applied as a form of attention mask M\mathbf{M} in the transformer, as shown in Figure 4. The spatial weights of the mask penalize samples depending on their spatial locations, the formula for weight generation is expressed in 3. The further away the generated sample is from the base sample, the larger the penalty is and the weight is smaller.

The spatial attention mask M\mathbf{M} is directly added onto the attention score in 2.

Computational optimization. Intuitively, the cyclic shifts increase the computational cost greatly, especially when the window size is large. To achieve computational efficiency, we optimize in three ways: (i) eliminating the cyclic shifts of the Query; (ii) halving the duplicated shifting periods; and (iii) adopting the programming optimization for matrix translation.

Suppose we have Q\mathbf{Q}, K\mathbf{K} and V\mathbf{V} of size (H,W,d)(H,W,d). Standard transformer flattens the features to (HW,d)(HW,d), two parts account for the time cost of attention computation are the attention score computation O(HW×dHW)O(HW\times dHW) and fusion feature computation O(HW×HW×d)O(HW\times HW\times d). After applying the cyclic shifts, the size of Q\mathbf{Q}, K\mathbf{K} and V\mathbf{V} are (Hr,Hr,2r−1,2r−1,r,r,d)(\frac{H}{r},\frac{H}{r},2r-1,2r-1,r,r,d) with rr as window size, and the complexity of computing the attention score increases to O((HrWr(2r−1)(2r−1))2×r2d)O((\frac{H}{r}\frac{W}{r}(2r-1)(2r-1))^{2}\times r^{2}d). We observe if QQ and KK perform the same shifting, computing attention scores is meaningless, so we just need to perform a cyclic operation on KK and keep QQ unchanged to achieve the same effect. In addition, note that the cyclic generated samples to the bottom-right and the top-left directions are repeated, we reduce the number of shifting periods by half for better efficiency. And we also apply a programming trick to improve the tracking speed, which is using permutations of matrix coordinates to perform cyclic shifts instead of direct translations on the matrix.

3 Tracking with Window Transformer

Multi-scale window transformer facilitates the tracking process by conducting accuracy-aware attention with windows at different scales. Therefore, the choice of the window sizes is extremely important. In our implementation, we set the number of heads nhn_{h} to 88 with window size ri=r_{i}= for head ii. Notice the second half of the heads have the same window size, that’s because we adopt feature map translation which displaces the backbone feature of the search image by (ri2,ri2)(\frac{r_{i}}{2},\frac{r_{i}}{2}) pixels. In this way, when the windows are partitioned in a non-overlapping manner, the contents of windows are complemented by each other to avoid the situation that the object has been segmented all the time.

In further, to improve the robustness of the tracking algorithm, we use two templates of the same size as the input of the transformer. One of which is fixed using the initial template, the other is online updated to the latest tracking result with high confidence, a score head is employed to control the updates, as designed in STARK.

In the training stage, we use the L1 loss and the generalized IoU loss to train the overall architecture in an end-to-end manner. During the inference, the template image and its corresponding backbone features are initialized in the first frame, and the search region is used as the input to the tracker during the tracking process in subsequent frames, with the predicted bounding box returned by the network as the final result.

Experiments

We train our model on the LaSOT, GOT-10k, and TrackingNet datasets. The image pairs are directly sampled from the same sequence and common data augmentation operations including brightness jitter and horizontal flip are applied. The size of the input template is 128×\times128 pixels, the search region is 525^{2} times of the target box area and further resized to 384×\times 384 pixels. We use ResNet-50 as the backbone, the parameters of which are initialized with ImageNet pretrained model. Other parameters in our model are initialized with Xavier Uniform. We use the λl1=5\lambda_{l1}=5 and λgiou=2\lambda_{giou}=2 as the loss weight for l1 loss and giou loss. The AdamW optimizer is employed with initial learning rates of 1e-5 and 1e-4 for backbone parameters and other parameters, respectively, and weight decay is set to 1e-4 for every 10 epochs after 500 epochs. We train our model on two Nvidia Tesla T4 GPUs for a total of 600 epochs, each epoch uses 4×1044\times 10^{4} images. The mini-batch size is set to 64 images with each GPU hosting 32 images. The training process of the update module is the same as . Our approach is implemented in Python 3.7 using PyTorch 1.6. CSWinTT operates about 12 frames per second (FPS) on a single GPU during the online tracking process.

2 State-of-the-art Comparison

We compare our proposed CSWinTT algorithm with the state-of-the-art trackers on five tracking benchmarks, including UAV123, LaSOT, TrackingNet, GOT-10k, and VOT2020.

UAV123: UAV123 gathers an application-specific collection of 123 sequences and captures from unmanned aerial vehicles video dataset. It adopts the Area Under the Curve (AUC) and Precision (P) as the evaluation metrics. The precision is used to measure the center distance and the AUC plot computes the intersection-over-union (IoU) score between the estimated bounding box and the ground-truth. As shown in Table 1, where the previous state-of-the-art trackers such as TrDiMP, TransT, and START are included for comparison, note that STARK-ST50 is chosen for the reason that it uses the same ResNet-50 backbone as our algorithm, which can more fairly compare the performance of transformer structure. Our CSWinTT outperform the aforementioned methods by a considerable margin and exhibits very competitive performance (70.5% AUC and 90.3% Precision) when compared to the best previous tracker STARK (69.2% AUC and 88.2% Precision)

LaSOT: LaSOT is a large-scale long-term dataset including 1400 sequences and distributed over 14 attributes, the testing subset of LaSOT contains 280 sequences with an average length of 2448 frames. Methods are ranked by the AUC, Precision, and Normalized Precision (PNorm). The evaluation results of compared tracking algorithms are shown in Table 1. Our model achieves the top-rank AUC score (66.2%) and Precision score (70.9%), which outperforms the previous best result by STARK-ST50, and also surpasses the other two transformer trackers TransT/TrDiMP for 1.3%/2.2% AUC score, respectively.

TrackingNet: TrackingNet is a large-scale tracking dataset consisting of 511 sequences for testing. The evaluation is performed on the online server. 1 shows that, compared with SOTA models, our CSWinTT performs better visual tracking quality and ranks at the first in AUC score of 81.9% and normalized precision of 86.7%. The specific gain is 0.7% relative improvement of the AUC score when compared with the TransT, which represents the previous best algorithm on this benchmark.

GOT-10k: GOT-10k is a large-scale dataset containing over 10k videos for training and 180 for testing. It forbid the trackers to use external datasets for training. We follow this protocol by retraining our trackers using only the GOT10k train split. As can be seen from Table 1, among previous transformer trackers, TrDiMP and STARK-ST50 provides the best performance, with an AUC score of 68.8% and 68.0%. Our approach has remarked improvement and obtains an AUC score of 69.4%, significantly outperforming the best existing tracker (TrDiMP) by 0.6%.

VOT2020: VOT2020 benchmark contains 60 challenging videos. The performance on this dataset is evaluated using the expected average overlap (EAO), which takes both accuracy (A) and robustness (R) into account. In addition, a new anchor-based evaluation protocol is proposed in VOT2020, the segmentation mask is adopted as the ground-truth. However, since our algorithm does not output a segmentation mask, trackers only predict bounding boxes are chosen as the comparisons to ensure a fair evaluation. It can be seen from the data in Table 2 that CSWinTT obtains an EAO of 0.304, ranking first in previous trackers.

3 Ablation Study

We conduct ablation analysis to evaluate the different components in our CSWinTT and evaluate the performance of diverse window sizes using the UAV123 dataset. Besides, we show the superiority of the three previously mentioned computational optimization strategies.

Effects of different components in our method. We evaluate the effect of components including multi-scale window attention (Win), cyclic shifts (CS), spatially regularized attention mask (SR), and relative position encoding (Pos) employed in our method. The ablation study result is shown in Tab. 3, #1 represents the performance of the original transformer. We can see that window-level attention alone (#2) is very ineffective as it greatly reduces the resolution of the attention mechanism, however, combining window-level attention with the cyclic shifting strategy can handle this drawback. It can be seen in #3, there is a 15.3% improvement in the AUC score after applying the cyclic shifts, and it outperforms the original transformer by 3.5%, which illustrates that the cyclic shifting strategy plays a key role on the window-level attention. #4 shows that the AUC score can be improved by 0.4% when employing the spatially regularized mask to cyclic shifting samples, which demonstrates that spatial regularity can alleviate the boundary artifacts to a certain extent and improve the performance of window attention. In addition, we test the effectiveness of relative position encoding in our method following the way of . The performance improves by 0.1% when the relative position encoding (#5) is used, the small improvement indicates that the position encoding in window-level attention is not very important, and confirms that window-level attention itself contains rich position information.

Different window sizes for our transformer. To explore the performance of diverse window sizes on cyclic shifting window attention, we designed a quantitative analysis experiment as shown in Table 4. The first four rows indicate that the same window size is used in all 8 heads in the case that cyclic shifting strategy is employed. It can see from the experimental results that the highest 70.0% AUC score is obtained in size 4×44\times 4 when using a single window size, as a matter of fact, the performance is really closer for all windows sizes. When adopting the multi-scale window size, the best AUC score of 70.5 is achieved, demonstrating that multi-scale windows can fuse information from different scales to improve the performance of the tracker.

Computation optimization and speed analysis. Cyclic shifting strategy brings a large computational burden, we improve the tracking speed by applying some optimization strategies, including removing the cyclic shifts of Query (RMQ), halving the shifts periods (Peri), and adopting the programming optimization for matrix translation (Prog), Table 5 shows the effect of each optimization method. The tracking speed is around 1 FPS with no optimization adopted, as shown in #1, which is almost an unusable state. With the cyclic shifts of Query are removed (#2), the tracking speed is greatly improved to 8.2 FPS, and it can be further improved by halving the shifts periods (#3). In addition, we also apply a PyTorch programming trick to use permutations of matrix coordinates to perform cyclic shifts instead of direct translations on the matrix, which also improves the tracking speed to some extent as shown in #4. Due to the absolute amount of computation introduced by the cyclic shifting window attention, the computing efficiency of our method is not as good as the original transformer (#5), but a satisfactory tracking speed of 12.4 FPS is achieved after our computational optimization.

4 Qualitative Analysis

Figure 5 shows the visual heat map of the attention, which exhibits the attention score of the last layer in the transformer matching module. The red area in the heat map indicates a high attention degree, while the blue area indicates a low attention degree. The first row shows the situation where the target object is obscured, the second and third rows show the scenario where the target is surrounded by similar distractors. From the visualization we can see that, compared to the pixel-level attention, the cyclic shifting window attention has a stronger discrimination ability of visual tracking, especially when the occlusion occurs or when there are similar distractors around the target object.

We further discuss why our proposed CSWinTT works. The strong discriminative ability mainly comes from two strategies: multi-scale window partition and cyclic shifts. After window partition, the target is split into multiple small blocks and each block contains the indivisible information of the object part. These blocks do not disrupt the pixels inside during attention, when some blocks are obscured and not visible, another part of the block can do attention without interference. Although there is no information exchange between different windows, the fusion of multi-scale windows can alleviate the problem, as well as be more robust to diverse sizes of occlusion areas. Additionally, the cyclic shifts can generate a more accurate attention score. For example, after window partition for a human body, there are two windows needed to do the attention. Suppose the first one is a window in the template that contains a head of the human body, which is in the center of the window. The second one is in the search region that contains the same head, as the human movement through the sequence, the head translates from the center to the edge of the window. At this point, a lower matching score will be obtained by window-level attention, which does not fully utilize the information in the windows. After employing the cyclic shifts, as shown in Figure 3, the head at the center of the template window and the head at the edge of the search region window can be finely matched. In addition, the position information in the attention can be obtained by the shift size, and this window-level position can better assist the tracking algorithm to distinguish the target object from the distractors.

Conclusion

In this work, we propose a transformer tracker with multi-scale cyclic shifting window attention, which is able to keep the integrity of the object and retain more location information when calculating the cross-window attention between the tracking target and the search area. Moreover, this new window attention is deliberated designed with two improvement schemes including spatially regularized attention mask and redundant computation removal to fully exploit the transformer structure for object tracking. Numerous experimental results on five challenging benchmarks demonstrate that our tracker performs better than previous state-of-the-art trackers. The proposed cyclic shifting window attention has stronger discrimination than the original pixel-level attention in the tracking field. Many other applications like image recognition and stereo matching may benefit from this window attention too.

Acknowledgement

This work is supported by the national key research and development program of China under Grant No.2020YFB1805601.

References