AiATrack: Attention in Attention for Transformer Visual Tracking
Shenyuan Gao, Chunluan Zhou, Chao Ma, Xinggang Wang, Junsong Yuan
Introduction
Visual tracking is one of the fundamental tasks in computer vision. It has gained increasing attention because of its wide range of applications . Given a target with bounding box annotation in the initial frame of a video, the objective of visual tracking is to localize the target in successive frames. Over the past few years, Siamese trackers , which regards the visual tracking task as a one-shot matching problem, have gained enormous popularity. Recently, several trackers have explored the application of the Transformer architecture and achieved promising performance.
The crucial components in a typical Transformer tracking framework are the attention blocks. As shown in Fig. 1, the feature representations of the reference frame and search frame are enhanced via self-attention blocks, and the correlations between them are bridged via cross-attention blocks for target prediction in the search frame. The Transformer attention takes queries and a set of key-value pairs as input and outputs linear combinations of values with weights determined by the correlations between queries and the corresponding keys. The correlation map is computed by the scaled dot products between queries and keys. However, the correlation of each query-key pair is computed independently, which ignores the correlations of other query-key pairs. This could introduce erroneous correlations due to imperfect feature representations or the existence of distracting image patches in a background clutter scene, resulting in noisy and ambiguous attention weights as visualized in Fig. 4.
To address the aforementioned issue, we propose a novel attention in attention (AiA) module, which extends the conventional attention by inserting an inner attention module. The introduced inner attention module is designed to refine the correlations by seeking consensus among all correlation vectors. The motivation of the AiA module is illustrated in Fig. 1. Usually, if a key has a high correlation with a query, some of its neighboring keys will also have relatively high correlations with that query. Otherwise, the correlation might be noise. Motivated by this, we introduce the inner attention module to utilize these informative cues. Specifically, the inner attention module takes the raw correlations as queries, keys, and values and adjusts them to enhance the appropriate correlations of relevant query-key pairs and suppress the erroneous correlations of irrelevant query-key pairs. We show that the proposed AiA module can be readily inserted into the self-attention blocks to enhance feature aggregation and into the cross-attention block to facilitate information propagation, both of which are very important in a Transformer tracking framework. As a result, the overall tracking performance can be improved.
How to introduce the long-term and short-term references is still an open problem for visual tracking. With the proposed AiA module, we present AiATrack, a streamlined Transformer framework for visual tracking. Unlike previous practices , which need an extra computational cost to process the selected reference frame during the model update, we directly reuse the cached features which are encoded before. An IoU prediction head is introduced for selecting high-quality short-term references. Moreover, we introduce learnable target-background embeddings to distinguish the target from the background while preserving the contextual information. With these designs, the proposed AiATrack can efficiently update short-term references and effectively exploit the long-term and short-term references for visual tracking.
We verify the effectiveness of our method by conducting comprehensive experiments on six prevailing benchmarks covering various kinds of tracking scenarios. Without bells and whistles, the proposed AiATrack sets new state-of-the-art results on these benchmarks with a real-time speed of 38 frames per second (fps).
In summary, the main contributions of our work are three-fold:
We propose a novel attention in attention (AiA) module, which can mitigate noise and ambiguity in the conventional attention mechanism and improve tracking performance by a notable margin.
We present a neat Transformer tracking framework with the reuse of encoded features and the introduction of target-background embeddings to efficiently and effectively leverage temporal references.
We perform extensive experiments and analyses to validate the effectiveness of our designs. The proposed AiATrack achieves state-of-the-art performance on six widely used benchmarks.
Related Work
Recently, Transformer has shown impressive performance in computer vision . It aggregates information from sequential inputs to capture global context by an attention mechanism. Some efforts have been made to introduce the attention structure to visual tracking. Recently, several works apply Transformer architecture to visual tracking. Despite their impressive performance, the potential of Transformer trackers is still limited by the conventional attention mechanism. To this end, we propose a novel attention module, namely, attention in attention (AiA), to further unveil the power of Transformer trackers.
How to adapt the model to the appearance change during tracking has also been investigated by previous works . A straightforward solution is to update the reference features by generation or ensemble . However, most of these methods need to resize the reference frame and re-encode the reference features, which may sacrifice computational efficiency. Following discriminative correlation filter (DCF) method , another family of approaches optimize the network parameters during the inference. However, they need sophisticated optimization strategies with a sparse update to meet real-time requirements. In contrast, we present a new framework that can efficiently reuse the encoded features. Moreover, a target-background embedding assignment mechanism is also introduced. Different from , our target-background embeddings are directly introduced to distinguish the target and background regions and provide rich contextual cues.
2 Attention Mechanism
Represented by non-local operation and Transformer attention , attention mechanism has rapidly received great popularity over the past few years. Recently, Transformer attention has been introduced to computer vision as a competitive architecture . In vision tasks, it usually acts as a dynamic information aggregator in spatial and temporal domains. There are some works that focus on solving existing issues in the conventional attention mechanism. Unlike these, in this paper, we try to address the noise and ambiguity issue in conventional attention mechanism by seeking consensus among correlations with a global receptive field.
3 Correlation as Feature
Treating correlations as features has been explored by several previous works . In this paper, we use correlations to refer to the matching results of the pixels or regions. They can be obtained by squared difference, cosine similarity, inner product, etc. Several efforts have been made to recalibrate the raw correlations by processing them as features through hand-crafted algorithms or learnable blocks . To our best knowledge, we introduce this insight to the attention mechanism for the first time, making it a unified block for feature aggregation and information propagation in Transformer visual tracking.
Method
where , , are different linear transformations. Here, , , and denote the linear transform weights for queries, keys, values, and outputs, respectively.
To address the aforementioned problem, we propose a novel attention in attention (AiA) module to improve the quality of the correlation map . Usually, if a key has a high correlation with a query, some of its neighboring keys will also have relatively high correlations with that query. Otherwise, the correlation might be a noise. Motivated by this, we introduce the AiA module to utilize the informative cues among the correlations in . The proposed AiA module seeks the correlation consistency around each key to enhance the appropriate correlations of relevant query-key pairs and suppress the erroneous correlations of irrelevant query-key pairs.
Specifically, we introduce another attention module to refine the correlation map before the softmax operation as illustrated in Fig. 2(b). As the newly introduced attention module is inserted into the conventional attention block, we call it an inner attention module, forming an attention in attention structure. The inner attention module itself is a variant of the conventional attention. We consider columns in as a sequence of correlation vectors which are taken as queries , keys and values by the inner attention module to output a residual correlation map.
Given the input , and , we first generate transformed queries and keys as illustrated in the right block of Fig. 2(b). To be specific, a linear transformation is first applied to reduce the dimensions of and to () for computational efficiency. After normalization , we add 2-dimensional sinusoidal encoding to provide positional cues. Then, and are generated by two different linear transformations. We also normalize to generate the normalized correlation vectors , i.e. . With , and , the inner attention module generates a residual correlation map by
where denotes linear transform weights for adjusting the aggregated correlations together with an identical connection.
Essentially, for each correlation vector in the correlation map , the AiA module generates its residual correlation vector by aggregating the raw correlation vectors. It can be seen as seeking consensus among the correlations with a global receptive field. With the residual correlation map, our attention block with AiA module can be formulated as
For a multi-head attention block, we share the parameters of the AiA module between the parallel attention heads. It is worth noting that our AiA module can be readily inserted into both self-attenion and cross-attention blocks in a Transformer tracking framework.
2 Proposed Framework
With the proposed AiA module, we design a simple yet effective Transformer framework for visual tracking, dubbed AiATrack. Our tracker is comprised of a network backbone, a Transformer architecture, and two prediction heads as illustrated in Fig. 3. Given the search frame, the initial frame is taken as a long-term reference and an ensemble of several intermediate frames are taken as short-term references. The features of the long-term and short-term references and the search frame are extracted by the network backbone and then reinforced by the Transformer encoder. We also introduce learnable target-background embeddings to distinguish the target from background regions. The Transformer decoder propagates the reference features as well as the target-background embedding maps to the search frame. The output of the Transformer is then fed to a target prediction head and an IoU prediction head for target localization and short-term reference update, respectively.
Transformer Architecture. The Transformer encoder is adopted to reinforce the features extracted by the convolutional backbone. For the search frame, we flatten its features to obtain a sequence of feature vectors and add sinusoidal positional encoding as in . The sequence of feature vectors is then taken by the Transformer encoder as its input. The Transformer encoder consists of several layer stacks, each of which is made up of a multi-head self-attention block and a feed-forward network. The self-attention block serves to capture the dependencies among all feature vectors to enhance the original features, and is equipped with the proposed AiA module. Similarly, this procedure is applied independently to the features of the reference frames using the same encoder.
The Transformer decoder propagates the reference information from the long-term and short-term references to the search frame. Different from the classical Transformer decoder , we remove the self-attention block for simplicity and introduce a two-branch cross-attention design as shown in Fig. 3 to retrieve the target-background information from long-term and short-term references. The long-term branch is responsible for retrieving reference information from the initial frame. Since the initial frame has the most reliable annotation of the tracking target, it is crucial for robust visual tracking. However, as the appearance of the target and the background change through the video, the reference information from the long-term branch may not be up-to-date. This could cause tracker drift in some scenes. To address this problem, we introduce the short-term branch to utilize the information from the frames that are closer to the current frame. The cross-attention blocks of the two branches have the identical structure following the query-key-value design in the vanilla transformer . We take the features of the search frame as queries and the features of the reference frames as keys. The values are generated by combining the reference features with target-background embedding maps, which will be described below. We also insert our AiA module into cross-attention for better reference information propagation.
Afterward, we attach the target-background embedding maps to the reference features and feed them to cross-attention blocks as values. The target-background embedding maps enrich the reused appearance features by providing contextual cues.
Prediction Heads. As described above, our tracker has two prediction heads. The target prediction head is adopted from . Specifically, the decoded features are fed into a two-branch fully-convolutional network which outputs two probability maps for the top-left and the bottom-right corners of the target bounding box. The predicted box coordinates are then obtained by computing the expectations of the probability distributions of the two corners.
To adapt the model to the appearance change during tracking, the tracker needs to keep the short-term references up-to-date by selecting reliable references which contain the target. Moreover, considering our embedding assignment mechanism in Eq. 4, the bounding box of the selected reference frame should be as accurate as possible. Inspired by IoU-Net and ATOM , for each predicted bounding box, we estimate its IoU with the ground truth via an IoU prediction head. The features inside the predicted bounding box are passed to a Precise RoI Pooling layer whose output is taken by a fully connected network to produce an IoU prediction. The predicted IoU is then used to determine whether to include the search frame as a new short-term reference.
We train the two prediction heads jointly. The loss of target prediction is defined by the combination of GIoU loss and L1 loss between the predicted bounding box and the ground truth. The training examples of the IoU prediction head are generated by sampling bounding boxes around the ground truths. The loss of IoU prediction is defined by mean squared error. We refer readers to the supplementary material for more details about training.
3 Tracking with AiATrack
Given the initial frame with ground truth annotation, we initialize the tracker by cropping the initial frame as long-term and short-term references and pre-computing their features and target-background embedding maps. For each subsequent frame, we estimate the IoU score of the bounding box predicted by target prediction head for model update. The update procedure is more efficient than the previous practices , as we directly reuse the encoded features. Specifically, if the estimated IoU score of the predicted bounding box is higher than the pre-defined threshold, we generate the target-background embedding map for the current search frame and store the embedding map in a memory cache together with its encoded features. For each new-coming frame, we uniformly sample several short-term reference frames and concatenate their features and embedding maps from the memory cache to update the short-term reference ensemble. The latest reference frame in the memory cache is always sampled as it is closest to the current search frame. The oldest reference frame in the memory cache will be popped out if the maximum cache size is reached.
Experiments
Our experiments are conducted with NVIDIA GeForce RTX 2080 Ti. We adopt ResNet-50 as network backbone which is initialized by the parameters pre-trained on ImageNet-1k . We crop a search patch which is times of the target box area from the search frame and resize it to a resolution of pixels. The same cropping procedure is also applied to the reference frames. The cropped patches are then down-sampled by the network backbone with a stride of 16. The Transformer encoder consists of 3 layer stacks and the Transformer decoder consists of only 1 layer. The multi-head attention blocks in our tracker have 4 heads with channel width of 256. The inner AiA module reduces the channel dimension of queries and keys to 64. The FFN blocks have 1024 hidden units. Each branch of the target prediction head is comprised of 5 Conv-BN-ReLU layers. The IoU prediction head consists of 3 Conv-BN-ReLU layers, a PrPool layer with pooling size of and 2 fully connected layers.
2 Results and Comparisons
We compare our tracker with several state-of-the-art trackers on three prevailing large-scale benchmarks (LaSOT , TrackingNet and and GOT-10k ) and three commonly used small-scale datasets (NfS30 , OTB100 and UAV123 ). The results are summarized in Tab. 1 and Tab. 2.
LaSOT. LaSOT is a densely annotated large-scale dataset, containing 1400 long-term video sequences. As shown in Tab. 1, our approach outperforms the previous best tracker KeepTrack by 1.9% in area-under-the-curve (AUC) and 3.6% in precision while running much faster (see Tab. 2). We also provide an attribute-based evaluation in Fig. 5 for further analysis. Our method achieves the best performance on all attribute splits. The results demonstrate the promising potential of our approach for long-term visual tracking.
TrackingNet. TrackingNet is a large-scale short-term tracking benchmark. It provides 511 testing video sequences without publicly available ground truths. Our performance reported in Tab. 1 is obtained from the online evaluation server. Our approach achieve 82.7% in AUC score and 87.8% in normalized precision score, surpassing all previously published trackers. It demonstrates that our approach is also very competitive for short-term tracking scenarios.
GOT-10k. To ensure zero overlaps of object classes between training and testing, we follow the one-shot protocol of GOT-10k and only train our model with the specified subset. The testing ground truths are also withheld and our result is evaluated by the official server. As demonstrated in Tab. 1, our tracker improves all metrics by a large margin, e.g. 2.3% in success rate compared with STARK and TrDiMP , which indicates that our tracker also has a good generalization ability to the objects of unseen classes.
NfS30. Need for Speed (NfS) is a dataset that contains 100 videos with fast-moving objects. We evaluate the proposed tracker on its commonly used version NfS30. As reported in Tab. 2, our tracker improves the AUC score by 2.7% over STARK and performs the best among the benchmarked trackers.
OTB100. Object Tracking Benchmark (OTB) is a pioneering benchmark for evaluating visual tracking algorithms. However, in recent years, it has been noted that this benchmark has become highly saturated . Still, the results in Tab. 2 show that our method can achieve comparable performance with state-of-the-art trackers.
UAV123. Finally, we report our results on UAV123 which includes 123 video sequences captured from a low-altitude unmanned aerial vehicle perspective. As shown in Tab. 2, our tracker outperforms KeepTrack by 0.9% and is suitable for UAV tracking scenarios.
3 Ablation Studies
To validate the importance of the proposed components in our tracker, we conduct ablation studies on LaSOT testing set and its new extension set , totaling 430 diverse videos. We summarize the results in Tab. 3, Tab. 4 and Tab. 5.
Target-Background Embeddings. In our tracking framework, the reference frames not only contain features from target regions but also include a large proportion of features from background regions. We implement three variants of our method to demonstrate the necessity of keeping the context and the importance of the proposed target-background embeddings. As shown in the 1st part of Tab. 3, we start from the variant (a), which is the implementation of the proposed tracking framework with both the target-background embeddings and the AiA module removed. Based on the variant (a), the variant (b) further discards the reference features of background regions with a mask. The variant (c) attaches the target-background embeddings to the reference features. Compared with the variant (a), the performance of the variant (b) drops drastically, which suggests that context is helpful for visual tracking. With the proposed target-background embeddings, the variant (c) can consistently improve the performance over the variant (a) in all metrics. This is because the proposed target-background embeddings further provide cues for distinguishing the target and background regions while preserving the contextual information.
Long-Term and Short-Term Branch. As discussed in Sec. 3.2, it is important to utilize an independent short-term reference branch to deal with the appearance change during tracking. To validate this, we implement a variant (d) by removing the short-term branch from the variant (c). We also implement a variant (e) by adopting a single cross-attention branch instead of the proposed two-branch design for the variant (c). Note that we keep the IoU prediction head for these two variants during training to eliminate the possible effect of IoU prediction on feature representation learning. From the 2nd part of Tab. 3, we can observe that the performance of variant (d) is worse than variant (c), which suggests the necessity of using short-term references. Meanwhile, compared with variant (c), the performance of variant (e) also drops, which validates the necessity to use two separate branches for the long-term and short-term references. This is because the relatively unreliable short-term references may disturb the robust long-term reference and therefore degrade its contribution.
Effectiveness of the AiA Module. We explore several ways of applying the proposed AiA module to the proposed Transformer tracking framework. The variant (f) inserts the AiA module into self-attention blocks in the Transformer encoder. Compared with the variant (c), the performance can be greatly improved on the two subsets of LaSOT. The variant (g) inserts the AiA module into the cross-attention blocks in the Transformer decoder, which also brings a consistent improvement. These two variants demonstrate that the AiA module generalizes well to both self-attention blocks and cross-attention blocks. When we apply the AiA module to both self-attention blocks and cross-attention blocks, i.e. the final model (i), the performance on the two subsets of LaSOT can be improved by 1.72.7% in all metrics compared with the basic framework (c).
Recall that we introduce positional encoding to the proposed AiA module (see Fig. 2). To verify its importance, we implement a variant (h) by removing positional encoding from the variant (i). We can observe that the performance drops accordingly. This validates the necessity of positional encoding, as it provides spatial cues for consensus seeking in the AiA module. More analysis about the components of the AiA module are provided in the supplementary material.
Superiority of the AiA Module. One may concern that the performance gain of the AiA module is brought by purely adding extra parameters. Thus, we design two other variants to demonstrate the superiority of the proposed module.
First, we implement a variant of our basic framework where each Attention-Add-Norm block is replaced by two cascaded ones. From the comparison of the first two rows in Tab. 4, we can observe that simply increasing the number of attention blocks in our tracking framework does not help much, which demonstrates that our AiA module can further unveil the potential of the tracker.
We also implement a variant of our final model by replacing the proposed inner attention with a convolutional bottleneck , which is designed to have a similar computational cost. From the comparison of the last two rows in Tab. 4, we can observe that inserting a convolutional bottleneck can also bring positive effects, which suggests the necessity of correlation refinement. However, the convolutional bottleneck can only perform a fixed aggregation in each local neighborhood, while our AiA module has a global receptive field with dynamic weights determined by the interaction among correlation vectors. As a result, our AiA module can seek consensus more flexibly and further boost the performance.
Visualization Perspective. In Fig. 4, we visualize correlation maps from the perspective of keys. This is because we consider the correlations of one key with queries as a correlation vector. Thus, the AiA module performs refinement by seeking consensus among the correlation vectors of keys. Actually, refining the correlations from the perspective of queries also works well, achieving 68.5% in AUC score on LaSOT.
Short-Term Reference Ensemble. We also study the impact of the ensemble size in the short-term branch. Tab. 5 shows that by increasing the ensemble size from 1 to 3, the performance can be stably improved. Further increasing the ensemble size does not help much and has little impact on the running speed.
Conclusion
In this paper, we present an attention in attention (AiA) module to improve the attention mechanism for Transformer visual tracking. The proposed AiA module can effectively enhance appropriate correlations and suppress erroneous ones by seeking consensus among all correlation vectors. Moreover, we present a streamlined Transformer tracking framework, dubbed AiATrack, by introducing efficient feature reuse and embedding assignment mechanisms to fully utilize temporal references. Extensive experiments demonstrate the superiority of the proposed method. We believe that the proposed AiA module could also be beneficial in other related tasks where the Transformer architecture can be applied to perform feature aggregation and information propagation, such as video object segmentation , video object detection and multi-object tracking .
Acknowledgment. This work is supported in part by National Key R&D Program of China No. 2021YFC3340802, National Science Foundation Grant CNS1951952 and National Natural Science Foundation of China Grant 61906119.
References
Additional Experiment Details
To make the tracking procedure in an end-to-end manner without tedious post-processing, we adopt the anchor-free prediction head proposed in , which outputs the probability maps and for the top-left and the bottom-right bounding box corners. The coordinates , , , of the predicted bounding box are then obtained by
2 Training Objective
With the predicted bounding box and predicted IoU , the whole network is jointly trained by minimizing prediction errors. The bounding box prediction loss is defined as the combination of GIoU loss and L1 loss. Together with the IoU prediction loss, the loss function can be written as
where and represent the ground truths of bounding box and IoU respectively and , , are the trade-off weights.
3 Training Strategy
Similar to previous works , we utilize the training splits of LaSOT , TrackingNet , GOT-10k , and COCO for offline training. As for the COCO image dataset, we apply data augmentation to generate synthetic video clips of diverse classes. During training, we randomly sample the search frame and reference frames such that the index of the search frame is larger than the indexes of reference frames. For training efficiency, we only sample one frame as the short-term reference. We also apply random affine transformations to jitter the sizes and locations of the short-term reference frame and search frame to simulate real tracking scenarios and avoid the influence of absolute position bias caused by padding . The network is trained with the AdamW optimizer . The learning rate is 1e-5 for the network backbone and 1e-4 for the other components. It decays by a factor of 10 during training. The parameters of the first convolutional layer and the first stage in the ResNet-50 backbone are fixed during training.
4 Different Structures of the AiA Module
Besides variant (h) in the paper, we also explore other structures of the AiA module, where the following components are studied: (1) Layer normalization applied to the value. (2) Linear transformation applied to the value. (3) Identical connection after the correlation aggregation. To evaluate their effect, we design two other structures of the AiA module, i.e. AiAv2 and AiAv3. The differences between these structures are shown in Fig. 6. Note that AiAv1 is the structure we implement in AiATrack and AiAv3 is a typical self-attention structure in the vanilla Transformer .
From the results in Tab. 6, we can observe that the layer normalization and the identical connection are not key components in our AiA module. Applying linear transformation to the value can further improve the performance, but we remove it for the trade-off between performance and computational cost. Besides the observations above, all the experimental results validate the effectiveness of correlation refinement in the conventional attention mechanism with an extra attention module.
5 Results on VOT
Different from previous reset-based evaluation protocol , VOT2020 proposes a new anchor-based evaluation protocol which is more reasonable. The same as STARK and DualTFR , we use Alpha-Refine to generate masks for evaluation since the ground truths of VOT2020 are annotated by the segmentation masks. The overall performance is ranked by the Expected Average Overlap (EAO). As shown in Tab. 7, our tracker exhibits very competitive performance, outperforming STARK with a margin of 5% in terms of EAO.
Additional Visualization Results
We also provide detailed attribute analysis on LaSOT . Fig. 7 shows that our tracker has an encouraging performance in various kinds of scenarios like background clutter, camera motion, and deformation. The results suggest the great potential of the proposed method when dealing with challenging scenarios.
2 Qualitative Comparisons
To qualitatively compare our tracker with the state-of-the-art trackers, we visualize our tracking results with two recent representative trackers: KeepTrack and STARK . Fig. 8 shows the tracking outputs for these trackers on some challenging video examples.