Siamese Transformer Pyramid Networks for Real-Time UAV Tracking

Daitao Xing, Nikolaos Evangeliou, Athanasios Tsoukalas, Anthony Tzes

Introduction

Unmanned Aerial Vehicle (UAV) tracking has drawn increasing attention in recent years given its enormous potential in diverse fields such as path planning , visual surveillance , and border security . While extensive advancements have been made towards powerful visual object tracking methods, the problem of real-time tracking has been overlooked. Moreover, the inherently limited power resources on lower performance compact devices further constraint the development of UAV tracking. Due to the optimization of both software and hardware on mobile devices and the progress of the lightweight but powerful backbone networks , the real-time applications based on visual classification, object detection, instance segmentation have been implemented on the CPU end. However, designing an efficient and effective object tracker for UAVs with limited computing resources, such as a single CPU-core, remains challenging. The lightweight backbones are insufficient for extracting robust discriminative features, which is vital for the tracking performance, especially under uncertainty scenarios. Thus, previous trackers try to address this problem by employing deeper networks , designing complex structures , or online updaters , which sacrifice the inference speed.

In this work, we alleviate the aforementioned problems, accommodate the lightweight backbone and build a real-time CPU-based tracker. Firstly, to complement the representative ability of lightweight backbone network, we integrate the Feature Pyramid Network (FPN) into the tracking pipeline. Although existing trackers also employ multi-scale features, most of them resort to a simple combination or use features for different tasks. We claim that this is fundamentally limited since a discriminative representation requires combining the contexts from multiple scales. Even though, FPN encodes the pyramid information from low/high level semantics, it only exploits contexts from local neighborhoods rather than explicitly modeling the global interactions. The perception of the FPN is constrained by the receptive field, which is limited on the shallower networks. Inspired by the development of Transformer and its ability to model global dependencies, recent works introduce attention-based modules and achieve profound results. However, the complexity of these models may cause computation/memory overhead which is not suited for pyramid architecture. Instead, we design a lightweight Transformer attention layer and embed it into pyramid network. The proposed Siamese Transformer Pyramid Network (named SiamTPN) augments the target features with lateral cross attention between pyramid features, producing robust target-specific appearance representation. Figure 2 illustrates the main difference between our tracker and existing ones. Moreover, our tracker based on a lightweight backbone network achieves state-of-the-art results while running at real-time speed on both GPU and CPU end, as shown in Figure 1. Our main contributions are summarized as follows:

We introduce a novel Transformer-based tracking framework for systems with limited computational resouces. These systems are typical encountered in UAVs with only CPU-support. To the best of our knowledge, this is the first deep learning based visual tracker running at real-time speed on UAVs using CPUs.

We propose a lightweight Transformer layer and integrate it into pyramid networks to build an efficient and effective framework.

Superior performance on multiple benchmarks as well as extensive ablation studies demonstrate the effectiveness of the proposed method. Particularly, our approach achieves state-of-the-art results and an AUC score of 58.1 on LaSOT with only a lightweight backbone while running at over 30 FPS on the CPU end. The field tests further validate the efficiency of SiamTPN in real world applications.

Related Work

With the requirement of running neural networks on mobile platforms, a series of lightweight models are proposed . AlexNet utilizes fully convolutional operations and achieves profound results on ImageNet classification tasks. MobileNet family proposes inverted residual block, depthwise separate convolution to save computation cost. The ShuffleNet family is another series of lightweight deep neural networks which introduce channel shuffle operation and optimize the network design for the target hardware. Feature Pyramid Network The feature pyramid (i.e. bottom-up feature pyramid) is the most common architecture in modern neural network design. The hierarchical structure of CNN encodes the contexts in the gradually increased receptive field. The Feature Pyramid Network (FPN) and Path Aggregation Network (PANet) are commonly used for the cross-scale feature interaction and multi-scale feature fusion. FPN includes a bottom-up as well as a top-down path to propagate semantic information into multi-level features.

2 Object Tracking

Discriminative correlation filter (DCF). DCFs have shown promising results for object tracking since MOSSE and KCF . After that, multi-channel features, color names and multi scale features are used to improve the tracking robustness. Further improvements are achieved with non-linear kernels , long-term memory and deep features . further improves the robustness and optimized DCF for UAV tracking. Deep learning based object tracking The popular Siamese network family based trackers address object tracking via similarity learning. SiamRPN introduces the region proposal network to jointly perform classification and regression. DaSiamRPN improves the discrimination power of the model with a distractor-aware module and SiamRPN++ further improves the performance with more powerful deep architectures. Recent works like SiamBAN , SiamFC++ and Ocean replace the RPN with an anchor-free mechanism and achieve faster tracking speed. DiMP and ATOM learn a discriminative classifier online to distinguish the target from the background. These methods require intense calculation which is not suitable for CPU-based tracking. Transformer. Transformer was first proposed for machine translation in and shows great potential in many sequential tasks. DETR first migrated Transformer into object detection tasks and achieves remarkable results. Recent works introduce an attention mechanism for improving the tracking performance. Motivated by DETR, make use of transformer to directly fuse correlation maps from different levels and obtains remarkable accuracy and speed for object tracking on UAVs. Instead of migrating the complex transformer encoder and decoder paradigm, in this work, we exploit the transformer encoder and design an attention based feature pyramid fusion network to learn the target-specific model more efficiently.

Proposed Method

As shown in Figure 2, the proposed SiamTPN, consists of three modules: one Siamese backbone network for feature extraction, one Transformer based feature pyramid network and one prediction head for per-pixel classification and regression.

where Γ\Gamma is the TPN module, and MM is a multi-channel correlation map and is adopted as the input to the classification and regression head. The overall architecture is shown in Figure 2.

2 Feature Fusion Network

Multi-head Attention. Generally, a Transformer has several encoder layers, and each encoder layer is composed of Multi-head attention (MHA) module and a multilayer perceptron (MLP) module. The attention function is operated on queries Q\mathbf{Q}, keys K\mathbf{K} and values V\mathbf{V} in the scale dot-production way, which can be expressed as:

where the CC is the key dimensionality to normalize the attention , and Pos\mathbf{Pos} is the positional encoding that are added to the input of each attention layer. The positional embedding in Transformer architectures is a location-dependent trainable parameter vector that is added to the token embeddings prior to inputting them to the Transformer blocks. The model representation capability is enhanced when extending the attention mechanism into multiple head way, which can be formulated as follows:

where nq=hqwq,nkv=hkvwkvn_{q}=h_{q}w_{q},n_{kv}=h_{kv}w_{kv}, w,hw,h is the resolution of input feature map. There exist three ways to reduce the computation cost: (1) reduce the query size, (2) reduce the dimension of CC, or (3) reduce the key and value size. However, reducing the query size also reduces the number of points for the prediction head, which eventually affects tracking accuracy. The same situation happens with the reduction of feature dimensionality. Since the feature maps with variable resolution are used as keys and values for fusion in TPN, we propose a pooling attention (RA) layer to reduce the spatial scale of K\mathbf{K} and V\mathbf{V}. Specifically, the K\mathbf{K} and V\mathbf{V} are fed into a pooling layer with both pooling and stride size of RR. To further reduce the computation cost of attention module, we remove the position encoding in original MHA for the following reasons: (1) the permutation of the input tokens is constrained by the final cross correlation. (2) Accessing and storage of the position embedding for each feature maps costs extra resources which is not suited for mobile devices. Overall, the mechanism of PA block (PAB) can be summarized as:

where MLP⁡\operatorname{MLP} is a fully connected feed-forward network, and Norm⁡\operatorname{Norm} is the LayerNorm to smooth the input feature. The structure comparison between MHA and PA module is shown in Figure 3.

3 Transformer Pyramid Network

To leverage the pyramid feature hierarchy Pi, i∈{3,4,5}P_{i},~{}i\in\{3,4,5\}, which has both low-level information and high-level semantics, a Transformer Pyramid Network (TPN) is proposed to build a blend feature with high-level semantics throughout. The TPN consists of stacked TPN blocks which takes pyramid features {P3,P4,P5}\{P_{3},P_{4},P_{5}\} and output new fusion feature {P3′,P4′,P5′}\{P_{3}^{{}^{\prime}},P_{4}^{{}^{\prime}},P_{5}^{{}^{\prime}}\}, as shown in Figure 4. The pyramid features are fed into a 1×11\times 1 convolution layer for dimension reduction, following a flatten operation before processing in the TPN. We fix the feature dimension (numbers of channels), denoted as CC in all the feature maps. The construction of the pyramid features involves a bottom-up pathway and centralized pathway. The bottom-up pathway is the feed-forward convolution from the backbone architecture and produces feature hierarchy {P3,P4,P5}\{P_{3},P_{4},P_{5}\}. Then a centralized pathway merges the feature hierarchy into a unified feature. Specifically, we use P4P_{4} as query for all feature hierarchy, yielding 3 combinations with different pooling scales which are processed by three parallel PAB locks. The outputs are directly added and fed into two self-attention PAB blocks to get the final semantic feature. The whole processing can be formulated as:

P3P_{3} and P5P_{5} are set as identity ones to avoid computation/memory overhead. Moreover, PA block design guarantee that the interdependencies among hierarchical features can be raised efficiently. The TPN Block repeats B times and produces the final representation for cross-correlation and the final prediction. Simplicity is central to our design and we have found that our model is robust to various design choices.

4 Prediction Head

The fusion features P4xP_{4}^{x} and P4zP_{4}^{z} are reshaped back to the original size before fed into the prediction head. Following , the Depth-wise Cross Correlation is performed between the search map and the template kernel to get a mult-channel correlation map. The correlation maps are fed into two separate branches. Each branch consists of 3 stacked convolution blocks to generate final outputs Aw×h×2clsA_{w\times h\times 2}^{cls} and Aw×h×2regA_{w\times h\times 2}^{reg}. Aw×h×2clsA_{w\times h\times 2}^{cls} represents the foreground and background scores for each point on feature maps and Aw×h×2regA_{w\times h\times 2}^{reg} predicts the distances from each feature point to the four sides of the bounding box. Overall, the objective function is

where Lcls\mathcal{L}_{cls} is the cross-entropy loss for classification, Liou\mathcal{L}_{iou} is GIOU loss between prediction boxes and ground truth box and Lreg\mathcal{L}_{reg} is the L1L1 loss for regression. Constants λcls\lambda_{cls}, λreg\lambda_{reg} and λiou\lambda_{iou} weight the losses.

Experimental Studies

This section first presents the implementation details and the comparisons between variants of the SiamTPN tracker, with the cross-correlation visualization results. Then, ablation studies are presented to analyze the effects of the key components. We further compare our method with the state-of-the-art methods both on aerial and prevalent benchmarks. Finally, we deployed our tracker on a UAV platform to test its effectiveness in real-world applications.

Model We apply our SiamTPN to three representative lightweight backbones, namely AlexNet , MobileNetV2 , ShuffleNetV2 . Using those networks as backbones enables us to adequately compare the effectiveness of proposed method. All backbones are pretrained on Imagenet. The details of backbone configuration for the different backbones are shown in Table 1. For ShuffleNet and MibileNet, we extract that the stages of spatial ratio equal to 18,116,132{\frac{1}{8},\frac{1}{16},\frac{1}{32}} respectively. For AlexNet, the last three layers are used for building feature pyramid.

Training Like the Siamese approaches, the network is trained offline with image pairs. The training data consists of the training splits from LaSOT , GOT10K , COCO and TrackingNet dataset. The image pairs are sampled from the videos with a maximum gap of 100 frames. The sizes of search images and templates are 256 × 256 pixels and 80 × 80 pixels respectively, corresponding to 424^{2} and 1.521.5^{2} times of the target box area, resulting in pyramid features {h3x=h3x=32,h4x=h4x=16,h5x=h5x=8}\{h_{3}^{x}=h_{3}^{x}=32,h_{4}^{x}=h_{4}^{x}=16,h_{5}^{x}=h_{5}^{x}=8\} and {h3z=h3z=10,h4z=h4z=5,h5z=h5z=3}\{h_{3}^{z}=h_{3}^{z}=10,h_{4}^{z}=h_{4}^{z}=5,h_{5}^{z}=h_{5}^{z}=3\}. Even though the lower input resolution brings additional speed increment, it is not the focus of this paper, so we set the aforementioned sizes for all the following experiments. The test images are augmented with some perturbation in the position and scale. For all backbones, the first layer and all BatchNorm layers are frozen during training. All experiments are trained for 100 epochs with 64 image pairs per batch. We use the ADAMW optimizer with initial learning rate of 10−510^{-5} for the backbone and 10−410^{-4} for the rest of the parts. The learning rate drops by a factor 0.1 decay on 90 epochs and the loss terms are weights with λcls=5\lambda_{cls}=5,λiou=5\lambda_{iou}=5,λreg=2\lambda_{reg}=2 respectively. During tracking, the scale penalty and Hanning windows is performed before selecting best prediction point from classification map Aw×h×2clsA_{w\times h\times 2}^{cls}. The final bounding box is given by adding the offsets predicted in Aw×h×2regA_{w\times h\times 2}^{reg} to the coordinates of the best prediction point.

2 Ablation Study

In this section, we verify the effectiveness of the proposed tracker from the following aspects: backbones choice, comparison with original Transformer and Convolution, the impact of TPN hyperparameters and the attention visualization. We follow the one-pass evaluation (Success and Precision) to compare different tracking configurations on the LaSOT test set and report the Success (AUC) scores. LaSOT is a large-scale long-term tracking benchmark which contains 280 videos for testing.

Backbones. The backbone network has the dominant impact on inference speed and accuracy. Modern architectures make use of residual skip connection, group/depth-wise convolution to design a competent network to learn more representative features, with even higher inference speed. We first compared the performance using different backbones. Similar to SiamFC , we remove all the feature fusion modules and predict results directly from P4P_{4}. We set CC=192 for all prediction layers. As shown in Table 2, the tracker with a simple backbone with prediction head achieves appreciable AUC scores on LaSOT with an average high inference speed on CPU end. Specifically, ShuffleNetV2 achieves AUC score of 34.1 with 48.1 FPS. A straight forward question is: Will more attached convolution layers help with tracking performance? We then stack additional convolution layers following P4P_{4} and Figure 5 shows the AUC changes along with the number of additional layers. Stacking more convolution layers improves the accuracy inefficiently and is worthless when compared with the speed drop. For ShufflenetV2, the speed drops over 30%30\% at a 15%15\% improvement on AUC score. We see that AlexNet is not suited for edge computing and ShuffleNetV2 and MobileNetV2 give comparable results both on accuracy and speed test. For the following experiments, we choose ShuffleNetV2 as the backbone. Comparison with original Transformer. To show the effect of our proposed TPN module and PA block, we design a tracker using the original Transformer. Similar to the setting of stacked convolution, we attach additional Transformer layers behind P4P_{4}. As shown in Figure 5, without the fusion of pyramid features, the tracker with only one additional transformer layer achieves better results than the tracker with six additional convolution layers. Moreover, the tracker with six transformer layers achieves an AUC score of 53.5 on LaSOT. Next, we implement an FPN using the same settings as TPN, but replacing the transformer layers with convolution and interpolation layers. The tracker with two stacked FPN learns more comprehensive representations from the interactions inside the feature pyramid and gets an AUC score of 47.2, demonstrating its advantage over the single layer architecture. However, the lack of the global dependencies become the bottleneck of improving accuracy. We further integrate Transformer layer into TPN blocks without using Pooling Attention layer. With the high-level semantics aggregated from pyramid features, the tracker achieves an state-of-the-art performance on LaSOT with an AUC score of 58.7. However, we see that the speed of tracker drops below 20 FPS which is not applicable for real-time tracking requirement. Finally, we test the results of TPN model with PA layers instead of transformer layer. Even the input size of queries and keys shrink with scale R, the tracker still achieves state-of-the-art performance . Nevertheless, the speed boosts up to 32.1 FPS with only 0.6 AUC score loss on LaSOT dataset, demonstrating the superiority of our method on both robustness and efficiency.

Impact of TPN hyperparameters. We discuss some architecture hyper parameters of the TPN model. Firstly, we examine the impact of the number of TPN blocks. With only one TPN block, the tracker produces a slight speed increment but suffers from AUC score drop from 58.1 to 52.8. Since the original transformer use 6 layers depth for both the encoder and decoder, we argue that 2 TPN blocks (depth=6) are enough for achieving robust tracking results. The number of heads in PA layer also plays an important role in tracking stability. For simplicity, we fix the head dimension=32, so we can test the input dimension C={128,192,256}C=\{128,192,256\} and head number N={4,6,8}N=\{4,6,8\} simultaneously. The tracker with 8 heads yields the best AUC score albeit at a cost of reducing in half of the inference time (FPS from 32.1 to 15.2). On the other hand, only using 4 heads is inefficient to learn an effective representation and only gives an AUC score of 46.2 on LaSOT. In practice, CC=192, NN=6, B=2 gives best balance between speed and accuracy.

Attention Visualization. The first three columns in Figure 6 show the response maps from the classification head with or without TPN module. Without TPN to learn discriminative features, the correlation results become dispersed and much easier to shift to distractors. The last three columns illustrate the attention maps between pyramid features. The attention between lower levels (P3P_{3} to P4P_{4}, P4P_{4} to P4P_{4}) distill more local information across the search area, while attention from high level (P5P_{5} to P4P_{4}) is more centralized on the semantics of the object target. All attention maps are calculated from the central feature point inside the bounding box with the whole key inputs.

3 Comparison with State-Of-The-Art Trackers

In this section, we compare our approach with 22 SOTA trackers. There are 4 anchor-based Siamese methods (SiamRPN , SiamRPN++ , DaSiamRPN , HiFT ), 5 anchor-free Siamese methods (SiamFC , SiamBAN , SiamCar , SiamFC++ , Ocean ), 10 DCF based methods (ECO , CCOT , KCF , ARCF , BACF , AutoTrack , CSRDCF , ROAM , DiMP , ATOM ), 2 attention based methods (CGACD , SiamAttn ) and 1 segmentaion based method, D3S . UAV123 . UAV123 is one of the largest UAV tracking benchmarks and adopts success and precision metrics for evaluation. As shown in Table 3, all trackers which achieve real-speed time on CPU are based on DCF, which rely on the handcraft features. This becomes a bottleneck of designing high-accuracy trackers. On the other hand, trackers relying on deeper networks like Resnet-50 can achieve high performance but are only applicable on GPU devices. Instead, our SiamTPN runs at real-time speed on the CPU while obtaining SOTA results. Specifically, SiamTPN gains a precision score of 85.8 and an AUC score of 66.04, outperforming the recent SOTA Siamese tracker SiamAttn. For a fair comparison, we develop a variant tracker based on AlexNet. While the AlexNet is not friendly on the CPU end, our tracker could run on the GPU at over 100 FPS while achieving consistent results with SiamRPN++. VOT2018 and OTB The VOT2018 dataset consists of 60 sequences with different challenge factors. The performance is compared in terms of EAO (Expected Average Overlap). OTB contains 100 sequences and evaluates performance with AUC score. Table 4 shows that our method achieves comparable results with SOTA algorithms on both VOT (second row) and OTB (third row) datasets. LaSOT . Figure 7 shows our SiamTPN achieves best results on the LaSOT test set, with an AUC score of 58.1 and beats all trackers based on deep Resnet trackers (DiMP, ATOM, OCEAN). Got10K is another large-scale dataset and employs Average Overlap (AO) as measurement. Following the requirement of generic object tracking, there is no overlap in object categories between the training set and test set, which is more challenging and requires a tracker with a powerful generalization ability. We follow their protocol and train the network with a training split. As demonstrated in Figure 1, SiamTPN achieves a relative 12%12\% higher performance on AO compared with the SOTA Siamese based tracker SiamRPN++ . On the other hand, our method exceeds all DCF based trackers while keeping the real-time inference speed on the CPU.

4 Real World Experimental Test

In this section, we verify the reliability of the proposed tracker in real-world UAV tracking. The hardware setup consists of a multi-copter UAV, an Embedded PC, a 3 axis Gimbal and a visual PTZ (pan-tilt-zoom) camera. We set up three different tracking scenarios to validate the tracking speed, generalization ability and robustness of SiamTPN. Specifically, the field tests include: (1) drone tracking with a ground stationary PTZ camera, as shown in Figure 8(a). (2) tracking and following a moving person with a drone and keeping the target within the field of view, as shown in Figure 8(b). (3) drone (evader) tracking with another drone (pursuer) with PTZ camera embedded, where two drones fly with custom trajectories but the parameters of PTZ camera are adjusted adaptively based on the position of evader, as shown in Figure 8(c). The position of drones are recorded with two GPS devices and shown in Figure 8(c)(I), where the red (blue) dots correspond to the pursuer (evader). Figure 8 shows the precise tracking results obtained under complex environments, exhibiting the robustness and practicability of tracker in real-world applications. We also compare the tracking speed variance under different bounding boxes size. Empirically, we split the bounding boxes into three categories based on pixel numbers, which is small (<1600<1600), medium (<10000<10000) and large (>10000>10000) ones. Figure 8(c)(II) demonstrates the steady inference speed under varies circumstances.

Conclusion

In this work, we propose a transformer pyramid network which aggregate semantics from different levels. The local interactions as well as the global dependencies are distilled from the cross attention among pyramid features. A pooling attention is further introduced to prevent the computation overhead. The comprehensive experiments demonstrate that our approach significantly improves the tracking results, while running at real-time speed on the CPU end.

References