AutoTrack: Towards High-Performance Visual Tracking for UAV with Automatic Spatio-Temporal Regularization

Yiming Li, Changhong Fu, Fangqiang Ding, Ziyuan Huang, Geng Lu

Introduction

Visual object tracking is one of the fundamental tasks in the computer vision community, aiming to localize the object sequentially only with the information given in the first frame. Endowing unmanned aerial vehicle (UAV) with visual tracking capability brings many applications, e.g., aerial cinematography , person following , aircraft tracking , and traffic patrolling .

There are currently two main research interests in this area: discriminative correlation filter (DCF)-based methods as well as deep learning-based approaches . In consideration of the limitation of power capacity and computational resources onboard UAVs, DCF framework is selected because of its high efficiency originating from calculation in the Fourier domain.

To improve DCF-based trackers, there are currently three directions: a) building more robust appearance model , b) mitigating boundary effect or imposing restrictions in learning , and c) mitigating filter degradation . Robust appearance can indeed boost performance, yet it leads to burdensome calculations. Filter degradation, on the other hand, is not improving it fundamentally. Most trackers try to improve performance using option b) by introducing regularization terms.

Recently, some attentions have been brought to using response maps generated in the detection phase to form the restrictions in learning . The intuition behind it is that the response map contains crucial information regarding the resemblance of current object and the appearance model. However, only exploits what we call the spatially global response map variations, while ignoring local response variation indicating credibility at different locations in the image: drastic local variation means low credibility and vice versa.

We fully exploit the local-global response variation to train our tracker with automatic spatio-temporal regularization, i.e., AutoTrack. While most parameters in regularization terms proposed by others are hyper-parameters that require large effort to tune, and would have a difficult time adjusting to new situations that the designers did not think of, we propose to learn some of the hyper-parameters automatically and adaptively. AutoTrack performs favorably against the state-of-the-art trackers, while running at ∼\sim6060 frames per second (fps) on a single CPU.

Our main contributions are summarized as follows:

We propose a novel spatio-temporal regularization term to simultaneously exploit local and global information hidden in response maps.

We develop a novel DCF-based tracker which can automatically tune the hyper-parameters of spatio-temporal regularization term on the fly.

We evaluate our tracker on 278 difficult UAV image sequences, and the evaluations have validated the state-of-the-art performance of our tracker compared to current CPU- and GPU-based trackers.

We introduce a novel application of visual object tracking in UAV localization and prove its effectiveness as well as generality in the practical scenarios.

Related Works

Tracking by detection: tracking-by-detection framework, which regards the tracking as a classification problem, is widely adopted in UAV . Among them, DCF has exhibited good performance with exceptional efficiency. The speed of traditional DCF-based trackers is around hundreds of fps on a single CPU, far exceeding the real-time requirement of UAV (3030 fps). Yet they are primarily subjected to the following issues.

a) Boundary effect: the circulant samples suffer from periodical splicing at the boundary, reducing filters’ discriminative power. Several works can mitigate boundary effect , but they used a constant spatial penalization which cannot adapt to various changes in different objects. K. Dai et al. optimized the spatial regularization in the temporal domain . Different to , we exploit the inherent information in DCF framework, so our method is more generic. Also, we have achieved better performance in the aerial scenarios in terms of speed and precision.

b) Filter degradation: the appearance model updated via a linear interpolation method cannot adapt to ubiquitous appearance change, leading to filter degradation. Some attempts are made to tackle the issue, e.g., training set management , temporal restriction , tracking confidence verification and over-fitting alleviation . Amongst them the temporal regularization is an effective and efficient way. Yet the non-adaptive regularization is prone to tracking drift once the filter is corrupted.

Tracking by deep learning: recently, deep learning-based tracking has caught wide attention due to its robustness, e.g., deep feature representation , reinforcement learning , residual learning and adversarial learning . However, for mobile robots, the above trackers cannot meet the requirement of real-time perception even with a high-end GPU. Currently, the state-of-the-art deep trackers are mostly built on siamese neural network . The pre-trained siamese trackers just need to traverse in a feed-forward way to get a similarity score for object localization, facilitating real-time implementation on GPU. However, on a mobile device solely with CPU, the speed of siamese-based trackers cannot satisfy the real-time needs. C. Huang et al. proposed a CPU-friendly deep tracker by training an agent working in a cascaded manner. It can run at near real-time speed by reducing calculation on easy frames. In summary, deep trackers can hardly meet real-time demands on CPU.

Vision-based localization: vision-based localization is crucial for UAV especially in GPS-denied environments. A. Breitenmoser et al. developed a monocular 6D pose estimation system based on passive markers in the visible spectrum . However, it performs worse in low-light environments. M. Faessler et al. presented a monocular localization system based on infrared LEDs to raise robustness in cluttered environments . Its generality, however, is limited since the system can only work in the infrared spectrum. Built on , we develop a localization system based on visual tracking. In light of robustness and generality of our tracker in various scenarios like illumination variation, occlusion and deformation, our localization system is more versatile compared to the infrared LED-based one .

Revisit STRCF

In this section, our baseline STRCF is revisited. The optimal filter Ht\mathbf{H}_{t} in frame tt is learned by minimizing the following objective function:

Although STRCF has achieved competent performance, it does have two limitations: a) the fixed spatial regularization failing to address appearance variation in the unforeseeable aerial tracking scenarios, b) the unchanged temporal penalty strength θ\theta (set as 15 in ) which is not general in all kinds of situations.

Automatic Spatio-Temporal Regularization

In this work, both local and global response variations are fully utilized to achieve simultaneous spatial and temporal regularizations, as well as automatic and adaptive hyper-parameter optimization.

First of all, we define local response variation vector Π=[∣Π1∣,∣Π2∣,...,∣ΠT∣]\mathbf{\Pi}=[|\Pi^{1}|,|\Pi^{2}|,...,|\Pi^{T}|], as can be seen in Fig. 1 for its 2D visualization in the object bounding box, in preparation for spatial regularization. Its ii-th element ∣Πi∣|\Pi^{i}| is defined as:

where [ψΔ][\psi_{\Delta}] is the shift operator to make two peaks in two response maps Rt{\mathcal{R}}_{t} and Rt−1{\mathcal{R}}_{t-1} coincide with each other, in order for removing the motion influence . Ri\mathcal{R}^{i} denotes the ii-th element in response map R\mathcal{R}.

where ζ\zeta and ν\nu denote hyper parameters. When the global variation is higher than the threshold ϕ\phi, it means that there are aberrances in response maps , so correlation filter ceases to learn. If it is lower than the threshold, the more dramatic the response map varies, the smaller the reference value will be, so that the restriction on temporal change of the correlation filters can be loosened and it can learn more rapidly in situations like large appearance variations.

Remark 1: Note that what we defined here is the reference value rather than the hyper-parameter itself. For the hyper-parameter of the temporal regularization, we use joint optimization to online estimate the value of it, so that the restriction can be online adaptively adjusted according to the response map variations. When appearance changes drastically, correlation filter learns more rapidly and vice versa.

2 Objective Optimization

Our objective function for joint optimization of filter as well as temporal regularization term can be written as:

By minimizing Eq. 6, an optimal solution can be obtained through alternating direction method of multipliers (ADMM) . The Augmented Lagrangian form of equation Eq. 6 can be formulated as:

Then we solve the following subproblems by ADMM.

Subproblem G^\widehat{\mathbf{G}}: given Ht,θt,V^t\mathbf{H}_{t},\theta_{t},\widehat{\mathbf{V}}_{t}, the optimal G^∗\widehat{\mathbf{G}}^{*} is:

Solving Eq. 9 directly is very difficult because of its complexity. So we decide to sample x^t\widehat{\mathbf{x}}_{t} across all KK channels in each pixel to simplify our formulation written by:

where the vector ρ\boldsymbol{\rho} takes the form ρ=Γj(X^t)y^j+θtΓj(G^t−1)−γΓj(V^t)+γΓj(TFHt)\boldsymbol{\rho}=\varGamma_{j}(\widehat{\mathbf{X}}_{t})\widehat{\mathbf{y}}_{j}+\theta_{t}\varGamma_{j}(\hat{\mathbf{G}}_{t-1})-\gamma\varGamma_{j}(\widehat{\mathbf{V}}_{t})+\gamma\varGamma_{j}(\sqrt{T}\mathbf{F}{\mathbf{H}}_{t}) for presentation.

Subproblem H\mathbf{H}: given θt,G^t,V^t\theta_{t},\widehat{\mathbf{G}}_{t},\widehat{\mathbf{V}}_{t}, we can optimize hk\mathbf{h}^{k} by:

The closed-form solution of hk\mathbf{h}^{k} can be written by:

Subproblem θt\theta_{t}: given other variables in Eq. 8, the optimal solution of θt\theta_{t} can be determined as:

Lagrangian multiplier update: after solving three subproblems above, we can update Lagrangian multipliers as:

where ii and i+1i+1 denotes the iteration index and the step size regularization constant γ\gamma (initially equals to 11) takes the form of γ(i+1)\gamma^{(i+1)}=min(γmax,βγi)min(\gamma_{max},\beta\gamma^{i}). (β=10\beta=10, γmax=10000\gamma_{max}=10000)

By iteratively solving the four subproblems above, we can optimize our objective function effectively and obtain the optimal filter G^t\widehat{\mathbf{G}}_{t} and temporal regularization parameter θt\theta_{t} in frame tt. Then G^t\widehat{\mathbf{G}}_{t} is used for detection in frame t+1t+1.

3 Object Localization

The tracked object is localized by searching for the maximum value of response map Rt\mathcal{R}_{t} calculated by:

where Rt\mathcal{R}_{t} is the response map in frame tt, F−1\mathscr{F}^{-1} denotes the inverse Fourier transform (IFT) operator and z^tk\widehat{\mathbf{z}}^{k}_{t} represents the Fourier form of extracted feature map in frame tt.

Localization by Tracking

Self-localization for UAV is essential for autonomous navigation. To develop a robust and universal localization system in dynamic and uncertain environments, we introduce visual object tracking into UAV localization for the first time. Specifically, we utilize the open-source software in , but employ AutoTrack to track four objects simultaneously instead of segmenting LEDs in the infrared spectrum. The main work-flow is briefly described below.

Prerequisites: the system requires the knowledge of four object configuration (non-symmetric), i.e., their positions in the world coordinate (observed in motion capture system), and intrinsic UAV-mounted camera parameters.

Initialization and tracking: after manually assigning four objects, AutoTrack starts to track them independently and output their location in the RGB image. Different to the system only applicable in infrared spectrum, our system can be used in versatile environments.

Correspondence search and pose optimization: correspondence between the tracked object configuration in the world coordinate and tracked results in image frames is firstly clarified, then the final 6D pose is optimized by fine-tuning the reprojection error .

Experiments

In this section, we firstly evaluate the tracking performance of AutoTrack with current state-of-the-art trackers on four difficult UAV benchmarks . Then, the proposed localization system is evaluated on Quanserhttps://www.quanser.com/products/autonomous-vehicles-research-studio/ platform in the indoor practical scenarios. The experiments of tracking performance evaluation are conducted using MATLAB R2018a on a PC with an i7-8700K processor (3.7GHz), 32GB RAM and NVIDIA GTX 2080 GPU. The tests of localization system are run on ROS using C++. For the hyper parameters of AutoTrack, we set δ=0.2\delta=0.2, ν=2×10−5\nu=2\times 10^{-5}, ζ=13\zeta=13. The threshold of ϕ\phi is 3000, ADMM iteration is set to 4. The sensitivity analysis of all the parameters can be found in the supplementary material.

For rigorous and comprehensive evaluation, the comparison between AutoTrack with the state-of-the-art methods is reported on four challenging and authoritative UAV benchmarks: DTB70 , UAVDT , UAV123@10fps and VisDrone2018-test-dev , with a total number of 119,830 frames. Noted that we use the same evaluation criteria with the four benchmarks .

DTB70: DTB70 , composed of 70 difficult UAV image sequences, primarily addresses the problem of severe UAV motion. In addition, various cluttered scenes and objects with different sizes as well as aspect ratios are included. We compare AutoTrack with nine state-of-the-art deep trackers, i.e., ASRCF , TADT , HCF , ADNet , CFNet , UDT+ , IBCCF , MDNet , MCPF , on DTB70, and the final results are reported in Fig. 2. Only with hand-crafted features, AutoTrack outperforms deep feature-based trackers (ASRCF , HCF , MCPF and IBCCF ) and pre-trained deep architecture-based trackers, i.e., MDNet , ADNet , UDT+ and CFNet . In summary, AutoTrack exhibits strong robustness against drastic UAV motion without losing efficiency, and also demonstrates a generality in tracking different objects in various scenes.

UAVDT: UAVDT mainly emphasizes vehicle tracking in various scenarios. Weather condition, flying altitude and camera view are three categories addressed by UAVDT. Compared to deep trackers including ASRCF , TADT , SiameseFC , DSiam , MCCT , ADNet , CFNet , DeepSTRCF , UDT+ , HCF , C-COT , ECO , IBCCF , MCPF and CREST , AutoTrack with a single CPU exhibits the best performance in terms of precision and speed, as shown in Table 1. In a word, AutoTrack has extraordinary performance in vehicle tracking despite omnipresent challenges.

1.2 Comparison with CPU-based trackers

Twelve real-time trackers (with a speed of >>3030fps), i.e., KCF , DCF , KCC fDSST , DSST , BACF , STAPLE-CA , STAPLE , MCCT-H , STRCF , ECO-HC , ARCF-H , and five non-real-time ones, i.e., SRDCF , SAMF , CSR-DCF , SRDCFdecon , ARCF-HC are used for comparison. The results of real-time trackers on four datasets are displayed in Fig. 3. Besides, the average performance of top ten CPU-based trackers in terms of speed and precision is demonstrated in the Table 2. It can be seen that AutoTrack is the best real-time tracker on CPU. Some tracking results are demonstrated in Fig. 4 and Fig. 6.

Overall performance evaluation: AutoTrack has outperformed all the CPU-based real-time trackers in both precision and success rate on DTB70 , UAVDT and UAV123@10fps . On VisDrone2018-test-dev , AutoTrack achieves comparable performance with the best tracker MCCT-H and ECO-HC in terms of precision and success rate. As for the average performance of top ten CPU-based trackers, AutoTrack has the best performance in precision with the second fast speed of 59.2fps, only slower than ECO-HC (69.5fps), however, we have achieved an average improvement of 4.8% in precision compared to ECO-HC. Moreover, AutoTrack has an advantage of 7.9% in precision and 108.5% in speed over the baseline STRCF.

Remark 2: M. Muller et al. created a 10fps dataset from the recorded 30fps one , thus the movement of tracked object between successive frames is larger, bringing more challenges. On UAV123@10fps, AutoTrack achieves a remarkable advantage of 5.8% in precision than the second best ECO-HC, proving its robustness against large motion.

Remark 3: Compared to ARCF-HC solely repressing the global response variation using a fixed parameter, we fully utilize the local-global information to fine-tune the spatio-temporal regularization term in an automatic manner. Extensive experiments have shown that AutoTrack achieves better performance while providing a much faster speed which is 3.1 times that of ARCF-HC.

Attribute-based evaluation: Success plots of eight attributes are exhibited in Fig. 5. In the normal appearance change scenarios (deformation, in-plane-rotation, viewpoint change), AutoTrack improves STRCF by 15.9%, 15.5% and 4.6% in success rate because the automatic temporal regularization can smoothly help filter adapt to new appearance. In illumination variation and large occlusion (aberrant appearance variation), AutoTrack has a superiority of 7.0% and 15.7% compared to STRCF in light of adaptive spatial regularization as well as aberrance monitoring mechanism which can stop training before contamination.

1.3 Ablation study

To validate the effectiveness of our method, AutoTrack is compared to itself with different modules enabled. The overall evaluation is presented in Table 3. With each module (automatic spatial regularization ASR, automatic temporal regularization ATR) added to the STRCF, the performance is smoothly improved. It is noted that ATR can also bring a gain in speed compared to ASR because we can reduce meaningless and detrimental training on contaminated samples. In addition, response maps of some frames are illustrated in Fig. 4. It can be clearly seen that response of our method is more reliable than that of baseline.

2 Evaluation of Localization System

We evaluate our localization system on six datasets covering 2,666 images, and in each dataset, the camera is moving at a distinct trajectory as UAV flies. The image is captured with a resolution of 1280×7201280\times 720 pixels at 10fps, using Intel RealSense (R200) camera looking ahead to perform building inspection, as shown in Fig. 7.

We adopt the UAV location in motion capture system as the ground truth. The mean position errors in x, y and z directions are reported in Table 4. Figure 8 exhibits the estimated position as well as respective error in every frame. The root-mean-square error (RMSE) of our method on 2,666 frames is 3.44 centimeters.

Remark 4: It is noted that our system is applicable in various scenarios because our tracker can track any arbitrary objects once given their information in the first frame. In summary, compared to LED-based localization system , our method is more versatile and can run at real-time frame rates in the real-world scenarios.

Conclusion

In this work, a generally applicable automatic spatio-temporal regularization framework is proposed for high-performance UAV tracking. Local response variation indicates local credibility, thus restricting local correlation filter learning. Global variation is able to control how much the correlation filter learns from the whole object. Comprehensive experiments have validated that AutoTrack is the best CPU-based tracker with a speed of ∼\sim6060fps, and even outperforms some state-of-the-art deep trackers on two UAV datasets . In addition, we try to bridge the gap between the theory and practice by utilizing visual tracking in UAV localization in the real world. Considerable tests proved the effectiveness and generality of our method. We strongly believe that our work can promote the development of visual tracking and its application in robotics.

Acknowledgment: This work is supported by the National Natural Science Foundation of China (No.61806148), the Fundamental Research Funds for the Central Universities (No.22120180009), and Tsinghua University Initiative Scientific Research Program.

References