VITAL: VIsual Tracking via Adversarial Learning

Yibing Song, Chao Ma, Xiaohe Wu, Lijun Gong, Linchao Bao, Wangmeng Zuo, Chunhua Shen, Rynson Lau, Ming-Hsuan Yang

Introduction

There has been an increasing need for tracking target objects in bounding boxes to understand video contents. Current state-of-the-art trackers are typically based on a two-stage tracking-by-detection framework. The first stage draws a sparse set of samples around the target object and the second stage classifies each sample as either the target object or as the background using a deep neural network. Despite the favorable performance on recent tracking benchmarks , the performance of the two-stage methods is limited by two aspects. First, the positive samples are spatially overlapped, and they cannot capture a variety of appearance changes over time. Second, the extreme foreground-background class imbalance negatively affects training the classification networks. It is of great importance to investigate how to eliminate these barriers to advance the tracking-by-detection framework in the deep learning era.

Prior trackers have made limited efforts on increasing the diversity of training data in learning deep classifiers. Since classifiers tend to learn a discriminative boundary between positive and negative samples, they emphasize on the most discriminative ones. However, as the target appearance varies frame-by-frame in the whole video sequence, the most discriminative samples in the current frame may not persist over a long temporal span. Typical examples of appearance changes caused by partial occlusion or out-of-plane rotation easily result in model overfitting, as current training samples may differ much from the previous ones. To alleviate this problem, existing trackers incrementally update the classifier through online sample collections. The noisy updates occur and bring tracker drift problem. Hence, a natural question is how we can augment positive samples in the feature space to capture target appearance variations in the temporal domain.

In this work, we take advantage of the recent progress in adversarial learning to augment training data to facilitate classifier training. For a deep classification network, such as the VGG-M model , we add a generative network between the last convolutional layer and the first fully connected layer. The generative network augments positive samples by generating weight masks randomly applied to the features, where each mask represents a specific type of appearance variation. Through adversarial learning, our network can identify the mask that maintains the most robust features of target appearance in the temporal domain. We show that the learned mask tends to decrease the weights of discriminative features, which tends to overfit in a single frame. Meanwhile, these features are hardly robust to appearance changes over the temporal span. In other words, adversarial learning helps our tracker exploit the most robust features over a long temporal span in classifier training, rather than overfitting to discriminative features in a single frame. Moreover, to mitigate the issue of class imbalance, we propose a high-order cost sensitive loss to decrease the effect of easy negative samples. Taking advantages of adversarial learning and high-order cost sensitive loss, our tracking method achieves favorable results against state-of-the-art trackers.

We summarize the main contributions of this work as follows:

We propose to use a generative adversarial network (GAN) to augment positive samples in the feature space to capture a variety of appearance changes over a temporal span.

We propose to use higher-order cost sensitive loss to mine hard negative samples to handle class imbalance.

We extensively validate our method on benchmark datasets with large-scale sequences. We show that our VITAL tracker performs favorably against state-of-the-art trackers.

Related Work

Visual tracking has long been an active research topic with extensive surveys over the last decade. In this section, we mainly discuss the representative visual trackers and the related issues on generative adversarial learning and class imbalance.

Visual tracking has a wide range of applications including action recognition , target analysis and augmented reality . State-of-the-art trackers are mainly based on the one-stage regression framework or the two-stage classification framework. As one of the most representative types of the one-stage regression framework, the correlation filter based trackers regress all the circular-shifted version of the input features into soft labels generated by a Gaussian function. By computing the correlation as an element-wise product in the Fourier domain, these trackers have received a lot of attention recently. Starting from the MOSSE tracker , many efforts have been made to improve the correlation filter for robust tracking. Extensions include, but are not limited to, kernelized correlation filters , scale estimation , re-detection , spatial regularization , ADMM optimization , sparse representation , CNN feature integrations and end-to-end CNN predictions .

In contrast, the two-stage classification framework poses the tracking task as a binary classification problem. The two-stage trackers emphasize on a discriminative boundary between the samples of the target object and background. Numerous learning schemes are proposed including P-N learning , multiple instance learning , structured SVMs , CNN-SVMs , domain adaptation , and ensemble learning . Unlike the existing two-stage tracking-by-detection trackers, our method, for the first time, takes advantage of the recent progress in generative adversarial learning to augment training samples in the feature space. The augmented samples capture a variety of appearance changes and thus strengthen the robustness of the classifier. In addition, we exploit hard negative samples to handle class imbalance limitation.

It is introduced in to generate realistic-looking images from random noise via the CNN. The generative adversarial network (GAN) consists of two subnetworks. One serves as a generator and the other as a discriminator. The generator aims at synthesizing images to fool the discriminator, while the discriminator tries to discriminates between real images and images synthesized by the generator. The generator and the discriminator are trained simultaneously by competing with each other. An advantage of adversarial learning is that the generator is trained to produce similar image statistics to those of the training samples so that the discriminator cannot differentiate. This manner is hardly achieved by existing empirical objective functions with supervised learning. The progress in generative adversarial learning has attracted a series of works on network training and computer vision applications, such as image generation , image stylization , object detection , and semantic segmentation . Unlike existing GANs that augment data in the image space, we apply adversarial learning to augment training samples in the feature space to capture appearance variations in temporal domain. In sum, our method exploits robust features over the long temporal span, instead of the discriminative features in individual frames.

This problem often exists in learning applications, where the amount of training data in one class (usually the positive class) is far less than that of another class (usually the negative class). A large portion of samples from the majority class are easy samples, which dominantly produce a large loss, and make the learning process unaware of the valuable samples from the minority class. Hard negative mining and reweighing training data are useful to alleviate the class imbalance problem to some extent. In visual tracking, class imbalance deteriorates the performance of the classifier, as the number of positive samples are extremely limited but the number of negative samples across the whole background is large. Unlike the aforementioned solutions for the class imbalance problem, we propose cost sensitive loss to decrease the effect from easy negative samples when training the classifier. This not only improves the tracking accuracy, but also accelerates the training convergence.

Proposed Algorithm

We build VITAL upon the CNN tracking-by-detection framework, which consists of feature extraction and classification. We interpret the classifier as the discriminator and propose a generator for adversarial learning . Unlike existing GAN-based methods, which expect to obtain generator mapping samples from one distribution to another after the training process, we expect to obtain a discriminator which is robust to target object variations. Fig. 2 shows the pipeline of our method, and the details are discussed below.

In the traditional adversarial learning , the generator GG takes a noise vector zz from a distribution Pnoise(z)P_{noise}(z) as an input and outputs an image G(z)G(z). The discriminator DD takes either G(z)G(z) or a real image xx with a distribution Pdata(x)P_{data}(x) as an input and outputs the classification probability. The generator GG is learned to maximize the probability of DD making a mistake. Using the standard cross entropy loss, the objective loss function for training GG and DD is defined as:

where the GG and DD networks are trained simultaneously. The training encourages GG to fit Pdata(x)P_{data}(x) so that DD will not be able to discriminate xx from G(z)G(z). Note that in Eq. 1, there are no ground truth annotations for zz and the learning process is unsupervised. After the training process, GG is removed and only DD is kept for inference.

Although GANs have been investigated in many computer vision tasks, a direct applying of Eq. 1 in the tracking-by-detection framework is not feasible. First, the input data to the framework are usually candidate object proposals rather than random noise. Second, we need to train the classifier via supervised learning using labeled samples rather than unlabeled ones. Third, we expect to use the classifier (i.e., DD) for inference rather than GG. These three factors limit the usage of GANs on visual tracking where both the input and learning strategy differ significantly.

We propose VITAL to narrow the gap between GANs and the tracking-by-detection framework. We add GG between feature extraction and the classifier as shown in Fig. 2. GG will predict a weight mask which operates on the extracted features. This mask is set randomly at the beginning and gradually identifies the discriminative features through adversarial learning. We define the input feature as CC, the mask generated by the GG network as G(C)G(C), the actual mask identifying the discriminative features as MM. We define the objective function as:

where the dot is the dropout operation on the feature CC. The mask contains only one channel and has the same resolution as CC. We express the predicted mask as M^\hat{M} and the value of the element (i,j)(i,j) as Mij^\hat{M_{ij}}. Meanwhile, we define the value of the element (i,j,k)(i,j,k) on feature CC as CijkC_{ijk}. The dropout operation is defined as follows:

In Eq. 2, we integrate the adversarial learning into the tracking-by-detection framework. We keep the input (i.e., the candidate object proposals) unchanged. When training DD (i.e, classifier), we extract features and enrich their representations in the feature space. Instead of empirically proposing data augmentation strategies, we let GG to identify the discriminative features, which are crucial for training DD. Initially, GG produces several random masks, which are akin to the random noise in Eq. 1. Each mask represents a specific type of appearance variation, and we expect these masks to cover the whole object variations. Through the adversarial learning process, GG will gradually identify the mask that degrades the classifier most. This indicates that the mask has identified the discriminative features. On the other hand, DD will gradually be trained without overfitting to the discriminative features from individual frames while relying on more robust features over a long temporal span. In each iteration of the adversarial learning, we first train DD and then GG. The detailed training procedure is presented in the following:

In one iteration of the training process, we pass the input feature through GG and obtain the predicted mask M^\hat{M}. We then conduct the dropout operation on this feature and sent the modified feature into DD. We keep the labels unchanged and train DD through supervised learning. Note that during this training process, there are multiple input features, GG will predict different masks according to different input features. It enables DD to focus on the temporal robust features without discriminative feature interference from single frames.

After training DD once, given an input feature, we create multiple output features based on several random masks. This feature diversifying process is performed through the dropout operation illustrated in Eq. 3. These features are passed onto DD, and we pick up the one with the highest loss. The corresponding mask of the selected feature is said to be effective in decreasing the impact of the discriminative features. We set this mask as MM in Eq. 2 and update GG accordingly.

Adversarial learning enables the classifier to focus on the temporal robust features instead of the discriminative ones in individual frames. Fig. 3 shows an example of how adversarial learning affects the classifier in practice. Fig. 3(a) shows the input frame with the ground truth annotation located at the face region. We use our VITAL tracker to represent the tracking-by-detection framework for illustration. We compute the entropy distribution based on the predicted probabilities from the classifier. The entropy measures the uncertainty of the prediction and is computed for binary classification as:

where pp is the predicted probability of the target object and 1−p1-p is the background. When p=0.5p=0.5, the value of the entropy HH is highest, which means that the classifier is uncertain to predict the label. When p=0p=0 or p=1p=1, the value of the entropy HH is lowest, which means that the classifier is certain about the prediction.

We compute the entropy distribution of Fig. 3(a) using VITAL without adversarial learning as shown in Fig. 3(b) and with adversarial learning as shown in Fig. 3(c). We note that these two distributions are similar despite some tiny variances. However, when the target undergoes partial occlusion and out-of-plane rotation as shown in Fig. 3(d), the entropy of VITAL without adversarial learning increases rapidly as shown in Fig. 3(e), which indicates that the classifier becomes uncertain around the target region. This is because the classifier is trained to focus on the discriminative features of the samples in the previous frames. As the target appearance varies in the following frames, these discriminative features vanish and decrease the classification accuracy. In comparison, the entropy distribution shown in Fig. 3(f) does not vary as significant as that in Fig. 3(e). It is because the classifier trained via diversified samples will not focus on the most discriminative features in individual frames. Instead, it tends to focus on more robust features over a long period of time. In sum, with the adversarial learning, VITAL becomes temporally robust while preserving the classification accuracy on individual frames.

2 Cost Sensitive Loss

We first revisit the cross entropy (CE) loss for binary classification. Formally, we define y∈{0,1}y\in\{0,1\} as the class labels and p∈p\in as the estimated probability for a class with label y=1y=1. Meanwhile, we define the probability for a class with label y=0y=0 as 1−p1-p. The CE loss is formulated as:

One notable problem of the CE loss is that easy negative samples, i.e., when p≪0.5p\ll 0.5 and y=0y=0, produce the loss with non-trivial magnitude. When summed over a large number of easy negative examples, these small loss values overwhelm the valuable rare positive class. In visual tracking, class imbalance lies between the limited positive samples and a substantial amount of negative samples across the whole background. Easy negative samples take over the majority of the CE loss and dominate the gradient.

Existing solutions to class imbalance include hard negative mining and training data reweighing . The simplest method to make a classifier cost sensitive involves a modification of the class importance. For example, when the ratio of positive and negative classes is 1:100, the importance factor of the negative class is set to be 0.01. Note that simply using a fixed factor to balance the importance of positive/negative examples does not identify the easiness or hardness of each example. We align our motivation to the recently proposed focal loss and add a modulating factor to the CE loss in terms of the network output probability pp. Formally, we build our cost sensitive loss upon the entropy loss as:

With the cost sensitive loss, we reformulate the objective function in Eq. 2 as:

where K1=1−D(M⋅C)K_{1}=1-D(M\cdot C) and K2=D(G(C)⋅C)K_{2}=D(G(C)\cdot C) are modulating factors that balance the training sample loss.

Tracking via VITAL

We illustrate how we perform VITAL for visual tracking. Note that we only involve GG when training the classifier and remove it in the test stage. The details are as follows:

We initialize our model through a two-stage training. In the first step we offline pretrain the model using positive and negative samples from the training data, which is from . In the second step we draw the samples from the first frame of the input sequence to finetune our model online. During offline pretraining, we randomly initialize DD and perform the training in a few iterations, then we involve GG for adversarial learning. See Sec. 3.1 for the details of the adversarial learning process where only positive samples are adopted. We mine the hard negative samples through the cost sensitive loss for training DD together with the diversified positive samples.

The online detection scheme is the same as existing tracking-by-detection approaches as we remove GG in this step. Given an input frame, we first generate multiple candidate proposals and extract their CNN features. We feed the CNN features of the candidate proposals into the classifier to get the probability scores.

We incrementally update our tracker frame-by-frame. Around the estimated position, we generate multiple samples and assign them with binary labels according to their intersection-over-union scores with the estimated bounding box. We use these training samples jointly train GG and DD during online update as illustrated in Sec. 3.1.

Experiments

In this section, we introduce the implementation details of VITAL and analyze the effects of adversarial learning and cost sensitive loss. Then we compare our VITAL tracker with state-of-the-art trackers on the benchmark datasets OTB-2013 , OTB-2015 and VOT-2016 for performance evaluation.

Our backbone feature extractor is based on the first three convolutional layers from the VGG-M model . When training GG, we prepare 9 random masks. The resolution of each mask is the same as that of the input features. We split this mask into 9 parts equally. We assign each part with label 1 in turn and the remaining parts with label 0. These masks are different from each other and cover all the parts in total. When training DD, we apply 9 masks to the input features independently to generate 9 diversified versions of each input feature. We then feed these diversified features into DD and select the one with the highest loss. The corresponding mask is denoted by MM as illustrated in Eq. 2 to train DD. During the adversarial learning, we iteratively apply the SGD solver to both GG and DD. We use 100 iterations to initialize both networks. The learning rate for training GG and DD are 10−310^{-3} and 10−410^{-4}, respectively. We update both networks every 10 frames using 10 iterations. Our VITAL tracker runs on a PC with an i7 3.6GHz CPU and a Tesla K40c GPU with the MatConvNet toolbox and the average speed is 1.5 FPS.

We follow the standard evaluation approaches. In the OTB-2013 and OTB-2015 datasets we use the one-pass evaluation (OPE) with precision and success plots metrics. The precision metric measures the frame locations rate within a certain threshold distance from ground truth locations. The threshold distance is set as 20 pixels. The success plot metric is set to measure the overlap ratio between the predicted bounding boxes and the ground truth. In the VOT-2016 dataset , we measure the performance in terms of Expected Average Overlap (EAO), Accuracy Ranks (Ar) and Robustness Ranks (Rr).

In VITAL, we train the classifier using the diversified positive samples with a cost sensitive loss. To validate the effectiveness of each component, we first implement a baseline algorithm by not enabling the adversarial training and using the standard cross entropy loss. We implement three alternative approaches based on the baseline algorithm. First, we train the classifier by generating random masks. Second, we train the classifier using adversarial learning (i.e., GAN). Third, we train the classifier using adversarial learning with the cost sensitive loss. Fig. 4 shows the results on the OTB-2013 dataset. We observe that using random masks deteriorates the classifier and results in inferior performance. It is because the spatial discriminative and temporal robust features are blocked randomly, which degrades the classifier to focus on either. In contrast, the mask predicted by adversarial learning effectively exploits the most robust features by blocking partial discriminative features in individual frames. The cost sensitive loss further improves the performance. However, the improvement of the cost sensitive loss is not as salient as that of the adversarial learning.

We compare VITAL with 29 trackers from the OTB-2013 benchmark and other 28 state-of-the-art trackers including DSST , KCF , TGPR , MEEM , RPT , LCT , MUSTer , HCFT , FCNT , SRDCF , CNN-SVM , DeepSRDCF , DAT , Staple , SRDCFdecon , CCOT , GOTURN , SINT , SiamFC , HDT , SCT , MDNet , DLS-SVM , ADNet , ECO , MCPF , CFNet and CREST . We evaluate all the trackers on 50 video sequences using the one-pass evaluation with distance precision and overlap success metrics.

Figure 5 shows the results from all compared trackers. For presentation clarity, we only show the top 10 trackers. The numbers listed in the legend indicate the AUC overlap success and 20 pixel distance precision scores. Overall, our VITAL tracker performs favorably against state-of-art trackers in both distance precision and overlap success. Figure 6 compares the performance under eight video attributes using one-pass evaluation. Our VITAL tracker handles large appearance variations well caused by deformation, in-plane and out-of-plane rotations. Compared to the representative tracking-by-detection tracker MDNet, we attribute our performance improvement by the diversified positive samples for training robust classifiers. The mask generated via adversarial learning captures a variety of object variations. It maskouts the discriminative features in individual frames while maintains the most robust features over a long temporal span. The advantage of exploiting the temporally robust features is clearly proved when dealing with occlusion. Through focusing on the persistently robust features, our VITAL tracker performs better than MDNet in a large margin. Meanwhile, our cost sensitive loss effectively decreases the loss from easy negative samples and forces the classifier to focus on hard ones. This facilitates discriminative classifiers to separate the target object from background. Our VITAL achieves leading performance in the presence of illumination variation and background clutter. However, for the low resolution sequences, our tracker does not perform as well as MDNet. This is because the target size of these sequences is small and the resolution of the weight masks predicted by adversarial learning is far low. For the scale variance sequence, the fixed size of weight mask cannot precisely maskout the discriminative features as the object size increases. Our future work will consider adaptively changing the size of the weight mask.

We compare our VITAL tracker on the OTB-2015 benchmark with the state-of-the-art trackers. Figure 7 shows that our VITAL tracker overall performs well. The ECO tracker achieves the best result in overlap success, while our VITAL ranks first in distance precision. Since the OTB-2015 dataset contains more videos with large scale changes and low resolution, our VITAL tracker does not perform as well as ECO in overlap success.

We compare our VITAL tracker with state-of-the-art trackers on the VOT-2016 benchmark, including Staple , MDNet , CCOT and ECO . VOT-2016 report shows that the strict state-of-the-art bound is 0.251 under EAO metric. Trackers whose EAO value exceeds this bound is defined as state-of-the-art. Table 1 shows that ECO performs best under the EAO metric. The performance of VITAL is comparable to that of CCOT and better than Staple and MDNet. According to the definition of the VOT report, all these trackers are state-of-the-art.

Fig. 8 qualitatively compare the results of the top performing trackers: CNN-SVM , CCOT , MDNet , ECO and VITAL on 12 challenging sequences. In a majority of these sequences, CNN-SVM fails to locate the target objects or estimates scale incorrectly because of the limited performance of the SVM classifier. MDNet improves CNN-SVM through an end-to-end CNN network formulation. It performs well on deformation (Trans), low resolution (Skiing) and fast motion (Diving). However, the classifier of MDNet is trained to focus on the discriminative features from individual frames, which may lead to overfitting in the presence of noisy update. It does not perform well in handling out-of-plane rotation (Ironman) and occlusion (Human4). The correlation filter based trackers (i.e., CCOT and ECO) extract CNN features and learn correlation filters independently. They do not take full advantage of the end-to-end deep architecture. In contrast, our VITAL tracker emphasizes on the most temporally robust features. The adversarial learning scheme makes the classifier aware a variety of appearance changes. The cost sensitive loss mines hard negative samples to further facilitate classifier learning. Our tracker VITAL performs favorably against state-of-the-art trackers.

Conclusion

In this paper we integrate adversarial learning into the tracking-by-detection framework to reduce overfitting on single frames. We adaptively dropout the discriminative features in single frame which draws the classifier attention. It enables the classifier to focus on the temporal robust features which are originally diminished during the training process. The adaptive dropout is achieved via adversarial learning to predict discriminative features according to different inputs. It enriches the target appearances in the feature space and augment the positive samples. Meanwhile, we use the cost sensitive loss to reduce the effect from easy negative samples. Extensive experiments on benchmarks demonstrate that our VITAL tracker performs favorably against state-of-the-art trackers.

References