Siamese Cascaded Region Proposal Networks for Real-Time Visual Tracking
Heng Fan, Haibin Ling
Introduction
Visual tracking is one of the most fundamental problems in computer vision, and has a long list of applications such as robotics, human-machine interaction, intelligent vehicle, surveillance and so forth. Despite great advances in recent years, visual tracking remains challenging due to many factor including occlusion, scale variation, etc.
Recently, Siamese network has drawn great attention in the tracking community owing to its balanced accuracy and speed. By formulating object tracking as a matching problem, Siamese trackers aim to learn offline a generic similarity function from a large set of videos. Among these methods, the work of proposes a one-stage Siamese-RPN for tracking by introducing the regional proposal network (RPN), originally used for object detection , into Siamese network. With the proposal extraction by RPN, this approach simultaneously performs classification and localization from multiple scales, achieving excellent performance. Besides, the use of RPN avoids applying the time-consuming pyramid for target scale estimation , resulting in a super real-time solution.
Despite having achieved promising result, Siamese-RPN may drift to the background especially in presence of similar semantic distractors (see Fig. 1). We identify two reasons accounting for this.
First, the distribution of training samples is imbalanced: (1) positive samples are far less than negative samples, leading to ineffective training of the Siamese network; and (2) most negative samples are easy negatives (non-similar non-semantic background) that contribute little useful information in learning a discriminative classifier . As a consequence, the classifier is dominated by the easily classified background samples, and degrades when encountering difficult similar semantic distractors.
Second, low-level spatial features are not fully explored. In Siamese-RPN (and other Siamese trackers), only features of the last layer, which contain more semantic information, are explored to distinguish target/background. In tracking, nevertheless, background distractors and the target may belong to the same category, and/or have similar semantic features . In such case, the high-level semantic features are less discriminative in distinguishing target/background.
In addition to the issues above, the one-stage Siamese-RPN applies a single regressor for target localization using pre-defined anchor boxes. These boxes are expected to work well when having a high overlap with the target. However, for model-free visual tracking, no prior information regarding the target object is known, and it is hard to estimate how the scale of target changes. Using pre-defined coarse anchor boxes in a single step regression is insufficient for accurate localization (see again Fig. 1).
The class imbalance problem is addressed in two-stage object detector (e.g., Faster R-CNN ). The first proposal stage rapidly filters out most background samples, and then the second classification stage adopts sampling heuristics such as a fixed foreground-to-background ratio to maintain a manageable balance between foreground and background. In addition, two steps of regressions achieve accurate localization even for objects with extreme shapes.
Motivated by the two-stage detector, we propose a multi-stage tracking framework by cascading a sequence of RPNs to solve the class imbalance problem, and meanwhile fully explore features across layers for robust visual tracking.
2 Contribution
As the first contribution, we present a novel multi-stage tracking framework, the Siamese Cascaded RPN (C-RPN), to solve the problem of class imbalance by performing hard negative sampling . C-RPN consists of a sequence of RPNs cascaded from the high-level to the low-level layers in the Siamese network. In each stage (level), an RPN performs classification and localization, and outputs the classification scores and the regression offsets for the anchor boxes in this stage. The easy negative anchors are then filtered out, and the rest, treated as hard examples, are utilized as training samples for the RPN of the next stage. Through such process, C-RPN performs stage by stage hard negative sampling. As a result, the distributions of training samples are sequentially more balanced, and the classifiers of RPNs are sequentially more discriminative in distinguishing more difficult distractors (see Fig. 1).
Another benefit of C-RPN is the more accurate target localization compared to the one-stage SiamRPN . Instead of using the pre-defined coarse anchor boxes in a single regression step, C-RPN consists of multiple steps of regressions due to multiple RPNs. In each stage, the anchor boxes (including locations and sizes) are adjusted by the regressor, which provides better initialization for the regressor of next stage. As a consequence, C-RPN progressively refines the target bounding box, leading to better localization as shown in Fig. 1.
Leverage features from different layers in the neural networks has been proven to be beneficial for improving model discriminability . To fully explore both the high-level semantic and the low-level spatial features for visual tracking, we make the second contribution by designating a novel feature transfer block (FTB). Instead of separately using features from a single layer in one RPN, FTB enables us to fuse the high-level features into low-level RPN, which further improves its discriminative power to deal with complex background, resulting in better performance of C-RPN. Fig. 2 illustrates the framework of C-RPN.
Last but not least, the third contribution is to implement a tracker based on the proposed C-RPN. In extensive experiments on six benchmarks, including OTB-2013 , OTB-2015 , VOT-2016 , VOT-2017 , LaSOT and TrackingNet , our C-RPN consistently achieves the state-of-the-art results and runs in real-time.
Related Work
Visual tracking has been extensively researched in recent decades. In the following we discuss the most related work, and refer readers to for recent surveys.
Deep tracking. Inspired by the successes in image classification , deep convolutional neural network (CNN) has been introduced into visual tracking and demonstrated excellent performances . Wang et al. propose a stacked denoising autoencoder to learn generic feature representation for object appearance modeling in tracking. Wang et al. introduce a fully convolutional neural network tracking (FCNT) approach by transferring the pre-trained deep features to improve tracking accuracy. Ma et al. replace hand-craft features in correlation filter tracking with deep features, achieving remarkable gains. Nam and Han propose a light architecture of CNNs with online fine-tuning to learn generic feature for tracking target. Fan and Ling extend this approach by introducing a recurrent neural network (RNN) to capture object structure. Song et al. apply adversary learning in CNN to learn richer representation for tracking. Danelljan et al. propose continuous convolution filters for correlation filter tracking, and later optimize this method in .
Siamese tracking. Siamese network has attracted increasing interest for visual tracking because of its balanced accuracy and accuracy. Tao et al. utilize Siamese network to off-line learn a matching function from a large set of sequences, then use the fixed matching function to search for the target in a local region. Bertinetto et al. introduce a fully convolutional Siamese network (SiamFC) for tracking by measuring the region-wise feature similarity between the target object and the candidate. Owing to its light structure and without model update, SiamFC runs efficiently at 80 fps. Held et al. propose the GOTURN approach by learning a motion prediction model with the Siamese network. Valmadre et al. use a Siamese network to learn the feature representation for correlation filter tracking. He et al. introduce a two-fold Siamese network for tracking. Wang et al. incorporate attention mechanism into Siamese network to learn a more discriminative metric for tracking. Notably, Li et al. combine Siamese network with RPN, and propose a one-stage Siamese-RPN tracker, achieving excellent performance. Zhu et al. introduce more negative samples to train a distractor-aware Siamese-RPN tracker. Despite improvement, this approach requires large extra training data from other domains.
Multi-level features. The features from different layers in the neural network contain different information. The high-level feature consists of more abstract semantic cues, while the low-level layers contains more detailed spatial information . It has been proven that tracking can be benefited using multi-level features. In , Ma et al. separately use features in three different layers for three correlation models, and fuse their outputs for the final tracking result. Wang et al. develop two regression models with features from two layers to distinguish similar semantic distractors.
Our approach. In this paper, we focus on solving the problem of class imbalance to improve model discriminability. Our approach is related but different from the Siamese-RPN tracker , which applies one-stage RPN for classification and localization and skips the data imbalance problem. In contrast, our approach cascades a sequence of RPNs to address the data imbalance by performing hard negative sampling, and progressively refines anchor boxes for better target localization using multi-regression. Our method is also related to using multi-level features for tracking. However, unlike in which multi-level features are separately used for independent models, we propose a feature transfer block to fuse features across layer for each RPN, improving its discriminative power in distinguishing the target object from complex background.
Siamese Cascaded RPN (C-RPN)
In this section, we detail the Siamese Cascaded RPN (referred to as C-RPN) as shown in Fig. 2.
C-RPN contains two subnetworks: the Siamese network and the cascaded RPN. The Siamese network is utilized to extract the features of the target template and the search region . Afterwards, C-RPN receives the features of and for each RPN. Instead of only using the features from one layer, we apply feature transfer block (FTB) to fuse the features from high-level layers for RPN. An RPN simultaneously performs classification and localization on the feature maps of . According to the classification scores and regression offsets, we filter out the easy negative anchors (e.g., an anchor whose negative confidence is larger than a preset threshold ), and refine the locations and sizes of the rest anchors, which are used for training RPN in the next stage.
As in , we adopt the modified AlexNet to develop our Siamese network. The Siamese network comprises two identical branches, the z-branch and the x-branch, which are employed to extract features from the target template and the search region , respectively (see Fig. 2). The two branches are designed to share parameters to ensure the same transformation applied to both and , which is crucial for the similarity metric learning. More details about the Siamese network can be referred to .
Different from that only uses the features from the last layer of the Siamese network for tracking, we leverage the features from multiple levels to improve model robustness. For convenience in next, we denote and as the feature transformations of and from the conv- layer in the Siamese network with layersFor notation simplicity, we name each layer in the Siamese network in an inverse order, i.e., conv-, conv-, , conv-, conv- for the low-level to the high-level layers..
2 One-Stage RPN in Siamese Network
To ensure classification and regression for each anchor, two convolution layers are utilized to adjust the channels of into suitable forms, denoted as and , for classification and regression, respectively. Likewise, we apply two convolution layers for but keep the channels unchanged, and obtain and . Therefore, the classification scores and the regression offsets for each anchor can be computed as
where is the anchor index, and denotes correlation between and where is served as the kernel. Each is a 2d vector, representing for negative and positive confidences of the anchor. Similarly, each is a 4d vector which represents the offsets of center point location and size of the anchor to groundtruth. Siamese RPN is trained with a multi-task loss consisting of two parts, i.e., the classification loss (i.e., softmax loss) and the regression loss (i.e., smooth loss). We refer readers to for further details.
3 Cascaded RPN
As mentioned earlier, previous Siamese trackers mostly ignore the problem of class imbalance, resulting in degenerated performance in presence of similar semantic distractors. Besides, they only use the high-level semantic features from the last layer, which does not fully explore multi-level features. To address these issues, we propose a multi-stage tracking framework by cascading a set of () RPNs.
For RPNl in the () stage, it receives fused features and of the conv- layer and the high-level layers from FTB, instead of features and from a single separate layer . The and are obtained as follows,
where , , and are derived by performing convolutions on and .
Let denote the anchor set in stage . With classification scores , we can filter out anchors in whose negative confidences are larger than a preset threshold , and the rest are formed into a new set of anchor , which is employed for training RPNl+1. For RPN1, is pre-defined. Besides, in order to provide a better initialization for regressor of RPNl+1, we refine the center locations and sizes of anchors in using the regression results in RPNl, thus generate more accurate localization compared to a single step regression in Siamese RPN , as illustrated in Fig. 4. Fig. 2 shows the cascade architecture of C-RPN.
where is the anchor index in of stage , a weight to balance losses, the label of anchor , and the true distance between anchor and groundtruth. Following , is a 4d vector, such that
where , , and are center coordinates of a box and its width and height. Variables and are for groundtruth and anchor of stage (likewise for , and ). It is worth noting that, different from using fixed anchors, the anchors in C-RPN are progressively adjusted by the regressor in the previous stage, and computed as
For the anchor in the first stage, , , and are pre-defined.
The above procedure forms the proposed cascaded RPN. Due to the rejection of easy negative anchors, the distribution of training samples for each RPN is gradually more balanced. As a result, the classifier of each RPN is sequentially more discriminative in distinguishing difficult distractors. Besides, multi-level feature fusion further improves the discriminability in handing complex background. Fig. 5 shows the discriminative powers of different RPNs by demonstrating detection response map in each stage.
4 Feature Transfer Block
To effectively leverage multi-level features, we introduce FTB to fuse features across layers so that each RPN is able to share high-level semantic feature to improve the discriminability. In detail, a deconvolution layer is used to match the feature dimensions of different sources. Then, different features are fused using element-wise summation, followed a ReLU layer. In order to ensure the same groundtruth for anchors in each RPN, we apply the interpolation to rescale the fused features such that the output classification maps and regression maps have the same resolution for all RPN. Fig. 6 shows the feature transferring for RPNl ().
5 Training and Tracking
Training. The training of C-RPN is performed on the image pairs that are sampled within a random interval from the same sequence as in . The multi-task loss function in Eq. (7) enables us to train C-RPN in an end-to-end manner. Considering that the scale of target changes smoothly in two consecutive frames, we employ one scale with different ratios for each anchor. The ratios of anchors are set to as in .
For each RPN, we adopt the strategy as in object detection to determine positive and negative training samples. We define the positive samples as anchors whose Intersection over union (IOU) with groundtruth is larger than a threshold , and negative samples as anchors whose IoU with groundtruth bounding box is less than a threshold . We generate at most 64 samples from one image pair.
Tracking. We formulate tracking as multi-stage detection. For each video, we pre-compute feature embeddings for the target template in the first frame. In a new frame, we extract a region of interest according to the result in last frame, and then perform detection using C-RPN on this region. In each stage, an RPN outputs the classification scores and regression offsets for anchors. The anchors with negative scores lager then are discarded, and the rest are refined and taken over by RPN in next stage. After the last stage , the remained anchors are regarded as target proposals, from which we determine the best one as the final tracking result using strategies in . Alg. 1 summarizes the tracking process by C-RPN.
Experiments
Implementation detail. C-RPN is implemented in Matlab using MatConvNet on a single Nvidia GTX 1080 with 8GB memory. The backbone Siamese network adopts the modified AlexNet by removing group convolutions. Instead of training from scratch, we borrow the parameters from the pretrained model on ImageNet . During training, the parameters of first two layers are frozen. The number of stages is set to 3. The thresholds , and are empirically set to 0.95, 0.6 and 0.3. C-RPN is trained end-to-end over 50 epochs using SGD, and the learning rate is annealed geometrically at each epoch from to . We train C-RPN using the training data from for experiment under Protocol II on LaSOT , and using VID and YT-BB for other experiments.
Note that the comparison with Siamese-RPN is fair since the same training data is used for training.
We conduct experiments on the popular OTB-2013 and OTB-2015 which consist of 51 and 100 fully annotated videos, respectively. C-RPN runs at around 36 fps.
Following , we adopt the precision plot in one-pass evaluation (OPE) to assess different trackers. The comparison with 14 state-of-the-art trackers (SiamRPN , DaSiamRPN , TRACA , ACT , BACF , ECO-HC , CREST , SiamFC , Staple , PTAV , SINT , CFNet , HDT and HCFT ) is shown in Fig. 7. C-RPN achieves the best performance on both two benchmarks. In specific, we obtain the 0.675 and 0.663 precision scores on OTB-2013 and OTB-2015, respectively. In comparison with the baseline one-stage SiamRPN with 0.658 and 0.637 precision scores, we obtain improvements by 1.9% and 2.6%, showing the advantages of multi-stage RPN in accurate localization. DaSiamRPN uses extra negative training data from other domains to improve the ability to handle similar distractors, and obtains 0.655 and 0.658 precision scores. Without using extra training data, C-RPN outperforms DaSiamRPN by 2.0% and 0.5%. More results and comparisons on OTB-2013 and OTB-2015 are shown in the supplementary material.
2 Experiments on VOT-2016 and VOT-2017
VOT-2016 consists of 60 sequences, aiming at assessing the short-term performance of trackers. The overall performance of a tracking algorithm is evaluated using Expected Average Overlap (EAO) which takes both accuracy and robustness into account. The speed of a tracker is represented with a normalized speed (EFO).
We evaluate C-RPN on VOT-2016, and compare it with 11 trackers including the baseline SiamRPN and other top ten approaches in VOT-2016. Fig. 8 shows the EAO of different trackers. C-RPN achieves the best results, significantly outperforming the baseline SiamRPN and other approaches. Tab. 1 lists the detailed comparisons of different trackers on VOT-2016. From Tab. 1, we can see that C-RPN outperforms other trackers in both accuracy and robustness, and runs efficiently.
VOT-2017 contains 60 sequences, which are developed by replacing the least 10 challenging videos in VOT-2016 with 10 difficult sequences. Different from VOT-2016 , VOT-2017 introduces a new real-time experiment by taking into both tracking performance and efficiency. We compare C-RPN with SiamRPN and other top ten approaches in VOT-2017 using the EAO of baseline and real-time experiments, as shown in Tab. 2. From Tab. 2, C-RPN achieves a EAO score of 0.289, which significantly outperforms the one-stage SiamRPN with EAO score of 0.243. In addition, compared with LSART and CFWCR , C-RPN shows competitive performance. In real-time experiment, C-RPN obtains the best result with EAO score of 0.273, outperforming all other trackers.
3 Experiment on LaSOT
LaSOT is a recent large-scale dataset aiming at both training and evaluating trackers. We compare C-RPN to 35 approaches, including ECO , MDNet , SiamFC , VITAL , StructSiam , TRACA , BACF and so forth. We refer readers to for more details about the compared trackers. We do not compare C-RPN to Siamese-RPN because neither its implementation nor results on LaSOT are available.
Following , we report the results of success (SUC) for different trackers as shown in Fig. 9. It shows that our C-RPN outperforms all other state-of-the-art trackers under two protocols. We achieve SUC scores of 0.459 and 0.455 under protocol I and II, outperforming the second best tracker MDNet with SUC scores 0.413 and 0.397 by 4.6% and 5.8, respectively. In addition, C-RPN runs at around 23 fps on LaSOT, which is more efficient than MDNet with around 1 fps. Compared with the Siamese network-based tracker SiamFC with 0.358 and 0.336 SUC scores, C-RPN gains the improvements by 11.1% and 11.9%. Due to limited space, we refer readers to supplementary material for more details about results and comparisons on LaSOT.
4 Experiment on TrackingNet
TrackingNet is proposed to assess the performance of a tracker in the wild. We evaluate C-RPN on its testing set with 511 videos. Following , we use three metrics precision (PRE), normalized precision (NPRE) and success (SUC) for evaluation. Tab. 3 demonstrates the comparison results to trackers with top PRE scoresThe result of C-RPN on TrackingNet is evaluated by the server provided by the organizer at http://eval.tracking-net.org/web/challenges/challenge-page/39/leaderboard/42. The results of compared trackers are reported from . Full comparison is shown in the supplementary material., showing that C-RPN achieves the best results on all three metrics. In specific, C-RPN obtains the PRE score of 0.619, NPRE score of 0.746 and SUC score of 0.669, outperforming the second best tracker MDNet with PRE score of 0.565, NPRE score of 0.705 and SUC score of 0.606 by 5.4%, 4.1% and 6.3%, respectively. Besides, C-RPN runs efficiently at a speed of around 32 fps.
5 Ablation Experiment
To validate the impact of different components, we conduct ablation experiments on LaSOT (Protocol II) and VOT-2017 .
Number of stages? As shown in Tab. 4, adding the second stage significantly improves one-stage baseline. The SUC on LaSOT is improved by 2.9% from 0.417 to 0.446, and the EAO on VOT-2017 is increased by 3.5% from 0.248 to 0.283. The third stage produces 0.9% and 0.6% improvements on LaSOT and VOT-2017, respectively. We observe that the improvement by the second stage is higher than that by the third stage. This suggests that most difficult background is handled in the second stage. Adding more stages may lead to further improvements, but also the computation (speed from 48 to 23 fps).
Negative anchor filtering? Filtering out the easy negatives aims to provide more balanced training samples for RPN in next stage. To show its effectiveness, we set threshold to 1 such that all refined anchors will be send to the next stage. Tab. 5 shows that removing negative anchors in C-RPN can improve the SUC on LaSOT by 1.6% from 0.439 to 0.455, and the EAO on VOT-2017 by 0.7% from 0.282 to 0.289, respectively, which evidences balanced training samples are crucial for training more discriminative RPN.
Feature transfer block? As demonstrated in Tab. 6, FTB improves the SUC on LaSOT by 1.3% from 0.442 to 0.455 without losing much efficiency, and the EAO on VOT-2017 by 1.1% from 0.278 to 0.289, validating the effectiveness of multi-level feature fusion in improving performance.
These studies show that each ingredient brings individual improvement, and all of them work together to produce the excellent tracking performance.
Conclusion
In this paper, we propose a novel multi-stage framework C-RPN for tracking. Compared with previous arts, C-RPN demonstrates more robust performance in handling complex background such as similar distractors by performing hard negative sampling within a cascade architecture. In addition, the proposed FTB enables effective feature leverage across layers for more discriminative representation. Moreover, C-RPN progressively refines the target bounding box using multiple steps of regressions, leading to more accurate localization. In extensive experiments on six popular benchmarks, C-RPN consistently achieves the state-of-the-art results and runs in real-time.