A Novel Performance Evaluation Methodology for Single-Target Trackers

Matej Kristan, Jiri Matas, Ales Leonardis, Tomas Vojir, Roman Pflugfelder, Gustavo Fernandez, Georg Nebehay, Fatih Porikli, Luka Cehovin

Introduction

Visual tracking is a rapidly evolving field that has been increasingly attracting attention of the vision community. It offers many scientific challenges and it emerges in other computer vision problems such as motion analysis, event detection and activity recognition. A steady increase of hardware performance and its price reduction have opened a vast application potential for tracking algorithms including surveillance systems, automotive systems, transport, sports analytics, medical imaging, mobile robotics, film post-production and human-computer interfaces.

The activity in the field is reflected in abundance of new tracking algorithms presented in journals and at conferences summarized in the many survey papers, e.g., . However, the boom in tracker proposals has not been accompanied by standardization of the methodology for their objective comparison.

One of the most influential performance analysis efforts for object tracking is PETS (Performance Evaluation of Tracking and Surveillance) . The first PETS workshop took place in 2000 aiming at evaluation of visual tracking algorithms for surveillance applications. Its focus gradually shifted to high-level event interpretation algorithms. Other frameworks and datasets have been presented since, but these focused on evaluation of surveillance systems and event detection, e.g., CAVIARhttp://homepages.inf.ed.ac.uk/rbf/CAVIARDATA1, i-LIDS http://www.homeoffice.gov.uk/science-research/hosdb/i-lids, ETISEOhttp://www-sop.inria.fr/orion/ETISEO, change detection , sports analytics (e.g., CVBASEhttp://vision.fe.uni-lj.si/cvbase06/), specialized on tracking specific objects like faces, e.g., FERET , or tracking for autonomous vehicles, e.g., KITTI . Recently, several works have been published in the broad area of model-free visual object tracking evaluation, eg., and following the success of the VOT challenges a performance evaluation benchmark for multiple target tracking was presented as well .

There are several important subfields in visual tracking, ranging from multi-camera, multi-target , to single-target trackers. These subfields are quite diverse, without a unified evaluation methodology and specific methodologies have to be tailored to each subfield.

In this paper, single-camera, single-target, model-free, causal trackers, applied to short-term tracking are considered. The model-free property means that the only supervised training example is provided by the bounding box in the first frame. The short-term tracking means that the tracker does not perform re-detection after the target is lost. Drifting off the target is considered a failure. The causality means that the tracker does not use any future frames to infer the object position in the current frame. The tracker output is specified by a rotated bounding box.

The evaluation of new tracking algorithms depends on three essential components: (1) performance evaluation measures, (2) a dataset and (3) an evaluation system. In the following, the requirements for these components are stated.

Performance measures. A wealth of performance measures have been proposed for single-object tracker evaluation, but there is no consensus on which measure should be preferred. Ideally, measures should clearly reflect different aspects of tracking. Apart from merely ranking, we also need to determine cases when two or more trackers are performing equally well. We require the following: The measures should allow an easy interpretation and should support tracker comparison with a well-defined tracker equivalence.

Datasets. The dataset should allow evaluation of trackers under diverse conditions like partial occlusion, clutter and illumination changes. One approach is to construct a very large dataset, but this does not guarantee diversity in visual attributes and it significantly slows down the process of evaluation. A better approach is to annotate each sequence with the visual attributes occurring in that sequence and perform clustering to reduce the size of the dataset, while keeping it diverse. Annotation is also important for per-attribute tracker analysis. A common approach is to annotate a sequence globally with an attribute if that attribute occurs anywhere in the sequence. The trackers can then be compared only on the sequences corresponding to a particular attribute. However, visual phenomena do not usually last throughout the entire sequence. For example, a partial occlusion might occur at the end of a sequence, while a tracker might fail due to some other effects occurring at the beginning of the sequence. In this case, the failure would be falsely attributed to the occlusion. A per-frame dataset labeling is thus required to facilitate a more precise analysis. This motivates the following requirements: (1) The dataset should be diverse in visual attributes. (2) Per-frame annotation of visual attributes is required.

Evaluation systems. For a rigorous evaluation, an evaluation system that performs the same experiment on different trackers using the same dataset is required. A wide-spread practice is to initialize the tracker in the first frame and let it run until the end of a sequence. However, the tracker might fail right at the beginning of the sequence due to some visual degradation, effectively meaning that the system utilized only the first few frames for evaluation of this tracker. Thus the first requirement for the system is that it fully uses the data. This means that once the tracker fails, the system has to detect the failure and reinitialize the tracker. Therefore, a certain level of interaction, that goes beyond simple running until the end of the sequence, is required. Furthermore, the evaluation system has to also account for the fact that the trackers are typically coded in various programming languages and often platform-dependent. This motivates the following set of requirements the evaluation system should meet: (1) Full use of the dataset. (2) Allow interaction with the tracker. (3) Support for multiple platforms. (4) Easy integration with trackers.

2 Our contributions

In this paper we present the following four contributions:

The first contribution is a novel tracker evaluation methodology based on two simple, easy interpretable, performance measures. A significant novelty of the proposed methodology is the use and first of its kind analysis of reinitializations at tracking failures. Reinitialization-based measures are compared theoretically and experimentally to standard counterparts that do not apply reinitialization. We propose a first of its kind tracker ranking methodology that addresses the concept of tracker equivalence and takes into account statistical significance as well as practical difference in tracking accuracy. A new visualization of ranks is proposed as well to aid comparative analysis.

The second contribution is a new dataset and evaluation system. The dataset is constructed by a novel video clustering approach based on visual properties. The dataset is fully annotated, all the sequences are labeled per-frame with visual attributes to facilitate in-depth analysis. The benefits of per-frame attribute annotation are analyzed theoretically and experimentally. The proposed evaluation system enjoys multi-platform compatibility and offers easy integration with trackers. The system has been tested in a large-scale distributed experiment on the VOT2013 and VOT2014 challenges.

The third contribution is a detailed comparative analysis of 38 trackers using the proposed methodology, making it the largest benchmark to date.

The forth contribution is a novel analysis of the sequences in the dataset from the perspective of tracking success.

Preliminary versions of some parts of this paper have been previously published (during the period 2013-2014) in three workshop papers .

The remainder of the paper is structured as follows: In Section 2 the most related work is reviewed and discussed. The new tracker evaluation methodology is presented and theoretically analyzed in Section 3, while the new dataset selection approach, the evaluation system and the results of the experimental analysis are presented in Section 4. Conclusions are drawn in Section 5.

Related work

A wealth of performance measures have been proposed for single-object tracker evaluation. These range from basic measures like center error , region overlap , tracking length , failure rate , F-score , pixel-based precision , to more sophisticated measures, such as CoTPS , which combine several measures. A nice property of the combined measures is that they provide a single score to rank the trackers. A downside is that they offer little insight into the tracker performance which limits their interpretability. All measures strongly depend on the experimental setup within which they are computed. For example, some evaluation protocols, like Wu et al. and Smeulders et al., initialize the trackers at the beginning of the sequence and let them run until the end. Measures computed in such a setup are inappropriate for short-term tracking evaluation, since the trackers are not expected to perform re-detection. The values of performance measures thus become irrelevant after the point of tracking failure. Including the frames past the point of failure in the computation of a global performance measure introduces significant distortions since failures closer to the beginning of the sequence are significantly more penalized than failures occurring later in the sequence.

While some authors choose several basic measures to compare their trackers, recent studies have shown that many measures are correlated and do not reflect diverse aspects of tracking performance. In this respect, choosing a large number of measures may in fact again bias results toward some particular aspects of tracking performance. Smeulders et al. propose using two measures: an F-score calculated at the Pascal region overlap criterion (threshold 0.50.5) and a center error. Note that the F-score based measure was originally designed for object detection. The threshold 0.50.5 is also rather high and there is no clear justification of why exactly this threshold should be used to compare trackers since it is hardly an indicator of tracking failure (see examples in Figure 1).

Since the center error becomes arbitrary high once the tracker fails, Wu et al. propose to measure the percentage of frames in which the center distance is within some prescribed threshold. However, this threshold significantly depends on the object size, which makes this particular measure quite brittle. A normalized center error measured during successful tracks may be used to alleviate the object size problem, however, the results in show that the trackers do not differ significantly under this measure which makes it less appropriate for tracker comparison. As an additional measure, propose an area under a ROC-like plot of thresholded overlaps. Recently, have shown that this is equivalent to the average region overlap measure computed from all frames of sequences. In fact, based on an extensive analysis of performance measures, Čehovin et al. argue that the region overlap is superior to the center error.

While it is important to study and evaluate the tracker performance separately in terms of several less correlated performance measures, it is sometimes required to rank trackers in a single rank list. In this case a convenient strategy is to combine these measures into rank averaging, similarly to what was done in the change detection challenge . In rank averaging, competing algorithms are ranked with respect to several performance measures and their ranks are averaged. This simulates competition of trackers with respect to different performance measures and assumes equal importance of all measures. The fact that trackers are ranked along each measure induces normalization of measures to a common scale prior to averaging.

Several authors propose to visually compare tracking performance via performance summarization plots. These plots show the percentage of frames for which the estimated object location is within some threshold distance of the ground truth. Most notable are precision plots , which measure the object location accuracy in terms of center error. Alternatively, success plots use the region overlap instead. Salti et al., implicitly account for variable threshold dependency by plotting the percentage of correctly tracked frames with respect to the mean region overlap within these frames. Čehovin et al. propose a similar visualization, but they apply a single, zero, threshold on the overlap. A tracker is thus represented as a single point in this 2D space, rather than a curve, which allows easier comparison. A drawback of performance plots is that they typically become cluttered when comparing several trackers on several sequences in the same plot. To address this, Smeulders et al. calculate a performance measure per sequence for a tracker and order these values from highest to lowest, thus obtaining a so-called survival curve. The performance of several trackers is then compared on the entire dataset by visualizing their survival curves.

2 Datasets

It is a common practice to compare trackers on many publicly-available sequences, which have became a de-facto standard in evaluation of new trackers. However, many of these sequences lack a standard ground truth labeling, which makes comparison of algorithms difficult. To sidestep this issue, Wu et al. have proposed a protocol for stochastic tracker evaluation on a selected dataset that does not require ground truth labels. A similar approach was adapted by to evaluate tracking algorithms on long sequences. Datasets with various visual phenomena equally represented are not usually used. In fact, many popular sequences are conceptually similar, which makes the results biased toward some particular types of the phenomena. To address this issue, Wu et al. annotated each sequence with several visual attributes and report tracker performance with respect to each attribute separately. However, a per-frame annotation is not provided and not all sequences are in color, which makes results skewed with proportions of color and gray sequences. Recently, Smeulders et al. , have presented a very large dataset called ‘Amsterdam Library of Ordinary Videos’ (ALOV). The dataset is composed of over three hundred sequences collected from published datasets and additional YouTube videos. The sequences are assigned to one of thirteen classes of difficulty and, with the exception of ten long sequences, are kept short to increase the diversity. The sequences are not annotated per-frame with visual attributes, some sequences contain cuts and ambiguously defined targets such as fireworks which makes the dataset inappropriate for short-term tracking evaluation.

3 Evaluation systems

The most notable and general evaluation systems are ODViS , VIVID , ViPER . The former two focus on the design of surveillance systems, while the latter is a set of utilities/scripts for annotation and computation of different types of performance measures. The recently proposed ViCamPEv toolkit is dedicated to testing a pre-determined set of OpenCV-based basic trackers. None of these systems support interaction with the tracker, which limits their applicability. Collecting the results from the existing publications is an alternative for benchmarking trackers. Pang et al. have proposed a page-rank-like approach to data-mine the published results and compile unbiased ranked performance lists. However, as the authors state in their paper, the proposed protocol is not appropriate for creating ranks of the recently published trackers due to the lack of sufficiently many publications that would compare these trackers.

Visual object tracker evaluation

The proposed methodology assumes that the evaluation system and the dataset fulfill the requirements stated in Section 1.1, i.e., (i) the dataset is per-frame annotated by visual attributes and the object positions are denoted by possibly rotated bounding boxes, (ii) trackers are run on each sequence of the dataset. Once the tracker drifts off the target, the system detects a tracking failure and re-initializes the tracker. All trackers are run multiple times to account for their possibly stochastic nature.

Based on the recent analysis of widely-used performance measures two weakly-correlated and easily interpretable measures were chosen: (i) accuracy and (ii) robustness. The accuracy at time-step tt measures how well the bounding box AtTA_{t}^{T} predicted by the tracker overlaps with the ground truth bounding box AtGA_{t}^{G} and is defined as the intersection-over-union

The robustness is the number of times the tracker failed, i.e., drifted from the target, and had to be reinitialized. A re-initialization is triggered when the overlap (Eq. 1) drops to zero.

In contrast to accuracy measurements, a single measure of robustness per experiment repetition is obtained. Let F(i,k)F(i,k) be the number of times the ii-th tracker failed in the experiment repetition kk over a set of frames. The average robustness of the ii-th tracker is then

The overall performance on the dataset can be estimated as the weighted average of the per-sequence performance measures, with weights proportional to the lengths of the sequences. Note that this is equivalent to concatenating the sequence of per-frame overlaps/failures from the entire dataset into a single super-sequence and calculating the two averages in (2) and (3). Similarly, per-visual-attribute performance can be evaluated for a specific attribute by collecting all the frames labelled as that attribute into an attribute super-sequence and calculating (2) and (3).

For a fair comparison, we propose a ranking-based methodology akin to but we introduce the concept of equally-ranked trackers. For each tracker, a group of so-called equivalent trackers containing trackers performing indistinguishably is determined and a corrected rank is then calculated. There are several choices for calculating the correction, e.g., one could take the min, max or mean of ranks in the group. The least conservative choice is max, since it always penalizes a tracker if the equivalency test cannot confirm the difference from a lower-ranked tracker, and on the other hand, the min is most conservative, since it always makes a correction in interest of the tracker. In the subsequent evaluation we use the mean of the ranks as a compromise between the two extrema. Note that the concept of equivalent trackers is not transitive, and should not be mistaken for the standard equivalence relation. For example, consider trackers T1T_{1}, T2T_{2} and T3T_{3}. It may happen that a tracker T2T_{2} performs indistinguishably from T1T_{1} and T3T_{3}, but this does not necessarily mean that T1T_{1} performs equally well as both, T2T_{2} and T3T_{3}. The equality of trackers should therefore be established for each tracker separately. Two types of tests for establishing performance equivalence are considered in the following.

A per-frame accuracy is available for each tracker. One way to gauge equivalence in this case is to apply a paired test to determine whether the difference in accuracies is statistically significant. When the differences are distributed normally, the Student’s t-test, which is often used in the aeronautic tracking research , is the appropriate choice. However, in a preliminary study we have applied Anderson-Darling tests of normality and have observed that the accuracies in frames are not always distributed normally, which might render the t-test inappropriate. As an alternative, the Wilcoxon Signed-Rank test as in is applied that tests a null hypothesis that differences come from a distribution with zero median (see for further details).

In case of robustness, several measurements of the number of tracker failures over the entire sequence in different runs is obtained. However, these cannot be paired, and the Wilcoxon Rank-Sum (also known as Mann-Whitney U-test) is used instead to test the difference in the average number of failures. This is a two-sided rank sum test which tests the null hypothesis that the number of failures of two trackers are independent samples from distributions with equal medians (see for further details).

1.2 Tests of practical differences

Note that statistical difference does not necessarily imply a practical difference , which is particularly important in equivalency tests for accuracy. The practical difference is a level of difference in accuracy that is considered negligibly small. This level can come from the noise in annotation, the fact that multiple ground truth annotations of bounding boxes might be equally valid, or simply from the fact that very small differences in tracking accuracy are negligible from a practical point of view. Therefore, a pair of trackers is considered to perform equally well in accuracy if their difference in performance is not statistically significant or if it fails the practical difference test.

where γt\gamma_{t} is the practical difference threshold corresponding to the tt-th frame.

1.3 Visualization of results

Results can be visualized by the accuracy-robustness plots proposed by in which a tracker is presented as a point in terms of accuracy and robustness. The accuracy is defined as in (2), while the robustness is converted into a probability of tracker failing after SS frames, thus scaling robustness into the range between zero and one. Since we have extended the methodology of to rankings, we also extend the visualization. In particular, the rank results can be displayed using the accuracy-robustness (AR) rank plots. Since each tracker is presented in terms of its rank with respect to robustness and accuracy, we can plot it as a single point on the corresponding 2D AR-rank plot. Trackers that perform well relative to the others are positioned in the top-right part of the plot, while the, relatively speaking, poorly-performing trackers occupy the bottom-left part.

2 Theoretical comparison to related works

The most related works to the performance evaluation methodology presented in this paper are the methodologies presented by Wu et al. and Smeulders et al. . In principle, all the methodologies use global averages based on the overlaps of tracker bounding boxes and ground truth. The main difference between and is that computes the average-overlap-based measure (like our approach), while computes an F-score at 0.5 overlap. For short-term tracking, the tracker is not required to re-detect the target after losing it. This means that the tracker is not required to report the target loss and the F-score from reduces to precision, i.e., the ratio of frames in which the overlap with ground truth is grater than 0.5. Applying such a high threshold reduces the strength of the performance measure. For example, consider a pair of trackers, tracker A and B: tracker A performs at 0.47 overlap, whereas tracker B performs at 0.1 overlap and none of the trackers ever drifts off the target. The F-score at overlap 0.5 is zero for both trackers, meaning that the measure cannot discern the performance among the trackers since their overlap is below 0.5. Furthermore, the measure would induce a large distinction between trackers A (F-score 0) and a tracker that performs at overlap 0.5 (F-score 1) even though the difference between both is only 0.03 overlap.

There are three notable differences between our methodology and . The first difference is that our methodology detects tracking failure and applies re-initializations, while the and do not re-initialize, nor detect a failure. The methodology from relies on compensating for this drawback by increasing the number of sequences to 50 and recently proposed using over 300 sequences. The second difference is that our methodology is based on per-frame visual-attribute annotation for per-visual attribute performance evaluation. On the other hand, globally annotate a sequence with all the appearing tributes. Per-visual attribute performance is then computed by using all frames of the sequences globally annotated by a particular attribute. The last difference relates to the ability to state that one tracker performs better than another. While all three methodologies produce ranks, only our methodology accounts for the practical as well as statistical difference and takes into account the noise in ground truth annotation to gauge equivalence of trackers.

The aim of the methodologies is to estimate the tracker overall or per-visual attribute performance and rank trackers according to this estimate. In this respect, the methodologies can be thought of as state estimators in which the hidden state is the tracker true performance (e.g., expected overlap). Thus, methodologies can be studied from the perspective of bias and variance of state estimators. In the following we apply this view to further analyze the properties of estimators in terms of applying re-initialization as well as per-frame visual attribute annotation.

Please see the outline of derivation in Appendix A.

2.2 The importance of per-frame annotation

Experimental evaluation

The tracker comparison methodology from Section 3 was applied to a large-scale experiment, organized as a Visual Object Tracking challengehttp://www.votchallenge.net/ (VOT2014). An annotated dataset (Section 4.2) was constructed and an evaluation system implemented in Matlab/Octave to fulfill the multi-platform, multi-programming language compatibility requirement from Section 1.1. A minimal API is defined to integrate a tracker with the system regardless of the programming language used to implement the tracker. The reader is referred to the evaluation kit document for further details. Researchers were invited to participate by downloading the evaluation kit, to integrate it into their trackers and to run it locally on their machines. The evaluation kit downloaded the VOT2014 dataset and performed a set of pre-defined experiments (Section 4.1.1). To ensure a fair analysis, the authors were instructed to select a single set of parameters for all experiments. This way, the authors of the trackers themselves were responsible for setting the proper parameters and removing possible errors from the tracker implementations. The raw results from the evaluation system were then submitted to the VOT2014 homepage, along with a short description of the trackers and optionally with the binaries or source code to allow the VOT2014 committee further verification of their results.

The VOT2014 challenge includes the following two experiments:

Experiment 1 (baseline) runs a tracker on all sequences in the VOT2014 dataset by initializing it on the ground truth bounding boxes.

Experiment 2 (bounding box perturbation) performs Experiment 1 with noisy bounding boxes. The noise affected the position and size by drawing perturbations uniformly from the ±10%\pm 10\% interval of the ground truth bounding box size and the rotation by drawing uniformly from the ±0.1\pm 0.1 radian range.

All the experiments were automatically performed by the evaluation kithttps://github.com/vicoslab/vot-toolkit. A tracker was run on each sequence 15 times to obtain a better statistics on its performance.

1.2 Tested trackers

In total 3838 trackers were considered in the challenge, most of which had been published in recent years and represent the state-of-the-art. These included 3333 original submissions and 55 baseline highly-cited trackers that were contributed by the VOT committee. We reference the unpublished trackers by the VOT2014 challenge report . For the interested readers a more detailed description of each tracker can be found in the supplementary material and a condensed summary of the trackers is available in Table V.

2 The VOT2014 Dataset

A usual approach to creating a diverse dataset is collecting all sequences from existing datasets. However, a large dataset does not necessarily mean being rich in visual properties. In fact, many sequences may be visually similar and would not contribute to the diversity while they would significantly slow down the evaluation process. We have therefore applied an approach that leads to a dataset that includes various visual phenomena while containing a small number of sequences.

The dataset was prepared as follows. The initial pool included 394394 sequences, including sequences used by various authors in the tracking community, the VOT2013 benchmark , the recently published ALOV dataset , the Online Object Tracking Benchmark and additional, so far unpublished, sequences. The set was manually filtered by removing sequences shorter than 200200 frames, grayscale sequences, sequences containing poorly defined targets (e.g., fireworks) and sequences containing cuts. The following global intensity (it) and spatial (sp) attributes were automatically computed for each of the 193193 remaining sequences:

Illumination change is defined as the average of the absolute differences between the object intensity in the first and remaining frames (it).

Object size change is the sum of averaged local size changes, where the local size change at frame tt is defined as the average of absolute differences between the bounding box area in frame tt and past fifteen frames (sp).

Object motion is the average of absolute differences between ground truth center positions in consecutive frames (sp).

Clutter is the average of per-frame distances between two histograms: one extracted from within the ground truth bounding box and one from an enlarged area (by factor 1.5) outside of the bounding box (it).

Camera motion is defined as the average of translation vector lengths estimated by key-point-based RANSAC between consecutive frames (sp).

Blur was measured by the Bayes-spectral-entropy camera focus measure (it).

Aspect-ratio change is defined as the average of per-frame aspect ratio changes. The aspect ratio change at frame tt is calculated as the ratio of the bounding box width and height in frame tt divided by the ratio of the bounding box width and height in the first frame (sp);

Object color change defined as the change of the average hue value inside the bounding box (it);

Deformation is calculated by dividing the images into 88 ×\times 88 grid of cells and computing the sum of squared differences of averaged pixel intensity over the cells in current and first frame (it).

Scene complexity represents the level of randomness (entropy) in the frames and it was calculated as e=∑i=0255bilog⁡bie=\sum_{i=0}^{255}b_{i}\log b_{i}, where bib_{i} is the number of pixels with value equal to ii (it).

In this way each sequence was represented as a 1010-dimensional feature vector. Sequences were clustered in an unsupervised way using affinity propagation into 1212 clustersThe parameters were automatically set. We checked that small perturbations did not result in different clusterings.. From these, 2525 sequences were manually selected such that the various visual phenomena like, occlusion, were still represented well within the selection.

The selected objects in each sequence are manually annotated by bounding boxes. For most sequences, the authors provide axis-aligned bounding boxes placed over the target. For most frames, the axis-aligned bounding boxes approximated the target well with large percentage of pixels within the bounding box (at least >60%>60\%) belonging to the target. Some sequences contained elongated, rotating or deforming targets and these were re-annotated by rotated bounding boxes. After inspecting all the bounding box annotations, sequences with misplaced original annotations were re-annotated.

Additionally, we labeled each frame in each sequence with five visual attributes that reflect a particular challenge in appearance degradation: (1) camera motion, (2) illumination change, (3) motion change, (4) size change and (5) occlusion. In case a particular frame had none of the five attributes, we labeled the frame as (6) neutral. A summary of sequence properties is presented in Figure 3. The average length of consecutive frames containing an attribute was 335.6335.6 for camera motion, 107.1107.1 for illumination change, 16.916.9 for occlusion, 27.727.7 for motion change, 34.534.5 for occlusion, and 99.599.5 for neutral frames.

The practical difference (Section 3.1.2) strongly depends on the target as well as the free parameters of the annotation model. Ideally, a per-frame estimate of γ\gamma would be required for each sequence, but that would present a significant undertaking. On the other hand, using a single threshold for the entire dataset is too restrictive as the properties of targets vary across the sequences. A compromise can be taken in this case by computing a single threshold per sequence. We propose selecting MM frames per sequence and have JJ expert annotators place the bounding boxes carefully KK times on each frame. In this way N=K×JN=K\times J bounding boxes are obtained per frame. One of the bounding boxes can be taken as a possible ground truth and N−1N-1 overlaps can be computed with the remaining ones. Since all annotations are considered “correct”, any two overlaps should be considered equivalent, therefore the difference between these two overlaps is an example of negligibly small difference. By choosing each of the bounding boxes as ground truth, M(N((N−1)2−N+1))/2M(N((N-1)^{2}-N+1))/2 samples of differences are obtained per sequence. The practical difference threshold per sequence is estimated as the average of these values.

Seven experts have annotated four frames per sequence three times. A single frame with an overlayed ground truth bounding box per sequence was displayed during annotation, serving as a guideline of what should be annotated. Thus a set of 15960 samples of differences was obtained per sequence and used to compute the per-sequence practical difference thresholds. The boxplots of the differences are shown in Figure 4 along with a few frames with overlaid annotations. It is clear that the threshold on practical difference varies over the sequences. For the sequences containing rigid objects, the practical difference threshold is small (e.g., ball), but becomes large for sequences with deformable/articulated objects (e.g., bolt).

3 Study of the methodology parameters

The effect of the burn-in period was further quantified by running several state-of-the-art trackers STRUCK , DSST , SAMF () and KCF and two trackers commonly used as baselines, CT and FRT on the VOT2014 dataset. Table I shows the average accuracy for different values of the burn-in period. The average accuracy is, as expected, slightly reduced when including the frames from the burn-in period. The extent of the drop in accuracy is larger for trackers that fail more often.

3.2 Influence of the re-initialization frame skipping

3.3 Influence of difference tests

The proposed methodology applies tests of performance equivalence by testing statistical and practical differences in tracker performance. In absence of these tests, trackers that perform slightly differently in average values of performance measures would be assigned different ranks even tough the difference in performance might not be statistically significant or below the annotation noise level (practical difference). To quantify the variations in ranks, we sampled 50 random sub-sets of 15 sequences from VOT2014 dataset, ranked DSST, KCF, SAMF, CT, FRT and Struck on all subsets and computed the average of the rank variances over all trackers. Table III reports the rank variations for sequence-pooled and attribute-normalized ranking. The difference tests consistently reduce the variance in both setups.

4 Comparison with related methodologies

Performance evaluation methodologies mainly differ in use of re-initialization and detail of visual attribute annotation in sequences. The theoretical predictions derived in Section 3.2 were again validated experimentally on the VOT2014 dataset using the trackers from previous section.

4.2 Importance of per-frame annotation

5 Application to tracker analysis on VOT2014

The results of the baseline and bounding box perturbation experiments described in Section 4.1.1 are visualized in Figure 7 and summarized in Table V. The AR-rank plots in Figure 7 are obtained by concatenating results of all sequences into a super-sequence, calculate the average performance measures and calculate the ranks from these. In Table V, these results are denoted as sequence-pooled ranking. In addition to rank plots, we show the accuracy/robustness raw plots (AR-raw) as proposed in as well. Note that the AR-raw plots compute the robustness as the probability of a tracker still tracking after SS frames. This parameter affects only scaling, but does not change the order of trackers. We chose S=100S=100 to fully utilize the horizontal space in the AR-raw plots.

In terms of accuracy, the top-performing trackers are DSST, SAMF, KCF and DGT. The DSST, SAMF and KCF are correlation-filter-based trackers derived from MOSSE that apply holistic models, i.e., a HOG . In fact, DSST and SAMF are extensions of the KCF tracker. The similarity in design is reflected in the AR plots (e.g., Figure 7). Note that these trackers form a cluster in the AR-rank space.

It is interesting to further study trackers that apply similar concepts for target localization. MatFlow extends Matrioska by applying a flock-of-trackers variant BDF. At a comparable accuracy ranks, the MatFlow by far outperforms the original Matrioska in robustness. The boost in robustness ranks might be attributed to addition of BDF, which is supported by the fact that BDF alone outperforms in robustness the flock-of-trackers tracker FoT as well as trackers based on variations of FoT, i.e., aStruck, HMM-TxD and dynMS. This speaks of resiliency to outliers in flock selection in BDF.

Two trackers combine color-based mean shift with flow, i.e., dynMS and HMM-TxD and obtain comparable ranks in robustness, however, the HMM-TxD achieves a significantly higher accuracy rank, which might be due to considerably more sophisticated tracker merging scheme in HMM-TxD. Both methods are outperformed in robustness by the scale-adaptive mean shift eASMS that applies motion prediction and colour space selection.

The set of evaluated trackers included the original Struck and two variations, TStruck and aStruck. TStruck is a CUDA-speeded-up TStruck and performs quite similarly to the original Struck in baseline and noise experiment. The aStruck applies the flock-of-trackers for scale adaptation in Struck and improves in robustness on the baseline experiment, but is ranked lower in the noise experiment. This implies that estimation of fewer parameters in Struck results in more accurate and robust performance in cases of poor initialization. This is consistent with the results of comparison of PLT trackers, which are derived from Struck. Note that these trackers by far outperform Struck, which further supports the importance of feature selection in PLT trackers.

Figure 9 shows the per-visual attribute normalized AR-rank plot for the baseline experiment. This plot was obtained by ranking trackers with respect to each attribute and averaging the ranking lists. In Table V, these results are denoted as per-attribute normalization. The AR-raw plot in Figure 9 was obtained by averaging per-attribute average raw performance measures. The general layout of the trackers is similar to the sequence-pooled AR plots in Figure 7, but there are differences in local ranks. The reason is that the sequence-pooled plots significantly depend on the distribution of the visual attributes in the dataset. This is confirmed by noting that the most strongly presented attributes in our dataset are camera motion and object motion (Figure 3) and by observing that the structure of the AR-rank plot for the baseline experiment (Figure 7) is very similar to the camera motion and object motion AR-rank plots from Figure 9. The attribute-normalized AR plots in Figure 8 removes this bias, giving equal importance to all the visual attributes. Averaging the accuracy and robustness ranks in the per-attribute normalization setup, the top performing trackers are DSST, SAMF, KCF, DGT and PLT trackers (see Table V). For reference, we also report the results for the sequence-normalized ranking which ranks trackers with respect to each sequence separately and averages the ranking lists. The resulting plots are shown in the bottom row of Figure 9. Observe that the general distribution of the trackers remains similar to the sequence-pooled plots Figure 7, reflecting the influence of the dominant visual attributes in the dataset. The most apparent difference is that the trackers are less dispersed in the AR-rank space. This is because 25 ranking lists are averaged, indicating that the tracker ranking lists vary over the individual sequences and are consequently pulled to the average rank by averaging.

Note that majority of the tested trackers are highly competitive. This is supported by the fact that the trackers, that are often used as baseline trackers, NCC, MIL, CT, FRT and IVT, occupy the bottom-left part of the AR-rank plots. Obviously these approaches vary in accuracy and robustness and are thus spread perpendicularly to the bottom-left-to-upper-right diagonal of AR-rank plots. In both experiments, the NCC is the least robust tracker. The Struck, which is often considered a state-of-the-art tracker is positioned in the middle of the AR plots, which further supports the quality of the tested trackers.

Next, we have ranked the individual types of visual degradation according to the tracking difficulty they present to the tested trackers. The expected number of failures per hundred frames was computed on each attribute for all trackers. The median of these per visual attribute was taken as a measure of tracking difficulty (see Table VI). The properties that present most difficulty are occlusion, motion change and size change, followed by camera motion and illumination change. Subsequences that do not contain any specific attribute (neutral) present little difficulty for the trackers in general as most trackers do not fail on such intervals.

6 Results of Sequence analysis

A further analysis was conducted to gain an insight into the dataset from a tracker perspective. For each sequence we have analyzed if a particular tracker failed at least once at a particular frame (Figure 10). By counting how many trackers failed at each frame, the level of difficulty can be visualized by the difficulty curve for each sequence (Figure 11). From these curves two measures of sequence difficulty are derived: area and max. The area is a sum of frame-wise values from the difficulty curve normalized by the number of frames, while the max is the maximum on this curve. The former indicates the average level of difficulty of a sequence, and the latter reflects the difficulty of the most difficult part in the sequence. Table VII summarizes the area and max values for all sequences. A high value of the area suggests that such sequence is challenging in a considerable number of frames. For example, the area for the david sequence is smaller than the area for the woman sequence, which suggests that david sequence is less challenging that the woman sequence. A large max indicates the presence of difficult frames. For example, a significant peak in the woman sequence (frame 566) suggests that this sequence contains a subsequence around this frame which is challenging to most of the trackers. In case of drunk sequence, the corresponding max value is 33 (see Table VII), thus almost all trackers successfully track the target.

Using the area measure the sequences were labeled by the following four levels of difficulty: Hard (area greater than 3.003.00), intermediate (area between 3.003.00 and 2.002.00), intermediate/easy (area between 1.001.00 and 1.001.00) and easy (area less than 1.001.00) (see Table VII). These levels were defined by manually clustering the areas into four clear clusters. Surprisingly, the david sequence (Figure 11) shows a small area in this study, although the sequence is usually considered in the community to be challenging and it is commonly referred in the literature. One explanation might be that the trackers are over-fitted to this sequence since it is so often used in evaluation and development. An alternative explanation might be that the sequence is actually not very challenging for tracking, but appears to be to a human observer. The popularity would then be explained by the fact that it is appealing to demonstrate good tracking performance on a sequence that appears difficult, even though it might not be. The analysis also shows that the motocross, hand2, diving, fish2, bolt and hand1 are the most challenging sequences. Most of the difficulties in these sequences arise from changes in camera and object motion as well as from rapid changes in object size. For example, motocross is hard because all three aforementioned nuisances occur simultaneously while the hand2 sequence shows challenging pose variations of the person’s hand. The diving sequence shows significant changes in object size, in bolt sequence both motions camera and object occur simultaneously, while the fish2 sequence shows challenging pose variations of the object.

Easy to intermediate sequences might remain valuable for tracker comparison as these sequences still conceal challenges in particular frames. These sequences are identified by considering max in Table VII. For example, almost all trackers fail at frame 7777 of the jogging sequence. A closer look at this frame and previous frames shows a complete occlusion of the object. Similarly, the woman sequence at frame 566566 (Figure 11) contains camera zooming which makes 1919 out of 3838 trackers fail. The bicycle sequence also shows a peak in the difficulty curve at frame 176176 (Figure 11). In this part of the sequence, an object is occluded, which is immediately followed by a shadow cast over the target. A significant peak is also present in the bolt sequence (Figure 11) at frame 1717, at which many trackers fail. A closer look at the frame and its neighbouring frames shows a significant object motion between the frames as a cause of failures.

Conclusion

In this paper a novel tracker performance evaluation methodology was presented. Requirements for the performance measures, the dataset and the evaluation system are defined and a new evaluation methodology is proposed which aims at a simple, easily interpretable, tracker comparison. The proposed methodology is the first of its kind to account for the tracker equivalence by considering statistical significance and practical differences. A new dataset and a cross-platform-compatible evaluation system were presented. The dataset consists of 25 color sequences, which are per-frame annotated by visual attributes and rotated boxes. Effects of re-initialization and per-frame annotation are studied theoretically and the theoretical predictions are verified with experiments. The novel performance evaluation was applied to comparison of 38 trackers, making it the largest benchmark to date. Using the benchmark, the dataset was analyzed from perspective of per-sequence and per-visual-attribute tracking difficulty. The raw results of all trackers are publicly available from the VOT homepage for reproduction of the results in this paper and to allow comparison with new trackers.

The results of an exhaustive analysis show that trackers tend to specialize either for robustness or accuracy. None of the trackers consistently outperformed the others by all measures at all sequence attributes. The top-performing trackers include trackers with holistic as well as part-based visual models. There is some evidence that robustness is achieved by discriminative learning where variants of structured SVM, e.g. PLT, seem promising. Variants of segmentation appear to play a beneficial role in tracking with noisy initializations. This is evident in favorable performance of trackers DGT and PLTs in the noise experiment. But relying strongly on segmentation reduces performance when color significantly changes which is seen in significant deterioration of the DGT on illumination change. Estimation of few parameters likely increases tracking robustness at reduced accuracy. Attribute-wise analysis shows that motion prediction significantly improves performance during dynamic target motion. Results show that evaluating trackers by pooling results from sequences largely depends on the types of attributes that dominate the dataset. A per-visual-attribute analysis and attribute normalization in final ranking is thus beneficial to remove this bias. Most of the tested trackers outperform standard baselines and perform favorably to common state-of-the-art such as Struck, making the benchmark quite challenging.

The per-attribute analysis of the new dataset showed that the visual attributes that are most challenging to trackers are occlusion, motion change and size change. Sequence-wise analysis showed that some sequences are challenging on average, other sequences are very challenging at particular frames, and some of them are well tackled by all the trackers. An interesting find is that one particular sequence (David), which is usually assumed challenging in the tracking community, seems not to be according to the presented analysis, as trackers rarely fail on this sequence.

Establishing standard datasets and evaluation methodology tends to result in significant short-term advances in the field, but it can also have negative effects, leading to empoverished specter of approaches that get put forward in the long run . Evaluation is often reduced to a single performance score, which might lead to degradation in research. The primary goal of the authors, i.e., coming up with new tracking concepts, shifts to increasing a single performance score, and this is further enforced by pre-occupied reviewers that may find appealing to base their decision on this single score as well. We would like to explicitly warn against this. In practical experiments we are in fact comparing performance of various implementations rather than concepts. Implementations sometimes contain tweaks that improve performance, while often being left out from the original papers in interest of purity of the theory.

We also point out that the notion of a ”best” tracker varies with the tracker application. For example, sports analytics applications, which sports scientists use for player accelerations and velocity analysis, crucially depend on the quality of the estimated player position and do not require autonomous real-time performance. Thus user intervention for tracker reinitialization is allowed at any point. In such applications a highly accurate tracker is required, but robustness is only desired, i.e., an accurate non-robust tracker would be preferred over a robust but inaccurate tracker. But other applications in which tracking autonomy is critical, a robust tracker would be preferred over an accurate but non-robust tracker. The presented methodology allows identifying these characteristics and their variation w.r.t. the visual attributes which goes beyond the related methodologies.

We believe that it is difficult to overfit a tracker to a visually diverse dataset, but tuning parameters may very likely contribute to higher ranks. Related works like suggest splitting the dataset into training and testing sequences, making all sequences available, but only providing the annotations for the training sequences. The evaluation is then performed by running the tracker on the test set and uploading the results to an online service that checks the results against the unpublished ground truth. One problem with such an approach is that re-initialization at failure becomes impossible, since the test-data ground truth is censored, thus reducing the strength of the performance measures. But a conceptual problem lies in the assumption that the ”unpublished” ground truth cannot be re-produced. In fact, if the annotation rules are followed faithfully, the researchers can easily annotate the ground truth in the censored part of the dataset and this annotation will be equally valid as the unpublished. So if overfitting would be possible, censoring the ground truth would introduce even a larger bias in the results in favor of researchers that simply spend time re-annotating the test dataset.

Because of the unavoidable dependence on implementation and efforts spent in adjusting the tracker parameters, care has to be taken when deciding for or against a new tracker based on performance scores. One approach might be to apply a comparative evaluation to position a new tracking approach against a set of standard baseline implementations using a single ranking experiment, use detailed analysis with respect to different visual attributes and put further focus on the theory.

Our future work will focus on revising and carefully enriching the dataset, continually improving the tracker evaluation methodology and, through further organization of the VOT challenges, pushing towards a standardised tracker comparison.

Appendix A Derivation of NOR and WIR statistics

Plugging these results into (13,14) yields equations (5) and (6) in the paper.

In the WIR scenario, the tracker is reset after failure and Δ\Delta frames after the reset are ignored in computation of the accuracy. It is easy to show the following equivalence

Plugging these into (13,14) yields equations (7) and (8) in the paper.

Acknowledgments

This work was supported in part by the following research programs and projects: Slovenian research agency research programs and projects P2-0095, P2-0214, J2-4284, J2-3607, the EU project EPiCS (grant agreement no 257906), the CTU Project SGS15/155/OHK3/2T/13 and by The Czech Science Foundation Project GACR P103/12/G084.

References