Tracking Every Thing in the Wild

Siyuan Li, Martin Danelljan, Henghui Ding, Thomas E. Huang, Fisher Yu

A BDD100K Per-class Evaluation Results

We provide per-class evaluation results using CLEARMOT [MOTA] and TETA metrics on the BDD100K [bdd100k] validation set in Table 1. Data distribution in BDD100k is long-tailed. The Car category consists of most of the tracks in the dataset. The rest of the categories are rare compared to the dominant ones. Thus, we characterized them as rare classes. TETer can achieve significant improvements across all rare classes on both established MOTA, IDF1, and our TETA metrics compared to the previous state-of-the-art QDtrack. In particular, TETer boosts MOTA of buses by over 7 points on the validation set and TETA by over 6 points. We also compare our TETer results with CEM with its class agnostic counterpart, where the model only uses the AET strategy without CEM. The result shows that our model gains significant improvements over rare classes where the class agnostic instance association cannot be well trained due to lacking annotations. For instance, we gain +3.8 MOTA on buses and +4.7 MOTA on riders. Further, we can observe improvements on the TETA score, where we gain +2 on train and +1.5 on motorcycle. This demonstrates that TETer can better handle tracking rare classes. With CEM, we exploit the semantic annotations offered by large-scale object detection datasets. It can integrate fine-grained cues required for classification (e.g. the difference between a big red bus and a red truck), which are difficult to learn effectively with class-agnostic appearance training on the long-tailed datasets.

B TAO Per Frequency Group Evaluation Results

We provide evaluation results per frequency group on the TAO validation set using TETA in Table 3. We first observe that TETA can effectively evaluate methods across different frequency groups, despite difficulties introduced by classification errors. Although ClsA drops significantly for both QDTrack [qdtrack] and TETer as categories become more rare, LocA and AssocA are relatively stable. This enables us to compare different methods even in large-scale, long-tailed settings where classification is the bottleneck.

Compared to QDTrack, TETer can obtain consistent improvements in TETA, LocA, and AssocA across all frequency groups, at the cost of a small degradation in ClsA. The improvements are more prominent on common and rare categories, where TETer can achieve over 3 points improvement in TETA. For rare categories, TETer achieves 1.8 points improvement in LocA and over 7 points in AssocA. Even on frequency categories, TETer can still improve AssocA by over 6 points. Table 3 also shows that the major differences between frequent and rare categories lies in classification. The localization and association capabilities of both trackers already generalize very well on rare categories.

C Exemplar-based Classification

Given an example object, exemplar-based classification means classifying objects by comparing with the given example to determine whether they belong to the same class. Given two neighboring frames t1t_{1} and t2t_{2} in a video sequence, all objects in t1t_{1} will be treated as exemplars. For each exemplar, we find all target objects in t2t_{2} that belong to the same class as the exemplar.

In this experiment, we compare our Class Exemplar Matching (CEM) with a hard prior baseline that matches objects with the same predicted class label. We evaluate both methods on the TAO validation set and compute precision-recall (PR) curves for comparison. A true positive (TP) match is a match between two objects that belong to the same category. A false positive (FP) match is a match between two objects that belong to different categories. A false negative (FN) is a non-match between two objects that belong to the same category. To compute the PR curve, we sample 10 thresholds from 0 to 0.99 with a fixed step size.

Figure 1 shows the results of the experiment. The hard prior baseline takes the argmax of the predictions of a softmax classifier from Faster R-CNN, thus there is only a single value in the PR curve. CEM significantly outperforms the hard prior baseline.

D TETA Details

We provide additional details regarding our TETA metric about how it disentangles classification and how it deals with evaluation on datasets with complete annotations.

The most direct way to disentangle classification is not to consider per-class performance and evaluate every object class-agnostically. However, on large-scale, long-tailed datasets, such evaluation will be dominated by objects of the few common categories, and the overall performance will not reflect the improvements on rare classes. On the other hand, per-class evaluation requires us to select prediction results for each class, which is sensitive to classification performance. If the classification is wrong, the contribution in localization and association will be ignored. TETA can naturally deal with this issue with the local cluster evaluation since we select predictions based on their location rather than class. To evaluate a particular class, we access predictions in the local clusters of ground-truth objects belonging to the chosen class. Thus, we can evaluate the localization and association performance even when the class predictions are wrong.

D.2 TETA with Complete Annotations

Multiple categories TETA can also work with complete annotations. First, the localization accuracy is not affected. In the case of incomplete annotations, we treat every unmatched predictions in each cluster as false positives. If we have exhaustive annotations, we still treat those unmatched predictions as false positives. The remaining question is how to penalize predictions that are not in any clusters. For such predictions, we know that they are not highly overlapped with any ground truth box, since we have exhaustive annotations. The predictions thus false classify background as one of the foreground classes, and so we treat them as classification false positives.

Single category For single category with exhaustive annotations, the classification term of TETA is meaningless and can be ignored. Also, since we do not need to perform per-class evaluation, the margin of the local cluster does not matter either. Thus, we can set the margin rr to 0. With these changes, TETA becomes similar to the HOTA [hota] metric with the only difference being that we use arithmetic mean instead of geometric mean.

D.3 Ablation Study of TETA

We provide an ablation study of the local cluster IoU margin rr of TETA. We perform this experiment on incomplete dataset TAO.

The results are shown in Table 2. The LocRe and LocPr represent the localization recall and precision. As we can see, with a larger rr, the LocPr increases since TETA becomes more conservative regarding identifying FPs. In the mean time, TETA makes fewer mistakes where the objects with no annotations are wrongly identified as FPs. In extreme crowded scenarios with incomplete annotations, it’s recommended to set a higher rr to avoid false punishment.

E Qualitative Results

We provide additional qualitative results of TETer.

We first compare the QDTrack [qdtrack] which uses class prediction as hard prior to associate objects with TETer which uses CEM. In Figure 3, we show an example of QDTrack producing ID Switches due to errors in classification, whereas TETer is more robust to such issues.

E.2 Class-agnostic vs. CEM

We further show the comparison between the class-agnostic association (AET baseline in Section 5.5) and association with CEM. We observe that most class agnostic association errors happen in rare classes where there are not enough videos to train the class-agnostic instance association module well. For instance, Figure 5 (a) shows the bicycle (16) is wrongly associated with the car with class-agnostic association, while using the CEM module helps to avoid the mistake. The CEM module utilizes the supervision from large-scale object detection datasets to learn fine-grained class appearance differences, which helps the association on rare classes.

E.3 Tracking results comparison: TAO metric vs. TETA

We provide results for cross-dataset analysis. In Figure 2, we show predictions from trackers that are optimized either for the TAO metric or TETA. The tracker optimized for the TAO metric generates more false positives that highly overlap, producing results that are difficult to use in practice. On the other hand, the tracker optimized for TETA produces cleaner results.

E.4 Rare class retrieval

We perform the class retrieval experiments in the rare classes on TAO to show the effectiveness of the CEM embeddings. We take the objects in the first frame of each ground truth track and use them as the retrieval templates to retrieve ground truth objects in the whole TAO validation set. The softmax prediction means we use the softmax confidence to retrieve objects that are predicted as the same class as the template. The CEM means we use the CEM embedding similarity to perform the retrieval. Figure 6 shows the CEM embedding can successfully retrieve the examples in the rare classes, while its softmax fails. Figure 7 shows some failures cases where the CEM module retrieves the wrong class due to occlusion or high visual similarities.

E.5 TETer failure cases

We also show some common failure cases of TETer on TAO in Figure 4. Note that TAO is annotated at 1 FPS. Thus, fast-moving objects usually have huge appearance changes in neighboring frames. Due to the large appearance and location variations, tracking is challenging on TAO. Also, TETer suffers from localization errors caused by occlusion.

F More Implementation Details

We provide more implementation and training details of our method and evaluation setup in different benchmarks.

We use the popular Faster R-CNN [ren2015faster] with ResNet as the backbone. Specifically, we use the ResNet-101 [he2016deep] for TETer on TAO and ResNet-50 on BDD100K. For the exemplar encoder, we use 4conv-3fc head with group normalization [wu2018group]. The final output channel numbers are 1230 for TETer on TAO [tao] and 256 for BDD100K [bdd100k]. We use the same network architecture for the instance appearance encoder but with only 1fc layers for the final output. The channel number of the instance appearance encoder is 256 by default on both datasets.

F.2 Training

TAO We train TETer following the TAO [dave2020tao] set up with a mixed LVISv0.5 [lvis] and COCO [coco] dataset. We set the batch size to 16 and the learning rate to 0.02. We train 24 epochs in total and decrease the learning by 0.1 after 16 and 22 epochs. For data augmentation, we randomly flip the images horizontally with a 0.5 ratio. We randomly resize the training images to keep their short edges between 640 to 800. We randomly sample images to form mini-batches with additionally repeat sampling for rare classes [lvis]. We set the repeat factor to 0.001. We train the instance appearance encoder on the TAO training set following the same setting as QDTrack [qdtrack].

BDD100K We use the same object detector as QDTrack [qdtrack]. For training the exemplar encoder, we freeze the object detector above and train with 8 BDD100K MOT categories using the BDD100K Detection set, which contains 70K images. For data augmentation, we randomly flip the images horizontally with a 0.5 ratio. We randomly resize the training images to keep their short edges between 640 to 800. We randomly sampled images to form mini-batches. We set the batch size to 128.

F.3 Inference and evaluation

TAO We evaluate our model on the TAO validation set with TETA. For the close-set setting, the TAO validation set contains 988 videos with 302 classes, a subset with LVIS classes. For the open-set setting, we merge the additional free-form classes [tao] as one unknown class. During inference, we use the fixed image scale with 800 at the short edge. We initialize a new track if the object has detection confidence higher than 0.0001.

BDD100K The BDD100K contains 200 videos (40k) for validation and 400 videos (80k) for testing. We use both the BDD100K validation set and the test set for evaluation. For the inference, we use the (1296,720)(1296,720) image scale. We use the best performed model in the validation set which is saved at 2 epoch. We initialize a new track if the detection confidence is higher than 0.7.