Chained-Tracker: Chaining Paired Attentive Regression Results for End-to-End Joint Multiple-Object Detection and Tracking

Jinlong Peng, Changan Wang, Fangbin Wan, Yang Wu, Yabiao Wang, Ying Tai, Chengjie Wang, Jilin Li, Feiyue Huang, Yanwei Fu

Introduction

Video-based scene understanding and human behavior analysis are important high-level tasks in computer vision with many valuable applications in real scene. They rely on many other tasks, within which Multiple-Object Tracking (MOT) is a significant one. However, MOT remains challenging due to the existence of occlusions, object trajectory overlap, possibly challenging background, etc., especially for crowded scenes.

Despite the great efforts and encouraging progress in the past years, there are two major problems of existing MOT solutions. One is that most methods are based on the tracking-by-detection paradigm [tracking-by-detection_ICCV09], which is plausible but suboptimal due to the infeasibility of global (end-to-end) optimization. It usually contains three sequential subtasks: object detection, feature extraction and data association. However, splitting the whole task into isolated subtasks may lead to local optima and more computation cost than end-to-end solutions. Moreover, data association heavily relies on the quality of object detection, which by itself is hard to generate reliable and stable results across frames as it discards the temporal relationships of adjacent frames.

The other problem is that recent MOT methods get more and more complex as they try to gain better performances. Re-identification and attention are two major points found to be helpful for improving the performance of MOT. Re-identification (or ID verification) is used to extract more robust features for data association. Attention helps the model to be more focused, avoiding the distraction by irrelevant yet confusing information (e.g. the complex background). Despite their effectiveness, the involvement of them in existing solutions greatly increases the model complexity and computational cost.

In order to solve the above problems, we propose a novel online tracking method named Chained-Tracker (CTracker), which unifies object detection, feature extraction and data association into a single end-to-end model. As can be seen in Fig. 1, our novel CTracker model is cleaner and simpler than the classical tracking-by-detection or partially end-to-end MOT methods. It takes adjacent frame pairs as input to perform joint detection and tracking in a single regression model that simultaneously regress the paired bounding boxes for the targets that appear in both of the two adjacent frames.

Furthermore, we introduce a joint attention module using predicted confidence maps to further improve the performance of our CTracker. It guides the paired boxes regression branch to focus on informative spatial regions with two other branches. One is the object classification branch, which predicts the confidence scores for the first box in the detected box pairs, and such scores are used to guide the regression branch to focus on the foreground regions. The other one is the ID verification branch whose prediction facilitates the regression branch to focus on regions corresponding to the same target. Finally, the bounding box pairs are filtered according to the classification confidence. Then, the generated box pairs belonging to the adjacent frame pairs could be associated using simple methods like IoU (Intersection over Union) matching [bochinski2017high] according to their boxes in the common frame. In this way, the tracking process could be achieved by chaining all the adjacent frame pairs (i.e. chain nodes) sequentially.

Benefiting from the end-to-end optimization of joint detection and tracking network, our model shows significant superiority over strong competitors while remaining simple. With the temporal information of the combined features from adjacent frames, the detector becomes more robust, which in turn makes data association easier, and finally results in better tracking performance.

The contribution of this paper can be summarized into the following aspects:

1. We propose an end-to-end online Multiple-Object Tracking model, to optimize object detection, feature extraction and data association simultaneously. Our proposed CTracker is the first method that converts the challenging data association problem to a pair-wise object detection problem.

2. We design a joint attention module to highlight informative regions for box pair regression and the performance of our CTracker is further improved.

3. Our online CTracker achieves state-of-the-art performance on the tracking result list with private detection of MOT16 and MOT17.

Related Work

Yu et. al [yu2016poi] proposed the POI algorithm, which conducted a high-performance detector based on Faster R-CNN [ren2015faster] by adding several extra pedestrian detection datasets. Chen et. al [chen2017enhancing] incorporated an enhanced detection model by simultaneously modeling the detection-scene relation and detection-detection relation, called EDMT. Furthermore, Henschel et. al [Henschel2017] added a head detection model to support MOT in addition to original pedestrian detection, which also needed extra training data and annotations. Bergmann et. al [bergmann2019tracking] proposed the Tracktor by exploiting the bounding box regression to predict the position of the pedestrian in the next frame, which was equal to modifying the detection box. However, the detection model and the tracking model in these detection-based methods are completely independent, which is complex and time-consuming. While our CTracker algorithm only needs one integrated model to perform detection and tracking, which is simple and efficient.

2 Partially End-to-end MOT Methods

Lu et. al [lu2020retinatrack] proposed RetinaTrack, which combined detection and feature extraction in the network and used greedy bipartite matching for data association. Sun et. al [sun2019deep] harnessed the power of deep learning for data association in tracking by jointly modeling object appearances and their affinities between different frames. Similarly, Chu et. al [chu2019famnet] designed the FAMNet to jointly optimize the feature extraction, affinity estimation and multi-dimensional assignment. Li et. al [li2019tracknet] proposed TrackNet by using frame tubes as input to do joint detection and tracking, however the links among tubes are not modeled which limits the trajectory lengths. Moreover, the model is designed and tested only for rigid object (vehicle) tracking, leaving its generalization ability questionable. Despite their differences, all these methods are just partially end-to-end MOT methods, because they just integrated some parts of the whole model, i.e. [lu2020retinatrack] combined the detection and feature extraction module in a network, [sun2019deep, chu2019famnet] combined the feature extraction and data association module. Differently, our CTracker is a totally end-to-end joint detection and tracking methods, unifying the object detection, feature extraction and data association in a single model.

3 Attention-assistant MOT Methods

Chu et. al [chu2017online] introduced a Spatial-Temporal Attention Mechanism (STAM) to handle the tracking drift caused by the occlusion and interaction among targets. Similarly, Zhu et. al [zhu2018online] proposed a Dual Matching Attention Networks (DMAN) with both spatial and temporal attention mechanisms to perform the tracklet data association. Gao et. al [gao2018osmo] also utilized an attention-based appearance model to solve the inter-object occlusion. All these attention-assistant MOT methods used a complex attention model to optimize data association in the local bounding box level. While our CTracker can improve both the detection and tracking performance through the simple object-attention and identity-attention in the global image level, which is more efficient.

Methodology

2 Chained-Tracker Pipeline

Node chaining. We use {Dt−1,D^t}\{\mathcal{D}_{t-1},\mathcal{\hat{D}}_{t}\} to represent \{(D_{t-1}^{i},\hat{D}_{t}^{i})\}_{i=1}^{{\color[rgb]{0,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,0}\pgfsys@color@gray@stroke{0}\pgfsys@color@gray@fill{0}n}_{t-1}} for convenience. The node chaining is done as follows. Firstly, in the node, every detected bounding box D1i∈D1D_{1}^{i}\in\mathcal{D}_{1} is initialized as a tracklet with a randomly assigned identity. Secondly, for any another node tt, we chain the adjacent nodes (Ft−1,Ft)(F_{t-1},F_{t}) and (Ft,Ft+1)(F_{t},F_{t+1}) by calculating the IoU (Intersection over Union) between the boxes in D^t\mathcal{\hat{D}}_{t} and Dt\mathcal{D}_{t} as shown in Fig. 2, where D^t\mathcal{\hat{D}}_{t} is the last boxes set of {Dt−1,D^t}\{\mathcal{D}_{t-1},{\mathcal{\hat{D}}_{t}}\} and Dt\mathcal{D}_{t} is the former boxes set of {Dt,D^t+1}\{\mathcal{D}_{t},{\mathcal{\hat{D}}_{t+1}}\}. Getting the IoU affinity, the detected boxes in D^t\mathcal{\hat{D}}_{t} and Dt\mathcal{D}_{t} are matched by applying the Kuhn-Munkres (KM) algorithm [kuhn1955hungarian]. For each matched box pair D^ti{\hat{D}_{t}^{i}} and Dtj{D_{t}^{j}}, the tracklet that D^ti{\hat{D}_{t}^{i}} belongs to is updated by appending Dtj{D_{t}^{j}}. Any unmatched box Dtk{D_{t}^{k}} is initialized as a new tracklet with a new identity. The chaining is done sequentially over all adjacent nodes and it builds long trajectories for individual targets.

Robustness enhancement (esp. against occlusions). To enhance the model’s robustness to serious occlusions (which can make detection fail in certain frames) and short-term disappearing (followed by quick reappearing), we retain the terminated tracklets and their identities for up to σ\sigma frames and continue finding matches for them in these frames, with the simple constant velocity prediction model [wojke2017simple, peng2020tpm] for motion estimation. In greater details, suppose target (Dt−1l,D^tl)(D_{t-1}^{l},\hat{D}_{t}^{l}) cannot find its match is node tt, we apply the constant velocity model to predict its bounding box Pt+τlP_{t+\tau}^{l} in frame t+τt+\tau (1<=τ<=σ1<=\tau<=\sigma) according to Dt−1lD_{t-1}^{l} (not the less reliable D^tl\hat{D}_{t}^{l}). When we chain node t+τ−1t+\tau-1 and node t+τt+\tau with {Dt+τ−1,D^t+τ}\{\mathcal{D}_{t+\tau-1},{\mathcal{\hat{D}}_{t+\tau}}\} and {Dt+τ,D^t+τ+1}\{\mathcal{D}_{t+\tau},{\mathcal{\hat{D}}_{t+\tau+1}}\}, the current set of all the predicted bounding boxes of retained targets denoted by Pt+τ\mathcal{P}_{t+\tau}, is appended to D^t+τ\mathcal{\hat{D}}_{t+\tau} for matching with Dt+τ\mathcal{D}_{t+\tau}. If Pt+τiP_{t+\tau}^{i} gets a match, its tracklet will be extended by linking to the new bounding boxes.

Effectiveness and limitations. Our model is effective for handling the cases when targets appear or disappear (i.e., enter or leave camera view), which are quite common for MOT. When a target is not in frame t−1t-1 but appears in frame tt, it is likely that no bounding box pair for it gets generated in the chain node (Ft−1,Ft)(F_{t-1},F_{t}). However, as long as this target continues to appear in frame t+1t+1, it will be detected in the next chain node (Ft,Ft+1)(F_{t},F_{t+1}) and get a new tracklet and identity there. Similarly, if a target is in the frame t−1t-1 but disappears from frame tt, it will not be detected in node (Ft,Ft+1)(F_{t},F_{t+1}), resulting the termination of its tracklet in node t−1t-1 or even t−2t-2. Note that the chaining operation itself cannot be fully parameterized and therefore it cannot be optimized together with the regressions. Since the regression model (as detailed below) does the major work and there is no need to get feedback for it from the chaining operation, we still use the “end-to-end” property to describe CTracker. A pure end-to-end trainable model requires a differentiable replacement to the current IoU matching based chaining strategy.

3 Network architecture

Overview. Our proposed CTracker network uses two adjacent frames as input and regresses the bounding box pair of the same target. To do this, we adopt ResNet-50 [he2016deep] as the backbone to extract high-level semantic features. It then integrates Feature Pyramid Networks (FPN) to generate multi-scale feature representation for subsequent prediction. In order to associate targets in adjacent frames, the scale-level feature maps from individual frames are firstly concatenated together, and then fed into the prediction network to regress bounding box pairs. As can be seen in Fig. 3, the paired boxes regression branch generates a box pair for each target, and the object classification branch predicts a score for each pair indicating the confidence of being foreground. To help the paired boxes regression branch to avoid the distraction by irrelevant yet confusing information, the object classification branch and the extra ID verification branch are used for attention guidance.

Paired Boxes Regression. Inspired by predicting the offsets relative to pre-defined (default) anchor boxes in object detection, we propose Chained-Anchors for the paired boxes regression branch to regress two boxes simultaneously. As a novel natural derivative of the anchors used in most object detection methods, Chained-Anchors are densely arranged on a spatial grid, each of them allows predicting two bounding boxes of the same object instance in two adjacent frames. In order to handle the large scale variation in real scenes, the K-means clustering as used in [yolov2] is conducted on all ground-truth bounding boxes in the dataset for getting the scales of chained-anchors. And each cluster is assigned to the corresponding level of FPN for later scale specific predictions. The detected bounding box pairs are firstly post-processed with soft-NMS [softnms] according to the IoU of the first box in each pair, and then filtered based on the confidence scores from the classification branch. Finally, the remaining box pairs are chained into the whole tracking trajectories using the method described in Sec. 3.2. To keep our model simple, both the paired boxes regression branch and the classification branch only stack four consecutive 3×\times3 Conv layers interleaved with ReLU activations before the final convolution layer.

Joint Attention Module. We design an attention mechanism based component called Joint Attention Module (JAM) to highlight local informative regions in the combined features before the regression branch. As shown from the right of Fig. 3, the ID verification branch is introduced to get confidence scores, indicating whether the two boxes in the detected pair belong to the same target. Then both the predicted confidence map of ID verification branch and object classification branch are used as attention maps. Note that the guidance from the two branches is complementary, the confidence maps from the classification branch focuses on foreground regions while the prediction from the ID verification branch is used to highlight the features of the same target.

Feature Reuse. Since the input of the network contains two adjacent frames, the common frame of two adjacent nodes has to be used twice in the tracking process. To avoid the nearly double cost of computation and memory in inference, we propose a Memory Sharing Mechanism (MSM) to temporarily save the extracted features of the current frame and reuse them until the next node is processed, as shown in Fig. 4. Besides, in order to make inference for the last node, we make a copy of frame NN as the hypothetical frame N+1N+1. To further avoid the repeated computation for the frame N+1N+1, we also apply the trick of feature resue to frame NN, and the feature of frame NN is copied as the feature of the hypothetical frame N+1N+1. We demonstrate that the proposed MSM can reduce almost half of the overall computation and time cost.

4 Label Assignment and Loss Design

where KtK_{t} is the total number of ground-truth bounding boxes for frame FtF_{t}.

With AtiA_{t}^{i}, suppose the predicted pair of bounding boxes are (Dti,D^t+1i)(D_{t}^{i},\hat{D}_{t+1}^{i}) and the corresponding ground-truth bounding boxes are (Gtj,Gt+1k)(G_{t}^{j},G_{t+1}^{k}) when they exist, the ID verification branch of CTracker shall get its ground-truth label as:

where I[⋅]\mathcal{I}[\cdot] represents the identity of the target in the bounding box.

We follow Faster R-CNN [faster] to regress offsets of (Dti,D^t+1i)(D_{t}^{i},\hat{D}_{t+1}^{i}) w.r.t. AtiA_{t}^{i}, where Dti=(xdt,i,ydt,i,wdt,i,hdt,i)D_{t}^{i}=(x^{t,i}_{d},y^{t,i}_{d},w^{t,i}_{d},h^{t,i}_{d}). Let (Δdt,i,Δd^t+1,i)(\Delta^{t,i}_{d},\Delta^{t+1,i}_{\hat{d}}) denote these offsets and (Δgt,j,Δgt+1,k)(\Delta^{t,j}_{g},\Delta^{t+1,k}_{g}) be the offsets for the ground-truths, we list the details of Δdt,i=(Δd,xt,i,Δd,yt,i,Δd,wt,i,\Delta^{t,i}_{d}=(\Delta^{t,i}_{d,x},\Delta^{t,i}_{d,y},\Delta^{t,i}_{d,w}, Δd,ht,i)\Delta^{t,i}_{d,h}) as an example (the others are similar):

The loss for the paired boxes regression branch is defined as follows:

where F(pclsi,cclsi)\mathcal{F}(p^{i}_{cls},c^{i}_{cls}) and F(pidi,cidi)\mathcal{F}(p^{i}_{id},c^{i}_{id}) are the focal losses [lin2017focal] for the classification branch and the ID verification branch (for mitigating the sample imbalance problem), respectively, with pclsip^{i}_{cls} and pidip^{i}_{id} denoting their predictions (confidence scores); α\alpha and β\beta are the weighting factors.

Experiment

We conduct the experiments on two public datasets: MOT16 [milan2016mot16] and MOT17. which contain the same image sequences including 7 training sequences and 7 test sequences. However, MOT16 and MOT17 contain different detection input, and different ground-truth labels (bounding boxes and identities), which would influence the training of CTracker. In public detection, MOT16 includes DPM [felzenszwalb2010object] detector while MOT17 includes DPM, Faster R-CNN [ren2015faster] and SDP [yang2016exploit] detectors. For a fair comparison with other methods, we trained two models separately using the training data from MOT16 and MOT17, and separately applied the two models on the MOT16 test set and MOT17 test set.

In the MOTChallenge benchmark, tracking performance is measured by the widely used CLEAR MOT Metrics [bernardin2008evaluating], including Multiple-Object Tracking Accuracy (MOTA), Multiple-Object Tracking Precision (MOTP), the total number of False Negatives (FN), False Positives (FP), Identity Switches (IDS), and the percentage of Mostly Tracked Trajectories (MT), Mostly Lost Trajectories (ML). ID F1 Score (IDF1) is also used to measure the trajectory identity accuracy. Among these metrics, MOTA is the primary metric to measure the overall detection and tracking performance. In addition, we use Tracker Speed in Frames Per Seconds (Hz) to measure the tracking speed of all methods.

2 Implementation Details

All the experiments are implemented on the PyTorch framework. During training, the ground-truth boxes with a visible score above 0.1 are selected to train the network. In order to avoid overfitting, we use several data augmentation strategies such as photometric distortions, random flip and random crop. The same augmentation operation is guaranteed to apply for each image in the same training pair. Then the augmented image pair are resized or padded to the half of their original images’ shorter side. We also add a novel data augmentation strategy in the temporal dimension to form chain nodes: instead of always choosing two adjacent frames, we sample two frames close to each other with a random temporal gap (1 to 3 frames).

As a speed-accuracy trade-off, we use the Resnet50 [he2016deep] network as the backbone in all the following experiments. All trainable weights except the BN parameters in Resnet50 are trained end-to-end using the Adam optimizer. We initialize the parameters for all the newly added convolutional layers with the Kaiming initialization method in [he2015delving] and set the initial learning rate to 5×e−55\times e^{-5}. The model training process takes 100 epochs with the batch size of 8 (4 training pairs). The weighting factors α\alpha and β\beta in the loss function are both set to 1. In the anchor matching stage, we use 0.5 for the positive threshold and 0.4 for the negative threshold. For paired boxes post-processing, we use a threshold of 0.7 for the soft-nms, and then further filter remaining pairs with the confidence threshold of 0.4. In the chaining stage, the IoU matching threshold is 0.5, and the retention threshold of σ\sigma is 10.

3 Ablation Study

Performance analysis. We compare the following models on MOT17 dataset to show the effectiveness of CTracker’s parts:

(1) Baseline. It only covers the classification branch and the paired boxes regression branch, without guidance from any attention map. This is the simplest implementation of our CTracker.

(2) Baseline+ObjAtten. In addition to the Baseline, the predicted confidence map of the object classification branch is used as an attention map, which is multiplied to the combined features before the paired boxes regression branch.

(3) Baseline+ObjAtten+IDVer. Except for the object classification branch with attention map and the paired boxes regression branch, we add the ID verification branch but do not use it as attention guidance.

(4) Baseline+JointAtten (CTracker). This is the full version of our approach.

(1) Baseline+ObjAtten performs significantly better than Baseline, which proves the effectiveness of the object attention operation. By applying the object classification branch as the attention map of the paired boxes regression branch, we can get more accurate bounding boxes. There is a significant improvement of MOTA, which increases from 64.4 to 66.0 and MOTP also increases from 78.2 to 78.8. The more accurate bounding boxes also result in better performance of data association, with IDF1 increasing from 51.6 to 55.7.

(2) Baseline+ObjAtten+IDVer performs slightly worse than Baseline+ObjAtten. Simply adding the independent ID verification branch is weak due to the lack of bounding boxes information. Reliable identification needs good bounding boxes.

(3) Baseline+JointAtten further outperforms Baseline+ObjAtten, indicating that the ID attention operation is also beneficial. By adding the ID verification branch and using it as another guidance of the paired boxes regression branch, the association of the regressed bounding boxes is more accurate. Though MOTA is only improved by 0.6, the IDF1 is improved by 1.7, and IDF1 can better reflect the accuracy of data association more clearly. On the other hand, by adding the ID attention, the model pays more attention to the data association and sacrifices slightly of the regression bounding box precision, thus the MOTP is decreased from 78.8 to 78.2. Qualitative results of CTracker are illustrated in Fig. 5.

Time cost analysis. We analyze the inference speed for each module in CTracker, displayed in Table 2. The time cost is measured for 1080×\times1920 images using single Tesla P40 and cuDNN v7 with Intel Xeon E5-2699v4@2.20GHz. In Table 2, CTracker-Det only predicts boxes for a single frame, which is the initial detection network of CTracker. Since nearly 70% of the forward time is spent on the backbone network, our original CTracker costs about double-time to perform joint detection and tracking compared with the initial detection network, the time increasing from 119.05 ms to 223.56 ms. With the help of the proposed Memory Sharing Mechanism (MSM) in Sec. 3.3, we achieve a faster joint detection and tracking model with only 29.05 ms extra cost compared with the detection network. There is just a small increase of time from 119.05 ms to 148.10 ms. To some extent, 29.05 ms per frame means the tracking module runs at 34.4 FPS, demonstrating the efficiency of our online approach.

4 Benchmark Evaluation

We compare our CTracker approach with other MOT methods on both MOT16 and MOT17 test datasets. For comparison, we trained our model separately using the MOT16 training data and MOT17 training data. Table 3 and Table 4 compare the tracking results of all the methods separately on MOT16 and MOT17 test dataset. From Table 3 and Table 4 we can find that:

(1) In the private detection part of both MOT16 and MOT17, our CTracker significantly outperforms existing online MOT methods in terms of MOTA. In MOT16, the MOTA of our approach is only 0.6 lower than the best offline method KDNT [yu2016poi], while it is 1.5 higher than its online version POI [yu2016poi]. In addition, KDNT and POI use many extra training data, including ETHZ pedestrian dataset [ess2008mobile], Caltech pedestrian dataset [dollar2009pedestrian] and their own collected surveillance dataset [yu2016poi]. While we only use the training data of MOT16. MOTA is the primary metric reflecting the overall detection and tracking performance, which proves the effectiveness of our approach.

(2) In the public detection part, Tracktor [bergmann2019tracking] performs the best in terms of MOTA. To have a comparison with Tracktor using the same detection result, we reproduce Tracktor using its code. Tracktor+CTdet in Table 4 is the tracking result of Tracktor using the detection result of our CTracker. Compared with the results of public detection, the MOTA of Tracktor+CTdet increases from 53.5 to 54.4 and IDF1 increases from 52.3 to 56.1, which indicates that the performance of our detection is better than the public detection. Besides, our CTracker outperforms Tracktor+CTdet in terms of all the metrics except IDS, which further proves the superior tracking performance of our CTracker.

(3) On the other hand, to keep the simplicity and efficiency of our CTracker, we abandon using the patch-level ReID features of the detected boxes like other MOT methods to enhance cross-frame data association. Thus, the IDF1 and IDS of our CTracker approach are lower than several methods. We conduct an extra experiment by adding features, introduced in the supplementary. To further prove the efficiency of our approach, we compare the time cost of CTracker with other state-of-the-art MOT methods on the MOT16 and MOT17 benchmark, as shown in the Hz column of Tabel 3 and Tabel 4. From Tabel 3 and Tabel 4 we can find that CTracker achieves the best tracking speed among all online MOT methods, although the fastest offline method runs at a similar tracking speed as our CTracker, but has a much lower MOTA than our CTracker, demonstrating the effectiveness and efficiency of our approach.

Conclusion

We designed a novel joint multiple-object detection and tracking framework named Chained-Tracker in this paper, which is the first totally end-to-end solution as far as we are aware. Different from existing methods, we use two adjacent frames as the input of our network, which is called a chain node. The network regresses a pair of bounding boxes for the same target in the two adjacent frames, guided by a simple yet novel joint attention module: an interplay of detection-driven object attention and ID verification-injected identity attention. Using the simple IoU information, two adjacent and overlapping nodes can be chained by their boxes in the common frame. The tracking trajectories can be generated by alternately applying the paired boxes regression and node chaining. Extensive experiments on widely used MOT benchmarks demonstrate the superiority of our approach in terms of both effectiveness and efficiency.