FAMNet: Joint Learning of Feature, Affinity and Multi-dimensional Assignment for Online Multiple Object Tracking
Peng Chu, Haibin Ling
Introduction
Tracking multiple objects in video is critical for many applications, ranging from vision-based surveillance to autonomous driving. A current popular framework to solve multiple object tracking (MOT) uses the tracking-by-detection strategy where target candidates generated from an external detector are associated and connected to form the target trajectories across frames . At the core of tracking-by-detection strategy lies the data association problem which is usually treated as three separate parts: feature extraction for candidate representation, affinity metric to evaluate the cost of each association hypothesis and association algorithm to find the optimal association. These parts involve multiple individual data-processing steps and are optimized differently from each other, which results in a complex method design and extensive tuning parameters to adapt different target categories and tracking scenarios.
Recently, deep neural network (DNN) has been investigated intensively to learn the association cost function in a unified architecture combining both feature extraction and affinity metric . Through training, the task and scenario prior can be automatically adapted by the candidate representation and estimation metric without manually tuning the hyper-parameters. However, the association algorithm still stands outside the network, which requires dedicated affinity samples to be manually fabricated from ground truth association for the training process. It is not guaranteed that training and inference phases share the same data distribution; consequently it may lead to the degraded generalizability of the trained model. Moreover, crowded targets, similar appearance and fast motion impose great ambiguity for the association only considering pairs of neighboring frames. Successful association requires global optimization across multiple frames, where higher-order discriminative clues such as appearance changes over time and motion context could be included. Learning the robust representation and affinity criteria without the cooperation from the association procedure in this circumstance is even more complicated.
Our objective in this paper is to formulate an end-to-end model for MOT: the Feature representation, Affinity model and Multi-dimensional assignment (MDA) are refined in a single deep network named FAMNet, which is optimized jointly to learn the task prior. In particular, feature sub-network is used to extract features for candidates on each frame, after which an affinity sub-network estimates the higher-order affinity for all association hypothesis. With the affinity, the MDA sub-network is to optimize globally and obtain the optimal assignments. By all layers in FAMNet designed differentiable, the feature and affinity sub-network can be trained directly referring to the assignment ground truth. To realize it, we make the following novelties to the FAMNet and its based tracking system:
We design an affinity sub-network that fuses discriminative higher-order appearance and motion information into the affinity estimation.
We propose an MDA sub-network, in which a modified rank-1 tensor approximation power iteration is designed differentiable and adapted for the deep learning architecture.
We integrate single object tracking into the data association-based MOT. Detections and tracking predictions are merged and selected optimally through MDA to construct the target trajectories.
We employ a target management scheme where a dedicated CNN network is used to refine the target bounding box to eliminate the noised candidates generated by external detector.
To show the effectiveness of the proposed approach, it is evaluated on the popular multiple pedestrian and vehicle tracking challenge benchmarks including MOT2015, MOT2017, KITTI-Car and UA-DETRAC. Our results show promising performance in comparison with other published works.
Related Work
Recently, deep learning is explored with increasing popularity in MOT with great success. Most recent solutions rely on it as a powerful discriminative technique . Tang et al. propose to use DNN based Re-ID techniques for affinity estimations. They include lift edges that connect two candidates spanning multiple frames to capture the long-term affinity. In , recurrent neural networks (RNN) and long short-term memory (LSTM) is adapted to model the higher-order discriminative clue. Those methods learn the networks in a separate process with the manually fabricated affinity training samples.
Some recent works have gone further to tentatively solve MOT in an entirely end-to-end fashion. Ondruska and Posner introduce the RNN for the task to estimate the candidate state. Although this work is demonstrated on the synthetic sensor data and no explicit data association is applied, it firstly shows the efficacy of using RNN for an end-to-end solution. Milan et al. propose an RNN-LSTM based online framework to integrate both motion affinity estimation and bipartite association into the deep learning network. They use LSTM to solve the data association target by target at each frame where the constrains in data association are not explicitly built into the network but learned from training data. For both works, only the occupancy status of targets are considered, the informative appearance clue is not utilized. Different from their methods, we propose an MDA sub-network which handles both the data association and the constrains, and our affinity fuses both the appearance and motion clue for better discriminability.
Overview
In this section, we first formulate the multiple object tracking (MOT) problem as a multi-dimensional assignment (MDA) form, and then provide an overview of our FAMNet-based tracking system (overview in Fig. 1).
2 Architecture Overview and Tracking Pipeline
For each association batch, the FAMNet based tracking system takes the image frames and corresponding detections provided by an external detector as input. Detection candidates are first used to generate the hypothesis trajectories. Image patches of target candidates together with the trajectory hypothesis are passed into FAMNet to compute the set of local assignments as shown in Fig. 1. Inside FAMNet, features of candidate patches are extracted through a feature sub-network. The affinity sub-network then calculates the affinity for all hypothesis trajectories on those features to form the affinity tensor as described in Sec. 4.1. With the affinity tensor, the optimal multi-dimension assignments are estimated by the MDA sub-network as explained in Sec. 4.2 and 4.3.
During training, the assignment ground truth is directly compared with the network output to compute the loss. The loss signal then back-propagates throughout the network to the feature and affinity sub-networks for learning, which is illustrated as the red paths in Fig. 1 and detailed in Sec. 4.4. In tracking phase, the output assignments together with single object tracking (SOT) predictions are used to update the trajectories of tracked targets through the target management scheme as described in Sec. 4.5 and 4.6.
We design our method in the online tracking framework intended for more casual applications. Under the constant velocity assumption, three frames are the minimum temporal span to calculate the motion affinity. Therefore, in the rest of paper, with two frames overlapping between association batches is used as to balance the computation cost and the sufficient depth of association to include the higher-order discriminative clues.
FAMNet
The affinity sub-network takes the features of candidates and hypothesis trajectories as input, and generates the affinity tensor as output.
Two levels of affinities are calculated for each hypothesis trajectory using the extracted feature set, as shown in Fig. 2. In detail, the affinity tensor is calculated as following:
For pair-wise affinity, the cross-correlation operation is used, such as
2 R1TA Power Iteration Layer
With the affinity tensor, we use R1TA power iteration to estimate the set of optimal assignments in Eq. 3. Solving the global optimum for MDA usually requires NP-hard probing. A sub-optimal approximation is usually guaranteed by a power iteration algorithm which can be expressed in the pure mathematic format.
where is a unit vector of the same dimension as and has elements equal to 1 only at and otherwise 0. In order to calculate Eq. 7 for all iterations, the loss gradients of assignment vectors at each iteration are also needed, which follows:
where can be calculated as, e.g. for , .
4 Training
During the training, the total loss is measured by the binary cross entropy between all predicted assignments \big{(}x_{i_{k-1}i_{k}}^{(k)}\in\big{)} and assignment ground truth \big{(}\bar{x}_{i_{k-1}i_{k}}^{(k)}\in\{0,1\}\big{)} , which is written as
With the total loss, gradients are calculated throughout the network back to the affinity and feature sub-networks.
5 Tracking by Integrating Detection and SOT
In the tracking phase, predictions using SOT techniques are included to recover missing candidates from the external detector. We add a virtual candidate to each candidate set to represent missing candidates and allow it to connect with any candidate in consecutive frames as shown in Fig. 3. Both real and virtual candidates are used to generate trajectory hypothesis. When calculating affinity, we choose the location maximizing the affinity in Eq. 5 as the center of the virtual candidates for each anchor candidate such that
Therefore, if an anchor candidate misses its detection in consecutive frame, it will connect with the virtual candidate which represents the location most similar to it in that consecutive frame, or, in terms of SOT, its tracking prediction. Each anchor candidate may have a different location predicted by SOT. We use in Eq. 10 to refer to the virtual candidate, its center coordinates may vary on different anchor candidates.
6 Target Management
On receiving the assignment results, target management handles target entering, exiting and updating. In assignment results, if multiple anchor candidates choose to associate with a virtual candidate, new candidates will be added into candidate sets accordingly. For a virtual candidate not associated with any anchor candidate in assignment results, it will be dropped from the candidate set. Furthermore, if the virtual candidate associated with an anchor candidate in this batch appears as an anchor candidate in the next batch, the appearance feature of the anchor candidate is reused in the next batch in case that the missing detection is caused by occlusion. This SOT process will continue until a confident real candidate is associated.
Experiment
We conduct experiments on four popular MOT datasets: MOT2015 and MOT2017 for pedestrian tracking, KITTI-Car and UA-DETRAC for vehicle tracking. All datasets are provided with referred detections from real detectors.
To evaluate the performance of the proposed method, the widely accepted CLEAR MOT metrics are reported, which include multiple object tracking precision (MOTP) and multiple object tracking accuracy (MOTA) that combines false positives (FP), false negatives (FN) and the identity switches (IDS). Additionally, we also report the percentage of mostly tracked targets (MT), the percentage of mostly lost targets (ML).
2 Evaluation Results
MOT2015. MOT2015 contains 11 different indoor and outdoor scenes of public places with pedestrians as the objects of interest, where camera motion, camera angle and imaging condition vary greatly. The dataset provides detections generated by the ACF-based detector . The numerical results on its test set are reported in Tab. 1. Our approach achieves clearly the state-of-the-art performance. In particular, our method achieves better performance in most metrics than the RNN based end-to-end online methods due to our discriminative higher-order affinity and the optimization method adapted. Our method also surpasses the same R1TA-based method which is with hand-crafted features and affinity metrics.
MOT2017. Similar to MOT2015, MOT2017 contains seven different sequences in both training and test datasets but with higher average target density (31.8 vs 10.6 on the test set), thus is more challenging. MOT2017 also focuses on evaluating the tracker performance on different detection quality. It provides three different detection inputs from DPM , Faster-RCNN and SDP , ranked in ascending order by AP. We train seven different sets of networks according to different scenes, without further fitting on the different detections. The numerical results are reported in Tab. 2. The performance of our method is better than or on par with other published state-of-the-art methods.
KITTI-Car. The KITTI dataset contains 21 video sequences in the training set and 29 in the test set for multiple vehicle tracking in street view, where videos are recorded through a camera mounting in front of a moving vehicle. Referred detections from the regionlet detector are used in our experiment. The numerical results on the dataset of our method along with other methods using the same detections are summarized in Tab. 3. Our method again surpasses the hand-crafted feature-based R1TA method, despite the fact that it uses a much larger association batch for off-line tracking. It is worth mentioning that motion affinity plays a more importance role in KITTI than in MOT2015 and MOT2017, since both targets and camera move faster and more regularly in KITTI.
UA-DETRAC. UA-DETRAC dataset is another multiple vehicle tracking dataset with 60 sequences for training and 40 sequences for testing. All sequences are recorded with static camera at a lift-up position near different drive ways in various of weather conditions. We use referred detection from CompACT detector in our experiment. UA-DETRAC reports the average of each MOT metric from a serials of results using different detection confidence thresholds (from 0 to 1.0 with 0.1 step). Comparison with other methods using the same detections are reported in Tab. 4. Proposed method achieves state-of-the-art performance among the published works. Our method also surpasses the IOU tracker which is an offline method and using a private detector.
3 Ablation Study
Conclusion
In this paper we proposed a novel deep architecture for MOT, which learns jointly, in an end-to-end fashion, features and high-order affinity directly from the ground truth trajectories. During tracking, predictions from SOT and a dedicated target management are include to further boost tracking robustness. Experiments on the MOT2015, MOT2017, KITTI-Car and UA-DETRAC datasets clearly show the effectiveness of proposed approach.