Online Multi-Object Tracking Using CNN-based Single Object Tracker with Spatial-Temporal Attention Mechanism
Qi Chu, Wanli Ouyang, Hongsheng Li, Xiaogang Wang, Bin Liu, Nenghai Yu
Introduction
Tracking objects in videos is an important problem in computer vision which has attracted great attention. It has various applications such as video surveillance, human computer interface and autonomous driving. The goal of multi-object tracking (MOT) is to estimate the locations of multiple objects in the video and maintain their identities consistently in order to yield their individual trajectories. MOT is still a challenging problem, especially in crowded scenes with frequent occlusion, interaction among targets and so on.
On the other hand, significant improvement has been achieved on single object tracking problem, sometimes called “visual tracking” in previous work. Most state-of-the-art single object tracking methods aim to online learn a strong discriminative appearance model and use it to find the location of the target within a search area in next frame . Since deep convolutional neural networks (CNNs) are shown to be effective in many computer vision applications , many works have explored the usage of CNNs to learn strong discriminative appearance model in single object tracking and demonstrated state-of-the-art performance recently. An intuitive thought is that applying the CNN based single object tracker to MOT will make sense.
However, problems are observed when directly using single object tracking approach for MOT.
First, single object tracker may learn from noisy samples. In single object tracking, the training samples for learning appearance model are collected online, where labels are based on tracking results. The appearance model is then used for finding the target in the next frame. When the target is occluded, the visual cue is unreliable for learning the appearance model. Consequently, the single object tracker will gradually drift and eventually fail to track the target. This issue becomes even more severe in MOT due to more frequent occlusion caused by interaction among targets. An example is shown in Figure 1, one target is occluded by another when they are close to each other, which makes the visual cues of the occluded target contaminated when this target is used for training. However, the tracking score of the occluded target is still relatively high at the beginning of occlusion. In this case, the corresponding single object tracker updates the appearance model with the corrupted samples and gradually drifts to the occluder.
Second, since a new single object tracker needs to be added into MOT system once a new target appears, the computational cost of applying single object trackers to MOT may grow intolerably as the number of tracked objects increases, which limits the application of computationally intensive single object trackers in MOT such as deep learning based methods.
In this work, we focus on handling the problems observed above. To this end, we propose a dynamic CNN-based framework with spatial-temporal attention mechanism (STAM) for online MOT. In our framework, each object has its own individual tracker learned online.
The contributions of this paper are as follows:
First, an efficient CNN-based online MOT framework. It solves the problem in computational complexity when simply applying CNN based single object tracker for MOT by sharing computation among multiple objects.
Second, in order to deal with the drift caused by occlusion and interactions among targets, spatial-temporal attention of the target is learned online. In our design, the visibility map of the target is learned and used for inferring the spatial attention map. The spatial attention map is applied to weight the features. Besides, the visibility map also indicates occlusion status of the target which is an important cue that needs to be considered in online updating process. The more severe a target is occluded, the less likely it should be used for updating corresponding individual tracker. It can be considered as temporal attention mechanism. Both the spatial and temporal attention mechanism help to help the tracker to be more robust to drift.
We demonstrate the effectiveness of the proposed online MOT algorithm, referred as STAM, using challenging MOT15 and MOT16 benchmarks.
Related Work
Multi-object Tracking by Data Associtation. With the development of object detection methods , data association has become popular for MOT. The main idea is that a pre-defined object detector is applied to each frame, and then trajectories of objects are obtained by associating object detection results. Most of these works adopt an off-line way to process video sequences in which the future frames are also utilized to deal with the problem. These off-line methods consider MOT as a global optimization problem and focus on designing various optimization algorithm such as network flow , continuous energy minimization , max weight independent set , k-partite graph , subgraph multi-cut and so on. However, offline methods are not suitable for causal applications such as autonomous driving. On the contrary, online methods generate trajectories only using information up to the current frame which adopt probabilistic inference or deterministic optimization (e.g. Hungarian algorithm used in ). One problem of such association based tracking methods is the heavy dependency on the performance of the pre-defined object detector. This problem has more influence for online tracking methods, since they are more sensitive to noisy detections. Our work focuses on applying online single object tracking methods to MOT. The target is tracked by searching for the best matched location using online learned appearance model. This helps to alleviate the limitations from imperfect detections, especially for missing detections. It is complementary to data association methods, since the tracking results of single object trackers at current frame can be consider as association candidates for data association.
Single Object Tracker in MOT. Some previous works have attempted to adopt single object tracking methods into MOT problem. However, single object tracking methods are often used to tackle a small sub-problem due to challenges mentioned in Sec. 1. For example, single object trackers are only used to generate initial tracklets in . Yu et al. partitions the state space of the target into four subspaces and only utilizes single object trackers to track targets in tracked state. There also exists a few works that utilize single object trackers throughout the whole tracking process. Breitenstein et al. use target-specific classifiers to compute the similarity for data association in a particle filtering framework. Yan et al. keep both the tracking results of single object trackers and the object detections as association candidates and select the optimal candidate using an ensemble framework. All methods mentioned above do not make use of CNN based single object trackers, so they can not update features during tracking. Besides, they do not deal with tracking drift caused by occlusion. Different from these methods, our work adopts online learned CNN based single object trackers into online multi-object tracking and focuses on handling drift caused by occlusion and interactions among targets.
Occlusion handling in MOT. Occlusion is a well-known problem in MOT and many approaches are proposed for handling occlusion. Most works aim at utilizing better detectors for handling partial occlusion. In this work, we attempt to handle occlusion from the perspective of feature learning, which is complementary to these detection methods. Specifically, we focus on learning more robust appearance model for each target using the single object tracker with the help of spatial and temporal attention.
Online MOT Algorithm
The overview of the proposed algorithm is shown in Figure 2. The following steps are used for tracking objects:
Step 1. At the current frame , the search area of each target is obtained using motion model. The candidates are sampled within the search area.
Step 2. The features of candidates for each target are extracted using ROI-Pooling and weighted by spatial attention. Then the binary classifier is used to find the best matched candidate with the maximum score, which is used as the estimated target state.
Step 3. The visibility map of each tracked target is inferred from the feature of corresponding estimated target state. The visibility map of the tracked target is then used along with the spatial configurations of the target and its neighboring targets to infer temporal attention.
Step 4. The target-specific CNN branch of each target is updated according to the loss of training samples in current and historical frames weighted by temporal attention. The motion model of each target is updated according to corresponding estimated target state.
Step 5. The object management strategy determines the initialization of new targets and the termination of untracked targets.
Step 6. If frame is not the last frame, then go to Step 1 for the next frame .
2 Dynamic CNN-based MOT Framework
We propose a dynamic CNN-based framework for online MOT, which consists of both shared CNN layers and target-specific CNN branches. As shown in Figure 3, the shared CNN layers encode the whole input frame as a large feature map, from which the feature representation of each target is extracted using ROI-Pooling . For computational efficiency, these shared layers are pre-trained on Imagenet Classification task , and not updated during tracking. All target-specific CNN branches share the same structure, but are separately trained to capture the appearance of different targets. They can be viewed as a set of single-object trackers.
The number of target-specific CNN branches varies with the number of existing targets. Once a new target appears, a new branch will be initialized and added to the model. If a target is considered to be disappeared, its corresponding branch will be removed from the entire model.
3 Online Tracking with STAM
The trajectory of an object can be represented by a series of states denoted by , where . and represent the center location of the target at frame . and denote the width and height of the target, respectively. Multi-object tracking aims to obtain the estimated states of all targets at each frame.
For the -th target to be tracked, its estimated state at frame is obtained by searching from a set of candidate states denoted by , which consists of two subsets:
3.2 Feature Extraction with Spatial Attention
The feature of candidate state is extracted from the shared feature map using ROI-Pooling and spatial attention mechanism. The ROI-Pooling from the shared feature map ignores the fact that the tracked targets could be occluded. In this case, the pooled features would be distorted by the occluded parts. To handle this problem, we propose a spatial attention mechanism which pays more attention to un-occluded regions for feature extraction.
Directly using spatial attention does not work well due to limited training samples in the online learning process. In our work, we first generate the visibility map which encodes the spatial visibility of the input samples. Then the spatial attention is derived from visibility map.
where, is the set of parameters. is modeled as two layers interleaved with ReLU layer. The first layer is a convolution layer which has the kernel size of and produces a feature map with 32 channels. The second layer is a fully connected layer with the output size of . Then the output is reshaped to a map with the size of . Each element in visibility map indicates the visibility of corresponding location in feature map . Some examples of generated visibility maps are shown in Figure 4.
where is implemented by a local connected layer followed by a spatial softmax layer and denotes the parameters. Then the spatial attention map is applied to weight the feature map as
where represents the channel-wise Hadamard product operation, which performs Hadamard product between and each channel of .
3.3 Target State Estimation Using Binary Classifier and Detection Results
Binary Classification. Given the refined feature representation , the classification score is obtained as follows:
where is the output of binary classifier which indicates the probability of candidate state belonging to target . is the parameter of the classifier for target . In our work, is modeled by two layers interleaved with ReLU layer. The first layer is a convolution layer which has the kernel size of and produces a feature map with 5 channels. The second layer is a fully connected layer with the output size of 1. Then a sigmoid function is applied to ensure the output to be in $$.
The primitive estimated state of target is obtained by searching for the candidate state with the maximum classification score as follows:
State Refinement. The primitive estimated state with too low classification score will bias the updating of the model. To avoid model degeneration, if the score is lower than a threshold , the corresponding target is considered as “untracked” in current frame . Otherwise, the primitive state will be further refined using the object detections states .
Specifically, the nearest detection state for is obtained as follows:
where calculates the bounding box IoU overlap ratio between and . Then the final state of target is refined as
where and is a pre-defined threshold.
4 Model Initialization and Online Updating
Each target-specific CNN branch comprises of visibility map, attention map and binary classifier. The parameters for visibility map are initialized in the first frame when the target appears and then all three modules are jointly learned.
For the initialization of parameters in obtaining visibility map, we synthetically generate training samples and the corresponding ground truth based on initial target state.
Feature Replacement. We replace the features of the sample with the features from another target or background at some region and set the ground truth for replaced region to . The replaced region is regarded as occluded. For each sample in the augmented set, the feature replacement is done using different targets/brackgrounds at different regions.
Given these training samples and ground truth visibility maps, the model is trained using cross-entropy loss.
4.2 Online Updating Appearance Model
After initialization in the initial frame, all three modules are jointly updated during tracking using back-propagation algorithm.
Training samples used for online updating are obtained from current frame and historical states. For tracked target, positive samples at current frame are sampled around the estimated target state with small displacements and scale variations. Besides, historical states are also utilized as positive samples. If the target is considered as ”untracked” at current frame, we only use historical states of the target as positive samples. All negative samples are collected at current frame . The target-specific branch needs to have the capability of discriminating the target from other targets and background. So both the estimated states of other tracked targets and the samples randomly sampled from background are treated as the negative samples.
For target , given the current positive samples set , historical positive samples set and the negative samples set , the loss function for updating corresponding target-specific branch is defined as
where, , , and are losses from negative samples, positive samples at current frame, and positive samples in the history, respectively. is the temporal attention introduced below.
Temporal Attention. A crucial problem for model updating is to balance the relative importance between current and historical visual cues. Historical samples are reliable positive samples collected in the past frames, while samples in current frame reflect appearance variations of the target. In this work, we propose a temporal attention mechanism, which dynamically pay attention to current and historical samples based on occlusion status.
Temporal attention of target is inferred from visibility map and the overlap statuses with other targets
where is the mean value of visibility map . is the maximum overlap between and all other targets in current frame . , and are learnable parameters. is the sigmoid function.
Since indicates the occlusion status of target . If is large, it means that target is undergoing severe occlusion at current frame . Consequently, the weight for positive samples at current frame is small according to Eq. 9. There, the temporal attention mechanism provides a good balance between current and historical visual cues of the target. Besides, if is smaller than a threshold , the corresponding target state will be added to the historical samples set of target .
4.3 Updating Motion Model
At frame , the velocity of target is updated as
where denotes the time gap for computing velocity. is the center location of target at frame . The variance of Gaussian noise is defined as
5 Object Management
In our work, a new target is initialized when a newly detected object with high detection score is not covered by any tracked targets. To alleviate the influence of false positive detections, the newly initialized target will be discarded if it is considered as “untracked” (Sec. 3.3.3) or not detected in any of the first frames. For target termination, we simply terminate the target if it is “untracked” for over successive frames. Besides, targets that exit the field of view are also terminated.
Experiments
In this section, we present the experimental results and analysis for the proposed online MOT algorithm.
The proposed algorithm is implemented in MATLAB with Caffe . In our implementation, we use the first ten convolutional layers of the VGG-16 network trained on Imagenet Classification task as the shared CNN layers. The threshold is set to 0.5, which determines whether the location found by single object tracker is covered by a object detection. The thresholds and are set to 0.7 and 0.3 respectively. For online updating, we collect positive and negative samples with and IoU overlap ratios with the target state at current frame, respectively. The detection scores are normalized to the range of $FT_{init}=0.2FT_{term}=2FT_{gap}=0.3F$ in motion model.
2 Datasets
We evaluate our online MOT algorithm on the public available MOT15 and MOT16 benchmarks containing 22 (11 training, 11 test) and 14 (7 training, 7 test) video sequences in unconstrained environments respectively. The ground truth annotations of the training sequences are released. We use the training sequences in MOT15 benchmark for performance analysis of the proposed method. The ground truth annotations of test sequences in both benchmarks are not released and the tracking results are automatically evaluated by the benchmark. So we use the test sequences in two benchmarks for comparison with various state-of-the-art MOT methods. In addition, these two benchmarks also provide object detections generated by the ACF detector and the DPM detector respectively. We use these public detections in all experiments for fair comparison.
3 Evaluation metrics
To evaluate the performance of multi-object tracking methods, we adopt the widely used CLEAR MOT metrics , including multiple object tracking precision (MOTP) and multiple object tracking accuracy (MOTA) which combines false positives (FP), false negatives (FN) and the identity switches (IDS). Additionally, we also use the metrics defined in , which consists of the percentage of mostly tracked targets (MT, a ground truth trajectory that are covered by a tracking hypothesis for at least 80% is regarded as mostly tracked), the percentage of mostly lost targets (ML, a ground truth trajectory that are covered by a tracking hypothesis for at most 20% is regarded as mostly lost), and the number of times a trajectory is fragmented (Frag).
4 Tracking Speed
The overall tracking speed of the proposed method on MOT15 test sequences is 0.5 fps using the 2.4GHz CPU and a TITAN X GPU, while the algorithm without feature sharing runs at 0.1 fps with the same environment.
5 Performance analysis
To demonstrate the effectiveness of the proposed method, we build five algorithms for components of different aspects of our approach. The details of each algorithm are described as follows:
p1: directly using single object trackers without the proposed spatial-temporal attention or motion model, which is the baseline algorithm;
p3: adding the spatial attention based on p2;
p4: adding the temporal attention based on p2;
p5: adding the spatial-temporal attention based on p2, which is the whole algorithm with all proposed components.
The performance of these algorithms on the training sequences of MOT15, in terms of MOTA which is a good approximation of the overall performance, are shown in Figure 5. The better performance of the algorithm p2 compared to p1 shows the effect of the using motion model in MOT. The advantages of the proposed spatial-temporal attention can be seen by comparing the performance of algorithm p5 and p2. Furthermore, compared to the algorithm p2, the performance improvement of p3 and p4 shows the effectiveness of spatial and temporal attention in improving tracking accuracy respectively. The improvement of p5 over both p3 and p4 shows that the spatial and temporal attention are complementary to each other. Algorithm p5 with all the proposed components achieves the best performance and improves 8% in terms of MOTA compared with the baseline algorithm p1, which demonstrates the effectiveness of our algorithm in handling the problems of using single object trackers directly.
6 Comparisons with state-of-the-art methods
We compare our algorithm, denoted by STAM, with several state-of-the-art MOT tracking methods on the test sequences of MOT15 and MOT16 benchmarks. All the compared state-of-the-art methods and ours use the same public detections provided by the benchmark for fair comparison. Table 1 presents the quantitative comparison results The quantitative tracking results of all these trackers are available at the website http://motchallenge.net/results/2D_MOT_2015/ and http://motchallenge.net/results/MOT16/..
MOT15 Results. Overall, STAM achieves the best performance in MOTA and IDS among all the online and offline methods. In terms of MOTA, which is the most important metric for MOT, STAM improves 4% compared with MDP, the best online tracking method that is peer-reviewed and published. Note that our method works in pure online mode and dose not need any training data with ground truth annotations. While MDP performs training with sequences in the similar scenario and its ground truth annotations for different test sequences. Besides, our method produce the lowest IDS among all methods, which demonstrates that our method can handle the interaction among targets well. Note that the CNNTCM and SiameseCNN also utilize CNNs to handle MOT problem but in offline mode. What’s more, their methods requir abundant training data for learning siamese CNN. The better performance compared to these CNN-based offline methods provides strong support on the effectiveness of our online CNN-based algorithm.
MOT16 Results. Similarly, STAM achieves the best performance in terms of MOTA, MT, ML, and FN among all online methods. Besides, the performance of our algorithm in terms of MOTA is also on par with state-of-the-art offline methods.
On the other hand, our method produces slightly more Frag than some offline methods, which is a common defect of online MOT methods due to long term occlusions and severe camera motion fluctuation.
Conclusion
In this paper, we have proposed a dynamic CNN-based online MOT algorithm that efficiently utilizes the merits of single object trackers using shared CNN features and ROI-Pooling. In addition, to alleviate the problem of drift caused by frequent occlusions and interactions among targets, the spatial-temporal attention mechanism is introduced. Besides, a simple motion model is integrated into the algorithm to utilize the motion information. Experimental results on challenging MOT benchmarks demonstrate the effectiveness of the proposed online MOT algorithm.
Acknowledgement: This work is supported by the National Natural Science Foundation of China (No.61371192), the Key Laboratory Foundation of the Chinese Academy of Sciences (CXJJ-17S044), the Fundamental Research Funds for the Central Universities (WK2100330002), SenseTime Group Limited, the General Research Fund sponsored by the Research Grants Council of Hong Kong (Project Nos. CUHK14213616, CUHK14206114, CUHK14205615, CUHK419412, CUHK14203015, CUHK14207814, and CUHK14239816), the Hong Kong Innovation and Technology Support Programme (No.ITS/121/15FX), and ONR N00014-15-1-2356.