Real-time Online Multi-Object Tracking in Compressed Domain
Qiankun Liu, Bin Liu, Yue Wu, Weihai Li, Nenghai Yu
Introduction
Multi-Object Tracking (MOT) is an important computer vision task which aims to estimate the trajectories of interested objects and maintain their identities across frames. It has various applications that with real-time and online requirements, such as autonomous driving and robot navigation. However, real-time online MOT still remains a challenging task.
Driven by advances in object detection, tracking-by-detection has become a popular strategy for MOT. Most existing methods focus on designing a complicate approach to tackle MOT in a data association manner. These methods can be divided into two categories: offline and online methods. The offline methods usually use future frames to track objects, which makes them impractical for casual applications. On the contrary, the online methods track objects based on the past and current frames and have achieved desirable performance, as shown in Figure 1. However, detection and data association are performed frame-by-frame in these online methods to ensure a good tracking performance, which is time-consuming and makes them unable to be applied in real-time applications. In order to re-recognize objects when occlusion happens, appearance affinity is commonly used in data association process . However, the appearance models utilized in these methods are independent from the detector and are therefore potentially sub-optimal.
In practical applications, videos are usually captured, stored and transmitted in the compressed domain. Implementations for some video tasks in the compressed domain are necessary and more suitable. Firstly, implementations in the compressed domain can take lower computational cost since not all frames need to be restored into RGB images (the RGB images in this paper denotes the regular colorful or gray images, which is used to distinguish them from the frames in the compressed domain). Secondly, the motion information is readily provided in the compressed domain which is helpful for video tasks. A few works, such as tracking , video object detection and action recognition , have been done in the compressed domain. Generally, there are two goals of the implementations in the compressed domain: (1) Feature propagation . Features are only extracted from the key frames which are restored into RGB images, and then propagated to the non-key frames. Computational cost can be saved since the frequency of feature extraction is reduced and the feature propagation is much more efficient than feature extraction. (2) Motion cues extraction . The motion cues of objects are directly extracted from the motion information (motion vectors and residuals, more specifically) without access to the RGB images, which are further used to handle the task.
Existing compressed domain based tracking algorithms focus on the motion cues extraction from the pixel level or the bounding-box level . For the pixel level, each pixel that locates in the bounding-box is shifted separately based on the motion vectors (MVs), then the smallest axis-aligned rectangle that includes all the shifted pixels is selected as the new bounding-box. For the bounding-box level, MVs that locate in the bounding-box are averaged to get the displacement, then the bounding-box is shifted. However, the scale variation can not be handled.
In this paper, we focus on real-time online multi-object tracking. To this end, we propose an Online MOT Tracker in Compressed Domain (OTCD). Since the adjacent frames are highly relevant, it is redundant to perform detection and data association in all frames. Motivated by this, we divide the frames into key and non-key frames respectively. For the key frames, detection and data association are performed. An Appearance CNN (A-CNN) which shares features with the detector is designed to assist data association, and it can be jointly trained with the detector. Note that RGB images are restored for the key frames since both detection and appearance feature extraction need to be performed on the RGB images. For the non-key frames, objects are directly propagated based on the motion cues which are extracted from the MVs and residuals by a Tracking CNN (T-CNN). Owning to the sparsity of key frames and the share of features between A-CNN and detector, our tracker achieves a great boost in tracking speed with little performance degradation.
To sum up, our contributions are as follows:
We develop an online unified MOT tracker to track objects in compressed domain for real-time applications.
We propose an appearance CNN to assist data association. The joint training of appearance CNN and detector helps to further promote the performance of our method.
We propose a tracking CNN to propagate objects through non-key frames while maintaining their identities without detection and data association, which accelerates our tracker greatly with little performance degradation.
The rest of this paper is organized as follows. Section 2 reviews the related work. Section 3 introduces the proposed tracker in detail and section 4 represents experimental results. Finally, section 5 makes a conclusion on our work in this paper.
Related work
In this section, we provide a brief overview about the usage of appearance features in MOT and the works implemented in the compressed domain.
Appearance Features in MOT. Appearance features can be used to improve tracking performance in crowded scenario where occlusion often happens. And various appearance features have been used in MOT, such as histogram of gradients , color histogram and integral channel features . Recently, the powerful deep features extracted by CNN have been introduced to MOT . Bae et al. proposed a deep appearance learning method to learn a discriminative appearance model in an online manner. Kieritz et al. designed an appearance model which was incrementally trained online for each object. Both and need to collect the training samples online, which is time-consuming. Chu et al. utilized single object tracker for MOT based on appearance features, but the online learned target-specific CNN layers need to be preserved for each target. Instead of learning the appearance model online, some researchers trained an appearance model offline , which can be used as a function to measure the affinity between different features while tracking online. Zhu et al. trained a spatial attention CNN which can focus on matching patterns of input image patches. The works in trained appearance models using person re-identification dataset and achieved great improvements.
The aforementioned methods separated data association from detection, and the utilized appearance models were isolated from the detector. Our work focuses on designing a compact appearance model, which shares features and can be jointly trained with the detector.
Works in the Compressed Domain. In order to use the motion information provided in the compressed domain freely, a few works have been done in the compressed domain. Ujiie et al. interpolated the bounding-boxes of objects in the bounding-box level for some frames to avoid detection and data association. Alvar et al. constructed approximate bounding-box of the target object in the pixel level based on the bounding-box in previous frame for single object tracking, but each frame needs to be restored into RGB image for detection. The work in fed MVs and residuals to a network to propagate the features across frames for object detection. However, the frames are processed in a batch manner which is inapplicable for online tasks. Wu et al. recognized different actions in the compressed domain and achieved great performance. Nevertheless, the MVs and residuals are traced back to the reference frame and accumulated on the way, which augments the computational cost.
Among these methods, the work in is the most related to our work. However, the work in cannot handle the scale variations of bounding-boxes since MVs are simply used (by averaging MVs that locates in the bounding-boxes) to predict the displacements of objects. While our method utilizes a tracking CNN to predict the velocities of objects based on MVs and residuals, in which the scale variation of bounding-boxes are considered.
Method
Frames in a compressed video are divided into Group of Pictures (GOP), and there are three types of frame generally: I-frame (intra-coded frame), P-frame (predictive frame) and B-frame (bi-directional frame). Among these three types of frame, I-frame can be treated as a regular RGB image, while P-frame and B-frame are encoded with MVs and residuals. The main difference between P-frame and B-frame is that P-frame is encoded in a predictive manner, while B-frame is encoded in a bi-directional manner. We share the same assumption with that the compressed video only contains I-frame and P-frame for simplicity. The reason is that the B-frames, which are encoded in a bi-directional manner, require special handling since we focus on online tracking, and we leave this for the future work. In this paper, we track objects in raw MPEG-4 videos which have an I-frame before every 11 P-frames. However, the proposed method is applicable for different compression techniques, such as MPEG-2 and H.264 . The reason is that different compression techniques usually use motion vectors and residuals for frame compression, thus the frames in the videos can be easily divided into key and non-key frames.
The overview of the proposed tracker OTCD is shown in Figure 2. The frames are divided into key and non-key frames in OTCD. Suppose there is a key frame every consecutive frames. Since MVs and residuals are required for the non-key frames and there are no MVs or residuals for I-frames in the compressed domain, I-frames are always regarded as the key-frames, which means should be a factor of GOP size.
Let and denote the sets of objects and detections in frame respectively. Note that is defined for the key frames only. For a key frame at time , the RGB image is restored and fed into a detector, which produces a set of detections . The objects in from last frame are associated with the detections in . The data association is solved by Hungarian algorithm based on the Intersection-over-Union (IoU) between bounding-boxes and the appearance affinity obtained by A-CNN. After then the birth and death of objects are managed. For each non-key frame, the corresponding MVs and residuals are fed into T-CNN to propagate objects from the previous frame to current frame.
2 Tracking in Key Frames
The tracking in key frames follows the footprint of per-frame approaches, including detection, data association and object management.
The detector is responsible for the detection of interested objects (pedestrian, particularly) and feature extraction. The R-FCN with ResNet-101 is used in our work directly since detection is beyond the scope of this paper. We take pedestrian as the foreground and others as the background for detection. The appearance features for detections are cropped from the feature map provided by the last convolutional layer on the conv4 stage of the backbone of detector.
2.2 Appearance CNN
Given the appearance feature of the detection and one appearance feature of the object , the probability of these two features belonging to the same object is used as the affinity between these two features:
The intuition of the Position-Sensitive (PS) layer is that we assume the features extracted from the same patches in different RGB images should be the same, but the patches may not well aligned due to the inaccurate detection, occlusion and pose change. The corresponding features in and may locate in different spatial positions. Hence, it is necessary to compare the feature vector from one spatial position in with the feature vectors from all spatial positions in , which produces a single channel feature map in .
Training of A-CNN. During the training process, each training sample contains two appearance features cropped from the feature map provided by the detector. The corresponding label is set to 0 (these two appearance features belong to different objects) or 1 (these two appearance features belong to the same object).
A-CNN is trained by the cross-entropy loss. Let and be the loss of A-CNN and detector respectively. Then A-CNN and the detector can be jointly trained via a multi-task loss , where is the weight to balance the loss.
2.3 Data Association
Given the set of detections in key frame , and the set of objects in the previous non-key frame , the data association process is divided into two steps.
Step : assign the detections in to confirmed objects based on the IoU cost between bounding-boxes. Let be the IoU cost between object and detection
and they will not be associated with each other if is greater than a threshold .
Step : assign the unmatched detections to objects in tentative state as well as those unmatched objects in step 1 based on the appearance cost. Let be the appearance cost between object and detection
where is the appearance affinity obtained by A-CNN. And they will not be associated if is greater than a threshold .
Suppose the detection is assigned to the object , the bounding-box of the -th object in frame is succeeded from the bounding-box of . The detection’s appearance feature is added to the feature set . The oldest feature will be abandoned if there are more than features in .
2.4 Object Management
The objects are managed by transforming their states between the pre-defined three states . Particularly:
An unmatched detection is initialized as a tentative object, and it will be confirmed if its detection confidence is larger than a threshold .
A confirmed object is transformed to tentative state if it has not been associated with any detections for more than consecutive key frames.
A tentative object will be confirmed if it has been associated with a detection for more than consecutive key frames.
A tentative object will be deleted if it has not been associated with any detections for more than consecutive key frames.
A deleted object remains at forever.
3 Tracking in Non-key Frames
The tracking in non-key frames is much more straightforward. The objects are directly propagated by T-CNN while maintaining their identities. The appearance features and state of each object are not changed during propagation since the RGB images are not restored for detection and data association. Let be the predicted velocity of the -th object in frame based on the bounding-box in frame :
where is the backbone CNN of T-CNN, and represents the input data for the non-key frame . is the velocity prediction function. As shown in Figure 4, is implemented by position-sensitive region-of-interest pooling layer (PSRoIPooling) proposed in R-FCN . Then the bounding-boxes in frame can be predicted easily by , where is the bounding-box prediction function that defined as
Obviously, the identity is maintained for each object during the propagation process.
Backbone CNN. The network used in T-CNN is much smaller than the network in detector, since the MVs and residuals only store the between two frames. Besides, a smaller network can reduce the computational cost. The backbone CNN is modified from ResNet-18 . Particularly, the last average pooling layer and fully connection layer are removed. As a common practice , the effective stride of ResNet-18 is reduced from pixels to pixels, which increases the resolution of feature maps. Then we can get three types of the modified ResNet-18 by changing the number of input channels in the first convolutional layer to 2, 3 and 5, which are denoted as ResNet2-18, ResNet3-18 and ResNet5-18, respectively.
In order to explore the tracking ability of T-CNN, four T-CNNs are designed, as shown in Figure 5:
T-CNNmv: only MVs are used to predict the velocities.
T-CNNres: only residuals are used to predicted the velocities.
T-CNNmv|res: MVs and residuals are both used, but they are concatenated together firstly to be fed into T-CNN.
T-CNNmv||res: MVs and residuals are both used, and they are fed into their corresponding branches. Then the outputs of these two branches are concatenated together.
For all prototypes of T-CNN, the convolutional layer (followed by a ReLU function) is used to produce a feature map with channels.
Training of T-CNN. The training of T-CNN is independent of the training of A-CNN and detector. The reason is that detection and appearance feature extraction need to be performed on RGB images, while T-CNN needs motion vectors and residuals to predict the velocities of objects. Given the ground-truth bounding-boxes of the -th object in frame and , the ground-truth velocity can be computed by , where is the inverse function of . The loss of T-CNN can be computed by
is the function defined in , and is the number of objects that appear in frames and frame . The objects that only appear in one frame will not contribute to the training process.
4 Time Consumption Analysis
Let , , and be the time consumption of detection, data association, object management and object propagation in each frame. Compared to the per-frame approaches, the speedup factor of our tracker depends on the sparsity of key frames:
Generally, and . In our implementation, . Then is about
For example, our tracker is faster approximately when . The tracking method is shown in Algorithm 1.
EXPERIMENTS
The proposed MOT tracker OTCD is implemented based on PyTorch library without optimization. Evaluation is on a workstation with GHz CPU and Nvidia TITAN Xp GPU.
We evaluate the proposed tracker on Citypersons , 2DMOT2015 , MOT16 and MOT17 . The sequences in MOT16 are the same with those in MOT17 but are provided with a less accurate ground-truth. Each sequence in 2DMOT2015 and MOT17 is compressed into a MPEG-4 video, and all images in Citypersons are compressed into a MPEG-4 video. All data used for training and testing is loaded from the compressed domain with the tool provided by . The loaded MVs and residuals can be treated as special (2 and 3 channels respectively) which have the same resolution with the video.
Sequences in 2DMOT2015 and MOT17 are divided into three sets. Testing set: the sequences in MOT17 test split. Validation set: MOT17-09 and MOT17-10. Training set: the rest sequences in MOT17 as well as those sequences in 2DMOT2015 train split but not included in MOT17.
2 Settings
The variable used in A-CNN and T-CNN is set to . The detections with confidence less than are abandoned. The detection confidence threshold is set to . The number of consecutive key frames , , are set to , , respectively. And the number of historical appearance features is set to . The thresholds and are set to and respectively.
Citypersons and training set are used to train R-FCN and A-CNN. The training samples for A-CNN are generated during training process. Particularly, for each ground-truth box in one image, two positive and one negative samples are collected with and IoU overlap ratios with this ground-truth box. We also randomly choose a ground-truth box belonging to other objects as another negative sample. The separate training of R-FCN and A-CNN are conducted with initial learning rate and respectively. The joint training of R-FCN and A-CNN is divided into 3 phases. Phase 1: train R-FCN for epochs with initial learning rate . Phase 2: train A-CNN for epochs with initial learning rate . Phase 3: train A-CNN and R-FCN jointly with initial learning rate . Here, we set to . T-CNN is trained on training set with initial learning rate . During all training process, the batch size is set to and the learning rate decades every epoches with exponential decay rate . All models are optimized by stochastic gradient descent until they converge.
3 Metrics
Trackers. We choose the following metrics to evaluate different trackers: Multi-Object Tracking Accuracy (MOTA) , Multi-Object Tracking Precision (MOTP) , how often an object is identified by the same ID (IDF1) , Mostly Tracked objects (MT), Mostly Lost objects (ML), number of False Positives (FP), number of False Negatives (FN), number of Identity Switches (IDS) , number of Fragments (Frag), and running speed (Hz).
Detector. Except FP and FN, we choose additional metrics to evaluate different detectors: Recall (Rcll), Precision (Prcn), Multiple Object Detection Accuracy (MODA) .
All metrics are evaluated by the toolkit provided by MOTChallenge benchmark .
4 Components Analyses
We first present the features produced by the PS layer in Figure 6. The PS layer actually divides one image patch into bins, and compares each bin with another image patch. A high similarity score will be produced if the compared image patterns are similar.
In order to demonstrate the effectiveness of the two steps data association procedure introduced in section 3.2.3 (denoted as IoUA-CNN), an one step data association procedure that simultaneously take IoU cost and appearance cost into consideration are also conducted (denoted as IoU+A-CNN). For the one step data association procedure, we use as the cost between object and detection , and they will not be associated with each other if is greater than the threshold . Several experiments are conducted by varying from to with the interval , and the best MOTA is achieved on validation set when , which is our default setting for the one step data association procedure. Note that and are two special cases of IoU+A-CNN which are denoted as IoU and A-CNN, respectively. An A-CNN without PS layer (A-CNN) is also trained to further demonstrate the effectiveness of the PS layer in A-CNN. The results are shown in Table I.
When the appearance cost is only used, MOTA is improved by and IDS is greatly reduced by with the help of the PS layer at the cost of tracking speed dropped from 3.2 Hz to 2.3 Hz. Favorable MOTA and IDS can be obtained when the IoU cost is only used. This is due to the fact that the bounding boxes of one object in the adjacent frames may be much closer, and the IoU cost is sufficient enough to associate them with each other. The best tracking speed is also possessed by the case where the IoU cost is only used. This is reasonable since appearance cost is no need to be computed, which is more computational expensive than IoU cost. However, MOTA is still improved by and IDS is further reduced by when A-CNN is used in IoUA-CNN with the little price of Hz dropped from 16.5 to 15.7.
Though MOTA is almost the same in IoU+A-CNN and IoUA-CNN, a better IDS and Hz are achieved by IoUA-CNN. Here is the explanation: (1) Hz. Both IoU cost and appearance cost need to be computed for each object-detection pair in IoU+A-CNN. While in IoUA-CNN, appearance cost (which is computational expensive) only needs to be computed for a small proportion of object-detection pairs. (2) IDS and MOTA. IoU cost and appearance cost may not be always both reliable. Simultaneously consideration of IoU cost and appearance cost will make them influence each other.
A case of occlusion is also shown in Figure 7. When data association procedure is conducted based on the IoU cost only, the tracker fails to re-recognize the occluded pedestrian. However, when the data association procedure is conducted based on both IoU and appearance cost as described in section 3.2.3, the tracker can re-recognize the occluded pedestrian successfully, which demonstrates the effectiveness of the proposed A-CNN and the data association method.
4.2 Joint Training of Appearance CNN and Detector.
The effectiveness of joint training of A-CNN and detector is shown in Table II. In terms of detection, all metrics are improved. Particularly, FP is greatly reduced about , Prcn and MODA are both improved by more than . In terms of tracking, FP is greatly reduced about , which leads to a higher overall tracking performance MOTA. Furthermore, IDF1 is improved by , which means the tracker can recognize an object with the same ID more often.
4.3 Tracking CNN
To demonstrate the effectiveness of the proposed T-CNN, we compare our method with some other trackers, including DeepSORT , SORT and IOU . The results are summarized in Figure 8 (b). The bounding-boxes of objects in the non-key frames are predicted by Kalman filter in DeepSORT and SORT, while they are copied from previous frame in IOU. For all trackers, the detection time (about ms) is considered and only detections in the key frames are provided for a fair comparison. However, the time consumption in the non-key frames are not considered for DeepSORT, SORT and IOU. As we can see, the tracking speed is limited due to the detection time consumption when . But all trackers achieve significant speedup with the descent tracking accuracy drop when . Thanks to the powerful tracking capability of T-CNN, our tracker possesses the slowest performance decline when compared with other trackers. For example, OTCDmv|res is accelerated from Hz to Hz at the cost of accuracy drop from to , while DeepSORT is accelerated from Hz to Hz at the cost of accuracy drop from to .
The effective performance gains brought by T-CNN does not hold when increases. Reasons are as follows: (1) Tracking accuracy. The birth and death of objects are not handled in the non-key frames, but objects may disappear or reappear during these frames, which leads to a poorer tracking performance when increases. (2) Tracking speed. Since we track objects without detection or data association on the non-key frames, the bounding-boxes of tracked objects may be imperfect, which results in more imperfect bounding-boxes in the following non-key frames. So more confirmed objects will not be associated with detections based on the IoU cost when the next one key frame arrives. Hence, more detections need to be assigned to objects based on the appearance cost, which is more time-consuming than the IoU cost. We choose for the balance of performance and speed.
5 MOTChallenge Benchmark
We compare the proposed tracker OTCD with other online state-of-the-art trackers in Table III on MOT16 and MOT17 test splits. During this evaluation, the detector is used for feature extraction, which means except for the backbone network of the detector, other parts of the detector are not used. The metrics of MVint(LinearK) are accessed from , while others are accessed from MOTChallenge leaderboards.
MOT16. We first compare OTCD with some trackers that track objects based on the detections provided by . Note that the tracking accuracy MOTA and speed Hz of DeepSORT is superior to those of OTCD, which is not the result reflected in Figure 8 (b), the reasons are as follows: (1) MOTA. The detections used in Figure 8 (b) and Table III are different. The quality of our own detections is inferior to the detections provided by . And a poor quality of detections has a negative impact on trackers. (2) Hz. Detection time consumption is considered for DeepSORT, SORT and OTCD in Figure 8 (b), while it is not considered in Table III.
Both MVint(LinearK) and OTCD are in the compressed domain, but OTCD achieves a better performance than MVint(LinearK) in all metrics except FP, IDS and Frag. Particularly, the tracking speed of OTCD is about faster than MVint(LinearK), while possessing a higher MOTA. The tracking results of OTCD based on our own detections are also provided. Note that there is a big gap of MOTA between OTCD and OTCD, which means the quality of detections has a significant impact on the tracking performance.
As for the public detections, OTCD achieves the fastest tracking speed among all trackers while maintaining a satisfying tracking accuracy. For example, the tracking speed of OTCD is faster than AMIR, which is the state-of-the-art method. Furthermore, OTCD achieves the best performance in Frag.
Interestingly, when is increased from to , both IDS and Frag are greatly reduced. This is reasonable since objects need to be recognized with a lower frequency when data association is only performed in sparse key frames.
MOT17. Overall, the tracking performance of OTCD is comparable with other trackers. Particularly, OTCD1 performs the best in FP, and OTCD3 takes the second place in Hz. Compared with MTDF17, which achieves the best in MOTA, our trackers run more than (OTCD1) and (OTCD3) faster but only with (OTCD1) and (OTCD3) performance degradation in MOTA. Compared with GM_PHD, which possesses the best tracking speed (38.4 Hz), OTCD3 tracks object at a lower speed (33.4 Hz), but with a 10.5% higher MOTA.
Conclusion
In this paper, we propose an online MOT tracker OTCD in compressed domain. The RGB images are restored in the key frames for detection and data association, while the MVs and residuals are directly fed into a tracking CNN to propagate objects through non-key frames, which can accelerate our tracker significantly. Furthermore, an appearance CNN which shares features with detector is introduced to assist data association, and it can be trained with detector jointly. Experimental results on MOTChallenge benchmark demonstrate the effectiveness of appearance CNN and tracking CNN, as well as the joint training of appearance CNN and detector.
Acknowledgments
This work is supported by the National Natural Science Foundation of China (Grant No. 61371192), the Key Laboratory Foundation of the Chinese Academy of Sciences (CXJJ-17S044) and the Fundamental Research Funds for the Central Universities (WK2100330002).