Real-time Online Multi-Object Tracking in Compressed Domain

Qiankun Liu, Bin Liu, Yue Wu, Weihai Li, Nenghai Yu

Introduction

Multi-Object Tracking (MOT) is an important computer vision task which aims to estimate the trajectories of interested objects and maintain their identities across frames. It has various applications that with real-time and online requirements, such as autonomous driving and robot navigation. However, real-time online MOT still remains a challenging task.

Driven by advances in object detection, tracking-by-detection has become a popular strategy for MOT. Most existing methods focus on designing a complicate approach to tackle MOT in a data association manner. These methods can be divided into two categories: offline and online methods. The offline methods usually use future frames to track objects, which makes them impractical for casual applications. On the contrary, the online methods track objects based on the past and current frames and have achieved desirable performance, as shown in Figure 1. However, detection and data association are performed frame-by-frame in these online methods to ensure a good tracking performance, which is time-consuming and makes them unable to be applied in real-time applications. In order to re-recognize objects when occlusion happens, appearance affinity is commonly used in data association process . However, the appearance models utilized in these methods are independent from the detector and are therefore potentially sub-optimal.

In practical applications, videos are usually captured, stored and transmitted in the compressed domain. Implementations for some video tasks in the compressed domain are necessary and more suitable. Firstly, implementations in the compressed domain can take lower computational cost since not all frames need to be restored into RGB images (the RGB images in this paper denotes the regular colorful or gray images, which is used to distinguish them from the frames in the compressed domain). Secondly, the motion information is readily provided in the compressed domain which is helpful for video tasks. A few works, such as tracking , video object detection and action recognition , have been done in the compressed domain. Generally, there are two goals of the implementations in the compressed domain: (1) Feature propagation . Features are only extracted from the key frames which are restored into RGB images, and then propagated to the non-key frames. Computational cost can be saved since the frequency of feature extraction is reduced and the feature propagation is much more efficient than feature extraction. (2) Motion cues extraction . The motion cues of objects are directly extracted from the motion information (motion vectors and residuals, more specifically) without access to the RGB images, which are further used to handle the task.

Existing compressed domain based tracking algorithms focus on the motion cues extraction from the pixel level or the bounding-box level . For the pixel level, each pixel that locates in the bounding-box is shifted separately based on the motion vectors (MVs), then the smallest axis-aligned rectangle that includes all the shifted pixels is selected as the new bounding-box. For the bounding-box level, MVs that locate in the bounding-box are averaged to get the displacement, then the bounding-box is shifted. However, the scale variation can not be handled.

In this paper, we focus on real-time online multi-object tracking. To this end, we propose an Online MOT Tracker in Compressed Domain (OTCD). Since the adjacent frames are highly relevant, it is redundant to perform detection and data association in all frames. Motivated by this, we divide the frames into key and non-key frames respectively. For the key frames, detection and data association are performed. An Appearance CNN (A-CNN) which shares features with the detector is designed to assist data association, and it can be jointly trained with the detector. Note that RGB images are restored for the key frames since both detection and appearance feature extraction need to be performed on the RGB images. For the non-key frames, objects are directly propagated based on the motion cues which are extracted from the MVs and residuals by a Tracking CNN (T-CNN). Owning to the sparsity of key frames and the share of features between A-CNN and detector, our tracker achieves a great boost in tracking speed with little performance degradation.

To sum up, our contributions are as follows:

We develop an online unified MOT tracker to track objects in compressed domain for real-time applications.

We propose an appearance CNN to assist data association. The joint training of appearance CNN and detector helps to further promote the performance of our method.

We propose a tracking CNN to propagate objects through non-key frames while maintaining their identities without detection and data association, which accelerates our tracker greatly with little performance degradation.

The rest of this paper is organized as follows. Section 2 reviews the related work. Section 3 introduces the proposed tracker in detail and section 4 represents experimental results. Finally, section 5 makes a conclusion on our work in this paper.

Related work

In this section, we provide a brief overview about the usage of appearance features in MOT and the works implemented in the compressed domain.

Appearance Features in MOT. Appearance features can be used to improve tracking performance in crowded scenario where occlusion often happens. And various appearance features have been used in MOT, such as histogram of gradients , color histogram and integral channel features . Recently, the powerful deep features extracted by CNN have been introduced to MOT . Bae et al. proposed a deep appearance learning method to learn a discriminative appearance model in an online manner. Kieritz et al. designed an appearance model which was incrementally trained online for each object. Both and need to collect the training samples online, which is time-consuming. Chu et al. utilized single object tracker for MOT based on appearance features, but the online learned target-specific CNN layers need to be preserved for each target. Instead of learning the appearance model online, some researchers trained an appearance model offline , which can be used as a function to measure the affinity between different features while tracking online. Zhu et al. trained a spatial attention CNN which can focus on matching patterns of input image patches. The works in trained appearance models using person re-identification dataset and achieved great improvements.

The aforementioned methods separated data association from detection, and the utilized appearance models were isolated from the detector. Our work focuses on designing a compact appearance model, which shares features and can be jointly trained with the detector.

Works in the Compressed Domain. In order to use the motion information provided in the compressed domain freely, a few works have been done in the compressed domain. Ujiie et al. interpolated the bounding-boxes of objects in the bounding-box level for some frames to avoid detection and data association. Alvar et al. constructed approximate bounding-box of the target object in the pixel level based on the bounding-box in previous frame for single object tracking, but each frame needs to be restored into RGB image for detection. The work in fed MVs and residuals to a network to propagate the features across frames for object detection. However, the frames are processed in a batch manner which is inapplicable for online tasks. Wu et al. recognized different actions in the compressed domain and achieved great performance. Nevertheless, the MVs and residuals are traced back to the reference frame and accumulated on the way, which augments the computational cost.

Among these methods, the work in is the most related to our work. However, the work in cannot handle the scale variations of bounding-boxes since MVs are simply used (by averaging MVs that locates in the bounding-boxes) to predict the displacements of objects. While our method utilizes a tracking CNN to predict the velocities of objects based on MVs and residuals, in which the scale variation of bounding-boxes are considered.

Method

Frames in a compressed video are divided into Group of Pictures (GOP), and there are three types of frame generally: I-frame (intra-coded frame), P-frame (predictive frame) and B-frame (bi-directional frame). Among these three types of frame, I-frame can be treated as a regular RGB image, while P-frame and B-frame are encoded with MVs and residuals. The main difference between P-frame and B-frame is that P-frame is encoded in a predictive manner, while B-frame is encoded in a bi-directional manner. We share the same assumption with that the compressed video only contains I-frame and P-frame for simplicity. The reason is that the B-frames, which are encoded in a bi-directional manner, require special handling since we focus on online tracking, and we leave this for the future work. In this paper, we track objects in raw MPEG-4 videos which have an I-frame before every 11 P-frames. However, the proposed method is applicable for different compression techniques, such as MPEG-2 and H.264 . The reason is that different compression techniques usually use motion vectors and residuals for frame compression, thus the frames in the videos can be easily divided into key and non-key frames.

The overview of the proposed tracker OTCD is shown in Figure 2. The frames are divided into key and non-key frames in OTCD. Suppose there is a key frame every KK consecutive frames. Since MVs and residuals are required for the non-key frames and there are no MVs or residuals for I-frames in the compressed domain, I-frames are always regarded as the key-frames, which means KK should be a factor of GOP size.

Let Ot={oit}i=1ItO^{t}=\{o_{i}^{t}\}_{i=1}^{I_{t}} and Dt={djt}j=1JtD^{t}=\{d_{j}^{t}\}_{j=1}^{J_{t}} denote the sets of objects and detections in frame tt respectively. Note that DtD^{t} is defined for the key frames only. For a key frame at time tt, the RGB image is restored and fed into a detector, which produces a set of detections DtD^{t}. The objects in Ot−1O^{t-1} from last frame are associated with the detections in DtD^{t}. The data association is solved by Hungarian algorithm based on the Intersection-over-Union (IoU) between bounding-boxes and the appearance affinity obtained by A-CNN. After then the birth and death of objects are managed. For each non-key frame, the corresponding MVs and residuals are fed into T-CNN to propagate objects from the previous frame to current frame.

2 Tracking in Key Frames

The tracking in key frames follows the footprint of per-frame approaches, including detection, data association and object management.

The detector is responsible for the detection of interested objects (pedestrian, particularly) and feature extraction. The R-FCN with ResNet-101 is used in our work directly since detection is beyond the scope of this paper. We take pedestrian as the foreground and others as the background for detection. The appearance features for detections are cropped from the feature map provided by the last convolutional layer on the conv4 stage of the backbone of detector.

2.2 Appearance CNN

Given the appearance feature fjtf_{j}^{t} of the detection djtd_{j}^{t} and one appearance feature fit−τ∈Fif_{i}^{t-\tau}\in F_{i} of the object oit−1o_{i}^{t-1}, the probability pi,jτp_{i,j}^{\tau} of these two features belonging to the same object is used as the affinity between these two features:

The intuition of the Position-Sensitive (PS) layer is that we assume the features extracted from the same patches in different RGB images should be the same, but the patches may not well aligned due to the inaccurate detection, occlusion and pose change. The corresponding features in fi′f^{\prime}_{i} and fj′f^{\prime}_{j} may locate in different spatial positions. Hence, it is necessary to compare the feature vector from one spatial position in fi′f^{\prime}_{i} with the feature vectors from all spatial positions in fj′f^{\prime}_{j}, which produces a single channel feature map in fj,if_{j,i}.

Training of A-CNN. During the training process, each training sample contains two appearance features cropped from the feature map provided by the detector. The corresponding label is set to 0 (these two appearance features belong to different objects) or 1 (these two appearance features belong to the same object).

A-CNN is trained by the cross-entropy loss. Let LAL_{A} and LDL_{D} be the loss of A-CNN and detector respectively. Then A-CNN and the detector can be jointly trained via a multi-task loss L=LD+λLAL=L_{D}+\lambda L_{A}, where λ\lambda is the weight to balance the loss.

2.3 Data Association

Given the set DtD^{t} of detections in key frame tt, and the set Ot−1O^{t-1} of objects in the previous non-key frame t−1t-1, the data association process is divided into two steps.

Step 11: assign the detections in DtD^{t} to confirmed objects based on the IoU cost between bounding-boxes. Let ci,jiouc_{i,j}^{iou} be the IoU cost between object oit−1o_{i}^{t-1} and detection djtd^{t}_{j}

and they will not be associated with each other if ci,jiouc_{i,j}^{iou} is greater than a threshold τiou\tau_{iou}.

Step 22: assign the unmatched detections to objects in tentative state as well as those unmatched objects in step 1 based on the appearance cost. Let ci,jappc_{i,j}^{app} be the appearance cost between object oit−1o_{i}^{t-1} and detection djtd^{t}_{j}

where pi,jτp_{i,j}^{\tau} is the appearance affinity obtained by A-CNN. And they will not be associated if ci,jappc_{i,j}^{app} is greater than a threshold τapp\tau_{app}.

Suppose the detection djtd^{t}_{j} is assigned to the object oit−1o_{i}^{t-1}, the bounding-box of the ii-th object in frame tt is succeeded from the bounding-box of djtd_{j}^{t}. The detection’s appearance feature fjtf_{j}^{t} is added to the feature set FiF_{i}. The oldest feature will be abandoned if there are more than lfl_{f} features in FiF_{i}.

2.4 Object Management

The objects are managed by transforming their states between the pre-defined three states {sT,sC,sD}\{s_{T},s_{C},s_{D}\}. Particularly:

An unmatched detection is initialized as a tentative object, and it will be confirmed if its detection confidence is larger than a threshold csT→sCc_{s_{T}\rightarrow s_{C}}.

A confirmed object is transformed to tentative state if it has not been associated with any detections for more than lsC→sTl_{s_{C}\rightarrow s_{T}} consecutive key frames.

A tentative object will be confirmed if it has been associated with a detection for more than lsT→sCl_{s_{T}\rightarrow s_{C}} consecutive key frames.

A tentative object will be deleted if it has not been associated with any detections for more than lsT→sDl_{s_{T}\rightarrow s_{D}} consecutive key frames.

A deleted object remains at sDs_{D} forever.

3 Tracking in Non-key Frames

The tracking in non-key frames is much more straightforward. The objects are directly propagated by T-CNN while maintaining their identities. The appearance features and state of each object are not changed during propagation since the RGB images are not restored for detection and data association. Let v^it=(v^i,xt,v^i,yt,v^i,wt,v^i,ht)\hat{v}_{i}^{t}=(\hat{v}_{i,x}^{t},\hat{v}_{i,y}^{t},\hat{v}_{i,w}^{t},\hat{v}_{i,h}^{t}) be the predicted velocity of the ii-th object in frame tt based on the bounding-box in frame t−1t-1:

where NT(⋅)\mathcal{N}_{T}(\cdot) is the backbone CNN of T-CNN, and ft\mathfrak{f}^{t} represents the input data for the non-key frame tt. V(⋅,⋅)\mathcal{V}(\cdot,\cdot) is the velocity prediction function. As shown in Figure 4, V(⋅,⋅)\mathcal{V}(\cdot,\cdot) is implemented by position-sensitive region-of-interest pooling layer (PSRoIPooling) proposed in R-FCN . Then the bounding-boxes in frame tt can be predicted easily by bit=B(v^it,bit−1)b_{i}^{t}=\mathcal{B}(\hat{v}_{i}^{t},b_{i}^{t-1}), where B(⋅,⋅)\mathcal{B}(\cdot,\cdot) is the bounding-box prediction function that defined as

Obviously, the identity is maintained for each object during the propagation process.

Backbone CNN. The network used in T-CNN is much smaller than the network in detector, since the MVs and residuals only store the changeschanges between two frames. Besides, a smaller network can reduce the computational cost. The backbone CNN is modified from ResNet-18 . Particularly, the last average pooling layer and fully connection layer are removed. As a common practice , the effective stride of ResNet-18 is reduced from 3232 pixels to 1616 pixels, which increases the resolution of feature maps. Then we can get three types of the modified ResNet-18 by changing the number of input channels in the first convolutional layer to 2, 3 and 5, which are denoted as ResNet2-18, ResNet3-18 and ResNet5-18, respectively.

In order to explore the tracking ability of T-CNN, four T-CNNs are designed, as shown in Figure 5:

T-CNNmv: only MVs are used to predict the velocities.

T-CNNres: only residuals are used to predicted the velocities.

T-CNNmv|res: MVs and residuals are both used, but they are concatenated together firstly to be fed into T-CNN.

T-CNNmv||res: MVs and residuals are both used, and they are fed into their corresponding branches. Then the outputs of these two branches are concatenated together.

For all prototypes of T-CNN, the 1×11\times 1 convolutional layer (followed by a ReLU function) is used to produce a feature map with 4m24m^{2} channels.

Training of T-CNN. The training of T-CNN is independent of the training of A-CNN and detector. The reason is that detection and appearance feature extraction need to be performed on RGB images, while T-CNN needs motion vectors and residuals to predict the velocities of objects. Given the ground-truth bounding-boxes of the ii-th object in frame t−1t-1 and tt, the ground-truth velocity vit=(vi,xt,vi,yt,vi,wt,vi,ht)v_{i}^{t}=(v_{i,x}^{t},v_{i,y}^{t},v_{i,w}^{t},v_{i,h}^{t}) can be computed by vit=B−1(bit−1,bit)v_{i}^{t}=\mathcal{B}^{-1}(b_{i}^{t-1},b_{i}^{t}), where B−1(⋅,⋅)\mathcal{B}^{-1}(\cdot,\cdot) is the inverse function of B(⋅,⋅)\mathcal{B}(\cdot,\cdot). The loss of T-CNN can be computed by

is the function defined in , and II is the number of objects that appear in frames tt and frame t−1t-1. The objects that only appear in one frame will not contribute to the training process.

4 Time Consumption Analysis

Let TdetT_{det}, TassT_{ass}, TmanT_{man} and TproT_{pro} be the time consumption of detection, data association, object management and object propagation in each frame. Compared to the per-frame approaches, the speedup factor ss of our tracker depends on the sparsity of key frames:

Generally, Tman≪TdetT_{man}\ll T_{det} and Tman≪TassT_{man}\ll T_{ass}. In our implementation, Tpro≈Tdet+Tass10T_{pro}\approx\frac{T_{det}+T_{ass}}{10}. Then ss is about

For example, our tracker is 2.5×2.5\times faster approximately when K=3K=3. The tracking method is shown in Algorithm 1.

EXPERIMENTS

The proposed MOT tracker OTCD is implemented based on PyTorch library without optimization. Evaluation is on a workstation with 2.62.6 GHz CPU and Nvidia TITAN Xp GPU.

We evaluate the proposed tracker on Citypersons , 2DMOT2015 , MOT16 and MOT17 . The sequences in MOT16 are the same with those in MOT17 but are provided with a less accurate ground-truth. Each sequence in 2DMOT2015 and MOT17 is compressed into a MPEG-4 video, and all images in Citypersons are compressed into a MPEG-4 video. All data used for training and testing is loaded from the compressed domain with the tool provided by . The loaded MVs and residuals can be treated as special imagesimages (2 and 3 channels respectively) which have the same resolution with the video.

Sequences in 2DMOT2015 and MOT17 are divided into three sets. Testing set: the sequences in MOT17 test split. Validation set: MOT17-09 and MOT17-10. Training set: the rest sequences in MOT17 as well as those sequences in 2DMOT2015 train split but not included in MOT17.

2 Settings

The variable mm used in A-CNN and T-CNN is set to 77. The detections with confidence less than 0.950.95 are abandoned. The detection confidence threshold csT→sCc_{s_{T}\rightarrow s_{C}} is set to 0.990.99. The number of consecutive key frames lsT→sCl_{s_{T}\rightarrow s_{C}}, lsC→sTl_{s_{C}\rightarrow s_{T}}, lsT→sDl_{s_{T}\rightarrow s_{D}} are set to 33, 22, 1010 respectively. And the number of historical appearance features lfl_{f} is set to 2424. The thresholds τiou\tau_{iou} and τapp\tau_{app} are set to 0.30.3 and 0.250.25 respectively.

Citypersons and training set are used to train R-FCN and A-CNN. The training samples for A-CNN are generated during training process. Particularly, for each ground-truth box in one image, two positive and one negative samples are collected with ≥0.7\geq 0.7 and ≤0.3\leq 0.3 IoU overlap ratios with this ground-truth box. We also randomly choose a ground-truth box belonging to other objects as another negative sample. The separate training of R-FCN and A-CNN are conducted with initial learning rate 10−310^{-3} and 10−410^{-4} respectively. The joint training of R-FCN and A-CNN is divided into 3 phases. Phase 1: train R-FCN for 1515 epochs with initial learning rate 10−310^{-3}. Phase 2: train A-CNN for 55 epochs with initial learning rate 10−410^{-4}. Phase 3: train A-CNN and R-FCN jointly with initial learning rate 10−410^{-4}. Here, we set λ\lambda to 11. T-CNN is trained on training set with initial learning rate 10−410^{-4}. During all training process, the batch size is set to 22 and the learning rate decades every 88 epoches with exponential decay rate 0.10.1. All models are optimized by stochastic gradient descent until they converge.

3 Metrics

Trackers. We choose the following metrics to evaluate different trackers: Multi-Object Tracking Accuracy (MOTA) , Multi-Object Tracking Precision (MOTP) , how often an object is identified by the same ID (IDF1) , Mostly Tracked objects (MT), Mostly Lost objects (ML), number of False Positives (FP), number of False Negatives (FN), number of Identity Switches (IDS) , number of Fragments (Frag), and running speed (Hz).

Detector. Except FP and FN, we choose additional metrics to evaluate different detectors: Recall (Rcll), Precision (Prcn), Multiple Object Detection Accuracy (MODA) .

All metrics are evaluated by the toolkit provided by MOTChallenge benchmark .

4 Components Analyses

We first present the features produced by the PS layer in Figure 6. The PS layer actually divides one image patch into m×mm\times m bins, and compares each bin with another image patch. A high similarity score will be produced if the compared image patterns are similar.

In order to demonstrate the effectiveness of the two steps data association procedure introduced in section 3.2.3 (denoted as IoU→\rightarrowA-CNN), an one step data association procedure that simultaneously take IoU cost and appearance cost into consideration are also conducted (denoted as IoU+A-CNN). For the one step data association procedure, we use ci,j=αci,jiou+(1−α)ci,jappc_{i,j}=\alpha c^{iou}_{i,j}+(1-\alpha)c^{app}_{i,j} as the cost between object oit−1o_{i}^{t-1} and detection djtd^{t}_{j}, and they will not be associated with each other if ci,jc_{i,j} is greater than the threshold τ=ατiou+(1−α)τapp\tau=\alpha\tau^{iou}+(1-\alpha)\tau^{app}. Several experiments are conducted by varying α\alpha from 0.10.1 to 0.90.9 with the interval 0.10.1, and the best MOTA is achieved on validation set when α=0.5\alpha=0.5, which is our default setting for the one step data association procedure. Note that α=1\alpha=1 and α=0\alpha=0 are two special cases of IoU+A-CNN which are denoted as IoU and A-CNN, respectively. An A-CNN without PS layer (A-CNNPS−{}^{-}_{\rm PS}) is also trained to further demonstrate the effectiveness of the PS layer in A-CNN. The results are shown in Table I.

When the appearance cost is only used, MOTA is improved by 1.2%1.2\% and IDS is greatly reduced by 37.3%37.3\% with the help of the PS layer at the cost of tracking speed dropped from 3.2 Hz to 2.3 Hz. Favorable MOTA and IDS can be obtained when the IoU cost is only used. This is due to the fact that the bounding boxes of one object in the adjacent frames may be much closer, and the IoU cost is sufficient enough to associate them with each other. The best tracking speed is also possessed by the case where the IoU cost is only used. This is reasonable since appearance cost is no need to be computed, which is more computational expensive than IoU cost. However, MOTA is still improved by 0.5%0.5\% and IDS is further reduced by 27.7%27.7\% when A-CNN is used in IoU→\rightarrowA-CNN with the little price of Hz dropped from 16.5 to 15.7.

Though MOTA is almost the same in IoU+A-CNN and IoU→\rightarrowA-CNN, a better IDS and Hz are achieved by IoU→\rightarrowA-CNN. Here is the explanation: (1) Hz. Both IoU cost and appearance cost need to be computed for each object-detection pair in IoU+A-CNN. While in IoU→\rightarrowA-CNN, appearance cost (which is computational expensive) only needs to be computed for a small proportion of object-detection pairs. (2) IDS and MOTA. IoU cost and appearance cost may not be always both reliable. Simultaneously consideration of IoU cost and appearance cost will make them influence each other.

A case of occlusion is also shown in Figure 7. When data association procedure is conducted based on the IoU cost only, the tracker fails to re-recognize the occluded pedestrian. However, when the data association procedure is conducted based on both IoU and appearance cost as described in section 3.2.3, the tracker can re-recognize the occluded pedestrian successfully, which demonstrates the effectiveness of the proposed A-CNN and the data association method.

4.2 Joint Training of Appearance CNN and Detector.

The effectiveness of joint training of A-CNN and detector is shown in Table II. In terms of detection, all metrics are improved. Particularly, FP is greatly reduced about 19.8%19.8\%, Prcn and MODA are both improved by more than 2%2\%. In terms of tracking, FP is greatly reduced about 36.6%36.6\%, which leads to a higher overall tracking performance MOTA. Furthermore, IDF1 is improved by 2.2%2.2\%, which means the tracker can recognize an object with the same ID more often.

4.3 Tracking CNN

To demonstrate the effectiveness of the proposed T-CNN, we compare our method with some other trackers, including DeepSORT , SORT and IOU . The results are summarized in Figure 8 (b). The bounding-boxes of objects in the non-key frames are predicted by Kalman filter in DeepSORT and SORT, while they are copied from previous frame in IOU. For all trackers, the detection time (about 6060 ms) is considered and only detections in the key frames are provided for a fair comparison. However, the time consumption in the non-key frames are not considered for DeepSORT, SORT and IOU. As we can see, the tracking speed is limited due to the detection time consumption when K=1K=1. But all trackers achieve significant speedup with the descent tracking accuracy drop when K>1K>1. Thanks to the powerful tracking capability of T-CNN, our tracker possesses the slowest performance decline when compared with other trackers. For example, OTCDmv|res is accelerated from 15.815.8 Hz to 46.046.0 Hz at the cost of accuracy drop from 38.4%38.4\% to 37.8%37.8\%, while DeepSORT is accelerated from 8.58.5 Hz to 34.034.0 Hz at the cost of accuracy drop from 38.4%38.4\% to 33.4%33.4\%.

The effective performance gains brought by T-CNN does not hold when KK increases. Reasons are as follows: (1) Tracking accuracy. The birth and death of objects are not handled in the non-key frames, but objects may disappear or reappear during these frames, which leads to a poorer tracking performance when KK increases. (2) Tracking speed. Since we track objects without detection or data association on the non-key frames, the bounding-boxes of tracked objects may be imperfect, which results in more imperfect bounding-boxes in the following non-key frames. So more confirmed objects will not be associated with detections based on the IoU cost when the next one key frame arrives. Hence, more detections need to be assigned to objects based on the appearance cost, which is more time-consuming than the IoU cost. We choose K=3K=3 for the balance of performance and speed.

5 MOTChallenge Benchmark

We compare the proposed tracker OTCD with other online state-of-the-art trackers in Table III on MOT16 and MOT17 test splits. During this evaluation, the detector is used for feature extraction, which means except for the backbone network of the detector, other parts of the detector are not used. The metrics of MVint(LinearK) are accessed from , while others are accessed from MOTChallenge leaderboards.

MOT16. We first compare OTCD with some trackers that track objects based on the detections provided by . Note that the tracking accuracy MOTA and speed Hz of DeepSORT is superior to those of OTCD1⋆{}^{\star}_{1}, which is not the result reflected in Figure 8 (b), the reasons are as follows: (1) MOTA. The detections used in Figure 8 (b) and Table III are different. The quality of our own detections is inferior to the detections provided by . And a poor quality of detections has a negative impact on trackers. (2) Hz. Detection time consumption is considered for DeepSORT, SORT and OTCD in Figure 8 (b), while it is not considered in Table III.

Both MVint(LinearK) and OTCD3⋆{}^{\star}_{3} are in the compressed domain, but OTCD3⋆{}^{\star}_{3} achieves a better performance than MVint(LinearK) in all metrics except FP, IDS and Frag. Particularly, the tracking speed of OTCD3⋆{}^{\star}_{3} is about 1.7×1.7\times faster than MVint(LinearK), while possessing a 4.7%4.7\% higher MOTA. The tracking results of OTCD based on our own detections are also provided. Note that there is a big gap of MOTA between OTCD1⋆{}^{\star}_{1} and OTCD1†{}^{\dagger}_{1}, which means the quality of detections has a significant impact on the tracking performance.

As for the public detections, OTCD achieves the fastest tracking speed among all trackers while maintaining a satisfying tracking accuracy. For example, the tracking speed of OTCD1⋆{}^{\star}_{1} is 17×17\times faster than AMIR, which is the state-of-the-art method. Furthermore, OTCD3⋆{}^{\star}_{3} achieves the best performance in Frag.

Interestingly, when KK is increased from 11 to 33, both IDS and Frag are greatly reduced. This is reasonable since objects need to be recognized with a lower frequency when data association is only performed in sparse key frames.

MOT17. Overall, the tracking performance of OTCD is comparable with other trackers. Particularly, OTCD1 performs the best in FP, and OTCD3 takes the second place in Hz. Compared with MTDF17, which achieves the best in MOTA, our trackers run more than 10×10\times (OTCD1) and 25×25\times (OTCD3) faster but only with 1.0%1.0\% (OTCD1) and 2.7%2.7\% (OTCD3) performance degradation in MOTA. Compared with GM_PHD, which possesses the best tracking speed (38.4 Hz), OTCD3 tracks object at a lower speed (33.4 Hz), but with a 10.5% higher MOTA.

Conclusion

In this paper, we propose an online MOT tracker OTCD in compressed domain. The RGB images are restored in the key frames for detection and data association, while the MVs and residuals are directly fed into a tracking CNN to propagate objects through non-key frames, which can accelerate our tracker significantly. Furthermore, an appearance CNN which shares features with detector is introduced to assist data association, and it can be trained with detector jointly. Experimental results on MOTChallenge benchmark demonstrate the effectiveness of appearance CNN and tracking CNN, as well as the joint training of appearance CNN and detector.

Acknowledgments

This work is supported by the National Natural Science Foundation of China (Grant No. 61371192), the Key Laboratory Foundation of the Chinese Academy of Sciences (CXJJ-17S044) and the Fundamental Research Funds for the Central Universities (WK2100330002).

References