Tracking Objects as Points

Xingyi Zhou, Vladlen Koltun, Philipp Krähenbühl

Introduction

In early computer vision, tracking was commonly phrased as following interest points through space and time . Early trackers were simple, fast, and reasonably robust. However, they were liable to fail in the absence of strong low-level cues such as corners and intensity peaks. With the advent of high-performing object detection models , a powerful alternative emerged: tracking-by-detection (or more precisely, tracking-after-detection) . These models rely on a given accurate recognition to identify objects and then link them up through time in a separate stage. Tracking-by-detection leverages the power of deep-learning-based object detectors and is currently the dominant tracking paradigm. Yet the best-performing object trackers are not without drawbacks. Many rely on slow and complex association strategies to link detected boxes through time . Recent work on simultaneous detection and tracking has made progress in alleviating some of this complexity. Here, we show how combining ideas from point-based tracking and simultaneous detection and tracking further simplifies tracking.

We present a point-based framework for joint detection and tracking, referred to as CenterTrack. Each object is represented by a single point at the center of its bounding box. This center point is then tracked through time (Figure 1). Specifically, we adopt the recent CenterNet detector to localize object centers . We condition the detector on two consecutive frames, as well as a heatmap of prior tracklets, represented as points. We train the detector to also output an offset vector from the current object center to its center in the previous frame. We learn this offset as an attribute of the center point at little additional computational cost. A greedy matching, based solely on the distance between this predicted offset and the detected center point in the previous frame, suffices for object association. The tracker is end-to-end trainable and differentiable.

Tracking objects as points simplifies two key components of the tracking pipeline. First, it simplifies tracking-conditioned detection. If each object in past frames is represented by a single point, a constellation of objects can be represented by a heatmap of points . Our tracking-conditioned detector directly ingests this heatmap and reasons about all objects jointly when associating them across frames. Second, point-based tracking simplifies object association across time. A simple displacement prediction, akin to sparse optical flow, allows objects in different frames to be linked. This displacement prediction is conditioned on prior detections. It learns to jointly detect objects in the current frame and associate them to prior detections.

While the overall idea is simple, subtle details matter in making this work. Tracked objects in consecutive frames are highly correlated. With the previous-frame heatmap given as input, CenterTrack could easily learn to repeat the predictions from the preceding frame, and thus refuse to track without incurring a large training error. We prevent this through an aggressive data-augmentation scheme during training. In fact, our data augmentation is aggressive enough for the model to learn to track objects from static images. That is, CenterTrack can be successfully trained on static image datasets (with “hallucinated” motion), with no real video input.

CenterTrack is purely local. It only associates objects in adjacent frames, without reinitializing lost long-range tracks. It trades the ability to reconnect long-range tracks for simplicity, speed, and high accuracy in the local regime. Our experiments indicate that this trade-off is well worth it. CenterTrack outperforms complex tracking-by-detection strategies on the MOT and KITTI tracking benchmarks. We further apply the approach to monocular 3D object tracking on the nuScenes dataset . Our monocular tracker achieves 28.3%28.3\% AMOTA@0.2, outperforming the monocular baseline by a factor of 3, while running at 22 FPS. It can be trained on labelled video sequences, if available, or on static images with data augmentation. Code is available at https://github.com/xingyizhou/CenterTrack.

Related work

Most modern trackers follow the tracking-by-detection paradigm. An off-the-shelf object detector first finds all objects in each individual frame. Tracking is then a problem of bounding box association. SORT tracks bounding boxes using a Kalman filter and associates each bounding box with its highest overlapping detection in the current frame using bipartite matching. DeepSORT augments the overlap-based association cost in SORT with appearance features from a deep network. More recent approaches focus on increasing the robustness of object association. Tang et al. leverage person-reidentification features and human pose features. Xu et al. take advantage of the spatial locations over time. BeyondPixel uses additional 3D shape information to track vehicles.

These methods have two drawbacks. First, the data association discards image appearance features or requires a computationally expensive feature extractor . Second, detection is separated from tracking. In our approach, association is almost free. Association is learned jointly with detection. Also, our detector takes the previous tracking results as an input, and can learn to recover missing or occluded objects from this additional cue.

Joint detection and tracking.

A recent trend in multi-object tracking is to convert existing detectors into trackers and combine both tasks in the same framework. Feichtenhofer et al. use a siamese network with the current and past frame as input and predict inter-frame offsets between bounding boxes. Integrated detection uses tracked bounding boxes as additional region proposals to enhance detection, followed by bipartite-matching-based bounding-box association. Tracktor removes the box association by directly propagating identities of region proposals using bounding box regression. In video object detection, Kang et al. feed stacked consecutive frames into the network and do detection for a whole video segment. And Zhu et al. use flow to warp intermediate features from previous frames to accelerate inference.

Our method belongs to this category. The difference is that all of these works adopt the FasterRCNN framework , where the tracked boxes are used as region proposals. This assumes that bounding boxes have a large overlap between frames, which is not true in low-framerate regimes. As a consequence, Tracktor requires a motion model for low-framerate sequences. Our approach instead provides the tracked predictions as an additional point-based heatmap input to the network. The network is then able to reason about and match objects anywhere in its receptive field even if the boxes have no overlap at all.

Motion prediction.

Motion prediction is another important component in a tracking system. Early approaches used Kalman filters to model object velocities. Held et al. use a regression network to predict four scalars for bounding box offset between frames for single-object tracking. Xiao et al. utilize an optical flow estimation network to update joint locations in human pose tracking. Voigtlaender et al. learn a high-dimensional embedding vector for object identities for simultaneous object tracking and segmentation. Our center offset is analogous to sparse optical flow, but is learned together with the detection network and does not require dense supervision.

Heatmap-conditioned keypoint estimation.

Feeding the model predictions as an additional input to a model works across a wide range of vision tasks , especially for keypoint estimation . Auto-context feeds the mask prediction back into the network. Iterative-Error-Feedback (IEF) takes another step by rendering predicted keypoint coordinates into heatmaps. PoseFix generates heatmaps that simulate test errors for human pose refinement.

Our tracking-conditioned detection framework is inspired by these works. A rendered heatmap of prior keypoints is especially appealing in tracking for two reasons. First, the information in the previous frame is freely available and does not slow down the detector. Second, conditional tracking can reason about occluded objects that may no longer be visible in the current frame. The tracker can simply learn to keep those detections from the prior frame around.

D object detection and tracking.

3D trackers replace the object detection component in standard tracking systems with 3D detection from monocular images or 3D point clouds . Tracking then uses an off-the-shelf identity association model. For example, 3DT detects 2D bounding boxes, estimates 3D motion, and uses depth and order cues for matching. AB3D achieves state-of-the-art performance by combining a Kalman filter with accurate 3D detections .

Preliminaries

Given an image with a set of annotated objects {p0,p1,…}\{\mathbf{p}_{0},\mathbf{p}_{1},\ldots\}, CenterNet uses a training objective based on the focal loss :

The Gaussian kernel σi\sigma_{i} is a function of the object size .

The size prediction is only supervised at the center locations. Let si\mathbf{s}_{i} be the bounding box size of the ii-th object at location pi\mathbf{p}_{i}. Size prediction is learned by regression

CenterNet further regresses to a refined center local location using an analogous L1 loss LlocL_{loc}. The overall loss of CenterNet is a weighted sum of all three loss terms: focal loss, size, and local location regression.

Tracking objects as points

We approach tracking from a local perspective. When an object leaves the frame or is occluded and reappears, it is assigned a new identity. We thus treat tracking as the problem of propagating detection identities across consecutive frames, without re-establishing associations across temporal gaps.

There are two main challenges here. The first is finding all objects in every frame – including occluded ones. The second challenge is associating these objects through time. We address both via a single deep network, trained end-to-end. Section 4.1 describes a tracking-conditioned detector that leverages tracked detections from the previous frame to improve detection in the current frame. Section 4.2 then presents a simple offset prediction scheme that is able to link detections through time. Finally, Sections 4.3 and 4.4 show how to train this detector from video or static image data.

As an object detector, CenterNet already infers most of the required information for tracking: object locations p^\hat{\mathbf{p}}, their size s^=S^p^\hat{\mathbf{s}}=\hat{S}_{\hat{\mathbf{p}}}, and a confidence measure w^=Y^p^\hat{w}=\hat{Y}_{\hat{\mathbf{p}}}. However, it is unable to find objects that are not directly visible in the current frame, and the detected objects may not be temporally coherent. One natural way to increase temporal coherence is to provide the detector with additional image inputs from past frames. In CenterTrack, we provide the detection network with two frames as input: the current frame I(t)I^{(t)} and the prior frame I(t−1)I^{(t-1)}. This allows the network to estimate the change in the scene and potentially recover occluded objects at time tt from visual evidence at time t−1t-1.

CenterTrack also takes prior detections {p0(t−1),p1(t−1),…}\{\mathbf{p}^{(t-1)}_{0},\mathbf{p}^{(t-1)}_{1},\ldots\} as additional input. How should these detections be represented in a form that is easily provided to a network? The point-based nature of our tracklets is helpful here. Since each detected object is represented by a single point, we can conveniently render all detections in a class-agnostic single-channel heatmap H(t−1)=R({p0(t−1),p1(t−1),…})H^{(t-1)}=\mathcal{R}(\{\mathbf{p}^{(t-1)}_{0},\mathbf{p}^{(t-1)}_{1},\ldots\}), using the same Gaussian render function as in the training of point-based detectors. To reduce the propagation of false positive detections, we only render objects with a confidence score greater than a threshold τ\tau. The architecture of CenterTrack is essentially identical to CenterNet, with four additional input channels. (See Figure 2.)

Tracking-conditioned detection provides a temporally coherent set of detected objects. However, it does not link these detections across time. In the next section, we show how to add one additional output to point-based detection to track objects through space and time.

2 Association through offsets

where pi(t−1)\mathbf{p}_{i}^{(t-1)} and pi(t)\mathbf{p}_{i}^{(t)} are tracked ground-truth objects. Figure 2 shows an example of this offset prediction.

With a sufficiently good offset prediction, a simple greedy matching algorithm can associate objects across time. For each detection at position p^\hat{p}, we greedily associate it with the closest unmatched prior detection at position p^−D^p^\hat{p}-\hat{D}_{\hat{p}}, in descending order of confidence w^\hat{w}. If there is no unmatched prior detection within a radius κ\kappa, we spawn a new tracklet. We define κ\kappa as the geometric mean of the width and height of the predicted bounding box for each tracklet. A precise description of this greedy matching algorithm is provided in supplementary material. The simplicity of this greedy matching algorithm again highlights the advantages of tracking objects as points. A simple displacement prediction is sufficient to link objects across time. There is no need for a complicated distance metric or graph matching.

3 Training on video data

CenterTrack is first and foremost an object detector, and trained as such. The architectural changed from CenterNet to CenterTrack are minor: four additional input channels and two output channels. This allows us to fine-tune CenterTrack directly from a pretrained CenterNet detector . We copy all weights related to the current detection pipeline. All weights corresponding to additional inputs or outputs are initialized randomly. We follow the CenterNet training protocol and train all predictions as multi-task learning. We use the same training objective with the addition of offset regression LoffL_{off}.

The main challenge in training CenterTrack comes in producing a realistic tracklet heatmap H(t−1)H^{(t-1)}. At inference time, this tracklet heatmap can contain an arbitrary number of missing tracklets, wrongly localized objects, or even false positives. These errors are not present in ground-truth tracklets {p0(t−1),p1(t−1),…}\{\mathbf{p}^{(t-1)}_{0},\mathbf{p}^{(t-1)}_{1},\ldots\} provided during training. We instead simulate this test-time error during training. Specifically, we simulate three types of error. First, we locally jitter each tracklet p(t−1)\mathbf{p}^{(t-1)} from the prior frame by adding Gaussian noise to each center. That is, we render pi′=(xi+r×λjt×wi,yi+r×λjt×hi)p_{i}^{\prime}=(x_{i}+r\times\lambda_{jt}\times w_{i},y_{i}+r\times\lambda_{jt}\times h_{i}), where r is sampled from a Gaussian distribution. We use λjt=0.05\lambda_{jt}=0.05 in all experiments. Second, we randomly add false positives near ground-truth object locations by rendering a spurious noisy peak pi′p_{i}^{\prime} with probability λfp\lambda_{fp}. Third, we simulate false negatives by randomly removing detections with probability λfn\lambda_{fn}. λfp\lambda_{fp} and λfn\lambda_{fn} are set according to the statistics of our baseline model. These three augmentations are sufficient to train a robust tracking-conditioned object detector.

In practice, I(t−1)I^{(t-1)} does not need to be the immediately preceding frame from time t−1t-1. It can be a different frame from the same video sequence. In our experiments, we randomly sample frames near tt to avoid overfitting to the framerate. Specifically, we sample from all frames kk where ∣k−t∣<Mf|k-t|<M_{f}, where Mf=3M_{f}=3 is a hyperparameter.

4 Training on static image data

Without labeled video data, CenterTrack does not have access to a prior frame I(t−1)I^{(t-1)} or tracked detections {p0(t−1),p1(t−1),…}\{\mathbf{p}^{(t-1)}_{0},\mathbf{p}^{(t-1)}_{1},\ldots\}. However, we can simulate tracking on standard detection benchmarks, given only single images I(t)I^{(t)} and detections {p0(t),p1(t),…}\{\mathbf{p}^{(t)}_{0},\mathbf{p}^{(t)}_{1},\ldots\}. The idea is simple: we simulate the previous frame by randomly scaling and translating the current frame. As our experiments will demonstrate, this is surprisingly effective.

5 End-to-end 3D object tracking

To perform monocular 3D tracking, we adopt the monocular 3D detection form of CenterNet . Specifically, we train output heads to predict object depth, rotation (encoded as an 88-dimensional vector ), and 3D extent. Since the projection of the center of the 3D bounding box may not align with the center of the object’s 2D bounding box (due to perspective projection), we also predict a 2D-to-3D center offset. Further details are provided in the supplement.

Experiments

We evaluate 2D multi-object tracking on the MOT17 and KITTI tracking benchmarks. We also evaluate monocular 3D tracking on the nuScenes dataset . Experiments on MOT16 can be found in the supplement.

MOT17 contains 7 training sequences and 7 test sequences , The videos were captured by stationary cameras mounted in high-density scenes with heavy occlusion. Only pedestrians are annotated and evaluated. The video framerate is 25-30 FPS. The MOT dataset does not provide an official validation split. For ablation experiments, we split each training sequence into two halves, and use the first half frames for training and the second for validation. Our main results are reported on the test set.

KITTI.

The KITTI tracking benchmark consists of 21 training sequences and 29 test sequences . They are collected by a camera mounted on a car moving through traffic. The dataset provides 2D bounding box annotations for cars, pedestrians, and cyclists, but only cars are evaluated. Videos are captured at 10 FPS and contain large inter-frame motions. KITTI does not provide detections, and all entries use private detection. We again split all training sequences into halves for training and validation.

nuScenes.

nuScenes is a newly released large-scale driving dataset with 7 object classes annotated for tracking . It contains 700 training sequences, 150 validation sequences, and 150 test sequences. Each sequence contains roughly 40 frames at 2 FPS with 6 slightly overlapping images in a panoramic 360°360\degree view, resulting in 168k training, 36k validation, and 36k test images. The videos are sampled at 12 FPS, but frames are only annotated and evaluated at 2 FPS. All baselines and CenterTrack only use keyframes for training and evaluation. Due to the low framerate, the inter-frame motion is significant.

Evaluation metrics.

We use the official evaluation metrics in each dataset. The common metric is multi-object tracking accuracy : MOTA=1−∑t(FPt+FNt+IDSWt)∑tGTtMOTA=1-\frac{\sum_{t}(FP_{t}+FN_{t}+IDSW_{t})}{\sum_{t}GT_{t}}, where GTtGT_{t}, FPtFP_{t}, FNtFN_{t}, and IDSWtIDSW_{t} are the number of ground-truth bounding boxes, false positives, false negatives, and identity switches in frame tt, respectively. MOTA does not rank tracklets according to confidence and is sensitive to the task-dependent output threshold θ\theta . The thresholds we use are listed in Section 5.2. The interplay between output threshold and true positive criteria matters. For 2D tracking , >0.5>0.5 bounding box IoU is a the true positive. For 3D tracking , bounding box center distance <2m<2m on the ground plane is the criterion for a true positive. When objects are successfully detected, but not tracked, they are identified as an identity switch (IDSW). The IDF1 metric measures the minimal cost change from predicted ids to the correct ids. In our ablation studies, we report false positve rate (FP) ∑tFPt∑tGTt\frac{\sum_{t}FP_{t}}{\sum_{t}GT_{t}}, false negative rate (FN) ∑tFNt∑tGTt\frac{\sum_{t}FN_{t}}{\sum_{t}GT_{t}}, and identity switches (IDSW) ∑tIDSWt∑tGTt\frac{\sum_{t}IDSW_{t}}{\sum_{t}GT_{t}} separately. In comparisons with other methods, we report the absolute numbers following the dataset convention . We also report the Most Tracked ratio (MT) for the ratio of most tracked (>80%>80\% time) objects and Most Lost ratio (ML) for most lost (<20%<20\% time) objects .

nuScenes adopts a more robust metric, AMOTA, which is a weighted average of MOTA across different output thresholds. Specifically,

where rr is a fixed recall threshold, P=∑tGTtP=\sum_{t}{GT_{t}} is the total number of annotated objects among all frames, and FPr=∑tFPr,tFP_{r}=\sum_{t}{FP_{r,t}} is the total number of false positive samples only considering the top confident samples that achieve the recall threshold rr. The hyperparameters n=40n=40 and α=0.2\alpha=0.2 (AMOTA@0.2), or α=1\alpha=1 (AMOTA@1) are set by the benchmark organizers. The overall AMOTA is the average AMOTA among all 7 categories.

2 Implementation details

Our implementation is based on CenterNet . We use DLA as the network backbone, optimized with Adam with learning rate 1.25e−41.25e-4 and batchsize 32. Data augmentations include random horizontal flipping, random resized cropping, and color jittering. For all experiments, we train the networks for 7070 epochs. The learning rate is dropped by a factor of 10 at the 60th epoch. We test the runtime on a machine with an Intel Core i7-8086K CPU and a Titan Xp GPU. The runtimes depend on the number of objects for rendering and the input resolution in each dataset.

The MOT dataset annotates each pedestrian as an amodal bounding box. That is, the bounding box always covers the whole body even when part of the object is out of the frame. In contrast, CenterNet requires the center of each inferred bounding box to be within the frame. To handle this, we separately predict the visible and amodal bounding boxes . Further details on this can be found in the supplement. We follow prior works to pretrain on external data. We train our network on the CrowdHuman dataset, using the static image training described in Section 4.4. Details on the CrowdHuman dataset and ablations of pretraining are in the supplement.

The default input resolution for MOT images is 1920×1080{1920\times 1080}. We resize and pad the images to 960×544960\times 544. We use random false positive ratio λfp=0.1\lambda_{fp}=0.1 and random false negative ratio λfn=0.4\lambda_{fn}=0.4. We only output tracklets that have a confidence of θ=0.4\theta=0.4 or higher, and set the heatmap rendering threshold to τ=0.5\tau=0.5. A controlled study of these hyperparameters is in the supplement.

For KITTI , we keep the original input resolution 1280×3841280\times 384 in training and testing. The hyperparameters are set at λfp=0.1\lambda_{fp}=0.1 and λfn=0.2\lambda_{fn}=0.2, with output threshold θ=0.4\theta=0.4 and rendering threshold τ=0.4\tau=0.4. We fine-tune our KITTI model from a nuScenes tracking model.

For nuScenes , we use input resolution 800×448{800\times 448}. We set λfp=0.1{\lambda_{fp}=0.1} and λfn=0.4{\lambda_{fn}=0.4}, and use output threshold θ=0.1{\theta=0.1} and rendering threshold τ=0.1{\tau=0.1}. We first train our nuScenes model for 140 epochs for just 3D detection and then fine-tune for 70 epochs for 3D tracking. Note that nuScenes evaluation is done per 360 panorama, not per image. We naively fuse all outputs from the 6 cameras together, without handling duplicate detections at the intersection of views .

Following common practice , we keep unmatched tracks “inactive” until they remain undetected for KK consecutive frames. Inactive tracks can be matched to detections and regain their ID, but not appear in the prior heatmap or output. The tracker stays online. Rebirth only matters for the MOT test set, where we use K=32K=32. For all other experiments, we found rebirth not to be required (K=0K=0).

3 Public detection

The MOT17 challenge only supports public detection. That is, participants are asked to use the provided detections. Public detection is meant to test a tracker’s ability to associate objects, irrespective of its ability to detect objects. Our method operates in the private detection mode by default. For the MOT challenge we created a public-detection version of CenterTrack that uses the externally provided (public) detections and is thus fairly compared to other participants in the challenge. This shows that the advantages of CenterTrack are not due to the accuracy of the detections but are due to the tracking framework itself.

Note that refining and rescoring the given bounding boxes is allowed and is commonly used by participants in the challenge . Following Tracktor , we keep the bounding boxes that are close to an existing bounding box in the previous frame. We only initialize a new trajectory if it is near a public detection. All bounding boxes in our results are either near a public detection in the current frame or near a tracked box in the previous frame. The algorithm’s diagram of this public-detection configuration can be found in the supplement. We use this public-detection configuration of CenterTrack for MOT17 test set evaluation and use the private-detection setting in our ablation studies.

4 Main results

All three datasets – MOT17 , KITTI , and nuScenes – host test servers with hidden annotations and leaderboards. We compare to all published results on these leaderboards. The numbers were accessed on Mar. 5th, 2020. We retrain CenterTrack on the full training set with the same hyperparameters in the ablation experiments.

Table 1 lists the results on the MOT17 challenge. We use our public configuration in Section 5.3 and do not pretrain on CrowdHuman . CenterTrack significantly outperforms the prior state of the art even when restricted to the public-detection configuration. For example CenterTrack improves MOTA by 5 points (an 8.6% relative improvement) over Tracktor v2 .

The public detection setting ensures that all methods build on the same underlying detector. Our gains come from two sources. Firstly, the heatmap input makes our tracker better preserve tracklets from the previous frame, which results in a much lower rate of false negatives. And second, our simple learned offset is effective. (See Section 5.6 for more analysis.) For reference, we also included a private detection version, where CenterTrack simultaneously detects and tracks objects (Table 1, bottom). It further improves the MOTA to 67.3%67.3\%, and runs at 17 FPS end-to-end (including detection).

For IDF1 and id-switch, our local model is not as strong as offline methods such as LSST17 , but is better than other online methods . We believe that there is an exciting avenue for future work in combining local trackers (such as our work) with stronger offline long-range models (such as SORT , LMP , and other ReID-based trackers ).

On KITTI , we submitted our best-performing model with flip testing . The model runs at 8282ms and yields 89.44%89.44\% MOTA, outperforming all published work (Table 2). Note that our model without flip testing runs at 4545ms with 88.7%88.7\% MOTA on the validation set (vs. 89.63%89.63\% with flip testing on the validation set). We avoid submitting to the test server multiple times following their test policy. The results again indicate that CenterTrack performs competitively with more complex methods.

On nuScenes , our monocular tracking method achieves an AMOTA@0.2 of 28.3%28.3\% and an AMOTA@1 of 4.6%4.6\%, outperforming the monocular baseline by a large margin. There are two main reasons. Firstly, we use a stronger and faster 3D detector (see the 3D detector comparison in the supplementary). More importantly, as shown in Table 6, the Kalman-filter-based 3D tracking baseline relies on hand-crafted motion rules , which are less effective in low-framerate regimes. Our method learns object motion from data and is much more stable at low framerates.

5 Ablation studies

We first ablate our two main technical contributions: tracking-conditioned detection (Section 4.1) and offset prediction (Section 4.2) on all three datasets. Specifically, we compare our full framework with three baselines.

runs a CenterNet detector at each individual frame and associates their identity only based on 2D center distance. This model does not use video data, but still uses two input images.

Without offset

uses just tracking-conditioned prediction with a predicted offset of zero. Every object is again associated to its closest object in the previous frame.

Without heatmap

predicts the center offset between frames and uses the updated center distance as the association metric, but the prior heatmap is not provided. The offset-based greedy association is used.

Table 4 shows the results. On all datasets, our full CenterTrack model performs significantly better than the baselines. Tracking-conditioned detection yields ∼2%\sim 2\% MOTA improvement on MOT and ∼3%\sim 3\% MOTA improvement on KITTI, with or without offset prediction. It produces more false positives but fewer false negatives. This is because with the heatmap prior, the network tends to predict more objects around the previous peaks, which are sometimes misleading. The merits of the heatmap outweigh the limitations and improve MOTA overall. Using the prior heatmap also significantly reduces IDSW on both datasets, indicating that the heatmap stabilizes detection.

Tracking offset prediction gives a huge boost on nuScenes and reduces IDSW consistently in MOT and KITTI. The effectiveness of the tracking offset appears to be related to the video framerate. When the framerate is high, motion between frames is small, and a zero offset is often a reasonable starting point for association. When framerate is low, as in the nuScenes dataset, motion between frames is large and static object association is considerably less effective. Our offset prediction scheme helps deal with such large inter-frame motion. Next, we ablate other components on MOT17.

Training with noisy heatmap.

The 2nd row in Table 5 shows the importance of injecting noise into heatmaps during training (Section 4.3). Without noise injection, the model fails to generalize and yields dramatically lower accuracy. In particular, this model has a large false negative rate. One reason is that in the first frame, the input heatmap is empty. This model had a hard time discovering new objects that were not indicated in the prior heatmap.

Training on static images.

We train a version of our model on static images only, as described in Section 4.4. The results are shown in Table 5 (3rd row, ‘Static image’). As reported in this table, training on static images gives the same performance as training on videos on the MOT dataset. Separately, we observed that training on static images is less effective on nuScenes, where framerate is low.

Matching algorithm.

We use a simple greedy matching algorithm based on the detection score, while most other trackers use the Hungarian algorithm. We show the performance of CenterTrack with Hungarian matching in the 4th row of Table 5. It does not improve performance. We choose greedy matching for simplicity.

Track rebirth.

We show CenterTrack with track rebirth (KK=3232) in the last row of Table 5. While the MOTA performance keeps similar, it significantly increases IDF1 and reduces ID switch. We use this setting for our MOT test set submission. For other datasets and evaluation metrics no rebirth was required (K=0K=0).

6 Comparison to alternative motion models

Our offset prediction is able to estimate object motion, but also performs a simple association, as current objects are linked to prior detections, which CenterTrack receives as one of its inputs. To verify the effectiveness of our learned association, we replace our offset prediction with three alternative motion models:

We set the offset to zeros. It is copied from Table 4 for reference only.

Kalman filter.

The Kalman filter predicts each object’s future state through an explicit motion model estimated from its history. It is the most widely used motion model in traditional real-time trackers . We use the popular public implementation from SORT .

Optical flow.

As an alternative motion model, we use FlowNet2 . The model was trained to estimate dense pixel motion for all objects in a scene. We run the strongest officially released FlowNet2 model (∼150\sim 150ms / image pair), and replace our learned offset with the predicted optical flow at each predicted object center.

The results are shown in Table 6. All models use the exact same detector. On the high-framerate MOT17 dataset, any motion model suffices, and even no motion model at all performs competitively. On KITTI and nuScenes, where the intra-frame motions are non-trivial, the hand-crafted motion rule of the Kalman filter performs significantly worse, and even the performance of optical flow degrades. This emphasizes that our offset model does more than just motion estimation. CenterTrack is conditioned on prior detections and can learn to snap offset predictions to exactly those prior detections. Our training procedure strongly encourages this through heavy data augmentation.

Conclusion

We presented an end-to-end simultaneous object detection and tracking framework. Our method takes two frames and a prior heatmap as input, and produces detection and tracking offsets for the current frame. Our tracker is purely local and associates objects greedily through time. It runs online (no knowledge of future frames) and in real time, and sets a new state of the art on the challenging MOT17, KITTI, and nuScenes 3D tracking benchmarks.

This work has been supported in part by the National Science Foundation under grant IIS-1845485.

References

Appendix 0.A Tracking algorithms

We adopt a simple greedy id association algorithm based on the center distance, shown in Algorithm 1. We use the same algorithm for both 2D tracking and 3D tracking.

A.2 Public tracking

For public tracking, we follow Tractor to extend a private tracking algorithm to public detection. The id association is exactly the same as private detection (Line 1 to Line 1). The difference lies in how a track can be created. In public detection, we only initialize a track if it is near a provided bounding box (Line 2 to Line 2).

Appendix 0.B Results on MOT16

MOT16 shares the same training and testing sequences with MOT17, but officially supports private detection. As is shown in Table 7, we rank 2nd among all published entries. We remark that all other entries use a heavy detector trained on private data and many rely on slow matching schemes . For example, LMP_p computes person-reidentification features for all pairs of bounding boxes using a Siamese network, requiring O(n2)O(n^{2}) forward passes through a deep network. In contrast, CenterTrack involves a single pass through a network and operates online at 17 FPS.

Appendix 0.C 3D detection

We follow CenterNet to regress to object depth D^∈RWR×HR\hat{D}\in R^{\frac{W}{R}\times\frac{H}{R}}, 3d extent Γ^∈RWR×HR×3\hat{\Gamma}\in R^{\frac{W}{R}\times\frac{H}{R}\times 3}, orientation (encoded as an 8-dimension vector) A^∈RWR×HR×8\hat{A}\in R^{\frac{W}{R}\times\frac{H}{R}\times 8}. The training loss for these are identical to CenterNet . Since the 2D bounding box center does not align with the projected 3D bounding box center due to perspective projection, we in addition regress to an offset from the 2D center to the projected 3D bounding box centerF^∈RWR×HR×2\hat{F}\in R^{\frac{W}{R}\times\frac{H}{R}\times 2}. We use L1Loss:

where fk∈R2f_{k}\in\mathcal{R}^{2} is the ground truth offset of object kk, and f^k=F^pk\hat{f}_{k}=\hat{F}_{\mathbf{p}_{k}} is the value in F^\hat{F} at location pk\mathbf{p}_{k}.

We show the 3D detection performance of CenterNet with the offset prediction in Table 8 for reference. The 3D detection performance is on-par with Mappilary and PointPillars , but far below the LiDAR based state-of-the-art Megvii .

Appendix 0.D Amodal bounding box regression

Appendix 0.E CrowdHuman dataset

CrowdHuman contains 15k training images with common pose annotations. The dataset is featured of high density and large occlusion. Both visible bounding box and the Amodal bounding box are annotated. We use the Amodal bounding box annotation in our experiments to align with MOT .

Appendix 0.F Pretraining experiments

For pretraining on CrowdHuman , we use input resolution 512×512512\times 512, false positive ratio λfp=0.1\lambda_{fp}=0.1, false negative ratio λfn=0.4\lambda_{fn}=0.4, random scaling ratio 0.050.05, and random translation ratio 0.050.05. The training follows Section.4.4 of the main paper. As shown in Table 9, the model trained on CrowdHuman achieves a decent 52.252.2 MOTA in MOT dataset, without seeing any MOT data.

Without CrowdHuman pretraining, our performance drops to 60.7%60.7\% MOTA on the validation set. Pretraining help improve detection quality by decreasing the false negatives. Note that most entries on MOT challenges use external data for pretraining, and some of them use private data . For reference, we also show our public detection results without pretraining in Table 9, last row. This model corresponds to the entry we submitted to MOT17 public detection challenge.

Appendix 0.G Additional experiments on KITTI

In Table 10, we show results of the same additional experiments (Section. 5.5 of the main paper) on KITTI dataset . The conclusions are the same as on MOT . Training on static images now performs slightly worse than training on video, mostly due to that KITTI has larger inter-frame motion than MOT. Training without random heatmap noise is much worse than the full model, with a high false-negative rate. And using the Hungarian algorithm works the same as using a greedy matching. Our model without nuScenes achieves 84.5%84.5\% MOTA on the validation set, this is on-par with other state-of-the-art trackers on KITTI with a heavy detector .

Appendix 0.H Output and rendering threshold

As the tracking evaluation metric (MOTA) does not consider the confidence of predictions, picking an output threshold is essential in all tracking algorithms (see discussion in AB3D ). In our case, we also need a threshold to render predictions to the prior heatmap. We search the optimal thresholds on MOT in Table 11. Basically, increasing both thresholds results in fewer outputs, thus increases the false negatives while decreases the false positives. We find a good balance at θ=0.4\theta=0.4 and τ=0.5\tau=0.5.