Objects as Points

Xingyi Zhou, Dequan Wang, Philipp Krähenbühl

Introduction

Object detection powers many vision tasks like instance segmentation , pose estimation , tracking , and action recognition . It has down-stream applications in surveillance , autonomous driving , and visual question answering . Current object detectors represent each object through an axis-aligned bounding box that tightly encompasses the object . They then reduce object detection to image classification of an extensive number of potential object bounding boxes. For each bounding box, the classifier determines if the image content is a specific object or background. One-stage detectors slide a complex arrangement of possible bounding boxes, called anchors, over the image and classify them directly without specifying the box content. Two-stage detectors recompute image-features for each potential box, then classify those features. Post-processing, namely non-maxima suppression, then removes duplicated detections for the same instance by computing bounding box IoU. This post-processing is hard to differentiate and train , hence most current detectors are not end-to-end trainable. Nonetheless, over the past five years , this idea has achieved good empirical success . Sliding window based object detectors are however a bit wasteful, as they need to enumerate all possible object locations and dimensions.

In this paper, we provide a much simpler and more efficient alternative. We represent objects by a single point at their bounding box center (see Figure 2). Other properties, such as object size, dimension, 3D extent, orientation, and pose are then regressed directly from image features at the center location. Object detection is then a standard keypoint estimation problem . We simply feed the input image to a fully convolutional network that generates a heatmap. Peaks in this heatmap correspond to object centers. Image features at each peak predict the objects bounding box height and weight. The model trains using standard dense supervised learning . Inference is a single network forward-pass, without non-maximal suppression for post-processing.

Our method is general and can be extended to other tasks with minor effort. We provide experiments on 3D object detection and multi-person human pose estimation , by predicting additional outputs at each center point (see Figure 4). For 3D bounding box estimation, we regress to the object absolute depth, 3D bounding box dimensions, and object orientation . For human pose estimation, we consider the 2D joint locations as offsets from the center and directly regress to them at the center point location.

The simplicity of our method, CenterNet, allows it to run at a very high speed (Figure 1). With a simple Resnet-18 and up-convolutional layers , our network runs at 142 FPS with 28.1%28.1\% COCO bounding box AP. With a carefully designed keypoint detection network, DLA-34 , our network achieves 37.4%37.4\% COCO AP at 52 FPS. Equipped with the state-of-the-art keypoint estimation network, Hourglass-104 , and multi-scale testing, our network achieves 45.1%45.1\% COCO AP at 1.4 FPS. On 3D bounding box estimation and human pose estimation, we perform competitively with state-of-the-art at a higher inference speed. Code is available at https://github.com/xingyizhou/CenterNet.

Related work

One of the first successful deep object detectors, RCNN , enumerates object location from a large set of region candidates , crops them, and classifies each using a deep network. Fast-RCNN crops image features instead, to save computation. However, both methods rely on slow low-level region proposal methods.

Faster RCNN generates region proposal within the detection network. It samples fixed-shape bounding boxes (anchors) around a low-resolution image grid and classifies each into “foreground or not”. An anchor is labeled foreground with a > ⁣ ⁣ ⁣0.7>\!\!\!0.7 overlap with any ground truth object, background with a < ⁣ ⁣ ⁣0.3<\!\!\!0.3 overlap, or ignored otherwise. Each generated region proposal is again classified . Changing the proposal classifier to a multi-class classification forms the basis of one-stage detectors. Several improvements to one-stage detectors include anchor shape priors , different feature resolution , and loss re-weighting among different samples .

Our approach is closely related to anchor-based one-stage approaches . A center point can be seen as a single shape-agnostic anchor (see Figure 3). However, there are a few important differences. First, our CenterNet assigns the “anchor” based solely on location, not box overlap . We have no manual thresholds for foreground and background classification. Second, we only have one positive “anchor” per object, and hence do not need Non-Maximum Suppression (NMS) . We simply extract local peaks in the keypoint heatmap . Third, CenterNet uses a larger output resolution (output stride of 44) compared to traditional object detectors (output stride of 1616). This eliminates the need for multiple anchors .

We are not the first to use keypoint estimation for object detection. CornerNet detects two bounding box corners as keypoints, while ExtremeNet detects the top-, left-, bottom-, right-most, and center points of all objects. Both these methods build on the same robust keypoint estimation network as our CenterNet. However, they require a combinatorial grouping stage after keypoint detection, which significantly slows down each algorithm. Our CenterNet, on the other hand, simply extracts a single center point per object without the need for grouping or post-processing.

3D bounding box estimation powers autonomous driving . Deep3Dbox uses a slow-RCNN style framework, by first detecting 2D objects and then feeding each object into a 3D estimation network. 3D RCNN adds an additional head to Faster-RCNN followed by a 3D projection. Deep Manta uses a coarse-to-fine Faster-RCNN trained on many tasks. Our method is similar to a one-stage version of Deep3Dbox or 3DRCNN . As such, CenterNet is much simpler and faster than competing methods.

Preliminary

Let I∈RW×H×3I\in R^{W\times H\times 3} be an input image of width WW and height HH. Our aim is to produce a keypoint heatmap Y^∈WR×HR×C\hat{Y}\in^{\frac{W}{R}\times\frac{H}{R}\times C}, where RR is the output stride and CC is the number of keypoint types. Keypoint types include C=17C=17 human joints in human pose estimation , or C=80C=80 object categories in object detection . We use the default output stride of R=4R=4 in literature . The output stride downsamples the output prediction by a factor RR. A prediction Y^x,y,c=1\hat{Y}_{x,y,c}=1 corresponds to a detected keypoint, while Y^x,y,c=0\hat{Y}_{x,y,c}=0 is background. We use several different fully-convolutional encoder-decoder networks to predict Y^\hat{Y} from an image II: A stacked hourglass network , up-convolutional residual networks (ResNet) , and deep layer aggregation (DLA) .

where α\alpha and β\beta are hyper-parameters of the focal loss , and NN is the number of keypoints in image II. The normalization by NN is chosen as to normalize all positive focal loss instances to 11. We use α=2\alpha=2 and β=4\beta=4 in all our experiments, following Law and Deng .

To recover the discretization error caused by the output stride, we additionally predict a local offset O^∈RWR×HR×2\hat{O}\in\mathcal{R}^{\frac{W}{R}\times\frac{H}{R}\times 2} for each center point. All classes cc share the same offset prediction. The offset is trained with an L1 loss

In the next section, we will show how to extend this keypoint estimator to a general purpose object detector.

Objects as Points

Let (x1(k),y1(k),x2(k),y2(k))(x_{1}^{(k)},y_{1}^{(k)},x_{2}^{(k)},y_{2}^{(k)}) be the bounding box of object kk with category ckc_{k}. Its center point is lies at pk=(x1(k)+x2(k)2,y1(k)+y2(k)2)p_{k}=(\frac{x_{1}^{(k)}+x_{2}^{(k)}}{2},\frac{y_{1}^{(k)}+y_{2}^{(k)}}{2}). We use our keypoint estimator Y^\hat{Y} to predict all center points. In addition, we regress to the object size sk=(x2(k)−x1(k),y2(k)−y1(k))s_{k}=(x_{2}^{(k)}-x_{1}^{(k)},y_{2}^{(k)}-y_{1}^{(k)}) for each object kk. To limit the computational burden, we use a single size prediction S^∈RWR×HR×2\hat{S}\in\mathcal{R}^{\frac{W}{R}\times\frac{H}{R}\times 2} for all object categories. We use an L1 loss at the center point similar to Objective 2:

We do not normalize the scale and directly use the raw pixel coordinates. We instead scale the loss by a constant λsize\lambda_{size}. The overall training objective is

We set λsize=0.1\lambda_{size}=0.1 and λoff=1\lambda_{off}=1 in all our experiments unless specified otherwise. We use a single network to predict the keypoints Y^\hat{Y}, offset O^\hat{O}, and size S^\hat{S}. The network predicts a total of C+4C+4 outputs at each location. All outputs share a common fully-convolutional backbone network. For each modality, the features of the backbone are then passed through a separate 3×33\times 3 convolution, ReLU and another 1×11\times 1 convolution. Figure 4 shows an overview of the network output. Section 5 and supplementary material contain additional architectural details.

At inference time, we first extract the peaks in the heatmap for each category independently. We detect all responses whose value is greater or equal to its 8-connected neighbors and keep the top 100100 peaks. Let P^c\hat{\mathcal{P}}_{c} be the set of nn detected center points P^={(x^i,y^i)}i=1n\hat{\mathcal{P}}=\{(\hat{x}_{i},\hat{y}_{i})\}_{i=1}^{n} of class cc. Each keypoint location is given by an integer coordinates (xi,yi)(x_{i},y_{i}). We use the keypoint values Y^xiyic\hat{Y}_{x_{i}y_{i}c} as a measure of its detection confidence, and produce a bounding box at location

where (δx^i,δy^i)=O^x^i,y^i(\delta\hat{x}_{i},\delta\hat{y}_{i})=\hat{O}_{\hat{x}_{i},\hat{y}_{i}} is the offset prediction and (w^i,h^i)=S^x^i,y^i(\hat{w}_{i},\hat{h}_{i})=\hat{S}_{\hat{x}_{i},\hat{y}_{i}} is the size prediction. All outputs are produced directly from the keypoint estimation without the need for IoU-based non-maxima suppression (NMS) or other post-processing. The peak keypoint extraction serves as a sufficient NMS alternative and can be implemented efficiently on device using a 3×33\times 3 max pooling operation.

1 3D detection

3D detection estimates a three-dimensional bounding box per objects and requires three additional attributes per center point: depth, 3D dimension, and orientation. We add a separate head for each of them. The depth dd is a single scalar per center point. However, depth is difficult to regress to directly. We instead use the output transformation of Eigen et al. and d=1/σ(d^)−1d=1/\sigma(\hat{d})-1, where σ\sigma is the sigmoid function. We compute the depth as an additional output channel D^∈WR×HR\hat{D}\in^{\frac{W}{R}\times\frac{H}{R}} of our keypoint estimator. It again uses two convolutional layers separated by a ReLU. Unlike previous modalities, it uses the inverse sigmoidal transformation at the output layer. We train the depth estimator using an L1 loss in the original depth domain, after the sigmoidal transformation.

The 3D dimensions of an object are three scalars. We directly regress to their absolute values in meters using a separate head Γ^∈RWR×HR×3\hat{\Gamma}\in\mathcal{R}^{\frac{W}{R}\times\frac{H}{R}\times 3} and an L1 loss.

Orientation is a single scalar by default. However, it can be hard to regress to. We follow Mousavian et al. and represent the orientation as two bins with in-bin regression. Specifically, the orientation is encoded using 88 scalars, with 44 scalars for each bin. For one bin, two scalars are used for softmax classification and the rest two scalar regress to an angle within each bin. Please see the supplementary for details about these losses.

2 Human pose estimation

Human pose estimation aims to estimate kk 2D human joint locations for every human instance in the image (k=17k=17 for COCO). We considered the pose as a k×2k\times 2-dimensional property of the center point, and parametrize each keypoint by an offset to the center point. We directly regress to the joint offsets (in pixels) J^∈RWR×HR×k×2\hat{J}\in\mathcal{R}^{\frac{W}{R}\times\frac{H}{R}\times k\times 2} with an L1 loss. We ignore the invisible keypoints by masking the loss. This results in a regression-based one-stage multi-person human pose estimator similar to the slow-RCNN version counterparts Toshev et al. and Sun et al. .

To refine the keypoints, we further estimate kk human joint heatmaps Φ^∈RWR×HR×k\hat{\Phi}\in\mathcal{R}^{\frac{W}{R}\times\frac{H}{R}\times k} using standard bottom-up multi-human pose estimation . We train the human joint heatmap with focal loss and local pixel offset analogous to the center detection discussed in Section. 3.

Implementation details

We experiment with 4 architectures: ResNet-18, ResNet-101 , DLA-34 , and Hourglass-104 . We modify both ResNets and DLA-34 using deformable convolution layers and use the Hourglass network as is.

The stacked Hourglass Network downsamples the input by 4×4\times, followed by two sequential hourglass modules. Each hourglass module is a symmetric 5-layer down- and up-convolutional network with skip connections. This network is quite large, but generally yields the best keypoint estimation performance.

Xiao et al. augment a standard residual network with three up-convolutional networks to allow for a higher-resolution output (output stride 44). We first change the channels of the three upsampling layers to 256,128,64256,128,64, respectively, to save computation. We then add one 3×33\times 3 deformable convolutional layer before each up-convolution with channel 256,128,64256,128,64, respectively. The up-convolutional kernels are initialized as bilinear interpolation. See supplement for a detailed architecture diagram.

Deep Layer Aggregation (DLA) is an image classification network with hierarchical skip connections. We utilize the fully convolutional upsampling version of DLA for dense prediction, which uses iterative deep aggregation to increase feature map resolution symmetrically. We augment the skip connections with deformable convolution from lower layers to the output. Specifically, we replace the original convolution with 3×33\times 3 deformable convolution at every upsampling layer. See supplement for a detailed architecture diagram.

We add one 3×33\times 3 convolutional layer with 256256 channel before each output head. A final 1×11\times 1 convolution then produces the desired output. We provide more details in the supplementary material.

We train on an input resolution of 512×512512\times 512. This yields an output resolution of 128×128128\times 128 for all the models. We use random flip, random scaling (between 0.6 to 1.3), cropping, and color jittering as data augmentation, and use Adam to optimize the overall objective. We use no augmentation to train the 3D estimation branch, as cropping or scaling changes the 3D measurements. For the residual networks and DLA-34, we train with a batch-size of 128 (on 8 GPUs) and learning rate 5e-4 for 140 epochs, with learning rate dropped 10×10\times at 90 and 120 epochs, respectively (following ). For Hourglass-104, we follow ExtremeNet and use batch-size 29 (on 5 GPUs, with master GPU batch-size 4) and learning rate 2.5e-4 for 50 epochs with 10×10\times learning rate dropped at the 40 epoch. For detection, we fine-tune the Hourglass-104 from ExtremeNet to save computation. The down-sampling layers of Resnet-101 and DLA-34 are initialized with ImageNet pretrain and the up-sampling layers are randomly initialized. Resnet-101 and DLA-34 train in 2.5 days on 8 TITAN-V GPUs, while Hourglass-104 requires 5 days.

We use three levels of test augmentations: no augmentation, flip augmentation, and flip and multi-scale (0.5, 0.75, 1, 1.25, 1.5). For flip, we average the network outputs before decoding bounding boxes. For multi-scale, we use NMS to merge results. These augmentations yield different speed-accuracy trade-off, as is shown in the next section.

Experiments

We evaluate our object detection performance on the MS COCO dataset , which contains 118k training images (train2017), 5k validation images (val2017) and 20k hold-out testing images (test-dev). We report average precision over all IOU thresholds (AP), AP at IOU thresholds 0.5(AP50AP_{50}) and 0.75 (AP75AP_{75}). The supplement contains additional experiments on PascalVOC .

Table 1 shows our results on COCO validation with different backbones and testing options, while Figure 1 compares CenterNet with other real-time detectors. The running time is tested on our local machine, with Intel Core i7-8086K CPU, Titan Xp GPU, Pytorch 0.4.1, CUDA 9.0, and CUDNN 7.1. We download code and pre-trained modelshttps://github.com/facebookresearch/Detectronhttps://github.com/pjreddie/darknet to test run time for each model on the same machine.

Hourglass-104 achieves the best accuracy at a relatively good speed, with a 42.2%42.2\% AP in 7.87.8 FPS. On this backbone, CenterNet outperforms CornerNet (40.6%40.6\% AP in 4.14.1 FPS) and ExtremeNet (40.3%40.3\% AP in 3.13.1 FPS) in both speed and accuracy. The run time improvement comes from fewer output heads and a simpler box decoding scheme. Better accuracy indicates that center points are easier to detect than corners or extreme points.

Using ResNet-101, we outperform RetinaNet with the same network backbone. We only use deformable convolutions in the upsampling layers, which does not affect RetinaNet. We are more than twice as fast at the same accuracy (CenterNet 34.8%34.8\%AP in 4545 FPS (input 512×512512\times 512) vs. RetinaNet 34.4%34.4\%AP in 1818 FPS (input 500×800500\times 800)). Our fastest ResNet-18 model also achieves a respectable performance of 28.1%28.1\% COCO AP at 142142 FPS.

DLA-34 gives the best speed/accuracy trade-off. It runs at 5252FPS with 37.4%37.4\%AP. This is more than twice as fast as YOLOv3 and 4.4%4.4\%AP more accurate. With flip testing, our model is still faster than YOLOv3 and achieves accuracy levels of Faster-RCNN-FPN (CenterNet 39.2%39.2\% AP in 2828 FPS vs Faster-RCNN 39.8%39.8\% AP in 1111 FPS).

We compare with other state-of-the-art detectors in COCO test-dev in Table 2. With multi-scale evaluation, CenterNet with Hourglass-104 achieves an AP of 45.1%45.1\%, outperforming all existing one-stage detectors. Sophisticated two-stage detectors are more accurate, but also slower. There is no significant difference between CenterNet and sliding window detectors for different object sizes or IoU thresholds. CenterNet behaves like a regular detector, just faster.

1.1 Additional experiments

In unlucky circumstances, two different objects might share the same center, if they perfectly align. In this scenario, CenterNet would only detect one of them. We start by studying how often this happens in practice and put it in relation to missing detections of competing methods.

In the COCO training set, there are 614614 pairs of objects that collide onto the same center point at stride 44. There are 860001860001 objects in total, hence CenterNet is unable to predict <0.1%<0.1\% of objects due to collisions in center points. This is much less than slow- or fast-RCNN miss due to imperfect region proposals (∼2%\sim 2\%), and fewer than anchor-based methods miss due to insufficient anchor placement (20.0%20.0\% for Faster-RCNN with 1515 anchors at 0.50.5 IOU threshold). In addition, 715715 pairs of objects have bounding box IoU >0.7>0.7 and would be assigned to two anchors, hence a center-based assignment causes fewer collisions.

To verify that IoU based NMS is not needed for CenterNet, we ran it as a post-processing step on our predictions. For DLA-34 (flip-test), the AP improves from 39.2%39.2\% to 39.7%39.7\%. For Hourglass-104, the AP stays at 42.2%42.2\%. Given the minor impact, we do not use it. Next, we ablate the new hyperparameters of our model. All the experiments are done on DLA-34.

During training, we fix the input resolution to 512×512512\times 512. During testing, we follow CornerNet to keep the original image resolution and zero-pad the input to the maximum stride of the network. For ResNet and DLA, we pad the image with up to 32 pixels, for HourglassNet, we use 128 pixels. As is shown in Table. 5(a), keeping the original resolution is slightly better than fixing test resolution. Training and testing in a lower resolution (384×384384\times 384) runs 1.71.7 times faster but drops 33AP.

We compare a vanilla L1 loss to a Smooth L1 for size regression. Our experiments in Table 5(c) show that L1 is considerably better than Smooth L1. It yields a better accuracy at fine-scale, which the COCO evaluation metric is sensitive to. This is independently observed in keypoint regression .

We analyze the sensitivity of our approach to the loss weight λsize\lambda_{size}. Table 5(b) shows 0.10.1 gives a good result. For larger values, the AP degrades significantly, due to the scale of the loss ranging from to output size w/Rw/R or h/Rh/R, instead of to 11. However, the value does not degrade significantly for lower weights.

By default, we train the keypoint estimation network for 140140 epochs with a learning rate drop at 9090 epochs. If we double the training epochs before dropping the learning rate, the performance further increases by 1.11.1 AP (Table 5(d)), at the cost of a much longer training schedule. To save computational resources (and polar bears), we use 140140 epochs in ablation experiments, but stick with 230230 epochs for DLA when comparing to other methods.

Finally, we tried a multiple “anchor” version of CenterNet by regressing to more than one object size. The experiments did not yield any success. See supplement.

2 3D detection

We perform 3D bounding box estimation experiments on KITTI dataset , which contains carefully annotated 3D bounding box for vehicles in a driving scenario. KITTI contains 78417841 training images and we follow standard training and validation splits in literature . The evaluation metric is the average precision for cars at 1111 recalls (0.00.0 to 1.01.0 with 0.10.1 increment) at IOU threshold 0.50.5, as in object detection . We evaluate IOUs based on 2D bounding box (AP), orientation (AOP), and Bird-eye-view bounding box (BEV AP). We keep the original image resolution and pad to 1280×3841280\times 384 for both training and testing. The training converges in 7070 epochs, with learning rate dropped at the 4545 and 6060 epoch, respectively. We use the DLA-34 backbone and set the loss weight for depth, orientation, and dimension to 11. All other hyper-parameters are the same as the detection experiments.

Since the number of recall thresholds is quite small, the validation AP fluctuates by up to 10%10\% AP. We thus train 55 models and report the average with standard deviation.

We compare with slow-RCNN based Deep3DBox and Faster-RCNN based method Mono3D , on their specific validation split. As is shown in Table 4, our method performs on-par with its counterparts in AP and AOS and does slightly better in BEV. Our CenterNet is two orders of magnitude faster than both methods.

3 Pose estimation

Finally, we evaluate CenterNet on human pose estimation in the MS COCO dataset . We evaluate keypoint AP, which is similar to bounding box AP but replaces the bounding box IoU with object keypoint similarity. We test and compare with other methods on COCO test-dev.

We experiment with DLA-34 and Hourglass-104, both fine-tuned from center point detection. DLA-34 converges in 320 epochs (about 3 days on 8GPUs) and Hourglass-104 converges in 150 epochs (8 days on 5 GPUs). All additional loss weights are set to 11. All other hyper-parameters are the same as object detection.

The results are shown in Table 5. Direct regression to keypoints performs reasonably, but not at state-of-the-art. It struggles particularly in high IoU regimes. Projecting our output to the closest joint detection improves the results throughout, and performs competitively with state-of-the-art multi-person pose estimators . This verifies that CenterNet is general, easy to adapt to a new task.

Figure 5 shows qualitative examples on all tasks.

Conclusion

In summary, we present a new representation for objects: as points. Our CenterNet object detector builds on successful keypoint estimation networks, finds object centers, and regresses to their size. The algorithm is simple, fast, accurate, and end-to-end differentiable without any NMS post-processing. The idea is general and has broad applications beyond simple two-dimensional detection. CenterNet can estimate a range of additional object properties, such as pose, 3D orientation, depth and extent, in one single forward pass. Our initial experiments are encouraging and open up a new direction for real-time object recognition and related tasks.

References

Appendix A: Model Architecture

See figure. 6 for diagrams of the architectures.

Appendix B: 3D BBox Estimation Details

Our network outputs maps for depths D^∈RWR×HR\hat{D}\in R^{\frac{W}{R}\times\frac{H}{R}}, 3d dimensions Γ^∈RWR×HR×3\hat{\Gamma}\in R^{\frac{W}{R}\times\frac{H}{R}\times 3}, and orientation encoding A^∈RWR×HR×8\hat{A}\in R^{\frac{W}{R}\times\frac{H}{R}\times 8}. For each object instance kk, we extract the output values from the three output maps at the ground truth center point location: d^k∈R\hat{d}_{k}\in R, γ^k∈R3\hat{\gamma}_{k}\in R^{3}, α^k∈R8\hat{\alpha}_{k}\in R^{8}. The depth is trained with L1 loss after converting the output to the absolute depth domain:

where dkd_{k} is the groud truth absolute depth (in meter). Similarly, the 3D dimension is trained with L1 Loss in absolute metric:

where γk\gamma_{k} is the object height, width, and length in meter.

The orientation θ\theta is a single scalar by default. Following Mousavian et al. , We use an 88-scalar encoding to ease learning. The 88 scalars are divided into two groups, each for an angular bin. One bin is for angles in B1=[−7π6,π6]B_{1}=[-\frac{7\pi}{6},\frac{\pi}{6}] and the other is for angles in B2=[−π6,7π6]B_{2}=[-\frac{\pi}{6},\frac{7\pi}{6}]. Thus we have 44 scalars for each bin. Within each bin, 22 of the scalars bi∈R2b_{i}\in R^{2} are used for softmax classification (if the orientation falls into to this bin ii). And the rest 22 scalars ai∈R2a_{i}\in R^{2} are for the sin⁡\sin and cos⁡\cos value of in-bin offset (to the bin center mim_{i}). I.e., α^=[b^1,a^1,b^2,a^2]\hat{\alpha}=[\hat{b}_{1},\hat{a}_{1},\hat{b}_{2},\hat{a}_{2}] The classification are trained with softmax and the angular values are trained with L1 loss:

where jj is the bin index which has a larger classification score.

Appendix C: Collision Experiment Details

We analysis the annotations of COCO training set to show how often the collision cases happen. COCO training set (train 2017) contains N=118287N=118287 images and M=860001M=860001 objects (with MS=356340M_{S}=356340 small objects, MM=295163M_{M}=295163 medium objects, and ML=208498M_{L}=208498 large objects) in C=80C=80 categories. Let the ii-th bounding box of image kk of category cc be bb(kci)=(x1(kci),y1(kci),x2(kci),y2(kci))bb^{(kci)}=(x_{1}^{(kci)},y_{1}^{(kci)},x_{2}^{(kci)},y_{2}^{(kci)}), its center after the 4×4\times stride is pkci=(⌊14⋅x1(kci)+x2(kci)2⌋,⌊14⋅y1(kci)+y2(kci)2⌋)p^{kci}=(\lfloor\frac{1}{4}\cdot\frac{x_{1}^{(kci)}+x_{2}^{(kci)}}{2}\rfloor,\lfloor\frac{1}{4}\cdot\frac{y_{1}^{(kci)}+y_{2}^{(kci)}}{2}\rfloor). And Let n(kc)n^{(kc)} be the number of object of category cc in image kk. The number of center point collisions is calculated by:

Similarly, we calculate the IoU based collision by

This gives NIoU@0.7=715N_{IoU@0.7}=715 and NIoU@0.5=5179N_{IoU@0.5}=5179.

RetinaNet assigns anchors to a ground truth bounding box if they have >0.5>0.5 IoU. In the case that a ground truth bounding box has not been covered by any anchor with IoU >0.5>0.5, the anchor with the largest IoU will be assigned to it. We calculate how often this forced assignment happens. We use 1515 anchors (55 size: 32, 64, 128, 256, 512, and 33 aspect-ratio: 0.5, 1, 2, as is in RetinaNet ) at stride S=16S=16. For each image, after resizing it as its shorter edge to be 800800 , we place these anchors at positions {(S/2+i×S,S/2+j×S)}\{(S/2+i\times S,S/2+j\times S)\}, where i∈[0,⌊(W−S/2)S⌋]i\in[0,\lfloor\frac{(W-S/2)}{S}\rfloor] and j∈[0,⌊(H−S/2)S⌋]j\in[0,\lfloor\frac{(H-S/2)}{S}\rfloor]. W, H are the image weight and height (the smaller one is equal to 800). This results in a set of anchors A\mathcal{A}. ∣A∣=15×⌊(W−S/2)S+1⌋×⌊(H−S/2)S+1⌋|\mathcal{A}|=15\times\lfloor\frac{(W-S/2)}{S}+1\rfloor\times\lfloor\frac{(H-S/2)}{S}+1\rfloor. We calculate the number of the forced assignments by:

RenitaNet requires Nanchor=170220N_{anchor}=170220 forced assignments: 125831125831 for small objects (35.3%35.3\% of all small objects), 1850518505 for medium objects (6.3%6.3\% of all medium objects), and 2588425884 for large objects (12.4%12.4\% of all large objects).

Appendix D: Experiments on PascalVOC

Pascal VOC is a popular small object detection dataset. We train on VOC 2007 and VOC 2012 trainval sets, and test on VOC 2007 test set. It contains 1655116551 training images and 49624962 testing images of 2020 categories. The evaluation metric is mean average precision (mAP) at IOU threshold 0.50.5.

We experiment with our modified ResNet-18, ResNet-101, and DLA-34 (See main paper Section. 5) in two training resolution: 384×384384\times 384 and 512×512512\times 512. For all networks, we train 7070 epochs with learning rate dropped 10×10\times at 4545 and 6060 epochs, respectively. We use batchsize 3232 and learning rate 1.25e1.25e-4 following the linear learning rate rule . It takes one GPU 77 hours/ 1010 hours to train in 384×384384\times 384 for ResNet-101 and DLA-34, respectively. And for 512×512512\times 512, the training takes the same time in two GPUs. Flip augmentation is used in testing. All other hyper-parameters are the same as the COCO experiments. We do not use Hourglass-104 because it fails to converge in a reasonable time (2 days) when trained from scratch.

The results are shown in Table. 6. Our best CenterNet-DLA model performs competitively with top-tier methods, and keeps a real-time speed.

Appendix E: Error Analysis

We perform an error analysis by replacing each output head with its ground truth. For the center point heatmap, we use the rendered Gaussian ground truth heatmap. For the bounding box size, we use the nearest ground truth size for each detection.

The results in Table 7 show that improving both size map leads to a modest performance gain, while the center map gains are much larger. If only the keypoint offset is not predicted, the maximum AP reaches 83.183.1. The entire pipeline on ground truth misses about 0.5%0.5\% of objects, due to discretization and estimation errors in the Gaussian heatmap rendering.