PP-YOLOv2: A Practical Object Detector

Xin Huang, Xinxin Wang, Wenyu Lv, Xiaying Bai, Xiang Long, Kaipeng Deng, Qingqing Dang, Shumin Han, Qiwen Liu, Xiaoguang Hu, Dianhai Yu, Yanjun Ma, Osamu Yoshie

Introduction

Object detection is a critical component of various real-world applications such as self-driving cars, face recognition, and person re-identification. In recent years, the performance of object detectors has been rapidly improved with the rise of deep convolutional neural networks (CNNs) . Although, recent works focus on novel detection pipeline (i.e., Cascade RCNN and HTC ), sophisticated network architecture design (DetectoRS and CBNET ) push forward the state-of-the-art object detection approaches, YOLOv3 is still one of the most widely used detector in industry. Because, in various practical applications, not only the computation resources are limited, but also the software support is insufficient. Without necessary technique support, two stage object detector(e.g. Faster RCNN , Cascade RCNN ) may excruciatingly slow. Meanwhile, a significant gap exists between the accuracy of YOLOv3 and two stage object detectors. Therefore, how to improve the effectiveness of YOLOv3 while maintaining the inference speed is an essential problem for practical use. To simultaneously satisfy two concerns, we add a bunch of refinements that almost not increase the infer time to improve the overall performance of the PP-YOLO . To note that, although a huge number of approaches claim to improve object detector’s accuracy independently, in practice, some methods are not effective when combined. Therefore, practical testing of combinations of such tricks is required. We follow the incremental manner to evaluate their effectiveness one by one. All our experiments are implemented based on PaddlePaddlehttps://github.com/PaddlePaddle/Paddle .

In fact, this paper is more like a TECH REPORT, which tells you how to build PP-YOLOv2 step by step. Theoretical justification of the failure cases is also involved. To this end, we achieve a better balance between effectiveness (49.5% mAP) and efficiency (69 FPS), surpassing existing robust detectors with roughly the same amount of parameters such as YOLOv4-CSP and YOLOv5lhttps://github.com/ultralytics/yolov5. Hopefully, our experience in building PP-YOLOv2 can help developers and researchers to think deeper in implementing object detectors for practical applications.

Revisit PP-YOLO

In this section, we will perform the implementation of our baseline model specifically.

Pre-Processing. Apply Mixup Training with a weight sampled from Beta(α,β)Beta({\alpha},{\beta}) distribution where α=1.5,β=1.5.\alpha=1.5,\beta=1.5. Then, RandomColorDistortion, RandomExpand, RandCrop and RandomFlip are applied one by one with probability 0.5. Next, Normalize RGB channels by subtracting 0.485, 0.456, 0.406 and dividing by 0.229, 0.224, 0.225, respectively. Finally, The input size is evenly drawn from .

Baseline Model. Our baseline model is PP-YOLO which is an enhanced version of YOLOv3. Specifically, it first replaces the backbone to ResNet50-vd. After that a total of 10 tricks which can improve the performance of YOLOv3 almost without losing efficiency are added to YOLOv3 such as Deformable Conv , SSLD , CoordConv , DropBlock , SPP and so on. The architecture of PP-YOLO is presented in the paper .

Training Schedule. On COCO train2017, the network is trained with stochastic gradient descent (SGD) for 500K iterations with a minibatch of 96 images distributed on 8 GPUs. The learning rate is linearly increased from 0 to 0.005 in 4K iterations, and it is divided by 10 at iteration 400K and 450K, respectively. Weight decay is set as 0.0005, and momentum is set as 0.9. Gradient clipping is adopted to stable the training procedure.

Selection of Refinements

Path Aggregation Network. Detecting objects at different scales is a fundamental challenge in object detection. In practice, a detection neck is developed for building high-level semantic feature maps at all scales. In PP-YOLO, FPN is adopted to compose bottom-up paths. Recently, several FPN variants have been proposed to enhance the ability of pyramid representation. For example, BiFPN , PAN , RFP and so on. We follow the design of PAN to aggregate the top-down information. The detailed structure of PAN is shown in Fig. 2.

Mish Activation Function. Mish activation function has been proved effective in many practical detectors, such as YOLOv4 and YOLOv5. They adopt the mish activation function in the backbone. However, we prefer to use pre-trained parameters because we have a powerful model which achieves 82.4% top-1 accuracy on ImageNet. To keep the backbone unchanged, we apply the mish activation function in the detection neck instead of the backbone.

Larger Input Size. Increasing the input size enlarges the area of objects. Thus, information of the objects on a small scale will be preserved easier than before. As a result, performance will be increased. However, a larger input size occupies more memory. To apply this trick, we need to decrease batch size. To be more specific, we reduce the batch size from 24 images per GPU to 12 images per GPU and expand the largest input size from 608 to 768. The input size is evenly drawn from .

IoU Aware Branch. In PP-YOLO, IoU aware loss is calculated in a soft weight format which is inconsistent with the original intention. Therefore, we apply a soft label format. Here is the IoU aware loss:

where tt indicates the IoU between the anchor and its matched ground-truth bounding box, pp is the raw output of IoU aware branch, σ(⋅)\sigma(\cdot) refers to the sigmoid activation function. To note that only positive samples’ IoU aware loss is computed. By replacing the loss function, IoU aware branch works better than before.

Experiments

COCO is a widely used benchmark in the field of object detection. In this work, we train all our models on the COCO train2017 which consists of 118k images across 80 classes. For evaluation, we evaluate our results on the COCO minival which consists of 5k testing images. Our evaluation metric also follows the standard COCO style mean Average Precision (mAP).

2 Ablation Studies

In this subsection, we present the effectiveness of each module in an incremental manner. Results are shown in Table 1, where infer time and FPS only consider the influence of the model in FP32-precision which does not include result decoder and NMS following YOLOv4.

A. First of all, we follow the original design of PP-YOLO to build our baseline. Since the heavy pre-processing on the CPU slows down the training, we decrease the images per GPU from 24 to 12. Reducing batch size drops mAP by 0.2%. Training settings are described in section 2 entirely.

A →\rightarrow B. The first refinement with a positive effect on PP-YOLO that we found was PAN. To stable the training process, we add several skip connections to our PAN module. The detailed structure of PAN is shown in Fig. 2. We can see that PAN and FPN are a group of symmetrical structures. When we perform it with Mish, it boosts the performance from 45.1% mAP to 47.1% mAP. Although model B is slightly slower than model A, such a significant gain promotes us to adopt PAN in our final model. For more details, please refer to our code.

B →\rightarrow C. Since the input size of YOLOv4 and YOLOv5 during evaluation is 640, we increase training and evaluation input size to 640 to build a fair comparison. The performance increases 0.6% mAP.

C →\rightarrow D. Keep increasing the input size should benefit more. However, it is impossible to use Larger Input Size and Larger Batch Size together. We train the model D with 12 images per GPU and Larger Input Size. It increases the mAP by 0.6% which brings more gains than Larger Batch Size. Therefore, we choose Larger Input Size in the final practice. The input size is evenly drawn from .

D →\rightarrow E. In the training phase, the modified IoU aware loss performs better than before. In the former version, the value of IoU aware loss will drop to 1e-5 in hundreds of iterations during training. After we modified the IoU aware loss, its value and the value of IoU loss are in the same order of magnitude, which is reasonable. After using this strategy, the mAP of model E increases to 49.1% without any loss of efficiency.

3 Comparison With Other State-of-the-Art Detectors

Comparison of the results on MS-COCO test split with other state-of-the-art object detectors is shown in Figure 1 and Table 2. We compare our method with YOLOv4-CSP and YOLOv5l because they have roughly the same amount of parameters as our model. It clearly shows that PP-YOLOv2 outperforms these two methods. With a similar FPS, PP-YOLOv2 outperforms YOLOv4-CSP by 2% mAP and surpasses YOLOv5l by 1.3% mAP. Besides, when we replace PP-YOLOv2’s backbone from ResNet50 to ResNet101, PP-YOLOv2 achieves comparable performance with YOLOv5x while it is 15.9% faster than YOLOv5x. Therefore, we can draw a conclusion that compared with other state-of-the-art methods, our PP-YOLOv2 has certain advantages in the balance of speed and accuracy.

Moreover, PP-YOLOv2 is implemented based on PaddlePaddle. As a deep learning framework, PaddlePaddle not only supports model implementation but also pays attention to model deployment. With official support, adapting TensorRT for PP-YOLOv2 is much easier than other detectors. Specifically, the Paddle inference engine with TensorRT, FP16-precision, and batch size = 1 further improves PP-YOLOv2’s infer speed. The speed-up ratios for PP-YOLOv2(R50) and PP-YOLOv2(R101) are 54.6% and 73%, respectively.

Things We Tried That Didn’t Work

Since it takes about 80 hours for training PP-YOLO with 8 V100 GPUs on COCO train2017, we involve COCO minitrain to speed up our analysis on ablation studies. COCO minitrain is a subset of the COCO train2017, containing 25K images. On COCO minitrain, the total iterations is 90K. We divide the learning rate by 10 at iteration 60k. Other settings are the same as training on COCO train2017.

We tried lots of stuff while we were working on PP-YOLOv2. Some of them have a positive effect on COCO minitrain while hinders the performance when training on COCO train2017. Due to the inconsistency, someone may doubt the experimental conclusion on COCO minitrain.The reason why we use COCO minitrain is that we want to seek refinements with universal features, which means that they should be useful on different scale datasets. It is also important to figure out the reason why they failed. Therefore, we discuss some of them in this section.

Cosine Learning Rate Decay. Different from linear step learning rate decay, cosine learning rate decay is exponentially decaying the learning rate. Therefore, the change of the learning rate is smooth, which benefits the training process. We follow the formula in Bag of Tricks to set learning rate at each epoch. Although cosine learning rate decay achieves better performance on COCO minitrain, it is sensitive to hyper-parameters such as initial learning rate, the number of warm up steps, and the ending learning rate. We tried several sets of hyper-parameters. However, we didn’t observe a positive effect on COCO train2017 eventually.

Backbone Parameter Freezing. When fine-tuning the ImageNet pre-trained parameters on downstream tasks, freezing parameters in the first two stages is a common practice. Since our pre-trained ResNet50-vd is much powerful than others(82.4% Top1 accuracy versus 79.3% Top1 accuracy), we are more motivated to adopt this strategy. On COCO minitrain, parameter freezing brings 1mAP gain, however, on COCO train2017 it decreases mAP by 0.8% . A possible reason for the inconsistency phenomena was speculated to be the different sizes of the two training sets. COCO minitrain is a fifth of COCO train2017. The ability to generalization of parameters that are trained on a small dataset may be worse than pre-trained parameters.

SiLU Activation Function. We tried using SiLU instead of Mish in detection neck. This increases 0.3% mAP on COCO minitrain but drops 0.5% mAP on COCO train2017. We are not sure about the reason.

Conclusions

This paper presents some updates to PP-YOLO, which forms a high-performance object detector called PP-YOLOv2. PP-YOLOv2 achieves a better balance between speed and accuracy than other famous detectors, such as YOLOv4 and YOLOv5. In this paper, we explore a bunch of tricks and show how to combine these tricks on the PP-YOLO detector and demonstrate their effectiveness. Moreover, with PaddlePaddle’s official support, the gap between model development and production deployment is narrowed. We hope this paper can help developers and researchers get better performance in practical scenes.

Acknowledgements

This work was supported by the National Key Research and Development Project of China (2020AAA0103500).

References

References