You Only Look One-level Feature
Qiang Chen, Yingming Wang, Tong Yang, Xiangyu Zhang, Jian Cheng, Jian Sun
Introduction
In state-of-the-art two-stage detectors and one-stage detectors , feature pyramids become an essential component. The most popular way to build feature pyramids is the feature pyramid networks (FPN) , which mainly brings two benefits: (1) multi-scale feature fusion: fusing multiple low-resolution and high-resolution feature inputs to obtain better representations; (2) divide-and-conquer: detecting objects on different levels regarding objects’ scales. A common belief for FPN is that its success relies on the fusion of multiple level features, inducing a line of studies of designing complex fusion methods manually , or via Neural Architecture Search (NAS) algorithms . However, the belief ignores the function of the divide-and-conquer in FPN. It leads to fewer studies on how these two benefits contribute to FPN’s success and may hinder new advances.
This paper studies the influence of FPN’s two benefits in one-stage detectors. We design experiments by decoupling the multi-scale feature fusion and the divide-and-conquer functionalities with RetinaNet . In detail, we consider FPN as a Multiple-in-Multiple-out (MiMo) encoder, which encodes multi-scale features from the backbone and provides feature representations for the decoder (the detection heads). We conduct controlled comparisons among Multiple-in-Multiple-out (MiMo), Single-in-Multiple-out (SiMo), Multiple-in-Single-out (MiSo), and Single-in-Single-out (SiSo) encoders in Figure 1. Surprisingly, the SiMo encoder, which only has one input feature C5 and does not perform feature fusion, can achieve comparable performance with the MiMo encoder (i.e., FPN). The performance gap is less than mAP. In contrast, the performance drops dramatically ( mAP) in MiSo and SiSo encoders. These phenomenons suggest two facts: (1) the C5 feature carries sufficient context for detecting objects on various scales, which enables the SiMo encoder to achieve comparable results; (2) the multi-scale feature fusion benefit is far away less critical than the divide-and-conquer benefit, thus multi-scale feature fusion might not be the most significant benefit of FPN, which is also demonstrated by ExFuse in semantic segmentation. Thinking one step deeper, divide-and-conquer is related to the optimization problem in object detection. It divides the complex detection problem into several sub-problems by object scales, facilitating the optimization process.
The above analysis suggests that the essential factor for the success of FPN is its solution to the optimization problem in object detection. The divide-and-conquer solution is a good way. But it brings memory burdens, slows down the detectors, and make detectors’ structure complex in one-stage detectors like RetinaNet . Given that the C5 feature carries sufficient context for detection, we show a simple way to address the optimization problem.
We propose You Only Look One-level Feature (YOLOF), which only uses one single C5 feature (with a downsample rate of 32) for detection. To bridge the performance gap between the SiSo encoder and the MiMo encoder, we first design the structure of the encoder properly to extract the multi-scale contexts for objects on various scales, compensating for the lack of multiple-level features; then, we apply a uniform matching mechanism to solve the imbalance problem of positive anchors raised by the sparse anchors in the single feature.
Without bells and whistles, YOLOF achieves comparable results with its feature pyramids counterpart RetinaNet but faster. In a single feature manner, YOLOF matches the performance of the recent proposed DETR while converging much faster (). With an image size of and other techniques , YOLOF achieve mAP running at fps on 2080Ti, which is faster than YOLOv4 . In a nutshell, the contributions of this paper are:
We show that the most significant benefits of FPN is its divide-and-conquer solution to the optimization problem in dense object detection rather than the multi-scale feature fusion.
We present YOLOF, which is a simple and efficient baseline without using FPN. In YOLOF, we propose two key components, Dilated Encoder and Uniform Matching, bridging the performance gap between the SiSo encoder and the MiMo encoder.
Extensive experiments on COCO benchmark indicates the importance of each component. Moreover, we conduct comparisons with RetinaNet , DETR and YOLOv4 . We can achieve comparable results with a faster speed on GPUs.
Related Works
It is a conventional technique to employ multiple features for object detection. Typical approaches to construct multiple features can be categorized into image pyramid methods and feature pyramid methods. Image pyramids based detector such as DPM dominates the detection in the pre-deep learning era. In CNN-based detectors, the image pyramids method also wins some researchers’ praise as it can achieve higher performance out of the box. However, the image pyramids method is not the only way to obtain multiple features; it is more efficient and natural to exploit feature pyramids’ power in CNN models. SSD first utilizes multiple-scale features and performs object detection on each scale for different scales objects. FPN follows SSD and UNet and constructs semantic-riched feature pyramids by combining shallow features and deep features. After that, several works follow FPN and focus on how to obtain better representations. FPN becomes an essential component and dominates modern detectors. It is also applied to popular one-stage detectors, such as RetinaNet , FCOS , and their variants . Another line of method to get feature pyramids is to use multi-branch and dilation convolution . Different from the above works, our method is a single-level feature detector.
Single-level feature detectors.
In early times, the R-CNN series and R-FCN only extract RoI features on a single feature, while their performances lag behind their multiple feature counterparts . Also, in one-stage detectors, YOLO and YOLOv2 only use the last output feature of the backbone. They can be super fast but have to bear a performance decline in detection. CornerNet and CenterNet follow this fashion and achieve competitive results while using a single feature with a downsample rate of 4 to detect all the objects. Using a high-resolution feature map for detection brings enormous memory cost and is not friendly to deployment. Recently, DETR introduces the transformer to detection and shows that it could achieve state-of-the-art results only use a single C5 feature. Due to the totally anchor-free mechanism and transformer learning phase, DETR needs a long training schedule for its convergence. The long training schedule characteristic is cumbersome for further improvements. Unlike these papers, we investigate the working mechanism of multiple-level detection. From the perspective of optimization, we provide an alternative solution to the widely used FPN. Moreover, YOLOF converges faster and achieves promising performance; thus, YOLOF can serve as a simple baseline for fast and accurate detectors.
Cost Analysis of MiMo Encoders
As mentioned in Section 1, the success of FPN in dense object detection is due to its solution to the optimization problem. However, the multi-level feature paradigm is inevitable to make detectors complex, brings memory burdens, and slows down the detector. In this section, we provide a quantitative study on the cost of MiMo encoders.
We design experiments based on RetinaNet with ResNet-50 . In detail, we format the pipeline for the detection task as a combination of three key parts: the backbone, the encoder, and the decoder (Figure 2). In this view, we show the FLOPs of each component in Figure 3. Compared with SiSo encoders, the MiMo encoder brings enormous memory burdens to the encoder and the decoder(134G vs. 6G) (Figure 3). Moreover, the detector with MiMo encoder runs much slower than the ones with SiSo encoders (13 FPS vs. 34 FPS) (Figure 3). The slow speed is caused by detecting objects on high-resolution feature maps in the detector with MiMo encoder, such as the C3 feature (with a downsample rate of 8). Given the above drawbacks of the MiMo encoder, we aim to find an alternative way to solve the optimization problem while keeping the detector simple, accurate, and fast simultaneously.
Method
Motivated by the above purpose and the finding that the C5 feature contains enough context for detecting numerous objects, we try to replace the complex MiMo encoder with the simple SiSo encoder in this section. But this replacement is nontrivial as the performance drops extensively when applying SiSo encoders according to the results in Figure 3. Given the situation, we carefully analyze the obstacles preventing SiSo encoders from getting a comparable performance with MiMo encoders. We find that two problems brought by SiSo encoders are responsible for the performance drop. The first problem is that the range of scales matching to the C5 feature’s receptive field is limited, which impedes the detection performance for objects across various scales. The second one is the imbalance problem on positive anchors raised by sparse anchors in the single-level feature. Next, we discuss these two problems in detail and provide our solutions.
Recognizing objects at vastly different scales is a fundamental challenge in object detection. One feasible solution to this challenge is to leverage multiple-level features. In detectors with MiMo or SiMo encoders, they construct multiple-level features with different receptive fields (P3-P7) and detect objects on the level with receptive field matching to their scales. However, the single-level feature setting changes the game. There is only one output feature in SiSo encoders, whose receptive field is a constant. As shown in Figure 4(a), the C5 feature’s receptive field can only cover a limited scale range, resulting in poor performance if the objects’ scales mismatches with the receptive field. To achieve the goal of detecting all objects with SiSo encoders, we have to find a way to generate an output feature with various receptive fields, compensating for the lack of multiple-level features.
We begin with enlarging the receptive field of the C5 feature by stacking standard and dilated convolutions . Although the covered scale range is enlarged to some extent, it still can not cover all object scales as the enlarging process multiplies a factor greater than to all originally covered scales. We illustrate the situation in Figure 4(b), where the whole scale range shifts to larger scales compare with the one in Figure 4(a). Then, we combine the original scale range and the enlarged scale range by adding the corresponding features, resulting in an output feature with multiple receptive fields covering all object scales (Figure 4(c)). The above operations can be easily achieved by constructing residual blocks with dilations on the middle convolution layer.
Based on the above designs, we propose our SiSo encoder in Figure 5, named as Dilated Encoder. It contains two main components: the Projector and the Residual Blocks. The projection layer first applies one convolution layer to reduce the channel dimension, then add one convolution layer to refine semantic contexts, which is the same as in the FPN . After that, we stack four successive dilated residual blocks with different dilation rates in the convolution layers to generate output features with multiple receptive fields, covering all objects’ scales.
Discussion:
Dilated convolution is a common strategy to enlarge the features’ receptive field in object detection. As reviewed in the Section 2, TridentNet use dilated convolution to generate multi-scale features. It deals with the scale variation problem in object detection via multi-branch structure and weight sharing mechanism, which is different from our single-level feature setting. Moreover, Dilated Encoder stack dilated residual blocks one by one without weight sharing. Although DetNet also successively applies dilated residual blocks, its purpose is to maintain the spatial resolution of the features and keep more details in the backbone’s outputs, while ours is to generate a feature with multiple receptive fields out of the backbone. The design of Dilated Encoder enables us to detecting all objects on single-level feature instead of on multiple-level features like TridentNet and DetNet .
2 Imbalance Problem on Positive Anchors
The definition of positive anchors is crucial for the optimization problem in object detection. In anchor-based detectors, strategies to define positive are dominated by measuring the IoUs between anchors and ground-truth boxes. In RetinaNet , if the max IoU of the anchor and ground-truth boxes is greater than a threshold , this anchor will be set as positive. We call it Max-IoU matching.
In MiMo encoders, the anchors are pre-defined on multiple levels in a dense paved fashion, and the ground-truth boxes generate positive anchors in feature levels corresponding to their scales. Given the divide-and-conquer mechanism, Max-IoU matching enables ground-truth boxes in each scale to generate a sufficient number of positive anchors. However, when we adopt the SiSo encoder, the number of anchors diminish extensively compare to the one in the MiMo encoder, from to , resulting in sparse anchorsIn SiSo encoders, we simply collapse multiple anchors on multiple-level features to single-level, e.g., we construct anchors with different anchor sizes of {32, 64, 128, 256, 512} on each position of the C5 feature.. Sparse anchors raise a matching problem for detectors when applying Max-IoU matching, as shown in Figure 6. Large ground-truth boxes induce more positive anchors than small ground-truth boxes in natural, which cause an imbalance problem for positive anchors. This imbalance makes detectors pay attention to large ground-truth boxes while ignoring the small ones when training.
To solve this imbalance problem in positive anchors, we propose an Uniform Matching strategy: adopting the k nearest anchor as positive anchors for each ground-truth box, which makes sure that all ground-truth boxes can be matched with the same number of positive anchors uniformly regardless of their sizes (Figure 6). Balance in positive samples makes sure that all ground-truth boxes participate in training and contribute equally. Besides, following Max-IoU matching , we set IoU thresholds in Uniform Matching to ignore large IoU () negative anchors and small IoU (<0.15) positive anchors.
Discussion: relation to other matching methods.
Applying topk in the matching process is not new. ATSS first select topk anchors for each ground-truth box on feature levels, then samples positive anchors among candidates by dynamic IoU thresholds. However, ATSS focuses on defining positives and negatives adaptively, while our uniform matching focuses on achieving balance on positive samples with sparse anchors. Although several previous methods achieve balance on positive samples, their matching processes are not designed for this imbalance problem. For example, YOLO and YOLOv2 match the ground-truth boxes with the best matching cell or anchor; DETR and apply Hungarian algorithm for matching. These matching methods can be view as top1 matching, which is a specific case of our uniform matching. More importantly, the difference between the uniform matching and the learning-to-match methods is that: the learning-to-match methods, such as FreeAnchor and PAA , adaptively separate anchors into positives and negatives according to the learning status, while uniform matching is fixed and does not evolve with training. The uniform matching is proposed to address the specific imbalance problem on positive anchors under the SiSo design. The comparison in Figure 6 and the results in Table 5e demonstrate the significance of the balance in positives in SiSo encoders.
3 YOLOF
Based on the solutions above, we propose a fast and straightforward framework with single-level feature, denoted as YOLOF. We format YOLOF into three parts: the backbone, the encoder, and the decoder. The sketch of YOLOF is shown in Figure 9. In this section, we give a brief introduction to the main components of YOLOF.
In all models, we simply adopt the ResNet and ResNeXt series as our backbone. All models are pre-trained on ImageNet. The output of the backbone is the C5 feature map which has 2048 channels and with a downsample rate of 32. To make a fair comparison with other detectors, all batchnorm layers in the backbone are frozen by default.
Encoder.
For the encoder (Figure 5), we first follow FPN by adding two projection layers (one and one convolution) after the backbone, resulting in a feature map with channels. Then, to enable the encoder’s output feature to cover all objects on various scales, we propose to add residual blocks, which consist of three consecutive convolutions: the first convolution apply channel reduction with a reduction rate of , then a convolution with dilation is used to enlarge the receptive field, at last, a convolution to recover the number of channels.
Decoder.
For the decoder, we adopt the main design of RetinaNet, which consists of two parallel task-specific heads: the classification head and the regression head (Figure 9). We only add two minor modifications. The first one is that we follow the design of FFN in DETR and make the number of convolution layers in two heads different. There are four convolutions followed by batch normalization layers and ReLU layers on the regression head while only have two on the classification head. The second is that we follow Autoassign and add an implicit objectness prediction (without direct supervision) for each anchor on the regression head. The final classification scores for all predictions are generated by multiplying the classification output with the corresponding implicit objectness.
Other Details.
As mentioned in the previous section, the pre-defined anchors in YOLOF are sparse, decreasing the match quality between anchors and ground-truth boxes. We add a random shift operation on the image to circumvent this problem. The operation shifts the image randomly with a maximum of 32 pixels in left, right, top, and bottom directions and aims to inject noises into the object’s position in the image, increasing the probability of ground-truth boxes matching with high-quality anchors. Moreover, we found that a restriction on the anchors’ center’s shift is also helpful to the final classification when using a single-level feature. We add a restriction that the centers’ shift for all anchors should smaller than pixels.
Experiments
We evaluate our YOLOF on the MS COCO benchmark and conduct comparisons with RetinaNet and DETR . Then, we provide a detailed ablation study of each component’s design with quantitative results and analysis. Finally, to give insights to further research on single-level detection, we provide error analysis and show the weaknesses of YOLOF compared with DETR . The details are as follows.
YOLOF is trained with synchronized SGD over 8 GPUs with a total of 64 images per mini-batch (8 images per GPU). All models are trained with an initial learning rate of . Moreover, following DETR , we set a smaller learning rate for the backbone, which is of the base learning rate. To stabilize the training at the beginning, we extend the number of warmup iterations from to . For training schedules, as we increase the batch size, the ’’ schedule setting in YOLOF is a total of iterations and with base learning rate decreased by 10 in the and the iteration. Other schedules are adjusted according to the principles in Detectron2 . For model inference, we employ NMS with a threshold of to post-process the results. For other hyperparameters, we follow the settings of RetinaNet .
1 Comparison with previous works
To make a fair comparison, we align RetinaNet with YOLOF by employing generalized IoU for the box loss, adding an implicit objectness prediction, and applying group normalization layers in heads (as there are only two images per GPU and both BN and SyncBN give poor results in RetinaNet https://github.com/facebookresearch/detectron2/blob/master/detectron2/modeling/meta_arch/retinanet.py#L532, we use GN instead of BN in the heads). The results are presented in Table 1. All ’’ models are trained with a single scale that the shorter side is set as 800 pixels and the longer side is at most 1333 . In the top section, we give RetinaNet baseline results trained with Detectron2 . In the middle section, we present the results of the improved RetinaNet baseline (with a ””), whose settings are aligned with YOLOF. In the last section, we show results from multiple YOLOF models. Thanks to the single-level feature, YOLOF achieves results on par with RetinaNet with a flops reduction (flops for each component in YOLOF are shown in Figure 3) and a speed up. Due to the large stride (32) of the C5 feature, YOLOF has an inferior performance () than RetinaNet on small objects. However, YOLOF achieves better performance on large objects (+3.3) as we add dilated residual blocks in the encoder. The comparison between RetinaNet and YOLOF with a ResNet-101 show similar evidence as well. Although YOLOF is inferior to RetinaNet on small objects when applying the same backbone, it can match small objects’ performance with a stronger backbone ResNeXt while running at the same speed. Moreover, to prove that our method is compatible and complementary to current technologies in object detection, we show results that training with multi-scale images and a longer schedule in the last two rows of Table 1. Finally, with the help of multi-scale testing, we obtain our final result of mAP and a competitive performance of mAP on small objects.
Comparison with DETR.
DETR is a recent proposed detector which introduces transformer to object detection. It achieves surprising results on the COCO benchmark and proves that by only adopting a single C5 feature, it can achieve comparable results with a multi-level feature detector (Faster R-CNN w/ FPN ) for the first time. Given this, one might expect that layers capture global dependencies such as transformer layers are required to achieve promising results in single-level feature detection. However, we show that a conventional network with local convolution layers can also achieve this goal. We compare DETR with global layers and YOLOF with local convolution layers in Table 2. The results show that YOLOF matches the DETR’s performance, and YOLOF gets more benefits from deeper networks than DETR (w/ ResNet-50 () vs. w/ ResNet-101 (+0.2)). Interestingly, we find that YOLOF outperforms DETR on small objects (+1.9 and +2.4) while lags behind DETR on large objects (-3.5 and -2.9). The finding is consistent with the local and global discussion above. More importantly, compared with DETR, YOLOF converge much faster (), making it more suitable than DETR to serve as a simple baseline for single-level detectors.
Comparison with YOLOv4.
YOLOv4 is an optimal speed and accuracy multi-level feature detector. It combines many tricks to achieve state-of-the-art results. As our purpose is to build a simple and fast baseline for single-level detectors, investigation on the bag of freebie tricks is outside of the scope of this work. Thus, we do not expect a rigidly aligned comparison on performance. To compare our YOLOF with YOLOv4, we apply the data augmentation methods as YOLOv4, adopt a three-phase training pipeline, modify the training settings accordingly, and add dilations on the last stage of the backbone (YOLOF-DC5 in Table 3). More technical details about the model and the training settings are given in the Appendix. As shown in Table 3, YOLOF-DC5 can run faster than YOLOv4 with a mAP improvement on overall performance. YOLOF-DC5 achieves less competitive results on small objects than YOLOv4 ( mAP vs. mAP) while outperforms it on large objects by a large margin ( mAP). The above results indicate that single-level detectors have great potential to achieve state-of-the-art speed and accuracy simultaneously.
2 Ablation Experiments
We run a number of ablations to analyze YOLOF. We first provide an overall analysis of the two proposed components. Then, we show the ablation experiments on detailed designs of each component. Results are shown in Table 4,5 and discussed in detail next.
Table 4 shows that both Dilated Encoder and Uniform Matching are necessary to YOLOF and bring considerable improvements. Specifically, Dilated Encoder has a significant impact on large objects (43.8 vs. 53.2) and slightly improves the results of small and medium objects. The results indicate that the limited scale range is a severe problem in the C5 feature (Section 4.1). Our Dilated Encoder provides a simple but effective solution to this problem. On the other side, the performance of small and medium objects drops significantly () without uniform matching, while the large objects’ performance is only lightly affected. The finding is consistent with the imbalance problem on positive anchors analyzed in Section 4.2. The positive anchors are dominated by large objects, resulting in poor results on small and medium objects. Finally, when we remove both Dilated Encoder and Uniform Matching, a single-level feature detector’s performance drops back to mAP like the results in Figure 1 and Figure 3.
Number of ResBlock:
YOLOF stacks residual blocks in the SiSo encoder. The results in Table 5a shows that stacking more blocks gives extensive improvements on large objects, which is due to the increment of the feature scale range. Although we observe continuous improvements with more blocks, we choose to add four residual blocks to keep YOLOF simple and neat.
Different dilations:
Following the analysis in Section 4.1, to enable the C5 feature to cover large scales, we replace the standard convolution layer in the residual blocks with its dilated counterpart. We show the results with different dilations in the residual blocks in Table 5b. Applying dilations to residual blocks bring improvements to YOLOF, while the improvements are saturated when using too large dilations. We conjecture that the reason for this phenomenon is that dilations of are enough to match object scales in all images.
Add shortcut or not:
Table 5c shows that shortcuts play an essential role in Dilated Encoder. The performance of all objects will drop significantly if we remove the shortcuts in residual blocks. According to Section 4.1, shortcuts combine different scale ranges. A largely and densely paved scale range covered by the feature is the critical factor for detecting all objects in a single-level feature manner.
Number of positives:
A comparison among the number of induced positive anchors by ground-truth boxes is conducted in Table 5d. Intuitively, more positive anchors can achieve better performance as the learning will be easier when given more samples. Thus, in our uniform matching manner, we empirically increase the number of positive anchors induced by each ground-truth box. As shown in Table 5d, the hyper-parameter is very robust for the performance when is larger than 1, which may suggest that the most important is the uniform matching manner in YOLOF. We set for our uniform matching as it is the best choice according to the results.
Uniform matching vs. other matchings:
We compare the uniform matching with other matching strategies for YOLOF and show results in Table 5e. The proposed uniform matching strategy can achieve the best results, compatible with the imbalance analysis in Figure 6. It worth noting that the Hungarian matching strategy can be roughly treated as Top1 matching (Table 5d) so that they get similar performance. The difference between them is that an anchor will only match one object in Hungarian matching while the Top1 matching does not have this constraint, and the experiments show that this is not important. The original ATSS find that top9 anchors are the best choice, while we find top15 anchors are much better in the single-level feature detector. By using top15 anchors, ATSS achieves a good result of 36.5 mAP while still lags behind our uniform matching by a 1.2 mAP gap.
3 Error Analysis
We add error analysis for YOLOF in this section to provide insights for future research in single-level feature detection. We adopt the recent proposed tool TIDE to compare YOLOF with DETR . As illustrated in Figure 7, DETR has a larger error in localization than YOLOF, which may be related to its regression mechanism. DETR regresses objects in a total anchor free manner and predicts the location globally in the image, which causes difficulties in localization. In contrast, YOLOF relies on pre-defined anchors, which is responsible for higher missing error than DETR in the predictions. According to the analysis in Section 4.2, the anchors of YOLOF are sparsely and not flexible enough in the inference stage. Intuitively, there are situations that there are no high-quality anchors pre-defined around a ground-truth box. Thus, introducing the anchor-free mechanism into YOLOF may help alleviate this problem, and we leave it for future work.
Conclusion
In this work, we identify that the success of FPN is due to its divide-and-conquer solution to the optimization problem in dense object detection. Given that FPN makes network structure complex, brings memory burdens, and slows down the detectors, we propose a simple but highly efficient method without using FPN to address the optimization problem differently, denoted as YOLOF. We prove its efficacy by making fair comparisons with RetinaNet and DETR. We hope our YOLOF can serve as a solid baseline and provide insight for designing single-level feature detectors in future research.
Appendix A: More Details
In Figure 8, we illustrate the detailed processes of generating outputs in encoders. The four encoders differ in the number of input features and output features. (a) The Multiple-in-Multiple-out (MiMo) encoder receives three levels from the backbone and output five levels. The structure of the MiMo encoder is the same as FPN in RetinaNet . (b) Single-in-Multiple-out (SiMo) only has one C5 feature for the input. As there are no other inputs, we remove the convolution layer designed for C3 and C4. (c) Multiple-in-Single-out (MiSo) receives three input features while only generate one output feature P5. To fully utilize the context in the input features, we adopt a structure similar to PANet in MiSo. (d) In the Single-in-Single-out (SiSo) encoder, we remove all other convolution layers and only keep the convolution layers in the level of C5.
Network Architecture of YOLOF
In Figure 9, we show a detailed network architecture of YOLOF. YOLOF detects objects on single-level feature, which is very simple. Our method consists of three components: the backbone, the encoder, and the decoder. The detailed design of these components are presented in Section 4.3.
Training Time & Memory:
In this section, we compare training time and training memory among YOLOF, DETR , and RetinaNet . As shown in Table 6, due to the long training schedule, DETR needs 112.5 hours to converge on COCO with eight 2080Ti GPUs, while YOLOF and RetinaNet only need 4.5 hours and 9.8 hours, respectively. As for training memory, YOLOF needs less memory than RetinaNet and DETR, which make YOLOF be trained with larger batch size and converge faster.
More Implementation Details:
The default training settings for YOLOF is a total of 64 images per mini-batch (8 images per GPU) with an initial learning rate of . While for ResNeXt-101 , we train with 4 images per GPU (batch size 32) and set the learning rate to following the linear rule . For multi-scale training, DETR apply random crop plus resize to simulate large image size during training. In YOLOF, we simply resize the image to large size. For multi-scale training, we follow HTC and adopt a strategy of random sample the image size between $$ with its largest edge no greater than 1600 pixels.
Detailed Settings to Compare with YOLOv4
To match the performance of YOLOv4, we first increase the number of dilated residual blocks in the dilated encoder from to . We adjust the dilations of these dilated residual blocks according to experimental results. We find that the dilations $0.049\times0.750.83\times0.53\times12\sim 1$ mAP improvement).
Appendix B: Additional Experimental Results
In RetinaNet , anchors are generated from multiple level features (P3-P7) with areas of to , respectively. At each level feature, RetinaNet paves anchors with sizes and aspect ratios . While in YOLOF, we only have a one-level feature to place anchors. To cover all objects’ scales, we add anchors with areas of , size , and aspect ratio in the single feature map, resulting in 5 anchors in each position. Moreover, we investigate the influence of more anchors in YOLOF. Following RetinaNet, we generate 45 anchors in each position with different sizes () and more aspect ratios (). All results are shown in Table 7. The results show that adding more aspect ratios does not change the performance of YOLOF, while the performance drops with more sizes. Thus, we choose to add a minimum of five anchors for YOLOF by default.
Hyper-parameter of ATSS:
Here, we provide the results of using different values of in ATSS in Table 8. The results show that the choice of used in the original paper is not the best choice in YOLOF. According to the results, we choose for ATSS in this paper.
Results with Dilated C5:
In this paper, we show that YOLOF performs well on the C5 feature. To boost the performance of YOLOF, we detect objects on a feature map with higher resolution than the C5 feature. Following DETR , we construct a backbone with dilation and without stride on its last stage. The backbone’s output feature is denoted as DC5, with a downsample rate of 16. In Table 9, we show the results of YOLOF-DC5 on COCO val split with ResNet-50 and ResNet-101 as the backbone. YOLOF-DC5 achieves higher performance than the original YOLOF but runs at a slower speed as the feature’s resolution is larger than C5. To achieve the results, we first add a smaller anchor, resulting in anchors per location (), then we increase the from to and change the ignore threshold for positive anchors from to . Other parameters are the same as before.
Acknowledgements
This work is supported by The National Key Research and Development Program of China (No. 2017YFA0700800), Beijing Academy of Artificial Intelligence (BAAI), National Natural Science Foundation of China (No.61972396, 61876182, 61906193), National Key Research and Development Program of China (No. 2020AAA0103402), the Strategic Priority Research Program of Chinese Academy of Sciences (No. XDB32050200), and The NSFC-General Technology Collaborative Fund for Basic Research (Grant No. U1936204).