Hit-Detector: Hierarchical Trinity Architecture Search for Object Detection

Jianyuan Guo, Kai Han, Yunhe Wang, Chao Zhang, Zhaohui Yang, Han Wu, Xinghao Chen, Chang Xu

Introduction

Object detection is a fundamental task in computer vision and has been widely applied in the real world, such as autonomous vehicles and surveillance video. The advancement of deep learning results in a number of convolutional neural network based solutions of object detection task. Typically, deep learning based detectors can be divided into two categories: (i) one-stage methods including YOLO and SSD which directly utilize CNNs to predict the bounding boxes of interest; and (ii) two-stage approaches such as Faster R-CNN that generates the bounding boxes after extracting region proposals upon a region proposal network (RPN). The advantage of single-stage methods lies in the high detection speed whereas two-stage methods dominate in detection accuracy.

A series of either one-stage or two-stage approaches have been developed to continuously boost the detection speed and accuracy. However, the manually designed architectures heavily rely on the expert knowledge while still might be suboptimal. Thus neural architecture search (NAS) that automates the design of network architectures and minimizes human labor has drawn much attention and made impressive progress, especially in image classification tasks . Compared with classification tasks that simply determine what the image is, detection tasks need to further figure out where the objects are. NAS for object detection therefore requires more careful design and is much more challenging.

Modern object detection systems usually consist of four components: (a) backbone for extracting semantic features, e.g. ResNet-50 and ResNeXt-101 ; (b) neck for fusing multi-level features, e.g. feature pyramid networks (FPN) ; (c) RPN for generating proposals (usually in two-stage detector); and (d) head for object classification and bounding box regression. Recently, there are works that explore NAS in object detection tasks to search for a good architecture of the backbone or FPN . With the searched backbone or FPN architectures, these works have achieved higher accuracy than the manually designed baselines with similar numbers of parameters and FLOPs.

Nevertheless, exploiting only one part of the detector at a time cannot fulfill the the potential of each component, and the separately searched backbone and neck may not be optimal or compatible with each other. As shown in Table 1, NAS-FPN for neck searching achieves a 38.9%38.9\% mAP which is higher than that of vanilla FPN, and DetNAS for backbone searching outperforms vanilla ResNet-50 backbone with 40.2%40.2\% mAP. However, a straightforward combination of NAS-FPN and DetNAS leads to a worse mAP, i.e., 39.4%39.4\%, let alone outperforms two models. This insightful observation motivates us to take the detector as a whole in NAS.

In this paper, we propose to simultaneously search all components of the detector in an end-to-end manner. Due to the differences among optimum space for each component and the difficulty in optimizing within large search space, we introduce a hierarchical way to mine the proper sub search space from the large volume of operation candidates. In particular, our proposed Hit-Detector framework consists of two key procedures as shown in Fig. 1. First, given a large search space containing all the operation candidates, we screen out the customized sub search space suitable for each part of detector with the help of group sparsity regularization. Secondly, we search the architectures for each part within the corresponding sub search space by adopting the differentiable manner. Extensive experiments demonstrate that our Hit-Detector achieves state-of-the-art results on the benchmark dataset, which validates the effectiveness of the proposed method.

Our main contributions can be summarized as follows:

This is the first time that architectures of backbone, neck and head are searched altogether in an end-to-end manner for object detection.

We show that different parts prefer different operations, and propose a hierarchical way to specify appropriate sub search space for different components in detection system to improve the sampling efficiency.

Our Hit-Detector outperforms either hand-crafted or automatically searched networks by a large margin with much less computational complexity.

Related Work

Object detection aims at determining what and where the object is when given an image. Riding the wave of convolutional neural networks, noticeable improvements in accuracy have been made in both one-stage and two-stage detectors. Generally, object detector consists of four parts: a backbone that extracts features from input images, a neck attached to backbone that fuses multi-level features, a region proposal network which generates prediction candidates on extracted featuresThere are only three parts (backbone, neck, and head) for one-stage detector. We take two-stage detector as an example here., and a head for classification and localization.

In the past few years, various methods in the literature have been proposed to tackle with detection task and attain significant progress. With NAS prospering automating design of model architecture, it has also boosted the probe into automatically searching for the best architecture for object detection, other than manually design. Here we briefly review some of the recent detectors in two dimensions:

The manual design of detectors in mainstream evolution of object detection solutions is promoted by several works. R-CNN is the first to show that a CNN could lead to dramatic performance improvement in object detection. Selective Search is used to generate proposals and SVM is applied to classify each region. Following R-CNN, Fast R-CNN is proposed to improve the speed by sharing computation of convolutional layers between proposals. Faster R-CNN replaces Selective Search with a novel RPN (region proposal network), further promoting the accuracy and makes it possible to train model in an end-to-end manner. In addition, Mask R-CNN extends Faster R-CNN mainly in instance segmentation task. Meanwhile, a series of proposal free detectors, i.e. one-stage detectors, has been proposed to speed up the detection. YOLO and YOLOv2 extract features from input images straightly for predicting bound boxes and associated class probabilities through a unified architecture. SSD further improves the mAP by predicting a set of bounding boxes from several feature maps with different scales.

In the mean time, some arts concentrate on improving specific parts such as backbone, neck, and head to improve the efficiency in object detectors. DetNet specifically designs a novel backbone network for object detection. FPN develops a top-down architecture to effectively encode features from different scales. PANet further modifies neck module to obtain a better fusion. Focal loss is proposed to solve the problem of class imbalance. MetaAnchor proposes a flexible mechanism which generates anchor from arbitrary prior boxes. Light-Head R-CNN designs a light head for two-stage detector to decrease the computation cost accordingly.

2 Neural architecture search

NAS (neutral architecture search) has attracted great attention recently. Several reinforcement learning based methods train a RNN controller to generate cell structure and form the network accordingly. Evolutionary algorithm based methods are also proposed to update architectures by mutating current ones. To speed up searching process, gradient based methods are proposed for continuous relaxation of search space, which allow the differentiable optimization in architecture search.

NAS for object detection

In addition to NAS works on classification, some recent works attempt to develop NAS for object detector. NATS claims that the effective receptive field of backbone is critical and uses NAS to search different dilation rates for each convolution layer in backbone. Similarly, DetNAS aims to search a better backbone for detection task. NAS-FPN targets at a better architecture of feature pyramid network for object detection, adopting NAS to discover a new feature pyramid architecture covering all cross-scale connections. Auto-FPN sequentially searches a better fusion of the multi-level features for neck and a better structure for head. However, there are two shortcomings in above methods: (i) the search space is defined by human prior and might be too naive for searching (e.g. four choices based on ShuffleNetV2 in DetNAS ); (ii) in each work, only one certain part is searched (e.g. backbone in , neck in ) can lead to suboptimal result in detection task. To tackle these two challenges, we propose Hit-Detector that filters proper search space for each parts hierarchically and searches every parts for a better detector in an end-to-end manner.

Hit-Detector

In this section, we introduce the proposed Hierarchical Trinity architecture search algorithm for object detection and the resulted detector, i.e. Hit-Detector. We first identify and analyze the problems in current NAS algorithms for object detection to clarify our motivation that we need to search all components together. Then we detail how to hierarchically filter the sub search space for each parts, and last, the end-to-end search process for Hit-Detector is depicted. The following statement of search algorithm is based on two-stage detection methods and can be easily applied to one-stage methods.

Two-stage detection system can be decoupled as four components: (i) Backbone. Commonly used backbones in detection system such as ResNet and ResNeXt are mostly manually designed for classification tasks. Usually, the largest proportion of parameters in a detector comes from backbone. For example, ResNet-101, the backbone of FPN , takes up 71% parameters of all, leaving a large potential for searching; (ii) Neck. Employing in-network feature pyramids to approximate different receptive fields can help detector localize objects better. Previous architectures of feature fusion neck are manually designed while NAS can exploit better operations for merging features from different scales; (iii) RPN. The typical RPN is lightweight and efficient: a convolution layer followed by two fully connected layers for region proposal classification and bounding box regression. We follow this design for its efficiency and do not search for this part; and (iv) Head. Detectors usually have a heavy head attached to former network. For example, Faster R-CNN employs the 5-th stage in ResNet and FPN uses two large fully connected layers (34% parameters of whole detector) to execute classification and regression, which are inefficient for detection. We call the backbone, the neck, and the head which are valuable to be searched as detector trinity in this paper.

Recently, several methods have been proposed to search either backbone or feature pyramid architecture for object detection. Denoting the search space A\mathcal{A} as a directed acyclic graph (DAG) where the nodes indicate the features while directed edges are associated with multifarious operations such as convolution layer and pooling layer. We can view each path from start point to end point in the graph as an architecture α∈A\alpha\in\mathcal{A}. Previous NAS works for object detection can be formulated as the optimization problem:

The purpose of NAS process is to find a specific architecture α∈A\alpha\in\mathcal{A} that minimizes the validation loss Lvaldet(α,wα∗)\mathcal{L}^{det}_{val}(\alpha,w^{*}_{\alpha}) with the trained weights wα∗w^{*}_{\alpha}. The above formulation can represent searching on backbone (e.g. A\mathcal{A} denotes the backbone search space) or feature pyramid network (e.g. A\mathcal{A} denotes the FPN search space).

However, the backbone, the neck (feature fusion network), and the head in object detection system should be highly consistent to each other. Only redesigning the backbone or neck is not enough, which can lead to suboptimal results. As shown in Table 1, combining the separately searched backbone in DetNAS and neck in NAS-FPN leads to a worse result. We claim that separately search backbone α\alpha, neck β\beta, and head γ\gamma in detector is inferior to search all these components end-to-end:

where α′,β′,γ′\alpha^{\prime},\beta^{\prime},\gamma^{\prime} are obtained by solving corresponding optimization problem as in Eq. 1 separately, while α∗,β∗,γ∗\alpha^{*},\beta^{*},\gamma^{*} are optimized through the end-to-end searching algorithm, and Ab,An,Ah\mathcal{A}_{b},\mathcal{A}_{n},\mathcal{A}_{h} are search spaces for backbone, neck, and head, respectively.

In this paper, we propose to search the trinity in detectors, i.e. backbone, neck, and head in an end-to-end manner:

As shown in Figure 1, we propose the hierarchical trinity architecture search framework to solve problems in Eq. 3, which includes two procedures: screening sub search space for each component and searching the trinity end-to-end.

2 Screening Sub Search Space

Search space is one of the key factors in neural architecture search. The search spaces in are manually designed with a limited number of operation candidates. Furthermore, there is usually the case that manually designed search space is not suitable for the architecture that needs to be optimized. In order to prevent insufficient search space, we start from the FBNet and traverse as many candidates as possible to form a large set of operation candidates, which are illustrated in Figure 2. Inverted residual block contains a 1×11\times 1 convolution, a k×kk\times k depthwise convolution, another 1×11\times 1 convolution and an expansion factor ee. Separable block contains a k×kk\times k depthwise convolution and a 1×11\times 1 convolution. If the output dimension is the same as input’s, we use a skip connection to add them together.

As shown in Figure 1(a), the whole search space consists of NN different operation candidates, e.g., N=32N=32 in our experimental setting. If we directly apply this large search space for NAS, the memory and computational overhead is so heavy that ordinary hardware cannot support the search process efficiently. Moreover, in Sec 4.4, we empirically show that the same operation can have different impacts on final result at different parts in detection system. In order to find the most suitable search space for each component and reduce the computational burden, we propose a screening scheme to hierarchically filter operation candidates for each component. As shown in Figure 1(b), every candidate is associated with a score. The candidates with higher scores are retained, and others with lower scores are deleted.

μ\mu is a trade-off hyper-parameter. In screening stage, the architecture parameter α\alpha is learned and the scores of candidates are ranked accordingly, the last several candidates will be removed from search space gradually until a search space with size NbN_{b} is obtained. The sub search spaces for backbone, neck, and head are hierarchically selected from the whole search space, namely, Ob\mathcal{O}_{b}, On\mathcal{O}_{n}, and Oh\mathcal{O}_{h}.

3 End-to-end Search in Hit-Detector

After obtaining the suitable sub search space for each component, we start the end-to-end search for object detector. We adopt the differentiable manner proposed in to solve Eq. 3, and represent the sub search space by a stochastic supernet. During searching, each intermediate node is computed as a weighted sum based on all candidates. For backbone, the ll-th node is formulated as

where xlx_{l} is the output of the ll-th layer, αlo\alpha_{l}^{o} is the parameter for operation o(⋅)o(\cdot), and Ob\mathcal{O}_{b} indicates the sub search space for backbone. The nodes in neck and head are similar to Eq. 5. This continuous relaxation makes the entire framework differentiable to both operation weights and architecture parameters, so that we can perform architecture search in an end-to-end manner. In testing phase, we can easily decode architectures from α,β,γ\alpha,\beta,\gamma by choosing operation with the highest score in each layer and construct detector using the selected operation.

The details of our detector supernet are described as follows. The basic structure of backbone consists of a 3×\times3 convolution head with stride of 2 followed by four stages that contain 4+4+8+4=204+4+8+4=20 blocks to be searched. Conventionally, we define the layers that producing feature map with the same spatial size belong to the same network stage. In each stage, the first block has a stride of 2 for downsampling, and the last feature map generated by the backbone has a downsample rate of 32 compared to the input image. The channel of each stage is set to {48,96,256,352}\{48,96,256,352\} respectively. We use {C1,C2,C3,C4}\{C_{1},C_{2},C_{3},C_{4}\} to represent feature maps of stride {4,8,16,32}\{4,8,16,32\} generated by backbone. Then we send the feature pyramid into the neck. Generally, high level features have better semantic information while low level features have more accurate location information. To propagate semantic signals and accurate information about localization to features from lower and higher level, we use both top-down and bottom-up path augmentation to enhance features from different levels inspired by . We use {P1,P2,P3,P4}\{P_{1},P_{2},P_{3},P_{4}\} and {N1,N2,N3,N4}\{N_{1},N_{2},N_{3},N_{4}\} to denote feature maps after top-down and bottom-up path, respectively. The whole neck contains 4+4=84+4=8 lateral connections to be searched, thus each feature level can search for different operations to assign proper receptive fields. Given the aligned feature maps generated by region proposal network, head is applied to predict final classification and refine the bounding box of the object. We design 44 blocks to be searched followed by a fully connected layer to form the detection head. The number of the output channels for each level in neck and head is 256 and the output of the fully connected layer is a 512 dimensional vector. In our experiments, we set Nb=Nn=Nh=8{N_{b}}={N_{n}}={N_{h}}=8. Even if we extract the sub search spaces carefully, the final search space contains 8(20+8+4)≈7.9×10288^{(20+8+4)}\approx 7.9\times 10^{28} possible architectures.

4 Optimization

In order to control the computational cost of the searched detector at the same time, we add the FLOPs constraint as a regularization term in the loss function and rewrite the Eq. 3 as:

where λ\lambda is the coefficient to balance accuracy and cost of the detector, C(α){\rm C}(\alpha) indicates FLOPs of the backbone part and can be decomposed as linear sum of each operations:

C(β){\rm C}(\beta) and C(γ){\rm C}(\gamma) can be calculated similarly.

It is clear that continuous relaxation from Eq. 5 is differentiable with respect to architecture parameters and operation weights, thus {α,β,γ,w}\{\alpha,\beta,\gamma,w\} can be optimized jointly using stochastic gradient descent. We adopt first-order approximation following and update architecture parameters and operation weights alternately. We first fix {α,β,γ}\{\alpha,\beta,\gamma\} and compute ∂L/∂w\partial\mathcal{L}/\partial w to train the network weights on 50% training data, then we fix the network weights and calculate ∂L/∂α\partial\mathcal{L}/\partial\alpha, ∂L/∂β\partial\mathcal{L}/\partial\beta, and ∂L/∂γ\partial\mathcal{L}/\partial\gamma to update the architecture parameters on the remaining 50% training data, where the loss function L\mathcal{L} is the localization and classification losses calculated on the detection mini-batch. The optimization alternates until the supernet converges.

Experiments

In this section, we investigate the effectiveness of proposed Hit-Detector by conducting elaborate experiments on COCO benchmark.

We conduct experiments on MS COCO 2014 dataset , which contains 80 object classes. Following , our training set is the union of 80k training images and 35k subset of validation images (trainval35k), and validation set is the remaining 5k validation images (minival). We consider Average Precision with different IoU thresholds from 0.5 to 0.95 with an interval of 0.05 as evaluation metric, i.e., mAP\rm mAP, AP50\rm{AP}_{50}, AP75\rm{AP}_{75}, APS\rm{AP}_{S}, APM\rm{AP}_{M} and APL\rm{AP}_{L}. The last three measure performance w.r.t. objects with different scales.

2 Implementation Details

Our implementation is based on mmdetection [mmdetection] with Pytorch framework . We firstly filter three sub search spaces for backbone, neck, and head, respectively. Then we search for our Hit-Detector following the algorithm in sec 3.4. Finally we train our searched model on the training set mentioned above. We use only horizontal flipping as data augmentation for training, and there is no data augmentation for testing. The experiments are conducted on 8 V100 GPUs.

We sequentially screen sub search spaces for different parts of the detector due to the GPU memory constraint. The number of operations in the whole search space O\mathcal{O} is 32 as shown in Fig. 2, and the number of operations to be screened for all three sub search spaces Ob\mathcal{O}_{b}, On\mathcal{O}_{n}, and Oh\mathcal{O}_{h} is set as 8. We halved the depth of backbone supernet to 2+2+4+2=102+2+4+2=10 for simplifying this process. We first pretrain the backbone supernet on ImageNet for 10 epochs with fixed architecture parameters, then fine-tune the entire detector supernet on COCO. SGD optimizer with momentum 0.9 and cosine schedule with initial learning rate 0.04 is used for learning model weights, while Adam optimizer with learning rate 0.0004 is adopted for updating architecture parameters. The μ\mu in Eq. 4 is empirically set to 0.1. We begin to optimize architecture parameters at the 6-th epochs and finish searching at 12 epochs, and we don’t use resource constraint in this stage.

Trinity Architecture searching.

After screening sub search spaces, we start to search the backbone, neck and head in an end-to-end manner. We pretrain the new backbone supernet on ImageNet based on the corresponding sub search space, and then search the detector on COCO dataset. The optimizers to learn the architecture parameters and weights are the same as we do when screening sub search spaces. The λ\lambda in Eq. 6 is empirically set to 0.01 to trade off the accuracy and the FLOPs constraint.

Training details.

We first pretrain the searched backbone on ImageNet for 300 epochs, then we fine-tune the whole detector on COCO training set. The input image is resized such that its shorter side has 800 pixels. We use SGD optimizer with a batch size of 4 images per GPU, and our model is trained for 12 epochs, known as 1×\times schedule. The initial learning rate is 0.04 and is divided by 10 at the 8-th and 11-th epoch. We set momentum as 0.9 and weight decay as 0.0001.

3 Main Results

FPN with ResNet-50 as backbone is the baseline model here. We replace the backbone in FPN with other excellent backbones, i.e. MobileNetV2 and ResNeXt-101 , and form two competitor models accordingly. As shown in Table 2, Hit-Detector surpasses baseline by a large margin with much less parameters. In addition, Our method is 1.1% higher on mAP compared to the ResNeXt based detector with less than one half of the parameters. And we outperform MobileNetV2 by 11.3% on mAP with only a bit more parameters. This demonstrates that our method can find a better architecture than hand-crafted baselines.

Comparisons with NAS based methods.

As shown in Table 2, we compare our method with both detector searched on COCO benchmark and detector that adopts NAS based model as backbone. FBNet is searched on ImageNet dataset and we directly apply it as the backbone of a detector, however, its performance on detection task is disappointing. DetNAS aims to search a better backbone directly on detection benchmark while leaves the rest parts unchanged. Our method outperforms DetNAS by 1.2% with less parameters and pretty much FLOPs. It can be seen that Hit-Detector surpasses all previous NAS based methods, which indicates that it is important to search the trinity in an object detector.

Comparison with State-of-the-arts on test-dev set.

We also compare the results of our Hit-Detector with other state-of-the-art methods on the COCO test-dev, and we summarize the comparisons in Table 3. Hit-Detector only applies horizontal flipping as data augmentation and 2x training scheme, achieves 44.5% mAP without bells and whistles. Our model has less parameters and performs even better compared to other detectors such as TridentNet and NAS-FPN. This demonstrates that our method can find a better architecture than hand-crafted or partly searched methods.

4 Ablation Studies

We use a toy example to demonstrate that different parts of the detector are sensitive to different operations, so that different parts need different sub search spaces. Here we simply choose 12 operations from the whole search space as shown in Figure 2. For example, “conv_k3_d1” denotes the convolution block with kernel size 3 and dilation rate 1, and “ir_k3_d1” denotes the inverted residual block with kernel size 3 and dilation rate 1, respectively. Take the left figure as an example, for each operation candidate, we randomly select a layer from backbone and replace original operation with the operation candidate to explore the influence of the selected operation candidate. For each part (backbone, neck, and head), we repeat the random process 6 times to ensure that operation candidate can be inserted in layers of different depths. We can find that (i) different parts prefer different operations. For example, operation “conv_k5_d3” achieves the highest mAP in backbone while “conv_k5_d1” performs the best in head; (ii) one operation has different performances in different parts, “conv_k3_d1” attains 30.4%-30.6% mAP in backbone but gets 30.7%-30.9% mAP in neck; and (iii) the performance of one operation within the same part is stable enough so that it’s reasonable for us to screen sub search spaces by different parts.

Column-sparse regularization.

To further evaluate the influence of the column-sparse regularization, we set the μ\mu in Eq. 4 to {0,0.01,0.1}\{0,0.01,0.1\} and randomly choose 4 operation candidates to draw the Figure 4. We can find that if there is no column-sparse regularization (μ=0\mu=0), probabilities of different operations are rather similar, which makes the screening process unstable. As μ\mu increases, differences between operations become more significant, which helps to screen proper sub search space easier. We set μ=0.1\mu=0.1 in the rest of our experiments.

Importance of screening sub search space.

We study the influence of screening different sub search spaces through the ablation study shown in Table 4. When three parts have the same sub search space, the mAP decreases from 41.4% to 40.1%, which indicates that different parts need to have their proper sub search space to achieve better performance.

Importance of trinity architecture searching.

We explore the impact of each searched part in Hit-Detector here. Taking ResNet-50 based FPN as baseline model, we replace one part of baseline model with our searched component each time and verify the new model on COCO minival. As shown in Table 5, our searched backbone architecture achieves 39.2% mAP, and the searched head also attains 38.5% mAP. With the searched trinity, the mAP of Hit-Detector increases to 41.4%, which is much higher than baseline model and the single part competitor. Another noteworthy finding is that the backbone and head can boost performance better than neck. We think the main reasons are that (i) object detection emphasizes more on perception to the location of each object in an image, thus the backbone designed for detection can perform better compared to the one designed for classification; and (ii) head aims to identify and refine the location of bounding boxes, thus searching more suitable convolution layers in head can bring more benefits to detection task.

Our searched Hit-Detector is depicted in Figure 5. We observed that the backbone prefers operations that have large kernel size. However, convolution layer with dilation 3 which achieves the best mAP in Figure 3 is not chosen. One main reason is that the ResNet-50 backbone in toy example only includes 3×33\times 3 convolution, thus operation with dilation 3 can boost the performance prominently. Except that, backbone and head prefer operations with bigger expansion rate such as inverted residual block which has more intermediate channels to increase the expression of feature, while the neck prefers bigger dilation rate for larger receptive fields.

Extension to one-stage detector.

To evaluate the generalization of our method, we apply it to RetinaNet for searching one-stage detector. The search algorithm is the same as mentioned in sec 3.4. As shown in Table 6, our model outperforms RetinaNet by 1.3% in terms of mAP and the model size is smaller than both VGG based SSD and ResNet-50 based RetinaNet.

Conclusion

In this work, we propose a hierarchical trinity architecture search scheme to address the problem that incomplete searching of detector would cause the inconsistency between different components and lead to the suboptimal performance. We reveal that different components prefer different operations and thus screen three sub search spaces to improve the searching efficiency. Then we search for all components of object detectors based on the sub search spaces in an end-to-end manner. The searched architecture, namely Hit-Detector, achieves state-of-the-art performance on COCO benchmark without bells and whistles. Acknowledgement This work of J. Guo and C. Zhang is supported by the National Nature Science Foundation of China under Grant 61671027 and the National Key R&D Program of China under Grant 2017YFB1002400, and C. Xu is supported by the Australian Research Council under Project DE180101438.

Appendix

In this supplementary material, we list the search space and corresponding sub search spaces for detector trinity in details, and we show the qualitative results of our Hit-Detector compared with other state-of-the-art methods.

The whole search space consists of N=32N=32 different operation candidates in our experimental setting. We list the candidates bellow:

where “ir”, “sep” and “conv” indicates the inverted residual block, separable block and convolution block, respectively, “k” indicates the kernel size, “d” indicates the dilation rate, “e” indicates the expansion rate of inverted residual block.

2 Sub search space

In our experiments, we set Nb=Nn=Nh=8N_{b}=N_{n}=N_{h}=8, and we list the top-8 candidate operations for backbone:

3 Qualitative results

We show the qualitative results of our Hit-Detector. We randomly sample some images from the COCO minival and show detection results with confidence bigger than 0.5. First column is the results of FPN , second column is the results of NAS-FPN implemented by ourselves, third column is the results of DetNAS , and the fourth column is the results of our Hit-Detector. Different box colors indicate different object categories.

References