Learning Spatial Fusion for Single-Shot Object Detection

Songtao Liu, Di Huang, Yunhong Wang

Introduction

Object detection is one of the most fundamental components in various downstream vision tasks. In recent years, the performance of object detectors has been remarkably improved thanks to the rapid development of deep convolutional neural networks (CNNs) and well-annotated datasets . However, handling multiple objects across a wide range of scales still remains a challenging problem. To achieve scale invariance, recent state-of-the-art detectors construct feature pyramids or multi-level feature towers .

The Single Shot Detector (SSD) is one of the first attempts to generate convolutional pyramidal feature representations for object detection. It reuses the multi-scale feature maps from different layers computed in the forward pass to predict objects of various sizes. However, this bottom-up pathway suffers from low accuracies on small instances as the shallow-layer feature maps contain insufficient semantic information. To address the disadvantage of SSD, Feature Pyramid Network (FPN) sequentially combines two adjacent layers in feature hierarchy in the backbone model with a top-down pathway and lateral connections. The low-resolution, semantically strong features are up-sampled and combined with high-resolution, semantically weak features to build a feature pyramid that shares rich semantics at all levels. FPN and other similar top-down structures are simple and effective, but they still leave much room for improvement. Indeed, many recent models with advanced cross-scale connections show accuracy gains through strengthening feature fusion. Besides the manually designed fusion structures, NAS-FPN applies Neural Architecture Search (NAS) techniques to pursue a better architecture, producing significant improvements upon many backbones.

Although these advanced studies deliver more powerful feature pyramids, they still leave room for scale-invariant prediction. Some evidences are given by SNIP , which adopts a scale normalization method that selectively trains and infers the objects of appropriate sizes in each image scale of the multi-scale image pyramids, achieving further improvements on the results of pyramidal feature based detectors with multi-scale testing. However, image pyramid solutions sharply increase the inference time, which makes them not applicable to real-world applications.

Meanwhile, compared to image pyramids, one main drawback of feature pyramids is the inconsistency across different scales, in particular for single-shot detectors. Specifically, when detecting objects with feature pyramids, a heuristic-guided feature selection is adopted: large instances are typically associated with upper feature maps and small instances with lower feature maps. When an object is assigned and treated as positive in the feature maps at a certain level, the corresponding areas in the feature maps of other levels are viewed as background. Therefore, if an image contains both small and large objects, the conflict among features at different levels tends to occupy the major part of the feature pyramid. This inconsistency interferes gradient computation during training and downgrades the effectiveness of feature pyramids. Some models adopt several tentative strategies to deal with this problem. set the corresponding areas of feature maps at adjacent levels as ignore regions (i.e. zero gradients), but this alleviation may increase inferior predictions at the the adjacent levels of features. TridentNet creates multiple scale-specific branches with different receptive fields for scale-aware training and inference. It breaks away from feature pyramids to avoid inconsistency, but also misses reusing its higher-resolution maps, limiting the accuracy of small instances.

In this paper, we propose a novel and effective approach, named adaptively spatial feature fusion (ASFF), to address the inconsistency in feature pyramids of single-shot detectors. The proposed approach enables the network to directly learn how to spatially filter features at other levels so that only useful information is kept for combination. For the features at a certain level, features of other levels are first integrated and resized into the same resolution and then trained to find the optimal fusion. At each spatial location, features at different levels are fused adaptively, i.e., some features may be filter out as they carry contradictory information at this location and some may dominate with more discriminative clues. ASFF offers several advantages: (1) as the operation of searching the optimal fusion is differential, it can be conveniently learned in back-propagation; (2) it is agnostic to the backbone model and it is applied to single-shot detectors that have a feature pyramid structure; and (3) its implementation is simple and the increased computational cost is marginal.

Experiments on the COCO benchmark confirm the effectiveness of our method. We first adopt the recent advanced training tricks and anchor-guiding pipeline to provide a solid baseline for YOLOv3 (i.e., 38.8% mAP with 50 FPS). We then employ ASFF to further improve this enhanced YOLOv3 and another strong single-stage detector, RetinaNet , equipped with different backbones by a large margin, while keeping computational cost under control. Especially, we boost the YOLOv3 baseline to 42.4% mAP with 45 FPS and 43.9% mAP with 29 FPS, which is a state-of-the-art speed and accuracy trade-off among all the existing detectors on COCO.

Related Work

Feature pyramid representations or multi-level feature towers are the basis of solutions of multi-scale processing in recent object detectors. SSD is one of the first attempts to predict class scores and bounding boxes from multiple feature scales in a bottom-up manner. FPN builds feature pyramid by sequentially combining two adjacent level of features with top-down pathway and lateral connections. Such connections effectively enhance feature representations and the rich semantics from depp and low-resolution features are shared at all levels.

Following FPN, many other models with similar top-down structures appear, which achieve substantial improvements for object detection. Recently, more advanced investigations have attempted to ameliorate such multi-scale feature representations. For instance, PANet proposes an additional bottom-up pathway based on FPN to increase the low-level information in deep layers. Chen et al. build a pyramid based on SSD that weaves features across different level of feature layers. DLA introduces iterative deep aggregation and hierarchical deep aggregation structures to better fuse semantic and spatial information. Kim et al. show a parallel feature pyramid by adopting spatial pyramid pooling and widening the network. Zhu et al. present a feature selective anchor-free module to dynamically choose the most suitable feature level for each instance. Kong et al. aggregate feature maps at all scales to a specific scale and then produce features at each scale by a global attention operation on the combined features. Libra R-CNN also integrates features at all levels to generate more balanced semantical features. In addition to manually designing the fusion structure, NAS-FPN applies the Neural Architecture Search algorithm to seek a more powerful fusion architecture, delivering the best single-shot detector.

In spite of competitive scores, those feature pyramid based methods still suffer from the inconsistency across different scales, which limits the further performance gain. To address this, set the corresponding regions of adjacent levels as ignore regions (i.e. zero gradients), but the relaxation in the adjacent levels tends to cause more inferior predictions as false positives. TridentNet drops out the structure of feature pyramids and creates multiple scale-specific branches with different receptive fields to adopt scale-aware training and inferencing, but the performance of small instances may suffer from the missing of its higher-resolution maps.

The proposed ASFF approach alleviates this issue by learning connections among different feature maps. Actually, this idea is not new within the domain of computer vision. adopts element-wise product in two adjacent feature maps to form a gate unit in a top-down manner for dense label prediction. The element-wise product operation reduces the categorical ambiguity from shallow layers and highlights the discriminability from deeper layers. This gating mechanism succeeds in semantic segmentation. However, the task of dense labeling does not need heuristic-guided feature selection required in object detection, since the features at all levels predict the same label map at different scales. It thus does not reduce spatial contradiction in object detection. proposes a sigmoid gating unit in the skip connection between convolutional and deconvolutional layers of features at each single level for visual counting. It optimizes the flow of information within the feature maps at the same level, but does not deal with the inconsistency in feature pyramids. ACNet employs a flexible way to switch global and local inference in processing the feature representations by adaptively determining the connection status among the feature nodes from CNNs, classical multi-layer perceptron and non-local network. In contrast to them, ASFF adaptively learns the import degrees for different levels of features on each location to avoid spatial contradiction.

Method

In this section, we instantiate our adaptively spatial feature fusion (ASFF) approach by showing how it works on the single-shot detectors with feature pyramids, such as SSD , RetinaNet , and YOLOv3 . Taking YOLOv3 as an example, we apply ASFF to it and demonstrate the resulting detector in the following steps. First, we push YOLOv3 to a baseline, much stronger than the origin , by adopting the recent advanced training tricks and anchor-free pipeline . Then, we present the formulation of ASFF and give a qualitative analysis of the consistency property of the pyramid feature fusion and ASFF. Finally, we display the details of training, testing, and implementing the models.

We take the YOLOv3 framework because it is simple and efficient. In YOLOv3, there are two main components: an efficient backbone (DarkNet-53) and a feature pyramid network of three levels. A recent work significantly improves the performance of YOLOv3 without modifying network architectures and bringing extra inference cost. Moreover, a number of studies indicate that the anchor-free pipeline contributes to considerably better performance with simpler designs. To better demonstrate the effectiveness of our proposed ASFF approach, we build a baseline, much stronger than the origin , based on these advanced techniques.

Following , we introduce a bag of tricks in the training process, such as the mixup algorithm , the cosine learning rate schedule, and the synchronized batch normalization technique . Besides those tricks, we further add an anchor-free branch to run jointly with anchor-based ones as does and exploit the anchor guiding mechanism proposed by to refine the results. Moreover, an extra Intersection over Union (IoU) loss function is employed on the original smooth L1 loss for better bounding box regression. More details can be found in the supplemental material.

With these advanced techniques mentioned above, we achieve 38.8% mAP on the COCO 2017 val set at a speed of 50 FPS (on Tesla V100), improving the original YOLOv3-608 baseline (33.0% mAP with 52 FPS ) by a large margin without heavy computational cost in inference.

2 Adaptively Spatial Feature Fusion

Different from the former approaches that integrate multi-level features using element-wise sum or concatenation, our key idea is to adaptively learn the spatial weight of fusion for feature maps at each scale. The pipeline is shown in Figure 2, and it consists of two steps: identically rescaling and adaptively fusing.

We denote the features of the resolution at level ll (l∈{1,2,3}l\in\{1,2,3\} for YOLOv3) as xl\mathbf{x}^{l}. For level ll, we resize the features xn\mathbf{x}^{n} at the other level n (n≠l)n\ (n\neq l) to the same shape as that of xl\mathbf{x}^{l}. Because the features at three levels in YOLOv3 have different resolutions as well as different numbers of channels, we accordingly modify the up-sampling and down-sampling strategies for each scale. For up-sampling, we first apply a 1×11\times 1 convolution layer to compress the number of channels of features to that in level ll, and then upscale the resolutions respectively with interpolation. For down-sampling with 1/21/2 ratio, we simply use a 3×33\times 3 convolution layer with a stride of 2 to modify the number of channels and the resolution simultaneously. For the scale ratio of 1/41/4, we add a 2-stride max pooling layer before the 2-stride convolution.

Adaptive Fusion.

Let xijn→l\mathbf{x}_{ij}^{n\rightarrow l} denote the feature vector at the position (i,j)(i,j) on the feature maps resized from level nn to level ll. We propose to fuse the features at the corresponding level ll as follows:

where yijl\mathbf{y}_{ij}^{l} implies the (i,j)(i,j)-th vector of the output feature maps yl\mathbf{y}^{l} among channels. αijl\alpha^{l}_{ij}, βijl\beta^{l}_{ij} and γijl\gamma^{l}_{ij} refer to the spatial importance weights for the feature maps at three different levels to level ll, which are adaptively learned by the network. Note that αijl\alpha^{l}_{ij}, βijl\beta^{l}_{ij} and γijl\gamma^{l}_{ij} can be simple scalar variables, which are shared across all the channels. Inspired by , we force αijl+βijl+γijl=1\alpha^{l}_{ij}+\beta^{l}_{ij}+\gamma^{l}_{ij}=1 and αijl,βijl,γijl∈\alpha^{l}_{ij},\beta^{l}_{ij},\gamma^{l}_{ij}\in, and define

Here αijl\alpha^{l}_{ij}, βijl\beta^{l}_{ij} and γijl\gamma^{l}_{ij} are defined by using the softmax function with λαijl\lambda^{l}_{\alpha_{ij}}, λβijl\lambda^{l}_{\beta_{ij}} and λγijl\lambda^{l}_{\gamma_{ij}} as control parameters respectively. We use 1×11\times 1 convolution layers to compute the weight scalar maps λαl\mathbf{\lambda}^{l}_{\alpha}, λβl\mathbf{\lambda}^{l}_{\beta} and λγl\mathbf{\lambda}^{l}_{\gamma} from x1→l\mathbf{x}^{1\rightarrow l}, x2→l\mathbf{x}^{2\rightarrow l} and x3→l\mathbf{x}^{3\rightarrow l} respectively, and they can thus be learned through standard back-propagation.

With this method, the features at all the levels are adaptively aggregated at each scale. The outputs {y1,y2,y3}\{\mathbf{y}^{1},\mathbf{y}^{2},\mathbf{y}^{3}\} are used for object detection following the same pipeline of YOLOv3.

3 Consistency Property

In this section, we analyze the consistency property of the proposed ASFF approach and the other alternatives of feature fusion. Without loss of generality, we focus on the gradient at a certain position (i,j)(i,j) of the unresized feature maps at level 1 x1\mathbf{x}^{1} in YOLOv3. Following the chain rule, the gradient is computed as:

It is worth to note that feature resizing usually uses interpolation for up-sampling and pooling for down-sampling. We thus assume that ∂xij1→l∂xij1≈1\frac{\partial\mathbf{x}_{ij}^{1\rightarrow l}}{\partial\mathbf{x}_{ij}^{1}}\approx 1 for simplicity. Then Eq. (3) can be written as:

For the two common fusion operations used in RetinaNet , YOLOv3 and other pyramidal feature based detectors (i.e. element-wise sum and concatenation), we can further simplify the equation to the following with ∂yij1∂xij1=1\frac{\partial\mathbf{y}_{ij}^{1}}{\partial\mathbf{x}_{ij}^{1}}=1 and ∂yijl∂xij1→l=1\frac{\partial\mathbf{y}_{ij}^{l}}{\partial\mathbf{x}_{ij}^{1\rightarrow l}}=1:

Suppose position (i,j)(i,j) at level 1 is assigned as the center of an object according to a certain scale matching mechanism and ∂L∂yij1\frac{\partial\mathcal{L}}{\partial\mathbf{y}_{ij}^{1}} is the gradient from the positive sample. As the corresponding positions are viewed as background in the other levels, ∂L∂yij2\frac{\partial\mathcal{L}}{\partial\mathbf{y}_{ij}^{2}} and ∂L∂yij3\frac{\partial\mathcal{L}}{\partial\mathbf{y}_{ij}^{3}} are the gradients from negative samples. This inconsistency disturbs the gradient of ∂L∂xij1\frac{\partial\mathcal{L}}{\partial\mathbf{x}_{ij}^{1}} and downgrades the training efficiency of the original feature maps x1\mathbf{x}^{1}.

One typical way to deal with this problem is to set the corresponding positions of the other levels as ignore regions (i.e. ∂L∂yij2=∂L∂yij3=0\frac{\partial\mathcal{L}}{\partial\mathbf{y}_{ij}^{2}}=\frac{\partial\mathcal{L}}{\partial\mathbf{y}_{ij}^{3}}=0) . However, although the conflict in xij1\mathbf{x}_{ij}^{1} is eliminated, the relaxation in yij2\mathbf{y}_{ij}^{2} and yij3\mathbf{y}_{ij}^{3} tends to cause more inferior predictions as false positives at the suboptimal levels.

For ASFF, it is straightforward to calculate the gradient from Eq. (1) and Eq. (4) as follows:

where αij1,αij2,αij3∈\alpha^{1}_{ij},\alpha^{2}_{ij},\alpha^{3}_{ij}\in. With these three coefficients, the inconsistency of gradient can be harmonized if αij2→0\alpha^{2}_{ij}\rightarrow 0 and αij3→0\alpha^{3}_{ij}\rightarrow 0. Since the fusion parameters can be learned by the standard back-propagation algorithm, a well-tuned training process can yield such effective coefficients (see some qualitative results in Figure 3 and Figure 4). Meanwhile, the supervision information of the background in ∂L∂yij2\frac{\partial\mathcal{L}}{\partial\mathbf{y}_{ij}^{2}} and ∂L∂yij2\frac{\partial\mathcal{L}}{\partial\mathbf{y}_{ij}^{2}} is kept, avoiding generating more false positives.

4 Training, Inference, and Implementation

Let Θ\Theta denote the set of network parameters (e.g., the weights of convolution filters) and Φ={λαl,λβl,λγl∣ l=1,2,3}\Phi=\{\mathbf{\lambda}^{l}_{\alpha},\mathbf{\lambda}^{l}_{\beta},\mathbf{\lambda}^{l}_{\gamma}|\ l=1,2,3\} be the set of fusion parameters that control the spatial fusion of each scale. We jointly optimize the two sets of parameters by minimizing a loss function L(Θ,Φ)\mathcal{L}(\Theta,\Phi), where L\mathcal{L} is the original YOLOv3 objective function plus the IoU regression loss for both anchor shape prediction and bounding box regression. Following , we apply mixup on the classification pretraining of DarkNet53, and all the new convolution layers are employed with the MSRA weight initialization method . To reduce the risk of overfitting and improve generalization of network predictions, we follow the approach of random shapes training as in YOLOv3 . More specifically, a mini-batch of NN training images is resized to N×3×H×WN\times 3\times H\times W, where H=WH=W is randomly picked in {320,352,384,416,448,480,512,544,576,608}\{320,352,384,416,448,480,512,544,576,608\}.

Inference.

During inference, the detection header at each level first predicts the shape of anchors, and then conducts classification and box regression following the same pipeline as that in YOLOv3 . Next, non-maximum suppression (NMS) with the threshold at 0.6 is applied to each class separately. For simplicity and fair comparison against other counterparts, we do not use the advanced testing tricks such as Soft-NMS or test-time image augmentations.

Implementation.

We implement the modified YOLOv3 as well as ASFF using the existing PyTorch v1.0.1 framework with CUDA 10.0 and CUDNN v7.1. The entire network is trained with stochastic gradient descent (SGD) on 4 GPUs (NVDIA Tesla V100) with 16 images per GPU. All models are trained for 300 epochs with the first 4 epochs of warmup and the cosine learning rate schedule from 0.001 to 0.00001. The weight decay is 0.0005 and the momentum is 0.9. We also follow the implementation of to turn off mixup augmentation for the last 30 epochs.

Experiments

We perform all the experiments on the bounding box detection track of the challenging MS COCO 2017 benchmark . We follow the common practice and use the COCO train-2017 split (consisting of 115k images) for training. We conduct ablation and sensitivity studies according to the evaluation on the val-2017 split (5k images). For our main results, we report COCO AP on the test-dev split (20k images), which has no public labels and requires uploading detection results to the evaluation server.

We first evaluate the contribution of several elements to our baseline detector for better reference. Results are reported in Table 1, where BoF denotes all the training tricks mentioned in , GA denotes the guided anchoring strategy , and IoU is the additional IoU loss in bounding box regression. From Table 1, we can see that all the techniques contribute to accuracy gain, and thanks to them, we deliver a final baseline which reaches an AP of 38.8%. It is worth to note that the improvement of almost all the components is cost free, as BoF and IoU do not add any additional computation and GA introduces only two 1×11\times 1 convolution layers for each level of feature maps. The final baseline achieves the speed of 50 FPS on a single Graphics Card of NVIDIA Tesla V100.

Effectiveness of Adjacent Ignore Regions.

To avoid gradient inconsistency, some works ignore the corresponding areas on the two adjacent levels of the chosen level for each target, and the ignored area is the same size as that of the positive one in the chosen level. In YOLOv3, only the center location of the chosen area is positive, and we thus ignore the corresponding center location at the two adjacent levels to follow the ignoring rule. Besides, we denote ϵignore\epsilon_{ignore} as the ratio of the widths and lengths of the ignored area to that of the target object area, and carry out some experiments with different values of ϵignore\epsilon_{ignore} to show the effectiveness of the ignoring strategy. Table 2 reports the study results. We can see that the larger ignore area indeed hurt the performance of the detector, by bringing more false positives.

Adaptively Spatial Feature Fusion.

ASFF significantly improves the box AP from 38.8% to 40.6% as shown in Table 3. To be more specific, most of the improvements come from APSAP_{S} and APMAP_{M}, yielding increases of 2.9% and 2.9% respectively compared with corresponding reference scores. It validates that the representation of high-resolution features is largely improved by the proposed adaptively fusion strategy. Moreover, ASFF only incurs 2 ms additional inference time, keeping the detector run efficiently with 46 FPS.

As described in Sec. 3.2, to adaptively fuse the features for each scale, the features at other levels are firstly resized to the same shape before fusion. To make fair comparison, we further report the accuracies of another two common fusion operations (i.e. element-wise sum and concatenation) with resized features in Table 3. We can see in the table that, these two operations improve the accuracy on APSAP_{S} and APMAP_{M} as ASFF does, but they both sharply downgrade the performance on APLAP_{L}. These results indicate that the inconsistency across different levels in feature pyramids brings negative influence on the training process and thus leaves the potential of pyramidal feature representation from being fully exploited.

2 Visual Analysis

In order to understand how the features are adaptively fused, we visualize some qualitative results in Figure 3 and 4. The detection results are in the left column. The heat maps of the learned weight scalars and the fused feature activation maps at each level are in the right column. For fused feature maps, we sum up the values among all the channels to visualize the activation maps. The red numbers near the boxes indicate the fused feature level that detects the object. Note that the actual resolutions of the three levels are different, and we resize them to a uniform size for better visualization.

Specifically, in Figure 3, we investigate how ASFF works when all objects in the image have roughly the same size. It is also worth to note that YOLOv3 only takes the center point of the object in the corresponding feature maps as a positive. For the image in the first row, all the three zebras are predicted from the fused feature maps of level 1. It indicates that their center areas are dominated by the original features of level 1 and the resized features within those areas from level 2 and 3 are filtered out. This filtering guarantees that the features of these three zebras at level 2 and 3 are treated as background and do not receive positive gradients in training. Regarding ASFF, in the fusion process of level 2 and 3, the central areas at the resized features from level 1 are also filtered out, and the original features of level 1 will receive no negative gradients in training. For the image in the second row, all the sheeps are predicted by the fused feature maps of level 3. We zoom in the heat maps of level 3 within the red box for better visualization. In fusion, the features from level 1 are kept in the object areas as they contain stronger semantic information, and the features from level 3 are extracted around each object since they are more sensitive for localization.

In Figure 4, we exhibit the images that have several objects of different sizes. Most of the fusion cases at the corresponding level are similar to the ones in Figure 3. Meanwhile, one may notice that the tennis racket in the second image is predicted from level 1, but the heat maps show that the main features within its central area are taken from the resized feature of level 2. We speculate that although the tennis racket is predicted from level 1 due to heuristic size selection, the features from level 2 are more discriminative in detecting it since they contain richer clues of lines and shapes. Thanks to our ASFF module, the final feature can be adaptively learned from optimal fusion, which contributes in particular to detecting challenging objects. Please see more visual results in the supplementary material.

3 Evaluation on Other Single-Shot Detectors

To better evaluate the performance of the proposed approach, we carry out additional experiments with another representative single-shot detector, namely RetinaNet . First, we directly adopt the official implementation to reproduce the baseline. We then add ASFF behind the pyramid feature maps from P3 to P5 on FPN, similar to Figure 2. As shown in Table 5, ASFF consistently increases the accuracy of RetinaNet with different backbones (i.e. ResNet-50 and ResNet-101).

4 Comparison to State of the Art

We evaluate our detector on the COCO test-dev split to compare with recent state-of-the-art methods in Table 4. Our final model is YOLOv3 with ASFF*, which is an enhanced ASFFversion by integrating other lightweight modules (i.e. DropBlock and RFB ) with 1.5×\times longer training time than the models in Section 4.1. Keeping the high efficiency of YOLOv3, we successfully uplift its performance to the same level as the state-of-the-art single-shot detectors (e.g., FCOS , CenterNet , and NAS-FPN ), as shown in Figure 1. Note that YOLOv3 can be evaluated at different input resolutions with the same weights, and when we lower the resolution of input images to pursue much faster detector, ASFF improves the performance more significantly.

Conclusion

This work identifies the inconsistency across different feature scales as a primary limitation for single-shot detectors with feature pyramids. To address this, we propose a novel ASFF strategy which learns the adaptive spatial fusion weight to filter out the inconsistency during training. It significantly improves strong baselines with tiny inference overhead and achieves a state-of-the-art speed and accuracy trade-off among all single-shot detectors.

References