Cascade RPN: Delving into High-Quality Region Proposal Network with Adaptive Convolution
Thang Vu, Hyunjun Jang, Trung X. Pham, Chang D. Yoo
Introduction
Object detection has received considerable attention in recent years for its applications in autonomous driving , robotics and surveillance . Given an image, object detectors aim to detect known object instances, each of which is assigned to a bounding box and a class label. Recent high-performing object detectors, such as Faster R-CNN , formulate the detection problem as a two-stage pipeline. At the first stage, a region proposal network (RPN) produces a sparse set of proposal boxes by refining and pruning a set of anchors, and at the second stage, a region-wise CNN detector (R-CNN) refines and classifies the proposals produced by RPN. Compared to R-CNN, RPN has received relatively less attention for improving its performance. This paper will focus on improving RPN by addressing its limitations that arise from heuristically defining the anchors and heuristically aligning the features to the anchors.
An anchor is defined by its scale and aspect ratio, and a set of anchors with different scales and aspect ratios are required to obtain a sufficient number of positive samples that have high overlap with the target objects. Setting appropriate scales and aspect ratios is important in achieving high detection performance, and it requires a fair amount of tuning .
An alignment rule is “implicitly” defined to set up a correspondence between the image features and the reference boxes. The input features of RPN and R-CNN should be well-aligned with the bounding boxes that are to be regressed. The alignment is guaranteed in R-CNN by the RoIPool or RoIAlign layer . The alignment in RPN is heuristically guaranteed: the anchor boxes are uniformly initialized, leveraging the observation that the convolutional kernel of the RPN uniformly strides over the feature maps. Such a heuristic introduces limitations for further improving detection performance as described below.
A number of studies have attempted to improve RPN by iterative refinement . Henceforth, this paper will refer to it as Iterative RPN. The motivation behind this idea is illustrated in Figure 2(a). Anchor boxes which are references for regression are uniformly initialized, and the target ground truth boxes are arbitrarily located. Thus, RPN needs to learn a regression distribution of high variance, as shown in Figure 2(a). If this regression distribution is perfectly learned, the regression distribution at stage 2 should be close to a Dirac Delta distribution. However, such a high-variance distribution at stage 1 is difficult to learn, requiring stage 2 regression. Stage 2 distribution has a lower variance compared to that of stage 1, and thus should be easier to learn but fails with Iterative RPN. The failure is implied by the observation in which the performance improvement of Iterative RPN is negligible compared to that of RPN, as shown in Figure 2(b). It is explained intuitively in Figure 2(c). Here, after stage 1, the anchor is regressed to be closer to the ground truth box; however, this breaks the alignment rule in detection.
This paper proposes an architecture referred to as Cascade RPN to systematically address the aforementioned problem arising from heuristically defining the anchors and aligning the features to the anchors. First, instead of using multiple anchors with different scales and aspect ratios, Cascade RPN relies on a single anchor and incorporates both anchor-based and anchor-free criteria in defining positive boxes to achieve high performance. Second, to benefit from multi-stage refinement while maintaining the alignment between anchor boxes and features, Cascade RPN relies on the proposed adaptive convolution that adapts to the refined anchors after each stage. Adaptive convolution serves as an extremely light-weight RoIAlign layer to learn the features sampled within the anchors.
Cascade RPN is conceptually simple and easy to implement. Without bells and whistles, a simple two-stage Cascade RPN achieves AR 13.4 points improvement compared to RPN baseline on the COCO dataset , surpassing any existing region proposal methods by a large margin. Cascade RPN can also be integrated into two-stage detectors to improve detection performance. In particular, integrating Cascade RPN into Fast R-CNN and Faster R-CNN achieves 3.1 and 3.5 points mAP improvement, respectively.
Related Work
Object detection can be roughly categorized into two main streams: one-stage and two-stage detection. Here, one-stage detectors are proposed to enhance computational efficiency. Examples falling in this stream are SSD , YOLO , RetinaNet , and CornerNet . Meanwhile, two-stage detectors aim to produce accurate bounding boxes, where the first stage generates region proposals followed by region-wise refinement and classification at the second stage, e.g., R-CNN , Fast R-CNN , Faster R-CNN , Cascade R-CNN , and HTC .
Region Proposals.
Region proposals have become the de-facto paradigm for high-quality object detectors . Region proposals serve as the attention mechanism that enables the detector to produce accurate bounding boxes while maintaining computation tractability. Early methods are based on grouping super-pixel (e.g., Selective Search , CPMC , MCG ) and window scoring (e.g., objectness in windows , EdgeBoxes ). Although these methods dominate the field of object detection in classical computer vision, they exhibit limitations as they are external modules independent of the detector and not computationally friendly. To overcome these limitations, Shaoqing et al. propose Region Proposal Network (RPN) that shares full-image convolutional features with the detection network, enabling nearly cost-free region proposals.
Multi-Stage RPN.
There have been a number of studies attempting to improve the performance of RPN . The general trend is to perform multi-stage refinement that takes the output of a stage as the input of the next stage and repeats until accurate localization is obtained, as presented in . However, this approach ignores the problem that the regressed boxes are misaligned to the image features, breaking the alignment rule required for object detection. To alleviate this problem, recent advanced methods rely on deformable convolution to perform feature spatial transformations and expect the learned transformations to align to the changes of anchor geometry. However, as there is no explicit supervision to learn the feature transformation, it is difficult to determine whether the improvement originates from conforming to the alignment rule or from the benefits of deformable convolution, thus making it less interpretable.
Anchor-based vs. Anchor-free Criterion for Sample Discrimination.
As a bounding box usually includes an object with some amount of background, it is difficult to determine if the box is a positive or a negative sample. This problem is usually addressed by comparing the Intersection over Union (IoU) between an anchor and a ground truth box to a predefined threshold; thus, it is referred to as the anchor-based criterion. However, as the anchor is uniformly initialized, multiple anchors with different scales and aspect ratios are required at each location to ensure that there are enough positive samples . The hyperparameters, such as scales and aspect ratios, are usually heuristically tuned and have a large impact on the final accuracy . Rather than relying on anchors, there have been studies that define positive samples by the distance between the prediction points and the center region of objects, referred to as anchor-free . This method is simple and requires fewer hyperparameters but usually exhibits limitations in dealing with complex scenes.
Region Proposal Network and Variants
Here, the regressor takes as input the image feature to output a prediction that minimizes the bounding box loss:
where is the robust loss defined in . The regressed anchor is simply inferred based on the inverse transformation of (1) as follows:
2 Iterative RPN and Variants
To alleviate this problem, recent advanced methods use deformable convolution to perform spatial transformations on the features as shown in Figure 3(c) and 3(d) and expect transformed features to align to the change in anchor geometry. However, this idea ignores the problem that there is no constraint to enforce the features to align with the changes in anchors: it is difficult to determine whether the deformable convolution produces feature transformation leading to alignment. Instead, the proposed Cascade RPN systematically ensures the alignment rule by using the proposed adaptive convolution.
Cascade RPN
Let denote the projection of anchor onto the feature map. The offset can be decoupled into center offset and shape offset (shown in Figure 3(e)):
where and is defined by the anchor shape and kernel size. For example, if kernel size is , then . As the offsets are typically fractional, sampling is performed with bilinear interpolation analogous to .
The illustrations of sampling locations in adaptive and other related convolutions are shown in Figure 4. Conventional convolution samples the features at contiguous locations with a dilation factor of 1. The dilated convolution increases the dilation factor, aiming to enhance the semantic scope with unchanged computational cost. The deformable convolution augments the spatial sampling locations by learning the offsets. Meanwhile, the proposed adaptive convolution performs sampling within the anchors to ensure alignment between the anchors and features. Adaptive convolution is closely related to the others. Adaptive convolution becomes dilated convolution if the center offsets are zeros. Deformable convolution becomes adaptive convolution if the offsets are deterministically derived from the anchors.
2 Sample Discrimination Metrics
Instead of using multiple anchors with predefined scales and aspect ratios, Cascade RPN relies on a single anchor per location and performs multi-stage refinement. However, this reliance creates a new challenge in determining whether a training sample is positive or negative as the use of anchor-free or anchor-based metric is highly adversarial. The anchor-free metric establishes a loose requirement for positive samples in the second stage and the anchor-based metric results in an insufficient number of positive training examples at the first stage. To overcome this challenge, Cascade RPN progressively strengthens the requirements through the stages by starting out with an anchor-free metric followed by anchor-based metrics in the ensuing stages. In particular, at the first stage, an anchor is a positive sample if its center is inside the center region of an object. In the following stages, an anchor is a positive sample if its IoU with an object is greater than the IoU threshold.
3 Cascade RPN
4 Learning
Cascade RPN can be trained in an end-to-end manner using multi-task loss as follows:
Here, is the regression loss at stage with the weight of , and is the classification loss. The two loss terms are balanced by . In the implementation, binary cross entropy loss and IoU loss are used as the classification loss and regression loss, respectively.
Experiments
The experiments are performed on the COCO 2017 detection dataset . All the models are trained on the train split (115k images). The region proposal performance and ablation analysis are reported on val split (5k images), and the benchmarking detection performance is reported on test-dev split (20k images).
Unless otherwise specified, the default model of the experiment is as follows. The model consists of two stages, with ResNet50-FPN being its backbone. The use of two stages is to balance accuracy and computational efficiency. A single anchor per location is used with size of , , , , and corresponding to the feature levels , , , , and , respectively . The first stage uses the anchor-free metric for sample discrimination with the thresholds of the center-region and ignore-region , which are adopted from , being 0.2 and 0.5. The second stage uses the anchor-based metric with the IoU threshold of 0.7. The multi-task loss is set with the stage-wise weight and the balance term . The NMS threshold is set to 0.8. In all experiments, the long edge and the short edge of the images are resized to 1333 and 800 respectively without changing the aspect ratio. No data augmentation is used except for standard horizontal image flipping. The models are implemented with PyTorch and mmdetection . The models are trained with 8 GPUs with a batch size of 16 (two images per GPU) for 12 epochs using SGD optimizer. The learning rate is initialized to 0.02 and divided by 10 after 8 and 11 epochs. It takes about 12 hours for the models to converge on 8 Tesla V100 GPUs.
The quality of region proposals is measured with Average Recall (AR), which is the average of recalls across IoU thresholds from 0.5 to 0.95 with a step of 0.05. The AR for 100, 300, and 1000 proposals per image are denoted as AR100, AR300, and AR1000. The AR for small, medium, and large objects computed at 100 proposals are denoted as ARS, ARM, and ARL, respectively. Detection results are evaluated with the standard COCO-style Average Precision (AP) measured at IoUs from 0.5 to 0.95. The runtime is measured on a single Tesla V100 GPU.
2 Benchmarking Results
The performance of Cascade RPN is compared to those of recent state-of-the-art region proposal methods, including RPN , SharpMask , GCN-NS , AttractioNet , ZIP , and GA-RPN . In addition, Iterative RPN and Iterative RPN+, which are referred to in Figure 3, are also benchmarked. The results of Sharp Mask, GCN-NS, AttractioNet, ZIP are cited from the papers. The results of the remaining methods are reproduced using mmdetection . Table 1 summarizes the benchmarking results. In particular, Cascade RPN achieves AR 13.4 points higher than that of the conventional RPN. Cascade RPN consistently outperforms the other methods in terms of AR under different settings of proposal thresholds and object scales. The alignment rule is typically missing or loosely conformed to in the other methods; thus, their performance improvements are limited. The alignment rule in Cascade RPN is systematically ensured such that the performance gain is greater and more reliable.
Detection Performance.
To investigate the benefit of high-quality proposals, Cascade RPN and the baselines are integrated into common two-stage object detectors, including Fast R-CNN and Faster R-CNN. Here, Fast R-CNN is trained on precomputed region proposals while Faster R-CNN is trained in an end-to-end manner. As studied in , despite high-quality region proposals, training a good detector is still a non-trivial problem, and simply replacing RPN by Cascade RPN without changes in the settings only brings limited gain. Following , the IoU threshold in R-CNN is increased and the number of proposals is decreased. In particular, the IoU threshold and the number of proposals are set to 0.65 and 300, respectively. The experimental results are reported in Table 2. Here, integrating RPN into Fast R-CNN and Faster R-CNN yields 37.0 and 37.1 mAP, respectively. From the results, the recall improvement is correlated with improvements in detection performance. As it has the highest recall, Cascade RPN boosts the performance for Fast R-CNN and Faster R-CNN to 40.1 and 40.6 mAP, respectively.
3 Ablation Study
To demonstrate the effectiveness of Cascade RPN, a comprehensive component-wise analysis is performed in which different components are omitted. The results are reported in Table 3. Here, the baseline is RPN with 3 anchors per location yielding AR1000 of 58.3. When the number of anchors per location is reduced to 1, the AR1000 drops to 55.8, implying that the number of positive samples dramatically decreases. Even when the multi-stage cascade is added, the performance is 58.0, which is still lower than that of the baseline. However, when adaptive convolution is applied to ensure alignment, the performance surges to 67.8, showing the importance of alignment in multi-stage refinement. The incorporation of anchor-free and anchor-based metrics for sample discrimination incrementally improves AR1000 to 68.6. The use of regression statistics (shown in Figure 2(a)) increases the performance to 71.5. Finally, applying IoU loss yields a slight improvement of 0.2 points. Overall, Cascade RPN achieves 16.5, 14.7, and 13.4 points improvement in terms of AR100, AR300, and AR1000 respectively, compared to the conventional RPN.
Acquisition of Alignment.
To demonstrate the effectiveness of the proposed adaptive convolution, the center and shape alignments, represented by the offsets in Eq. (7), are progressively applied. Here, the center and shape offsets maintain the alignments in position and semantic scope, respectively. Table 5 shows that the AR1000 improves from 58.0 to 64.1 using only the center alignment. When both the center and shape alignments are ensured, the performance increases to 67.8.
Sample Discrimination Metrics.
The experimental results with different combinations of sample discrimination metrics are shown in Table 5. Here, AF and AB denote that the anchor-free and anchor-based metrics are applied for all stages, respectively. Meanwhile, AFAB indicates that the anchor-free metric is applied at stage 1 followed by anchor-based metric at stage 2. Here, AF and AB yield the AR1000 of 66.4 and 67.8 respectively, both of which are significantly less than that of AFAB. It is noted that the thresholds for each metric are already adapted through stages. The results imply that applying only one of either anchor-free or anchor-based metric is highly adversarial. The both metrics should be incorporated to achieve the best results.
Qualitative Evaluation.
The examples of region proposal results at the first and second stages are illustrated in the first and second row of Figure 5, respectively. The results show that the output proposals at the second stage are more accurate and cover a larger number of objects.
Number of Stages.
Table 7 shows the proposal performance on different number of stages. In the 3-stage Cascade RPN, an IoU threshold of 0.75 is used for the third stage. The 2-stage Cascade RPN achieves the best trade-off between AR1000 and inference time.
Extension with Cascade R-CNN.
Table 7 reports the detection results of the Cascade R-CNN with different proposal methods. The Cascade RPN improves AP by 0.8 points compared to RPN. The improvement is mainly from AP75, where the objects have high IoU with the ground truth.
Conclusion
This paper introduces Cascade RPN, a simple yet effective network architecture for improving region proposal quality and object detection performance. Cascade RPN systematically addresses the limitations that conventional RPN heuristically defines the anchors and aligns the features to the anchors. A simple implementation of a two-stage Cascade RPN achieves AR 13.4 points higher than the baseline, surpassing any existing region proposal methods. When adopting to Fast R-CNN and Faster R-CNN, Cascade RPN can improve the detection mAP by 3.1 and 3.5 points, respectively.
This work was supported by Institute for Information & communications Technology Planning & Evaluation(IITP) grant funded by the Korea government (MSIT) (2017-0-01780, The technology development for event recognition/relational reasoning and learning knowledge-based system for video understanding) and (No. 2019-0-01396, Development of framework for analyzing, detecting, mitigating of bias in AI model and training data)