UnitBox: An Advanced Object Detection Network

Jiahui Yu, Yuning Jiang, Zhangyang Wang, Zhimin Cao, Thomas Huang

Introduction

Visual object detection could be viewed as the combination of two tasks: object localization (where the object is) and visual recognition (what the object looks like). While the deep convolutional neural networks (CNNs) has witnessed major breakthroughs in visual object recognition , the CNN-based object detectors have also achieved the state-of-the-arts results on a wide range of applications, such as face detection , pedestrian detection and etc .

Currently, most of the CNN-based object detection methods could be summarized as a three-step pipeline: firstly, region proposals are extracted as object candidates from a given image. The popular region proposal methods include Selective Search , EdgeBoxes , or the early stages of cascade detectors ; secondly, the extracted proposals are fed into a deep CNN for recognition and categorization; finally, the bounding box regression technique is employed to refine the coarse proposals into more accurate object bounds. In this pipeline, the region proposal algorithm constitutes a major bottleneck in terms of localization effectiveness, as well as efficiency. On one hand, with only low-level features, the traditional region proposal algorithms are sensitive to the local appearance changes, e.g., partial occlusion, where those algorithms are very likely to fail. On the other hand, a majority of those methods are typically based on image over-segmentation or dense sliding windows , which are computationally expensive and have hamper their deployments in the real-time detection systems.

To overcome these disadvantages, more recently the deep CNNs are also applied to generate object proposals. In the well-known Faster R-CNN scheme , a region proposal network (RPN) is trained to predict the bounding boxes of object candidates from the anchor boxes. However, since the scales and aspect ratios of anchor boxes are pre-designed and fixed, the RPN shows difficult to handle the object candidates with large shape variations, especially for small objects.

Besides, to balance the bounding boxes with varied scales, DenseBox requires the training image patches to be resized to a fixed scale. As a consequence, DenseBox has to perform detection on image pyramids, which unavoidably affects the efficiency of the framework.

The paper proposes a highly effective and efficient CNN-based object detection network, called UnitBox. It adopts a fully convolutional network architecture, to predict the object bounds as well as the pixel-wise classification scores on the feature maps directly. Particularly, UnitBox takes advantage of a novel Intersection over Union (IoUIoU) loss function for bounding box prediction. The IoUIoU loss directly enforces the maximal overlap between the predicted bounding box and the ground truth, and jointly regress all the bound variables as a whole unit (see Figure 1). The UnitBox demonstrates not only more accurate box prediction, but also faster training convergence. It is also notable that thanks to the IoUIoU loss, UnitBox is enabled with variable-scale training. It implies the capability to localize objects in arbitrary shapes and scales, and to perform more efficient testing by just one pass on singe scale. We apply UnitBox on face detection task, and achieve the best performance on FDDB among all published methods.

IoU Loss Layer

where x~t\widetilde{x}_{t}, x~b\widetilde{x}_{b}, x~l\widetilde{x}_{l}, x~r\widetilde{x}_{r} represent the distances between current pixel location (i,j)(i,j) and the top, bottom, left and right bounds of ground truth, respectively. For simplicity, we omit footnote i,ji,j in the rest of this paper. Accordingly, a predicted bounding box is defined as x=(xt,xb,xl,xr)\boldsymbol{x}=(x_{t},x_{b},x_{l},x_{r}), as shown in Figure 1.

where L\mathcal{L} is the localization error.

2 IoU Loss Layer: Forward

In the following, we present a new loss function, named the IoUIoU loss, which perfectly addresses above drawbacks. Given a predicted bounding box x\boldsymbol{x} (after ReLU layer, we have xt,xb,xl,xr≥0x_{t},x_{b},x_{l},x_{r}\geq 0) and the corresponding ground truth x~\boldsymbol{\widetilde{x}}, we calculate the IoUIoU loss as follows:

In Algorithm 1, x~≠0\boldsymbol{\widetilde{x}}\neq\boldsymbol{0} represents that the pixel (i,j)(i,j) falls inside a valid object bounding box; XX is area of the predicted box; X~\widetilde{X} is area of the ground truth box; IhI_{h}, IwI_{w} are the height and width of the intersection area II, respectively, and UU is the union area.

3 IoU Loss Layer: Backward

To deduce the backward algorithm of IoUIoU loss, firstly we need to compute the partial derivative of XX w.r.t. xx, marked as ∇xX\nabla_{x}X (for simplicity, we notate xx for any of xtx_{t}, xbx_{b}, xlx_{l}, xrx_{r} if missing):

To compute the partial derivative of II w.r.t xx, marked as ∇xI\nabla_{x}I:

Finally we can compute the gradient of localization loss L\mathcal{L} w.r.t. xx:

From Eqn. 7, we can have a better understanding of the IoUIoU loss layer: the ∇xX\nabla_{x}X is the penalty for the predict bounding box, which is in a positive proportion to the gradient of loss; and the ∇xI\nabla_{x}I is the penalty for the intersection area, which is in a negative proportion to the gradient of loss. So overall to minimize the IoUIoU loss, the Eqn. 7 favors the intersection area as large as possible while the predicted box as small as possible. The limiting case is the intersection area equals to the predicted box, meaning a perfect match.

UnitBox Network

Based on the IoUIoU loss layer, we propose a pixel-wise object detection network, named UnitBox. As illustrated in Figure 2, the architecture of UnitBox is derived from VGG-16 model , in which we remove the fully connected layers and add two branches of fully convolutional layers to predict the pixel-wise bounding boxes and classification scores, respectively. In training, UnitBox is fed with three inputs in the same size: the original image, the confidence heatmap inferring a pixel falls in a target object (positive) or not (negative), and the bounding box heatmaps inferring the ground truth boxes at all positive pixels.

To predict the confidence, three layers are added layer-by-layer at the end of VGG stage-4: a convolutional layer with stride 11, kernel size 512×3×3×1512\times 3\times 3\times 1; an up-sample layer which directly performs linear interpolation to resize the feature map to original image size; a crop layer to align the feature map with the input image. After that, we obtain a 1-channel feature map with the same size of input image, on which we use the sigmoid cross-entropy loss to regress the generated confidence heatmap; in the other branch, to predict the bounding box heatmaps we use the similar three stacked layers at the end of VGG stage-5 with convolutional kernel size 512 x 3 x 3 x 4. Additionally, we insert a ReLU layer to make bounding box prediction non-negative. The predicted bounds are jointly optimized with IoUIoU loss proposed in Section 2. The final loss is calculated as the weighted average over the losses of the two branches.

Some explanations about the architecture design of UnitBox are listed as follows: 1) in UnitBox, we concatenate the confidence branch at the end of VGG stage-4 while the bounding box branch is inserted at the end of stage-5. The reason is that to regress the bounding box as a unit, the bounding box branch needs a larger receptive field than the confidence branch. And intuitively, the bounding boxes of objects could be predicted from the confidence heatmap. In this way, the bounding box branch could be regarded as a bottom-up strategy, abstracting the bounding boxes from the confidence heatmap; 2) to keep UnitBox efficient, we add as few extra layers as possible. Compared to DenseBox in which three convolutional layers are inserted for bounding box prediction, the UnitBox only uses one convolutional layer. As a result, the UnitBox could process more than 10 images per second, while DenseBox needs several seconds to process one image; 3) though in Figure 2 the bounding box branch and the confidence branch share some earlier layers, they could be trained separately with unshared weights to further improve the effectiveness.

With the heatmaps of confidence and bounding box, we can now accurately localize the objects. Taking the face detection for example, to generate bounding boxes of faces, firstly we fit the faces by ellipses on the thresholded confidence heatmaps. Since the face ellipses are too coarse to localize objects, we further select the center pixels of these coarse face ellipses and extract the corresponding bounding boxes from these selected pixels. Despite its simplicity, the localization strategy shows the ability to provide bounding boxes of faces with high accuracy, as shown in Figure 3.

Experiments

In this section, we apply the proposed IoUIoU loss as well as the UnitBox on face detection task, and report our experimental results on the FDDB benchmark . The weights of UnitBox are initialized from a VGG-16 model pre-trained on ImageNet, and then fine-tuned on the public face dataset WiderFace . We use mini-batch SGD in fine-tuning and set the batch size to 10. Following the settings in , the momentum and the weight decay factor are set to 0.9 and 0.0002, respectively. The learning rate is set to 10−810^{-8} which is the maximum trainable value. No data augmentation is used during fine-tuning.

2 Performance of UnitBox

To demonstrate the effectiveness of the proposed method, we compare the UnitBox with the state-of-the-arts methods on FDDB. As illustrated in Section 3, here we train an unshared UnitBox detector to further improve the detection performance. The ROC curves are shown in Figure 6. As a result, the proposed UnitBox has achieved the best detection result on FDDB among all published methods.

Except that, the efficiency of UnitBox is also remarkable. Compared to the DenseBox which needs seconds to process one image, the UnitBox could run at about 12 fps on images in VGA size. The advantage in efficiency makes UnitBox potential to be deployed in real-time detection systems.

Conclusions

References