Attend Refine Repeat: Active Box Proposal Generation via In-Out Localization

Spyros Gidaris, Nikos Komodakis

Introduction

Category agnostic object proposal generation is a computer vision task that has received an immense amount of attention over the last years. Its definition is that for a given image a small set of instance segmentations or bounding boxes must be generated that will cover with high recall all the objects that appear in the image regardless of their category. In object detection, applying the recognition models to such a reduced set of category independent location hypothesis [Girshick et al.(2014)Girshick, Donahue, Darrell, and Malik] instead of an exhaustive scan of the entire image [Felzenszwalb et al.(2010)Felzenszwalb, Girshick, McAllester, and Ramanan, Sermanet et al.(2013)Sermanet, Eigen, Zhang, Mathieu, Fergus, and LeCun], has the advantages of drastically reducing the amount of recognition model evaluations and thus allowing the use of more sophisticated machinery for that purpose. As a result, proposal based detection systems manage to achieve state-of-the-art results and have become the dominant paradigm in the object detection literature [Girshick et al.(2014)Girshick, Donahue, Darrell, and Malik, He et al.(2015a)He, Zhang, Ren, and Sun, Girshick(2015), Gidaris and Komodakis(2015), Gidaris and Komodakis(2016), Ren et al.(2015)Ren, He, Girshick, and Sun, Zagoruyko et al.(2016)Zagoruyko, Lerer, Lin, Pinheiro, Gross, Chintala, and Dollár, Bell et al.(2015)Bell, Zitnick, Bala, and Girshick, Shrivastava et al.(2016)Shrivastava, Gupta, and Girshick]. Object proposals have also been used in various other tasks, such as weakly-supervised object detection [Cinbis et al.(2015)Cinbis, Verbeek, and Schmid], exemplar 2D-3D detection [Massa et al.(2016)Massa, Russell, and Aubry], visual semantic role labelling [Gupta and Malik(2015)], caption generation [Karpathy and Fei-Fei(2015)] or visual question answering [Shih et al.(2015)Shih, Singh, and Hoiem].

In this work we focus on the problem of generating bounding box object proposals rather than instance segmentations. Several approaches have been proposed in the literature for this task [Van de Sande et al.(2011)Van de Sande, Uijlings, Gevers, and Smeulders, Arbeláez et al.(2014)Arbeláez, Pont-Tuset, Barron, Marques, and Malik, Zitnick and Dollár(2014), Krähenbühl and Koltun(2015), Krähenbühl and Koltun(2014), Alexe et al.(2012)Alexe, Deselaers, and Ferrari, Lu et al.()Lu, Javidi, and Lazebnik, Cheng et al.(2014)Cheng, Zhang, Lin, and Torr, Hayder et al.(2016)Hayder, He, and Salzmann, Chen et al.(2015)Chen, Ma, Wang, and Zhao, Dai et al.()Dai, He, Li, Ren, and Sun]. Among them our work is most related to the CNN-based objectness scoring approaches [Kuo et al.(2015)Kuo, Hariharan, and Malik, Ghodrati et al.(2015)Ghodrati, Diba, Pedersoli, Tuytelaars, and Van Gool, Pinheiro et al.(2015)Pinheiro, Collobert, and Dollar] that recently have demonstrated state-of-the-art results [Pinheiro et al.(2015)Pinheiro, Collobert, and Dollar, Pinheiro et al.(2016)Pinheiro, Lin, Collobert, and Dollár].

In the objectness scoring paradigm, a large set of image boxes is ranked according to how likely it is for each image box to tightly enclose an object — regardless of its category — and then this set is post-processed with a non-maximum-suppression step and truncated to yield the final set of object proposals. In this context, Kuo et al\bmvaOneDot [Kuo et al.(2015)Kuo, Hariharan, and Malik] with their DeepBox system demonstrated that training a convolutional neural network to perform the task of objectness scoring can yield superior performance over previous methods that were based on low level cues and they provided empirical evidence that it can generalize to unseen categories. In order to avoid evaluating the computationally expensive CNN-based objectness scoring model on hundreds of thousands image boxes, which is necessary for achieving good localization of all the objects in the image, they use it only to re-rank the proposals generated from a faster but less accurate proposal generator thus being limited by its localization performance. Instead, more recent CNN-based approaches apply their models only to ten of thousands image boxes, uniformly distributed in the image, and jointly with objectness prediction they also infer the bounding box of the closest object to each input image box. Specifically, the Region Proposal Network in Faster-RCNN [Ren et al.(2015)Ren, He, Girshick, and Sun] performs bounding box regression for that purpose while the DeepMask method predicts the foreground mask of the object centred in the image box and then it infers the location of the object’s bounding box by extracting the box that tightly encloses the foreground pixels. The latter has demonstrated state-of-the-art results and was recently extended with a top-down foreground mask refinement mechanism that exploits the convolutional feature maps at multiple depths of a neural network [Pinheiro et al.(2016)Pinheiro, Lin, Collobert, and Dollár].

Our work is also based on the paradigm of having a CNN model that given an image box it jointly predicts its objectness and a new bounding box that is better aligned on the object that it contains. However, we opt to advance the previous state-of-the-art in box proposal generation in two ways: (1) improving the object’s bounding box prediction step (2) actively generating the set of image boxes that will be processed by the CNN model.

Regarding the bounding box inference step we exploit the recent advances in object detection where Gidaris and Komodakis [Gidaris and Komodakis(2016)] showed how to improve the object-specific localization accuracy. Specifically, they replaced the bounding box regression step with a localization module, called LocNet, that given a search region it infers the bounding box of the object inside the search region by assigning membership probabilities to each row and each column of that region and they empirically proved that this localization task is easier to be learned from a convolutional neural network thus yielding more accurate box predictions during test time. Given the importance of having accurate bounding box locations in the proposal generation task, we believe that it would be of great interest to develop and study a category agnostic version of LocNet for this task.

Our second idea for improving the box proposal generation task stems from the following observation. Recent state-of-the-art box proposal methods evaluate only a relatively small set of image boxes (in the order of 10k10k) uniformly distributed in the image and rely on the bounding box prediction step to fix the localization errors. However, depending on how far an object is from the closest evaluated image box, both the objectness scoring and the bounding box prediction for that object could be imperfect. For instance, Hosang et al\bmvaOneDot [Hosang et al.(2015)Hosang, Benenson, Dollár, and Schiele] showed that in the case of the detection task the correct recognition of an object from an image box is correlated with how well the box encloses the object. Given how similar are the tasks of category-specific object detection and category-agnostic proposal generation, it is safe to assume that a similar behaviour will probably hold for the latter one as well. Hence, in our work we opt for an active object localization scheme, which we call Attend Refine Repeat algorithm, that starting from a set of seed boxes it progressively generates newer boxes that are expected with higher probability to be on the neighbourhood or to tightly enclose the objects of the image. Thanks to this localization scheme, our box proposal system is capable to both correct initially imperfect bounding box predictions and to give higher objectness score to candidate boxes that are more well localized on the objects of the image.

To summarize, our contributions with respect to the box proposal generation task are:

We developed a box proposal system that is based on an improved category-agnostic object location refinement module and on an active box proposal generation strategy that behaves as an attention mechanism that focus on the promising image areas in order to propose objects. We call the developed box proposal system AttractioNet: (Att)end (R)efine Repeat: (Act)ive Box Proposal Generation via (I)n-(O)ut Localization (Net)work.

We exhaustively evaluate our system both on PASCAL and on the more challenging COCO datasets and we demonstrate significant improvement with respect to the state-of-the-art on box proposal generation. Furthermore, we provide strong evidence that our object location refinement module is capable of generalizing to unseen categories by reporting results for the unseen categories of ImageNet detection task and NYU-Depth dataset.

Finally, we evaluate our box proposal generation approach in the context of the object detection task using a VGG16-Net based detection system and the achieved average precision performance on the COCO test-dev set manages to significantly surpass all other VGG16-Net based detection systems while even being on par with the ResNet-101 based detection system of He et al\bmvaOneDot [He et al.(2015b)He, Zhang, Ren, and Sun].

The remainder of the paper is structured as follows: We describe our box proposal methodology in section §2, we show experimental results in section §3 and we present our conclusions in section §4.

Our approach

The active box proposal generation strategy that we employ in our work, which we call Attend Refine Repeat algorithm, starts from a set of seed boxes, which only depend on the image size, and it then sequentially produces newer boxes that will better cover the objects of the image while avoiding the "objectless" image areas (see Figure 1). At the core of this algorithm lies a CNN-based box proposal model that, given an image II and the coordinates of a box BB, executes the following operations:

this operation scores the box BB based on how likely it is to tightly enclose an object, regardless of its category.

The pseudo-code of the Attend Refine Repeat algorithm is provided in Algorithm 1. Specifically, it starts by initializing the set of candidate boxes C to the empty set and then creates a set of seed boxes B0\textbf{B}^{0} by uniformly distributing boxes of various fixed sizes in the image (similar to Cracking Bing [Zhao et al.(2014)Zhao, Liu, and Yin]). Then on each iteration tt it estimates the objectness Ot\textbf{O}^{t} of the boxes generated in the previous iteration, Bt−1\textbf{B}^{t-1}, and it refines their location (resulting in boxes Bt\textbf{B}^{t}) by attempting to predict the bounding boxes of the objects that are closest to them. The results {Bt,Ot}\{\textbf{B}^{t},\textbf{O}^{t}\} of those operations are added to the candidates set C and the algorithm continues. In the end, non-maximum-suppression [Felzenszwalb et al.(2010)Felzenszwalb, Girshick, McAllester, and Ramanan] is applied to the candidate box proposals C and the top KK box proposals, set P, are returned.

The advantages of having an algorithm that sequentially generates new box locations given the predictions of the previous stage are two-fold:

Attention mechanism: First, it behaves as an attention mechanism that, on each iteration, focuses more and more on the promising locations (in terms of box coordinates) of the image (see Figure 1). As a result of this, boxes that tightly enclose the image objects are more likely to be generated and to be scored with high objectness confidence.

Robustness to initial boxes: Furthermore, it allows to refine some initially imperfect box predictions or to localize objects that might be far (in terms of center location, scale and/or aspect ratio) from any seed box in the image. This is illustrated via a few characteristic examples in Figure 2. As shown in each of these examples, starting from a seed box, the iterative bounding box predictions gradually converge to the closest (in terms of center location, scale and/or aspect ratio) object without actually being affected from any nearby instances.

2 CNN-based box proposal model

In this section we describe in more detail the object localization and objectness scoring modules of our box proposal model as well as the CNN architecture that implements the entire Attend Refine Repeat algorithm that was presented above.

In order for our active box proposals generation strategy to be effective, it is very important to have an accurate and robust category agnostic object location refinement module. Hence we follow the paradigm of the recently introduced LocNet model [Gidaris and Komodakis(2016)] that has demonstrated superior performance in the category specific object detection task over the typical bounding box regression paradigm [Gidaris and Komodakis(2015), Girshick(2015), Sermanet et al.(2013)Sermanet, Eigen, Zhang, Mathieu, Fergus, and LeCun, Ren et al.(2015)Ren, He, Girshick, and Sun] by formulating the problem of bounding box prediction as a dense classification task. Here we use a properly adapted version of that model for the task at hand.

We note that in contrast to the original LocNet model that is optimized to yield a different set of probability vectors of each category in the training set, here our category-agnostic version is designed to yield a single set of probability vectors that should accurately localize any object regardless of its category (see also section §\S2.2.3 that describes in detail the overall architecture of our proposed model). It should be also mentioned that this is a more challenging task to learn since, in this case, the model should be able to localize the target objects even if they are in crowded scenes with other objects of the same appearance and/or texture (see the two left-most examples of Figure 3) without exploiting any category supervision during training that would help it to better capture the appearance characteristics of each object category. On top of that, our model should be able to localize objects of unseen categories. In the right-most example of Figure 3, we provide an indicative result produced by our model that verifies this test case. In this particular example, we apply a category-agnostic refinement module trained on PASCAL to an object whose category ("clock") was not present in the training set and yet our trained model had no problem of confidently predicting the correct location of the object. In section 3.2 of the paper we also provide quantitative results about the generalization capabilities of the location refinement module.

2.2 Objectness scoring module

The functionality of the objectness scoring module is that it gets as input a box BB and yields a single probability pobjp_{obj} of whether or not this box tightly encloses an object, regardless of what the category of that object might be.

2.3 AttractioNet architecture

We call the overall network architecture that implements the Attend Refine Repeat algorithm with its In-Out object location refinement module and its objectness scoring module, AttractioNetAttractioNet : (Att)end (R)efine Repeat: (Act)ive Box Proposal Generation via (I)n-(O)ut Localization (Net)work. Given an image II, our AttractioNet model will be required to process multiple image boxes of various sizes, by two different modules and repeat those processing steps for several iterations of the Attend Refine Repeat algorithm. So, in order to have an efficient implementation we follow the SPP-Net [He et al.(2015a)He, Zhang, Ren, and Sun] and Fast-RCNN [Girshick(2015)] paradigm and share the operations of the first convolutional layers between all the boxes, as well as across the two modules and all the Attend Refine Repeat algorithm repetitions (see Figure 4). Specifically, our AttractioNet model first forwards the image II through a first sequence of convolutional layers (conv. layers of VGG16-Net [Simonyan and Zisserman(2014)]) in order to extract convolutional feature maps FIF_{I} from the entire image. Then, on each iteration tt the box-wise part of the architecture, which we call Attend & Refine Network, gets as input the image convolutional feature maps FIF_{I} and a set of box locations Bt−1\textbf{B}^{t-1} and yields the refined bounding box locations Bt\textbf{B}^{t} and their objectness scores Ot\textbf{O}^{t} using its object location refinement module sub-network and its objectness scoring module sub-network respectively. In Figure 5 we provide the work-flow of the Attend & Refine Network when processing a single input box BB. The architecture of its two sub-networks is described in more detail in the rest of this section:

Object location refinement module sub-network. This module gets as input the feature map FIF_{I} and the search region RR and yields the probability vectors pxp_{x} and pyp_{y} of that search region with a network architecture similar to that of LocNet. Key elements of this architecture is that it branches into two heads, the X and Y, each responsible for yielding the pxp_{x} or the pyp_{y} outputs. Differently from the original LocNet architecture, the convolutional layers of this sub-network output 128 feature channels instead of 512, which speeds up the processing by a factor of 4 without affecting the category-agnostic localization accuracy. Also, in order to yield a fixed size feature for the RR region, instead of region adaptive max-pooling this sub-network uses region bilinear pooling [Dai et al.(2015)Dai, He, and Sun, Johnson et al.(2016)Johnson, Karpathy, and Fei-Fei] that in our initial experiments gave slightly better results. Finally, our version is designed to yield two probability vectors of size MMHere we use M=56M=56., instead of C×2C\times 2 vectors of size MM (where CC is the number of categories), given that in our case we aim for category-agnostic object location refinement.

Objectness scoring module sub-network. Given the image feature maps FIF_{I} and the window BB it first performs region adaptive max pooling of the features inside BB that yields a fixed size feature (7×7×5127\times 7\times 512). Then it forwards this feature through two linear+ReLU hidden layers of 40964096 channels each (fc_6 and fc_7 layers of VGG16) and a final linear+sigmoid layer with one output that corresponds to the probability pobjp_{obj} of the box BB tightly enclosing an object. During training the hidden layers are followed by Dropout units with dropout probability p=0.5p=0.5.

3 Training procedure

Training loss: During training the following multi-task loss is optimized:

where θ\theta are the learnable network parameters, {Bk,Tk,Ik}k=1NL\{B_{k},T_{k},I_{k}\}_{k=1}^{N^{L}} are NLN^{L} training triplets for learning the localization task and {Bk,yk,Ik}k=1NO\{B_{k},y_{k},I_{k}\}_{k=1}^{N^{O}} are NON^{O} training triplets for learning the objectness scoring task. Each training triple {B,T,I}\{B,T,I\} of the localization task includes the image II, the box BB and the target localization probability vectors T={Tx,Ty}T=\{T_{x},T_{y}\}. If (Bl∗,Bt∗)(B^{*}_{l},B^{*}_{t}) and (Br∗,Bb∗)(B^{*}_{r},B^{*}_{b}) are the top-left and bottom-right coordinates of the target box B∗B^{*} then the target probability vectors Tx ⁣= ⁣{Tx,i}i=1M{T}_{x}\!=\!\{T_{x,i}\}_{i=1}^{M} and Ty ⁣= ⁣{Ty,i}i=1M{T}_{y}\!=\!\{T_{y,i}\}_{i=1}^{M} are defined as:

The loss Lloc(θ∣B,T,I)L_{loc}(\theta|B,T,I) of this triplet is the sum of binary logistic regression losses:

where pap_{a} are the output probability vectors of the localization module for the image II, the box BB and the network parameters θ\theta. The training triplet {B,y,I}\{B,y,I\} for the objectness scoring task includes the image II, the box BB and the target value y∈{0,1}y\in\{0,1\} of whether the box BB contains an object (positive triplet with y=1y=1) or not (negative triplet with y=0y=0). The loss Lobj(θ∣B,y,I)L_{obj}(\theta|B,y,I) of this triplet is the binary logistic regression loss ylog⁡(pobj)+(1−y)log⁡(1−pobj)y\log(p_{obj})+(1-y)\log(1-p_{obj}), where pobjp_{obj} is the objectness probability for the image II, the box BB and the network parameters θ\theta.

Creating training triplets: In order to create the localization and objectness training triplets of one image we first artificially create a pool of boxes that our algorithm is likely to see during test time. Hence we start by generating seed boxes (as the test time algorithm) and for each of them we predict the bounding boxes of the ground truth objects that are closest to them using an ideal object location refinement module. This step is repeated one more time using the previous ideal predictions as input. Because of the finite search area of the search region RR the predicted boxes will not necessarily coincide with the ground truth bounding boxes. Furthermore, to account for prediction errors during test time, we repeat the above process by jittering this time the output probability vectors of the ideal location refinement module with 20%20\% noise. Finally, we merge all the generated boxes (starting from the seed ones) to a single pool. Given this pool, the positive training boxes in the objectness localization task are those that their IoUIoU with any ground truth object is at least 0.50.5 and the negative training boxes are those that their maximum IoUIoU with any ground truth object is less than 0.40.4. For the localization task we use as training boxes those that their IoUIoU with any ground truth object is at least 0.50.5.

Optimization: To minimize the objective we use stochastic gradient descent (SGD) optimization with an image-centric strategy for sampling training triplets. Specifically, in each mini-batch we first sample 4 images and then for each image we sample 64 training triplets for the objectness scoring task (50%50\% are positive and 50%50\% are negative) and 32 training triplets for the localization task. The momentum is set to 0.90.9 and the learning schedule includes training for 320k320k iterations with a learning rate of lr=0.001l_{r}=0.001 and then for another 260k260k iterations with lr=0.0001l_{r}=0.0001. The training time is around 7 days (although we observed that we could have stopped training on the 5th day with insignificant loss in performance).

Scale and aspect ratio jittering: During test time our model is fed with a single image scaled such that its shortest dimension to be 10001000 pixels or its longest dimension to not exceed the 14001400 pixels. However, during training each image is randomly resized such that its shortest dimension to be one of the following number of pixels {300 ⁣: ⁣50 ⁣: ⁣1000}\{300\!:\!50\!:\!1000\} (using Matlab notation) taking care, however, the longest dimension to not exceed 10001000 pixels. Also, with probability 0.50.5 we jitter the aspect ratio of the image by altering the image dimensions from W×HW\times H to (αW)×H(\alpha W)\times H or W×(αH)W\times(\alpha H) where the value of α\alpha is uniformly sampled from 2.2:.2:1.02^{.2:.2:1.0} (Matlab notation). We observed that this type of data augmentation gives a slight improvement on the results.

Experimental results

In this section we perform an exhaustive evaluation of our box proposal generation approach, which we call AttractioNet, under various test scenarios. Specifically, we first evaluate our approach with respect to its object localization performance by comparing it with other competing methods and we also provide an ablation study of its main novel components in §3.1. Then, we study its ability to generalize to unseen categories in §3.2, we evaluate it in the context of the object detection task in §3.3 and finally, we provide qualitative results in §3.4.

Training set: In order to train our AttractioNet model we use the training set of MS COCO [Lin et al.(2014)Lin, Maire, Belongie, Hays, Perona, Ramanan, Dollár, and Zitnick] detection benchmark dataset that includes 80k80k images and it is labelled with 8080 different object categories. Note that the MSCOCO dataset is an ideal candidate for training our box proposal model since: (1) it is labelled with a descent number of different object categories and (2) it includes images captured from complex real-life scenes containing common objects in their natural context. The aforementioned training set properties are desirable for achieving good performance on "difficult" test images (a.k.a. images in the wild) and generalizing to "unseen" during training object categories.

Implementation details: In the active box proposal algorithm we use 10k10k seed boxes generated with a similar to Cracking Bing [Zhao et al.(2014)Zhao, Liu, and Yin] techniqueWe use seed boxes of 3 aspect ratios, 1:21:2, 2:12:1 and 1:11:1, and 9 different sizes of the smallest seed box dimension {16,32,50,72,96,128,192,256,384}\{16,32,50,72,96,128,192,256,384\}.. To reduce the computational cost of our algorithm, after the first repetition we only keep the top 2k2k scored boxes and we continue with this number of candidate box proposals for four more extra iterations. In the non-maximum-suppression [Felzenszwalb et al.(2010)Felzenszwalb, Girshick, McAllester, and Ramanan] (NMS) step the optimal IoU threshold (in terms of the achieved AR) depends on the desired number of box-proposals. For example, for 10, 100, 1000 and 2000 proposals the optimal IoU thresholds are 0.550.55, 0.750.75, 0.900.90 and 0.950.95 respectively (note that the aforementioned IoU thresholds were cross validated on a set different from the one used for evaluation). For practical purposes and in order to have a unified NMS process, we first apply NMS with the IoU threshold equal to 0.950.95 and get the top 2000 box proposals, and then follow a multi-threshold NMS strategy that re-orders this set of 2000 boxes such that for any given number KK, the top KK box proposals in the set better cover (in terms of achieved AR) the objects in the image (see appendix A).

Here we evaluate our AtractioNet method in the end task of box proposal generation. For that purpose, we test it on the first 5k5k images of the COCO validation set and the PASCAL [Everingham et al.(2010)Everingham, Van Gool, Williams, Winn, and Zisserman] VOC2007 test set (that also includes around 5k5k images).

Evaluation Metrics: As evaluation metric we use the average recall (AR) which, for a fixed number of box proposals, averages the recall of the localized ground truth objects for several Intersection over Union (IoU) thresholds in the range .5:.05:.95 (Matlab notation). The average recall metric has been proposed from Hosang et al\bmvaOneDot [Hosang et al.(2015)Hosang, Benenson, Dollár, and Schiele, Hosang et al.(2014)Hosang, Benenson, and Schiele] where in their work they demonstrated that it correlates well with the average precision performance of box proposal based object detection systems. In our case, in order to evaluate our method we report the AR results for 10, 100 and 1000 box proposals using the notation AR@10, AR@100 and AR@1000 respectively. Also, in the case of 100 box proposals we also report the AR of the small (α<322\alpha<32^{2}), medium (322≤α≤96232^{2}\leq\alpha\leq 96^{2}) and large (α>962\alpha>96^{2}) sized objects using the notation AR@100-Small, AR@100-Medium and AR@100-Large respectively, where α\alpha is the area of the object. For extracting those measurements we use the COCO API (https://github.com/pdollar/coco).

In Table 1 we report the average recall (AR) metrics of our method as well as of other competing methods in the COCO validation set. We observe that the average recall performance achieved by our method exceeds all the previous work in all the AR metrics by a significant margin (around 10 absolute points in the percentage scale). Similar gains are also observed in Table 2 where we report the average recall results of our methods in the PASCAL VOC2007 test set. Furthermore, in Figure 6 we provide for our method the recall as a function of the IoU overlap of the localized ground truth objects. We see that the recall decreases relatively slowly as we increase the IoU from 0.5 to 0.75 while for IoU above 0.85 the decrease is faster.

Comparison with previous state-of-the-art. In Figure 7 we compare the box proposals generated from our AttractioNet model (Ours entry) against those generated from the previous state-of-the-art [Pinheiro et al.(2016)Pinheiro, Lin, Collobert, and Dollár] (entries SharpMask, SharpMaskZoom and SharpMaskZoom2) w.r.t. the recall versus IoU trade-off and average recall versus proposals number trade-off that they achieve. Also, in Table 1 we report the AR results both for our method and for the SharpMask entries. We observe that the model proposed in our work has clearly superior performance over the SharpMask entries under all test cases.

1.2 Ablation study

Here we perform an ablation study of our two key ideas for improving the state-of-the-art on the bounding box proposal generation task:

Object location refinement module. In order to assess the importance of our object location refinement module we evaluated two test cases for generating box proposals: (1) simply applying the objectness scoring module on a set of 18k18k seed boxes (first row of Table 3) and (2) applying both the objectness scoring module and the object location refinement module on the same set of 18k18k seed boxes (second row of Table 3). Note that in none of them is the active box generation strategy being used. The average recall results of those two test cases are reported in the first two rows of Table 3. We observe that without the object location refinement module the average recall performance of the box proposal system is very poor. In contrast, the average recall performance of the test case that involves the object location refinement module but not the active box generation strategy is already better than the previous state-of-the-art as reported in Table 1, which demonstrates the very good localization accuracy of our category agnostic location refinement module.

Active box generation strategy. Our active box generation strategy, which we call Attend Refine Repeat algorithm, attends in total 18k18k boxes before it outputs the final list of box proposals. Specifically, it attends 10k10k seed boxes in the first repetition of the algorithm and 2k2k actively generated boxes in each of the following four repetitions. A crucial question is whether actively generating those extra 8k8k boxes is really essential in the task or we could achieve the same average recall performance by directly attending 18k18k seed boxes and without continuing on the active box generation stage. We evaluated such test case and we report the average recall results in Table 3 (see rows 2 and 3). We observe that employing the active box generation strategy (3rd row in Table 3) offers a significant boost in the average recall performance (between 3 and 6 absolute points in the percentage scale) thus proving its importance on yielding well localized bounding box proposals. Also, in the right side of Figure 8 we plot the average recall metrics as a function of the repetitions number of our active box generation strategy. We observe that the average recall measurements are increased as we increase the repetitions number and that the increase is more steep on the first repetitions of the algorithm while it starts to converge after the 4th repetition.

1.3 Run time

In the current work we did not focus on providing an optimized implementation of our approach. There is room for significantly improving computational efficiency. For instance, just by using SVD decomposition on the fully connected layers of the objectness module at post-training time (similar to Fast-RCNN [Girshick(2015)]) and early stopping a sequence of bounding box location refinements in the case it has already converged A sequence of bounding box refinements is considered that it has converged when the IoU between the two lastly predicted boxes in the sequence is greater than 0.9., the runtime drops from 4.0 seconds to 1.63 seconds without losing almost no accuracy (see Table 4). There are also several other possibilities that we have not yet explored such as tuning the number of feature channels and/or network layers of the CNN architecture (similar to the DeepBox [Kuo et al.(2015)Kuo, Hariharan, and Malik] and the SharpMask [Pinheiro et al.(2016)Pinheiro, Lin, Collobert, and Dollár] approaches).

In the remainder of this section we will use the fast version of our AttractioNet approach in order to provide experimental results.

2 Generalization to unseen categories

So far we have evaluated our AttractioNet approach — in the end task of object box proposal generation — on the COCO validation set and the PASCAL VOC2007 test set that are labelled with the same or a subset of the object categories "seen" in the training set. In order to assess the AttractioNet’s capability to generalize to "unseen" categories, as it is suggested by Chavali et al\bmvaOneDot [Chavali et al.(2015)Chavali, Agrawal, Mahendru, and Batra], we evaluate our AttractioNet model on two extra datasets that are labelled with object categories that are not present in its training set ("unseen" object categories).

From COCO to ImageNet [Russakovsky et al.(2015)Russakovsky, Deng, Su, Krause, Satheesh, Ma, Huang, Karpathy, Khosla, Bernstein, et al.]. Here we evaluate our COCO trained AttractioNet box proposal model on the ImageNet [Russakovsky et al.(2015)Russakovsky, Deng, Su, Krause, Satheesh, Ma, Huang, Karpathy, Khosla, Bernstein, et al.] ILSVRC2013 detection task validation set that is labelled with 200 different object categories and we report average recall results in Table 5. Note that among the 200 categories of ImageNet detection task, 60 of them, as we identified, are also present in the AttractioNet’s training set (see Appendix C). Thus, for a better insight on the generalization capabilities of AttractioNet, we divided the ImageNet detection task categories on two groups, the "seen" by AttractioNet categories and the "unseen" categories, and we report the average recall results separately for those two groups of object categories in Table 5. For comparison purposes we also report the average recall performance of a few indicative other box proposal methods that their code is publicly available. We observe that, despite the performance difference of our approach between the "seen" and the "unseen" object categories (which is to be expected), its average recall performance on the "unseen" categories is still quite high and significantly better than the other box proposal methods. Note that even the non-learning based approaches of Selective Search and EdgeBoxes exhibit a performance drop on the "unseen" by AttractioNet group of object categories, which we assume is because this group contains more intrinsically difficult to discover objects.

From COCO to NYU Depth dataset [Nathan Silberman and Fergus(2012)]. The NYU Depth V2 dataset [Nathan Silberman and Fergus(2012)] provides 1449 images (recorded from indoor scenes) that are densely pixel-wise annotated with 864 different categories. We used the available instance-wise segmentations to create ground truth bounding boxes and we tested our COCO trained AttractioNet model on them (see Table 6). Note that among the 864 available pixel categories, a few of them are "stuff" categories (e.g. wall, floor, ceiling or stairs) or in general non-object pixel categories that our object box proposal method should by definition not recall. Thus, during the process of creating the ground truth bounding boxes, those non-object pixel segmentation annotations were excluded (see Appendix D). In Table 6 we report the average recall results of our AttractioNet method as well as of a few other indicative methods that their code is publicly available. We again observe that our method surpasses all other approaches by a significant margin. Furthermore, in this case the superiority of our approach is more evident on the average recall of the small and medium sized objects.

To conclude we argue that our learning based AttractioNet approach exhibits good generalization behaviour. Specifically, its average recall performance on the "unseen" object categories remains very high and is also much better than other competing approaches, including both learning-based approaches such as the MSG and hand-engineered ones such as the Selective Search or the EdgeBoxes methods. A performance drop is still observed while going from "seen" to "unseen" categories, but this is something to be expected given that any machine learning algorithm will always exhibit a certain performance drop while going from "seen" to "unseen" data (i.e. training set accuracy versus test set accuracy).

3 AttractioNet box proposals evaluation in the context of the object detection task

Here we evaluate our AttractioNet box proposals in the context of the object detection task by training and testing a box proposal based object detection system on them (specifically we use the fast version of AttractioNet that is described in section 3.1.3).

Detection system. Our box proposal based object detection system consists of a Fast-RCNN [Girshick(2015)] category-specific recognition module and a LocNet Combined ML [Gidaris and Komodakis(2016)] category-specific bounding box refinement module (see Appendix B for more details). As post-processing we use a non-max-suppression step (with IoU threshold of 0.35) that is enhanced with the box voting technique described in the MR-CNN system [Gidaris and Komodakis(2015)] (with IoU threshold of 0.75). Note that we did not include iterative object localization as in the LocNet [Gidaris and Komodakis(2016)] or MR-CNN [Gidaris and Komodakis(2015)] papers, since our bounding box proposals are already very well localized and we did not get any significant improvement from running the detection system for extra iterations. Using the same trained model we provide results for two test cases: (1) using a single scale of 600 pixels during test time and (2) using two scales of 500 and 1000 pixels during test time.

Detection evaluation setting. The detection evaluation metrics that we use are the average precision (AP) for the IoU thresholds of 0.500.50 (AP@0.50@0.50), 0.750.75 (AP@0.75@0.75) and the COCO style of average precision (AP@0.50:0.95@0.50:0.95) that averages the traditional AP over several IoU thresholds between 0.500.50 and 0.950.95. Also, we report the COCO style of average precision with respect to the small (AP@@Small), medium (AP@@Medium) and large (AP@@Large) sized objects. We perform the evaluation on 5k5k images of COCO 2014 validation set and we provide final results on the COCO 2015 test-dev set.

Detection results. In Figure 9 we provide plots of the achieved average precision (AP) as a function of the used box proposals number and in Table 7 we provide the average precision results for 10, 100, 1000 and 2000 box proposals. We observe that in all cases, the average precision performance of the detection system seems to converge after the 200 box proposals. Furthermore, for single scale test case our best COCO-style average precision is 0.320 and for the two scales test case our best COCO-style average precision is 0.337. By including horizontal image flipping augmentation during test time our COCO-style average precision performance is increased to 0.343. Finally, in Table 8 we provide the average precision performance in the COCO test-dev 2015 set where we achieve a COCO-style AP of 0.341. By comparing with the average precision performance of the other competing methods, we observe that:

Comparing with the other VGG16-Net based object detection systems (ION [Bell et al.(2015)Bell, Zitnick, Bala, and Girshick] and MultiPath [Zagoruyko et al.(2016)Zagoruyko, Lerer, Lin, Pinheiro, Gross, Chintala, and Dollár] systems), our detection system achieves the highest COCO-style average precision with its main novelties w.r.t. the Fast R-CNN [Girshick(2015)] baseline being (1) the use of the AttractioNet box proposals that are introduced in this paper and (2) the LocNet [Gidaris and Komodakis(2016)] category specific object location refinement technique that replaces the bounding box regression step.

Comparing with the ION [Bell et al.(2015)Bell, Zitnick, Bala, and Girshick] detection system, which is also VGG16-Net based, our approach is better on the COCO-style AP metric (that favours good object localization) while theirs is better on the typical AP@@0.50 metric. We hypothesize that this is due to the fact that our approach targets to mainly improve the localization aspect of object detection by improving the box proposal generation step while theirs the recognition aspect of object detection. The above observation suggests that many of the novelties introduced on the ION [Bell et al.(2015)Bell, Zitnick, Bala, and Girshick] and MultiPath [Zagoruyko et al.(2016)Zagoruyko, Lerer, Lin, Pinheiro, Gross, Chintala, and Dollár] systems w.r.t. object detection could be orthogonal to our box proposal generation work.

The achieved average precision performance of our VGG16-Net based detection system is close to the state-of-the-art ResNet-101 based Faster R-CNN+++ detection system [He et al.(2015b)He, Zhang, Ren, and Sun] that exploits the recent successes in deep representation learning introduced — under the name Deep Residual Networks — in the same work by He et al\bmvaOneDot [He et al.(2015b)He, Zhang, Ren, and Sun]. Presumably, our overall detection system could also be benefited by being based on the Deep Residual Networks [He et al.(2015b)He, Zhang, Ren, and Sun] or the more recent wider variant called Wide Residual Networks [Zagoruyko and Komodakis(2016)].

Finally, our detection system has the highest average precision performance w.r.t. the small sized objects, which is a challenging problem, surpassing by a healthy margin even the ResNet-101 based Faster R-CNN+++ detection system [He et al.(2015b)He, Zhang, Ren, and Sun]. This is thanks to the high average recall performance of our box proposal method on the small sized objects.

4 Qualitative results

In Figure 10 we provide qualitative results of our AttractioNet box proposal approach on images coming from the COCO validation set. Note that our approach manages to recall most of the objects in an image, even in the case that the depicted scene is crowded with multiple objects that heavily overlap with each other.

Conclusions

In our work we propose a bounding box proposals generation method, which we call AttractioNet, whose key elements are a strategy for actively searching of bounding boxes in the promising image areas and a powerful object location refinement module that extends the recently introduced LocNet [Gidaris and Komodakis(2016)] model on localizing objects agnostic to their category. We extensively evaluate our method on several image datasets (i.e. COCO, PASCAL, ImageNet detection and NYU-Depth V2 datasets) demonstrating in all cases average recall results that surpass the previous state-of-the-art by a significant margin while also providing strong empirical evidence about the generalization ability of our approach w.r.t. unseen categories. Even more, we show the significance of our AttractioNet approach in the object detection task by coupling it with a VGG16-Net based detector and thus managing to surpass the detection performance of all other VGG16-Net based detectors while even being on par with a heavily tuned ResNet-101 based detector. We note that, apart from object detection, there exist several other vision tasks, such as exemplar 2D-3D detection [Massa et al.(2016)Massa, Russell, and Aubry], visual semantic role labelling [Gupta and Malik(2015)], caption generation [Karpathy and Fei-Fei(2015)] or visual question answering [Shih et al.(2015)Shih, Singh, and Hoiem], for which a box proposal generation step can be employed. We are thus confident that our AttractioNet approach could have a significant value with respect to many other important applications as well.

Acknowledgements

This work was supported by the ANR SEMAPOLIS project. We would like to thank Pedro O. Pinheiro for help with the experimental results and Sergey Zagoruyko for helpful discussions. We would also like to thank the authors of SharpMask [Pinheiro et al.(2016)Pinheiro, Lin, Collobert, and Dollár] (Pedro O. Pinheiro, Tsung-Yi Lin, Ronan Collobert and Piotr Dollar) for providing us with its box proposals.

References

Appendix A Multi-threshold non-max-suppresion re-ordering

As already described in section 2.1, at the end of our active box proposal generation strategy we include a non-maximum-suppression [Felzenszwalb et al.(2010)Felzenszwalb, Girshick, McAllester, and Ramanan] (NMS) step that is applied on the set C of scored candidate box proposals in order then to take the final top KK output box proposals (see algorithm 1). However, the optimal IoU threshold (in terms of the achieved AR) for the NMS step depends on the desired number KK of output box-proposals. For example, for 10, 100, 1000 and 2000 proposals the optimal IoU thresholds are 0.550.55, 0.750.75, 0.900.90 and 0.950.95 respectively. Since our plan is to make our box proposal system publicly available, we would like to make its use easier for the end user. For that purpose, we first apply on the set C of scored candidate box proposals a simple NMS step with IoU threshold equal to 0.950.95 in order to then get the top 2000 box proposals and then we follow a multi-threshold non-max-suppression technique that re-orders this set of 2000 box proposals such that for any given number KK the top KK box proposals in the set better cover (in terms of achieved AR) the objects in the image.

Appendix B Detection system

In this section we provide further implementation details about the object detection system used in §3.3.

Architecture. Our box proposal based object detection network consists of a Fast-RCNN [Girshick(2015)] category-specific recognition module and a LocNet Combined ML [Gidaris and Komodakis(2016)] category-specific bounding box refinement module that share the same image-wise convolutional layers (conv1_1 till conv5_3 layers of VGG16-Net).

Training. The detection network is trained on the union of the COCO train set that includes around 80k80k images and on a subset of the COCO validation set that includes around 35k35k images (the remaining 5k5k images of COCO validation set are being used for evaluation). For training we use our AttractioNet box proposals and we define as positives those that have IoU overlap with any ground truth bounding box at least 0.5 and as negatives the remaining proposals. For training we use SGD where each mini-batch consists of 4 images with 64 box proposals each (256 boxes per mini-batch in total) and the ratio of negative-positive boxes is 3:1. We train the detection network for 500k500k SGD iterations starting with a learning rate of 0.0010.001 and dropping it to 0.00010.0001 after 320k320k iterations. We use the same scale and aspect ratio jittering technique that is used on AttractioNet and is described in section 2.3.

Appendix C Common categories between ImageNet and COCO

In this section we list the ImageNet detection task object categories that we identified to be present also in the COCO dataset. Those are: airplane, apple, backpack, baseball, banana, bear, bench, bicycle, bird, bowl, bus, car, chair, cattle, computer keyboard, computer mouse, cup or mug, dog, domestic cat, digital clock, elephant, horse, hotdog, laptop, microwave, motorcycle, orange, person, pizza, refrigerator, sheep, ski, tie, toaster, traffic light, train, zebra, racket, remote control, sofa, tv or monitor, table, watercraft, washer, water bottle, wine bottle, ladle, flower pot, purse, stove, koala bear, volleyball, hair dryer, soccer ball, rugby ball, croquet ball, basketball, golf ball, ping-pong ball, tennis ball.

Appendix D Ignored NUY Depth dataset categories

In this section we list the 12 most frequent non-object categories that we identified on the NUY Depth V2 dataset: curtain, cabinet, wall, floor, ceiling, room divider, window shelf, stair, counter, window, pipe and column.