WSOD^2: Learning Bottom-up and Top-down Objectness Distillation for Weakly-supervised Object Detection

Zhaoyang Zeng, Bei Liu, Jianlong Fu, Hongyang Chao, Lei Zhang

Introduction

The capability of recognizing and localizing objects in an image reveals a deep understanding of visual information, and has attracted many attentions in recent years. Significant progresses have been achieved with the development of convolutional neural network (CNN) . However, current state-of-the-art object detectors mostly rely on a large scale of training data which requires manually annotated bounding boxes (e.g., PASCAL VOC 2007/2012 , MS COCO , Open Images ). To relieve the heavy labeling effort and reduce cost, weakly-supervised object detection paradigm has been proposed by leveraging only image-level annotations .

To address weakly-supervised object detection (WSOD) task, most previous works adopt multiple instance learning method to transform WSOD into multi-label classification problems . Later on, online instance classifier refinement (OICR) and proposal cluster learning (PCL) are proposed to learn more discriminative instance classifiers by explicitly assigning instance labels. Both OICR and PCL adopt the idea of utilizing the outputs of initial object detector as pseudo ground truths, which has been shown benefits in improving the classification ability of WSOD. However, a classification model often targets at detecting the existence of objects for a category, while it is not able to predict the location, size and the number of objects in images. This weakness usually results in the detection of partial or oversized bounding boxes, as shown in the first and third rows in Figure 1. The performances of OICR and PCL heavily rely on the accuracy of the initial object detection results, which limit further improvement with large margins. Also, they neglect learning bounding box regression, which plays an important role in the design of modern object detectors . C-WSL integrates bounding box regressors into OICR framework to reduce localization errors, however, it relies on a greedy ground truths selection strategy which requires additional counting annotations .

Existing works that rely on the initial weakly-supervised object detection results try to learn the object boundary from feature maps by convolutional neural network (CNN). Although CNN is an expert to learn discriminative local features of an object with image-level labels in a top-down fashion (we call it top-down classifiers in this work), it performs poorly in detecting whether a bounding box contains a complete object without the ground truth for supervision.

Some low-level feature based object evidences (e.g. color contrast and superpixels straddling ) have been proposed to measure a generic objectness that quantifies how likely a bounding box contains an object of any class in a bottom-up way. Inspired by these bottom-up object evidences, in this work, we explore to use their advantage for improving the capability of a CNN model in capturing objectness in images. We propose to integrate these bottom-up evidences that are good at discovering boundary and CNN with powerful representation ability in a single network.

We propose a WSOD framework with Objectness Distillation (WSOD2) to leverage bottom-up object evidences and top-down classification output with a novel training mechanism. First, given an input image with thousands of region proposals (e.g., generated by Selective Search ), we learn several instance classifiers to predict classification probabilities of each region proposal. Each of these classifiers can help to select multiple high-confident bounding boxes as possible object instances (i.e., pseudo classification and bounding box regression ground truths). Second, we incorporate a bounding box regressor to fine-tune the location and size of each proposal. Third, as each bounding box cannot capture precise object boundaries by CNN features alone, we combine bottom-up object evidences and top-down CNN confidence scores in an adaptive linear combination way to measure the objectness of each candidate bounding box, and assign labels for each region proposal to train the classifiers and regressor.

For some discriminative small bounding boxes that CNN prefers, the bottom-up object evidence (e.g., superpixels straddling) tends to be very low. WSOD2 can regulate pseudo ground truths to satisfy both higher CNN confidence and low-level object completeness. In addition, a bounding box regressor is integrated to reduce the localization error, and augment the effect of bottom-up object evidences during training at the same time. We design an adaptive training strategy to make the guidance gradually distilled, which enables that a CNN model can be trained strong enough to represent both discriminative local and boundary information of objects when the model converges.

To the best of our knowledge, this work is the first to explore bottom-up object evidences in weakly-supervised object detection task. The contribution can be summarized as follows:

We propose to combine bottom-up object evidences with top-down class confidence scores in weakly-supervised object detection task.

We propose WSOD2 (WSOD with objectness distillation) to distill object boundary knowledge in CNN by a bounding box regressor and an adaptive training mechanism.

Our experiments on PASCAL VOC 2007/2012 and MS COCO datasets demonstrate the effectiveness of the proposed WSOD2.

Related Work

Weakly-supervised object detection has attracted many attentions in recent years. Most existing works adopt the idea of multiple-instance learning to transform weakly-supervised object detection into multi-label classification problems. Bilen et al. proposes WSDDN which performs multiplication on the score of classification and detection branches, so that high-confident positive samples can be selected. Tang et al. and Tang et al. find that online transforming image-level label into instance-level supervision is an effective way to boost the accuracy, and thus propose to online refine several branches of instance classifiers based on the outputs of previous branches. As class activation map produced by a classifier can roughly localize the object , Wei et al. tries to utilize it to generate course detection results, and use them as reference for the later refinement. Most previous works rely heavily on pseudo ground truths mining, either online (inside training loop) or offline (after training). Such pseudo ground truths are determined by classification confidence or hand-crafted rules , which are not accurate to measure the objectness of regions.

2 Bounding Box Regression

Bounding box regression is proposed in , and is adopted by almost all recent CNN-based fully-supervised object detectors since it can reduce the localization errors of predicted boxes. However, only a few works introduce bounding box into weakly-supervised object detection due to the lack of supervision. Some works consider bounding box regression as a post-processing module. Among which, OICR directly uses the detection results of training set to train Fast R-CNN. W2F designs some strategies to offline select pseudo ground truth with high precision, based on the output of OICR. Differently, Gao et al. integrate bounding box regressors into OICR inside training loop which leverage addition counting information to help selecting pseudo ground truths.

In this paper, we integrate bounding box regressor into weakly-supervised detector, and assign regression targets by novelly leveraging bottom-up object evidence.

Approach

The overview of our proposed weakly-supervised object detector with objectness distillation (WSOD2) is illustrated in Figure 2. We first adopt a based multiple instance detector (i.e. Cls 0) to obtain the initial detected object bounding boxes. Based on the localization of each proposed bounding box, we compute the bottom-up object evidence. Such evidence serves as guidance to transform image-level labels into instance-level supervision. We optimize the whole network in an end-to-end and adaptive fashion. In this section, we will introduce WSOD2 in detail.

In weakly-supervised object detection, only image-level annotations are available. To better understand semantic information inside an image, we need to go deep into region-level, and analyze the characteristic of each box. We first build a base detector to obtain initial detection result. We follow WSDDN to adopt the idea of multiple instance learning to optimize the base detector by transforming WSOD into multi-label classification problem. Specifically, given an input image, we first generate region proposals RR by Selective Search and extract region features x by a CNN backbone, an RoI Pooling layer and two fully-connected layers.

where [σc]ij\left[\sigma^{c}\right]_{ij} denotes the prediction of ithi^{th} class label for jthj^{th} region proposal, and [σd]ij\left[\sigma^{d}\right]_{ij} is the weight learned of jthj^{th} region proposal for ithi^{th} class. We compute the proposal scores by element-wise product s=σc⊙σds=\sigma^{c}\odot\sigma^{d}, and aggregate over the region dimensions to obtain image-level score vector ϕ=[ϕ1,ϕ2,⋯ ,ϕC]\phi=[\phi_{1},\phi_{2},\cdots,\phi_{C}] by ϕc=∑r=1∣R∣[s]cr\phi_{c}=\sum_{r=1}^{|R|}{\left[s\right]_{cr}}. In such way, we can utilize the image-level class label as supervision and apply binary cross-entropy loss to optimize the base detector. The base loss function is denoted as:

where ϕc^=1\hat{\phi_{c}}=1 indicates that the input image contains cthc^{th} class, and ϕc^=0\hat{\phi_{c}}=0 otherwise. The prediction score ss is considered as initial detection result. However, it is not precise enough and can be further refined as discussed in .

2 Bottom-up and Top-Down Objectness

The essence of an object detector is a bounding box ranking function, in which objectness measurement is an important factor. It is common to consider classification confidence as objectness score in recent CNN-based detectors . However, such strategy has a flaw in weakly-supervised scenario that it is difficult for trained detectors to distinguish complete objects from discriminate object parts or irrelevant background. To relieve this issue, we explore bottom-up object evidences (e.g., superpixels straddling) which play important roles in traditional object detection.

As stated in , objects are standalone things with well-defined boundaries and centers. Thus, we expect a box with a complete object to have a higher objectness score than a partial, oversized or background box. Bottom-up object evidence summarizes the boundary characteristic of common objects, which can help make up for the boundary discovering weakness of CNN.

We propose to integrate bottom-up object evidence to train weakly-supervised object detectors. Specifically, inspired by OICR , we build KK instance classifiers on top of x, consider the output of kthk^{th} classifier as the supervision of (k+1)th(k+1)^{th} one, and exploit bottom-up object evidence to guide the network training. Each classifier is implemented by a fully-connected layer and a softmax layer along C+1C+1 categories (we consider background as 0th0^{th} class). Formally, for kthk^{th} classifier, we define the refinement loss function of kthk^{th} classifier as:

where prkp^{k}_{r} denotes the {C+1}\{C+1\}-dim output class probability of proposal rr, and p^rk\hat{p}^{k}_{r} indicates its ground truth one-hot label. CE(prk,p^rk)=−∑c=0Cp^rcklog(prck)CE(p^{k}_{r},\hat{p}^{k}_{r})=-\sum_{c=0}^{C}\hat{p}^{k}_{rc}log(p^{k}_{rc}) is a standard cross entropy function. Since the real instance-level ground truth labels are unavailable, we use an online strategy to dynamically select pseudo ground truth labels of each proposal in training loop, which will be further explained in Sec 3.4. We online assign loss weight wrkw_{r}^{k} based on the objectness of proposal rr. Specifically, we first extract bottom-up evidence of rr and denote it as Obu(r)O_{bu}(r), then integrate Obu(r)O_{bu}(r) with Otdk(r)O_{td}^{k}(r), which is the class confidence produced by kthk^{th} classifier. wrkw_{r}^{k} is a linear combination of bottom-up evidence and top-down confidence as follows:

where α\alpha denotes the impact factor of bottom-up object evidence. Three terms in Eqn 4 are defined as follows:

Bottom-up object evidence ObuO_{bu}. We mainly adopt Superpixels Straddling(SS) as bottom-up evidence in this work, and we also explore other three evidences: textbfMulti-scale Saliency(MS), Color Constrast(CC) and Edge Density(ED). Experiment details of these evidences can be found in Sec. 4.2.

Top-down Class Confidence OtdO_{td}. We compute top-down confidence of current branch based on the output of previous branch. Specifically, once we obtain class probability prk−1p^{k-1}_{r} of (k−1)th(k-1)^{th} branch, top-down class confidence of kthk^{th} branch is computed as:

Since p^k\hat{p}^{k} is a one-hot vector, only one value of pk−1p^{k-1} will be picked to computed Otdk(r)O^{k}_{td}(r).

Impact Factor α\alpha. α\alpha is the impact factor to balance the effect of bottom-up object evidence and top-down class confidence, which is computed by some weight decay functions. Such design enables boundary knowledge to be distilled into CNN, which will be detailed discussed in Sec. 3.4.

As bottom-up object evidence and top-down class confidence can measure how likely a box contain a object from the perspective of boundary and semantic information, we consider these two representations as bottom-up and top-down objectness, respectively.

3 Bounding Box Regression

Bottom-up object evidence is capable to discovery object boundary, so we explore how to make it guide the pre-computed bounding boxes updated during training. An intuitive idea is to integrate bounding box regression to refine the positions and sizes of proposals.

Bounding box regression is a necessary component in typical fully-supervised object detector, as it is able to reduce localization errors. Although bounding box annotations are unavailable in weakly-supervised object detection, some existing works shows that online or offline mining pseudo ground truths and regressing them can boost the performance a lot. Inspired by this idea, we integrate a bounding box regressor on the top of x, and make it can be online updated. The bounding box regressor has the same formulation as in Fast R-CNN . For region proposal rr, the regressor predicts offsets of locations and sizes tr=(trx,try,trw,trh)t_{r}=(t^{x}_{r},t^{y}_{r},t^{w}_{r},t^{h}_{r}), and is further optimized as follows:

where t^r\hat{t}_{r} is computed by the coordinates and sizes difference between rr and r^\hat{r} as described in , where r^\hat{r} indicates the regression reference. RposR_{pos} indicates positive (non-background) regions, which will be explained in Sec. 3.4. smoothL1smooth_{L1} function is the same function as defined in . wrKw_{r}^{K} denotes the regression loss weights computed by the last classification branch. We compute pseudo regression reference r^\hat{r} based on the influence of wrKw_{r}^{K} which evaluates the objectness of a proposal as we stated in Sec. 3.2:

where MM is positive sample mining function which will be explained in Sec 3.4, and TiouT_{iou} is a specific IoU threshold. Eqn 7 enables each positive region sample to approach a nearby box which has the high objectness.

We adopt bounding box regression to augment the box prediction during training. We update Eqn 4 as:

where r′r^{\prime} is rr offset by trt_{r}. We keep Otdk(r)O^{k}_{td}(r) unchanged because OtdkO^{k}_{td} contains a RoI feature warping operation, which will be affected by bounding box prediction. In this new formulation, the localization of proposals is online updated. The updated boxes may achieve higher objectness, which means more precise and complete regression targets have higher probability to be selected.

4 Objectness Distillation

Eqn 3 has the similar formulation as knowledge distillation , where the external knowledge comes from bottom-up and top-down objectness. Inside which, α\alpha is a weight to balance each knowledge. At the beginning of training, top-down classifiers are not reliable enough, so we expect bottom-up evidences to take the dominant place in the combination (i.e. Eqn 4). With the guidance of bottom-up evidences, the network will try to regulate the confidence distribution of top-down classifiers to comply with bottom-up evidences. We call this process objectness distillation.

As the training proceeds, the reliability of OtdO_{td} increases, and OtdO_{td} inherits the boundary decision ability from ObuO_{bu}, while it still keeps the semantic understanding ability because of the classification supervision. Therefore, α\alpha can gradually move the attention from bottom-up object evidences to top-down CNN confidences. Specifically, α\alpha is computed by some weight decay functions. We survey several weight decay functions including polynomial, cosine and constant functions, and we will compare the effectiveness of different functions in Sec 4.2.

Except α\alpha, to enable objectness distillation, we also need to determine p^rk\hat{p}^{k}_{r}. We want to leverage bottom-up evidences to enhance boundary representation while keep the semantic recognition ability, thus we utilize output from previous branch of classifier to mine positive proposals.

Given the output from (k−1)th(k-1)^{th} classifier, we mine pseudo ground truths by following steps:

We apply Non-Maximum Suppression (NMS) on RR based on class probability prk−1p^{k-1}_{r} of each proposal rr using a pre-defined threshold TnmsT_{nms}. We denote the kept boxes as RkeepR_{keep}.

For each category c(c>0)c(c>0), if ϕ^c=1\hat{\phi}_{c}=1, we seek all boxes from RkeepR_{keep} whose class confidences on category cc are greater than another pre-defined threshold TconfT_{conf}, and assign these boxes category label cc. Specially, if no box is selected, we seek the one with highest score. The set of all seek boxes is denoted as RseekR_{seek}.

For each seed box in RseekR_{seek}, we seek all its neighbor boxes in RR. Here we consider a box is the neighbor of another box if their Intersection over Union (IoU) is greater than a threshold TiouT_{iou}. We denote the set of all neighbor boxes as RneighborR_{neighbor}. All neighbor boxes will be assigned the same class label as their seed boxes. Other non-seed and non-neighbor boxes will be considered as background. We transform the assigned labels to one-hot vector to obtain all p^rk\hat{p}^{k}_{r}.

Finally, we consider the union set of RseekR_{seek} and RneighborR_{neighbor} as the positive proposals: Rpos=Rseek∪RneighborR_{pos}=R_{seek}\cup R_{neighbor}.

We group the above operations into function M(k,R)M(k,R) which will return the set of positive proposals, as we mentioned in Sec 3.2 and Sec 3.3. By such way, close positive samples will be assigned same category label, while sample with high objectness will receive high weight. Such information will be distilled into CNN by optimization, thus CNN will gradually increase the ability of discovering object boundary.

5 Training and Inference Details

Training. The overall learning target is formulated as:

where λ1\lambda_{1} and λ2\lambda_{2} are hyper-parameters to balance loss weights.. We adopt λ1=1\lambda_{1}=1 and λ2=0.3\lambda_{2}=0.3, and follow to set K=3K=3. Since the supervision of all KK classifiers comes from previous branches, we set α=0\alpha=0 in the first 2,0002,000 iterations for warm-up. When mining pseudo ground truths, typically we follow to set Tnms=0.3,Tconf=0.7,Tiou=0.5T_{nms}=0.3,T_{conf}=0.7,T_{iou}=0.5.

Inference. Our model have KK refinement classifiers and one bounding box regressor. For each predicted box, we follow to average the outputs from all KK classifiers to produce the class confidence, and adjust its position and size using the bounding box regressor. Finally, we apply NMS with threshold 0.30.3 to remove redundant detected boxes.

Experiments

Datasets and evaluation metrics. We evaluate our approach on three object detection benchmarks: PASCAL VOC 2007 & 2012 and MS COCO . After removing the bounding box annotations provided by these datasets, we only use images and their label information for training. PASCAL VOC 2007 and 2012 consists of 9,9629,962 and 22,53122,531 images of 2020 categories, respectively. For PASCAL VOC, we train on trainval split (5,0115,011 images for 2007 and 11,54011,540 for 2012), report mean average precision (mAP) on test split, and also adopt correct localization (CorLoc) on trainval split to measure the localization accuracy. Both two metrics are performed under the condition of IoU>0.5IoU>0.5 as a standard setting. MS COCO contains 8080 categories. We train on train2014 split and evaluate on val2014 split, which consists of 82,78382,783 and 40,50440,504 images, respectively. We report AP@.50AP@.50 and AP@[.50:.05:.95]AP@\left[.50:.05:.95\right] on val2014.

Implementation details. We adopt VGG16 as the CNN backbone, and use parameters pre-trained on ImageNet for initialization. We randomly initialize the weights of all new layers using Gaussian distributions with -mean and standard deviations 0.010.01 (except 0.0010.001 for bounding box regressor), and initialize all new biases to . We follow a widely-used setting to use Selective Search to generate about 2,0002,000 proposals for each image. The whole network is end-to-end optimized using SGD with an initial learning rate of 10−310^{-3}, weight decay of 0.00050.0005 and momentum of 0.90.9. The overall iteration step number is set to 80,00080,000 on VOC 2007, and the learning rate will be divided by 1010 at 40,000th40,000^{th} step. For VOC 2012 we double the iteration step number and learning rate decay step is also doubled to 80,000th80,000^{th} step. For MS COCO we set iteration step number to 360,000360,000, and make learning rate decay at 180,000th180,000^{th} step. We follow to adopt multi-scale settings in training. Specifically, the short edge of the input image will be randomly re-scaled to a scale in {480,576,588,864,1280}\{480,576,588,864,1280\}, and we restrict the length of the long edge not greater than 20002000. Besides, horizontal flip of all training images will be also used for training. We report single-scale testing results for ablation study, and report multi-scale testing results when comparing with previous works. All our experiments are implemented based on PyTorch on 44 NVIDIA P100 GPUs.

2 Ablation Study

We conduct ablation studies to demonstrate the effectiveness of WSOD2 on PASCAL VOC 2007.

Bottom-up evidences. For bottom-up object evidence, we test the effect of four evidences in both individual and combined ways. The four evidences are list as follows:

1) Multi-scale Saliency(MS) which summarizes the saliency over several scales;

2) Color Constrast(CC) which computes the color distribution difference with immediate surrounding area;

3) Edge Density(ED) which computes the density of edges in the inner rings;

4) Superpixels Straddling(SS) which analyzes the straddling of all superpixels.

Since the value ranges of different evidences are inconsist, we normalize the computed value to [0−1]\left[0-1\right]. For CC, ED and MS, we fix their parameters by setting θMS=0.2,θCC=2,θED=2\theta^{MS}=0.2,\theta^{CC}=2,\theta^{ED}=2 empirically due to the lack of supervision. For SS, we follow to set θσSS=0.8,θkSS=300\theta^{SS}_{\sigma}=0.8,\theta^{SS}_{k}=300. We refer the readers to for more details of these four evidences and the meaning of θMS,θCC,θED,θSS\theta^{MS},\theta^{CC},\theta^{ED},\theta^{SS}.

To easier analyze the effect of these bottom-up evidences, we simply keep α=1\alpha=1 in this ablation experiment for all settings that include these evidences, and α=0\alpha=0 for the method that does not involve any bottom-up evidence as a baseline for comparison. We also test the combination of these four evidences by their average. As discussed in , linear combination is not a good way to combine them, we conduct this experiment only for evaluating the effectiveness of bottom-up evidences and inspiring future works.

The results are shown in Table 1. From the comparison with the baseline, we can find that the performance can increase significantly with the guidance of bottom-up evidences. Table 1 also includes AP on all categories, from which we find that different evidences may favor different categories. For example, for single evidence, ED favors to “boat”, while not performs good on “tv”. Moreover, we can find that this result also agrees with the performance that measures objectness of each evidence as reported in , which indicates that these bottom-up evidences are positive correlated to object detection performance. From the result of their combination, we can find it achieves better performance than all single evidences except SS. We believe that linear average is not a correct way to combine these evidence, and better ways can be explored in the future. We adopt SS as bottom-up object evidence in later experiments.

Impact factor α\alpha. We test several weight decay functions, including constant (α=0,0.5,1\alpha=0,0.5,1), polynomial(α=−(n/N)γ+1\alpha=-(n/N)^{\gamma}+1, where γ=2,3,1,1/2,1/3\gamma=2,3,1,1/2,1/3) and cosine (α=(1+cos(nπ/N)/2\alpha=(1+cos(n\pi/N)/2) functions where nn and NN indicate current step and total step number, respectively. The results are shown in Figure 3. From the comparison of the first three lines, we find that bottom-up evidences will help the model learn the boundary representation and results in better object detection result. Among different designs, linear decay (i.e., α=−(n/N)+1\alpha=-(n/N)+1) performs best and the later experiments are conducted based on this setting. We remain exploration of the best parameters for future study.

Effect of each component. Table 2 shows the effectiveness of each component. We can find that the bounding box regressor brings at least 2.62.6 mAP improvement. Settings that do not use NMS means directly consider the highest-confident box for each category as seed box as OICR . NMS can also improve 0.80.8 mAP. Details of bottum-up evidences (BU) and α\alpha decay function are discussed above, where both bottom-up evidences and α\alpha decay function can bring 2.22.2 mAP improvement.

3 Comparisons with State-of-the-Arts

We evaluate WSOD2 on PASCAL VOC 2007 & 2012 and MS COCO datasets , report the performances and compare with state-of-the-art weakly-supervised detectors. As most of our compared approaches adopt multi-scale testing, we report our multi-scale testing results.

AP evaluation on PASCAL VOC. From Table 4 we can find that WSOD2 achieves 53.653.6 mAP on PASCAL VOC 2007, which significantly outperforms other end-to-end trainable models with at least 5.35.3 mAP. WSOD2 is also robust on PASCAL VOC 2012 and achieves 47.247.2 mAP, which is shown in Table 5.

Besides, we follow the common setting in fully-supervised object detection to train WSOD2 on PASCAL VOC 07+12 trainval splits, and denote it as WSOD2∗{}^{2}*. Such setting achieves a surprising mAP score 56.156.1 as shown in the last row of Table4.

CorLoc evaluation on PASCAL VOC. CorLoc evaluates the localization accuracy of detectors on training set. We report results on PASCAL VOC 2007 and 2012 trainval split in Table 4 and Table 5, respectively. We can find that WSOD2 significantly surpasses outperforms other end-to-end trainable models on both PASCAL VOC 2007 and 2012.

AP evaluation on MS COCO. We report results on MS COCO dataset in Table 6. Since few works report results on MS COCO dataset, we only compare performance with and . We can find that WSOD2 outperforms compared works by at least 22 AP.

4 Visualization and Case Study

We make a qualitative analysis of the effectiveness of WSOD2 compared with OICR. We extract the conv5 features of trained models, and visualize some cases in Figure 4. The highlighted parts indicate the high response area of the input image in CNN. Compared with OICR, WSOD2 can gradually transfer the response area from discriminate parts to complete objects.

Figure 5 exhibits some successful and failure cases of WSOD2. We obverse that WSOD2 can well handle multiple discrete instances, while there still remains a challenge to solve detection problem in dense scenarios. We also find that for “person” class, most weakly-supervised object detectors tend to find human faces. The reason is that in the current datasets, human face is the most common pattern of “person”, while other parts are often missed in the image. This remains a challenging problem and we can consider leveraging human structure prior in the future.

Conclusion

In this paper, we propose a novel weakly-supervised object detection with bottom-up and top-down objectness distillation (i.e., WSOD2) to improve the deep objectness representation of CNN. Bottom-up object evidence, which could measures the probability of a bounding box including a complete object, is utilized to distill boundary features in CNN in an adaptive training way. We also propose a training strategy that integrates bounding box regression and progressive instance classifier in an end-to-end fashion. We conduct experiments on some standard datasets and settings for WSOD task with our approach. Results demonstrate the effectiveness of our proposed WSOD2 in both quantitative and qualitative way. We also make a thorough analysis on the challenges and possible improvement (e.g., for “person” class) of WSOD problem.

Acknowledgments

This work is partially supported by NSF of China under Grant 61672548, U1611461, 61173081, and the Guangzhou Science and Technology Program, China, under Grant 201510010165.

References