Deep Patch Learning for Weakly Supervised Object Classification and Discovery

Peng Tang, Xinggang Wang, Zilong Huang, Xiang Bai, Wenyu Liu

Introduction

In this paper, we study the problems of weakly supervised object classification and discovery, which are with great importance in computer vision community. As shown in the top and middle of Fig. 1, given an input image and its category labels (e.g., image-level annotations), object classification is to learn object classifiers for classifying which object classes (e.g., person) appear in testing images.We refer this task as weakly supervised object classification since it does not require patch-level annotations for training. Similar to object detection, object discovery is to learn object detectors for detecting the location of objects in input images, as shown in the bottom of Fig. 1. Different from the fully supervised object detection task that requires exhaustive patch-level/bounding-box annotations for training, object discovery is weakly supervised, i.e., only image-level annotations are necessary to train object discovery models, as shown in the top of Fig. 1.Object discovery is also called weakly supervised object detection, common object detection, etc., in other papers. Nowadays, large scale datasets with patch-level annotations are available , and many object classification and detection methods are benefited from such fine-grained annotations . However, compared with the great amounts of images with only image-level annotations (e.g., using image search queries to search on the Internet), the amount of exhaustively annotated images is still relatively small. This inspires us to explore methods that can deal with only image-level annotations.

A popular solution for weakly supervised learning is Multiple Instance Learning (MIL) . In MIL, a set of bags and bag labels are given, and each bag consists of a collection of instances, where instances labels are unknown for training. MIL has two constraints: 1) if a bag is positive, at least one instance in the bag should be positive; 2) if a bag is negative, all instances in the bag should be negative. It is natural to treat images as bags and patches as instances. In addition, patch-level feature has wide applications in computer vision community, like image classification , object detection , and object discovery . Then we can combine the patch-level feature with MIL for object classification and discovery.

Specifically, the connections between MIL methods and weakly supervised object classification and discovery are introduced as follows. As defined in , there are three paradigms for MIL: instance space based MIL methods learn an instance classifier, bag space based MIL methods learn the similarity among bags, and embedded space based MIL methods map bags to representations. For object classification, many methods aggregate extracted patch features into a vector for each image as image representation, and use the representation to train a classifier , which is similar to embedded space based MIL methods as in the top-right of Fig. 2. Meanwhile for object discovery, instance space based MIL methods are directly applied on patch features to find object of interest , as shown in the bottom of Fig. 2.

Recently, deep Convolutional Neural Networks (CNNs) have obtained great success on image classification . However, conventional CNNs without patch-level image features are unsuitable to recognize complex images and unable to obtain state-of-the-art performance on challenging datasets, e.g., PASCAL VOC datasets . There are many reasons: Unlike ImageNet, which have millions object centered images, in PASCAL VOC, (1) There is a limited number of training images; (2) The images have complex structure, and objects have large spatial transformation and scale variation; (3) The images have multiple labels.

Now the state-of-the-art object classification methods for complex datasets are based on local image patches and CNNs . And as shown in Fig. 2, it is natural to treat object classification in complex images as a MIL problem. Thus, it is important to combine deep CNNs with MIL. There are a few early attempts. For example, similar to embedded space based MIL methods, Cimpoi et al. combines CNN-based patch features with Fisher Vector to learn image representations. The Hypotheses CNN Pooling (HCP) and Deep Multiple Instance Learning (DMIL) find the most representative patches in images. These examples show that the patch-based CNN has its advantage over plain CNNs. Also, for object discovery, many methods use CNN to extract patch features, and discover objects by instance space based MIL methods . All these methods are patch-based, and they are more preferable on complex datasets than plain CNNs.

However, these methods have some limitations. First, they separately feed each patch into CNN models for feature extraction, ignoring the fact that computation on convolutional layers for overlapping patches can be shared, thus reduced, in both training and testing procedures. Second, they treat patch feature extraction, image representation learning, and object classification, or discovery as separate stages. During training, every stage requires its own training data, taking up a lot of disk space for storage. At the same time, treating these stages separately may harm performance, as the stages may not be independent. Therefore it is better to integrate them into a unified framework. Third, features of patches are extracted using pre-trained models, i.e., they can not learn dataset or task specific patch features. Last, they treat object classification and discovery as independent tasks, which have been demonstrated to be complementary by our experiments. Inspired by these facts, we propose a novel framework, called Deep Patch Learning (DPL), which integrates patch feature learning, image representation learning, object classification and discovery into a unified framework.

Inspired by the fully supervised object detection methods SPPnet and Fast R-CNN , our DPL reduces the training and testing time by sharing the computation on convolutional layers for different patches. Meanwhile, it combines different stages of object classification and discovery to form an end-to-end framework for classification and discovery. That is, DPL optimizes the patch feature learning, image representation learning, and image classifying jointly by backpropagation, which mainly focuses on object classification. In the meantime, it uses a MIL loss for each patch feature, and trains a deep MIL network end-to-end, which can discover the most representative patches in images. These two blocks (object classification block and MIL based discovery block) are combined via a multi-task learning framework, which boosts the performance for each task. Moreover, as images may have multiple labels, the MIL loss is adapted to make it suitable for the multi-class case. Notice that for both object classification and discovery, only image-level annotations are utilized for training, which makes our method quite different from the fully supervised methods that require detailed patch-level supervisions.

To demonstrate the effectiveness of our method, we perform elaborate experiments on the PASCAL VOC 2007 and 2012 datasets. The DPL achieves state-of-the-art performance on object classification, and very competitive results on object discovery. Moreover, it takes only 1.851.85s and 2.82.8s for each image during testing, using AlexNet and VGG16 CNN backend, respectively, which is much faster than the previous best performed method HCP .

To summarize, the main contributions of our work are as follows.

We propose a weakly supervised learning framework to integrate different stages of object classification into a single deep CNN framework, in order to learn patch features for object classification in an end-to-end manner. The proposed object classification network is much more effective and efficient than previous patch-based deep CNNs.

We novelly integrate the two MIL constraints into the loss of our deep CNN framework to train instance classifiers, which can be applied for object discovery.

We embed two tasks object classification and discovery into a single network, and perform classification and discovery simultaneously. We also demonstrate that the two tasks are complementary to some extent. To the best of our knowledge, it is the first time to demonstrate that classification and discovery can be complementary to each other in an end-to-end neural network. We think this reveals new phenomenon that makes sense to this community.

Our method achieves state-of-the-art performance on object classification, and very competitive results on discovery, with faster testing speed on PASCAL VOC datasets.

The rest of this paper is organized as follows. In Section 2, related work is listed. In Section 3, the detailed architecture of our DPL is described. In Section 4, we present some experiments on several object classification and discovery benchmarks. In Section 5, some discussions of experimental setups are presented. Section 6 concludes the paper.

Related Work

MIL was first proposed by Dietterich et al. for drug activity prediction. Then many methods have emerged in the MIL community . Our method can be regarded as a MIL based method as we treat images as bags and patches as instances. Meanwhile, learning image representations can be viewed as embedded space based MIL method and learning instance classifier can be viewed as instance space based MIL method. However, traditional MIL methods mainly focus on the problem that bags only have one single label, while in real-world tasks each bag may be associated with more than one class label, e.g., an image may contains multiple objects from different classes. A solution for the multi-class problem is to adapt the MIL by training a binary classifier for each class through the one-vs.-all strategy . And the Multi-Instance Multi-Label (MIML) problem also have been proposed instead of the single label MIL problem. As many images in the PASCAL VOC datasets have multiple objects from different classes, our method is also based on the MIML. Similar to the one-vs.-all strategy, we train some binary classifiers using the multi-class sigmoid cross entropy loss. But instead of training these binary classifiers separately, we train all classifiers jointly and share features among these classifiers, just like the multi-task learning . Moreover, different from previous MIL methods, we integrate the MIL constraints into the popular deep CNN, and apply our method to object classification and discovery.

There are also many other computer vision methods benefit from the MIL. Wei et al. and Wu et al. have combined the CNN and MIL for end-to-end object classification. Their methods are also be end-to-end trainable and can learn patch features. However, their methods have to resize patches to a special size and feed all patches into the CNN models separately, as shown in Figure 3 (b). This will result in huge time consumption for training and testing due to ignoring the fact that computation on convolutional layers for overlapping patches can be shared. Meanwhile, use instance space based MIL methods for solution, which means they train an instance classifier under the MIL constraints, and aggregate instance scores by max-pooling as bag scores. Then they classify bags (images) by these pooled bag scores. Different from their methods, as shown in Figure 3 (d), we share computation of convolutional layers among different patches, and combine both embedded space and instance space based MIL methods into a single network, which can achieve much promising results.

MIL is also a prevalent method for object discovery. Cinbis et al. and Wang et al. have used MIL for object discovery, and have achieved some state-of-the-art performance. But their methods separate patch feature extraction and MIL into two separate stages, which may limit their performance.

2 Patch-based Image Classification

Patch-based methods are popular for image classification as its robustness for spatial transformation, scale variation, and cluttered background. BoF is a very popular pipeline for image classification. It extracts a set of local features like SIFT or HOG from patches, and then uses some unsupervised ways , or weakly supervised methods to aggregate patch features for image representation. These image representations are used for image classification. To consider the spatial layout of images, the Spatial Pyramid Matching (SPM) is employed to enhance the performance. But their pipeline treats patch feature extraction, image representation and classification as independent stages, whereas our method integrates these into a single network and trains the network end-to-end.

Recently, Lobel et al. and Parizi et al. have proposed a method to combine the last two stages, i.e., image representation and classification. They learn patterns of patches and image classifier jointly, and the results show they have improved the performance significantly. Sydorov et al. have proposed a method to learn the parameters of Fisher Vector and image classifier end-to-end. But as a matter of fact, they do not perform real end-to-end classification. That is, although they can learn the image representation and classifier jointly, they still treat patch feature extraction as an independent part. This will lead to a large consumption of time and space for computing and saving the patch features. Different from their methods, our method achieves real end-to-end learning.

Yang et al. also proposes to learn local patch level information for object classification. They propose a multi-view MIL framework, and chooses the Fisher Vector to aggregate patch features. But their method is also not end-to-end, and requires fine-grained bounding-box annotations for training, whereas our method is end-to-end and weakly supervised.

3 Fully Supervised Object Detction

Inspired by the SPPnet and the great success of CNN for image classification , Girshick have proposed a Fast R-CNN method for fast proposal classification method in fully supervised setting. Their method can also learn patch features. Our method follows the path of this work to share computation on convolutional layers among all patches. But as shown in Figure 3 (c) and (d), the differences between our method and are multi-fold: 1) Fast R-CNN focuses on supervised object detection, whereas the proposed DPL focuses on weakly supervised image classification and object discovery. 2) Fast R-CNN requires bounding-box annotations, whereas DPL only requires image-level annotations. Annotating object bounding-boxes is labor- and time-consuming, whereas image-level annotations are easier to obtain. 3) In summary, Fast R-CNN is a fully supervised object detection framework; DPL is a weakly supervised deep learning framework for joint image classification and object discovery.

The Architecture of Deep Patch Learning

The architecture of Deep Patch Learning (DPL) is shown in Figure 4. Given an image and some patches, DPL first passes the image through some convolutional (conv) layers to generate conv feature maps for the whole image, and the size of feature maps is decided by the size of input image. After that, the Spatial Pyramid Pooling (SPP) layer can be employed for each patch to produce some fixed-size feature maps. Then each feature map can be fed into several fully connected (fc) layers, which will output a set of patch features. At last, these patch features are branched into two different streams with two different tasks: one jointly learns the image representation and classifier focusing on object classification (the classification block), and the other finds most positive patches focusing on object discovery (the discovery block). Only image-level annotations are used as supervisions to train the two streams. In this section, we will introduce these steps referred above.

Our method is patch-based, so it is necessary to generate patches first. The simplest and fastest way is sliding window, i.e., sliding a set of fixed-size windows over the image. But objects only cover a small portion of images and may have various shape, thus patches by fixed-size sliding window are always with low recall. Some methods propose to generate patches based on some visual cues, like segmentation and edge . Here we choose the “fast” mode of Selective Search (SS) to generate patches due to its fast speed and high recall.

2 Pre-trained CNN Models

Using CNN models which were trained on large scale datasets like ImageNet to fine-tune on target dataset has achieved marvelous performance. We fine-tune our model on two widely used models AlexNet and VGG16 .

3 CNN and Convolutional Feature Maps

As we stated in Section 3.2, we choose two CNN models AlexNet and VGG16. All the two models have conv layers with some max-pooling layers and three fc layers. Conv and max-pooling layers are implemented in a sliding window manner. Actually, all conv and max-pooling layers can deal with inputs of arbitrary sizes, and their outputs maintain roughly the same aspect ratio as input images. Meanwhile, conv and pooling operations will not change the relative spatial distribution of input images. Outputs from conv layers are known as conv feature maps . Though conv and max-pooling layers have the ability to handle arbitrary sized input images, the two CNN models require fix-sized input images because fc layers demand fixed-length input vectors.

4 SPP Layer

As fc layers take fixed-length input vector, the pre-trained CNN models require fixed-size input image. Therefore, the changing of image size and aspect ratio may somehow leads to loss in the performance. To handle this problem, the Fast R-CNN uses a SPP layer to realize fast proposal classification. Our work follows this path. In special, we replace the last max-pooling layer by the SPP layer. That is, given ii-th patch RiR_{i} and its coordinate is (lix,liy,rix,riy)(l^{x}_{i},l^{y}_{i},r^{x}_{i},r^{y}_{i}) that indicate the horizontal/vertical ordinates of the top left and bottom right points, suppose the feature maps size is 1/n1/n of original image size (e.g., 1/161/16 for VGG16), we can project the coordinate of RiR_{i} to (lix/n,liy/n,rix/n,riy/n)(l^{x}_{i}/n,l^{y}_{i}/n,r^{x}_{i}/n,r^{y}_{i}/n) that corresponds to the coordinate of patch ii on feature maps. Then we can obtain feature maps of patch ii by cropping the portion of whole image feature maps inside RiR_{i} and resizing it to fixed-size. The size of resized feature maps is depended on the pre-trained CNN model (e.g., 6×66\times 6 for AlexNet and 7×77\times 7 for VGG16). Taking VGG16 as an example, suppose the jj-th cropped feature map of RiR_{i} is xij\mathbf{x}_{ij}, we can divide the xij\mathbf{x}_{ij} into 7×77\times 7 uniform grids. Then the output oijko^{k}_{ij} from the kk-th grid RikR^{k}_{i} is as Eq. (1). This procedure will produce fixed-size feature maps for each patch, which can be transmitted to the following fc layers. More details can be found in .

5 Multi-task Learning Loss

As shown in Figure 4 and referred above, our DPL will produce two different scores for two different tasks respectively, one for object classification, and the other for object discovery. Therefore, we replace the last fc layer and the softmax layer of pre-trained models by our multi-task loss. Here we denote the classification loss as LclsL_{cls} and discovery loss as LdisL_{dis}, and the total loss is as follows.

where XiX_{i}, YiY_{i} are the input image and its image-level label respectively. Here we will introduce these two losses in detail.

To learn the parameters of part filters W\mathbf{W} and image classifier Ucls\mathbf{U}_{cls}, the derivative ∂Lcls/∂Ucls\partial L_{cls}/\partial\mathbf{U}_{cls} and ∂Lcls/∂W\partial L_{cls}/\partial\mathbf{W} is required to be computed. This can be easily achieved by standard backpropagation, as shown in Eq. (3) and Eq. (4).

where II is the batch size per-iteration and JiJ_{i} is the patch number of image XiX_{i}. Actually the connections between patch features and encoded patches, image representation and predicted scores are the matrix multiplication, which can be performed by fc layer and is a standard layer in CNN, so we do not give the detailed derivatives of ∂scls_ic/Ucls\partial s_{cls\_ic}/\mathbf{U}_{cls}, ∂scls_ic/∂Fir\partial s_{cls\_ic}/\partial F_{ir}, and ∂Eijn/∂W\partial E_{ijn}/\partial\mathbf{W}. The derivative of the SPM with max-pooling layer is computed by

Where the mod is the operation that computes the remainder, and mm is the mm-th grid satisfying m=ceil(r/N)m=ceil(r/N). Through the backpropagation, an end-to-end system for patch feature learning, image representation, and classification can be obtained.

5.2 Object Discovery

Different from the object classification, which aims at finding some important parts to compose the object, the object discovery is to find the patch that can locate the object exactly. That is, object classification tends to learn the local information of an object, and object discovery tends to learn the global information of an object. The two tasks are complementary in some degree, so here we also perform object discovery, as shown in the discovery block of Figure 4.

Object discovery and instance space based MIL method have similar targets. That is, the former wants to find the most representative patches of an object in the image, and the latter wants to find positive instances in the positive bag. If we treat image as bag and patches as instances, these two concepts may be equivalent. There is other work that utilizes instance space based MIL methods to realize object discovery . Our object discovery method also adopts this method to find the most positive patch of the object, just as the MI-SVM .

To learn the parameters of patch classifier Udis\mathbf{U}_{dis}, the derivative ∂Ldis/∂Udis\partial L_{dis}/\partial\mathbf{U}_{dis} is required to be computed, which can be easily achieved by the backpropagation, as shown in Eq. (6).

The connection between patch features and patch scores can also be achieved by the fc layer, so we only give the derivative of the max-pooling layer as Eq. (7).

Through the backpropagation, end-to-end object discovery can thus be performed.

5.3 Loss

where σ(x)\sigma(x) is the sigmoid function σ(x)=1/(1+exp⁡(−x))\sigma(x)=1/(1+\exp(-x)). Using the Eq. (8), we train CC binary classifiers each of which distinguishes images are with/without one object class, just similar to the one-vs.-all strategy for multi-class classification. After that, the derivative of Eq. (8) can be obtained as follows.

Then all the derivatives of parameters can be derived. We can observe that only image-level labels Yi\mathbf{Y}_{i} are necessary to optimize the loss in Eq. (8), which confirms our method is totally weakly supervised.

Experiments

In this section we will show the experiments of our DPL method for object classification and discovery.

As stated in Section 3.2, we choose two popular CNN architectures AlexNet and VGG16 , which are pre-trained on the ImageNet . These pre-trained models can be downloaded from the Caffe model zoohttps://github.com/BVLC/caffe/wiki/Model-Zoo. We replace the last max-pooling layer, the final fc layer, and the softmax loss layer by the layers defined in Section 3. The dimension of encoded patch is set to 256256 (i.e., N=256N=256). Then we choose three different SPM scales {1×1,2×2,3×1}\{1\times 1,2\times 2,3\times 1\} for the SPM with max-pooling layer after the patch encoding layer. The fc layers for patch encoding, image and patch score prediction are initialized using Gaussian distributions with -mean and standard deviations 0.010.01. Biases are initialized to be . The mini-batch size is set to 22. For AlexNet, learning rates of all layers are set to 0.0010.001 in the first 3030K mini-batch iterations and 0.00010.0001 in the later 1010K iterations. For VGG16, as it is very deep and hard to train, we first only train the layers after the second fc layer 55K iterations with learning rate 0.0010.001, and then train another 4040K iterations as for AlexNet. The momentum and weight decay are set to 0.90.9 and 0.00050.0005 respectively.

1.2 Datasets and Evaluation Measures

We test our DPL method on two famous object classification and discovery benchmarks PASCAL VOC 2007 and PASCAL VOC 2012 , which have 9,9629,962 and 22,53122,531 images respectively with 20 different object categories. The datasets are split into standard train, val and test sets. We use the trainval set (5,0115,011 images for VOC 2007 and 11,54011,540 images for VOC 2012) with only image-level labels to train our models. During the testing procedure, for object classification, we compute Average Percision (AP) and the mean of AP (mAP) as the evaluation metric to test our model on the test setFor VOC 2012, the evaluation is performed online via the PASCAL VOC evaluation server (http://host.robots.ox.ac.uk:8080/). (4,9524,952 images for VOC 2007 and 10,99110,991 images for VOC 2012). For object discovery, we report the CorLoc on the trainval set as in , which computes the percentage of the correct location of objects under the PASCAL criteria (Intersection over Union (IoU) >0.5>0.5 between the ground truths and predicted bounding boxes).

1.3 Patch Generation Protocols

There are many different methods to generate patches, like proposal based methods Selective Search (SS) and EdgeBoxes , or sliding widow, which need 1.51.5s, 0.250.25s, and less than 0.010.01s respectively (we use the “fast” mode of SS). To get patches, we choose SS to produce 11-33K patches for each image. For data augmentation, we use five image scales {480,576,688,864,1200}\{480,576,688,864,1200\} (resize the longest side to one of the scales and maintain the aspect ratio of images) with their horizontal flips to train the model. For testing, we use the same five scales without flips and compute the mean score of these scales.

1.4 Experimental Platform

Our code is written by C++ and Python, based on the Caffe and the publicly available implementation of Fast R-CNN . All of our experiments are running on a NVIDIA GTX TitanX GPU with 1212GB memory. Codes for reproducing the results are available at https://github.com/ppengtang/dpl.

2 Object Classification

We first report our results for object classification. Even though the discovery block in Figure 4 mainly focuses on object discovery, it can also produce image-level scores. So for object classification, we compute the mean score of two different tasks. The results on VOC 2007 and VOC 2012 are shown in Table 1 and Table 2.The results of our method on VOC 2012 are also available on http://host.robots.ox.ac.uk:8080/anonymous/PRKWXL.html and http://host.robots.ox.ac.uk:8080/anonymous/PWADSM.html.

From the results, we can observe that our method outperforms other CNN-based methods using single model. Specially, our method is better than other patch-based methods for quite a lot. For example, the method in extract ten different scale patch features from pre-trained VGG19 model with Fisher Vector. In , patch features are extracted from five different scales with mean-pooling. HCP combines the MIL constraints and CNN models to find the most representative patches in images. Our method even outperforms the FeV+LV-20 that utilizes bounding-box annotations during training, which shows the potential for combining CNNs with weakly supervised methods (e.g., MIL). As shown in Table 1 and Table 2, our method achieves 1.8%1.8\% and 2.0%2.0\% incresement on VOC 2007 and VOC 2012 respectively. These results show that our DPL method can achieve the state-of-the-art performance on object classification. On VOC 2012, the best result was reported in literature is the combination of HCP-VGG16 and NUS-PSL , which achieves 93.2%93.2\% mAP, but it just simply averages the predicted scores by two methods.

Some patterns from our patch encoding method are also visualized in Figure 5. We can observe that, though only image-level annotations are avaliable during training, our method can learn patterns with great semantic information. For example, “Pattern 5” corresponds to head of person; “Pattern 7” corresponds to wheel of bicycle; “Pattern 141” corresponds to screen of tvmonitor; and so on.

To train our DPL model, it takes 66 hours in AlexNet and 2828 hours in VGG16. During testing, our DPL only takes 1.851.85s and 2.82.8s per-image in AlexNet and VGG16 respectively. It is much faster comparing with the HCP (33s and 1010s per-image in AlexNet and VGG16 respectively) that has achieved the state-of-the-art performance on object classification previously.

3 Object discovery

We also perform some object discovery experiments. For object discovery, we only use the predicted patch scores from the discovery block in Figure 4 and choose the patch with maximum score. The results on VOC 2007 and VOC 2012 are shown in Table 3 and Table 4.

From the results, we can observe that our method can achieve quite competitive performance on object discovery. It outperforms other MIL-based methods like , but a little weaker than the method in . The method in finds a compact cluster for object and some clusters for the background. Except for being sensitive to the number of clusters, it is a must to tune parameters for each class tediously. As other weakly supervised methods do not compute their CorLoc on VOC 2012, we only compare our method with unsupervised object co-localization methods in Table 4. We can observe that our method outperforms the co-localization methods on VOC 2012. It is not surprise as our method benefits from image-level annotations, whereas are unsupervised (without image-level labels during training).

Figure 6 shows some success and failure discovery cases on VOC 2007. As we can observe, though failure cases do not perform that well, they can still find the representative part of the whole object (e.g., the face for person), or the box not only including the object but also containing its adjacent similar objects and can locate the object exactly.

Discussion

In this section, we will discuss the influence factors of the multi-task learning, the image scales, and the method to generate patches. Without loss generality, we only choose AlexNet to perform experiments on the PASCAL VOC 2007 dataset. If not specified, all the reported testing time in this section does not includes that of the patch generation procedure.

Multi-task learning may improve the performance for each task as different tasks can influence each other by the shared representation . Here we test the influence of different tasks. The results are shown in Table 5. As we can see, the multi-task learning can improve the classification mAP by 0.3%0.3\% and the discovery CorLoc by 3.1%3.1\%. These results demonstrate that the two tasks are complementary to some degree.

2 Multi-scale vs. Single-Scale

To evaluate the influence of image scales, we conduct a single-scale experiment that only uses one scale 600600 to compare with the five scales experiment. The results are shown in Table 6. We can observe that, multi-scale can improve the classification and discovery results evidently (+2.6+2.6 and +3.6+3.6 respectively) but with the additional testing time. Notably, using a multi-scale approach allows one to increase the accuracy for both tasks. Even though this approach increases the testing time slightly, it could be of interest for applications in which accuracy remains more important than response time during system operation.

3 The Influence of Different Patch Generation Methods

In the previous experiments, we choose SS to extract patches. Here we will compare three different methods to generate patches, including SS , EdgeBoxes , and Sliding Window (SW). For EdgeBoxes, we generate 256256 patches for each image, so it can accelerate the testing speed (we also test the performance when increase the patch number, but the results show that the performance is only improved a little but the speed slow down a lot). For SW method, we extract patches from 77 different scales widow 32×{2,3,...,8}32\times\{2,3,...,8\} with step size 3232. This operation will generate 500500 to 10001000 patches per-image. The results are shown in Table 7. From the results, we can observe that the method to extract patches affects the performance greatly, especially for object discovery. What is more, SS is the best method for both object classification and discovery. It is interesting that the SW method can achieve similar classification mAP comparing with SS with less testing time. Notice that the time to generate SW patches is negligible, so during testing, the SW method is about 2×2\times, 8×8\times, and 13×13\times faster than EdgeBoxes, SS, and HCP, respectively. For systems only focusing on object classification, the SW method is preferable as it reduces the testing time significantly with no cost of performance.

Conclusions

In this paper, a novel DPL method is proposed, which integrates the patch feature learning, image representation learning, object classification and discovery into a unified framework. The DPL explicitly optimizes patch-level image representation, which is totally different from conventional CNNs. It also combines the CNN based patch-level feature learning with MIL methods, thus can be trained in a weakly supervised manner. The excellent performance of DPL on object classification and discovery confirms its effectiveness. These inspiring results show that learning good patch-level image presentation and combining CNNs with MIL are very promising directions to explore in various vision problems. In the future, we will study how to apply DPL for other visual recognition problems, including introducing DPL into solve very large scale problems.

Acknowledgements

This work was primarily supported by National Natural Science Foundation of China (NSFC) (No. 61503145, No. 61572207, and No. 61573160) and the CAST Young Talent Supporting Program.

References

References