Generative Prompt Model for Weakly Supervised Object Localization

Yuzhong Zhao, Qixiang Ye, Weijia Wu, Chunhua Shen, Fang Wan

Introduction

Weakly supervised object localization (WSOL) is a challenging task when provided with image category supervision but required to learn object localization models. As a pioneered WSOL method, Class Activation Map (CAM) defines global average pooling (GAP) to generate semantic-aware localization maps based on a discriminatively trained activation model. Such a fundamental method, however, suffers from partial object activation while often missing full object extent, Fig. 1(upper). The nature behind the phenomenon is that discriminative models are born to pursue compact yet discriminative features while ignoring representative ones .

Many efforts have been proposed to alleviate the partial activation issue by introducing spatial regularization terms , auxiliary localization modules , or adversarial erasing strategies . Nevertheless, the fundamental challenge about how to use a discriminatively trained classification model to generate precise object locations remains.

In this study, we propose a generative prompt model (GenPromp), Fig. 1(lower), which formulates WSOL as a conditional image denoising procedure, solving the fundamental partial object activation problem in a new and systematic way. During training, for each category (e.g.goldfish in Fig. 1) in the predetermined category set, GenPromp converts each category label to a learnable prompt embedding (frf_{r}) through a pre-trained vision-language model (CLIP) . The learnable prompt embedding is then fed to a transformer encoder-decoder to conditionally recover the noisy input image. Through multi-level denoising, the representative features of input images are back-propagated from the transformer decoder to the learnable prompt embedding, which is updated to the representative embedding frf_{r}.

During inference, GenPromp linearly combines learned representative embeddings (frf_{r}) with discriminative embeddings (fdf_{d}) to obtain both object generative and discrimination capability, Fig. 2. fdf_{d} is queried from a pre-trained vision-language model (CLIP), which incorporates the correspondence between text (e.g.category labels) with vision feature embeddings. The combined embedding (fcf_{c}) is used to generate attention maps at multiple levels and timestamps, which are aggregated to object activation maps through a post-processing strategy. On CUB-200-2011 and ImageNet-1K , GenPromp respectively outperforms the best discriminative models by 5.2% (87.0% vs.vs. 81.8% ) and 5.6% (65.2% vs.vs. 59.6% ) in Top-1 Loc.

We propose a generative prompt model (GenPromp), providing a systematic way to solve the inconsistency between the discriminative models with the generative localization targets by formulating a conditional image denoising procedure.

We propose to query discriminative embeddings from an off-the-shelf vision-language model using image labels as input. Combining the discriminative embeddings with the generative prompt model facilitates localizing objects while depressing backgrounds.

GenPromp significantly outperforms its discriminative counterparts on commonly used benchmarks, setting a solid baseline for WSOL with generative models.

Related Work

Weakly Supervised Object Localization. As a simple-yet-effective method, CAM localizes objects by introducing global average pooling (GAP) to an image classification network. CAM is also extended from WSOL to weakly supervised detection and segmentation . However, CAM suffers from partial activation, i.e.i.e., activating the most discriminative parts instead of full object extent. The reason lies in the inconsistency between the discriminative models (i.e.i.e., classification model) with the generative target (object localization).

To solve the partial activation problem, adversarial erasing, discrepancy learning, online localization refinement, classifier-localizer decoupling, and attention regularization methods are proposed. As a spatial regularization method, adversarial erasing online removes significantly activated regions within feature maps to drive learning the missed object parts. With a similar idea, spatial discrepancy learning leverages adversarial classifiers to enlarge object areas. Through classifier-localizer decoupling, PSOL partitions the WSOL pipeline into two parts: class-agnostic object localization and object classification. For class-agnostic localization, it uses class-agnostic methods to generate noisy pseudo annotations and then perform bounding box regression on them without class labels. While online refinement of low-level features improves activation maps for WSOL , BAS specifies a background activation suppression strategy to assist the learning of WSOL models. C2AM generates class-agnostic activation maps using contrastive learning without category label supervision. FAM optimizes the object localizer and classifiers jointly via object-aware and part-aware attention modules. TS-CAM and LCTR utilize the attention maps defined on long-range feature dependency of transformers to localize objects.

Despite the progress, existing methods typically ignore the fundamental challenge of WSOL, i.e.i.e. discriminative models are required to perform representative (localization) tasks.

Vision-Language Models. Vision-language models have demonstrated increasing importance for vision tasks. In the early years, great efforts are paid to label image-text pairs which are important for vision-language model training . In recent years, the born-ed association relations of image and text on the Web facilitate collecting a massive quantity of image-text pairs , which requires much lower annotation cost compared to those manually annotated datasets (See supplementary for details). Such image-text pairs enable building the association between image category labels and visual feature embedding, which is the foundation of this study. Based on the massive quantity of image-text pairs, we pre-train two components of GenPromp (i.e.Stable diffusion , CLIP ) with two large image-text pair datasets. In specific, we pre-train the image denoising model using LAION-5B and the CLIP model on WIT .

Generative Vision Models. As the foundation of most generative vision models, GAN defines an adversarial training process, where it simultaneously trains a generative model and a discriminative model. Many efforts have been made to improve GAN such as better optimization strategies , conditional image generation and improved architecture . Recently, a generative vision paradigm, i.e.i.e., denoising diffusion probabilistic models (DDPM) become popular and have the potential to surpass GAN in several vision generation tasks such as image generation and image editing . DDPM have also been adapted to some perception tasks such as object detection and image segmentation , which inspires this study for WSOL.

Viewing categories labels as prompt embeddings is an important feature of GenPromp, which originates from conditional image generation tasks. This follows DDPM which feeds a language description encoded by the language model (e.g.BERT ) to the generative model for sematic-aware image generation. Existing works have proposed to learn the representative embeddings from user-provided images that contains a new object or a new style, which drive DDPM to generate synthesis images that contains that object or style. Inspired by this learning paradigm, GenPromp is proposed to learn the representative embeddings of each category, which are crucial for object localization.

Preliminaries

Diffusion models learn a data distribution p(x)p(x) by gradually denoising a normally distributed variable, which corresponds to learning a reverse process of a Markov Chain . Stable diffusion leverages well-trained perceptual compression model E\mathcal{E} to transfer the denoising diffusion process from the high-dimensional pixel space to a low-dimensional latent space, which reduces the computational burden and increases efficiency. The objective function of Stable diffusion is defined as

where ϵ∼N(0,1)\epsilon\sim N(0,1) is sampled from the normal distribution and ϵθ(∘,t,f)\epsilon_{\theta}(\circ,t,f) is the neural backbone, which is implemented as an attention-unet conditioned on time tt and the prompt embedding ff. E\mathcal{E} is implemented as a VQGAN encoder. ztz_{t} is the noisy latent of the input image xx at time t,t∈T={1,2,⋯ ,1000}t,t\in T=\{1,2,\cdots,1000\}. {α‾t}t∈T\{\overline{\alpha}_{t}\}_{t\in T} denotes a set of hyperparameters that steer the levels of noise added.

Generative Prompt Model

GenPromp consists of a training stage (Fig. 3) and a finetuning stageThe architecture is similar as Fig. 3 but using a different training set. Please refer to the supplementary material for details.. In the training stage, the discriminative embedding fdf_{d} is queried from a pre-trained CLIP model (Querying Discriminative Embeddings), meanwhile a representative embedding frf_{r} is specified for each category (Learning Representative Embedding.). In the finetuning stage, fdf_{d} and frf_{r} are used to prompt the backbone network finetuning (Model Finetuning). The trained models and combined prompt embeddings (Combining Embeddings) are used to predict attention maps for WSOL (Fig. 4).

Querying Discriminative Embeddings. The embedding fdf_{d} is queried from a pre-trained CLIP model , which uses a prompt string (pp) as input and outputs a language feature vector. As CLIP is pre-trained with a discriminative loss (i.e.i.e., contrastive loss), fdf_{d} is therefore discriminative.

During training, for a given category (e.g.goldfish in Fig. 3), the prompt is obtained by filling the category label string to a template (e.g.“a photo of a goldfish”). During inference, however, the image category labels are unavailable. An intuitive way is using a pre-trained classifier to predict the category of the input image. Unfortunately, some category labels correspond to strings of multiple words, which imply multiple tokens and multiple attention maps. Such multiple attention maps with noise would deteriorate the localization performance. To address this issue, we heuristically select the last token of the first string to initialize the prompt, which is referred to as the meta token. For example, we select goldfish for category string “goldfish, Carassius auratus” and select ray for category string “electric ray, crampfish, numbfish, torpedo”. The meta token of a category is typically the name of its superclass.

To query discriminative embedding fdf_{d} from the CLIP model, the prompt pp is converted to an ordered list of numbers (e.g.[[a→\rightarrow302, photo→\rightarrow1125, of→\rightarrow539, a→\rightarrow302, goldfish→\rightarrow806])])⟨BOS⟩\langle\texttt{BOS}\rangle,⟨EOS⟩\langle\texttt{EOS}\rangle and padding tokens are omitted for clarity.. This procedure is performed by looking up the dictionary of the Tokenizer. The list of numbers is used to index the embedding vectors v(p)=[v302,v1125,v539,v302,v806]v(p)=[v_{302},v_{1125},v_{539},v_{302},v_{806}] through an index-based Embedding layer. Such embedding vectors (language vectors) are input to a pre-trained vision-language model (CLIP) to generate the discriminative language embedding fdf_{d}, as

Learning Representative Embedding. Using solely the discriminative embedding fdf_{d}, GenPromp could miss representative yet less discriminative features. As an example in the second column and second row of Fig. 2, the activation map of the tower obtained by prompting the network with fdf_{d} suffers from partial object activation. The representative embedding frf_{r} is thereby introduced as a prompt learning procedure, as shown in Fig. 3. In specific, the dictionary of the Tokenizer is extended by involving a new token (i.e.i.e., ⟨goldfish⟩\langle{\rm\texttt{goldfish}}\rangle), which is referred to as the concept token. The prompt with the concept token is encoded to the representative embedding frf_{r} through the pre-trained CLIP model.

Before learning, the representative embedding frf_{r} is initialized as the discriminative embedding fdf_{d}. When learning representative embedding, images belonging to a same category are collected to form the training set. Each input image is encoded to the latent variable z0z_{0} by the VQGAN encoder, Fig. 3. Different levels of noises are added to z0z_{0} to have a noisy latent ztz_{t} through Eq. 2. The noisy latent ztz_{t} is fed to the attention-unet ϵθ\epsilon_{\theta} to produce multi-layer feature maps Fl,tF_{l,t}, where l∈Ll\in L index the encoded features of different layers. The Tokenizer Embedding layer in Fig. 3 incorporate embeddings for each word/token, where the embeddings (i.e.i.e., vrv_{r}) of the corresponding concept tokens (e.g.e.g., ⟨goldfish⟩\langle{\rm\texttt{goldfish}}\rangle) are trainable. By optimizing a denoising procedure defined by stable diffusion, the representative embedding frf_{r} is learned, as

Benefiting from the property of the generative denoising model, fr∗f_{r}^{*} can identify the common features among objects in the training set, learning representative features that define each category. As shown in the second column of Fig. 2, the activation map of tower obtained by prompting the network with frf_{r} activates the full object area.

Model Finetuning. After obtaining the representative embeddings fr∗f_{r}^{*} for all image categories, the backbone network (attention-unet parameterized by ϵθ\epsilon_{\theta}) is finetuned on the WSOL dataset, as

which further optimizes the diffusion model using both fdf_{d} and fr∗f_{r}^{*} as prompts to reduce the domain gap between the model and the target dataset.

Combining Embeddings. After model finetuning, the discriminative and representative embeddings (fdf_{d} and frf_{r}) are linearly combined, as

where w∈w\in is an experimentally determined weighted factor. As shown in Fig 2, for the category towel, a large ww activates background noise. Meanwhile, a small ww could cause partial object activation. An appropriate ww can balance the discriminative features and representative features of the categories, achieving the best localization performance.

Weakly Supervised Object Localization

As shown in Fig. 4, WSOL is defined as a conditional image denoising procedure. Similar to the training procedure, each input image is encoded to the latent variable z0z_{0} by the VQGAN encoder. Different levels of noise are added to z0z_{0} to generate noisy latent ztz_{t} through Eq. 2. The noisy latent ztz_{t} is fed to the finetuned attention-unet ϵθ∗\epsilon_{\theta^{*}} to produce multi-layer feature maps Fl,tF_{l,t}, where l∈Ll\in L index the encoded features of different layers. Meanwhile, the input image is classified with a classifier pre-trained on the target dataset to obtain the category label, which is used to initialize two prompts (e.g.goldfish and ⟨goldfish⟩\langle\texttt{goldfish}\rangle). According to Eq. 3, the two prompts are encoded to discriminative embedding fdf_{d} and representative embedding frf_{r}. The two embeddings are then combined to fcf_{c}, which is the condition for the image denoising model. Through performing denoising, cross attention maps are generated by using Fl,tF_{l,t} as the Query vector and fcf_{c} as the Key vector, as

One notable capability of GenPromp is to generate attention maps {ml,t}l∈L,t∈T\{m_{l,t}\}_{l\in L,t\in T} which exhibit distinct characteristics based on the specific layer ll at time tt, as shown in Fig. 5. The characteristics of these attention maps can be concluded as follows: (1) Attention maps with higher resolution can provide more detailed localization clues but introduce more noise. (2) Attention maps of different layers can focus on different parts of the target object. (3) Smaller tt provides a less noisy background but tends to partial object activation. (4) Larger tt activates the target object more completely but introduces more background noise. Based on these observations, we propose to aggregate attention maps at multiple layers and timesteps to obtain a unified activation map, as

Experimentally, we find that aggregating the attention maps of spatial resolutions 8×88\times 8 and 16×1616\times 16 at time steps 11 and 100100 produces the best localization performance. A thresholding approach is then applied to predict the object locations based on the unified activation map.

Experiments

Datasets. We evaluate GenPromp on two commonly used benchmarks, i.e.i.e., CUB-200-2011 and ImageNet-1K. CUB-200-2011 is a fine-grained bird dataset that contains 200 categories of birds with 5994 training images and 5794 test images. ImageNet-1K is a large-scale visual recognition dataset containing 1,000 categories with 1.2 million training images and 50,000 validation images.

Evaluation Metrics. We follow the previous methods and use Top-1 localization accuracy (Top-1 Loc), Top-5 localization accuracy (Top-5 Loc), and GT-known localization accuracy (GT-known Loc) as the metrics. For localization, a bounding box prediction is positive when it satisfies: (1) the predicted category label is correct, and (2) the IoU between the bounding box prediction and one of the ground-truth boxes is greater than 50%\%. GT-known indicates that it considers only the IoU constraint regardless of the classification result.

Implementation Details. GenPromp is implemented based the Stable Diffusion model , which is pre-trained on LAION-5B . The text encoder, i.e.i.e., CLIP, is pre-trained on WIT . During training, we resize the input image to 512×\times512 and augment the training data with RandomHorizontalFlip and ColorJitter. We then optimize the network using AdamW with ϵ\epsilon=1ee−-8, β1\beta_{1}=0.9, β2\beta_{2}=0.999 and weight decay of 1ee−-2 on 8 RTX3090. In the training stage, we optimize GenPromp for 2 epochs with learning rate 5ee−-5 and batch size 8 for each category in CUB-200-2011 and ImageNet-1K. In the finetuning stage, we train GenPromp for 100,000 iterations with learning rate 5ee−-8 and batch size 128.

2 Main Results

Performance Comparison with SOTA Methods. In Table 1, the performance of the proposed GenPromp is compared with the state-of-the-art (SOTA) models. On CUB-200-2011 dataset, GenPromp achieves localization accuracy of Top-1 87.0%, Top-5 96.1%, which surpasses the SOTA methods by significant margins. Specifically, GenPromp achieves surprisingly 98.0% localization accuracy under GT-known metric, which shows the effectiveness of introducing generative model for WSOL. GenPromp outperforms the SOTA method C2AM\text{C}^{2}\text{AM} by 5.2% (87.0% vs.vs. 81.8%) and 5.0% (96.1% vs.vs. 91.1%) under Top-1 Loc and Top-5 Loc metrics respectively. When solely considering the localization performance (using GT-known Loc metric), GenPromp significantly outperforms the SOTA method SCM by 1.4% (98.0% vs.vs. 96.6%). On the more challenging ImageNet-1K dataset, GenPromp also significantly outperforms the SOTA method C2AM\text{C}^{2}\text{AM} and BAS by 5.6% (65.2% vs.vs. 59.6%), 6.0% (73.4% vs.vs. 67.4%), and 3.2% (75.0% vs.vs. 71.8%) under Top-1 Loc, Top-5 Loc and GT-known Loc metrics respectively. Such strong results clearly demonstrate the superiority of the generative model over conventional discriminative models for weakly supervised object localization.

Localization Results with respect to Prompt Embeddings. The localization results of GenPromp are shown in Fig. 6. In the first row of Fig. 6(upper), for the image with category spaniel, the discriminative prompt embedding fdf_{d} fails to activate the legs, while the representative prompt embedding frf_{r} activates full object extent but suffering from the background noise. By combining fdf_{d} and frf_{r}, fcf_{c} fully activates the object regions while maintaining low background noise. In the second row of Fig. 6(upper), we show the localization maps with prompt embeddings generated by different categories. Interestingly, categories that are highly related to spaniel (i.e.,i.e., dog, husky) can also correctly activate the foreground object, which indicates that GenPromp is robust to the classifiers. When using the categories that are less related to spaniel (i.e.i.e., man, tower) will introduce too much background noise and fail to localize the objects.

As shown in Fig. 6(lower), for the test images which contain multiple objects from various classes, GenPromp is able to generate high quality localization maps when given the corresponding prompt embeddings (generated using corresponding categories). The result demonstrates GenPromp can not only generate representative localization maps but also be able to discriminate object categories, revealing the potential of extending GenPromp to more challenging weakly supervised object detection or segmentation task.

Statistical Result with respect to Token Embeddings. In Fig. 7, the embedding vectors of the meta token are uniformly distributed in the two-dimensional feature space (yellow dots). After learning the representative embeddings/features of the categories, the embedding vectors of the concept token become uneven (blue dots), i.e.i.e. less discriminative indicated by less deviation σx\sigma_{x} and σy\sigma_{y}, which reveals the inconsistency between representative and discriminative embeddings.

3 Ablation Study

Baseline. We build the baseline method by using solely the discriminative embedding fdf_{d} (Line 1 of Table 2). The performance (61.2% Top-1 Loc Acc.) of the baseline outperforms the SOTA methods, which indicates the great advantage of introducing the generative model for WSOL.

Representative Embedding. When using the trained representative embedding frf_{r} initialized by fdf_{d} (Line 3 of Table 2), GenPromp outperforms the baseline by 2.8% (64.0% vs.vs. 61.2%) under Top-1 Loc metric, which validates the importance of the representative features of categories for object localization. We also train representative embedding frf_{r} which is random initialized (Line 2 of Table 2). The performance of frf_{r} drops to very low-level suffering from the local optimal embeddings.

Embedding Combination. By combining frf_{r} and fdf_{d}, the performance of GenPromp can be further boost to 64.5% in Top-1 Loc (Line 5 of Table 2). We also conduct experiments to show the effect of the combination weight ww on ImageNet 1K dataset. The performance keeps consistent when the combination weight ww falls to [0.4,0.8][0.4,0.8], as shown in Fig. 8. With a proper ww, GenPromp is able to generate discriminative while representative prompt embeddings and active full object extent with the least background noise.

Model Finetuning. By filling the domain gap through finetuning the backbone network (attention-unet), a performance gain of 0.6%0.6\% (65.1% vs.vs. 64.5%) can be achieved in Top-1 Loc (Line 8 of Table 2).

Model Size and Training Data. In Table 3, we re-implement TS-CAM with larger backbone (e.g.ViT-H) and more training data (e.g.LAION-2B). The re-implemented TS-CAM achieves higher classification accuracy while much lower localization accuracy compared to the Deit-S based TS-CAM. We attribute this to the inherent flaw of the discriminatively trained classification model, i.e.i.e., local discriminative regions are capable of minimizing image classification loss, but experience difficulty in accurate object localization. A larger backbone and more training data even make this phenomenon even more serious.

Effect of Timesteps and Resolutions. In Tables 4 and 5, we evaluate the performance by aggregating attention maps under different timesteps and spatial resolutions. It can be seen that aggregating the attention maps of spatial resolutions 8×88\times 8 and 16×1616\times 16 at time steps 11, 100100 produces the best localization performance.

Conclusion and Future Remark

To solve the partial object activation brought by discriminatively trained activation models in a systematic way, we propose GenPromp, which formulates WSOL as a conditional image denoising procedure. During training, we use the learnable embeddings to conditionally recover the input image with the aim to learn the representative embeddings. During inference, GenPromp first combines the trained representative embeddings with another discriminative embeddings from the pre-trained CLIP model. The combined embeddings are then used to generate attention maps at multiple timesteps and resolutions, which facilitate localizing the full object extent. GenPromp not only sets a solid baseline for WSOL with generative models but also provides fresh insight for handling vision tasks using vision-language models.

When claiming the advantages of GenPromp, we also realize its disadvantages. One major disadvantage is its dependency on large-scale pre-trained vision-language models, which could slow down the inference speed and raise the requirement for GPU memory cost. Such a disadvantage should be solved in future work.

References

Appendix A Annotation Cost

As shown in Table 6, we compare the size and data collection methods of commonly used datasets with three types of annotations: image category labels, text descriptions, and bounding boxes. It can be observed that datasets with accurate bounding box labels, such as Cityscapes and COCO, are usually small in size due to the high cost of manual annotation. However, when using image category labels, the dataset size can be increased to 14 million (e.g., ImageNet). For datasets with a huge size, such as JFT-3B, WIT, and LAION-5B, manual annotation becomes impractical. Instead, semi-automatic annotation methods or web crawler algorithms are used to extensively collect noisy annotated data. Thanks to the rapid development of the Internet, a large number of image-text pairs can be found in websites, forums, and libraries, which are naturally annotated by citizens and can be easily obtained by crawler algorithms. Since collecting image-text pairs hardly requires human participation, their annotation cost is negligible. In this paper, the proposed method GenPromp is implemented based the Stable Diffusion model, which is pre-trained on LAION-5B. Accordingly, GenPromp hardly introduces additional annotation cost for a weakly supervised learning system.

Appendix B Finetuning

As described in Section 4 in the main document, after obtaining the representative embeddings fr∗f_{r}^{*} for all image categories, we finetune the attention-unet (parameterized by ϵθ\epsilon_{\theta}) to reduce the domain gap between the model and the target dataset. We demonstrate the pipeline of finetuning in Fig. 9. The finetuning pipeline is different from the training pipeline (shown in Fig. 3 of the main document) in two aspects: (1) Each training batch contains images of multiple categories, (2) The prompt embeddings are frozen while the attention-unet is trainable.

Appendix C Prompt Ensemble

To further improve the performance of GenPromp, we propose a prompt ensemble strategy. As shown in Fig. 10, during training, we random select a template from a template set. Then, we respectively fill the meta token (goldfish) and the concept token (⟨goldfish⟩\langle\texttt{goldfish}\rangle) into the template to obtain the two input prompts, which are used to learn the representative embedding frf_{r}. During inference, for each template in the template set, we combines it with the two tokens to form the input prompts. Then, all the prompts are encoded into prompt embeddings by the pre-trained CLIP model. After that, the discriminative embedding fdf_{d} is obtained by averaging the discriminative embeddings generated by different templates (e.g.fd1,fd2,fd3,fd4f_{d_{1}},f_{d_{2}},f_{d_{3}},f_{d_{4}} in Fig. 10), the representative embedding frf_{r} is obtained by averaging the representative embeddings generated by different templates (e.g.fr1,fr2,fr3,fr4f_{r_{1}},f_{r_{2}},f_{r_{3}},f_{r_{4}} in Fig. 10). Finally, fdf_{d} and frf_{r} are combined to fcf_{c}, which is fed into the network to generate attention maps. In experiments, we use a template set that consists of 7 templates:

Appendix D Additional Experimental Results

Complete Performance Comparison with SOTA Methods. Table 9 shows the complete performance comparison of the proposed GenPromp and the state-of-the-art (SOTA) models (extension of Table 1 in the main document). On CUB-200-2011 and ImageNet-1K dataset, GenPromp surpasses the SOTA methods by significant margins. Such strong results clearly demonstrate the superiority of the generative model over conventional discriminative models for weakly supervised object localization.

Localization Error Analysis. To further reveal the effect of the proposed prompt embeddings (e.g.fdf_{d}, frf_{r}, fcf_{c}), following TS-CAM , we evaluate the localization errors of: multi-instance error (M-Ins), localization part error (Part), and localization more error (More). They are respectively defined as follows.

M-Ins indicates that the predicted bounding box intersects with at least two ground-truth boxes, and IoG>0.3\text{IoG}>0.3.

Part indicates that the predicted bounding box only cover the parts of object, and IoP>0.5\text{IoP}>0.5.

More indicates that the predicted bounding box is larger than the ground truth bounding box by a large margin, and IoG>0.7\text{IoG}>0.7.

where IoG and IoP are defined as Intersection over Ground truth box and Intersection over Predict bounding box, respectively (similar to IoU (Intersection over Union)). Each metric calculates the percentage of images belonging to corresponding error in the validation/test set. Please refer to TS-CAM for a detailed definition of the three metrics. Table 7 lists localization error statistics of M-Ins, Part, and More. Compare to the discriminative embedding fdf_{d}, the learned representative embedding frf_{r} reduces both Part and More errors by 1.4% (3.8% vs.vs. 2.4%) and 0.9% (8.2% vs.vs. 7.3%) respectively, demonstrating that the representative embedding alleviates the partial object activation problem. By combining the representative embedding frf_{r} with the discriminative embedding fdf_{d}, the More errors drop 0.4% (7.3% vs.vs. 6.9%) while the Part errors increase 0.6% (3.0% vs.vs. 2.4%) compared to frf_{r}. This demonstrates that fcf_{c} can further depress the background noise while keeping relatively low Part errors.

Effect of Noise ϵ\epsilon. In Table 8, we evaluate the performance by setting the noise ϵ\epsilon (in Eq. 1 and Eq. 2 of the main document) to 0 during inference. Without noise ϵ\epsilon, the performance of GenPromp drops 0.3% in Top-1 Loc in average. Similar to the methods based on adversarial erasing, the input noise in GenPromp can also alleviate the part activation issue, which drives the network to mine the representative yet less discriminative object parts.

Additional Restults with respect to Model Size and Training Data. In Table 10, we re-implement TS-CAM with larger backbone (e.g.Deit-B, ViT-L, ViT-H) and more training data (e.g.LAION-2B). As the model size getting larger, the performance of TS-CAM becomes worse on CUB-200-2011 under GT-known Loc metric, Table 10(upper). As shown in Table 10(lower), by finetuning ViT-H-based TS-CAM for 3 epochs on ImageNet-1K, it achieves higher classification accuracy (74.7% vs.vs. 74.3% on Top-1 Cls) while much lower localization accuracy (53.2% vs.vs. 67.6% on GT-known Loc) compared to the Deit-S-based TS-CAM. By finetuning the model for more epochs (e.g.6 epochs), it achievea higher classification accuracy (77.4% vs.vs. 74.7% on Top-1 Cls) but lower localization accuracy (52.2% vs.vs. 53.2% under GT-known Loc metric), demonstrating that more epochs can not improve the localization performance of TS-CAM. We attribute this phenomenon to the inherent flaw of the discriminatively trained classification model, i.e.i.e., local discriminative regions are capable of minimizing image classification loss but experience difficulty in accurate object localization. A larger backbone and more training data make this phenomenon even more serious.

Detailed Ablation Study. Table 11 provides a detailed ablation of the performance contribution of each component and their combinations, with respect to Multi-resolution, Multi-timesteps, Prompt ensemble, Prompt embedding and Finetuning.

Appendix E Additional Visualization Results

In Fig. 11, we visualize the localization results of GenPromp and compared them with the discriminatively trained model (e.g.CAM ). The object Localization maps of CAM (column b) suffer from partial object activation. Localization maps of GenPromp (column d) with sole representative embeddings (frf_{r}) covers more object extent but introducing background noise. Those of GenPromp (column e) with combined embeddings (fcf_{c}) not only activate full object extent but also depress background noise for precise object localization.

We also provide additional visualization results of Fig. 2, Fig. 5 and Fig. 6 in the main document. The results are shown in Fig. 12, Fig. 13 and Fig. 14 respectively.