DiffuMask: Synthesizing Images with Pixel-level Annotations for Semantic Segmentation Using Diffusion Models

Weijia Wu, Yuzhong Zhao, Mike Zheng Shou, Hong Zhou, Chunhua Shen

Introduction

Semantic segmentation is a fundamental task in vision, and existing data-hungry semantic segmentation models usually require a large amount of data with pixel-level annotations to achieve significant progress. Unfortunately, pixel-wise mask annotation is a labor-intensive and expensive process. For example, labeling a single semantic urban image in Cityscapes can take up to 60 minutes, underscoring the level of difficulty involved in this task Additionally, in some cases, it may be challenging or even impossible to collect images due to existing privacy and copyright. To reduce the cost of annotation, weakly-supervised learning has become a popular approach in recent years. This approach involves training strong segmentation models using weak or cheap labels, such as image-level labels , points , scribbles , and bounding boxes . Although these methods are free of pixel-level annotations, still suffer from several disadvantages, including low-performance accuracy, complex training strategy, indispensable extra annotation cost (e.g., edge), and image collection cost.

With the great development of computer graphics (e.g., generative model), an alternative way is to utilize synthetic data, which is largely available from the virtual world, and the pixel-level ground truth can be freely and automatically generated. DatasetGAN firstly exploits the feature space of a trained GAN and trains a shallow decoder to produce pixel-level labeling. BigDatasetGAN extends DatasetGAN to handle the large class diversity of ImageNet. However, both methods suffer from certain drawbacks, the need for a small number of pixel-level labeled examples to generalize to the rest of the latent space and suboptimal performance due to imprecise generative masks.

Recently, large-scale language-image generation (LLIG) models, such as DALL-E , and Stable Diffusion , have shown phenomenal generative semantic and compositional power, as shown in Fig. 1. Given one language description, the text-conditioned image generation model can create corresponding semantic things and stuff, where visual and textual embedding are fused using spatial cross-attention. We dive deep into the cross-attention layers and explore how they affect the generative semantic object and structure of the image. We find that cross-attention maps are the core, which binds visual pixels and text tokens of the prompt text. Also, the cross-attention maps contain rich class (text token) discriminative spatial localization information, which critically affects the generated image.

Can the attention map be used as mask annotation? Consider semantic segmentation —a ‘good’ pixel-level semantic mask annotation should satisfy two conditions: (a) class-discriminative (i.e., localize and distinguish the categories in the image); (b) high-resolution, precise mask (i.e., capture fine-grained detail). Fig. 2 presents a visualization of cross attention map between text token and vision. 8×88\times 8, 16×1616\times 16, 32×3232\times 32, and 64×6464\times 64, as four different resolutions, are extracted from different layers of the U-Net of Stable Diffusion . 8×88\times 8 feature map is the lowest resolution, including obvious class-discriminative location. 32×3232\times 32 and 64×6464\times 64 feature maps include high-resolution and highlight fine-grained details. The average map shows the possibility for us to use for semantic segmentation, where it is class-discriminative and fine-grained. To further validate the potential of the attention map of the generative task, we convert the probability map to a binary map with fixed thresholds γ\gamma, and refine them with Dense CRF , as shown in Fig. 2. With the 0.350.35 threshold, the mask presents a wonderful precision on fine-grained details (e.g., foot, ear of the ‘horse’).

Based on the above observation, we present DiffuMask, an automatic procedure to generate a massive high-quality image with a pixel-level semantic mask. Unlike DatasetGAN and BigDatasetGAN , DiffuMask does not require any pixel-level annotations. This approach takes full advantage of powerful zero-shot text-to-image generative models such as Stable Diffusion , which are trained on web-scale image-text pairs. DiffuMask mainly includes two advantages for two challenges: 1) Precise Mask. An adaptive threshold of binarization is proposed to convert the probability map (attention map) to a binary map, as the mask annotation. Besides, noise learning is used to filter noisy labels. 2) Domain Gap: retrieval-based prompt (various and verisimilar prompt guidance) and data augmentations (e.g., Splicing ), as two effective solutions, are designed to reduce the domain gap via enhancing the diversity of data. With the above advantages, DiffuMask can generate infinite images with pixel-level annotation for any class without human effort. These synthetic data can then be used for training any semantic segmentation architecture (e.g., mask2former ), replacing real data.

To summarize, our contributions are three-folds:

We show a novel insight that it is possible to automatically obtain the synthetic image and mask annotation from a text-supervised pre-trained diffusion model.

We present DiffuMask, an automatic procedure to generate massive image and pixel-level semantic annotation without human effort and any manual mask annotation, which exploits the potential of the cross-attention map between text and image.

Experiments demonstrate that segmentation methods trained on DiffuMask perform competitively on real data, e.g., VOC 2012. For some classes, e.g., dog, the performance is close to that of training with real data (within 3% gap). Moreover, in the open-vocabulary segmentation (zero-shot) setting, DiffuMask achieves new SOTA results on the Unseen classes of VOC 2012.

Related Work

Reducing Annotation Cost. Various ways can be explored to reduce the segmentation data cost, including interactive human-in-the-loop annotation , nearest-neighbor mask transfer , or weak/cheap mask annotation supervision in different levels, such as image-level labels , points , scribbles , and bounding boxes . Among the above-related works, image-level label supervised learning presents the lowest cost, and its performance is unacceptable. Bounding boxes annotation usually shows a competitive performance than pixel-wise supervised methods, but its annotation cost is the most expensive. By comparison, synthetic data presents many advantages, including lower data cost without image collection, and infinite availability for enhancing the diversity of data.

Image Generation. Image generation is a basic and challenging task in computer vision. There are several mainstream methods for the task, including Generative Adversarial Networks (GAN) , Variational autoencoders (VAE) , flow-based models , and Diffusion Probabilistic Models (DM) . Recently, the diffusion model has drawn lots of attention due to its wonderful performance. GLIDE used pre-trained language model (CLIP ) and the cascaded diffusion structure for text-to-image generation. Similarly, DALL-E 2 of OpenAI Imagen obtain the corresponding text embedding with CLIP and adopted a similar hieratical structure to generate images. To increase accessibility and reduce significant resource consumption, Stable Diffusion of Stability AI introduced a novel direction in which the model diffuses on VAE latent spaces instead of pixel spaces.

Synthetic Dataset Generation. Prior works for dataset synthesis mainly utilize 3D scene graphs to render images and their labels. 2D methods, i.e., Generative Adversarial Networks (GAN) mainly is used to solve domain adaptation task , which leverages image-to-image translation to reduce the domain gap. Recently, inspired by the success of generative model (e.g., DALL-E 2, Stable Diffusion), some works further try to explore the potential of synthetic data to replace real data as the training data in many downstream tasks, including image classification , object detection , image segmentation , 3D Rendering . DatasetGAN utilized a few labeled real images to train a segmentation mask decoder, leading to an infinite synthetic image and mask generator. Based on DatasetGAN, BigDatasetGAN scale the class diversity to ImageNet size, which generates 1k classes with manually annotated 5 images per class. With Stable diffusion and Mask R-CNN pre-trained on COCO dataset, Li et al. design and train a grounding module to generate images and segmentation masks. Different from the above methods, we go one step further and synthesize accurate semantic labels by exploiting the potential of cross attention map between text and image. One significant advantage of the DiffuMask is that it does not require any manual localization annotations (i.e., box and mask) and only rely on text supervision.

Methodology

In this paper, we explore simultaneously generating images and the semantic mask described in the text prompt with the existing pre-trained diffusion model. Using the synthetic data to train the existing segmentation methods, and apply them to the real images.

The core is to exploit the potential of the cross-attention map in the generative model and domain gap between synthetic and real data, providing corresponding new insights, solutions, and analysis. We introduce the preliminary of cross attention in Sec. 3.1, Mask generation and refinement with cross-attention map in text-conditioned diffusion models in Sec. 3.2, data diversity enhancement with prompt engineering in Sec. 3.4, data augmentation in Sec. 3.5.

2 Mask Generation and Refinement

Based on Equ. 1, we can obtain the corresponding cross attention map Ajs,t\mathcal{A}_{j}^{s,t}. ss denotes the attention map from ss-th layer of U-Net, and corresponding to four different resolutions, i.e., 8×88\times 8, 16×1616\times 16, 32×3232\times 32, and 64×6464\times 64, as shown in Fig. 2. tt denotes tt-th diffusion step (time). Then the average cross-attention map can be calculated by aggregating the multi-layer and multi-time attention maps as follows:

where SS and TT refer to the total steps and the number of layers (i.e., four for U-Net). Normalization is necessary due the value of the attention map from the output of Softmax is not a probability between 0 and 1.

The above method is not practical and effective, while the optimal threshold of each image and each category are not exactly the same. To explore the relationship between threshold and binary mask quality, we set a simple analysis experiment. Stable Diffusion is used to generate 1k images and corresponding attention maps for each class. The prediction of Mask2former pre-trained on Pascal-VOC 2012 as the ground truth is adopted to calculate the quality of mask quality (mIoU), as shown in Fig. 3. The optimal threshold of different classes usually are different, e.g., around 0.480.48 for ‘Bottle’ class, different from that (i.e., around 0.390.39) of ‘Dog’ class. To achieve the best quality of the mask, the adaptive threshold is a feasible solution for the various binarization for each image and class.

2.2 Adaptive Threshold for Binarization

It is challenging to determine the optimal threshold for binarizing the probability maps because of the variation in shape and region for each object class. The image generation relies on text-supervision, which does not provide a precise definition of the shape and region of object classes. For example, the masks with 0.45γ0.45\gamma and that with 0.35γ0.35\gamma in Fig. 2, the model can not judge which one is better, while no location information as supervision and reference is provided by human effort.

where Lmatch(B^,Bγ){\cal L}_{\rm match}(\hat{B},B_{\gamma}) is a pair-wise matching cost of IoU between affinity map B^\hat{B} and a binary map from attention map with threshold γ\gamma. As a result, an adaptive threshold γ^\hat{\gamma} can be obtained for each image of each class. The red points in Fig. 3 represent the corresponding threshold from matching with the affinity map. They are usually close to the optimal threshold.

3 Noise Learning

Although refined mask Bγ^B_{\hat{\gamma}} presents a competitive result, there are still existing noisy labels with low precision. Fig. 5 provides the probability density distribution of IoU for the ‘Horse’ and ‘Bird’ classes. The masks with IoU under 80%80\% account for a non-negligible proportion and may cause a significant performance drop. Inspired by noise learning for the classification task, we design a simple, yet effective noise learning (NL) strategy to prune the noise labels for the segmentation task.

NL improves the data quality by identifying and filtering noisy labels. The main procedure (see Fig. 4) comprises two steps: (1) Count: estimating the distribution of label noise QBγ^,B∗{Q_{B_{\hat{\gamma}},B^{*}}} to characterize pixel-level label noise, B∗B^{*} refers to the prediction of model. (2) Rank, and Prune: filter out noisy examples and train with errors removed data. Formally, given massive generative images and annotations {(I,Bγ^)}\{(\mathcal{I},B_{\hat{\gamma}})\}, a segmentation model θ\bm{\theta} (e.g., Mask2former , Mask-RCNN ) is used to predict out-of-sample probabilities of segmentation result θ:I→Mc(Bγ^;I,θ)\bm{\theta}:\mathcal{I}\rightarrow\bm{M}_{c}(B_{\hat{\gamma}};\mathcal{I},\bm{\theta}) by cross-validation. Then we can estimate the joint distribution of noisy labels Bγ^B_{\hat{\gamma}} and true labels, QBγ^,B∗c=ΦIoU(Bγ^,B∗){Q^{c}_{B_{\hat{\gamma}},B^{*}}}=\Phi_{\text{IoU}}(B_{\hat{\gamma}},B^{*}), where cc denotes cc-th class. With QBγ^,B∗c{Q^{c}_{B_{\hat{\gamma}},B^{*}}}, some interpretable and explainable ranking methods, such as loss reweighting can be used for CL to find label errors using. In this paper, we adopt a simple and effective modularized rank and prune method, i.e., Prune by Class, which decouples the model and data cleaning procedure. For each class, select and prune α%\alpha\% examples with the lowest self-confidence QBγ^,B∗c{Q^{c}_{B_{\hat{\gamma}},B^{*}}} as the noisy data, and train model θ\bm{\theta} with the remaining clean data. While α%\alpha\% is set to 50%50\%, the probability density distribution of IoU from the remaining clean data is presented in Fig. 5 (yellow). CL can bring an obvious gain for the mask precision, which further taps the potential of attention map as mask annotation.

4 Prompt Engineering

Previous works have shown the effectiveness of prompt engineering on diversity enhancement of generative data. These studies utilize a variety of prompt modifiers to influence the generated images, e.g., GPT3 used by ImaginaryNet . Unlike generation-based or modification-based prompts, we design two practical, reality-based prompt strategies.

Prompt with Sub-Classes. Simple text prompts, such as ‘Photo of a bird’, often results in monotony for generative images, as depicted in Fig. 6 (upper), they fail to capture the diverse range of objects and scenes found in the real world. To address this challenge, we incorporate ‘sub-classes’ for each category to improve diversity. To achieve this, we select KK sub-classes for each category from Wikihttps://en.wikipedia.org/wiki/Main_Page and integrate this information into the prompt templates. Fig. 6 (down) presents an example for ‘bird’ category. Given KK sub-classes, i.e., Golden Bullul, Crane, this allows us to obtain KK corresponding text prompts ‘Photo of a [sub-class] bird’, denoted by {P^1,P^2,...,P^K}\{\mathcal{\hat{P}}_{1},\mathcal{\hat{P}}_{2},...,\mathcal{\hat{P}}_{K}\}.

Retrieval-based Prompt. The prompt P^\mathcal{\hat{P}} still is a handcrafted sentence template, we expect to develop it into a real language prompt in the human community. One feasible solution for that is through prompt retrieval . As shown in Fig. 4, given a prompt P^\mathcal{\hat{P}}, i.e., ‘Photo of a [sub-class] car in the street’, Clipretrieval pre-trained on Laion5B is used to retrieve top NN real images and captions, where the captions as the final prompt sets. Using this approach, we can collect a total of K×NK\times N text prompts, denoted by ∑i=1K×NP^i\sum_{i=1}^{K\times N}\mathcal{\hat{P}}_{i}, for our synthetic data. During inference, we randomly sample a prompt from this set to generate each image.

5 Data Augmentation

To further reduce the domain gap between the generated images and the real-world images in terms of size, blur, and occlusion, data augmentations Φ(⋅)\Phi(\cdot) (e.g., Splicing ), as the effective strategies are used, as shown in Fig. 7. Splicing. Synthetic image usually present normal size for the foreground (object), i.e., objects typically occupy the majority of image. However, real-world images often contain objects of varying resolutions, including small objects in datasets such as Cityscapes . To address this issue, we use Splicing augmentation. Fig. 7 (a) presents one example for the image splicing (2×22\times 2). In the experiment, six scales of image splicing are used, i.e., 1×21\times 2, 2×12\times 1, 2×22\times 2, 3×33\times 3, 5×55\times 5, and 8×88\times 8, and the images are sampled from train set randomly. Gaussian Blur. Synthetic images typically exhibit a uniform level of blur, whereas real images exhibit varying degrees of blur due to motion, focus, and artifact issues. Gaussian Blur is used to increase the diversity of blur, where the length of Gaussian Kernel is randomly sampled from a range of 66 to 2222. Occlusion. Similar to CutMix , to make the model focus on discriminative parts of objects, patches of another image are cut and pasted among training images where the corresponding labels are also mixed proportionally to the area of the patches. Perspective Transform. Similar to the above augmentations, perspective transform is used to improve the diversity of the generated images by simulating different viewpoints.

Experiments

Datasets and Task. Datasets. Following the previous works for semantic segmentation, Pascal-VOC 2012 , ADE20k and Cityscapes are used to evaluate DiffuMask. Tasks. Three tasks are adopted in our experiment, i.e., semantic segmentation, open-vocabulary segmentation, and domain generalization.

Implementation Details The pre-trained Stable Diffusion , the text encoder of CLIP , AffinityNet are adopted as the base components. We do not finetune the Stable Diffusion and only train AffinityNet for each category. The corresponding parameter optimization and setting (e.g., initialization, data augmentation, batch size, learning rate) all are similar to that of the original paper. Synthetic data for training. For each category on Pascal-VOC 2012 , we generate 10k10k images and set α\alpha of noise learning to 0.70.7 to filter 7k7k images. As a result, we collect 60k60k synthetic data for 2020 classes as the final training set, and the spatial resolution is 512×512512\times 512. For Cityscapes , we only evaluate 22 important classes, i.e., ‘Human’ and ‘Vehicle’, including six sub-classes, person, rider, car, bus, truck, train, and generate 30k30k images for each sub-category, where 10k10k images are selected as the final training data by noise learning. Considering the relationship between rider and motorbike/bicycle, we set the two classes to be ignored, while evaluating the ‘Human’ class on Table 2 and Table 6. In our experiment, only a single object for an image is considered. Multi-categories generation usually causes the unstable quality of the images, limited by the generation ability of Stable Diffusion. Mask2Former is used as the baseline to evaluate the dataset. 8 Tesla V100 GPUs are used for all experiments.

Evaluation Metrics. Mean intersection-over-union (mIoU) , as the common metric of semantic segmentation, is used to evaluate the performance. For open-vocabulary segmentation, following the prior , the mIoU averaged on seen classes, unseen classes, and their harmonic mean are used.

Mask Smoothness. The mask Bγ^B_{\hat{\gamma}} generated by the Dense CRF often contains jagged edges and numerous small regions that do not correspond to distinct objects in the image. To address these issues, we trained a segmentation model θ\bm{\theta} (i.e. Mask2Former), using the mask Bγ^B_{\hat{\gamma}} generated by the Dense CRF as input. We then used this model to predict the pseudo labels for the training set of synthetic data, resulting in a final semantic mask annotation

Cross Validation for Noise Learning. In the experiment, we performed the three-fold cross-validation for each class. The five-fold cross-validation (CV) is a process in which all data is randomly split into kk folds, in our case kk == 33, and then the model is trained on the k−1k-1 folds, while one fold is left to test the quality.

2 Protocol-I: Semantic Segmentation

VOC 2012. Table 1 presents the results of semantic segmentation on the VOC 2012. The existing segmentation methods trained on synthetic data (DiffuMask) can achieve a competitive performance, i.e., 70.6%70.6\% v.s.v.s. 84.3%84.3\% for mIoU with Swin-B backbone. A point worth emphasizing is that our synthetic data does not need any manual localization and mask annotation, while real data need humans to perform a pixel-wise mask annotation. For some categories, i.e., bird, cat, cow, horse, sheep, DiffuMask presents a powerful performance, which is quite close to that of training on real (within 5%5\% gap). Besides, finetune on few real data, the results can be improved further, and exceed that of training on full real data, e.g., 84.9%84.9\% mIoU finetune on 5.05.0k real data v.sv.s 83.4%83.4\% mIoU training on full real data (10.610.6k).

Cityscapes. Table 2 presents the results on Cityscapes. Urban street scenes of Cityscapes are more challenging, including a mass of small objects and complex backgrounds. We only evaluate two classes, i.e., Vehicle and Human, which are the two most important categories in the driving scene. Compared with training on real images, DiffuMask presents a competitive result, i.e., 79.6%79.6\% vs.vs. 90.8%90.8\% mIoU.

ADE20K ADE20K, as one more challenging dataset, is also used to evaluate the DiffuMask. Table 5 presents the results of three categories (bus, car, person) on ADE20K. With fewer synthetic images (66k), we achieve a competitive performance than that of a mass of real images (20.220.2k). Compared with the other two categories, Class car achieves the best performance, with 73.4%73.4\% mIoU.

3 Protocol-II: Open-vocabulary Segmentation

As shown in Fig. 1, it is natural and seamless to extend the text-driven synthetic data (our DiffuMask) to the open-vocabulary (zero-shot) task. As shown in Table 3, compared with priors training on real images with manually annotated mask, DiffuMask can achieve a SOTA result on Unseen classes. It is worth mentioning that DiffuMask is pure synthetic/fake data and supervised by text, while priors all must need the real image and corresponding manual mask annotation. Li et al., as one contemporaneous work, use the segmentation model pre-trained on COCO to predict the pseudo label of the synthetic image, which is high-cost.

4 Protocol-III: Domain Generalization

Table 6 presents the results for cross-dataset validation, which can evaluate the generalization of data. Compared with real data, DiffuMask show powerful effectiveness on domain generalization, e.g., 69.5%69.5\% with DiffuMask v.sv.s 68.068.0 with ADE20K on VOC 2012 val. The domain gap between real datasets sometimes is bigger than that among synthetic and real data. For Motorbike class, model training with Cityscapes only achieves 28.9%28.9\% mIoU, but that of DiffuMask is 63.2%63.2\% mIoU. We argue that the main reason is domain shift in foreground and background domains, i.e., Cityscapes contains images of city roads, with the majority of Motorbike objects being small in size. But VOC 2012 is an open-set scenario, where Motorbike objects vary greatly in size and include close-up shots.

5 Ablation Study

Compared with Attention Map. Table 4(a) presents the comparison with the attention map and the impact of binarization threshold γ\gamma. It is clear that the optimal threshold for different categories is different, even various for different images of the same category. Sometimes it is sensitive for some categories, such as Dog. The mIoU of 0.40.4 γ\gamma is better than that of 0.60.6 γ\gamma around 40%40\% mIoU, which can not be neglectful. By contrast, our adaptive threshold is robust. Fig. 3 also shows it is close to the optimal threshold.

Prompt Engineering. Table 4(b) provides the related ablation study for prompt strategies. Retrieval-based and sub-classes prompt all can bring an obvious gain. For dog, 1010 sub-classes prompt brings a 7.7%7.7\% mIoU improvement, which is quite significant. It is reasonable, the fine-grained prompts can directly enhance the diversity of generative images, as shown in Fig. 6.

Noise Learning. Table 4(c) presents the impact of prune threshold α\alpha. 10k10k synthetic images for each class are used in this experiment. The gain is considerable while α\alpha changes from 0.30.3 to 0.50.5. In other experiments, we set the α\alpha to 0.70.7 for each category.

Data Augmentation. The ablation study for the four augmentations is shown in Table 4(d). Compared with the other three augmentations, the gain of image splicing is the biggest. One main reason is that the synthetic images are all 512×512512\times 512 resolution and the size of the object usually is normal, image splicing can enhance the diversity of scale.

What causes the performance gap between synthetic and real data. Domain gap and mask precision are the main reasons for the performance gap between synthetic and real data. Table 8 is set To further explore the problem. Li et al. shows that the pseudo mask of the synthetic image from Mask2former pre-trained on VOC 2012 is quite accurate, and can as the ground truth. Thus, we also use the pseudo label from the pre-trained Mask2former to train the model. As shown in Table 8, mask precision cause 6.4%6.4\% mIoU gap, and the domain gap of images causes 4.5%4.5\% mIoU gap. Notably, for the bird class, the use of synthetic data with a pseudo label resulted in better results than the corresponding real images. This observation suggests that there may be no domain gap for the bird class in the VOC 2012 dataset.

Backbone Table 7 presents the ablation study for the backbone. For some classes, e.g. sheep, the stronger backbone can bring obvious gains, i.e. Swin-B achieves 27.5%27.5\% mIoU improvement than that of ResNet 50. And the mIoU of all classes with Swin-B achieves 19.2%19.2\% mIoU improvements. It is an interesting and novel insight that a stronger backbone can reduce the domain gap between synthetic and real data. To give a further analysis for that, we present some results comparison of visualizations, as shown in Fig. 8. Swin-B brings an obvious improvement in classification, False Negatives, and mask precision.

Conclusion

A new insight is presented in this paper, demonstrating that the accurate semantic mask of generative images can be automatically obtained through the use of a text-driven diffusion model. To achieve this goal, we present DiffuMask, an automatic procedure to generate image and pixel-level semantic annotation. The existing segmentation methods training on synthetic data of DiffuMask can achieve a competitive performance over the counterpart of real data. Besides, DiffuMask shows the powerful performance for open-vocabulary segmentation, which can achieve a promising result on Unseen category. We hope DiffuMask can bring new insights and inspiration for bridging generative data and real-world data in the community.

Acknowledgements

W. Wu, C. Shen’s participation was supported by the National Key R&D Program of China (No. 2022ZD0118700). W. Wu, H. Zhou’s participation was supported by the National Key Research and Development Program of China (No. 2022YFC3602601), and the Key Research and Development Program of Zhejiang Province of China (No. 2021C02037). M. Shou’s participation was supported by the National Research Foundation, Singapore under its NRFF Award NRF-NRFF13-2021-0008, and his Start-Up Grant from National University of Singapore. Thank you to Runlong Liao for pointing out some citation errors.

References