Explore In-Context Segmentation via Latent Diffusion Models

Chaoyang Wang, Xiangtai Li, Henghui Ding, Lu Qi, Jiangning Zhang, Yunhai Tong, Chen Change Loy, Shuicheng Yan

Introduction

In-context learning provides a new perspective for cross-task modeling for vision and NLP. It enables the model to learn and predict according to the prompts. GPT-3 firstly defines in-context learning, which is interpreted as inferring on unseen tasks conditioning on some input-output pairs given as contexts. Several works also explore in-context learning in vision, where the prompts are the visual task input-outputs.

As for the segmentation field, it plays the same role as the few-shot segmentation (FSS) . Most approaches calculate the matching distance between query images and support images (known as visual prompts for in-context learning). To overcome the strict constraints on data volume and category in FSS, and to enable generalization across different tasks, several works recently extend the concept to in-context segmentation and formulate it as a mask generalization task (Fig. 1(b)). This is fundamentally different from those matching or prototype-based discriminative models (Fig. 1(a)) since it directly generates image masks via mask decoding. However, these approaches always need massive datasets to learn such correspondence.

Most recently, latent diffusion models (LDM) demonstrate great potential for generative tasks. Several works show excellent performance in conditional image content creation. Although LDM was initially proposed for generation, there have been some attempts to use it for perceptual tasks. Fig. 2(a) illustrates the mainstream pipeline for LDM-based segmentation. They typically rely on textual prompts for semantic guidance and additional neural networks to assist with LDM. However, the former may not always be available in real-world scenarios, and the latter hinders exploring the LDM’s own segmentation capabilities. Reliance on these auxiliary components might lead to poor model performance in their absence. We design a baseline model, which adopts the LDM as a feature extractor, and test this hypothesis in Sec. 4.4. Moreover, although models the segmentation process without these additional designs, it lacks the capability for in-context segmentation. Motivated by the above analysis, we argue whether in-context segmentation can be formulated as an image mask generation process that can fully explore the generation potential of LDMs.

In this paper, for the first time, we explore the potential of diffusion models for in-context segmentation, as shown in Fig. 1(c). We aim to answer the following novel questions: 1) Whether LDMs can perform in-context segmentation and achieve good enough results? 2) Which meta-architecture in LDM is best suited for in-context segmentation? 3) How do in-context instructions and output alignment affect the LDM’s performance?

To address the questions above, we propose a minimalist LDM-based in-context segmentation framework, Ref LDM-Seg, as shown in Fig. 2(b). Ref LDM-Seg relies on visual prompts for guidance without any subsequent neural networks. We analyze three significant factors: instruction extraction, output alignment, and meta-architectures. First, we propose a simple but effective instruction extraction strategy. The experiments show that the instructions obtained in this manner can provide effective guidance, and our model demonstrates robustness to incorrect instructions. Next, to align with the binary segmentation mask and 3-channel image, we design a new output alignment target via pseudo masking modeling. Then, we propose two meta-architectures, namely Ref LDM-Seg-f and Ref LDM-Seg-n. They differ in input formulation, denoising steps, and optimization targets, as shown in Fig. 3. In particular, we design two optimization targets for Ref LDM-Seg-f, respectively, in pixel and latent space. The experiments demonstrate the importance of output alignment. Unlike methods , our focus is primarily on the impact of architecture rather than data. As such, we seek to keep a certain amount of training data, which is larger than the few-shot training dataset but much smaller than that of the foundation model. To this end, finally, we propose a size-limited in-context segmentation benchmark consisting of image semantic segmentation, video object segmentation, and video semantic segmentation. We conduct comprehensive ablation studies and broadly compare our method with previous works to demonstrate its effectiveness.

There exists a task gap between generation and segmentation in diffusion models. However, LDM can still act as an effective minimalist for in-context segmentation.

The visual prompts and output alignment both play an essential role in LDM-based in-context segmentation. The success of the segmentation is determined by the former, while the latter affects its quality.

During the denoising process, the LDM-based model outputs low-frequency information first, followed by those with higher frequency.

Our proposed joint dataset avoids over-fitting and does not compromise the generalization capability to out-of-domain data.

Related Work

Diffusion Model. Diffusion models have shown remarkable performance on generation tasks, such as image generation , image editing , image super resolution , video generation , and point cloud . Although the diffusion model is initially designed for generation tasks, several works employ it for segmentation through two pipelines. The first pipeline treats the diffusion model as a feature extractor. These works typically rely on a decoder head for post-processing. Conversely, the second pipeline extracts features through a pre-trained backbone, then employs the diffusion model as the decoder head. These additional neural networks greatly influence our judgment of the true capabilities of diffusion model itself in the segmentation task. Moreover, many of these works require textual prompts for guidance. However, the textual prompts, such as categories or captions, are not always available in real-world scenarios.

In-context Learning. GPT-3 firstly defines in-context learning, which is interpreted as inferring on unseen tasks conditioning on some input-output pairs given as contexts, also known as prompts. As a new concept in computer vision, in-context learning motivates several attempts . The work is the first to adopt masked image modeling (MIM) as a visual in-context framework. Painter and SegGPT follow the same spirit but scale up with massive training data. Different from MIM, we aim to explore in-context segmentation with the latent diffusion model to explore the potential of condition generation.

Few-shot Segmentation. Few-shot segmentation aims to segment query images given support samples. The current works typically draw on the idea of metric learning by matching spatial location features with semantic centroids. Furthermore, two-branch conditional networks , 4D dense convolution , data augmentation and transformer-based architecture are also widely adopted by researchers. Although few-shot segmentation and in-context segmentation share a similar episode paradigm, the dataset used in few-shot segmentation is typically very small, making the model prone to overfitting. This may influence the evaluation of the generalization ability.

Parameter Efficient Tuning. These methods aim to fine-tune only a tiny portion of parameters to adapt the pre-trained foundation models to various downstream tasks. They maintain the pre-trained knowledge of the foundation models. However, they suffer from inadequate expressive power. In our experiments, we adopt the low-rank adaptation (LoRA) , a tool widely used in diffusion models, to demonstrate the task gap between generation and segmentation. Although some works try to avoid the dilemma by fine-tuning the prompts only, it is time-costly to learn and restore a new embedding for each prompt. In contrast, we employ a prompt encoder to extract in-context instructions from prompts.

Method

In this section, we first revisit the diffusion model, clarify our setting, and provide notations (Sec. 3.1). We then introduce our framework, Ref LDM-Seg, in Sec. 3.2. We discuss three crucial components: instruction extraction, output alignment, and meta-architecture.

Diffusion Model. Diffusion models belong to probabilistic generative models that define a chain of forward and backward processes. In the forward process, the model gradually corrupts the data sample z0z_{0} into a noisy latent ztz_{t} for t∈1...Tt\in{1...T}: q(zt∣z0)=N(zt;α‾tz0,(1−α‾t)I)q(z_{t}|z_{0})=\mathcal{N}(z_{t};\sqrt{\overline{\alpha}_{t}}z_{0},(1-\overline{\alpha}_{t})I) , where α‾t=∏i=0Tαs=∏i=0T(1−βs)\overline{\alpha}_{t}=\prod_{i=0}^{T}\alpha_{s}=\prod_{i=0}^{T}(1-\beta_{s}) and β\beta is the noise schedule. During the training, the model learns to predict the noise ϵθ(zt,t)\epsilon_{\theta}(z_{t},t) under the supervision of L2L_{2} loss as L=12∣∣ϵθ(zt,t)−ϵ(t)∣∣2\mathcal{L}=\frac{1}{2}||\epsilon_{\theta}(z_{t},t)-\epsilon(t)||^{2}. During the inference, the model starts from a random noise zT ∼N(0,1)z_{T}~{}\sim\mathcal{N}(0,1), and gradually predicts the noise. Then, it reconstructs the original data z0z_{0} with TT steps based on the estimated noise.

Task Setting. We define in-context segmentation (ICS) as a task of combining few-shot segmentation (FSS) and video semantic segmentation (VSS). Our ICS setting differs from standard few-shot in the following three aspects:

Dataset Size. Compared with few-shot segmentation, our proposed ICS uses a larger dataset to improve generalization and avoid over-fitting on specific datasets.

Data Category. Realistic scenarios do not strictly distinguish between base classes and novel classes. Our proposed in-context segmentation is more practical.

Data Type. Our proposed in-context segmentation incorporates images and videos into the same framework, which conforms better to the trend of generalist models.

2 Framework

The LDM is initially designed for generative tasks. Most of the work that applies LDM to segmentation requires subsequent neural networks to handle intermediate features or imperfect segmentation results. However, as a generation model, adopting such a design does not realize the generation potential of LDM. To this end, we chose Stable Diffusion as the base model with minimum changes to explore such potential. This subsection explores three key factors influencing the process: instruction extraction, output alignment, and meta-architectures.

Instruction Extraction. Instructions play an important role in LDM. They act as the compressed representation of the prompt, controlling the model to focus on the region of interest and guiding the denoising process. Stable Diffusion encodes textual prompts into instructions using CLIP. To align with pre-trained weights, we use CLIP ViT as our prompt encoder. The instructions are the output of the last hidden layers, except for the first token. We use a linear layer as an adapter to transform the dimension of the instructions. The annotation mask MsM_{s} is used as an attention map in cross-attention layers, forcing the model to focus on those tokens that belong to the foreground regions:

where FiF_{i} indicates the ithi^{th} adapter we use and τi\tau_{i} are the corresponding instructions for ithi^{th} cross-attention layer.

Output Alignment. It is important to note that we employ LDM for the segmentation task, so the inconsistency between 1-channel masks and 3-channel images is non-negligible. A pseudo mask must be designed to align the gaps as an intermediate step toward the binary segmentation mask. We argue that a good pseudo mask strategy should meet the following requirements: 1) Binary masks can be obtained from the pseudo masks through simple arithmetic operations. 2) The pseudo masks should be informative beyond binary values.

An intuitive method follows a mapping rule that transforms the binary masks MM to 3-channel pseudo masks PMvPM_{v}:

where bgbg and fgfg indicate background and foreground, respectively. MiM_{i} is the value in position ii. aa, bb are both scalar, indicating the value of a specific channel in the pseudo masks. We set a<ba<b.

It is easy to recover the binary segmentation mask with simple arithmetic operations:

Beyond the vanilla design, we also propose an augmented strategy to fuse the information of images into pseudo masks. Denote the image as II, and the augmented pseudo masks are formulated as follows:

where α\alpha controls the strength of the information of image, and IiI_{i} is the value in position ii. We set α>1\alpha>1 and a<ba<b.

Meta-architectures. As shown in Fig. 3, we explore two representative meta-architectures, namely Ref LDM-Seg-f and Ref LDM-Seg-n. The difference between them mainly lies in the optimization target and denoising time steps.

Ref LDM-Seg-f indicates one-step denoising process and the optimization target is the segmentation mask itself. As shown in Fig. 3(a), a noise variant ϵt\epsilon_{t} is added to the latent variant zqz_{q}. Our ICS model ff, U-Net, takes as input the noisy latent zt=zq+ϵtz_{t}=z_{q}+\epsilon_{t}, where ϵt\epsilon_{t} is the output of the noise scheduler, tt controls the noise strength.

We adopt the low-rank adaptation strategy to maintain the knowledge of pre-trained weights and avoid catastrophic forgetting. All parameters are frozen except those q,k,v,o projections in attention layers. We propose two optimization strategies that align the model outputs with the ground truth in the pixel space (10) or in the latent space (11), respectively. We employ the L2 loss, typically used in LDM, rather than any explicit segmentation loss.

In the inference stage, Ref LDM-Seg-f conducts only a one time step and outputs the segmentation (pseudo) masks. The video is treated as a sequence of images. The first frame and its annotation are used as prompts, and subsequent frames are inferred conditioned on it. For videos containing multiple categories, we first calculate the probability of each category as a foreground in turn and then select the category with the highest probability:

Ref LDM-Seg-n indicates multi-step denoising process and employs an indirect optimization strategy. Unlike Ref LDM-Seg-f, it starts from Gaussian noise and gradually denoises to get the final segmentation mask.

where the noise scheduler determines tt.

Similar to Ref LDM-Seg-f, the ICS model ff takes as input the latent variable ztz_{t} and the instructions τ\tau, but outputs the estimation of the noise rather than the pseudo mask. We also adopt L2 loss as follows:

To minimize the randomness brought by initial noise, strengthen the guidance of in-context instructions, and maintain the consistency between the outputs and the queries, we adopt classifier-free guidance (CFG) . In the training stage, the query latent zqz_{q} and condition τ\tau are randomly set to null embedding with probability p=0.05p=0.05.

where γq\gamma_{q} and γτ\gamma_{\tau} control the guidance of query and in-context instruction, respectively.

Empirical Study

In this section, we explore the Ref LDM-Seg on in-context segmentation empirically. We first demonstrate the experimental settings in Sec. 4.1. Then, we conduct a comprehensive analysis of how frameworks (Sec. 4.2) and data (Sec. 4.3) influence our model, respectively. Finally, we report the results compared with other methods in Sec. 4.4.

Benchmark Details. As mentioned in Sec. 3.1, the in-context segmentation model aims to solve multiple tasks with one model, regardless of data type and domain. To this end, we adopt several popular public datasets as part of our benchmark (Tab. 1), including PASCAL , COCO , DAVIS-16 and VSPW . The training data for VSPW is sampled every four frames for each video. All ‘stuff’ categories are annotated as background. The combined dataset has around 100K images for training. Note that our interest focuses on the model’s performance with limited training data rather than an infinite expansion. The scaling law is not our concern.

Implementation Details. We utilize the Stable Diffusion 1.5 model as the initialization and set the resolution as 256×256256\times 256. Our model is jointly trained on the combined dataset for 80K iterations with a batch size of 16. We employ an AdamW optimizer and a linear learning rate scheduler. CLIP ViT/L-14 is adopted as the prompt encoder. We set the CFG coefficient for query and instructions as 1.5 and 7, respectively. By default, the rank of LoRA is equal to 4 in all experiments. We use the VAE of Stable Diffusion 1.5 without any fine-tuning.

Our training follows the spirit of episodic learning. For the image dataset, images with the same semantic labels are considered as a pair of queries and prompts. The video dataset follows the image dataset but additionally requires that the query and prompt come from the same video.

In Sec. 4.2 and Sec. 4.3, our model is trained on the combined dataset but evaluated on the COCO dataset. The results for all datasets are reported in Sec. 4.4.

Evaluation Metrics. We adopt the class mean intersection over union as our evaluation metric in empirical study, which is formulated as mIoU=1C∑i=1CIoUimIoU=\frac{1}{C}\sum_{i=1}^{C}IoU_{i}. CC is the number of classes except background. We also report the foreground-background IoU for other image-level tasks in our benchmark following . For video-level tasks, we respectively adopt their evaluation metrics in DAVIS-16 and VSPW .

2 Study on Framework Design

In this subsection, we investigate the impact of two meta-architectures and the corresponding output alignment and optimization strategies from the architectural perspective.

Meta-architecture. Tab. 3(a) presents a comparison of meta-architecture. We also consider low-rank adaptation (LoRA). Overall, Ref LDM-Seg-f performs better than Ref LDM-Seg-n. In the best case, the former can achieve a performance of 59.6 mIoU, while the latter can only reach 39.3 mIoU. Interestingly, the two meta-architectures demonstrate distinct characteristics depending on whether LoRA is utilized. When using LoRA, Ref LDM-Seg-f gains around seven mIoU improvement, but the performance of Ref LDM-Seg-n drastically decreases.

Output Alignment. The output of Ref LDM-Seg-n at different time steps is presented in Fig. 4, aligned with the augmented pseudo mask (Equ. 8). Interestingly, the target’s relative position is determined at the beginning of the denoising process. The low-frequency signal is generated first. Then, our model gradually generates high-frequency information like shape or texture as the denoising process proceeds. We quantitatively study the effect of output alignment for Ref LDM-Seg-n in Tab. 3(b). The augmented pseudo mask (Equ. 8) significantly outperforms the vanilla one (Equ. 4) and improves the performance from 24.1 mIoU to 31.9 mIoU. As the vanilla pseudo mask (Equ. 4) only has two values (foreground and background), we hypothesize that the trivial strategy reduces the capability of Ref LDM-Seg-n to capture the semantics of the image. In this case, the model may learn some shortcuts.

Optimization Space. In Tab. 3(c), we study the effects of optimization space for Ref LDM-Seg-f. Optimization in the pixel space (Equ. 10) benefits the model and brings about 3.7 mIoU improvement. Moreover, LoRA seems more critical for optimization space. Ref LDM-Seg-f drops 3.3 mIoU without LoRA, even though it is optimized in the pixel space. The performance of Ref LDM-Seg-f with different LoRA ranks is reported in Tab. 3(d). All these models use latent space optimization. Ref LDM-Seg-f (LO, rank=8) in Tab. 3(d) achieves a mIoU of 57.9, which is close to Ref LDM-Seg-f (PO, rank=4) in Tab. 3(c).

Discussion. Our experiments show that there exists a task gap between generation and segmentation for LDM. Intuitively, Ref LDM-Seg-f outputs the segmentation mask directly, aligning with mainstream segmentation models. In contrast, Ref LDM-Seg-n gradually denoises in the latent space. Experiments related to LoRA also support the hypothesis above. LoRA retains most of the pre-trained knowledge while limiting its expressive power, which hinders the performance of Ref LDM-Seg-n.

3 Study on Various Datasets

In-context Instruction. Our model segments the target region based on in-context instructions. Fig. 5 illustrates some examples. In the first and second cases, our model accurately segments the target region based on the single instruction provided. The third case shows the model can accept several different instructions without performance degradation. In the fourth case, the provided instruction becomes ineffective when it conflicts with the query.

Tab. 4(c) reports the results with the number of instructions from 1 to 10. The model’s performance improves as the number of instructions increases, saturating at 63.3 mIoU with five instructions. Further increasing the number of instructions results in only a slight further improvement.

Out-of-domain Dataset. We also test our model on an out-of-domain dataset, FSS1000 . This dataset includes 1000 categories that were not present in the training data. Additionally, it contains numerous ‘stuff’ categories labeled as background during training. We compare with two representative generalist methods, Painter and PerSAM . Our proposed Ref LDM-Seg-f outperforms these methods and achieves an mIoU of 69.8.

Dataset Combination. Tab. 4(a) presents the outcomes of Ref LDM-Seg-f trained on a single dataset. The model exhibits overfitting when trained solely on a small dataset like PASCAL. Conversely, the model trained exclusively on the COCO dataset demonstrates superior results on COCO, but it lacks generalization compared to the combined dataset.

Pre-trained Weight. We also try to use SDXL for initialization or use a larger resolution, as shown in Tab. 4(d). Surprisingly, the two architectures exhibit different characteristics. A larger model like SDXL or a larger resolution is only beneficial for Ref LDM-Seg-f but unsuitable for Ref LDM-Seg-n. This may be due to the task gap discussed in Sec. 4.2.

4 Comparison with Previous Methods

We compare our methods with related works on our benchmark and report the performance under the one-shot segmentation dataset.

Benchmark Results. We compare our method with previous baselines in Tab. 4. Specifically, the specialist models belong to the family of discriminative models. Prompt Diffusion uses a ControlNet architecture. Painter belongs to the family of masked image modeling. It is a visual foundation model trained on massive data. PerSAM employs a strong foundation model, SAM , for segmentation. We also design LDM-FE, a specialist baseline that employs an LDM to extract features and a PFENet decoder for prediction. In the ICS task, category labels are not available. Therefore, the textual prompt in LDM-FE is set as null.

On image tasks, it can be seen that Ref LDM-Seg-f achieves the best results compared with these methods. Ref LDM-Seg-n also achieves a decent performance of 62.8 mIoU on PASCAL and 39.3 mIoU on COCO. LDM-FE performs similarly to other specialist models but is inferior to Ref LDM-Seg-f, possibly due to the lack of in-context instructions. On video tasks, our models show comparable performance compared with the generalist models. Considering the huge amount of training data used in Painter and SAM, it is acceptable that our method cannot exceed these foundation models on specific metrics.

Fig. 6 shows the visual comparison between our model and previous works on the COCO dataset. It is evident from these results that our model successfully segments the target region, whereas other methods fail to establish the semantic connection between prompt and query, leading to missing and false-positive predictions. Fig. 7 presents more visualizations on video semantic segmentation of the VSPW dataset. From these visual results, our model exhibits robustness to various scenes and categories in both image and video segmentation tasks.

Results on One-shot Segmentation. We also test our model under the few-shot segmentation setting. As shown in Tab. 5, our model achieves decent or better performance on almost all folds of COCO-20i{20^{i}} dataset. Specifically, it significantly outperforms some recently proposed generalist model, PerSAM by around 30% mIoU.

Conclusion

For the first time, we explore in-context segmentation via latent diffusion models. We propose two meta-architectures and design several output alignment strategies and optimization methods. We observe a task gap between generation and segmentation in diffusion models, but the LDM itself can be an effective minimalist for in-context segmentation tasks. We also propose an in-context segmentation benchmark and achieve comparable or even better results than specialists or vision foundation models.

Broader Impact. Our work makes a first-step exploration of conditional mask generation for in-context segmentation. Our proposed method provides a unified and general framework. We hope our exploration will attract the community’s attention to the unification of visual generation and perception tasks via diffusion models.

Acknowledgement. This work is supported by the National Key Research and Development Program of China (No. 2023YFC3807600).

References

Appendix

Overview. In this supplementary, we present more results and details:

Appendix B. More empirical studies and analysis.

Appendix C. More visual results and discussion.

Appendix A More Details on Method

Implementation Details. In the training stage, we use the PNDM noise scheduler and set the whole time step as 1000. The input images are randomly resized and cropped to 256×256256\times 256. Additionally, they are flipped with a probability of 0.5. We set the time step in Ref LDM-Seg-f as 0. The AdamW optimizer use a learning rate of 1e-4 and weight decay of 1e-2. Our experiments are mainly conducted on 4 NVIDIA A100-80G GPUs. In the inference stage, the denoising time steps for Ref LDM-Seg-n is 20. The number of trainable parameters is 872M and 13M for the model without LoRA and with LoRA, respectively.

Appendix B More Empirical Studies

Classifier-free Guidance. Fig. 9 qualitatively demonstrates the effect of classifier-free guidance (CFG). When γq=1\gamma_{q}=1 and γτ=1\gamma_{\tau}=1, the denoising process degenerates into a trivial form, resulting in more missing or false-positive predictions. Moreover, the in-context instructions are essential in our model (the third and fourth columns). Ref LDM-Seg-n fails to segment the target regions accurately with weak visual guidance.

Comparison with SegGPT. Tab. 6 compares our Ref LDM-Seg-f with SegGPT on PASCAL and COCO dataset.

Appendix C More Visual Results

Denoising Process of Ref LDM-Seg-n. Fig. 8 provides more examples to demonstrate how Ref LDM-Seg-n generates pseudo masks in the inference stage. Starting from Gaussian noise, Ref LDM-Seg-n gradually denoises and differentiates the foreground and background regions. The outputs align with the augmented pseudo masks PMaPM_{a} (the sixth column). It is evident that the information of query images is fused into the segmentation mask.

More Visualization on images and videos. Fig. 11 and Fig. 12 provide more visual results on COCO and VSPW. Our model performs well across different scenarios and categories.

Failure Cases Analysis. We show some failure cases in Fig. 10. These samples exhibit significant changes in attitude, focus, or size from the prompts to the queries. Our model fails in these hard examples.

Limitations and Future Work. According to previous works , huge amounts of training data are essential to exploit the potential of LDM fully. To this end, we will scale up the training data in our future work. We also consider scaling up the parameters if more training data is available. Moreover, we will explore more advanced prompt encoder architectures and prompt engineering methods.