InstructCV: Instruction-Tuned Text-to-Image Diffusion Models as Vision Generalists
Yulu Gan, Sungwoo Park, Alexander Schubert, Anthony Philippakis, Ahmed M. Alaa
Introduction
Recent work on text-to-image models has achieved impressive performance in image synthesis . Particularly, diffusion models have demonstrated remarkable capabilities of transforming diverse text prompts into realistic images, even for novel concepts. Models like DALL·E and Stable Diffusion highlight this progress, now finding use in real-world applications. However, despite these impressive results, generative text-to-image models have so far not been exploited as a unified basis for standard visual recognition tasks. Instead, the predominant approach for these tasks is to design dedicated task-specific architectures and loss functions , foregoing the opportunity to learn generalizable representations across heterogeneous problem domains and data landscapes.
Previous attempts to create unified models for computer vision tasks have predominantly relied on prompt tuning approaches in conjunction with sequence-to-sequence architectures . This general framework enables conditioning on input images as well as task-specific prompts by representing image pixels and (trainable) prompts as sequences of discrete tokens. When trained on multi-task datasets, the resulting output tokens align with the desired outcomes for the respective tasks prompted. One illustrative example of this approach is Pix2Seq , which follows an autoregressive language modeling approach for processing tokenized image pixels and task identifier codes. Another example includes a class of methods based on visual prompting, which defines task-specific prompts in pixel space, unifying multiple vision tasks within a common “inpainting” framework . In both examples, the task-specific prompts steer a single architecture to execute multiple vision tasks. However, these prompts consist of (uninterpretable) numerical values derived from specific training datasets, which may limit their ability to generalize to new datasets, tasks, or categories.
In this paper, we propose a unified model for computer vision tasks that conducts a given task by following natural language instructions (Fig. 1). Our framework, dubbed InstructCV, repurposes generative text-to-image models to create a universal language interface for vision tasks. It does so by casting multiple computer vision tasks as text-to-image generation problems, where textual prompts (instructions) serve as explicit task descriptors, guiding the generation process to produce the visual task output corresponding to the input image. By conditioning on natural language descriptions of vision tasks, InstructCV enhances the representation of semantic coherence between images and language prompts, improving the model’s generalization capabilities to new human-written instructions and new categories compared to prior “generalist” vision models .
To train InstructCV, we follow an instruction tuning approach applied to a pretrained conditional diffusion model (Stable Diffusion). We generate the instruction tuning data by constructing a multi-modal, multi-task training dataset that comprises tuples of textual instructions, input images and visually-encoded task outputs. We do so by first combining several standard computer vision datasets across multiple tasks including segmentation, object detection, depth estimation, and classification. Next, in order to create heterogeneous and semantically rich textual instructions, we use a large language model (LLM) to paraphrase prompt templates for each vision task. Finally, we encode the output of the vision task associated with each instruction in the form of an output image (e.g., a masking pattern for semantic segmantation). Using this dataset, we utilize the InstructPix2Pix architecture to instruction-tune a text-to-image diffusion model, transforming its functionality from a generative image synthesis model into an instruction-guided multi-task vision learner.
Our experiments demonstrate that InstructCV achieves competitive results compared to other vision generalist and task-specific vision models. Particularly, InstructCV displays compelling generalization properties, surpassing the performance of state-of-the-art vision generalist models on external datasets as well as on unseen prompts in open-vocabulary segmentation tasks.
InstructCV
The InstructCV framework comprises two key steps: (a) construction of a multi-modal and multi-task instruction-tuning dataset, and (b) finetuning a pretrained conditional diffusion model using the dataset generated in step (a). (See Fig. 2 for a pictorial depiction.) Details of each step are provided below.
We combine four widely-used computer vision datasets (MS-COCO , ADE20K , Oxford-III-Pets and NYUv2 ) covering four vision tasks (semantic segmentation, object detection, monocular depth estimation and classification), into a single multi-task dataset in which is the input image, is the task output and is the task identifier ( in our setup). We convert into a multi-modal instruction-tuning dataset , in which the task identifier for training point is expressed as a natural language instruction , and the label is represented in a visual format . We construct the dataset through the following steps.
LLM-based instruction generation. For each vision task under consideration, we pick a prompt template that describes the task, e.g., “Segment the %category%” for the semantic segmentation task. We then attach a prompt to each training data point by inserting the category within the image in the corresponding task template, e.g., = “Segment the cat”. We consider two incarnations of our instruction-tuning dataset. First, we consider a baseline dataset that only uses the deterministic task-specific prompt templates described above. We refer to as this dataset as the fixed prompts (FP) instruction-tuning dataset \mathcal{D}^{\mbox{FP}}_{\mbox{\mathcal{I}}}=\{({\bf x}_{i},{\bf v}({\bf y}_{i}),\mathcal{I}^{m_{i}}_{\mbox{temp}}({\bf x}_{i},{\bf y}_{i}))\}_{i}. In addition, we use a T5-based paraphrasing LLM to generate rephrased versions of the prompt template to create a diverse range of instructions. As shown in Fig. 2(a) (top), the LLM takes as an input the prompt template (e.g., “Segment the %category%”) and produces a wide range of paraphrased variants (e.g., “Highlight the %category%”). This procedure ensures that our instruction set is varied yet firmly tied to the core intent of the original prompt. We use the LLM to sample a rephrased variant of the prompt template for each training data point in . We refer to the resulting instruction tuning dataset as the rephrased prompt (RP) dataset \mathcal{D}^{\mbox{RP}}_{\mbox{\mathcal{I}}}=\{({\bf x}_{i},{\bf v}({\bf y}_{i}),\mathcal{I}_{i}({\bf x}_{i},{\bf y}_{i}))\}_{i}.
Visual encoding of task outputs. We format the target label of each task to represent it in the same RGB image space as the input image through a “visual encoding” function (Fig. 2(a) (bottom)). This enables casting all tasks in a unified text-to-image generation framework, leveraging Pix2Pix architectures. That is, given an image and instruction , InstructCV produces an image that encodes the task output. In the following, we provide the definition of for all tasks under study.
(1) Semantic Segmentation. The target output of this task is typically an assignment of a label or category to every pixel in an image. A natural choice of for the semantic segmentation task is a binary mask that labels pixels in the input image belonging to the prompted category in .
(2) Object Detection. Here, the goal is to identify the spatial position of a category in an image using a bounding box, i.e., the label comprises bounding box coordinates . We define for object detection as the image with a bounding box overlaid according to the coordinates in .
(3) Monocular Depth Estimation. The target of this task is the depth value (i.e., distance relative to the camera) of each pixel in the RGB image . For this task, we define the visually-encoded target as an RGB image in which pixel colors encode the depth values. This encoding is done by converting depth values ranging from to meters (based on depth ranges in the NYUv2 dataset ) into the discrete space for RGB image representation, i.e., . We then apply the same value across all three RGB channels to create a visual depth map.
(4) Image Classification. In multi-class image classification, the target label is a categorical value indicating the object depicted in the image . To represent image classification in a Pix2Pix format, we resort to a color-coding methodology. To this end, we use a prompt template of the form: “Display %color% if the image contains %category%”. We sample random colors when filling in the template for individual training points. This steers the text-to-image model to produce an image consisting of the pure color block if the category specified in the prompt is visible in . (Note that this approach only enables us to predict if a specific category is in the input image . For multi-class classification, we need to use a series of prompts specifying all categories of interest one at a time.)
2 Instruction-tuning a Latent Diffusion Model
We use our instruction-tuning dataset \mathcal{D}_{\mbox{\mathcal{I}}}=\{({\bf x}_{i},{\bf v}({\bf y}_{i}),\mathcal{I}_{i})\}_{i} to train a (conditional) diffusion model that conducts the vision task specified in the instruction on the input image , producing a visually-encoded task output . By finetuning the text-to-image diffusion model using \mathcal{D}_{\mbox{\mathcal{I}}}, we steer its functionality from a generative model to a language-guided multi-task vision learner. We use a training procedure similar to that of InstructPix2Pix , which applies uses a similar multi-modal dataset (pairs of images and editing instructions) to train an instruction-guided image editing model.
Diffusion models generate data by gradually denoising a normally distributed random variable; this process amounts to learning the reverse dynamics of a Markov chain with fixed length . Latent diffusion models applies this approach within the latent space of a pretrained variational autoencoder with encoder and decoder . Training a diffusion model involves a forward diffusion and a reverse denoising process. During the forward process, the image is transformed to its latent representation v(y), which is then injected with Gaussian noise over steps:
where the time-varying constants control how much noise is added at each timestep and are chosen such that roughly converges to a standard Gaussian vector. This forward process does not contain any trainable parameters and can be described as . In the reverse diffusion process, the objective is to learn a model to progressively denoise the latents to recover the initial latent . The target image can then be reconstructed as . The reverse diffusion process can be written as:
where the means are typically parameterized using neural networks and the variances are predetermined constants. Such a denoising model is learned by optimizing a reweighted variant of the variational lower bound on the data distribution , i.e.,
where is a standard normal random variable and , where and are based on the diffusion distribution . The noise predictor is obtained from the parameterization . The model is trained to predict the noise vector at each time-step in order to denoise the latent variable .
Instruction-tuning via image and text conditioning. The objective in (3) can be further refined to condition on both the input image and instruction to generate the desired output —this can be achieved by learning a conditional noise predictor , minimizing the following loss function:
We use a pretrained Stable Diffusion checkpoint, which exhibits strong text-to-image generation capabilities, as the backbone architecture of our model. For text conditioning, we adopt the same methodology as in , utilizing the instruction instead of image captions as textual inputs. For image conditioning, we concatenate the encoded input image with the latent , which are then fed into input to the first layer of the noise predictor .
Related Work
Repurposing diffusion models for vision tasks. Diffusion models have achieved impressive performance in image generation , text-to-image synthesis , as well as generation of other modalities such as video and audio . The idea of repurposing diffusion models to tackle standard computer vision tasks has been considered before to develop models for object detection and open-vocabulary panoptic segmentation . These approaches were limited to single-task settings with specialized loss functions and (unimodal) architectures. Contrarily, InstructCV provides a unified architecture for multi-task learning, with a natural language interface that enhances generalization to new datasets and categories. The idea of “instruction-tuning” text-to-image diffusion models was introduced in with the objective of finetuning the model to follow editing instructions. InstructCV builds on this framework to adapt text-to-image models for performing conventional visual recognition tasks. To our knowledge, this is one of the earliest efforts in this direction.
Vision Generalists. Several prior attempts have aimed to develop unified models capable of executing multiple vision tasks within a single, shared architecture. Motivated by successes of LLMs, recent work has attempted to design such generalist models based on sequence-to-sequence architectures. Among these, models such as Florence , OFA , CoCa and BEiT-3 , learn general representation encoder, which require individual finetuning to each specific downstream task. Methods such as Unified-IO and Pix2Seq-v2 build single architectures that are capable of performing multiple vision tasks via prompt tuning. However, the sequence-based operation of these models results in slow inference speeds, and the tuned prompts may not generalize to unseen datasets/categories. Another line of work proposes vision transformer-based architectures that frame different vision tasks as inpainting problems . This work focuses on in-context learning based on visual prompts and does not consider language-based instructions, which we believe is a more natural interface for general-purpose models. To the best of our knowledge, the only generalist model that supports a language-based interface for vision tasks similar to that of InstructCV is the VisionLLM model developed in . This is an LLM-based framework that treats images as a foreign language and aligns vision-centric tasks with language tasks that can be flexibly defined using language-based instructions. VisionLLM and InstructCV share a common objective but use different approaches. VisionLLM finetunes a pre-trained LLM using vision-centric tasks, whereas InstructCV finetunes a text-to-image model by substituting image captions with instructional text. We were unable to empirically compare the two models as the code for VisionLLM was not available at the time of writing this paper.
Experiments
We evaluate InstructCV across the four vision tasks under study (semantic segmentation, object detection, monocular depth estimation and image classification). For this purpose, we consider widely used datasets for each task: ADE20k for semantic segmentation, MS-COCO for object detection, NYUv2 for depth estimation and Oxford-IIIT Pet for classification. In what follows, we explain the processing steps and evaluation procedure for all tasks under consideration.
Semantic Segmentation. ADE20K covers semantic categories and comprises images of which we use for training, for validation, and for testing. We follow the same protocol as suggested in to implement the training/test split. At inference time, we average the outputs of the three channels of the output image to obtain the final segmentation mask. We evaluate the accuracy of segmentation masks using the Mean Intersection over Union (mIoU) metric.
Object Detection. MS-COCO contains training and validation images with labels for different categories. We follow the same protocol as in Pix2Seq to set up the training/test split. At inference time, we follow the post-processing steps in Appendix A.3 to derive the coordinates and category of each Region of Interest (RoI) from the output image . We then aggregate the results for all categories in order to calculate the Mean Average Precision (mAP).
Depth Estimation. The NYUv2 dataset consists of indoor scenes captured by a Microsoft Kinect camera. We follow the official training/test split, with image-depth pairs used for training, and used for testing. For the test images we report the Root Mean Square Error (RMSE), absolute mean relative error (A.Rel), and the share of interior pixels with a different threshold . During inference, we take the average across the three channels of the output image and apply the inverse of the linear transformation used in training to obtain a depth estimate in the range of $$ meters.
Image Classification. As mentioned in Section 2.1, we implement image classification by asking the InstructCV model if a category is visible in the input image using the following template prompt: “Display %color_1% if the image contains %category%, else display %color_2%”. We evaluate the accuracy of InstructCV for classification by assessing whether contains the color block corresponding to the correct category. We do so by generating image pairs based on the Oxford-III Pet dataset for binary classification and augment these with negative pairs, where the category mentioned in the language instruction is not present. We then evaluate a classification score defined as: , i.e., the Euclidean distance between the pixel-wise colors of the output image and target color block specified in the task instruction .
Our pooled multi-modal/multi-task instruction-tuning dataset comprises 180,285 images. We create two versions of the dataset, \mathcal{D}^{\mbox{FP}}_{\mbox{\mathcal{I}}} & \mathcal{D}^{\mbox{RP}}_{\mbox{\mathcal{I}}}, with fixed and rephrased prompts as described in Section 2.1.
External Datasets. Since the generalist baselines and InstructCV were trained on different datasetsUnified-IO has been trained on datasets while Pix2SeqV2 was trained on MS-COCO., we consider additional external datasets that are outside of the training distribution of all baselines. To this end, we consider the following datasets: ImageNet for classification, SUNRGB-D for object detection and VOC for segmentation and monocular depth estimation tasks.
Implementation Details. We train InstructCV for 20 epochs on NVIDIA A GPUs over 10 hours. The training involves images with a resolution of and incorporates data augmentation including random horizontal flipping and cropping with a batch size of . The proposed model is initialized with EMA weights obtained from the Stable Diffusion checkpoint, and trained with a learning rate without any warm-up stage. Further details can be found in appendix A.1. We refer to the models trained on \mathcal{D}^{\mbox{FP}}_{\mbox{\mathcal{I}}} and \mathcal{D}^{\mbox{RP}}_{\mbox{\mathcal{I}}} as InstructCV-FP and InstructCV-RP, respectively.
2 Results
Table 1 presents a quantitative comparison between InstructCV and task-specific as well as generalist vision models in both in-distribution and out-of-distribution datasets. (Fig. 4 displays illustrative examples of InstructCV outputs.) For depth estimation, we compare InstructCV with DepthFormer , BinsFormer and UviM . For semantic segmentation, our baselines include Mask2Former and SSA . For classification, we consider baseline classifiers with ResNet and ViT backbones. Lastly, for object detection, we consider DETR and Mask R-CNN . Task-specific models for semantic segmentation, classification, and object detection have not been assessed on datasets beyond their training distribution. This is due to the fact that these new datasets introduce categories absent in the model’s original training, which precludes zero-shot generalization. We consider Unified-IO and Pix2SeqV2 as generalist vision baselines. Note that we only report object detection results for Pix2SeqV2 , as this model does not cover the other tasks involved in the development of InstructCV. Additional tasks Pix2SeqV2 is able to perform, such as keypoint detection, have not been integrated into the current version of InstructCV.
Performance comparisons. Overall, Instruct CV performs competitively compared to both generalist and task-specific baselines across all four tasks. For depth estimation within in-distribution data, InstructCV achieves a 10% improvement in RMSE compared to the second best model, BinsFormer . Notably, InstructCV demonstrates strong generalization performance to unseen datasets, surpassing all baselines by a large margin, with the exception of classification tasks. For instance, for the task of depth estimation the task-specific models Binsformer and DepthFormer experience high performance drops of and , respectively, while InstructCV’s performance improved by resulting in a lower RMSE than the best task-specific model. Similarly, for the task of object detection, the performance of the Pix2SeqV2 generalist model drops by when evaluated on VOC—InstructCV outperforms this generalist model by a in mAP@0.5. Among all baselines, only Unified-IO demonstrates comparable generalization properties to unseen datasets. However, InstructCV outperforms Unified-IO on all tasks except for classification. Notably, for semantic segmentation, InstructCV surpasses the performance of Unified-IO by in mIOU. Classification is the task where InstructCV exhibited its weakest performance compared to baselines.
Generalization to unseen categories. Most existing task-specific and multi-task models are built based on a fixed category pool determined by their training data, and do not exhibit zero-shot capabilities in detecting, segmenting or classifying categories outside of this set . Because InstructCV leverages semantically meaningful instructions and a pre-trained text-to-image generative model to guide learning, we expect its task-specific capabilities to generalize to new categories. To investigate its generalization capabilities to unseen categories, we evaluate InstructCV on an open-vocabulary segmentation task using the FSS-1000 dataset . This dataset comprises object classes, many of which have not been previously annotated in other computer vision datasets. Because the baselines in Table 1 do not accommodate many of these categories, we instead compare InstructCV with generalist vision models that exhibit zero-shot capabilities—the visual prompting by Inpainting model in and the Generalist Painter model in . Both methods rely on visual prompting approaches, where the input prompt is a set of pixels and the vision tasks are all represented within a unified inpainting framework. Table 2 presents the results. Overall, InstructCV outperforms Generalist Painter and Inpainting by and in mIoU, respectively. This demonstrates the value of repurposing text-to-image models and language-based prompts to improve the generalization capabilities of generalist approaches to computer vision tasks.
Generalization to new user-written instructions. The InstructCV-RP model undergoes training using the dataset \mathcal{D}^{\mbox{RP}}_{\mbox{\mathcal{I}}}, which comprises a variety of instructions, all aimed at conveying a common underlying intent (i.e., describing a specific visual task). Through this diverse training data that encompasses a broad spectrum of phrasings and descriptions of the same task, we expect that InstructCV-RP will be able to extrapolate its learning to novel user-generated instructions at inference time. To test this, we compared the performance of the InstructCV-FP and Instruct-RP variants on the semantic segmentation task. Both models were tested using manually-selected prompts (unseen in training data) on 200 images in the ADE20k test data. The results in Table 3 indicate that the InstructCV-RP model exhibits consistent performance and more robustness to variations in the phrasing of user instructions compared to InstructCV-FP. For example, testing InstructCV-FP with the instruction “Please highlight the image segment containing %category%.” instead of the template prompt “Segment %category%.” led to a drop in mIOU. Conversely, InstructCV-RP only incurred a performance reduction of with this instruction compared to the template prompt. This suggests that our LLM-based prompt rephrasing approach effectively enhances the ability of InstructCV to generalize to new user-generated prompts that convey descriptions of tasks similar to those seen during training.
Computational costs. Finally, we note that InstructCV was trained in an end-to-end fashion, with only 2,000 finetuning steps. The inference time of InstructCV on a single NVIDIA A100 GPU is 5 seconds (for a 256x256 image). This is a significant improvement over comparable generalist models, such as Unified-IO , which was trained from scratch using 1.5 million steps and takes around 40 seconds for inference on a single NVIDIA A100 GPU. Notably, InstructCV not only simplifies the training process but also outperforms Unified-IO in various tasks. These improvements can be attributed to the already impressive capabilities of the underlying text-to-image model. By instruction-tuning a generative model with a relatively small number of steps and a moderately-sized dataset, we are able to steer its functionality with performance that is competitive with bespoke generalist models.
Conclusion
In this paper, we introduce a unified language interface for computer vision tasks, dubbed InstructCV, eliminating the need for task-specific design choices and allowing for task execution based on natural language instructions. InstructCV frames various computer vision tasks as text-to-image generation problems. In this setup, textual instructions describe the task, and the resulting image serves as a visual representation of the task output. Following the InstructPix2Pix architecture, we curate a multi-task and multi-modal dataset to instruction-tune a pre-trained text-to-image diffusion model, steering its function from a generative model to an instruction-guided multi-task vision learner. By harnessing semantically meaningful language instructions to drive the learning process, our model demonstrates compelling generalization capabilities across unseen data, categories, and user instructions.
Limitations. While InstructCV improves upon the computational cost of existing generalist models, the inference speed of our model lags behind specialized task-specific models and falls short of meeting the real-time inference requirements for tasks such as object detection and segmentation. Additionally, the Pix2Pix formulation can lead to inadmissible output images for a given vision task, e.g., failure to generate a box-shaped output for the object detection task. Furthermore, the semantic flexibility of InstructCV is constrained by the richness and diversity of our instruction-tuning dataset, which is currently generated by rephrasing a limited set of template prompts. This raises questions for future work: can this learning paradigm accommodate instructions that introduce more nuanced conditions? For example, an instruction might cap the count of objects to be detected. Exploring such ideas might require the integration of strategies such as learning from human feedback, which could enable more versatile generalist models by improving alignment of task outputs with more complex prompts.
Acknowledgements
The authors would like to thank David Sontag (MIT) for insightful feedback and discussions.
References
Appendix A Appendix
We train our multi-task vision model across 20 epochs for 10 hours on an array of 8 80GB NVIDIA A100 GPUs. Our training utilizes images of 256 × 256 resolution and a batch size of 128. Augmentation techniques applied include random horizontal flipping and crop augmentation. For the latter, images are first subjected to random resizing between 256 and 288 pixels before being cropped to a 256-pixel size. We set our learning rate at 10e-4 without incorporating a learning rate warm-up phase. Model initialization is performed using the EMA weights from the Stable Diffusion v1.5 checkpoint, and we adopt other training settings from the public Stable Diffusion repository. For the results presented in this paper, we operate at a 256-pixel resolution with 100 denoising steps. We employ an Euler ancestral sampler with a denoising variance schedule as proposed by . On an NVIDIA A100 GPU, our model takes approximately 10 seconds to solve a given vision task.
A.2 Examples of the rephrased prompts.
Figure 6-(a)/(b) qualitatively compares the robustness of InstructCV to changes in instruction wording when trained using the fixed prompt or the more diverse rephrased prompt dataset. The model trained using the fixed prompt data shows reasonable performance on segmentation tasks, even for prompts that slightly deviate from the standard instruction wording. However, as these deviations increase misclassfications become more common. Instead, the model trained on the rephrased prompt data appears more robust to such changes in task formulation. Notably, the model appears to show basic semantic understanding as it is in the last prompt example able to infer the correct intent despite the simultaneous occurrence of the potential object detection targets ’spider man’ and ’face’.
A.3 Post-processing steps for object detection tasks
We employ image processing techniques to derive the bounding box coordinates from the output image. First, we apply median and bilateral filters to the image in order to mitigate noise and enhancing features. Following, we convert the image from RGB to HSV space to isolate the red region within the target object, which is then extracted from the original image. Next, we identify closed contours in the image by converting it to a grayscale map, performing threshold segmentation based on grayscale, and ultimately, binarizing the image. However, naively applying this approach could potentially result in the removal of some accurate bounding boxes due to the disruptions induced by complex image backgrounds. To circumvent this issue, we cross-reference the bounding boxes with the dataset annotations and retain any bounding box predicted by the model that exhibits an Intersection over Union (IOU) greater than 0.5. We exclude disturbances that are not rectangular or that contain numerous red dots within the contour. The coordinates of the remaining contours are subsequently added to the prediction list.