Images Speak in Images: A Generalist Painter for In-Context Visual Learning

Xinlong Wang, Wen Wang, Yue Cao, Chunhua Shen, Tiejun Huang

Introduction

Training one generalist model that can execute diverse tasks simultaneously, and even can perform a new task given a prompt and very few examples, is one important step closer to artificial general intelligence. In NLP, the emergence of in-context learning presents a new path in this direction, which uses language sequences as the general interface and allows the model to rapidly adapt to various language-centric tasks with only a handful of prompts and examples.

Thus far, in computer vision, in-context learning is rarely explored and remains unclear how to achieve that. This can be attributed to the differences between two modalities. One difference is that NLP tasks mainly consist of language understanding and generation, so their output spaces can be unified as sequences of discrete language tokens. But vision tasks are the abstractions of raw visual input to varied granularity and angles. Thus, vision tasks vary significantly in output representations, leading to various task-specific loss functions and architecture designs. The second difference is that the output space of NLP tasks is even the same as the input. Thus, the task instruction and the example’s input/output, which are all language-token sequences, can be directly used as the input condition (also denoted as the task prompt), which can be processed straightforwardly by the large language model. However, in computer vision, it is unclear how to define general-purpose task prompts or instructions that the vision model can understand and transfer to out-of-domain tasks. Several recent attempts tackle these difficulties by following the solutions in NLP. They more-or-less convert vision problems into NLP ones via discretizing the continuous output spaces of vision tasks, and using the language or specially-designed discrete tokens as the task prompts.

However, we believe that images speak in images, i.e., image itself is a natural interface for general-purpose visual perception. In this work, we address the above obstacles with a vision-centric solution. The core observation is that most dense-prediction vision problems can be formulated as image inpainting, i.e.:

Given an input image, prediction is to inpaint the desired but missing output “image”.

Thus, we need a representation of 3-channel tensor that appears as an “image” for the output of the vision tasks, and specify the task prompts using a pair of images. Here we showcase several representative vision tasks for training, including depth estimation, human keypoint detection, semantic segmentation, instance segmentation, image denoising, image deraining, and image enhancement, and unify their output spaces using a 3-channel tensor, a.k.a. “output image”. We carefully design the data format for each task, such as instance mask and per-pixel discrete labels of panoptic segmentation, per-pixel continuous values of depth estimation, and high-precision coordinates of pose estimation. Including more tasks is very straightforward, as we only need to construct new data pairs and add them to the training set, without modifications to either the model architecture or loss function.

Based on this unification, we train a generalist Painter model with an extremely simple training process. During training, we stitch two images from the same task into a larger image, and so do their corresponding output images. Then we apply masked image modeling (MIM) on pixels of the output image, with the input image being the condition. With such a learning process, we enable the model to perform tasks conditioned on visible image patches, that is, the capability of in-context prediction with the visual signal as context.

Thus the trained model is capable of the in-context inference. That is, we directly use the input/output paired images from the same task as the input condition to indicate which task to perform. Examples of in-context inference are illustrated in Figure LABEL:fig:teaser, consisting of seven in-domain examples (seven rows at top) and three out-of-domain examples (three rows at bottom). This definition of task prompts does not require deep understanding of language instructions as need by almost all previous approaches, and makes it very flexible for performing both in-domain and out-of-domain vision tasks.

Without bells and whistles, our model can achieve competitive performance compared to well-established task-specific models, on several fundamental vision tasks across high-level visual understanding to low-level image processing, namely, depth estimation on NYUv2 , semantic segmentation on ADE-20K , human keypoint detection on COCO , panoptic segmentation on COCO , and three low-level image restoration tasks. Notably, on depth estimation of NYUv2, our model achieves state-of-the-art performance, outperforming previous best results by large margins which have heavy and specialized designs on architectures and loss functions. Compared to other generalist models, Painter yields significant improvements on several challenging tasks.

Related Work

The emergence of Transformer provides the possibility to share the basic modeling module across different modalities. Until now, Transformers are widely-adopted in language , vision , speech and multimodal domains. Perceiver and Perceiver-IO are the first attempts to use the exact same Transformer architecture in different domains, such as natural language and visual understanding processing, StarCraft II, and multi-modal domains. If the input could be transformed to a sequence of tokens, one can adopt Transformer for modeling the relationships between different tokens.

Vision Generalist

Due to the general modeling capability of Transformer, there are some efforts to unify different tasks in vision domains, resulting in several vision generalists . DETR first adopted Transformer as the task specific head for object detection. Based on this, Pix2Seq defined the output space of object detection as a discrete space, and conduct this task in an auto-regressive manner. Due to the fundamental nature of object detection, Pix2Seq provides a direction for unifying different vision tasks using discrete spaces, thus motivating a lot of following work. Unified-IO and OFA both homogenize the diverse inputs and outputs to a sequence of discrete tokens, perform joint modeling in a sequence-to-sequence manner over vision, vision & language and NLP tasks, and use T5-style architectures with billions of parameters, where Unified-IO unifies more tasks than OFA with larger size of models. Pix2Seq v2 unified object detection, instance segmentation, keypoint estimation and image captioning in the same defined discrete spaces as Pix2Seq. UViM unified pixel-labeling tasks with the same modeling approach but trained separate models for different tasks, such as panoptic segmentation, depth estimation and colorization.

Notably, from our point of view, the input of visual signals is continuous in nature, thus we try to make the output space of several representative vision tasks as continuous as images to reduce the quantization error caused by discretization and further enable the in-context visual learning with masked image modeling.

In-Context Learning

For the first time, GPT-3 defined a new learning paradigm, in-context learning, where a series of NLP tasks can be formulated as the text completion task given prompts and examples. In-context learning grants models new capabilities to perform on-the-fly computational reasoning or novel-pattern recognition that is unlikely to have occurred in training. Flamingo extended the input of the large language models to not only texts but also images and videos, but still used languages as the general interface, such that the model can perform many visual-linguistic tasks given prompts and examples, such as image captioning, visual question answering, optical character recognition (OCR), etc. In other domains, it appears non-trivial to directly introduce the in-context learning capability. AD uses algorithm distillation to combine in-context capability with reinforcement learning. In computer vision, a concurrent work performs inpainting on the figures and infographics from vision articles, but only works for predicting on discrete space as the language domain, and shows the in-context capability on foreground segmentation, single object detection and colorization. While the work proves the concept of in-context learning for vision tasks, no results were reported on standard benchmark datasets. Thus, it remains unclear how the method performs on real-world datasets. In contrast, our model works well with masked image modeling on pixels on seven diverse and challenging vision tasks, including depth estimation, keypoint estimation, semantic segmentation, panoptic segmentation, image denoising, image deraining, and image enhancement, and also shows highly competitive performance on these tasks.

Approach

The core idea of our framework is to reformulate most vision tasks such as depth estimation, semantic segmentation, instance segmentation, keypoint detection and image restoration as an image inpainting problem. To do so, we redefine the output space of those tasks as “images”.

We denote an input image as x\bf{x} with the size of H×W×3H\times W\times 3, and standard definition of the task ground truth as yt{\bf y}^{t} which has various sizes for different task tt, and we redefine these task outputs still in the image space, denoting as y^t{\hat{\bf y}}^{t} with the size of H×W×3{H}\times{W}\times 3. Our philosophy is to keep the spatial relationships between pixels to be intact, and each pixel of the output image still represents the output for this task of the corresponding input image pixel but in the RGB space. That is, y^i,jt{\hat{\bf y}}^{t}_{i,j} with three dimensions denotes the corresponding ground truth of input pixel xi,j{\bf x}_{i,j}. We select seven representative vision tasks with diverse types of outputs, such as depth estimation, semantic segmentation, keypoint detection and panoptic segmentation. Here we show how we redefine the per pixel ground-truth for each task as a 3-channel tensor, similar to the RGB space. Note that in theory, a fixed number of output channels can serve our purpose. We choose 3 channels to make an image.

is a dense prediction task with weak semantics, to estimate the per-pixel depth value (distance relative to the camera) given an input RGB image. For NYUv2 , the per-pixel ground-truth depth yi,jt{\bf y}^{t}_{i,j} is a real value in the range of meters.Herewemaptheground−truthvaluefromreal−valuedrangemeters. Here we map the ground-truth value from real-valued range to the integer space with range $,,{\hat{\bf y}}^{t}_{i,j,0}=\lfloor{\bf y}^{t}_{i,j}\times\frac{255}{10}\rfloor,andletthethreechannels, and let the three channels{\hat{\bf y}}^{t}_{i,j,0},,{\hat{\bf y}}^{t}_{i,j,1}andand{\hat{\bf y}}^{t}_{i,j,2}bethesamegroundtruth.Ininference,wedirectlyaveragetheoutputsofthethreechannelsandthenperformtheinverselineartransformationofthetrainingtoobtainadepthestimateintherangeofbe the same ground truth. In inference, we directly average the outputs of the three channels and then perform the inverse linear transformation of the training to obtain a depth estimate in the range of$.

Semantic segmentation

is a dense prediction task with strong semantics, to predict the per-pixel semantic label given an input image. Given a semantic segmentation task with LL categories, we let the RGB space to represent these LL categories with the same margin in each space. We represent LL as a 3-digit number with bb-base system, where b=⌈L13⌉b={\lceil L^{\frac{1}{3}}\rceil}, and y^i,j,0t{\hat{\bf y}}^{t}_{i,j,0}, y^i,j,1t{\hat{\bf y}}^{t}_{i,j,1} and y^i,j,2t{\hat{\bf y}}^{t}_{i,j,2} represent their hundreds, tens and ones places, with a margin defined as m=⌊256b⌋m=\lfloor\frac{256}{b}\rfloor. For example, ADE-20K has 150 semantic categories with one background class, thus we set the base as b=6b=6, and the margin as m=42m=42. The output channels are defined as y^i,j,0t=⌊lb2⌋×m{\hat{\bf y}}^{t}_{i,j,0}=\lfloor\frac{l}{b^{2}}\rfloor\times m, y^i,j,1t=⌊lb⌋mod  b×m{\hat{\bf y}}^{t}_{i,j,1}=\lfloor\frac{l}{b}\rfloor\mod b\times m, and y^i,j,2t=lmod  b×m{\hat{\bf y}}^{t}_{i,j,2}=l\mod b\times m, where ll denotes the corresponding category and is an integer value in range of [0,L)[0,L). In inference, we discretize the output of each pixel with the margin mm and obtain its corresponding category.

Keypoint detection

is a fine-grained localization task to simultaneously detect the objects (e.g., human) and localize their detailed keypoints. We follow recent heatmap-based top-down pipeline , thus it is defined as a 17-category point-localization task for human keypoint detection. For each keypoint, we need to localize it to its corresponding pixel, which is very fine-grained. We decouple this task into a combination of 17-category keypoint classification using two channels of y^i,j,1t{\hat{\bf y}}^{t}_{i,j,1} and y^i,j,2t{\hat{\bf y}}^{t}_{i,j,2}, and class-agnostic keypoint localization using another channel of y^i,j,0t{\hat{\bf y}}^{t}_{i,j,0}. For the 17-category classification task, we define each keypoint as a 9×\times9 pixel square, and define the color of each square using the approach in semantic segmentation. For the class-agnostic point localization task, we define 17 squares, each of which is a 17×\times17 heatmap with Gaussian distribution that the center pixel with largest value of 255 is the position of the ground truth keypoint. In inference, we obtain the category and location for each keypoint as the final results from the 3-channel output image.

Panoptic segmentation

is a combination of semantic segmentation task and instance segmentation task. Thus, we perform these two tasks separately for ease of redefinition of output space and optimization, and then combine their results to obtain the results for panoptic segmentation.

Here we introduce the redefinition of output space of class-agnostic instance segmentation. For this task, we directly change the color of each instance mask in the image to the same one, thus different instances use different colors. In theory, we can randomly choose colors for different instances, but we find that this setting would make the model hard to optimize. To address this optimization issue, we follow SOLO to assign the color of each instance mask according to the absolute position of its center in the image. We conceptually divide the image into 16×\times20×\times20 blocks, corresponding to three channels respectively. We assign a fixed color to each block, then color the mask accordingly if its center locates in that block. In inference, we adopt each color as a kernel to compute the distance with each pixel in the image, and then set a threshold to get the final masks. To get the category for each instance mask, we directly assign the majority category in the semantic segmentation result within each instance mask as its category.

Image restoration

takes the corrupted image as input and outputs the corresponding clean image. In this paper, we investigate three representative image restoration tasks, including image denoising, image deraining, and low-light image enhancement. Since both the input and output are inherently defined in the RGB space, these tasks can be seamlessly unified in our Painter model without any transformation.

2 A Masked Image Modeling Framework

Based on the redefinition of the output spaces of above representative vision tasks, the input and output of these tasks are all images. Thus, here we directly apply a standard masked image modeling (MIM) pipeline for training, illustrated in Figure 1. This framework consists of three major components: input format, architecture and loss function.

During training, each input sample is the concatenation of two pairs of images from the same task, which have already been applied data augmentations separately, as shown in the left part of Figure 1. Each image pair consists of one image and its corresponding task output, which is also redefined as an image. We randomly mask the task output image and train the model to reconstruct the missing pixels. For the masked area, we follow the NLP community and previous works to use a learnable token vector to replace each masked patch. We adopt the block-wise masking strategy and find the masking ratio as 75% to work well.

Architecture

We adopt a vanilla vision Transformer (ViT) as the encoder, which consists of stacked Transformer blocks, e.g., 24 blocks for ViT-large. We concatenate 4 feature maps evenly sampled from these blocks and use a simple three-layer head to map the features of each patch to its original resolution, e.g., 16×16×316\times 16\times 3. Specifically, the head consists of a linear (1×\times1 convolution) layer, a 3×\times3 convolution layer, and another linear layer.

Since each input sample includes both the input image and its output image, the input resolution can be larger than the traditional training process, and also the computation cost. To address this problem, we propose to reduce the computation cost by merging the early features of the input image and the output image. Specifically, we feed the input image and the output image to the model in parallel, then add their features patch by patch after a few blocks, e.g., 3 blocks by default. This design saves nearly half of the computation costs, but we find no degradation in performance.

Loss Function

3 In-Context Inference

At the first time, we design an in-context inference procedure, which is very flexible for performing both in-domain and out-of-domain vision tasks. As the input and output spaces of vision tasks have been unified as images, we can directly use the input/output paired images from the same task as the input condition (task prompts) to indicate which task to perform, and concatenate them with the input image and a masked image for completing the corresponding task. Examples are shown in Figure LABEL:fig:teaser. This definition of task prompt does not require deep understanding of language instructions like previous approaches, but uses the visual signal as context which can be understood by the vision model and well matches the nature of the visual domain.

Also, different task prompts would lead to different results. Thus, how to select or generate a more suitable task prompt can be a new direction to explore. Here we present two simple baselines and we leave more explorations as the future work. The first baseline is to obtain a better prompt via selection, that we traverse the whole training set in a heuristic manner and select the best-performing example pair for each task. The second baseline is to generate a task prompt. We define the task prompt as the learnable tensors, freeze the whole model, and then use the training loss to optimize the task prompts. We compare these two solutions with the random counterpart in the experiments with the visualizations, as shown in §4.3.

Experiments

NYUv2 dataset consists of 464 indoor scenes captured by a Microsoft Kinect camera. The official training split (24K images) is used for training, and we report the Root Mean Square Error (RMSE), absolute mean relative error (A.Rel) and the percentage of inside pixels with different thresholds of δ\delta on the 654 testing images from 215 indoor scenes.

ADE20K is a widely-used semantic segmentation dataset, covering a broad range of 150 semantic categories. It has 25K images in total, with 20K for training, 2K for validation, and another 3K for testing. We adopt the widely-used metric of mean IoU (mIoU) for evaluation.

MS-COCO contains approximately 118K training images and 5K validation images used for evaluation, with 80 “things” and 53 “stuff” categories. Panoptic segmentation task is evaluated on the union of “things” and “stuff” categories. During training, we generate the output images of semantic segmentation with 133 categories, and the class-agnostic instance segmentation with only “things” categories. During inference, we perform inference twice for each input image to obtain the results of semantic segmentation and class-agnostic instance segmentation, respectively, then merge them together to get the results of panoptic segmentation. We report panoptic quality as the measure.

For human keypoint detection, we use the standard person detection results from Simple Baseline , which use the standard splits on COCO with 15K training samples and validation set for evaluation, and report the AP based on OKS as the evaluation metric .

Image restoration tasks are evaluated on several popular benchmarks, including SIDD for image denoising, LoL for low-light image enhancement, and the merged deraining dataset for deraining. More details about these datasets are provided in Table S1.

Training details

During training, we employ an AdamW optimizer with a cosine learning rate scheduler, and train for 54K iterations. The training hyper-parameters are: the batch size as 2048, base learning rate as 1ee−3-3, weight decay as 0.05, β1=0.9\beta_{1}=0.9, β2=0.999\beta_{2}=0.999, drop path ratio as 0.1, warm-up for 5.4K iterations. A light data augmentation strategy is used: random resize cropping with scale range of [0.3,1][0.3,1] and a aspect ratio range of [\nicefrac34,\nicefrac43][\nicefrac{{3}}{{4}},\nicefrac{{4}}{{3}}], followed by a random flipping. The input image size is 448×448448\times 448 by default. The sampling weight for each task is 0.10.1 (NYUv2 depth estimation), 0.20.2 (ADE-20K semantic segmentation), 0.150.15 (COCO class-agnostic instance segmentation), 0.250.25 (COCO semantic segmentation), 0.20.2 (COCO human keypoint detection), 0.150.15 (image denoising), 0.050.05 image deraining, and 0.050.05 (low-light image enhancement).

2 Results

With the corresponding task prompts, we compare our approach, Painter, with recent best vision generalist models and specialized models on seven representative tasks, shown in Table 1. Without task-specific design, our Painter sets new records on NYUv2 depth estimation, outperforming not only vision generalists, such as Unified-IO and UViM , but also the state-of-the-art specialized model, e.g., BinsFormer . For COCO keypoint detection, Painter significantly outperforms the generalist model Pix2Seq v2 by 7.3 AP. For ADE-20K semantic segmentation, COCO panoptic segmentation, and the three low-level image processing tasks, our model achieves comparable performance to those well-designed task-specific models. There is still much room for boosting our approach compared to other well-designed specialized models. For example, our default input image size is 448×448448\times 448 while those specialized panoptic segmentation models use a much larger resolution, e.g., 1024×10241024\times 1024 in Mask2Former . Achieving state-of-the-art performance on every task is not the major goal of this paper, and we leave this as future work.

Joint training vs. separate training

We compare two training settings, i.e., joint training and separate training, in Table 2. For the separate training setting, we train each model separately on each task with the same number of iterations and architectures as joint training on this task. We can observe that models with joint training generally outperform that with separate training on most of the tasks, indicating that our in-context training approach with the unification of output spaces can somewhat benefit them from each other. But conflict may still exist, e.g., joint training performs slightly worse on keypoint detection. Exploring the relationships between tasks in our simple and unified framework could be an interesting and promising direction.

Qualitative results

To demonstrate the capability of our generalist model in an intuitive perspective, we visualize the task output of the selected images from the validation set of several in-domain tasks, such as semantic segmentation, depth estimation, instance segmentation, human keypoint detection, image denoising, image deraining, and low-light image enhancement. As shown in Figure 2, Painter can make very accurate predictions on all these tasks.

3 Prompt Tuning

Different task prompts would lead to varied results. Here we compare three simple baselines, ‘random’, ‘searched’ and ‘learned’ prompts. Random prompts denote to randomly select one example from that task as the prompt. Searched prompt denotes to traverse the training set in a heuristic manner and select the best-performing example pair for each task. Learned prompt denotes that we define the task prompt as the learnable tensors, freeze the whole model, and then use the training loss to optimize the task prompts.

The qualitative comparisons are shown in Table 3 . We can see that the model using the searched and learned prompts perform better than that with random ones, which indicates that the optimization process helps the task to find a more suitable prompt. Also, the model with random prompts works relatively well, indicating the robustness of our model. In Figure S6, we show the visualizations of the learned prompt images for different tasks.

It is very promising that the performance of a task can be improved by optimizing the input prompt. This provides the possibility that if a category or task does not appear in the training data, we can still achieve relatively good performance on this task by optimizing the prompts, without tuning the parameters of the model, which is also one core advantage of in-context learning.

4 Generalization

The core property of in-context learning is that it allows the model to rapidly adapt to various tasks with only a handful of prompts and examples. Here, we explore this capability via visualizations, shown in Figure LABEL:fig:teaser, Figure 3, and Figure S5. From these visualizations, we find that our model can perform the task seen in training but with input images that their categories are unseen in training, such as open-vocabulary keypoint detection (e.g., monkey, horse and broom), object segmentation (koala) and instance segmentation (tablets). In addition, we quantitatively evaluate Painter on a few-shot segmentation benchmark which requires to segment objects of 1k novel classes. Painter largely outperforms a concurrent work , as reported in Table 4,

Discussion and Conclusion

In this work, for the first time, we explore how to perform in-context visual learning and present a vision-centric solution in which the context is defined as visual signals. Thus, the models are granted with the capability to perform tasks conditioned on the paired data from that task, that is, in-context inference. Our approach achieves highly competitive performance on a representative and diverse set of seven tasks.

This work is not without drawbacks. First, there is still much room for boosting our approach especially on the difficult task of panoptic segmentation, comparing to the specialized models. In addition, since our approach is designed based on visual signals as contexts, this general interface does not seem natural for modeling language signals. How to model discrete language signals as continuous ones seems to be an impressive direction, and some work has started to emerge recently.

While there are previous approaches that hope to use general interfaces to solve multiple vision tasks, we may be the first to grant models the ability to learn and complete tasks in context, which does have the opportunity to handle out-of-domain tasks. We hope this work will draw attention to this promising direction, and we believe that the best GPT-3 moment in the vision field is yet to come.

Acknowledgement

This project is supported by the National Key R&D Program of China (2020AAA0105200). We would like to thank Hanxiao Qu, Yemin Shi, Yan Tian, and Xigang Cao for their help on the GPU resources, Xinxin Liu and Kai Lu for beautifying the figures, as well as other colleagues at Beijing Academy of Artificial Intelligence for support throughout this project.

References

Appendix

Appendix A Additional Implementation Details

In this section, we provide additional details of the data preparation and pose-processing for different tasks. PyTorch-style pseudo-code is provided to better illustrate the implementation details. They can be done with a few operations in implementation. For each task, Painter doesn’t involve more complex post-processing compared to the specialist methods.

As described in Section 3.1 of the main paper, we formulate different semantic categories using different colors in RGB space. To this end, we define the background and ignore areas as black, i.e., pixels in color (0, 0, 0), and generate the colors for foreground categories using the pseudo-code elaborated in Figure S1.

During inference, to decode the output image to a single-channel ID map where each pixel represents a class ID, we compute the L1 distance between each output pixel and the pre-defined colors for each semantic category, and take the ID of the closest color as the predicted category. The pseudo-code for the post-processing is illustrated in Figure S2.

Keypoint detection

For keypoint detection, the output image consists of the R channel which denotes the class-agnostic heatmaps and the G/B channels that represent the keypoint categories. As illustrated in Figure S3, we convert the output image to a 17-channel heatmap, and follow the commonly used post-processing to obtain the final keypoint locations.

Panoptic segmentation

As described in Section 3.2, we decompose the panoptic segmentation task into semantic segmentation and class-agnostic instance segmentation. During training, the semantic segmentation sub-task uses the same setting as the semantic segmentation on ADE-20K , except that we set the base b=7b=7 when assigning colors. Similar to the color generation process of semantic segmentation, we generate colors for each location category used in class-agnostic instance segmentation. The color of each instance mask is determined by the location of its center.

During inference, the semantic ID map can be obtained using the post-processing described in Figure S2, while the class-agnostic instance masks are generated by thresholding the distance between predicted colors and the pre-defined colors for location categories. Matrix NMS is adopted to remove duplicate instance predictions. We apply the majority vote of pixels from the semantic prediction to get the semantic class for each instance mask, as illustrated in Figure S4. Finally, we follow Panoptic FPN to merge the semantic segmentation and the instance segmentation predictions to obtain the panoptic segmentation results.

Image restoration

The detailed statistics of the datasets that are used for image restoration are shown in Table S1.

Appendix B Additional Results

We report the results of ablation experiments on several components of our framework, with a shorter schedule of 3k iterations and other hyper-parameters unchanged on semantic segmentation of ADE-20K.

During training, each input sample consists of both the input image and output image, which results in high memory cost and significantly slows down the training process. We reduce nearly half of the computation costs by merging the early features of the input image and the output image, i.e., adding their features patch by patch after a three blocks. Table LABEL:tab:ablation-merge shows that this new design even incurs performance increase. We argue that this design further provides the pixel-to-pixel correspondence between the input and its output via stacking them together. But in the original setting, these relationships need to be learned by the model, which will make the optimization more difficult especially in a short schedule.

Encoder

We adopt standard Vision Transformer (ViT) with different model sizes as the encoder, including ViT-base and ViT-large. Results are shown in Table LABEL:tab:ablation-encoder. We find that the model with ViT-L outperforms that with ViT-B by very large margins. This observation is intuitive, that generally larger models yield better performance. For generalist models, they can use more data but with less task-specific prior on method design, thus may require more model capacity than task-specific models.

Head

We use a light three-layer head that consists of a linear (1×\times1 convolution) layer, a 3×\times3 convolution layer, and another linear layer, to map the feature of each patch to its original resolution, e.g., 16×16×316\times 16\times 3. The feature of each patch is the concatenation of the 4 feature maps evenly sampled from the transformer blocks. As shown in Table LABEL:tab:ablation-decoder, the light head achieves clear gains over the baseline with only a linear layer.

Loss function

Appendix C Additional Visualization

In this section, we provide more visualizations. As shown in Figure S5, Painter performs in-context inference according to different prompt images. Note that Painter is never trained to solve these tasks during training, e.g., keypoint detection of potato, object segmentation of bee, and instance segmentation of tomato.