Generalized Decoding for Pixel, Image, and Language
Xueyan Zou, Zi-Yi Dou, Jianwei Yang, Zhe Gan, Linjie Li, Chunyuan Li, Xiyang Dai, Harkirat Behl, Jianfeng Wang, Lu Yuan, Nanyun Peng, Lijuan Wang, Yong Jae Lee, Jianfeng Gao
Introduction
Visual understanding at different levels of granularity has been a longstanding problem in the vision community. The tasks span from image-level tasks (e.g., image classification , image-text retrieval, image captioning , and visual question answering (VQA) ), region-level localization tasks (e.g., object detection and phrase grounding ), to pixel-level grouping tasks (e.g., image instance/semantic/panoptic segmentation ). Until recently, most of these tasks have been separately tackled with specialized model designs, preventing the synergy of tasks across different granularities from being exploited. In light of the versatility of transformers , we are now witnessing a growing interest in building general-purpose models that can learn from and be applied to a diverse set of vision and vision-language tasks, through multi-task learning , sequential decoding , or unified learning strategy . While these works have shown encouraging cross-task generalization capabilities, most target the unification of image-level and region-level tasks, leaving the important pixel-level understanding underexplored. In , the authors attempt to unify segmentation into a decoding of a coordinate sequence or a color map, which, however, produces suboptimal performance and limited support for open-world generalization.
Arguably, understanding images down to the pixel level is one of the most important yet challenging problems in that: (1) pixel-level annotations are costly and undoubtedly much more scarce compared to other types of annotations; (2) grouping every pixel and recognizing them in an open-vocabulary manner is less studied; and (3) more importantly, it is non-trivial to learn from data at two substantially different granularities while also obtaining mutual benefits. Some recent efforts have attempted to bridge this gap from different aspects. In , Chen et al. propose a unified architecture Mask2Former that tackles all three types of segmentation tasks but in a closed set. To support open vocabulary recognition, a number of works study how to transfer or distill rich semantic knowledge from image-level vision-language foundation models such as CLIP and ALIGN to specialist models . However, all these initial explorations focus on specific segmentation tasks of interest and do not show generalization to tasks at different granularities. In this work, we take one step further to build a generalized decoder called X-DecoderHere, ‘X’ denotes versatile, and also represents ‘piXel’. towards the unification of pixel-level and image-level vision-language understanding, as shown in Figure 1.
A generalized decoding framework. We formulate all tasks including pixel-level image segmentation, image-level retrieval and vision-language tasks into a generic decoding procedure. Specifically, X-Decoder is built on top of a vision backbone and a transformer encoder for extracting multi-scale image features, following the framework of Mask2Former . The key novelty lies in the decoder design. First, it takes two sets of queries as input: () generic non-semantic queries that aim to decode segmentation masks for universal segmentation, similar to Mask2Former , and () newly introduced textual queries to make the decoder language-aware for a diverse set of language-related vision tasks. Second, it predicts two types of outputs: pixel-level masks and token-level semantics, and their different combinations can seamlessly support all tasks of interest. Third, we use a single text encoder to encode the textual corpus involved in all tasks, including concepts in segmentation, phrases in referring segmentation, tokens in image captioning and questions in VQA, etc. As a result, our X-Decoder can naturally facilitate the synergy across tasks and advocate the learning of a shared visual-semantic space, while respecting the heterogeneous nature of different tasks.
An end-to-end learning paradigm. With our generalized decoder design, we propose an end-to-end pretraining method to learn from all granularities of supervision. We unite three types of data: panoptic segmentation, referring segmentation, and image-text pairs. Unlike previous works that use pseudo-labeling techniques to extract fine-grained supervision from image-text pairs , X-Decoder directly groups and proposes a few meaningful segmentation candidates, so that it can map the regions easily to the contents described in the captions on the fly. Meanwhile, the referring segmentation task bridges generic segmentation and image captioning by sharing the pixel-level decoding with the former and semantic queries with the latter.
Strong zero-shot and task-specific transferability to a wide range of segmentation and VL tasks. Pre-trained with a limited amount of segmentation data and millions of image-text pairs, our X-Decoder supports a diversity of tasks in a zero-shot and open-vocabulary manner. Concretely, our model can be directly applied for all three types of segmentation tasks in a wide range of domains, establishing new state-of-the-art on ten settings of seven datasets. When transferred to specific tasks, our model also exhibits consistent superiority to previous works. Finally, we observe some intriguing properties in our model that it can support some novel task compositions and efficient finetuning, thanks to the flexibility endowed by our model design.
From Specialist to Generalist Models
Pixel-level image understanding, also known as image segmentation, has been a long-standing problem .
Generic Segmentation. There are mainly three well-defined tasks for pixel-level understanding, including semantic , instance , and panoptic segmentation. Semantic segmentation cares about the per-pixel semantic within an image , whereas instance segmentation groups pixels of the same semantic meaning into object instances. Models for both tasks have evolved from CNN-based architectures to transformer-based ones , and from two-stage models to one-stage models and to the recent query-based approaches . With the capability of per-pixel and instance-level understanding, a natural step was taken to formulate panoptic segmentation . Most recently, Mask2Former proposed to address all three tasks with a unified encoder-decoder architecture. Nevertheless, all these works cope with a limited number of categories, i.e., models can hardly recognize concepts absent in the training set. In MSeg , the authors manually merge different datasets and train a more generalized model on the composite set, which is still limited to being a closed set.
Open-Vocabulary Segmentation. Recently, a number of works opt to transfer or distill the rich visual-semantic knowledge from foundation models like CLIP and ALIGN to specific segmentation tasks. Prominent examples include LSeg , OpenSeg , and . Instead of using existing models, GroupViT performed language-image pretraining from scratch with a bottom-up grouping ViT , while DenseCLIP demonstrated the superiority of foundation models in finetuning settings compared with supervised models. Recently, MaskCLIP proposed to tackle open-vocabulary panoptic and semantic segmentation by leveraging CLIP, and achieved SoTA performance on ADE20K and PASCAL .
Referring Segmentation by nature is open-vocabulary in that it does not presume a fixed number of phrases in the training and inference times. Models are usually designed specifically to learn from target datasets using various multimodal fusion strategies . Since the emergence of vision transformers, works like LAVT enhance the cross-modal interactions from the very beginning, which led to SoTA on RefCOCO , RefCOCO+ and G-Ref . CLIPSeg extended the textual query to a visual query and showed superior performance not only on referring segmentation but also on semantic segmentation.
In this work, we propose X-Decoder, which is the first model to tackle generic and referring segmentation tasks all in one model. Furthermore, the generalized decoder jointly learns from segmentation data and image-text pairs end-to-end, and thus can augment the synergy across tasks for rich pixel-level and image-level understanding.
2 Vision-Language Understanding
Vision-language (VL) pretraining has proven to be effective for various VL tasks . The field has evolved from a transformer fusion model with pre-extracted object features to end-to-end transformers , that directly learn from raw image pixels. Recently, researchers have found that image-text data at scale can be helpful for visual representation learning (e.g., enabling zero-shot image classification and action recognition ). VL pre-trained models can be further extended to region-level tasks, such as phrase grounding and open-vocabulary object detection , and unified frameworks that aim to combine image-text pairs with region-level data have also been proposed . A comprehensive review on this topic is provided in .
We are witnessing a clear trend from building specialist models to generalist ones. Early efforts build a multi-task learning paradigm to accommodate a diversity of tasks. However, the interactions among different tasks in these works are less studied, and the combination usually leads to performance degradation compared with specialist models. Recently, a number of works aim to reformulate the tasks into a unified sequential decoding process . In this work, instead of developing a unified interface for vision and VL tasks, our X-Decoder builds a generalized decoding paradigm that can seamlessly connect the tasks by taking the common (e.g., semantic) but respecting the natural differences (e.g., spatial mask v.s. sequential language), leading to significant improvements for different segmentation and VL tasks across the board.
X-Decoder
Our model follows the generic design of encoder-decoder architecture as shown in Fig. 2. Given an input image , we first use an image encoder to extract features . Afterwards, we use the text encoder to encode a textual query into of length . The visual features, textual queries and the non-semantic or latent queries are fed to our X-Decoder to predict the outputs:
where and are the pixel-level masks and token-level semantics, respectively. In the above formula, we note three critical designs to empower the generalization ability of our X-Decoder to a variety of vision and vision-language tasks.
We define two types of queries and outputs for X-Decoder. As discussed earlier, the queries for the decoder are categorized into latent queries and text queries , which undertake generic vision and vision-language tasks, respectively, and their combinations can further support various language-aware tasks such as referring segmentation, VQA, etc. Likewise, the output is categorized into pixel-level mask and semantic embedding . By simply using different combinations, we can adapt our X-Decoder to various tasks with the same suite of parameters.
We employ a single text encoder to encode the textual corpus from all tasks. The common text encoder is used to encode referring phrases, text descriptions, image captions in the task of referring segmentation, image-text retrieval and image captioning, respectively. Furthermore, we reformulate the mask classification in segmentation into a mask-text matching problem between and the textual embeddings of prompted textual concepts similar to . Sharing the text encoder for all textual corpus could maximally exchange knowledge from different tasks and learn a richer and more coherent semantic space.
We fully decouple the image and text encoder. In many previous unified encoder-decoder models , the image and text are fused in the encoder side. This design makes it intractable not only for global image-text contrastive learning , but also generative pretraining . In contrast, by fully decoupling the image and text encoder and using the outputs all as queries, X-Decoder can learn from both intra-image supervisions and inter-image ones, which is essential to learn stronger pixel-level representations and support different granularity of tasks.
2 Unification of Tasks
Based on the above designs, X-Decoder can be used to seamlessly unify different vision and vision-language tasks, simply with different combinations of queries as inputs. Generic Segmentation. For this task, there are no textual queries as inputs. Hence, Eq. (1) becomes:
where , have the same size of . Eq. (2) reduces to Mask2former , but with open-vocabulary capacity since we use mask-text matching for mask classification.
Referring Segmentation. It requires both latent and text queries as inputs, thus shares the same formula as Eq. (1). Similar to generic segmentation, we only use the first decoded outputs corresponding to the latent queries. Compared with Eq. (2), referring segmentation can be regarded as language-conditioned generic segmentation.
Image-Text Retrieval. The decoupled image and text encoder in our X-Decoder makes it straightforward for inter-image retrieval tasks. Specifically, we only feed the latent queries to the decoder and obtain the semantic representation of an image:
where has the same length as , and the last (-th) token in is then used to compute the similarities between images and texts.
Image Captioning and VQA. For both tasks, X-Decoder takes both latent and text queries and decodes the outputs:
where correspondingly has equal size to , and no masks are predicted. There are two slight differences between the two tasks. First, the caption prediction follows a causal masking strategy while VQA does not. Second, we use all the outputs in for captioning, but only the last one to predict the answer for VQA.
The adaptation of our X-Decoder to each task is further depicted in Fig. 3. Based on this unification, we can pretrain our X-Decoder jointly with all tasks using a proper combination of queries and losses, and further finetune for individual tasks without any extra heads.VQA is used for pretraining following common practice. As discussed earlier, a lineup of works exploited a sequential decoding interface for the unification . However, in this work, we advocate the unification by functionality rather than interface, namely, we maximally share the common parts of different tasks while keeping the remaining unchanged for individual tasks.
3 Unified Architecture
We follow Mask2Former to build our decoder architecture. Given an image , we extract hierarchical visual features from layers:
where and is the size of feature map at level and is the feature dimension. These hierarchical feature maps are important for pixel-level understanding at different scales.
One Decoder for All Tasks. Given the visual features , X-Decoder uses a stack of transformer layers to refine the queries and render the outputs. At layer , it first cross-attends the visual features and then performs self-attention among latent and text queries:
In Eq. (6), we let all queries cross-attend the visual features. For latent queries, we use a masked cross-attention mechanism as in , and full attention for the textual queries. In Eq. (7), we specifically design the self-attention mechanism to prompt the synergy of tasks: we use the last latent query to extract the global image representation and the remaining for generic segmentation; for image captioning, each textual query can attend itself, its predecessors and all latent queries; for referring segmentation, latent queries will attend all text queries to use it as the language condition.
Based on these rules, the resulting self-attention in our X-Decoder is shown in Fig. 4.
The output of our X-Decoder is also categorized into two types: 1) pixel-wise mask and 2) semantic outputs. X-Decoder always produces the masks only for the latent queries, i.e., for all the latent queries. As for the semantic outputs, X-Decoder predicts the outputs for both latent and text queries, i.e., , to cover both mask recognition and caption generation.
One Encoder for All Texts. Our text encoder consists of a number of transformer layers. Given the raw text such as a phrase or caption, we convert it to discrete tokens using an off-the-shelf tokenizer and then send it to the text encoder. We apply causal masking to ensure its outputs are compatible with caption decoding. For segmentation, we follow to convert the class name into a phrase with a text prompt (e.g., “dog” “an image of dog”), and encode the phrase as above.
4 End-to-End Pre-training
We train our X-Decoder in an end-to-end manner with two types of losses corresponding to the outputs.
Semantic Loss. There are three losses on the semantic outputs corresponding to three tasks. For image-text retrieval, we compute the language-image contrastive loss as . We take the last valid token feature of from the text encoder to represent a text as and take the last entry in derived from X-Decoder as . As a result, we obtain pairs of features for a minibatch of image-text pairs. Afterwards, we compute the dot-product between these feature pairs to obtain an affinity matrix , and compute the bidirectional cross-entropy loss:
where are the class labels corresponding to diagonal entries in , and is the transpose of .
For mask classification, we encode all class names including “background” into text queries and take the last valid token feature from each to represent the concept. Afterward, we take the decoder outputs corresponding to the first latent queries and compute the dot-product between these outputs and concept embeddings to obtain an affinity matrix and compute the loss , with the ground-truth class .
For image captioning, we first extract the embeddings for all tokens in the vocabulary of size from the text encoder. Given the last semantic outputs from X-Decoder, we compute the dot-product with all token embeddings to obtain an affinity matrix . Then we compute the cross-entropy loss , with the ground-truth next-token id .
Mask Loss. Given the predictions derived from latent queries, we use Hungarian matching to find the matched entries of first outputs to ground-truth annotations. Afterward, we follow to use binary cross-entropy loss and dice loss to compute the loss for masks. We combine the above four losses to pretrain our X-Decoder. More details can be found in Appendix.
Experiments
Datasets and Settings. We pretrain X-Decoder on three types of data including panoptic segmentation, image-text pairs (itp), and referring segmentation. For panoptic and referring segmentation, we use COCO2017 with segmentation annotations and exclude the validation sets of Ref-COCOg UMD and COCO Karpathy . In total, there are 104k images for segmentation pretraining, out of which 30k images are with referring segmentation annotations. For image-text pairs, we use the standard 4M corpora, including Conceptual Captions , SBU Captions , Visual Genome , and COCO Captions . We broadly evaluate our models on all tasks covered by pretraining, including generic (Semantic/Instance/Panoptic) segmentation, referring segmentation, image-text retrieval, and image captioning. In particular, we benchmark on 10 settings of 7 datasets covering a wide range of domains. Moreover, we finetune and report results on VQA for fine-grained visual reasoning.
Implementation Details. Our visual encoder follows to use 100 latent queries and 9 decoder layers for segmentation, and we add one additional latent query for image-level task. However, we do not adopt a deformable encoder as it does not generalize well to open-vocabulary settings (see in Appendix). We adopt Focal-T and DaViT-B/L as the vision encoder and a transformer text encoder with causal masking as language encoder. The models are pretrained on large-scale image-text data (Base or Large) or UniCL for the tiny model. During pretraining, we set a minibatch for segmentation to and image-text pairs to . The image resolution is set to for segmentation and for image-text data respectively. We follow a similar balanced sampling strategy in to ensure the segmentation data are always observed for a consistent number of epochs, regardless of the total number of image-text pairs. Based on this, we pretrain all models for 50 epochs using AdamW as the optimizer. During finetuning, we have task-specific designs, please refer to details in Appendix.
2 Task-Specific Transfer
Without any architecture change except adding a head for VQA, we directly finetune X-Decoder to demonstrate its task transfer capability. Table 1 presents the comparisons with previous specialized and generalized models.
Comparison with segmentation models. We list the most recent models for individual tasks, including Mask2Former , Panoptic SegFormer , KMaX-DeepLab for generic segmentation, and LAVT for referring segmentation. Notably, our 25 epoch finetuned X-Decoder (L) establishes a new SoTA on ADE20k dataset that outperforms the current SoTA KMaX-DeepLab (L) on ADE Panoptic Segmentation (our model trained with 1024 resolution achieves 51.0 PQ), as well as Instance Segmentation SoTA, Mask2Former-L. On COCO, our model attains comparable performance to Mask2Former and kMaX-DeepLab. There are three reasons to explain minor inferiority. First, we do not use deformable attention in X-Decoder, which typically benefits supervised settings but hurts open-vocabulary performance. Second, we use the language-image pretrained model as the backbone, which can understand richer semantics but lags behind the supervised model for classification tasks . Third, we use 100 latent queries for segmentation, which is half of that in Mask2Former (L). Finally, we compare with LAVT on COCO G-ref. It is worth pointing out that with lightweight finetuning, our tiny model already outperforms LAVT-Base (61.9 v.s. 61.2). Further increasing the model size can bring additional gains by 2.6 and 2.7 points respectively, which helps to set a new record on this benchmark.
Comparison with VL models. We compare with a set of VL models on image-text retrieval, image captioning and VQA in Table 1. X-Decoder achieves competitive performance across the board. Specifically, X-Decoder outperforms strong baseline UNITER and rivals VinVL on COCO retrieval, and even beats all the methods on Flickr30k . Unlike all these works, the image and text encoders are fully decoupled in X-Decoder, which leads to a much faster inference speed. On captioning and VQA, our models also demonstrate superior performance to their counterparts. For example, it outperforms VinVL by 1.3 and 1.7 on CIDEr and BLEU, respectively. Note that most of these works use sophisticatedly designed training objectives, such as masked data modeling, image-text matching and hard-negative mining . In contrast, X-Decoder is pretrained with image-text contrastive and image captioning, along with the segmentation losses. The simplicity and effectiveness imply a great potential of using X-Decoder as a general pretraining paradigm for VL.
Comparison with generalist models. We further compare with prior arts that explore general-purpose vision models. Limited works report the generic segmentation performance. Our model outperforms UViM and Pix2Seq v2 significantly on COCO panoptic (56.7 v.s. 45.8) and instance segmentation (46.7 v.s. 38.2), respectively. With the same amount of segmentation data, these margins strongly justify our model design, i.e., unifying functionality without any tweaks for individual tasks. When compared with GLIPv2 , our model achieves comparable performance. Note that GLIPv2 uses over 10M pretraining data, including around 2M with box supervision. Despite the huge gap in pretraining data, X-Decoder outperforms GLIPv2 on both captioning and VQA. Furthermore, X-Decoder also beats other general-purpose models like UniT , GPV , UniTAB and Unified-IO .
Efficient Finetuning. Finally, we study whether our pretrained X-Decoder can be finetuned for segmentation with a low cost. In Table 3, we show that we can simply finetune the class embedding layer, mask embedding layer or the whole decoder to reach a decent segmentation performance and surpass the fully finetuned tiny SoTA models like kMaX-DeepLab . These results imply an efficient way of using our pretrained X-Decoder models.
3 Zero-Shot Transfer
Without any change in model weights, X-Decoder can be directly applied to various segmentation tasks and datasets after pretraining. In Table 2, we evaluate our model in a zero-shot manner on seven commonly used segmentation datasets in 10 different settings from diverse domains, including common indoor (e.g., ADE20K and Pascal ), outdoor (e.g., Cityscapes ) and self-driving scenarios (e.g., BDD ). We report PQ, mAP and mIoU for panoptic, instance and semantic segmentation respectively. And we visualize the predicted open-vocabulary segmentation result on each dataset in Fig. 5.
Comparison with baselines. We build two X-Decoder variants: (1) X-Decoder-Seg, which is only trained with COCO panoptic segmentation using a text encoder for class names; and (2) X-Decoder-Seg+, where we take the heuristic way to extract noun phrases from COCO captions and use them as extra supervision on top of the matched decoder outputs. First, X-Decoder-Seg shows clear advantages on open-vocabulary segmentation over MSeg , that manually conducts label mapping across different datasets. Second, the extra supervision from COCO captions improves model performance on 9 out of 15 metrics, which indicates the benefit of joint learning with image-level supervision. Third, when pretraining with the full X-Decoder, the performance is significantly boosted. Notably, the mIoU metric is improved by 7.4, 3.4 and 2.6 on SUN, ADE-150 and PC-459, respectively.
Comparison with state-of-the-art. We further compare with the most advanced methods for open-vocabulary image segmentation in Table 2. Clearly, our models achieve the best results across all datasets. Among the base-sized models, X-Decoder (B) outperforms OpenSeg (B) on two challenging datasets, ADE-150 and PC-459 for semantic segmentation. Scaling X-Decoder to large size further improves mIoU by 2.4 and 1.4 on these two datasets. Among prior arts, MaskCLIP is the first proposed for open-vocabulary panoptic segmentation by combining Mask2Former with CLIP models. With COCO caption supervisions, our simple baseline X-Decoder-Seg+ already performs comparably. The full version of our tiny model X-Decoder (T) surpasses MaskCLIP across the board except A-847. We note that these comparisons are not strictly fair in terms of supervision, settings and models used. However, these results demonstrate the effectiveness of our X-Decoder to learn from the different granularity of supervisions end-to-end for open-vocabulary segmentation, which leads to new SoTA on 10 settings of 7 datasets across three segmentation tasks.
4 Model Inspection
Pretraining Tasks. By default, we exploit four pretraining tasks including generic and referring segmentation, captioning and retrieval. In Table 7, we keep the generic segmentation while ablating the importance of the other pretraining tasks. Accordingly, we have the following observations:
Image-text retrieval can help open-vocabulary segmentation. On ADE, the mIoU drops from 23.4 to 21.8, and PQ drops 0.7 without image-text retrieval. Since we share the same semantic space for both tasks, a good visual-semantic alignment learned from the retrieval task can directly benefit the recognition of novel concepts.
Image captioning helps referring segmentation and vice versa. We observe a drop of 2.0 pts on COCO g-Ref without captioning task, and a 3.2 pts drop of CIDEr from removing referring task. The two tasks share the same text encoder for text queries. Joint training, therefore, improves the understanding of text inputs.
Image captioning and retrieval can mutually benefit each other. When removing captioning during pretraining, the image retrieval R@1 drops by 0.8, and the captioning CIDEr drops significantly by 3.2 pts from removing retrieval task. Our X-Decoder promotes harmony of generative and contrastive learning.
The above observations verify that the unified design of X-Decoder can prompt the synergy of different tasks.
Query Interactions. The interaction among tasks is highly dependent on the interaction between latent and text queries. We have described how the queries interact with each other by default in Fig. 4. Here, we investigate how our model behaves with different interactions. In Table 5, we show the performance across tasks with ablated versions and have the following takeaways:
Image captioning requires both fine-grained and global image information. Comparing the first with the second and third row in the table, we find the CIDEr score significantly drops if we cut off the information flow from the global latent query or other latent queries to text queries (82.0 78.6 and 78.9, respectively).
Language-condition is important for referring segmentation. In the last row, we turn off the interaction from text queries to latent queries. This significantly hurts referring segmentation (59.7 57.6). On the one hand, this indicates that we can convert generic segmentation to referring segmentation using post-hoc matching with referring texts. On the other hand, sending the text phrase as input to X-Decoder is essential to modulate our model to specifically decode the targets.
VL Batch Size & Dataset The default batch size of VL task is , here we explore the gradual decreasing of VL batch size. In addition, each VL dataset is removed individually to investigate the pre-trained performance on different tasks.
Decreasing VL batch size hurts VL tasks and open-vocab Segmentation performance. As shown in Table. 5, decreasing the VL task batch size from to significantly hurts the retrieval and captioning tasks’ performance, where ir@1, tr@1, CIDEr decrease by 3.2, 4.3, and 8.9 points respectively. Further, the open-vocabulary performance also drops 0.3 points on each metric.
VG dataset hurts pretraining VL tasks performance but improves open-vocab segmentation. As shown in Table 7, removing the visual genome from the pretraining VL dataset significantly improves captioning task with 22.1 points during pretraining, but only 0.2 points after finetuning. Moreover, open-vocabulary semantic segmentation drops around 0.8 points.
5 Task Composition
X-Decoder has the unique benefit of task interaction, thanks to the sophisticated architecture design on latent and text queries as well as the decoder architecture. It enables joint task inference and iterative task inference with a single set of weights. In Fig. 6, we show our model can perform region-based retrieval and referring based captioning without any architecture/weight change. For example, given a set of animal images (row 1, Fig. 6) and text query, our model first retrieves the correct image (flamingo and giraffe) and then grounds the query with pixel-level predictions. Further, our model can easily adapted to referring captioning by first localizing a given word and then modulating the predicted mask in the cross-attention layers. Lastly, we also integrate X-Deocder with diffusion model to do referring image editing demonstrated in the latter half of the second row in Fig. 6.
Conclusion
We present X-Decoder, a model that seamlessly supports pixel-level and image-level vision-language understanding. With a simple and generalized design, X-Decoder can unite and support generic segmentation, referring segmentation and VL tasks effortlessly, achieving strong generalizability and competitive or even SoTA performance. We hope this work can shed a light on the design of the next-generation general-purpose vision system.
We appreciated the constructive discussion with Haotian Zhang. This work was also supported in part by NSF CAREER IIS2150012, the Wisconsin Alumni Research Foundation, and the Institute of Information & communications Technology Planning & Evaluation (IITP) grant funded by the Korea government (MSIT) (No. 2022- 0-00871, Development of AI Autonomy and Knowledge Enhancement for AI Agent Collaboration).
References
Appendix A Experiment Settings
In the main paper, all the pre-trained models are trained with 50 epochs of COCO data and roughly 45 epochs of 10 million image-text pairs. The batch size of COCO images and image text pairs are 32 and 1024 respectively. And 32 GPUs are used for pretraining. The AdamW optimizer is used in pretraining with the initial learning rate 1e-4. A step-wise scheduler is used to decay the learning rate by 0.1 on the fraction of training steps.
A.2 Finetuning
Image-Text Retrieval. For both COCO and Flickr30k image-text retrieval, we finetune the models for 10 epochs using AdamW as the optimizer. We set the image resolution to 384 and the batch size to 2048. The learning rates are 3e-5 for the X-Decoder part and 3e-6 for the vision and language backbones.
Image Captioning. Similar to image-text retrieval, we finetune the captioning models for 10 epochs using AdamW as the optimizer. We set the image resolution to 480 and the batch size to 256. The learning rates are 2e-5 for the X-Decoder part and 2e-6 for the vision and language backbones. We use beam search during caption generation with the beam size set to 5. We do not use CIDEr optimization for our captioning models.
VQA. For VQA, we add a new classification layer on the top of the model and finetune the models for 10 epochs using AdamW as the optimizer. We set the image resolution to 640 and the batch size to 256. The learning rates are 1e-4 for the X-Decoder part, 1e-5 for the vision and language backbones, and 1e-3 for the VQA classification layer.
Generic Segmentation. For generic segmentation, we finetune the pretrained checkpoint with 24 epochs with start learning rate 1e-4. We decay the learning rate by factor 10 at epoch 21 and 23, respectively. The batch size of ADE20k is 64, and 32 for COCO.
Referring Segmentation. For referring segmentation, we also finetune the pretrained checkpoint with 24 epochs. However, as RefCOCO has been used in pretraining, thus the initial learning rate is 1e-5. It also decays twice at 21 and 23 epochs. We use a batch size of 64 during training. Further, in addition to the normal setting that multiple backbone and language encoder learning rates with 0.1, here we also multiply the transformer encoder learning rate by 0.1.
Appendix B Open-Vocab Segmentation Benchmark
We propose an open vocabulary segmentation benchmark on 9 datasets with different evaluation metrics. The goal of this benchmark is to provide a comprehensive and standard evaluation protocol for open-vocabulary segmentation on different vocabulary sizes and image domains.
Table 8 shows the dataset statistics in the benchmark. It supports all generic segmentation tasks including semantic/instance/panoptic segmentation. It covers a variety of scopes ranging from 20 to 847 classes. In addition, the evaluation scene includes common objects, in-door scenes as well as autonomous driving scenarios. To enable a better understanding of the open-vocabulary ability on the training/evaluation datasets. We evaluate the coverage of training datasets captions and evaluation datasets concepts in Fig. 10-16 (we split the caption into single words and phrases to find mappings in categories). The major results of the open-vocabulary segmentation are evaluated in the main paper, Tab. 2.
Appendix C Extra Ablation Studies
In our main paper, we observed that the vision-language pretraining objectives including image-text contrastive learning and image captioning have clear benefits to image segmentation, particularly in the zero-shot setting. Here, we further study the role of segmentation objectives in vision-language understanding. To investigate, we remove the segmentation data (COCO panoptic segmentation and referring segmentation) and only pretrain X-Decoder on the four million image-text pairs, denoted by X-Decoder-VL. Afterwards, we transfer the model to downstream VL tasks. As we can see from Table 9, the performance significantly drops across all tasks after removing the segmentation data for pretraining. We suspect that segmentation data can help models to learn more fine-grained visual understanding and consequently benefit vision-language tasks. Along with our findings in the main paper, we conclude that pixel-level segmentation and vision-language learning are complementary to each other for zero-shot and task-specific transfer.
C.2 Model Architecture Inspection
In Table. 10, we report the results using three different vision backbone architectures, including Swin , FocalNet and DaViT . All models in the first block are with tiny size and trained on the combination of image-label and image-text pairs, following the settings in UniCL . In the second block, all the models are initialized with Florence pre-trained DaVit-d5 model. Through the comparisons, we have the following observations: (1) FocalNet and DaViT achieve better performance than Swin across all metrics. Particularly, FocalNet achieves the best performance on generic and referring segmentation, while DaViT is better on the zero-shot vision-language evaluations; (2) After adding the deformable attention, we can see a boost on supervised segmentation but significant (especially large model) degradation on the open-vocabulary segmentation on ADE20K dataset. Based on these experimental results, we make the design choices as mentioned in our main submission: (1) we remove deformable attention in the favor of open-vocabulary segmentation; (2) we use FocalNet as the tiny vision encoder and train it by ourselves using UniCL, while using DaViT as the base and large vision encoder.
C.3 Open-Vocabulary Generic Segmentation Settings Inspection
In Tab. 11, we study the progressive enrichment of data and training settings as well as the pre-trained model usage. X-Decoder-Seg is the baseline of adding a text encoder to Mask2Former with a learnable language encoder. X-Decoder-Seg+ takes use of caption nouns for Hungarian matching to enrich the vocabulary size. In addition to the main paper, we add row 3 in Tab. 11 to demonstrate the performance of X-Decoder with only coco image text pairs. Comparing 3rd row and 4th row, we find adding extra image-text pairs for pretraining clearl improve open-vocabulary segmentation performance especially when the vocabulary size is large (e.g. ADE-150, CONTEXT-59/459). The way of pretraining vision backbone also matters. Comparing the last two rows side by side, though the backbone model sizes are similar, using ImageNet-21K for pretraining leads to inferior performance on most of the datasets except for CONTEXT-459 which contains most number of categories. These results demonstrate the benefits of using more image-text pairs for pretraining the vision backbone or our X-Decoder.
Appendix D Segmentation In the Wild Benchmark
As shown in the main submission, our X-Decoder exhibits a strong generalization ability to segment images in ten settings of seven datasets from different domains, without any dataset-specific finetuning. Inspired by the object detection in the wild setting proposed in GLIP , we resort to more domain-specific datasets on the web to further examine the generality of our model. Specifically, we download 55 instance segmentation datasets from Roboflow https://roboflow.com/. Afterward, we clean the datasets by excluding those containing visually undetectable categories (e.g. Different species of plant) or categories labeled with other languages. In the end, we compile 25 datasets that are suitable for evaluation into segmentation in the wild (SegInW) benchmark and report instance segmentation mAP. The dataset meta information is listed in Tab. 12, and examplar images are shown in Fig. 7.
On the SegInW benchmark, we evaluate zero-shot, few-shot, and fine-tuned segmentation for five models (X-Decoder-Seg+ as baselines, and X-Decoder with different visual backbone) on three different tuning scales. In Fig. 8, we report the zero-shot instance segmentation performance on 25 datasets separately in a descending order. Accordingly, X-Decoder shows reasonably good generalization ability to a wide range of visual and concept domains. Specifically, it achieves higher mAP on common objects like fruits and animals but lower ones on fine-grained datasets like toolkits and rare concepts like rail and brain tumor. In Fig. 9, we further show the line chars for few-shot learning and fully-finetuning, and observe that:
X-Decoder has privilege on small-scale tuning. As shown in Fig. 9 (a-b), comparing with X-Decoder-Seg+ that only extract noun phrase to increase vocabulary size, X-Decoder performs much better with few-shot/finetune setting. Although X-Decoder (B) and X-Decoder-Seg+ (B) have similar zero-shot performance, the gap increases with the number of images tuned. However, as the number of parameters tuned increased by a large margin Fig. 9 (c), the performance gap between X-Decoder and X-Decoder-Seg+ is shrunk to a small margin.
Zero-Shot gap could be bridged by tuning. X-Decoder (L) and X-Decoder (L-IN21K) are initialized with different pre-trained image backbones. Specifically, X-Decoder (L) is initialized by Florence pre-trained Davit-d5, whereas X-Decoder (L-IN21K) is initialized with FocalNet-L pretrained on ImageNet-21k . As shown in Fig. 9 (a-c), although the gap between X-Decoder-L and X-Decoder-L-IN21K on the zero-shot setting is relatively large. However, the gap on 5/10/full finetuned settings is much smaller and even cross in some settings.
Tuning class embedding is enough for few-shot settings. As shown in Fig. 9 (e-h), on the smaller scale backbone including (T/B), although tuning the full decoder has a better result, the gap is not obvious on 0-10 shots. And on larger scale models including L/L-IN21K, tuning with class embedding has similar/better results on 0-10 shots.
We show more detailed results in Table 13, Table 14 and Table 15. Similar to Table 3, we report the number of parameters tuned in each setting.
Appendix E Extra Visualization
In this part, we demonstrate the generalization ability to video datasets and flexibility to support task compositions for X-Decoder with more qualitative visualizations.
Open-vocabulary generic segmentation is one of the main advantages of X-Decoder. We also apply generic segmentation in a zero-shot manner to the YoutubeVOS dataset. As shown in Fig. 17, our model can be well generalized to video zero-shot generic segmentation and make predictions that are consistent across frames. As a result, our model can be used in video segmentation directly or a good initialization for further finetuning.
E.2 Zero-Shot Referring Video Segmentation
Besides the generic segmentaton on video frames, our X-Decoder can be easily adapted to referring video segmentation as well without any architectural change or finetuning. In Fig. 18, we visualize some examples of referring video segmentation on the YoutubeVOS dataset in a zero-shot manner. We can see that our model can generate rather accurate outputs given various referring phrases. Notably, in addition to the strong segmentation performance for given concepts, the model can also correctly distinguish the spatial locations (e.g., left v.s. right in the first row), and object attributes (e.g., a baby gorilla instead of an adult gorilla in the second row) in these unseen videos.
E.3 Zero-Shot Image Captioning
To test the generalization ability of X-Decoder, we also ask the model generate image captions on the YoutubeVOS dataset, which is in a different domain from the image data. As we can see from the examples in Fig. 19, the model can correctly predict the object, activity, and environment in an image. Interestingly, the captions for the first 6 images sampled from 3 different videos show that our approach can correctly differentiate the movements from similar scenarios (e.g., a man playing vs. a man standing in the first two samples.).
E.4 Zero-Shot Referring Captioning
In compensating for the visualization of the main paper, we add more referring captioning samples in Fig 20. The phrase before “:” is the referring phrase, and the sentence after “:” is the generated caption. The grounding mask of the referring phrase is highlighted in pink. Clearly, our model can simultaneously segments the referred region and generates a region-specific caption. Complementary to regular image captioning systems, such a novel functionality provides a way of interpreting images in a more fine-grained manner. Note that our X-Decoder was never trained to generate such regional captions.
E.5 Zero-Shot Referring Image Editing
Finally, given the high-quality referring segmentation results with X-Decoder, we can effortlessly combine it with off-the-shelf Stable-Diffusion image inpainting model and perform zero-shot referring image editing. As shown in Fig. 21, the model first performs referring segmentation, then the original image and the segmentation mask are fed into the inpainting model to generate the inpainted image. For example, given “change bird to squirrel”, it first extracts the bird segment (blue region) from the input image and then replace the segmented region with a generated squirrel. Likewise in other samples, we can see all the generated images look natural and follow the inpainting instructions very well. These impressive plug-and-play results imply a great potential of combining our X-Decoder and advanced generative AI models for fine-grained precise image editing.
Appendix F Discussions
Future Directions. The extensive quantitative and qualitative results have demonstrated the strong performance and generalization ability of our X-Decoder for a variety of vision and vision-language tasks at different granularities. Upon the current X-Decoder design, we see two directions worth future explorations: (1) Pretrain the whole model in one stage effectively and efficiently. Currently, the model still requires a separate pretraining for the image and text encoders. However, since our model supports large-scale image-text contrastive learning thanks to the decoupled design, we can easily unify the CLIP-style pretraining with the decoder pretraining in an end-to-end manner. (2) Unify all level of supervisions. Due to high annotation costs, the pixel-level segmentation annotations by nature are much less than the region-level box and image-level annotations. It is worth building a more unified learning paradigm to jointly learn from pixel-level, region-level and image-level supervision to attain a more powerful unified model.
Social Impact. This work is mainly focused on the design of a generalized decoder for various vision and vision-language tasks. We have used a pretrained image and text encoder and further pretrained the models on a combination of various datasets and tasks. Since the models are trained on large-scale webly-crawled image-text pairs, the negative impact might arise due to the potential offensive or biased content in the data. To mitigate this issue, we need to have a careful sanity check on the training data and model predictions before deploying it in practical scenarios.