Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation

Feng Li, Hao Zhang, Huaizhe xu, Shilong Liu, Lei Zhang, Lionel M. Ni, Heung-Yeung Shum

Introduction

Object detection and image segmentation are fundamental tasks in computer vision. Both tasks are concerned with localizing objects of interest in an image but have different levels of focus. Object detection is to localize objects of interest and predict their bounding boxes and category labels, whereas image segmentation focuses on pixel-level grouping of different semantics. Moreover, image segmentation encompasses various tasks including instance segmentation, panoptic segmentation, and semantic segmentation with respect to different semantics, e.g., instance or category membership, foreground or background category.

Remarkable progress has been achieved by classical convolution-based algorithms developed for these tasks with specialized architectures, such as Faster RCNN for object detection, Mask RCNN for instance segmentation, and FCN for semantic segmentation. Although these methods are conceptually simple and effective, they are tailored for specialized tasks and lack the generalization ability to address other tasks. The ambition to bridge different tasks gives rise to more advanced methods like HTC for object detection and instance segmentation and Panoptic FPN , K-net for instance, panoptic, and semantic segmentation. Task unification not only helps simplify algorithm development but also brings in performance improvement in multiple tasks.

Recently, DETR-like models developed based on Transformers have achieved inspiring progress on many detection and segmentation tasks. As an end-to-end object detector, DETR adopts a set-prediction objective and eliminates hand-crafted modules such as anchor design and non-maximum suppression. Although DETR addresses both the object detection and panoptic segmentation tasks, its segmentation performance is still inferior to classical segmentation models. To improve the detection and segmentation performance of Transformer-based models, researchers have developed specialized models for object detection , image segmentation , instance segmentation , panoptic segmentation , and semantic segmentation .

Among the efforts to improve object detection, DINO takes advantage of the dynamic anchor box formulation from DAB-DETR and query denoising training from DN-DETR , and further achieves the SOTA result on the COCO object detection leaderboard for the first time as a DETR-like model. Similarly, for improving image segmentation, MaskFormer and Mask2Former propose to unify different image segmentation tasks using query-based Transformer architectures to perform mask classification. Such methods have achieved remarkable performance improvement on multiple segmentation tasks.

However, in Transformer-based models, the best-performing detection and segmentation models are still not unified, which prevents task and data cooperation between detection and segmentation tasks. As an evidence, in CNN-based models, Mask-R-CNN and HTC are still widely acknowledged as unified models that achieve mutual cooperation between detection and segmentation to achieve superior performance than specialized models. Though we believe detection and segmentation can help each other in a unified architecture in Transformer-based models, the results of simply using DINO for segmentation and using Mask2Former for detection indicate that they can not do other tasks well, as shown in Table 2 and 2. Moreover, trivial multi-task training can even hurt the performance of the original tasks. It naturally leads to two questions: 1) why cannot detection and segmentation tasks help each other in Transformer-based models? and 2) is it possible to develop a unified architecture to replace specialized ones?

To address these problems, we propose Mask DINO, which extends DINO with a mask prediction branch in parallel with DINO’s box prediction branch. Inspired by other unified models for image segmentation, we reuse content query embeddings from DINO to perform mask classification for all segmentation tasks on a high-resolution pixel embedding map (1/4 of the input image resolution) obtained from the backbone and Transformer encoder features. The mask branch predicts binary masks by simply dot-producting each content query embedding with the pixel embedding map. As DINO is a detection model for region-level regression, it is not designed for pixel-level alignment. To better align features between detection and segmentation, we also propose three key components to boost the segmentation performance. First, we propose a unified and enhanced query selection. It utilizes encoder dense prior by predicting masks from the top-ranked tokens to initialize mask queries as anchors. In addition, we observe that pixel-level segmentation is easier to learn in the early stage and propose to use initial masks to enhance boxes, which achieves task cooperation. Second, we propose a unified denoising training for masks to accelerate segmentation training. Third, we use a hybrid bipartite matching for more accurate and consistent matching from ground truth to both boxes and masks.

Mask DINO is conceptually simple and easy to implement under the DINO framework. To summarize, our contributions are three-fold. 1) We develop a unified Transformer-based framework for both object detection and segmentation. As the framework is extended from DINO, by adding a mask prediction branch, it naturally inherits most algorithm improvements in DINO including anchor box-guided cross attention, query selection, denoising training, and even a better representation pre-trained on a large-scale detection dataset. 2) We demonstrate that detection and segmentation can help each other through a shared architecture design and training method. Especially, detection can significantly help segmentation tasks, even for segmenting background "stuff" categories. Under the same setting with a ResNet-50 backbone, Mask DINO outperforms all existing models compared to DINO (+0.8+0.8 AP on COCO detection) and Mask2Former (+2.6+\textbf{2.6} AP, +1.1+\textbf{1.1} PQ, and +1.5+\textbf{1.5} mIoU on COCO instance, COCO panoptic, and ADE20K semantic segmentation). 3) We also show that, via a unified framework, segmentation can benefit from detection pre-training on a large-scale detection dataset. After detection pre-training on the Objects365 dataset with a SwinL backbone, Mask DINO significantly improves all segmentation tasks and achieves the best results on instance (54.5 AP on COCO), panoptic (59.4 PQ on COCO), and semantic (60.8 mIoU on ADE20K) segmentation among models under one billion parameters.

Related Work

Detection: Mainstream detection algorithms have been dominated by convolutional neural network-based frameworks, until recently Transformer-based detectors achieve great progress. DETR is the first end-to-end and query-based Transformer object detector, which adopts a set-prediction objective with bipartite matching. DAB-DETR improves DETR by formulating queries as 44D anchor boxes and refining predictions layer by layer. DN-DETR introduces a denoising training method to accelerate convergence. Based on DAB-DETR and DN-DETR, DINO proposes several new improvements on denoising and anchor refinement and achieves new SOTA results on COCO detection. Despite the inspiring progress, DETR-like detection models are not competitive for segmentation. Vanilla DETR incorporates a segmentation head in its architecture. However, its segmentation performance is inferior to specialized segmentation models and only shows the feasibility of DETR-like detection models to deal with detection and segmentation simultaneously. Segmentation: Segmentation mainly includes instance, semantic, and panoptic segmentation. Instance segmentation is to predict a mask and its corresponding category for each object instance. Semantic segmentation requires to classify each pixel including the background into different semantic categories. Panoptic segmentation unifies the instance and semantic segmentation tasks and predicts a mask for each object instance or background segment. In the past few years, researchers have developed specialized architectures for the three tasks. For example, Mask-RCNN and HTC can only deal with instance segmentation because they predict the mask of each instance based on its box prediction. FCN and U-Net can only perform semantic segmentation since they predict one segmentation map based on pixel-wise classification. Although models for panoptic segmentation unifies the above two tasks, they are usually inferior to specialized instance and semantic segmentation models. Until recently, some image segmentation models are developed to unify the three tasks with a universal architecture. For instance, Mask2Former improves MaskFormer by introducing masked-attention to Transformer. Mask2Former has a similar architecture as DETR to probe image features with learnable queries but differs in using a different segmentation branch and some specialized designs for mask prediction. However, while Mask2Former shows a great success in unifying all segmentation tasks, it leaves object detection untouched and our empirical study shows that its specialized architecture design is not suitable for predicting boxes. Unified Methods: As both object detection and segmentation are concerned with localizing objects, they naturally share common model architectures and visual representations. A unified framework not only helps simplify the algorithm development effort, but also allows to use both detection and segmentation data to improve representation learning. There have been several previous works to unify segmentation and detection tasks, e.g., Mask RCNN , HTC , and DETR . Mask RCNN extends Faster RCNN and pools image features from Region Of Interest (ROI) proposed by RPN. HTC further proposes an interleaved way of predicting boxes and masks. However, these two models can only perform instance segmentation. DETR predicts boxes and masks together in an end-to-end manner. However, its segmentation performance largely lags behind other models. According to Table 2, adding DETR’s segmentation head to DINO results in inferior instance segmentation results. How to attain mutual assistance between segmentation and detection has long been an important problem to solve.

Mask DINO

Mask DINO is an extension of DINO . On top of content query embeddings, DINO has two branches for box prediction and label prediction. The boxes are dynamically updated and used to guide the deformable attention in each Transformer decoder. Mask DINO adds another branch for mask prediction and minimally extends several key components in detection to fit segmentation tasks. To better understand Mask DINO, we start by briefly reviewing DINO and then introduce Mask DINO.

DINO is a typical DETR-like model, which is composed of a backbone, a Transformer encoder, and a Transformer decoder. The framework is shown in Fig. 1 (the blue-shaded part without red lines). Following DAB-DETR , DINO formulates each positional query in DETR as a 4D anchor box, which is dynamically updated through each decoder layer. Note that DINO uses multi-scale features with deformable attention . Therefore, the updated anchor boxes are also used to constrain deformable attention in a sparse and soft way. Following DN-DETR , DINO adopts denoising training and further develops contrastive denoising to accelerate training convergence. Moreover, DINO proposes a mixed query selection scheme to initialize positional queries in the decoder and a look-forward-twice method to improve box gradient back-propagation.

2 Why a universal model has not replaced the specialized models in DETR-like models?

Remarkable progress has been achieved by Transformer-based detectors and segmentation models. For instance, DINO and Mask2Former have achieved the best results on COCO detection and panoptic segmentation, respectively. Inspired by such progress, we attempted to simply extend these specialized models for other tasks but found that the performance of other tasks lagged behind the original ones by a large margin, as shown in Table 2 and 2. It seems that trivial multi-task training even hurts the performance of the original task. However, in convolution-based models, it has shown effective and mutually beneficial to combine detection and instance segmentation tasks. For example, detection models with Mask R-CNN head is still ranked the first on the COCO instance segmentation. We will take DINO and Mask2Former as examples to discuss the challenges in unifying Transformer-based detection and segmentation. −-What are the differences between specialized detection and segmentation models? Image segmentation is a pixel-level classification task, while object detection is a region-level regression task. In DETR-based model, the decoder queries are responsible for these tasks. For example, Mask2Former uses such decoder queries to dot-product the high-resolution feature maps to produce segmentation masks, while DINO uses them to regress boxes. However, as such queries in Mask2Former only have to compare per-pixel similarity with the image features, they may not be aware of the region-level position of each instance. On the contrary, queries in DINO are not designed to interact with such low-level features to learn pixel-level representation. Instead, they encode rich positional information and high-level semantics for detection. −-Why cannot Mask2Former do detection well? The Transformer decoder of Mask2Former is designed for segmentation tasks and does not suit detection for three reasons. First, its queries follow the design in DETR without being able to utilize better positional priors as studied in Conditional DETR , Anchor DETR , and DAB-DETR . For example, its content queries are semantically aligned with the features from the Transformer encoder, whereas its positional queries are just learnable vectors as in vanilla DETR instead of being associated with a single-mode position We refer the interested readers to discussions in Sec. 3 in DAB-DETR . If we remove its mask branch, it reduces to a variant of DETR , whose performance is inferior to recently improved DETR models. Second, Mask2Former adopts masked attention (multi-head attention with attention mask) in Transformer decoders. The attention masks predicted from a previous layer are of high resolution and used as hard-constraints for attention computation. They are neither efficient nor flexible for box prediction. Third, Mask2Former cannot explicitly perform box refinement layer by layer. Moreover, its coarse-to-fine mask refinement in decoders fails to use multi-scale features from the encoder. As shown in Table 2, the generated box AP from mask is 4.5 AP worse than DINO and trivial multi-task learning by adding a detection head is not working We also notice there are issues in official Mask2Former Github (https://github.com/facebookresearch/Mask2Former/issues/43) that fail to make Mask2Former work well by adding a detection head.. −-Why cannot DETR/DINO do segmentation well? As shown in Table 2, simply 1) adding DETR’s segmentation head or 2) adding Mask2Former’s segmentation head result in inferior performance compared to Mask2Former. We analyze the reasons as follows. The reason for 1) is that DETR’s segmentation head is not optimal. The vanilla DETR lets each query embedding dot-product with the smallest feature map to compute attention maps and then upsamples them to get the mask predictions. This design lacks an interaction between queries and larger feature maps from the backbone. In addition, the head is too heavy to use mask auxiliary loss for mask refinement. The reason for 2) is that features in improved detection models are not aligned with segmentation. For example, DINO inherits many designs from like query formulation, denoising training, and query selection. However, these components are designed to strengthen region-level representation for detection, which is not optimal for segmentation.

3 Our Method: Mask DINO

Mask DINO adopts the same architecture design for detection as in DINO with minimal modifications. In the Transformer decoder, Mask DINO adds a mask branch for segmentation and extends several key components in DINO for segmentation tasks. As shown in Fig. 1, the framework in the blue-shaded part is the original DINO model and the additional design for segmentation is marked with red lines.

4 Segmentation branch

Following other unified models for image segmentation, we perform mask classification for all segmentation tasks. Note that DINO is not designed for pixel-level alignment as its positional queries are formulated as anchor boxes and its content queries are used to predict box offset and class membership. To perform mask classification, we adopt a key idea from Mask2Former to construct a pixel embedding map which is obtained from the backbone and Transformer encoder features. As shown in Fig. 1, the pixel embedding map is obtained by fusing the 1/41/4 resolution feature map CbC_{b} from the backbone with an upsampled 1/81/8 resolution feature map CeC_{e} from the Transformer encoder. Then we dot-product each content query embedding qcq_{c} from the decoder with the pixel embedding map to obtain an output mask mm.

where M\mathcal{M} is the segmentation head, T\mathcal{T} is a convolutional layer to map the channel dimension to the Transformer hidden dimension, and F\mathcal{F} is a simple interpolation function to perform 2x upsampling of CeC_{e}. This segmentation branch is conceptually simple and easy to implement in the DINO framework, as shown in Fig. 1.

5 Unified and Enhanced Query Selection

Unified query selection for mask: Query selection has been widely used in traditional two-stage models and many DETR-like models to improve detection performance. We further improve the query selection scheme in Mask DINO for segmentation tasks.

The encoder output features contain dense features, which can serve as better priors for the decoder. Therefore, we adopt three prediction heads (classification, detection, and segmentation) in the encoder output. Note that the three heads are identical to the decoder heads. The classification score of each token is considered as the confidence to select top-ranked features and feed them to the decoder as content queries. The selected features also regress boxes and dot-product with the high-resolution feature map to predict masks. The predicted boxes and masks will be supervised by the ground truth and are considered as initial anchors for the decoder. Note that we initialize both the content and anchor box queries in Mask DINO whereas DINO only initializes anchor box queries. Mask-enhanced anchor box initialization: As summarized in Sec 3.2, image segmentation is a pixel-level classification task while object detection is a region-level position regression task. Therefore, compared to detection, though segmentation is a more difficult task with fine-granularity, it is easier to learn in the initial stage. For example, masks are predicted by dot-producting queries with the high-resolution feature map, which only needs to compare per-pixel semantic similarity. However, detection requires to directly regress the box coordinates in an image. Therefore, in the initial stage after unified query selection, mask prediction is much more accurate than box (the qualitative AP comparison between mask prediction and box prediction in different stages is also shown in Table 10 and 10). Therefore, after unified query selection, we derive boxes from the predicted masks as better anchor box initialization for the decoder. By this effective task cooperation, the enhanced box initialization can bring in a large improvement to the detection performance.

6 Segmentation Micro Design

Unified denoising for mask: Query denoising in object detection has shown effective to accelerate convergence and improve performance. It adds noises to ground-truth boxes and labels and feed them to the Transformer decoder as noised positional queries and content queries. The model is trained to reconstruct ground truth objects given their noised versions. We also extend this technique to segmentation tasks. As masks can be viewed as a more fine-grained representation of boxes, box and mask are naturally connected. Therefore, we can treat boxes as a noised version of masks, and train the model to predict masks given boxes as a denoising task. The given boxes for mask prediction are also randomly noised for more efficient mask denoising training. The detailed noise and its hyperparameters used in our model are shown in Appendix B.2. Hybrid matching: Mask DINO, as in some traditional models , predicts boxes and masks with two parallel heads in a loosely coupled manner. Hence the two heads can predict a pair of box and mask that are inconsistent with each other. To address this issue, in addition to the original box and classification loss in bipartite matching, we add a mask prediction loss to encourage more accurate and consistent matching results for one query. Therefore, the matching cost becomes λclsLcls+λboxLbox+λmaskLmask\lambda_{cls}\mathcal{L}_{cls}+\lambda_{box}\mathcal{L}_{box}+\lambda_{mask}\mathcal{L}_{mask}, where Lcls,Lbox\mathcal{L}_{cls},\mathcal{L}_{box}, and Lmask\mathcal{L}_{mask} are the classification, box, and mask loss and λ\lambda are their corresponding weights. The detailed losses used in our model and their corresponding weights are shown in Appendix B.1. Decoupled box prediction: For the panoptic segmentation task, box prediction for "stuff" categories is unnecessary and intuitively inefficient. For example, many "stuff" categories are background like "sky", whose GT mask-derived boxes are highly irregular and often cover the whole image. Therefore, box prediction for these categories can mislead the instance-level ("thing") detection and segmentation. To address this problem, we remove box loss and box matching for "stuff" categories. More specifically, the box prediction pipeline remains the same for "stuff" to locate meaningful regions and extract features with deformable attention. However, we do not count their box prediction loss. In our hybrid matching, the box loss for "stuff" is set to the mean of "thing" categories. This decoupled design can accelerate training and yield additional gains for panoptic segmentation.

Experiments

We conduct extensive experiments and compare with several specialized models for four popular tasks including object detection, instance, panoptic, and semantic segmentation on COCO , ADE20K , and Cityscapes . For all experiments, we use batch size 16 and A100 GPUs with 40GB memory. We use a ResNet-50 and a SwinL backbone for our main results and SOTA model. Under ResNet-50, we use 44 A100 GPUs for all tasks without extra data. The implementation details are in Appendix A.

Instance segmentation and object detection. In Table 3, we compare Mask DINO with other instance segmentation and object detection models. Mask DINO outperforms both the specialized models such as Mask2Former and DINO and hybrid models such as HTC under the same setting. Especially, the instance segmentation results surpass the strong baseline Mask2Former by a large margin (+2.7 AP and +2.6 AP) on the 12-epoch and 50-epoch settings. Moreover, Mask DINO significantly improves the convergence speed, outperforming Mask2Former with less than half training epochs (44.2{44.2} AP in 24 epochs). In addition, after using mask-enhanced box initialization, our detection performance has been significantly improved (+1.2 AP), which even outperforms DINO by 0.8 AP. These results indicate that task unification is beneficial. Without bells and whistles, we achieve the best detection and instance segmentation performance among DETR-like model with a SwinL backbone without extra data.

Panoptic segmentation. We compare Mask DINO with other models in Table 4. Mask DINO outperforms all previous best models on both the 1212-epoch and 5050-epoch settings by 1.0 PQ and 1.1 PQ, respectively. This indicates Mask DINO has the advantages of both faster convergence and superior performance. One interesting observation is that we outperform Mask2Former in terms of both PQThPQ^{Th} and PQStPQ^{St}. However, instead of using dense and hard-constrained masked attention, we predict boxes and then use them in deformable attention to extract query features. Therefore, our box-oriented deformable attention also works well with "stuff" categories, which makes our unified model simple and efficient. In addition, we improve the mask APpanTh{}_{pan}^{Th} by 2.6 to 44.344.3 AP, which is 0.60.6 higher than the specialized instance segmentation model Mask2Fomer (43.743.7 AP). Semantic segmentation. In Table 6 and 6, we show the performance of semantic segmentation with a ResNet-50 backbone. We use 100100 queries for these small datasets. We outperform Mask2Former on both ADE20K and Cityscapes by 1.6{1.6} and 0.6{0.6} mIoU on the reported performance.

2 Comparison with SOTA Models

In Table 7, we compare Mask DINO with SOTA models on three image segmentation tasks to show its scalability. We use the SwinL backbone and pre-train DINO on the Objects365 detection dataset. Even without using extra data, we outperform Mask2Former on all three tasks, especially on instance segmentation (+2.5 AP). As Mask DINO is an extension of DINO, the pre-trained DINO model can be used to fine-tune Mask DINO for segmentation tasks. After fine-tuning Mask DINO on the corresponding tasks, we achieve the best results on instance (54.5 AP), panoptic (59.4 PQ), and semantic (60.8 mIoU) segmentation among model under one billion parameters. Compared to SwinV2-G , we significantly reduce the model size to 1/15 and backbone pre-training dataset to 1/5. Our detection pre-training also significantly helps all segmentation tasks including panoptic and semantic with "stuff" categories. However, previous specialized segmentation models such as Mask2Former can not use detection datasets and adding a detection head to it results in poor performance as shown in Table 2, which severely limits the data scalability. By unifying four tasks in one model, we only need to pre-train one model on a large-scale dataset and finetune on all tasks for 10 to 20 epochs (Mask2Former needs 100 epochs), which is more computationally efficient and simpler in model design.

3 Ablation Studies

We conduct ablation studies using a ResNet-50 backbone to analyze Mask DINO on COCO val2017. Unless otherwise stated, our experiments are based on object detection and instance segmentation without Mask-enhanced anchor box initialization. Query selection. Table 10 shows the results of our query selection for instance segmentation, where we additionally provide the performance of different decoder layers in one single model. Mask2Former also predicts the masks of learnable queries as initial region proposals. However, their performance lags behind Mask DINO by a large margin (-38.5AP\textbf{-38.5}AP). With our effective query selection scheme, the mask performance achieves 39.6 AP without using the decoder. In addition, our mask performance at layer six is already comparable to the final results with 9 layers. In Table 10, we show that in query selection the predicted box is inferior to mask, which indicates segmentation is easier to learn in the initial stage. Therefore, our proposed mask-enhanced box initialization enhances boxes with masks in query selection to provide better anchor boxes (+15.6 AP) for the decoder, which results in +1.2 AP improvement in the final detection performance. Feature scales. Mask2Former shows that concatenating multi-scale features as input to Transformer decoder layers does not improve the segmentation performance. However, in Table 10, Mask DINO shows that using more feature scales in the decoder consistently improves the performance. Object detection and segmentation help each other. To validate task cooperation in Mask DINO, we use the same model but train different tasks and report the 12 epoch and 50 epoch results. As shown in Table 13, only training one task will lead to a performance drop. Although only training object detection results in faster convergence in the early stage for box prediction, the final performance is still inferior to training both tasks together. Decoder layer number. In DINO, increasing the decoder layer number to nine will decrease the performance of box. In Table 13, the result indicates that increasing the number of decoder layers will contribute to both detection and segmentation in Mask DINO. We hypothesize that the multi-task training become more complex and require more decoders to learn the needed mapping function. Matching. In Table 13, we show that only using boxes or masks to perform bipartite matching is not optimal in Mask DINO. A unified matching objective makes the optimization more consistent. Decoupled box prediction. In Table 15, we show the effectiveness of our decoupled box prediction for panoptic segmentation. This decoupled design of "thing" and "stuff" accelerates training in the early stage (12-epoch setting) and improves the final performance (50-epoch setting). Effectiveness of the algorithm components. In Table 15, we remove each algorithm component at a time and show that each component contributes to the final performance. In addition, after removing all the proposed components, both detection and segmentation performance drop by a large margin. This result indicates that if we trivially add detection and segmentation tasks in one DETR-based model, the features are not aligned for detection and segmentation tasks to achieve mutual cooperation.

We also present visualization analysis in Appendix A.

Conclusion

In this paper, we have presented Mask DINO as a unified Transformer-based framework for both object detection and image segmentation. Conceptually, Mask DINO is a natural extension of DINO from detection to segmentation with minimal modifications on some key components. Mask DINO outperforms previous specialized models and achieves the best results on all three segmentation tasks (instance, panoptic, and semantic) among models under one billion parameters. Moreover, Mask DINO shows that detection and segmentation can help each other in query-based models. In particular, Mask DINO enables semantic and panoptic segmentation to benefit from a better visual representation pre-trained on a large-scale detection dataset. We hope Mask DINO can provide insights for enabling task cooperation and data cooperation towards designing a universal model for more vision tasks. Limitations: Different segmentation tasks fail to achieve mutual assistance in Mask DINO in COCO panoptic segmentation. For example, in COCO panoptic segmentation, the mask AP still lags behind the model only trained with instances. In addition, under the large-scale setting, we have not achieved a new SOTA detection performance as the segmentation head requires additional GPU memory. To accommodate this memory limitation, for the large-scale setting, we have to use smaller image size and less number of queries compared with DINO, which impacts the final performance of object detection. In the future, we will further optimize the implementation to develop a more universal and efficient model to promote task cooperation.

References

Appendix A Visualization analysis

There has been a trend to unify detection and segmentation tasks using convolution-based models, which not only simplifies model design but also promotes mutual cooperation between detection and segmentation. There are mainly three motivations for us to propose Mask DINO. First, DINO has achieved SOTA results on object detection. Previous works such as Mask RCNN , HTC , and DETR have shown that a detection model can be extended to do segmentation and help design better segmentation models. Second, detection is a relatively easier task than instance segmentation. As shown in Table 3 (and other previous studies), Box AP is usually 4+4+ AP higher than mask AP. Therefore, box prediction can guide attention to focus on more meaningful regions and extract better features for mask prediction. Third, the new improvements in DINO and other DETR-like models such as query selection and deformable attention can also help segmentation tasks. For example, Mask2Former adopts learnable decoder queries, which cannot take advantage of the position information in the selected top KK features from the encoder to guide mask predictions. Fig. 2(a)(b)(c) show that the output of Mask2Former in the -th decoder layer is far away from the GT mask while Mask DINO outputs much better masks as region proposals. Mask2Former also adopts specialized masked attention to guide the model to attend to regions of interest. However, masked attention is a hard constraint which ignores features outside a provided mask and may overlook important information for following decoder layers. In addition, deformable attention is also a better substitute for its high efficiency allowing attention to be applied to multi-scale features without too much computational overhead. Fig. 2(d)(e) show a predicted mask of Mask2Former in its 11-st decoder layer and the corresponding output of Mask DINO. The prediction of Mask2Former only covers less than half of the GT mask, which means that the attention can not see the whole instance in the next decoder layer. Moreover, a box can also guide deformable attention to a proper region for background stuff, as shown in Fig. 2(f)(g).

Appendix B Implementation details

The code is available in the supplementary materials. We also provide some detailed descriptions of our implementation here.

Dataset and metrics: We evaluate Mask DINO on two challenging datasets: COCO 2017 for object detection, instance segmentation, and panoptic segmentation; ADE20K for semantic segmentation. They both have "thing" and "stuff" categories, therefore we follow the common practice to evaluate object detection and instance segmentation on the "thing" categories and evaluate panoptic and semantic segmentation on the union of the "thing" and "stuff" categories. Unless otherwise stated, all results are trained on the train split and evaluated on the validation split. For object detection and instance segmentation, the results are evaluated with the standard average precision (AP) and mask AP result. For panoptic segmentation, we evaluate the results with the panoptic quality (PQ) metric . We also report APpanThAP_{pan}^{Th} (AP on the "thing" categories) and APpanStAP_{pan}^{St} (AP on the "stuff" categories). For semantic segmentation, the results are evaluated with the mean Intersection-over-Union (mIoU) metric .

Backbone: We report results with two public backbones: ResNet-50 and SwinL . To achieve SOTA performance using a large model with the SwinL backbone, we use Objects365 to pre-train an object detection model and then fine-tune the model on the corresponding datasets for all tasks. Though we only pre-train for object detection, our model generalizes well to improve the performance of all segmentation tasks. Loss function: As we train detection and segmentation tasks jointly, there are totally three kinds of losses, including classification loss Lcls\mathcal{L}_{cls}, box loss Lbox\mathcal{L}_{box}, and mask loss Lmask\mathcal{L}_{mask}. Among them, box loss (L1 loss LL1\mathcal{L}_{L1} and GIOU loss Lgiou\mathcal{L}_{giou}) and classification loss (focal loss ) are the same as DINO . For mask loss, we adopt cross-entropy Lce\mathcal{L}_{ce} and dice loss Ldice\mathcal{L}_{dice}. We also follow to use point loss in mask loss for efficiency. Therefore, the total loss is a linear combination of three kinds of losses: λclsLcls+λL1LL1+λgiouLgiou+λceLce+λdiceLdice\lambda_{cls}\mathcal{L}_{cls}+\lambda_{L1}\mathcal{L}_{L1}+\lambda_{giou}\mathcal{L}_{giou}+\lambda_{ce}\mathcal{L}_{ce}+\lambda_{dice}\mathcal{L}_{dice}, where we set λcls=4,λL1=5,λgiou=2,λce=5\lambda_{cls}=4,\lambda_{L1}=5,\lambda_{giou}=2,\lambda_{ce}=5, and λdice=5\lambda_{dice}=5. Basic hyper-parameters: Mask DINO has the same architecture as DINO , which is composed of a backbone, a Transformer encoder, and a Transformer decoder. Compared to DINO, we increase the number of decoder layers from six to nine and use 300300 queries. We follow Mask-RCNN and Mask2Former to setup the training and inference settings for segmentation tasks. We use batch size 1616 and train 50 epoch for COCO segmentation tasks (instance and panoptic), 160K iteration for ADE20K semantic segmentation, and 90K90K iterations for Cityscapes semantic segmentation. We set the initial learning rate (lr) as 1×10−41\times 10^{-4} and adopt a simple lr scheduler, which drops lr by multiplying 0.1 at the 11-th epoch for the 12-epoch setting and the 20-th epoch for the 24-epoch setting. For the other segmentation settings, we drop the lr at 0.9 and 0.95 fractions of the total number of training steps by multiplying 0.1. Under the ResNet-50 backbone, we use 44 A100 GPUs each with 40GB memory for all tasks. We report the frames-per-second (fps) tested on the same A100 NVIDIA GPU for Mask2Former and Mask DINO by taking the average computing time with batch size 1 on the entire validation set. Augmentations and Multi-scale setting: We use the same training augmentations as in Mask2Former , where the major difference from DINO on COCO is that we use large-scale jittering (LSJ) augmentation and a fixed size crop to 1024×10241024\times 1024, which also works well for detection tasks. We use the same multi-scale setting as in DINO to use 4 scales in ResNet-50-based models and 5 scales in SwinL-based models.

B.2 Denoising training

Following DN-DETR , we train the model to reconstruct the ground-truth objects given the noised ones. These noised objects will be concatenated with the original decoder queries during training, but will be removed during inference. We add noise to both the bounding box and labels, which will serve as positional embedding and content embedding input to decoder queries. As a box can be viewed as a noised version of a segmentation mask, our unified denoising training will reconstruct the masks given the noised boxes, which improves segmentation training. Label noise: For label noise, we use label flip, which randomly flips a ground-truth label into another possible label in the dataset with probability pp. After adding noise, all the labels will go through a label embedding to construct high-dimensional vectors, which will be the content queries of the decoder. pp is set to 0.2 in our model. Box noise: A box can be formulated as (x,y,w,h)(x,y,w,h), which is also the positional query of DINO . We add two kinds of noise to the box including center shifting and box scaling. For center shifting, we sample a random perturbation (Δx,Δy)(\Delta x,\Delta y) to the box center. The sampled noise is constrained to ∣Δx∣<λ1w2|\Delta x|<\frac{\lambda_{1}w}{2} and ∣Δy∣<λ1h2|\Delta y|<\frac{\lambda_{1}h}{2}, where λ1∈(0,1)\lambda_{1}\in(0,1) is a hyperparameter to control the maximum shifting. For box scaling, the width and height of the box are randomly scaled to [(1−λ2),(1+λ2)]\left[(1-\lambda_{2}),(1+\lambda_{2})\right] of the original ones, where λ2\lambda_{2} is also a hyperparameter to control the scaling. In our model, we set λ1=λ2=0.4\lambda_{1}=\lambda_{2}=0.4.

Appendix C Large models setting

For large models with the SwinL backbone, we follow the same setting of DINO to pre-train a model on the Objects365 dataset for object detection. Then we finetune the pre-trained model on COCO instance and panoptic segmentation for 24 epochs and on ADE20K semantic segmentation for 160k iterations. For training settings on instance and panoptic segmentation on COCO, we use 1.2×1.2\times larger scale (1280×12801280\times 1280) and 1616 A100 GPUs. For training settings on ADE20K semantic, we use 3×3\times more queries (900900) and 88 A100 GPUs. We also use Exponential Moving Average (EMA) in this setting, which helps in ADE20K semantic segmentation.

Appendix D SOTA Results on COCO test-dev

We show the COCO test-dev results in Table 16.