Instances as Queries

Yuxin Fang, Shusheng Yang, Xinggang Wang, Yu Li, Chen Fang, Ying Shan, Bin Feng, Wenyu Liu

Introduction

Instance segmentation is a fundamental yet challenging computer vision task that requires an algorithm to assign a pixel-level mask with a category label for each instance of interest in image. Prevalent state-of-the-art instance segmentation methods are based on high performing object detectors and follow a multi-stage paradigm. Among which, the Mask R-CNN family is the most successful one, where the regions-of-interest (RoI) for instance segmentation is extracted via a region-wise pooling operation (e.g., RoIPool or RoIAlign ) based on the box-level localization information from the region proposal network (RPN) , or the previous stage bounding-box prediction . The final instance mask is obtained via feeding the RoI feature into the mask head, which is a small fully convolutional network (FCN) .

Recently, DETR is proposed to reformulate object detection as a query based direct set prediction problem, whose input is mere 100100 learned object queries. Follow-up works in object detection improve this query based approach and achieve comparable performance with state-of-the-art detectors such as Cascade R-CNN . The results show that query based instance-level perception is a very promising research direction. Thus, enabling query based detection framework to perform instance segmentation is highly desirable. However, we find that it is inefficient to integrate the previous successful practices in Cascade Mask R-CNN and HTC , which are state-of-the-art mask generation solutions in the non-query based paradigm, directly into query based detectors for instance mask generation. Therefore, an instance segmentation method tailored for the query based end-to-end framework is urgently needed.

To bridge this gap, we propose QueryInst (Instances as Queries), a query based end-to-end instance segmentation method driven by parallel supervision on dynamic mask heads . The key insight of QueryInst is to leverage the intrinsic one-to-one correspondence in object queries across different stages, and one-to-one correspondence between mask RoI features and object queries in the same stage. Specifically, we set up dynamic mask heads in parallel with each other, which transform each mask RoI feature adaptively according to the corresponding query, and are simultaneously trained in all stages. The mask gradient not only flows back to the backbone feature extractor, but also to the object query, which is intrinsically one-to-one interlinked in different stages. The queries implicitly carry the multi-stage mask information, which is read by RoI features in dynamic mask heads for final mask generation. There is no explicit connection between different stage mask heads or mask features. Moreover, the queries are shared between object detection and instance segmentation sub-networks in each stage, enabling cross-task communications that one task can take advantage of the information from the other task. We demonstrate that this shared query design can fully leverage the synergy between object detection and instance segmentation. When the training is completed, we throw away all the dynamic mask heads in the intermediate stages and only use the final stage predictions for inference. Under such a scheme, QueryInst surpasses the state-of-the-art HTC in terms of AP while runs much faster. Concretely, our main contributions are summarized as follows:

We attempt to solve instance segmentation from a new perspective that uses parallel dynamic mask heads in the query based end-to-end detection framework. This novel solution enables such a new framework to outperform well-established and highly-optimized non-query based multi-stage schemes such as Cascade Mask R-CNN and HTC in terms of both accuracy and speed (see Fig. 1). Specifically, using ResNet-101101-FPN backbone , QueryInst obtains 48.148.1 APbox and 42.842.8 APmask on COCO test\mathtt{test}-dev\mathtt{dev}, which is 22 point higher than HTC in terms of both box AP and mask AP, while runs 2.4×2.4\times faster. Without bells and whistles, our best model achieves 50.450.4 APbox and 46.646.6 APmask on COCO test\mathtt{test}-dev\mathtt{dev}.

We set up a task-joint paradigm for query based object detection and instance segmentation by leveraging the shared query and multi-head self-attention design. This paradigm establishes a kind of communication and synergy between detection and segmentation tasks, which encourages this two tasks to benefits from each other. We demonstrate that our architecture design can also significantly improve the object detection performance.

We extend the QueryInst to video instance segmentation task (VIS) task by simply adding a vanilla track head. Experiments on YouTube-VIS dataset indicate that with same tracking approach, our methods outperforms MaskTrack R-CNN and SipMask-VIS by a large margin. QueryInst-VIS can even outperform well-designed VIS approaches such as STEm-Seg and VisTR .

Related Work

Recently, query based methods emerged to tackle the set-prediction problems. Concretely, DETR first introduces the query based methods with transformer architecture to object detection. Deformable DETR , UP-DETR , ACT and TSP improve the performance on the top of DETR. The recently proposed Sparse R-CNN builds a query based set-prediction framework upon R-CNN based detector. For segmentation, VisTR introduces a query based sequence matching and segmentation method to video instance segmentation, building a fully end-to-end framework for instance segmentation in video. Max-DeepLab presents the first box-free end-to-end panoptic segmentation model with a global memory as external query. Trackformer and Transtrack build a query based multiple object tracktor upon DETR and Deformable DETR, respectively, and attain comparable results to the non-query based methods. AS-Net introduces a query based set-prediction pipeline to human object interaction and obtains promising results. Despite query based set-prediction method is being widely used to many computer vision tasks, few efforts are conducted to build a successful query based instance segmentation framework. We aim to achieve this goal in this paper.

Object detection is a fundamental computer vision task which aims to detect visual objects with bounding boxes. With the propose of R-CNN , Fast R-CNN and Faster R-CNN , anchor based methods dominate object detection for a long period. CenterNet and FCOS establish anchor-free detectors with competitive detection performance. Recently, with the proposed DETR , query based set-prediction methods catch lots of attentions. Deformable DETR introduces deformable convolution to the DETR framework, achieving better performance with faster training convergence. UP-DETR extends DETR to unsupervised scenarios. ACT and TSP introduce the adaptive clustering module and a new bipartite matching method to DETR. Sparse R-CNN build a query based detector on top of R-CNN architecture, while OneNet and DeFCN are end-to-end detector built upon the one-stage FCOS . In this work, we present a query based instance segmentation method on the top of the query based Sparse R-CNN detector.

Instance segmentation is a fundamental yet challenging computer vision task that requires an algorithm to assign a pixel-level mask with a category label for each instance of interest in image. Mask R-CNN introduces a fully convolutional mask head to Faster R-CNN detector. Casacde Mask R-CNN simply combine the Casacde R-CNN with Mask R-CNN. HTC presents interleaved execution and mask information flow and achieves state-of-the-art performance. In addition to R-CNN based methods, YOLACT , SipMask , CondInst and SOLO build one-stage instance segmentation framework on the top of one-stage framework, achieving comparable results with favorable inference speed. Following the R-CNN based methods, we present a query based instance segmentation framework.

Instances as Queries

We propose QueryInst (Instances as Queries), a query based end-to-end instance segmentation method. QueryInst consists of a query based object detector and six dynamic mask heads driven by parallel supervision. Our key insight is to leverage the intrinsic one-to-one correspondence in queries across different stages. This correspondence exists in all query based framework regardless of the specific instantiations and applications. The overall architecture of QueryInst is illustrated in Fig. 2 (c).

QueryInst can be built on any multi-stage query based object detector . We choose Sparse R-CNN as our default instantiation, which has six query stages. The object detection pipeline is depicted in Fig. 2 (a) and can be formulated as follows:

2 Mask Head Architecture

For instance mask prediction, we first adopt the widely used vanilla mask head architecture design in Mask R-CNN as our instance segmentation baseline. The model architecture is depicted in Fig. 2 (b). Based on the object detection pipeline described in Sec. 3.1, the mask generation process can be expressed as follows:

where bt\boldsymbol{b}_{t} is the bounding box predictions from the object detector. Pmask\mathcal{P}^{\mathtt{mask}} denotes a region-wise pooling operator for mask RoI features extraction. Mt\mathcal{M}_{t} indicates the mask FCN head consisting of a stack of four consecutive conv-layers, one dconv-layer and one 1×11\times 1 conv-layer for mask generation . mt\boldsymbol{m}_{t} is the current stage mask predictions.

Overall, this vanilla design is an analogy of Cascade Mask R-CNN in a query based framework. However, we find that this design is not as effective as the original Cascade Mask R-CNN. Moreover, establishing explicit mask flow following HTC on top of this design (Fig. 2 (b)) can only bring moderate improvements at a cost of large drops in both training and inference speed. Part of the reasons may be the number of queries in our framework is much smaller than the number of proposals in Cascade Mask R-CNN and HTC, resulting in limited availability of training samples.

2.2 Dynamic Mask Head

3 Per-mask Information Flow with Parallel Supervision

In query based models such as , the model learns different specialization for each query slot , i.e., qt[s]\boldsymbol{q}_{t}[s] is the transformed and refined version of previous stage qt−1[s]\boldsymbol{q}_{t-1}[s] in the same ss-th slot. Moreover, xtmask[s]\boldsymbol{x}^{\mathtt{mask}}_{t}[s] corresponds to and is refined by qt[s]\boldsymbol{q}_{t}[s] . Therefore there is an one-to-one correspondence across different stage queries inherent in these frameworks, as well as one-to-one correspondence between mask RoI features and object queries in the same stage.

During training, the per-mask information (i.e., the mask gradient) not only flows back to mask RoI features xtmask\boldsymbol{x}_{t}^{\mathtt{mask}}, but also to the object query qt−1∗\boldsymbol{q}_{t-1}^{*}, which is intrinsically one-to-one interlinked in different stages. Therefore the per-mask information flow is naturally established by leveraging the inherent properties of query based frameworks, with no additional connection needed. After the training is completed, the information for mask prediction is stored in queries.

4 Shared Query and MSA for Joint Detection and Segmentation

5 Comparisons with Cascade Mask R-CNN and HTC

Cascade Mask R-CNN presents a multi-stage architecture to resample with higher intersection over union (IoU) thresholds for the latter stages, and progressively refine the training distributions of region proposals . This resampling operation guarantees the availability of a large number of accurate localized proposals for the final stage. Despite outperforming non-cascade counterparts, the effectiveness of Cascade Mask R-CNN mainly stems from the progressively refined proposal recall. Whereas, the mask heads across different stages are isolated, and the input feature for mask head always come from the same FPN feature regardless of the stage.

To mitigate the aforementioned drawback of Cascade Mask R-CNN, HTC improves Cascade Mask R-CNN by introducing direct and explicit connections across the mask heads at different stages. The current stage mask features are combined with the accumulated mask features from all previous stages. Equivalently, the mask head of the final stage is 3×3\times deeper than the first stage. Therefore the final stage mask prediction can benefit from deeper features. Establishing direct mask information flow as HTC can alleviate the issue in Cascade Mask R-CNN to some extent. However, this explicit connection across mask heads at different R-CNN stages results in inefficient training and inference.

There are some potential issues inherent in the aforementioned non-query based instance segmentation paradigms. For Cascade Mask R-CNN and HTC, the quality of proposals in different stages is refined in the statistical sense . For each stage, the number and distribution of training samples are quite different, and there is no explicit and intrinsic correspondence for each individual proposal across different stages . Moreover, there is also a mismatch between the training and inference sample distribution . Therefore, introducing direct connections at the architecture-level is necessary for the mask heads in different stages to explicitly learn the correspondence .

Our method does not directly solve the aforementioned issues, but bypasses them. For QueryInst, the connections across stages are naturally established by one-to-one correspondence inherent in queries. This approach eliminates the explicit multi-stage mask head connection and the proposal distribution inconsistency issues. We show that the proposed new paradigm can surpass Cascade Mask R-CNN and HTC in terms of both accuracy and speed.

6 QueryInst-VIS for Video Instance Segmentation

Video instance segmentation (VIS) is a highly relevant task to still-image instance segmentation that aims at detecting, classifying, segmenting and tracking visual instances over video frames. We demonstrate that QueryInst can be easily extended to VIS with minimal modifications by simply adding the vanilla track head in MaskTrack R-CNN baseline . The proposed model coined as QueryInst-VIS can perform video instance segmentation in an online manner while operating at real-time. The total training and inference pipeline keep the same as MaskTrack R-CNN. We evaluate QueryInst-VIS on the challenging YouTube-VIS benchmark to demonstrate its effectiveness.

Experiments

COCO. Most of our experiments are conducted on the challenging COCO dataset . Following the common practice, we use the COCO train2017\mathtt{train2017} split (115k115k images) for training and the val2017\mathtt{val2017} split (5k5k images) as validation for our ablation study. We report our main results on the test\mathtt{test}-dev\mathtt{dev} split (20k20k images).

Cityscapes. Cityscapes is an ego-centric street-scene dataset with 88 categories, 29752975 train images, and 500500 validation images for instance segmentation. The images are with higher resolution (1024×20481024\times 2048 pixels) compared with COCO, and have more pixel-accurate ground-truth.

YouTube-VIS. In addition to static-image instance segmentation, we demonstrate the effectiveness of our QueryInst on video instance segmentation. YouTube-VIS is a challenging dataset for video instance segmentation task, which has a 4040-category label set, 4,8834,883 unique video instances and 131k131k high-quality manual annotations. There are 2,2382,238 training videos, 302302 validation videos, and 343343 test videos.

2 Implementation Details

Training Setup. Our implementation is based on MMDetection and Detectron2 . Following , the default training schedule is 3636 epochs and the initial learning rate is set to 2.5×10−52.5\times 10^{-5}, divided by 1010 at 2727-th epoch and 3333-th epoch, respectively. We adopt AdamW optimizer with 1×10−41\times 10^{-4} weight decay. Hyper-parameters, configurations as well as the label assignment procedures follow the setting in . In total, the R-CNN head of QueryInst contains 66 stages in parallel as . The mask head is trained by minimizing dice loss . Without special mentioning, we adopt QueryInst model trained with 100100 queries and ResNet-5050-FPN as backbone in our experiments in the ablation studies.

3 Main Results

Comparisons on COCO Instance Segmentation.

The comparison of QueryInst with the state-of-the-art instance segmentation methods on COCO test\mathtt{test}-dev\mathtt{dev} are listed in Tab. 1. We have tested different backbones and data augmentations. CondInst (with auxiliary semantic branch) and SOLOv2 are the latest state-of-the-art instance segmentation approach based on dynamic convolutions. A 55-stage QueryInst trained with 100100 queries outperforms them with over 1.11.1 mask AP gain under similar inference speed. QueryInst trained with 100100 queries can also surpass Cascade Mask R-CNN by 1.51.5 mask AP while runs with the same FPS. For fair comparisons with HTC , we train HTC using the 3636 epochs training schedule and multi-scale data augmentations following the standard setting in , yielding ∼1\sim 1 higher mask AP than original results reported in . Under same experimental conditions, QueryInst outperforms the state-of-the-art HTC in terms of both accuracy and speed. Moreover, QueryInst outperforms HTC in terms of AP at different IoU thresholds (AP50 and AP75) as well as AP at different scales (APS, APM and APL), regardless of the experimental configuration. We also find that compared with Cascade Mask R-CNN and HTC, the query based QueryInst can benefit more from stronger data argumentation used in We experimentally study the effects of different training schedules and data augmentations to the Mask R-CNN family in the Appendix.. Specifically, using ResNet-101101-FPN backbone and stronger multi-scale data argumentation with random crop, QueryInst surpasses HTC by 2.02.0 mask AP and 1.81.8 box AP while runs 2.4×2.4\times faster. Further, QueryInst with deformable ResNeXt-101101-FPN backbone achieves 44.644.6 mask AP and 50.450.4 box AP without bells and whistles.

We demonstrate that the instance segmentation performance of QueryInst is not simply come from the accurate bounding box provided by Sparse R-CNN object detector. On the contrary, QueryInst can largely improve the detection performance. The best result of Sparse R-CNN (ResNet-101101-FPN, 300300 queries, 480∼800480\sim 800 w/ crop, 3636 epochs) reported in is 46.346.3 box AP. Under the same experimental setting, QueryInst can achieve 48.148.1 box AP, which outperform Sparse R-CNN by 1.81.8 box AP. We also show in the ablation study that QueryInst can outperform Cascade Mask R-CNN and HTC based on a weaker query based detector.

We also apply QueryInst to the recent state-of-the-art Swin Transformer backbone without further modifications, and we find the proposed model is quite capable of adapting with Swin-L. Without bells and whistles, QueryInst can achieve the art performance in instance segmentationhttps://paperswithcode.com/sota/instance-segmentation-on-coco as well as object detectionhttps://paperswithcode.com/sota/object-detection-on-coco. For the first time, we demonstrate that an end-to-end query based framework driven by parallel supervision is competitive with well-established and highly-optimized methods in instance-level recognition tasks.

Comparisons on Cityscapes Instance Segmentation.

We also conduct experiments on Cityscapes dataset to demonstrate the generalization of QueryInst. Following the standard setting in , all models are first pre-trained on COCO train2017\mathtt{train2017} split then finetuned on Cityscapes using fine\mathtt{fine} annotations for 24k24k iterations with batch size 88 (11 image per GPU). The initial learning rate is linearly scaled to 1.25×10−51.25\times 10^{-5} and is reduced by a factor of 1010 at step 18k18k.

The results are shown in Tab. 2. QueryInst achieves 39.439.4 AP on val\mathtt{val} split and 34.434.4 AP on test\mathtt{test} split, surpassing several strong baselines. Notably, compared to the dynamic convolution based method CondInst , QueryInst with ResNet-5050 backbone outperforms CondInst with both ResNet-101101-DCN-BiFPN backbone and semantic branch. Overall, our QueryInst achieves leading results on Cityscapes dataset without bells and whistles.

Video Instance Segmentation Results on YouTube-VIS.

Tab. 3 shows the video instance segmentation results on YouTube-VIS val\mathtt{val} set. Following the standard setting in , we first pre-train the instance segmentation model on COCO train2017\mathtt{train2017}, then we finetune the corresponding VIS model on YouTube-VIS train\mathtt{train} set for 1212 epochs. The maximum number of instances in one frame in YouTube-VIS dataset is 1010, so we set the number of queries to 1010 in QueryInst-VIS. The setting enables the model to operate at real-time (>30>30 FPS).

As mentioned in Sec. 3.6, QueryInst-VIS adopts the vanilla track method of MaskTrack R-CNN and SipMask-VIS , while it obtains 4.34.3 AP improvement compared to MaskTrack R-CNN and 2.12.1 AP improvement compared to SipMask-VIS. Moreover, QueryInst can outperform many well-established and highly-optimized VIS approaches, such as STEm-Seg, CompFeat and VisTR in terms of both accuracy and speed.

4 Ablation Study

Tab. 6 studies the impact of using shared query and MSA. As expected in Sec. 3.4, using shared query and MSA simultaneously establishes a kind of communication and synergy between detection and segmentation tasks, which encourages this two tasks to benefits from each other and achieves the highest box AP and mask AP. Moreover, this configuration consumes minimal parameters and computation budgets. Therefore we choose using shared query and MSA as the default instantiation of our QueryInst.

Tab. 4 studies the impact of different mask head architectures on query and non-query based frameworks. All stages are simultaneously trained. For non-query based frameworks, the 11-st row is the results of Cascade Mask R-CNN and the 22-nd row is HTC . We have the following 33 major observations.

First, we find that directly integrating cascade mask head and HTC mask flow into the query based model is not as effective as in its original framework. When cascade mask head is applied (33-th row), the query based model is 0.50.5 APbox and 0.60.6 APmask lower than the original Cascade Mask R-CNN (11-th row). When HTC mask flow is applied (44-th row), the query based model is 0.60.6 APbox and 0.40.4 APmask lower than the original HTC (22-th row). These results demonstrate that previous successful empirical practice from non-query based multi-stage models is possibly inadequate for query based models (Sec. 3.2.1).

Conclusion

In this paper, we propose an efficient query based end-to-end instance segmentation framework, QueryInst, driven by parallel supervision on dynamic mask heads. To our knowledge, QueryInst is the first query based instance segmentation method that outperforms previous state-of-the-art non-query based instance segmentation approaches. Extensive study proves that parallel mask supervision can bring great performance improvement without any decent for inference speed, while dynamic mask head with both shared query and MSA joints two sub-tasks of detection and segmentation naturally. We hope this work can strength the understanding of query based frameworks and facilitate future research.

Appendix

We study the impact of using shared query and shared MSA in the paper. Fig. 5 gives an illustration for our ablation study. Fig. 5 from left to right corresponds to configurations in Tab. 6 in our paper from top to bottom. Using the shared query and shared MSA configuration, QueryInst achieves the best performance in terms of both box AP and mask AP with the least additional overhead.

Training Time of QueryInst

Here, we compare the training time of QueryInst with the state-of-the-art instance segmentation method HTC . As shown in Tab 7, under the same experimental configuration, QueryInst outperforms HTC using less training time.

Effects of Different Training Schedules and Data Augmentations to the Mask R-CNN Family

In Tab. 1, we find that for Cascade Mask R-CNN and HTC, stronger data augmentation cannot bring significant improvements under the 3×3\times training schedule. Here we give a detailed experimental study in Tab. 8.

We choose Mask R-CNN with ResNet-101101-FPN backbone as a representative. Using modest data augmentation (640∼800640\sim 800), the 3×3\times schedule works well for the model to converge to near optimum, the performance begins to degenerate under longer training schedules. These findings are consistent with .

Using stronger data augmentation (480∼800480\sim 800, w/ crop) cannot bring further improvements under the 3×3\times schedule, but can boost the performance as the training time becomes longer.

Compared with the Mask R-CNN family (Mask R-CNN, Cascade Mask R-CNN & HTC), the proposed QueryInst can benefit from stronger data augmentation even under the 3×3\times schedule. We will conduct further study of these properties in the future.

Qualitative Results on COCO

We provide some qualitative results on COCO val\mathtt{val} split in Fig. 6.

Qualitative Results on Cityscapes

Fig. 7 gives some qualitative results on Cityscapes test\mathtt{test} split.

Additional Visualization of Dynamic Mask Feature

References