Instances as Queries
Yuxin Fang, Shusheng Yang, Xinggang Wang, Yu Li, Chen Fang, Ying Shan, Bin Feng, Wenyu Liu
Introduction
Instance segmentation is a fundamental yet challenging computer vision task that requires an algorithm to assign a pixel-level mask with a category label for each instance of interest in image. Prevalent state-of-the-art instance segmentation methods are based on high performing object detectors and follow a multi-stage paradigm. Among which, the Mask R-CNN family is the most successful one, where the regions-of-interest (RoI) for instance segmentation is extracted via a region-wise pooling operation (e.g., RoIPool or RoIAlign ) based on the box-level localization information from the region proposal network (RPN) , or the previous stage bounding-box prediction . The final instance mask is obtained via feeding the RoI feature into the mask head, which is a small fully convolutional network (FCN) .
Recently, DETR is proposed to reformulate object detection as a query based direct set prediction problem, whose input is mere learned object queries. Follow-up works in object detection improve this query based approach and achieve comparable performance with state-of-the-art detectors such as Cascade R-CNN . The results show that query based instance-level perception is a very promising research direction. Thus, enabling query based detection framework to perform instance segmentation is highly desirable. However, we find that it is inefficient to integrate the previous successful practices in Cascade Mask R-CNN and HTC , which are state-of-the-art mask generation solutions in the non-query based paradigm, directly into query based detectors for instance mask generation. Therefore, an instance segmentation method tailored for the query based end-to-end framework is urgently needed.
To bridge this gap, we propose QueryInst (Instances as Queries), a query based end-to-end instance segmentation method driven by parallel supervision on dynamic mask heads . The key insight of QueryInst is to leverage the intrinsic one-to-one correspondence in object queries across different stages, and one-to-one correspondence between mask RoI features and object queries in the same stage. Specifically, we set up dynamic mask heads in parallel with each other, which transform each mask RoI feature adaptively according to the corresponding query, and are simultaneously trained in all stages. The mask gradient not only flows back to the backbone feature extractor, but also to the object query, which is intrinsically one-to-one interlinked in different stages. The queries implicitly carry the multi-stage mask information, which is read by RoI features in dynamic mask heads for final mask generation. There is no explicit connection between different stage mask heads or mask features. Moreover, the queries are shared between object detection and instance segmentation sub-networks in each stage, enabling cross-task communications that one task can take advantage of the information from the other task. We demonstrate that this shared query design can fully leverage the synergy between object detection and instance segmentation. When the training is completed, we throw away all the dynamic mask heads in the intermediate stages and only use the final stage predictions for inference. Under such a scheme, QueryInst surpasses the state-of-the-art HTC in terms of AP while runs much faster. Concretely, our main contributions are summarized as follows:
We attempt to solve instance segmentation from a new perspective that uses parallel dynamic mask heads in the query based end-to-end detection framework. This novel solution enables such a new framework to outperform well-established and highly-optimized non-query based multi-stage schemes such as Cascade Mask R-CNN and HTC in terms of both accuracy and speed (see Fig. 1). Specifically, using ResNet--FPN backbone , QueryInst obtains APbox and APmask on COCO -, which is point higher than HTC in terms of both box AP and mask AP, while runs faster. Without bells and whistles, our best model achieves APbox and APmask on COCO -.
We set up a task-joint paradigm for query based object detection and instance segmentation by leveraging the shared query and multi-head self-attention design. This paradigm establishes a kind of communication and synergy between detection and segmentation tasks, which encourages this two tasks to benefits from each other. We demonstrate that our architecture design can also significantly improve the object detection performance.
We extend the QueryInst to video instance segmentation task (VIS) task by simply adding a vanilla track head. Experiments on YouTube-VIS dataset indicate that with same tracking approach, our methods outperforms MaskTrack R-CNN and SipMask-VIS by a large margin. QueryInst-VIS can even outperform well-designed VIS approaches such as STEm-Seg and VisTR .
Related Work
Recently, query based methods emerged to tackle the set-prediction problems. Concretely, DETR first introduces the query based methods with transformer architecture to object detection. Deformable DETR , UP-DETR , ACT and TSP improve the performance on the top of DETR. The recently proposed Sparse R-CNN builds a query based set-prediction framework upon R-CNN based detector. For segmentation, VisTR introduces a query based sequence matching and segmentation method to video instance segmentation, building a fully end-to-end framework for instance segmentation in video. Max-DeepLab presents the first box-free end-to-end panoptic segmentation model with a global memory as external query. Trackformer and Transtrack build a query based multiple object tracktor upon DETR and Deformable DETR, respectively, and attain comparable results to the non-query based methods. AS-Net introduces a query based set-prediction pipeline to human object interaction and obtains promising results. Despite query based set-prediction method is being widely used to many computer vision tasks, few efforts are conducted to build a successful query based instance segmentation framework. We aim to achieve this goal in this paper.
Object detection is a fundamental computer vision task which aims to detect visual objects with bounding boxes. With the propose of R-CNN , Fast R-CNN and Faster R-CNN , anchor based methods dominate object detection for a long period. CenterNet and FCOS establish anchor-free detectors with competitive detection performance. Recently, with the proposed DETR , query based set-prediction methods catch lots of attentions. Deformable DETR introduces deformable convolution to the DETR framework, achieving better performance with faster training convergence. UP-DETR extends DETR to unsupervised scenarios. ACT and TSP introduce the adaptive clustering module and a new bipartite matching method to DETR. Sparse R-CNN build a query based detector on top of R-CNN architecture, while OneNet and DeFCN are end-to-end detector built upon the one-stage FCOS . In this work, we present a query based instance segmentation method on the top of the query based Sparse R-CNN detector.
Instance segmentation is a fundamental yet challenging computer vision task that requires an algorithm to assign a pixel-level mask with a category label for each instance of interest in image. Mask R-CNN introduces a fully convolutional mask head to Faster R-CNN detector. Casacde Mask R-CNN simply combine the Casacde R-CNN with Mask R-CNN. HTC presents interleaved execution and mask information flow and achieves state-of-the-art performance. In addition to R-CNN based methods, YOLACT , SipMask , CondInst and SOLO build one-stage instance segmentation framework on the top of one-stage framework, achieving comparable results with favorable inference speed. Following the R-CNN based methods, we present a query based instance segmentation framework.
Instances as Queries
We propose QueryInst (Instances as Queries), a query based end-to-end instance segmentation method. QueryInst consists of a query based object detector and six dynamic mask heads driven by parallel supervision. Our key insight is to leverage the intrinsic one-to-one correspondence in queries across different stages. This correspondence exists in all query based framework regardless of the specific instantiations and applications. The overall architecture of QueryInst is illustrated in Fig. 2 (c).
QueryInst can be built on any multi-stage query based object detector . We choose Sparse R-CNN as our default instantiation, which has six query stages. The object detection pipeline is depicted in Fig. 2 (a) and can be formulated as follows:
2 Mask Head Architecture
For instance mask prediction, we first adopt the widely used vanilla mask head architecture design in Mask R-CNN as our instance segmentation baseline. The model architecture is depicted in Fig. 2 (b). Based on the object detection pipeline described in Sec. 3.1, the mask generation process can be expressed as follows:
where is the bounding box predictions from the object detector. denotes a region-wise pooling operator for mask RoI features extraction. indicates the mask FCN head consisting of a stack of four consecutive conv-layers, one dconv-layer and one conv-layer for mask generation . is the current stage mask predictions.
Overall, this vanilla design is an analogy of Cascade Mask R-CNN in a query based framework. However, we find that this design is not as effective as the original Cascade Mask R-CNN. Moreover, establishing explicit mask flow following HTC on top of this design (Fig. 2 (b)) can only bring moderate improvements at a cost of large drops in both training and inference speed. Part of the reasons may be the number of queries in our framework is much smaller than the number of proposals in Cascade Mask R-CNN and HTC, resulting in limited availability of training samples.
2.2 Dynamic Mask Head
3 Per-mask Information Flow with Parallel Supervision
In query based models such as , the model learns different specialization for each query slot , i.e., is the transformed and refined version of previous stage in the same -th slot. Moreover, corresponds to and is refined by . Therefore there is an one-to-one correspondence across different stage queries inherent in these frameworks, as well as one-to-one correspondence between mask RoI features and object queries in the same stage.
During training, the per-mask information (i.e., the mask gradient) not only flows back to mask RoI features , but also to the object query , which is intrinsically one-to-one interlinked in different stages. Therefore the per-mask information flow is naturally established by leveraging the inherent properties of query based frameworks, with no additional connection needed. After the training is completed, the information for mask prediction is stored in queries.
4 Shared Query and MSA for Joint Detection and Segmentation
5 Comparisons with Cascade Mask R-CNN and HTC
Cascade Mask R-CNN presents a multi-stage architecture to resample with higher intersection over union (IoU) thresholds for the latter stages, and progressively refine the training distributions of region proposals . This resampling operation guarantees the availability of a large number of accurate localized proposals for the final stage. Despite outperforming non-cascade counterparts, the effectiveness of Cascade Mask R-CNN mainly stems from the progressively refined proposal recall. Whereas, the mask heads across different stages are isolated, and the input feature for mask head always come from the same FPN feature regardless of the stage.
To mitigate the aforementioned drawback of Cascade Mask R-CNN, HTC improves Cascade Mask R-CNN by introducing direct and explicit connections across the mask heads at different stages. The current stage mask features are combined with the accumulated mask features from all previous stages. Equivalently, the mask head of the final stage is deeper than the first stage. Therefore the final stage mask prediction can benefit from deeper features. Establishing direct mask information flow as HTC can alleviate the issue in Cascade Mask R-CNN to some extent. However, this explicit connection across mask heads at different R-CNN stages results in inefficient training and inference.
There are some potential issues inherent in the aforementioned non-query based instance segmentation paradigms. For Cascade Mask R-CNN and HTC, the quality of proposals in different stages is refined in the statistical sense . For each stage, the number and distribution of training samples are quite different, and there is no explicit and intrinsic correspondence for each individual proposal across different stages . Moreover, there is also a mismatch between the training and inference sample distribution . Therefore, introducing direct connections at the architecture-level is necessary for the mask heads in different stages to explicitly learn the correspondence .
Our method does not directly solve the aforementioned issues, but bypasses them. For QueryInst, the connections across stages are naturally established by one-to-one correspondence inherent in queries. This approach eliminates the explicit multi-stage mask head connection and the proposal distribution inconsistency issues. We show that the proposed new paradigm can surpass Cascade Mask R-CNN and HTC in terms of both accuracy and speed.
6 QueryInst-VIS for Video Instance Segmentation
Video instance segmentation (VIS) is a highly relevant task to still-image instance segmentation that aims at detecting, classifying, segmenting and tracking visual instances over video frames. We demonstrate that QueryInst can be easily extended to VIS with minimal modifications by simply adding the vanilla track head in MaskTrack R-CNN baseline . The proposed model coined as QueryInst-VIS can perform video instance segmentation in an online manner while operating at real-time. The total training and inference pipeline keep the same as MaskTrack R-CNN. We evaluate QueryInst-VIS on the challenging YouTube-VIS benchmark to demonstrate its effectiveness.
Experiments
COCO. Most of our experiments are conducted on the challenging COCO dataset . Following the common practice, we use the COCO split ( images) for training and the split ( images) as validation for our ablation study. We report our main results on the - split ( images).
Cityscapes. Cityscapes is an ego-centric street-scene dataset with categories, train images, and validation images for instance segmentation. The images are with higher resolution ( pixels) compared with COCO, and have more pixel-accurate ground-truth.
YouTube-VIS. In addition to static-image instance segmentation, we demonstrate the effectiveness of our QueryInst on video instance segmentation. YouTube-VIS is a challenging dataset for video instance segmentation task, which has a -category label set, unique video instances and high-quality manual annotations. There are training videos, validation videos, and test videos.
2 Implementation Details
Training Setup. Our implementation is based on MMDetection and Detectron2 . Following , the default training schedule is epochs and the initial learning rate is set to , divided by at -th epoch and -th epoch, respectively. We adopt AdamW optimizer with weight decay. Hyper-parameters, configurations as well as the label assignment procedures follow the setting in . In total, the R-CNN head of QueryInst contains stages in parallel as . The mask head is trained by minimizing dice loss . Without special mentioning, we adopt QueryInst model trained with queries and ResNet--FPN as backbone in our experiments in the ablation studies.
3 Main Results
Comparisons on COCO Instance Segmentation.
The comparison of QueryInst with the state-of-the-art instance segmentation methods on COCO - are listed in Tab. 1. We have tested different backbones and data augmentations. CondInst (with auxiliary semantic branch) and SOLOv2 are the latest state-of-the-art instance segmentation approach based on dynamic convolutions. A -stage QueryInst trained with queries outperforms them with over mask AP gain under similar inference speed. QueryInst trained with queries can also surpass Cascade Mask R-CNN by mask AP while runs with the same FPS. For fair comparisons with HTC , we train HTC using the epochs training schedule and multi-scale data augmentations following the standard setting in , yielding higher mask AP than original results reported in . Under same experimental conditions, QueryInst outperforms the state-of-the-art HTC in terms of both accuracy and speed. Moreover, QueryInst outperforms HTC in terms of AP at different IoU thresholds (AP50 and AP75) as well as AP at different scales (APS, APM and APL), regardless of the experimental configuration. We also find that compared with Cascade Mask R-CNN and HTC, the query based QueryInst can benefit more from stronger data argumentation used in We experimentally study the effects of different training schedules and data augmentations to the Mask R-CNN family in the Appendix.. Specifically, using ResNet--FPN backbone and stronger multi-scale data argumentation with random crop, QueryInst surpasses HTC by mask AP and box AP while runs faster. Further, QueryInst with deformable ResNeXt--FPN backbone achieves mask AP and box AP without bells and whistles.
We demonstrate that the instance segmentation performance of QueryInst is not simply come from the accurate bounding box provided by Sparse R-CNN object detector. On the contrary, QueryInst can largely improve the detection performance. The best result of Sparse R-CNN (ResNet--FPN, queries, w/ crop, epochs) reported in is box AP. Under the same experimental setting, QueryInst can achieve box AP, which outperform Sparse R-CNN by box AP. We also show in the ablation study that QueryInst can outperform Cascade Mask R-CNN and HTC based on a weaker query based detector.
We also apply QueryInst to the recent state-of-the-art Swin Transformer backbone without further modifications, and we find the proposed model is quite capable of adapting with Swin-L. Without bells and whistles, QueryInst can achieve the art performance in instance segmentationhttps://paperswithcode.com/sota/instance-segmentation-on-coco as well as object detectionhttps://paperswithcode.com/sota/object-detection-on-coco. For the first time, we demonstrate that an end-to-end query based framework driven by parallel supervision is competitive with well-established and highly-optimized methods in instance-level recognition tasks.
Comparisons on Cityscapes Instance Segmentation.
We also conduct experiments on Cityscapes dataset to demonstrate the generalization of QueryInst. Following the standard setting in , all models are first pre-trained on COCO split then finetuned on Cityscapes using annotations for iterations with batch size ( image per GPU). The initial learning rate is linearly scaled to and is reduced by a factor of at step .
The results are shown in Tab. 2. QueryInst achieves AP on split and AP on split, surpassing several strong baselines. Notably, compared to the dynamic convolution based method CondInst , QueryInst with ResNet- backbone outperforms CondInst with both ResNet--DCN-BiFPN backbone and semantic branch. Overall, our QueryInst achieves leading results on Cityscapes dataset without bells and whistles.
Video Instance Segmentation Results on YouTube-VIS.
Tab. 3 shows the video instance segmentation results on YouTube-VIS set. Following the standard setting in , we first pre-train the instance segmentation model on COCO , then we finetune the corresponding VIS model on YouTube-VIS set for epochs. The maximum number of instances in one frame in YouTube-VIS dataset is , so we set the number of queries to in QueryInst-VIS. The setting enables the model to operate at real-time ( FPS).
As mentioned in Sec. 3.6, QueryInst-VIS adopts the vanilla track method of MaskTrack R-CNN and SipMask-VIS , while it obtains AP improvement compared to MaskTrack R-CNN and AP improvement compared to SipMask-VIS. Moreover, QueryInst can outperform many well-established and highly-optimized VIS approaches, such as STEm-Seg, CompFeat and VisTR in terms of both accuracy and speed.
4 Ablation Study
Tab. 6 studies the impact of using shared query and MSA. As expected in Sec. 3.4, using shared query and MSA simultaneously establishes a kind of communication and synergy between detection and segmentation tasks, which encourages this two tasks to benefits from each other and achieves the highest box AP and mask AP. Moreover, this configuration consumes minimal parameters and computation budgets. Therefore we choose using shared query and MSA as the default instantiation of our QueryInst.
Tab. 4 studies the impact of different mask head architectures on query and non-query based frameworks. All stages are simultaneously trained. For non-query based frameworks, the -st row is the results of Cascade Mask R-CNN and the -nd row is HTC . We have the following major observations.
First, we find that directly integrating cascade mask head and HTC mask flow into the query based model is not as effective as in its original framework. When cascade mask head is applied (-th row), the query based model is APbox and APmask lower than the original Cascade Mask R-CNN (-th row). When HTC mask flow is applied (-th row), the query based model is APbox and APmask lower than the original HTC (-th row). These results demonstrate that previous successful empirical practice from non-query based multi-stage models is possibly inadequate for query based models (Sec. 3.2.1).
Conclusion
In this paper, we propose an efficient query based end-to-end instance segmentation framework, QueryInst, driven by parallel supervision on dynamic mask heads. To our knowledge, QueryInst is the first query based instance segmentation method that outperforms previous state-of-the-art non-query based instance segmentation approaches. Extensive study proves that parallel mask supervision can bring great performance improvement without any decent for inference speed, while dynamic mask head with both shared query and MSA joints two sub-tasks of detection and segmentation naturally. We hope this work can strength the understanding of query based frameworks and facilitate future research.
Appendix
We study the impact of using shared query and shared MSA in the paper. Fig. 5 gives an illustration for our ablation study. Fig. 5 from left to right corresponds to configurations in Tab. 6 in our paper from top to bottom. Using the shared query and shared MSA configuration, QueryInst achieves the best performance in terms of both box AP and mask AP with the least additional overhead.
Training Time of QueryInst
Here, we compare the training time of QueryInst with the state-of-the-art instance segmentation method HTC . As shown in Tab 7, under the same experimental configuration, QueryInst outperforms HTC using less training time.
Effects of Different Training Schedules and Data Augmentations to the Mask R-CNN Family
In Tab. 1, we find that for Cascade Mask R-CNN and HTC, stronger data augmentation cannot bring significant improvements under the training schedule. Here we give a detailed experimental study in Tab. 8.
We choose Mask R-CNN with ResNet--FPN backbone as a representative. Using modest data augmentation (), the schedule works well for the model to converge to near optimum, the performance begins to degenerate under longer training schedules. These findings are consistent with .
Using stronger data augmentation (, w/ crop) cannot bring further improvements under the schedule, but can boost the performance as the training time becomes longer.
Compared with the Mask R-CNN family (Mask R-CNN, Cascade Mask R-CNN & HTC), the proposed QueryInst can benefit from stronger data augmentation even under the schedule. We will conduct further study of these properties in the future.
Qualitative Results on COCO
We provide some qualitative results on COCO split in Fig. 6.
Qualitative Results on Cityscapes
Fig. 7 gives some qualitative results on Cityscapes split.